Contextual assistant using mouse pointing or touch cues

The system enhances digital assistant accuracy by using spatial input on a graphical user interface to resolve query ambiguity and identify screen objects, providing precise responses to ambiguous queries.

JP7893889B2Active Publication Date: 2026-07-22GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2023-04-10
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Digital assistants struggle to provide accurate answers to ambiguous queries without additional context, particularly when the user's intent is unclear from the linguistic content.

Method used

A system that combines voice input with spatial input on a graphical user interface to resolve query ambiguity by identifying the object of interest on the screen, using techniques such as cursor position, touch input, or highlighting, and then obtaining information about that object.

Benefits of technology

Enables digital assistants to provide accurate responses to ambiguous queries by leveraging both audio and spatial input to uniquely identify objects on the screen, improving context awareness and response accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893889000001
    Figure 0007893889000001
  • Figure 0007893889000002
    Figure 0007893889000002
  • Figure 0007893889000003
    Figure 0007893889000003
Patent Text Reader

Abstract

The method (400) includes receiving audio data (202) corresponding to a query (104) spoken by a user, accepting a user input indication in a graphical user interface (300) displayed on the screen indicating a spatial input (212) applied at a first location (114) on the screen, and processing the audio data to determine a transcription (214) of the query. The method also includes performing a query interpretation on the transcription to determine that the query references an object (116) displayed on the screen without uniquely identifying the object and requests information about the object. The method further includes disambiguating the query to uniquely identify the object to which the query refers using the user input indication indicating the spatial input applied at the first location on the screen, obtaining information about the object requested by the query, and providing a response (252) to the query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a context assistant that uses mouse pointing or touch cues.

Background Art

[0002] In a voice-responsive environment, a user can speak out a query, and a digital assistant executes an action to obtain an answer to the query. The digital assistant is particularly effective in providing accurate answers to queries on common topics, and the query itself generates the information necessary for the digital assistant to obtain an answer to the query. However, when the query is ambiguous, the digital assistant requires additional context before obtaining an answer to the query. In some cases, additional context required to obtain an answer to the query is provided by identifying the user's interest when the user spoke out the query. Therefore, a digital assistant that receives a query must have some way of identifying additional context of the user who spoke the query.

Summary of the Invention

[0003] One aspect of the present disclosure provides a computer implementation method for causing data processing hardware to perform operations when executed by the data processing hardware, wherein these operations include receiving audio data corresponding to a query spoken by a user and captured by an assistant-enabled device associated with the user. The operations also include receiving a user input display indicating spatial input applied at a first position on the screen in a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, and processing the audio data to determine a transcription of the query using a speech recognition model. The operations also include interpreting the query transcription to determine that the query refers to an object displayed on the screen without uniquely identifying the object, and that it requests information about the object displayed on the screen. The operations also include resolving the ambiguity of the query to uniquely identify the object referred to by the query using a user input display indicating spatial input applied at a first position on the screen, and obtaining information about the object requested by the query in response to uniquely identifying the object. The operations also include providing a response to the query, which includes the obtained information about the object.

[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the operation also includes detecting a trigger event and, in response to detecting the trigger event, activating a GUI displayed on a screen to enable detection of spatial input and activating a speech recognition model to enable speech recognition to be performed on incoming audio data captured by an assistant-enabled device. In these embodiments, detecting a trigger event includes detecting the presence of a hotword in the received audio data by a hotword detector. Alternatively, detecting a trigger event may include accepting user input indications indicating the selection of a graphic element in a GUI displayed on a screen, accepting user input indications indicating the selection of a physical button located on an assistant-enabled device, detecting a predefined gesture made by the user, or detecting a predefined movement / pose of an assistant-enabled device.

[0005] In some examples, receiving a user input display indicating spatial input applied at a first position includes detecting that the cursor position is displayed in the GUI at the first position when the user speaks a query, detecting touch input received in the GUI at the first position when the user speaks a query, or detecting a box selection action performed in the GUI at the first position when the user speaks a query. In these examples, resolving the ambiguity of the query to uniquely identify an object includes receiving image data containing multiple candidate objects displayed in the GUI and the corresponding positions of the multiple candidate objects displayed in the GUI, and identifying a candidate object from the multiple candidate objects having the corresponding position closest to the first position as the object referred to by the query.

[0006] In additional examples, accepting user input displays indicating spatial input applied at a first position includes accepting user input displays indicating spatial input applied at a first position, and resolving query ambiguity to uniquely identify an object includes uniquely identifying the string underlined by the underline action as the object referenced by the query. In other examples, accepting user input displays indicating spatial input applied at a first position includes detecting highlighting actions performed in the GUI that highlight the string displayed in the GUI at a first position, and resolving query ambiguity to uniquely identify an object includes uniquely identifying the string highlighted by the highlighting action as the object referenced by the query.

[0007] In some embodiments, obtaining information about an object requested by a query includes querying a search engine using a uniquely identified object and one or more terms in the query transcription to obtain a list of results in response to the query, and displaying the list of results in response to the query within a GUI displayed on a screen. Displaying the list of results in response to the query may further include generating a graphic element representing the highest-ranking result in the list of results in response to the query, and displaying the list of results in response to the query at a first position on the screen within the GUI displayed on the screen. Optionally, the operation may further include determining that the uniquely identified object contains text in a first language, and therefore obtaining information about an object requested by a query includes obtaining a translation of the text in a second language different from the first language.

[0008] Other aspects of this disclosure provide a system including data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform an operation which includes receiving audio data corresponding to a query spoken by a user and captured by an assistant-enabled device associated with the user. The operation also includes receiving a user input display indicating spatial input applied at a first position on the screen in a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, and processing the audio data to determine a transcription of the query using a speech recognition model. The operation also includes interpreting the query transcription to determine that the query refers to an object displayed on the screen without uniquely identifying the object, and that it requests information about the object displayed on the screen. The operation also includes resolving the ambiguity of the query to uniquely identify the object the query refers to using a user input display indicating spatial input applied at a first position on the screen, and obtaining information about the object requested by the query in response to uniquely identifying the object. The operation also includes providing a response to the query which includes the obtained information about the object.

[0009] This embodiment may include one or more of the following optional features. In some embodiments, the operation also includes detecting a trigger event and, in response to detecting the trigger event, activating a GUI displayed on a screen to enable detection of spatial input and activating a speech recognition model to enable speech recognition to be performed on incoming audio data captured by an assistant-enabled device. In these embodiments, detecting a trigger event includes detecting the presence of a hotword in the received audio data by a hotword detector. Alternatively, detecting a trigger event may include accepting user input indications indicating the selection of a graphic element in a GUI displayed on a screen, accepting user input indications indicating the selection of a physical button located on an assistant-enabled device, detecting a predefined gesture made by the user, or detecting a predefined movement / pose of an assistant-enabled device.

[0010] In some examples, receiving a user input display indicating spatial input applied at a first position includes detecting that the cursor position is displayed in the GUI at the first position when the user speaks a query, detecting touch input received in the GUI at the first position when the user speaks a query, or detecting a box selection action performed in the GUI at the first position when the user speaks a query. In these examples, resolving the ambiguity of the query to uniquely identify an object includes receiving image data containing multiple candidate objects displayed in the GUI and the corresponding positions of the multiple candidate objects displayed in the GUI, and identifying a candidate object from the multiple candidate objects having the corresponding position closest to the first position as the object referred to by the query.

[0011] In additional examples, accepting user input displays indicating spatial input applied at a first position includes accepting user input displays indicating spatial input applied at a first position, and resolving query ambiguity to uniquely identify an object includes uniquely identifying the string underlined by the underline action as the object referenced by the query. In other examples, accepting user input displays indicating spatial input applied at a first position includes detecting highlighting actions performed in the GUI that highlight the string displayed in the GUI at a first position, and resolving query ambiguity to uniquely identify an object includes uniquely identifying the string highlighted by the highlighting action as the object referenced by the query.

[0012] In some embodiments, obtaining information about an object requested by a query includes querying a search engine using a uniquely identified object and one or more terms in the query transcription to obtain a list of results in response to the query, and displaying the list of results in response to the query within a GUI displayed on a screen. Displaying the list of results in response to the query may further include generating a graphic element representing the highest-ranking result in the list of results in response to the query, and displaying the list of results in response to the query at a first position on the screen within the GUI displayed on the screen. Optionally, the operation may further include determining that the uniquely identified object contains text in a first language, and therefore obtaining information about an object requested by a query includes obtaining a translation of the text in a second language different from the first language.

[0013] Details of one or more embodiments of this disclosure are described in the accompanying drawings and the following description. Other embodiments, features, and advantages will become apparent from the description and drawings and the claims. [Brief explanation of the drawing]

[0014] [Figure 1] This is a schematic diagram of an exemplary system that includes a contextual assistant using mouse pointing or touch cues. [Figure 2] This is a schematic diagram of an example component of the context assistant. [Figure 3A] This is an exemplary graphical user interface (GUI) that is rendered on the user device screen, including the context assistant. [Figure 3B] This is an exemplary graphical user interface (GUI) that is rendered on the user device screen, including the context assistant. [Figure 3C] This is an exemplary graphical user interface (GUI) that is rendered on the user device screen, including the context assistant. [Figure 4] This is an illustrative flowchart of the behavior of an array of methods for resolving query ambiguity using mouse pointing or touch queuing. [Figure 5] This is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. [Modes for carrying out the invention]

[0015] Similar reference symbols in various drawings refer to the same elements.

[0016] The way users interact with assistant-enabled devices is designed to be primarily through voice input, if not exclusively. While assistant-enabled devices are effective at obtaining answers to queries on general topics (e.g., "What is the capital of Michigan?"), context-driven queries require the assistant-enabled device to acquire additional information to provide accurate answers. For example, an assistant-enabled device may have difficulty providing a confident / accurate answer to a query like "Show me more of this" without further context.

[0017] In scenarios where additional context beyond a voice query is needed to answer a query, it is beneficial for an assistant-enabled device to capture image data from its screen. For example, a user might naturally query an assistant-enabled device by saying, "Show me more windows like that." Here, the voice query identifies the user as looking for windows similar to a certain object, but is ambiguous because the object is unclear from the linguistic content of the query. By using image data from the assistant-enabled device's screen, the device can narrow down the windows being searched from the entire screen showing the city to a separate sub-region containing a specific building within the city, where user input applied at a specific location on the screen, in conjunction with the voice query, can be detected. By capturing input data and image data along with the query, the assistant-enabled device can generate a response to a query about a building within the city, even though the user needs to explicitly identify the building in the voice query.

[0018] Figure 1 shows an example of a system 100 that includes a user device 10 and / or a remote system 60 that communicates with the user device 10 via a network 40. The user device 10 and / or the remote system 60 run a point assistant 200 that a user 102 can interact with through voice and spatial input, enabling the point assistant 200 to generate responses to queries that refer to objects displayed on the screen of the user device 10, even though the query cannot uniquely identify the object for which it is seeking information. In the illustrated example, the user device 10 corresponds to a smartphone, but the user device 10 may include, but is not limited to, other computing devices that have a display screen or communicate with a display screen, such as a tablet, smart display, desktop / laptop, smartwatch, smart appliance, smart glasses / headset, or vehicle infotainment device. The user device 10 includes data processing hardware 12 and memory hardware 14 that stores instructions that cause the data processing hardware 12 to operate when executed on the data processing hardware 12. The remote system 60 (e.g., a server, a cloud computing environment) also includes data processing hardware 62 and memory hardware 64 that stores instructions causing the data processing hardware 62 to operate when executed on it. As will be described in more detail below, the point assistant 200 running on the user device 10 and / or the remote system 60 includes a speech recognition unit 210 and a response generator 250, and can utilize one or more information sources 240 stored in the memory hardware 14, 64. In some examples, the execution of the point assistant 200 is shared across the user device 10 and the remote system 60.

[0019] The user device 10 includes an array of one or more microphones 16 configured to capture sound, such as voice, directed towards the user device 10. The user device 10 also runs a graphical user interface (GUI) 300 configured to capture user input display via one of the following: touch, gesture, gaze, and / or input devices (e.g., mouse, trackpad, or stylus) to control the functions of the user device 10 for display on a screen 18 that communicates with data processing hardware 12. The GUI 300 may be an interface associated with an application 50 running on the user device 10 that presents multiple objects within the GUI 300. The user device 10 may further include, or communicate with, an audio output device (e.g., a speaker) 19 that can output audio, such as music and / or synthesized speech, from the point assistant 200. The user device 10 may also include a physical button 17 located on the user device 10 and configured to receive a haptic selection by the user 102 to invoke the point assistant 200.

[0020] The user device 10 may include an audio subsystem 106 for extracting audio data 202 (Figure 2) from the query 104. For example, referring to Figure 1, the audio subsystem 106 may receive streaming audio captured by one or more microphones 16 of the user device 10 corresponding to the utterance 106 of the query 104 spoken by user 102, and extract audio data (e.g., acoustic frames) 202. The audio data 202 may include acoustic features such as Mel-frequency cepstrum coefficients (MFCCs) or filter bank energies calculated across a window of the audio signal. In the illustrated example, the query 104 spoken by user 102 includes, "Hey Google, what is this?"

[0021] The user device 10 can run a hotword detector 20 (i.e., on the data processing hardware 12) configured to detect the presence of a hotword 105 in the streaming audio without performing semantic analysis or speech recognition processing on the streaming audio. The hotword detector 20 can run on the audio subsystem 106. The hotword detector 202 can receive audio data 202 and determine whether the utterance 106 contains a specific hotword 105 spoken by the user 102 (e.g., "Hey, Google"). That is, the hotword detector 20 can be trained to detect the presence of the hotword 105 (e.g., "Hey, Google") or one or more other variations of that hotword (e.g., "Okay, Google") in the audio data 202. Detecting the presence of a hotword 105 in audio data 202 may correspond to a trigger event that invokes the point assistant 200 to activate the GUI 300 displayed on the screen 18 to enable detection of spatial input 112, and to activate the speech recognition unit 210 to perform speech recognition on the audio data 202 corresponding to the utterance 106 of the hotword 105 and / or one or more other terms characterizing the query 104 that follows the hotword. In some examples, the hotword 105 is spoken in the utterance 106 that follows the query 105, and a portion of the audio data 202 characterizing such a query 104 is buffered and retrieved by the speech recognition unit 210, which retrieves a portion of the audio data 202 when the hotword 105 is detected in the audio data 202. In some embodiments, the trigger event includes receiving a user input indication in the GUI 300 indicating a selection of a graphic element 21 (e.g., a graphical microphone). In other embodiments, the trigger event includes receiving a user input indication indicating a selection of a physical button 17 located on the user device 10.In other embodiments, the trigger event includes detecting a predefined gesture made by user 102 (e.g., via an image and / or radar sensor), or detecting a predefined movement / pose of user device 10 (e.g., using one or more sensors such as an accelerometer and / or gyroscope).

[0022] User device 10 may further include an image subsystem 108 configured to extract the position 114 (e.g., X-Y coordinate position) on screen 18 of the spatial input 112 applied in GUI 300. For example, user 102 may provide a user input display 110 indicating the spatial input 112 within GUI 300 at position 114 on the screen. Image subsystem 108 may additionally extract image data (e.g., pixels) 204 corresponding to one or more objects 116 currently displayed on screen 18. In the illustrated example, GUI 300 receives a user input display 110 indicating the spatial input 112 applied at a first position 114 on screen 18, and the image data 202 includes an object (i.e., a golden retriever) 116 displayed on screen 18 in proximity to the first position 114.

[0023] Continuing to refer to the system 100 in Figure 1 and the point assistant 200 in Figure 2, the speech recognition unit 210 receives audio data 202 as input and runs an automatic speech recognition (ASR) model (e.g., a speech recognition model) 212 that generates / predicts a corresponding transcription 214 of the query 104 as output. In the illustrated example, the query 104 includes the phrase "What is this?" which requests information 246 about an object 116 displayed on the GUI 300 on the screen without uniquely identifying the object 116. As will be described in more detail below, the point assistant 200 uses a spatial input 112 applied at a first position 114 on the screen 118 to resolve the ambiguity of the query 104 in order to uniquely identify the object 116 that the query 104 refers to. Once the object 116 is uniquely identified, the point assistant 200 can obtain information 246 about the object and generate a response 252 to the query 104 that includes the obtained information 246 about the object 116. The response generator 250 may generate a text representation of the response 252 to the query 104. Here, the point assistant 200 instructs the user device 10 to display the response 252 in the GUI 300 for the user 102 to read. In the illustrated example, the point assistant 200 generates the text representation of the response 252, "It is a golden retriever," for display in the GUI 300. As will be described in more detail below, the point assistant 200 may require additional context extracted by the image subsystem 108 (i.e., the user 102 applied the spatial input 112 at a first position 114 corresponding to the object 116) so that the object 116 referred to by the query 104 can be uniquely identified in order to obtain the information 246 to include in the response 252. In some examples, the response generator 250 uses a text-to-speech (TTS) system 260 to convert the text representation of the response 252 into synthesized speech.In these examples, in addition to or instead of displaying the text representation of response 252 in GUI 300, point assistant 200 generates a synthetic voice for audible output from speaker 19 of user device 10.

[0024] Referring to FIG. 2, point assistant 200 further includes a natural language understanding (NLU) module 220 configured to perform query interpretation on the corresponding transcription 214 in order to ultimately determine the meaning behind the transcription 214. NLU module 220 can also receive context information 201 to assist in the interpretation of transcription 214. Context information 201 can indicate the application 50 (FIG. 1) currently running on user device 10, previous query 104 from user 102, the detection of a specific hot word 105, or any other information that can be utilized by NLU module 220 to interpret query 104. Continuing with the example, context information 201 can indicate that the user is interacting with a web-based application 50 running on user device 102, and NLU module 230 performs query interpretation to determine to specify an action 232 to obtain description / information about some object 116 displayed in GUI 300 that user 102 is likely viewing. However, NLU module 230 determines that query 104 is ambiguous because object 116 is not explicitly identified in transcription 214 other than by the term "this". In other words, the query interpretation performed by NLU module 230 determines that query 104 refers to object 116 displayed on screen 18 without uniquely identifying object 116 and specifies an action 232 to request information 246 about object 116.

[0025] To satisfy query 104, the NLU module 220 needs to resolve the ambiguity of query 104 so that it uniquely identifies the object 116 that query 104 refers to. For example, in a scenario where query 104 includes the corresponding transcription 214 “Show me similar bicycles” while multiple bicycles are currently displayed on screen 18, the NLU module 220 can perform query interpretation on the corresponding transcription 214 so that it can determine that user 102 is referring to an object (i.e., a bicycle) 116 displayed on GUI 300 without uniquely identifying object 116 and without requesting information 246 about object 116 (i.e., other objects similar to bicycle 116). In this example, the NLU module 220 determines that query 104 specifies an action 232 to retrieve an image of a bicycle similar to one of the bicycles displayed on the screen, but cannot satisfy query 104 because the bicycle that the query refers to cannot be identified from transcription 214.

[0026] The NLU module 220 can use a user input display showing a spatial input 112 applied at a first position 114 on the screen as additional context to resolve ambiguity in the query 104 in order to uniquely identify the object 116 referred to by the query. The NLU module 220 can additionally use image data 204 to resolve ambiguity in the query 104, where the image data 204 may include a plurality of candidate objects displayed in the GUI and the corresponding positions of the plurality of candidate objects displayed in the GUI. The image data 204 may be extracted by the image subsystem 108 from graphic content that is rendered and displayed in the GUI 300. The image data 204 may include labels that identify the candidate objects. In some examples, the image subsystem 108 performs one or more object recognition techniques on the graphic content to identify the candidate objects. By using the image data 204 and the received user input display, which is a spatial input 118 applied at a first position 114, the NLU module 220 may be able to uniquely identify an object as an object rendered for display in the GUI 300 closest to the first position 114 of the spatial input 118. In some examples, the contents of the transcription 214 can further narrow down the possibility of the object referred to by the query by describing at least the type of object or by indicating one or more features / characteristics of the object referred to by the query. Once object 116 is uniquely identified, the point assistant 200 adds object 116 to perform action 232 to obtain information 246 about object 116 requested by query 104. Once the point assistant 200 has obtained information 246 about object 116 requested by query 104, the response generator 250 provides a response 252 to query 104, which includes the obtained information 246 about object 116.

[0027] Referring to Figure 3A, in some embodiments, receiving a user input display 110 indicating a spatial input 112 at a first position 114 includes detecting that the position of the cursor 310 is displayed in the GUI 300a at the first position 114 when the user 102 speaks a query 104. In these embodiments, the NLU module 220 further receives image data 204 including a plurality of candidate objects 320, 320a to 320c displayed in the GUI 300. Each of the plurality of candidate objects 320 includes a corresponding position 322, 322a to 322c in the GUI 300a displayed on the screen. These positions may be characterized quantified or otherwise using one or more coordinate systems, such as Cartesian coordinates using a pixel coordinate system with the origin defined by the lower left of the GUI 300a, or a polar coordinate system.

[0028] Furthermore, each of the candidate objects 320 may be spatially defined by a bounding box 330a, 330a-330c, or a box having the smallest dimensions that contains all of the candidate objects 320. From among the multiple candidate objects 320, the NLU module 220 may identify the candidate object 320c having the corresponding position 322c closest to the first position 114 as the object 116 referenced by the query 104. In some examples, if the bounding boxes 330 of two or more candidate objects 320 overlap, the NLU module 220 can use a best intersection technique to calculate the overlap between the two or more bounding boxes 330 in order to identify the object 116 referenced by the query 104. In the illustrated example, the position of the cursor 310 indicates that the spatial input 112 is applied at position 114 where the object 116 containing the sun is displayed.

[0029] In other embodiments (not shown), the user input display 110 indicating spatial input 112 at a first position 114 includes detecting touch input received within the GUI 300 at the first position 114 when user 102 speaks query 104. Alternatively, the user input display 110 indicating spatial input 112 at the first position 114 includes detecting a box selection action performed within the GUI 300 at the first position 114 when user 102 speaks query 104.

[0030] Referring to Figure 3B, in some embodiments, receiving a user input display 110 indicating a spatial input 112 at a first position 114 includes detecting a confinement selection action performed within the GUI 300b at the first position 114. In response to detecting the confinement selection action, the NLU module 220 cuts out a subset of the image data 204 that is contained within the region identified by the confinement selection action and located at the first position 114, and uses the first position 114 in the image data 204 to uniquely identify the object 116 that the query 104 is referring to. In the illustrated example, the object within the region of the confinement selection action includes a building.

[0031] Referring to Figure 3C, in some embodiments, receiving a user input display 110 indicating a spatial input 112 at a first position includes detecting an underlining action performed within the GUI 300c at the first position 114. In these embodiments, the query 104 may target a string displayed within the GUI 300c at the first position 114 (e.g., "Bienvenue au cours de francais!"). For example, the query 104 may include the phrase "What does this mean?". As shown in Figure 3A, the NLU module 220 may identify a candidate object 320 (e.g., the underlined string) having the corresponding position 322 closest to the first position 114 as the object 116 referred to by the query 104. In other embodiments (not shown), a user input display 110 indicating a spatial input 112 at a first position 114 includes detecting a highlighting action performed by GUI300c that highlights a string (e.g., "Bienvenue au cours de francais!") at the first position 114. In these embodiments, the ambiguation model 230 ambiguates the query 104 to uniquely identify the object 116 referenced by the query as the string highlighted by the highlighting action.

[0032] Referring again to Figure 2, once NLU 220 resolves the ambiguity of query 104 and uniquely identifies object 116, NLU module 220 performs action 232 to insert object 116 into the missing object slot of action 232 and obtain information 246 about the uniquely identified object 116 requested by query 104. In some embodiments, point assistant 200 performs identified action 232 to obtain information 246 about the object 116 requested by query 104 by querying information source 240. In these embodiments, information source 240 may include search engine 242, and point assistant 200 queries search engine 242 using the uniquely identified object 116 and one or more terms in the transcription 214 of query 104 to obtain information 246 about the object 116 requested by query 104. For example, the point assistant 200 queries a search engine to obtain information 238 that includes one or more words in the transcription 214 "What is this?", as well as a description of a golden retriever uniquely identified as the object 116 requested by the query. The information source may include an object recognition engine 244 that applies image processing techniques to detect and recognize a pattern (i.e., a golden retriever) in the image data 204 in order to classify object 116 as a golden retriever and obtain information 238 that provides information about golden retrievers. The information may include a link to a content source (e.g., a web page). That is, the information source 240 can use the image data 204 along with the transcription 214 of query 104 to obtain the information 246 requested by query 104. The response generator 250 receives the information 246 requested by query 104 and generates a response 252 "It is a golden retriever". As described above, the response generator 250 can generate a response 252 to the query 104 as a text representation 19 displayed on the GUI 300 on the screen of the user device 10.

[0033] In other examples, the point assistant 200 queries the search engine 242 to obtain a list of results in response to query 104. In these examples, query 104 may also be a similarity query 104, in which case the user 102 seeks a list of results that are visually similar to an object 116 in the GUI 300 on the screen of the user device 10. Once the information source 240 returns information 246 containing the list of results, the response generator 250 may generate a response 252 to query 104 as a text representation 19 containing the list of results displayed on the GUI 300 on the screen of the user device 10. When the point assistant 200 displays the response 252, it may further generate a graphic element representing the highest-ranked result in the list of results in response to query 104, where the highest-ranked result is displayed more prominently than the rest of the results in the ranked list (e.g., in a larger font, highlighted color, at a first position 114).

[0034] In some embodiments, the point assistant 200 determines that a uniquely identified object 116 contains text in a first language (e.g., French). Here, the user 102 who spoke query 104 may speak only a second language (e.g., English) different from the first language. For example, as shown in Figure 3C, the uniquely identified object 116 contains the text "Bienvenue au cours de francais!" in the first language. When the point assistant queries the information source 240 for information 246 about the object, the information source 240 may obtain the translation of the uniquely identified object 116 in the second language, "Welcome to French class!". For example, the information source 240 may include a text-to-text machine translation model.

[0035] Figure 4 is a flowchart of an exemplary sequence of operations of Method 400 for a contextual assistant using mouse pointing or touch cues. Method 400 includes, in operation 402, receiving audio data 202 corresponding to a query 104 spoken by user 102 and captured by an assistant-enabled device (e.g., user device) 10 associated with user 102. Method 400 further includes, in operation 404, receiving a user input display 110 indicating a spatial input 112 applied at a first position 114 on a screen in a graphical user interface 300 displayed on a screen that communicates with data processing hardware 12. In operation 406, Method 400 includes processing the audio data 202 using a speech recognition model 212 to determine a transcription 214 of the query 104.

[0036] In operation 408, method 400 also includes query interpretation of the transcription 214 of query 104 and determining that query 104 refers to an object 116 displayed on a screen without uniquely identifying the object 116, and requests information 256 about the object 116 displayed on the screen. In operation 410, method 400 further includes resolving the ambiguity of query 104 in order to uniquely identify the object 116 that query 104 refers to, using a user input display 110 indicating a spatial input 112 applied at a first position 114 on the screen. In operation 412, in response to uniquely identifying the object 116, method 400 includes obtaining information 246 about the object 116 requested by query 104. In operation 414, method 400 further includes providing a response 252 to query 104 containing the obtained information 246 about the object 116.

[0037] Figure 5 is a schematic diagram of an exemplary computing device 500 that can be used to implement the systems and methods described in this document. The computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the embodiments of the invention described and / or claimed in this document.

[0038] The computing device 500 includes a processor 510, memory 520, a storage device 530, a high-speed interface / controller 540 connected to memory 520 and a high-speed expansion port 550, and a low-speed bus 570 and a low-speed interface / controller 560 connected to storage device 530. Each component 510, 520, 530, 540, 550, and 560 is interconnected using various buses and may be mounted on a common motherboard or otherwise present as needed. The processor 510 (e.g., data processing hardware 12 or data processing hardware 62 in Figure 1) processes instructions for execution within the computing device 500, including instructions stored in memory 520 or storage device 530, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 580 connected to the high-speed interface 540. In other embodiments, multiple processors and / or multiple buses may be used as needed, along with multiple memories and memory types. Additionally, multiple computing devices 500 may be connected, with each device performing some of the necessary operations (for example, as a server bank, a group of blade servers, or a multiprocessor system).

[0039] Memory 520 (for example, memory hardware 14 or memory hardware 64 in Figure 1) stores information non-temporarily within the computing device 500. Memory 520 may be computer-readable media, volatile memory units, or non-volatile memory units. Non-temporarily memory 520 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.

[0040] The storage device 530 can provide high-capacity storage to the computing device 500. In some embodiments, the storage device 530 is a computer-readable medium. In various different embodiments, the storage device 530 may be a device array including a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or a storage area network or other configuration device. In additional embodiments, the computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that perform one or more of the above-described methods at runtime. The information carrier is a computer-readable medium or machine-readable medium such as memory 520, the storage device 530, or memory on the processor 510.

[0041] The high-speed controller 540 manages the bandwidth-intensive operation of the computing device 500, and the low-speed controller 560 manages the low-bandwidth-intensive operation. Such role assignments are merely examples. In some embodiments, the high-speed controller 540 is coupled to memory 520, a display 580 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some embodiments, the low-speed controller 560 is coupled to a storage device 530 and a low-speed expansion port 590. The low-speed expansion port 590 may include various communication ports (such as USB, Bluetooth, Ethernet, and wireless Ethernet) that can connect to one or more input / output devices such as a keyboard, pointing device, scanner, or network devices such as switches and routers via a network adapter, etc.

[0042] The computing device 500 can be implemented in many different forms, as shown in the figure. For example, it may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0043] Various embodiments of the systems and technologies described herein can be realized in digital electronic and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs executable and / or interpretable on a programmable system comprising at least one programmable processor, at least one input device, and at least one output device, which may be specialized or general-purpose, coupled to receive data and instructions from and transmit data and instructions to a storage system.

[0044] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application may be called an “application,” “app,” or “program.” Exemplary applications include, but are not limited to, system diagnostic applications, system administration applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0045] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transient computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor that includes a machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0046] Non-temporary memory may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by a computing device. Non-temporary memory may also be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random-access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), phase-change memory (PCM), and disk or tape.

[0047] The processes and logical flows described herein can be performed by one or more programmable processors, also called data processing hardware, executing one or more computer programs to perform functions by acting on input data and producing outputs. Processes and logical flows can also be performed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Processors suitable for executing computer programs include, for example, one or more processors from both general-purpose and special-purpose microprocessors, and from either type of digital computer. Generally, processors receive instructions and data from read-only memory, random-access memory, or both. The basic elements of a computer are a processor for executing instructions, and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely connected to receive data from or transmit data to them, or both. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices including EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be complemented by or integrated into dedicated logic circuits.

[0048] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal screen) monitor or touchscreen, and optionally a keyboard and pointing device, such as a mouse or trackball, by which the user can input to the computer. Other types of devices can also be used to bring about user interaction, for example, feedback given to the user may be any form of sensory feedback, such as visual feedback, auditory feedback or haptic feedback, and input from the user may be received in any form, such as acoustic input, voice input or haptic input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the user's device, for example, by sending a web page to the user's client device's web browser in response to a request received from a web browser.

[0049] Several embodiments have been described. Nevertheless, it is understood that various modifications can be made without departing from the spirit and scope of this disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A method performed by data processing hardware, Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, A graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware accepts a user input display indicating a spatial input applied at a first position on the screen, The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity in a query in order to uniquely identify the object referred to by the query, using the user input display that shows the spatial input applied at the first position on the screen, Receiving image data including a plurality of candidate objects displayed in the GUI and the corresponding positions of the plurality of candidate objects displayed in the GUI, Based on the proximity of the distance between each of the plurality of candidate objects and the first position, and the characteristics of the object described by the content of the transcription, the object is uniquely identified from the plurality of candidate objects. This includes, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, Methods that include...

2. A method performed by data processing hardware, Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, In a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, the system receives a user input display indicating spatial input applied at a first position on the screen, and includes detecting an underlining action performed in the GUI, which involves underlining a string displayed in the GUI at the first position. The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity in a query to uniquely identify the object referred to by the query using the user input display indicating the spatial input applied at the first position on the screen, including uniquely identifying the string underlined by the underline action as the object referred to by the query, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, Methods that include...

3. A method performed by data processing hardware, Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, In a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, the system receives a user input display indicating spatial input applied at a first position on the screen, and includes detecting a highlighting action performed in the GUI that highlights a string displayed in the GUI at the first position. The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity of a query to uniquely identify the object referred to by the query using the user input display indicating the spatial input applied at the first position on the screen, including uniquely identifying the string highlighted by the highlight action as the object referred to by the query, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, Methods that include...

4. A method performed by data processing hardware, Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, A graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware accepts a user input display indicating a spatial input applied at a first position on the screen, The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Using the user input display that shows the spatial input applied at the first position on the screen, the ambiguity of the query is resolved in order to uniquely identify the object that the query refers to, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, Includes, The method further includes determining that the uniquely identified object contains text in a first language. A method for obtaining the information relating to the object requested by the query, which includes obtaining a translation of the text in a second language different from the first language.

5. Detecting a trigger event and responding to the detection of the trigger event, To enable the detection of spatial input, the GUI displayed on the screen, The speech recognition model enables speech recognition to be performed on input audio data captured by the assistant-enabled device. To start up, The method according to any one of claims 1 to 4, further comprising:

6. The method according to claim 5, wherein detecting the trigger event includes detecting the presence of a hotword in the received audio data using a hotword detector.

7. Detecting the aforementioned trigger event means The system accepts user input indicating the selection of graphic elements displayed on the aforementioned screen. The system accepts user input indicating the selection of a physical button located on the aforementioned assistant-enabled device. To detect a predefined gesture performed by the user, or To detect predefined movements / poses of the aforementioned assistant-enabled device, The method according to claim 5, comprising one of the following.

8. Receiving the user input display indicating the spatial input applied at the first position means that When the user speaks the query, it is detected that the cursor position is displayed in the GUI at the first position. When the user speaks the query, the GUI detects the touch input received at the first location, or The method according to claim 1 or 4, further comprising detecting a box selection action performed in the GUI at a first location when the user speaks the query.

9. To obtain the information relating to the object requested by the aforementioned query, To obtain a list of results in response to the query, the process includes querying a search engine using the uniquely identified object and one or more terms in the transcription of the query. The method according to any one of claims 1 to 4, wherein providing the response to the query, which includes the acquired information, includes displaying a list of the results in response to the query within the GUI displayed on the screen.

10. Displaying the list of results in response to the aforementioned query is: To generate a graphic element representing the highest-ranking result in the list of results obtained in response to the query, In the GUI displayed on the screen, the list of results in response to the query is displayed at the first position on the screen, The method according to claim 9, further comprising:

11. A computer program for causing the data processing hardware to perform the method according to any one of claims 1 to 4.

12. Data processing hardware and Memory hardware that communicates with the data processing hardware, which stores instructions that cause the data processing hardware to perform an operation when executed by the data processing hardware, and the operation is Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, A graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware accepts a user input display indicating a spatial input applied at a first position on the screen, The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity in a query in order to uniquely identify the object referred to by the query, using the user input display that shows the spatial input applied at the first position on the screen, Receiving image data including a plurality of candidate objects displayed in the GUI and the corresponding positions of the plurality of candidate objects displayed in the GUI, Based on the proximity of the distance between each of the plurality of candidate objects and the first position, and the characteristics of the object described by the content of the transcription, the object is uniquely identified from the plurality of candidate objects. This includes, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, This includes memory hardware, A system that includes this.

13. Data processing hardware and Memory hardware that communicates with the data processing hardware, which stores instructions that cause the data processing hardware to perform an operation when executed by the data processing hardware, and the operation is Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, In a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, the system receives a user input display indicating spatial input applied at a first position on the screen, and includes detecting an underlining action performed in the GUI, which involves underlining a string displayed in the GUI at the first position. The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity in a query to uniquely identify the object referred to by the query using the user input display indicating the spatial input applied at the first position on the screen, including uniquely identifying the string underlined by the underline action as the object referred to by the query, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, This includes memory hardware, A system that includes this.

14. Data processing hardware and Memory hardware that communicates with the data processing hardware, which stores instructions that cause the data processing hardware to perform an operation when executed by the data processing hardware, and the operation is Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, In a graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware, the system receives a user input display indicating spatial input applied at a first position on the screen, and includes detecting a highlighting action performed in the GUI that highlights a string displayed in the GUI at the first position. The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Resolving ambiguity of a query to uniquely identify the object referred to by the query using the user input display indicating the spatial input applied at the first position on the screen, including uniquely identifying the string highlighted by the highlight action as the object referred to by the query, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, This includes memory hardware, A system that includes this.

15. Data processing hardware and Memory hardware that communicates with the data processing hardware, which stores instructions that cause the data processing hardware to perform an operation when executed by the data processing hardware, and the operation is Receiving audio data corresponding to queries spoken by the user and captured by an assistant-enabled device associated with the user, A graphical user interface (GUI) displayed on a screen that communicates with the data processing hardware accepts a user input display indicating a spatial input applied at a first position on the screen, The process involves using a speech recognition model to process the audio data and determine the transcription of the query. The query interpretation is performed on the transcription of the aforementioned query, and the query is, The object displayed on the aforementioned screen is being referred to without uniquely identifying the object, The request is for information about the object displayed on the aforementioned screen, To decide, Using the user input display that shows the spatial input applied at the first position on the screen, the ambiguity of the query is resolved in order to uniquely identify the object that the query refers to, In response to uniquely identifying the object, obtain the information relating to the object requested by the query, To provide a response to the query that includes the acquired information relating to the object, This includes memory hardware, Includes, The operation further includes determining that the uniquely identified object contains text in a first language, and obtaining the information relating to the object requested by the query includes obtaining a translation of the text in a second language different from the first language.

16. The above operation involves detecting a trigger event and, in response to detecting the trigger event, To enable the detection of spatial input, the GUI displayed on the screen, The speech recognition model enables speech recognition to be performed on input audio data captured by the assistant-enabled device. To start up, The system according to any one of claims 12 to 15, further comprising:

17. The system according to claim 16, wherein detecting the trigger event includes detecting the presence of a hotword in the received audio data by a hotword detector.

18. Detecting the aforementioned trigger event means The system accepts user input indicating the selection of graphic elements displayed on the aforementioned screen. The system accepts user input indicating the selection of a physical button located on the aforementioned assistant-enabled device. To detect a predefined gesture performed by the user, or To detect predefined movements / poses of the aforementioned assistant-enabled device, The system according to claim 16, comprising one of the following.

19. Receiving the user input display indicating the spatial input applied at the first position means that When the user speaks the query, it is detected that the cursor position is displayed in the GUI at the first position. When the user speaks the query, the GUI detects the touch input received at the first location, or This includes one of the following: detecting a box selection action performed in the GUI at the first location when the user speaks the query; Resolving the ambiguity of the query in order to uniquely identify the aforementioned object is Receiving image data including a plurality of candidate objects displayed in the GUI and the corresponding positions of the plurality of candidate objects displayed in the GUI, The object referred to by the query is to identify the candidate object from among the plurality of candidate objects having the corresponding position closest to the first position, The system according to any one of claims 13 to 15, including the system described in any one of claims 13 to 15.

20. To obtain the information relating to the object requested by the aforementioned query, To obtain a list of results in response to the query, the process includes querying a search engine using the uniquely identified object and one or more terms in the transcription of the query. The system according to any one of claims 12 to 15, wherein providing the response to the query, which includes the acquired information, includes displaying a list of the results in response to the query within the GUI displayed on the screen.

21. Displaying the list of results in response to the aforementioned query is: To generate a graphic element representing the highest-ranking result in the list of results obtained in response to the query, In the GUI displayed on the screen, the list of results in response to the query is displayed at the first position on the screen, The system according to claim 20, further comprising: