Virtual assistant model using gaze context information
By integrating gaze-tracking and vocal input with a language model, the device enhances virtual assistant applications on wearable devices, addressing the limitation of relying solely on spoken or typed input, achieving contextually aware and efficient interactions.
Patent Information
- Application Number
- PCT/US2025/025211
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-17
- Publication Date
- 2025-10-23
AI Technical Summary
Virtual assistant applications on wearable devices like XR devices are limited by their inability to utilize context information associated with the user's desired task, as they rely solely on spoken or typed input, lacking the ability to incorporate gaze and environmental context.
A device integrates gaze-tracking technology to determine the user's focus and combines it with vocal input, using a language model to generate contextually relevant responses, and employs an evaluation model to assess and update the response quality.
Enables more contextually aware and efficient interactions by grounding responses in the user's current focus and saliency history, providing accurate and coherent outputs through gaze-enhanced input processing.
Smart Images

Figure US2025025211_23102025_PF_FP_ABST
Abstract
Description
VIRTUAL ASSISTANT MODEL USING GAZECONTEXT INFORMATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 636.597, filed on Apnl 19, 2024. entitled “VIRTUAL ASSISTANT MODEL USING GAZE CONTEXT INFORMATION,” the disclosure of which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Virtual assistants, such as chatbots or bots, on augmented reality (AR) devices, virtual reality (VR) devices, extended reality (XR) devices, and other devices provide potential as agents that can enable better productivity associated with the user of the device. The virtual assistants use a model (e.g., a large language model (LLM)) by processing input text from users and generating responses utilizing the knowledge and patterns learned from large training data sets. These models enable virtual assistants to understand natural language, infer context, and generate coherent and contextually appropriate responses (or actions), enhancing the quality and conversational ability of the assistant.SUMMARY
[0003] This disclosure relates to systems and methods for providing an assistant application on a computing device (e.g. an XR headset, a wearable such as a head-wom device). In some implementations, a device can be configured to provide an assistant application that performs tasks, provides information, or automates services based on user input through text or voice commands. In some examples, to supplement a user input via voice or text, the device can be configured to identify content displayed by the device and select a subset of the content based on the user’s gaze. From the subset of content and the user input, the assistant application can be configured to apply a model (i.e., language model) that determines a response to the user input. The response can then be provided to the user. In some implementations, the response can include a visual response provided via the display of the device. In some implementations, the response can be provided via at least one speaker on the device. In some examples, the response can be generated via at least one search to a search engine generated from the user input and the context.
[0004] In some implementations, the model can be evaluated using an additional model that compares the input and response to additional inputs and responses that are associated with user evaluations. In some examples, the additional model can determine a projected user evaluation for the input and response using the information from the additional inputs and responses. In some implementations, the evaluation can be used to select the quantity or amount of displayed content selected for processing by the model. In some implementations, the evaluation can be used to update the model.
[0005] In some aspects, the techniques described herein relate to a (e.g., computer- implemented) method including: identifying (e.g., by a computing system such as a wearable computing device) an input from a user; in response to the input, identifying (e.g., by the computing system) a gaze position associated with the user; identifying (e.g.. by the computing system) content on a display associated with the gaze position; determining (e.g., by the computing system) a response based on an application of a model to the input and the content; and providing (e.g. by the computing system) the response to the user.
[0006] In some aspects, the techniques described herein relate to a computing system (e.g. a wearable device such as a head-wom device, e.g., an XR device) including: at least one processor; a computer-readable storage medium operatively coupled to the at least one processor; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing system to perform a method, the method including: identifying an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
[0007] In some aspects, the techniques described herein relate to a computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method including: identify ing an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user. In some aspects, the techniques described herein relate to a computer program carrying said program instructions.
[0008] The accompanying drawings and the description below outline the details of one or more implementations. Other features will be apparent from the description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG. 1 illustrates a computing environment to provide an assistant on a device according to an implementation.
[0010] FIG. 2 illustrates a method of operating a device to provide an assistant according to an implementation.
[0011] FIG. 3 illustrates an operational scenario of selecting content on a display to support a user request according to an implementation.
[0012] FIG. 4 illustrates an operational scenario of providing a response to a user of an assistant application according to an implementation.
[0013] FIG. 5 illustrates an operational scenario of evaluating an assistant application according to an implementation.
[0014] FIG. 6 illustrates a method of evaluating an assistant application according to an implementation.
[0015] FIG. 7 illustrates an operational scenario of updating an assistant application according to an implementation.
[0016] FIG. 8 illustrates a computing system that implements an assistant application and evaluates the assistant application according to an implementation.DETAILED DESCRIPTION
[0017] Computing devices, such as wearable computing devices, can employ a virtual assistant application designed to help users perform tasks or access information using natural language, often through voice or text commands. These applications can schedule appointments, send messages, control smart devices, answer questions, and more. Virtual assistant applications use natural language processing (NLP) to understand user input, whether spoken or typed. Once the input is interpreted, the assistant uses artificial intelligence and machine learning to determine the best response or action. It then executes tasks like setting reminders, retrieving information, or controlling devices.
[0018] For example, a user of an XR device can generate input, voice or typed, to “Identify person X.” The virtual assistant application can identify the user’s intent (i.e., identifying biographical information associated with person X). Once the intent is determined, the application can retrieve relevant data from a trusted knowledge base or online sources, focusing on key facts like occupation, notable achievements, and background details. It then composes an informative summary to present to the user. The device can present the information audibly or visually on a display.
[0019] In another example, the user can provide input to “Generate a new note.” The virtual assistant application can interpret the intent as a request to create and store a piece of information. It then prompts the user (if needed) for content or a title, determines the appropriate format or location for storing the note (such as a note-taking app, local storage, or the cloud), and creates a new note with the provided by the user or default content. Finally, the device can confirm the note has been made (e.g., a visual notification) and may offer follow-up options like editing, sharing, or setting a reminder. However, one technical problem with limiting the input to spoken or typed information on a wearable device, such as an XR device, is that the virtual assistant cannot use context information associated with the user’s desired task.
[0020] In at least one technical solution, a device may employ a model, e.g. a language model such as an LLM that combines information about the user’s gaze and vocal input to provide contextually relevant actions on a computing device, such as initiating an application, initiating an operation, responding to a query, or some other action, e.g. on a computing device. The gaze can be determined through technology integrated into the device’s hardware that identifies the movement and focus of the user’s eyes, allowing the device to understand where the user is looking within the environment and the content viewed by the user. The device may use cameras, infrared sensors, accelerometers, or some other sensor to determine the gaze associated with the user. Once the gaze is defined for the user, the device can determine the displayed content relevant to a text query from the user and generate an action using the model, the text from the query, and the context derived from the displayed content. The context can comprise image-to-text recognition, text extraction from a document or image, or other contextual information.
[0021] As an illustrative example, a user of an XR device may generate a query’ using voice or typed text for “What animal is this?” In response to the query, the system will identify contextual information associated with the term “this” using the gaze information of the user. In some examples, the device will determine the content that is the focus of the gaze and then extract relevant textual information that can provide context to the user's query. Accordingly, if the user’s focus is an image of a tiger, the context can be added to the input of the model, such as a language model such as an LLM, to determine a response that indicates that the animal is a tiger. In some examples, the device can identify images (and a text descriptor associated with the images), text in the user’s gaze, or some other content related to the gaze.
[0022] In some implementations, a user provides verbal input that is detected by one or more microphones on the device and converted to text. In response to the request, the device uses contextual memory associated with the user’s gaze to identify relevant content items for the verbal input. The contextual information can include image recognition, text extraction from a document, text identified from an image, or some other contextual information. The text information (i.e., the contextual) information is combined with the verbal input and provided to the model, e.g. the language model such as an LLM processing model. Once the LLM processing model processes the text from the verbal input and the text derived from the content of the user's gaze, the device may initiate an action. Here, the action includes an audio output and / or a user interface (UI) action, however, other actions may include adding updating portions of applications (e.g., updating a calendar), taking an action in association with an application (e.g., opening the application, generating an email, and the like), or some other action based on the output from the model.
[0023] In one implementation, a device leverages multiple input and output modalities to enable seamless communication between the user and the virtual assistant agent. The user can initiate a verbal request by pressing a key and then speaking, can initiate a request using a command word or phrase, can initiate the request by pressing a button on the device, or can initiate the voice input in some other manner. The user’s speech can be converted to text, which returns the transcribed text that may then be displayed in a user interface, providing visual feedback to the user. The device may comprise any (or all) of the necessary input / output systems, such as e.g. microphones, keyboards, etc. to facilitate the multiple input and output modalities.
[0024] Once the final transcription is identified, it can be processed by a language model application for natural language understanding and generation. The system may maintain a chat history to preserve the interaction context, allowing the assistant application to generate relevant and coherent responses. The response is then displayed on a user interface or display of the device, and the device converts the text into speech. This response is generated when no contextual information is required in association with the user gaze.
[0025] In one implementation, in addition to using the user’s voice input, the device further uses contextual information or contextual memory. By correlating the spatial position of the text with the eye-tracking data from the XR headset, the device can determine which text the user is currently paying attention to. The system maintains a buffer of the user’s saliency history, preserving the order of the text to ensure coherence.
[0026] To differentiate between saccades (i.e.. quick eye movements) and fixations, the device may register text the user has fixated on for a minimum time threshold. The contextual memory on the device may be variable and may include different quantities of words. However, the difference in the size of the contextual memory (i.e., the number of words permitted) may affect the performance of the LLM. Accordingly, in some implementations, the system can select different quantities of contextual data in the user's gaze to provide performance and effective responses to user requests.
[0027] When the user makes a verbal request, the contextual memory is combined with the user’s query' and sent to the model for processing. As at least one technical effect, this approach allows the assistant application to generate responses grounded in the user's current focus and saliency history, enabling more contextually aware and efficient interactions. After each request, the contextual memory buffer may be cleared and prepared to capture new context for the next interaction.
[0028] In some implementations, the device may provide visual feedback associated with the user query. For example, when the user generates a query, the device may generate text, a symbol, or an avatar near the location of the gaze associated with the query. The device may also highlight the area corresponding to the LLM processing, the window on the display, or provide some other indication of the context area associated with the query7. Once the query is processed, the device may also provide feedback associated with the action from the query. The feedback may include a text answer for the request (which also may be provided through speakers), summarize the action taken in an application (e g., adding an event to the calendar), or provide some other visual feedback in association with the query7. The feedback may be audio in some examples.
[0029] In addition to providing the model that uses context information associated with the gaze, the included technical solution may also evaluate the model using an evaluating model or LLM and the results related to a previous user study. The evaluating LLM is presented with the same description of the task as the human participants in the previous user study, along with the formatted interaction logs, then the LLM is asked to provide a data interchange format output (i.e.. including key-value pairs) where each field corresponds to a question from the survey. The survey asks questions that permit the surveys respondents to indicate whether they found the model helpful, whether they found the model performed to their satisfaction, or some other rating associated with it.
[0030] When determining whether the model is helpful or satisfies the user’s expectations, the evaluating model is provided with the dialogue associated with a userinteraction with the model. The dialogue is then processed (i.e., compared to previous dialogues with users) to determine an expected helpfulness rating or satisfaction rating associated with the model. For example, with a first dialogue, the evaluating LLM may determine that it aligns with a first happiness score or rating. However, with a second dialogue, the evaluating LLM may determine that it aligns with a second happiness score or rating. The evaluating LLM is used to determine happiness or satisfaction ratings based on the dialogue from the user. Advantageously, rather than performing additional user studies, the evaluating LLM provides a technical solution that allows a computing system or device to assess a model without performing a direct user survey. This can be used to generate updates that can provide better services to the user or users of the model. In the example of the model that offers actions based on user voice input and gaze information, the model may determine how well the model performed in association with a user’s happiness and satisfaction. In some implementations, the evaluation model can be used to update the model to provide different responses or adjust parameters associated with the model. In some implementations, the evaluation model can indicate potential response issues related to the model and offer them to the application developer. In some implementations, the evaluation model can increase or decrease the context from the user gaze to more accurately respond to the user input. For example, the evaluation model can reduce the context (e.g., text or images) selected as part of the context associated with the user input.
[0031] FIG. 1 illustrates a computing environment 100 to provide an assistant on a device according to an implementation. Computing environment 100 includes user 1 10, device 130, external devices 138, and user view7140. Device 130 includes display 131, sensors 132, camera 133, and assistant application 126. User view- 140 represents the view associated with user 110, and includes gaze 154, content 190. content 191, content 192, and response 195. User 110 provides input 152, representing a voice prompt or request to device 130. Device 130 is an example of a wearable device, such as an XR device, smart glasses, or another w earable device. External devices 138 can include one or more desktop computers, server computers, tablets, smartphones, or other systems that can support assistant application 126. Although demonstrated as a wearable device in computing environment 100, device 130 can represent any device capable of receiving inputs from a user and tracking the user’s gaze to identify relevant context associated with the inputs.
[0032] In some implementations, device 130 executes assistant application 126 to answer user inputs and perform tasks through text or voice. For example, the user may request information about text on display 131 by providing '‘What does this mean?” Inresponse to the request, assistant application 126 can identify the text referenced by identifying the user’s gaze and implement a search to obtain a search result. Once received, assistant application 126 can provide a response based on the received search result, where the response can comprise a visual or audible response.
[0033] In another example, user 110 can request, “Can you add the ingredients for this dish to my shopping list?” The virtual assistant accesses the content the user is viewing, specifically the recipe on the page (identified via tracking the user’s gaze), extracts the listed ingredients, and adds them directly to the user’s shopping list in their grocery application, providing a seamless and contextual response to the request.
[0034] In computing environment 100, device 130 includes display 131, which is a screen or projection surface that presents immersive visual content to user 110. merging virtual elements with the real world. Display 131 can include optical see-through displays (e.g., AR headsets) or video pass-through (e.g., MR / VR devices). Device 130 further includes sensors 132, such as accelerometers, gyroscopes, magnetometers, microphones, depth, infrared, and proximity sensors. The sensors can be used to monitor the user’s physical movement, identify depth information for other objects, identify eye movement for the user, or provide some other operation. Device 130 also includes camera 133, which can capture the real or physical environment to overlay virtual objects (e.g., application interfaces) for identifying the movements of user 110 and surroundings to enable accurate interaction within the augmented or virtual space. In some examples, camera 133 can be positioned as an outward view to capture the physical world associated with the user’s gaze. Display 131 can receive updates from assistant application 126 to overlay a response 195 related to input 152. Sensors 132 and camera 133 provide data to assistant application 126 that can be used to identify’ the user’s gaze (or gesture) and identify displayed content associated with input 152.
[0035] In some implementations, user 110 provides a voice input 152. In some examples, device 130 can identify input 152 in response to user 110 selecting a physical or virtual button. In some examples, device 130 can identify input 152 in response to user 110 providing a keyword or phrase. In some examples, device 130 can identify input 152 based on the understanding of the language from user 110. In response to input 152, assistant application 126 determines a gaze position associated with gaze 154 and identifies relevant content from content 191 as associated with the gaze. In some examples, the relevant content corresponds to one or more images can be processed using image-to-text. Image-to-text, also known as optical character recognition (OCR), is the process of extracting readable and editable text from images using machine learning or computer vision techniques. It enablessystems to interpret text within scanned documents, photos, or screenshots for indexing, searching, or further processing. For example, an image of an apple can return a text descriptor of the apple. In some examples, the relevant content corresponds to text associated with a document, webpage, or other text content provided on the device. At least a portion of the text can be selected based on gaze 154 to provide context associated with voice input 152.
[0036] After the context from gaze 154 is determined, assistant application 126 can use a language model to determine response 195. In some implementations, the language model can be configured to tokenize input 152 and the context (i.e., break it dowTi into smaller units like w ords or sub-w ords) and convert the input and context into embeddings that the model can understand. The language model can be configured to analyze the context and structure of the input using neural network layers, particularly attention mechanisms, to determine the most relevant and coherent response 195. An attention mechanism assigns different w eights to each part of the input sequence, allowing the model to identify the most pertinent words when generating a response. In some implementations, generating response 195 is based at least in part on information from external devices 138 (e.g., a search engine, database, and the like). In some implementations, at least a portion of the model can be employed on external devices 138, where external devices can provide additional processing and memory resources for device 130. Response 195 can be provided visually via display 131 or audibly via a speaker on device 130. In some implementations, response 195 can provide a status associated with input 152. For example, if a user requests an action, such as generating a note for a portion of text, response 195 can provide a visual notification indicating that the operation associated with the input w as completed.
[0037] FIG. 2 illustrates method 200 of operating a device to provide an assistant according to an implementation. Method 200 is described below, referencing systems and elements of computing environment 100 of FIG. 1. In some implementations, method 200 can be performed by device 130 of FIG. 1 . In some implementations, at least a portion of method 200 is performed using one or more additional devices, such as external devices 138.
[0038] Method 200 includes identifying an input from a user at step 201. The device can identify the voice input by monitoring for a wake word or words or activation signal (e g., button press). The device can also identify an input by monitoring the user’s language to identify a potential request or prompt to the device. Method 200 further includes, in response to the voice input, identifying a gaze position associated with the user at step 202 and identifying content on a display associated with the gaze position at step 203.
[0039] In some implementations, device 130 identifies gaze 154 using integrated technology, such as infrared (IR) cameras and light sources embedded within the wearable device. The IR light illuminates the eyes, and the cameras capture reflections from the cornea and pupil. By analyzing these reflections, device 130 can calculate the user’s point of gaze through vector-based algorithms. This data is then mapped onto the virtual environment or displayed content to determine where the user is looking. In some implementations, based on the determined gaze 154, device 130 can select a subset of content that is displayed and relevant to the user. In some implementations, device 130 can select a full document (e.g., a webpage, word document, and the like). In some examples, device 130 can select a portion of the document. In some implementations, the content includes images, text, or some other visible content to the user. For example, device 130 can select a paragraph based on gaze 154 corresponding to the content. As another example, device 130 can perform image-to-text processing to provide a text descriptor associated with an image in gaze 154. In some implementations, the amount of text selected is determined by the device based on its proximity to the user’s gaze. In some implementations, the device can use a combination of images and text. In some implementations, the device can be configured to prefer either text or images in the user’s gaze to select the context from the display.
[0040] Method 200 further provides determining a response based on an application of a model to the input and the content at step 204 and providing the response to the user at step 205. In some implementations, the model comprises a language model. In some examples, the language model processes input 152 by first converting the user’s input and context from the user’s gaze (e.g., text or text descriptor) into a structured format it can analyze. This contextual information and any verbal or textual input form a more complete input or prompt. The model then tokenizes the input (breaking it into smaller parts like words or sub-words), interprets the meaning using its trained parameters, and generates a coherent, relevant response based on patterns it learned during training configuration.
[0041] The model’s response generation is guided by both the explicit input and the implicit context provided by the user’s gaze. For example, if a user is looking at a virtual object and says, “What is this?” the gaze data helps disambiguate “this” by identifying the object of interest. The language model integrates this spatial and semantic context to produce a more accurate and context-aware response. Using the example in computing environment 100, the language model of assistant application 126 can process input 152 with context derived from the context associated with gaze 154 to provide a contextually relevant response 195. In some implementations, the response is displayed. The response can includeinformation directly related to the object or element the user is focusing on, as well as contextual details such as functionality, instructions, or suggestions for interaction. It may also adapt the language or format of the reply based on the user's attention span or engagement level inferred from gaze behavior. In some implementations, the response can be generated at least partly based on preferences of context associated with previous requests and the content included in the requests and responses. In some examples, the response can be generated, at least in part, using information derived from external sources, such as search engines. For example, a web search can be used to provide a result that is provided via a natural language summary7as response 195.
[0042] FIG. 3 illustrates an operational scenario 300 of selecting content on a display to support a user request according to an implementation. Operational scenario 300 includes content 310, content 311, content 312, gaze 340, and selected content 350. Content 311 includes text 330, 331, 332, 333, and 334.
[0043] In operational scenario 300, a user of a device can generate an input (voice or typed). In response to the input, the device can be configured to identify the user’s gaze 340 and identify selected content 350 associated with gaze 340. In some implementations, the user gaze can be identified via one or more sensors on the device that determine eye movement associated with the user. Eye movement can be detected using integrated eyetracking systems that typically employ IR light sources and high-speed cameras to illuminate and capture reflections from the user’s eyes. By analyzing the position and movement of features such as the pupil and corneal reflection, the system can calculate gaze direction and focal points, such as gaze 340.
[0044] In response to identifying gaze 340, the device can be configured to identify selected content 350. In some examples, selected content 350 corresponds to a perimeter or area around gaze 340. In some implementations, selected content 350 corresponds to a threshold quantity of text near gaze 340. In some implementations, selected content 350 corresponds to an image at the focus of gaze 340. In some implementations, selected content 350 corresponds to an image-to-text descriptor associated with an image identified in gaze 340. Here, selected content 350 corresponds to a subset of text (text 332. 333, and 334) identified from gaze 340. Once the context is identified as selected content 350, a natural language model can be applied to the input and the selected content 350 to respond as described herein.
[0045] In some implementations, the selected content can correspond to an area around a focus of the user’s gaze. For example, the selected content can include text orimages that are within a region of the user’s gaze. In some implementations, the selected content can correspond to a window or application displayed for the user that corresponds to the user’s gaze. In some implementations, the selected content can correspond to a document in the user’s gaze. In some implementations, the area or size of the context can be determined from the word choice or intent identified in the input. For example, a first input for a definition of a word may identify a smaller quantity of context than a second input for a summary of a document. The device can be configured to process the language of the input in some examples and determine the quantify (or type) of context required to respond to the user input. For example, the user can provide, “What is in this image?” The device can be configured to identify' the intent is associated with an image and identify an image around the focus of the user’s gaze.
[0046] FIG. 4 illustrates an operational scenario 400 of responding to a user of an assistant application according to an implementation. Operational scenario 400 includes user 410, verbal input 412, gaze information 414, context 416, language model 420, external sources 430, and output 440. Language model 420 can be implemented by a wearable device, such as device 130 of FIG. 1. In some implementations, at least a portion of language model 420 can be implemented at least partially on a second device, such as a companion device (e.g., smartphone or computer).
[0047] In operational scenario 400, user 410 provides verbal input 412. In response to verbal input 412. the device obtains gaze information 414 and context 416 associated with gaze information 414. In some implementations, the device can be configured to use one or more sensors to determine a gaze location associated with the user. The device can then be configured to map the gaze location to a subset of content on a display of the device. The subset of content corresponds to context 416, which can include text, images, or some other media or content. In some implementations, the content can be processed (e.g., image-to-text) to provide a text descriptor for context 416 viewed by the user. Language model 420 can process verbal input 412 and context 416, e.g., context 416 derived gaze information 414, to generate output 440.
[0048] In some implementations, language model 420 can use a voice prompt or input and the content in the user's gaze by integrating speech recognition with eye-tracking technology. The voice prompt provides the language input, which the model processes for intent and context. At the same time, the user's gaze data highlights specific content on a screen, such as a word, image, or section of text, that the user is focusing on. By combining these two inputs, the model can more accurately infer the user's intent and generate arelevant, context-aware response. This multimodal interaction enables a more intuitive and efficient communication experience. In some implementations, language model 420 can retrieve additional information from web searches or databases to provide additional context associated with the user’s search. For example, if the user provides an input of “What is this?” referring to text within a document. A w eb search can be performed with the terms identified in the user’s gaze, and a response can be provided to the user. In some implementations, language model 420 can respond using visual text or audio. A natural language response from a language model is generated by analyzing the user’s input, predicting the most likely following words based on patterns learned during training, and assembling those w ords into a coherent, contextually appropriate response. The model uses probabilities to choose each word to produce fluent and relevant text based on its training or configuration.
[0049] FIG. 5 illustrates an operational scenario 500 of evaluating an assistant application according to an implementation. Operational scenario 500 includes user 510, input 512, response 514, language model 520. logs 516. evaluation model 530, training data 535, and evaluation 540. In some examples, operational scenario 500 can be performed by device 130 of FIG. 1. In some implementations, the steps of operational scenario 500 can be performed at least partially on a secondary' device, such as a server, to evaluate a language model.
[0050] Operational scenario 500 includes applying language model 520 to provide response 514 for input 512. In some examples, language model 520 is provided as part of a virtual assistant that can provide various operations for the user of a device. A virtual assistant can use language model 520 to understand and interpret user inputs and generate appropriate responses based on context derived from viewed content or based on context and intent. It may also access tools, databases, or APIs to provide more accurate or real-time information. The assistant aims to simulate natural conversation while helping the user complete tasks or find information. In some implementations, language model 520 can use the input 512 provided by the user and context derived from content viewed by the user to generate the response 514. As responses are generated for the user inputs, language model 520 also generates logs 51 , such as records that indicate the inputs provided by the user and the responses provided to the inputs. In some examples, logs 516 can further include context information (e.g., text or images) derived from the user’s gaze.
[0051] Logs 516 can be provided to a second model or evaluation model 530 that is used to evaluate how- successful language model 520 is in responding to user inputs. In someimplementations, evaluation model 530 is configured or trained using training data 535 to associate input and response language to evaluations by one or more users as part of a study. In at least one implementation, users can provide inputs to a language model and receive responses to the inputs. The users then indicate whether the inputs and corresponding responses were successful. The success can be affected by the amount of context gathered from the user gaze or the language model 520 itself. For example, when the language model 520 requires additional training, the user can evaluate the language model 520. indicating that additional training could be necessary. In another example, when the evaluation indicates that the response was not on topic, the evaluation can indicate that different amounts of context (i.e., text or images) from the user gaze are required.
[0052] Evaluation model 530 identifies the language in logs 516 and applies evaluation model 530 to generate evaluation 540. The evaluation model 530 can identify word choice and structure of the interactions in logs 516 to determine whether language model 520 should be reviewed favorably or unfavorably. For example, if logs 516 indicate that the user has to request a response multiple times to receive a desired result, then evaluation model 530 can indicate that the language model 520 did not successfully respond to the user inputs. Alternatively, if logs 516 include language that indicates that responses were accurately provided to the user inputs, then evaluation model 530 can indicate that language model 520 successfully acted as a virtual assistant to the user. The language of logs 516 is evaluated against the language of training data 535 to determine how successful language model 520 is in responding to user inputs. In some implementations, evaluation 540 can adjust the amount of context gathered in association with a user input (e.g., increase or decrease the quantity of text or images identified in association with the user gaze). In some implementations, evaluation 540 can adjust language model 520 by changing weights and other parameters to improve language model 520 for future inputs. As at least one technical effect, rather than performing additional studies with users evaluating the language model 520, the language model 520 can be evaluated based on the language in future interactions that is compared with the previous evaluations of the language model 520.
[0053] FIG. 6 illustrates method 600 of evaluating an assistant application according to an implementation. Method 600 includes steps that can be provided by device 130 of FIG. 1 or any other device capable of implementing an evaluation of a language model as described herein. Such devices can include desktop computers, smartphones, servers, or other computing devices.
[0054] Method 600 includes obtaining a record including at least one input and at least one response from a language model at step 601. In some implementations, the language model can maintain records or logs that indicate inputs or requests to the language model and responses provided by the language model to the request. In some implementations, the records can further indicate context (e.g., text and images) identified from the user’s gaze. Method 600 further includes applying a second model, e.g.. the evaluation model 530 of Fig. 5, to the record to determine an evaluation associated with the language model at step 602. In some implementations, the second model is trained or configured to receive input text associated with inputs and responses and determine an evaluation based on records and assessments from a user study. When the language in the obtained record indicates a positive user evaluation of a language model, the evaluation can be positive. Alternatively, the evaluation can be negative when the language in the obtained record indicates a negative user evaluation of a language model (e.g., multiple requests to obtain a desired result).
[0055] Method 600 further includes updating the language model based on the evaluation at step 603. In some examples, the language model can be updated using feedback through fine-tuning or reinforcement learning. In fine-tuning, the language model is retrained on examples labeled as good or bad to reinforce desirable responses. With reinforcement learning, the language model learns to favor positively rated outputs by adjusting its behavior to maximize reward signals from evaluators. In some implementations, the adjustments can be used to change various parameters in the language model to provide a more desirable response. Although demonstrated as updating the model, the evaluation can adjust the quantity and type of context identified in association with a user input. The adjustments can include increasing or decreasing the context selected, prioritizing images or text, or some other adjustment associated with the context.
[0056] FIG. 7 illustrates an operational scenario 700 of updating an assistant application according to an implementation. Operational scenario 700 includes user 710, input 712, response 714, language model 720, logs 716, evaluation model 730, updated language model 721. input 752, and response 754. In some examples, operational scenario 700 can be performed by device 130 of FIG. 1. In some implementations, the steps of operational scenario 700 can be performed at least partially on a secondary device, such as a server, to evaluate the language model 720.
[0057] Operational scenario 700 includes generating logs 716 from language model 720 that receives input 712 and provides response 714. Logs 716 include recorded entries that capture the user’s input (i.e., prompts to the virtual assistant application) and the system’soutput (responses). In some examples, logs 716 may further include timestamps associated with the requests (e.g., inputs 712), context information derived from the user’s gaze, or some other information associated with inputs and responses for language model 720. Language model 720 can be configured to use natural language processing and machine learning techniques to interpret user inputs and generate appropriate responses. When an input is received, the assistant application processes the input through a language model trained on datasets to understand context, intent, and semantics. It then uses this understanding to predict and construct a coherent, contextually relevant response. In some implementations, language model 720 can use context associated with the user’s gaze (i.e., text and images in the user’s gaze). In some implementations, language model 720 can be supplemented via web searches and database searches to identify relevant information associated with the user’s input.
[0058] Once logs 716 are generated, evaluation model 730 can be configured to process logs 716 and determine an update to language model 720. In some implementations, evaluation model 730 is trained using a set of additional inputs and responses and corresponding user evaluations to the inputs and responses. The training permits evaluation model 730 to identify relevant language in the input and response that corresponds to the effectiveness of the language model. For example, if a user requires multiple inputs to receive a desired result, evaluation model 730 can be configured to determine that changes are required to provide better responses.
[0059] Here, evaluation model 730 provides updated language model 721. The updates can be used to refine model parameters that can provide better responses to the user inputs. For example, to test updated language model 721, the updated model can be tested using the same inputs to determine whether improved responses are provided to the inputs. This can reduce the inputs required to provide the user with desired information. The updated language model 721 can be tested, and its logs can be compared to the original logs 716 to determine whether the responses are improved to the original inputs. Once updated language model 721 is created, user 710 can generate input 752 and receive response 754 using updated language model 721.
[0060] Although demonstrated as updating the language model 720, evaluation model 730 can update the context identified in association with a user input. The update can be used to retrieve additional context (e.g.. additional text or images based on the user’s gaze), identify less content (e.g., less text or images based on the user’s gaze), or some other change in the context identified from the user’s gaze. For example, when evaluation model 730determines that the language model 720 is not successfully providing responses based on logs 716, the device executing the assistant application can adjust the context (e.g., increase text) provided in association with the inputs to improve the response provided. This process can be iteratively executed to improve the responses associated with the user inputs.
[0061] In some implementations, for the logs or records maintained of the user interactions with the virtual assistant application, the user may be provided with controls allowing the user to make an election as to both if and when systems, programs, or features described herein may enable collection of user information (e.g., information about a user’s social network, social actions, or activities, profession, a user’s preferences, or a user’s current location), and if the user is sent content or communications from a server. In addition, specific data may be treated in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be treated so that no personally identifiable information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. The queries to the agent, i.e.. the assistant application, can remove any personal identifiable information associated with the user in some examples. Thus, the user may have control over what information is collected about the user and how that information is used to provide the context and multimodal query responses described herein.
[0062] FIG. 8 illustrates a computing system 800 that implements an assistant application 824 and evaluates the assistant application 824 according to an implementation. Computing system 800 represents any apparatus, computing system, or systems with which the various operational architectures, processes, scenarios, and sequences are disclosed herein for providing an assistant application 824 can be implemented. Computing system 800 can be an example of an XR device, wearable device, or other computing device capable of the operations described herein. Computing system 800 is an example of device 130 from FIG. 1. Computing system 800 can be distributed across multiple computing devices in some examples. Computing system 800 includes storage system 845, processing system 850, communication interface 860, and input / output (I / O) device(s) 870. Processing system 850 is operatively linked to communication interface 860, I / O device(s) 870, and storage system 845. In some implementations, communication interface 860 and / or I / O device(s) 870 may be communicatively linked to storage system 845. Computing system 800 may further include other components such as a battery and enclosure that are not shown for clarity.
[0063] Communication interface 860 comprises components that communicate over communication links, such as network cards, ports, radio frequency, processing circuitry (and corresponding software), or some other communication devices. Communication interface 860 may be configured to communicate over metallic, wireless, or optical links. Communication interface 860 may be configured to use Time Division Multiplex (TDM), Internet Protocol (IP), Ethernet, optical networking, wireless protocols, communication signaling, or some other communication format - including combinations thereof. Communication interface 860 may be configured to communicate with external devices, such as servers, user devices, or other computing devices.
[0064] I / O device(s) 870 may include peripherals of a computer that facilitate the interaction between the user and computing system 800. Examples of I / O device(s) 870 may include keyboards, mice, trackpads, monitors, displays, printers, cameras, microphones, external storage devices, sensors, and the like. In some implementations, I / O device(s) 870 include sensors capable of determining the gaze of the user and mapping the gaze to a portion of content displayed by a display of computing system 800.
[0065] Processing system 850 comprises microprocessor circuitry' (e.g., at least one processor) and other circuitry that retrieves and executes operating software (i.e., program instructions) from storage system 845. Storage system 845 may include volatile and nonvolatile, removable, and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Storage system 845 may be implemented as a single storage device but may also be implemented across multiple storage devices or sub-systems. Storage system 845 may comprise additional elements, such as a controller to read operating software from the storage systems. Examples of storage media (also referred to as computer-readable storage media or a computer-readable storage medium) of the storage system 845 include random access memory, read-only memory, magnetic disks, optical disks, and flash memory', as well as any combination or variation thereof, or any other ty pe of storage media. In some implementations, the storage media may be non- transitory’. In some instances, at least a portion of the storage media may be transitory. In no case is the storage media a propagated signal.
[0066] Processing system 850 is ty pically mounted on a circuit board that may also hold the storage system 845. The operating software of storage system 845 comprises computer programs, firmware, or some other form of machine-readable program instructions. The operating software of storage system 845 comprises assistant application 824. Theoperating software on storage system 845 may further include an operating system, utilities, drivers, network interfaces, applications, or some other type of software. When read and executed by processing system 850 the operating software on storage system 845 directs computing system 800 to operate as described herein. In at least one implementation, the operating software can provide method 200 described in FIG. 2. In at least one implementation, the operating software can provide method 600 of FIG. 6. The operating software can provide or cause the at least one processor to identify inputs and provide responses by applying a language model to the inputs and context identified from the user’s gaze. In some examples, the operating software can provide or cause the at least one processor to generate a record of the inputs and responses and evaluate the language model using a second model.
[0067] In at least one implementation, assistant application 824 directs processing system 850 to identify an input from a user. In response to the input, assistant application 824 can further direct processing system 850 to determine a gaze position associated with the user and identify content on a display associated with the gaze position. Assistant application 824 further directs processing system 850 to determine a response based on an application of a language model to the input and the content and provide a response to the user. The response can be provided via a display or audible feedback for the user.
[0068] In at least one implementation, assistant application 824 can direct processing system 850 to generate or determine a record, including at least one input and at least one response from a user. Assistant application 824 can further direct processing system 850 to apply a second model to the record to determine an evaluation associated with the language model, the second model configured based on at least one additional record associated with at least one additional evaluation. In some examples, the second model is trained on an association of inputs and responses to user feedback or user evaluations. The second model can determine language and other characteristics associated with different user evaluations (e.g., repeating questions to receive a desired response). The language can then be mapped from the records (or logs) to a predicted user evaluation associated with the log. In some implementations, the evaluation can be used to update the language model and change one or more parameters associated with the language model. In some implementations, the evaluation can be used to update the context identified in association with the user gaze (e.g., increase or decrease the quantity of context identified for an input).
[0069] Example clauses are provided below. Although these are examples, these clauses should not be considered exhaustive.
[0070] Clause 1. A method comprising: identifying an input from a user; in response to the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
[0071] Clause 2. The method of clause 1, wherein identify ing the content on the display associated with the gaze position comprises: identifying an image on the display associated with the gaze position; identifying text that describes the image; and identifying the text as the content.
[0072] Clause 3. The method of clauses 1 or 2, further comprising: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
[0073] Clause 4. The method of any one of the preceding clauses, wherein the content comprises a first quantify, and the method further comprising: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and determining a second quantity based on the evaluation.
[0074] Clause 5. The method of any one of the preceding clauses, further comprising: identify ing a second input from the user; in response to identifying the second input, identify ing a second gaze position associated with the user; identify ing second content of the second quantify' on the display associated with the second gaze position; determining a second response based on a second application of a model to the input and the second content; and providing the second response to the user.
[0075] Clause 6. The method of any one of the preceding clauses, wherein identifying the content on the display associated with the gaze position comprises identifying text on the display associated with the gaze position.
[0076] Clause 7. The method of any one of the preceding clauses, wherein identifying the content on the display associated with the gaze position comprises identifying at least one image on the display associated with the gaze position.
[0077] Clause 8. The method of any one of the preceding clauses, further comprising: generating a search for a search engine based on the content; and receiving a search responseto the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
[0078] Clause 9. A computing system comprising: at least one processor; a computer- readable storage medium operatively coupled to the at least one processor; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing system to perform a method, the method comprising: identifying an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
[0079] Clause 10. The computing system of clause 9, wherein identifying the content on the display associated with the gaze position comprises: identifying an image on the display associated with the gaze position; identifying text that describes the image; and identifying the text as the content.
[0080] Clause 11. The computing system of clause 9 or 10, wherein the method further compnses: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
[0081] Clause 12. The computing system of any one of clauses 9 to 1 1. wherein the content comprises a first quantity, and wherein the method further comprises: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and determining a second quantify based on the evaluation.
[0082] Clause 13. The computing system of any one of clauses 9 to 12, wherein the method further comprises: identifying a second input from the user; in response to identifying the second input, identifying a second gaze position associated with the user; identifying second content of the second quantify on the display associated with the second gaze position; determining a second response based on a second application of a model to the input and the second content; and providing the second response to the user.
[0083] Clause 14. The computing system of any one of clauses 9 to 13, wherein identifying the content on the display associated with the gaze position comprises identifying text on the display associated with the gaze position.
[0084] Clause 15. The computing system of any one of clauses 9 to 14, wherein identifying the content on the display associated with the gaze position comprises identifying at least one image on the display associated with the gaze position.
[0085] Clause 16. The computing system of any one of clauses 9 to 15, wherein the method further comprises: generating a search for a search engine based on the content; and receiving a search response to the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
[0086] Clause 17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising: identifying an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
[0087] Clause 18. The computer-readable storage medium of clause 17, wherein identifying the content on the display associated with the gaze position comprises: identifying an image on the display associated with the gaze position; identifying text that describes the image; and identifying the text as the content.
[0088] Clause 19. The computer-readable storage medium of clause 17 or 18, wherein the method further comprises: generating a record including at least the input and the response: applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
[0089] Clause 20. The computer-readable storage medium of any one of clauses 17 to 20, wherein the method further comprises: generating a search for a search engine based on the content; and receiving a search response to the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
[0090] In this specification and the appended claims, the singular forms “a,” “an” and “the” do not exclude the plural reference unless the context dictates otherwise. Further, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B. Further, connecting lines or connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physicalconnections, or logical connections may be present in a practical device. Moreover, no item or component is essential to the practice of the implementations disclosed herein unless the element is specifically described as “essential” or '‘critical.’’
[0091] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that a precise value or range thereof is not required and need not be specified. As used herein, the terms discussed above will have ready and instant meaning to one of ordinary skill in the art.
[0092] Moreover, the use of terms such as up, down, top, bottom, side, end, front, back, etc. herein are used concerning a currently considered or illustrated orientation. If they are considered concerning another orientation, such terms must be correspondingly modified.
[0093] Further, in this specification and the appended claims, the singular forms “a,” “an” and '‘the” do not exclude the plural reference unless the context dictates otherwise. Moreover, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B.
[0094] Although certain example methods, apparatuses, and articles of manufacture have been described herein, the scope of coverage of this patent is not limited thereto. It is to be understood that the terminology employed herein is to describe aspects and is not intended to be limiting. On the contrary', this patent covers all methods, apparatus, and articles of manufacture fairly falling within the scope of the claims of this patent.
Claims
WHAT IS CLAIMED IS:
1. A method comprising: identifying an input from a user; in response to the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
2. The method of claim 1, wherein identifying the content on the display associated with the gaze position comprises: identifying an image on the display associated with the gaze position; identifying text that describes the image; and identifying the text as the content.
3. The method of claims 1 or 2, further comprising: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
4. The method of any one of claims 1 to 3, wherein the content comprises a first quantity of content, and the method further comprising: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and determining a second quantity of content based on the evaluation.
5. The method of claim 4, further comprising: identifying a second input from the user:in response to identifying the second input, identifying a second gaze position associated with the user; identifying second content of the second quantity of content on the display associated with the second gaze position; determining a second response based on a second application of the model to the second input and the second content; and providing the second response to the user.
6. The method of any one of claims 1 to 5, wherein identifying the content on the display associated with the gaze position comprises identifying text on the display associated with the gaze position.
7. The method of any one of claims 1 to 6, wherein identifying the content on the display associated with the gaze position comprises identifying at least one image on the display associated with the gaze position.
8. The method of any one of claims 1 to 7, further comprising: generating a search for a search engine based on the content; and receiving a search response to the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
9. A computing system comprising: at least one processor; a computer-readable storage medium operatively coupled to the at least one processor; and program instructions stored on the computer-readable storage medium that, w hen executed by the at least one processor, direct the computing system to perform a method, the method comprising: identifying an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position;determining a response based on an application of a model to the input and the content; and providing the response to the user.
10. The computing system of claim 9, wherein identifying the content on the display associated with the gaze position comprises: identifying an image on the display associated with the gaze position; identify ing text that describes the image; and identify ing the text as the content.
11. The computing system of claim 9 or 10, wherein the method further comprises: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
12. The computing system of any one of claims 9 to 11, wherein the content comprises a first quantity of content, and wherein the method further comprises: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and determining a second quantity of content based on the evaluation.
13. The computing system of claim 12, wherein the method further comprises: identifying a second input from the user; in response to identifying the second input, identifying a second gaze position associated with the user; identify ing second content of the second quantity of content on the display associated with the second gaze position; determining a second response based on a second application of the model to the second input and the second content; and providing the second response to the user.
14. The computing system of any one of claims 9 to 13. wherein identifying the content on the display associated with the gaze position comprises identifying text on the display associated with the gaze position.
15. The computing system of any one of claims 9 to 14, wherein identifying the content on the display associated with the gaze position comprises identifying at least one image on the display associated with the gaze position.
16. The computing system of any one of claims 9 to 15, wherein the method further comprises: generating a search for a search engine based on the content; and receiving a search response to the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising: identifying an input from a user; in response to identifying the input, identifying a gaze position associated with the user; identifying content on a display associated with the gaze position; determining a response based on an application of a model to the input and the content; and providing the response to the user.
18. The computer-readable storage medium of claim 17, wherein identifying the content on the display associated with the gaze position comprises: identify ing an image on the display associated with the gaze position; identify ing text that describes the image; and identifying the text as the content.
19. The computer-readable storage medium of claim 17 or 18, wherein the method further comprises: generating a record including at least the input and the response; applying a second model to the record to determine an evaluation associated with the model, the second model configured based on at least one additional record associated with at least one additional evaluation; and updating the model based on the evaluation.
20. The computer-readable storage medium of any one of claims 17 to 19, wherein the method further comprises: generating a search for a search engine based on the content; and receiving a search response to the search, wherein determining the response based on the application of the model to the input and the content is further based on the search response.
Citation Information
Patent Citations
Method and apparatus for evaluating user intention understanding satisfaction, electronic device and storage medium
US20210383802A1