Display results on a wearable device based on context

By determining the state of immersiveness and using machine learning, wearable devices deliver contextually relevant results for voice inputs, addressing inconsistent interactions and improving user experience.

WO2026039741A1PCT designated stage Publication Date: 2026-02-19GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042186
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-19
Filing Date
2025-08-15
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Wearable devices struggle to provide contextually appropriate results for voice requests based on the device's state of immersiveness, such as AR, MR, or VR, leading to inconsistent and ineffective user interactions.

Method used

The device determines its state of immersiveness and uses machine learning models to process voice inputs, associating them with context-aware responses, including data storage for previous interactions, to provide tailored results based on the user's environment and immersion level.

Benefits of technology

Enables contextually relevant visual responses to voice inputs, adapting the format and content based on the device's state, enhancing user experience by providing appropriate immersive or informative outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042186_19022026_PF_FP_ABST
    Figure US2025042186_19022026_PF_FP_ABST
Patent Text Reader

Abstract

According to at least one implementation, a method includes receiving audio input from a user of a device. The method further includes determining whether the device is in a first state or a second state based on content displayed by the device in response to the audio. When in the first state, the method includes determining a first result based on the audio and the first state and causing display of the first result. When in the second state, the method includes determining a second result based on the audio and the second state and causing display of the second result.
Need to check novelty before this filing date? Find Prior Art

Description

Atty Docket No. 0120-1234W01DISPLAY RESULTS ON A WEARABLE DEVICEBASED ON CONTEXTCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 808,444, filed on May 19, 2025, and U.S. Provisional Application No. 63 / 683,543, filed on August 15, 2024, the disclosures of which are incorporated herein by reference in their entireties.BACKGROUND

[0002] A wearable device, such as an extended reality (XR) device or smart glasses, is a computing system that enables a user to perceive and interact with digital content within a physical environment by combining elements of virtual reality (VR), augmented reality (AR), and / or mixed reality (MR). To display content, the XR device uses a combination of sensors (e g., cameras, inertial measurement units, eye trackers) to determine the user’s pose, position, and surrounding context. The device can be configured to render 2D or 3D virtual imagery that is spatially aligned with the real world. The content can then be provided or projected onto transparent or opaque displays within the device’s optical system, enabling immersive or contextually overlaid visuals.SUMMARY

[0003] This disclosure relates to systems and methods for managing display results on a wearable device. In some implementations, a wearable device, such as an XR device or smart glasses, receives input from a user of a device. In response to the input, the device can be configured to determine whether the device is in a first state or a second state based on content displayed by the device. In some implementations, the first state corresponds to a first amount of immersion provided by the content, and the second state corresponds to a second amount of immersion provided by the content. In some examples, the first state can include a mixed reality' state, and the second state can include an augmented reality' state. When the device is in the first state, the device can determine a first result based on the input and the first state and cause display of the first result. When the device is in the second state, theAtty Docket No. 0120-1234W01 device can determine a second result based on the input and the second state and cause display of the second result.

[0004] In some aspects, the techniques described herein relate to a (computer- implemented) method including: receiving audio from a user of a device; in response to (e.g., reception of) the audio (e.g., by the device), determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0005] In some aspects, the techniques described herein relate to a system including: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the system to perform a method, the method including: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0006] In some aspects, the techniques described herein relate to a computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method including: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0007] The accompanying drawings and the description below outline the details of one or more implementations. Other features will be apparent from the description, drawings, and claims.Atty Docket No. 0120-1234W01BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG. 1 illustrates a computing environment to manage the display of results according to an implementation.

[0009] FIG. 2A illustrates a method of determining a result for a voice request based on device state according to an implementation.

[0010] FIG. 2B illustrates a method of using a data store to determine and display results according to an implementation.

[0011] FIG. 3 A illustrates an operational scenario 300 of providing content to a user based on device state according to an implementation.

[0012] FIG. 3B illustrates an operational scenario 350 of providing results to a user according to an implementation.

[0013] FIG. 4 illustrates an operational scenario of selecting a display of content based on device state according to an implementation.

[0014] FIG. 5 illustrates an operational scenario of using a data store to determine display results according to an implementation.

[0015] FIG. 6 illustrates a block diagram of a device according to an implementation.

[0016] FIG. 7 illustrates a first example implementation of a spatial action model.

[0017] FIG. 8A is a block diagram illustrating a second example implementation of a spatial action model.

[0018] FIG. 8B is an example semantic graph that may be used in the implementation of FIG. 8 A.

[0019] FIG. 8C illustrates a first example of page sparsification that may be used in the implementation of FIG. 8A.

[0020] FIG. 8D illustrates a second example of page sparsification that may be used in the implementation of FIG. 8A.

[0021] FIG. 9A is a block diagram illustrating a third example implementation of a spatial action model.

[0022] FIG. 9B is a block diagram illustrating a fourth example implementation of a spatial action model.

[0023] FIG. 10 is a block diagram illustrating a fifth example implementation of a spatial action model.

[0024] FIG. 11 is a block diagram illustrating a sixth example implementation of a spatial action model.Atty Docket No. 0120-1234W01

[0025] FIG. 12 illustrates a computing system to provide display results to an implementation.DETAILED DESCRIPTION

[0026] A wearable device, such as an extended reality (XR) device or smart glasses, includes hardware worn on the body that enables users to experience and interact with digital content overlaid on or integrated into the real world. These devices can be configured to support Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) by using sensors, cameras, and displays to track the user’s environment and movements. This allows for immersive or interactive experiences where digital content can either replace the physical world entirety or enhance it with contextual information.

[0027] In some implementations, users can use voice requests to interact with a wearable device. The device captures speech using one or more built-in microphones and processes it with one or more speech recognition models to convert spoken words into text or actions. The system can then interpret the meaning and trigger the appropriate response, such as opening an application, navigating a menu, or controlling virtual objects on the device. For example, a user can request directions to the nearest coffee shop. In response to the request, the device can process the speech and determine the intent (i.e., find the closest coffee shop). Once the intent is determined, the device can act, such as opening a maps application and searching for the nearest coffee shop to provide directions to the end user. However, while the device can give results to voice requests, at least one technical problem exists in providing the results in an effective format for the user.

[0028] In at least one technical solution, a device can be configured to provide a result to a voice request based on context or state associated with the device. In some implementations, the context can include a state of immersiveness (i.e., an amount of immersion) associated with the device (e.g., AR state or MR state). The amount of immersion can refer to the degree to w hich a user perceives digital content as part of their environment. This can range from content overlaid on the physical world (e.g., augmented reality) to a fully enclosed, computer-generated environment (e.g., virtual reality). An AR state may refer to a mode where a device overlays computer-generated information, such as text or 2D graphics, onto the user’s view of the real world. A VR state may refer to a mode where a device provides a fully immersive experience that replaces the user’s real-world environment with a simulated one. An MR state may refer to a mode that blends real and virtual worlds, allowingAtty Docket No. 0120-1234W01 digital objects to interact with the physical environment in real-time (e.g., placing a virtual couch in a living room of the user).

[0029] For example, the device can provide a first result when the device is in a first state (e.g., AR state) and a second result when the device is in a second state (e.g., MR state). When the user generates a voice request, the device can perform natural language processing on the request. For example, the user can provide a request to “take me to Paris.” The device can be configured to use speech recognition to convert the spoken request into text, then natural language processing interprets the intent. Intent can be determined by separating the user’s speech into meaningful parts. The device can be configured to use models, such as machine learning models, to identify what action the user wants. The device can match the words and phrases to a defined intent, such as executing a maps application.

[0030] In some technical solutions, to determine the action (or result), the device can identify' additional context associated with content displayed on the device, the environment of the user, or another context characteristic. When the user provides speech input, the device can identify’ the context and determine the result for the speech input based on the context. In some implementations, the user can provide the same voice or audio input in different contextual environments and receive different results based on the context. For example, suppose the user provides “take me to Paris'’ in a first contextual state (e.g., a first state of immersiveness). In that case, the device can display first content associated with the voice input. Alternatively, when the user provides the same speech input in a second contextual state (e.g., a second state of immersiveness), the device can give second content as a response to the voice input. As at least one technical effect, the contextual information can provide different results (e.g., visual responses) to the same voice input. A “state,” as used herein, can refer to the overall operational context of a device at a given time, determined from one or more factors. These one or more factors may include, but are not limited to, the immersion level (amount of immersion) for displayed content (such as augmented reality, mixed reality, or virtual reality ), the geographical location of the device, the time of day, the user's physical surroundings and movements, and / or the applications currently executing on the device. The determined state is used to select a corresponding result or response to a user’s input.

[0031] As an illustrative example, the user can say, “Take me to the Eiffel Towner in Paris.” In response to the request, the device can determine the contextual state associated with the device. The contextual state can include the level of immersiveness (e.g., whether the user is in an MR or VR mode), the geographic location of the user (e.g., in the user’s living room), other content displayed by the device, the physical surroundings of the user, orAtty Docket No. 0120-1234W01 other contextual information. If the user is in a VR mode, the device can be configured to retrieve and render a 3D scene of Paris that includes the Eiffel Tower. These images can be provided from a search engine and correspond to the user’s intent. Thus, based on the voice request and the context, content can be displayed that reflects the current state of the device. In an alternative example, if the user is walking and in a MR mode for the device, the device can query a search engine based on the voice request for flights, directions, or other travel information to take the user to the desired location. The information can be displayed in 2D format, permitting the user to continue interacting with objects and other elements in the physical space. As a technical effect, based on the immersion status associated with the display, the device can be configured to provide different result (i.e., content) to the same verbal request. A ‘"result” is the information or content that is generated and displayed by the device in response to a user's input. The format, content, and level of immersion of the result are determined based on the device’s current state, allowing for a contextually appropriate visual response, such as a two-dimensional information card, a three-dimensional virtual object, or a fully immersive scene.

[0032] In another illustrative example, a device can include at least one outwardly facing camera that captures information about the physical environment of the user. The device can be configured with a memory store that stores information about the physical environment, such as objects observed in the physical environment, tags for the objects, text from the objects, and the like. For example, the user can provide a prompt to remember a sofa in a store. The device can be configured to identify the name, image, or other information associated with the sofa. In some implementations, at least a portion of the information can be obtained via an image search associated with the sofa. In some implementations, at least a portion of the information can be derived from text identified in one or more images captured of the sofa (e.g., a price tag or descriptor). In some implementations, the user requests that the device store information about the sofa. In some implementations, the device can store information associated with objects based on the time the user views the object, speech context associated with the user when viewing the object, or other factors. For example, the device can store information about the sofa based on the user’s gaze focusing on the sofa for a threshold period.

[0033] After the device stores information about an object, the user can query information about the object, and the query results can be provided based on the context associated with the device. The context can include the user’s location, content displayed by the device, an immersion state related to the display of the device, or other context. TheAtty Docket No. 0120-1234W01“immersion state” or “amount of immersion” can refer to the degree to which a system presents digital content that overlays, blends with, and / or replaces a user’s perception of the physical environment. For example, the amount of immersion can range from displaying two- dimensional information on a transparent display to generating a fully virtual world. For example, if the user requests, “Can I see that sofa?” The device can identify the user's location (e.g., living room), the current immersion state for the device, and the like. If the user is in their house, the device can use MR to display a version of the sofa in the house, permitting the user to view a rendering of the couch over the physical environment. If the user is in another store, the device can provide a list of search results or an image of the couch, allowing the user to compare another sofa, identify additional information about the sofa, or access other associated details. As at least one technical effect, the results that are provided to the user can be based on the contextual information of the user’s environment and immersion state.

[0034] In some implementations, the device can be configured to use a machine learning model to respond to the user input. In some implementations, the model can be configured to receive inputs associated with the user’s speech and the device’s state information and provide a response based on the inputs. The model can be configured (i.e., trained) using a combination of supervised learning for speech recognition and natural language understanding, using labeled datasets of audio input and their meanings. As used herein, “audio” can refer to an electronic signal generated by one or more microphones that captures sound waves corresponding to a user’s spoken utterance. This signal is configured for processing by a speech recognition or natural language processing system to extract an actionable command or user intent. For context-awareness, the model can be configured with rule-based logic or configured using context-rich datasets that pair user inputs with relevant device state information and correct responses. The configuration permits the model to associate audio inputs and device states to desired responses. Thus, while some inputs can provide 2D results (e.g., a web result), other inputs can provide 3D immersive results. The immersive results can include VR results (e.g., placing the user in a requested location), MR results (e.g., displaying an object in the user’s environment), or AR results (e.g., displaying a 2D search result for the audio input).

[0035] In at least one illustrative example, a user can request to store data associated with one or more entities (e.g.. physical objects, events, people, businesses, and the like). For example, a user can generate a request to store information about a table in a department store. The device can be configured to capture various data about the table, such as the nameAtty Docket No. 0120-1234W01 or identifier for the table, the size of the table, the color of the table, or any other data associated with the table. In some implementations, at least a portion of the data can be obtained from the user’s request (e.g., '‘Store the ‘'X” table for later”, where ‘'X” is the identifier for the table). In some implementations, at least a portion of the data can be obtained from one or more sensors on the wearable device. The data can be obtained from captured images, audio, or some other sensor on the device. For example, the device can capture an image of the table and perform a reverse image search to determine an identifier for the table. The device can also identify the size, manufacturer, or any other information associated with the object.

[0036] Once stored, either locally on the device or in a remote user profile (i.e., in a cloud system), the user can provide an input. The input can include an audio input in some examples. In other examples, the input can comprise a typed input. For example, referring to storing data about the table, the user can provide voice input “Let me see the table.” In response to the request, the device can be configured to perform natural language processing and determine the user’s intent (i.e., to display the table previously stored by the user). In some examples, the intent can be determined using a model configured to relate intent to data previously stored (i.e., data about the table). The device can further be configured to determine a current state associated with the device. The term “state” can denote a set of parameters or a data structure that represents the current condition of the device and / or its environment. This data structure is populated with inputs from various sources, such as display mode information, sensor data (e g., location from GPS, user movement from a movement sensor), and system information (e.g., system clock, running processes). This state can be evaluated to provide a contextually appropriate response to a user request. The state can be based on whether the device is in a MR, VR. or AR mode. The state can be based at least partially on the location of the user. The state can be based on the current time. The state can be based on a combination of the aforementioned factors.

[0037] For example, if the user requested to see the table at home and the screen state for the device permits MR displays, a representation of the table can be displayed, permitting the user to view the table in their environment. As a technical effect, the user can test objects within a current environment when the state of the device satisfies at least one criterion. The state of the device can be based on the display state of the device, location of the device, time of day, or any other factor to determine how information should be displayed for a user.

[0038] As an alternative example, if the user requested to see the table in another furniture store, the device can determine the state of the device (e.g., based on a display state.Atty Docket No. 0120-1234W01 the device location, the time, etc.) and provide an AR response that displays information about the saved table to the user. The information can include the name or identifier of the table, the manufacturer of the table, size information for the table, or other information about the table (i.e., object). In some examples, the device can store a first portion of datum (e.g., a unique identifier for the table), and use a search engine to obtain additional information to provide to the user. The additional information can include the listed information or can include a webpage in some examples that corresponds to the object. As a technical effect, the user can be provided with potential information in a display format based on the device state.

[0039] FIG. 1 illustrates a computing environment 100 to manage the display of results according to an implementation. Computing environment 100 includes user 110, device 130, voice input 144, and display view 105. Device 130 includes display 131. sensors 132, camera 133, and response application 126. Display view 105 represents the view of user 110 using device 130. Display view 105 includes result 1 0 displayed in response to voice input 144 (i.e., audio input). Device 130 is an example of a wearable device, such as an XR device, smart glasses, or another wearable device. Voice input 144 can include audio that denotes any sound-based input from a user, such as a voice request or spoken command, that is received by the device and serves as a trigger for initiating a process. The system is configured to interpret this audio to determine a user's intent and provide a context-dependent result.

[0040] In at least one implementation, device 130 receives voice input 144 and processes voice input 144 using response application 126. Response application 126 can be configured to perform natural language processing on the speech of voice input 144 to determine intent and provide a result or response to voice input 144. In addition to using voice input 144, response application 126 further determines context associated with device 130 to determine the result or result 160 associated with voice input 144. The context can include an immersion state for the display of the device (i.e., the amount of immersion), the location of the device, the surroundings identified via camera 133, or some other information about the device. From the context and voice input 144. result 160 is determined and displayed via display 131. In the example of computing environment 100, result 160 is provided as a 2D response. For example, voice input 144 can include a request for the best restaurant nearby. In response to the request, response application 126 can process the natural language of voice input 144 and the state information associated with device 130 to determine the results provided to the user.Atty Docket No. 0120-1234W01

[0041] In computing environment 100, device 130 includes display 131, a screen or projection surface that presents immersive visual content to user 110. merging virtual elements with the real world. Display 131 can include optical see-through displays or video pass-through. Device 130 further includes sensors 132, such as accelerometers, gyroscopes, magnetometers, depth, infrared, and proximity sensors. The sensors can be used to monitor the physical movement of the user, identify depth information for other objects, identify eye movement for the user, or perform other operations. Device 130 also includes camera 133. which can capture the real or physical environment to overlay virtual objects (e.g., application interfaces) for identifying movements of user 110 and surroundings to enable accurate interaction within the augmented or virtual space. In some examples, camera 133 can be positioned as an outward view to capture the physical world associated with the user’s gaze. Display 131 can receive updates from response application 126 to display content associated with voice input 144. Sensors 132 and camera 133 provide data associated with the context for device 130, including location, environmental characteristics, and the like. The data can provide context associated with the voice input 144.

[0042] In some implementations, user 110 provides voice input 144, which is processed by response application 126. In addition to performing natural language processing, response application 126 can include a model (e.g., a machine learning model) that determines result 160 based on context associated with device 130. The model processes voice input 144 using a language understanding component to derive the intent. Additionally, the mode can receive contextual information from device 130, like location, environment information, executing applications, and / or screen state (e.g., AR, MR, or VR), which can be encoded into a structured format. The tw o streams of information, voice input 144 and context, can be combined using mechanisms like a neural network or attention-based model to determine the response or result 160. In some examples, the model determines what information to display and how the information is displayed. For example, for the same voice request, device 130 can determine that VR content should be displayed based on a first state of the device and that MR content should be displayed based on a second state of the device.

[0043] In some implementations, device 130 and response application 126 can maintain a data store associated with one or more earlier interactions between user 110 and device 130. In some implementations, the data store can maintain logs, images, or other information associated with earlier interactions between user 110 and device 130. For example, user 110 can interact with response application 126 in association with a trip to Paris. User 110 can conduct searches for (and identify ) potential restaurants in the city andAtty Docket No. 0120-1234W01 indicate preferences associated with returned restaurants. For example, the user can provide feedback to a search indicating, "We should check out this restaurant.” Device 130 can update the data store with the restaurant (and information about the restaurant). The information about the restaurant can include the location, menu items highlighted for the restaurant, hours of operation, or other information associated with the restaurant. Later, such as when the user visits Paris, the user can provide voice input 144, indicating “What was the restaurant I was looking at?” Device 130 and response application 126 can be configured to identify information associated with the restaurant in the data store and use the information to provide a response or result to the user. In some implementations, the selected information can be based on context associated with the device, such as the location of the device, environmental information gathered from an outward facing camera (e.g., a view of a menu), the screen state of the device, or some other contextual information associated with the device.

[0044] Although demonstrated as initiating result 160 in response to voice input 144, response application 126 can provide result 160 based on the context associated with device 130. For example, user 110 can provide first interactions with device 130, wherein device 130 can store at least a portion of the interactions. In some examples, objects of interest are tagged in association with the interaction, such as images, dates, locations, or other information associated with the interaction. For example, user 110 can provide a first interaction with device 130 to identify the best bagel in New York. Later, when user 110 visits New York or is in a location near a bagel shop identified during the first interaction, response application 126 can generate a result 160 (e.g., a notification). The notification can indicate the bagel shop and the reasoning for providing the information to the user (e.g., the previous interaction). In some examples, the notification can be displayed. In other examples, the notification can be provided via audio and speakers located on device 130. In some examples, the information provided in association with result 160 refers to the first interaction (e.g., user preferences for bagels, pricing, etc.). In some examples, the user may also manually request this information using voice input 144. As a technical effect, device 130 can maintain information about previous interactions between user 110 and device 130, permitting the device to provide context during future interactions. Result 160 can be provided based on voice input 144, the current context associated with device 130, and previous interactions between user 110 and device 130.

[0045] In at least one example, a system can store at least one datum associated with an entity (e.g., an object, an event, a business, a person, and the like) based on a userAtty Docket No. 0120-1234W01 command. As used herein, "at least one datum" refers to one or more discrete pieces of information associated with an entity or interaction. A datum may include, but is not limited to, an identifier, a descriptive attribute (e.g., name, size, color), contextual metadata (e.g., a timestamp or location), or a log of a user interaction, which is stored for later retrieval and processing. In at least one example, the system can store at least one datum associated with an object based on contextual information associated with the object. The contextual information can include the user viewing the object for a threshold period, the user discussing the object with a second person, the user providing context indicating interest, or some other factor. For example, the device can be configured to identify an object based on the user’s speech, such as “I really like that table.’' In response to expressly receiving a request to store information about the object or an implicit determination to store information about the object, the device can store at least one datum associated with the object. In some implementations, the at least one datum can be stored locally on the device. In some implementations, the at least one datum can be stored remotely, such as on a companion device (e.g., smartphone coupled via wireless connection) or a server system (e.g., cloud storage system). The remote device can be communicatively coupled to the wearable device using any wireless or wired protocol.

[0046] Returning to the example of the table, the user can provide speech that indicates an interest in the table (e.g., “I really like this table’'). Based on natural language process and a model that indicates that at least one datum should be stored with the table, the device can store the at least one datum associated with the object. The at least one datum can include an identifier for the table (i.e., object), manufacturer, dimensions, or any other datum associated with the object. In some implementations, the at least one datum is expressly provided by the user. For example, the user can provide an identifier or model name associated with the table. In other implementations, the at least one datum can be determined from one or more sensors on the wearable device. For example, the device can use cameras or speakers to identify' information about the object. Referring to the table, the device can capture an image of the table and perform a reverse image search to identify the at least one datum. The device can then store the information for the object for a later interaction with the user. In some implementations, the at least one datum is stored indefinitely. In other implementations, the at least one datum is stored for a threshold period (e.g., a number of objects or a time period) before it is removed from storage. As an example, the device can remove data associated with the table following a threshold period when the user does not reference a table in a following request.Atty Docket No. 0120-1234W01

[0047] FIG. 2A illustrates method 200 of determining a result for a voice request based on device state according to an implementation. In some examples, method 200 can be performed by a wearable device, such as device 130 of FIG. 1 or computing system 1200 of FIG. 12.

[0048] Method 200 includes receiving audio input from a user of a device at step 201. In some implementations, a device can be configured with one or more microphones configured to receive voice requests or speech (i.e.. audio) from the device user. Method 200 further includes, in response to the audio input, determining whether the device is in a first state or a second state based on content displayed by the device at step 202. In some implementations, the display of the device can be configured to provide content using VR that immerses a user in a 3D environment or AR that adds digital content onto the real world. When the content is in VR the device can be determined to be in a first state, while when in AR the device can be determined to be in a second state. As used in this specification, a "state" is a snapshot of a device's current operational and / or environmental context, used to dynamically tailor the content and format of a response to a user. The state can be determined by one or more factors such as the display mode, device location, and current time, allowing the same user input to generate different results under different conditions. To determine whether the device is in the first state or the second state based on the content displayed by the device, a predefined rule (e.g., an assignment table) or model may be used, which is provided with the content displayed by the device as input and provides an indication of the state as output.

[0049] Method 200 further includes determining a first result based on the audio input and the first state in response to determining that the device is in the first state, and displaying the first result at step 203. Method 200 also includes determining a second result based on the audio input and the second state in response to determining that the device is in the second state, and displaying the second result at step 204. As used herein, a “result” can be a set of renderable data, such as a two-dimensional user interface element or a three-dimensional model, that is selected or generated by a response application. The selection and generation process uses the user’s audio input and the device's determined state as inputs to provide a context-aware output for display. As a technical effect, a device can be configured to receive the same voice input and display differing results based on the content displayed by the device. For example, if the user provides voice input to “Take me to Las Vegas.” If the user is in a state associated with AR, then the device can open a mapping application and provideAtty Docket No. 0120-1234W01 directions to Las Vegas. The device can further provide other search results associated with traveling to Las Vegas, including flights, trains, etc.

[0050] In contrast, if the user is in a state of VR, the device can request and receive from a search engine 3D display information associated with the city and generate a VR rendering of the city. This enables the user to view or be immersed in the requested location. In some examples, the device can be configured to determine the level of immersion (e.g., AR or VR) currently provided by the device and determine the result based on the level of immersion.

[0051] In some examples, the device can use a machine learning model that receives the voice input and the state associated with the content (i.e., level of immersiveness). The model can be configured (i.e., trained) to associate input parameters with corresponding visual outputs. Thus, while a map can be provided in association with a first state, a more immersive presentation can be provided to the user when the device is in a second state. In some examples, the device can consider additional factors, including preferences of the user, location of the device, environmental surroundings of the user (e.g., sitting on the couch or walking), the types of applications that are available to support the voice request, or other context information associated with the device. Based on the context, the device can determine what result is provided to a requesting user. For example, based on an available application and the immersive state, the device can provide a first result using the available application. However, if the application is unavailable (not installed or not available in the immersive state), the device can provide a different result (e g., opening another application).

[0052] FIG. 2B illustrates method 250 of using a data store to determine and display results according to an implementation. In some examples, method 250 can be performed by a wearable device, such as device 130 of FIG. 1 or computing system 1200 of FIG. 12.

[0053] Method 250 includes storing at least one datum associated with a first interaction between a user and a device at step 251. The term “datum” can denote any piece of information, such as a name, an attribute, or a timestamp, that is stored in association with an entity (e.g., an object or event) to be used as a reference in future interactions. In some implementations, a user can interact with a device by providing speech (i.e., queries) and receiving results to the speech. For example, a user can request the best burger places in Denver. In response to the request, the device can provide search results. As part of the interaction, the user can provide additional speech input to identify specific restaurants of interest, menu items of interest, pictures of menu items, locations of restaurants, and the like. In some examples, the device can store various types of information, such as the spokenAtty Docket No. 0120-1234W01 words (transcript or log), the user’s intent (what they want to do like identify the best burger location), context (like time, location, or current application used for the interaction), system response, and metadata such as confidence scores, speech duration, and background noise. The device may also store user preferences, corrections, or feedback.

[0054] After the data is stored in association with the first interaction, method 250 further includes determining context associated with the device at step 252. The context can include the device’s location, a level of immersion associated with content displayed by the device, a user’s movement associated with the device, applications executing on the device, or other contexts, including combinations thereof. For example, returning to the burger example, the device can determine that the user is walking in Denver and has a mapping application displayed on the device.

[0055] Method 250 further includes determining a result based on the context and at least one datum at step 253 and displaying the result on the device at step 254. For example, if the user is walking in Denver, the device can search the stored information about the first interaction to determine relevant information to provide to the user. The information can include locations identified as part of the best burger search, a map of the locations, menus, or other options. In some implementations, the device can be configured to provide the user with a suggestion about burgers (e.g., “Would you like help finding that burger you were looking for?"’). In response to the selection of the notification, the user can be provided with additional information derived from the stored data of the first interaction. In some examples, the result can be updated as the user moves or provides input to the device. For example, the device can be configured to provide a map to the burger location and, when the location of the device reaches the restaurant, can provide a user with a menu. As a technical effect, the device can change the result provided to the user based on the context associated with the device.

[0056] In some implementations, the device can further use speech input from the user to determine the results that the user provides. For example, the user can give input requesting, “What was the burger place that I was looking up?” The device can be configured to access the data store and determine a name, location, menu, and the like associated with the burger place and provide at least a portion of the information to the user. Based on the context of the device (e.g., the location), the device can provide directions if the user is near the restaurant. As a result, the device can use saved data associated with interactions with the user, context from the device, and speech input of the user to provide contextually relevant information to the user.Atty Docket No. 0120-1234W01

[0057] In another illustrative example, the device can be configured to provide a first interaction with a user, which may include voice input (via a microphone) or text input (via a controller, keyboard, or another input device). The user can provide direct requests (e.g., “store information about ‘X’”) or determine intent to store data about an object (e.g., a user focusing on an object for a threshold period determined through one or more sensors on the device). For example, the user can focus on a couch for a specified period, and the device can perform a reverse image search in association with an image captured from an outwardfacing camera. In some examples, the object of focus can be determined based on the user’s gaze or head position.

[0058] Once the object is identified, the device can store at least one datum associated with the object in a data store to be used with a future interaction with the user. For example, when the user views an object (e.g., a table) for a threshold period, the system can perform a reverse image search and store at least one datum associated with the table (e.g., a unique identifier). In some implementations, the datum can be provided by a reverse image search and a search engine. After being stored, the user can perform a second interaction to receive information associated with the table visually. The information can be provided in a variety of formats or configurations based on the device state. For example, when in AR, the device can provide first content visually to the user. However, when in MR, the device can provide second content to the user. The content can reflect the intention of the user, which can be inferred at least in part on the device state. In some implementations, the device can be configured to use additional parameters or factors, such as the location of the user, the time of day, or other factors in determining the state of the device. As a result, in a first state, the device can provide results in a first format or configuration and provide results in a second format or configuration in a second state. In some implementations, the device can provide different information depending on the device state. For example, when in MR, the device can provide a visual representation of the object, while in AR, the device can provide a summary of information about the object (e.g., a webpage using a search from the at least one datum). The information provided to the user in association with the request can be determined from the intent (i.e., inferred from the device state). In some examples, the format or configuration of the information can be inferred from the device state. As a technical effect, the device can be configured to provide first content when the device is in a first state and second content when the device is in a second state.

[0059] For example, a user can provide a verbal request to the device to “remember this concert” when viewing a poster of the concert. The device can then store at least oneAtty Docket No. 0120-1234W01 datum about the concert (i.e., entity) in a data store. Once stored the user can then generate a second request during a second iteration, such as "‘show me that concert I was looking at.” In response to the request, the device can process the natural language and determine the intent of the user based at least in part on the device state, and provide a response based on the device state.

[0060] In at least one implementation, the device can determine the state based on content displayed. The content displayed can indicate whether the user is in an AR, MR. or VR state. In some implementations, the device can determine the location of the device and determine the state of the device based on the location. For example, the device can be in a first state when the user is at home and a second state when the user is in the office or in a public space. In some implementations, the device can determine the state based on the current time, wherein different times can correspond to different states.

[0061] Returning to the example of the user requesting information associated with the concert. The device can determine a first state when the user is viewing VR content, and display a result based on the first state and the request. For example, the device can generate a display of the concert that immerses the user in a song performed at the concert. The device can further be configured to determine a second state when the user is viewing AR content or is in an AR mode associated with the device. In response to determining that the device is in the second state, the device can determine information associated with the concert and provide the information to the user (e.g., as a summery). In at least one implementation, the information can be determined based on the datum stored for the event (e.g., event name, date, and the like). In some examples, at least some of the information provided to the user as a result can be determined from a search engine using the at least one datum stored for the event. For example, the device can store the name of the event based on text extracted from an outward-facing camera on the device. The device can then perform a web search using the event name to determine additional information that can be provided to the user as part of the result. As a technical effect, the immersiveness of the result provided to the user can depend on the displayed content on the device, the user's location, the time of day, or other state information associated with the device.

[0062] FIG. 3 A illustrates an operational scenario 300 of providing content to a user based on device state according to an implementation. Operational scenario 300 includes display view 305 and display view 306. Display view 305 and display view 306 include application 311 and application 312. Display view 306 further includes result 320. Display view 306 is representative of an updated view of display view 305.Atty Docket No. 0120-1234W01

[0063] In operational scenario 300, a wearable device identifies an audio input from a user at step 330. The wearable device further identifies the status of the device and the display at step 331. In some implementations, the status of the display includes determining whether the user is viewing content in VR or AR. Operational scenario 300 further comprises generating a result based on the status of the display and the voice input at step 332. In some implementations, the device provides different responses based on the state of the display. The user can provide the same voice input, but be provided with two different types of results based on the state of the display. For example, while result 320 is provided in operational scenario 300, the result could include a VR response based on the state of the display. In some implementations, the system can use a machine learning model to determine result 320. The model can be configured based on pairing screen state and voice inputs to corresponding output responses. Thus, while an audio input may correspond to a first output display (e.g., VR display) based on first context for the device, the voice input may correspond to a second output display (e.g., AR display) based on second context for the device. Although it was demonstrated as using the display state to determine the result, a device can use additional context to determine the result. The additional context can include the environment of the user, the location of the user, movement associated with the user, or other information associated with the device.

[0064] As an illustrative example, the user can be sitting at the desk in a first display mode (e.g.. VR) and provide a request to "take me to Paris.” In response to the request, the device determines a display state and determines a result based on the display state or content on the display. For example, if the user is in a VR mode, then the device can display content in VR (e.g., an immersive view of the city). If the user is in an AR mode, then the device can display a map or other information in the city using a 2D presentation. The device can further consider other state factors associated with the device.

[0065] FIG. 3B illustrates an operational scenario 350 of providing results to a user according to an implementation. Operational scenario 350 includes display view 356 and result 370. In operational scenario 350, a device stores interaction data at step 380. The interaction data may include logs associated with the user interaction (user requests and responses, including search results), can consist of objects or locations of interest, images, labels, or other information associated with the user interaction with the objects. For example, the interaction can include information about a user attempting to identify the best bagel in New York. The information can include search results, restaurant names, locations, menus, or other information associated with the interaction.Atty Docket No. 0120-1234W01

[0066] Operational scenario 350 further includes identifying context from the device at step 381 and generating a result based on the context and the stored interaction data at step 382. In some examples, the context includes the location of the device, the movement of the device, voice information from the user, or other information associated with the device. For example, returning to the example of the bagel interaction, when the user is near a bagel restaurant, the device can determine the location of the device and provide the user with information corresponding to the bagel restaurant referenced in the previous search. The information can include directions, a menu, or other information associated with the restaurant. In some examples, the information can be updated as the context changes for the device (e.g., the user approaches the restaurant). In some implementations, the user can expressly request information associated with the previous interaction, and the device can provide information associated with the request. At least a portion of the result can be provided based on the device’s context (e.g., the location, movement, environment, etc.).

[0067] FIG. 4 illustrates an operational scenario 400 of selecting a display of content based on device state according to an implementation. Operational scenario 400 includes user 410, device 430. voice input 444, display view 401 (representative of an MR view) with result 402, and display view 405 (representative of an AR view) with result 406.

[0068] In operational scenario 400, user 410 provides voice input 444 that is processed by device 430. Device 430 can represent a wearable device, such as an XR device or smart glasses. Device 430 can process voice input 444 to determine an action associated with either display view 401 or display view 405. In some implementations, device 430 can be configured to perform natural language processing to determine an intent associated with user 410. In some examples, device 430 can determine user intent for an action from voice input by capturing commands using an onboard microphone, converting the speech into text using a speech recognition model, and analyzing the resulting text with a natural language processing application. The system interprets the meaning and context of the spoken words to infer the user’s desired action, such as displaying or interacting with digital content. The inferred intent is then matched to a corresponding command or function on the device to execute the requested task.

[0069] In some examples, device 430 can be configured to determine the content and the display of the content based on the device’s state. In some implementations, the state can be determined based on the display mode of the device (e.g., MR, XR, or VR). For example, the device 430 is capable of operating in an AR state and / or the device 430 is capable of operating in an MR state and / or the device 430 is capable of operating in a VR state (e.g., ARAtty Docket No. 0120-1234W01 and MR, or AR and VR, or MR and VR, or AR, MR and VR). In some examples, the state can be determined by additional factors (or context), such as location, time of day, and the like. In some examples, the content can be determined based on the state. For example, the device can determine first content and a first format (result 402) in display view 401 and second content and a second format (result 406) in display view 405.

[0070] For example, in display view 401 and result 402 can represent an MR result that blends a 3D object or interactive element in the physical environment. For example, if voice input 444 included a request to ‘’Let me see that mousepad,” device 430 can identify the device state (i.e., an MR state) and display result 402 that provides a virtual representation of the mousepad. In some implementations, the 3D representation can be based on information associated with the object (e.g.. expressly provided by the user or obtained via one or more sensors), or can be determined based on at least one web search. In some examples, the user can expressly indicate the object or objects for display (e.g., let me see model X mousepad). In other implementations, the user can refer to stored data associated with objects on the device.

[0071] For example, the device can store at least one datum associated with entities (e g., objects, events, people, and the like). The device can store the datum based on a user request (e.g., a verbal request). In some examples, the device can infer the intent of the user to store information about an entity based on sensor data, user speech, and the like. For example, the user can speak about an object, such as a car, and the device can store at least one datum associated with the car based on the user's speech. If the user provides voice input of "I really like this car", then the device can store at least one datum about the car, such as an identifier, manufacturer, or other information about the car. In some implementations, the at least one datum can be expressly provided via the user, while in other implementations, the at least one datum can be determined from other sources, such as a web search associated with the object. In some examples, the information can be identified using sensor data from at least one sensor, such as a camera. The gathered sensor data can be processed to determine the at least one datum or generate a web search to obtain the at least one datum. For example, the device can capture an image of the car and generate a reverse image search with the image to identify the manufacturer of the car. The device can then store the search results as the at least one datum associated with the entity (i.e., the car).

[0072] Referring to operational scenario 400, in display view 405 and result 406, when device 430 determines that the device is in a second state (e.g., an AR state), result 406 can be determined based on voice input 444 and the device state. In some implementations.Atty Docket No. 0120-1234W01 result 406 can be different than result 402 (displayed in display view 401) based on the device’s state. For example, result 406 in display view 405 can represent a text-based summary presented as part of an AR application window for the user. As a technical effect, when the user provides voice input 444, device 430 can select the appropriate result 402 (represented in display view 401) or result 406 (represented in display view 405) based on device state. The different results can provide different information or a different visual representation to better reflect the cunent state of the device and the user's intent. For example, the number of dimensions (e.g., 2D or 3D) of the displayed result 402, 406 can be selected based on the state.

[0073] In another illustrative example, the user can request that the device store information associated with a car (e.g., a model or manufacturer for the car). The information can be obtained from the user’s input or from sensor data, such as an outward-facing camera. Once the information is identified, the information can be stored for the user. The information or data can be stored locally on the device or can be stored on an external device, such as a server. In some implementations, the user can expressly request the entity be stored (e.g., a voice request as part of an interaction with the device) or can be inferred based on the language or actions of the user. For example, the user can focus their gaze on the car for a threshold period. When stored, the at least one datum associated with the car can include one or more attributes for the car (e.g., make, model, etc.), the location associated with adding the car (e.g., a car dealership, town, coordinates, and the like), a time stamp, or other information associated with storing the information about the vehicle. The device can be configured to use the location information, time information, and the like to identify a user’s reference to the entity. For example, the user can provide voice input for "What was the car I was looking at yesterday?” The device can then search the data store for a car that was viewed by the user yesterday to identify the at least one datum associated with the car (e.g.., using the timestamp for the car). The device can then be configured to generate a result based on the state of the device, the request of the user, and the identified datum in the data store.

[0074] In some implementations, the result can include a first level of immersiveness, such as displaying the car in the user’s driveway when at home (e.g., mixed reality). The display of the result (and the information provided) can be based on currently displayed content and the location of the device in some examples. In another implementation, the result can include a second level of immersiveness, such as displaying a list of information about the vehicle as part of an application window in AR. The result can be determined based at least on currently displayed content, and may further be determined based on the user’sAtty Docket No. 0120-1234W01 location, time of day, or other factors. In some implementations, the device can be configured to provide the information in different visual formats to suit better the current state of the device (location, time, content displayed, etc.)

[0075] FIG. 5 illustrates an operational scenario 500 of using a data store to determine display results according to an implementation. Operational scenario 500 includes interactions 520, 521, and 522, entity data 530, 531. and 532, data store 540, user 510, voice input 544, and response 560. The steps of operational scenario 500 can be performed by a wearable device to support user requests. In some implementations, at least a portion of the operations can be performed using a combination of a wearable device and a companion device (e g., smartphone or tablet).

[0076] In operational scenario 500, a device can support interactions with a user to request information (e.g., provide a search), store information, or offer other services to the device. In some implementations, the device can store information based on user interactions 520, 521, 522. The interactions may include a direct request from the user to store data about an entity or can comprise an inferred request to store data about an entity. For example, the user can expressly indicate “Remember this show7’ when viewing a poster. The device can be configured to use an outward-facing camera to capture an image of the poster and extract at least a portion of data about the entity7(i.e., show) from the poster. The information can include an identifier, time, date, or any other information derived from the poster. The entity data (e.g., entity data 530) can further include information derived from at least one search (e.g., a web search). For example, the device can perform a search of a band name identified in the poster to derive a portion of the entity7data.

[0077] In some implementations, the entity7data is stored for a threshold period (e.g., two weeks) and then removed from data store 540. In some examples, the data can be purged from data store 540 when the user does not use the data within a threshold period. In some examples, the data can be purged based on a user request.

[0078] After entity data 530, 531, and 532 is stored in data store 540, the device can receive voice input 544 from user 510. The device can perform request processing 550 to identify relevant context in data store 540. For example, the user can provide input associated with “Tet me see the show that I was looking at yesterday.” The device can perform natural language processing on voice input 544, then search data store 540 for the relevant information. In at least one example, the device can use the timestamps to identify entities saved from yesterday7and labels or data associated with the entities to identify the entity relevant to the show. The device can be configured to identify7the entity within the period andAtty Docket No. 0120-1234W01 the entity that corresponds to a show. Once the entity data is identified, request processing 550 can generate response 560 based on the entity data. For example, if entity data 530 is identified in association with voice input 544, then the device can generate response 560 based on entity data 530.

[0079] In some implementations, the device can be configured to format or configure the display of response 560 based on the content currently displayed by the device. Response 560 can be formatted based on whether the device is in AR. VR, or MR state in some examples. In some implementations, the format or configuration can be based on the location of the user (e.g., work or home). In some examples, the format or configuration can be based on the time of day. In some examples, the format or configuration can provide different information about the object. For example, in a VR state, the device can display an immersive presentation of the show (e.g., like being at the show). Alternatively, in an AR state, the device can provide an application window that provides information about the show, such as date, time, location, etc. At least some of the information can be derived from a search in some examples.

[0080] In some implementations, the examples described here can include systems and techniques that enable user interface (UI) control via a spatial action machine learning (ML) model. For example, the systems and methods described herein allow use of one or more ML models to utilize voice inputs, screen states, sensor inputs, and / or semantic graph(s) of applications to determine and provide (e.g.. execute) an action(s) desired by a requesting user with respect to one or more UI(s).

[0081] At least one technical problem solved by the described techniques includes controlling UIs of applications. At least one technical problem solved by the described techniques includes providing user interface control for XR devices. At least one technical problem solved by the described techniques includes integrating or otherwise using multiple sensors to determine a user intent with respect to one or more UIs, while minimizing a number of steps required to be taken by the user to obtain a desired result.

[0082] At least one solution to the above and other technical problems includes providing one or more ML models and associated components that are enabled to interpret various types of inputs and determine associated actions, and that are enabled and authorized to execute the determined actions. Such solution(s) include determining a semantic graph(s) that represents an application(s) and / or UI(s), including representing a name, semantics, and function(s) of Ul / application elements using corresponding nodes of the semantic graphs, while representing relationships between such nodes using edges of the semantic graphs.Atty Docket No. 0120-1234W01Then, the ML model(s) may be configured to process graphical inputs, to thereby process the semantic graphs in conjunction with various other potential inputs. Such potential inputs may include, e.g., voice requests from a user, motion / position data from motion sensors, image data from image sensors, gaze tracking data from gaze tracking sensors, or screen state information from an operating system of a device providing the UI(s) / application(s).

[0083] Described solutions thereby provide users with desired results, while requiring a minimum of information and explicit instruction from users. For example, while looking at a particular image, UI screen, or real world object, a user may ask “what is that?”, and described techniques may determine the object of the query', determine an associated application / UI to use to provide further information, and enter a relevant query into the determined application / UI to thereby display a desired response. As another example, the system can identify a user request, determine a device state based on content displayed by the device, and generate a result based on the state. In some examples, different states can provide different results to the same user input, where the different results can include a different level of immersiveness in some examples. The different results can also provide different information or visuals associated with the same entity (e.g.. object, event, and the like).

[0084] In other examples, described techniques may be used to provide results directly , effectively skipping over intermediate steps that would otherwise be required for the user to perform to obtain a desired result, and even if the user is not aware of the steps that would be needed to obtain the desired result. For example, a ML model constructed using described techniques may utilize a semantic graph of a system file structure. Described techniques may thus directly provide any available outcome within the file structure. The system can also use a semantic graph to associate and identify display configurations for different sets of applications.

[0085] For example, in a simplified example, the file structure may include a path for “connections” that includes “Bluetooth” and that further includes “connecting a new Bluetooth device.” The user may thus simply specify a request to connect a particular Bluetooth device, and the described techniques may then be used to establish the requested connection, and to do so in one step from the point of view of the user. In other words, the user is not required to navigate through, be aware of, or even be able to navigate through the various steps needed to establish a new' Bluetooth connection. Similar operations can also be performed to select a display configuration for a set of applications.Atty Docket No. 0120-1234W01

[0086] As illustrated by the preceding examples, and as described in more detail, below, described techniques may thus be used to infer a nature of a user’s request from a minimum of cues or other input or detail from the user. Described techniques may be further configured to provide a desired result to the user with a minimum of effort and / or knowledge being required from the user, including, e.g., executing one or more UI actions on the behalf of the user. Accordingly, users may be provided with experiences in which the users are able to reach desired outcomes in a fast and efficient manner.

[0087] In conventional systems, interactions in spatial computing devices (e g., extended reality (XR) headsets) are enabled by a breadth of input / output capabilities (e.g., hand tracking or eye tracking), but nonetheless introduce high user friction, such as requirements on high physical motion and limited precision and / or bandwidth of inputs.

[0088] Described techniques provide interaction paradigms for spatialized computing, including using the uniqueness of the I / O (multimodal signals in real time) and solving for friction and precision. Such interaction paradigms redefine interactions with computing systems, including shifting away from sequential interactions to accomplish tasks, towards enabling users to directly move to those end states immediately.

[0089] One input paradigm is based on voice inputs, which may use less physical effort than other input modalities, in combination with a backend that enables semantic understanding of a relevant system(s) and interface(s), and is thus able to take context-aware actions and retain memory to perform complex tasks and interactions in support of an end to end user interaction(s).

[0090] Existing action-oriented voice interfaces may rely on a hardw are-first verticalized approach, which exist as standalone applications or interfaces that do not contain context of the system, connection points across applications, or context across user interactions. Such conventional approaches are thus limited in their capabilities based on the architecture(s) of the solution(s) and therefore may not be able to perform complex tasks that enable broader user interactions and new' interaction paradigms.

[0091] In contrast, described techniques redefine a human computer interaction paradigm, including spatial computing systems, and extending beyond to other form factors and platforms. Specifically, described techniques cover the usage of voice input pow ered by artificial intelligence (Al) that supports, e.g., conversational latency, extensive context length, and action-capable outputs, and that may be integrated deeply into the operating system of a device to support interactions on behalf of the user.Atty Docket No. 0120-1234W01

[0092] Relevant components may include, e.g., a voice interface, including transcription and parsing or prompt tuning. Components may include embedded Al, including spatial action models, as described herein, and which may support various inputs (such as, e.g., a spatial scene, context images, user behavior, and pointing locations) in addition to text input transcribed from voice inputs. With these and / or other inputs, a spatial action model may provide outputs as machine language that are mapped into interactions for the system experience.

[0093] In particular, action-based outputs, embedded into the system, may be provided by including wrapper models around or within the spatial action model(s) to be able to effectively output interactions. Deep integration points within the operating system may be provided to provide support for variants of interactions, including, e.g.. in-app multi-stage user interactions, system interactions happening in parallel to user actions, and interactions that happen asynchronously to the user’s supervision, but with the approval of the user.

[0094] For example, such interactions may be described in various interaction tiers, for the sake of illustration and example. For example, a single panel, single application, single shot input may result in reaching a defined goal in the context of a multi action end state. For example, a user may search a video in a video application from no initial state. In another example, a user may search a restaurant within a maps application, merely by looking at a panel or screen of the maps application and saying, “I want to book a restaurant for Sunday around <time range>, and I like this type of food <context>.”

[0095] In another example, a user of an email application may request an identification and summary of most relevant emails, including, where appropriate, initiating draft response emails, or opening relevant documents or slides. In other examples, a user maywish to plan a trip in an efficient manner and may request assistance in planning the trip to a specified location, for a specified time, and including activities that are based on the user’s private preferences, e.g., “make me a trip plan to go to Italy- for 5 days based on my preferences. Here are some things I want to make sure I cover: ().” In these examples, the spatial action model described herein may take multiple actions across multiple websites and applications, thereby queuing up a summary for approval with a single shot booking proposal by the user.

[0096] In another example, a user may wish to purchase an item if / when a price of the item reaches a defined price. For example, the user may input, “I am looking for a new pair of sneakers, but they are currently out of my budget. Can you keep track of these shoes for me and buy them on my behalf across any site, once they have reached below <price>. OrAtty Docket No. 0120-1234W01 alternatively, give me a prompt when a new pair is released and purchase if <conditions> are met.

[0097] In another example, a user can provide a request associated with an entity (e.g., an object), such as “Let me see that couch I was looking at.” The device can be configured to process the speech input to identify user intent and generate a result that corresponds to the request. In some implementations, the result can be determined based at least one current content displayed (or type of content, such as AR. XR, or MR). In some implementations, the result can be determined based on the location associated with the request or the time of the request. In some examples, the screen state, location, time, and the like can be used to define a device state, and the response can be generated based on the device state. In some examples, context for the request can be defined via a data store or sensor data to identify ambiguous or referenced terms of the user, such as “that couch.” The display format or configuration for the response can be based on the device state in some examples.

[0098] In order for a multimodal spatial action model as described herein to provide these and other types of control of a system UI, the model may utilize a persistent memory of a relevant UI framework. For example, for one or more UIs and / or applications, a semantic graph may be maintained in which nodes represent UI states and edges represent input events. Graph nodes may then be associated with high-level semantics (e g., learned offline).

[0099] In addition to the graph, visual and voice multimodal input may be mapped to these semantics to enable interactions in the system. Interactions may also be triggered by application intents, representing one instantiation of a depth-1 graph / tree. Such semantic graphs can be constructed using various techniques, some of which are described below as examples.

[0100] In some implementations, a wearable device, such as computing system 1200 of FIG. 12, can include an action manager application (or an action manager) that can be configured to execute actions using a user interface in response to requests from the user. In some implementations, the actions can include responding to user requests as described herein (i.e.. in different visual formats based on device state). In some implementations, the device can be configured to identify device state based on displayed content (e.g., AR, VR, or MR content), the location of the device, time of day, or another factor. Based on the state of the device, the system can determine what content to be provided to the user and / or the visual appearance of the content for the user. For example, MR content can be provided in some state conditions, while AR content can be provided as a response in other state conditions.Atty Docket No. 0120-1234W01

[0101] In the present description, the term action can be understood to represent or include any functionality of a relevant user interface. Such actions may therefore include, for example, selecting a user interface element (e.g., clicking on a clickable element), opening an application and corresponding user interface, inputting text or other data into a corresponding field (e g., a text entry7box), or interacting with any interactable element of the user interface. Actions may also include, for example, summarizing, translating, explaining, searching, or otherwise processing input from the and / or portions of one or more relevant user interfaces.

[0102] Actions executed by the action manager may thus replace, augment, or supersede input primitives commonly used for human-computer interaction (HCI). In the context of conventional HCI, the user may be required to follow a sequence of such input primitives to reach a desired state of a user interface, or other desired result. For example, in conventional contexts, if the user wishes to order takeout food from a nearby restaurant, the user might have to, e.g., open a web browser, enter text to search for nearby restaurants, navigate a menu of a selected restaurant, select a food delivery7service, and finalize a purchase. In contrast, using the action manager, the user may simply specify, “order a French baguette from Panera Bread using Uber Eats,’7and the action manager may execute the series of actions referenced above to thereby7ultimately provide the user with a specified food item in a shopping cart of the specified food delivery7service. In some implementations, if the user desires, the action manager may be provided with the authority to execute the transaction to consummate the purchase, while in other implementations, the action of finalizing the transaction may be required to be executed or approved by the user.

[0103] As may be observed from the preceding example, the action manager may control actions among multiple user interfaces and associated applications. For example, the action manager may use an output or result of a first action of the user interface as an input at the user interface.

[0104] It will be appreciated that conventional, existing systems for executing actions in response to a user request, such as a voice request, generally use a static dictionary7or rulebook that enables implementation of an action in direct response to a user request. For example, a request such as “select the enter button” may be required to be implemented in a one-to-one fashion with execution of an action, so that a user is required to provide a series of instructions for actions that mirror the actions that such a user would have to take if using graphical user interface elements of the user interface.Atty Docket No. 0120-1234W01

[0105] In contrast, described techniques may provide a requested result, without requiring each intermediate action to be user-specified. As a result, user friction is reduced, and desired results may be obtained in a fast and efficient manner.

[0106] By virtue of the preceding examples, and following examples, the action manager may be understood to represent or provide a virtual assistant. Although various types of conventional virtual assistants exist, it will be appreciated that conventional virtual assistants do not provide the type of executed actions provided by the action manager. Such conventional virtual assistants further fail to provide the type of requested end results described herein, with the as-described ability' to reduce or omit intervening user interface interactions in providing such requested end results.

[0107] Further, the action manager may utilize various types of multi-modal inputs to infer or otherwise determine an action to be executed. For example, inputs may include voice or other audio inputs from audio sensors, images from image sensors, pose / position data from motion sensors, or gaze data from gaze tracking devices. Inputs may further include screen captures of the screen (including screen states), as well as application data, system data, or other data stored using the memory of the device. For example, such data may include graph data representing one of the applications on the device and / or corresponding UIs or graph data representing personal preferences of the user.

[0108] These and various other inputs may be used by the action manager to determine an action(s) desired by the user. For example, as described in detail, below, the action manager may include a spatial action model that models actions, e.g., determination, selection, and implementation of such actions, with respect to the various types of inputs just referenced, and other ty pes of inputs.

[0109] As a result, the user may be provided with desired actions quickly and efficiently, with minimal effort and minimal knowledge required. Consequently, the user may experience improvements, e.g., in learning, productivity', entertainment, communication, and health / wellness, some examples of which are provided herein.

[0110] FIG. 6 illustrates a block diagram of a device 600 according to an implementation. In the example of FIG. 6. a device 600 represents any device that may be used to implement the action manager 620. For example, such devices may include smartphones, laptops, smartwatches, and many other devices, and combinations thereof. In some examples, action manager can be implemented by computing system 1200 of FIG. 12.

[0111] In FIG. 6. the device 600 is illustrated as including a processor 622. a memory 624, sensors 626, and a screen 628. Device 600 further includes memory 624, sensors 626,Atty Docket No. 0120-1234W01 and screen 628. The processor 622 should be understood to represent any suitable processor(s) that may be used to execute instructions stored using the memory 624, including any relevant applications and the action manager 620.

[0112] The action manager 620 is illustrated as including a spatial action model 630. The spatial action model 630 should be understood to represent any suitable machine learning model and associated input / output layers, adapters, or wrappers that are trained to provide the type of multi-modal input processing described herein to thereby determine corresponding actions to be taken with respect to one or more user interfaces (and / or underlying application(s)). In some implementations, the system can change the display of applications or application windows on the device. In some implementations, the system can identify results to a query based on device state. The result can represent a different visual representation in some examples.

[0113] A screen state detector 632 may be configured to determine a current state of content of the screen 628. For example, the screen state detector 632 may determine one or more current user interfaces displayed using the screen 628, as well as current content being displayed, functions being provided, or instructions being executed. In the case of 3D XR environments, immersive applications may execute in a background state and may not be visibly displayed or rendered at a given point in time but may be captured and characterized by the screen state detector 632, as well. Moreover, as the screen 628 may display a current real-world view (e.g., in a passthrough mode of an XR device, or when viewing an image capture screen of a smartphone), the screen state detector 632 should be understood to capture or include real -world objects / views, as well as rendered content.

[0114] A saliency detector 634 may be configured to process a screen state determined by the screen state detector 632, in order to assess most-relevant or most-salient state data to be supplied to the spatial action model at a given point in time. For example, in a 3D XR immersive environment, the user may be provided with a 360-degree view and may potentially view multiple panels simultaneously or may view a single user interface that spans the entire field of view of the user.

[0115] The saliency detector 634 may be configured to determine, e.g.. based on various other inputs, including, e.g., voice inputs from the user, most-relevant portions of the screen 628 and / or provided screen content or data. For example, the user may express an otherwise ambiguous request, such as, “what is that?’", or “what do I see here?'’, and the saliency detector 634 may utilize gaze-tracking data from a gaze-tracking device of the sensors 626 to identify a restricted field of view most likely to correspond to the request. InAtty Docket No. 0120-1234W01 this way, the saliency detector 634 may effectively disambiguate the request in a fast and efficient manner and increase an accuracy of a provided response.

[0116] A user interface graph 636 refers to a graphical representation of one or more user interfaces and / or applications currently stored / available, e.g., using the memory 624. For example, all applications currently open, in use, or available for use may be included, including applications that may be executing on a second or remote device.

[0117] In the present description, a user interface graph refers to any graph that represents available or potential user interface states and associated actions that may be performed with respect to the user interface(s). Multiple user interfaces may be associated with one application, and multiple user interfaces across multiple applications may be included in the user interface graph 636. In the present description, user interfaces graphs may include, or be referred to as, action graphs, state graphs, or semantic graphs, which are examples or instances of user interface graphs and therefore may vary somewhat in terms of, e.g., how different types of such graphs are constructed and what types of user interface information is included in such graphs.

[0118] For example, the action manager 620 may have access to an applicationspecific semantic graph for many different applications that may be executable by the device 600, across many different users. Each such application-specific semantic graph may include varying types of data and / or levels of detail used to characterize a corresponding application.

[0119] For example, some commonly used applications and / or smaller applications may include detailed semantic graphs that characterize all or virtually all application functions or aspects. Semantic graphs for other applications may include a high level of detail with respect to top or high level application functions or user interface screens, with less detail included for lower-level functions / screens that are used less often. Further, for a single application layer or UI screen, some portions that are static or relatively less likely to change may be mapped and graphed in detail, while other, more dynamic portions may be graphed in less detail.

[0120] Such application mappings may also change over time. For example, as a new application becomes available or is added, an initial semantic graph may be added, as well. As one or more users use the application, the semantic graph may be updated and extended over time. Further, as a given application experiences updates or other changes, the corresponding semantic graph may be updated, as well.

[0121] In some cases, existing application graphs may be accessed, stored, or otherwise leveraged to construct a corresponding semantic graph for the correspondingAtty Docket No. 0120-1234W01 applications. For example, widely used applications may provide, or make accessible, publicly available application graphs that assist application developers in developing new, compatible applications, and such available application graphs may be used, e.g., augmented, to construct semantic graphs that are usable by the spatial action model 630. In some examples, the graphs can be used to determine how to generate a result for a user request, where the format or configuration of the result can be based on the device state.

[0122] The specific user, at a given point in time, may have some subset of applications stored using the memory 624 or accessed via a network, and currently open, active, or available. The user interface graph 636 (which may be a semantic graph or global semantic graph) may combine or simultaneously access corresponding application semantic graphs on behalf of the user, for processing by the spatial action model 630. Accordingly, the user interface graph 636 may be referred to as a global semantic graph or a user-specific semantic graph, to indicate inclusion of multiple application-specific semantic graphs. The graphs can be used to determine how to provide a result based on the device state (e.g., selecting a first state

[0123] For example, as described herein, the spatial action model 630 may utilize multiple applications to achieve a desired result or action. For example, the spatial action model 630 may traverse the user interface graph 636 using the output of a first application, obtained by executing a first action as input to a second application, thereby executing a second action. Thus, by traversing various portions of the user interface graph 636. the spatial action model 630 is capable of providing a desired result across multiple levels of multiple applications.

[0124] In addition, a preference graph 638 may be constructed with respect to the individual user. The preference graph 638 may provide, for example, preferences of the user with respect to desirable content, actions to be provided or avoided, and various manner(s) in which provided actions may be provided.

[0125] As a result, for example, the spatial action model 630 may provide actions even without a direct query’ from the user. For example, if image sensors and / or the screen 628 provide images of content that may be of interest to the user based on the preference graph 638, the spatial action model 630 may automatically select an application(s) and execute an action(s) to obtain relevant or useful information for the user.

[0126] Such information may be surfaced to the user immediately, or stored for later use, e.g., using an action memory 640. For example, if the user is wearing the device as AR glasses, the AR glasses may capture an item in the field of view' of the user that the user mayAtty Docket No. 0120-1234W01 not actively notice or designate, but that the spatial action model 630 determines to be relevant to the preference graph 638.

[0127] In these and other scenarios, the spatial action model 630 may thus output relevant information, e.g., by opening a relevant search application, entering text or images related to the recognized content, and storing received search results. The spatial action model 630 may execute these actions in real time and demonstrated to the user in a rendered user interface(s) or may do so in a background process not visible to the user. In the latter case, the spatial action model 630 may store search results and related information in the action memory' 640, which may then be reviewed later by the user in a batch fashion.

[0128] The action memory 640 may be used for other purposes. For example, the action memory 640 may be used to store session information for one or more action sessions of the user. For example, the spatial action model 630 may thus be provided with an ability' to use earlier action executions and related information as input(s) to determining a current action to be executed. For example, the user may ask an initial question about a book viewed using the screen 628 early in a session, and then later in the session, might ask a related question, such as, ‘’can you provide a summary of the book I asked about earlier?”, without having to view or otherwise designate the book again at that time. Accordingly, the user may again be provided with fast, efficient access to desired actions by the action manager 620.

[0129] FIG. 7 illustrates a first example implementation of a spatial action model. In the example of FIG. 7. at a device / hardware abstraction layer (HAL), one or more devices (e.g., XR devices) provide multi-channel audio, wide screen capture, and other ty pes of sensor data.

[0130] Sensor models corresponding to the various illustrated multi-modal inputs may then process the corresponding ones of the inputs. For example, a sensor model may process the multi-channel audio to denoise the audio and otherwise process and filter the audio to isolate user voice commands.

[0131] A sensor model may process widescreen capture data. The sensor model may relate visual / screen data to corresponding user interfaces and applications, and thus to corresponding application semantic graphs. Techniques for related visual / screen data to graph data are described in more detail below.

[0132] A sensor model(s) may process sensor data from one or more of the various ty pes of usable sensors, including image sensors, motion sensors (e.g., inertial measurement units (IMUs)), and gaze-tracking sensors. For example, sensor data may be filtered or otherwise processed, including various types of sensor fusion.Atty Docket No. 0120-1234W01

[0133] Corresponding encoders may be configured to input the processed sensor data for subsequent input to a visual language model (VLM). For example, an encoder, e.g., including a conformer, may be configured to model local / global audio sequence dependencies of audio sequences received from the underlying audio sensor model processing the multi-channel audio.

[0134] An encoder may receive screen / visual / graph data for tokenization and input to the VLM. The encoder may be implemented, for example, using a contrastive captioner (Coca) in PyTorch with vector quantization (VQ). The encoder may be implemented as one or more translators for a large language model (LLM), such as the VLM, to tokenize graph data, where the encoder weights may be fine-tuned using, e g., XR / AR specific data (e g., images, screens). Further, various types of weight modification, such as Low-Rank Adaption (LoRA) may be used to effectively reduce a number of parameters required for training and inference, thereby providing efficient fine-tuning without retraining an entire model, and reducing associated memory requirements for storing the trained encoder.

[0135] One or more customer encoders may be used to convert processed sensor data for input to the VLM. For example, different encoders may be used when underlying sensor data incudes image, pose / position, and / or gaze sensor(s).

[0136] In FIG. 7, the VLM may thus represent any suitable visual large language model compatible with inputs from the encoders. Also in FIG. 7, prompt data may be used so that prompts for the VLM are customized and engineered to obtain desired results. For example, prompt templates for commonly requested or needed actions may be provided. Further, the VLM may be trained to provide desired actions based desired prompts or prompt templates. Outputs of the various encoders may be combined with, or modified by, the stored prompts to improve outputs of the VLM.

[0137] Retrieval augmented generation (RAG) data may also be used to improve or facilitate operations of the VLM in conjunction with outputs of the various encoders. RAG data, for example, may include application data for many different applications, some or all of which may be included in a global semantic graph of a user, such as the user. In this way, as described in more detail, below, requirements for creating the global semantic graph(s) may be reduced, and existing application data / graphs may be leveraged to ensure availability and use of many different applications in the system of FIG. 7.

[0138] A customer wrapper may be provided that translates textual and / or image outputs of the VLM into the types of actions described herein, including, e.g., including opening applications, entering text into user interfaces, making selections, generating virtualAtty Docket No. 0120-1234W01 assets, or providing pointers to real world objects. For example, the customer wrapper may be implemented by translating textual outputs of the VLM to Python or other suitable code, as input to a suitable application program interface (API), and / or as a generated API.

[0139] To determine an appropriate action, the wrapper may be implemented as a decoder that is a classifier that classifies outputs of the VLM into at least one action class. In such examples, the w rapper may be implemented as a small footprint neural netw ork that converts text strings into one of a plurality of action classes.

[0140] In the example of FIG. 7, the various encoders may each have a control token(s) when inputting to the VLM. Thus, any of the encoders may invoke an action. For example, the graph encoder encoding semantic graphs of associated user interfaces may include a control token, so that, as referenced above, input of a screen or other image, by itself or in conjunction w ith other data (but not requiring voice data or other direct or synchronous input from the user), may invoke an action(s). For example, as referenced, the preference graph and a captured image / screen may be sufficient to trigger an action such as a search for data related to the captured image / screen, even without a corresponding request from the user.

[0141] FIG. 8A is a block diagram illustrating a second example implementation of a spatial action model. In FIG. 8A, a UI parser may be implemented, e.g., using unsupervised or supervised techniques. For example, in the latter case, appli cations / user interfaces may be human annotated for processing.

[0142] For example, as showai in the examples of FIGS. 8C and 8D, UI screens may undergo a process of sparsification, in which extraneous images and data are removed and the remaining data is converted to formatted text, e.g., in extensible Markup Language (XML), and stored in XML files. Meanwhile, screen / image frames may be captured at a given frame rate per second (fps), e.g., using circular prediction mapping (CPM) for coding image sequences.

[0143] For example, image frames may be captured, and relevant positions and elements, such as those appropriate for action inference, of user interfaces may be identified. For example, a clickable or selectable element may be identified, and an associated position, e.g., in pixel coordinates, may be captured. Thus, for each image / screen, action-relevant portions may be quantified, identified, localized, and otherwise characterized.

[0144] A UI graph encoder may transform the XML files into graph-based representations, such as graphical embeddings or tokens, that are compatible with a large VLM. In this context, the large VLM should be understood to be of a size (in terms of theAtty Docket No. 0120-1234W01 number of parameters, memory' footprint, and associated processing resources) that can handle multiple different applications. For example, the large VLM may be implemented as a server-side VLM. In conjunction with one or more semantic graph reconstruction prompts, the large VLM may thus be configured to determine a semantic graph capable of representing any combination of processed applications.

[0145] As shown in FIG. 8B, the semantic graph may include many nodes, connected by edges representing relationships between the nodes. For example, each node may consist of a name of an element, including clickable elements or other controller elements. Each node may include semantics of its corresponding element, such as “select this option if. .. A node may contain information regarding neighboring nodes. A node may also contain coordinates, e.g., 2D or 3D coordinates, within a relevant user interface.

[0146] Further in FIG. 8A, a UI graph to text token encoder may interface with the semantic graph. A small VLM, e.g., operable on a user device of a user, may receive an audiovisual or other multimodal query from the user.

[0147] Thus, the small VLM may process relevant aspects of the semantic graph together with the multimodal query and action decoder, as an example of the type of customer wrapper, may be configured to determine a relevant action. In the example, the semantic graph may include a connection file structure for a device, including various ty pes of wireless connections. Therefore, like the example provided above, the user may provide the query, “establish a Bluetooth connection for this device”, and the semantic graph may be used to move to the appropriate connection screen and connect the device in the requested manner, without requiring the user to navigate through the connection file structure.

[0148] FIG. 8A illustrates that applications may be onboarded in a manner that does not require input or assistance from application developers, e.g., does not require an application to be constructed in a particular manner or include any particular metadata. XML representations of application pages with relevant pixel coordinates determined from synchronized CPM frames (which may include 3D or stereoscopic image frames) enable a graphical, 3D spatial representation of action-relevant user interface aspects, to thereby enable construction and updating of the semantic graph.

[0149] Many additional or alternative techniques may7be used to generate and maintain the semantic graph. For example, heuristics-based approaches may be used in which existing data on application is polled to determine application / user interface features and aspects with associated levels of confidence.Atty Docket No. 0120-1234W01

[0150] For example, as referenced above, and as may be observed with respect to FIGS. 8C and 8D, portions of some user interfaces may be determined to be relatively static over time and with respect to other portions of the same user interfaces. For example, in the context of an email application as shown in FIG. 8D, the structure of the page may be generally or relatively static, while content of individual emails may change rapidly / dynamically .

[0151] In many cases, an application may be onboarded with a minimum level of mapping to a corresponding semantic graph, and then improvements may be made over time to enable a more complete semantic graph for the application. For example, semantic graph updates may occur in conjunction with tracking user interactions across multiple users, and / or in conjunction with application updates.

[0152] Initial application onboarding may also be performed by application type, so that multiple applications of a similar type may be onboarded quickly. For example, again with respect to the example of FIG. 8D, multiple email applications may share similar features, e.g., related to an inbox, sending / receiving email, etc., and such similarities may be leveraged in generating an initial semantic graph that applies (with appropriate modifications) to multiple applications of the application type.

[0153] In some examples, RAG application data may also be leveraged. For example, application structural data may be available from application providers and can be used, e.g., as an external RAG data source, which can then be combined with the semantic graph (or subsets thereof) to determine appropriate actions.

[0154] During query7processing and other action determinations, the small VLM is thus enabled to perform various degrees of graph traversal to determine a correct and desired action. For example, when the semantic graph (including any relevant RAG data) includes a graphical path to a desired action, then the action decoder may determine, provide, and potentially execute a desired action without requiring the user to navigate through intervening graph nodes (e.g., user interface elements). Such traversals may occur, for example, when navigating through fully static frameworks, such as the example of establishing a Bluetooth connection.

[0155] In other examples, vary ing degrees of navigation steps may be requested to obtain a desired action. For example, in response to a user request, the small VLM and the action decoder may determine the farthest node available in an initial graph traversal. Then, the user may be presented with screen(s) in which the user may proceed by specifying moreAtty Docket No. 0120-1234W01 specific input primitives, such as, e.g., “select this button"’, until a desired end point action is reached.

[0156] As noted above, screen states may thus be relevant to decisions made by the small VLM in conjunction with relevant portions of the semantic graph. For example, a current state of a user interface screen may dictate a starting point for subsequent graph traversals.

[0157] In some examples, traversals may be initiated based on user-expressed intentions, without reference to any existing screen state. For example, even when no user interface is present or no application is opened, the user may request booking of a reservation at a specified time, date, and location, and the small VLM may. e.g., open a website or scheduling application, enter requested parameters, open other websites (e.g., nearby restaurants of relevant food types), select open times, and otherwise schedule the reservation.

[0158] In many cases, actions may be classified with respect to a degree of impact or other parameter that may be relevant to executing actions. Then, high-impact actions may be indicated to request approval from the user prior to executing a corresponding action. For example, actions related to purchases may require express approval from the user, while actions related to executing a search may not.

[0159] FIG. 8B is an example semantic graph that may be used in the implementation of FIG. 8A. FIG. 8C illustrates a first example of page sparsification that may be used in the implementation of FIG. 8A. FIG. 8D illustrates a second example of page sparsification that may be used in the implementation of FIG. 8 A. As shown in FIGS. 8B and 8C, static, structural elements of a user interface may be extracted, and individual pieces of content may be removed to be able to learn page structure that can then be embedded into the type of graph illustrated in FIG. 8B.

[0160] FIG. 9A is a block diagram illustrating a third example implementation of a spatial action model. In the example of FIG. 9A, a UI rulebook may be constructed offline that captures a set of desired and available actions. In contrast with the example of FIG. 8 A, the UI rulebook is not required to include a full semantic graph, but rather may include a set of prompts that correspond to specified actions and. when submitted to a multi-modal VLM in conjunction with screen video and a received voice query, cause the multi-modal VLM to perform the desired action(s) as corresponding UI events.

[0161] For example, the UI rulebook may include textual descriptions of actions expressed as. or describing, machine language (e.g.. an API call) for performing a corresponding action, such as opening a particular application. The few-shot prompting mayAtty Docket No. 0120-1234W01 include all actions that the user may want to invoke, with instructions to execute only those actions of the UI rulebook that relate to the screen video / voice query inputs.

[0162] Advantageously, the implementation of FIG. 9A can be constructed quickly and easily for various specific use cases and does not require the same type or extent of encoders, transformers, decoders, or wrappers in conjunction with the VLM. For example, the UI rulebook may be constructed entirely textually, without requiring a graph embedding to be processed accurately by the multi-modal VLM. Consequently, the approach of FIG. 9A may be deployed rapidly. On the other hand, the approach of FIG. 9 A may be more difficult to adapt or update over time, as new applications, features, or actions are developed, and may be more difficult to deploy widely across many different contexts and use cases.

[0163] FIG. 9B is a block diagram illustrating a fourth example implementation of a spatial action model. In contrast to FIG. 9A, the implementation of FIG. 9B utilizes a semantic graph or state graph, which may be processed by a graph-to-text decoder and provided to a multimodal VLM agent. The multi-modal VLM agent also receives a CPM stream of screen frames at a given frame rate (fps) and text converted from received speech of a user.

[0164] A VLM may thus process these inputs received view the multi-modal VLM agent, and following additional processing by a back-end multi-modal VLM agent, a text-to- action decoder may be configured to provide a traversal strategy defined with respect to the original semantic graph or state graph.

[0165] In contrast to FIGS. 8A and 8B, the implementation of FIG. 9B may not require or utilize a full semantic graph that specifies pixel coordinates of elements, e.g., in an XR environment. For example, the state graph of FIG. 9B may represent an available state graph of a given application or operating system feature(s), such as in the Bluetooth examples provided above. The state graph of FIG. 9B may grow or be modified or updated over time, based on user interactions with the state graph. Also, in contrast to FIGS. 8A and 8B, the implementation of FIG. 9B may require little or no special prompting or prompt engineering, so that it is not necessary to have a set of textual descriptions of (aspects of) desired actions.

[0166] FIG. 10 is a block diagram illustrating a fifth example implementation of a spatial action model. In the example of FIG. 10, a single model is constructed that includes a large multimodal model and an action model. The model of FIG. 10 may thus receive any number of XR action tokens associated with input events, visual tokens, and query tokens, and directly output or execute an action, e.g., input event for a relevant user interface.Atty Docket No. 0120-1234W01

[0167] In some examples, the implementation of FIG. 10 does not require separate custom encoders, decoders, or wrappers, nor does it require prompt engineering. On the other hand, the implementation of FIG. 10 may require comparatively greater efforts in training the large multimodal and action model.

[0168] FIG. 11 is a block diagram illustrating a sixth example implementation of a spatial action model. FIG. 11 utilizes the type of RAG approach described above, in which an external knowledge base is used to supplement or support semantic graph construction and processing.

[0169] For example, as referenced above, many applications may ship with, or otherwise have available, a graphical representation of their core functions and application states. By themselves, such graphs may not be suitable or sufficient for the types of action graphs described herein and may not be easily or sufficiently adaptable to provide described features. Moreover, even if such graphs can be adapted or generated, it may be problematic to incorporate all such applications that a given user(s) may desire to utilize in the context of a single global semantic graph.

[0170] Instead, as shown in FIG. 11. an initial inference of a desired action, or a desired subset of a graph, may be performed using RAG techniques. Then, a RAG-identified subgraph may be loaded and used to perform the types of further graph analysis described above. In various implementations, more or less local memory resources may thus be used, with corresponding decreases / increases in action latency, depending on user or administrator preferences.

[0171] FIG. 12 illustrates a computing system 1200 to provide display results according to an implementation. Computing system 1200 is representative of any computing system or systems with which the various operational architectures, processes, scenarios, and sequences disclosed herein can be implemented to change a display configuration based on user input. Computing system 1200 may represent a wearable computing device, such as an XR device or smart glasses. Computing system 1200 can include multiple computing devices in some examples (e.g., a wearable device and a companion device, such as a smartphone or tablet). Computing system 1200 includes storage system 1245, processing system 1250, communication interface 1260, and input / output (I / O) device(s) 1270. Processing system 1250 is operatively linked to communication interface 1260, I / O device(s) 1270, and storage system 1245. In some implementations, communication interface 1260 and / or I / O device(s) 1270 may be communicatively linked to storage system 1245. Computing system 1200 mayAtty Docket No. 0120-1234W01 further include other components, such as a battery and enclosure, that are not shown for clarity.

[0172] Communication interface 1260 comprises components that communicate over communication links, such as network cards, ports, radio frequency, processing circuitry and software, or some other communication devices. Communication interface 1260 may be configured to communicate over metallic, wireless, or optical links. Communication interface 1260 may be configured to use Time Division Multiplex (TDM). Internet Protocol (IP), Ethernet, optical networking, wireless protocols, communication signaling, or some other communication format, including combinations thereof. Communication interface 1260 may be configured to communicate with external devices, such as servers, user devices, or some other computing device.

[0173] I / O device(s) 1270 may include computer peripherals that facilitate the interaction between the user and computing system 1200. Examples of I / O device(s) 1270 may include keyboards, mice, trackpads, monitors, displays, printers, cameras, microphones, external storage devices, and the like.

[0174] Processing system 1250 comprises microprocessor circuitry (e.g.. at least one processor) and other circuitry that retrieves and executes operating software from storage system 1245. Storage system 1245 may include volatile and nonvolatile, removable, and nonremovable media implemented in any method or technology for information storage, such as computer-readable instructions, data structures, program modules, or other data. Storage system 1245 may be implemented as a single storage device but may also be implemented across multiple storage devices or sub-systems. Storage system 1245 may comprise additional elements, such as a controller to read operating software from the storage systems. Examples of storage media (also referred to as computer-readable storage media) include random access memory, read-only memory, magnetic disks, optical disks, and flash memory, as well as any combination or variation thereof or any other type of storage media. In some implementations, the storage media may be non-transitory. In some instances, at least a portion of the storage media may be transitory. In no case is the storage media a propagated signal.

[0175] Processing system 1250 is typically mounted on a circuit board that may hold the storage system. The operating software of storage system 1245 comprises computer programs, firmware, or another form of machine-readable program instructions. The operating software of storage system 1245 comprises result application 1224. The operating software on storage system 1245 may include an operating system, utilities, drivers, networkAtty Docket No. 0120-1234W01 interfaces, applications, or other types of software. When read and executed by processing system 1250 the operating software on storage system 1245 directs computing system 1200 to operate as described in the previously described FIGs 1-11.

[0176] Below are example clauses associated with the present disclosure. The described clauses should not be considered exhaustive.

[0177] Clause 1. A method comprising: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0178] Clause 2. The method of clause 1, wherein the first state includes a mixed reality7state, and the second state includes an augmented reality state.

[0179] Clause 3. The method of clause 1, wherein determining whether the device is in the first state or the second state is further based on a location of the device.

[0180] Clause 4. The method of clause 1 further comprising: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

[0181] Clause 5. The method of clause 4, wherein the at least one datum comprises an identifier for an object.

[0182] Clause 6. The method of clause 1, wherein determining whether the device is in the first state or the second state is further based on a time for the audio.

[0183] Clause 7. The method of clause 1 further comprising: receiving a request to store an entity; storing at least one datum associated with the entity; wherein the first state includes a mixed reality state; wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an alternate reality state; wherein the second result comprises a summary of the entity.

[0184] Clause 8. The method of clause 1 further comprising: capturing an image of an entity; determining at least one datum associated with the entity based on the image; storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.Atty Docket No. 0120-1234W01

[0185] Clause 9. A system comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the system to perform a method, the method comprising: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0186] Clause 10. The system of clause 9. wherein the first state includes a mixed reality state, and the second state includes an augmented reality state.

[0187] Clause 11. The system of clause 9, wherein determining whether the device is in the first state or the second state is further based on a location of the device.

[0188] Clause 12. The system of clause 9, wherein the method further comprises: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

[0189] Clause 13. The system of clause 12, wherein the at least one datum comprises an identifier for an object.

[0190] Clause 14. The system of clause 9, wherein determining whether the device is in the first state or the second state is further based on a time for the audio.

[0191] Clause 15. The system of clause 9, wherein the method further comprises: receiving a request to store an entity; storing at least one datum associated with the entity: wherein the first state includes a mixed reality state; wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an alternate reality state; wherein the second result comprises a summary' of the entity.

[0192] Clause 16. The system of clause 9. wherein the method further comprises: capturing an image of an entity; determining at least one datum associated with the entity based on the image; and storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.

[0193] Clause 17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processorAtty Docket No. 0120-1234W01 to perform a method, the method comprising: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

[0194] Clause 18. The computer-readable storage medium of clause 17, wherein the method further comprises: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

[0195] Clause 19. The computer-readable storage medium of clause 17, wherein the method further comprises: receiving a request to store an entity; storing at least one datum associated with the entity; wherein the first state includes a mixed reality state; wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an alternate reality state; wherein the second result comprises a summary of the entity.

[0196] Clause 20. The method of clause 17, wherein the method further comprises: capturing an image of an entity; determining at least one datum associated with the entity based on the image; storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.

[0197] In accordance with aspects of the disclosure, implementations of various techniques and methods described herein may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Implementations may be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine-readable storage device, a computer-readable medium, a tangible computer-readable medium), for processing by. or to control the operation of. data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). In some implementations, a tangible computer-readable storage medium may be configured to store instructions that when executed cause a processor to perform a process. A computer program, such as the computer program(s) described above, may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module.Atty Docket No. 0120-1234W01 component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed to be processed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0198] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is. therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the scope of the implementations. They have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components and / or features of the different implementations described.

[0199] It will be understood that, in the foregoing description, when an element is referred to as being on, connected to, electrically connected to, coupled to, or electrically coupled to another element, it may be directly on, connected or coupled to the other element, or one or more intervening elements may be present. In contrast, when an element is referred to as being directly on. directly connected to or directly coupled to another element, there are no intervening elements present. Although the terms directly on, directly connected to. or directly coupled to may not be used throughout the detailed description, elements that are shown as being directly on, directly connected or directly coupled can be referred to as such. The claims of the application, if any, may be amended to recite exemplary relationships described in the specification or shown in the figures.

[0200] As used in this specification, a singular form may, unless definitively indicating a particular case in terms of the context, include a plural form. Spatially relative terms (e.g., over, above, upper, under, beneath, below, lower, and so forth) are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. In some implementations, the relative terms above and below can, respectively, include vertically above and vertically below. In some implementations, the term adjacent can include laterally adjacent to or horizontally adjacent to.

Claims

Atty Docket No. 0120-1234W01WHAT IS CLAIMED IS:

1. A method comprising: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

2. The method of claim 1 , wherein the first state includes a mixed reality state, and the second state includes an augmented reality state.

3. The method of claim 1 or 2, wherein determining whether the device is in the first state or the second state is further based on a location of the device.

4. The method of one of claims 1 to 3 further comprising: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

5. The method of claim 4, wherein the at least one datum comprises an identifier for an object.

6. The method of one of claims 1 to 5, wherein determining whether the device is in the first state or the second state is further based on a time for the audio.

7. The method of one or claims 1 to 6 further comprising: receiving a request to store data about an entity; storing at least one datum associated with the entity;Atty Docket No. 0120-1234W01 wherein the first state includes a mixed reality state; wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an augmented reality state; wherein the second result comprises a summary of the entity'.

8. The method of one of claims 1 to 7 further comprising: capturing an image of an entity: determining at least one datum associated with the entity7based on the image; storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.

9. A system comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the system to perform a method, the method comprising: receiving audio from a user of a device; in response to the audio, determining yvhether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state: determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

10. The system of claim 9. wherein the first state includes a mixed reality state, and the second state includes an augmented reality state.

11. The system of claim 9 or 10, wherein determining whether the device is in the first state or the second state is further based on a location of the device.Atty Docket No. 0120-1234W0112. The system of one of claims 9 to 11, wherein the method further comprises: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

13. The system of claim 12, wherein the at least one datum comprises an identifier for an object.

14. The system of one of claims 9 to 13, wherein determining whether the device is in the first state or the second state is further based on a time for the audio.

15. The system of one of claims 9 to 14, wherein the method further comprises: receiving a request to store data about an entity: storing at least one datum associated with the entity: wherein the first state includes a mixed reality state: wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an augmented reality state; wherein the second result comprises a summary of the entity.

16. The system of one of claims 9 to 15, wherein the method further comprises: capturing an image of an entity; determining at least one datum associated with the entity based on the image; and storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.

17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising: receiving audio from a user of a device; in response to the audio, determining whether the device is in a first state or a second state based on content displayed by the device; in response to determining that the device is in the first state:Atty Docket No. 0120-1234W01 determining a first result based on the audio and the first state; and causing display of the first result; and in response to determining that the device is in the second state: determining a second result based on the audio and the second state; and causing display of the second result.

18. The computer-readable storage medium of claim 17, wherein the method further comprises: storing at least one datum associated with an interaction between the user and the device, the interaction comprising at least a first audio from the user, wherein determining the first result based on the audio and the first state is further based on the at least one datum.

19. The computer-readable storage medium of claim 17 or 18, wherein the method further comprises: receiving a request to store data about an entity; storing at least one datum associated with the entity; wherein the first state includes a mixed reality' state; wherein the first result comprises a virtual representation of the entity in a physical environment for the user; wherein the second state includes an alternate reality' state; wherein the second result comprises a summary of the entity'.

20. The computer-readable storage medium of one of claims 17 to 19, wherein the method further comprises: capturing an image of an entity; determining at least one datum associated with the entity' based on the image; storing the at least one datum associated with the entity; wherein determining the first result is further based on the at least one datum.

Citation Information

Patent Citations

  • Contextual voice commands

    US20100312547A1

  • Human-computer interaction method, and electronic device and system

    WO2022052776A1