Context menus for physical objects

The system enhances XR devices by identifying physical objects and overlaying relevant actions, enabling context-aware interactions and seamless integration with the real world, addressing the challenge of effective action provision in XR environments.

WO2025184554A1PCT designated stage Publication Date: 2025-09-04GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017919
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing XR devices struggle to effectively provide actions associated with physical objects in the user's environment, lacking seamless integration of digital and physical elements and efficient interaction methods.

Method used

A system configured to receive images of objects, identify them using computer vision and deep learning models, and overlay relevant actions on a display, allowing users to interact with physical objects through gestures, voice commands, or controllers, maintaining context for different objects.

Benefits of technology

Enables immersive experiences by providing context-aware interactions with physical objects, allowing users to view and execute actions seamlessly integrated with the real world, enhancing user engagement and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017919_04092025_PF_FP_ABST
    Figure US2025017919_04092025_PF_FP_ABST
Patent Text Reader

Abstract

According to at least one implementation, a method includes receiving an image of an object and determining an identifier associated with the object. The method further determines at least one action based on the identifier associated with the object in response to a selection of the object. The method also includes displaying at least one descriptor for the at least one action on a display with a view of the object.
Need to check novelty before this filing date? Find Prior Art

Description

CONTEXT MENUS FOR PHYSICAL OBJECTSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 559,043, filed on February 28, 2024, entitled “REAL-WORLD OBJECT INTERACTIONS USING INSTANCE-BASED MODELS,” and U.S. Provisional Patent Application No. 63 / 559,032, filed on February 28, 2024, entitled “CONTEXT MENUS FOR REAL-WORLD OBJECTS,” the disclosures of which are incorporated herein by reference in their entirety.BACKGROUND

[0002] An extended reality (XR) device incorporates a spectrum of technologies that blend physical and virtual worlds, including virtual reality (VR), augmented reality (AR), and mixed reality (MR). These devices immerse users in digital environments, either by blocking out the real world (VR), overlaying digital content onto the real world (AR), or blending digital and physical elements seamlessly (MR). XR devices include headsets, glasses, or screens equipped with sensors, cameras, and displays that track the movement of users and their surroundings to deliver immersive experiences across various applications such as gaming, education, healthcare, and industrial training.SUMMARY

[0003] This disclosure relates to systems and methods for providing actions and action menus for physical objects. In some implementations, a device can be configured to receive an image of an object and determine an identifier associated with the object. In some implementations, the device can be wearable. In some implementations, the device can provide at least one indicator on a display to identify the object. The device can further be configured to identify a selection of the object and determine at least one action based on the identifier associated with the object. The device can be configured to display at least one indicator for the at least one action on a display with a view of the object. In some implementations, such as a wearable device, the at least one indicator is overlaid on a view of the object. In some implementations, the user can view the object via a see-through display or video pass-through display, and the at least one indicator can be overlaid and visible to the user via the display. The user can provide input to select an action based on the at least oneindicator provided. In some examples, the indicator can comprise a text-based indicator, the text can include a text summary of the potential action for the user. In some examples, the text can include text attributes (e.g., name of the action). In some examples, the indicator can comprise an image-based indicator, the image (i.e., symbol or icon) indicating the potential action. In some examples, the indicators can include text and images.

[0004] In some implementations, the device can be configured to identify a second object in the received image and determine a second identifier associated with the second object. The device can be configured to associate the object with a first instance of a model to respond to user input associated with the object and associate the second object with a second instance of the model to respond to user input associated with the second object. In some examples, the first and second instances can maintain separate contexts (e.g., conversation history) to respond to requests related to the corresponding object.

[0005] In some aspects, the techniques described herein relate to a method including: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator of the at least one action on a display with a view of the object.

[0006] In some aspects, the techniques described herein relate to a computing apparatus including: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing apparatus to perform a method, the method including: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator of the at least one action on a display with a view of the object.

[0007] In some aspects, the techniques described herein relate to a computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, cause the at least one processor to execute a method, the method including: receiving an image from a camera; identifying a first object and a second object in the image; associating the first object with a first instance of a model, the first instance configured to respond to user input associated with the first object; and associating the second object with a second instance of the model, the second instance configured to respond to user input associated with the second object.

[0008] The details of one or more implementations are outlined in the accompanying drawings and the description below. Other features will be apparent from the description and drawings and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 illustrates a computing environment to provide actions associated with identified objects according to an implementation.

[0010] FIG. 2 illustrates a method of providing actions associated with identified objects according to an implementation.

[0011] FIG. 3 illustrates an operational scenario of identifying objects using an image and providing potential actions according to an implementation.

[0012] FIG. 4 illustrates a method of using models to implement actions associated with multiple objects according to an implementation.

[0013] FIG. 5 illustrates an operational scenario of using models to implement actions associated with objects according to an implementation.

[0014] FIGs. 6A and 6B illustrate an operational scenario of user interactions with physical objects according to an implementation.

[0015] FIG. 7 illustrates a computing system to manage actions and menus for physical objects according to an implementation.DETAILED DESCRIPTION

[0016] Computing devices, such as wearable devices and extended reality (XR) devices, provide users with an effective tool for gaming, training, education, healthcare, mobile computing, and more. An XR device merges the physical and virtual worlds, encompassing virtual reality (VR), augmented reality (AR), and mixed reality (MR) experiences. These devices can include headsets or glasses equipped with sensors, cameras, and displays that track users’ movements and surroundings, allowing them to interact with digital content in real time. XR devices offer immersive experiences by either completely replacing the real world with a virtual one (VR), overlaying digital information onto the real world (AR), or seamlessly integrating digital and physical elements (MR). Input to XR devices may be provided through gestures, voice commands, controllers, and eye movements. Users interact with the virtual environment by manipulating objects, navigating menus, and triggering actions using these input methods, which are translated by the device’s sensors and algorithms into corresponding digital interactions within the XR space. However, at least onetechnical problem exists in effectively providing actions associated with physical objects in the user’ s environment.

[0017] In at least one technical solution, a system comprising one or more computing devices can be configured to receive an image of one or more objects. In some examples, the device can be wearable with a camera that captures outward-facing images of the user’s environment. For example, a wearable device can include an outward-facing camera near the temple area of the device to capture images of a user’s environment. The system can further be configured to determine identifiers or metadata associated with one or more objects. The one or more objects can be identified using computer vision techniques, such as deep learning models like convolutional neural networks (CNNs). These models process images by analyzing pixel patterns, edges, shapes, and textures to recognize and classify objects based on a vast amount of previously trained data. More advanced object detection algorithms, such as You Only Look Once (YOLO) and Faster R-CNN, can identify objects and their locations by defining bounding boxes around them. These systems can be further enhanced with techniques like semantic segmentation, which assigns labels to individual portions of the images, allowing for more precise recognition and differentiation between objects in complex scenes. For example, from an image of a shelf in a grocery store, the system can identify a set of products on the shelf. In some implementations, the identifier for a physical object identified from an image could be a descriptive label, such as “red wooden chair,” or a structured tag like "Chair_00123." It may also be a product name, serial number, barcode, or classification to uniquely classify the object or classify the object into a group of objects. In some implementations, the identifier indicates a type of object, and the identifier can be a text-based indicator of the type of object (e.g., grocery object).

[0018] In some implementations, the system is further configured to identify a selection of an object and determine at least one action based on the identifier associated with the object. Returning to the example of the grocery store shelf, the system can identify a selection of the object and display a set of one or more actions related to the identity of the object. In some examples, the system can display one or more indicators for the one or more objects identified in the image. The indicators can include arrows, highlights, color changes, text descriptors, or other indicators associated with the objects identified in the image. For example, when the system identifies the objects on the grocery store shelf, the system can highlight the identified objects, indicating the available objects for user selection. Once an object is selected, using a voice command, gesture, gaze, controller, touchpad, or another input mechanism, the system identifies available actions based on the object identifier anddisplays the available actions. In some implementations, the action can be overlaid on the image in the example of a two-dimensional image. In some examples, the system can overlay the actions on a display for the device, such as the display for an XR device described below. As at least one technical effect, the user can view the physical world using optical see- through displays (like in AR headsets) or video pass-through (in MR / VR devices) and view the available actions overlaid on the physical object. The set of actions can be provided as a list, arranged around the object, or formatted in another manner. In some examples, the actions are positioned near the selected object and are placed over at least one other physical object. In some implementations, the displayed actions comprise descriptors or indicators of the potential actions that the user can implement. In some examples, the indicators comprise text indicators that indicate and summarize an action to be implemented (e.g., text attributes). The text indicators can indicate a voice command to execute the corresponding action in some implementations. In some examples, the displayed actions can be provided as an image (e.g., icon or symbol), indicating the potential action that can be implemented. The user can then select an action (using a voice command, gesture, and the like).

[0019] In some examples, the system can relate object identifiers to different actions. For example, a grocery object can be associated with a first set of actions. In contrast, a text display (poster or other display) can be associated with different actions. The first set of actions can include purchasing options for the object, ingredients for the object, recipes for the object, or some other action related to grocery elements. The second set of actions can include options for adding an event to the calendar, performing a translation (if necessary), or other actions associated with the text. As at least one technical effect, different sets of actions can be provided, depending on the object identified. This lets the user quickly view and select actions most relevant to the identified object.

[0020] In some XR or wearable device implementations, the system identifies physical objects using a combination of computer vision, depth sensing, and spatial mapping. The system can be configured to use technologies like Simultaneous Localization and Mapping (SLAM), which uses cameras and sensors to detect surfaces, edges, and spatial features. Depth sensors (e.g., LiDAR or structured light) can help estimate distance and object shape, while machine learning models can recognize specific objects by comparing them to a database of known items. In some XR systems, the device can use fiducial markers, spatial anchors, or other elements to track objects more precisely in augmented or mixed reality experiences. For example, a wearable device user can view a set of objects positioned on a table. The device can identify the objects from at least one image captured by the deviceand establish a location associated with the objects. When the user selects an object, using a gesture, voice command, controller, or some other mechanism, a set of actions can be displayed as an overlay for the physical environment. In some examples, the device overlays digital content (i.e., the available actions) onto the real-world view using a combination of sensors, cameras, and display technologies. The device can capture the physical environment using depth sensors, LiDAR, or SLAM to understand spatial positioning. The device then processes this data and aligns virtual objects (e.g., a set of potential actions) with real-world surfaces or markers. The digital content can be rendered in real-time, matching the user’s perspective and movements using optical see-through displays (like in AR headsets) or video pass-through (in MR / VR devices). In some examples, the system can use advanced techniques like occlusion, lighting estimation, and physics simulation to improve realism, making the digital elements appear integrated with the real world. The potential actions can be displayed as a list, positioned around the object, or arranged in another manner. In some examples, the actions are placed near the selected object so the user can view the object and the available actions. The selected object can be highlighted, promoted, or otherwise indicated to the user with the actions.

[0021] In some implementations, the system can implement a model, such as a language model, that can process user requests associated with a particular object. For example, the system can identify an object and determine a set of potential actions related to the object. Once chosen, the set of potential actions can be displayed for the user, and the user can generate verbal requests associated with the various actions. As part of the model, the system can take voice requests and implement the associated action by combining automatic speech recognition (ASR), natural language processing (NLP), and command execution systems. In some examples, ASR converts the spoken input into text. The NLP component then processes this text to understand the intent and extract relevant parameters. Based on the recognized command, the device can be configured to map the request to an action, such as searching for information, executing software functions, or providing another operation. Once the appropriate action is determined, the device interacts with application programming interfaces (APIs), hardware controllers, or cloud services to complete the request. The system can then be configured to provide feedback through text, voice, or visual confirmation to ensure the action was successfully executed. An action implemented by the device can be a system-executed operation that interacts with, enhances, or provides information about the object. This action can include highlighting the object, retrieving data for the object, displaying information, initiating a virtual interaction (e.g., rotating, scaling, or simulatingusage), or triggering real-world commands such as activating loT-connected devices. For example, the user may view a product on the shelf of a store and request reviews associated with the product from the set of potential actions. From the request, the system can retrieve the corresponding reviews and provide the reviews to the user (visually or audibly).

[0022] In some implementations, the system can execute multiple instances of the model for the objects identified in the environment. As at least one technical effect, the different objects can maintain different action or conversation histories (i.e., context), permitting the system to differentiate between user action requests associated with a first object and user requests related to a second object. In some implementations, the device can maintain conversational or contextual history associated with the different objects. As a result, the user can move between objects (i.e., select a first object and then a second object) and maintain separate contextual interactions associated with the objects. This can permit the user to refer to previous actions and / or previous responses from the system.

[0023] In some implementations, the system can execute at least one model instance that is used to compare objects identified from the image of the environment. For example, the user may request a comparison between two objects identified in a store (e.g., calorie count between food items). The system can be configured to identify natural language from the user, indicating the request to compare objects. In some implementations, the system can identify phrases or contextual information indicating the request to compare objects. For example, the user may gesture to multiple objects in the environment and provide voice input to compare the calories of the objects. The model can identify the voice input and the contextual information from the gesture to perform the comparison operation and provide the information requested by the user. In some implementations, the model can be configured to perform comparisons associated with various attributes across objects of the same or similar type. For example, with food objects, the device can compare cost, calories, or other attributes associated with the objects. As at least one technical effect, one or more models can provide information about a single object or a comparison between multiple objects.

[0024] FIG. 1 illustrates a computing environment 100 to provide actions associated with identified objects according to an implementation. Computing environment 100 includes user 110, device 130, external system 138, user gaze 140, and user view 141. Device 130 includes display 131, sensors 132, camera 133, and action application 126. User view 141 is representative of the view for user 110 and includes gesture 142, desk 151, object 152, object 153, object 154, and actions 155. Device 130 is an example of a wearable device, such as an XR device. Although demonstrated as a wearable device, similar operations can beperformed by other devices, including smartphones, tablets, or other computing systems. External system 138 can include one or more desktop computers, server computers, tablets, smartphones, or other systems that can assist action application 126 in providing the operations as described herein.

[0025] In computing environment 100, device 130 includes display 131, which is a screen or projection surface that presents immersive visual content to user 110, merging virtual elements with the real world. Display 131 can include optical see-through displays (e.g., AR headsets) or video pass-through (e.g., MR / VR devices). Device 130 further includes sensors 132, such as accelerometers, gyroscopes, magnetometers, depth, infrared, and proximity sensors. The sensors can be used to monitor the physical movement of the user, identify depth information for other objects, identify eye movement for the user, or provide some other operation. Device 130 also includes camera 133, which can capture the real or physical environment to overlay virtual objects (e.g., application interfaces) seamlessly and for tracking movements of user 110 and surroundings to enable accurate interaction within the augmented or virtual space. In some examples, camera 133 can be positioned as an outward view to capture the physical world associated with the user’s gaze. Display 131 can receive an update 181 from action application 126 to overlay a set of one or more potential actions associated with an object identified in the physical environment. Sensors 132 and camera 133 provide data 170 and data 171 to action application 126 that can be used to identify objects in the physical environment, the location of the objects in the physical environment, or some other information about the user and / or environment.

[0026] In the example of computing environment 100, user view 141 represents the field of view for user 110. User view 141 includes gesture 142, desk 151, object 152, object 153, object 154, and actions 155. In at least one implementation, action application 126 receives one or more images from camera 133 corresponding to the physical environment for the user (e.g., from an outward-facing camera). Action application 126 identifies objects 152, 153, and 154 from the images. In some examples, the objects are identified using computer vision techniques, primarily through deep learning models like CNNs. These models analyze pixel patterns to detect shapes, textures, and colors, distinguishing objects from the background. Object detection algorithms such as YOLO and Faster R-CNN can refine this process by drawing bounding boxes around identified objects and classifying them. Semantic segmentation techniques go a step further by labeling each pixel in an image to understand object boundaries more precisely.

[0027] In some examples, object detection can be performed locally on device 130.In some examples, at least some object identification can be performed using external system 138 (e.g., a server or set of servers). In some implementations, when the objects are identified, the objects are assigned an identifier or object type. In some implementations, the identifier can comprise a text identifier or the individual object and / or a group identifier for the object (e.g., grocery object, signage object, and the like). For example, edible objects or products can be classified with a first identifier, while objects with text, such as posters, can be classified with a second identifier. For another example, the identifier can comprise a text indicator and an object type (e.g., grocery item). In some implementations, as objects 152, 153, and 154 are identified, indicators can be added to the display to indicate the identified objects. The indicators can include shading, arrows, highlights, boundaries, or other indicators associated with the object. For example, a circle can be added around object 154 when the object is identified. In some examples, device 130 assigns coordinates to physical objects within a 3D space. Using SLAM and depth-sensing technologies (like LiDAR or stereo cameras), the device can be configured to map objects relative to itself, for example, using a Cartesian coordinate system (X, Y, Z). These coordinates are then used to anchor digital overlays, enable object interaction, and maintain spatial consistency across sessions. Some devices can be configured to use world-locked anchors to ensure persistent positioning in augmented reality experiences. In some examples, the device can be configured to use anchors or world-locked anchors in a two-dimensional image that maps the location (or center) of an object in space.

[0028] In some implementations, after one or more objects are identified, action application 126 determines, for an object, one or more actions based on the identifier for the object. For example, object 154 can represent a bottle of milk, while object 152 represents a poster. Action application 126 can determine a different set of actions based on the identifier associated with the object. In some implementations, the actions can be based on the object type (e.g., produce, signage, and the like). When the user selects an available object 154 using a gesture, a controller, a touchpad, gaze, or some other mechanism, including combinations thereof, device 130 can display actions 155 that correspond to the selected object. In some implementations, device 130 can identify gestures using cameras, depth sensors, and hand tracking. It detects key hand points (e.g., finger joints) using computer vision and machine learning, mapping movements to gestures like pinching or swiping. In some implementations, the object can be selected via the user’s gaze focusing on the object.

[0029] Here, the user provides gesture 142 to select object 154, and actions 155 associated with the object are provided to support the selection. The displayed actions can comprise image or text-based indicators of potential actions that can be employed in association with the physical object. The actions can include providing a web search of the manufacturer, searching for the calorie count (when a food object), searching for the price, or some other action based on the identifier for the object. From the available actions, the user can request an action, wherein the request can include selecting an action using a gesture, providing voice input, or providing some other selection from the available actions. In some examples, the actions comprise suggestions for the user, and the user may give alternative requests not offered as part of the action suggestions. In some implementations, actions 155 are displayed as an overlay on the physical environment. In some examples, device 130 detects the object’s position and orientation in space, then anchors virtual elements (i.e., the actions) to it using spatial mapping and world-locked coordinates. This allows augmented content, such as annotations and potential actions for the object, to appear fixed to the object, which can maintain realism and perspective as the user moves.

[0030] In some implementations, action application 126 can associate a model with each of the objects. The model can identify voice requests using speech recognition and NLP. The model can capture audio through microphones, process the input using voice recognition, and interpret the user’s intent. The system then maps the request to defined actions, such as triggering a search, saving the object identifier, or other actions. With gaze tracking and gesture recognition, context-awareness can further refine the response, ensuring accurate execution of user commands relative to the various actions.

[0031] In some examples, the device can maintain context associated with the various requests from the user and the responses provided. This permits the device to identify context for future requests and resolve ambiguities related to the request, such as identifying references to ambiguous terms. In some examples, each identified object can be associated with a different model that provides the action suggestions and supports the requests related to each object. In at least one implementation, a comparison model can be used to compare objects and provide information associated with the comparison (e.g., comparison between grocery store objects).

[0032] Although demonstrated as a wearable device in computing environment 100, device 130 can comprise devices such as smartphones, tablets, and the like. The device can capture one or more images of the environment and identify objects in the images. Potential actions can be overlaid over the objects (e.g., on the images) based on a user’s objectselection. In some implementations, the device’s screen can provide a video feed of the environment, and the content (e.g., actions) can be overlaid in the video feed consistent with the identified objects. The overlaying can be based on the spatial location or anchors of the objects, like the operations for a wearable device.

[0033] FIG. 2 illustrates method 200 of providing actions associated with identified objects according to an implementation. The steps of method 200 can be performed by a system of one or more devices, such as device 130 of FIG. 1. However, method 200 can be performed using multiple devices in some examples (e.g., device 130 with external system 138 from FIG. 1). The steps of method 200 are referenced parenthetically in the paragraphs that follow.

[0034] Method 200 includes receiving (201) an image of an object. In some implementations, the image is captured via a camera, such as an outward-facing camera on an XR device. Method 200 further includes determining (202) an identifier associated with the object. The identifier can be a unique label, descriptor, or digital tag assigned based on visual, spatial, or metadata attributes to distinguish and interact with the physical object in the XR environment. The identifier can label the individual object or the object as an object type (e.g., grocery object, signage, etc.). In some implementations, the method can determine additional metadata, such as the location of the object within the environment, an object type from the identifier, or some other information. For example, the image can include a laptop computer, and the method can determine the identifier (i.e., laptop computer), classify the device as an electronics device, and determine the object’s location in the environment. In some examples, the system can be configured to recognize objects in images using computer vision and deep learning models like CNNs. It extracts features such as edges and textures, compares them to a trained dataset, and assigns probabilities to identify objects. When an object is identified, the system can be configured to map its position and orientation using computer vision and spatial tracking technologies, such as SLAM. It can create a digital reference point tied to real-world coordinates, ensuring virtual content remains aligned with the object.

[0035] Method 200 further includes, in response to a selection of an object, determining (203) at least one action based on the identifier associated with the object. In some implementations, the system can maintain one or more data structures that associate device identifiers with corresponding actions. For example, a first object type (e.g., grocery item) can be associated with a first set of actions, while a second object type or identifier (e.g., signage) can be associated with a second set of actions. An action can be a system-executed operation that interacts with, enhances, or provides information about the object. Method 200 further includes displaying (204) the at least one action on a display with a view of the object. In some implementations, the at least one action can be overlaid onto the view of the physical environment based on the anchor associated with the object. For example, a grocery item, such as a smoothie, can be anchored in physical space. The one or more potential actions are then overlaid using the display, such that the digital content appears with the physical object. The one or more actions can be provided around the object, as a list with the object, or in some other manner. In some implementations, the user can select the object using a voice command where available objects can be displayed with an identifier that the user can request using voice. In some implementations, the user can select the object using a controller, gestures, gaze, or some other means to select the object. In some examples, the potential actions are provided as text or image-based indicators that indicate or suggest actions to the user. In some implementations, the indicators can indicate voice commands or gestures that can trigger the corresponding action. For example, the system can display a voice command to trigger a recipe associated with the object.

[0036] Once the actions are displayed for the object, the user can request an action. In some examples, the user can give a voice request corresponding to the action. For example, the actions can include an option to search for the object via the internet and the user can provide a voice input to “search.” In other examples, the user can provide gestures, use a touchscreen, touchpad, or other to select an available action. In response to the request, the system can implement the corresponding action.

[0037] In some implementations, the system can be configured to associate and execute a model for the identified objects, wherein each object is associated with a different instance of the model. The model is used to identify requests for actions and process the requests to generate a response. In some implementations, the system can process the request using NLP or gesture recognition. If the request is verbal, the device converts speech to text, applies NLP to understand intent, and maps it to an action. If the input is a gesture or touch, computer vision and sensor data interpret the movement and correlate it with a command. In some examples, the system can use a combination of gesture and voice to identify the intent associated with the user request. In some examples, the device can maintain context information, such as a dialogue of action requests and responses, and use the context information to respond to future requests. The context information can be used to resolve ambiguities in some implementations. The context information can be separated per instance of the model, permitting the user to request an action associated with a first object andinformation associated with a second objection with the context separated between the objects.

[0038] FIG. 3 illustrates a user perspective 300 of identifying objects using an image and providing potential actions according to an implementation. User perspective 300 includes gesture 342, display 350, desk 351, poster 352, poster 353, objects 354, 355, and 356, and actions 357. Actions 357 can include descriptors or indicators that indicate potential or suggested actions to a user and can comprise images, text, or some combination thereof. In some examples, actions 357 can indicate voice or gestures to trigger the corresponding action.

[0039] User perspective 300 represents a user’s perspective of a physical environment with a wearable device like an XR device. The XR device can capture at least one image of the physical environment and identify objects 354, 355, and 356 in the image. In some implementations, the objects identified include objects relevant to the user, where the system may not recognize other objects. When objects 354, 355, and 356 are identified, the system can highlight or otherwise indicate the identified objects in the user perspective 300. In some implementations, the device can include an optical see-through display or a video pass- through display, permitting the user to view both the physical environment and the content identified for the objects.

[0040] Once the objects are identified from the image or images of the environment, the system further identifies a selection of the object by the user. In some implementations, such as that depicted in user perspective 300, the user can provide a gesture that selects an available object 354. In some implementations, the device can be configured to identify a gesture selection using computer vision and sensor data from cameras, infrared sensors, or LiDAR. It tracks hand movements, finger positions, and gestures in real time, comparing them to a trained model of predefined gestures. The data can be processed to identify user intent and selection. The device can alternatively use controllers, touchpads, or other input mechanisms to select the object, where a selector or cursor can be overlaid on the display for selection. In some implementations, the system can be configured to identify voice input associated with the object. For example, each of the objects 354, 355, and 356 can be labeled (e.g., a name of the product identified from the image). The user can provide voice input with the label to select the corresponding object, wherein the device can provide NLP to process the voice input.

[0041] In some implementations, the system can provide the user with one or more potential actions associated with the object in response to the selection. The actions caninclude, for a sample grocery item, instructions for the object, nutrition for the object, purchasing the object, or some other action associated with the object. In some implementations, the recommended actions can be determined using the imaging information for the object, the location of the user, or some other contextual relevant information about the object. In some implementations, the system can maintain a record of actions taken with similar objects by the user and select the one or more actions for display based on the record. For example, when a grocery object is identified, the system can include ten potential actions to be taken in association with the object. The system can promote a subset of the ten potential actions based on the frequency that the various actions were requested by the user. The promotion can include displaying only the subset of actions, displaying the actions with a larger font, displaying the promoted actions in a different location, or another promotion of the actions. The user can provide input (voice, gesture, and the like) to implement a corresponding action associated with the object.

[0042] FIG. 4 illustrates method 400 of using models to implement actions associated with multiple objects according to an implementation. The steps of method 400 are referenced parenthetically in the paragraphs that follow. The steps of method 400 can be implemented via a system of one or more devices, such as device 130 of FIG. 1.

[0043] Method 400 includes receiving (401) an image from a camera and identifying (402) a first object and a second object in the image. In some implementations, the image corresponds to a portion of a physical environment for a user. In some implementations, the image is captured via an outward-facing camera. The system can be configured to process the image to identify relevant objects for the user. For example, the system can be configured to identify grocery objects identified in the image. In some implementations, the system processes an image to identify objects using computer vision and models. It first captures the image through a camera and pre-processes it (e.g., resizing, noise reduction) before passing it through a CNN or another trained model. The model extracts features, classifies objects, and provides bounding boxes or segmentation masks to identify and localize multiple objects within the image.

[0044] Once the objects are identified, method 400 includes associating (403) a first model with the first object the first model configured to process a plurality of object types. Method 400 further includes associating (404) a second model with the second object, the second model being a second instance of the first model. In some implementations, by providing different instances of the same model, the system can provide contextually relevant responses to user requests. For example, for each instance, the system can maintain aconversation history with action requests and responses, which can provide contextually relevant responses associated with the object. In some implementations, the user can transition from requesting information associated with a first object to requesting information associated with a second object. By separating the instances, the system can provide contextually relevant responses associated with the individual object.

[0045] In some implementations, the model comprises a natural language model. The model can be configured to identify user requests by processing input text (or speech) through a series of steps, including tokenization, parsing, and semantic analysis. It can use deep learning algorithms, such as transformer-based architectures, to understand context and intent. The model can then generate a relevant response by predicting the most appropriate text based on training data and learned patterns. It can refine the responses through contextual understanding (i.e., conversation history), user feedback, and predefined rules or integrations with external databases

[0046] In some implementations, the system can be configured to determine how frequently the user requests different actions associated with different object types. From the frequency information, the system can provide or display potential actions to the user. For example, the device can identify the most frequent actions provided by a user in association with grocery objects. In response to identifying a new grocery object, the system can provide a suggested set of actions based on the identified frequency.

[0047] FIG. 5 illustrates an operational scenario 500 of using models to implement actions associated with objects according to an implementation. Operational scenario 500 includes objects 510, 511, and 512, model instances 520, 521, and 522, and compare instance 530. The operations associated with model instances 520, 521, and 522, and compare instance 530 may be performed by one or more devices, such as device 130 of FIG. 1 or computing system 700 of FIG. 7.

[0048] In operational scenario 500, a system identifies objects 510, 511, and 512 from at least one image. In some implementations, the system identifies a subset of objects from the environment that are relevant to the user (e.g., grocery items). As the objects are identified, the system assigns a model instance to each of the objects. Here, the system assigns model instance 520 to object 510, model instance 521 to object 511, and model instance 522 to object 512. Model instances 520, 521, and 522 can be representative of language models that can process user requests (text or verbal). Language models can process user requests by analyzing input text, identifying intent, and using deep learning algorithms to generate contextually relevant responses. They can leverage vast training data,tokenization, and transformer architectures to predict the best reply based on learned patterns. In some examples, the models can be configured to respond to requests associated with the object type (e.g., recipe requests, calories, ingredients, and the like for grocery items). When a user selects an object in the physical environment (e.g., using a gesture, voice command, controller, etc.), the system executes the model to interact with the user requests and provide responses. For example, model instance 520 can receive a first request to identify the calories associated with object 510. Subsequently, when the user selects object 512, model instance 522 can separately process any requests associated with that item. As at least one technical effect, the context for each of the objects can be separated. In some implementations, each of the model instances can use visual cues (e.g., gestures) to further provide context and respond to user requests.

[0049] In addition to the model instances 520, 521, and 522 associated with each of the objects 510, 511, and 512, operational scenario 500 further includes compare instance 530. Compare instance 530 is executed to provide a comparison between two objects. Compare instance 530 can be used when the language of the user, gesture of the user, or other input operation indicates an intent to compare two objects. In some examples, the system can select the objects for comparison based on the objects location in the image, the gesture of the user toward the objects, the user holding the objects, or based on some other factor. In some examples, the system can select the objects for comparison based on the user verbally indicating the objects (e.g., using the identifier allocated to the objects or displayed for the objects. For example, the user can provide voice input for “Which of these has the most calories?” and provide a gesture that corresponds to a selection of object 510 and object 511. Compare instance 530, which can be configured to process the request based on the terminology voice input, can process the request as a language model that retrieves the required information and provides a response to the user (e.g., displaying the result). In some examples, the compare instance 530 can also maintain context associated with the requests from the user, including conversation history (requests and responses), gestures, imaging, and other context to respond to requests from the user. In some implementations, the contexts for different comparisons can be maintained separately, like the model instances 520, 521, and 522. In some examples, the

[0050] FIG. 6A and FIG. 6B illustrate an operational scenario of user interactions with physical objects according to an implementation. FIG. 6 A provides user perspective 600 and FIG. 6B includes user perspective 601. FIG. 6 A includes gesture 642, objects 654 and 655, and actions 657. FIG. 6B includes gesture 643, objects 654 and 655, and actions 658.Actions 657 and 658 are representative of indicators that describe or indicate the potential actions associated with a corresponding action. The operations associated with FIG. 6A and 6B can be performed by a device or system of devices, such as device 130 from FIG. 1 or computing system 700 of FIG. 7. The operations can be performed by an XR device in some examples.

[0051] Referring first to FIG. 6A, a system can capture an image of a physical environment and identify one or more objects in the physical environment. In some implementations, the system is configured to identify objects of interest to the users, such as grocery items. In some implementations, as the objects are identified, an indicator is provided in association with each of the objects. The indicator can comprise a highlighted portion over the object, an arrow, a text identifier for the object, or some other indicator for the identified objects. For example, the system can be configured to overlay digital content onto physical environments by using sensors, cameras, and spatial mapping to track the real world and render virtual elements in real-time, aligning them with the user’s perspective. As a result, content is provided by the system between the physical object and the user’s view but is arranged such that it is overlaid with the physical object (e.g., on a display of an XR device).

[0052] Once the objects are identified, the user can select an object using voice, gesture, a touchpad (and selector), or some other means. Here, the user provides gesture 642 to select object 654. In response to the selection, the system can indicate the selection by highlighting the object, displaying a pointer at the object, displaying a boundary for the object, displaying a name for the object, or other indicators for the various objects. For example, an XR device can use an outward-facing camera to capture at least one image of a physical environment. The objects identified from the at least one image can be anchored in space, and the anchors can be used to display the potential actions (i.e., descriptors or indicators of actions) on the device’s display. In some implementations, the user can view the object via a see-through display or video pass-through display, and the actions 657 (i.e., action indicators) can be overlaid and visible to the user via the display. In some examples, the actions can be determined based on the identifier associated with object 654. In some examples, the action can be determined at least in part on previous action requests provided by the user, wherein more frequent requests can be promoted or suggested for the user. The actions can be determined in association with similar object types (e.g., grocery objects). Once actions 657 are displayed with object 654, the user can provide input or requests associated with the object. In some implementations, the requests are processed via a model instance, wherein the model can use natural language processing or a language model torespond to the user requests. In some examples, the model instance for object 654 differs from object 655.

[0053] Turning FIG. 6B, the user provides gesture 643 to transition from object 654 to object 655. In some examples, rather than using the gesture, the change can be based on gaze, voice input, a controller, or some other input. Suppose the system cannot determine the object the user is referencing. In that case, the system can request additional information from the user, including a gesture, the user touching the object, voice input, or some other clarification associated with the object. In response to selecting object 655, actions 658 associated with object 655 are displayed in for user perspective 601. In some examples, the physical environment is visually passed or video passed through to the user as part of a display, and actions 658 are overlaid as part of digital content on the display. The display of actions 658 can be based on the anchor (spatial location) for the object. Actions 658 can be assembled as a menu around the object, a list menu next to the object, or in some other manner. In some implementations, actions 658 are different from actions 657 provided for object 654. Actions 658 can be determined based on the identifier for the object, historical action requests from the user, or some other factor. Once displayed, the user can select an action from actions 658, where the selection can be verbal, via gesture, or other means.

[0054] FIG. 7 illustrates a computing system 700 to manage actions and menus for physical objects according to an implementation. Computing system 700 represents any apparatus, computing system, or systems with which the various operational architectures, processes, scenarios, and sequences are disclosed herein for managing actions associated with physical objects identified in an environment. Computing system 700 can be an example of an XR device, wearable device, or other computing device capable of the operations described herein. Computing system 700 is an example of device 130 from FIG. 1. Computing system 700 includes storage system 745, processing system 750, communication interface 760, and input / output (VO) device(s) 770. Processing system 750 is operatively linked to communication interface 760, I / O device(s) 770, and storage system 745. In some implementations, communication interface 760 and / or I / O device(s) 770 may be communicatively linked to storage system 745. Computing system 700 may further include other components such as a battery and enclosure that are not shown for clarity.

[0055] Communication interface 760 comprises components that communicate over communication links, such as network cards, ports, radio frequency, processing circuitry (and corresponding software), or some other communication devices. Communication interface 760 may be configured to communicate over metallic, wireless, or optical links.Communication interface 760 may be configured to use Time Division Multiplex (TDM), Internet Protocol (IP), Ethernet, optical networking, wireless protocols, communication signaling, or some other communication format - including combinations thereof. Communication interface 760 may be configured to communicate with external devices, such as servers, user devices, or some other computing device.

[0056] I / O device(s) 770 may include peripherals of a computer that facilitate the interaction between the user and computing system 700. Examples of I / O device(s) 770 may include keyboards, mice, trackpads, monitors, displays, printers, cameras, microphones, external storage devices, sensors, and the like. In some implementations, VO device(s) 770 include at least one outward-facing camera configured to capture images associated with the physical environment. In some implementations, VO device(s) 770 consists of a see-through or video pass-through display providing a view of the physical environment. The display can provide overlaid content, such as potential actions and responses to requests from users of computing system 700.

[0057] Processing system 750 comprises microprocessor circuitry (e.g., at least one processor) and other circuitry that retrieves and executes operating software (i.e., program instructions) from storage system 745. Storage system 745 may include volatile and nonvolatile, removable, and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Storage system 745 may be implemented as a single storage device but may also be implemented across multiple storage devices or sub-systems. Storage system 745 may comprise additional elements, such as a controller to read operating software from the storage systems. Examples of storage media (also referred to as computer-readable storage media or a computer-readable storage medium) include random access memory, readonly memory, magnetic disks, optical disks, and flash memory, as well as any combination or variation thereof, or any other type of storage media. In some implementations, the storage media may be non-transitory. In some instances, at least a portion of the storage media may be transitory. In no case is the storage media a propagated signal.

[0058] Processing system 750 is typically mounted on a circuit board that may also hold the storage system. The operating software of storage system 745 comprises computer programs, firmware, or some other form of machine-readable program instructions. The operating software of storage system 745 comprises action application 724. The operating software on storage system 745 may further include an operating system, utilities, drivers, network interfaces, applications, or some other type of software. When read and executed byprocessing system 750 the operating software on storage system 745 directs computing system 700 to operate as described herein. In at least one implementation, the operating software can provide method 200 described in FIG. 2 or method 400 described in FIG. 4. The operating software can provide or cause the at least one processor to manage actions with physical objects as described herein.

[0059] Example claim clauses are provided below. Although these are examples, these clauses should not be considered exhaustive.

[0060] Clause 1. A method comprising: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator of the at least one action on a display with a view of the object.

[0061] Clause 2. The method of clause 1, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object in the image; associating the first object with a first instance of a model to respond to user input associated with the first object; and associating the second object with a second instance of the model to respond to user input associated with the second object.

[0062] Clause 3. The method of clause 2, further comprising: receiving an input from a user in association with the first object; processing the input using the first instance of the model to determine an action from the at least one action; and executing the action.

[0063] Clause 4. The method of clause 1, further comprising: receiving a request for a first action of the at least one action; and in response to the request, executing the first action.

[0064] Clause 5. The method of clause 1, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object captured in the image; in response to a second selection of the second object, determining a first action for the second object based on the second identifier associated with the second object, the first action different from the at least one action; and displaying the first action on the display with a view of the second object.

[0065] Clause 6. The method of clause 5, further comprising: removing the at least one action from the display.

[0066] Clause 7. The method of clause 1, further comprising: identifying a set of actions executed in association with a set of objects related to the object; wherein determining the at least one action is further based on the set of actions.

[0067] Clause 8. The method of clause 1, wherein the object is a first object, and wherein the method further comprises: determining a second identifier associated with a second object captured in the image; receiving a request to compare a first attribute of the first object with a second attribute of the second object; in response to the request, generating a response based on at least one model; and displaying the response.

[0068] Clause 9. A computing apparatus comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing apparatus to perform a method, the method comprising: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator of the at least one action on a display with a view of the object.

[0069] Clause 10. The computing apparatus of clause 9, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object in the image; associating the first object with a first instance of a model to respond to user input associated with the first object; and associating the second object with a second instance of the model to respond to user input associated with the second object.

[0070] Clause 11. The computing apparatus of clause 10, wherein the method further comprises: receiving an input from a user associated with the first object; processing the input using the first instance of the model to determine an action from the at least one action; and executing the action.

[0071] Clause 12. The computing apparatus of clause 9, wherein the method further comprises: receiving a request for a first action of the at least one action; and in response to the request, executing the first action.

[0072] Clause 13. The computing apparatus of clause 9, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object captured in the image; in response to a second selection of the second object, determining a first action for the second object based on the second identifier associated with the second object, the first action different from the at least one action; and displaying the first action on the display with a view of the second object.

[0073] Clause 14. The computing apparatus of clause 13, wherein the method further comprises: removing the at least one action from the display.

[0074] Clause 15. The computing apparatus of clause 9, wherein the method further comprises: identifying a set of actions executed in association with a set of objects related to the object; wherein determining the at least one action is further based on the set of actions.

[0075] Clause 16. The computing apparatus of clause 9, wherein the object is a first object, and wherein the method further comprises: determining a second identifier associated with a second object captured in the image; receiving a request to compare a first attribute of the first object with a second attribute of the second object; in response to the request, generating a response based on at least one model; and displaying the response.

[0076] Clause 17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, cause the at least one processor to execute a method, the method comprising: receiving an image from a camera; identifying a first object and a second object in the image; associating the first object with a first instance of a model, the first instance configured to respond to user input associated with the first object; and associating the second object with a second instance of the model, the second instance configured to respond to user input associated with the second object.

[0077] Clause 18. The computer-readable storage medium of clause 17, wherein the method further comprises: determining at least one action associated with the first object; and in response to a selection of the first object, displaying at least one indicator of the at least one action on a display with a view of the first object.

[0078] Clause 19. The computer-readable storage medium of clause 18, wherein the selection comprises a first selection, and wherein the method further comprises: determining at least one additional action associated with the second object; in response to a second selection of the second object; displaying at least one additional indicator of the at least one additional action associated with the second object with a view of the second object; and removing the at least one indicator from the display.

[0079] Clause 20. The computer-readable storage medium of clause 17, wherein the method further comprises: receiving a request from a user associated with the first object; generating a response to the response based on the first instance of the model; and providing the response to the request via a display.

[0080] In this specification and the appended claims, the singular forms “a,” “an” and “the” do not exclude the plural reference unless the context dictates otherwise. Further, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictatesotherwise. For example, “A and / or B” includes A alone, B alone, and A with B. Further, connecting lines or connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physical connections, or logical connections may be present in a practical device. Moreover, no item or component is essential to the practice of the implementations disclosed herein unless the element is specifically described as “essential” or “critical.”

[0081] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that a precise value or range thereof is not required and need not be specified. As used herein, the terms discussed above will have ready and instant meaning to one of ordinary skill in the art.

[0082] Moreover, the use of terms such as up, down, top, bottom, side, end, front, back, etc. herein are used concerning a currently considered or illustrated orientation. If they are considered concerning another orientation, such terms must be correspondingly modified.

[0083] Further, in this specification and the appended claims, the singular forms “a,” “an” and “the” do not exclude the plural reference unless the context dictates otherwise. Moreover, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B.

[0084] Although certain example methods, apparatuses, and articles of manufacture have been described herein, the scope of coverage of this patent is not limited thereto. It is to be understood that the terminology employed herein is to describe aspects and is not intended to be limiting. On the contrary, this patent covers all methods, apparatus, and articles of manufacture fairly falling within the scope of the claims of this patent.

Claims

WHAT IS CLAIMED IS:

1. A method comprising: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator associated with the at least one action on a display with a view of the object.

2. The method of claim 1, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object in the image; associating the first object with a first instance of a model to respond to user input associated with the first object; and associating the second object with a second instance of the model to respond to user input associated with the second object.

3. The method of claim 2, further comprising: receiving an input from a user in association with the first object; processing the input using the first instance of the model to determine an action from the at least one action; and executing the action.

4. The method of claim 1, further comprising: receiving a request for a first action of the at least one action; and in response to the request, executing the first action.

5. The method of claim 1, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object captured in the image; in response to a second selection of the second object, determining a first action for the second object based on the second identifier associated with the second object, the first action different from the at least one action; anddisplaying a first indicator associated the first action on the display with a view of the second object.

6. The method of claim 5, further comprising: removing the at least one indicator from the display.

7. The method of claim 1, further comprising: identifying a set of actions executed in association with a set of objects related to the object; wherein determining the at least one action is further based on the set of actions.

8. The method of claim 1, wherein the object is a first object, and wherein the method further comprises: determining a second identifier associated with a second object captured in the image; receiving a request to compare a first attribute of the first object with a second attribute of the second object; in response to the request, generating a response based on at least one model; and displaying the response.

9. A computing apparatus comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing apparatus to perform a method, the method comprising: receiving an image of an object; determining an identifier associated with the object; in response to a selection of the object, determining at least one action based on the identifier associated with the object; and displaying at least one indicator associated with the at least one action on a display with a view of the object.

10. The computing apparatus of claim 9, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object in the image; associating the first object with a first instance of a model to respond to user input associated with the first object; and associating the second object with a second instance of the model to respond to user input associated with the second object.

11. The computing apparatus of claim 10, wherein the method further comprises: receiving an input from a user associated with the first object; processing the input using the first instance of the model to determine an action from the at least one action; and executing the action.

12. The computing apparatus of claim 9, wherein the method further comprises: receiving a request for a first action of the at least one action; and in response to the request, executing the first action.

13. The computing apparatus of claim 9, wherein the object comprises a first object in a physical environment, and the method further comprises: determining a second identifier associated with a second object captured in the image; in response to a second selection of the second object, determining a first action for the second object based on the second identifier associated with the second object, the first action different from the at least one action; and displaying a first indicator associated the first action on the display with a view of the second object.

14. The computing apparatus of claim 13, wherein the method further comprises: removing the at least one indicator from the display.

15. The computing apparatus of claim 9, wherein the method further comprises: identifying a set of actions executed in association with a set of objects related to the object; wherein determining the at least one action is further based on the set of actions.

16. The computing apparatus of claim 9, wherein the object is a first object, and wherein the method further comprises: determining a second identifier associated with a second object captured in the image; receiving a request to compare a first attribute of the first object with a second attribute of the second object; in response to the request, generating a response based on at least one model; and displaying the response.

17. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, cause the at least one processor to execute a method, the method comprising: receiving an image from a camera; identifying a first object and a second object in the image; associating the first object with a first instance of a model, the first instance configured to respond to user input associated with the first object; and associating the second object with a second instance of the model, the second instance configured to respond to user input associated with the second object.

18. The computer-readable storage medium of claim 17, wherein the method further comprises: determining at least one action associated with the first object; and in response to a selection of the first object, displaying at least one indicator associated with the at least one action on a display with a view of the first object.

19. The computer-readable storage medium of claim 18, wherein the selection comprises a first selection, and wherein the method further comprises: determining at least one additional action associated with the second object; in response to a second selection of the second object; displaying at least one additional indicator of the at least one additional action associated with the second object with a view of the second object; and removing the at least one indicator from the display.

20. The computer-readable storage medium of claim 17, wherein the method further comprises: receiving a request from a user associated with the first object; generating a response to the response based on the first instance of the model; and providing the response to the request via a display.

Citation Information

Patent Citations

  • Curated contextual overlays for augmented reality experiences

    US20220358689A1