User interaction with a remote device

JP2024540828A5Pending Publication Date: 2025-10-03QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024519903
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-12
Filing Date
2022-10-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Users often face difficulties interacting with remote devices due to inaccessible user interfaces, language barriers, poor visibility, or lack of contextual understanding, especially when multiple devices are present in a scene, leading to confusion and inefficiency in managing device interactions.

Method used

An extended reality (XR) system is used to facilitate user interaction with remote devices by providing contextual guidance through virtual overlays and user interfaces, leveraging input data and context information to simplify and manage interactions, even with devices lacking external controls or displays.

Benefits of technology

The XR system enhances user interaction by providing intuitive guidance, reducing clutter, and improving accessibility, allowing users to effectively control devices under various conditions, including poor lighting or interface inaccessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system, method, and non-transitory medium for presenting information associated with at least one input option are provided. An exemplary method can include receiving data identifying one or more input options associated with a first device in a scene, and determining information relevant to at least one of the scene, the first device, and a user associated with a second device, including using at least one memory, and outputting user guidance data corresponding to the input option for which relevant contextual information was determined based on the one or more input options and the information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates generally to interactions with remote devices. For example, aspects of the present disclosure include filtering and / or suggesting virtual content for user interaction with a remote device. [Background technology]

[0002] Extended reality technologies can be used to present virtual content to a user and / or combine real and virtual environments from the physical world to provide an extended reality experience to a user. The term extended reality can encompass virtual reality, augmented reality, mixed reality, and the like. Each of these forms of extended reality allows a user to experience or interact with an immersive virtual environment or content. For example, an extended reality experience may allow a user to interact with a real or physical environment that is augmented or extended with virtual content.

[0003] Extended reality technologies can be implemented to enhance user experiences in a wide range of contexts, including entertainment, healthcare, retail, education, social media, among others. Summary of the Invention

[0004] Disclosed are systems, apparatus, methods, and computer-readable media for determining user interaction data for remote device interactions (e.g., interactions between one or more remote devices, such as an Extended Reality (XR) device and an Internet-of-Things (Internet-of-Things) device). According to at least one example, a method is provided for presenting information associated with at least one input option. The method includes receiving data identifying one or more input options associated with a device in a scene, determining information relevant to at least one of the scene, the device, and a user associated with the electronic device, including using at least one memory, and outputting user guidance data corresponding to the input option for which relevant contextual information was determined based on the one or more input options and the information.

[0005] In another example, an apparatus is provided for presenting information associated with at least one input option, the apparatus including at least one memory and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory. The at least one processor may be configured and configured to receive data identifying one or more input options associated with a device in a scene, determine information relevant to at least one of the scene, the device, and a user associated with the electronic device, including using the at least one memory, and output user guidance data corresponding to an input option for which relevant contextual information was determined based on the one or more input options and the information.

[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to receive data identifying one or more input options associated with a device in a scene, determine, including using at least one memory, information relevant to at least one of the scene, the device, and a user associated with the electronic device, and output user guidance data corresponding to the input option for which relevant contextual information has been determined based on the one or more input options and the information.

[0007] In another example, an apparatus is provided for presenting information relevant to at least one input option, the apparatus including means for receiving data identifying one or more input options associated with a device in a scene, means for determining, including using at least one memory, information relevant to at least one of the scene, the device, and a user associated with the electronic device, and means for outputting user guidance data corresponding to the input option for which relevant contextual information was determined based on the one or more input options and the information.

[0008] In some aspects, the aforementioned methods, non-transitory computer-readable media, and apparatus may include predicting a user interaction with the device based on the information, and presenting user guidance data corresponding to the input option based on the one or more input options and the predicted user interaction.

[0009] In some examples, the user guidance data may include at least one of a user input element associated with the input option, a virtual overlay on a physical object associated with the input option, and / or a cue indicating how to provide an input associated with the input option.

[0010] In some examples, the device may include a connected device with network communication capabilities, and the aforementioned methods, non-transitory computer-readable media, and apparatus may include determining a hand gesture representing a predicted user interaction based on the information and the one or more input options, and presenting user guidance data. In some examples, the predicted user interaction may include a predicted user input to the device. In some cases, the user guidance data may include an indication of a hand gesture that, when detected, invokes an actual user input at the device.

[0011] In some aspects, presenting the user guidance data can include rendering, on a display associated with the electronic device, a virtual overlay configured to appear to be located on a surface of the device. In some examples, the virtual overlay can include user interface elements associated with the input options. In some cases, the user interface elements can include at least one of a virtual user input object associated with the input option and a visual indication of a physical control object on the device configured to receive an input corresponding to the input option.

[0012] In some examples, the information includes at least one of a user's gaze and a user's posture, and the aforementioned methods, non-transitory computer-readable media, and apparatus may include predicting a user interaction with the device based on at least one of the user's gaze and the user's posture, detecting an actual user input associated with the input option after presenting the user guidance data, the actual user input representing the predicted user interaction, and sending a command to the device corresponding to the actual user input associated with the input option.

[0013] In some examples, outputting user guidance data corresponding to the input options can include displaying the user guidance data. In some examples, outputting user guidance data corresponding to the input options can include outputting audio data representative of the user guidance data.

[0014] In some examples, outputting user guidance data corresponding to the input option may include displaying the user guidance data and outputting audio data associated with the displayed user guidance data.

[0015] In some aspects, the methods, non-transitory computer-readable media, and apparatus described above can include receiving data from the device identifying one or more input options associated with the device. In some aspects, the methods, non-transitory computer-readable media, and apparatus described above can include receiving data from a server identifying one or more input options associated with the device.

[0016] In some cases, the device does not have an external user interface for receiving one or more user inputs.

[0017] In some aspects, the methods, non-transitory computer readable media, and apparatus described above may include refraining from presenting further user guidance data associated with the device based on the information.

[0018] In some aspects, the methods, non-transitory computer-readable media, and apparatus may include, after presenting the user guidance data, obtaining a user input associated with the input option and transmitting instructions to the device corresponding to the user input. In some cases, the instructions may be configured to control one or more operations of the device.

[0019] In some aspects, the device comprises a camera, a mobile device (e.g., a mobile phone, i.e., a so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR), an augmented reality (AR), or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the device includes a camera or cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the devices described above may include one or more sensors.

[0020] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0021] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0022] Exemplary embodiments of the present application are described in detail below with reference to the following drawings: [Brief description of the drawings]

[0023] [Figure 1] FIG. 1 is a block diagram illustrating an example extended reality (XR) system, according to some examples of the present disclosure. [Diagram 2] FIG. 1 illustrates exemplary landmark points on a hand that may be used to track hand position and hand interaction with a virtual environment, according to some examples of the present disclosure. [Diagram 3] FIG. 1 illustrates an example of an XR system worn by a user, according to some examples of the present disclosure. [Figure 4A] FIG. 1 illustrates an example of a user using an XR system to interact with an elevator control panel, according to some examples of the present disclosure. [Figure 4B] FIG. 1 illustrates an example of a user using an XR system to interact with an elevator control panel, according to some examples of the present disclosure. [Figure 4C] FIG. 1 illustrates an example of a user using an XR system to interact with an elevator control panel, according to some examples of the present disclosure. [Figure 5A] FIG. 1 illustrates an example of a user using an XR system to interact with a remote control capable of controlling a television, according to some examples of the present disclosure. [Figure 5B] FIG. 1 illustrates an example of a user using an XR system to interact with a remote control capable of controlling a television, according to some examples of the present disclosure. [Figure 6A] FIG. 1 illustrates an example of a user using an XR system to interact with a thermostat, according to some examples of the present disclosure. [Figure 6B] FIG. 1 illustrates an example of a user using an XR system to interact with a thermostat, according to some examples of the present disclosure. [Figure 7A] FIG. 1 illustrates an example of a user using an XR system that can determine whether to provide user interface input options for interacting with an image frame, according to some examples of the present disclosure. [Figure 7B] FIG. 1 illustrates an example of a user using an XR system that can determine whether to provide user interface input options for interacting with an image frame, according to some examples of the present disclosure. [Figure 8A] FIG. 2 illustrates an example of a user using an extended reality system to interact with one or more devices when there are multiple devices in the scene, according to some examples of the present disclosure. [Figure 8B]FIG. 2 illustrates an example of a user using an extended reality system to interact with one or more devices when there are multiple devices in the scene, according to some examples of the present disclosure. [Figure 8C] FIG. 2 illustrates an example of a user using an extended reality system to interact with one or more devices when there are multiple devices in the scene, according to some examples of the present disclosure. [Figure 8D] FIG. 2 illustrates an example of a user using an extended reality system to interact with one or more devices when there are multiple devices in the scene, according to some examples of the present disclosure. [Figure 9] 1 is a flow diagram illustrating an example of a process for presenting information associated with at least one input option according to some examples of the present disclosure. [Figure 10] 1 illustrates an exemplary computing system in accordance with some examples of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0024] Specific aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purposes of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0025] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0026] A user often interacts with different devices that can provide specific functionality of interest to the user. For example, a user may interact with smart devices (e.g., Internet of Things (IoT) or other network-connected devices), mobile devices, control devices (e.g., remote controls for televisions, appliances, speakers, etc.), system control panels, appliances, etc. In various illustrative examples, a user may interact with a network-connected television to manage content viewing or to change the power settings of the network-connected television, may interact with a network-connected light bulb to control the light emitted by the network-connected light bulb or the operation of the network-connected light bulb, may interact with a network router to configure the operation and settings of the network router, may interact with a network-connected thermostat to control the temperature or configuration settings of the network-connected thermostat, etc.

[0027] In some cases, a device may include a hardware user interface or may present a graphical user interface (e.g., by displaying the user interface using a display or other technology) that a user can use to interact with the device. However, in some cases, a user may have difficulty interacting with the device through the device's user interface. In one example, the user interface may be out of reach of the user, which may prevent (or make it difficult) the user from interacting with the user interface associated with the device. As another example, content displayed by the user interface may be in a language that the user does not understand, may be too small for the user to recognize, or may otherwise make it cumbersome for the user to use the user interface. In some examples, input options (e.g., supported inputs, supported input methods, etc.) associated with the device may not be readily apparent or easily understood by the user (e.g., gestures or voice commands, as opposed to user interface inputs, etc.), which may make it difficult for the user to interact with the device. In other examples, the user interface may not have accessibility settings (or appropriate accessibility settings) to assist users with visual, speech, and / or hearing impairments. In some cases, a device may not have outward / external controls that a user can see or otherwise access, and / or may not include a display to present a user interface that a user can use to interact with the device. Furthermore, often a scene may include multiple remote devices, such as, for example, a connected light bulb, a television, a connected plug, a connected speaker, etc. When there are multiple devices in a scene, it may be difficult for a user device to interact with and / or manage interactions with a particular remote device from the multiple devices in the scene.

[0028] For example, a room (e.g., kitchen, bedroom, office, living room, etc.) may have multiple devices with connectivity and / or interaction capabilities. To initiate and / or manage communication and / or interaction with one of the multiple devices, a user device may have difficulty identifying a specific device from the multiple devices, managing and / or simplifying user interactions and / or data associated with that specific device, managing content relevant to that specific device, etc. In some cases, a user device may obtain interaction data (e.g., input options / capabilities, inputs, outputs, graphical user interfaces, user interaction assistance data, etc.) from multiple devices in a scene. If the user device does not have sufficient knowledge and / or understanding of the scene and / or current context and / or does not receive clear instructions from the user, the user device may be overloaded with data (e.g., interaction data, device data, etc.). In some examples, a user device may have difficulty managing content and / or interaction with one or more of the multiple devices in a scene. However, the systems and techniques described herein may enable a user device to restrict, limit, filter, declutter, etc., virtual content associated with remote devices in a scene based on context information. In some cases, using the systems and techniques described herein, a user device can refine the content that is processed and / or presented for communication / interaction with remote devices in a scene and / or can present the content in a manner that best suits / fits the context (e.g., large, small, overlay, world-locked, head- or device-locked, etc.). In some examples, a user device can use the context information to understand how to interact with a particular remote device in a scene, the interaction data and / or virtual content associated with that particular remote device, how to manage interactions with any of the remote devices in the scene, etc.

[0029] As described further herein, in many cases, the systems and techniques described herein can enable a device to consolidate and / or simplify user interaction with remote devices in a scene. In such cases, the device can better manage, consolidate, simplify, and / or facilitate communication and / or interaction with remote devices in a scene. In some cases, the device can facilitate and / or support user interaction with one or more remote devices even in more challenging scenes and / or situations. To illustrate, a user and / or a device associated with the user may have difficulty interacting with the device. For example, a user may have difficulty interacting with a television remote control under poor lighting conditions, when the remote control is out of the user's reach, or when the buttons on the remote control are difficult for the user to see / understand (e.g., due to the user's disability, the size of the buttons, the language of the button labels, etc.). As another example, a user may have difficulty interacting with a network router or connected device that does not have external controls (e.g., a network-connected thermostat, light bulb, speaker, camera, appliance, switch, etc.), especially if the user does not have access to a user interface for interacting with the network router or IoT device. As yet another example, a user may have difficulty interacting with a control panel, such as a vehicle or elevator control panel, if the control panel (or a particular control on the control panel) is out of reach of the user or if the user does not know which control to use for the desired operation / interaction.

[0030] As described in more detail herein, systems, apparatus, methods (also referred to as processes, and computer-readable media (collectively referred to herein as "systems and techniques") are described herein to improve, integrate, simplify, and / or facilitate user interaction with remote devices. In some examples, electronic devices may be connected devices (e.g., network-connected devices), mobile devices, devices lacking outward facing / external controls, devices lacking a display and / or user interface, devices with particular characteristics that present one or more challenges to a user wishing to interact with such devices (e.g., devices having interfaces in different languages ​​that are not recognized / understood by the user, limited accessibility options, etc.), devices that are out of reach of the user, or devices that are not accessible to the user. The electronic device may integrate, simplify, and / or facilitate interaction with other devices, such as devices with simple controls / interfaces, and / or any other devices. In some examples, the electronic device configured to facilitate interaction with other devices may include a smartphone, a smart wearable device (e.g., a smart watch, smart earphones, etc.), an extended reality (XR) system or device (e.g., smart glasses, head mounted display (HMD), etc.), etc. Although examples are described herein using an XR system as an example of an electronic device that may implement the techniques described herein, the techniques may be performed using other electronic devices (e.g., a mobile device, a smart wearable device, etc.).

[0031] In general, an XR system or device can provide virtual content to a user and / or combine a real-world or physical environment with a virtual environment (composed of virtual content) to provide an XR experience to a user. The real-world environment can include real-world objects (also referred to as physical objects), such as books, people, vehicles, buildings, tables, chairs, and / or other real-world or physical objects. An XR system or device can facilitate interaction with different types of XR environments (e.g., a user may use an XR system or device to interact with an XR environment). An XR system can include a virtual reality (VR) system that facilitates interaction with a VR environment, an augmented reality (AR) system that facilitates interaction with an AR environment, a mixed reality (MR) system that facilitates interaction with an MR environment, and / or other XR systems. As used herein, the terms XR system and XR device are used interchangeably. Examples of XR systems or devices include, among others, HMDs, smart glasses (e.g., network-connected glasses that can communicate using a communication network).

[0032] AR is a technology that provides virtual or computer-generated content (called AR content) over a user's view of a physical real-world scene or environment. AR content can include virtual content such as video, images, graphic content, location data (e.g., Global Positioning System (GPS) data or other location data), sound, any combination thereof, and / or other augmented content. AR systems or devices are designed to enhance (or augment) a real person's current perception, rather than replacing it. For example, a user can see an actual stationary or moving physical object through an AR device display (e.g., the lenses of AR glasses), but the user's visual perception of the physical object may be augmented or enhanced by a virtual image of that object (e.g., a real-world car replaced by a virtual image of a Delorean), by AR content added to the physical object (e.g., virtual wings added to a live animal), by AR content displayed against the physical object (e.g., informational virtual content displayed near a sign on a building, a virtual coffee cup virtually fixed to a real-world table in one or more images, etc.), and / or by displaying other types of AR content. Various types of AR systems can be used for gaming, entertainment, and / or other applications.

[0033] In some cases, two types of AR systems that can be used to provide AR content include video see-through (also called video pass-through) displays and optical see-through displays. Video see-through and optical see-through displays can be used to enhance a user's visual perception of real-world or physical objects. In a video see-through system, a live video of a real-world scenario is displayed (e.g., including one or more objects augmented or augmented on the live video). A video see-through system can be implemented using a mobile device (e.g., video on a mobile phone display), an HMD, or other suitable device that can display video and computer-generated objects on the video.

[0034] An optical see-through system with AR capabilities can display AR content directly on a view of a real-world scene (e.g., without displaying video content of the real-world scene). For example, a user can see a physical object in the real-world scene through a display (e.g., glasses or lenses), and the AR system can display AR content (e.g., projected or otherwise displayed) on the display to provide the user with an enhanced visual perception of one or more real-world objects. Examples of optical see-through AR systems or devices are AR glasses, HMDs, another AR headset, or other similar devices that can include a lens or glass in front of each eye (or a single lens or glass on both eyes), allowing a user to directly view a real-world scene with a physical object, while also allowing an augmented image of the object or additional AR content to be projected onto the display to augment the user's visual perception of the real-world scene.

[0035] VR provides a fully immersive experience in a three-dimensional computer-generated VR environment or video that depicts a virtual version of a real-world environment. The VR environment can be interacted with in a seemingly realistic or physical manner. As a user experiencing the VR environment moves in the real world, the images rendered in the virtual environment also change, giving the user the perception that they are moving within the VR environment. For example, the user can turn left or right, look up or down, and / or move forward or backward, thus changing the user's perspective of the VR environment. The VR content presented to the user can change accordingly, so that the user's experience is as seamless as the real world. VR content can, in some cases, include VR video, which can be captured and rendered with extremely high quality, potentially providing a truly immersive virtual reality experience. Virtual reality applications can include games, training, education, sports videos, online shopping, among others. VR content can be rendered and displayed using a VR system or device, such as a VRHMD or other VR headset that completely covers the user's eyes during the VR experience.

[0036] MR technology can combine aspects of VR and AR to provide an immersive experience for the user, for example, in an MR environment, real-world and computer-generated objects can interact (e.g., a real person can interact with a virtual person as if the virtual person were a real person).

[0037] In some cases, the XR system may track a part of the user (e.g., the user's hand and / or fingertip) to enable the user to interact with an item of virtual content. An XR system, such as smart glasses or an HMD, may implement a camera and / or one or more sensors to track the position of the XR system and other objects in the physical environment in which the XR system is located. The XR system may use such tracking information to provide a realistic XR experience to a user of the XR system. For example, the XR system may enable a user to experience or interact with an immersive virtual environment or content. To provide a realistic XR experience, some XR systems or devices may integrate virtual content with the physical world. In some cases, the XR system or device may match the relative pose and movement of objects and devices. For example, the XR system may use the tracking information to calculate the relative pose of the device, the object, and / or a map of the real-world environment to match the relative positions and movements of the device, the object, and / or other parts of the real-world environment. Using the pose and movement of one or more devices, objects, and / or other parts of the real-world environment, the XR system can anchor the content to the real-world environment in a way that appears realistic to a user of the XR system. The relative pose information can be used to match the virtual content with the user's perceived movements and the spatiotemporal states of the devices, objects, and other parts of the real-world environment.

[0038] In some examples, the XR system can be used to provide user guidance data for interacting with one or more other devices, such as one or more connected devices, a remote control, a control panel, a mobile device, etc. According to the systems and techniques described herein, the XR system can be leveraged to enable more intuitive and natural interactions with content and / or other devices. In some examples, the XR system can detect other devices in a scene and facilitate and / or manage interactions with the other devices. In some cases, the XR system can have pre-established connections (e.g., pairing, etc.) with other devices, which allows the XR system to detect other devices in the scene. In other examples, the XR system can maintain a map of the scene including other devices located in the scene, which can be used to detect other devices when the XR system is in the scene. In some examples, the XR system can use one or more sensors, such as image sensors, audio sensors, radar sensors, LIDAR sensors, etc., that the XR system can use to detect other devices in the scene. In some cases, the XR system can use context information associated with the scene to determine that other devices are in the scene. For example, the XR system may include context information about the scene and / or other devices that indicates to the XR device that other devices are present in the scene. When the XR system is in the scene, the XR system can determine that other devices are in the scene based on the context information. The XR system can obtain information from other devices detected in the scene to facilitate, manage, etc., interactions with the other devices.

[0039] For example, the XR system may obtain input data from other devices detected in the scene that indicates or identifies one or more input options for the other devices, and / or context information associated with the scene, the other devices, and / or the XR system. The XR system may process the input data and / or context information and present input options to a user using (e.g., wearing) the XR system for interacting with (e.g., controlling, accessing content, accessing status information, accessing output, etc.) the user interface, virtual content, and / or device. In some cases, the XR system may also be used to control other devices in the scene based on the input data and / or context information, as described further herein.

[0040] In some examples, the XR system can obtain or receive input data indicating one or more input options from other devices, from a server (e.g., a cloud-based server responsible for the operation of other devices), and / or from another source. The input data can indicate some or all of the input options that can be used to interact with the other device, such as an input type (e.g., gesture-based, voice-based, touch-based, etc.), a function based on a particular input (e.g., a swipe to the right can cause a thermostat to turn up the temperature, etc.), among other information. In some examples, if the device does not communicate input data to the XR system (e.g., after the XR system has sent a request to the device), the XR system can determine or infer that the device does not have the capability to communicate with the XR system (e.g., over a wireless network). In such examples, the XR system can present instructions or other information to help the user determine how to interact with the device (e.g., highlight a particular button on the device).

[0041] The XR system can use the context information to determine the content, input options, user interface, and / or modality to output for the user to use to interact with other devices. In some examples, the XR system can obtain or receive the context information locally (e.g., using one or more sensors of the XR system, such as one or more cameras, one or more inertial measurement units (IMUs), etc.) and / or from one or more remote sources (e.g., a server, the cloud, the Internet, other devices, one or more sensors, etc.). The context information can relate to other devices (with which the user may interact using the XR system), the scene or environment in which the XR system and / or other devices are located, a user of the XR system attempting to interact with other devices, and / or other context at any given time. For example, the context information may include the intended user interaction with other devices, one or more actions of the user within the scene, characteristics or personal information associated with the user, historical information associated with the user and other devices (e.g., past uses of the devices by the user, etc.), user interface capabilities of the other devices (e.g., whether they have outward / external controls that the user can see or otherwise access), information associated with the other devices (e.g., how far the other devices are from the XR system), information associated with the scene (e.g., lighting, noise, etc.), and / or other information.

[0042] The context information can provide the XR system with contextual awareness of a situation / context associated with the user, the XR system, other devices, a scene, etc. Using the context information and data identifying input options associated with the other devices, the XR system can output (e.g., present, provide, generate, etc.) user interaction data corresponding to one or more input options that enable user interaction with the other devices. In some cases, the XR system can present visual content / data corresponding to one or more input options. For example, the user interaction data can include one or more user interface elements associated with the input option, cues indicating how to provide input associated with the input option, and / or other data. In some cases, the XR system can additionally or alternatively output non-visual data corresponding to one or more input options. For example, the XR system can output haptic and / or audio information, such as audio cues or instructions corresponding to one or more input options.

[0043] The context information may enable the XR system to output content (e.g., visual virtual content, audio content, haptic feedback, etc.) that is contextually appropriate given the situation / context associated with the user, the XR system, other devices, the scene, and / or that facilitates interaction with other devices. In some examples, the XR system may use the context information to simplify user interaction and / or associated data / content. For example, in some cases, other devices may have multiple input features / capabilities. The XR system may use the context information to filter input options associated with the XR system and / or devices that are relevant to the current context. The XR system may output a filtered / reduced number of input options to simplify user interaction, user interaction data / content, etc.

[0044] As previously described, the XR system can use the input data and context information to render a user interface and / or input options for a user to interact with other devices from the XR system. In some cases, the XR system can communicate with other devices to provide input / commands to the other devices based on user interactions with the user interface rendered by the XR system. For example, the XR system can display a user interface to be seen by the user as an overlay on the other device or a portion of the other device (e.g., a control device, a surface, a display, a panel, etc.) to facilitate user interaction with the other device. The XR system can detect and translate user interactions with the user interface to generate control commands and / or user interaction instructions. For example, the XR system can detect user gestures and / or inputs via the device (e.g., via the XR system and / or a controller) and interpret / translate such user gestures and / or inputs as interactions with the user interface. The XR system can then generate commands / instructions to control the other devices based on the interpreted / translated user gestures and / or inputs. Interaction with the overlaid user interface may be used by the XR system to control other devices based on user input and / or access data / output from the other devices. In some cases, to facilitate and / or improve user interaction with the other devices, the XR system may locate and / or map the other devices such that the user interface can be accurately rendered on the other devices or portions of the other devices.

[0045] In some examples, the XR system may present a user interface overlaid on the scene in a world-locked or screen-locked manner to provide user interaction guidance data and / or input options to the user via the user interface. In some cases, the overlay may include one or more graphical user interface elements with guidance information that indicates to the user how to interact with the one or more graphical user interface elements.

[0046] Further details regarding the generation of virtual private spaces are provided herein with respect to various figures. FIG. 1 is a diagram illustrating an example extended reality (XR) system 100 according to some aspects of the disclosure. The XR system 100 can execute (or execute) XR applications and perform XR actions. In some examples, the XR system 100 can perform tracking and locating, mapping, and positioning and rendering of virtual content on the display 109 as part of an XR experience. For example, the XR system 100 can generate a map (e.g., a three-dimensional (3D) map) of a scene in the physical world, track the attitude (e.g., position and location) of the XR system 100 relative to the scene (e.g., relative to the 3D map of the scene), place and / or anchor virtual content at a particular location on the map of the scene, and render the virtual content on the display 109 such that the virtual content appears to be at a location in the scene that corresponds to the particular location on the map of the scene where the virtual content is placed and / or anchored. The display 109 may include glasses, a screen, one or more lenses, a projector, and / or other display mechanisms that allow a user to view the real-world environment and further display XR virtual content thereon.

[0047] In this illustrative example, the XR system 100 includes one or more image sensors 102, accelerometers 104, gyroscopes 106, storage devices 107, compute components 110, an XR engine 120, an input options engine 122, a context management engine 123, an image processing engine 124, and a rendering engine 126. It should be noted that the components 102-126 shown in FIG. 1 are non-limiting examples provided for illustrative and descriptive purposes, and other examples may include more, fewer, or different components than those shown in FIG. 1. For example, in some cases, the XR system 100 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, audio sensors, etc., other than the accelerometer 104 and gyroscope 106), one or more display devices, one or more other processing engines, one or more other hardware components, and / or one or more other software and / or hardware components not shown in FIG. 1. An example architecture and example hardware components that may be implemented by the XR system 100 are further described below in conjunction with FIG.

[0048] Further, for simplicity and illustrative purposes, one or more image sensors 102 will be referred to herein (e.g., in the singular) as image sensor 102. However, as one of ordinary skill in the art will appreciate, XR system 100 may include a single image sensor or multiple image sensors. Additionally, reference to any of the components (e.g., 102-126) of XR system 100 in the singular or plural should not be construed as limiting the number of such components implemented by XR system 100 to one or more than one. For example, reference to accelerometer 104 in the singular should not be construed as limiting the number of accelerometers implemented by XR system 100 to one. As one of ordinary skill in the art will appreciate, for any of the components 102-126 shown in FIG. 1, XR system 100 may include only one of such components or more than two of such components.

[0049] The XR system 100 includes or is in communication with (wired or wireless) an input device 108. The input device 108 may include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, any combination thereof, and / or other input devices. In some cases, the image sensor 102 may capture images that can be processed to interpret the gesture commands.

[0050] The XR system 100 may be part of or implemented by a single computing device or multiple computing devices. In some examples, the XR system 100 may be part of an electronic device, such as an extended reality head mounted display (HMD) device, extended reality glasses (e.g., extended reality glasses or AR glasses), a camera system (e.g., digital camera, IP camera, video camera, security camera, etc.), a telephone system (e.g., smartphone, mobile phone, conferencing system, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a smart television, a display device, a game console, a video streaming device, an Internet of Things (IoT) device, and / or any other suitable electronic device.

[0051] In some implementations, one or more of the image sensor 102, the accelerometer 104, the gyroscope 106, the storage 107, the compute component 110, the XR engine 120, the input options engine 122, the context management engine 123, the image processing engine 124, and the rendering engine 126 may be part of the same computing device. For example, in some cases, one or more of the image sensor 102, the accelerometer 104, the gyroscope 106, the storage 107, the compute component 110, the XR engine 120, the input options engine 122, the context management engine 123, the image processing engine 124, and the rendering engine 126 may be incorporated into an HMD, extended reality glasses, a smartphone, a laptop, a tablet computer, a gaming system, and / or any other computing device. However, in some implementations, one or more of the image sensor 102, accelerometer 104, gyroscope 106, storage 107, compute component 110, XR engine 120, input options engine 122, context management engine 123, image processing engine 124, and rendering engine 126 may be part of two or more separate computing devices. For example, in some cases, some of the components 102-126 may be part of or implemented by one computing device, and remaining components may be part of or implemented by one or more other computing devices.

[0052] The storage device 107 may be any storage device for storing data. Furthermore, the storage device 107 may store data from any component of the XR system 100. For example, the storage device 107 may store data from the image sensor 102 (e.g., image or video data), data from the accelerometer 104 (e.g., measurements), data from the gyroscope 106 (e.g., measurements), data from the compute component 110 (e.g., processing parameters, preferences, virtual content, rendering content, scene maps, tracking and localization data, object detection data, privacy data, XR application data, face recognition data, occlusion data, etc.), data from the XR engine 120, data from the input options engine 122, data from the context management engine 123, data from the image processing engine 124, and / or data from the rendering engine 126 (e.g., output frames). In some examples, the storage device 107 may include a buffer for storing frames for processing by the compute component 110.

[0053] The one or more compute components 110 may include a central processing unit (CPU) 112, a graphics processing unit (GPU) 114, a digital signal processor (DSP) 116, and / or an image signal processor (ISP) 118. The compute components 110 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality (e.g., tracking, localization, pose estimation, mapping, content fixation, content rendering, etc.), image / video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), machine learning, filtering, and any of the various operations described herein. In this example, the compute components 110 implement an XR engine 120, an input options engine 122, a context management engine 123, an image processing engine 124, and a rendering engine 126. In other examples, the compute components 110 may also implement one or more other processing engines.

[0054] Image sensor 102 can include any image and / or video sensor or capture device. In some examples, image sensor 102 can be part of a multi-camera assembly, such as a dual camera assembly, a three-camera assembly, a four-camera assembly, or other number of cameras. In some examples, image sensor 102 can include any combination of one or more visible light cameras (e.g., configured to capture monochrome or color images, such as red-green-blue or RGB images), one or more infrared (IR) cameras and / or near-infrared (NIR) cameras, one or more depth sensors, and / or other types of image sensors or cameras.

[0055] The image sensor 102 can capture images and / or video content (e.g., raw images and / or video data), which can be processed by the compute component 110, the XR engine 120, the input options engine 122, the context management engine 123, the image processing engine 124, and / or the rendering engine 126, as described herein. For example, the image sensor 102 can capture image data, generate frames based on the image data, and / or provide image data or frames to the compute component 110, the XR engine 120, the input options engine 122, the context management engine 123, the image processing engine 124, and / or the rendering engine 126 for processing. A frame can include a video frame of a video sequence or a still image. A frame can include an array of pixels representing a scene. For example, the frame may be a Red-Green-Blue (RGB) frame, having red, green, and blue color components per pixel, a Luminance, Red-Difference, Blue-Difference (YCbCr) frame, having one Luminance component and two Chrominance (color) components (Red Chrominance and Blue Chrominance) per pixel, or any other suitable type of color or monochrome image.

[0056] In some cases, the image sensor 102 (and / or other cameras of the XR system 100) can be configured to also capture depth information. For example, in some implementations, the image sensor 102 (and / or other cameras) can include an RGB-depth (RGB-D) camera. In some examples, the XR system 100 can include one or more depth sensors (not shown) that are separate from the image sensor 102 (and / or other cameras) and can capture depth information. For example, such depth sensors can obtain depth information independent of the image sensor 102. In some examples, the depth sensor can be physically located in the same general location as the image sensor 102, but can operate at a different frequency or frame rate than the image sensor 102. In some examples, the depth sensor can take the form of a light source that can project a structured or textured light pattern, which may include one or more narrow bands of light, onto one or more objects in the scene. The depth information can then be obtained by exploiting the geometric distortion of the projected pattern caused by the surface shape of the object. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered with a camera (eg, an RGB camera).

[0057] The XR system 100 also includes one or more sensors other than the image sensor 102. The one or more sensors may include one or more accelerometers (e.g., accelerometer 104), one or more gyroscopes (e.g., gyroscope 106), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the compute component 110. For example, the accelerometer 104 may detect acceleration by the XR system 100 and generate acceleration measurements based on the detected acceleration. In some cases, the accelerometer 104 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that may be used to determine a position or attitude of the XR system 100. The gyroscope 106 may detect and measure the orientation and angular velocity of the XR system 100. For example, the gyroscope 106 may be used to measure the pitch, roll, and yaw of the XR system 100. In some cases, the gyroscope 106 can provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 102 and / or the XR engine 120 can use measurements obtained by the accelerometer 104 (e.g., one or more translation vectors) and / or the gyroscope 106 (e.g., one or more rotation vectors) to calculate the attitude of the XR system 100. As previously mentioned, in other examples, the XR system 100 can also include other sensors such as an inertial measurement unit (IMU), a magnetometer, gaze and / or gaze tracking sensors (e.g., gaze tracking cameras), machine vision sensors, smart scene sensors, voice recognition sensors, impact sensors, shock sensors, position sensors, tilt sensors, etc.

[0058] In some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure a specific force, angular velocity, and / or orientation of the XR system 100. In some examples, the one or more sensors may output measurement information associated with the capture of an image captured by the image sensor 102 (and / or other cameras of the XR system 100) and / or depth information obtained using one or more depth sensors of the XR system 100.

[0059] The output of one or more sensors (e.g., accelerometer 104, gyroscope 106, one or more other types of IMUs, and / or other sensors) may be used by the extended reality engine 120 to determine the pose of the XR system 100 (also called head pose) and / or the pose of the image sensor 102 (or other camera of the XR system 100). In some cases, the pose of the XR system 100 and the pose of the image sensor 102 (or other camera) may be the same. The pose of the image sensor 102 refers to the position and orientation of the image sensor 102 with respect to a reference frame (e.g., with respect to the object 202). In some implementations, the pose of the camera can be determined in six degrees of freedom (6DOF), which refers to three translational components (e.g., which may be given in X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as an image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame).

[0060] In some cases, a device tracker (not shown) can track the pose (e.g., 6DOF pose) of the XR system 100 using measurements from one or more sensors and image data from the image sensor 102. For example, the device tracker can fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and movement of the XR system 100 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 100, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map of the scene. The 3D map updates can include, for example, but not limited to, new or updated features and / or features or landmark points associated with the scene and / or the 3D map of the scene, location updates that identify or update the position of the XR system 100 within the scene and the 3D map of the scene, etc. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The XR system 100 can use the mapped scene (e.g., a scene of the physical world represented by and / or associated with the 3D map) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0061] In some aspects, the pose of the image sensor 102 and / or the entire XR system 100 can be determined and / or tracked by the compute component 110 using a visual tracking solution based on images captured by the image sensor 102 (and / or other cameras of the XR system 100). For example, in some examples, the compute component 110 can perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, the compute component 110 can perform SLAM or can be in communication (wired or wireless) with a SLAM engine (not shown). SLAM refers to a class of techniques in which a map of an environment (e.g., a map of the environment being modeled by the XR system 100) is created while simultaneously tracking the pose of the camera (e.g., the image sensor 102) and / or the XR system 100 relative to that map. The map may be referred to as a SLAM map and may be three-dimensional (3D). SLAM techniques can be performed using color or grayscale image data captured by the image sensor 102 (and / or other cameras of the XR system 100) and can be used to generate estimates of 6DOF attitude measurements of the image sensor 102 and / or the XR system 100. Such SLAM techniques configured to perform 6DOF tracking can be referred to as 6DOF SLAM. In some cases, the output of one or more sensors (e.g., the accelerometer 104, the gyroscope 106, one or more IMUs, and / or other sensors) can be used to estimate, correct, and / or adjust the estimated attitude.

[0062] In some cases, 6DOF SLAM (e.g., 6DOF tracking) can associate features observed from a particular input image from the image sensor 102 (and / or other camera) to a SLAM map. For example, 6DOF SLAM can use the association of feature points from the input image to determine the pose (position and orientation) of the image sensor 102 and / or the XR system 100 relative to the input image. 6DOF mapping can also be performed to update the SLAM map. In some cases, a SLAM map maintained using 6DOF SLAM can include 3D feature points triangulated from two or more images. For example, key frames can be selected from the input images or video stream to represent the observed scene. For every key frame, a respective 6DOF camera pose associated with the image can be determined. The pose of the image sensor 102 and / or the XR system 100 can be determined by projecting features from the 3D SLAM map onto the image or video frame and updating the camera pose from the verified 2D-3D correspondence.

[0063] In one illustrative example, the compute component 110 may extract feature points from all input images or from each key frame. As used herein, a feature point (also called a registration point) is a distinctive or identifiable part of an image, such as a part of a hand, an edge of a table, among others. The features extracted from the captured image may represent distinctive feature points along a three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and every feature point may have an associated feature location. The feature points in a key frame may match (be the same as or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations that are used to examine one or more pixels of an image to determine whether a feature is present at a particular pixel. Feature detection may be used to process the entire captured image or some portion of an image. For each image or key frame, once a feature is detected, a local image patch near the feature may be extracted. Features may be extracted using any suitable technique, such as Scale Invariant Feature Transform (SIFT) (which locates features and generates their descriptions), Speed ​​Up Robust Features (SURF), Gradient Oriented Position Histogram (GLOH), Normalized Cross Correlation (NCC), or other suitable technique.

[0064] In some cases, the XR system 100 may also track the user's hands and / or fingers to enable the user to interact with and / or control virtual content within the virtual environment (e.g., virtual content displayed in a virtual private space). For example, the XR system 100 may track the pose and / or movement of the user's hands and / or fingertips to identify or translate user interactions with the virtual environment. User interactions may include, for example, without limitation, moving an item of virtual content, resizing an item of virtual content and / or a location in a virtual private space, selecting an input interface element in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interface), providing input through a virtual user interface, etc.

[0065] 2 is a diagram illustrating example landmark points of a hand 200 that may be used to track the position of the hand 200 and its interaction with a virtual environment, such as virtual content displayed within a virtual private space as described herein. The landmark points illustrated in FIG. 2 correspond to different parts of the hand 200, including a landmark point 235 on the palm of the hand 200, a landmark point on the thumb 230 of the hand 200, a landmark point on the index finger 232 of the hand 200, a landmark point on the middle finger 234 of the hand 200, a landmark point on the ring finger 236 of the hand 200, and a landmark point on the pinky finger 238 of the hand 200. The palm of the hand 200 can move in three translational directions (e.g., measured in the X, Y, and Z directions relative to a plane, such as the image plane) and in three rotational directions (e.g., measured in yaw, pitch, and roll relative to the plane), thus providing six degrees of freedom (6DOF) that may be used for registration and / or tracking. The 6 DOF movements of the palm are shown as squares in FIG. 2, as indicated in legend 240.

[0066] Different joints of the fingers of the hand 200 allow different degrees of movement, as shown in legend 240. As shown by the diamond shapes (e.g., diamond 233) in FIG. 2, the base of each finger (corresponding to the metacarpophalangeal joint (MCP) between the proximal phalanges and the metacarpals) has two degrees of freedom (2DOF) corresponding to flexion and extension, as well as abduction and adduction. As shown by the circle shapes (e.g., circle 231) in FIG. 2, the top joint of each finger (corresponding to the interphalangeal joint between the distal, middle, and proximal phalanges) has one degree of freedom (1DOF) corresponding to flexion and extension. As a result, the hand 200 provides 26 degrees of freedom (26DOF) from which to track the hand 200 and for the hand 200 to interact with virtual content rendered by the XR system 100.

[0067] The XR system 100 can use one or more of the landmark points on the hand 200 to track the hand 200 (e.g., track the pose and / or movement of the hand 200) and track interactions with the virtual environment rendered by the XR system 100. As described above, as a result of detection of one or more landmark points on the hand 200, the pose of the landmarks (and thus the hand and fingers) in a relative physical position to the XR system 100 can be established. For example, a landmark point (e.g., landmark point 235) on the palm of the hand 200 can be detected in an image, and the location of the landmark point can be determined relative to the image sensor 102 of the XR system 100. A point (e.g., a center point such as a center of gravity or other center point) of an item of virtual content rendered by the XR system 100 can be translated to a position on (or rendered on) a display (e.g., display 109 of FIG. 1) of the XR system 100 relative to the position determined for the landmark point on the palm of the hand 200.

[0068] As described below, the XR system 100 can also register the virtual content and / or the hand 200 to points in the real world (as detected in one or more images) and / or other parts of the user. For example, in some implementations, in addition to determining the physical pose of the XR system 100 (or the XR system 100) and / or the hand 200 relative to an item of virtual content, the XR system 100 can determine the location of other landmarks, such as distinct points (called feature points) on a wall, one or more corners of an object, features on a floor, points on a human face, points on a nearby device, among others. In some cases, the XR system 100 can place the virtual content within a particular location with respect to feature points detected in the environment, which can correspond, for example, to detected objects and / or people in the environment.

[0069] In some examples, the pose of the XR system 100 (and / or the user's head) can be determined using, for example, image data from the image sensor 102 and / or measurements from one or more sensors, such as the accelerometer 104, the gyroscope 106, and / or one or more other sensors (e.g., one or more magnetometers, one or more inertial measurement units (IMUs), etc.). The head pose can be used to determine the position of virtual content, the hands 200, and / or objects and / or people in the environment.

[0070] The operations of the XR engine 120, the input options engine 122, the context management engine 123, the image processing engine 124, and the rendering engine 126 (and any image processing engines) may be performed by any of the compute components 110. In an illustrative example, the operations of the rendering engine 126 may be realized by the GPU 114, and the operations of the XR engine 120, the input options engine 122, the context management engine 123, and the image processing engine 124 may be realized by the CPU 112, the DSP 116, and / or the ISP 118. In some cases, the compute components 110 may include other electronic circuitry or hardware, computer software, firmware, or any combination thereof, to perform any of the various operations described herein.

[0071] In some examples, the XR engine 120 can perform XR operations to generate an XR experience based on data from one or more sensors on the XR system 100, such as the image sensor 102, the accelerometer 104, the gyroscope 106, and / or one or more IMUs, radar, etc. In some examples, the XR engine 120 can perform tracking, localization, pose estimation, mapping, content fixation operations, and / or any other XR operations / functions. The XR experience may include use of the XR system 100 to present XR content (e.g., virtual reality content, augmented reality content, mixed reality content, etc.) to a user during a virtual session. In some examples, the XR content and experiences may be provided by the XR system 100 via an XR application (e.g., executed or implemented by the XR engine 120) that provides a particular XR experience, such as, for example, an XR gaming experience, an XR classroom experience, an XR shopping experience, an XR entertainment experience, an XR activity (e.g., an action, a troubleshooting action, etc.), among others. During an XR experience, a user can view and / or interact with virtual content using XR system 100. In some cases, a user can view and / or interact with the virtual content while also viewing and / or interacting with the physical environment around the user, allowing the user to have an immersive experience between the physical environment and the virtual content mixed or integrated with the physical environment.

[0072] The XR engine 120, the input options engine 122, and the context management engine 123 can perform various operations to determine (and manage) how, where, and / or when to render particular virtual content to one or more other devices. For example, the XR engine 120, the input options engine 122, and the context management engine 123 can facilitate interaction with other devices, such as connected devices (e.g., network-connected cameras, speakers, light bulbs, hubs, locks, plugs, thermostats, displays, alarm systems, televisions (TVs), gadgets, appliances, etc.), mobile devices, devices lacking outward / external controls, devices lacking a display and / or user interface, devices with certain characteristics that present one or more challenges to a user wishing to interact with such devices (e.g., devices having controls / interfaces in a language not understood by the user, controls / interfaces not recognized / understood by the user, limited accessibility options, input options not easily recognized / understood by the user, etc.), devices with controls / interfaces that are beyond the reach of the user, and / or any other devices. For example, the input options engine 122 can obtain or receive input data indicating or identifying one or more input options of another device. The input options engine 122 can transmit the input data to the XR engine 120. The context management engine 123 can obtain or receive information (e.g., context information, etc.) regarding other devices (with which a user may interact using the XR system 100), a scene or environment in which the XR system and / or other devices are located, a user of the XR system intending to interact with other devices, and / or other contexts. The context management engine 123 can transmit the context information to the XR engine 120.

[0073] The XR engine 120 can use the input data indicating one or more input options for the device and / or the context information to cause the rendering engine 126 to present relevant information to the user. For example, the XR engine 120 can cause the rendering engine 126 to output guidance data corresponding to the input options for which the relevant context information has been determined. The guidance data can inform the user which input options are available for other devices, how to provide such input, etc. In some examples, the guidance data can filter out input options (and / or related information) that may be less relevant or unavailable given the current context (e.g., given information relevant to the XR system 100, the scene, other devices, and / or the user associated with the XR system 100). In some examples, the input data indicating one or more input options for the device and / or the context information can cause the rendering engine 126 to present user interfaces and / or input options for interacting with the device (e.g., control, access content, access status information, access output, etc.).

[0074] For example, based on the context information and the input data identifying input options associated with the other device, the XR engine 120 can cause the rendering engine 126 to present user interaction data corresponding to one or more input options that enable user interaction with the other device. For example, the user interaction data can include one or more user interface elements (e.g., selectable control options, etc.) associated with the input option, cues (e.g., highlights, arrows, text, etc.) indicating how to provide input associated with the input option, and / or other data. The context information can enable the XR engine 120 to present content (e.g., virtual content, audio content, user interface content, etc.) that is contextually appropriate given the situation / context associated with the user, the XR system 100, the other devices, the scene, and / or that otherwise facilitates interaction with the other devices. In some cases, the XR engine 120 can cause the rendering engine 126 to not present content or to reduce or filter the amount of virtual content to display based on the context (e.g., to display a subset of user interface options).

[0075] In some cases, the XR engine 120 can leverage its XR capabilities to facilitate interaction with other devices. For example, the XR system 100 can have AR capabilities, such as the ability to display virtual content on the display of the XR system 100 while also allowing the user to view the real-world environment through the display. The XR engine 120 can leverage its AR capabilities to render a user interface for the user to directly interact with the other device, or by providing input to the XR engine 120 (e.g., using the input device 108). For example, the rendering engine 126 can render a user interface on the display such that the user interface appears to the user as an overlay on the other device or a portion of the other device (e.g., a control, a surface, a display, a panel, etc.) to facilitate user interaction with the other device. Interaction with the overlaid user interface can be used by the XR engine 120 (or other components of the XR system 100) to control the other device based on the user input and / or to access data / output from the other device. In some cases, the XR system 100 may locate and / or map the other device to facilitate and / or improve user interaction with the other device. For example, the XR system 100 may locate the other device and use the location information to render a user interface overlaid on the other device or a portion of the other device.

[0076] The XR system 100 can register or anchor (e.g., be positioned relative to) detected feature points in a scene. For example, the input options engine 122, the context management engine 123, and / or the image processing engine 124, in conjunction with the XR engine 120 and / or the rendering engine 126, can anchor virtual content of a user interface to feature points of a surface on which the virtual content will be displayed.

[0077] In some examples, the XR system 100 can communicate with one or more other devices to provide input / commands to the other devices based on user interactions with a user interface rendered by the rendering engine 126. For example, the XR engine 120 can cause the XR system 100 to send commands (using a transmitter or transceiver) to the devices based on user input received via the rendered user interface and / or input options. The commands can cause the devices to perform one or more functions based on the user input.

[0078] In other examples, the XR engine 120 can leverage one or more interaction modes (e.g., visual, audio / voice, gesture-based, motion-based, etc.) to facilitate user interaction with other devices. For example, the XR engine 120 can use hand tracking and / or gesture recognition capabilities to enable a user to interact (e.g., control, access, etc.) with other devices using gestures and other interactions. As another example, the XR engine 120 can use voice recognition to enable a user to use voice commands to interact with other devices.

[0079] As previously discussed, the input data can indicate some or all of the input options that can be used to interact with another device. For example, the input options can include, among other information, the types of input supported by the device (e.g., gesture-based, voice-based, touch-based, etc.), functions the device can perform based on a particular input (e.g., a swipe right can cause a thermostat to turn up the temperature, etc.).

[0080] In some examples, the input options engine 122 can obtain or receive input data indicating one or more input options of another device from another device, from a server (e.g., a cloud-based server responsible for the operation of the other device), and / or from another source. For example, the input options engine 122 can send (or cause a transmitter or transceiver to send) a request for input options to the device. In one example, the XR system 100 can detect or sense a device based on a previous network pairing with the device, based on a periodic beacon signal transmitted (e.g., broadcast) by the device indicating its presence, based on detection of the device in one or more images provided by the image sensor 102, etc. In response to detecting or sensing the device, the XR system 100 can request input data from the device. In response to the request, the device can respond with input data indicating any input options associated with the device. In another example, the input options engine 122 can request input options for a given device from a server associated with the device (e.g., a Google™ server associated with a Google Home™ device). The server can respond with input data indicating the input options associated with the device.

[0081] In some cases, the input options engine 122 can determine one or more input options of the other device by processing one or more images of the other device captured by the image sensor 102. For example, using an elevator control panel as an illustrative example of a user interface of a device (elevator), the input options engine 122 can receive an image of the elevator control panel from the image sensor 102. Using machine learning (e.g., using one or more neural network based object detectors or classifiers), computer vision (e.g., using computer vision based object detectors or classifiers), or other image analysis techniques, the input options engine 122 can determine that the control panel includes 15 numbers corresponding to floors of a building, an open door button, a close door button, an emergency button, and / or other physical or virtual buttons.

[0082] In some examples, if the device does not communicate input data to the XR system (e.g., after the XR system has sent a request to the device), the XR system may determine or infer that the device does not have the capability to communicate with the XR system (e.g., over a wireless network). In such examples, the XR system may present instructions or other information to help the user determine how to interact with the device (e.g., highlighting a particular button on the device). For example, in the elevator example above, the elevator may not have the capability to communicate with the XR system (e.g., over a communication network). In such examples, instead of sending one or more commands to control the elevator based on received user input, the XR system may present virtual content to help the user interact with the elevator.

[0083] The context management engine 123 can send context information to the XR engine 120. The context information provides the XR engine 120 with contextual awareness of a situation / context associated with a user, the XR system 100, other devices, a scene, etc. The XR engine 120 can use the context information to manage, adjust, and / or determine content, input options, user interfaces, and / or modalities to present to the user for interacting with other devices. For example, as described above, the context information can be used by the XR engine 120 to present content (e.g., virtual content, audio content, user interface content, etc.) that is contextually relevant to the situation / context associated with the user, the XR system 100, other devices, a scene, and / or that facilitates interaction with other devices.

[0084] The context information may relate to other devices (with which the user may interact using the XR system 100), a scene or environment in which the XR system 100 and / or other devices are located, a user of the XR system 100 attempting to interact with the other device, and / or other context at any given time. In some examples, the context information may include an intended interaction of the user with the other device, such as an intent to interact with the device, an intent to interact with a particular input option of the device's user interface (e.g., a particular user interface control element), and / or other intended user interaction. The context management engine 123 may infer the intended user interaction based on the gaze, a particular gesture being performed, a user holding the other device, a user walking towards the device, and / or other information, etc. For example, the context management engine 123 may determine that the user is gazing at a thermostat and, based on the determined gaze, determine that the user is attempting to interact with the thermostat. In another example, the context information may include one or more actions of the user in the scene. For example, the one or more actions may include, among others, a user walking towards the device, walking towards a door in the scene, a user sitting in a chair where one typically watches television, etc. Other examples of context information include characteristics associated with the user (e.g., visual quality, spoken language, etc.), historical information associated with the user and other devices (e.g., past use of the device by the user, the user's experience level with the device or similar devices, etc.), user interface capabilities of the other devices (e.g., whether they have outward / external controls that the user can see or otherwise access), information associated with the other devices (e.g., how far away the other devices are from the XR system 100), information associated with the scene (e.g., lighting, noise levels such as ambient sounds, objects or other obstructions between the XR system 100 and the device, whether other users are present in the scene, etc.), and / or other information.

[0085] In some examples, the context management engine 123 can determine, obtain, or receive context information locally. For example, the context management engine 123 can obtain sensor information from one or more sensors of the XR system 100 (e.g., the image sensor 102, the accelerometer 104, the gyroscope 106, and / or other sensors of the XR system 100). The context management engine 123 can process the sensor information to determine context information such as one or more intended interactions of the user with other devices, one or more actions of the user in the scene, user interface capabilities of the other devices, information associated with the other devices (e.g., the distance of the device from the XR system 100 and thus the user), information associated with the scene (e.g., lighting, noise, objects or obstacles between the XR system 100 and the device, etc.), and / or other context information. In one illustrative example, the context management engine 123 can receive an image from the image sensor 102 indicating that the user is looking at the oven, and can also receive sensor data from the accelerometer 104 and / or the gyroscope 106 indicating that the user is walking toward the oven. Based on the image and sensor data, the context management engine 123 can determine that a user is attempting to interact with the oven.

[0086] In some examples, the context management engine 123 may obtain or receive context information from one or more remote sources (e.g., a server, a cloud, the Internet, other devices, one or more sensors, etc.). For example, the context management engine 123 may have access to a user profile stored in a network-based or cloud-based system associated with the user that indicates characteristics associated with the user (e.g., the user's visual quality, the language the user speaks, etc.) and / or historical information associated with the user and other devices (e.g., the user's experience level with one or more devices, how long the user has owned one or more devices, etc.). In some cases, the user profile may be stored locally on the XR system 100.

[0087] As previously discussed, the contextual information provides the XR engine 120 with contextual awareness so that the XR system 100 can present content that is contextually relevant to the particular situation in which the XR system 100 is being used. For example, the XR engine 120 can use the contextual information, including characteristics of the user, other devices, etc., when determining what AR content to present to the user to guide and / or assist the user interaction with other devices and / or how to present the AR content. In one example, the XR engine 120 can consider the language understood / spoken by the user to ensure that the AR content presented is in a language understood / spoken by the user. As another example, the XR engine 120 can align user interface guidance with expectations regarding the user's knowledge of the device. In one illustrative example, if the user is estimated to be a novice in using the device (e.g., based on the amount of time the user has owned the device, the number of times the user has used the device, and / or based on other factors), the XR engine 120 can output further support for the user (e.g., by presenting instructions that can be used to control the device using input). In another example, if the user is presumed to be an expert in using the device (e.g., the user is presumed to have at least a threshold amount of experience / familiarity), the XR engine 120 may not present further support for the user or may reduce / minimize such support. As another example, if the user is wearing one or more items that may affect interactions with other devices, the XR engine 120 may adjust what user interface, control, and / or guidance is output to the user to take into account what the user is wearing.In one illustrative example, the XR engine 120 can adjust the user interface, controls, and / or guidance it provides if the user is wearing gloves that may impede the user's ability to select / touch a control, if the user is wearing glasses (or no glasses) or sunglasses that may impede visibility, if the user has a medical device that limits the user's movement and ability to interact with certain controls, etc.

[0088] As another example, the XR engine 120 can use contextual information (e.g., environmental factors) associated with the scene when determining how to present AR content and / or what AR content to present to the user. For example, the XR engine 120 can take into account lighting conditions, such as poorly lit or bright lighting conditions (e.g., can suggest using audio and / or visual cues), ambient sounds (e.g., can suggest using visual cues over audio cues), the presence of other people (e.g., this can suggest a need to be discreet or private, such as not providing audio output, not presenting input options that include hand or gestural movements, etc.).

[0089] In some examples, the XR engine 120 can determine which of multiple devices the user intends to interact with. In some cases, the XR engine 120 can cause the rendering engine 126 to present a user interface that allows the user to interact with the particular device with which the user intends to interact. In some cases, the XR engine 120 can output guidance (e.g., as visual or audio content) on how to interact with and / or control the particular device. For example, in a kitchen with a smart home assistant that plays music, a connected refrigerator, and a connected stove, the user may be interested in changing the temperature of the stove. Based on the context information received from the context management engine 123, the XR engine 120 can determine that the user intends to interact with the connected stove rather than the smart home assistant or the connected refrigerator. For example, based on the context information, the XR engine 120 can determine the user's gaze, posture, and / or movement and determine that the user intends to interact with the connected stove as opposed to the smart home assistant or the connected refrigerator. In some examples, the XR engine 120 can determine a number of possible gestures that a user can perform to interact with a connected stove. For example, the XR engine 120 can determine that a user can use a knob rotation gesture to change the temperature of the stove. The XR engine 120 can provide an output that informs the user (e.g., via a visual cue, an audio cue, etc.) that the user can use a knob rotation gesture to change the temperature of the stove. In some cases, the XR engine 120 can detect a knob rotation gesture from the user and communicate an associated input to the connected stove. In some cases, the connected stove can directly detect the knob rotation gesture.

[0090] The image processing engine 124 can perform one or more image processing operations related to the virtual user interface content being presented. For example, the image processing engine 124 can perform image processing operations based on data from the image sensor 102. In some cases, the image processing engine 124 can perform image processing operations such as, for example, filtering, demosaicing, scaling, color correction, color conversion, segmentation, noise reduction filtering, spatial filtering, artifact correction, etc. The rendering engine 126 can obtain image data generated and / or processed by the compute component 110, the image sensor 102, the XR engine 120, the input options engine 122, the context management engine 123, and / or the image processing engine 124, and can render video and / or image frames for presentation on a display device.

[0091] Although the XR system 100 is shown to include certain components, one skilled in the art will appreciate that the XR system 100 may include more or less components than those shown in Figure 1. For example, the XR system 100 may also optionally include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more network interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or other hardware or processing devices not shown in Figure 1. Illustrative examples of computing systems and hardware components that may be implemented with the XR system 100 are described below with respect to Figure 9.

[0092] FIG. 3 illustrates an example of an extended reality system 300 worn by a user 301. In some cases, the extended reality system 300 is similar to the XR system 100 of FIG. 1 and can perform similar operations. The extended reality system 300 can include any suitable type of XR device or system, such as AR or MR glasses, AR, VR, or MRHMD, or other XR devices. Some examples described below can be described using AR for illustrative purposes. However, aspects described below can be applied to other types of XR, such as VR and MR. The extended reality system 300 shown in FIG. 3 can include a light see-through AR device, which allows the user 301 to view the real world while wearing the extended reality system 300.

[0093] For example, a user 301 may view an object 303 in a real-world environment on a plane 304 at a distance from the user 301. As shown in FIG. 3, the extended reality system 300 includes an image sensor 302 and a display 309. As described above, the display 309 may include glasses, a screen, lenses, and / or other display mechanisms that allow the user 301 to view the real-world environment and also allow AR content to be displayed thereon. AR content (e.g., images, videos, graphics, virtual or AR objects, or other AR content) may be projected or displayed on the display 309. In one example, the AR content may include an augmented version of the object 303. In another example, the AR content may include additional AR content related to the object 303 and / or related to one or more other objects in the real-world environment. Although one image sensor 302 and one display 309 are shown in FIG. 3, the extended reality system 300 can, in some implementations, include multiple cameras and / or multiple displays (e.g., a display for the right eye and a display for the left eye).

[0094] As discussed above with respect to FIG. 1, the XR engine 122 can utilize input data and context data indicating one or more input options for a device to cause the rendering engine 126 to present input options for interacting with a user interface and / or device (e.g., control, access content, access status information, access output, etc.). In an illustrative example, the XR system 100 can assist a user in interacting with a control panel, such as an elevator control panel, a vehicle control panel, etc. FIGS. 4A, 4B, and 4C are diagrams illustrating an example of a user using the XR system 400 to interact with an elevator control panel 410. For example, when a user wearing the XR system 400 enters an elevator, the XR system 400 can determine the user's gaze (e.g., via eye gaze scanning, based on an eye gaze camera, etc.) and / or one or more gestures, such as gestures performed using the user's hand 405. Based on the gaze and / or gestures, the XR system 400 can determine that the user is unable or having difficulty finding a floor number to press on the elevator control panel 410. For example, the XR system 400 may determine that the user's eyes are moving back and forth (as if the user is searching for the correct number).

[0095] The XR system 400 can determine that the user is searching for a control button associated with a particular floor based on context information obtained by the XR system 400. For example, the context information can include check-in information at the hotel (e.g., obtained by the XR system 400 from a server associated with the hotel), voice commands provided by the user to the XR system 400, user preferences / inputs (e.g., the user may input a floor number into the XR system 100), recognized voice detected from the user (e.g., the XR system 400 recognizes a user uttering the words "Floor 16," such as using an always-on voice), a detected room number (e.g., detected from one or more images captured by the image sensor 102, etc.).

[0096] The XR system 400 can present cues associated with a particular floor control button on an elevator control panel to guide the user to the correct button. For example, the XR system 400 can determine that the user is looking for a control button corresponding to floor 16 (associated with the button numbered "16" in FIGS. 4A and 4B). As shown in FIG. 4B, the XR system 400 can display virtual content 412 highlighting the control button corresponding to floor 16. The virtual content 412 appears as if it is overlaid on the actual control button corresponding to floor 16, helping the user easily identify the control button they are searching for. As shown in FIG. 4C, the XR system 400 can display virtual content 412 highlighting the control button corresponding to floor 16 and can also present text 414 ("Here is the button for floor 16") and an arrow icon that provides further information to help the user identify the correct control button. 4B and 4C show highlighting and text as examples of visual cues, the cues may additionally or alternatively include other types of visual cues, audio cues that indicate to the user the identity or location of the correct button (e.g., to guide the user's hand based on hand tracking), and / or other types of cues, which may be particularly useful for users with disabilities or who have difficulty remembering or learning the correct buttons.

[0097] In some cases, after multiple trips in the elevator, the XR system 400 may know that the user can find the correct button quickly and / or without help. In response, the XR system 400 may decide to stop providing virtual content to help identify the control button. In other cases, the XR system 400 may continue to provide such assistance to the user (e.g., if context information with the user's characteristics indicates that the user has a disability). In some cases, the XR system 400 may similarly assist the user in interactions with other control panels, devices, industrial machinery, etc.

[0098] In another illustrative example, the XR system 100 can help a user interact with a remote control that can be used to control a device. FIGS. 5A and 5B are diagrams illustrating an example of a user using an XR system 500 to interact with a remote control 510 that can control a television 511. The XR system 500 can determine that the user is likely to have difficulty using the remote control 510 based on various factors. For example, the XR system 500 can determine that the lighting conditions in the room in which the television set 511 and the user are located are poor (e.g., based on one or more images of the room acquired using the image sensor 102, based on an ambient light sensor of the XR system 500, etc.). Furthermore, based on context information acquired by the XR system 500 that is indicative of a characteristic of the user, the XR system 500 can determine that the user has poor myopia. Additionally or alternatively, the XR system 500 can determine that a label or button on the remote control 510 is difficult to read, is in a language the user does not understand, and / or makes it difficult for the user to interact with the remote control 510. Based on this context information, the XR system 500 can determine that the user is likely having difficulty seeing the controls / labels on the remote control 510. In some cases, the XR system 500 can additionally or alternatively detect that the user is having difficulty interacting with the remote control. For example, the XR system 500 can use eye tracking to determine that the user is squinting and / or scanning the remote control for the correct button.

[0099] The XR system 500 can obtain further contextual information indicating that the user intends to switch to a particular channel (channel 34) that is playing an event of interest to the user. For example, the XR system 500 can determine that the user has scheduled a sports game on the user's digital calendar. In another example, the XR system 500 can determine that the user always watches a particular channel on a particular night. With contextual knowledge that the remote control 510 is likely to be difficult to interact with and that the user intends to switch to a particular channel (e.g., because the user frequently watches an event on a particular channel at the current time), the XR system 500 can display virtual data 512 that highlights the correct "3" and "4" buttons on the remote control 510 to help the user identify the buttons and select the buttons to switch to the particular channel 34. In some examples, the XR system 500 can sequentially highlight button "3" and then button "4" so that the user knows to press button "3" before pressing button "4."

[0100] In some cases, in addition to or instead of highlighting or highlighting numbers as described above, the XR system 500 may use audio to confirm the options (e.g., the event and / or the associated button / channel) determined by the XR system 500. For example, the XR system 500 may provide an audio prompt asking the user to confirm that they wish to switch to a particular channel playing the event. The XR system 500 may receive a confirmation from the user (e.g., based on user input provided via an input device such as the input device 108) and continue to assist. In one example, in response to receiving the confirmation, the XR system 500 may highlight the correct buttons (e.g., the "3" and "4" buttons) on the remote control 510 that correspond to the channel (e.g., channel 34). In another example, in response to receiving the confirmation, the XR system 500 may automatically send a command to the television 511 and / or the remote control 510 to cause the television 511 to change the channel (e.g., channel 34).

[0101] In another illustrative example, the XR system 100 can help a user interact with a thermostat. FIGS. 6A and 6B are diagrams illustrating an example of a user using the XR system 600 to interact with a thermostat 610. For example, the thermostat 610 can be configured to interpret one or more gestures using gesture recognition and can perform one or more functions based on the detected gestures. However, the user may not know the correct gesture command that can be used to cause the thermostat 610 to perform a particular function. Similar to what has been discussed above, the XR system 600 can use the obtained context information to determine that the user is having trouble interacting with the thermostat 610. For example, the XR system 600 can use eye tracking to determine that the user is gazing at the thermostat 610 and / or can process one or more images to determine that the user is performing a gesture (e.g., using the hand 605) but the thermostat is not performing any function based on the gesture. The XR system 600 can use any other contextual information to determine that the user is having difficulty interacting with the thermostat 610.

[0102] In some cases, the user can issue a voice command or other input that identifies the interaction they want to have with the thermostat 610. For example, the user can state, "set the temperature to 68 degrees," and the XR system 600 can recognize the voice command. The XR system 600 can determine gesture commands that the thermostat 610 is configured to interpret, for example, from the thermostat 610, from a server associated with the thermostat 610 (e.g., a Nest™ server associated with a Nest™ thermostat).

[0103] As shown in FIG. 6B, upon recognition of the voice command, receipt of another input indicating the user's desired setting, and / or a determination that the user is having difficulty interacting with the thermostat 610, the XR system 600 can present the user with a set of gesture commands (including gesture command 612, gesture command 614, and gesture command 616) that can be applied (e.g., with or without a voice command) to cause the thermostat 610 to perform the desired function. In some cases, as shown in FIG. 6B, the gesture commands can have corresponding numbers indicating the order in which the gesture commands 612, 614, 616 should be performed to cause the thermostat 610 to perform the desired function. For example, the user can perform gesture command 612 to cause the thermostat 610 to enter a temperature adjustment mode. The user can then perform gesture command 614 to cause the thermostat 610 to increase the temperature. For example, each time the user performs a "thumbs up" gesture command 614, the thermostat 610 may increase the temperature by one degree F. The user may then perform a gesture command 616 to cause the thermostat 610 to exit the temperature adjustment mode.

[0104] In another illustrative example, the XR system 100 can help a user interact with a digital picture frame. Figures 7A and 7B are diagrams illustrating an example of a user 701 using an XR system 700 that can determine whether to provide user interface input options for interacting with a digital picture frame 710. For example, the digital picture frame 710 can be configured to display metadata (e.g., a description of the displayed technology, background, people, etc.) related to the content displayed by the digital picture frame 710. For example, as shown in Figure 7B, the digital picture frame 710 can display metadata 712 next to a photo of the two dancing with the caption "This is Nancy and Bob on her birthday."

[0105] The metadata may not be of interest to the user 701 if the user 701 is not gazing at the digital picture frame 710, if the user is walking quickly by the digital picture frame 710, and / or is otherwise not likely to be interested in the metadata 712. The XR system 700 may obtain context information such as a detected gaze indicating that the user is not looking at the digital picture frame 710, a detected motion indicating that the user is walking, the user's 701 calendar (e.g., indicating that the user has an upcoming appointment at another location), the user's 701 preferences, user history data, user communications, and / or other context information. Based on the context information obtained by the XR system 700, the XR system 700 may determine that the user is walking by the digital picture frame 710, the user is not looking at the digital picture frame 710, and / or the user is not likely to be interested in the metadata 712. The XR system 700 can then send a command to the digital picture frame 710 to prevent it from presenting the metadata 712 (which could potentially distract the user).

[0106] 8A is a diagram illustrating an example of a user 802 using an XR system 800 to interact with one or more devices when multiple devices are present in a scene. In this example, the scene includes a digital picture frame 810, a remote control 812, and a digital thermostat 814. The user 802 can use the XR system 800 to interact with any of the devices in the scene including the digital picture frame 810, the remote control 812, and / or the digital thermostat 814. The XR system 800 can receive input options and / or related data from the digital picture frame 810, the remote control 812, and the digital thermostat 814, and can present user guidance data associated with the input options corresponding to the digital picture frame 810, the remote control 812, and / or the digital thermostat 814.

[0107] In some cases, when a scene includes multiple remote devices as shown in FIG. 8A, the XR system 800 may have difficulty determining which of the multiple devices in the scene the user 802 wants to interact with and / or with which the XR system 800 should present guidance data. In some examples, the XR system 800 may receive input options from a digital picture frame 810, a remote control 812, and a digital thermostat 814. The input options may include information regarding the types of inputs available / allowable on the digital picture frame 810, the remote control 812, and the digital thermostat 814. However, when receiving such data from multiple devices in the scene, the XR system 800 may become overloaded with data from the multiple devices. Data overload may make it difficult for the XR system 800 to present relevant information to the user 802, present information to the user without significant confusion, manage data and / or interactions with the devices, etc.

[0108] For example, referring to FIG. 8B, the XR system 800 can display data 820 related to the digital picture frame 810, data 822 related to the remote control 812, and data 824 related to the digital thermostat 814. The data 820, 822, and 824 can become overwhelming for the XR system 800 and / or the user 802. For example, as shown in FIG. 8B, when presented by the XR system 800, the data 820, 822, and 824 can become cluttered, and the information rendered by the XR system 800 can become overloaded and difficult to parse, manage, understand, etc. In some examples, the XR system 800 can filter data from the digital picture frame 810, the remote control 812, and the digital thermostat 814 to simplify and / or consolidate the data presented by the XR system 800 for the digital picture frame 810, the remote control 812, and the digital thermostat 814. In some cases, the XR system 800 can limit the data presented to data that corresponds to the particular device that is relevant.

[0109] To illustrate, and referring to FIG. 8C , the XR system 800 can predict that the user will want or will interact with the remote device 812. The XR system 800 can use this information to filter out data 820 associated with the digital picture frame 810 and data 824 associated with the digital thermostat 814 to simplify the data presented by the XR system 800. The XR system 800 can present data 822 associated with the remote device 812 that is predicted to be relevant to the user 802 and / or the current context. In some examples, the data 822 can include input options that the user can use to interact with the remote control 812 and / or an indication of the input options that inform the user how to interact with the remote control 812. In some cases, the data 822 can include user guidance data to facilitate user interaction with the remote device 812. In some examples, the presentation shown in FIG. 8C can declutter and / or simplify the data presented by the XR system 800.

[0110] In some cases, the XR system 800 can use the context information to predict which of the digital picture frame 810, the remote control 812, and the digital thermostat 814 that the user 802 will interact with are relevant to the user 802, etc., and / or what data from the digital picture frame 810, the remote control 812, and / or the digital thermostat 814 to present to the user. In some cases, the XR system 800 can present simplified information related to the digital picture frame 810, the remote control 812, and the digital thermostat 814, which the XR system 800 and the user 802 can use to filter out less relevant information and reduce the amount of information presented to the user 802 to the information most relevant to the user 802.

[0111] For example, referring to FIG. 8D , the XR system 800 can present a menu 840 of input options. The menu 840 can show various devices available for user interaction and provide the user 802 with the ability to select a particular device of interest to the user. If the user selects a particular device, the XR system 800 can present data, such as input options, that correspond to that particular device. For example, if the user 802 selects a digital frame 810 from the menu 840, the XR system 800 can present data relevant to the digital frame 810 and exclude other data relevant to the remote control 812 and / or thermostat 814.

[0112] In another illustrative example, the XR system 100 can assist a user in using an automobile control device. For example, a user driving a vehicle can pull over to the side of the road. The XR system 100 can detect that the user has pulled over to the side of the road (e.g., based on motion information, audio data from the user, data from the vehicle, etc.). Based on determining that the user has pulled over to the side of the road, the XR system 100 can determine that the user should activate the vehicle's hazard lights. The user can scan the vehicle's dashboard for a hazard light button. The XR system 100 can use eye tracking, image analysis, and / or other techniques to determine that the user is unable to locate or is having difficulty locating the hazard light button. The XR system 100 can use such contextual information to determine that the user needs help locating the hazard light button. The XR system 100 can locate / identify the hazard button and overlay AR content around the hazard button to guide / assist the user in locating the hazard button.

[0113] Although specific illustrative examples of XR system 100 (and other XR systems) using input data and context information to determine input options to present to a user are described above, XR system 100 may perform any other function based on the input data and context information to assist a user of XR system 100 in interacting with one or more other devices.

[0114] 9 is a flow chart illustrating an example of a process 900 for presenting information associated with at least one input option using one or more of the techniques described herein. At block 802, the process 900 may include receiving data identifying one or more input options associated with a device in a scene. For example, an XR system (e.g., XR system 100) may receive data identifying which input options are available at one or more remote devices in a scene.

[0115] At block 904, the process 900 may include determining, including using at least one memory, information relevant to at least one of a scene, a device, and a user (e.g., the XR system 100) associated with the electronic device. In some examples, the information may include context information. The context information may provide information about, for example, the scene, the user, the device, and / or the electronic device.

[0116] At block 906, the process 900 may include outputting user guidance data corresponding to the input option for which the relevant contextual information was determined based on the one or more input options and the information. In some examples, the user guidance data may include at least one of a user input element associated with the input option, a virtual overlay on a physical object associated with the input option, and / or a cue indicating how to provide an input associated with the input option.

[0117] In some aspects, process 900 may include predicting a user interaction with the device based on the information and presenting user guidance data corresponding to the input option based on the one or more input options and the predicted user interaction.

[0118] In some examples, the device may include a connected device having network communication capabilities, and the process 900 may include determining a hand gesture representing a predicted user interaction based on the information and the one or more input options and presenting user guidance data. In some examples, the predicted user interaction may include a predicted user input to the device. In some cases, the user guidance data may include an indication of a hand gesture that, when detected, invokes an actual user input at the device.

[0119] In some aspects, presenting the user guidance data can include rendering, on a display associated with the electronic device, a virtual overlay configured to appear to be placed on a surface of the device. In some examples, the virtual overlay can include user interface elements associated with the input options. In some cases, the user interface elements can include at least one of a virtual user input object associated with the input option and a visual indication of a physical control object on the device configured to receive an input corresponding to the input option.

[0120] In some examples, the information includes at least one of a user gaze and a user posture, and process 900 may include predicting a user interaction with the device based on at least one of the user gaze and the user posture, detecting an actual user input associated with the input option after presenting the user guidance data, the actual user input representing the predicted user interaction, and sending a command to the device corresponding to the actual user input associated with the input option.

[0121] In some examples, outputting user guidance data corresponding to the input options can include displaying the user guidance data. In some examples, outputting user guidance data corresponding to the input options can include outputting audio data representative of the user guidance data.

[0122] In some examples, outputting user guidance data corresponding to the input option may include displaying the user guidance data and outputting audio data associated with the displayed user guidance data.

[0123] In some aspects, process 900 may include receiving data from the device identifying one or more input options associated with the device. In some aspects, process 900 may include receiving data from a server identifying one or more input options associated with the device.

[0124] In some cases, the device does not have an external user interface for receiving one or more user inputs. In some aspects, process 900 may include refraining from presenting further user guidance data associated with the device based on the information.

[0125] In some aspects, the process 900 may include, after presenting the user guidance data, obtaining a user input associated with the input option and sending instructions to the device corresponding to the user input. In some cases, the instructions may be configured to control one or more operations of the device.

[0126] In some examples, the processes described herein (e.g., process 900 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 900 may be performed by the XR system 100 of FIG. 1. In another example, process 900 may be performed by a computing device having the computing system 1000 shown in FIG. 10. For example, a computing device having the computing architecture shown in FIG. 10 may include the components of the XR system 100 of FIG. 1 and may perform the operations of FIG. 9.

[0127] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device, having resource capabilities to perform the processes described herein, including process 800. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0128] Components of a computing device may be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0129] Process 900 is illustrated as a logical flow diagram, whose operations represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0130] Additionally, process 900 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0131] Fig. 10 illustrates an example of a system for implementing some aspects of the present technology. In particular, Fig. 10 illustrates an example of a computing system 1000, which may be, for example, an internal computing system, a remote computing system, a camera, or any computing device comprising any of the components thereof, in which the components of the system communicate with each other using a connection 1005. The connection 1005 may be a physical connection using a bus, or a direct connection to a processor 1010, such as in a chipset architecture. The connection 1005 may also be a virtual connection, a network connection, or a logical connection.

[0132] In some embodiments, computing system 1000 is a distributed system in which the functionality described in this disclosure may be distributed across a data center, multiple data centers, a peer network, etc. In some embodiments, one or more of the system components described represent many components, each performing some or all of the functionality that is the subject of the component description. In some embodiments, the components may be physical or virtual devices.

[0133] The exemplary system 1000 includes at least one processing unit (CPU or processor) 1010 and connections 1005 coupling various system components to the processor 1010, including system memory 1015, such as read only memory (ROM) 1020 and random access memory (RAM) 1025. The computing system 1000 may include a cache of high speed memory 1012, either directly connected to the processor 1010, in close proximity to the processor 1010, or integrated as part of the processor 1010.

[0134] The processor 1010 may include any general purpose processor, as well as hardware or software services, such as services 1032, 1034, and 1036, stored in a storage device 1030 and configured to control the processor 1010, as well as special purpose processors whose software instructions are embedded in the actual processor design. The processor 1010 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0135] To enable user interaction, computing system 1000 includes input device(s) 1045, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 1000 may also include output device(s) 1035, which can be one or more of several output mechanisms. In some cases, a multi-modal system may enable a user to provide multiple types of input / output to communicate with computing system 1000. Computing system 1000 may generally include a communication interface 1040, which can govern and manage user input and system output.The communications interface may be any of the following: audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, Apple® Lightning® port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, radio-frequency identification (RFID) wireless signal transmission, near-field communications (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WLAN), and Bluetooth® wireless signal transmission. The wireless communication device may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Access), infrared (IR) communications wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or any combination thereof.The communication interface 1040 may include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the position of the computing system 1000 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite system (BDS), and the European Galileo GNSS. There is no constraint to operate with any particular hardware arrangement, and therefore the basic features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.

[0136] The storage device 1030 may be a non-volatile and / or non-transitory and / or computer readable memory device, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disk (DVD) optical disk, a blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (ICC), a circuit, IC chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random-access memory (RRAM),The memory may be a hard disk or other type of computer readable medium capable of storing data that is accessible by a computer, such as a memory, such as RRAM / ReRAM, phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0137] The storage device 1030 may include software services, servers, services, etc., that cause the system to perform functions when code defining such software is executed by the processor 1010. In some embodiments, hardware services that perform a particular function may include software components stored in a computer-readable medium in association with necessary hardware components, such as the processor 1010, the connections 1005, the output devices 1035, etc., to perform the function. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media that can store, store, or convey instruction(s) and / or data. Computer-readable media may include non-transitory media on which data is stored and does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memories, or memory devices. A computer-readable medium may have code and / or machine-executable instructions stored on it, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0138] In some embodiments, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when mentioned, non-transitory computer readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0139] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For clarity of explanation, in some cases, the technology may be presented as including individual functional blocks comprising devices, device components, and steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0140] Particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowcharts may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when the operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0141] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions, or in some cases configure a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network-attached storage devices, etc.

[0142] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program product) to perform the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may be embodied in a peripheral device or an add-in card. Such functionality may be implemented on a circuit board among different chips, or on different processes executing in a single device, as further examples.

[0143] The instructions, media for carrying such instructions, computing resources for executing such instructions, and other structures for supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0144] In the above description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0145] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terminology used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the present specification.

[0146] When a component is described as being "configured to" perform a particular operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0147] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0148] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, a claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. As another example, a claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can additionally include unrecited items within the set of A and B.

[0149] The various exemplary logic blocks, modules, circuits, and algorithm steps described with respect to the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0150] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as separate but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium including program code including instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, which may carry or communicate program code in the form of instructions or data structures and which may be accessed, read, and / or executed by a computer.

[0151] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, or alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein.

[0152] Illustrative aspects of the present disclosure include the following. Aspect 1. An apparatus for outputting information associated with at least one input option, comprising at least one memory and at least one processor coupled to the at least one memory, the at least one processor configured to receive data identifying one or more input options associated with a device in a scene, determine information relevant to at least one of the scene, the device, and a user associated with the device, including using the at least one memory, and output user guidance data corresponding to an input option for which relevant contextual information has been determined based on the one or more input options and the information.

[0153] Example 2. The apparatus of example 1, wherein to receive data identifying one or more input options, at least one processor is configured to perform object recognition within the scene to identify one or more input options for operating the object.

[0154] Aspect 3. The apparatus of aspect 2, wherein the object is within at least one of a threshold proximity to the user and within a field of view (FOV) of the user.

[0155] Aspect 4. The apparatus of any of aspects 1 to 3, wherein at least one processor is configured to detect one or more further devices in the scene, determine a confidence value predicting whether a user will interact with the device or the one or more further devices based on the information, and predict user interaction with the device in response to the determined confidence value exceeding a threshold.

[0156] Aspect 5. The apparatus of aspect 4, wherein the at least one processor is configured to filter content associated with one or more additional devices in response to a prediction of user interaction with the device, and output content associated with the device.

[0157] Example 6. The apparatus of any of Examples 4 or 5, wherein the at least one processor is configured to identify the device and any of the one or more additional devices with which the user is expected to interact.

[0158] Aspect 7. The device of any of aspects 4-6, wherein the at least one processor is configured to simplify presentation of user interface content to avoid at least one of content overload, confusion, and content clutter.

[0159] Example 8. The apparatus of any of Examples 1-7, wherein the device comprises a connected device with network communication capabilities, and wherein at least one processor is configured to determine hand gestures representing predicted user interactions including predicted user inputs to the device based on the information and one or more input options, and, when detected, present user guidance data including instructions for hand gestures that invoke actual user inputs at the device.

[0160] Aspect 9. The device of any of aspects 1-8, wherein the user guidance data includes at least one of a user input element associated with the input option, a virtual overlay on a physical object associated with the input option, and a cue indicating how to provide input associated with the input option.

[0161] Aspect 10. The apparatus of aspect 7, wherein the at least one processor is configured to predict user interaction with the device based on the information and present user guidance data corresponding to the input option based on the one or more input options and the predicted user interaction.

[0162] Aspect 11. The device of any of aspects 1-10, wherein to present user guidance data, at least one processor is configured to render a virtual overlay on a display associated with the device, the virtual overlay being configured to appear to be placed on a surface of the device, the virtual overlay including user interface elements associated with the input options, the user interface elements including at least one of a virtual user input object associated with the input option and a visual indication of a physical control object on the device configured to receive input corresponding to the input option.

[0163] Aspect 12. The apparatus of any of aspects 1-11, wherein the information includes at least one of a user's gaze and a user's posture, and the at least one processor is configured to predict a user interaction with the device based on at least one of the user's gaze and the user's posture, and after presenting the user guidance data, detect an actual user input associated with the input option, the actual user input representing the predicted user interaction, and send a command to the device corresponding to the actual user input associated with the input option.

[0164] Aspect 13. The apparatus of any of aspects 1-12, wherein to output user guidance data corresponding to the input options, the at least one processor is configured to display the user guidance data.

[0165] Aspect 14. The apparatus of any of aspects 1-13, wherein to output user guidance data corresponding to the input option, the at least one processor is configured to output audio data representing the user guidance data.

[0166] Aspect 15. The device of any of aspects 1-14, wherein to output user guidance data corresponding to the input option, at least one processor is configured to display the user guidance data and output audio data associated with the displayed user guidance data.

[0167] Example 16. The apparatus of any of Examples 1-15, wherein the at least one processor is configured to receive data from the device identifying one or more input options associated with the device.

[0168] Example 17. The apparatus of any of Examples 1-16, wherein the at least one processor is configured to receive, from the server, data identifying one or more input options associated with the device.

[0169] Example 18. The apparatus of any of Examples 1-17, wherein the device does not have an external user interface for receiving one or more user inputs.

[0170] Example 19. The apparatus of any of Examples 1-18, wherein the at least one processor is configured to refrain from presenting further user guidance data associated with the device based on the information.

[0171] Embodiment 20. The device of any one of embodiments 1 to 19, wherein the device is an extended reality device.

[0172] Embodiment 21. The device of any of embodiments 1 to 20, further comprising a display.

[0173] Aspect 22. The device of aspect 21, wherein the display is configured to display at least the user guidance data.

[0174] Aspect 23. The apparatus of any of aspects 1 to 22, wherein the at least one processor is configured to, after presenting the user guidance data, obtain a user input associated with the input option, and transmit instructions to the device corresponding to the user input, the instructions being configured to control one or more operations of the device.

[0175] Aspect 24. The apparatus of any of aspects 1 to 23, wherein the information relevant to at least one of the scene, the device, and the user includes at least one of a predicted user interaction with the device, one or more actions of the user in the scene, characteristics associated with the user, historical information associated with the user and the device, user interface capabilities of the device, information associated with the device, and information associated with the scene.

[0176] Aspect 25. The apparatus of any of aspects 1 to 24, wherein at least one processor is configured to detect one or more further devices in a scene, determine a confidence value indicative of a likelihood that a user will interact with the device or the one or more further devices based on the context information, and predict a user interaction in response to the determined confidence value exceeding a threshold.

[0177] Example 26. The apparatus of any of Examples 1-25, wherein the at least one processor is further configured to receive confirmation from the user regarding the user interaction data, and to interact with the device in response to the confirmation from the user.

[0178] Aspect 27. The device of aspect 26, wherein the confirmation is an audio confirmation.

[0179] Aspect 28. The device of any of aspects 26 or 27, wherein the confirmation is a user input received at the device.

[0180] Aspect 29. A method for outputting information associated with at least one input option, comprising receiving data identifying one or more input options associated with a device in a scene, determining, including using at least one memory, information relevant to at least one of the scene, the device, and a user associated with the electronic device, and outputting user guidance data corresponding to an input option for which relevant contextual information has been determined based on the one or more input options and the information.

[0181] Example 30. The method of example 29, wherein receiving data identifying one or more input options includes performing object recognition within the scene to identify one or more input options for operating the object.

[0182] Aspect 31. The method of aspect 30, wherein the object is within at least one of a threshold proximity to the user and within the user's field of view (FOV).

[0183] Aspect 32. The method of any of aspects 29 to 31, further comprising: detecting one or more additional devices within the scene; determining a confidence value based on the information to predict whether a user will interact with the device or one or more additional devices; and predicting user interaction with the device in response to the determined confidence value exceeding a threshold.

[0184] Aspect 33. The method of any of aspects 29 to 32, further comprising filtering content associated with one or more additional devices in response to a prediction of user interaction with the device, and outputting content associated with the device.

[0185] Example 34. The method of any of Examples 29-33, further comprising identifying any of the device and one or more additional devices with which the user is expected to interact.

[0186] Aspect 35. The method of aspect 34, further comprising simplifying the presentation of user interface content to avoid at least one of content overload, confusion, and content clutter.

[0187] Aspect 36. The method of any of aspects 29 to 35, wherein the user guidance data includes at least one of a user input element associated with the input option, a virtual overlay on a physical object associated with the input option, and a cue indicating how to provide input associated with the input option.

[0188] Aspect 37. The method of any of aspects 29 to 36, further comprising: predicting a user interaction with the device based on the information; and presenting user guidance data corresponding to the input option based on the one or more input options and the predicted user interaction.

[0189] Aspect 38. The apparatus of aspect 37, wherein the device comprises a connected device with network communication capabilities, and the method further includes determining hand gestures representing predicted user interactions including predicted user input to the device based on the information and one or more input options, and presenting user guidance data including indications of hand gestures that invoke actual user input at the device that are detected.

[0190] Aspect 39. The method of any of aspects 29-38, wherein presenting the user guidance data includes rendering a virtual overlay on a display associated with the electronic device, the virtual overlay being configured to appear to be placed on a surface of the device, the virtual overlay including user interface elements associated with the input options, the user interface elements including at least one of a virtual user input object associated with the input option and a visual indication of a physical control object on the device configured to receive input corresponding to the input option.

[0191] Aspect 40. The method of any of aspects 29 to 39, wherein the information includes at least one of a user's gaze and a user's posture, and the method includes predicting a user interaction with the device based on at least one of the user's gaze and the user's posture, detecting an actual user input associated with the input option after presenting the user guidance data, the actual user input representing the predicted user interaction, and sending a command to the device corresponding to the actual user input associated with the input option.

[0192] Aspect 41. The method of any of aspects 29-40, wherein outputting user guidance data corresponding to the input option includes displaying the user guidance data.

[0193] Aspect 42. The method of any of aspects 29-41, wherein outputting user guidance data corresponding to the input option includes outputting audio data representing the user guidance data.

[0194] Aspect 43. The method of any of aspects 29 to 42, wherein outputting user guidance data corresponding to the input option includes displaying the user guidance data and outputting audio data associated with the displayed user guidance data.

[0195] Example 44. The method of any of Examples 29-43, further comprising receiving data from the device identifying one or more input options associated with the device.

[0196] Example 45. The method of any of Examples 29-44, further comprising receiving data from a server identifying one or more input options associated with the device.

[0197] Example 46. The method of any of Examples 29 to 45, wherein the device does not have an external user interface for receiving one or more user inputs.

[0198]

[0046] Example 47. The method of any of Examples 29-46, further comprising: refraining from presenting further user guidance data associated with the device based on the information.

[0199] Aspect 48. Any of the methods of aspects 29 to 47, further comprising: after presenting the user guidance data, obtaining a user input associated with the input option; and sending instructions to the device corresponding to the user input, the instructions being configured to control one or more operations of the device.

[0200] Aspect 49. The method of any of aspects 29 to 48, wherein the information relevant to at least one of the scene, the device, and the user includes at least one of a predicted user interaction with the device, one or more actions of the user in the scene, characteristics associated with the user, historical information associated with the user and the device, user interface capabilities of the device, information associated with the device, and information associated with the scene.

[0201] Aspect 50. The method of any of aspects 29 to 49, further comprising: detecting one or more additional devices within the scene; determining a confidence value indicative of a likelihood that a user will interact with the device or the one or more additional devices based on the context information; and predicting a user interaction in response to the determined confidence value exceeding a threshold.

[0202]

[0036] Embodiment 51. The method of any of embodiments 29-50, further comprising receiving a confirmation from the user regarding the user interaction data, and interacting with the device in response to the confirmation from the user.

[0203] Aspect 52. The method of aspect 51, wherein the confirmation is an audio confirmation.

[0204] Aspect 53. The method of any of aspects 51 or 52, wherein the confirmation is a user input received at the device.

[0205] Aspect 54. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of aspects 29-53.

[0206] Embodiment 55. An apparatus comprising means for carrying out the method of any one of embodiments 29 to 53.

Claims

1. 1. An extended reality head-mounted device for outputting information associated with at least one input option, comprising: at least one memory; at least one processor coupled to the at least one memory, the at least one processor comprising: receiving data identifying multiple input options associated with a device in a scene; determining, including using at least one memory, context information relevant to at least one of the scene, the device, and a user associated with the apparatus; filtering the plurality of input options associated with the device based on the context information to determine one or more input options associated with the context information; outputting user guidance data corresponding to the one or more input options associated with the context information; The extended reality head-mounted device is configured to:

2. 2. The extended reality head-mounted device of claim 1, wherein the user guidance data includes at least one of a user input element associated with the one or more input options, a virtual overlay on a physical object associated with the one or more input options, and a cue indicating how to provide input associated with the one or more input options.

3. the at least one processor: predicting user interactions with the device based on the information; presenting the user guidance data corresponding to the one or more input options based on the plurality of input options and the predicted user interaction.

10. The extended reality head-mounted device of claim 1, wherein the extended reality head-mounted device is configured to:

4. The device comprises a connected device with network communication capabilities, and the at least one processor: determining a hand gesture representing a predicted user interaction including a predicted user input to the device based on the information and the plurality of input options; presenting the user guidance data, which, when detected, includes an indication of the hand gesture that invokes actual user input at the device.

10. The extended reality head-mounted device of claim 1, wherein the extended reality head-mounted device is configured to:

5. To present the user guidance data, the at least one processor: configured to render, on a display associated with the apparatus, a virtual overlay configured to appear to be placed on a surface of the device, the virtual overlay including user interface elements associated with the one or more input options, the user interface elements including at least one of virtual user input objects associated with the one or more input options and visual indications of physical control objects on the device configured to receive input corresponding to the one or more input options; 10. The extended reality head-mounted device of claim 1.

6. The information includes at least one of a line of sight of the user and a posture of the user, and the at least one processor: predicting a user interaction with the device based on at least one of the gaze of the user and the posture of the user; detecting an actual user input associated with the one or more input options after presenting the user guidance data, the actual user input representing the predicted user interaction; sending a command to the device corresponding to the actual user input associated with the one or more input options; 10. The extended reality head-mounted device of claim 1, wherein the extended reality head-mounted device is configured to:

7. To output the user guidance data corresponding to the one or more input options, the at least one processor: displaying the user guidance data; 10. The extended reality head-mounted device of claim 1, wherein the extended reality head-mounted device is configured to:

8. the at least one processor: obtaining a user input associated with the one or more input options after presenting the user guidance data; transmitting instructions to the device corresponding to the user input, the instructions being configured to control one or more operations of the device; 10. The extended reality head-mounted device of claim 1, wherein the extended reality head-mounted device is configured to:

9. A method for outputting information associated with at least one input option in an extended reality head-mounted device, comprising: receiving data identifying a plurality of input options associated with a device in a scene; determining, including using at least one memory, context information relevant to at least one of the scene, the device, and a user associated with the electronic device; filtering the plurality of input options associated with the device based on the context information to determine one or more input options associated with the context information; outputting user guidance data corresponding to the one or more input options associated with the context information; A method comprising:

10. 10. The method of claim 9, wherein the user guidance data includes at least one of a user input element associated with the one or more input options, a virtual overlay on a physical object associated with the one or more input options, and a cue indicating how to provide input associated with the one or more input options.

11. predicting a user interaction with the device based on the information; and 10. The method of claim 9, further comprising: presenting the user guidance data corresponding to the one or more input options based on the plurality of input options and the predicted user interaction.

12. wherein the device comprises a connected device with network communication capabilities, and the method comprises: determining a hand gesture representing a predicted user interaction including a predicted user input to the device based on the information and the plurality of input options; 12. The method of claim 11, further comprising: presenting the user guidance data including an indication of the hand gesture that, when detected, invokes actual user input at the device.

13. presenting the user guidance data rendering, on a display associated with the electronic device, a virtual overlay configured to appear to be placed on a surface of the device, the virtual overlay including user interface elements associated with the one or more input options, the user interface elements including at least one of virtual user input objects associated with the one or more input options and visual indications of physical control objects on the device configured to receive input corresponding to the one or more input options; 10. The method of claim 9.

14. The information includes at least one of a line of sight of the user and a posture of the user, and the method further comprises: predicting a user interaction with the device based on at least one of the gaze of the user and the posture of the user; detecting an actual user input associated with the one or more input options after presenting the user guidance data, the actual user input representing the predicted user interaction; 10. The method of claim 9, further comprising: sending a command to the device corresponding to the actual user input associated with the one or more input options.

15. obtaining user input associated with the one or more input options after presenting the user guidance data; 10. The method of claim 9, further comprising: transmitting instructions to the device corresponding to the user input, the instructions configured to control one or more operations of the device.