Cooperative Tracking

JP2024529228A5Pending Publication Date: 2025-05-19QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023575401
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-18
Filing Date
2022-06-08
Publication Date
2025-05-19

AI Technical Summary

Technical Problem

XR devices face challenges in hand tracking due to computational expense and battery drain, and lose accuracy when hands leave or are occluded from the field of view, necessitating efficient resource management and external device integration for continuous tracking.

Method used

Systems and techniques that utilize data streams from multiple devices to merge image data with external sensors for continuous object tracking, including hand tracking, by detecting conditions such as resource availability and occlusion, and offloading processing to external devices to maintain accuracy and conserve battery.

Benefits of technology

Enhances hand tracking accuracy and extends battery life by leveraging external devices for computational tasks, ensuring seamless interaction with virtual environments even when hands are outside the XR device's field of view or occluded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The imaging system can receive an image of a portion of an environment. The environment can include an object, such as a hand or a display. The imaging device can identify the data stream from the external device, for example, by detecting the data stream in the image or by wirelessly receiving the data stream from the external device. The imaging device can detect a condition based on the image and / or the data stream, for example, by detecting that an object is missing from the image, by detecting low resources at the imaging device, and / or by detecting visual media content displayed by the display in the image. Upon detecting the condition, the imaging device automatically determines a location of the object (or a portion thereof) using the data stream and / or the image. The imaging device generates and / or outputs content based on the location of the object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001]

[0001] The present disclosure relates generally to image processing. For example, aspects of the present disclosure relate to systems and techniques for combining data from multiple devices to perform object tracking within an environment and provide output based on the tracking. [Background technology]

[0002]

[0002] An extended reality (XR) device is a device that displays an environment to a user, for example via a head-mounted display (HMD), glasses, a mobile handset, or other device. The environment is at least partially distinct from the real-world environment in which the user and device are located and may include, for example, virtual content. A user can generally interactively change the user's view of the environment, for example by tilting or moving the XR device. Virtual reality (VR), augmented reality (AR), and mixed reality (MR) are examples of XR.

[0003]

[0003] An XR device may include one or more image sensors, for example in one or more cameras. For example, a camera in an XR device may be used to capture image data of a real-world environment in the direction a user is looking and from the perspective of the user's position. An image sensor in an XR device may also be used to capture image data for tracking purposes (e.g., hand tracking, head tracking, body tracking, etc.).

[0004]

[0004] An XR device can display a representation of a user's hands in an environment that the XR device displays to the user, so that the user feels as if they are in that environment. Hand tracking can enable the XR device to accurately represent the user's hands in the environment and can enable the user to interact with real or virtual objects in the environment. However, hand tracking generally requires that the user keep the user's hands within the field of view (FOV) of the image sensor of the XR device. The XR device can suffer from errors if the user's hands move out of the FOV or are occluded. Hand tracking is generally a computationally expensive process that can quickly drain battery power. Summary of the Invention

[0005] In some embodiments, systems and techniques for feature tracking based on data from multiple devices are described. An imaging device, such as an XR device, can utilize one or more data streams from one or more external devices. For example, an image can be received from an image sensor of the imaging device. The image can be an image of a portion of an environment. The environment includes objects such as a user's hand or a display screen, but the objects may or may not be present in the portion of the environment shown in the image. The imaging device can identify the data stream from the external device, for example, based on the image (e.g., by identifying a data stream shown in the image, such as visual media content displayed on an external display device shown in the image), based on one or more transmissions of the data stream from the external device to the imaging device (e.g., over a wireless or wired network), based on user input, and / or based on other factors. The imaging device can detect a condition, such as based on the image, the data stream, an operating condition of the imaging device, any combination thereof, and / or based on other factors.In some examples, the state may be based on the imaging device losing sight of an object, the imaging device being low on computational resources (e.g., based on low power and / or other operating conditions of the device), the imaging device detecting visual media content (or a representation thereof) in an image, based on user input or settings that request the use of an external device rather than the imaging device (e.g., an XR device) when available for a particular function (e.g., displaying content, tracking an object such as a user's hand, head, or body), based on user input or settings that indicate a preference that a device (e.g., an external device) be used for a particular function when plugged into the imaging device, or based on privacy and / or security factors. The timing may be based on: (which may also be based on user input or settings); based on user input (e.g., user input requesting offloading resources to an external device, such as a user input requesting to turn off the imaging device, a user input requesting to turn on or off an external device, such as a light via a home automation application running on the imaging device); based on the capabilities of the imaging device's image sensor (e.g., when an infrared (IR) sensor on one device is useful when ambient lighting is insufficient, when the object being tracked is moving at high speed and an image sensor with a higher frame rate is more appropriate, etc.); or any combination thereof.

[0006]

[0006] In some cases, the imaging device may merge data from the data stream with an image captured by the image sensor, resulting in a merged data set. Based on detecting the condition, the imaging device may determine a position of at least a portion of an object in the environment based on the data stream, the image, and / or the merged data set. The imaging device may generate an output (e.g., content, commands for controlling the imaging device, commands for controlling an external device, etc.). The imaging device may output content based on the position of at least a portion of the object in the environment. In one embodiment, if the object is a user's hand, the content generated and / or output by the imaging device may accurately position a virtual object held by the user's hand based on the position of the user's hand (determined based on the data stream, the image, and / or the merged data set), even if the user's hand is not shown in the image. If the object is a display screen and / or visual content displayed on the display screen, the content generated and / or output by the imaging device may position the virtual content adjacent to the position of the display screen.

[0007]

[0007] In one embodiment, an apparatus for image processing is provided. The apparatus includes a memory and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured and capable of: receiving an image of a portion of the environment, the portion including an object, captured by an image sensor; identifying a data stream from an external device; detecting a condition based on at least one of the image, the data stream, and an operational condition of the apparatus; determining a position of the object in the environment based on at least one of the image and the data stream in response to detecting the condition; and generating an output based on the position of the object in the environment.

[0008]

[0008] In another embodiment, a method of image processing is provided that includes receiving, by a device, an image of a portion of the environment captured by an image sensor, the environment including an object, identifying a data stream from an external device, detecting a condition based on at least one of the image, the data stream, and an operational state of the device, determining a position of the object in the environment based on at least one of the image and the data stream in response to detecting the condition, and generating an output based on the position of the object in the environment.

[0009]

[0009] In another embodiment, a non-transitory computer readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to receive an image of a portion of the environment captured by an image sensor, the environment including an object, identify a data stream from an external device, detect a condition based on at least one of the image, the data stream, and an operational state of the device, determine a position of the object in the environment based on at least one of the image and the data stream in response to detecting the condition, and generate an output based on the position of the object in the environment.

[0010]

[0010] In another embodiment, an apparatus for image processing is provided, the apparatus including: means for receiving an image of a portion of the environment captured by an image sensor, the environment including an object, means for identifying a data stream from an external device, means for detecting a condition based on at least one of the image, the data stream, and an operational state of the apparatus, means for determining a position of the object in the environment based on at least one of the image and the data stream in response to detecting the condition, and means for generating an output based on the position of the object in the environment.

[0011]

[0011] In some aspects, to detect a condition based on an image, the above-mentioned method, apparatus, and computer-readable medium further include determining that an object is missing from a portion of the environment in the image.

[0012]

[0012] In some aspects, to determine that an object is missing from a portion of an environment in an image, the above-mentioned methods, apparatus, and computer-readable media include determining that at least a portion of the object is occluded in the image.

[0013] In some aspects, the external device includes a second image sensor. In some cases, the data stream includes a second image of a second portion of the environment. In such a case, determining a position of the object in the environment can be based at least in part on a depiction of the object in the second image. In some aspects, the portion of the environment and the second portion of the environment overlap.

[0014] In some aspects, for detecting a condition based on an operating state of the device, the above-mentioned methods, devices, and computer-readable media include determining that resource availability is below a threshold. In some aspects, for determining that resource availability is below a threshold, the above-mentioned methods, devices, and computer-readable media include determining that a battery level of a battery is below a battery level threshold.

[0015]

[0015] In some aspects, to determine that resource availability falls below a threshold, the methods, apparatus, and computer-readable media described above include determining that available bandwidth falls below a bandwidth threshold.

[0016]

[0016] In some aspects, to detect a state based on an operational state of the device, the above-mentioned methods, devices, and computer-readable media include receiving a user input corresponding to offloading processing to an external device.

[0017] In some aspects, to generate the output, the methods, devices, and computer-readable media described above include generating content. In some cases, the methods, devices, and computer-readable media described above include one or more processors configured to output the content based on a position of an object in an environment.

[0018] In some aspects, to output the content, the methods, apparatus, and computer-readable media described above include transmitting the content to be displayed (eg, to a display of an apparatus or device).

[0019]

[0019] In some aspects, the methods, apparatus, and computer-readable media described above include detecting an additional condition based on at least one of an additional image captured by the image sensor, a data stream, and an operational state of the apparatus, and performing a function previously performed by the external device in response to detecting the additional condition.

[0020] In some aspects, the methods, devices, and computer-readable media described above include controlling a device based on user input to generate an output.

[0021] In some aspects, to detect conditions based on an image, the methods, apparatus, and computer-readable media described above include determining one or more lighting conditions in the image.

[0022]

[0022] In some aspects, to determine one or more lighting conditions within an image, the above-mentioned methods, devices, and computer-readable media include determining that one or more light values ​​of the image fall below an illumination threshold.

[0023]

[0023] In some aspects, to determine the position of an object within an environment, the above-mentioned methods, apparatus, and computer-readable media include sending a request to an external device identifying the position of the object within the environment, and receiving a response from the external device identifying the position of the object within the environment.

[0024] In some aspects, the object is a display of an external display device.

[0025]

[0025] In some aspects, to detect a condition based on an image, the above-mentioned methods, apparatus, and computer-readable media include identifying within the image visual media content displayed on a display of an external display device.

[0026] In some aspects, to generate the output, the methods, apparatus, and computer-readable media described above include generating content, in some cases, the content virtually extending the display of the external display device.

[0027] In some aspects, to generate the output, the methods, apparatus, and computer-readable media described above include generating the content at least in part by overlaying virtual content over a region of the image, in some cases the region of the image is based on a position of an object in the environment.

[0028] In some aspects, the object is a display of an external display device. In some cases, a region of the image is adjacent to a representation of the display of the external display device in the image.

[0029] In some aspects, the object is a hand of a user of the device. In some cases, the hand is at least partially adjacent to a region of the image.

[0030] In some aspects, the methods, apparatus, and computer-readable media described above include, in response to detecting the condition, generating a merged dataset by combining at least data from the data stream with an image captured by the image sensor. In some cases, determining a location of the object is based at least in part on the merged dataset.

[0031] In some aspects, to generate an output, the methods, apparatus, and computer-readable media described above include generating the content. In some cases, the output, methods, apparatus, and computer-readable media described above include transmitting or sending the content to an audio output device (e.g., of an apparatus or device) to be played.

[0032] In some aspects, each of the above-mentioned apparatus or devices may be, be part of, or include an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a smart device or assistant, a vehicle, a mobile device (e.g., a mobile phone, or a so-called "smartphone", or other mobile device), a wearable device, a personal computer, a laptop computer, a tablet computer, a server computer, or other device. In some aspects, the apparatus or device includes an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, the apparatus or device includes one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus or device includes one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, the above-mentioned apparatus or device may include one or more sensors. In some cases, one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level, and / or other state), and / or for other purposes.

[0033]

[0033] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0034]

[0034] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings. [Brief description of the drawings]

[0035]

[0035] Exemplary embodiments of the present application are described in detail below with reference to the following drawings. [Figure 1]

[0036] FIG. 1 is a block diagram illustrating an example architecture of an image capture and processing system, according to some embodiments. [Diagram 2]

[0037] FIG. 1 is a block diagram illustrating an example architecture of an extended reality (XR) system, according to some embodiments. [Figure 3A]

[0038] FIG. 1 is a perspective view of a head mounted display (HMD) used as an XR system, according to some embodiments. [Figure 3B]

[0039] FIG. 3B is a perspective view illustrating the head mounted display (HMD) of FIG. 3A being worn by a user, according to some embodiments. [Figure 4A]

[0040] FIG. 1 is a perspective view showing the front side of a mobile handset that includes a front-facing camera and is used as an XR system, according to some embodiments. [Figure 4B]

[0041] FIG. 2 is a perspective view showing the back of a mobile handset that includes a rear-facing camera and is used as an XR system, according to some embodiments. [Diagram 5]

[0042] FIG. 1 is a perspective view of a user wearing a head mounted display (HMD) used as an XR system and performing hand tracking to determine gesture-based inputs based on hands that are in the field of view (FOV) of the HMD, according to some embodiments. [Figure 6A]

[0043] FIG. 1 is a perspective view of a user wearing a head mounted display (HMD) used as an XR system and performing hand tracking to determine gesture-based inputs based on the position of the user's hands based on the hands being in the field of view (FOV) of an external camera, even if the hands are outside the HMD's FOV, according to some embodiments. [Figure 6B]

[0044] FIG. 1 is a perspective view of a user wearing a head mounted display (HMD) used as an XR system and performing hand tracking to determine gesture-based inputs based on the position of the user's hands based on the hands being in the field of view (FOV) of an external camera, even when occlusions block the hands from the HMD's FOV, according to some embodiments. [Figure 7]

[0045] FIG. 1 is a perspective view illustrating an external head mounted display (HMD) device that provides assistance by hand tracking of the hands of a user of an HMD used as an XR system due to a low battery condition in the HMD, according to some embodiments. [Figure 8A]

[0046] FIG. 1 is a perspective view showing a user wearing a head mounted display (HMD) used as an XR system to position virtual content based on the position of the display within the FOV of the HMD and / or visual content displayed on the display. [Figure 8B]

[0047] FIG. 1 is a perspective view illustrating a user wearing a head mounted display (HMD) used as an XR system to position a virtual representation of visual content displayed on a display based on the position of the display and / or visual content even if the display and / or visual content is outside the field of view (FOV) of the HMD, according to some embodiments. [Figure 9]

[0048] FIG. 4 is a flow diagram illustrating operations for processing image data, according to some embodiments. [Figure 10]

[0049] FIG. 1 illustrates an example of a computing system for implementing some aspects described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0036]

[0050] Some aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments can be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0037]

[0051] The following description merely provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the scope of the present application as set forth in the appended claims.

[0038]

[0052] A camera is a device that uses an image sensor to capture image frames, such as still images or video frames, upon receiving light. The terms "image," "image frame," and "frame" are used interchangeably herein. A camera can be configured with various image capture and image processing settings. Different settings result in images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, f-stop, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to an image sensor for capturing one or more image frames. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color changes, can constitute post-processing of one or more image frames. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) for processing one or more image frames captured by the image sensor.

[0039]

[0053] An extended reality (XR) device is a device that displays an environment to a user, e.g., via a head mounted display (HMD), glasses, a mobile handset, or other device. The displayed environment is at least partially different from the real-world environment in which the user and device are located and may, e.g., include virtual content. In some cases, the environment that an XR device displays to a user can be at least partially virtual. A user can generally interactively change the user's view of the environment that the XR device displays to the user, e.g., by tilting the XR device and / or moving the XR device translationally or laterally. Tilting the XR device can include tilting or rotating along a pitch axis, a yaw axis, a roll axis, or a combination thereof. Translational / lateral movement of the XR device can include movement along a path drawn in a three-dimensional volume having three perpendicular axes, such as the X-axis, the Y-axis, and the Z-axis. An XR device that tracks only rotational movement of the XR device can be referred to as an XR device with three degrees of freedom (3DoF). An XR device that tracks both rotational and translational movements of the XR device may be referred to as an XR device with six degrees of freedom (6DoF) tracking capability.

[0040]

[0054] An XR device may include sensors such as image sensors, accelerometers, gyroscopes, inertial measurement units (IMUs), or combinations thereof. An XR device may use data captured by these sensors to detect movement of the XR device within a real-world environment, such that, for example, the XR device may interactively update a user's view of the environment based on rotational and / or translational movement of the XR device. Some XR devices may also use data captured by these sensors to detect and / or track features of one or more objects, such as a user's hands. Even XR devices that otherwise display a completely virtual VR environment to a user may still display a representation of the user's own hands within the environment. Displaying a representation of the user's hands within the environment may increase a user of the XR device's immersion in the environment, helping the user feel that they are truly within the environment. Displaying a representation of the user's hands within the environment may also enable the user's hands to interact with virtual objects and / or interfaces (e.g., menus) within the environment displayed by the XR device.

[0041]

[0055] An XR device may perform object tracking, which may be useful to enable a user to use the user's hands to interact with virtual objects and / or interfaces displayed by the XR device. For example, the XR device may track one or more hands of a user of the XR device to determine the pose (e.g., position and orientation) of one or more hands. Hand tracking may be useful to ensure that the pose of a representation of the user's hands used by the XR device (e.g., to determine gesture-based inputs, to display one or more hand representations, etc.) is accurately synchronized with the real-world position of the user's hands. Other types of tracking may also be performed, including head tracking, body tracking, torso tracking, tracking of controllers used to interact with the XR device, and / or tracking of other objects. In one embodiment, hand tracking may be useful to enable the XR device to accurately render occlusion of the environment by the user's hands, occlusion of hands by one or more real objects in the environment or virtual objects displayed by the XR device, occlusion of any real or virtual objects by the hand(s) based on the user holding the real or virtual object in the user's hand, etc. In some cases, hand tracking can stop functioning properly if the user's hand moves out of the field of view of the XR device's sensors, as shown, for example, in Figure 6A, described below. In other cases, hand tracking can stop functioning properly if the user's hand is blocked from the field of view of the XR device's sensors, as shown, for example, in Figure 6B.

[0042]

[0056] Object tracking (e.g., hand tracking, head tracking, body tracking, etc.) is a computationally expensive process that can quickly drain the battery of an XR device. Therefore, it may be useful to offload some hand tracking tasks based on the operating state of the XR device, such as when the XR device is low on battery power or other computational resources (e.g., as shown in FIG. 7). In some XR systems, it may also be useful to track other types of objects. For example, in some XR systems, it may be useful to track the display screen, e.g., as shown in FIGS. 8A-8B.

[0043]

[0057] Described herein are techniques for an imaging device (e.g., an XR device) to utilize one or more data streams from one or more external devices. For example, an image can be received from an image sensor of the imaging device. The image may be an image of a portion of an environment that includes an object. The object may or may not be present in the portion of the environment shown in the image. The object may be, for example, a hand of a user of the imaging device, a head of a user, a body of a user, another body part of a user of the imaging device, a display screen, image media content displayed on a display screen, video media content displayed on a display screen, a person, an animal, a vehicle, a plant, another XR device (in addition to the imaging device, which may be an XR device), another object, or a combination thereof.

[0044]

[0058] The imaging device can identify the data stream from the external device. For example, the imaging device can identify the data stream from the external device based on an image received from an image sensor (e.g., by identifying a data stream shown in the image, such as media content displayed on an external display device shown in the image), based on one or more transmissions of the data stream from the external device to the imaging device (e.g., over a wireless network or a wired connection), based on user input, and / or based on other factors. The imaging device can detect the condition, such as based on the image, the data stream, an operating state of the imaging device, any combination thereof, and / or based on other factors.In some examples, the state may be based on the imaging device losing track of the object (e.g., because the tracked object has moved outside the FOV of the imaging device, is obstructed from the field of view of the imaging device by a real-world or virtual object, etc.), the imaging device being low on computational resources (e.g., based on low power and / or other operating conditions of the device), the imaging device detecting visual media content (or a representation thereof) in an image, based on user input or settings that request the use of an external device rather than the imaging device (e.g., an XR device) when available for a particular function (e.g., displaying content, tracking an object such as a user's hands, head, or body), or a preference that a device (e.g., an external device) be used for a particular function when plugged into the imaging device. the imaging device is based on a user input or setting indicating that the imaging device is not being tracked, that privacy and / or security are factors (which may also be based on user input or setting), based on user input (e.g., user input requesting offloading resources to an external device, such as a user input requesting that the imaging device be turned off, user input requesting that an external device, such as a light, be turned on or off via a home automation application running on the imaging device, etc.), based on the capabilities of the imaging device's image sensor (e.g., when an infrared (IR) sensor on one device is useful when ambient lighting is insufficient, when the object being tracked is moving at a high speed and an image sensor with a higher frame rate is more appropriate, etc.), or any combination thereof.

[0045]

[0059] In response to detecting the condition, the imaging device can generate an output. For example, based on detecting the condition, the imaging device can generate a merged dataset by merging or combining data from the data stream with an image captured by the image sensor. In some cases, in response to detecting the condition, the imaging device can determine a position of at least a portion of an object in the environment based on the data stream, the image, the merged dataset, or any combination thereof. The imaging device can generate and output content based on the position of at least a portion of the object in the environment. For example, if the object is a user's hand, the content generated and / or output by the imaging device can accurately position a virtual object held by the user's hand based on the position of the user's hand even if the user's hand is not shown in the image. If the object is a display screen and / or visual content displayed on the display screen, the content generated and / or output by the imaging device can position the virtual content adjacent to or in some other predetermined relative position to the position of the display screen and / or visual content displayed on the display screen. The content output by the imaging device can include at least a portion of the merged dataset. The imaging device and the external device can perform privacy negotiation. For example, the external device can identify to the imaging device what data streams from the external device the imaging device can and cannot use, and vice versa.

[0046]

[0060] In a first exemplary embodiment, the external device includes an external camera, and the data stream from the external device includes a camera feed (e.g., one or more images) from the external camera. The external camera may be a camera from another imaging device (e.g., another XR device) or a camera from another camera. The external camera may be in the same environment as the imaging device and / or have the same environment in its FOV as the imaging device has in its FOV. The condition may include, for example, that the imaging device has lost sight of the user's hand(s) and is unable to properly perform hand tracking. For example, the user may have moved the user's hand(s) outside the field of view of the imaging device (e.g., as in FIG. 6A ) and / or an occlusion may have blocked the user's hand(s) from the perspective of the imaging device's camera(s) (e.g., as in FIG. 6B ). However, the user's hand(s) may be shown in the camera feed from the external camera. The imaging device can use the camera feed from the external camera to assist in identifying where the user's hands are relative to content shown in the image captured by the image sensor of the imaging device. In some cases, the external device can include a processor that can perform preliminary processing, for example, by performing hand detection and / or hand tracking using images from the camera feed from the external camera. The external device can transmit the image(s) from the camera feed and / or data corresponding to the preliminary processing to the imaging device. Content generated and / or output by the imaging device can include modifications to the image based on hand tracking, such as incorporating virtual content into the image. The virtual content can be placed on (or relative to) the display of the imaging device based on the position(s) of the user's hand(s).

[0047]

[0061] In a second exemplary embodiment, the external device includes an external camera, and the data stream from the external device includes a camera feed (e.g., one or more images) from the external camera. The external camera may be a camera from another imaging device (e.g., another XR device) or a camera from another camera. The external camera may be in the same environment as the imaging device and / or have the same environment in its FOV as the imaging device has in its FOV. In such an embodiment, the state may be based on an operating state of the XR device. For example, the state may be based on detecting that the imaging device is low on battery power, data bandwidth, processing bandwidth, another computational resource, or a combination thereof. The imaging device may use the camera feed from the external camera to help perform hand tracking or other function(s), or a combination thereof, that may be battery-intensive, bandwidth-intensive, processing-intensive, or otherwise use a large amount of computational resources. As in the first exemplary embodiment, the external device in the second exemplary embodiment may perform preliminary processing (e.g., by performing hand detection and / or tracking on images from the camera feed from the external camera). The external device can send the (pre-processed) image(s) from the camera feed and / or data corresponding to the pre-processing to the imaging device. The content generated and / or output by the imaging device can include modifications to the image based on hand tracking, such as incorporating virtual content into the image based on the hand position(s).

[0048]

[0062] In a third illustrative example, the external device includes a display screen. The external device, in this example, may be a television, a laptop computer, a desktop computer monitor, a smart home device or assistant, a video game console monitor, a mobile handset having a display screen, a wearable device having a display screen, a television display screen, another device having a display screen, the display screen itself, or a combination thereof. The data stream from the external device may include visual media content displayed on the display screen. The image captured by the imaging device may include a representation of the display screen of the external device and therefore may include a representation of the visual media content displayed on the display screen of the external device. The state may include detection of a representation of the display screen and / or a representation of the visual media content displayed on the display screen in an image captured by an image sensor of the imaging device. For example, a user of the imaging device may view the external device displaying visual media content on its display screen through the user's imaging device. For example, the visual media content may be a television program, a movie, a video game, a slideshow, another type of image, another type of video, or some combination thereof. Merging the data from the data stream (the visual media content) with the image may include adding information to the representation of the visual media content in the image. The added information may include, for example, information about actors in a television program or movie scene, information about deleted scenes, information about video game statistics such as health, and / or other information. To a user of the imaging device, the added information may appear adjacent to, overlaid on, or otherwise positioned relative to the representation of the visual media content.

[0049]

[0063] 1 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components used to capture and process an image of a scene (e.g., an image of a scene 110). The image capture and processing system 100 can capture a standalone image (or photo) and / or can capture a video that includes multiple images (or video frames) in a particular sequence. A lens 115 of the system 100 faces a scene 110, such as a portion of a real-world environment, and receives light from the scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0050]

[0064] The one or more controls 120 can control exposure, focus, and / or zoom based on information from image sensor 130 and / or based on information from image processor 150. The one or more controls 120 may include multiple mechanisms and components, for example, the control 120 may include one or more exposure controls 125A, one or more focus controls 125B, and / or one or more zoom controls 125C. The one or more controls 120 may also include additional controls beyond those shown, such as controls to control analog gain, flash, high dynamic range (HDR), depth of field, and / or other image capture characteristics.

[0051]

[0065] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some embodiments, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or servo, thereby adjusting the focus. In some cases, additional lenses, such as one or more microlenses over each photodiode of the image sensor 130, may be included in the system 100, each of which bends light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may be referred to as image capture settings and / or image processing settings.

[0052]

[0066] An exposure control 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control 125A can control the size of the aperture (e.g., aperture size or f-stop), the duration the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0053]

[0067] The zoom control 125C of the control mechanism 120 can obtain the zoom setting. In some embodiments, the zoom control 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some embodiments, the lens assembly can include a parfocal zoom lens or a variable focus zoom lens. In some embodiments, the lens assembly can include a focusing lens (which may be the lens 115 in some cases) that first receives light from the scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before the light reaches the image sensor 130. In some cases, an afocal zoom system can include two positive (e.g., converging, convex) lenses of equal or similar focal lengths (e.g., within a threshold difference) with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control 125C moves one or more of the lenses in the afocal zoom system, such as one or both of the negative and positive lenses.

[0054]

[0068] Image sensor 130 includes one or more arrays of photodiodes or other light receiving elements. Each photodiode measures an amount of light that ultimately corresponds to a particular pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters, thus measuring light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, and each pixel of the image is generated based on red light data from at least one photodiode covered by a red color filter, blue light data from at least one photodiode covered by a blue color filter, and green light data from at least one photodiode covered by a green color filter. Other types of color filters can use yellow, magenta, and / or cyan (also called "emerald") color filters instead of or in addition to the red, blue, and / or green color filters. Some image sensors may be completely devoid of color filters and instead use different photodiodes (possibly stacked vertically) throughout the pixel array. Different photodiodes across the pixel array can have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore no color depth.

[0055]

[0069] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching some photodiodes or portions of some photodiodes at some times and / or from some angles, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying an analog signal output by the photodiode, and / or an analog to digital converter (ADC) for converting the analog signal output of the photodiode (and / or amplified by the analog gain amplifier) ​​to a digital signal. In some cases, instead or in addition, some components or functions described with respect to one or more of control mechanisms 120 may be included in image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active-pixel sensor (APS), a complimentary metal-oxide semiconductor (CMOS), an n-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0056]

[0070] Image processor 150 may include one or more processors, such as one or more image signal processors (ISP) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other types of processors 5010 described with respect to computing device 5000. Host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes host processor 152 and ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) ports 156), central processing units (CPUs), graphics processing units (GPUs), broadband modems (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth, Global Positioning System (GPS), etc.), any combination thereof, and / or other components.The I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (e.g., MIPI CSI-2), a physical (PHY) layer port or interface, an Advanced High-performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one exemplary embodiment, the host processor 152 may communicate with the image sensor 130 using an I2C port and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0057]

[0071] Image processor 150 may perform several tasks such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, object detection, object tracking, descriptor generation, receiving input, managing output, managing memory, or any combination thereof. Image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 / 5020, read-only memory (ROM) 145 / 5025, a cache, a memory unit, another storage device, or any combination thereof.

[0058]

[0072] Various input / output (I / O) devices 160 may be connected to image processor 150. I / O device 160 may include a display screen, a keyboard, a keypad, a touch screen, a track pad, a touch-sensitive surface, a printer, any other output device 5035, any other input device 5045, or some combination thereof. In some cases, captions may be entered into image processing device 105B via a physical keyboard or keypad of I / O device 160 or via a virtual keyboard or keypad of a touch screen of I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable a wireless connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. The peripheral devices may include any of the types of I / O devices 160 described above, and may themselves be considered I / O devices 160 when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0059]

[0073] In some cases, image capture and processing system 100 may be a single device. In some cases, image capture and processing system 100 may be two or more separate devices including image capture device 105A (e.g., a camera) and image processing device 105B (e.g., a computing device coupled to a camera). In some implementations, image capture device 105A and image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some implementations, image capture device 105A and image processing device 105B may be separate from one another.

[0060]

[0074] 1, a vertical dashed line divides image capture and processing system 100 of FIG. 1 into two portions representing image capture device 105A and image processing device 105B, respectively. Image capture device 105A includes lens 115, control mechanism 120, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, some components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included in image capture device 105A.

[0061]

[0075] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smartphone, a mobile phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a smart home device or assistant, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some embodiments, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or any combination thereof. In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or another computing device.

[0062]

[0076] Although the image capture and processing system 100 is shown as including several components, one skilled in the art will appreciate that the image capture and processing system 100 may include many more components than those shown in FIG. 1. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware that may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform various operations described herein. The software and / or firmware may include one or more instructions stored in a computer-readable storage medium and executable by one or more processors of an electronic device that implements the image capture and processing system 100.

[0063]

[0077] Described herein are systems, apparatus, processes, and computer readable media for identifying and tracking the position of an object in one or more images. Each of the images can be captured using an image sensor 130 of an image capture device 150A, an image capture and processing system 100, or a combination thereof. Each of the images can be processed using an image processing device 105B, an image capture and processing system 100, or a combination thereof. The image capture and processing system 100 can be part of an XR system or an XR device, such as the XR system 210 of FIG. 2. The image capture and processing system 100 can be a sensor of the XR system or an XR device, such as the sensor 215 of the XR system 210 of FIG. 2. The image capture and processing system 100 can be part of an external device, such as the external device 220 of FIG. 2. The image capture and processing system 100 can be a sensor of an external device, such as the sensor 225 of the external device 220 of FIG. 2.

[0064]

[0078] 2 is a block diagram 200 illustrating an example architecture of an extended reality (XR) system 210. The XR system 210 of FIG. 2 includes one or more sensors 215, a processing engine 205, an output content generation engine 280, and an output device 290.

[0065]

[0079] The processing engine 205 of the XR system 210 can receive sensor data from one or more sensors 215 of the XR system 210. The one or more sensors 215 of the XR system 210 can include, for example, one or more image sensors 130, one or more accelerometers, one or more gyroscopes, one or more inertial measurement units (IMUs), one or more light detection and ranging (LIDAR) sensors, one or more radio detection and ranging (RADAR) sensors, one or more sound detection and ranging (SODAR) sensors, one or more sound navigation and ranging (SONAR) sensors, one or more time-of-flight (ToF) sensors, one or more structured light sensors, one or more microphones, one or more other sensors described herein, or combinations thereof. In some embodiments, the one or more sensors 215 can be coupled to the processing engine 205 via one or more wired and / or wireless sensor connectors. In some embodiments, the sensor data can include one or more images. The one or more images may include a still image, a video frame of one or more videos, or a combination thereof. The one or more images may be referred to as a still image, an image frame, a video frame, a frame, or a combination thereof. A dashed box is shown around one or more sensors 215 of the XR system 210 to indicate that the one or more sensors 215 may be considered part of the XR system 210 and / or the processing engine 205.

[0066]

[0080] The processing engine 205 of the XR system 210 can receive sensor data from one or more sensors 225 of the external device 220. The one or more sensors 225 of the external device 220 can include, for example, one or more image sensors 130, one or more accelerometers, one or more gyroscopes, one or more IMUs, one or more LIDAR sensors, one or more RADAR sensors, one or more SODAR sensors, one or more SONAR sensors, one or more ToF sensors, one or more structured light sensors, one or more microphones, one or more other sensors described herein, or combinations thereof. In some examples, the external device 220 and / or the one or more sensors 225 can be coupled to the processing engine 205 via one or more wired and / or wireless connections. The one or more images may be referred to as still images, image frames, video frames, frames, or combinations thereof.

[0067]

[0081] The processing engine 205 of the XR system 210 includes an inter-device negotiation engine 230 that can negotiate with the external device 220. The inter-device negotiation engine 230 can include a communication transceiver 235. The communication transceiver 235 can include one or more wired communication transceivers, one or more wireless communication transceivers, or a combination thereof. The inter-device negotiation engine 230 of the XR system 210 can receive sensor data from the sensor 225 of the external device 220 using the communication transceiver 235. The inter-device negotiation engine 230 of the XR system 210 can also use the communication transceiver 235 to send negotiation data to the external device 220 and / or receive negotiation data from the external device 220 as part of one or more negotiations, such as a synchronization negotiation, a security negotiation, a privacy negotiation, or a combination thereof.

[0068]

[0082] The inter-device negotiation engine 230 of the XR system 210 may include a synchronization negotiation engine 240 that synchronizes sensor data received from one or more sensors 225 of the external device 220 with sensor data received from one or more sensors 215 of the XR system 210. For example, the sensor data received from one or more sensors 225 of the external device 220 may be tagged with a timestamp at which individual elements of the sensor data (e.g., individual images) were captured by one or more sensors 225 of the external device 220. Similarly, the sensor data received from one or more sensors 215 of the XR system 210 may be tagged with a timestamp at which individual elements of the sensor data (e.g., individual images) were captured by one or more sensors 215 of the XR system 210. The synchronization negotiation engine 240 may match elements of the sensor data from one or more sensors 225 of the external device 220 with corresponding elements of the sensor data from one or more sensors 215 of the XR system 210 based on the corresponding timestamps matching as closely as possible. In an example embodiment, one or more sensors 215 of the XR system 210 may capture an image having a timestamp of 4:30.3247, and one or more sensors 225 of the external device 220 may capture images having timestamps of 4:29.7930, 4:30.0139, 4:30.3923, and 4:30.8394. The synchronization negotiation engine 240 may identify that the timestamp of 4:30.3923 from the sensor data of the one or more sensors 225 of the external device 220 most closely matches the timestamp of 4:30.3247 from the sensor data of the one or more sensors 215 of the XR system 210. Thus, the synchronization negotiation engine 240 can synchronize an image corresponding to a timestamp of 4:30.3923 from the sensor data of one or more sensors 225 with an image corresponding to a timestamp of 4:30.3247 from the sensor data of one or more sensors 215 of the XR system 210.In some embodiments, the synchronization negotiation engine 240 can send a request to the external device 220 for sensor data that most closely matches the timestamp of the sensor data from one or more sensors 215 of the XR system 210. The synchronization performed by the synchronization negotiation engine 240 can be based on sensor capabilities. For example, if the sensor 215 of the XR system 210 captures images at 90 frames per second (fps) and the sensor 225 of the external device 220 captures images at 30 fps, the synchronization negotiation engine 240 can synchronize every third image captured by the sensor 215 of the XR system 210 with the image captured by the sensor 225 of the external device 220.

[0069]

[0083] The inter-device negotiation engine 230 of the XR system 210 may include a security negotiation engine 245. The security negotiation engine 245 may perform a security handshake between the XR system 210 and the external device 220. The security handshake may include, for example, a transport layer security (TLS) handshake, a secure sockets layer (SSL) handshake, or a combination thereof. The security handshake may identify the version of the encryption protocol used between the XR system 210 and the external device 220, determine the cipher suite used between the XR system 210 and the external device 220, and authenticate the identity of the XR system 210 and / or the external device 220 using one or more digital signatures (and / or one or more certificate authorities). The security handshake may generate a session key to use symmetric encryption after the handshake is completed. The security handshake may generate or retrieve an asymmetric key pair for each of the XR system 210 and the external device 220, and may transfer the public keys from each key pair from the device where they are generated or retrieved to the other device. The XR system 210 and the external device 220 may then communicate via encrypted communications using asymmetric and / or symmetric encryption following the security handshake.

[0070]

[0084] The inter-device negotiation engine 230 of the XR system 210 can include a privacy negotiation engine 247. The privacy negotiation engine 247 can request sensor data from the sensor 225 of the external device 220 for use for an identified purpose, for example, for hand tracking as in FIG. 6A, FIG. 6B, or FIG. 7. The external device 220 can allow or deny the XR system 210 access to the sensor data from the sensor 225 of the external device 220 for the identified purpose. In some examples, the external device 220 can include a whitelist of purposes for which the external device 220 can allow sharing of the sensor data from the sensor 225 of the external device 220. In some examples, the external device 220 can include a blacklist of purposes for which the external device 220 cannot allow (but should instead deny) sharing of the sensor data from the sensor 225 of the external device 220. In some examples, the privacy negotiation engine 247 may request sensor data from the sensor 225 of the external device 220 for use for multiple purposes, but the external device 220 may respond to indicate that the external device 220 will allow sharing of sensor data from the sensor 225 of the external device 220 for only a subset of the multiple purposes. The privacy negotiation engine 247 may respect any limitations the external device 220 identifies on the purposes for which the sensor data from the sensor 225 of the external device 220 may be used.

[0071]

[0085] In some embodiments, the external device 220 can make certain requests or demands of the XR system 210 when sensor data is sent from the sensor 225 of the external device 220 to the XR system 210, and the privacy negotiation engine 247 can agree to it and perform a corresponding action. For example, in some embodiments, the external device 220 can request that the XR system 210 delete the sensor data from the sensor 225 of the external device 220 immediately after use or a predetermined period of time after use. The privacy negotiation engine 247 can agree to this requirement and ensure that the XR system 210 deletes the sensor data from the sensor 225 of the external device 220 immediately after use or a predetermined period of time after use. In some embodiments, the external device 220 can request that the XR system 210 not use, discard, or replace certain aspects of the sensor data from the sensor 225 of the external device 220. For example, the external device 220 may request that the XR system 210 not use or anonymize names, faces, or other sensitive information in the sensor data from the sensors 225 of the external device 220. The privacy negotiation engine 247 may agree to this requirement and ensure that the XR system 210 does not use, discards, or replaces certain portions of aspects of the sensor data from the sensors 225 of the external device 220.

[0072]

[0086] The processing engine 205 of the XR system 210 includes a feature management engine 250. The feature management engine 250 receives sensor data from one or more sensors 215 of the XR system 210. The feature management engine 250 receives sensor data from one or more sensors 225 of the external device 220. The inter-device negotiation engine 230 can synchronize the sensor data from the one or more sensors 215 of the XR system 210 with the sensor data from the one or more sensors 225 of the external device 220 prior to or simultaneously with the receipt of the sensor data by the feature management engine 250. The inter-device negotiation engine 230 can identify any security and / or privacy restrictions, constraints, and / or requirements prior to or simultaneously with the receipt of the sensor data by the feature management engine 250.

[0073]

[0087] The feature management engine 250 includes a feature extraction engine 255. The feature extraction engine 255 can detect and / or extract features from sensor data from one or more sensors 215 of the XR system 210. In some cases, the feature extraction engine 255 can detect and / or extract features from sensor data from one or more sensors 225 of the external device 220. For example, if the sensor data includes an image, the feature extraction engine 255 can detect and / or extract visual features. Visual features can include distinctive, unique, and / or identifiable portions of an image, such as corners, edges, gradients, and / or portions of an image that exhibit blobs. A blob can be defined as a region in which one or more image characteristics (e.g., brightness, color, tone, hue, saturation, or combinations thereof) are constant or nearly constant. To detect and / or extract features in an image, the feature extraction engine 255 can perform a scale space search, and the feature extraction engine 255 can use a frame buffer for the scale space search. To detect features in an image, the feature extraction engine 255 can use edge detection, corner detection, blob detection, ridge detection, affine invariant feature detection, or a combination thereof. Edge detection can include, for example, Canny detection, Deriche detection, difference detection, Sobel detection, Prewitt detection, and / or Roberts cross edge detection. Corner detection can include, for example, features from the Harris operator, Shi and Tomasi, level curve curvature, Hessian feature intensity measure, smallest univalue segment assimilating nucleus (SUSAN), and / or accelerated segment test (FAST) corner detection.Blob detection can include, for example, Laplacian of Gaussians (LoG), Difference of Gaussians (DoG), Determinant of Hessian (DoH), Maximum Stable Extremum Region, and / or Principal Curvature-based Region Detector (PCBR) blob detection. Affine-invariant feature detection can include Affine Shape Adaptive, Harris Affine, and / or Hessian Affine feature detection.

[0074]

[0088] To extract a feature, the feature extraction engine 255 can generate a feature descriptor. The feature descriptor can be generated based on extracting a local image patch around the feature and a description of the feature shown in the local image patch. The feature descriptor can describe the feature, for example, as a collection of one or more feature vectors. Features may be extracted using any suitable technique, such as Scale Invariant Feature Transform (SIFT), Learned Invariant Feature Transform (LIFT), Speed ​​Up Robust Features (SURF), Gradient Location-Orientation Histogram (GLOH), Histogram of Oriented Gradients (HOG), Oriented Fast and Rotated Brief (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoint (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross Correlation (NCC), Descriptor Matching, another suitable technique, or a combination thereof. In some embodiments, feature detection and / or extraction using feature extraction engine 255 may include identifying the location of a feature within an image, identifying the location of a feature within a 3D environment, or both.

[0075]

[0089] The feature management engine 250 includes a feature tracking engine 260. The feature extraction engine 255 can track features detected and / or extracted by the feature extraction engine 255 from one image to another. The feature tracking performed by the feature tracking engine 260 can include frame-to-frame tracking, box tracking, Kanade-Lucas-Tomasi (KLT) feature tracking, mean shift feature tracking, or combinations thereof. Some features represent parts of objects in the environment, such as a hand or a display screen. The feature tracking engine 260 can track the movement of objects in the environment by tracking features of the objects in the environment relative to features of the environment.

[0076]

[0090] The feature management engine 250 includes a data fusion engine 265. In some embodiments, the data fusion engine 265 can match features detected and / or extracted by the feature extraction engine 255 from sensor data received from one or more sensors 215 of the XR system 210 with features detected and / or extracted by the feature extraction engine 255 from sensor data received from one or more sensors 225 of the external device 220. In some cases, the one or more sensors 215 of the XR system 210 and the one or more sensors 225 of the external device 220 can be arranged such that there is at least some overlap between scenes of the real-world environment captured (in the case of imaging sensors) and / or sensed (in the case of non-imaging sensors) by the respective sensors. In some embodiments, the data fusion engine 265 can match features tracked by the feature tracking engine 260 from sensor data from one or more sensors 215 of the XR system 210 with features tracked by the feature tracking engine 260 from sensor data from one or more sensors 225 of the external device 220. For example, the data fusion engine 265 can identify a single three-dimensional point (having a three-dimensional set of coordinates) of a particular feature detected, extracted, and / or tracked in both the sensor data from the one or more sensors 215 of the XR system 210 and the sensor data from the one or more sensors 225 of the external device 220. By matching some features common in both sets of sensor data, the data fusion engine 265 can also map features in one set of sensor data but not in the other to features that are in both sets of sensor data. Thus, the data fusion engine 265 can locate features in the sensor data from the one or more sensors 225 of the external device 220 that are not present in the sensor data from the one or more sensors 215 of the XR system 210 to features that are present in the sensor data from the one or more sensors 215 of the XR system 210.Similarly, the data fusion engine 265 may locate features in the sensor data from one or more sensors 215 of the XR system 210 that are not present in the sensor data from one or more sensors 225 of the external device 220 relative to features present in the sensor data from one or more sensors 225 of the external device 220. In some embodiments, some operations described herein as being performed by the data fusion engine 265, such as feature mapping, may be performed regardless of whether the processing engine 205 of the XR system 210 receives sensor data from one or more sensors 225 of the external device 220. In some embodiments, some operations described herein as being performed by the data fusion engine 265, such as feature mapping, may be performed by another portion of the feature extraction engine 255, the feature tracking engine 260, or the feature management engine 250.

[0077]

[0091] In some examples, the feature management engine 250 can perform pose estimation of the pose of the XR system 210 (and / or each of the sensors 215 of the XR system 210) within the real-world environment in which the XR system 210 resides. The pose can include a position in a three-dimensional space, such as a set of three-dimensional translation coordinates (e.g., in the horizontal (x), vertical (y), and depth (z) directions). Additionally or alternatively, the pose can include an orientation (e.g., pitch, yaw, and / or roll). The feature management engine 250 can estimate the pose based on features detected and / or extracted by the feature extraction engine 255, based on features tracked by the feature tracking engine 260, based on features fused and / or mapped by the data fusion engine 265, or based on a combination thereof. In some aspects, the feature management engine 250 can perform stereo matching on the features, for example, when the sensors 215 and / or sensors 225 include a group (e.g., a pair) of image sensors representing a multi-scope view of the same scene. In some aspects, the feature management engine 250 can perform mapping such as map densification, keyframe addition, keyframe removal, bundle adjustment, loop closure detection, relocation, and / or one or more other simultaneous localization and mapping (SLAM) operations. In some examples, the pose of the XR system 210 (and / or each of the sensors 215 and / or 225) can be determined independent of feature detection and / or extraction. For example, the pose can be determined using a positioning procedure, such as using positioning reference signals (PRS), beacon signals, ToF measurements, etc. In the case of a stationary sensor or external device, the pose can be retrieved from a memory of the sensor or external device, or from a separate server where it may have been previously stored (e.g., during a calibration process, during device setup based on user input indicating the location of the sensor or external device, etc.).

[0078]

[0092] The feature management engine 250 can output feature information 270 based on features detected, extracted, and / or tracked from sensor data from one or more sensors 215 of the XR system 210 using the feature extraction engine 255 and / or the feature tracking engine 260. The feature management engine 250 can output extended feature information 275 based on features detected, extracted, tracked, and / or merged (combined) from both sensor data from one or more sensors 215 of the XR system 210 and sensor data from one or more sensors 225 of the external device 220 using the feature extraction engine 255 and / or the feature tracking engine 260, or using the feature extraction engine 255, the feature tracking engine 260, and / or the data fusion engine 265. In some cases, the extended feature information 275 can identify additional features not included in the feature information 270 and thus can represent a more complete feature mapping of the environment represented within the sensor data from one or more sensors 215 of the XR system 210 and / or the sensor data from one or more sensors 225 of the external device 220. The extended feature information 275 can identify more accurate feature locations than the feature information 270 and therefore can represent a more accurate feature mapping of the environment represented in the sensor data from the one or more sensors 215 of the XR system 210 and / or the sensor data from the one or more sensors 225 of the external device 220.

[0079]

[0093] The XR system 210 can include an output content engine 280. The output content engine 280 can generate output content 285 based on sensor data from one or more sensors 215 of the XR system 210, sensor data from one or more sensors 225 of the external device 220, and / or virtual content. In some embodiments, the output content 285 can include an output image that is a modified version of an input image from sensor data from one or more sensors 215 of the XR system 210, modified to add virtual content positioned based on augmented feature information 275 (including feature information extracted from sensor data from one or more sensors 225 of the external device 220). For example, a feature corresponding to a particular object, such as a hand or a display screen in the environment, may be in the augmented feature information 275 but not in the feature information 270 if the object is within the field of view of one or more sensors 225 of the external device 220 but not within the field of view of one or more sensors 215 of the XR system 210.

[0080]

[0094] The XR system 210 can output the output content 285 to an output device 290 of the XR system 210. The output device 290 can include, for example, a display, an audio output device, any of the output devices 1035 of FIG. 10, or a connector that can couple the XR system 210 to one of the types of output devices listed above. In some embodiments, the output content 285 can include one or more images and / or one or more videos that the XR system 210 can display using a display of the output device 290. The display can include a display screen, such as a liquid crystal display (LCD) display, a plasma display, a light emitting diode (LED) display, an organic LED (OLED) display, an electronic paper display, an electronic ink display, or a combination thereof. The display can include a projector and / or a projection surface onto which the projector projects an image. The projection surface can be opaque, transparent, or translucent. The display may be a display of a head mounted display (HMD) 310, a display of XR glasses (e.g., AR glasses), a display 345 of a mobile handset 410, and / or other devices. In some embodiments, the output content 285 may include one or more images of a video that the XR system 210 may display using a display of the output device 290. In some embodiments, the output content 285 may include one or more audio clips that the XR system 210 may play using an audio output device of the output device 290. The audio output device may include, for example, a speaker, headphones, or a combination thereof.

[0081]

[0095] In some examples, the XR system 210 receives the sensor data of the sensor 225 of the external device 220 directly from the external device 220. In some examples, the XR system 210 receives the sensor data of the sensor 225 of the external device 220 indirectly from an intermediate device. Examples of intermediate devices can include, for example, a server and / or a cloud service to which the external device 220 uploads its sensor data. The negotiation described herein as being performed between the inter-device negotiation engine 230 of the XR system 210 and the external device 220 can, in some cases, instead be performed between the inter-device negotiation engine 230 of the XR system 210 and the intermediate device.

[0082]

[0096] FIG. 3A is a perspective view 300 illustrating a head mounted display (HMD) 310 used as an extended reality (XR) system 210. The HMD 310 may be, for example, an augmented reality (AR) headset (e.g., AR glasses or smart glasses), a virtual reality (VR) headset, a mixed reality (MR) headset, another type of XR headset, or some combination thereof. The HMD 310 may be an embodiment of the XR system 210 or may be part of the XR system 210. The HMD 310 includes a first camera 330A and a second camera 330B along the front of the HMD 310. The first camera 330A and the second camera 330B may be embodiments of the sensor 215 of the XR system 210. In some embodiments, the HMD 310 may have only a single camera. In some embodiments, the HMD 310 may include one or more additional cameras, in addition to the first camera 330A and the second camera 330B, which may be embodiments of the sensors 215 of the XR system 210. In some embodiments, the HMD 310 may include one or more additional sensors, in addition to the first camera 330A and the second camera 330B, which may be embodiments of the sensors 215 of the XR system 210.

[0083]

[0097] The HMD 310 may include one or more displays 340 visible to a user 320 wearing the HMD 310 on the user's 320 head. The one or more displays 340 of the HMD 310 may be an embodiment of the output device 290 of the XR system 210. In some embodiments, the HMD 310 may include one display 340 and two viewfinders. The two viewfinders may include a left viewfinder for the left eye of the user 320 and a right viewfinder for the right eye of the user 320. The left viewfinder may be oriented so that the left eye of the user 320 sees the left side of the display. The right viewfinder may be oriented so that the right eye of the user 320 sees the right side of the display. In some embodiments, the HMD 310 may include two displays 340, including a left display that displays content to the left eye of the user 320 and a right display that displays content to the right eye of the user 320.

[0084]

[0098] FIG. 3B is a perspective view 350 showing the head mounted display (HMD) of FIG. 3A being worn by a user 320. The user 320 wears the HMD 310 on the user's 320 head over the user's 320 eyes. The HMD 310 can capture images with a first camera 330A and a second camera 330B. In some embodiments, the HMD 310 displays one or more output images to the user's 320 eyes. The output images may be examples of the output content 285. The output images can be based on the images captured by the first camera 330A and the second camera 330B. The output images can provide a stereoscopic view of the environment, possibly with information overlaid and / or other modifications. For example, the HMD 310 can display a first display image based on the image captured by the first camera 330A to the right eye of the user 320. The HMD 310 can display a second display image based on the image captured by the second camera 330B to the left eye of the user 320. For example, the HMD 310 can provide overlay information in the display image overlaid on top of the images captured by the first camera 330A and the second camera 330B.

[0085]

[0099] 4A is a perspective view 400 showing a front side of a mobile handset 410 including a forward-facing camera and used as an extended reality (XR) system 210. The mobile handset 410 may be an example of an XR system 210. The mobile handset 410 may be, for example, a mobile phone, a satellite phone, a portable gaming console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop, a mobile device, any other type of computing device or computing system 1100 described herein, or a combination thereof. The front side 420 of the mobile handset 410 includes a display 440. The front side 420 of the mobile handset 410 may include a first camera 430A and a second camera 430B. The first camera 430A and the second camera 430B may be examples of sensors 215 of the XR system 210. The first camera 430A and the second camera 430B are shown in a bezel around the display 440 on the front 420 of the mobile handset 410. In some embodiments, the first camera 430A and the second camera 430B can be located in a notch or cutout cut out of the display 440 on the front 420 of the mobile handset 410. In some embodiments, the first camera 430A and the second camera 430B can be under-display cameras located between the display 440 and the remainder of the mobile handset 410 such that light passes through a portion of the display 440 before reaching the first camera 430A and the second camera 430B. The first camera 430A and the second camera 430B in the perspective view 400 are forward-facing cameras. The first camera 430A and the second camera 430B face in a direction perpendicular to the plane of the front 420 of the mobile handset 410. The first camera 430A and the second camera 430B may be two of one or more cameras of the mobile handset 410. In some embodiments, the front face 420 of the mobile handset 410 may have only a single camera.In some embodiments, the mobile handset 410 may include one or more additional cameras in addition to the first camera 430A and the second camera 430B, which may be embodiments of the sensors 215 of the XR system 210. In some embodiments, the mobile handset 410 may include one or more additional sensors in addition to the first camera 430A and the second camera 430B, which may be embodiments of the sensors 215 of the XR system 210. The front face 420 of the mobile handset 410 also includes a display 440. In some cases, the front face 420 of the mobile handset 410 includes two or more displays 440. The one or more displays 440 on the front face 420 of the mobile handset 410 may be embodiments of the output devices 290 of the XR system 210.

[0086]

[0100] 4B is a perspective view 450 showing the back of a mobile handset that includes a rear-facing camera and is used as an extended reality (XR) system 210. The mobile handset 410 includes a third camera 430C and a fourth camera 430D on a rear surface 460 of the mobile handset 410. The third camera 430C and the fourth camera 430D in the perspective view 450 are rear-facing. The third camera 430C and the fourth camera 430D may be examples of sensors 215 of the XR system 210. The third camera 430C and the fourth camera 430D are oriented perpendicular to the plane of the rear surface 460 of the mobile handset 410. The rear surface 460 of the mobile handset 410 does not have a display 440 as shown in the perspective view 450, although in some embodiments, the rear surface 460 of the mobile handset 410 may include one or more rear displays. In embodiments where the back 460 of the mobile handset 410 includes one or more rear displays, the one or more rear displays may be embodiments of the output device 290 of the XR system 210. When the back 460 of the mobile handset 410 includes one or more rear displays, any arrangement layout of the third camera 430C and the fourth camera 430D relative to the one or more rear displays may be used, as described with respect to the first camera 430A and the second camera 430B relative to the display 440 of the front 420 of the mobile handset 410. The third camera 430C and the fourth camera 430D may be two of the one or more cameras of the mobile handset 410. In some embodiments, the back 460 of the mobile handset 410 may have only a single camera. In some embodiments, the mobile handset 410 may include one or more additional cameras in addition to the first camera 430A, the second camera 430B, the third camera 430C, and the fourth camera 430D, which may also be embodiments of the sensor 215 of the XR system 210.In some embodiments, the mobile handset 410 may include, in addition to the first camera 430A, the second camera 430B, the third camera 430C, and the fourth camera 430D, one or more additional sensors that may be embodiments of the sensors 215 of the XR system 210.

[0087]

[0101] 5 is a perspective view of a user wearing a head mounted display (HMD) 310 used as an extended reality (XR) system 210 and performing hand tracking to determine gesture-based input based on the position of the hand 525 of the user 320 within the field of view (FOV) 520 of the HMD 310. In another embodiment, the HMD 310 can be used to place a virtual object based on the position of the hand 525 within the FOV 520 of the HMD 310. The first camera 330A and / or the second camera 330B of the HMD 310 are used as the sensor 215 of the XR system 210. The FOV 520 of the HMD 310 represents the FOV of the first camera 330A and / or the second camera 330B. The FOV 520 of the HMD 310 is shown using dashed lines. The hand 525 of the user 320 is within the FOV 520 of the sensor 215 of the HMD 310. Thus, the XR system 210 of the HMD 310 detects, extracts, and / or tracks features of the hand 525 of the user 320 relative to other features of the real-world environment in which the user 320 and the HMD 310 are located to identify a pose of the hand 525 of the user 320 relative to the real-world environment in which the user 320 and the HMD 310 are located. The pose of the hand 525 may include a position of the hand and / or an orientation (e.g., pitch, yaw, and / or roll) of the hand 525. Based on the pose of the hand 525, the HMD 310 may determine a gesture-based input, such as for controlling a user interface (UI) of the HMD 310.

[0088]

[0102] As mentioned above, in some cases, the HMD 310 can determine where to display a virtual object relative to the hand 525 based on the determined pose of the hand 525. The virtual object represents a virtual object that the HMD 310 displays to the user 320 using the display 340, but does not exist in the real-world environment in which the user 320 and the HMD 310 are present. In one exemplary embodiment, the virtual object is a sword and can be displayed by the HMD 310 as if it were being held by the hand 525 of the user 320. The pose (position and orientation) of the virtual object depends on the pose of the hand 525. The output content generation engine 280 of the XR system 210 of the HMD 310 can add the virtual object 540 to the output content 285 before the output content 285 is displayed on the display(s) 340 (output on the output device 290).

[0089]

[0103] 6A is a perspective view 600 of a user 320 wearing a head mounted display (HMD) 310 used as an extended reality (XR) system 210 and performing hand tracking to determine gesture-based inputs based on the position of the user's 320 hands 525 even if the hands 525 are outside the field of view (FOV) 620 of the HMD 310. The HMD 310 can perform hand tracking based on the hands 515 being within the FOV 615 of an external camera 610 even if the hands 525 are outside the FOV 602. The FOV 620 of the HMD 310 represents the FOV of one or more cameras and / or other sensors of the HMD 310. The FOV 620 of the HMD 310 is shown using a dashed line. The hands 525 of the user 320 are not within the FOV 620 of the HMD 310 because the user 320 has moved the hands 525 too far away from the FOV 620 of the HMD 310. Thus, using its own camera and / or other sensors, the HMD 310 would not be able to identify and / or track the position of the user's hand 525 in that position of Figure 6A. Even though the hand 525 of the user 320 is not within the FOV 620 of the HMD 310, the hand 525 can still be tracked to determine any gesture-based input, to determine where to display a virtual object relative to the hand 525 when at least a portion of the virtual object is still displayed within the FOV 620 of the HMD 310 (depending on the illustrated pose of the hand 525 of the user 320), and / or to perform some other function based on the tracked pose of the hand 525.

[0090]

[0104] The XR system 210 of the HMD 310 losing sight of the hand 525 (or another object being tracked by the XR system 210) may be a condition that the XR system 210 detects and uses to determine when to perform one or more other functions. The XR system 210 of the HMD 310 may detect this condition in the situation shown in FIG. 6A due to the hand 525 moving out of the FOV 620 of the HMD 310 or due to no longer detecting the hand 525 within the FOV 620. The XR system 210 of the HMD 310 may send a request 640 for assistance with hand tracking to the external camera 610. The external camera 610 may be an example of the external device 220 of FIG. 2. For example, the external camera 610 may be part of an external device such as a laptop computer, a desktop computer, a television, a smart home device or assistant, a mobile device (e.g., a smartphone), a tablet computer, or other external device. One or more image sensors and / or other sensors of the external camera 610 may be an embodiment of the sensor 225 of the external device 220. The XR system 210 of the HMD 310 may perform inter-device negotiation with the external camera 610 as described with respect to the inter-device negotiation engine 230. In response, the external camera 610 may transmit hand tracking data 645 as part of a data stream to the XR system 210 of the HMD 310. The hand tracking data 645 may include sensor data captured by one or more sensors of the external camera 610, such as one or more image sensors. The FOV 615 of the external camera 610 is shown using dashed and dotted lines. The FOV 615 of the external camera 610 includes the hands 525 of the user 325. In some examples, before the external camera 610 transmits the hand tracking data 645 to the XR system 210 of the HMD 310, the hand tracking data 645 may be at least partially processed by the external camera 610, e.g., to detect features, extract features, track features, and / or perform one or more other operations of the feature management engine 250, which may reduce computational resources (e.g., battery consumption of the HMD 310, amount of processing resources being used, etc.).The XR system 210 of the HMD 310 can use the hand tracking data 645 to identify a pose of the hand 525 of the user 320, even though the hand 525 is not within the FOV 620 of the HMD 310. Even though the hand 525 is not within the FOV 620 of the HMD 310, the XR system 210 of the HMD 310 can use the hand pose determined based on the hand tracking data 645 to determine one or more gesture-based inputs being performed by the user (e.g., to control a UI of the HMD 310, such as an application running on the HMD 310), determine where to display a virtual object within the FOV 620 of the HMD 310 in a precise pose based on the pose of the hand 525 of the user 320, and / or perform one or more other functions.

[0091]

[0105] 6B is a perspective view 650 of a user 320 wearing a head mounted display (HMD) 310 used as an extended reality (XR) system 210 and performing hand tracking to determine gesture-based inputs based on the position of the user's 320's hand 525 when an occlusion 660 (e.g., a real-world object) occludes the hand 525 within the field of view (FOV) 670 of the HMD 310. The HMD 310 can perform hand tracking even when the hand 525 is occluded based on the hand 525 being within the FOV 615 of an external camera 610. The FOV 670 of the HMD 310 represents the FOV of one or more cameras and / or other sensors of the HMD 310. The FOV 670 of the HMD 310 is shown using dashed lines. The hand 525 of the user 320 is within the FOV 670 of the HMD 310, but is occluded from the field of view of the HMD 310 because the FOV 670 is partially occluded by the occlusion 660. The occlusion 660 occludes the hand 525 within the FOV 670 of the HMD 310. Thus, by itself, the HMD 310 would not be able to identify and / or track the position of the user's hand 525 at its position in FIG. 6B. Even though the hand 525 of the user 320 is occluded within the FOV 670 of the HMD 310, the hand 525 can still be tracked to determine any gesture-based input, determine where to display a virtual object relative to the hand 525 when at least a portion of the virtual object is still displayed within the FOV 670 of the HMD 310 (depending on the illustrated pose of the hand 525 of the user 320), and / or perform some other function based on the tracked pose of the hand 525.

[0092]

[0106] The XR system 210 of the HMD 310 losing sight of the hand 525 (or another object being tracked by the XR system 210) may be a condition that the XR system 210 detects and uses to determine when to perform one or more other functions. The XR system 210 of the HMD 310 may detect this condition in the situation shown in FIG. 6B because an occlusion 660 is occluding the hand 525 in the FOV 670 of the HMD 310. Similar to FIG. 6A, the XR system 210 of the HMD 310 in FIG. 6B may send a request 640 for hand tracking assistance to the external camera 610. The XR system 210 of the HMD 310 may perform inter-device negotiation with the external camera 610 as described with respect to the inter-device negotiation engine 230. In response, the external camera 610 may send hand tracking data 645 as part of a data stream to the XR system 210 of the HMD 310. The hand tracking data 645 may include sensor data captured by one or more sensors of the external camera 610, such as one or more image sensors. The FOV 615 of the external camera 610 is shown using dashed and dotted lines. The FOV 615 of the external camera 610 includes the hands 525 of the user 325. In some embodiments, the hand tracking data 645 may be at least partially processed by the external camera 610 to, for example, detect features, extract features, track features, and / or perform one or more other operations of the feature management engine 250 before the external camera 610 transmits the hand tracking data 645 to the XR system 210 of the HMD 310, which may reduce computational resources (e.g., battery consumption of the HMD 310, amount of processing resources being used, etc.). Using the hand tracking data 645, the XR system 210 of the HMD 310 can identify the pose of the hand 525 of the user 320 even though the hand 525 is not within the FOV 620 of the HMD 310, even though the hand 525 is occluded within the FOV 670 of the HMD 310.The determined hand pose can be used to determine one or more gesture-based inputs being performed by the user (e.g., to control the UI of the HMD 310, such as an application running on the HMD 310), to determine where to display virtual objects within the FOV 620 of the HMD 310 in a precise pose based on the pose of the user's 320 hands 525, and / or to perform one or more other functions.

[0093]

[0107] In some examples, the external camera 610 may be a standalone camera device, such as a security camera, as shown in Figures 6A and 6B. In some examples, the external camera 610 of Figures 6A and 6B may be one or more cameras of another HMD 710 (as in Figure 7), a mobile handset 410, a laptop computer, a desktop computer, or any other type of external device 220.

[0094]

[0108] 7 is a perspective view 700 showing an external HMD 710 device providing assistance by hand tracking the hands 525 of a user 320 of a head mounted display (HMD) 310 used as an extended reality (XR) system 210 due to a low battery state 735 (as an example of an operating state of an XR device) in the HMD 310. The FOV (not shown) of the HMD 310 may be the FOV of one or more cameras and / or one or more sensors of the HMD 310. The FOV (not shown) of the HMD 310 may include the hands 525 or may lack the hands 525. The FOV (not shown) of the external HMD 710 may be the FOV of one or more cameras and / or one or more sensors of the external HMD 710. The FOV (not shown) of the external HMD 710 may include the hands 525 or may lack the hands 525.

[0095]

[0109] The XR system 210 of the HMD 310 can detect a condition in the HMD 310 corresponding to a level of the computational resources of the HMD 310 meeting or being below a threshold level. The XR system 210 of the HMD 310 can detect a condition in the HMD 310 corresponding to a usage level of the computational resources of the HMD 310 meeting or exceeding a threshold level. For example, FIG. 7 illustrates the HMD 310 detecting a low battery condition 735 indicating that the battery level of one or more batteries of the HMD 310 meets or is below a threshold battery level (e.g., 50% of full battery level, 40% of full battery level, or other level). An alert 730 is illustrated based on the HMD 310 detecting the low battery condition 735. The XR system 210 of the HMD 310 can send a request 740 for hand tracking assistance to the external HMD 710. The external HMD 710 may be an embodiment of the external device 220 of FIG. 2. One or more image sensors and / or other sensors of the external HMD 710 may be examples of sensors 225 of the external device 220. The XR system 210 of the HMD 310 may perform inter-device negotiation with the external HMD 710 as described with respect to the inter-device negotiation engine 230. In response, the external HMD 710 may transmit hand tracking data 745 as part of a data stream to the XR system 210 of the HMD 310. The hand tracking data 745 may include sensor data captured by one or more sensors of the external HMD 710, such as one or more image sensors. In some examples, before the external HMD 710 transmits the hand tracking data 745 to the XR system 210 of the HMD 310, the hand tracking data 745 may be at least partially processed by the external HMD 710, e.g., to detect features, extract features, track features, and / or perform one or more other operations of the feature management engine 250, and to reduce computational resources (e.g., reduce battery consumption of the HMD 310, reduce the amount of processing resources being used, etc.).The XR system 210 of the HMD 310 can use the hand tracking data 745 to identify the posture of the hand 525 of the user 320 and / or whether the hand 525 is within the FOV (not shown) of the HMD 310.

[0096]

[0110] Because the HMD 310 can offload at least some of its hand tracking tasks to the external HMD 710, the HMD 310 can reduce its battery load and use the battery longer, and therefore last longer despite its low battery state 735. In some examples, the HMD 310 can turn off or otherwise disable its camera and / or other sensors. In some examples, the HMD 310 can reduce the capture quality or rate of sensor data from its sensors, for example, from 90 fps image capture to 30 fps capture. In some examples, the HMD 310 can partially or fully rely on the camera and / or other sensors of the external HMD 710. In some examples, the HMD 310 can at least partially turn off or otherwise disable at least some of the functionality of the feature management engine 250, such as the feature extraction engine 255, the feature tracking engine 260, and / or the data fusion engine 265. In some embodiments, the HMD 310 may partially or fully rely on an external HMD 710 to perform at least some of the functions of the feature management engine 250, such as the feature extraction engine 255, the feature tracking engine 260, and / or the data fusion engine 265. In some embodiments, the HMD 310 may turn off or otherwise disable the display 340 of the HMD 310. In some embodiments, the HMD 310 may send its output content 285 to another display device, such as a smartwatch, a laptop, or another display device. These adjustments to the operation of the XR system 210 of the HMD 310 may enable the HMD 310 to reduce its battery load and use the battery longer and therefore last longer despite its low battery state 735.

[0097]

[0111] In some examples, the XR system 210 of the HMD 310 can detect other conditions besides the low battery condition 735 of FIG. 7. For example, detecting the condition can include detecting a level of other computational resources of the HMD 310 that meets or is below a threshold level. Detecting the condition can include detecting a usage level of computational resources of the HMD 310 that meets or exceeds a threshold level. For example, the condition can be that the available memory (e.g., memory 1015, ROM 1020, and / or RAM 1025) of the HMD 310 meets or is below a threshold memory level. The condition can be that the available storage space (e.g., on the storage device 1030) of the HMD 310 meets or is below a threshold level. The condition can be that the available network bandwidth of the HMD 310 meets or is below a threshold network bandwidth level. The condition can be that the available processor bandwidth of the HMD 310 meets or is below a threshold processor bandwidth level. The condition may be that the processor usage of the HMD 310 meets or exceeds a threshold processor usage level.

[0098]

[0112] In some embodiments, the external HMD 710 of Figure 7 may be an HMD as shown in Figure 7. In some embodiments, the external HMD 710 may instead be a standalone camera device (e.g., a security camera) (as in external camera 610 of Figures 6A and 6B), a mobile handset 410, or any other type of external device 220.

[0099]

[0113] 8A is a perspective view 800 showing a user 320 wearing a head mounted display (HMD) 310 used as an extended reality (XR) system 210 and placing virtual content 815 within an image displayed by a display(s) 340 of the HMD 310 based on the location of an external display 810 (external to the HMD 310) within a FOV 835 of the HMD 310 and / or visual (media) content 812 displayed on the external display 810. As shown in FIG. 8A, the user 320 wearing the HMD 310 faces an external display 810 displaying visual (media) content 812. The external display 810 includes a camera 814. The FOV 835 of the HMD 310 represents the FOV of one or more cameras and / or other sensors of the HMD 310. The FOV 835 of the HMD 310 is shown using dashed lines. The external display 810 and the visual (media) content 812 displayed on the display 810 are both within the FOV 835 of the HMD 310.

[0100]

[0114] The XR system 210 of the HMD 310 can detect the external display 810 (e.g., in one or more images captured by one or more cameras and / or other sensors of the HMD 310) and / or can detect visual (media) content 812 displayed on the external display 810. Detection of the external display 810 and / or detection of the visual (media) content 812 displayed on the external display 810 may be a state that the XR system 210 of the HMD 310 detects and uses to determine when to perform one or more other functions (e.g., determine the position of other objects in the external display 810 and / or the environment surrounding the HMD 310, perform a function based on that position, etc.). The XR system 210 of the HMD 310 can detect this state in the situation shown in FIG. 8A because the display 810 and visual (media) content 812 are within the FOV 835 of the HMD 310.

[0101]

[0115] In some examples, in response to detecting the condition, the XR system 210 of the HMD 310 can send a request 840 for additional (media) content 845 to one or more servers 847. In some examples, the request 840 can be based on a particular visual (media) content 812 detected by the XR system 210 of the HMD 310, for example, based on a media recognition system of the XR system 210 of the HMD 310. The request 840 can identify the visual (media) content 812 detected by the XR system 210 of the HMD 310. The one or more servers 847 can provide the additional (media) content 845 to the XR system 210 of the HMD 310. The additional (media) content 845 can be specific to the visual (media) content 812. In some cases, the request 840 may include a representation of the visual (media) content 812 captured by a sensor of the HMD 310, and the one or more servers 847 may recognize the particular visual (media) content 812 based on a media recognition system of the one or more servers 847. The XR system 210 of the HMD 310 may generate the virtual content 815 using the additional (media) content 845. The XR system 210 of the HMD 310 may determine the pose (e.g., position and / or orientation) of the virtual content 815 within the FOV 835 of the HMD 310 within the output content 285 based on the pose (e.g., position and / or orientation) of the display 810 and / or the visual (media) content 812 within the FOV 835 of the HMD 310. The virtual content 815 may include a title 820 of the visual (media) content 812, identified in FIG. 8A as “Speedy Tracking”. Title 820 may be displayed adjacent to and above display 810 and visual (media) content 812. In one embodiment, virtual content 815 may include display extension 825 that extends display 810 adjacent to and to the right of display 810 and visual (media) content 812, for example based on additional widescreen video data in additional (media) content 845.The virtual content 815 may include metadata 830 about the virtual content 815 adjacent to and to the left of the display 810 and the visual (media) content 812. The metadata 830 may identify the release date of the virtual content 812 (1998) and that the visual (media) content 815 stars a famous actor. In some embodiments, the virtual content 815 may include additional information or content related to the visual (media) content 812, such as deleted scenes. In some embodiments, at least some of the virtual content 815 may be overlaid on top of the display 810 and / or the visual (media) content 812. For example, the virtual content 815 may be used to highlight or circle a particular actor or object in the visual (media) content 812. For example, if the visual (media) content 812 is a sports game, the virtual content 815 may highlight or circle an important object that is difficult to see, such as a ball or a hockey puck.

[0102]

[0116] 2, the external display 810 can function as the external device 220, and the visual (media) content 812 can function as a data stream from the external device 220, akin to sensor data from the sensor 225. In some cases, instead of or in addition to displaying the visual content 812, the display 810 can transmit the visual (media) content 812 to the XR system 210 of the HMD 310, so that the XR system 210 of the HMD 310 can more easily detect and / or recognize the visual (media) content 812 in images and / or other sensor data captured by the image sensor and / or other sensors of the HMD 310. In some examples, one or more servers 847 can function as the external device 220, and the additional (media) content 845 can function as a data stream from the external device 220, akin to sensor data from the sensor 225.

[0103]

[0117] In another example, a user wearing the HMD 310 can face the external display 810 such that the external display 810 is within the FOV of one or more cameras and / or other image sensors. One or more cameras (and / or other image sensors) of the HMD 310 and the camera 814 (and / or other image sensors) of the external display 810 can be used for object tracking. Similar to what was described with respect to Figures 6A and 6B, based on detecting such conditions, the HMD 310 can determine whether to use the camera / image sensor(s) of the HMD 310, the camera / image sensor(s) of the external display 810, or the camera / image sensor(s) of both the HMD and the external display 810 for tracking purposes.

[0104]

[0118] 8B is a perspective view 850 showing a user 320 wearing a head mounted display (HMD) 310 used as an extended reality (XR) system 210 that places a virtual representation 860 of visual (media) content 812 displayed on the display 810 within an image displayed by a display(s) 340 of the HMD 310 based on the position of the display 810 and / or visual (media) content 812 even though the display 810 and / or visual (media) content 812 are outside the field of view (FOV) 890 of the HMD 310. The user 320 wearing the HMD 310 is no longer facing the display 810 displaying the visual (media) content 812. The FOV 890 of the HMD 310 represents the FOV of one or more cameras and / or other sensors of the HMD 310. The FOV 890 of the HMD 310 is shown using dashed lines. The display 810 and the visual (media) content 812 displayed on the display 810 are not within (and therefore are missing from) the FOV 890 of the HMD 310.

[0105]

[0119] In one embodiment, the XR system 210 of the HMD 310 may detect the presence of the display 810 in the vicinity of the HMD 310 (e.g., detected within wireless communication range of the HMD 310 or within the FOV of the HMD 310 at an earlier time), which may be a state that the XR system 210 of the HMD 310 detects and uses to determine when to perform one or more other functions. In one embodiment, the XR system 210 of the HMD 310 may determine that it has lost sight of the display 810 and / or the visual (media) content 812 (e.g., based on determining that the display 810 and / or the visual content 812 is no longer within the FOV 890 of the HMD 310), which may be a state that the XR system 210 of the HMD 310 detects and uses to determine when to perform one or more other functions. The XR system 210 of the HMD 310 may detect such a condition in the situation shown in FIG. 8B due to, for example, the user 320 turning the user's head and / or body to the right such that the display 810 and the visual (media) content 812 are no longer within the FOV 890 of the HMD 310. In response to detecting the condition, the XR system 210 of the HMD 310 may automatically send a request 880 for the visual (media) content 812 to the display 810 and / or one or more computing devices associated with the display 810 (e.g., an entertainment device, a media center device, or a computing system 1000 connected to the display 810). The display 810 and / or one or more computing devices associated with the display 810 may respond to the request 880 by providing the visual (media) content 812 as part of a data stream. The XR system 210 of the HMD 310 may generate a virtual representation 860 of the visual (media) content 812 as virtual content 815 within the FOV 890 of the HMD 310. In some cases, the XR system 210 of the HMD 310 can generate a directional indicator 870 as virtual content 815 within the FOV 890 of the HMD 310.The directional indicator 870 points to the position of the display 810 displaying the visual (media) content 812. The virtual representation 860 of the visual content 812 can enable the user 320 of the HMD 310 to continue to see the visual (media) content 812 even if the user 320 rotates away from the display 810. Thus, the user 320 does not have to miss any of the visual (media) content 812 even if the user 320 needs to temporarily change orientation. The directional indicator 870 pointing left can inform the user 320 to rotate left to re-orient to the display 810 displaying the visual (media) content 812. Additional virtual content 815 based on additional (media) content 845 from one or more servers 847, such as titles 820 of the virtual (media) content 812, can also be displayed within the FOV 890 of the HMD 310.

[0106]

[0120] 2, the display 810 can function as the external device 220 and the visual (media) content 812 can function as a data stream from the external device 220, akin to sensor data from the sensor 225. In some cases, instead of or in addition to displaying the visual (media) content 812, the display 810 can transmit the visual (media) content 812 to the XR system 210 of the HMD 310, so that the XR system 210 of the HMD 310 can more easily detect and / or recognize the visual (media) content 812 in images and / or other sensor data captured by the image sensor and / or other sensors of the HMD 310. In some examples, one or more servers 847 can function as the external device 220 and the additional (media) content 845 can function as a data stream from the external device 220, akin to sensor data from the sensor 225.

[0107]

[0121] Other examples of conditions that may cause the HMD 310 to perform one or more functions (e.g., determining the position of an object, requesting an external device to offload resources, requesting assistance from an external device with hand tracking, etc.) may include a user input or setting that requests the use of an external device rather than an imaging device (e.g., an XR device) when available for a particular function (e.g., displaying content, tracking an object such as a user's hand, head, or body), a user input or setting that indicates a preference that a device (e.g., an external device) be used for a particular function when plugged into the imaging device, privacy and / or security are factors (which may also be based on user input or setting), based on user input (e.g., a user input that requests that resources be offloaded to an external device, such as a user input requesting that the imaging device be turned off, a user input requesting that an external device, such as a light be turned on or off via a home automation application running on the imaging device, etc.), based on the capabilities of the imaging device's image sensor (e.g., when an infrared (IR) sensor on one device is useful when ambient lighting is insufficient, when the object being tracked is moving at a high speed and an image sensor with a higher frame rate is more appropriate, etc.), or any combination thereof.

[0108]

[0122] For example, the HMD 310 or an application running on the HMD 310 may be programmed with settings (e.g., based on user input provided to the HMD 310 and / or application, set by default, etc.) indicating a preference to use an external device for a particular function when the external device is available (e.g., physically or wirelessly connected to the HMD 310) and / or when the external device is capable of performing the function. In one example, an external display (e.g., a television, laptop computer, smart home device or assistant, tablet computer, desktop computer, external XR device, etc.) connected to the HMD 310 may be used to display content for the HMD 310 based on such settings being selected by a user or otherwise enabled (or possibly set by default). In another example, one or more cameras and / or other sensors of an external device connected to the HMD 310 may be used to track an object (e.g., a user's hand, head, or body, an additional external device other than the external device performing the tracking) based on such settings being selected by a user or otherwise enabled (or possibly set by default).

[0109]

[0123] In some examples, the HMD 310 or an application running on the HMD 310 may be programmed with privacy or security settings (e.g., based on user input provided to the HMD 310 and / or application, set by default, etc.) that indicate a preference for using an external device when security and / or privacy may be compromised by using the HMD 310. For example, based on a privacy or security setting selected or otherwise enabled by a user (or possibly set by default), the HMD 310 may determine that content displayed on the HMD 310 is viewable by other people and / or cameras and is therefore not private or secure. In response to determining that the content is not private / secure, the HMD 310 may send a command to the external device requesting that the external device display the content.

[0110]

[0124] In some cases, the HMD 310 can request assistance from an external device based on the capabilities and / or components of the external device. For example, the external device can include an image sensor not present on the HMD 310. In one example, the image sensor can include an IR sensor capable of performing object tracking (e.g., hand tracking, head tracking, body tracking, etc.) when ambient lighting is insufficient (e.g., in a low light condition). In such an example, the HMD 310 can detect when a low light condition exists (e.g., based on analyzing an image captured by a camera of the HMD 310), such as when one or more light values ​​of an image are below an illumination threshold (e.g., below a particular luminance, lux, or other illumination value, such as 3 lux or less). In response to detecting a low light condition, the HMD 310 can send a command to the external device requesting that the external device either capture an image using the IR sensor and / or any other sensor and perform object tracking using the image (in which case the external device can send pose information to the HMD 310), or send the image to the HMD 310 to perform tracking. In another embodiment, the image sensor may include a camera capable of capturing images at a high frame rate that may be used to track objects moving at high speeds. In such an embodiment, the HMD 310 may detect that an object is moving at high speeds and may send a command to the external device requesting that the external device either capture images using the high frame rate camera and / or any other sensor and use the images to perform object tracking or send the images to the HMD 310 to perform tracking.

[0111]

[0125] In some examples, a user can provide user input (e.g., gesture input, pressing a virtual or physical button, etc.) to control whether the HMD 310 or an external device executes a particular function. In one example, a user can provide user input to the HMD 310 requesting that the HMD 310 offload object tracking functions (e.g., hand tracking, head tracking, body tracking, etc.) to an external device (e.g., a television, laptop computer, smart home device or assistant, tablet computer, desktop computer, external XR device, etc.) even when the battery of the HMD 310 is above a threshold and the hand is within the FOV of the HMD 310. For example, a user can plan to use the HMD 310 for an extended period of time (e.g., playing a game for an extended period of time) that will require a battery-based handoff to an external device at some point. In another example, a user may prefer to use the HMD 310 for a function even when the function drains the battery if the execution of the function can be better performed by the HMD 310 rather than an external device (e.g., based on one or more capabilities or components of the HMD 310). In such an embodiment, the user can provide user input to the HMD 310 to override the handoff of functionality to the external device.

[0112]

[0126] In some cases, the HMD 310 can detect a condition that indicates that an external device is needed to perform a function or that the HMD 310 is needed to perform a function. In one exemplary embodiment, while performing hand tracking of a user's hand of the HMD 310, the HMD 310 can determine (e.g., based on past use or the nature of the task) that the hand is moving toward the edge of the HMD 310's FOV and therefore that the user will continue to move the hand beyond the HMD 310's FOV. Before or when the hand moves past the FOV, the HMD 310 can send a command to the external device to turn on one or more cameras and begin capturing images or videos of the hand. The HMD 310 can request that the external device perform object tracking and send hand pose information to the HMD 310 or send images / videos to the HMD 310 so that the HMD 310 can perform tracking. In such an embodiment, the HMD 310 can resume performing tracking once the hand returns within the known FOV of one or more cameras of the HMD 310. In another example embodiment, the HMD 310 may determine that a user is moving (or will be moving) away from the FOV of one or more sensors (e.g., cameras or other sensors) that are fixed in place (e.g., a camera on a laptop) and that are being used for object tracking. Based on determining that the user will be moving out of the FOV of the one or more sensors, the HMD 310 may transition to performing tracking using its own camera or other sensor (in which case the HMD 310 sends a command to the external device to stop performing tracking using that sensor). In some cases, when the HMD 310 and / or the external device determines that it will not use one or more sensors (e.g., cameras) for tracking, the HMD 310 and / or the external device may turn off the sensor, which may save power, improve privacy / security, etc.

[0113]

[0127] In some examples, the HMD 310 can detect additional conditions that can trigger the HMD 310 to execute a function or resume execution of a function previously offloaded to an external device. For example, as described with respect to the example of FIG. 7, the HMD 310 can offload one or more object tracking tasks (e.g., hand tracking, head tracking, body tracking, etc.) to an external device based on an operational state of the HMD 310 (e.g., when the battery of the HMD 310 is low on power or other computational resources, such as below a threshold battery level). The HMD 310 can then charge the battery of the HMD 310 such that the battery level is greater than the threshold battery level. Based on detecting that the battery level has exceeded the threshold battery level, the HMD 310 can send a command to the external device requesting that one or more object tracking tasks be at least partially executed by the HMD 310. In response to the command, the external device can stop execution of the object tracking task(s) and the HMD 310 can start or resume execution of the object tracking task(s).

[0114]

[0128] 9 is a flow diagram illustrating a process 900 for processing image data. The process 900 may be performed by an imaging system. In some embodiments, the imaging system may be the XR system 210 of FIG. 2. In some embodiments, the imaging system may include, for example, the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the XR system 210, the processing engine 205, the inter-device negotiation engine 230, the feature management engine 250, the output content generation engine 280, the output device 290, a head mounted display (HMD) device (e.g., HMD 310), a mobile handset 410, an external HMD device 710, one or more servers 847, a computing system 1000, or a combination thereof.

[0115]

[0129] At operation 905, the process 900 includes receiving, by a device (e.g., an imaging system), an image of a portion of an environment captured by an image sensor (e.g., an image sensor of the device). The environment includes an object. At operation 910, the process 900 includes identifying a data stream from an external device. Examples of external devices may include the external device 220, the sensor 225 of the external device 220, the HMD 310 of FIG. 3, the mobile handset 410, the external camera 610, the external HMD 710, the display 810, one or more servers 847, the computing system 1000, or a combination thereof.

[0116]

[0130] At operation 915, the process 900 includes detecting a condition based on an image, a data stream, a device operational state, or any combination thereof. In some cases, detecting the condition based on the image includes determining that an object is missing from a portion of an environment in the image. In one embodiment, determining that an object is missing from a portion of an environment in the image includes determining that at least a portion of the object is occluded in the image (e.g., as shown in FIG. 6B). In some cases, detecting the condition based on the device operational state includes determining that resource availability is below a threshold. In one embodiment, determining that resource availability is below a threshold includes determining that a battery level of a battery is below a battery level threshold. In another embodiment, determining that resource availability is below a threshold includes determining that an available bandwidth is below a bandwidth threshold. In some cases, detecting the condition based on the device operational state includes receiving a user input corresponding to offloading processing to an external device. For example, as described above, a user can provide user input (e.g., gesture input, pressing a virtual or physical button, etc.) to control whether the HMD 310 or the external device performs a particular function.

[0117]

[0131] In some examples, detecting the condition based on the image includes determining one or more lighting conditions (e.g., low light conditions) in the image. In some cases, determining the one or more lighting conditions in the image can include determining that one or more light values ​​of the image are below an illumination threshold (e.g., an illumination threshold of 3 lux).

[0118]

[0132] In some examples, the object is a display of an external display device. In some cases, process 900 includes detecting the state based on the image, at least in part, by identifying within the image visual media content displayed on a display of the external display device.

[0119]

[0133] At operation 920, the process 900 includes, in response to detecting the condition, determining a position of the object within the environment based on at least one of the image and the data stream. In some instances, the external device includes a second image sensor. In some instances, the data stream includes a second image of a second portion of the environment, and determining the position of the object within the environment is based at least in part on a depiction of the object in the second image. In some examples, the portion of the environment in the image and the second portion of the environment overlap.

[0120]

[0134] In some examples, determining the location of the object within the environment includes sending a request to an external device identifying the location of the object within the environment. In some examples, process 900 can include receiving a response from the external device identifying the location of the object within the environment.

[0121]

[0135] In some examples, in response to detecting the condition, process 900 may include generating a merged data set by combining at least data from the data stream with an image captured by the image sensor. In such examples, determining the location of the object may be based at least in part on the merged data set.

[0122]

[0136] At operation 925, process 900 includes generating an output based on the position of the object in the environment. In some embodiments, generating the output includes generating content. In some cases, process 900 includes outputting the content based on the position of the object in the environment. For example, outputting the content includes, and can include, transmitting or sending the content to a display of the device to be displayed. In some embodiments, the content virtually extends the display of an external display device. In some cases, process 900 can include sending the content to be played to an audio output device.

[0123]

[0137] In some embodiments, generating the output includes controlling a device based on a user input. For example, the HMD 310 can receive a user input for controlling a device or the HMD 310 (e.g., a user input requesting to turn on or off an external device such as a light via a home automation application running on the imaging device, a user input requesting to turn off the HMD 310, etc.).

[0124]

[0138] In some examples, generating the output includes generating the content at least in part by overlaying virtual content over a region of the image. In such examples, the region of the image is based on a position of an object in the environment. If the object is a display of an external display device, the region of the image is adjacent to a representation of the display of the external display device in the image. In some examples, the object is a hand of a user of the device, and the hand is at least partially adjacent to the region of the image.

[0125]

[0139] In some examples, the process 900 can include detecting an additional condition based on at least one of an additional image captured by the image sensor, a data stream, and an operational state of the device. In response to detecting the additional condition, the process 900 can include performing a function previously performed by the external device. For example, the HMD 310 described above can detect an additional condition that can trigger the HMD 310 to perform a function or resume performance of a function previously offloaded to the external device (e.g., hand tracking, head tracking, body tracking, etc.).

[0126]

[0140] In some examples, the processes described herein (e.g., process 900 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 900 may be performed by the XR system 210 of FIG. 2. In another example, process 900 may be performed by a computing device having the computing system 1000 shown in FIG. 10. For example, a computing device having the computing system 1000 shown in FIG. 10 may include components of the image processing engine 205 of the XR system 210 and may perform the operations of FIG. 10.

[0127]

[0141] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device, having resource capabilities to perform the processes described herein, including process 900. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) configured to perform steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0128]

[0142] Components of a computing device may be implemented in circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0129]

[0143] Process 900 is illustrated as a logical flow diagram, whose operations represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement a process.

[0130]

[0144] Additionally, process 900 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored in a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0131]

[0145] 10 illustrates an example of a system for implementing some aspects of the present technology. Specifically, FIG. 10 illustrates an example of a computing system 1000, which may be any computing device, such as an internal computing system, a remote computing system, a camera, or any component thereof, where the components of the system communicate with each other using a connection 1005. The connection 1005 may be a physical connection using a bus, or a direct connection to a processor 1010, such as in a chipset architecture. The connection 1005 may also be a virtual connection, a network connection, or a logical connection.

[0132]

[0146] In some embodiments, computing system 1000 is a distributed system in which the functionality described in this disclosure may be distributed across a data center, multiple data centers, a peer network, etc. In some embodiments, one or more of the described system components represent many components that each perform some or all of the functionality that is the subject of the component description. In some embodiments, the components may be physical or virtual devices.

[0133]

[0147] The exemplary system 1000 includes at least one processing unit (CPU or processor) 1010 and connections 1005 coupling various system components to the processor 1010, including system memory 1015, such as read only memory (ROM) 1020 and random access memory (RAM) 1025. The computing system 1000 may include a cache of high speed memory 1012, either directly connected to the processor 1010, in close proximity to the processor 1010, or integrated as part of the processor 1010.

[0134]

[0148] The processor 1010 may include any general purpose processor, as well as hardware or software services, such as services 1032, 1034, and 1036 stored in a storage device 1030, configured to control the processor 1010, and special purpose processors where the software instructions are integrated into the actual processor design. The processor 1010 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0135]

[0149] To enable user interaction, computing system 1000 includes input devices 1045, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 1000 may also include output devices 1035, which may be one or more of several output mechanisms. In some cases, a multi-modal system may enable a user to provide multiple types of input / output to communicate with computing system 1000. Computing system 1000 may include a communication interface 1040, which may generally govern and manage user input and system output.The communications interface may be any of the following: audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, Apple® Lightning® port / plug, Ethernet® port / plug, fiber optic port / plug, proprietary wired port / plug, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, radio-frequency identification (RFID) wireless signal transmission, near-field communications (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WLAN), and the like. The present invention may be capable of performing or facilitating the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Access), Infrared (IR) communications radio signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network radio signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, radio signal transmission along the electromagnetic spectrum, or any combination thereof.The communication interface 1040 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the position of the computing system 1000 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo GNSS. There is no constraint to operate on any particular hardware configuration, and therefore the basic features herein may be easily substituted for improved hardware or firmware configurations as they are developed.

[0136]

[0150] The storage device 1030 may be a non-volatile and / or non-transitory and / or computer readable memory device, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disk (DVD) optical disk, a blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (ICC), a circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROMIt may also be a hard disk or other type of computer readable medium capable of storing data that is accessible by a computer, such as EPROM, FLASHEPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0137]

[0151] The storage device 1030 may include software services, servers, services, etc. that cause the system to perform functions when code defining such software is executed by the processor 1010. In some embodiments, hardware services that perform particular functions may include software components stored in a computer-readable medium in association with the necessary hardware components, such as the processor 1010, connections 1005, output devices 1035, etc., to perform the functions.

[0138]

[0152] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instruction(s) and / or data. Computer-readable media can include non-transitory media that can store data and do not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memories, or memory devices. A computer-readable medium can have code and / or machine-executable instructions stored thereon, which can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0139]

[0153] In some embodiments, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when stated, non-transitory computer readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0140]

[0154] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments can be practiced without these specific details. For clarity of explanation, in some cases, the present technology may be presented as including individual functional blocks, including devices, device components, steps or routines in a method implemented in software, or functional blocks comprising a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0141]

[0155] Particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowcharts may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0142]

[0156] The processes and methods according to the above-described embodiments can be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions can include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions, or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions. Portions of the computer resources used can be accessible over a network. The computer-executable instructions can be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described embodiments include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, etc.

[0143]

[0157] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program product) to perform the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor or processors may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in a peripheral device or an add-in card. Such functionality may also be implemented on a circuit board between different chips, or on different processes running on a single device, as further examples.

[0144]

[0158] The instructions, media for carrying such instructions, computational resources for executing such instructions, and other structures for supporting such computational resources are exemplary means for providing the functionality described in this disclosure.

[0145]

[0159] In the above description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the concepts of the present invention can be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above can be used individually or jointly. Moreover, the embodiments can be utilized in any number of environments and applications other than those described herein without departing from the scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0146]

[0160] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the present specification.

[0147]

[0161] When a component is described as being "configured to" perform some operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0148]

[0162] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0149]

[0163] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, a claim language reciting "at least one of A and B" means A, B, or A and B. In another example, a claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, a claim language reciting "at least one of A and B" can mean A, B, or A and B, and can further include items not listed in the set of A and B.

[0150]

[0164] Various exemplary logic blocks, modules, circuits, and algorithm steps described with respect to the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0151]

[0165] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0152]

[0166] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or may be incorporated into a combined video encoder-decoder (CODEC).

[0153]

[0167] Exemplary aspects of the present disclosure include the following.

[0154]

[0168] Aspect 1: An apparatus for processing image data, the apparatus comprising at least one memory and one or more processors coupled to the memory, the one or more processors configured to: receive an image of a portion of the environment captured by an image sensor, the environment including an object, identify a data stream from an external device, detect a condition based on at least one of the image, the data stream, and an operational state of the apparatus, determine a position of the object in the environment based on at least one of the image and the data stream in response to detecting the condition, and generate an output based on the position of the object in the environment.

[0155]

[0169] Aspect 2: The apparatus of aspect 1, wherein to detect a state based on an image, the one or more processors are configured to determine that an object is missing from a portion of an environment in the image.

[0156]

[0170] Aspect 3: The apparatus of aspect 2, wherein to determine that an object is missing from a portion of an environment in an image, one or more processors are configured to determine that at least a portion of the object is occluded in the image.

[0157]

[0171] Aspect 4: The apparatus of aspect 2 or 3, wherein the external device includes a second image sensor, the data stream includes a second image of a second portion of the environment, and determining the position of the object within the environment is based at least in part on a depiction of the object in the second image.

[0158]

[0172] Embodiment 5: The apparatus of embodiment 4, wherein the part of the environment and the second part of the environment overlap.

[0159]

[0173] Aspect 6: The apparatus of any one of aspects 1 to 5, wherein to detect the condition based on an operating state of the apparatus, the one or more processors are configured to determine when resource availability falls below a threshold.

[0160]

[0174] Aspect 7: The apparatus of aspect 6, wherein to determine that resource availability falls below a threshold, the one or more processors are configured to determine that a battery level of the battery falls below a battery level threshold.

[0161]

[0175] Aspect 8: The apparatus of aspect 6 or 7, wherein to determine that resource availability falls below a threshold, the one or more processors are configured to determine that available bandwidth falls below a bandwidth threshold.

[0162]

[0176] Aspect 9: The apparatus of any one of aspects 1 to 8, wherein one or more processors are configured to receive user input corresponding to offloading processing to an external device to detect a state based on an operating state of the apparatus.

[0163]

[0177] Example 10: The apparatus of any one of Examples 1 to 9, wherein the one or more processors are configured to generate content to generate the output.

[0164]

[0178] Aspect 11: The apparatus of aspect 10, wherein the one or more processors are configured to output content based on a position of an object within the environment.

[0165]

[0179] Aspect 12: The apparatus of aspect 11, further comprising a display, wherein the one or more processors are configured to send content to be displayed to the display to output the content.

[0166]

[0180] Aspect 13: The apparatus of any one of aspects 1 to 12, wherein the one or more processors are configured to detect an additional condition based on at least one of an additional image captured by the image sensor, a data stream, and an operational state of the apparatus, and in response to detecting the additional condition, perform a function previously performed by the external device.

[0167]

[0181] Example 14: The device of any one of Examples 1 to 13, wherein the one or more processors are configured to control the device based on user input to generate the output.

[0168]

[0182] Aspect 15: An apparatus described in any one of aspects 1 to 14, wherein to detect a condition based on an image, one or more processors are configured to determine one or more lighting conditions in the image.

[0169]

[0183] Aspect 16: The device of aspect 15, wherein to determine one or more lighting conditions in the image, the one or more processors are configured to determine that one or more light values ​​of the image are below an illumination threshold.

[0170]

[0184] Aspect 17: An apparatus described in any one of aspects 1 to 16, wherein to determine a position of an object within the environment, the one or more processors are configured to send a request to an external device identifying a position of the object within the environment and receive a response from the external device identifying the position of the object within the environment.

[0171]

[0185] Example 18: The apparatus of any one of examples 1 to 17, wherein the object is a display of an external display device.

[0172]

[0186] Aspect 19: The apparatus of aspect 18, wherein to detect a state based on an image, the one or more processors are configured to identify, within the image, visual media content displayed on a display of the external display device.

[0173]

[0187] Aspect 20: The apparatus of aspect 18 or 19, wherein to generate the output, the one or more processors are configured to generate content, the content virtually extending the display of the external display device.

[0174]

[0188] Aspect 21: The device of any one of aspects 1 to 20, wherein to generate the output, one or more processors are configured to generate content, at least in part, by overlaying virtual content over a region of the image, the region of the image being based on a position of an object within the environment.

[0175]

[0189] Aspect 22: The apparatus of aspect 21, wherein the object is a display of an external display device and the region of the image is adjacent to a representation of the display of the external display device in the image.

[0176]

[0190] Aspect 23: The device of aspect 21, wherein the object is a hand of a user of the device, the hand being at least partially adjacent to the region of the image.

[0177]

[0191] Example 24: An apparatus described in any one of examples 1 to 21, wherein the object is visual content displayed on a display.

[0178]

[0192] Example 25: An apparatus described in any one of examples 1 to 21, wherein the object is the head of a user of the apparatus.

[0179]

[0193] Example 26: An apparatus according to any one of examples 1 to 21, wherein the object is the body of a user of the apparatus.

[0180]

[0194] Aspect 27: The apparatus of any one of aspects 1 to 26, wherein the one or more processors are further configured to generate a merged dataset in response to detecting the state by at least combining data from the data stream with an image captured by the image sensor, and determining the position of the object is based at least in part on the merged dataset.

[0181]

[0195] Embodiment 28: An apparatus described in any one of embodiments 1 to 27, wherein the apparatus is a head mounted display (HMD).

[0182]

[0196] Aspect 29: The apparatus of any one of aspects 1 to 28, further comprising an audio output device, wherein to generate the output, the one or more processors are configured to generate content, and the one or more processors are configured to send content to be played to the audio output device.

[0183]

[0197] Aspect 30: A method for processing image data, the method including: receiving an image of a portion of the environment captured by an image sensor, the environment including an object; identifying, by the device, a data stream from an external device; detecting a condition based on at least one of the image, the data stream, and an operational state of the device; in response to detecting the condition, determining a position of the object in the environment based on at least one of the image and the data stream; and generating an output based on the position of the object in the environment.

[0184]

[0198] Aspect 31: The method of aspect 30, wherein detecting a state based on the image includes determining that an object is missing from a portion of an environment in the image.

[0185]

[0199] Aspect 32: The method of aspect 31, wherein determining that the object is missing from a portion of the environment in the image includes determining that at least a portion of the object is occluded in the image.

[0186]

[0200] Aspect 33: The method of aspect 31 or 32, wherein the external device includes a second image sensor, the data stream includes a second image of a second portion of the environment, and determining the position of the object within the environment is based at least in part on the depiction of the object in the second image.

[0187]

[0201] Embodiment 34: The method of embodiment 33, wherein the part of the environment and the second part of the environment overlap.

[0188]

[0202] Aspect 35: The method of any one of aspects 30 to 34, wherein detecting a state based on an operational state of the device includes determining that resource availability falls below a threshold.

[0189]

[0203] Aspect 36: The method of aspect 35, wherein determining that resource availability falls below a threshold includes determining that a battery level of the battery falls below a battery level threshold.

[0190]

[0204] Aspect 37: The method of aspect 35 or 36, wherein determining that resource availability is below a threshold comprises determining that the available bandwidth is below a bandwidth threshold.

[0191]

[0205] Aspect 38: The method of any one of aspects 30 to 37, wherein detecting a state based on an operational state of the device includes receiving a user input corresponding to offloading processing to an external device.

[0192]

[0206] Aspect 39: The method of any one of aspects 30 to 38, wherein generating the output includes generating content.

[0193]

[0207] Aspect 40: The method of aspect 39, further comprising outputting content based on a position of an object within the environment.

[0194]

[0208] Aspect 41: The method of aspect 40, wherein outputting the content includes transmitting the content to a display of the device to be displayed.

[0195]

[0209] Aspect 42: The method of any one of aspects 30 to 41, further comprising: detecting an additional condition based on at least one of an additional image captured by the image sensor, a data stream, and an operational state of the device; and performing a function previously performed by the external device in response to detecting the additional condition.

[0196]

[0210] Embodiment 43: The method of any one of embodiments 30 to 42, wherein generating the output includes controlling a device based on user input.

[0197]

[0211] Example 44: A method according to any one of examples 30 to 43, wherein detecting a condition based on an image includes determining one or more lighting conditions within the image.

[0198]

[0212] Aspect 45: The method of aspect 44, wherein determining one or more lighting conditions in the image includes determining that one or more light values ​​of the image are below an illumination threshold.

[0199]

[0213] Aspect 46: A method according to any one of aspects 30 to 45, wherein determining the position of the object within the environment includes sending a request to an external device identifying the position of the object within the environment, and receiving a response from the external device identifying the position of the object within the environment.

[0200]

[0214] Embodiment 47: The method of any one of embodiments 30 to 46, wherein the object is a display of an external display device.

[0201]

[0215] Aspect 48: The method of aspect 47, wherein detecting a state based on the image includes identifying within the image visual media content displayed on a display of the external display device.

[0202]

[0216] Aspect 49: The method of aspect 47 or 48, wherein generating the output includes generating content, the content virtually extending the display of the external display device.

[0203]

[0217] Aspect 50: A method described in any one of aspects 30 to 49, wherein generating the output includes generating content, at least in part, by overlaying virtual content over a region of the image, the region of the image being based on a position of an object within the environment.

[0204]

[0218] Aspect 51: The method of aspect 50, wherein the object is a display of an external display device and a region of the image is adjacent to a representation of the display of the external display device in the image.

[0205]

[0219] Aspect 52: The method of aspect 50, wherein the object is the hand of a user of the device and the hand is at least partially adjacent to a region of the image.

[0206]

[0220] Embodiment 53: A method according to any one of embodiments 30 to 50, wherein the object is visual content displayed on a display.

[0207]

[0221] Embodiment 54: A method according to any one of embodiments 30 to 50, wherein the object is the head of a user of the device.

[0208]

[0222] Aspect 55: A method described in any one of aspects 30 to 50, wherein the object is the body of a user of the device.

[0209]

[0223] Aspect 56: A method according to any one of aspects 30 to 55, further comprising, in response to detecting the state, generating a merged dataset by combining at least data from the data stream with an image captured by the image sensor, and determining the position of the object is based at least in part on the merged dataset.

[0210]

[0224] Aspect 57: The method of any one of aspects 30 to 56, wherein generating the output includes generating content and further includes sending the content to be played to an audio output device.

[0211]

[0225] Aspect 58: A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations recited in any one of aspects 1 to 57.

[0212]

[0226] Example 59: An apparatus comprising means for performing the operations described in any one of examples 1 to 57.

[0213]

[0227] Aspect 60: An apparatus for processing image data, the apparatus comprising at least one memory and one or more processors coupled to the memory. The one or more processors are configured to receive an image of a portion of the environment, the portion of the environment including an object, captured by an image sensor, detect a condition related to resource availability, and in response to detecting the condition, determine a position of at least a portion of the object in the environment based on at least a data stream from the device, and output content based on the position of at least a portion of the object in the environment.

[0214]

[0228] Aspect 61: The apparatus of aspect 60, wherein to detect the condition, the one or more processors are configured to determine that resource availability falls below a threshold.

[0215]

[0229] Aspect 62: The apparatus of aspect 61, wherein to determine that resource availability falls below a threshold, the one or more processors are configured to determine that a battery level of the battery falls below a battery level threshold.

[0216]

[0230] Aspect 63: The apparatus of aspect 61 or 62, wherein to determine that resource availability falls below a threshold, the one or more processors are configured to determine that the available bandwidth falls below a bandwidth threshold.

[0217]

[0231] Aspect 64: An apparatus described in any one of aspects 60 to 63, wherein to determine a position of at least a portion of an object within the environment, one or more processors are configured to send a request to the device identifying a position of at least a portion of the object within the environment and receive a response from the device identifying a position of at least a portion of the object within the environment.

[0218]

[0232] Aspect 65: An apparatus described in any one of aspects 60 to 64, wherein the one or more processors are further configured to generate content, at least in part, by overlaying virtual content over a region of the image, the region of the image being based on a position of at least a portion of an object within the environment.

[0219]

[0233] Aspect 66: The device of aspect 65, wherein the object is a hand of a user of the device, the hand being at least partially adjacent to the region of the image.

[0220]

[0234] Aspect 67: An apparatus described in any one of aspects 60 to 66, wherein the one or more processors are further configured to generate a merged dataset by merging at least data from the data stream with an image captured by the image sensor in response to detecting the state, and determining a position of at least a portion of the object is based on the merged dataset.

[0221]

[0235] Embodiment 68: An apparatus described in any one of embodiments 60 to 67, wherein the apparatus is a head mounted display (HMD).

[0222]

[0236] Aspect 69: The device of any one of aspects 60 to 68, further comprising a display, wherein to output the content, the one or more processors are configured to transmit content to the display for display by the display.

[0223]

[0237] Aspect 70: The apparatus of any one of aspects 60 to 69, further comprising an audio output device, wherein to output the content, the one or more processors are configured to transmit content to the audio output device to be played by the audio output device.

[0224]

[0238] Aspect 71: A method for processing image data, comprising the operation described in any one of aspects 60 to 70.

[0225]

[0239] Aspect 72: A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations described in any one of aspects 60 to 70.

[0226]

[0240] Example 73: An apparatus comprising means for performing the operations described in any one of examples 60 to 70.

[0227]

[0241] Aspect 74: An apparatus for processing image data, the apparatus comprising at least one memory and one or more processors coupled to the memory, the one or more processors configured to receive an image of a portion of the environment captured by an image sensor, the portion of the environment including an object, detect a condition based on the image, generate content based on at least a data stream from the device in response to detecting the condition, and output the content based on a position of at least a portion of the object in the environment.

[0228]

[0242] Aspect 75: The apparatus of aspect 74, wherein to detect the condition, the one or more processors are configured to determine that an object is missing from a portion of the environment in the image.

[0229]

[0243] Aspect 76: The apparatus of aspect 74, wherein the object is a display of an external device.

[0230]

[0244] Aspect 77: The apparatus of aspect 76, wherein to detect the state, the one or more processors are configured to identify, within the image, a depiction of visual media content displayed on a display of the external device.

[0231]

[0245] Aspect 78: The device of aspect 76, wherein to detect the condition, the one or more processors are configured to detect the presence of a display in the vicinity of the device.

[0232]

[0246] Aspect 79: The device described in aspect 76, wherein the one or more processors are further configured to generate a directional indicator indicating a position of the display.

[0233]

[0247] Aspect 80: The apparatus of any one of aspects 76 to 79, wherein the content virtually extends the display of the external device.

[0234]

[0248] Aspect 81: An apparatus described in any one of aspects 74 to 80, wherein one or more processors are configured to generate content, at least in part, by overlaying virtual content over a region of an image, the region of the image being based on a position of at least a portion of an object within the environment.

[0235]

[0249] Aspect 82: The apparatus of aspect 81, wherein the object is a display of an external device and the region of the image is adjacent to a representation of the display of the external device in the image.

[0236]

[0250] Aspect 83: An apparatus described in any one of aspects 74 to 82, wherein one or more processors are configured to generate a merged dataset by at least merging data from the data stream with an image captured by the image sensor in response to detecting a state, and content is generated based on the merged dataset.

[0237]

[0251] Embodiment 84: An apparatus described in any one of embodiments 74 to 83, wherein the apparatus is a head mounted display (HMD).

[0238]

[0252] Aspect 85: The apparatus of any one of aspects 74 to 84, further comprising a display, wherein to output the content, the one or more processors are configured to send content to the display for display by the display.

[0239]

[0253] Aspect 86: The apparatus of any one of aspects 74 to 85, further comprising an audio output device, wherein to output the content, the one or more processors are configured to transmit content to be played by the audio output device to the audio output device.

[0240]

[0254] Aspect 87: A method for processing image data, comprising the operation described in any one of aspects 74 to 86.

[0241]

[0255] Aspect 88: A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations described in any one of aspects 74 to 86.

[0242]

[0256] Example 89: An apparatus comprising means for performing the operations described in any one of examples 74 to 86.

Claims

1. 1. An apparatus for processing image data, comprising: At least one memory; at least one processor coupled to the at least one memory, receiving first image data of an environment captured by a first image sensor, wherein the environment includes an object; receiving second image data of the environment from an external device, the second image data being captured by a second image sensor of the external device; Detecting a transition from the object being within a field of view of the first image sensor to the object being no longer within the field of view of the first image sensor based on at least the first image data; responsive to detecting the transition, determining a position of the object within the environment based on at least the second image data; generating an output based on the position of the object within the environment; At least one processor configured to An apparatus comprising:

2. A method for processing image data in at least one processor, comprising: receiving a first image of an environment captured by a first image sensor, wherein the environment includes an object; receiving second image data of the environment from an external device, the second image data being captured by a second image sensor of the external device; detecting a transition from the object being within a field of view of the first image sensor to the object being no longer within the field of view of the first image sensor based on at least the first image; responsive to detecting the transition, determining a position of the object within the environment based on at least the second image data; generating an output based on the position of the object within the environment; A method comprising:

3. For detecting the transition, the at least one processor is configured to determine that the object is missing from a first portion of the environment shown in the first image data, and for determining that the object is missing from the first portion of the environment in the image, the at least one processor is configured to determine that at least a portion of the object is occluded in the first image data; or the second image data illustrating a second portion of the environment, and determining the location of the object within the environment is based at least in part on a depiction of the object in the second image data, and the first portion of the environment and the second portion of the environment overlap.

3. An apparatus according to claim 1 or a method according to claim 2.

4. To detect the transition, the at least one processor is configured to determine when resource availability falls below a threshold, and to determine when the availability of the resource falls below the threshold, the one or more processors are configured to determine when a battery level of a battery falls below a battery level threshold or determine when available bandwidth falls below a bandwidth threshold.

3. An apparatus according to claim 1 or a method according to claim 2.

5. The apparatus of claim 1 or the method of claim 2, wherein the at least one processor is configured to receive user input corresponding to offload processing to the external device to detect the transition.

6. Preferably, the device of claim 1 further comprises a display, or the method of claim 2, To generate the output, the at least one processor is configured to generate content, preferably to output the content based on the position of the object within the environment; To output the content, the at least one processor is configured to transmit the content to the display for display.

3. An apparatus according to claim 1 or a method according to claim 2.

7. The at least one processor: Detecting a state based on at least one of additional image data captured by the first image sensor, the second image data, or an operational state of the device; responsive to detecting the condition, performing, by the at least one processor, a function previously configured to be performed by the external device. It is configured as follows:

3. An apparatus according to claim 1 or a method according to claim 2.

8. The at least one processor, 3. The apparatus of claim 1 or the method of claim 2, configured to control the apparatus based on user input to generate the output.

9. The method of claim 8, wherein the at least one processor is configured to determine one or more lighting conditions within the first image data to detect the transition; The apparatus of claim 1 or the method of claim 2, wherein to determine the one or more lighting conditions in the image, the one or more processors are preferably configured to determine that one or more light values ​​of the image fall below an illumination threshold.

10. The at least one processor, sending a request to the external device identifying the location of the object within the environment; receiving a response from the external device identifying the location of the object within the environment to determine the location of the object within the environment; It is configured as follows:

3. An apparatus according to claim 1 or a method according to claim 2.

11. The object is a display of an external display device, preferably the at least one processor is configured to identify visual media content displayed on the display of the external display device to detect the transition within the first image data; or the at least one processor is configured to generate content that virtually extends the display of the external display device to generate the output.

3. An apparatus according to claim 1 or a method according to claim 2.

12. The at least one processor, configured to overlay virtual content over a region of the first image data to generate the output, the region of the first image data being based on the position of the object within the environment; Preferably, the object is a display of an external display device and the region of the first image data is adjacent to a representation of the display of the external display device in the first image data, or preferably, the object is a hand of a user of the apparatus and the hand is at least partially adjacent to the region of the first image data.

3. An apparatus according to claim 1 or a method according to claim 2.

13. The at least one processor: responsive to detecting the transition, combining second image data with the first image data to generate a merged data set, and determining the location of the object based on the merged data set to determine the location of the object. The method is further configured as follows:

3. An apparatus according to claim 1 or a method according to claim 2.

14. The device of claim 1 , wherein the device is a head mounted display (HMD).

15. 3. The apparatus of claim 1 further comprising an audio output device, the method comprising: the at least one processor is configured to generate content and to transmit the content to be played to the audio output device; 3. An apparatus according to claim 1 or a method according to claim 2.