Processing image data in augmented reality system

By identifying and encoding regions of interest and non-regions of interest in images within an extended reality system, the problem of bandwidth and power waste is solved, enabling efficient image data transmission and processing.

CN121128170APending Publication Date: 2025-12-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480028481.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-01
Filing Date
2024-03-12
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In extended reality systems with a split architecture, existing technologies struggle to effectively limit the size of image data transmission, leading to excessive bandwidth consumption and power waste. Furthermore, synchronizing and processing image data from multiple regions of interest simultaneously incurs overhead and carries the risk of data loss.

Method used

By identifying the region of interest (ROI) in an XR device and encoding the ROI and non-ROI regions using different encoding parameters, the encoded data can be sent to reduce bandwidth consumption. Specific methods include using higher quantization parameters to encode the non-ROI regions.

Benefits of technology

This effectively reduces the transmission bandwidth of image data, saves power consumption, ensures the data integrity and synchronization of the region of interest, and reduces the risk of data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128170A_ABST
    Figure CN121128170A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing image data are described herein. For example, a method for processing image data is provided. The method may include capturing an image; determining at least one region of interest of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encoding a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one region of interest; encoding a second portion of the graph according to a second parameter to generate second encoded data; and transmitting the first encoded data and the second encoded data to a computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates throughout to processing image data in extended reality systems. For example, aspects of this disclosure include systems and techniques for capturing image data at a first device, encoding the image data for transmission to a second device, and transmitting the encoded data to the second device. Some aspects relate to receiving the transmitted encoded data and processing the encoded data at the second device. Background Technology

[0002] Extended reality (XR) systems (e.g., virtual reality (VR), augmented reality (AR), and / or mixed reality (MR)) can provide a virtual experience to a user by displaying virtual content that fills most or all of the user’s field of view at a display or by overlaying virtual content onto or next to the user’s real-world field of view (e.g., using a see-through or pass-through display).

[0003] XR systems typically include a display (e.g., a head-mounted display (HMD) or smart glasses), an image capture device located close to the display, and a processing device. In such XR systems, the image capture device captures an image that indicates the user's field of view, the processing device generates virtual content based on the user's field of view, and the display shows the virtual content within the user's field of view.

[0004] In some XR systems (e.g., split-architecture XR systems), the processing device may be separate from the display and / or image capture device. For example, the processing device may be part of an accessory device (e.g., a smartphone, tablet, laptop, personal computer, or server), while the display and image capture device may be part of an XR device (such as an HMD, smart glasses, or other types of device).

[0005] In this type of split-architecture XR system, the XR device can send image data (captured by the image capture device) to a companion device, and the companion device can determine or generate virtual content data based on the image data. The companion device can then send the virtual content data to the XR device for display on a monitor.

[0006] It may be desirable to limit the size of image data sent from the XR device to the companion device. Limiting the size of the data sent saves bandwidth available for communication between the XR device and the companion device. Bandwidth can be measured in bit rate, which is the number of bits that can be sent during a given period of time (e.g., bits per second). Saving bandwidth can save power (e.g., by sending less data) and / or allow the saved bandwidth to be used to send other data. Summary of the Invention

[0007] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description presented below.

[0008] Systems and techniques for processing image data are described. According to at least one example, a method for processing image data is provided. The method includes: capturing an image; determining at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encoding a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one ROI; encoding a second portion of the image according to a second parameter to generate second encoded data; and transmitting the first encoded data and the second encoded data to a computing device.

[0009] In another example, an apparatus for processing image data is provided, the apparatus including at least one memory and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: cause an image capture device to capture an image; determine at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encode a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one ROI; encode a second portion of the image according to a second parameter to generate second encoded data; and cause at least one transmitter to transmit the first encoded data and the second encoded data to a computing device. In some cases, the apparatus includes the image capture device to capture the image. In some cases, the apparatus includes the at least one transmitter to transmit the first encoded data and the second encoded data to a computing device.

[0010] In another example, a non-transitory computer-readable medium is provided, on which instructions are stored, which, when executed by one or more processors, cause the one or more processors to: configure the at least one processor to: cause an image capturing device to capture an image; determine at least one region of interest in the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encode a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one region of interest; encode a second portion of the image according to a second parameter to generate second encoded data; and cause at least one transmitter to transmit the first encoded data and the second encoded data to a computing device.

[0011] In another example, an apparatus for processing image data is provided. The apparatus includes: components for capturing an image; components for determining at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; components for encoding a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one ROI; components for encoding a second portion of the image according to a second parameter to generate second encoded data; and components for transmitting the first encoded data and the second encoded data to a computing device.

[0012] Systems and techniques for processing image data are described. According to at least one example, a method for processing image data is provided. The method includes: receiving, at a computing device, first data encoded in a first image from an image capture device; determining at least one region of interest (ROI) of the first image; sending an indication of the at least one ROI from the computing device to the image capture device; receiving second data encoded in a second image, a first portion of the second image encoded according to a first parameter corresponding to the at least one ROI, and a second portion of the second image encoded according to a second parameter; decoding the second data to generate a reconstructed instance of the second image; and tracking objects in the reconstructed instance of the second image.

[0013] In another example, an apparatus for processing image data is provided, the apparatus including at least one memory and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: receive first data encoded in a first image from an image capture device; determine at least one region of interest (ROI) of the first image; send an indication of the at least one ROI to the image capture device; receive second data encoded in a second image, a first portion of the second image encoded according to a first parameter corresponding to the at least one ROI, and a second portion of the second image encoded according to a second parameter; decode the second data to generate a reconstructed instance of the second image; and track objects in the reconstructed instance of the second image. In some cases, the apparatus includes the at least one transmitter to send the indication of the at least one ROI to the image capture device.

[0014] In another example, a non-transitory computer-readable medium is provided, on which instructions are stored, which, when executed by one or more processors, cause the one or more processors to: receive first data encoding a first image from an image capture device; determine at least one region of interest (ROI) of the first image; send an indication of the at least one ROI to the image capture device; receive second data encoding a second image, a first portion of which is encoded according to a first parameter corresponding to the at least one ROI, and a second portion of which is encoded according to a second parameter; decode the second data to generate a reconstructed instance of the second image; and track objects in the reconstructed instance of the second image.

[0015] In another example, an apparatus for processing image data is provided. The apparatus includes: components for receiving first data encoded in a first image from an image capture device at a computing device; components for determining at least one region of interest (ROI) of the first image; components for sending an indication of the at least one ROI from the computing device to the image capture device; components for receiving second data encoded in a second image, wherein a first portion of the second image is encoded according to a first parameter corresponding to the at least one ROI, and a second portion of the second image is encoded according to a second parameter; components for decoding the second data to generate a reconstructed instance of the second image; and components for tracking objects in the reconstructed instance of the second image.

[0016] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles (or computing devices or systems of vehicles), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, television sets (e.g., network-connected television sets), robotic devices or systems, or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.

[0017] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0018] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0019] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0020] Figure 1 These are illustrations of examples of extended reality (XR) systems according to various aspects of this disclosure;

[0021] Figure 2 This is a diagram illustrating the architecture of an example XR system according to some aspects of this disclosure;

[0022] Figure 3 This is a block diagram illustrating another example XR system according to various aspects of this disclosure;

[0023] Figure 4 These are illustrations of example images that can be processed according to various aspects of this disclosure;

[0024] Figure 5 This is an illustration of another example image that can be processed according to various aspects of this disclosure;

[0025] Figure 6 This is yet another example image illustrating how it can be processed according to various aspects of this disclosure;

[0026] Figure 7 This is a flowchart illustrating an example process for processing image data according to various aspects of this disclosure;

[0027] Figure 8 This is a flowchart illustrating another example process for processing image data according to various aspects of this disclosure;

[0028] Figure 9 Example computing device architectures are illustrated, showing example computing devices that can implement the various technologies described herein. Detailed Implementation

[0029] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0030] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0031] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0032] Some extended reality (XR) systems may employ computer vision and / or perception processes, which may include detection algorithms, recognition algorithms, and / or tracking algorithms. For example, a computer vision process may receive an image, detect (and / or recognize) real-world objects in the image (e.g., people, hands, vehicles, etc.), and track real-world objects in the image.

[0033] In some cases, when an XR system (including a discrete architecture XR system) implements computer vision and / or perception processes, most or all of these processes are implemented at an accessory device of the XR system, rather than within the XR device itself. For example, an XR device may capture images and provide those images to an accessory device that implements detection, recognition, and / or tracking algorithms.

[0034] As mentioned earlier, in a split-architecture XR system, it may be desirable to limit the size of image data sent from the XR device to the companion device, for example, to limit the power consumption of the XR device and / or save bandwidth for other purposes.

[0035] Detection and / or recognition algorithms can operate on the entire image to detect real-world objects within it. Tracking algorithms can focus on (or only need) the portion of the image representing a real-world object. For example, a tracking algorithm can operate using pixels that include the bounding box of the real-world object to be tracked, and does not need to find the entire image in which the bounding box is located. In this disclosure, the term "bounding box" can refer to a plurality of image pixels surrounding and including an object represented in an image pixel. Object detection or object tracking algorithms can define a bounding box around an object.

[0036] One solution to limit the size of transmission from an XR device to a companion device in an XR system involves identifying regions of interest (ROIs) within a captured image (for example, one or more bounding boxes may be identified as corresponding ROIs). The XR device can transmit image data for the ROIs but not for the non-ROI portions of the image. One problem with this type of solution is that it may require transmitting separate image data for each of several individual ROIs. In cases where the captured image is part of a series of images (e.g., frames of a video), it will be necessary to synchronize the transmission of such separate image data to ensure that the relative timing of the ROI data is synchronized across the series of images (e.g., because ROIs across frames may arrive or be processed out of order, for example, based on transmission and / or processing delays). Such solutions may also require additional overhead in transmission (e.g., establishing a separate data stream for each ROI) and / or processing (e.g., requiring separate encoding / decoding sessions). Additionally, such solutions may allow multiple (e.g., two, three, etc.) transmissions of portions of an image (e.g., when portions of two or more ROIs partially overlap). Furthermore, in cases where the identification of a ROI is lost, such solutions may result in the lost ROIs becoming unavailable for tracking.

[0037] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Technologies”) for processing image data in an XR system. The systems and technologies described herein may include an XR device comprising an image capture device capable of capturing images. The XR device may determine one or more regions of interest within an image and may encode the image into encoded data for transmission. When encoding the image into encoded data, the XR device may use parameters different from those used by the XR device to encode one or more regions of interest in the image (e.g., using a higher quantization parameter (QP)) to encode non-regions of interest (e.g., portions of the image not included in the regions of interest). By using different parameters (e.g., using a larger QP) to encode the non-regions of interest, the systems and technologies can encode them using fewer bits than when using the parameters used to encode the regions of interest. Encoding image data with fewer bits saves bandwidth when transmitting the encoded data.

[0038] For example, the portion of coded data representing a non-region of interest (NROI) might be more densely encoded per bit than the portion representing a ROI (e.g., representing more pixels with fewer bits). The terms "data density," "bit density," and similar terms refer to the number of image pixels represented by each bit of coded data used to represent the number of pixels. More dense encoding can lead to loss of image data when reconstructing an image from the coded data. However, by selecting the NROI portion of the image for more dense encoding, object detection, object recognition, and / or object tracking operations may not be affected by this loss of image data.

[0039] As an example of data density, first data representing an image encoded using a first QP may have a first density. Second data representing the same image encoded using a second QP (e.g., twice the size of the first QP) may have a second density greater than (e.g., twice the size of the first density) because the second data may be smaller than the first data (e.g., half the size of the first data), while still encoding the same image (albeit at a lower image quality). In another example, half of an image (e.g., a portion of the image including the region of interest) may be encoded using the first QP, while the other half of the image (e.g., the portion of the image not containing the region of interest) may be encoded using the second QP. The resulting data may be denser than the first data (e.g., 150% of the density of the first data) because the resulting data may be smaller than the first data (e.g., 75% the size of the first data), while still encoding the same image (albeit at a lower image quality for half of the image).

[0040] By increasing data density, systems and technologies can use less bandwidth to transmit image data (e.g., between XR devices and companion devices in an XR system). Reducing the bandwidth used to transmit data saves power (e.g., the power of the XR device) and / or allows bandwidth to be used to transmit other data.

[0041] Various aspects of this application will be described with reference to the accompanying drawings.

[0042] Figure 1 This is a diagram illustrating an example of an extended reality (XR) system 100 according to various aspects of the present disclosure. As shown, the XR system 100 includes an XR device 102, an accessory device 104, and a communication link 106 between the XR device 102 and the accessory device 104. In some cases, the XR device 102 typically implements aspects of extended reality display, image capture, and / or view tracking, including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. In some cases, the accessory device 104 typically implements aspects of extended reality computation. For example, the XR device 102 may capture images of the environment of user 108 and provide these images to the accessory device 104 (e.g., via communication link 106). The accessory device 104 may render virtual content (e.g., in relation to the captured environmental images) and provide the virtual content to the XR device 102 (e.g., via communication link 106). The XR device 102 may display the virtual content to user 108 (e.g., within user 108's field of view 110).

[0043] Typically, XR device 102 may display virtual content to be viewed by user 108 within field of view 110. In some examples, XR device 102 may include a transparent surface (e.g., optical glass) such that virtual objects can be displayed on the transparent surface (e.g., by generating or projecting onto the transparent surface) to overlay virtual content onto real-world objects viewed through the transparent surface (e.g., in a perspective configuration). In some cases, XR device 102 may include a camera and may display real-world objects (e.g., as frames or images captured by the camera) and virtual objects overlaid on the displayed real-world objects (e.g., in a pass-through configuration). In various examples, XR device 102 may include virtual reality headsets, smart glasses, real-time feed video cameras, GPUs, one or more sensors (e.g., such as one or more inertial measurement units (IMUs), image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, displays, smart glasses, etc.), etc.

[0044] The companion device 104 can render virtual content to be displayed by the companion device 104. In some examples, the companion device 104 may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device or a mobile device acting as a server device), any other computing device and / or combinations thereof.

[0045] The communication link 106 can be a wired or wireless connection according to any suitable wireless protocol, such as, for example, Universal Serial Bus (USB), Ultra Wideband (UWB), Institute of Electrical and Electronics Engineers (IEEE) 802.11 (Wi-Fi), IEEE 802.15, or Bluetooth. ® In some cases, communication link 106 may be a direct wireless connection between XR device 102 and companion device 104. In other cases, communication link 106 may be via one or more intermediate devices, such as routers or switches and / or across a network.

[0046] Depending on various aspects, XR device 102 can capture images and provide the captured images to accessory device 104. Accessory device 104 can implement detection, recognition, and / or tracking algorithms based on the captured images.

[0047] Figure 2 This is a schematic diagram illustrating the architecture of an example extended reality (XR) system 200 according to some aspects of this disclosure. The XR system 200 can execute XR applications and implement XR operations.

[0048] In this exemplary example, the XR system 200 includes one or more image sensors 202, accelerometers 204, gyroscopes 206, storage devices 208, input devices 207, displays 212, computing components 214, XR engines 224, image processing engines 226, rendering engines 228, and communication engines 230. It should be noted that... Figure 2 The components 202 to 230 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include those with... Figure 2 The components shown may be more, fewer, or different than those shown. For example, in some cases, the XR system 200 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2One or more other software and / or hardware components not shown. While various components of the XR system 200 (such as image sensor 202) may be referred to herein in the singular, it should be understood that the XR system 200 may include multiple components of any of the components discussed herein (e.g., multiple image sensors 202).

[0049] Display 212 may be or may include glass, screen, lens, projector and / or other display mechanisms that allow users to see a real-world environment and also allow XR content to be overlaid, superimposed, blended or otherwise displayed on it.

[0050] XR system 200 may include or may communicate (wired or wirelessly) with input device 210. Input device 210 may include any suitable input device, such as a touchscreen, pen or other pointing device, keyboard, mouse, buttons or keys, microphone for receiving voice commands, gesture input device for receiving gesture commands, video game controller, steering wheel, joystick, set of buttons, trackball, remote control, any other input device discussed herein, or any combination thereof. In some cases, image sensor 202 may capture images that can be processed to interpret gesture commands.

[0051] The XR system 200 can also communicate with one or more other electronic devices (wired or wireless). For example, the communication engine 230 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 230 may correspond to... Figure 9 The communication interface is 926.

[0052] In some embodiments, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be part of the same device. For example, in some cases, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be integrated into an HMD, extended reality glasses, a smartphone, a laptop, a tablet, a gaming system, and / or any other computing device. However, in some embodiments, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be part of two or more separate computing devices. For example, in some cases, some of the components 202-230 may be part of or implemented by a computing device, and the remaining components may be part of or implemented by one or more other computing devices. For example, in a discrete sensing XR system, XR system 200 may include a first device (e.g., an XR device such as...). Figure 1 The first device (XR device 102) includes a display 212, an image sensor 202, an accelerometer 204, a gyroscope 206, and / or one or more computing components 214. The XR system 200 may also include a second device that includes additional computing components 214 (e.g., implementing an XR engine 224, an image processing engine 226, a rendering engine 228, and / or a communication engine 230). In such examples, the second device may generate virtual content based on information or data (e.g., images, sensor data, such as measurements from the accelerometer 204 and gyroscope 206) and may provide this virtual content to the first device for display on the first device. The second device may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server equipment (e.g., an edge or cloud-based server, a personal computer acting as a server equipment, or a mobile device acting as a server equipment), any other computing device, and / or combinations thereof.

[0053] Storage device 208 can be any storage device used for storing data. Furthermore, storage device 208 can store data from any component of the XR system 200. For example, storage device 208 can store data from image sensor 202 (e.g., image or video data), data from accelerometer 204 (e.g., measurement results), data from gyroscope 206 (e.g., measurement results), data from computing component 214 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 224, data from image processing engine 226, and / or data from rendering engine 228 (e.g., output frames). In some examples, storage device 208 may include a buffer for storing frames processed by computing component 214.

[0054] Computing component 214 may be or may include a central processing unit (CPU) 216, a graphics processing unit (GPU) 218, a digital signal processor (DSP) 220, an image signal processor (ISP) 222, and / or other processors (e.g., a neural network processing unit (NPU) implementing one or more trained neural networks). Computing component 214 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, localization, pose estimation, mapping, content anchoring, content rendering, prediction, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, computing component 214 may implement (e.g., control, operate, etc.) an XR engine 224, an image processing engine 226, and a rendering engine 228. In other examples, computing component 214 may also implement one or more other processing engines.

[0055] Image sensor 202 may include any image and / or video sensor or capture device. In some examples, image sensor 202 may be part of a multi-camera assembly, such as a dual-camera assembly. Image sensor 202 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by computing component 214, XR engine 224, image processing engine 226, and / or rendering engine 228, as described herein.

[0056] In some examples, image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on that image data and / or may provide the image data or frame to XR engine 224, image processing engine 226, and / or rendering engine 228 for processing. The image or frame may include a video frame in a video sequence or a still image. The image or frame may include an array of pixels representing a scene. For example, the image may be: a red-green-blue (RGB) image with red, green, and blue color components per pixel; a lightness, redness, and blueness (YCbCr) image with a lightness component and two chromaticity (redness and blueness) components per pixel; or any other suitable type of color or monochrome image.

[0057] In some cases, image sensor 202 (and / or other cameras of XR system 200) may also be configured to capture depth information. For example, in some implementations, image sensor 202 (and / or other cameras) may include an RGB depth (RGB-D) camera. In some cases, XR system 200 may include one or more depth sensors (not shown) that are separate from image sensor 202 (and / or other cameras) and capable of capturing depth information. For example, such depth sensors may acquire depth information independently of image sensor 202. In some examples, depth sensors may be physically mounted in the same approximate location or position as image sensor 202, but may operate at a different frequency or frame rate than image sensor 202. In some examples, depth sensors may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Depth information can then be obtained by utilizing the geometric deformation of the projected pattern caused by the surface shape of the objects. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).

[0058] The XR system 200 may also include other sensors among its one or more sensors. These one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. These one or more sensors may provide velocity, orientation, and / or other positioning-related information to the computing component 214. For example, accelerometer 204 may detect the acceleration of the XR system 200 and may generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that can be used to determine the positioning or attitude of the XR system 200. Gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 224 may use measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or measurements obtained by the gyroscope 206 (e.g., one or more rotation vectors) to calculate the attitude of the XR system 200. As previously mentioned, in other examples, the XR system 200 may also include other sensors such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye-tracking sensor, a machine vision sensor, a smart scene sensor, a voice recognition sensor, a shock sensor, a vibration sensor, a positioning sensor, a tilt sensor, etc.

[0059] As noted above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure the specific force, angular velocity, and / or orientation of the XR system 200. In some examples, the one or more sensors may output measured information associated with the capture of an image by the image sensor 202 (and / or other cameras of the XR system 200) and / or depth information obtained using one or more depth sensors of the XR system 200.

[0060] XR engine 224 can use the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs and / or other sensors) to determine the attitude of XR system 200 (also referred to as head attitude) and / or the attitude of image sensor 202 (or other cameras of XR system 200). In some cases, the attitude of XR system 200 and the attitude of image sensor 202 (or other cameras) can be the same. The attitude of image sensor 202 refers to the attitude of image sensor 202 relative to a reference frame (e.g., about). Figure 1The camera attitude is determined by the positioning and orientation of the field of view (110). In some implementations, the camera attitude can be determined with respect to 6 degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame (such as the image plane)) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera attitude can be determined with respect to 3 degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).

[0061] In some cases, a device tracker (not shown) may use measurements from one or more sensors and image data from image sensor 202 to track the pose (e.g., 6DoF pose) of the XR system 200. For example, the device tracker may fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the XR system 200 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map for that scene. 3D map updates may include, for example, but not limited to, new or updated features and / or features or landmarks associated with the scene and / or the 3D map of the scene, positioning updates identifying or updating the position of the XR system 200 within the scene and the 3D map of the scene, etc. The 3D map provides a digital representation of the scene in the real / physical world. In some examples, 3D maps can anchor location-based objects and / or content to real-world coordinates and / or objects. XR system 200 can use mapped scenes (e.g., scenes in the physical world represented by and / or associated with 3D maps) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0062] In some aspects, computing component 214 may use a visual tracking solution to determine and / or track the pose of image sensor 202 and / or the XR system 200 as a whole, based on images captured by image sensor 202 (and / or other cameras of XR system 200). For example, in some examples, computing component 214 may perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, computing component 214 may perform SLAM or may communicate (wired or wirelessly) with a SLAM system (not shown). SLAM refers to a class of techniques that simultaneously track the pose of a camera (e.g., image sensor 202) and / or XR system 200 relative to an environment (e.g., a map of the environment modeled by XR system 200) while creating a map of the environment. This map may be called a SLAM map and may be three-dimensional (3D). SLAM technology can be performed using color or grayscale image data captured by image sensor 202 (and / or other cameras of XR system 200) and can be used to generate estimates of 6DoF attitude measurements of image sensor 202 and / or XR system 200. Such SLAM technology configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated attitude.

[0063] Figure 3 This is a block diagram illustrating an example extended reality (XR) system 300 according to various aspects of this disclosure. The XR system 300 may include an XR device 302 and an accessory device 322. The XR device 302 may be a head-mounted device (e.g., an HMD, smart glasses, etc.). Figure 1 Example of XR device 102. Companion device 322 may be a computing device, may be included in a computing device, or may be implemented in a computing device, such as a mobile phone, tablet, laptop, personal computer, server, computing system of a vehicle, or other computing device. Companion device 322 may be... Figure 1 Example of supporting equipment 104.

[0064] XR device 302 includes an image capturing device 304 capable of capturing one or more images 306 (e.g., the image capturing device can continuously capture images 306). Images 306 may be or may include single-view images (e.g., monocular images) or multi-view images (e.g., stereo pairs). Images 306 may include one or more regions of interest (ROIs) 308 and one or more non-ROI portions 310. When capturing images 306, XR device 302 may or may not distinguish between ROIs 308 and non-ROI portions 310. According to a first example, XR device 302 may identify ROIs 308 (e.g., based on the user's gaze, which is based on another camera pointed at the user's eye). Figure 3 (Image captured, not illustrated in the image below). According to the second example, the accessory device 322 may identify a region of interest 308 within an image 306 using one or more techniques (described in more detail below) and provide the XR device 302 with ROI information 330 indicating the region of interest 308. The XR device 302 may parse the newly captured image 306 based on the region of interest 308 determined by the accessory device 322 based on the previously captured image 306. For example, the XR device 302 may identify pixels in the newly captured image 306 that are associated with the region of interest 308 identified based on the previously captured image 306.

[0065] XR device 302 may process image 306 at image processing engine 312. Image processing engine 312 may be circuitry or a chip (e.g., a field-programmable gate array (FPGA) or an image processor). Image processing engine 312 may filter image 306 (e.g., to remove noise), etc. In some cases, image processing engine 312 may receive ROI information 330 and apply a low-pass filter to the non-ROI region 310 of image 306. Applying a low-pass filter can remove high-frequency spatial content from image data, which allows image data to be encoded using fewer bits per pixel (e.g., by encoder 314). Applying a low-pass filter to an image may have the effect of blurring the image. Because the low-pass filter is applied to the non-ROI region 310 instead of the ROI 308, the accompanying device 322 may not be affected in its ability to detect, identify, and / or track objects in the ROI 308 of image 306.

[0066] Image processing engine 312 provides processed image data to encoder 314 (which may be a combined encoder-decoder device, also known as a codec). Encoder 314 may be a circuit or chip (e.g., an FPGA or a processor), or may be implemented in a circuit or chip. Encoder 314 may encode the processed image data for transmission (e.g., as separate data packets for sequential transmission). In one exemplary example, encoder 314 may encode the image data based on a video decoding standard such as High Efficiency Video Decoding (HEVC), Universal Video Decoding (VVC), or another video decoding standard. In another exemplary example, encoder 314 may encode the image data using a machine learning system trained to encode images (e.g., trained using supervised, semi-supervised, or self-supervised learning techniques).

[0067] Encoder 314 can receive ROI information 330 and can use different parameters (e.g., different quantization parameters (QPs)) when encoding the region of interest 308 and the non-ROI portion 310 of image 306, respectively. Encoder 314 can support quantization parameter maps with block granularity. For example, encoder 314 can use a first QP to encode the region of interest 308 and a second QP (e.g., higher than the first QP) to encode the non-ROI portion 310 of image 306. By using the second (e.g., higher) QP to encode the non-ROI portion 310 of the image data, encoder 314 can generate more dense (e.g., composed of fewer bits) encoded data than when encoding the entire content of each image 306 using the first QP. For example, because the image data is encoded using a higher QP to encode the non-ROI portion 310 of image 306, the encoded data can represent image 306 using fewer bits compared to encoding the entire content of each image 306 using the first QP. Identifying the region of interest 308 and not using a higher QP for the region of interest 308 ensures that the region of interest 308 retains its original image quality, thereby preserving the object detection, recognition, and / or tracking capabilities of the accessory device 322.

[0068] Additionally or alternatively, the image processing engine 312 or encoder 314 may apply a mask to the non-interest region 310 of image 306 before encoding the image data. Such a mask renders the non-interest region 310 as a uniform value (e.g., the average intensity of image 306). Masking the non-interest region 310 of image 306 with a uniform value allows the resulting image data to be encoded using fewer bits per pixel, for example, because a skip mode can be used to encode the uniform value.

[0069] Filtering or masking the image data can provide additional beneficial effects if the data is subsequently encoded using a different QP. For example, applying a different QP during encoding may introduce artifacts into the image (e.g., at quantization difference boundaries). Applying a low-pass filter or mask can limit or reduce such artifacts.

[0070] Alternatively or additionally, pixels of the region of interest 308 may be filled, which may reduce artificial discontinuities and / or enhance the compression gain and / or subjective quality of the region of interest 308 in the reconstructed image. Alternatively or additionally, the non-region of interest portion 310 may be intra-coded, which may reduce dynamic random access memory traffic.

[0071] In some cases, if the object to be tracked is very close to the image capture device 304, the object may occupy a large portion of the image 306. The tracker algorithm may be able to handle lower-quality images of the object (e.g., images encoded with relatively high QP and / or filtered images) because the object's features may be easily detected and / or tracked, since the object occupies a large portion of the image 306. In this case, the portion of image 306 occupied by the object can be encoded using a higher QP and / or filtered to save bandwidth.

[0072] Additionally or alternatively, the QP (and / or low-pass filter passband) may be determined based on an inverse relationship with the distance between the object represented by the region of interest 308 and the image capture device 304. The distance between the object and the image capture device 304 may be determined by the accompanying device 322 (e.g., based on a stereo image and / or a distance sensor of the accompanying device 322). As an example, the farther the object is from the image capture device 304, the lower the QP likely to be selected for encoding the region of interest 308 representing the object. As another example, the farther the object is from the image capture device 304, the larger the passband of the low-pass filter selected for filtering the region of interest 308 representing the object may be. In some cases, the QP and / or passband may be determined by the recognition and / or tracking engine 326 (e.g., enabling the detection, recognition, and / or tracking of objects in the region of interest 308 of the reconstructed image).

[0073] After encoding the image data, the XR device 302 can send the encoded data to the companion device 322 (e.g., using...). Figure 3(Communication engine not illustrated). Encoded data may include relatively few bits (e.g., low-pass filtering of image data, encoding portions of image data using relatively high quantization parameters, or masking image data). In other words, encoded data may include fewer bits compared to encoding the entire image using a low QP, without filtering, and without masking. Less bandwidth can be used to transmit encoded data comprising relatively fewer bits compared to the bandwidth required to transmit data encoded without low-pass filtering, using relatively high QP, and / or masking for portions of image data. Saving bandwidth at XR device 302 saves power at XR device 302.

[0074] The accessory device 322 can receive encoded data (e.g., using...) Figure 3 (A communication engine not illustrated) and provides encoded data to decoder 324. The lines between encoder 314 and decoder 324 are illustrated using dashed lines to indicate that communication of encoded image data between encoder 314 and decoder 324 can be wired or wireless, for example, according to any suitable communication protocol, such as USB, UWB, Wi-Fi, IEEE 902.15, or Bluetooth. ® Similarly, dashed lines are used to illustrate other lines between the XR device 302 and the accessory device 322 (including the line between the ROI information 330 and the image processing engine 312, the line between the ROI information 330 and the encoder 314, and the line between the encoder 334 and the decoder 316) to indicate that the communication represented by these lines can be wired or wireless.

[0075] Decoder 324 (which may be a codec) decodes the encoded image data. Decoder 324 may be a circuit or chip (e.g., an FPGA or a processor), or may be implemented in a circuit or chip. The decoded image data may differ from image 306. For example, based on image processing engine 312 applying a low-pass filter to the image data and / or applying a mask before encoding the image data and / or based on decoder 324 applying a different QP to the image data during encoding, the decoded image data may differ from image 306. However, based on image processing engine 312 filtering and / or masking the non-interest region 310 instead of the interest region 308, and / or based on encoder 314 using a relatively low QP when encoding the interest region 308, the interest region 308 in the decoded image data may be substantially the same as in image 306.

[0076] The recognition and / or tracking engine 326 (which may be a circuit or chip (e.g., an FPGA or a processor), or may be implemented in a circuit or chip) receives decoded image data and uses the decoded image data to perform operations related to object detection, object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, and / or other computer vision tasks. For example, the recognition and / or tracking engine 326 may identify a region of interest 308 based on object recognition techniques (e.g., identifying an object represented in image 306 and tracking the location of that object across multiple images 306). As another example, the recognition and / or tracking engine 326 may identify the region of interest 308 based on hand tracking techniques (e.g., identifying a hand as the region of interest 308 and / or using a hand as an indicator to identify the region of interest 308, such as a hand pointing to the region of interest 308). As yet another example, the recognition and / or tracking engine 326 may identify the region of interest 308 based on semantic segmentation techniques or saliency detection techniques (e.g., determining significant regions of image 306).

[0077] The identification and / or tracking engine 326 can identify a region of interest 308, enabling it to track objects within the region of interest 308. The region of interest 308 can be associated with objects detected and / or tracked by the identification and / or tracking engine 326. For example, the region of interest 308 can be a bounding box that includes detected and / or tracked objects.

[0078] The recognition and / or tracking engine 326 may generate ROI information 330 indicating the determined region of interest 308 and provide the ROI information 330 to the image processing engine 312 and / or encoder 314. Additionally or alternatively, the recognition and / or tracking engine 326 may determine the object pose 328. The object pose 328 may indicate the location and / or orientation of the object detected and / or tracked by the recognition and / or tracking engine 326.

[0079] Renderer 332 (which may be circuitry or a chip (e.g., an FPGA or a processor), or may be implemented in circuitry or a chip) may receive object pose 328 from recognition and / or tracking engine 326 and may render an image based on object pose 328 for display by XR device 302. For example, renderer 332 may determine where to display virtual content in display 320 of XR device 302 based on object pose 328. As an example, renderer 332 may determine to display virtual content to overlay tracked real-world objects within the user's field of view.

[0080] Renderer 332 provides a rendered image to encoder 334. In some cases, encoder 334 and decoder 324 may be included in the same circuitry or chip. In other cases, encoder 334 may be independent of decoder 324. In any case, encoder 334 may be a circuitry or chip (e.g., an FPGA or processor), or may be implemented in a circuitry or chip. Encoder 334 may encode image data from renderer 332 for transmission (e.g., as separate data packets for sequential transmission). In one exemplary example, encoder 334 may encode image data based on a video decoding standard (such as HEVC, VVC, or another video decoding standard). In another exemplary example, encoder 334 may encode image data using a machine learning system trained to encode images (e.g., trained using supervised, semi-supervised, or self-supervised learning techniques).

[0081] After encoding the image data, the accessory device 322 can send the encoded data to the XR device 302 (e.g., using...). Figure 3 (Communication engine not illustrated). XR device 302 can receive encoded data (e.g., using...) Figure 3 The encoder 314 (not illustrated in the diagram) decodes the encoded data at decoder 316. In some cases, decoder 316 and encoder 314 may be included in the same circuit or chip. In other cases, decoder 316 may be independent of encoder 314. In any case, decoder 316 may be a circuit or chip (e.g., an FPGA or a processor), or may be implemented in a circuit or chip.

[0082] Image processing engine 318 may receive and process decoded image data from decoder 316. For example, image processing engine 318 may perform one or more of the following operations: color conversion, error hiding, and / or image distortion for head posture during display (which may also be referred to in the art as post-reprojection). Display 320 may receive and display the processed image data from image processing engine 318.

[0083] In some cases, XR device 302 may periodically transmit additional image data fully encoded using a QP (e.g., a relatively low QP) without low-pass filtering or masking. Such images may allow identification and / or tracking engine 326 to detect objects and / or identify or update additional regions of interest 308. Additionally or alternatively, in some cases, identification and / or tracking engine 326 may request XR device 302 to capture and transmit one or more images 306 encoded using a relatively low QP and / or without low-pass filtering. Identification and / or tracking engine 326 may request such images 306 based on determining the likelihood that a new object might be represented in such images 306.

[0084] Figure 4 This is an illustration of an example image 402 that can be processed according to various aspects of this disclosure. For example, image 402 can be displayed on an XR device (e.g., Figure 3 XR device 302 or Figure 1 The image capture device (e.g., XR device 102) Figure 3 The image is captured at the image capture device 304. Two regions of interest can be identified in the image 402: region of interest 404 and region of interest 406. According to some examples, region of interest 404 and region of interest 406 can be captured by an accessory device (e.g., Figure 1 Supporting equipment 104 or Figure 3 The accessory device 322) is identified. As described above, regions of interest 404 and 406 in image 402 can be identified based on previous images in a series of images including image 402. As described above, regions of interest 404 and 406 can be identified based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision. Pixels of image 402 not included in regions of interest 404 or 406 can be non-region of interest portions 408 of image 402.

[0085] In some examples, different parameters can be used to encode image 402 when encoding region of interest 404, region of interest 406, and non-region of interest portion 408. For example, a first quantization parameter (QP) can be used to encode region of interest 404 and / or region of interest 406, and a second QP (e.g., greater than the first QP) can be used to encode non-region of interest portion 408.

[0086] Alternatively or additionally, the region of interest 408 may be filtered (e.g., using a low-pass filter) before encoding. Alternatively or additionally, the region of interest 408 may be masked, for example, by modifying image 402 so that it is represented by a uniform value. The uniform value may be the average value (or average intensity value) of image 402 or the average value (or average intensity value) of the region of interest 408.

[0087] By applying a relatively high QP (e.g., higher than the QP used to encode regions of interest 404 and / or 406) when encoding the non-interest region 408, and / or by filtering or masking the non-interest region 408, the encoded data for image 402 can be smaller than the data without utilizing a relatively high QP, filtering, and / or masking. Smaller data may require less bandwidth to transmit (e.g., from an XR device to a companion device).

[0088] In some respects, if an image includes two (or more) regions of interest that are close to or overlap each other, then the system and techniques can merge the two (or more) regions of interest and treat the merged regions of interest as a whole.

[0089] Figure 5 This is an illustration of an example image 502 that can be processed according to various aspects of this disclosure. For example, image 502 can be processed in an XR device (e.g., Figure 3 XR device 302 or Figure 1 The image capture device (e.g., XR device 102) Figure 3 The image is captured at the image capture device 304. The region of interest 504 can be identified in the image 502. According to some examples, the region of interest 504 can be captured by an accessory device (e.g., Figure 1 Supporting equipment 104 or Figure 3 The associated equipment 322) is identified. As described above, the region of interest 504 in image 502 can be identified based on a previous image in a series of images including image 502. As described above, the region of interest 504 can be identified based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision.

[0090] Pixels of image 502 that are not included in region of interest 504 may be a non-region of interest portion 510 of image 502. The non-region of interest portion 510 may include one or more regions surrounding region of interest 504, including, for example, surrounding region 506 and surrounding region 508.

[0091] In some examples, different parameters can be used to encode image 502 when encoding region of interest 504, surrounding region 506, surrounding region 508, and non-region of interest portion 510. For example, a first QP can be used to encode region of interest 504, a second QP (e.g., greater than the first QP) can be used to encode surrounding region 506, a third QP (e.g., greater than the second QP) can be used to encode surrounding region 508, and a fourth QP (e.g., greater than the third QP) can be used to encode non-region of interest portion 510.

[0092] Additionally or alternatively, the surrounding region 506, surrounding region 508, and / or the region of non-interest portion 510 may be filtered (e.g., using one or more low-pass filters) prior to encoding. In some examples, different filters (e.g., with different passbands) may be used to encode each of the surrounding region 506, surrounding region 508, and the region of non-interest portion 510. For example, a first low-pass filter with a first passband may be applied to the surrounding region 506, a second low-pass filter with a second passband (e.g., smaller than the first passband) may be applied to the surrounding region 508, and a third low-pass filter with a third passband (e.g., smaller than the second passband) may be applied to the region of non-interest portion 510.

[0093] Alternatively or additionally, the region of interest 510 (including surrounding regions 506 and / or 508) may be masked, for example, by modifying image 502 so that the region of interest 510 is represented by a uniform value. The uniform value may be the average value (or average intensity value) of image 502 or the average value (or average intensity value) of the region of interest 510.

[0094] By applying one or more different parameters when encoding the non-interest region 510, and / or by filtering or masking the non-interest region 510, the encoded data of image 502 can be smaller than the data without using different parameters, filtering, and / or masking. Smaller data may require less bandwidth to transmit (e.g., from an XR device to a companion device).

[0095] Figure 6 This is an illustration of an example image 602 that can be processed according to various aspects of this disclosure. For example, image 602 can be processed in an XR device (e.g., Figure 3 XR device 302 or Figure 1 The image capture device (e.g., XR device 102) Figure 3 The image is captured at the image capture device 304. The region of interest 604 can be identified in the image 602. According to some examples, the region of interest 604 can be captured by an accessory device (e.g., Figure 1 Supporting equipment 104 or Figure 3 The associated equipment 322) is identified. As described above, the region of interest 604 in image 602 can be identified based on a previous image in a series of images including image 602. As described above, the region of interest 604 can be identified based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision.

[0096] Pixels of image 602 that are not included in region of interest 604 may be a non-region of interest portion 606 of image 602. Non-region of interest portion 606 may include pixels located at multiple different corresponding distances from region of interest 604. For example, non-region of interest portion 606 may include pixels located at a substantially distance 608 from region of interest 604 (e.g., a rectangular pixel ring), pixels located at a substantially distance 610 from region of interest 604 (e.g., a rectangular pixel ring), pixels located at a substantially distance 612 from region of interest 604 (e.g., a rectangular pixel ring), and pixels located at a substantially distance 614 from region of interest 604 (e.g., a rectangular pixel ring).

[0097] In some examples, different parameters can be used to encode the image 602 when encoding the region of interest 604 and the non-region of interest portion 606 based on the distance between the region of interest 604 and the encoded pixels.

[0098] For example, a first quantization parameter (QP) can be used to encode the region of interest 604, a second QP (e.g., greater than the first QP) can be used to encode pixels that are substantially 608 away from the region of interest 604, a third QP (e.g., greater than the second QP) can be used to encode pixels that are substantially 610 away from the region of interest 604, a fourth QP (e.g., greater than the third QP) can be used to encode pixels that are substantially 612 away from the region of interest 604, and a fifth QP (e.g., greater than the fourth QP) can be used to encode pixels that are substantially 614 away from the region of interest 604. For example, the QP used to encode a given pixel can be determined based on the distance between the given pixel and the region of interest 604.

[0099] Alternatively or additionally, the non-interest region 606 may be filtered (e.g., using one or more low-pass filters) prior to encoding. In some examples, different filters (e.g., with different passbands) may be used to encode different pixels of the non-interest region 606. For example, a first low-pass filter with a first passband may be applied to pixels that are substantially 608 away from the region of interest 604, a second low-pass filter with a second passband (e.g., smaller than the first passband) may be applied to pixels that are substantially 610 away from the region of interest 604, a third low-pass filter with a third passband (e.g., smaller than the second passband) may be applied to pixels that are substantially 612 away from the region of interest 604, and a fourth low-pass filter with a fourth passband (e.g., smaller than the third passband) may be applied to pixels that are substantially 614 away from the region of interest 604.

[0100] Alternatively or additionally, the region of interest 606 may be masked, for example, by modifying image 602 so that the region of interest 606 is represented by a uniform value. The uniform value may be the average value (or average intensity value) of image 602 or the average value (or average intensity value) of the region of interest 606.

[0101] By applying one or more different parameters when encoding the non-interest region 606, and / or by filtering or masking the non-interest region 606, the encoded data of image 602 can be smaller than the data without using different parameters, filtering, and / or masking. Smaller data may require less bandwidth to transmit (e.g., from an XR device to a companion device).

[0102] Figure 7 This is a flowchart of a process 700 for processing image data according to various aspects of this disclosure. One or more operations of process 700 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, or other types of computing devices. One or more operations of process 700 may be implemented as software components that execute and run on one or more processors.

[0103] At box 702, a computing device (or one or more components thereof) can capture an image. For example, Figure 3 The XR device 302 can use the image capture device 304 to capture images.

[0104] At box 704, the computing device (or one or more components thereof) may determine at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision. In some aspects, the computing device (or one or more components thereof) may receive an indication of the at least one ROI from another computing device to determine the at least one ROI. For example, XR device 302 may receive an indication of the at least one ROI from another computing device. Figure 3 Tracking engine 326 receives Figure 3 The ROI information 330. The XR device 302 (e.g., another computing device) and the XR device 302 can determine the region of interest based on the ROI information 330.

[0105] In some aspects, a region of interest can be determined to track at least one corresponding object represented in at least one region of interest. For example, accessory device 322 can determine a region of interest to track an object represented in that region of interest.

[0106] In some aspects, in order to determine the at least one region of interest, the at least one processor may determine at least two regions of interest in the image. A first portion of the image may correspond to the at least two regions of interest. For example, the first portion may correspond to... Figure 4 Region of interest 404 and region of interest 406.

[0107] At frame 706, a computing device (or one or more components thereof) may encode a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one region of interest. For example, XR device 302 may use the first parameter to encode a first portion of the image captured at frame 702 (the first portion corresponding to the region of interest determined at frame 704).

[0108] At box 708, the computing device (or one or more components thereof) may encode a second portion of the image according to a second parameter to generate second encoded data. For example, XR device 302 may use the second parameter to encode a second portion of the image captured at box 702 (which does not correspond to the region of interest defined at box 704).

[0109] In some respects, the first parameter may be or may include a first quantization parameter, and the second parameter may be or may include a second quantization parameter, which is greater than the first quantization parameter.

[0110] In some aspects, a computing device (or one or more components thereof) may encode a third portion of an image according to a third quantization parameter to generate third coded data. The third portion of the image may surround a first portion of the image. The third quantization parameter may be greater than the first quantization parameter and less than the second quantization parameter. For example, the first quantization parameter may be used to encode a region of interest 504, a second quantization parameter greater than the first quantization parameter may be used to encode a region of non-interest 510, and a third quantization parameter between the first and second quantization parameters may be used to encode the surrounding region 506.

[0111] In some aspects, the second quantization parameter may be or may include multiple quantization parameters. The second portion of the image may include multiple portions of the image. A computing device (or one or more components thereof) may use a corresponding quantization parameter from the multiple quantization parameters to encode each of the multiple portions of the image to generate multiple encoded data. Each corresponding quantization parameter from the multiple quantization parameters used to encode each of the multiple portions of the image may be based on the distance between the first portion of the image and each corresponding portion of the multiple portions of the image. For example, a first quantization parameter may be used to encode pixels that are substantially 608 distances from the region of interest 604, a second quantization parameter may be used to encode pixels that are substantially 610 distances from the region of interest 604, a third quantization parameter may be used to encode pixels that are substantially 612 distances from the region of interest 604, and a fourth quantization parameter may be used to encode pixels that are substantially 614 distances from the region of interest 604.

[0112] In some aspects, the computing device (or one or more components thereof) may compress the second encoded data while encoding the second portion of the image to generate second encoded data. In some aspects, the computing device (or one or more components thereof) may blur the second portion of the image before encoding it. In some aspects, the computing device (or one or more components thereof) may filter the second portion of the image using a low-pass filter before encoding it. In some aspects, the computing device (or one or more components thereof) may mask the second portion of the image using representative values ​​of the image before encoding it.

[0113] In some aspects, the computing device (or one or more components thereof) may determine the second parameter based on a bandwidth threshold such that the transmission of the first and second encoded data does not exceed the bandwidth threshold. In some aspects, the computing device (or one or more components thereof) may determine the second parameter based on an object detection threshold.

[0114] At block 710, a computing device (or one or more components thereof) may send first encoded data and second encoded data to the computing device. For example, XR device 302 may send data encoded at blocks 706 and 708 to accessory device 322.

[0115] Figure 8This is a flowchart of a process 800 for processing image data according to various aspects of this disclosure. One or more operations of process 800 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, or other types of computing devices. One or more operations of process 800 may be implemented as software components that execute and run on one or more processors.

[0116] At box 802, the computing device (or one or more components thereof) may receive first data encoded in the first image from the image capture device. For example, Figure 3 The supporting equipment 322 can be obtained from Figure 3 The XR device 302 receives data that encodes the image.

[0117] At frame 804, a computing device (or one or more components thereof) may determine at least one region of interest (ROI) of the first image. For example, an accessory device 322 may determine the ROI of an image received at frame 802.

[0118] In some aspects, the at least one region of interest may be determined based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision. In some aspects, a computing device (or one or more components thereof) may determine the region of interest based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision.

[0119] In some aspects, a computing device (or one or more components thereof) can determine at least two regions of interest (ROIs) of the first image. A first portion of the second image may correspond to these at least two ROIs. For example, ROIs 404 and ROIs 406 may be determined.

[0120] At box 806, the computing device (or one or more components thereof) may send an indication of the at least one region of interest to the image capturing device. For example, the accessory device 322 may send ROI information 330, which may indicate the region of interest determined at box 804.

[0121] At block 808, a computing device (or one or more components thereof) may receive second data encoding a second image, wherein a first portion of the second image is encoded according to a first parameter, the first portion of the second image corresponding to the at least one region of interest, and a second portion of the second image is encoded according to a second parameter. For example, an accessory device 322 may receive the second image from an XR device 302. Different parameters may be used to encode the second image for different regions. For example, the first parameter used for the region of interest identified at block 804 may be different from the parameter used to encode portions of the second image outside the region of interest.

[0122] At box 810, a computing device (or one or more components thereof) may decode the second data to generate a reconstructed instance of the second image. For example, an accessory device 322 may decode the data to reconstruct an instance of the second image.

[0123] At box 812, a computing device (or one or more components thereof) can track objects in a reconstructed instance of the second image. For example, an accessory device 322 can track objects represented in a reconstructed instance of the second image.

[0124] In some examples, the methods described herein (e.g., process 700, process 800, and / or other methods described herein) may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 3 XR system 300, Figure 3 XR equipment 302, Figure 3 The accompanying equipment 322 or another system or device performs the operation. In another example, one or more of these methods may be performed by... Figure 9 The computing device architecture 900 shown is implemented wholly or partially. For example, it has Figure 9 The computing device of the computing device architecture 900 shown may include or be included in Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 3 XR system 300, Figure 3 XR equipment 302, Figure 3 It is used in the supporting equipment 322 or as a component of another system or equipment, and can enable the operation of process 700, process 800 and / or other processes described herein.

[0125] Computing devices may include any suitable device, such as a vehicle or vehicle-mounted computing device, mobile device (e.g., mobile phone), desktop computing device, tablet computing device, wearable device (e.g., VR headset, AR headset, AR glasses, network-connected watch or smartwatch, or other wearable device), server computer, robotic device, television set, and / or any other computing device with the resource capability to perform the processes described herein (including process 800 and / or other processes described herein). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0126] A component capable of implementing a computing device in a circuit. For example, a component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0127] Processes 700, 800, and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0128] Additionally, processes 700, 800, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or implemented in a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0129] Figure 9 An example computing device architecture 900 is illustrated, illustrating example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 900 may include, implement, or be included in... Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 3 XR system 300, Figure 3 XR equipment 302, Figure 3 The supporting equipment 322 or any or all of another system or equipment.

[0130] The components of the computing device architecture 900 are shown to communicate electrically with each other using a connection 912, such as a bus. The example computing device architecture 900 includes a processing unit (CPU or processor) 902 and a computing device connection 912 that couples various computing device components, including computing device memories 910 (such as read-only memory (ROM) 908 and random access memory (RAM) 906), to the processor 902.

[0131] The computing device architecture 900 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 902. The computing device architecture 900 may copy data from memory 910 and / or storage device 914 to cache 904 for fast access by the processor 902. In this way, the cache can provide performance improvements by avoiding latency for the processor 902 while waiting for data. These and other engines can control or be configured to control the processor 902 to perform various actions. Other computing device memories 910 may also be used. Memory 910 may include various different types of memory with different performance characteristics. The processor 902 may include any general-purpose processor and hardware or software services configured to control the processor 902 (such as services 1 916, service 2 918, and service 3 920 stored in storage device 914), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 902 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0132] To enable user interaction with the computing device architecture 900, input device 922 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 924 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some instances, a multi-mode computing device allows the user to provide multiple types of input to communicate with the computing device architecture 900. Communication interface 926 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0133] Storage device 914 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 906, read-only memory (ROM) 908, and hybrid forms thereof. Storage device 914 may include services 916, 918, and 920 for controlling processor 902. Other hardware or software engines or modules are envisioned. Storage device 914 may be connected to computing device connection 912. In one aspect, a hardware engine or module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 902, connection 912, output device 924, etc., to perform the function.

[0134] The term "substantially," when referring to a given parameter, property, or condition, can mean the degree to which a given parameter, property, or condition is satisfied within a small range of deviations, as would be understood by one of ordinary skill in the art, such as, for example, within acceptable manufacturing tolerances. For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0135] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet computer, laptop computer, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0136] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or a particular aspect. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0137] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects.

[0138] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0139] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0140] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as optical discs (CDs) or digital versatile discs (DVDs)), flash memory, USB devices provided with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0141] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.

[0142] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or intercalation cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed in a single device.

[0143] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0144] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0145] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this specification.

[0146] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0147] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., via a wired or wireless connection and / or other suitable communication interface).

[0148] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, the claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language that states "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the claim language that states "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0149] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0150] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0151] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0152] The exemplary aspects of this disclosure include:

[0153] Aspect 1. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: cause an image capture device to capture an image; determine at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encode a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one ROI; encode a second portion of the image according to a second parameter to generate second encoded data; and cause at least one transmitter to transmit the first encoded data and the second encoded data to a computing device. In some cases, the apparatus includes the image capture device to capture the image. In some cases, the apparatus includes the at least one transmitter to transmit the first encoded data and the second encoded data to a computing device.

[0154] Aspect 2. The apparatus according to aspect 1, wherein the at least one region of interest is determined to track at least one corresponding object represented in the at least one region of interest.

[0155] Aspect 3. The apparatus according to any one of Aspect 1 or 2, wherein: in order to determine the at least one region of interest, the at least one processor is configured to determine at least two regions of interest in the image; and the first portion of the image corresponds to the at least two regions of interest.

[0156] Aspect 4. The apparatus according to any one of aspects 1 to 3, wherein, in order to determine the at least one region of interest, the at least one processor is configured to receive an indication of the at least one region of interest from the computing device.

[0157] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the first parameter includes a first quantization parameter, and the second parameter includes a second quantization parameter, the second quantization parameter being greater than the first quantization parameter.

[0158] Aspect 6. The apparatus according to aspect 5, wherein the at least one processor is further configured to: encode a third portion of the image according to a third quantization parameter to generate third coded data, the third portion of the image surrounding the first portion of the image, the third quantization parameter being greater than the first quantization parameter and less than the second quantization parameter.

[0159] Aspect 7. The apparatus according to aspect 6, wherein the second quantization parameter comprises a plurality of quantization parameters, and wherein the second portion of the image comprises a plurality of portions of the image, wherein the at least one processor is further configured to: encode each of the plurality of portions of the image using a corresponding quantization parameter among the plurality of quantization parameters to generate a plurality of encoded data, wherein each corresponding quantization parameter among the plurality of quantization parameters used to encode each of the plurality of portions of the image is based on the distance between the first portion of the image and each corresponding portion among the plurality of portions of the image.

[0160] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein the at least one processor is further configured to: compress the second encoded data while encoding the second portion of the image to generate the second encoded data.

[0161] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the at least one processor is further configured to: blur the second portion of the image before encoding the second portion of the image.

[0162] Aspect 10. The apparatus according to any one of aspects 1 to 9, wherein the at least one processor is further configured to: filter the second portion of the image using a low-pass filter before encoding the second portion of the image.

[0163] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the at least one processor is further configured to: mask the second portion of the image using a representative value of the image before encoding the second portion of the image.

[0164] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the at least one processor is further configured to: determine the second parameter based on a bandwidth threshold such that the transmission of the first encoded data and the second encoded data does not exceed the bandwidth threshold.

[0165] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the at least one processor is further configured to determine the second parameter based on an object detection threshold.

[0166] Aspect 14. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive first data encoding a first image from an image capture device; determine at least one region of interest (ROI) of the first image; cause at least one transmitter to transmit an indication of the at least one ROI to the image capture device; receive second data encoding a second image, a first portion of the second image being encoded according to a first parameter, the first portion of the second image corresponding to the at least one ROI, and a second portion of the second image being encoded according to a second parameter; decode the second data to generate a reconstructed instance of the second image; and track objects in the reconstructed instance of the second image. In some cases, the apparatus includes the at least one transmitter to transmit the first encoded data and the second encoded data to a computing device.

[0167] Aspect 15. The apparatus according to aspect 14, wherein the at least one region of interest is determined based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision. In some cases, the processor can be configured to determine the at least one region of interest based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision.

[0168] Aspect 16. The apparatus according to any one of Aspects 14 or 15, wherein: in order to determine the at least one region of interest, the at least one processor is configured to determine at least two regions of interest in the first image; and the first portion of the second image corresponds to the at least two regions of interest.

[0169] Aspect 17. A method for processing image data, the method comprising: capturing an image; determining at least one region of interest (ROI) of the image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; encoding a first portion of the image according to a first parameter to generate first encoded data, the first portion of the image corresponding to the at least one ROI; encoding a second portion of the image according to a second parameter to generate second encoded data; and transmitting the first encoded data and the second encoded data to a computing device.

[0170] Aspect 18. The method according to aspect 17, wherein the at least one region of interest is determined to track at least one corresponding object represented in the at least one region of interest.

[0171] Aspect 19. The method according to any one of Aspects 17 or 18, wherein: determining the at least one region of interest includes determining at least two regions of interest in the image; and the first portion of the image corresponds to the at least two regions of interest.

[0172] Aspect 20. The method according to any one of aspects 17 to 19, wherein determining the at least one region of interest includes receiving an indication of the at least one region of interest from the computing device.

[0173] Aspect 21. The method according to any one of Aspects 17 to 19, wherein the first parameter includes a first quantization parameter, and the second parameter includes a second quantization parameter, the second quantization parameter being greater than the first quantization parameter.

[0174] Aspect 22. The method according to aspect 21, the method further comprising: encoding a third portion of the image according to a third quantization parameter to generate third coded data, the third portion of the image surrounding the first portion of the image, the third quantization parameter being greater than the first quantization parameter and less than the second quantization parameter.

[0175] Aspect 23. The method according to aspect 22, wherein the second quantization parameter comprises a plurality of quantization parameters, and wherein the second portion of the image comprises a plurality of portions of the image, the method further comprising: encoding each of the plurality of portions of the image using a corresponding quantization parameter of the plurality of quantization parameters to generate a plurality of encoded data, wherein each corresponding quantization parameter of the plurality of quantization parameters used to encode each of the plurality of portions of the image is based on a distance between the first portion of the image and each corresponding portion of the plurality of portions of the image.

[0176] Aspect 24. The method according to any one of aspects 17 to 23, the method further comprising: compressing the second encoded data while encoding the second portion of the image to generate the second encoded data.

[0177] Aspect 25. The method according to any one of aspects 17 to 24, the method further comprising: blurring the second portion of the image before encoding the second portion of the image.

[0178] Aspect 26. The method according to any one of aspects 17 to 25, the method further comprising: filtering the second portion of the image using a low-pass filter before encoding the second portion of the image.

[0179] Aspect 27. The method according to any one of aspects 17 to 26, the method further comprising: masking the second portion of the image using a representative value of the image before encoding the second portion of the image.

[0180] Aspect 28. The method according to any one of Aspects 17 to 27, the method further comprising: determining the second parameter based on a bandwidth threshold such that the transmission of the first encoded data and the second encoded data does not exceed the bandwidth threshold.

[0181] Aspect 29. The method according to any one of aspects 17 to 29, the method further comprising: determining the second parameter based on an object detection threshold.

[0182] Aspect 30. A method for processing image data, the method comprising: receiving, at a computing device, first data encoding a first image from an image capturing device; determining at least one region of interest (ROI) of the first image; sending from the computing device to the image capturing device an indication of the at least one ROI; receiving, second data encoding a second image, a first portion of the second image being encoded according to a first parameter, the first portion of the second image corresponding to the at least one ROI, and a second portion of the second image being encoded according to a second parameter; decoding the second data to generate a reconstructed instance of the second image; and tracking objects in the reconstructed instance of the second image.

[0183] Aspect 31. A method for processing image data by an extended reality system, the method comprising: capturing a first image at an image capture device of the extended reality system; encoding the first image into first data; transmitting the first data from the image capture device to a computing device of the extended reality system; determining at the computing device at at least one region of interest (ROI) of the first image based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; transmitting an indication of the at least one ROI from the computing device to the image capture device; capturing a second image at the image capture device; encoding a first portion of the second image according to a first parameter to generate first encoded data, the first portion of the second image corresponding to the at least one ROI; encoding a second portion of the second image according to a second parameter to generate second encoded data; transmitting the first encoded data and the second encoded data from the image capture device to the computing device; decoding the first encoded data and the second encoded data to generate a reconstructed instance of the second image; and tracking objects in the reconstructed instance of the second image.

[0184] Aspect 32. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one of aspects 17 to 31.

[0185] Aspect 33. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 17 to 31.

Claims

1. An apparatus for processing image data, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: Enables the image capture device to capture an image; At least one region of interest in the image is determined based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; The first portion of the image is encoded according to the first parameter to generate first encoded data, wherein the first portion of the image corresponds to the at least one region of interest; The second portion of the image is encoded according to the second parameter to generate second encoded data; as well as At least one transmitter sends the first encoded data and the second encoded data to a computing device.

2. The apparatus of claim 1, wherein the at least one region of interest is determined to track at least one corresponding object represented in the at least one region of interest.

3. The apparatus according to claim 1, wherein: To determine the at least one region of interest, the at least one processor is configured to determine at least two regions of interest in the image; and The first portion of the image corresponds to the at least two regions of interest.

4. The apparatus of claim 1, wherein, in order to determine the at least one region of interest, the at least one processor is configured to receive an indication of the at least one region of interest from the computing device.

5. The apparatus of claim 1, wherein the first parameter includes a first quantization parameter, and the second parameter includes a second quantization parameter, the second quantization parameter being greater than the first quantization parameter.

6. The apparatus of claim 5, wherein the at least one processor is further configured to: encode a third portion of the image according to a third quantization parameter to generate third coded data, the third portion of the image surrounding the first portion of the image, the third quantization parameter being greater than the first quantization parameter and less than the second quantization parameter.

7. The apparatus of claim 6, wherein the second quantization parameter comprises a plurality of quantization parameters, and wherein the second portion of the image comprises a plurality of portions of the image, wherein the at least one processor is further configured to: Each of the plurality of portions of the image is encoded using a corresponding quantization parameter among the plurality of quantization parameters to generate a plurality of encoded data, wherein each corresponding quantization parameter among the plurality of quantization parameters used to encode each of the plurality of portions of the image is based on the distance between the first portion of the image and each corresponding portion among the plurality of portions of the image.

8. The apparatus of claim 1, wherein the at least one processor is further configured to: compress the second encoded data while encoding the second portion of the image to generate the second encoded data.

9. The apparatus of claim 1, wherein the at least one processor is further configured to: blur the second portion of the image before encoding the second portion of the image.

10. The apparatus of claim 1, wherein the at least one processor is further configured to: filter the second portion of the image using a low-pass filter before encoding the second portion of the image.

11. The apparatus of claim 1, wherein the at least one processor is further configured to: mask the second portion of the image using a representative value of the image before encoding the second portion of the image.

12. The apparatus of claim 1, wherein the at least one processor is further configured to: determine the second parameter based on a bandwidth threshold such that the transmission of the first encoded data and the second encoded data does not exceed the bandwidth threshold.

13. The apparatus of claim 1, wherein the at least one processor is further configured to determine the second parameter based on an object detection threshold.

14. An apparatus for processing image data, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: Receive first data encoded from the image capture device; Determine at least one region of interest in the first image; At least one transmitter sends an indication of the at least one region of interest to the image capturing device; Receive second data that encodes a second image, wherein a first portion of the second image is encoded according to a first parameter, the first portion of the second image corresponds to the at least one region of interest, and a second portion of the second image is encoded according to a second parameter; The second data is decoded to generate a reconstructed instance of the second image; as well as Track the objects in the reconstructed instance of the second image.

15. The apparatus of claim 14, wherein the at least one region of interest is determined based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision.

16. The apparatus according to claim 14, wherein: To determine the at least one region of interest, the at least one processor is configured to determine at least two regions of interest in the first image; and The first portion of the second image corresponds to the at least two regions of interest.

17. A method for processing image data, the method comprising: Capture images; At least one region of interest in the image is determined based on at least one of object recognition, object tracking, hand tracking, semantic segmentation, saliency detection, or computer vision; The first portion of the image is encoded according to the first parameter to generate first encoded data, wherein the first portion of the image corresponds to the at least one region of interest; The second portion of the image is encoded according to the second parameter to generate second encoded data; as well as The first encoded data and the second encoded data are sent to the computing device.

18. The method of claim 17, wherein the at least one region of interest is determined to track at least one corresponding object represented in the at least one region of interest.

19. The method of claim 17, wherein: Determining the at least one region of interest includes determining at least two regions of interest in the image; and The first portion of the image corresponds to the at least two regions of interest.

20. The method of claim 17, wherein determining the at least one region of interest comprises receiving an indication of the at least one region of interest from the computing device.

21. The method of claim 17, wherein the first parameter includes a first quantization parameter, and the second parameter includes a second quantization parameter, the second quantization parameter being greater than the first quantization parameter.

22. The method according to claim 21, further comprising: The third portion of the image is encoded according to a third quantization parameter to generate third coded data, the third portion of the image surrounding the first portion of the image, the third quantization parameter being greater than the first quantization parameter and less than the second quantization parameter.

23. The method of claim 22, wherein the second quantization parameter comprises a plurality of quantization parameters, and wherein the second portion of the image comprises a plurality of portions of the image, the method further comprising: Each of the plurality of portions of the image is encoded using a corresponding quantization parameter among the plurality of quantization parameters to generate a plurality of encoded data, wherein each corresponding quantization parameter among the plurality of quantization parameters used to encode each of the plurality of portions of the image is based on the distance between the first portion of the image and each corresponding portion among the plurality of portions of the image.

24. The method according to claim 17, further comprising: When encoding the second portion of the image to generate the second encoded data, the second encoded data is compressed.

25. The method according to claim 17, further comprising: The second portion of the image is blurred before it is encoded.

26. The method according to claim 17, further comprising: Before encoding the second portion of the image, the second portion of the image is filtered using a low-pass filter.

27. The method of claim 17, further comprising: Before encoding the second portion of the image, the second portion of the image is masked using a representative value of the image.

28. The method according to claim 17, further comprising: The second parameter is determined based on a bandwidth threshold, such that the transmission of the first encoded data and the second encoded data does not exceed the bandwidth threshold.

29. The method of claim 17, further comprising: The second parameter is determined based on the object detection threshold.

30. A method for processing image data, the method comprising: The computing device receives first data encoded from the image capture device, which is used to encode the first image. Determine at least one region of interest in the first image; The computing device sends an indication of the at least one region of interest to the image capturing device; Receive second data that encodes a second image, wherein a first portion of the second image is encoded according to a first parameter, the first portion of the second image corresponds to the at least one region of interest, and a second portion of the second image is encoded according to a second parameter; The second data is decoded to generate a reconstructed instance of the second image; and objects in the reconstructed instance of the second image are tracked.