Efficient processing of image data based on region of interest

By providing ROI indication at the image sensor and coordinating the operation of the image sensor and ISP, image data is processed at full resolution only within the ROI, solving the problem of resource waste in the prior art and achieving high efficiency in image data processing and improved system performance.

CN121890102APending Publication Date: 2026-04-17QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-08-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently save power, bandwidth, and processing time when processing image data, especially in extended reality devices, particularly in the coordination between image sensors and image signal processors, leading to resource waste and inefficiency.

Method used

By providing an indication of the region of interest (ROI) at the image sensor, the operation of the image sensor and the front-end image signal processor (ISP) is coordinated, and image data is processed at full resolution only within the ROI, while it is processed at a lower resolution in non-ROI areas. This reduces the size of the image data, thereby saving power, bandwidth and processing time.

Benefits of technology

It achieves efficient coordination between the image sensor and the ISP, reduces image data processing time and resource consumption, and improves the efficiency of image data processing and the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121890102A_ABST
    Figure CN121890102A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing data are described herein. For example, an apparatus for processing data is provided. In some embodiments, an apparatus may include an image signal processor (ISP) configured to: receive image data and an indication of a region of interest (ROI) from an image sensor; determining an image processing setting for processing the image data based on the ROI; and processing the image data based on the image processing settings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates throughout to the efficient processing of image data based on regions of interest. For example, aspects of this disclosure include systems and techniques for: receiving image data from an image sensor and an indication of a region of interest (ROI) of that image data at an image signal processor; determining image processing settings based on the ROI; and processing the image data based on those image processing settings. Background Technology

[0002] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). Cameras can be configured with various image capture settings and / or image processing settings to alter the appearance of the captured images. Image capture settings can be determined and applied before and / or during image capture, such as ISO, exposure time (also known as exposure, exposure duration, or shutter speed), aperture size (also known as aperture value), focus, and gain (including analog and / or digital gain). Furthermore, image processing settings can be configured for post-processing of the image, such as changes to contrast, brightness, saturation, sharpness, levels, curves, and color. Summary of the Invention

[0003] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following outline presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description that follows.

[0004] Systems and techniques for processing data are described. According to at least one example, an apparatus for processing data is provided. The apparatus includes: an image signal processor (ISP) configured to: receive image data and an indication of a region of interest (ROI) from an image sensor; determine image processing settings for processing the image data based on the ROI; and process the image data based on the image processing settings.

[0005] In another example, a method for processing data is provided. The method includes: receiving image data and an indication of a region of interest (ROI) from an image sensor at an image signal processor (ISP); determining image processing settings for processing the image data at the ISP based on the ROI; and processing the image data at the ISP based on the image processing settings.

[0006] In another example, an apparatus for processing data is provided, the apparatus including at least one memory and at least one processor (e.g., an image signal processor (ISP)) (e.g., configured in a circuit), the at least one processor being coupled to the at least one memory. The at least one processor is configured to: receive image data and indication of a region of interest (ROI) from an image sensor; determine image processing settings for processing the image data based on the ROI; and process the image data based on the image processing settings.

[0007] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors (e.g., one or more image signal processors (ISPs)), cause the one or more processors to: receive image data from an image sensor and an indication of a region of interest (ROI); determine image processing settings for processing the image data based on the ROI; and process the image data based on the image processing settings.

[0008] In another example, an apparatus for processing data is provided. The apparatus includes: components for receiving image data from an image sensor and an indication of a region of interest (ROI); components for determining image processing settings for processing the image data based on the ROI; and components for processing the image data based on the image processing settings.

[0009] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles (or computing devices or systems of vehicles), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, televisions (e.g., network-connected televisions), robotic devices or systems, or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0012] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0013] Figure 1 These are illustrations of examples of extended reality (XR) systems according to various aspects of this disclosure;

[0014] Figure 2 This is a block diagram illustrating the architecture of an example extended reality (XR) system according to some aspects of this disclosure;

[0015] Figure 3 This is a block diagram illustrating an example architecture of an image processing system according to various aspects of this disclosure;

[0016] Figure 4These are illustrations of example systems that can be used to efficiently process image data according to various aspects of this disclosure;

[0017] Figure 5 This is a block diagram illustrating an example system for processing image data;

[0018] Figure 6A This is a block diagram illustrating an example system for efficiently processing image data according to various aspects of this disclosure;

[0019] Figure 6B This is a block diagram illustrating an example system for efficiently processing image data according to various aspects of this disclosure;

[0020] Figure 7 This is an illustration of example groupings that may include image data and ROI indicators according to various aspects of this disclosure;

[0021] Figure 8 These are illustrations of two example groups that may include image data and ROI indicators, according to various aspects of this disclosure;

[0022] Figure 9 This is a flowchart illustrating another example process for efficiently processing image data according to various aspects of this disclosure;

[0023] Figure 10 This is a block diagram illustrating examples of deep learning neural networks that can be used to implement a perception module and / or one or more verification modules, based on some aspects of the disclosed techniques.

[0024] Figure 11 This is a block diagram illustrating examples of convolutional neural networks (CNNs) according to various aspects of this disclosure; and

[0025] Figure 12 This is a block diagram illustrating an example computing device architecture that can implement the various technologies described herein. Detailed Implementation

[0026] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0027] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with descriptions that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0028] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0029] A camera is a device that uses an image sensor to receive light and capture image data (e.g., still image frames or video data frames). Electronic devices (e.g., mobile phones, wearable devices (e.g., smartwatches, smart glasses, etc.), tablet computers, extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, etc.), connected devices, laptop computers, etc.) are increasingly equipped with camera hardware to capture image frames, such as still images and / or video frames, for consumption. For example, electronic devices may include cameras that allow the electronic devices to capture video or images of scenes, people, objects, etc. Additionally, cameras themselves are used in a variety of configurations (e.g., handheld digital cameras, digital single-lens reflex (DSLR) cameras, wearable cameras (including body-mounted cameras and head-mounted cameras), fixed cameras (e.g., for security and / or monitoring), vehicle-mounted cameras, etc.).

[0030] The camera can be configured with various image capture and image processing settings to alter the appearance of an image. Image camera settings can be determined and applied before or during image capture, such as ISO, exposure time (also known as exposure duration and / or shutter speed), aperture size (also known as aperture value), focus, and gain. Image processing settings can be configured for post-processing of the image, such as changes to contrast, brightness, saturation, sharpness, levels, curves, and color.

[0031] In some examples, the camera may include one or more processors, such as an image signal processor (ISP), that process image data captured by the image sensor. For example, raw image frames captured by the image sensor may be processed by the camera's ISP (e.g., to alter the contrast, brightness, saturation, sharpness, level, curves, and / or color of the raw image frames) to generate final image frames. In some cases, the electronics implementing the camera may further process the captured images or videos for certain effects (e.g., compression, image enhancement, image restoration, scaling, frame rate conversion, etc.) and / or certain applications (e.g., computer vision, extended reality (e.g., augmented reality, virtual reality, etc.), object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, and automation, etc.).

[0032] The camera may include an image capture device (or portion) comprising one or more image sensors capable of capturing light and generating raw image frames based thereon. The camera may include an image processing device (or portion) comprising one or more image signal processors (ISPs) capable of processing the raw image frames. The image processing device (or portion) may include a front-end section and a back-end section. The image capture device (or portion) may stream raw image data to the front-end section of the image processing device (or portion). The front-end section may perform one or more operations on the raw image data (e.g., when the raw image data is received). For example, the front-end section may perform one or more operations related to bad pixel correction (BPC), lens correction, lens shading correction, phase detection pixel correction, demosaic, lateral chromatic aberration correction, Bayer filtering, adaptive Bayer filtering, tone mapping, noise reduction, etc. The front-end section may provide the processed image frame to the back-end section, for example, by writing the processed image frame to memory (such as Double Data Rate (DDR) Synchronous Dynamic Random Access Memory (SDRAM) or any other memory device). The back-end can retrieve the processed image frame from memory and further process it. For example, the back-end can perform motion stabilization on the processed image frame.

[0033] Image processing and / or write and / or read operations between the front-end and back-end portions can result in significant power, bandwidth, and / or time consumption. One technique for saving power, bandwidth, and / or processing time involves processing an image frame based on a corresponding region of interest (ROI) within the frame. For example, such a technique could process the ROI of the image frame at full resolution and process non-ROI regions (e.g., outside the ROI) at a lower resolution than the full resolution, instead of processing the entire image frame at full resolution. Processing non-ROI regions at a lower resolution may mean processing fewer pixels (and / or writing to and reading from memory between the front-end and back-end portions), which saves power, bandwidth, and / or processing time. The final image may have lower resolution in non-ROI regions, but these regions may be less of interest than the ROIs, so the reduction in resolution may be insignificant or unimportant. For example, the ROI may be based on the user's gaze. The user may expect high resolution where they are looking, but may not notice lower resolution in areas outside their gaze.

[0034] Generally speaking, in an ISP processing pipeline, the earlier the resolution of a portion of the image is reduced, the greater the savings in power, bandwidth, and / or processing time. For example, if the front-end ISP includes three ISP engines (e.g., arranged in series and each ISP engine performs a separate operation), reducing the resolution of a non-ROI region at or before the first of the three ISP engines can save more power and / or processing time than reducing the resolution of a non-ROI region at the third ISP engine.

[0035] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, “Systems and Techniques”) for efficiently processing image data based on Region of Interest (ROI). The systems and techniques described herein may include ROI-based imaging techniques implemented at an image sensor. For example, an image sensor may generate image data at full resolution within the ROI of an image frame and image data at a lower resolution outside the ROI within the image frame. The lower-resolution image data outside the ROI may include smaller image data (e.g., fewer pixels) than the image data when the entire image is at full resolution. These systems and techniques save power, bandwidth, and / or processing time by processing smaller image data compared to other systems capable of processing full-resolution image data. Furthermore, these systems and techniques save even more power, bandwidth, and / or processing time than other systems by reducing the size of the image data at the image sensor rather than partially relying on an ISP pipeline.

[0036] One challenge in implementing ROI-based techniques (e.g., reducing the resolution of non-ROI regions) at the image sensor is coordinating the operation of the image sensor and the ISP. For example, the image sensor generates raw image data and provides it to the front-end ISP. The front-end ISP can then process the raw image data upon receipt. If a first portion of the raw image data (e.g., corresponding to an ROI) has a relatively high resolution, and a second portion (e.g., corresponding to a non-ROI region) has a relatively low resolution, the front-end ISP knows that the location and size of the data received by it in the full-resolution coordinates of the entire frame may be important. One method for coordinating the operation of the image sensor and the front-end ISP involves providing an indication of the ROI to both the image sensor and the front-end ISP. However, in such methods, the timing of providing the ROI indication to the front-end ISP is critical, and timing delays or resets may disable timing, which can lead to errors when processing image data at the front-end ISP.

[0037] These systems and technologies enable ROI-based techniques at the image sensor level by providing an indication of the ROI to the front-end ISP. For example, these systems and technologies allow the image sensor to provide both an ROI indicator and image data to the front-end ISP. The front-end ISP can receive the image data and the ROI indicator substantially simultaneously (e.g., in the same data packet) and can thus easily determine the correlation between the image data and the ROI. In this way, these systems and technologies enable ROI-based techniques to be implemented at the image sensor level, which can save power, bandwidth, and / or processing time throughout the ISP pipeline.

[0038] Various aspects of this application will be described below with reference to the accompanying drawings.

[0039] Figure 1This is a diagram illustrating an example of an extended reality (XR) system 100 according to various aspects of the present disclosure. As shown, the XR system 100 includes an XR device 102, an accessory device 104, and a communication link 106 between the XR device 102 and the accessory device 104. In some cases, the XR device 102 typically implements aspects of extended reality (including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc.) display, image capture, and / or view tracking. In some cases, the accessory device 104 typically implements aspects of extended reality computation. For example, the XR device 102 may capture an image of the environment of user 108 and provide the image to the accessory device 104 (e.g., via communication link 106). The accessory device 104 may render virtual content (e.g., in relation to the captured image of the environment) and provide the virtual content to the XR device 102 (e.g., via communication link 106). The XR device 102 may display the virtual content to user 108 (e.g., within user 108's field of view 110).

[0040] Typically, XR device 102 may display virtual content to be viewed by user 108 within field of view 110. In some examples, XR device 102 may include a transparent surface (e.g., optical glass) such that virtual objects can be displayed on the transparent surface (e.g., by projection onto the transparent surface) to overlay virtual content onto real-world objects viewed through the transparent surface (e.g., in a perspective configuration). In some cases, XR device 102 may include a camera and may display both real-world objects (e.g., as frames or images captured by the camera) and virtual objects overlaid on the displayed real-world objects (e.g., in a pass-through configuration). In various examples, XR device 102 may include aspects of a virtual reality headset, smart glasses, a live-feed video camera, a GPU, one or more sensors (e.g., such as one or more inertial measurement units (IMUs), image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, displays, smart glasses, etc.).

[0041] The companion device 104 can render virtual content to be displayed by the companion device 104. In some examples, the companion device 104 may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, or a mobile device acting as a server device), any other computing device and / or combinations thereof.

[0042] The communication link 106 can be based on any suitable wireless protocol, such as, for example, IEEE 802.11 (Wi-Fi), IEEE 802.15, or Bluetooth. ®The communication link 106 can be a direct wireless connection between the XR device 102 and the companion device 104 in some cases. In other cases, the communication link 106 can be via one or more intermediate devices, such as routers or switches and / or across a network.

[0043] Improving the efficiency of processing image data captured by the camera of XR device 102 can be beneficial. For example, the camera of XR device 102 can capture image data. One or more image signal processors (ISPs) on XR device 102 and / or one or more ISPs on accessory device 104 can be used to process the image data. Reducing the size of the captured image data (e.g., by reducing the resolution of the region of interest (ROI) portion of the image data) to save power, bandwidth, and / or processing time within XR system 100 can be beneficial to the operation of XR system 100. Reducing the size of the image data within the ISP of XR device 102 to save power of XR device 102 (e.g., the XR device can be powered by a relatively small battery) can be particularly beneficial.

[0044] Figure 2 This is a diagram illustrating the architecture of an example extended reality (XR) system 200 according to some aspects of this disclosure. The XR system 200 can execute XR applications and implement XR operations. The architecture of the XR system 200 can be... Figure 1 An example of the architecture of the XR system 100.

[0045] In this exemplary example, the XR system 200 includes one or more image sensors 202, accelerometers 204, gyroscopes 206, storage devices 208, input devices 210, displays 212, computing components 214, XR engines 224, image processing engines 226, rendering engines 228, and communication engines 230. It should be noted that... Figure 2 The components 202 to 230 shown are non-limiting examples provided for illustrative and explanatory purposes, and other examples may include those with... Figure 2 The components shown may be more, fewer, or different than those shown. For example, in some cases, the XR system 200 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radar, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 2One or more other software and / or hardware components not shown herein. While various components of the XR system 200 (such as image sensor 202) may be referred to in the singular herein, it should be understood that the XR system 200 may include multiple components discussed herein (e.g., multiple image sensors 202).

[0046] Display 212 may be or may include glass, screen, lens, projector and / or other display mechanism that allows users to see a real-world environment and also allows XR content to be overlaid on, overlapped with, mixed with or otherwise displayed on the real-world environment.

[0047] XR system 200 may include input device 210 or be able to communicate with such input device (wired or wireless). Input device 210 may include any suitable input device, such as a touchscreen, pen or other pointing device, keyboard, mouse, buttons or keys, microphone for receiving voice commands, gesture input device for receiving gesture commands, video game controller, steering wheel, joystick, set of buttons, trackball, remote control, any other input device discussed herein, or any combination thereof. In some cases, image sensor 202 may capture images that can be processed to interpret gesture commands.

[0048] The XR system 200 can also communicate with one or more other electronic devices (wired or wireless). For example, the communication engine 230 can be configured to manage connections and communicate with one or more electronic devices. In some cases, the communication engine 230 may correspond to... Figure 12 The communication interface is 1226.

[0049] In some implementations, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be part of the same computing device. For example, in some cases, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be integrated into an HMD, extended reality glasses, a smartphone, a laptop computer, a tablet computer, a gaming system, and / or any other computing device. However, in some implementations, the image sensor 202, accelerometer 204, gyroscope 206, storage device 208, display 212, computing component 214, XR engine 224, image processing engine 226, and rendering engine 228 may be part of two or more independent computing devices. For example, in some cases, some of the components 202-230 may be part of or implemented by a computing device, and the remaining components may be part of or implemented by one or more other computing devices. For example, such as in a discrete sensing XR system, XR system 200 may include a first device (e.g., an HMD such as...). Figure 1 The first device (XR device 102) includes a display 212, an image sensor 202, an accelerometer 204, a gyroscope 206, and / or one or more computing components 214. The XR system 200 may also include a second device, which includes additional computing components 214 (e.g., [missing information - likely a specific component name]). Figure 1 The accompanying device 104 may implement an XR engine 224, an image processing engine 226, a rendering engine 228, and / or a communication engine 230. In such examples, the second device may generate virtual content based on information or data (e.g., images, sensor data, such as measurements from accelerometer 204 and gyroscope 206) and may provide the virtual content to the first device for display at the first device. The second device may be or may include a smartphone, laptop computer, tablet computer, personal computer, gaming system, server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, or a mobile device acting as a server device), any other computing device, and / or combinations thereof.

[0050] Storage device 208 can be any storage device used for storing data. Furthermore, storage device 208 can store data from any component of the XR system 200. For example, storage device 208 can store data from image sensor 202 (e.g., image or video data), data from accelerometer 204 (e.g., measurements), data from gyroscope 206 (e.g., measurements), data from computing component 214 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from XR engine 224, data from image processing engine 226, and / or data from rendering engine 228 (e.g., output frames). In some examples, storage device 208 may include a buffer for storing frames processed by computing component 214.

[0051] Computing component 214 may be or may include a central processing unit (CPU) 216, a graphics processing unit (GPU) 218, a digital signal processor (DSP) 220, an image signal processor (ISP) 222, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). Computing component 214 may perform various operations such as image enhancement, computer vision, graphics rendering, extended reality operations (e.g., tracking, localization, pose estimation, map building, content anchoring, content rendering, prediction, etc.), image and / or video processing, sensor processing, recognition (e.g., text recognition, face recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, computing component 214 may implement (e.g., control, operate, etc.) an XR engine 224, an image processing engine 226, and a rendering engine 228. In other examples, computing component 214 may also implement one or more other processing engines.

[0052] Image sensor 202 may include any image and / or video sensor or capture device. In some examples, image sensor 202 may be part of a multi-camera assembly, such as a dual-camera assembly. Image sensor 202 may capture image and / or video content (e.g., raw image and / or video data), which may then be processed by computing component 214, XR engine 224, image processing engine 226, and / or rendering engine 228, as described herein.

[0053] In some examples, image sensor 202 may capture image data and may generate an image (also referred to as a frame) based on that image data and / or provide the image data or frame to XR engine 224, image processing engine 226, and / or rendering engine 228 for processing. The image or frame may include a video frame in a video sequence or a still image. The image or frame may include an array of pixels representing a scene. For example, the image may be: a red-green-blue (RGB) image with red, green, and blue color components per pixel; a lightness, redness, and blueness (YCbCr) image with a lightness component and two chromaticity (redness and blueness) components per pixel; or any other suitable type of color or monochrome image.

[0054] In some cases, image sensor 202 (and / or other cameras of XR system 200) may be configured to also capture depth information. For example, in some implementations, image sensor 202 (and / or other cameras) may include an RGB depth (RGB-D) camera. In some cases, XR system 200 may include one or more depth sensors (not shown) that are separate from image sensor 202 (and / or other cameras) and can capture depth information. For example, such depth sensors may acquire depth information independently of image sensor 202. In some examples, depth sensors may be physically mounted in the same general location or orientation as image sensor 202, but may operate at a different frequency or frame rate than image sensor 202. In some examples, depth sensors may take the form of a light source that projects a structured or textured light pattern (which may include one or more narrowband lights) onto one or more objects in a scene. Depth information can then be obtained by utilizing the geometric deformation of the projected pattern caused by the surface shape of the objects. In one example, depth information may be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).

[0055] The XR system 200 may also include other sensors among its one or more sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 204), one or more gyroscopes (e.g., gyroscope 206), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other position-related information to the computing component 214. For example, accelerometer 204 may detect the acceleration of the XR system 200 and may generate an acceleration measurement based on the detected acceleration. In some cases, accelerometer 204 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward) that can be used to determine the position or attitude of the XR system 200. Gyroscope 206 may detect and measure the orientation and angular velocity of the XR system 200. For example, gyroscope 206 may be used to measure the pitch, roll, and yaw of the XR system 200. In some cases, gyroscope 206 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 202 and / or the XR engine 224 may use measurements obtained by the accelerometer 204 (e.g., one or more translation vectors) and / or measurements obtained by the gyroscope 206 (e.g., one or more rotation vectors) to calculate the attitude of the XR system 200. As previously mentioned, in other examples, the XR system 200 may also include other sensors such as an inertial measurement unit (IMU), a magnetometer, gaze and / or eye-tracking sensors, machine vision sensors, smart scene sensors, voice recognition sensors, impact sensors, vibration sensors, position sensors, tilt sensors, etc.

[0056] As described above, in some cases, one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or orientations of the XR system 200. In some examples, one or more sensors may output measurement information associated with the capture of images by the image sensor 202 (and / or other cameras of the XR system 200) and / or depth information obtained using one or more depth sensors of the XR system 200.

[0057] XR engine 224 can use the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs and / or other sensors) to determine the pose of XR system 200 (also referred to as head pose) and / or the pose of image sensor 202 (or other cameras of XR system 200). In some cases, the pose of XR system 200 and the pose of image sensor 202 (or other cameras) can be the same. The pose of image sensor 202 refers to the pose of image sensor 202 relative to (e.g., relative to...) Figure 1The camera pose is determined by the positioning and orientation of the field of view (110) reference frame. In some implementations, the camera pose can be determined for 6 degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame such as the image plane) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera pose can be determined for 3 degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).

[0058] In some cases, a device tracker (not shown) may use measurements from one or more sensors and image data from image sensor 202 to track the pose (e.g., 6DoF pose) of the XR system 200. For example, the device tracker may fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of the XR system 200 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of the XR system 200, the device tracker may generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map for that scene. 3D map updates may include, for example, but not limited to, new or updated features and / or features or landmarks associated with the scene and / or the 3D map of that scene, location updates identifying or updating the positioning of the XR system 200 within the scene and the 3D map of that scene, etc. The 3D map provides a digital representation of the scene in the real / physical world. In some examples, 3D maps can anchor location-based objects and / or content to real-world coordinates and / or objects. XR system 200 can use mapped scenes (e.g., scenes in the physical world represented by a 3D map and / or scenes associated with that 3D map) to merge the physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0059] In some aspects, computing component 214 may use a visual tracking solution to determine and / or track the pose of image sensor 202 and / or the XR system 200 as a whole, based on images captured by image sensor 202 (and / or other cameras of XR system 200). For example, in some examples, computing component 214 may use computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques to perform tracking. For example, computing component 214 may perform SLAM or may communicate (wired or wirelessly) with a SLAM engine (not shown). SLAM refers to a class of techniques that create a map of an environment (e.g., a map of the environment modeled by XR system 200) while tracking the pose of the camera (e.g., image sensor 202) and / or XR system 200 relative to that map. This map may be called a SLAM map and may be three-dimensional (3D). SLAM technology can be performed using color or grayscale image data captured by image sensor 202 (and / or other cameras of XR system 200) and can be used to generate an estimate of 6DoF attitude measurement of image sensor 202 and / or XR system 200. Such SLAM technology configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 204, gyroscope 206, one or more IMUs and / or other sensors) can be used to estimate, correct, and / or otherwise adjust the estimated attitude.

[0060] Improving the efficiency of processing image data captured by image sensor 202 can be beneficial. For example, image sensor 202 can capture raw image data, and ISP 222 (which can be implemented on HMD and / or an accessory device) can process the raw image data. Reducing the size of the captured image data (e.g., by reducing the resolution of the non-ROI portion of the image data) to save power, bandwidth, and / or processing time within XR system 200 can be beneficial to the operation of XR system 200.

[0061] Figure 3 This is a block diagram illustrating an example architecture of an image processing system 300 according to various aspects of this disclosure. The image processing system 300 includes various components for capturing and processing images, such as an image of scene 306. The image processing system 300 can capture image frames (e.g., still images or video frames). In some cases, a lens 308 and an image sensor 318 (which may include an analog-to-digital converter (ADC)) may be associated with an optical axis. In one exemplary example, both the photosensitive area of ​​the image sensor 318 (e.g., a photodiode) and the lens 308 may be centered on the optical axis.

[0062] The image processing system 300 may be part of or implemented by a single computing device or multiple computing devices. In some examples, the image processing system 300 may be part of electronic devices (or multiple electronic devices), such as camera systems (e.g., digital cameras, IP cameras, video cameras, security cameras, etc.), telephone systems (e.g., smartphones, cellular phones, conferencing systems, etc.), laptops or notebook computers, tablet computers, set-top boxes, smart TVs, display devices, game consoles, XR devices (e.g., HMDs, smart glasses, etc.), IoT (Internet of Things) devices, smart wearable devices, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. For example, the image processing system 300 may be... Figure 1 In the XR system 100 or Figure 2 This is implemented in the XR system 200. For example, the image capture device 302 can be... Figure 2 An example of an image sensor 202, and an image processing device 304 may be used. Figure 2 It is implemented in the computing component 214.

[0063] In some examples, the lens 308 of the image processing system 300 faces the scene 306 and receives light from the scene 306. The lens 308 bends the incident light from the scene toward the image sensor 318. The light received by the lens 308 then passes through the aperture of the image processing system 300. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 310. In other cases, the aperture may have a fixed size.

[0064] One or more control mechanisms 310 may control exposure, focus, and / or zoom based on information from image sensor 318 and / or image processor 324. In some cases, one or more control mechanisms 310 may include multiple mechanisms and components. For example, control mechanism 310 may include one or more exposure control mechanisms 312, one or more focus control mechanisms 314, and / or one or more zoom control mechanisms 316. One or more control mechanisms 310 may also include, in addition to Figure 3 Additional control mechanisms beyond those illustrated herein. For example, in some cases, one or more control mechanisms 310 may include controls for controlling analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0065] The focus control mechanism 314 of the control mechanism 310 can obtain focus settings. In some examples, the focus control mechanism 314 stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 314 can adjust the position of the lens 308 relative to the position of the image sensor 318. For example, based on the focus settings, the focus control mechanism 314 can adjust the focus by moving the lens 308 closer to or further away from the image sensor 318 via an actuating motor or servo system (or other lens mechanism). In some cases, the image processing system 300 may include additional lenses. For example, the image processing system 300 may include one or more microlenses on each photodiode of the image sensor 318. These microlenses can each bend light received from the lens 308 toward the corresponding photodiode before the light reaches the photodiode.

[0066] In some examples, focus settings may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. Focus settings may be determined using control mechanism 310, image sensor 318, and / or image processor 324. Focus settings may be referred to as image capture settings and / or image processing settings. In some cases, lens 308 may be fixed relative to the image sensor and focus control mechanism 314.

[0067] Exposure control mechanism 312 of control mechanism 310 can obtain exposure settings. In some cases, exposure control mechanism 312 stores exposure settings in a memory register. Based on the exposure settings, exposure control mechanism 312 can control the aperture size (e.g., aperture size or aperture value), the duration of aperture opening (e.g., exposure time or shutter speed), the duration of light collection by the sensor (e.g., exposure time or electronic shutter speed), the sensitivity of image sensor 318 (e.g., ISO speed or film speed), the analog gain applied by image sensor 318, or any combination thereof. Exposure settings may be referred to as image capture settings and / or image processing settings.

[0068] The zoom control mechanism 316 of the control mechanism 310 can obtain zoom settings. In some examples, the zoom control mechanism 316 stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 316 can control the focal length of an assembly (lens assembly) of lens elements including lens 308 and one or more additional lenses. For example, the zoom control mechanism 316 can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 308) that first receives light from scene 306, where the light then passes through a focusing zoom system between the focusing lens (e.g., lens 308) and image sensor 318 before reaching image sensor 318. In some cases, a focusing zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference between them), with a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, zoom control mechanism 316 moves one or more lenses in the focusing zoom system, such as a negative lens and one or both positive lenses. In some cases, zoom control mechanism 316 can control zoom by capturing images from an image sensor (e.g., including image sensor 318) among a plurality of image sensors at a zoom setting corresponding to the zoom setting. For example, image processing system 300 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, zoom control mechanism 316 may capture images from the corresponding sensor based on the selected zoom setting.

[0069] Image sensor 318 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by image sensor 318. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters, and thus light matching the color of the filter covering the photodiode can be measured. Various color filter arrays can be used, such as, for example, and not limited to, Bayer color filter arrays, four-color filter arrays (QCFA), and / or any other color filter array.

[0070] In some cases, image sensor 318 may optionally or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, UV cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). Image sensor 318 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed relative to one or more control mechanisms in control mechanism 310 may be alternatively or additionally included in image sensor 318. The image sensor 318 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0071] Image processor 324 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 328), one or more host processors (including host processor 326), and / or related to Figure 12 The computing device architecture 1200 may include one or more processors of any other type discussed. The host processor 326 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 324 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes the host processor 326 and the ISP 328. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 330), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 330 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industry Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer ports or interfaces), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 326 may communicate with image sensor 318 using the I2C port, and ISP 328 may communicate with image sensor 318 using the MIPI port.

[0072] Image processor 324 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 324 can store image frames and / or processed images in random access memory (RAM) 320, read-only memory (ROM) 322, cache, memory unit, another storage device, or some combination thereof.

[0073] Various input / output (I / O) devices 332 may be connected to the image processor 324. I / O devices 332 may include a display screen, keyboard, keypad, touchscreen, touchpad, touch-sensitive surface, printer, any other output device, any other input device, or any combination thereof. In some cases, text may be entered into the image processing device 304 via the physical keyboard or keypad of the I / O device 332, or via the virtual keyboard or keypad of the touchscreen of the I / O device 332. I / O devices 332 may include one or more ports, jacks, or other connectors that enable wired connections between the image processing system 300 and one or more peripheral devices, through which the image processing system 300 may receive data from and / or send data to one or more peripheral devices. I / O devices 332 may include one or more wireless transceivers that enable wireless connections between the image processing system 300 and one or more peripheral devices, through which the image processing system 300 may receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 332 discussed earlier, and they can be considered I / O devices 332 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.

[0074] In some cases, the image processing system 300 may be a single device. In other cases, the image processing system 300 may be two or more independent devices, including an image capture device 302 (e.g., a camera) and an image processing device 304 (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 302 and the image processing device 304 may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 302 and the image processing device 304 may be disconnected from each other.

[0075] like Figure 3 As shown, the vertical dashed line will Figure 3The image processing system 300 is divided into two parts, namely image capture device 302 and image processing device 304. Image capture device 302 includes a lens 308, a control mechanism 310, and an image sensor 318. Image processing device 304 includes an image processor 324 (including an ISP 328 and a host processor 326), RAM 320, ROM 322, and I / O devices 332. In some cases, certain components illustrated in image capture device 302 (such as ISP 328 and / or host processor 326) may be included in image capture device 302. In some examples, image processing system 300 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof).

[0076] Although the image processing system 300 is shown as including certain components, those skilled in the art will understand that the image processing system 300 may include more than [other components]. Figure 3 The components shown herein are additional components. Components of the image processing system 300 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image processing system 300 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image processing system 300.

[0077] In some examples, Figure 12 The computing device architecture 1200 shown and further described below may include an image processing system 300, an image capture device 302, an image processing device 304, or a combination thereof.

[0078] Improving the efficiency of processing image data captured by image capture device 302 may be beneficial. For example, image capture device 302 can capture image data, and image processor 324 can process the image data. Reducing the size of the captured image data (e.g., by reducing the resolution of the region of interest (ROI) portion of the image data) to save power, bandwidth, and / or processing time within image processing system 300 may be beneficial to the operation of image processing system 300.

[0079] Figure 4 This is an illustration of an example system 400 that can be used to efficiently process image data according to various aspects of this disclosure. For example, image sensor 404 may capture image data 406 and image data 408. Both image sensor 404 and image data 406 may represent scene 402. Image data 406 may represent a region of interest (ROI) of scene 402 and may have a higher resolution than image data 408. Image data 408 may represent a larger field of view of scene 402 and may have a lower resolution than image data 406. System 400 may process image data 406 and image data 408 at an image signal processor (ISP) 410, a graphics processing unit (GPU) 412, and / or a data processing unit (DPU) 414. By processing image data 406 and image data 408 at the ISP 410, GPU 412 and / or DPU 414 (instead of processing an image representing the field of view of image data 408 at the resolution of image data 406), system 400 can save power, bandwidth and / or processing time.

[0080] More specifically, light from scene 402 can be focused onto image sensor 404 (e.g., through a lens). Image sensor 404 can generate image data (e.g., image data 406 and image data 408) representing the field of view of image sensor 404 of scene 402. In a conventional system, the image sensor can generate a single image frame representing the full field of view of the scene at a single resolution (e.g., at the highest resolution of the image sensor). As an example, in a conventional system, the image sensor can consist of 16 million individual light sensors. The image sensor can generate an image frame with a single resolution, which can be 16 megabytes in size (e.g., one byte for each light sensor in the individual light sensors). Thus, a conventional system can generate a 16-megabyte full-frame image. In contrast, image sensor 404 can generate image data 406 representing the ROI of scene 402 (and having a first resolution) and image data 408 representing the area surrounding the ROI (and having a second, lower resolution). For example, image data 406 can represent a quarter of the field of view of image sensor 404 of scene 402 at full resolution. For example, if image sensor 404 consists of 16 million individual light sensors, then image data 406 could be a 4-megabyte image representing a quarter of the field of view of image sensor 404 of scene 402. Furthermore, image data 408 could represent the full field of view of image sensor 404 of scene 402 at a quarter of full resolution. For example, image data 408 could be a 4-megabyte image representing the full field of view of image sensor 404 of scene 402. In such cases, image data 408 could include one byte of data from every four individual light sensors. For example, intensity data from the group of four light sensors could be downsampled or averaged to reduce the resolution of image data 408. According to such an example, system 400 could process (e.g., at ISP 410 and above) 8 megabytes of image data (e.g., image data 406 (which could be a 4-megabyte image) and image data 408 (which could be a 4-megabyte image)). In contrast, a typical example system could process 16 megabytes of image data. In this way (e.g., by capturing a full-resolution image of the ROI and a quarter-resolution full-frame image), the system 400 can reduce the image data processing load of the system 400 compared to conventional image capture and processing systems.

[0081] The scaling provided in the above example is illustrative. Other scaling may be used in other cases. For example, image data 406 may represent any portion of the field of view of image sensor 404 of scene 402 (e.g., image data 406 may represent half, one-third, one-fifth, one-sixth, etc., of the field of view of image sensor 404 of scene 402). Additionally or alternatively, image data 408 may have any resolution smaller than the full resolution that image sensor 404 is capable of having. For example, image data 408 may be 75%, 50%, 25%, 12.5%, etc., of the full resolution of image sensor 404. Additionally or alternatively, in some cases, image data 408 may exclude the ROI. For example, a portion of the field of view of scene 402 captured by image data 406 may be omitted from image data 408. Additionally or alternatively, in some cases, image sensor 404 may capture more than two images of scene 402. For example, image sensor 404 can capture a first image of the ROI (e.g., image data 406) at 100% resolution, a second image of the peripheral region (e.g., the edge of the field of view around the ROI but not extending into the scene 402, such as a ring or frame around the ROI) at 50% resolution, and a third image of the remainder of the field of view of scene 402 at 25% resolution.

[0082] System 400 can process image data 406 and image data 408 at ISP 410. ISP 410 can be or may include any number of individual ISPs. Figure 2 ISP 222, Figure 3 Image processor 324 and / or Figure 3 ISP 328 can be an example of ISP 410. ISP 410 may include a front-end portion that receives image data 406 and image data 408 from image sensor 404 and processes image data 406 and image data 408 upon receipt (e.g., processing line by line upon receipt, or processing portion by portion upon receipt). ISP 410 may also include a back-end portion and memory. The front-end portion may store processed image data in memory, and the back-end portion may read processed image data from memory and further process the read image data. By reducing the total size of the image data to be processed by ISP 410 (by reducing the frame size of image data 406 compared to a full-frame image and reducing the resolution of image data 408 compared to a full-resolution image), system 400 can save power, bandwidth, and / or processing time at ISP 410.

[0083] In some cases, system 400 may process image data 406 and image data 408 at GPU 412 (e.g., after processing at ISP 410). Figure 2 GPU 218 can be an example of GPU 412. Similar to the description regarding ISP 410, system 400 can save power, bandwidth, and / or processing time at GPU 412 by reducing the total size of the image data to be processed by GPU 412. GPU 412 is optional in system 400. For example, in some cases, system 400 may omit GPU 412. In other cases, ISP 410 may be omitted, and GPU 412 may process image data 406 and image data 408, and power, bandwidth, and / or processing time may be saved based on the reduced total size of the processed image data.

[0084] Additionally or alternatively, in some cases, system 400 may process image data 406 and image data 408 at DPU 414 (e.g., after being processed at ISP 410 and / or GPU 412). Figure 2 The computing component 214 may include an instance of DPU 414. Similar to that described with respect to ISP 410, system 400 can save power, bandwidth, and / or processing time at DPU 414 by reducing the total size of the image data to be processed by DPU 414. Similar to ISP 410, GPU 412 is optional in system 400. For example, in some cases, system 400 may omit DPU 414. In other cases, ISP 410 and / or GPU 412 may be omitted, and DPU 414 may process image data 406 and image data 408, and power, bandwidth, and / or processing time may be saved based on the reduced total size of the processed image data.

[0085] In some cases, DPU 414 may generate image data 420 (which may be a composite image) including data from image data 406 and data from image data 408. For example, DPU 414 may include a scaler 416 and a mixer 418. Scaler 416 may scale image data 408 (e.g., as processed by ISP 410 and / or GPU 412). For example, image data 408 may only have a fraction of the total number of pixels in a full-frame, full-resolution image (e.g., based on image data 408 captured with a fraction of the full-resolution capability of image sensor 404). Scaler 416 may enlarge image data 408 by increasing the number of pixels in image data 408. Mixer 418 may combine image data 406 (e.g., as processed by ISP 410 and / or GPU 412) with image data 408 (e.g., as enlarged by scaler 416). For example, mixer 418 can insert image data 406 into image data 408 (such as when zoomed by scaler 416). In addition, mixer 418 can blend pixels at the edges of image data 406 (e.g., making the interface between the edges of image data 406 and image data 408 less noticeable).

[0086] System 400 may display image data 420 (e.g., at a display 422), store image data 420 (e.g., for later display and / or processing), send image data 420 (e.g., for display, storage, and / or processing by another system or device), and / or process image data 420. Processing image data 420 may include using image data 420 to perform operations related to object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, and automation.

[0087] As mentioned above, it may be important to coordinate the operation of image sensor 404 (when generating image data 406 and image data 408) with the operation of ISP 410 (when processing image data 406 and image data 408). For example, ISP 410 may be configured to process image data (e.g., image data 406 and image data 408) when it receives image data from image sensor 404. For ISP 410 to process image data 406 and image data 408 correctly, it may be important to inform ISP 410 of certain information about image data 406 and image data 408. For example, some operations of ISP 410 may be based at least in part on the position of the ROI of image data 406 within a larger frame of image data 408. For example, lens shading correction operations may be based at least in part on the position of the pixel being processed within a frame. As an example, if the Region of Interest (ROI) is on the left side of an image frame (e.g., if image data 406 represents the portion on the left side of image data 408), then lens correction shading may operate differently for pixels on the left side of the ROI than for pixels on the right side of the ROI. For image data 406 and / or image data 408 to be processed correctly, it may be important for the ISP 410 to know the relationship between image data 406 and image data 408 (e.g., the position of the ROI on which image data 406 is based relative to image data 408).

[0088] Figure 5 This is a block diagram illustrating an example system 500 for processing image data. System 500 provides an ROI indicator 520 to an ISP 518, enabling the ISP 518 to correctly process image data 516. Image sensor 514 may be... Figure 4 An example of an image sensor 404. Image data 516 can be... Figure 4 Examples of image data 406 and / or image data 408. ISP 518 can be... Figure 4 Example of ISP 410.

[0089] System 500 includes a gaze engine 502 that can determine the ROI based on the viewer's gaze. For example, gaze engine 502 can receive data from a gaze tracking sensor and determine and / or predict the viewer's gaze. Gazing engine 502 can generate an ROI indicator 504 that indicates the ROI based on the viewer's gaze.

[0090] The gaze engine 502 may provide an ROI indicator 504 to the camera driver 506. The camera driver 506 may control at least some operations of the image sensor 514 and / or the ISP 518. For example, the camera driver 506 may determine and / or set image capture settings used by the image sensor 514 to capture image data 516. Additionally or alternatively, the camera driver 506 may determine and / or set image processing settings used by the ISP 518 to process the image data 516.

[0091] Camera driver 506 may provide ROI indicator 508 (which may be the same as ROI indicator 504, or may be a reformulated version of ROI indicator) to camera control interface 510. Camera control interface 510 may be an interface between camera driver 506 and image sensor 514. Camera control interface 510 may control the operation of image sensor 514 (e.g., in the direction of camera driver 506 and / or more directly than camera driver 506). Camera control interface 510 may provide ROI indicator 512 (which may be the same as ROI indicator 508, or may be a reformulated version of ROI indicator) to image sensor 514.

[0092] Image sensor 514 can generate image data 516 based on ROI indicator 512. For example, ROI indicator 512 can indicate an ROI within the field of view of image sensor 514. Image sensor 514 can capture image data representing the ROI at a first resolution (e.g., at the highest resolution of image sensor 514). Additionally, image sensor 514 can capture image data outside the ROI and within the field of view of image sensor 514 at a second, lower resolution. For example, image sensor 514 can capture image data based on ROI indicator 512. Figure 4 Image data 406 and image data 408. Image sensor 514 can provide image data 516 (including image data captured based on ROI indicator 512) to ISP 518.

[0093] The ISP 518 can process image data 516 to generate image data 522. The ISP 518 may include a front-end portion that processes the image data 516 upon receipt from the image sensor 514. The camera driver 506 may provide the ISP 518 with a ROI indicator 520 (which may be the same as ROI indicator 504, or may be a reformulated version of ROI indicator 504). The ISP 518 can process the image data 516 based on the ROI indicator 520. For example, the ISP 518 may perform operations related to lens shading correction, and the ROI indicator 520 can be used to determine how to perform lens shading correction on the image data 516.

[0094] The camera driver 506, which provides the ROI indicator 520 to the ISP 518, is an example of informing the ISP 518 of the ROI so that the ISP 518 can process image data 516 based on the ROI. However, there are inherent challenges in the example of system 500. Due to the multi-threaded nature of the software and the fact that many tasks can run in parallel (e.g., in the ISP 518), one inherent challenge in the example of system 500 is ensuring that image data 516 matches the configuration of the ISP 518 (e.g., based on the ROI indicator 520). This challenge is particularly significant when image sensor 514 captures image data 516 at a high frame rate. For example, delays in capturing and / or processing image data 516 (e.g., at the ISP engine of image sensor 514 and / or ISP 518) may cause image data 516 to become out of sync with the ROI indicator 520. For example, the ISP 518 may adjust its image processing settings based on the ROI indicator 520 that does not correspond to the image data 516 arriving at the ISP 518.

[0095] Figure 6A This is a block diagram illustrating an example system 600A for efficiently processing image data 610 according to various aspects of the present disclosure. System 600A includes an image sensor 602 that can generate image data 610 based on a region of interest (ROI). Furthermore, system 600A includes an image signal processor (ISP) 612 that can process image data 610 based on the ROI.

[0096] For example, image sensor 602 may receive ROI indicator 604. ROI indicator 604 may be based on the viewer's gaze (e.g., via...). Figure 5 The image sensor 602 generates the image data 610 based on the gaze engine 502. The ROI indicator 604 can be an indication of the ROI within the field of view of the image sensor 602. The image sensor 602 can generate image data 610 based on the ROI. The image sensor 602 can be used with... Figure 5Image sensor 514 similar to and / or capable of performing the same functions as Figure 5 The image sensor 514 operates essentially the same. For example, image sensor 602 can generate a first image of the ROI at a first resolution (e.g., the full resolution of image sensor 602) and a second image of the area surrounding the ROI at a second, lower resolution. For example, image sensor 602 can generate... Figure 4 Image data 406 and image data 408. Image data 610 can be used with... Figure 5 Image data 516 (e.g., including image data 406 and image data 408) is the same as or may be the same as... Figure 5 The image data 516 is basically similar.

[0097] However, unlike image sensor 514, image sensor 602 can provide ROI indicator 608 to ISP 612. ROI indicator 608 can be an indication of the position of the ROI relative to the field of view of image sensor 602. For example, ROI indicator 608 can be an indication of the position of the ROI within the field of view of image sensor 602. Additionally or alternatively, ROI indicator 608 can describe the relationship between image frames. For example, in the case where image data 610 includes image data 406 and image data 408, ROI indicator 608 can describe the relationship between image data 406 and image data 408. In some cases, ROI indicator 608 can be the same as or substantially similar to ROI indicator 604. In other cases, ROI indicator 608 can include the same information as ROI indicator 604, but can be formatted differently.

[0098] ISP 612 can process image data 610 based on an ROI as indicated by ROI indicator 608. ISP 612 can perform one or more operations related to Bad Pixel Correction (BPC), lens correction, lens shading correction, phase detection pixel correction, demosaic, lateral chromatic aberration correction, Bayer filtering, adaptive Bayer filtering, tone mapping, noise reduction, etc. ISP 612 can perform some of these operations based on an ROI as indicated by ROI indicator 608. For example, ISP 612 can generate a setting 614 based on the ROI and process image data 610 based on the setting 614. ISP 612 can generate and / or output image data 616 (which may be the processed image data 610). Additionally or alternatively, ISP 612 can output an ROI indicator 618 (which may be the same as ROI indicator 608 or a reformatted version of ROI indicator 608).

[0099] Figure 6BThis is a block diagram illustrating an example system 600B for efficiently processing image data 610 according to various aspects of this disclosure. System 600B may be... Figure 6A Examples of system 600A. For example, system 600B may be the same as system 600A, may be substantially similar to system 600A and / or may perform the same or substantially the same operations as system 600A, but system 600B may include examples of details based on an example aspect of system 600A.

[0100] For example, according to an example aspect of system 600B, image sensor 602 may provide ROI indicator 608 and image data 610 to ISP 612 in packet 606. For example, image sensor 602 may generate packet 606 including image data 610 and ROI indicator 608, and provide packet 606 to ISP 612. Image sensor 602 may be coupled to ISP 612 at an interface (e.g., Mobile Industry Processor Interface (MIPI)). Therefore, packet 606 may be a MIPI packet.

[0101] According to an example aspect of system 600B, ISP 612 may include a Camera Serial Interface (CSI) decoder (CSID 620). CSID 620 may receive image data 610 and ROI indicator 608 (e.g., in packet 606) from image sensor 602. CSID 620 may parse packet 606 (e.g., to extract ROI indicator 608, which may include position and size information) and provide image data 624 (which may be a decoded version of the image data) and ROI indicator 622 (which may be the same as ROI indicator 608 or may be a reformatted version of ROI indicator 608) to ISP engine 626 of ISP 612. To reduce cabling costs, the control plane may be interleaved with the data plane.

[0102] Additionally or alternatively, based on the example aspects of System 600B, ISP 612 may include multiple ISP engines. ISP 612 may include any number of ISP engines. For simplicity, regarding... Figure 6B Two ISP engines are illustrated and described. Specifically, ISP 612 includes ISP engine 626 and ISP engine 634. However, ISP 612 may include any number of ISP engines (e.g., one, three, four, or more). An interface may exist between each of these ISP engines (e.g., between ISP engine 626 and ISP engine 634). Such an interface may be internal to IPS 612 and may be flexible.

[0103] Each ISP engine in ISP 612 can receive image data from a previous ISP engine (or from CSID 620), process the received image data, and provide the processed image data to a subsequent ISP engine (or at the output of ISP 612). Furthermore, each of these ISP engines can receive a ROI indicator from a previous ISP engine (or from CSID 620) and process the image data based on the ROI indicator. For example, each of these ISP engines can receive the ROI indicator, determine the corresponding image processing settings for the ISP engine, and process the image data based on the corresponding image processing settings. Additionally, each of these ISP engines can provide the ROI indicator to a subsequent ISP engine (or at the output of ISP 612). For example, ISP engine 626 can receive ROI indicator 622 and image data 624 from CSID 620. ISP engine 626 can generate settings 628 based on ROI indicator 622 and process image data 624 based on settings 628. Furthermore, ISP engine 626 can provide ROI indicator 630 (which may be the same as ROI indicator 622) and image data 632 (which may be image data 624 as processed by ISP engine 626) to ISP engine 634. As another example, ISP engine 634 can receive ROI indicator 630 and image data 632 from ISP engine 626. ISP engine 634 can generate setting 636 based on ROI indicator 630 and process image data 632 based on setting 636. Additionally, ISP engine 634 can provide ROI indicator 618 (which may be the same as ROI indicator 630) and image data 616 (which may be image data 632 as processed by ISP engine 634) at the output of ISP 612 (or at a corresponding output).

[0104] Figure 7 This is an illustration of an example group 700 that may include image data and ROI indicators according to various aspects of this disclosure. Group 700 may be... Figure 6B Example of group 606. Group 700 could be a MIPI group.

[0105] Packet 700 may include a start-of-frame (FOF) 702, which may include one or more bits to indicate the start of packet 700. Following FOF 702 may be a header 704. Header 704 may include a data identifier, a word count field, and an 8-bit error correction code (ECC). Header 704 may conform to the MIPI protocol. Additionally, header 704 may include a Region of Interest (ROI) indicator. The ROI indicator may be associated with image data in payload 706. The ROI indicator may indicate the location and / or size of an ROI in the image data. Following header 704 may be payload 706. Payload 706 may include image data. Following payload 706 may be a trailer 708, which may include a checksum or cyclic redundancy check (CRC). Following trailer 708 may be end-of-frame (FOF) 710, which may include one or more bits to indicate the end of packet 700. According to some terminology, each of the start of frame 702, header 704, payload 706, trailer 708, and end of frame 710 can be individually referred to as a packet.

[0106] Figure 8 This is an illustration of two example groups 800 that may include image data and ROI indicators according to various aspects of this disclosure. Groups 802 and 822 may be collectively referred to as group 800. Groups 800 may jointly provide Figure 6A The transmission of ROI indicator 608 and image data 610.

[0107] Each of groups 802 and 822 can be with Figure 7 The packet 700 is substantially similar. For example, packet 802 may include start of frame 804, header 806, payload 808, trailer 810, and end of frame 812, and packet 822 may include start of frame 824, header 826, payload 828, trailer 830, and end of frame 832. Start of frame 804 and start of frame 824 may be the same as or substantially similar to start of frame 702. Payload 808 and payload 828 may be the same as or substantially similar to payload 706. End of frame 812 and end of frame 832 may be the same as or substantially similar to end of frame 710. According to some terminology, each of start of frame 804, header 806, payload 808, trailer 810, end of frame 812, start of frame 824, header 826, payload 828, trailer 830, and end of frame 832 may be individually referred to as a packet.

[0108] Headers 806 and 826 may be similar to header 704. However, headers 806 and 826 may or may not include the ROI indicator present in header 704. Tail headers 810 and 830 may be similar to tail header 708. However, tail headers 810 and 830 may include ROI indicators. For example, each group in group 800 may include an ROI indicator for subsequent image frames in the tail of the respective group. For example, if group 802 precedes group 822, group 802 in tail header 810 may include an ROI indicator indicating the location and / or size of the ROI of the image data in payload 828. Furthermore, group 822 in tail header 830 may include an ROI indicator indicating the location and / or size of the ROI of the image data in subsequent frames. Figure 8 not shown in the example).

[0109] Figure 9 This is a flowchart of a process 900 for efficiently processing image data according to various aspects of this disclosure. One or more operations of process 900 may be performed by a computing device (or apparatus) or a component of such computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform process 900. One or more operations of process 900 may be implemented as software components that execute and run on one or more processors.

[0110] At box 902, a computing device (or one or more components thereof) (e.g., an image signal processor (ISP)) may receive image data and indications of regions of interest (ROIs) from an image sensor. For example, the ISP 612 of FIG. 6 may receive packet 606 of FIG. 6, which may include image data 610 and ROI indication 608.

[0111] In some aspects, the computing device (or one or more components thereof) (e.g., the ISP) may include a Camera Serial Interface (CSI) decoder. The CSI decoder may: receive packets from an image sensor including image data and indications of the Region of Interest (ROI); parse the packets; and provide the image data and indications of the ROI to a subsequent ISP engine in one or more ISP engines. For example, ISP 612 may include CSID 620. CSID 620 may receive packet 606 from image sensor 602. Packet 606 may include an ROI indicator 608 and image data 610. CSID 620 may parse packet 606 and provide an ROI indicator 622 and image data 624 to ISP engine 626.

[0112] In some aspects, a computing device (or one or more components thereof) (e.g., an ISP) may receive image data and an indication of ROI in a packet, and the ISP may parse the indication of ROI from the packet header. For example, ISP 612 may receive ROI indicator 608 and image data 610 in packet 606, and may parse the indication of ROI from the header of packet 606 (e.g., from...). Figure 7 The header 704 of packet 700 is parsed to resolve the ROI indicator 608. In some respects, the packet may be or may include a Mobile Industry Processor Interface (MIPI) packet. For example, packet 606 may be a MIPI packet.

[0113] In some aspects, a computing device (or one or more components thereof) (e.g., an ISP) may receive an indication of a Region of Interest (ROI) in the tail of a first packet and parse the indication of the ROI from the tail of the first packet, wherein the ISP is configured to receive image data in the payload of a second packet. For example, ISP 612 may receive multiple packets 606 and may parse an ROI indicator 608 from a first packet of multiple packets 606 and image data 610 from a second packet of multiple packets 606. For example, ISP 612 may receive packet 800 and parse an ROI indicator 608 from the tail 810 of packet 802 and image data 610 from the payload 828 of packet 822.

[0114] In some respects, the indication of ROI can be a first indication of ROI. The image sensor (or one or more components thereof) of the computing device (e.g., coupled to the ISP) can: receive a second indication of ROI; generate image data based on the ROI; and provide the image data and the first indication of ROI to the image signal processor (ISP). For example, Figure 5The gaze engine 502 can determine a ROI indicator 504 and provide the ROI indicator 504 to a camera driver 506, which in turn can provide an ROI indicator 508 to a camera control interface 510, which in turn can provide an ROI indicator 512 to an image sensor 514. The image sensor 514 can capture image data 516 based on the ROI indicator 512 and can provide the image data 516 to an ISP 518. Additionally, the camera driver 506 can provide an ROI indicator 520 (which may correspond to ROI indicators 512 and 504) to the ISP 518. In some aspects, a computing device (or one or more components thereof) or a processor coupled thereto can determine the ROI based on data from a gaze tracking sensor. For example, the gaze engine 502 can determine the ROI indicator 504 based on a gaze detected by the gaze tracking sensor.

[0115] In some aspects, the image sensor may generate packets that include a second indication of the ROI in the header of the packet and include image data in the payload of the packet. For example, image sensor 602 may generate packet 606 and include an ROI indicator 608 in the header of packet 606 (e.g., in header 704 of packet 700). In some aspects, the packet may be or may include MIPI packets. In some aspects, the image sensor may: generate a first packet that includes a second indication of the ROI in the tail of the first packet; and generate a second packet that includes image data in the payload of the second packet. For example, image sensor 602 may generate a plurality of packets 606, and include an indicator 608 in a first packet of the plurality of packets 606, and include image data 610 in a second packet of the plurality of packets 606. For example, image sensor 602 may generate packet 800, and include an ROI indicator 608 in the tail 810 of packet 822, and include image data 610 in the payload 828 of packet 802.

[0116] In some aspects, the image sensor (or one or more components thereof) of the computing device (e.g., coupled to an ISP) can: generate a first portion of an image corresponding to a Region of Interest (ROI) at a first resolution; and generate a second portion of the image outside the ROI at a second resolution, wherein the first resolution is greater than the second resolution. For example, Figure 4 The image sensor 404 can generate image data 406 with a first resolution and image data 408 with a second resolution.

[0117] At box 904, a computing device (or one or more components thereof) (e.g., an ISP) may determine image processing settings for processing image data based on the ROI. For example, ISP 612 may determine settings 628 based on ROI indicator 608.

[0118] At box 906, a computing device (or one or more components thereof) (e.g., an ISP) may process image data based on image processing settings. For example, ISP engine 626 may process image data 624 based on settings 628.

[0119] In some aspects, a computing device (or one or more components thereof) (e.g., an ISP) may include one or more ISP engines. Each of the one or more ISP engines may: determine appropriate image processing settings based on the ROI; and process image data based on the appropriate image processing settings. For example, ISP 612 may include ISP engine 626 and ISP engine 634. ISP engine 626 may determine setting 628 based on ROI indicator 622, and may process image data 624 based on setting 628. Additionally or alternatively, ISP engine 634 may determine setting 636 based on ROI indicator 630, and may process image data 632 based on setting 636.

[0120] In some aspects, a computing device (or one or more components thereof) (e.g., an ISP) may include one or more ISP engines. Each of the one or more ISP engines may: receive image data and indications of ROI from an image sensor or from a previous ISP engine among the one or more ISP engines; and provide the image data and indications of ROI to a subsequent ISP engine among the one or more ISP engines or to the output of the ISP. For example, ISP 612 may include ISP engine 626 and ISP engine 634. ISP engine 626 may determine setting 628 based on ROI indication 622 and may process image data 624 based on setting 628. Additionally or alternatively, ISP engine 634 may determine setting 636 based on ROI indication 630 and may process image data 632 based on setting 636.

[0121] In some respects, a computing device (or one or more components thereof) (e.g., an ISP) may perform operations associated with at least one of the following: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; and / or noise reduction.

[0122] In some aspects, a computing device (or one or more components thereof) (e.g., an ISP) may include a first ISP engine and / or a second ISP engine. The first ISP engine may: determine a first image processing setting associated with a first ISP operation based on the ROI; and perform the first ISP operation based on the first image processing setting. The second ISP engine may: determine a second image processing setting associated with a second ISP operation based on the ROI; and perform the second ISP operation based on the second image processing setting. For example, ISP 612 may include ISP engine 626 and ISP engine 634. ISP engine 626 may determine setting 628 based on a first ISP operation of ISP engine 626, and may perform a first IPS operation on image data 624 based on setting 628. Additionally or alternatively, ISP engine 634 may determine setting 636 based on a second ISP operation of ISP engine 634, and may perform a second IPS operation on image data 632 based on setting 636. The first ISP operation and the second ISP operation may be associated with at least one of the following: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; or noise reduction.

[0123] In some respects, a computing device (or one or more components thereof) (e.g., an ISP) may process image data when it is received from an image sensor. For example, when image data 610 is received from image sensor 602, ISP 612 may process image data 610.

[0124] In some examples, as previously noted, the methods described herein (e.g., Figure 9 The process 900 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 2 Image sensor 202, Figure 2 Computing component 214 Figure 3 Image processing system 300 Figure 3 Image capture device 302, Figure 3 Image processing equipment 304 Figure 3 Image processor 324, Figure 4 System 400 Figure 4 Image sensor 404, Figure 4 ISP 410, Figure 5 System 500 Figure 5 Image sensor 514, Figure 5 ISP 518 Figure 6A System 600A, Figure 6B System 600B Figure 6A or Figure 6B Image sensor 602, Figure 6A or Figure 6B The ISP 612 can be used to execute these methods, or they can be executed by another system or device. In another example, these methods (e.g., Figure 9 One or more of the processes 900 and / or other methods described herein may be used by Figure 12 The computing device architecture 1200 shown is implemented wholly or partially. For example, it has Figure 12 The computing device of the computing device architecture 1200 shown may include Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 2 Image sensor 202, Figure 2 Computing component 214 Figure 3 Image processing system 300 Figure 3 Image capture device 302, Figure 3 Image processing equipment 304 Figure 3 Image processor 324, Figure 4 System 400 Figure 4 Image sensor 404, Figure 4 ISP 410, Figure 5 System 500 Figure 5 Image sensor 514, Figure 5 ISP 518 Figure 6A System 600A, Figure 6B System 600B Figure 6A or Figure 6B Image sensor 602, Figure 6A or Figure 6BThe ISP 612 is a component of, or included therein, and enables the operation of process 900 and / or other processes described herein. In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0125] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0126] Process 900 and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0127] Additionally, process 900 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0128] As noted above, various aspects of this disclosure may utilize machine learning models or systems.

[0129] Figure 10 This is an exemplary example of a neural network 1000 (e.g., a deep learning neural network) that can be used to implement machine learning-based feature segmentation, implicit neural representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, and / or automation. The neural network 1000 can be... Figure 5 An example of the gaze engine 502, or one that can be implemented. Figure 5 The gaze engine 502. Furthermore, the neural network 1000 can... Figure 4 Image data 420 Figure 5 Image data 522 Figure 6A and Figure 6B Image data 616 and / or ROI indicator 618 are used as input.

[0130] Input layer 1002 includes input data. In an exemplary example, input layer 1002 may include data representing an image of a user's eye, image data 420, image data 522, image data 616, and / or ROI indicator 618. Neural network 1000 includes multiple hidden layers 1006a, 1006b to 1006n. Hidden layers 1006a, 1006b to 1006n include "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. Neural network 1000 also includes an output layer 1004, which provides the output produced by the processing performed by hidden layers 1006a, 1006b to 1006n. In an exemplary example, output layer 1004 provides... Figure 5 The ROI indicator is 504.

[0131] The neural network 1000 may be or may include a multi-layer neural network with interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, the neural network 1000 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, the neural network 1000 may include a recurrent neural network, which may have loops that allow information to be carried across nodes when reading input.

[0132] Information can be exchanged between nodes through node-to-node interconnects between layers. Nodes in input layer 1002 can activate the node set in the first hidden layer 1006a. For example, as shown, each input node in input layer 1002 is connected to each node in the first hidden layer 1006a. Nodes in the first hidden layer 1006a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 1006b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable functions. The output of hidden layer 1006b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 1006n can activate one or more nodes in output layer 1004, at which the output is provided. In some cases, although a node in neural network 1000 (e.g., node 1008) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.

[0133] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 1000. Once the neural network 1000 is trained, it can be called a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about the nodes learned in the interconnection. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), thereby allowing the neural network 1000 to adapt to the input and learn as more and more data is processed.

[0134] The neural network 1000 can be pre-trained to process features from the data in the input layer 1002 using different hidden layers 1006a, 1006b to 1006n, so as to provide an output through the output layer 1004. In an example where the neural network 1000 is used to identify features in an image, the neural network 1000 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating features in the image (for feature segmentation machine learning systems) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the label of the image may be [0 0 1 0 0 0 0 0 0 0].

[0135] In some cases, the Neural Network 1000 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the Neural Network 1000 is trained well enough to accurately tune the weights of each layer.

[0136] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 1000. The weights are initially randomized before training the neural network 1000. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0137] As noted above, for the first training iteration of a neural network 1000, the output will likely include values ​​due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 1000 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function includes mean squared error (MSE), which is defined as... The loss can be set to equal to The value of .

[0138] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. The Neural Network 1000 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as...) can be calculated. ,in These are the weights at a specific layer, used to determine the weights that contribute the most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... ,in Indicates weight, This represents the initial weights, and This represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0139] Neural Network 1000 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 1000 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0140] Figure 11 This is an exemplary example of a Convolutional Neural Network (CNN) 1100. The input layer 1102 of the CNN 1100 includes data representing an image or frame. For example, the data may include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at that location in the array. Using the previous example from above, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 1104, an optional non-linear activation layer, a pooling hidden layer 1106, and a fully connected layer 1108 (which may be hidden) to obtain the output at the output layer 1110. Although... Figure 11 Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in the CNN 1100. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0141] The first layer of CNN 1100 can be a convolutional hidden layer 1104. The convolutional hidden layer 1104 analyzes the image data from the input layer 1102. Each node in the convolutional hidden layer 1104 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 1104 can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 1104. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of the filter. In an exemplary example, if the input image comprises a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 1104. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 1104 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.

[0142] The convolutional property of the convolutional hidden layer 1104 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 1104 may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 1104. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 1104. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thereby determining a sum value for each node of the convolutional hidden layer 1104.

[0143] The mapping from the input layer to the convolutional hidden layer 1104 is called an activation map (or feature map). An activation map includes node-specific values ​​representing the filter results at each location of the input volume. Activation maps may include arrays containing various sums of values ​​produced by the filter for each iteration over the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 1104 may include several activation maps to identify multiple features in the image. Figure 11 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 1104 can detect three different types of features, each of which is detectable across the entire image.

[0144] In some examples, nonlinear hidden layers can be applied after convolutional hidden layer 1104. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can add nonlinearity to CNN 1100 without affecting the receptive field of convolutional hidden layer 1104.

[0145] A pooling hidden layer 1106 can be applied after the convolutional hidden layer 1104 (and, when used, after the non-linear hidden layer). The pooling hidden layer 1106 is used to simplify the information in the output of the convolutional hidden layer 1104. For example, the pooling hidden layer 1106 can take each activation map output from the convolutional hidden layer 1104 and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 1106 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 1104. Figure 11 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 1104.

[0146] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 1104. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node being a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 1104, the output from pooling hidden layer 1106 will be an array of 12×12 nodes.

[0147] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0148] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of the image and discard the exact location information. This can be done without affecting the results of feature detection because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of the CNN 1100.

[0149] The final connection in the network is a fully connected layer that connects each node from the pooling hidden layer 1106 to each output node in the output layer 1110. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 1104 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 1106 comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region on each of the three feature maps. Extending this example, the output layer 1110 may comprise ten output nodes. In this example, each node of the 3×12×12 pooling hidden layer 1106 is connected to each node of the output layer 1110.

[0150] The fully connected layer 1108 takes the output of the previous pooling hidden layer 1106 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 1108 can determine the high-level features most relevant to a particular class and may include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 1108 and the pooling hidden layer 1106 can be computed to obtain the probabilities for different classes. For example, if the CNN 1100 is used to predict that the object in the image is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0151] In some examples, the output from output layer 1110 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes from which CNN 1100 must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector representing objects of ten different classes is [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0152] Figure 12 An example computing device architecture 1200 is illustrated, illustrating example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device of a vehicle), or other devices. For example, computing device architecture 1200 may include, implement, or be included in any of the following: Figure 1 XR system 100, Figure 1 XR equipment 102, Figure 1 Supporting equipment 104 Figure 2 XR system 200, Figure 2 Computing component 214 Figure 3 Image processing system 300 Figure 3 Image processor 324, Figure 4 System 400 Figure 5 System 500 Figure 6ASystem 600A and / or Figure 6B The system is 600B.

[0153] The components of computing device architecture 1200 are shown to communicate electrically with each other using a connection 1212, such as a bus. Example computing device architecture 1200 includes a processing unit (CPU or processor) 1202 and a computing device connection 1212 that couples various computing device components, including computing device memories 1210 (such as read-only memory (ROM) 1208 and random access memory (RAM) 1206), to the processor 1202.

[0154] The computing device architecture 1200 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1202. The computing device architecture 1200 may copy data from memory 1210 and / or storage device 1214 to cache 1204 for fast access by the processor 1202. In this way, the cache can provide performance improvements by avoiding latency for the processor 1202 while waiting for data. These and other modules may control or be configured to control the processor 1202 to perform various actions. Other computing device memory 1210 may also be available. Memory 1210 may include various different types of memory with different performance characteristics. The processor 1202 may include any general-purpose processor and hardware or software services configured to control the processor 1202 (such as services 11216, 1218, and 31220 stored in storage device 1214), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1202 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.

[0155] To enable user interaction with the computing device architecture 1200, input device 1222 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1224 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some instances, multi-mode computing devices allow users to provide multiple types of input to communicate with computing device architecture 1200. Communication interface 1226 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0156] Storage device 1214 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 1206, read-only memory (ROM) 1208, and hybrid forms thereof. Storage device 1214 may include services 1216, 1218, and 1220 for controlling processor 1202. Other hardware or software modules are envisioned. Storage device 1214 may be connected to computing device connection 1212. In one aspect, a hardware module performing a particular function may include software components for performing that function stored in a computer-readable medium connected to necessary hardware components such as processor 1202, connection 1212, output device 1224, etc.

[0157] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0158] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0159] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.

[0160] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0161] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. Although flowcharts can describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0162] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0163] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network sending, etc.

[0164] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.

[0165] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0166] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0167] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that various inventive concepts may be embodied and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above may be used individually or in combination. Furthermore, aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0168] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this description.

[0169] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0170] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0171] The language of a claim stating “at least one of” and / or “one or more of” in a set, or other languages, indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language stating “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, the claim language stating “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of “at least one of” and / or “one or more of” in a set does not limit the set to the items listed in the set. For example, the claim language stating “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0172] Claim language or other languages ​​that state "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," etc., indicate that one or more processors (in any combination) are capable of performing associated operations. For example, claim language that states "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks of operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that states "at least one processor, the at least one processor being configured to: X, Y, and Z" may mean that any single processor can perform only a subset of operations X, Y, and Z.

[0173] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0174] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).

[0175] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.

[0176] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0177] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration). Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0178] The exemplary aspects of this disclosure include:

[0179] Aspect 1. An apparatus for processing data, the apparatus comprising: an image signal processor (ISP), the image signal processor (ISP) being configured to: receive image data and an indication of a region of interest (ROI) from an image sensor; determine image processing settings for processing the image data based on the ROI; and process the image data based on the image processing settings.

[0180] Aspect 2. The apparatus according to aspect 1, wherein the ISP includes one or more ISP engines, wherein each of the one or more ISP engines is configured to: determine appropriate image processing settings based on the ROI; and process the image data based on the appropriate image processing settings.

[0181] Aspect 3. The apparatus according to any one of Aspects 1 or 2, wherein the ISP comprises one or more ISP engines, wherein each of the one or more ISP engines is configured to: receive the image data and the indication to the ROI from the image sensor or from a previous ISP engine among the one or more ISP engines; and provide the image data and the indication to the ROI to a subsequent ISP engine among the one or more ISP engines or at the output of the ISP.

[0182] Aspect 4. The apparatus according to aspect 3, wherein the one or more ISP engines include a camera serial interface (CSI) decoder, the camera serial interface (CSI) decoder being configured to: receive from the image sensor a packet including the image data and the indication to the ROI; parse the packet; and provide the image data and the indication to the ROI to a subsequent ISP engine among the one or more ISP engines.

[0183] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the ISP is configured to perform operations associated with at least one of: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; or noise reduction.

[0184] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the ISP comprises at least one of: a first ISP engine configured to: determine a first image processing setting related to a first ISP operation based on the ROI; and perform the first ISP operation based on the first image processing setting; and a second ISP engine configured to: determine a second image processing setting related to a second ISP operation based on the ROI; and perform the second ISP operation based on the second image processing setting.

[0185] Aspect 7. The apparatus according to aspect 6, wherein each of the first ISP operation and the second ISP operation is associated with at least one of the following: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; or noise reduction.

[0186] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein, in order to process the image data, the ISP is configured to process the image data upon receiving the image data from the image sensor.

[0187] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the ISP is configured to receive the image data and the indication to the ROI in a packet, and the ISP is configured to parse the indication to the ROI from the header of the packet.

[0188] Aspect 10. The apparatus according to aspect 9, wherein the group includes a Mobile Industry Processor Interface (MIPI) group.

[0189] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the ISP is configured to receive the indication for the ROI in the tail of a first packet, wherein the ISP is configured to parse the indication for the ROI from the tail of the first packet, and wherein the ISP is configured to receive the image data in the payload of a second packet.

[0190] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the indication to the ROI is a first indication to the ROI, and wherein the apparatus further comprises: the image sensor, wherein the image sensor is configured to: receive a second indication to the ROI; generate the image data based on the ROI; and provide the image data and the first indication to the ROI to an image signal processor (ISP).

[0191] Aspect 13. The apparatus according to aspect 12, wherein, in order to generate the image data, the image sensor is configured to: generate a first portion of the image corresponding to the ROI at a first resolution; and generate a second portion of the image outside the ROI at a second resolution, wherein the first resolution is greater than the second resolution.

[0192] Aspect 14. The apparatus according to any one of Aspects 12 or 13, the apparatus further comprising at least one processor configured to determine the ROI based on data from a gaze tracking sensor.

[0193] Aspect 15. The apparatus according to any one of Aspects 12 to 14, wherein, in order to provide the image data and the second indication to the ROI to the ISP, the image sensor is configured to generate packets, the packets including the second indication to the ROI in the header of the packets and including the image data in the payload of the packets.

[0194] Aspect 16. The apparatus according to aspect 15, wherein the packet includes a Mobile Industry Processor Interface (MIPI) packet.

[0195] Aspect 17. The apparatus according to any one of Aspects 12 to 16, wherein: in order to provide the second indication of the ROI to the ISP, the image sensor is configured to generate a first packet, the first packet including the second indication of the ROI in the tail of the first packet; and in order to provide the image data to the ISP, the image sensor is configured to generate a second packet, the second packet including the image data in the payload of the second packet.

[0196] Aspect 18. A method for processing data, the method comprising: receiving image data and an indication of a region of interest (ROI) from an image sensor at an image signal processor (ISP); determining image processing settings for processing the image data at the ISP based on the ROI; and processing the image data at the ISP based on the image processing settings.

[0197] Aspect 19. The method according to aspect 18, wherein the ISP includes one or more ISP engines, and the method further includes: determining appropriate image processing settings at each of the one or more ISP engines based on the ROI; and processing the image data at each of the one or more ISP engines based on the appropriate image processing settings.

[0198] Aspect 20. The method according to any one of Aspects 18 or 19, wherein the ISP comprises one or more ISP engines, and the method further comprises: receiving, at each of the one or more ISP engines, the image data and the indication to the ROI from the image sensor or from a previous ISP engine among the one or more ISP engines; and providing, at each of the one or more ISP engines, the image data and the indication to the ROI to a subsequent ISP engine among the one or more ISP engines or at the output of the ISP.

[0199] Aspect 21. The method according to aspect 20, the method further comprising: receiving, at a camera serial interface (CSI) decoder, a packet including the image data and the indication to the ROI from the image sensor; parsing the packet at the CSI decoder; and providing, at the CSI decoder, the image data and the indication to the ROI to a subsequent ISP engine among the one or more ISP engines.

[0200] Aspect 22. The method according to any one of Aspects 18 to 21, the method further comprising performing operations associated with at least one of: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; or noise reduction.

[0201] Aspect 23. The method according to any one of Aspects 18 to 22, the method further comprising at least one of: determining, at a first ISP engine, a first image processing setting related to a first ISP operation based on the ROI; and performing the first ISP operation at the first ISP engine based on the first image processing setting; or determining, at a second ISP engine, a second image processing setting related to a second ISP operation based on the ROI; and performing the second ISP operation at the second ISP engine based on the second image processing setting.

[0202] Aspect 24. The method according to aspect 23, wherein each of the first ISP operation and the second ISP operation is associated with at least one of the following: lens shading correction; bad pixel correction (BPC); phase detection pixel correction; demosaic; lateral chromatic aberration correction; Bayer filtering; adaptive Bayer filtering; tone mapping; or noise reduction.

[0203] Aspect 25. The method according to any one of Aspects 18 to 24, wherein processing the image data includes processing the image data upon receiving the image data from the image sensor.

[0204] Aspect 26. The method according to any one of aspects 18 to 25, the method further comprising receiving the image data and the indication to the ROI in a packet, and parsing the indication to the ROI from the header of the packet.

[0205] Aspect 27. The method according to aspect 26, wherein the group includes the Mobile Industry Processor Interface (MIPI) group.

[0206] Aspect 28. The method according to any one of Aspects 18 to 27, the method further comprising receiving the indication for the ROI in the tail of a first packet, wherein the ISP is configured to parse the indication for the ROI from the tail of the first packet, and wherein the ISP is configured to receive the image data in the payload of a second packet.

[0207] Aspect 29. The method according to any one of Aspects 18 to 28, wherein the indication to the ROI is a first indication to the ROI, and the method further comprises: receiving a second indication to the ROI at an image sensor; generating image data based on the ROI at the image sensor; and providing the image data and the first indication to the ROI from the image sensor to an image signal processor (ISP).

[0208] Aspect 30. The method according to aspect 29, the method further comprising: generating a first portion of an image corresponding to the ROI at a first resolution at the image sensor; and generating a second portion of the image outside the ROI at a second resolution at the image sensor, wherein the first resolution is greater than the second resolution.

Claims

1. An apparatus for processing data, the apparatus comprising: Image signal processor (ISP), the image signal processor (ISP) is configured to: Receive image data and indications of regions of interest (ROI) from the image sensor; Based on the ROI, determine the image processing settings for processing the image data; and The image data is processed based on the image processing settings.

2. The apparatus of claim 1, wherein the ISP comprises one or more ISP engines, and each of the one or more ISP engines is configured to: Based on the ROI, determine the corresponding image processing settings; and The image data is processed based on the corresponding image processing settings.

3. The apparatus of claim 1, wherein the ISP comprises one or more ISP engines, and each of the one or more ISP engines is configured to: The image data and the indication of the ROI are received from the image sensor or from a previous ISP engine in one or more ISP engines; and The image data and the indication of the ROI are provided to a subsequent ISP engine in one or more ISP engines or at the output of the ISP.

4. The apparatus of claim 3, wherein the one or more ISP engines include a camera serial interface (CSI) decoder, the camera serial interface (CSI) decoder being configured to: Receive from the image sensor a packet including the image data and the indication of the ROI; Parse the group; and The image data and the indication of the ROI are provided to a subsequent ISP engine in one or more ISP engines.

5. The apparatus of claim 1, wherein the ISP is configured to perform operations associated with at least one of the following: Lens shadow correction; Bad pixel correction (BPC); Phase detection pixel correction; Remove the mosaic; Lateral color difference correction; Bayer filtering; Adaptive Bayer filtering; Pitch mapping; or Noise reduction.

6. The apparatus of claim 1, wherein the ISP comprises at least one of the following: The first ISP engine is configured as follows: Based on the ROI, determine the first image processing settings related to the first ISP operation; and The first ISP operation is performed based on the first image processing settings; and The second ISP engine is configured as follows: Based on the ROI, determine the second image processing settings related to the second ISP operation; and The second ISP operation is performed based on the second image processing settings.

7. The apparatus of claim 6, wherein each of the first ISP operation and the second ISP operation is associated with at least one of the following: Lens shadow correction; Bad pixel correction (BPC); Phase detection pixel correction; Remove the mosaic; Lateral color difference correction; Bayer filtering; Adaptive Bayer filtering; Pitch mapping; or Noise reduction.

8. The apparatus of claim 1, wherein, in order to process the image data, the ISP is configured to process the image data upon receiving the image data from the image sensor.

9. The apparatus of claim 1, wherein the ISP is configured to receive the image data and the indication to the ROI in a packet, and the ISP is configured to parse the indication to the ROI from the header of the packet.

10. The apparatus of claim 9, wherein the packet includes a Mobile Industry Processor Interface (MIPI) packet.

11. The apparatus of claim 1, wherein the ISP is configured to receive the indication for the ROI in the tail of a first packet, wherein the ISP is configured to parse the indication for the ROI from the tail of the first packet, and wherein the ISP is configured to receive the image data in the payload of a second packet.

12. The apparatus of claim 1, wherein the indication of the ROI is a first indication of the ROI, and wherein the apparatus further comprises: The image sensor, wherein the image sensor is configured to: Receive a second instruction for the ROI; The image data is generated based on the ROI; as well as The image data and the first indication of the ROI are provided to the image signal processor (ISP).

13. The apparatus of claim 12, wherein, in order to generate the image data, the image sensor is configured to: Generate a first portion of the image corresponding to the ROI at a first resolution; and A second portion of the image outside the ROI is generated at a second resolution, wherein the first resolution is greater than the second resolution.

14. The apparatus of claim 12, further comprising at least one processor configured to determine the ROI based on data from a gaze tracking sensor.

15. The apparatus of claim 12, wherein, in order to provide the image data and the second indication to the ROI to the ISP, the image sensor is configured to generate packets, the packets including the second indication to the ROI in the header of the packets and including the image data in the payload of the packets.

16. The apparatus of claim 15, wherein the packet includes a Mobile Industry Processor Interface (MIPI) packet.

17. The apparatus according to claim 12, wherein: In order to provide the second indication of the ROI to the ISP, the image sensor is configured to generate a first packet, the first packet including the second indication of the ROI in the tail of the first packet; and In order to provide the image data to the ISP, the image sensor is configured to generate a second packet, the second packet including the image data in its payload.

18. A method for processing data, the method comprising: Image data and indications of regions of interest (ROIs) are received from the image sensor at the image signal processor (ISP); At the ISP, image processing settings for processing the image data are determined based on the ROI; as well as The image data is processed at the ISP based on the image processing settings.

19. The method of claim 18, wherein the ISP comprises one or more ISP engines, and the method further comprises: At each of the one or more ISP engines, the corresponding image processing settings are determined based on the ROI; as well as The image data is processed at each of the one or more ISP engines based on the corresponding image processing settings.

20. The method of claim 18, wherein the ISP comprises one or more ISP engines, and the method further comprises: The image data and the indication of the ROI are received at each of the one or more ISP engines from the image sensor or from a previous ISP engine in the one or more ISP engines; as well as The image data and the indication of the ROI are provided to a subsequent ISP engine in the one or more ISP engines or at the output of the ISP at each ISP engine.

21. The method according to claim 20, further comprising: At the camera serial interface (CSI) decoder, a packet including the image data and the indication of the ROI is received from the image sensor; The packet is parsed at the CSI decoder; as well as The image data and the indication of the ROI are provided at the CSI decoder to a subsequent ISP engine in one or more ISP engines.

22. The method of claim 18, further comprising performing an operation associated with at least one of the following: Lens shadow correction; Bad pixel correction (BPC); Phase detection pixel correction; Remove the mosaic; Lateral color difference correction; Bayer filtering; Adaptive Bayer filtering; Pitch mapping; or Noise reduction.

23. The method of claim 18, further comprising at least one of the following: At the first ISP engine, a first image processing setting related to the first ISP operation is determined based on the ROI; as well as The first ISP operation is performed at the first ISP engine based on the first image processing settings; or At the second ISP engine, second image processing settings related to the second ISP operation are determined based on the ROI; as well as The second ISP operation is performed at the second ISP engine based on the second image processing settings.

24. The method of claim 23, wherein each of the first ISP operation and the second ISP operation is associated with at least one of the following: Lens shadow correction; Bad pixel correction (BPC); Phase detection pixel correction; Remove the mosaic; Lateral color difference correction; Bayer filtering; Adaptive Bayer filtering; Pitch mapping; or Noise reduction.

25. The method of claim 18, wherein processing the image data includes processing the image data upon receiving the image data from the image sensor.

26. The method of claim 18, further comprising receiving the image data and the indication to the ROI in a packet, and parsing the indication to the ROI from the header of the packet.

27. The method of claim 26, wherein the group comprises a Mobile Industry Processor Interface (MIPI) group.

28. The method of claim 18, further comprising receiving the indication for the ROI in the tail of a first packet, wherein the ISP is configured to parse the indication for the ROI from the tail of the first packet, and wherein the ISP is configured to receive the image data in the payload of a second packet.

29. The method of claim 18, wherein the indication of the ROI is a first indication of the ROI, and the method further comprises: Receive a second indication of the ROI at the image sensor; The image data is generated at the image sensor based on the ROI; as well as The image data and the first indication of the ROI are provided from the image sensor to the image signal processor (ISP).

30. The method according to claim 29, further comprising: At the image sensor, a first portion of the image corresponding to the ROI is generated at a first resolution; as well as A second portion of the image outside the ROI is generated at the image sensor at a second resolution, wherein the first resolution is greater than the second resolution.