Efficient image data processing

By downsampling image data and upsampling using machine learning models, the problem of high computational complexity in existing technologies is solved, achieving efficient image data processing and generating high-resolution image data similar to the original image data.

CN121532793APending Publication Date: 2026-02-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480047394.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-02
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing image processing technologies are computationally complex and resource-intensive, especially when processing high-resolution image data, leading to increased power consumption and processor usage, making it difficult to efficiently generate full-resolution image data.

Method used

By downsampling image data to generate low-resolution downsampled image data, and then using a machine learning model to process the downsampled image data and generate upsampled image data, the effect of full-resolution image data can be achieved, saving computing resources.

Benefits of technology

When processing image data, it reduces computational complexity and resource consumption, while generating high-resolution image data similar to the original image data, thus achieving efficient image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121532793A_ABST
    Figure CN121532793A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing data are described herein. For example, a method for processing data is provided. The method may include: obtaining image data having a first resolution; downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; processing the down-sampled image data to generate processed down-sampled image data; and generating up-sampled image data based on the processed down-sampled image data and the image data using a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to efficient image data processing. For example, some aspects of this disclosure include systems and techniques for downsampling image data, processing downsampled image data, and upsampling processed image data. As another example, some aspects of this disclosure include systems and techniques for receiving first image data and second image data, processing the first image data, and generating processed second image data based on the processed first and second image data. Background Technology

[0002] A camera can focus light onto an image sensor, which can generate image data representing the light. Image data can represent images, such as still images and / or video frames. An image signal processor (ISP) can receive image data (e.g., raw image data from the image sensor) and process it, for example, to perform operations related to: demosaicing, color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), phase detection autofocus (PDAF), automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, or some combination thereof. Once processed, the image data can be displayed, stored (e.g., for display at a later time or for use by another system, such as a computer vision system), and / or transmitted (e.g., for display by another device or for use by another system, such as a computer vision system). Summary of the Invention

[0003] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceptual aspects, nor should it be considered to identify key or decisive elements relating to all conceptual aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description that follows.

[0004] This document describes systems and techniques for processing data. According to at least one example, an apparatus for processing data is provided. The apparatus includes a memory and one or more processors coupled to the memory. The one or more processors are configured to: acquire image data having a first resolution; downsample the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; process the downsampled image data to generate processed downsampled image data; and use a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0005] In another example, a method for processing data is provided. The method includes: obtaining image data having a first resolution; downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; processing the downsampled image data to generate processed downsampled image data; and using a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0006] In another example, a non-transitory computer-readable medium having instructions stored thereon is provided, which, when executed by at least one processor, cause the at least one processor to: obtain image data having a first resolution; downsample the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; process the downsampled image data to generate processed downsampled image data; and use a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0007] As another example, an apparatus is provided. The apparatus includes: components for acquiring image data having a first resolution; components for downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; components for processing the downsampled image data to generate processed downsampled image data; and components for generating upsampled image data based on the processed downsampled image data and the image data using a machine learning model.

[0008] As another example, an apparatus for processing data is provided. The apparatus includes a memory and one or more processors coupled to the memory. The one or more processors are configured to: acquire first image data; process the first image data to generate processed first image data; acquire second image data; and generate processed second image data based on the processed first image data and the second image data using a machine learning model.

[0009] In another example, a method for processing data is provided. The method includes: obtaining first image data; processing the first image data to generate processed first image data; obtaining second image data; and using a machine learning model to generate processed second image data based on the processed first image data and the second image data.

[0010] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by at least one processor, cause the at least one processor to: obtain first image data; process the first image data to generate processed first image data; obtain second image data; and use a machine learning model to generate processed second image data based on the processed first image data and the second image data.

[0011] As another example, an apparatus is provided. The apparatus includes: components for acquiring first image data; components for processing the first image data to generate processed first image data; components for acquiring second image data; and components for generating processed second image data using a machine learning model based on the processed first image data and the second image data.

[0012] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices), extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), vehicles (or computing devices or systems of vehicles), smart or connected devices (e.g., Internet of Things (IoT) devices), wearable devices, personal computers, laptop computers, video servers, televisions (e.g., network-connected televisions), robotic devices or systems, or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.

[0013] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0014] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0015] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0016] Figure 1 This is a block diagram illustrating an example architecture of an image processing system according to various aspects of this disclosure;

[0017] Figure 2A These are illustrations of example systems capable of efficiently processing image data according to various aspects of this disclosure;

[0018] Figure 2B This is an illustration illustrating examples of downsampling image data according to various aspects of this disclosure;

[0019] Figure 2C This is an illustration illustrating examples of rearranging image data according to various aspects of this disclosure;

[0020] Figure 2D This is an illustration illustrating examples of rearranging image data according to various aspects of this disclosure;

[0021] Figure 3A This is a diagram illustrating an example neural network capable of processing image data according to various aspects of this disclosure;

[0022] Figure 3B This is an illustration illustrating examples of rearranging image data according to various aspects of this disclosure;

[0023] Figure 4 This is an illustration of another example system that can efficiently process image data according to various aspects of this disclosure;

[0024] Figure 5 This is a diagram illustrating another example neural network capable of processing image data according to various aspects of this disclosure;

[0025] Figure 6A This is an illustration of another example system that can efficiently process image data according to various aspects of this disclosure;

[0026] Figure 6BThis is an illustration of another example system that can efficiently process image data according to various aspects of this disclosure;

[0027] Figure 7 This is a flowchart illustrating an example process for processing image data according to various aspects of this disclosure;

[0028] Figure 8 This is a flowchart illustrating another example process for processing image data according to various aspects of this disclosure;

[0029] Figure 9 This is a block diagram illustrating examples of deep learning neural networks that can be used to implement a perception module and / or one or more verification modules, based on some aspects of the disclosed techniques.

[0030] Figure 10 This is a block diagram illustrating examples of convolutional neural networks (CNNs) according to various aspects of this disclosure; and

[0031] Figure 11 This is a block diagram illustrating an example computing device architecture that can implement the various technologies described herein. Detailed Implementation

[0032] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation to provide a thorough understanding of the various aspects of this application. However, it will be apparent, however, that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0033] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0034] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0035] An image sensor can capture raw image data. This raw image data can indicate the intensity of light detected by the individual photodetectors in the sensor's photodetector array. In some cases, the photodetector array can be coupled to a filter (e.g., a Bayer filter) that filters out certain wavelengths, allowing the individual photodetectors to detect different colors of light. For example, one photodetector out of every four can detect red light, two out of every four can detect green light, and one out of every four can detect blue light. The photodetector array can include H*W photodetectors, where H is the number of photodetectors along the height dimension of the array, and W is the number of photodetectors along the width dimension of the array. The raw image data can include one data point for each photodetector. Therefore, the raw image data can include H*W data points.

[0036] Some image processing systems can process raw image data, for example, to perform sensor-related processing (e.g., defect pixel correction and demosaicing), lens processing (e.g., shadow correction), color correction, tone and gamma processing, noise reduction, sharpening, color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, or some combination thereof. Some image processing systems can process raw image data at full resolution (e.g., processing H*W pixels of raw image data).

[0037] Image processing can be computationally complex or expensive. For example, processing image data can result in significant power consumption and processor usage (which may prevent the processor from performing other tasks). As image data increases in size, the computational complexity of performing image processing on it also increases. For instance, processing raw image data at full resolution may be more computationally expensive than processing it at a reduced resolution (such as half resolution) (e.g., discarding or ignoring half of the original image data). However, in many cases, a processed full-resolution image (e.g., for display) is desired rather than a processed reduced-resolution image (e.g., a half-resolution image).

[0038] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, "systems and techniques") for efficiently processing image data. In some aspects, the systems and techniques described herein can obtain image data with a first resolution. The systems and techniques can downsample the image data to generate downsampled image data with a second resolution lower than the first resolution. The systems and techniques can process the downsampled image data to generate processed downsampled image data, and use a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data. The machine learning model can be trained to receive full-resolution image data and the processed lower-resolution image data as input, and generate full-resolution image data.

[0039] Full-resolution image data is essentially similar to image data that would be generated by processing image data at a first resolution (e.g., full resolution), which would require higher image processing complexity. By processing downsampled image data, this system and technique can save computational resources (e.g., power and / or processing time) compared to processing image data at the first resolution. By upsampling the processed image data, this system and technique can generate a version of image data with a first resolution (e.g., original resolution) that is substantially similar to the image data if the image data had already been processed. In this way, this system and technique can save computational resources while still generating full-resolution image data that is substantially similar to the image data if the image data had already been processed.

[0040] In some aspects, the systems and techniques described herein can obtain first image data and second image data. The first and second image data can be sequential images of a series of images (e.g., frames of video data). The first and second image data can be similar. The systems and techniques can process the first image data to generate processed first image data. The systems and techniques can use a machine learning model to generate processed second image data based on the processed first and second image data. The machine learning model can be trained to receive image data and the first processed image data as input and generate the second processed image data. For example, the machine learning model can be trained to generate second image data that is substantially similar to the second image data given that the second image data has been processed. By not processing the second image data, the systems and techniques can save computational resources compared to processing the first and second image data. By generating processed second image data based on the first and second image data, the systems and techniques can generate a version of second image data that is substantially similar to the second image data given that the second image data has been processed. In this way, the systems and techniques can save computational resources while still generating processed first image data and second image data that is substantially similar to the second image data given that the second image data has been processed.

[0041] Various aspects of this application will be described below with reference to the accompanying drawings.

[0042] Figure 1 This is a block diagram illustrating an example architecture of an image processing system 100 according to various aspects of this disclosure. The image processing system 100 includes various components for capturing and processing images, such as an image of scene 106. The image processing system 100 can capture image frames (e.g., still images or video frames). In some cases, a lens 108 and an image sensor 118 (which may include an analog-to-digital converter (ADC)) may be associated with an optical axis. In one exemplary example, both the photosensitive area of ​​the image sensor 118 (e.g., a photodiode) and the lens 108 may be centered on the optical axis.

[0043] In some examples, the lens 108 of the image processing system 100 faces the scene 106 and receives light from the scene 106. The lens 108 refracts the incident light from the scene onto the image sensor 118. The light received by the lens 108 then passes through the aperture of the image processing system 100. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 110. In other cases, the aperture may have a fixed size.

[0044] One or more control mechanisms 110 may control exposure, focus, and / or zoom based on information from image sensor 118 and / or image processor 124. In some cases, one or more control mechanisms 110 may include multiple mechanisms and components. For example, control mechanism 110 may include one or more exposure control mechanisms 112, one or more focus control mechanisms 114, and / or one or more zoom control mechanisms 116. One or more control mechanisms 110 may also include, in addition to Figure 1 Additional control mechanisms beyond those illustrated herein. For example, in some cases, one or more control mechanisms 110 may include controls for controlling analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0045] The focus control mechanism 114 of the control mechanism 110 can obtain focus settings. In some examples, the focus control mechanism 114 stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 114 can adjust the position of the lens 108 relative to the position of the image sensor 118. For example, based on the focus settings, the focus control mechanism 114 can adjust the focus by moving the lens 108 closer to or further away from the image sensor 118 via an actuated motor or servo system (or other lens mechanism). In some cases, additional lenses may be included in the image processing system 100. For example, the image processing system 100 may include one or more microlenses on each photodiode of the image sensor 118. These microlenses can each bend light received from the lens 108 toward the corresponding photodiode before the light reaches the photodiode.

[0046] In some examples, focus settings may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. Focus settings may be determined using control mechanism 110, image sensor 118, and / or image processor 124. Focus settings may be referred to as image capture settings and / or image processing settings. In some cases, lens 108 may be fixed relative to the image sensor and focus control mechanism 114.

[0047] Exposure control mechanism 112 of control mechanism 110 can obtain exposure settings. In some cases, exposure control mechanism 112 stores exposure settings in a memory register. Based on this exposure setting, exposure control mechanism 112 can control the aperture size (e.g., aperture size or aperture value), the duration of aperture opening (e.g., exposure time or shutter speed), the duration of light collection by the sensor (e.g., exposure time or electronic shutter speed), the sensitivity of image sensor 118 (e.g., ISO speed or film speed), the analog gain applied by image sensor 118, or any combination thereof. Exposure settings may be referred to as image capture settings and / or image processing settings.

[0048] The zoom control mechanism 116 of the control mechanism 110 can obtain zoom settings. In some examples, the zoom control mechanism 116 stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 116 can control the focal length of an assembly (lens assembly) of lens elements including lens 108 and one or more additional lenses. For example, the zoom control mechanism 116 can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 108) that first receives light from scene 106, where the light then passes through a focusing zoom system between the focusing lens (e.g., lens 108) and image sensor 118 before reaching image sensor 118. In some cases, a focusing zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference between them), with a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, zoom control mechanism 116 moves one or more lenses in the focusing zoom system, such as a negative lens and one or both positive lenses. In some cases, zoom control mechanism 116 can control zoom by capturing images from an image sensor (e.g., including image sensor 118) among a plurality of image sensors at a zoom setting corresponding to the zoom setting. For example, image processing system 100 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, zoom control mechanism 116 may capture images from the corresponding sensor based on the selected zoom setting.

[0049] Image sensor 118 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 118. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters, and thus light matching the color of the filter covering the photodiode can be measured. Various color filter arrays can be used, such as, for example, and not limited to, Bayer color filter arrays, four-color filter arrays (QCFA), and / or any other color filter array.

[0050] In some cases, image sensor 118 may optionally or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, UV cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). Image sensor 118 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed relative to one or more control mechanisms in control mechanism 110 may be alternatively or additionally included in image sensor 118. Image sensor 118 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0051] Image processor 124 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 128), one or more host processors (including host processor 126), and / or related to Figure 11The computing device architecture 1100 may include one or more processors of any other type discussed. The host processor 126 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 124 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes the host processor 126 and the ISP 128. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 130), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™ This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 130 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer port or interface), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 126 may communicate with image sensor 118 using the I2C port, and ISP 128 may communicate with image sensor 118 using the MIPI port.

[0052] Image processor 124 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 124 can store image frames and / or processed images in random access memory (RAM) 120, read-only memory (ROM) 122, cache, memory unit, another storage device, or some combination thereof.

[0053] Various input / output (I / O) devices 132 may be connected to the image processor 124. I / O devices 132 may include displays, keyboards, keypads, touchscreens, touchpads, touch-sensitive surfaces, printers, any other output devices, any other input devices, or any combination thereof. In some cases, text may be input into the image processing device 104 via the physical keyboard or keypad of the I / O device 132, or via a virtual keyboard or keypad on the touchscreen of the I / O device 132. I / O devices 132 may include one or more ports, jacks, or other connectors that enable wired connections between the image processing system 100 and one or more peripheral devices, through which the image processing system 100 may receive data from and / or send data to one or more peripheral devices. I / O devices 132 may include one or more wireless transceivers that enable wireless connections between the image processing system 100 and one or more peripheral devices, through which the image processing system 100 may receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 132 discussed earlier, and they can be considered I / O devices 132 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.

[0054] In some cases, the image processing system 100 may be a single device. In other cases, the image processing system 100 may be two or more separate devices, including an image capture device 102 (e.g., a camera) and an image processing device 104 (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 102 and the image processing device 104 may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 102 and the image processing device 104 may be disconnected from each other.

[0055] like Figure 1 As shown, the vertical dashed line will Figure 1The image processing system 100 is divided into two parts, namely image capture device 102 and image processing device 104. Image capture device 102 includes a lens 108, a control mechanism 110, and an image sensor 118. Image processing device 104 includes an image processor 124 (including an ISP 128 and a host processor 126), RAM 120, ROM 122, and I / O devices 132. In some cases, certain components illustrated in image capture device 102 (such as ISP 128 and / or host processor 126) may be included in image capture device 102. In some examples, image processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof).

[0056] Image processing system 100 may be part of or implemented by a single computing device or multiple computing devices. In some examples, image processing system 100 may be part of electronic devices (or multiple electronic devices), such as camera systems (e.g., digital cameras, IP cameras, video cameras, security cameras, etc.), telephone systems (e.g., smartphones, cellular phones, conferencing systems, etc.), laptops or notebook computers, tablet computers, set-top boxes, smart TVs, display devices, game consoles, XR devices (e.g., HMDs, smart glasses, etc.), IoT (Internet of Things) devices, smart wearable devices, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices.

[0057] Although the image processing system 100 is shown as including certain components, those skilled in the art will understand that the image processing system 100 may include more than [other components]. Figure 1 The components shown herein are additional components. Components of the image processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image processing system 100 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image processing system 100.

[0058] In some examples, Figure 11The computing device architecture 1100 shown and further described below may include an image processing system 100, an image capture device 102, an image processing device 104, or a combination thereof.

[0059] Figure 2A This is an illustration of an example system 200 capable of efficiently processing image data 204 according to various aspects of the present disclosure. For example, system 200 may receive image data 204 from sensor 202. System 200 may include: a downsampler 206 that can downsample image data 204 to generate image data 208; an image processor 210 that can process image data 208 to generate image data 212; a rearranger 214 that can rearrange image data 204 to generate image data 216; and a machine learning model 218 that generates image data 220 based on image data 212 and image data 216.

[0060] Sensor 202 can be Figure 1 Example of image sensor 118 or Figure 1 An example of an image capture device 102 is provided. System 200 may include sensor 202. Alternatively, sensor 202 may not be part of system 200, but sensor 202 may provide image data 204 to system 200. Sensor 202 may include a photodetector array and may generate image data 204 based on light illuminating the photodetector array. Image data 204 may indicate the intensity of light detected by the individual photodetectors of the photodetector array of sensor 202. The photodetector array of sensor 202 may include H*W photodetectors, where H is the number of photodetectors along the height dimension of the array and W is the number of photodetectors along the width dimension of the array. Image data 204 may include a data point for each photodetector in the array. Thus, image data 204 may include H*W data points (which may be pixels). In some cases, the photodetector array may be coupled to a filter (e.g., a Bayer filter) that can filter out certain wavelengths, allowing the individual photodetectors to detect different colors of light. For example, one of every four photodetectors can detect red light, two of every four photodetectors can detect green light, and one of every four photodetectors can detect blue light. Therefore, one of every four data points in image data 204 can represent the intensity of red light, two of every four data points in image data 204 can represent the intensity of green light, and one of every four data points in image data 204 can represent the intensity of blue light.

[0061] Downsampler 206 can Figure 1The image data 204 may be implemented in the image processor 124, or in another processor or circuit of the image processing device 104. The downsampler 206 may downsample the image data 204 to generate image data 208 (which may be downsampled image data). As an example, Figure 2B This is a diagram illustrating examples of downsampling data according to various aspects of this disclosure. Figure 2B The example demonstrates downsampling 4x4 image data to generate 2x2 image data. Return to... Figure 2A Image data 204 may have H*W data points (e.g., arranged in an H*W grid). Therefore, image data 204 may be full-resolution image data. Downsampler 206 may downsample image data 204 to obtain image data 208, which may include fewer data points than image data 204. For example, downsampler 206 may merge image data 204 by a factor of 2 (in both the horizontal and vertical directions) to generate image data 208 with H / 2*W / 2 data points. Therefore, image data 208 may be quarter-resolution image data (e.g., representing the same light intensity as image data 204, but with a quarter-resolution). Downsampler 206 may consider the Bayer pattern by, for example, merging groups of four pixels with adjacent groups of four pixels. Therefore, image data 208 may still include data according to the Bayer pattern. Merging by a factor of 2 is given as an example. Other factors and downsampling techniques are within the scope of this disclosure.

[0062] Image processor 210 can Figure 1 Implemented in the image processor 124, or Figure 1An example of image processor 124. Image processor 210 can process image data 208 to generate image data 212 (which may be the processed image data). Image processor 210 can perform multiple tasks such as sensor-related processing (e.g., defect pixel correction and demosaicing), lens processing (e.g., shadow correction), color correction, hue and gamma processing, noise reduction, sharpening color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, or some combination thereof. In addition, image processor 210 can perform demosaicing on image data 208 when generating image data 212, such that image data 212 includes data in red, green, and blue (RGB) format. Therefore, image data 212 may include three data points for each data point of image data 208 (e.g., representing the intensity of red light, the intensity of green light, and the intensity of blue light). Therefore, since image data 208 has been downsampled by a factor of 2 (to have H / 2*W / 2 data points), image data 212 may include H / 2*W / 2*3 data points. These H / 2*W / 2*3 data points can be arranged into a three-dimensional tensor with dimensions H / 2*W / 2*3.

[0063] The rearranger 214 can be implemented in the image processor 124, or in another processor or circuitry of the image processing device 104. The rearranger 214 can rearrange local pixel groups of image data 204 to generate image data 216 (which may be the rearranged image data). Rearranging the image data may include changing the dimensions of the image data without losing image data. As an example, Figure 2C and Figure 2D This is an illustration illustrating examples of rearranging image data according to various aspects of this disclosure. Figure 2C Examples include a left-to-right S2D 2x rearrangement (e.g., transforming 2x2 image data on the left to 1x1x4 image data on the right) and a right-to-left depth-to-space (D2S) 2x rearrangement (e.g., transforming 1x1x4 image data on the right to 2x2 image data on the left). Figure 2D Examples of left-to-right S2D2x rearrangements (e.g., transforming 8x8 image data on the left to 4x4x4 image data on the right) and right-to-left depth-to-space (D2S) 2x rearrangements (e.g., transforming 4x4x4 image data on the right to 8x8 image data on the left). Return to Figure 2AAs previously mentioned, image data 204 may have H*W data points (e.g., arranged in a two-dimensional grid). Furthermore, one of every four data points may represent the intensity of red light, two of every four data points may represent the intensity of green light, and one of every four data points may represent the intensity of blue light. Rearranger 214 can rearrange image data 204 by arranging the data points representing red light into groups or two-dimensional grids, arranging the data points representing green light into two groups or two-dimensional grids, and arranging the data points representing blue light into groups or two-dimensional grids. Therefore, image data 216 may include four groups or two-dimensional grids, each group or two-dimensional grid having a dimension of H / 2*W / 2. The four H / 2*W / 2 two-dimensional grids of image data 216 can be arranged into a three-dimensional tensor with a dimension of H / 2*W / 2*4. Rearranger 214 can implement a transformation that is known in the art as a space-to-depth (S2D) transformation.

[0064] Machine learning model 218 can be implemented in image processor 124 or in another processor of image processing device 104. Machine learning model 218 can be trained to receive full-resolution image data and processed lower-resolution image data as input and generate processed full-resolution image data. For example, machine learning model 218 can be trained using backpropagation training techniques by receiving full-resolution image data (e.g., the original image) and processed lower-resolution image data (e.g., the original image that has been downsampled and then processed) as input. Machine learning model 218 can then generate full-resolution image data as output. The original image data can be processed separately, and the processed original image data can be compared with the output of machine learning model 218 to determine the loss based on the difference between the processed original image data and the output of machine learning model 218. Machine learning model 218 can be tuned (e.g., the weights or parameters of machine learning model 218 can be adjusted) to reduce the loss in further iterations of the training process. The backpropagation process can be repeated any number of times to train machine learning model 218.

[0065] During inference (e.g., after training), machine learning model 218 may receive image data 212 (which may be a tensor with dimensions H / 2*W / 2*3) and image data 216 (which may be a tensor with dimensions H / 2*W / 2*4) as input. In some cases, image data 212 and image data 216 may be concatenated to form a single tensor with dimensions H / 2*W / 2*7. Machine learning model 218 may include one or more convolutional layers. The size of the convolutional layers may be set to operate according to the dimensions of the single tensor (H / 2*W / 2*7). Based on image data 212 and image data 216, machine learning model 218 may generate image data 220. Image data 220 may have dimensions H*W (e.g., image data 220 may be full-resolution image data). Image data 220 may be substantially similar to the image data that will be generated by processing image data 204 using image processor 210. For example, image data 220 can be approximated by the result of processing image data 204 using image processor 210. However, because in system 200, image processor 210 operates on image data 208 (which is one-quarter the size of image data 204), system 200 saves computational resources (processing time and power) compared to a system that uses image processor 210 to process image data 204.

[0066] Image data 220 can be displayed (e.g., on a display of a system or device including system 200 or by another system or device). For example, image data 220 can be displayed on a display of a smartphone including system 200, on a display of an extended reality (XR) system including system 200, or on a display (e.g., a computer or television display) after being generated by system 200 in a separate system or device. Additionally or alternatively, image data 220 can be stored (e.g., in the memory of a system or device including system 200). For example, image data 220 can be stored in memory by a camera including system 200. At a later time, image data 220 can be displayed or provided to another system or device (e.g., for display or analysis by another system or device). Additionally or alternatively, image data 220 can be transmitted (e.g., from a system or device including system 200 to another system or device). For example, a camera can generate image data 220 and can transmit image data 220 to another system or device (e.g., for display or analysis). In some respects, image data 220 can be analyzed (e.g., through machine learning models) without being displayed. For example, image data 220 can be used by object tracking algorithms.

[0067] As an example, system 200 is illustrated and described as downsampling image data 204 at downsampler 206 by a factor of 2 and upsampling image data 212 at machine learning model 218. In other respects, system 200 may downsample and upsample image data by other factors (e.g., 1.5, 3, 4, etc.).

[0068] Figure 3A This is a diagram illustrating an example neural network 300 capable of processing image data 302 and image data 304 according to various aspects of this disclosure. The neural network 300 may be... Figure 2A An example of machine learning model 218. The neural network 300 can receive image data 302 and image data 304 as input, rearrange the image data 302 and image data 304, and feed the rearranged image data to multiple convolutional layers (e.g., convolutional layer 308, convolutional layer 310, convolutional layer 312, convolutional layer 314, and convolutional layer 316), and then to a rearranger 318, which can rearrange the image data to generate image data 320.

[0069] Image data 302 can be processed downsampled image data. For example, image data 302 can be... Figure 2A Example of image data 212. For example, the original image could be full-resolution image data with H*W data points, where H is related to the height dimension and W is related to the width dimension. The original image could be downsampled (e.g., by factoring by two for horizontal and vertical merging) (e.g., as per [reference to image data 212]). Figure 2B (As illustrated and described). Downsampled image data can have H / 2*W / 2 data points. Downsampled image data can be processed (e.g., with...). Figure 2A The image data 302 may be the same as or substantially similar to the image processor 210 and / or a processor that performs the same or substantially the same operations as the image processor. The image data 302 may be the resulting processed downsampled image data. Processing may include demosaicing, which may convert the image data 302 to RGB format. The image data 302 may be arranged as a tensor with dimensions H / 2*W / 2*3.

[0070] Image data 304 can be rearranged image data. For example, image data 304 could be... Figure 2A Example of image data 216. For example, the original image data (from which image data 302 is derived) can be rearranged to generate image data 304 (e.g., as per [reference to image data 304]). Figure 2C and Figure 2D(As illustrated and described). For example, the original image data may have a dimension of H*W and may be based on a Bayer pattern (e.g., one out of every four data points represents the intensity of red light, two out of every four data points represent the intensity of green light, and one out of every four data points represents the intensity of blue light). The original data may be rearranged into a tensor such that color is the dimension. For example, the data points representing red light may be arranged into a two-dimensional grid, the data points representing green light may be arranged into two two-dimensional grids, and the data points representing blue light may be arranged into a two-dimensional grid. The two-dimensional grids may be stacked to form a tensor with a dimension of H / 2*W / 2*4.

[0071] The rearranger 306 can receive image data 302 and image data 304 as input. The rearranger 306 can combine (e.g., cascade) image data 302 and image data 304 to form a single tensor with dimensions H / 2*W / 2*7. The rearranger 306 can rearrange the tensor to form new two-dimensional grid layers by selecting one of every four data points in each two-dimensional grid layer. For example, an H / 2*W / 2 two-dimensional grid representing data points of red light can be transformed into four H / 4*W / 4 two-dimensional grids that can be rearranged into a tensor with dimensions H / 4*W / 4*4. In this way, the H / 2*W / 2*7 tensor can be rearranged into an H / 4*W / 4*28 tensor. The rearranger 306 can implement a space-to-depth (S2D) transformation. Rearranging the H / 2*W / 2*7 tensor to generate the H / 4*W / 4*28 tensor can be analogous to... Figure 2C and Figure 2D The illustrated and described S2D rearrangement from left to right, however, the dimensions (e.g., the start dimension and the end dimension) can be different.

[0072] Convolutional layer 308 can receive H / 4*W / 4*28 tensors from rearranger 306. Convolutional layer 308 can convolve the H / 4*W / 4*28 tensors using a kernel (e.g., a 3*3 kernel). In addition, convolutional layer 308 may include linear units (e.g., rectified linear units (ReLU) or parametric rectified linear units (PReLU)).

[0073] Each of convolutional layers 310, 312, 314, and 316 can sequentially receive tensors and convolve them. For example, each of convolutional layers 310, 312, and 314 may include a corresponding kernel (e.g., a 3x3 kernel or a 1x1 kernel), and each of convolutional layers 310, 312, and 314 can sequentially convolve tensors using that corresponding kernel. Furthermore, each of convolutional layers 310, 312, 314, and 316 may include a corresponding linear unit (e.g., ReLU or PReLU). Convolutional layer 310 can receive the H / 4*W / 4*64 tensor output by convolutional layer 308, convolve the H / 4*W / 4*64 tensor (e.g., using a 3*3 kernel), and output the H / 4*W / 4*64 tensor (with essentially the same dimensions as the tensor received by convolutional layer 310, ignoring any size reduction due to convolution) to convolutional layer 312. Similarly, convolutional layer 312 can receive the H / 4*W / 4*64 tensor output by convolutional layer 310, convolve the H / 4*W / 4*64 tensor (e.g., using a 3*3 kernel), and output the H / 4*W / 4*64 tensor (ignoring any size reduction due to convolution) to convolutional layer 314. Similarly, convolutional layer 314 can receive the H / 4*W / 4*64 tensor output by convolutional layer 312, convolve the H / 4*W / 4*64 tensor (e.g., using a 3*3 kernel), and output the H / 4*W / 4*64 tensor (ignoring any size reduction due to convolution) to convolutional layer 316. Convolutional layer 316 can receive the H / 4*W / 4*64 tensor output by convolutional layer 314, convolve the H / 4*W / 4*64 tensor (e.g., using a 1*1 kernel), and output the H / 4*W / 4*48 tensor to rearranger 318.

[0074] Rearranger 318 can receive H / 4*W / 4*48 tensors from convolutional layer 316 and rearrange the H / 4*W / 4*48 tensors into image data 320. Rearranger 318 can rearrange the H / 4*W / 4*48 tensors into H*W*3 tensors (e.g., by combining four H / 4*W / 4 two-dimensional grids into one H*W two-dimensional grid). Rearranger 318 can implement a transformation that is known in the art as a depth-to-space (D2S) transformation. Rearranging H / 4*W / 4*48 tensors to generate H*W*3 tensors can be analogous to... Figure 2C and Figure 2D The illustrated and described right-to-left D2S rearrangement, however, may differ in the dimensions (e.g., the start and end dimensions) and the scaling of the rearrangement (e.g., ...). Figure 2C and Figure 2D This example demonstrates 2x scaling, while rearranger 318 performs 4x scaling. Figure 3B This is an illustration illustrating examples of rearranging image data according to various aspects of this disclosure. For example, Figure 3B Examples of left-to-right S2D 4x rearrangement (e.g., transforming 4x4 image data on the left to 1x1x16 image data on the right) and right-to-left D2S 4x rearrangement (e.g., transforming 1x1x16 image data on the right to 4x4 image data on the left). Return to Figure 3A Each of the H*W two-dimensional grids can represent a color (e.g., red, green, and blue). Therefore, the H*W*3 tensor of image data 320 can represent an image in RGB format. Thus, image data 320 can be full-resolution image data in RGB format.

[0075] As an example, neural network 300 is illustrated and described as upsampling image data 302 by a factor of 2. In other respects, neural network 300 may upsample image data by other factors (e.g., 1.5, 3, 4, etc.).

[0076] As mentioned above Figure 2A As described, system 200 can efficiently process image data 204 by reducing the size of image data 204 to generate image data 208 and then processing image data 208 using image processor 210. Processing image data 208 at image processor 210 instead of image data 204 saves computational resources compared to processing image data 204 at image processor 210. Furthermore, system 200 can use machine learning model 218 (… Figure 3A The system 200 can generate image data 220 (where neural network 300 can be an example of this machine learning model) to produce image data 220, which can be substantially similar to the result of processing image data 204 at image processor 210. In this way, the system 200 can save computational resources while producing substantially the same result compared to processing image data 204 at image processor 210. The system 200 can use machine learning model 218 (where neural network 300 can be an example of this machine learning model) to reduce the amount of image data processed by image processor 210 by reducing the size of the image data before processing it, thereby saving computational resources.

[0077] Figure 4This is an illustration of an example system 400 that can efficiently process image data 404 and image data 410 according to various aspects of this disclosure. For example, system 400 may receive image data 404 and image data 410 from sensor 402. System 400 may include: an image processor 406 for processing image data 404 to generate image data 408; and a machine learning model 412 for generating image data 414 based on image data 408 and image data 410. Image data 404 and image data 410 may be sequential images in a series of images (e.g., frames of video data). By processing image data 404 at image processor 406 instead of image data 410 at image processor 406, system 400 can save computational resources compared to processing both image data 404 and image data 410 at image processor 406. Furthermore, system 400 can use machine learning model 412 to generate image data 414, which can be substantially similar to the result of processing image data 410 at image processor 406. In this way, system 400 can save computational resources while generating substantially the same result compared to processing image data 410 at image processor 406. System 400 can use machine learning model 412 to reduce the amount of image data processed by image processor 406 by reducing the number of frames of image data processed by image processor 406, thereby saving computational resources.

[0078] Sensor 402 can be Figure 1 Example of image sensor 118 or Figure 1 An example of an image capture device 102. Sensor 402 can generate image data 404 and image data 410 based on light illuminating a photodetector array. Image data 404 and image data 410 can be in any suitable format. For example, image data 404 and image data 410 can be raw image data, RGB image data, or lightness, blue projection, red projection (YUV) image data. As mentioned above, image data 404 and image data 410 can be sequential images in a series of images (e.g., frames of video data). For example, image data 404 can be the first frame of video data, and image data 410 can be the second frame of video data (e.g., the immediately following frame). Based on the frame capture rate of sensor 402 and based on the scene captured by sensor 402, many pixels in the pixels of image data 410 can be the same as or substantially similar to the corresponding pixels of image data 404.

[0079] Image processor 406 can Figure 1 Implemented in the image processor 124, or Figure 1An example of image processor 124. Image processor 406 can process image data 404 to generate image data 408 (which may be processed image data). Image processor 406 can perform a variety of tasks, such as sensor-related processing (e.g., defect pixel correction and demosaicing), lens processing (e.g., shadow correction), color correction, tone and gamma processing, noise reduction, sharpening color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, or some combination thereof.

[0080] Machine learning model 412 can be implemented in image processor 124 or in another processor of image processing device 104. Machine learning model 412 can be trained to receive processed first image data and second image data as input and generate processed second image data. For example, machine learning model 412 can be trained using backpropagation training techniques by receiving processed image data (e.g., a processed first frame) and unprocessed image data (e.g., an unprocessed second frame) as input. Machine learning model 412 can then generate processed second image data as output. The second image data can be processed separately, and the processed second image data can be compared with the output of machine learning model 412 to determine the loss based on the difference between the processed image data and the output of machine learning model 412. Machine learning model 412 can be tuned (e.g., the weights or parameters of machine learning model 412 can be adjusted) to reduce the loss in further iterations of the training process. The backpropagation process can be repeated any number of times to train machine learning model 412.

[0081] During inference (e.g., after training), machine learning model 412 may receive image data 408 (e.g., which may be processed by image processor 406) and image data 410 as input. Machine learning model 412 may include one or more convolutional layers. Based on image data 408 and image data 410, machine learning model 412 may generate image data 414. Image data 414 may be substantially similar to the image data that would be generated by processing image data 410 using image processor 406. For example, image data 414 may approximate the result of processing image data 410 using image processor 406. However, because image processor 406 does not process image data 410 in system 400, system 400 saves computational resources (processing time and power) compared to a system that uses image processor 406 to process image data 410.

[0082] Both image data 408 and image data 414 can be output by system 400. For example, image data 408 (which can be processed by image processor 406) and image data 414 (which can be substantially similar to the result of processing image data 410 at image processor 406) can be output by system 400. Similar to the above regarding... Figure 2A The content described by image data 220 can be displayed, stored, sent, and / or analyzed, including image data 408 and / or image data 414.

[0083] System 400 is illustrated and described as generating image data 414 based on two frames (for example, a processed frame (image data 408) and an unprocessed frame (image data 410)). In other aspects, system 400 may generate image data 414 based on another number of frames. For example, image data 404 may be the first frame of video data, and image data 410 may be the second frame of video data. The third frame of video data (… Figure 4 (Not illustrated) can be received and processed by image processor 406. In this case, machine learning model 412 can be based on the processed first frame (image data 408), the unprocessed second frame (image data 410), and the processed third frame (not illustrated). Figure 4 (Not illustrated) to generate image data 414. As another example, system 400 may generate image data 414 based on three or four processed frames and one unprocessed frame.

[0084] Figure 5 This is a diagram illustrating an example neural network 500 capable of processing image data 502 and image data 504 according to various aspects of this disclosure. The neural network 500 can be... Figure 4 An example of machine learning model 412. Neural network 500 can receive image data 502 and image data 504 as input, and feed image data 502 and image data 504 to multiple convolutional layers (e.g., convolutional layer 506, convolutional layer 508, and convolutional layer 510). The multiple convolutional layers can generate image data 512.

[0085] Image data 502 can be processed image data. For example, image data 502 can be... Figure 4 Example of image data 408. For example, the first frame of video data can be processed (e.g., with...). Figure 4 Image data 502 may be the obtained processed image data. Image data 504 may be unprocessed image data. For example, image data 504 may be... Figure 4Example of image data 410. For example, image data 504 could be the second frame of video data (e.g., immediately following the first frame).

[0086] Each of convolutional layers 506, 508, and 510 can sequentially receive image data and convolve it. For example, each of convolutional layers 506, 508, and 510 may include a corresponding kernel, and each of them can sequentially convolve the image data using that corresponding kernel. Furthermore, each of convolutional layers 506, 508, and 510 may include a corresponding linear unit (e.g., ReLU or PReLU). Convolutional layer 510 can generate image data 512.

[0087] As an example, neural network 500 is described and illustrated as having three convolutional layers. In other respects, neural network 500 may include any number of convolutional layers, such as 2, 4, 5, 6, etc.

[0088] Figure 6A This is a diagram illustrating an example system 600A capable of efficiently processing image data 604 and image data 622 according to various aspects of the present disclosure. For example, system 600A may receive image data 604 and image data 622 from sensor 602. System 600A may include: a downsampler 606 that can downsample image data 604 to generate image data 608; an image processor 610 for processing image data 608 to generate image data 612; a rearranger 614 for rearranging image data 604 to generate image data 616; a first machine learning model 618 for generating image data 620 based on image data 612 and image data 616; and a second machine learning model 624 for generating image data 626 based on image data 620 and image data 622.

[0089] System 600A can implement or include system 200 and can benefit from the processing efficiency of system 200. For example, system 600A's downsampler 606, image processor 610, rearranger 614, and machine learning model 618 can implement system 200. For example, image data 604 can play a role in system 600A similar to... Figure 2A The image data 204 plays the same role in system 200. Downsampler 606 can be used with... Figure 2A The downsampler 206 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operation as, it. Image data 608 can be... Figure 2AThe image data 208 is the same as, or can be substantially similar to, that of the image processor 610. Figure 2A The image processor 210 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as, it. Image data 612 can be... Figure 2A The image data 212 is the same as, or can be substantially similar to, it. The rearranger 614 can be with... Figure 2A The rearranger 214 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as. Image data 616 can be with Figure 2A The image data 216 is the same, or can be substantially similar to, it. Machine learning model 618 can be... Figure 2A The machine learning model 218 is the same as it, can be substantially similar to it, and / or can perform the same or substantially the same operations as it. Figure 3A The neural network 300 can be Figure 6A An example of machine learning model 618. Image data 620 can be used with... Figure 2A The image data 220 is the same, or can be substantially similar to it.

[0090] System 600A can implement or include system 400 and can benefit from the processing efficiency of system 400. For example, system 600A can implement downsampler 606, image processor 610, rearranger 614, and machine learning model 618. Figure 4 The image processor 406, and the machine learning model 624 can be used with... Figure 4 The machine learning model 412 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as, it. For example, image data 604 can play the same role in system 600A as... Figure 4 The image data 404 plays the same role in system 400. The downsampler 606, image processor 610, rearranger 614, and machine learning model 618 can play the same role in system 600A. Figure 4 The image processor 406 plays the same role in system 400. Image data 620 can play the same role in system 600A. Figure 4 Image data 408 plays the same role in system 400. Image data 622 can play the same role in system 600A. Figure 4 The image data 410 plays the same role in system 400. Machine learning model 624 can be used with... Figure 4 The machine learning model 412 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as, it. Image data 626 can play a similar role in system 600A. Figure 4The image data 414 plays the same role in system 400.

[0091] By processing image data 608 at image processor 610 instead of image data 604 at image processor 610, system 600A can save computational resources compared to processing image data 604 at image processor 610. Furthermore, system 600A can utilize machine learning model 618 ( Figure 3A The neural network 300 (which may be an example of this machine learning model) generates image data 620, which can be substantially similar to the result of processing image data 604 at image processor 610. In this way, system 600A can save computational resources while generating substantially the same result compared to processing image data 604 at image processor 610. In this way, system 600A can use machine learning model 618 (where neural network 300 may be an example of this machine learning model) to reduce the amount of image data processed by image processor 610 by reducing the size of the image data before processing it, thereby saving computational resources.

[0092] Additionally, image data 604 and image data 622 can be sequential images in a series of images (e.g., frames of video data). By processing image data 608 (which is based on image data 604) at image processor 610 instead of processing image data 622 at image processor 610, system 600A can save computational resources compared to processing both image data 604 and image data 622 at image processor 610. Furthermore, system 600A can use machine learning model 624 ( Figure 5 The neural network 500 (which can be an example of this machine learning model) generates image data 626, which can be substantially similar to the result of processing image data 622 at image processor 610. In this way, system 600A can save computational resources compared to processing image data 622 at image processor 610, while generating substantially the same result. In this way, system 600A can use machine learning model 624 ( Figure 5 The neural network 500 (which can be an example of this machine learning model) reduces the amount of image data processed by the image processor 610 by reducing the number of frames of image data processed by the image processor 610, thereby saving computational resources.

[0093] Figure 6BThis is an illustration of an example system 600B that can efficiently process image data 604 and image data 622 according to various aspects of the present disclosure. For example, system 600B can receive image data 604 and image data 622 from sensor 602. System 600B may include: a downsampler 606 for downsampling image data 604 to generate image data 608; and a rearranger 628 for rearranging image data 622 to generate image data 630. Alternatively, system 600B may include: an image processor 610 for processing image data 612 to generate image data 608; and a machine learning model 632 for generating image data 634 based on image data 612 and image data 630.

[0094] System 600B can implement or include system 200 and can benefit from the processing efficiency of system 200. For example, system 600B's downsampler 606, image processor 610, rearranger 628, and machine learning model 632 can implement system 200. For example, image data 604 can play a role in system 600B similar to... Figure 2A The image data 204 plays the same role in system 200. Downsampler 606 can be used with... Figure 2A The downsampler 206 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operation as, it. Image data 608 can be... Figure 2A The image data 208 is the same as, or can be substantially similar to, that of the image processor 610. Figure 2A The image processor 210 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as, it. Image data 612 can be... Figure 2A The image data 212 is the same as, or can be substantially similar to, it. The rearranger 628 can be with... Figure 2A The rearranger 214 is the same as, can be substantially similar to, and / or can perform the same or substantially the same operations as. However, the rearranger 628 can rearrange the image data 622, which can be the second frame in a series of sequential frames (e.g., after image data 604). In addition, the machine learning model 632 can perform operations similar to... Figure 2A The machine learning model 218 operates essentially the same.

[0095] System 600B can implement or include system 400 and can benefit from the processing efficiency of system 400. For example, system 600B can implement downsampler 606, image processor 610, rearranger 614, and machine learning model 632. Figure 4The image processor 406, and in addition, the machine learning model 632 can perform operations related to... Figure 4 The machine learning model 412 operates essentially the same way. For example, image data 604 can play a similar role in system 600B as... Figure 4 The image data 404 plays the same role in system 400. The downsampler 606, image processor 610, rearranger 614, and machine learning model 632 can play the same role in system 600B as... Figure 4 The image processor 406 plays the same role in system 400. Image data 612 can play the same role in system 600B. Figure 4 Image data 408 plays the same role in system 400. Image data 630 can play the same role in system 600B. Figure 4 The image data 410 plays the same role in system 400. In addition, the machine learning model 632 can perform the same... Figure 4 The machine learning model 412 operates essentially the same way. Image data 634 can play a similar role in system 600B. Figure 4 The image data 414 plays the same role in system 400.

[0096] For example, sensor 602 can provide image data 604 and image data 622. Image data 604 and image data 622 can be sequential images in a series of images (e.g., frames of video data). Downsampler 606 can downsample image data 604 to generate image data 608. Image data 608 can be smaller than image data 604. For example, image data 604 can have a size of H*W, and image data 608 can have a size of H / 2*W / 2. Image processor 610 can process image data 612 to generate image data 612 (which can have a size of H / 2*W / 2*3). Rearranger 628 can rearrange image data 622 to generate image data 630. Image data 630 can have the same size as image data 622, but can be rearranged. For example, image data 622 can have a dimension of H*W, and image data 630 can have a dimension of H / 2*W / 2*4. Machine learning model 632 can process image data 612 and image data 630 to generate image data 634.

[0097] Machine learning model 632 can perform with Figure 2A System 200 machine learning model 218 and Figure 4The machine learning model 412 of system 400 operates substantially similarly. Machine learning model 632 can be implemented in image processor 124 or in another processor of image processing device 104. Machine learning model 412 can be trained to receive processed lower-resolution first image data and rearranged second image data as input, and generate processed second image data. For example, machine learning model 632 can be trained using backpropagation training techniques by receiving processed lower-resolution image data (e.g., a processed first frame with a resolution lower than the original) and unprocessed rearranged image data (e.g., an unprocessed second frame) as input. Machine learning model 632 can then generate processed full-resolution second image data as output. Separately, the second image data can be processed (at full resolution), and the processed second image data can be compared with the output of machine learning model 632 to determine the loss based on the difference between the processed image data and the output of machine learning model 632. Machine learning model 632 can be tuned (e.g., the weights or parameters of machine learning model 632 can be adjusted) to reduce the loss in further iterations of the training process. The backpropagation process can be repeated any number of times to train the machine learning model 632.

[0098] During inference (e.g., after training), machine learning model 632 may receive image data 612 (e.g., which may be processed by image processor 610) (which may be a tensor with dimensions H / 2*W / 2*3) and image data 630 (which may be a tensor with dimensions H / 2*W / 2*4) as input. In some cases, image data 612 and image data 630 may be concatenated to form a single tensor with dimensions H / 2*W / 2*7. Machine learning model 632 may include one or more convolutional layers. The size of the convolutional layers may be set to operate according to the dimensions of the single tensor (H / 2*W / 2*7). Based on image data 612 and image data 630, machine learning model 632 may generate image data 634. Image data 634 may be substantially similar to the image data that would be generated by processing image data 622 using image processor 610. Image data 634 may have dimensions H*W (e.g., image data 634 may be full-resolution image data). For example, image data 634 can be approximated by the result of processing image data 622 using image processor 610. However, because image processor 610 does not process image data 622 in system 600B, system 600B saves computational resources (processing time and power) compared to a system that uses image processor 610 to process image data 622.

[0099] By processing image data 608 at image processor 610 instead of image data 604 at image processor 610, system 600B can save computational resources compared to processing image data 604 at image processor 610. Additionally, image data 604 and image data 622 can be sequential images in a series of images (e.g., frames of video data). By processing image data 608 (which is based on image data 604) at image processor 610 instead of image data 622 at image processor 610, system 600B can save computational resources compared to processing both image data 604 and image data 622 at image processor 610. Furthermore, system 600B can use machine learning model 632 to generate image data 634, which can be substantially similar to the result of processing image data 622 at image processor 610. In this way, system 600B can save computational resources while generating substantially the same result compared to processing image data 622 at image processor 610. In this way, system 600B can use machine learning model 632 to reduce the amount of image data processed by image processor 610 by reducing the number of frames of image data processed by image processor 610, thereby saving computing resources.

[0100] Figure 7 This is a flowchart of a process 700 for processing image data according to various aspects of this disclosure. One or more operations of process 700 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a camera, a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, or other types of computing devices. One or more operations of process 700 may be implemented as software components that execute and run on one or more processors.

[0101] At box 702, the computing device (or one or more components thereof) can acquire image data with a first resolution. For example, Figure 2A The system 200 can acquire image data 204. The image data 204 may have a first resolution (e.g., H*W).

[0102] At box 704, a computing device (or one or more components thereof) may downsample image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution. For example, downsampler 206 of system 200 may downsample image data 204 to generate image data 208. Image data 208 may have a second resolution (e.g., H / 2 * W / 2). In some aspects, in order to downsample image data, the computing device (or one or more components thereof) may merge the image data (e.g., as per [reference to image data]). Figure 2B (As illustrated and described).

[0103] At box 706, a computing device (or one or more components thereof) can process downsampled image data to generate processed downsampled image data. For example, image processor 210 of system 200 can process image data 208 to generate image data 212.

[0104] In some aspects, processing the image data at box 706 may or may include: applying one or more filters to the downsampled image data; reducing noise in the downsampled image data; sharpening the downsampled image data; demosaicing the downsampled image data; and / or remosaicing the downsampled image data. In the same aspect, at box 706, the computing device (or one or more components thereof) may perform sensor-related processing (e.g., defective pixel correction and demosaicing), lens processing (e.g., shading correction), color correction, hue and gamma processing, noise reduction, sharpening color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, or some combination thereof.

[0105] At box 708, a computing device (or one or more components thereof) may use a machine learning model to generate upsampled image data based on processed downsampled image data and image data. For example, machine learning model 218 of system 200 may generate image data 220 based on image data 212 and image data 216 (which may be rearranged image data 204).

[0106] In some aspects, the upsampled image data may have a first resolution. For example, the image data may be full-resolution image data, and the upsampled image data may also be full-resolution image data. In some aspects, the image data may be or may include raw image data captured by the device's image sensor, and the downsampled image data may be processed by the device's image signal processor (ISP). For example, image data 204 may have already been captured by sensor 202 of system 200, and image data 212 may be processed by image processor 210 of system 200.

[0107] In some aspects, the computing device (or one or more components thereof) can rearrange image data to generate rearranged image data; and provide the rearranged image data and downsampled image data as input to a machine learning model. For example, the rearranger 214 of system 200 can rearrange image data 204 to generate image data 216, and provide image data 216 and image data 212 to a machine learning model 218. In some aspects, in order to rearrange the image data, the computing device (or one or more components thereof) can rearrange the image data into a tensor having dimensions related to the dimensions of the downsampled image data. For example, the rearranger 214 can rearrange image data 204 to have dimensions H / 2*W / 2*4 based on image data 208 having dimensions H / 2*W / 2*3 and / or based on image data 212 having dimensions H / 2*W / 2*3.

[0108] In some aspects, a machine learning model can be trained to receive first image data with a first resolution and second image data with a second resolution as input, and generate third image data with the first resolution. In some aspects, the machine learning model can be or may include a neural network comprising two or more convolutional layers.

[0109] In some respects, computing devices (or one or more components thereof) can display, store, and / or transmit upsampled image data.

[0110] In some aspects, the image data may be first image data, and a computing device (or one or more components thereof) may acquire second image data; and processed second image data is generated based on the upsampled image data and the second image data. For example, Figure 6ASystem 600A can acquire image data 604 (which may be first image data) (e.g., at box 702). System 600A can downsample image data 604 at downsampler 606 to generate image data 608 (e.g., at box 704). 600A can process image data 608 at image processor 610 to generate image data 612 (e.g., at box 706). System 600A can use machine learning model 618 to generate image data 620 based on image data 612 and image data 616 (which may be rearranged image data 604). System 600A can acquire image data 622 (which may be second image data). System 600A can feed image data 620 and image data 622 to machine learning model 624 to generate image data 626 (e.g., at box 708).

[0111] In such aspects, the first image data may be a first frame of video data, and the second image data may be a second frame of image data. In some aspects, a computing device (or one or more components thereof) may further acquire third image data and process the third image data to generate processed third image data. In such aspects, processed second image data may be further generated based on the processed third image data.

[0112] In some aspects, a computing device (or one or more components thereof) may acquire second image data having a first resolution; downsample the second image data to generate downsampled second image data having a second resolution; process the downsampled second image data to generate processed downsampled second image data; and use a machine learning model to generate upsampled second image data based on the processed downsampled second image data and the second image data. For example, after processing image data 204 to generate image data 220, system 200 may receive the second image data and may downsample the second image data at downsampler 206 to generate second downsampled image data. System 200 may process the second downsampled image data at image processor 210 to generate second processed image data. System 200 may provide the second processed image data and the second rearranged image data to machine learning model 218 to generate second upsampled image data.

[0113] Figure 8This is a flowchart of a process 800 for processing image data according to various aspects of this disclosure. One or more operations of process 800 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a camera, a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capability to perform process 800. One or more operations of process 800 may be implemented as software components that execute and run on one or more processors.

[0114] At box 802, the computing device (or one or more components thereof) can obtain the first image data. For example, Figure 4 The system 400 can obtain image data 404.

[0115] At block 804, a computing device (or one or more components thereof) may process the first image data to generate processed first image data. For example, system 400 may process image data 404 at image processor 406 to generate image data 408.

[0116] In some aspects, processing the image data at box 804 may or may include: applying one or more filters to the downsampled image data; reducing noise in the downsampled image data; sharpening the downsampled image data; demosaicing the downsampled image data; and / or remosaicing the downsampled image data. In the same aspect, at box 804, the computing device (or one or more components thereof) may perform sensor-related processing (e.g., defective pixel correction and demosaicing), lens processing (e.g., shading correction), color correction, hue and gamma processing, noise reduction, sharpening color space conversion, pixel interpolation, automatic exposure control (AEC), automatic gain control (AGC), contrast detection autofocus (CDAF), automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, or some combination thereof.

[0117] At box 806, the computing device (or one or more components thereof) may obtain the second image data. For example, system 400 may obtain image data 410. In some aspects, the first image data may be a first frame of video data; and the second image data may be a second frame of video data.

[0118] At box 808, a computing device (or one or more components thereof) may use a machine learning model to generate processed second image data based on processed first image data and second image data. For example, system 400 may provide image data 408 and image data 410 to machine learning model 412, and machine learning model 412 may generate image data 414 based on image data 408 and image data 410.

[0119] In some respects, a machine learning model can be trained to receive image data and first processed image data as input, and generate second processed image data. In other respects, a machine learning model can be or may include a neural network comprising two or more convolutional layers.

[0120] In some aspects, a computing device (or one or more components thereof) can acquire third image data and process the third image data to generate processed third image data. In such aspects, second image data can be further generated based on the processed third image data. For example, system 400 can acquire third image data and process the third image data at image processor 406 to generate processed third image data. System 400 can provide the processed third image data, together with image data 408 and image data 410, to machine learning model 412 to generate image data 414.

[0121] In some aspects, to process the first image data, a computing device (or one or more components thereof) may downsample the first image data to generate downsampled first image data; process the downsampled first image data to generate processed downsampled image data; and generate upsampled first image data based on the processed downsampled image data and the first image data. In such aspects, processed second image data may be generated based on the upsampled first image data. For example, Figure 6A System 600A can acquire image data 604 (which may be first image data) (e.g., at box 802). System 600A can downsample image data 604 at downsampler 606 to generate image data 608, and process image data 608 at image processor 610 to generate image data 612 (e.g., at box 804). System 600A can use machine learning model 618 to generate image data 620 based on image data 612 and image data 616 (which may be rearranged image data 604) (e.g., at box 804). System 600A can acquire image data 622 (which may be second image data) (e.g., at box 806). System 600A can feed image data 620 and image data 622 to machine learning model 624 to generate image data 626 (e.g., at box 808).

[0122] In some examples, as previously noted, the methods described herein (e.g., Figure 7 The process 700 Figure 8 The process 800 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of these methods may be performed by... Figure 1 Image processing equipment 104 Figure 1 Image processor 124 Figure 1 The host processor 126 Figure 1 ISP 128 Figure 2A System 200 Figure 3A 300 neural networks Figure 4 System 400 Figure 5 Neural network 500, Figure 6A System 600A, Figure 6B The system 600B or another system or device performs these methods. In another example, these methods (e.g., Figure 7 The process 700 Figure 8 One or more of the processes 800 and / or other methods described herein may be used by Figure 11 The computing device architecture 1100 shown is implemented wholly or in part. For example, it has Figure 11 The computing device of the computing device architecture 1100 shown may include Figure 1 Image processing equipment 104 Figure 1 Image processor 124 Figure 1 The host processor 126 Figure 1 ISP 128 Figure 2A System 200 Figure 3A 300 neural networks Figure 4 System 400 Figure 5 Neural network 500, Figure 6A System 600A, Figure 6B The system 600B may include or be incorporated in components of it, and may perform the operation of processes 700, 800, and / or other processes described herein. In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0123] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0124] Processes 700, 800, and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0125] Additionally, processes 700, 800, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or implemented in combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0126] As noted above, various aspects of this disclosure may utilize machine learning models or systems.

[0127] Figure 9 This is an exemplary example of a neural network 900 (e.g., a deep learning neural network) that can be used to implement the machine learning-based image processing, image generation, feature segmentation, implicit neural representation generation, rendering, and / or classification described above. The neural network 900 can be... Figure 2A Machine learning model 218 Figure 3A 300 neural networks Figure 4Machine learning model 412 Figure 5 Neural network 500, Figure 6A Machine learning model 618 Figure 6A Machine learning model 624 and / or Figure 6B Examples of machine learning models 632, or those that can be implemented.

[0128] Input layer 902 includes input data. In an exemplary example, input layer 902 may include image data (e.g., Figure 2A Image data 212, Figure 2A Image data 216 Figure 3A Image data 302, Figure 3A Image data 304 Figure 4 Image data 408 Figure 4 Image data 410 Figure 5 Image data 502, Figure 5 Image data 504 Figure 6A Image data 612 Figure 6A Image data 616 Figure 6A Image data 620 Figure 6A Image data 622 Figure 6B Image data 612 and / or Figure 6B Image data 630). The neural network 900 includes multiple hidden layers 906a, 906b to 906n. Hidden layers 906a, 906b to 906n comprise "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. The neural network 900 also includes an output layer 904, which provides the output produced by the processing performed by the hidden layers 906a, 906b to 906n. In an exemplary example, output layer 904 can provide image data (e.g., image data). Figure 2A Image data 220 Figure 3A Image data 320, Figure 4 Image data 414 Figure 5 Image data 512 Figure 6A Image data 620 Figure 6A Image data 626 and / or Figure 6B Image data 634).

[0129] The neural network 900 may be or may include a multi-layer neural network with interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information while processing it. In some cases, the neural network 900 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, the neural network 900 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.

[0130] Information can be exchanged between nodes through node-to-node interconnects between layers. Nodes in input layer 902 can activate the node set in the first hidden layer 906a. For example, as shown, each input node in input layer 902 is connected to each node in the first hidden layer 906a. Nodes in the first hidden layer 906a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 906b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 906b can then activate nodes in the next hidden layer, and so on. Finally, the output of hidden layer 906n can activate one or more nodes in output layer 904, providing the output at those nodes. In some cases, although a node in neural network 900 (e.g., node 908) is shown as having multiple output lines, the node has a single output and all lines shown as outputs from the node represent the same output value.

[0131] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 900. Once the neural network 900 is trained, it can be called a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. This interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 900 to adapt to the input and learn as more and more data is processed.

[0132] The neural network 900 can be pre-trained to process features from the data in the input layer 902 using different hidden layers 906a, 906b to 906n, so as to provide an output through the output layer 904. In an example where the neural network 900 is used to identify features in an image, the neural network 900 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating features in the image (for feature segmentation machine learning systems) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the label of the image may be [0 0 1 0 0 0 0 0 0 0].

[0133] In some cases, the neural network 900 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the neural network 900 is trained well enough to accurately tune the weights of each layer.

[0134] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 900. The weights are initially randomized before training the neural network 900. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0135] As noted above, for the first training iteration of a neural network 900, the output may include values ​​due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 900 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function is mean squared error (MSE), which is defined as... The loss can be set to equal E. total The value of .

[0136] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. Neural Network 900 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0137] Neural Network 900 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 900 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0138] Figure 10 This is an exemplary example of a Convolutional Neural Network (CNN) 1000. The input layer 1002 of the CNN 1000 includes data representing an image or frame. For example, the data may include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 1004, an optional non-linear activation layer, a pooling hidden layer 1006, and a fully connected layer 1008 (which may be hidden) to obtain an output at the output layer 1010. Although... Figure 10 Only one hidden layer from each hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in a CNN 1000. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0139] The first layer of a CNN 1000 can be a convolutional hidden layer 1004. The convolutional hidden layer 1004 analyzes the image data from the input layer 1002. Each node in the convolutional hidden layer 1004 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 1004 can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 1004. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of the filter. In an exemplary example, if the input image comprises a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 1004. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 1004 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.

[0140] The convolutional property of the convolutional hidden layer 1004 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 1004 may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 1004. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 1004. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 1004.

[0141] The map construction from the input layer to the convolutional hidden layer 1004 is called an activation map (or feature map). An activation map includes values ​​for each node representing the filter results at each location in the input volume. Activation maps may include arrays containing various sums of values ​​produced by the filter on each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 1004 may include several activation maps to identify multiple features in the image. Figure 10 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 1004 can detect three different types of features, each of which is detectable across the entire image.

[0142] In some examples, nonlinear hidden layers can be applied after the convolutional hidden layer 1004. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can increase the nonlinearity of the CNN 1000 without affecting the receptive field of the convolutional hidden layer 1004.

[0143] A pooling hidden layer 1006 can be applied after the convolutional hidden layer 1004 (and, in use, after the non-linear hidden layer). The pooling hidden layer 1006 is used to simplify the information in the output of the convolutional hidden layer 1004. For example, the pooling hidden layer 1006 can take each activation map output from the convolutional hidden layer 1004 and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 1006 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 1004. Figure 10 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 1004.

[0144] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 1004. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node being a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 1004, the output from pooling hidden layer 1006 will be an array of 12×12 nodes.

[0145] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0146] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere within a region of an image, discarding the exact location information. This can be done without affecting the results of feature detection, because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 1000.

[0147] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 1006 to each output node in the output layer 1010. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 1004 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 1006 comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region in each of the three feature maps. Extending this example, the output layer 1010 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 1006 is connected to each node of the output layer 1010.

[0148] The fully connected layer 1008 takes the output of the previous pooling hidden layer 1006 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 1008 can determine the high-level features most relevant to a particular class and may include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 1008 and the pooling hidden layer 1006 can be computed to obtain the probabilities for different classes. For example, if the CNN 1000 is used to predict that the object in an image is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0149] In some examples, the output from output layer 1010 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes from which CNN 1000 must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector represents objects of ten different classes as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0150] Figure 11 Example computing device architecture 1100 illustrates example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1100 may include, implement Figure 1 Image processing equipment 104 Figure 1 Image processor 124 Figure 1 The host processor 126 Figure 1 ISP 128 Figure 2A System 200 Figure 3A 300 neural networks Figure 4 System 400 Figure 5 Neural network 500, Figure 6A System 600A, Figure 6BAny or all of the 600B systems, or included in them.

[0151] The components of computing device architecture 1100 are shown to communicate electrically with each other using a connection 1112 (such as a bus). The example computing device architecture 1100 includes a processing unit (CPU or processor) 1102 and a computing device connection 1112 that couples various computing device components, including computing device memories 1110 (such as read-only memory (ROM) 1108 and random access memory (RAM) 1106), to the processor 1102.

[0152] The computing device architecture 1100 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1102. The computing device architecture 1100 may copy data from memory 1110 and / or storage device 1114 to cache 1104 for fast access by the processor 1102. In this way, the cache can provide performance improvements by avoiding latency for the processor 1102 while waiting for data. These and other modules may control or be configured to control the processor 1102 to perform various actions. Other computing device memories 1110 may also be used. Memory 1110 may include various different types of memory with different performance characteristics. The processor 1102 may include any general-purpose processor and hardware or software services configured to control the processor 1102 (such as services 1 1116, 2 1118, and 3 1120 stored in storage device 1114), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1102 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.

[0153] To enable user interaction with computing device architecture 1100, input device 1122 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1124 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some instances, multi-mode computing devices allow users to provide multiple types of input to communicate with computing device architecture 1100. Communication interface 1126 typically governs and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0154] Storage device 1114 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 1106, read-only memory (ROM) 1108, and hybrid forms thereof. Storage device 1114 may include services 1116, 1118, and 1120 for controlling processor 1102. Other hardware or software modules are envisioned. Storage device 1114 may be connected to computing device connection 1112. In one aspect, a hardware module performing a specific function may include software components for performing that function stored in a computer-readable medium connected to necessary hardware components such as processor 1102, connection 1112, output device 1124, etc.

[0155] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0156] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0157] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or a particular aspect. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0158] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0159] The various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While a flowchart may describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the diagrams. A process can correspond to a method, function, process, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0160] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0161] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0162] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.

[0163] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0164] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0165] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including such variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0166] Those skilled in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of this description.

[0167] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0168] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0169] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, the claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language set "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in that set. For example, the claim language stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0170] Claims using phrases such as "at least one processor, the at least one processor being configured to," or other languages ​​indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim stating "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks involving operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim stating "at least one processor, the at least one processor being configured to: X, Y, and Z" could mean that any single processor can perform only a subset of operations X, Y, and Z.

[0171] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0172] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0173] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0174] The exemplary aspects of this disclosure include:

[0175] Aspect 1. An apparatus for processing data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: acquire image data having a first resolution; downsample the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; process the downsampled image data to generate processed downsampled image data; and use a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0176] Aspect 2. The apparatus according to aspect 1, wherein, in order to process the downsampled image data, the at least one processor is configured to perform at least one of the following: applying one or more filters to the downsampled image data; reducing noise in the downsampled image data; sharpening the downsampled image data; de-mosaicing the downsampled image data; or re-mosaicing the downsampled image data.

[0177] Aspect 3. The apparatus according to any one of Aspect 1 or 2, wherein the machine learning model is trained to receive first image data having the first resolution and second image data having the second resolution as input, and to generate third image data having the first resolution.

[0178] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

[0179] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the upsampled image data has the first resolution.

[0180] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the image data includes raw image data captured by the image sensor of the apparatus, and wherein the downsampled image data is processed by the image signal processor (ISP) of the apparatus.

[0181] Aspect 7. The apparatus according to any one of Aspects 1 to 6, wherein the at least one processor is further configured to: rearrange the image data to generate rearranged image data; and provide the rearranged image data and the downsampled image data as input to the machine learning model.

[0182] Aspect 8. The apparatus according to aspect 7, wherein, in order to rearrange the image data, the at least one processor is configured to rearrange the image data into a tensor having a dimension related to the dimension of the downsampled image data.

[0183] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein the at least one processor is further configured to perform at least one of: displaying, storing or transmitting the upsampled image data.

[0184] Aspect 10. The apparatus according to any one of aspects 1 to 9, wherein, in order to downsample the image data, the at least one processor is configured to merge the image data.

[0185] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the image data includes first image data, and wherein the at least one processor is further configured to: obtain second image data; and generate processed second image data based on the upsampled image data and the second image data.

[0186] Aspect 12. The apparatus according to aspect 11, wherein: the first image data includes a first frame of video data; and the second image data includes a second frame of the video data.

[0187] Aspect 13. The apparatus according to any one of Aspects 11 or 12, wherein the at least one processor is further configured to: obtain third image data; and process the third image data to generate processed third image data; wherein the processed second image data is further generated based on the processed third image data.

[0188] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein the image data comprises first image data, and wherein the at least one processor is further configured to: obtain second image data having the first resolution; downsample the second image data to generate downsampled second image data, wherein the downsampled second image data has the second resolution; process the downsampled second image data to generate processed downsampled second image data; and use the machine learning model to generate upsampled second image data based on the processed downsampled second image data and the second image data.

[0189] Aspect 15. A method for processing data, the method comprising: obtaining image data having a first resolution; downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; processing the downsampled image data to generate processed downsampled image data; and using a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0190] Aspect 16. The method according to aspect 15, wherein processing the downsampled image data includes at least one of: applying one or more filters to the downsampled image data; reducing noise in the downsampled image data; sharpening the downsampled image data; de-mosaicing the downsampled image data; or re-mosaicing the downsampled image data.

[0191] Aspect 17. The method according to any one of Aspects 15 or 16, wherein the machine learning model is trained to receive first image data having the first resolution and second image data having the second resolution as input, and to generate third image data having the first resolution.

[0192] Aspect 18. The method according to any one of Aspects 15 to 17, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

[0193] Aspect 19. The method according to any one of Aspects 15 to 18, wherein the upsampled image data has the first resolution.

[0194] Aspect 20. The method according to any one of Aspects 15 to 19, wherein the image data includes raw image data captured by an image sensor of the device, and wherein the downsampled image data is processed by an image signal processor (ISP) of the device.

[0195] Aspect 21. The method according to any one of Aspects 15 to 20, the method further comprising: rearranging the image data to generate rearranged image data; and providing the rearranged image data and the downsampled image data as input to the machine learning model.

[0196] Aspect 22. The method according to aspect 21, wherein rearranging the image data includes rearranging the image data into a tensor having dimensions related to the dimensions of the downsampled image data.

[0197] Aspect 23. The method according to any one of Aspects 15 to 22, the method further comprising at least one of: displaying, storing or transmitting the upsampled image data.

[0198] Aspect 24. The method according to any one of Aspects 15 to 23, wherein downsampling the image data includes merging the image data.

[0199] Aspect 25. The method according to any one of Aspects 15 to 24, wherein the image data includes first image data, and wherein the method further includes: obtaining second image data; and generating processed second image data based on the upsampled image data and the second image data.

[0200] Aspect 26. The method according to aspect 25, wherein: the first image data includes a first frame of video data; and the second image data includes a second frame of the video data.

[0201] Aspect 27. The method according to any one of Aspects 25 or 26, the method further comprising: obtaining third image data; and processing the third image data to generate processed third image data; wherein the processed second image data is further generated based on the processed third image data.

[0202] Aspect 28. The method according to any one of Aspects 15 to 27, wherein the image data includes first image data, and wherein the method further comprises: obtaining second image data having the first resolution; downsampling the second image data to generate downsampled second image data, wherein the downsampled second image data has the second resolution; processing the downsampled second image data to generate processed downsampled second image data; and using the machine learning model to generate upsampled second image data based on the processed downsampled second image data and the second image data.

[0203] Aspect 29. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to: obtain image data having a first resolution; downsample the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; process the downsampled image data to generate processed downsampled image data; and use a machine learning model to generate upsampled image data based on the processed downsampled image data and the image data.

[0204] Aspect 30. An apparatus for processing data, the apparatus comprising: means for obtaining image data having a first resolution; means for downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; means for processing the downsampled image data to generate processed downsampled image data; and means for generating upsampled image data based on the processed downsampled image data and the image data using a machine learning model.

[0205] Aspect 31. An apparatus for processing data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: acquire first image data; process the first image data to generate processed first image data; acquire second image data; and generate processed second image data using a machine learning model based on the processed first image data and the second image data.

[0206] Aspect 32. The apparatus according to aspect 31, wherein the machine learning model is trained to receive image data and first processed image data as input, and to generate second processed image data.

[0207] Aspect 33. The apparatus according to any one of Aspects 31 or 32, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

[0208] Aspect 34. The apparatus according to any one of aspects 31 to 33, wherein: the first image data includes a first frame of video data; and the second image data includes a second frame of the video data.

[0209] Aspect 35. The apparatus according to any one of aspects 31 to 34, wherein, in order to process the first image data, the at least one processor is configured to perform at least one of the following: applying one or more filters to the first image data; reducing noise in the first image data; sharpening the first image data; de-mosaicing the first image data; or re-mosaicing the first image data.

[0210] Aspect 36. The apparatus according to any one of aspects 31 to 35, wherein the at least one processor is further configured to: obtain third image data; and process the third image data to generate processed third image data; wherein the processed second image data is further generated based on the processed third image data.

[0211] Aspect 37. The apparatus according to any one of Aspects 31 to 36, wherein: in order to process the first image data, the at least one processor is configured to: downsample the first image data to generate downsampled first image data; process the downsampled first image data to generate processed downsampled image data; and generate upsampled first image data based on the processed downsampled image data and the first image data; and generate the processed second image data based on the upsampled first image data.

[0212] Aspect 38. A method for processing data, the method comprising: obtaining first image data; processing the first image data to generate processed first image data; obtaining second image data; and using a machine learning model to generate processed second image data based on the processed first image data and the second image data.

[0213] Aspect 39. The method according to aspect 38, wherein the machine learning model is trained to receive image data and first processed image data as input, and to generate second processed image data.

[0214] Aspect 40. The method according to any one of Aspects 38 or 39, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

[0215] Aspect 41. The method according to any one of Aspects 38 to 40, wherein: the first image data includes a first frame of video data; and the second image data includes a second frame of the video data.

[0216] Aspect 42. The method according to any one of Aspects 38 to 41, wherein processing the first image data includes at least one of: applying one or more filters to the first image data; reducing noise in the first image data; sharpening the first image data; de-mosaicing the first image data; or re-mosaicing the first image data.

[0217] Aspect 43. The method according to any one of aspects 38 to 42, the method further comprising: obtaining third image data; and processing the third image data to generate processed third image data; wherein the processed second image data is further generated based on the processed third image data.

[0218] Aspect 44. The method according to any one of Aspects 38 to 43, wherein: processing the first image data includes: downsampling the first image data to generate downsampled first image data; processing the downsampled first image data to generate processed downsampled image data; generating upsampled first image data based on the processed downsampled image data and the first image data; and generating the processed second image data based on the upsampled first image data.

[0219] Aspect 45. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to perform any one of aspects 15 to 28.

[0220] Aspect 46. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 15 to 28.

[0221] Aspect 47. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to perform any one of aspects 38 to 44.

[0222] Aspect 48. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 38 to 44.

Claims

1. An apparatus for processing data, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: Obtain image data with a first resolution; The image data is downsampled to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; The downsampled image data is processed to generate processed downsampled image data; as well as Machine learning models are used to generate upsampled image data based on the processed downsampled image data and the image data.

2. The apparatus according to claim 1, wherein, In order to process the downsampled image data, the at least one processor is configured to perform at least one of the following: Apply one or more filters to the downsampled image data; Reduce noise in the downsampled image data; Sharpen the downsampled image data; The downsampled image data is then de-mosaiced. or The downsampled image data is then re-mosaicized.

3. The apparatus of claim 1, wherein the machine learning model is trained to receive first image data having the first resolution and second image data having the second resolution as input, and to generate third image data having the first resolution.

4. The apparatus of claim 1, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

5. The apparatus of claim 1, wherein the upsampled image data has the first resolution.

6. The apparatus of claim 1, wherein the image data includes raw image data captured by the image sensor of the apparatus, and wherein the downsampled image data is processed by the image signal processor (ISP) of the apparatus.

7. The apparatus of claim 1, wherein the at least one processor is further configured to: The image data is rearranged to generate rearranged image data; and The rearranged image data and the downsampled image data are provided as input to the machine learning model.

8. The apparatus according to claim 7, wherein, In order to rearrange the image data, the at least one processor is configured to rearrange the image data into a tensor having dimensions related to the dimensions of the downsampled image data.

9. The apparatus of claim 1, wherein the at least one processor is further configured to perform at least one of the following: displaying, storing, or transmitting the upsampled image data.

10. The apparatus according to claim 1, wherein, In order to downsample the image data, the at least one processor is configured to merge the image data.

11. The apparatus of claim 1, wherein the image data includes first image data, and wherein the at least one processor is further configured to: Obtain second image data; and The processed second image data is generated based on the upsampled image data and the second image data.

12. The apparatus according to claim 11, wherein: The first image data includes the first frame of the video data; and The second image data includes the second frame of the video data.

13. The apparatus of claim 11, wherein the at least one processor is further configured to: Obtain third image data; and Process the third image data to generate processed third image data; The processed second image data is further generated based on the processed third image data.

14. The apparatus of claim 1, wherein the image data includes first image data, and wherein the at least one processor is further configured to: Obtain second image data having the first resolution; The second image data is downsampled to generate downsampled second image data, wherein the downsampled second image data has the second resolution; The downsampled second image data is processed to generate processed downsampled second image data; as well as The machine learning model is used to generate upsampled second image data based on the processed downsampled second image data and the second image data.

15. A method for processing data, the method comprising: Obtain image data with a first resolution; The image data is downsampled to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; The downsampled image data is processed to generate processed downsampled image data; as well as Machine learning models are used to generate upsampled image data based on the processed downsampled image data and the image data.

16. The method of claim 15, wherein processing the downsampled image data comprises at least one of the following: Apply one or more filters to the downsampled image data; Reduce noise in the downsampled image data; Sharpen the downsampled image data; Perform demosaic processing on the downsampled image data; or The downsampled image data is then re-mosaicized.

17. The method of claim 15, wherein the machine learning model is trained to receive first image data having the first resolution and second image data having the second resolution as input, and to generate third image data having the first resolution.

18. The method of claim 15, wherein the machine learning model comprises a neural network, the neural network comprising two or more convolutional layers.

19. The method of claim 15, wherein the upsampled image data has the first resolution.

20. The method of claim 15, wherein the image data comprises raw image data captured by the image sensor of the device, and wherein the downsampled image data is processed by the image signal processor (ISP) of the device.

21. The method according to claim 15, further comprising: The image data is rearranged to generate rearranged image data; as well as The rearranged image data and the downsampled image data are provided as input to the machine learning model.

22. The method of claim 21, wherein rearranging the image data comprises rearranging the image data into a tensor having dimensions related to the dimensions of the downsampled image data.

23. The method of claim 15, further comprising at least one of: displaying, storing, or transmitting the upsampled image data.

24. The method of claim 15, wherein downsampling the image data includes merging the image data.

25. The method of claim 15, wherein the image data includes first image data, and wherein the method further comprises: Obtain the second image data; as well as The processed second image data is generated based on the upsampled image data and the second image data.

26. The method of claim 25, wherein: The first image data includes the first frame of the video data; and The second image data includes the second frame of the video data.

27. The method of claim 25, further comprising: Obtain third image data; as well as Process the third image data to generate processed third image data; The processed second image data is further generated based on the processed third image data.

28. The method of claim 15, wherein the image data includes first image data, and wherein the method further comprises: Obtain second image data having the first resolution; The second image data is downsampled to generate downsampled second image data, wherein the downsampled second image data has the second resolution; The downsampled second image data is processed to generate processed downsampled second image data; as well as The machine learning model is used to generate upsampled second image data based on the processed downsampled second image data and the second image data.

29. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to: Obtain image data with a first resolution; The image data is downsampled to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; The downsampled image data is processed to generate processed downsampled image data; and Machine learning models are used to generate upsampled image data based on the processed downsampled image data and the image data.

30. An apparatus for processing data, the apparatus comprising: A component used to obtain image data with a first resolution; A component for downsampling the image data to generate downsampled image data, wherein the downsampled image data has a second resolution lower than the first resolution; Components for processing the downsampled image data to generate processed downsampled image data; and A component for generating upsampled image data using a machine learning model based on the processed downsampled image data and the image data.