Machine learning-based image adjustment

Machine learning-based image adjustment using neural networks generates spatially varying maps for dynamic and content-aware image processing, addressing the inefficiencies of conventional ISPs by enhancing image quality and reducing manual parameter adjustments.

JP7758685B2Active Publication Date: 2025-10-22QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022566701
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-10
Filing Date
2021-05-13
Publication Date
2025-10-22
Estimated Expiration
2041-05-13

AI Technical Summary

Technical Problem

Conventional image signal processors (ISPs) require manual adjustment of numerous parameters, which is time-consuming and inflexible, leading to inefficient and non-content-based image processing.

Method used

Implement machine learning-based image adjustment using trained neural networks to generate spatially varying maps for image processing functions, allowing dynamic and content-based adjustments.

Benefits of technology

Enables high-precision, content-aware image processing with reduced manual intervention, improving efficiency and quality of image adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758685000004
    Figure 0007758685000004
  • Figure 0007758685000005
    Figure 0007758685000005
  • Figure 0007758685000006
    Figure 0007758685000006
Patent Text Reader

Abstract

The imaging system can acquire image data, for example, from an image sensor. The imaging system can provide the image data as input data to a machine learning system, which can generate one or more maps based on the image data. Each map can specify the strength with which an image processing function should be applied to each pixel of the image data. Different maps may be generated for different image processing functions, such as noise reduction, sharpening, or color saturation. The imaging system can generate a modified image based on the image data and the one or more maps, for example, by applying each of the one or more image processing functions according to each of the one or more maps. The imaging system can provide the image data and the one or more maps to a second machine learning system to generate the modified image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to image processing, and more particularly to systems and techniques for performing machine learning-based image adjustments. [Background technology]

[0002] A camera is a device that uses an image sensor to receive light and capture image frames, such as still images or video frames. A camera may include a processor, such as an image signal processor (ISP), that can receive and process one or more image frames. For example, raw image frames captured by a camera sensor may be processed by an ISP to generate a final image. A camera may be configured with various image capture and image processing settings to alter the appearance of an image. Some camera settings, such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain, are determined and applied before or during the capture of a photograph. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color modifications, may constitute post-processing of a photograph. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Andrew Aitken et al., "Checkerboard artifact free sub-pixel convolution" Summary of the Invention [Problem to be solved by the invention]

[0004] A conventional image signal processor (ISP) has separate, individual blocks that address different segments of an image-based problem space. For example, a typical ISP has separate function blocks that each apply specific operations to raw camera sensor data to produce a final output image. Such function blocks may include blocks for demosaicing, noise reduction (denoising), color processing, and tone mapping, among other image processing functions. Each of these function blocks contains numerous manually adjusted parameters, resulting in an ISP with a large number of manually adjusted parameters (e.g., 10,000 or more) that must be readjusted according to each customer's adjustment preferences. Manual adjustment of such parameters is very time-consuming and expensive, and is therefore typically performed only once. Once adjusted, a conventional ISP typically uses the same parameters for every image. [Means for solving the problem]

[0005] In some examples, systems and techniques are described for performing machine learning-based image adjustment using one or more machine learning systems. An imaging system can acquire image data, for example from an image sensor. The imaging system can provide the image data as input data to the machine learning system, which can generate one or more maps based on the image data. Each map can specify the strength with which an image processing function should be applied to each pixel of the image data. Different maps may be generated for different image processing functions, such as noise reduction, sharpening, or color saturation. The imaging system can generate a modified image based on the image data and the one or more maps, for example, by applying each of the one or more image processing functions according to each of the one or more maps. The imaging system can provide the image data and the one or more maps to a second machine learning system to generate the modified image.

[0006] In one example, an apparatus for image processing is provided. The apparatus includes a memory and one or more processors (e.g., implemented as circuits) coupled to the memory. The one or more processors are configured and capable of: acquiring image data; using the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps associated with a respective image processing function; and generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0007] In another example, a method of image processing is provided, the method including: acquiring image data, using the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps associated with a respective image processing function, and generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0008] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to obtain image data; use the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps associated with a respective image processing function; and generate an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0009] In another example, an apparatus for image processing is provided, the apparatus including: means for acquiring image data; means for generating one or more maps using the image data as input to one or more trained neural networks, each map of the one or more maps associated with a respective image processing function; and means for generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0010] In some embodiments, a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, each value of the plurality of values ​​in the map indicating a strength with which to apply the image processing function to a corresponding region of the image data, which in some embodiments corresponds to a pixel of the image.

[0011] In some embodiments, the one or more maps include a plurality of maps, a first map of the plurality of maps associated with a first image processing function, and a second map of the plurality of maps associated with a second image processing function. In some embodiments, the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function. In some embodiments, the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used in applying the first image processing function to a corresponding region of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used in applying the second image processing function to a corresponding region of the image data.

[0012] In some aspects, the image data includes luminance channel data corresponding to the image, and using the image data as input to the one or more trained neural networks includes using the luminance channel data corresponding to the image as input to the one or more trained neural networks. In some aspects, generating the image based on the image data includes generating the image based on the luminance channel data and chrominance data corresponding to the image.

[0013] In some aspects, the one or more trained neural networks output one or more affine coefficients based on using the image data as input to the one or more trained neural networks, and generating the one or more maps includes at least generating the first map by transforming the image data using the one or more affine coefficients. In some aspects, the image data includes luminance channel data corresponding to an image, and transforming the image data using the one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients. In some aspects, the one or more affine coefficients include a multiplier, and transforming the image data using the one or more affine coefficients includes multiplying luminance values ​​of at least a subset of the image data by the multiplier. In some aspects, the one or more affine coefficients include an offset, and transforming the image data using the one or more affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset. In some aspects, the one or more trained neural networks output the one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first map with one or more gradients in the image data.

[0014] In some embodiments, generating the image based on the image data and the one or more maps includes using the image data and the one or more maps as inputs to a second set of one or more trained neural networks that are separate from the one or more trained neural networks. In some embodiments, generating the image based on the image data and the one or more maps includes demosaicing the image data using the second set of one or more trained neural networks.

[0015] In some aspects, each map of the one or more maps is spatially varied based on different types of objects depicted in the image data. In some aspects, the image data includes an input image having multiple color components for each pixel of the multiple pixels of the image data. In some aspects, the image data includes raw image data from one or more image sensors, the raw image data including at least one color component for each pixel of the multiple pixels of the image data.

[0016] In some aspects, the methods, apparatus, and computer-readable media described above further include an image sensor that captures image data, and acquiring the image data includes acquiring the image data from the image sensor. In some aspects, the methods, apparatus, and computer-readable media described above further include a display screen, and the one or more processors are configured to display the image on the display screen. In some aspects, the methods, apparatus, and computer-readable media described above further include a communications transceiver, and the one or more processors are configured to transmit the image to a recipient device using the communications transceiver.

[0017] In some aspects, the apparatus comprises a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the apparatus includes a camera or multiple cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data.

[0018] This Summary does not identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter, which subject matter should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0019] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings.

[0020] Exemplary embodiments of the present application are described in detail below with reference to the following drawings: [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is a block diagram illustrating an example architecture of an image capture and processing system, according to some examples. [Figure 2] FIG. 1 is a block diagram illustrating an example system including an image processing machine learning (ML) system, according to some examples. [Figure 3A] FIG. 1 is a conceptual diagram illustrating an example of an input image including multiple pixels labeled P0 through P63, according to some examples. [Figure 3B] FIG. 1 is a conceptual diagram illustrating an example of a spatial tuning map, according to some examples. [Figure 4] FIG. 1 is a conceptual diagram illustrating an example of an input noise map that is applied to adjust the strength of noise reduction in an input image to generate a modified image, according to some examples. [Figure 5] 10A-10C are conceptual diagrams illustrating examples of the effect of applying a sharpness adjustment image processing function at different strengths to an example image, in accordance with some examples. [Figure 6] 1A-1C are conceptual diagrams illustrating examples of input sharpness maps that are applied to adjust the strength of sharpening in an input image to generate a modified image, according to some examples. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example gamma curve and various examples of the effect of different values ​​in a tone map translating into different strengths of tone adjustment applied to an example image, according to some examples. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example of an input tone map applied to the strength of a tone adjustment in an input image to generate a modified image, according to some examples. [Figure 9] FIG. 1 is a conceptual diagram illustrating saturation levels in luminance-chrominance (YUV) color space, according to some examples. [Figure 10] 1 is a conceptual diagram illustrating processed variants of an exemplary image, each processed using different alpha (α) values ​​to adjust color saturation, according to some examples. [Figure 11] 1A-1C are conceptual diagrams illustrating examples of input saturation maps applied to the strength and direction of saturation adjustments in an input image to generate a modified image, according to some examples. [Figure 12A] FIG. 1 is a conceptual diagram illustrating a hue, saturation, value (HSV) color space, according to some examples. [Figure 12B] 1 is a conceptual diagram illustrating modification of a hue vector in a luminance-chrominance (YUV) color space, according to some examples. [Figure 13] 1A-1C are conceptual diagrams illustrating examples of input hue maps that are applied to the strength of hue adjustments in an input image to generate a modified image, according to some examples. [Figure 14A]FIG. 1 is a conceptual diagram illustrating an example system including an image processing system that receives an input image and multiple spatially-varying tuning maps, according to some examples. [Figure 14B] FIG. 1 is a conceptual diagram illustrating an example system including an image processing system that receives an input image and a plurality of spatially-varying tuning maps, and an auto-tuning machine learning (ML) system that receives the input image and generates a plurality of spatially-varying tuning maps, according to some examples. [Figure 14C] FIG. 1 is a conceptual diagram illustrating an example system including a self-tuning machine learning (ML) system that receives an input image and generates a plurality of spatially varying tuning maps, according to some examples. [Figure 14D] FIG. 1 is a conceptual diagram illustrating an example system, according to some examples, including an image processing system that receives an input image and multiple spatially-varying tuning maps, and an auto-tuning machine learning (ML) system that receives a downscaled version of the input image and generates smaller spatially-varying tuning maps that are upscaled into the multiple spatially-varying tuning maps. [Figure 15] FIG. 1 is a block diagram illustrating an example of a neural network that may be used by an image processing system and / or a self-tuning machine learning (ML) system, according to some examples. [Figure 16A] FIG. 1 is a block diagram illustrating an example of training an image processing system (e.g., an image processing ML system), according to some examples. [Figure 16B] FIG. 1 is a block diagram illustrating an example of training an auto-tuning machine learning (ML) system, according to some examples. [Figure 17] FIG. 1 is a block diagram illustrating an example system including a self-tuning machine learning (ML) system that generates a spatially-varying tuning map from luminance channel data by generating affine coefficients that modify the luminance channel data according to local linearity constraints, in accordance with some examples. [Figure 18] FIG. 1 is a block diagram illustrating example details of an auto-tuning ML system, according to some examples. [Figure 19] FIG. 1 is a block diagram illustrating an example neural network architecture of a local neural network of an auto-tuning ML system, according to some examples. [Figure 20] FIG. 1 is a block diagram illustrating an example neural network architecture of a global neural network for an auto-tuning ML system, according to some examples. [Figure 21A] FIG. 1 is a block diagram illustrating an example neural network architecture for an auto-tuning ML system, according to some examples. [Figure 21B] FIG. 10 is a block diagram illustrating another example of a neural network architecture for an auto-tuning ML system, according to some examples. [Figure 21C] FIG. 1 is a block diagram illustrating an example neural network architecture for a spatial attention engine, according to some examples. [Figure 21D] FIG. 1 is a block diagram illustrating an example neural network architecture for a channel attention engine, according to some examples. [Figure 22] FIG. 1 is a block diagram illustrating an example of a preconditioned image signal processor (ISP), according to some examples. [Figure 23] FIG. 1 is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP), according to some examples. [Figure 24] FIG. 1 is a block diagram illustrating an example neural network architecture for a machine learning (ML) image signal processor (ISP), according to some examples. [Figure 25A] 1 is a conceptual diagram illustrating an example of the strength of a first tone adjustment applied to an exemplary input image to generate a modified image, according to some examples. [Figure 25B] FIG. 25B is a conceptual diagram illustrating an example of the strength of a second tone adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 26A]FIG. 25B is a conceptual diagram illustrating an example of a strength of a first detail adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 26B] FIG. 25B is a conceptual diagram illustrating an example of the strength of a second detail adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 27A] FIG. 25B is a conceptual diagram illustrating an example of the strength of a first color saturation adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 27B] FIG. 25B is a conceptual diagram illustrating an example of the strength of a second color saturation adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 27C] FIG. 25B is a conceptual diagram illustrating an example of the strength of a third color saturation adjustment applied to the example input image of FIG. 25A to generate a modified image, according to some examples. [Figure 28] 1 is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) receiving as input various tuning parameters used to adjust the ML ISP, according to some examples. [Figure 29] FIG. 1 is a block diagram illustrating examples of particular tuning parameter values ​​that may be provided to a machine learning (ML) image signal processor (ISP), according to some examples. [Figure 30] FIG. 10 is a block diagram illustrating additional examples of specific tuning parameter values ​​that may be provided to a machine learning (ML) image signal processor (ISP), according to some examples. [Figure 31] FIG. 1 is a block diagram illustrating examples of objective functions and various losses that may be used during training of a machine learning (ML) image signal processor (ISP), according to some examples. [Figure 32] FIG. 10 is a conceptual diagram illustrating an example of patch-by-patch model inference that results in non-overlapping output patches at a first image location, according to some examples. [Figure 33]FIG. 33 is a conceptual diagram illustrating an example of patch-by-patch model estimation of FIG. 32 at a second image location, according to some examples. [Figure 34] FIG. 33 is a conceptual diagram illustrating an example of patch-by-patch model estimation of FIG. 32 at a third image location, according to some examples. [Figure 35] FIG. 33 is a conceptual diagram illustrating an example of model estimation for each patch of FIG. 32 at a fourth image location, according to some examples. [Figure 36] FIG. 1 is a conceptual diagram illustrating an example of a spatially fixed tone map applied at the image level, according to some examples. [Figure 37] FIG. 1 is a conceptual diagram illustrating an exemplary application of a spatially varying map to process input image data to generate an output image in which the strength of the saturation adjustment varies spatially, according to some examples. [Figure 38] FIG. 10 is a conceptual diagram illustrating an exemplary application of a spatially-varying map to process input image data to generate an output image in which the strength of the tone adjustment and the strength of the spatially-varying detail adjustment vary spatially, according to some examples. [Figure 39] 1 is a conceptual diagram illustrating an automatically adjusted image generated by an image processing system using one or more tuning maps to adjust an input image, according to some examples. [Figure 40A] FIG. 1 is a conceptual diagram illustrating an output image generated by an image processing system using one or more spatially varying tuning maps generated using a self-tuning machine learning (ML) system, according to some examples. [Figure 40B] FIG. 40B is a conceptual diagram illustrating an example of a spatially-varying tuning map that may be used to generate the output image shown in FIG. 40A from the input image shown in FIG. 40A, according to some examples. [Figure 41A] FIG. 1 is a conceptual diagram illustrating an output image generated by an image processing system using one or more spatially varying tuning maps generated using a self-tuning machine learning (ML) system, according to some examples. [Figure 41B]FIG. 41B is a conceptual diagram illustrating an example of a spatially-varying tuning map that may be used to generate the output image shown in FIG. 41A from the input image shown in FIG. 41A, according to some examples. [Figure 42A] 1 is a flow diagram illustrating an example process for processing image data, according to some examples. [Figure 42B] 1 is a flow diagram illustrating an example process for processing image data, according to some examples. [Figure 43] FIG. 1 illustrates an example computing system for implementing certain aspects described herein. DETAILED DESCRIPTION OF THE INVENTION

[0022] Some aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for purposes of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0023] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with an enabling description for practicing the exemplary embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the present application, as set forth in the appended claims.

[0024] A camera is a device that uses an image sensor to receive light and capture image frames, such as still images or video frames. The terms “image,” “image frame,” and “frame” are used interchangeably herein. A camera may be configured with various image capture and image processing settings. Different settings result in images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters may be applied to an image sensor to capture one or more image frames. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color modifications, may constitute post-processing of one or more image frames. For example, settings or parameters may be applied to a processor (e.g., an image signal processor or ISP) to process one or more image frames captured by the image sensor.

[0025] A camera may include a processor, such as an ISP, that can receive one or more image frames from an image sensor and process the one or more image frames. For example, raw image frames captured by a camera sensor may be processed by the ISP to generate a final image. In some examples, the ISP may process the image frames using multiple filters or processing blocks that are applied to the captured image frames, such as demosaicing, gain adjustment, white balance adjustment, color balance or color correction, gamma compression, tone mapping or adjustment, noise removal or filtering, edge enhancement, contrast adjustment, intensity adjustment (such as darkening or lightening), among others. In some examples, the ISP may include a machine learning system (e.g., one or more trained neural networks, one or more trained machine learning models, one or more artificial intelligence algorithms, and / or one or more other machine learning components) that can process the image frames and output processed image frames.

[0026] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") are described for performing machine learning-based automatic image adjustment using one or more machine learning systems. An imaging system can acquire image data, for example, from an image sensor. The imaging system can provide the image data as input data to the machine learning system. The machine learning system can generate one or more maps based on the image data. Each map of the one or more maps can specify a strength with which an image processing function should be applied to each pixel of the image data. Different maps may be generated for different image processing functions, such as noise reduction, noise addition, sharpening, blurring, detail enhancement, detail reduction, tone adjustment, color desaturation, color saturation enhancement, hue adjustment, or combinations thereof. The one or more maps may vary spatially such that different pixels of the image correspond to different strengths due to application of the image processing function in each map. The imaging system can generate a modified image based on the image data and the one or more maps, for example, by applying each of the one or more image processing functions according to each of the one or more maps. In some examples, the imaging system can provide the image data and the one or more maps to a second machine learning system to generate a modified image. The second machine learning system may be separate from the machine learning system.

[0027] In some camera systems, a host processor (HP), sometimes referred to as an application processor (AP), is used to dynamically configure the image sensor with new parameter settings. The HP may be used to dynamically configure the parameter settings of the ISP pipeline to match the exact settings of the input image sensor frame so that the image data is processed correctly. In some examples, the HP may be used to dynamically configure the parameter settings based on one or more maps. For example, the parameter settings may be determined based on the values ​​in the maps.

[0028] Generating maps and using them for image processing can provide various technical advantages to an image processing system. Spatially varying image processing maps can enable an imaging system to apply image processing functions with different strengths in different regions of an image. For example, a region of an image depicting sky may be processed differently with respect to one or more image processing functions than another region of the same image depicting grass. Multiple maps can be generated in parallel, increasing the efficiency of image processing.

[0029] In some examples, the machine learning system can generate one or more affine coefficients, such as multipliers and offsets, which the imaging system can use to modify components of the image data (e.g., the luminance channel of the image data) to generate each map. Local linearity constraints can ensure that one or more gradients in the map align with one or more gradients in the image data, thereby reducing halo effects when applying image processing functions. Using affine coefficients and / or local linearity constraints to generate the maps can produce higher quality spatially-varying image modifications than systems that do not use affine coefficients and / or local linearity constraints to generate the maps, e.g., due to better alignment between the image data and the maps and reduced halo effects at boundaries of depicted objects.

[0030] 1 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components used to capture and process images of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture standalone images (or photographs) and / or can capture videos that include multiple images (or video frames) in a particular order. A lens 115 of the system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0031] The one or more controls 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more controls 120 may include multiple mechanisms and components. For example, the controls 120 may include one or more exposure controls 125A, one or more focus controls 125B, and / or one or more zoom controls 125C. The one or more controls 120 may also include additional controls beyond those shown, such as analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0032] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can actuate a motor or servo to move the lens 115 closer to or farther from the image sensor 130, thereby adjusting the focus. In some cases, the system 100 may include additional lenses, such as one or more microlenses above each photodiode of the image sensor 130, each of which bends light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus setting may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may also be referred to as image capture settings and / or image processing settings.

[0033] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or f / stop), the length of time the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may also be referred to as an image capture setting and / or an image processing setting.

[0034] The zoom control 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable-focus zoom lens. In some examples, the lens assembly may include a focusing lens (which may in some cases be the lens 115) that first receives light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference) and a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses.

[0035] Image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image produced by image sensor 130. In some cases, different photodiodes may be covered by different color filters and may therefore measure light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes red, blue, and green filters, and each pixel of the image is generated based on red light data from at least one photodiode covered by a red color filter, blue light data from at least one photodiode covered by a blue color filter, and green light data from at least one photodiode covered by a green color filter. Other types of color filters may use yellow, magenta, and / or blue-green (also called "emerald") color filters instead of or in addition to red, blue, and / or green color filters. Some image sensors may lack color filters entirely and instead use different photodiodes (possibly stacked vertically) throughout the pixel array. Different photodiodes across the pixel array may have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore lack color depth.

[0036] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching some photodiodes or portions of some photodiodes at certain times and / or from certain angles, which may be used for phase-detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier to amplify the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) and convert the analog signal output of the photodiode (and / or amplified by the analog gain amplifier) ​​to a digital signal. In some cases, instead or in addition, some components or functions discussed with respect to one or more of control mechanisms 120 may be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0037] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISPs 154), one or more host processors (including host processor 152), and / or one or more of any other types of processors 5010 discussed with respect to computing device 5000. Host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, image processor 150 is a single integrated circuit or chip (called a system-on-chip or SoC) that includes host processor 152 and ISPs 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) ports 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth™, Global Positioning System (GPS), etc.), any combination of these, and / or other components. I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a serial peripheral interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (e.g., MIPI CSI-2), a physical (PHY) layer port or interface, an Advanced High-performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, host processor 152 can communicate with image sensor 130 using an I2C port, and ISP 154 can communicate with image sensor 130 using a MIPI port.

[0038] Image processor 150 may perform several tasks such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 / 5020, read only memory (ROM) 145 / 5025, a cache, a memory unit, another storage device, or some combination thereof.

[0039] Various input / output (I / O) devices 160 may be connected to image processor 150. I / O devices 160 may include a display screen, a keyboard, a keypad, a touchscreen, a trackpad, a touch-sensitive surface, a printer, any other output device 5035, any other input device 5045, or some combination thereof. In some cases, captions may be entered into image processing device 105B through a physical keyboard or keypad of I / O device 160 or through a virtual keyboard or keypad of a touchscreen of I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or send data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable a wireless connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. Peripheral devices may include any of the types of I / O devices 160 previously discussed, and may themselves be considered I / O devices 160 when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0040] In some cases, image capture and processing system 100 may be a single device. In some cases, image capture and processing system 100 may be two or more separate devices including image capture device 105A (e.g., a camera) and image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, image capture device 105A and image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electronic connectors and / or wirelessly via one or more wireless transceivers. In some implementations, image capture device 105A and image processing device 105B may be separate from one another.

[0041] 1 into two portions representing image capture device 105A and image processing device 105B, respectively. Image capture device 105A includes lens 115, control mechanism 120, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, some components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included in image capture device 105A.

[0042] Image capture and processing system 100 may include an electronic device such as a mobile or fixed telephone handset (e.g., a smartphone, a mobile phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, image capture device 105A and image processing device 105B may be different devices. For example, image capture device 105A can include a camera device, and image processing device 105B can include a computing device, such as a mobile handset, a desktop computer, or other computing device.

[0043] While image capture and processing system 100 is shown as including several components, those skilled in the art will understand that image capture and processing system 100 may include many more components than those shown in FIG. 1 . The components of image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of image capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing image capture and processing system 100.

[0044] Conventional camera systems (e.g., image sensors and ISPs) are adjusted using parameters and process images according to the adjusted parameters. The ISPs are typically adjusted during manufacturing using fixed adjustment methods. Camera systems (e.g., image sensors and ISPs) also typically perform global image adjustments based on predetermined conditions such as light level, color temperature, exposure time, among others. Conventional camera systems are also adjusted using coarse-grained heuristic-based adjustments (e.g., window-based local tone mapping). As a result, conventional camera systems are not capable of enhancing images based on the content contained in the image.

[0045] Systems, devices, processes, and computer-readable media for performing machine learning-based image adjustment are described herein. Machine learning-based image adjustment techniques can be applied to processed images (e.g., output by an ISP and / or an image post-processing system) and / or to raw image data from an image sensor. Machine learning-based image adjustment techniques can provide dynamic adjustments for each image (rather than the fixed adjustments of traditional camera systems) based on the scene content contained in the image. Machine learning-based image adjustment techniques can also provide the ability to incorporate additional semantic context-based adjustments (e.g., segmentation information), as opposed to providing only heuristic-based adjustments.

[0046] In some examples, one or more tuning maps can also be used to perform machine learning-based image adjustment techniques. In some examples, the tuning map can have the same resolution as the input image, with each value in the tuning map corresponding to a pixel in the input image. In some examples, the tuning map can be based on a downsampled version of the input image and therefore have a lower resolution than the input image, with each value in the tuning map corresponding to more than one pixel in the input image (e.g., if the downsampled version of the input image has half the length and width of the input image, there are four or more adjacent pixels in a square or rectangular structure). A tuning map can also be referred to as a spatial tuning map or a spatially-varying tuning map, which refers to the fact that the values ​​in the tuning map can vary spatially with respect to each location in the tuning map. In some examples, each value in the tuning map can correspond to a predetermined subset of the input image. In some examples, each value in the tuning map can correspond to the entire input image, in which case the tuning map may be referred to as a spatially fixed tuning map.

[0047] The tuning map provides the image processing machine learning system with the ability to make pixel-level adjustments for each image, enabling high-precision image adjustments rather than just global image adjustments. In some examples, the tuning map may be automatically generated by image processor 150, host processor 152, ISP 154, or a combination thereof. In some examples, when using a machine learning system, the tuning map may be generated automatically. In some examples, when using a machine learning system, the tuning map may have pixel-level precision. For example, a first machine learning system can be used to generate pixel-level tuning maps, and a second machine learning system can process the input image and one or more tuning maps to generate a corrected image (also referred to as an adjusted image or a processed image). In some examples, the first machine learning system may be at least partially executed by and / or interface with image processor 150, host processor 152, ISP 154, or a combination thereof. In some examples, the second machine learning system may be at least partially executed by and / or interface with image processor 150, host processor 152, ISP 154, or a combination thereof.

[0048] 2 is a block diagram illustrating an example of a system including an image processing machine learning (ML) system 210. The image processing ML system 210 receives as input one or more input images and one or more spatial tuning maps 201. An exemplary input image 202 is shown in FIG. 2. The image processing ML system 210 can process any type of image data. For example, in some examples, the input images provided to the image processing ML system 210 may include images captured by the image sensor 130 of the image capture device 105A of FIG. 1 and / or processed by an ISP (e.g., ISP 154 of FIG. 1) and / or any post-processing components of the camera system. The image processing ML system 210 may be implemented using one or more convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or a combination thereof. The image processing ML system 210 and / or any of the above-listed machine learning elements may be trained using supervised learning (which may be part of the image processing ML system 210), unsupervised learning, reinforcement learning, deep learning, or a combination thereof.

[0049] In some examples, the input image 202 may include multiple saturation and / or color components for each pixel of the image data (e.g., a red (R) color component or sample, a green (G) color component or sample, and a blue (B) color component or sample). In some cases, the device may include multiple cameras, and the image processing ML system 210 may process images acquired by one or more of the multiple cameras. In one illustrative example, a dual-camera mobile phone, tablet, or other device may be used to capture larger images with a wider angle (e.g., a larger field of view (FOV)), capture more light (resulting in greater sharpness and clarity, among other benefits), generate 360-degree (e.g., virtual reality) video, and / or perform other functions that are improved over those achieved by a single-camera device.

[0050] In some examples, the image processing ML system 210 can process raw image data provided by one or more image sensors. For example, an input image provided to the image processing ML system 210 may include raw image data generated by an image sensor (e.g., image sensor 130 of FIG. 1). In some cases, the image sensor may include an array of photodiodes that can capture frames of raw image data. Each photodiode can represent a pixel location and generate a pixel value for that pixel location. The raw image data from the photodiodes may include a single brightness or grayscale value for each pixel location in the frame. For example, a color filter array may be integrated with the image sensor or may be used in conjunction with the image sensor (placed over the photodiodes) to convert monochrome information to brightness. One illustrative example of a color filter array is a Bayer pattern color filter array (or Bayer color filter array), which enables the image sensor to capture frames of pixels having a Bayer pattern with one of a red, green, or blue filter at each pixel location. In some cases, a device may include multiple image sensors, in which case image processing ML system 210 may process raw image data acquired by the multiple image sensors. For example, a device with multiple cameras may capture image data using multiple cameras, and image processing ML system 210 may process raw image data from the multiple cameras.

[0051] Various types of spatial tuning maps may be provided as input to the image processing ML system 210. Each type of spatial tuning map may be associated with a respective image processing function. The spatial tuning maps may include, for example, a noise tuning map, a sharpness tuning map, a tone tuning map, a saturation tuning map, a hue tuning map, any combination thereof, and / or other tuning maps for other image processing functions. For example, a noise tuning map may include values ​​corresponding to the strength with which noise augmentation (e.g., noise reduction, noise addition) should be applied to different portions (e.g., different pixels) of the input image that are mapped to locations (e.g., coordinates) in the different portions of the input image. A sharpness tuning map may include values ​​corresponding to the strength with which sharpness augmentation (e.g., sharpening) should be applied to different portions (e.g., different pixels) of the input image that are mapped to locations (e.g., coordinates) in the different portions of the input image. A tone tuning map may include values ​​corresponding to the strength with which tone augmentation (e.g., tone mapping) should be applied to different portions (e.g., different pixels) of the input image that are mapped to locations (e.g., coordinates) in the different portions of the input image. The saturation tuning map may include values ​​corresponding to the strength with which saturation augmentation (e.g., increasing saturation or decreasing saturation) should be applied to different portions (e.g., different pixels) of the input image that map to locations (e.g., coordinates) in the different portions of the input image. The hue tuning map may include values ​​corresponding to the strength with which hue augmentation (e.g., shifting hue) should be applied to different portions (e.g., different pixels) of the input image that map to locations (e.g., coordinates) in the different portions of the input image.

[0052] In some examples, the values ​​in the tuning map may range from a value of 0 to a value of 1. The range may be inclusive, and thus include 0 and / or 1 as possible values. The range may not be inclusive, and thus do not include 0 and / or 1 as possible values. For example, each location in the tuning map may include every value between 0 and 1. In some examples, one or more of the tuning maps may include values ​​other than those between 0 and 1. Various tuning maps are described in more detail below. An input saturation map 203 is shown in FIG. 2 as an example of a spatial tuning map.

[0053] Although all of the figures herein are shown in black and white, input image 202 represents a colored image depicting a pink-red flower in the foreground, green leaves in one portion of the background, and a white wall in another portion of the background. Input saturation map 203 includes a value of 0.5 corresponding to the area of ​​pixels in input image 202 depicting the pink-red flower in the foreground. This area with a value of 0.5 is shown as gray in input saturation map 203. A value of 0.5 in input saturation map 203 indicates that the saturation of that area should remain the same, neither increased nor decreased. Input saturation map 203 includes a value of 0.0 corresponding to the area of ​​pixels in input image 202 depicting the background (green leaves and white wall). This area with a value of 0.0 is shown as black in input saturation map 203. A value of 0.0 in input saturation map 203 indicates that the area should be completely desaturated. The corrected image 216 represents the image in which the foreground flowers are still as saturated as in the input image 202 (still pink and red), but the background (both the green leaves and the white wall) is completely desaturated and therefore depicted in grayscale.

[0054] Each tuning map (e.g., input saturation map 203) may have the same size and / or resolution as the size and / or resolution of one or more input images (e.g., input image 202) provided as input to image processing ML system 210 to be processed to generate a modified image (e.g., modified image 216). Each tuning map (e.g., input saturation map 203) may have the same size and / or resolution as the size and / or resolution of a modified image (e.g., modified image 216) generated based on the tuning map. In some examples, a tuning map (e.g., input tuning map 203) may be smaller or larger than the input image (e.g., input image 202), in which case either the tuning map or the input image (or both) may be downscaled, upscaled, and / or upsampled before modified image 216 is generated.

[0055] In some examples, tuning map 303 may be stored as an image, such as a grayscale image. In tuning map 303, a shade of gray at a pixel of tuning map 303 may indicate a value for that pixel. In an illustrative example, black may represent a value of 0.0, medium gray or neutral gray may represent a value of 0.5, and white may represent a value of 1.0. Other shades of gray may represent values ​​between those described above. For example, light gray can represent values ​​between 0.5 and 1.0 (e.g., 0.6, 0.7, 0.8, 0.9). Dark gray can represent values ​​between 0.5 and 0.0 (e.g., 0.1, 0.2, 0.3, 0.4). In some examples, the opposite scheme may be used, where white represents a value of 0.0 and black represents a value of 1.0.

[0056] 3A is a conceptual diagram illustrating an example input image 302 including a number of pixels labeled P0 through P63, according to some examples. The input image is 7 pixels wide and 7 pixels high. The pixels are numbered sequentially from P0 through P63, counting from left to right within each row, starting from the top row and counting up to the bottom row.

[0057] 3B is a conceptual diagram illustrating an example of a spatial tuning map 303, according to some examples. The tuning map 303 includes a plurality of values ​​labeled V0 through V63. The tuning map is 7 pixels wide and 7 pixels high. The pixels are numbered sequentially from V0 to V63, counting from left to right within each row, starting from the top row and counting up to the bottom row.

[0058] Each value in tuning map 303 corresponds to a pixel in input image 302. For example, value V0 in tuning map 303 corresponds to pixel P0 in input image 302. The values ​​in tuning map 303 are used to adjust or modify the corresponding pixel in input image 302. In some examples, each value in adjustment map 303 indicates the strength and / or direction to use when applying an image processing function to the corresponding pixel of image data. In some examples, each value in tuning map 303 indicates the amount of image processing function to apply to the corresponding pixel. For example, a first value for V0 in tuning map 303 (e.g., a value of 0, a value of 1, or another value) may indicate that the image processing function of the tuning map should be applied with zero strength (not applied at all) to corresponding pixel P0 in input image 302. In another example, a second value for V15 in tuning map 303 (e.g., a value of 0, a value of 1, or another value) indicates that the image processing function should be applied with maximum strength (maximum amount of image processing function) to corresponding pixel P15 in input image 302. Values ​​in different types of tuning maps can indicate different levels of applicability of each image processing function. In one illustrative example, tuning map 303 can be a saturation map, where a value of 0 in the saturation map can indicate that the corresponding pixel is fully desaturated, resulting in a grayscale or monochrome pixel (maximum intensity in the desaturation or negative saturation direction), a value of 0.5 in the saturation map can indicate that no saturation or desaturation effect is applied to the corresponding pixel (saturation intensity of 0), and a value of 1 in the saturation map can indicate maximum saturation (maximum intensity in the saturation or positive saturation direction). Values ​​between 0 and 1 in the saturation map indicate various varying levels of saturation between the levels described above. For example, a value of 0.2 represents slight desaturation (low intensity in the desaturation or negative saturation direction), and a value of 0.8 represents slight saturation enhancement (low intensity in the saturation or positive saturation direction).

[0059] The image processing ML system 210 processes one or more input images and one or more spatial tuning maps 201 to generate a modified image 215. An exemplary modified image 216 is shown in FIG. 2. The modified image 216 is a modified version of the input image 202. As discussed above, the saturation map 203 includes a value of 0.5 for areas in the saturation map 203 that correspond to pixels depicting the pink-red flower shown in the foreground of the input image 202. The saturation map 203 includes a value of 0 for locations that correspond to pixels in the image 202 that do not depict the foreground flower of the image 202 (and instead depict the background). Values ​​of 0 in the saturation map 203 indicate that the pixels corresponding to those values ​​are fully desaturated (maximum strength in the desaturation or negative saturation direction). As a result, all pixels in the modified image 215, except for pixels representing flowers, are depicted in grayscale (black and white). A value of 0.5 in the saturation map 203 indicates no change in saturation (a saturation strength of 0) for the pixels corresponding to those values.

[0060] The image processing ML system 210 may include one or more neural networks that are trained to process one or more input images and one or more spatial tuning maps 201 to generate modified images 215. In one illustrative example, supervised learning techniques may be used to train the image processing ML system 210. For example, a backpropagation training process may be used to adjust the node weights (and in some cases, other parameters such as biases) of the neural network of the image processing ML system 210. Backpropagation includes a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. This process is repeated a certain number of times for each set of training data until the parameter weights of the neural network are precisely adjusted. Further details regarding the image processing ML system 210 are provided herein.

[0061] As described above, various types of spatial tuning maps may be provided for use by image processing ML system 210. One example of a spatial tuning map is a noise map that indicates the amount of denoising (noise reduction) to apply to pixels of an input image.

[0062] 4 is a conceptual diagram illustrating an example of an input noise map 404 that is applied to adjust the strength of noise reduction in an input image 402 to produce a modified image. The portion of the input image 402 in the lower left corner, labeled in parentheses as "smooth," has little to no visible noise. The remainder of the input image 402, labeled in parentheses as "noisy," contains visible noise (along with a grainy appearance).

[0063] The noise map key 405 is shown in FIG. 4. The noise map key 405 indicates that a value of 0 in the noise map 404 indicates that no noise reduction is applied, and a value of 1 indicates that maximum noise reduction is applied. As shown, the input noise map 404 contains values ​​of 0 for black areas of the noise map 404 that correspond to "smooth" areas of pixels in the lower left region of the input image 402 that have little or no noise (pixels in the lower left region of the image 402). The input noise map 404 contains values ​​of 0.8 for the remaining light gray shade areas in the noise map 404 that correspond to "noisy" areas of pixels in the input image 402 that have noise. The image processing ML system 210 can process the input image 402 and the input noise map 404 to produce a modified image 415. As shown, the corrected image 415 is noise-free based on the noise removal performed by the image processing ML system 210 for pixels in the input image 402 that correspond to locations in the noise map 404 having a value of 0.8.

[0064] Another example of a spatial tuning map is a sharpening map that indicates the amount of sharpening to be applied to pixels of an input image, for example, a blurry image can be brought into focus by sharpening the pixels of the image.

[0065] 5 is a conceptual diagram illustrating an example of the effect of applying a sharpness adjustment image processing function at different intensities to an example image, according to some examples. In some examples, a value of 0 in the sharpness map indicates that no sharpening is applied to the corresponding pixel in the input image to generate the corresponding output image, and a value of 1 in the sharpness map indicates that maximum sharpening is applied to the corresponding pixel in the input image to generate the corresponding output image. When the value in the sharpness adjustment map is set to 0, image 510 is generated by image processing ML system 210, with no sharpening applied to image 510 (or sharpening with a strength of 0). When the value in the sharpness map is set to 0.4, image 511 is generated by image processing ML system 210, with a medium amount of sharpening (medium strength sharpening) applied to the pixels of image 511. When the value in the sharpness map is set to 1.0, image 512 is generated, with the maximum amount of sharpening (maximum strength sharpening) applied to the pixels of image 512.

[0066] An enhanced image (or sharpened image) may be generated by the image processing ML system 210 based on the input image and the sharpness map. For example, the image processing ML system 210 can use the sharpness map to sharpen the input image, resulting in a sharpened image (also referred to as an enhanced image). The sharpened (or enhanced) image depends on the image detail as well as the alpha (α) parameter. For example, the following formula may be used to generate the sharpened (or enhanced) image: enhanced image = original input image + α × detail. The values ​​in the sharpness map may cause the image processing ML system 210 to modify both the alpha (α) parameter and the image detail parameter to obtain the sharpened image. As an illustrative example of generating a sharpened (or enhanced) image, the image processing ML system 210 can apply edge-preserving filtering to the input image, such as by applying an edge-preserving nonlinear filter. An example of an edge-preserving nonlinear filter is a bilateral filter. In some cases, the bilateral filter has two hyperparameters, including spatial sigma and range sigma. Spatial sigma controls the filter window size (larger window sizes result in greater smoothing). Range sigma controls the filter size along the intensity dimension (larger values ​​blur different intensities together). In some examples, image details may be obtained by subtracting the filtered image (the smoothed image resulting from edge-preserving filtering) from the input image. A sharpened (or enhanced) image may then be obtained by adding a portion of the image details (represented by the alpha (α) parameter) to the original input image. Using the image details and alpha (α), the sharpened (enhanced) image may be obtained as follows: Enhanced Image = Original Input Image + α × Detail

[0067] FIG. 6 is a conceptual diagram illustrating an example of an input sharpness map 604 applied to adjust the strength of sharpening in an input image 602 to generate a modified image 615, according to some examples. Clearly, the input image 602 has a blurred appearance. The sharpness map key 605 specifies that a value of 0 in the sharpness map 604 indicates that no sharpening is applied to the corresponding pixel of the input image 602 (sharpening is applied at a strength of 0), and a value of 1 indicates that maximum sharpening is applied to the corresponding pixel (sharpening is applied at a strength of 1). As shown, the input sharpness map 604 includes a value of 0 for the black-shaded area in the lower left portion of the sharpness map 604. The input sharpness map 604 includes a value of 0.9 for the remaining area of ​​the sharpness map 604 (light gray shades). The image processing ML system 210 can process the input image 602 and the input sharpness map 604 to produce the modified image 615. As shown, pixels in the modified image 615 corresponding to locations in the sharpness map 604 with a value of 0 are blurred (because sharpening was applied at a strength of 0), while pixels in the modified image 615 corresponding to locations in the sharpness map 604 with a value of 0.9 are sharp (not blurred) due to the sharpening performed at a high strength by the image processing ML system 210 in that area.

[0068] A tone map is another example of a spatial tuning map. A tone map indicates the amount of brightness adjustment to be applied to pixels in an input image, resulting in a darker or lighter tone image. A gamma value controls the amount of brightness to a pixel, which is called gamma correction. For example, gamma correction can be done by:

number

[0069] 7 is a conceptual diagram illustrating an example gamma curve 715 and various examples of the effect that different values ​​in a tone map translate into different strengths of tone adjustment applied to an example image, according to some examples. In some examples, the gamma (γ) ranges from 10 -0.4 From 10 0.4 , which corresponds to values ​​from 0.4 to 2.5. Gamma exponents from -0.4 to 0.4 may be mapped to the range 0 to 1 for calculating the tone tuning map. For example, a value of 0 in the tone map corresponds to a gamma (γ) value of 0.4 (10 -0.4 ) is applied to the corresponding pixel in the input image, and a value of 0.5 in the tone map corresponds to a gamma (γ) value of 1.0 (10 0.0 ) is applied to the corresponding pixel in the input image, and a value of 1 in the tone map indicates a gamma (γ) value of 2.5 (corresponding to 10 0.47 shows that a gamma adjustment (corresponding to a gamma of 0.4) is applied to corresponding pixels in the input image. Output image 710, output image 711, and output image 712 are all generated by applying different strengths of tone adjustment to the same input image. Image 710 is generated by image processing ML system 210 when a gamma value of 0.4 is applied, resulting in an increased brightness and therefore a brighter image than the input image. Image 711 is generated when a gamma value of 1.0 is applied, resulting in no change in brightness. Thus, output image 711 is identical in tone to the input image. Image 712 is generated when a gamma value of 2.5 is applied, resulting in a decreased brightness and therefore a darker image than the input image.

[0070] 8 is a conceptual diagram illustrating an example of an input tone map 804 being applied to the strength of tone adjustments in an input image 802 to generate a modified image 815, according to some examples. The tone map key 805 specifies that a value of 0 in the tone map 804 indicates that a gamma (γ) value of 0.4 is applied to the corresponding pixel in the input image 802 (a tone adjustment of strength 0.4), a value of 0.5 in the tone map 804 indicates that a gamma (γ) value of 1.0 is applied to the corresponding pixel in the input image 802 (a tone adjustment of strength 1.0 indicating no change in tone), and a value of 1 in the tone map 804 indicates that a gamma (γ) value of 2.5 is applied to the corresponding pixel in the input image 802 (a tone adjustment of maximum strength 2.5). The input tone map 804 includes a value of 0.12 for a group of locations in the left half of the tone map 804 (areas with darker gray shades). The value of 0.12 corresponds to an approximate gamma (γ) value of 0.497. For example, the exponent can be determined as Exponent=((0.12-0.5)*0.4) / 0.5=-0.304, and the gamma (γ) value can be determined as Gamma=10 Exponent =10 (-0.304)=0.497. Tone map 804 contains a value of 0.5 for the remaining group of positions (areas with medium gray shades) in the right half of tone map 804. A value of 0.5 corresponds to a gamma (γ) value of 1.0.

[0071] The image processing ML system 210 can process the input image 802 and the input tone map 804 to generate a modified image 815. As shown, pixels in the left half of the modified image 815 have a lighter tone (compared to corresponding pixels in the left half of the input image 802 and / or compared to pixels in the right half of the modified image 815) due to the use of a gamma (γ) value of 0.497 in the gamma conversion based on a tone map value of 0.12. Pixels in the right half of the modified image 815 have the same tone as corresponding pixels in the right half of the input image 802 (and / or a darker tone compared to pixels in the left half of the modified image 815) due to the use of a gamma (γ) value of 1.0 in the gamma conversion based on a tone map value of 0.5.

[0072] A saturation map is another example of a spatial tuning map. A saturation map indicates the amount of saturation to apply to pixels of an input image, resulting in an image with modified color values. In a saturation map, saturation refers to the difference between RGB pixels (or pixels defined using another color space, such as YUV) from a corresponding grayscale image (e.g., as shown in Equations 1 through 3 below). This saturation may sometimes be referred to as the "saturation difference." For example, increasing the saturation of an image may make the colors in the image more intense, while decreasing the saturation may tone down the colors. If the saturation is reduced sufficiently, the image may become desaturated, resulting in a grayscale image. As described below, an Alpha Blending technique may sometimes be performed, where an alpha (α) value determines the saturation of the colors.

[0073] Saturation adjustment may be applied using different techniques: In the YUV color space, the values ​​of the U color difference component (blue-prone) and the Y color difference component (red-prone) of the image's pixels may be adjusted to set different saturation levels for the image.

[0074] Figure 9 is a conceptual diagram 915 illustrating saturation levels in YUV color space, where the x-axis represents the U color difference component (emphasis toward blue) and the y-axis represents the V color difference component (emphasis toward red). Luminance (Y) is constant at a value of 128 in the schematic diagram of Figure 9. Desaturation is represented at the center of the diagram, where the U value is 128 and the V value is 128, and the image is a grayscale (desaturated) image. The greater the distance (higher or lower U and / or V values), the greater the saturation. For example, a greater distance from the center corresponds to a higher saturation.

[0075] In another example, an Alpha Blending technique may be used to adjust the saturation. The mapping between luminance (Y) and the R, G, B channels may be expressed using the following equation: R'=α×R+(1-α)×Y Equation (1) G'=α×G+(1-α)×Y Equation (2) B'=α×B+(1-α)×Y Equation (3)

[0076] FIG. 10 is a conceptual diagram showing processed transformations of an exemplary image, each processed using different alpha (α) values ​​to adjust color saturation, according to some examples. At values ​​of α less than 1 (α<1), the image becomes less saturated. For example, when α=0, all of the R′, G′, and B′ values ​​are equal to the luminance (Y) value (α=0: R′=G′=B′=Y), resulting in grayscale (desaturation). Output images 1020, 1021, and 1022 are all generated by applying different strengths of saturation adjustment to the same input image. The input image shows three female faces next to each other. Because FIG. 10 is shown in grayscale rather than color, the red channels of output images 1020, 1021, and 1022 are shown in FIG. 10 to illustrate the changes in red saturation. Increased red saturation is visible as brighter areas, while decreased red saturation is visible as darker areas. When α = 1, there is no effect on saturation. For example, when α = 1, equations (1)-(3) result in R' = R, G', G, and B' = B. Output image 1021 in FIG. 10 is an example of an image without the effect of saturation based on an alpha (α) value set to a value of 1.0. Thus, output image 1021 is identical to the input image in terms of color saturation. Output image 1020 in FIG. 10 is an example of a grayscale (de-saturated) image resulting from an alpha (α) value of 0.0. Because FIG. 10 shows the red channel of output image 1020, a darker face in output image 1020 (compared to output image 1021 and therefore the input image) indicates less saturation of the red color in the face. For values ​​of α greater than 1 (α > 1), the image becomes more saturated. The R, G, and B values ​​may be limited to a maximum intensity. Output image 1022 in Figure 10 is an example of a highly saturated image resulting from an alpha (α) value of 2.0. Because Figure 10 shows the red channel of output image 1020, the brighter face in output image 1022 (compared to output image 1021 and therefore the input image) indicates a higher saturation of the red color in the face.

[0077] 11 is a conceptual diagram illustrating an example of an input image 1102 and an input saturation map 1104. A saturation map key 1105 specifies that a value of 0 in the saturation map 1104 indicates desaturation, a value of 0.5 indicates no saturation effect, and a value of 1 indicates saturation. The input saturation map 1104 includes a value of 0.5 for a group of positions in the lower left portion of the saturation map 1104, indicating that no saturation effect will be applied to the corresponding pixels in the input image 1102. The saturation map 1104 includes a value of 0.1 for a group of positions outside the lower left portion of the saturation map 1104. A value of 0.1 results in a desaturated (mostly desaturated) pixel.

[0078] The image processing ML system 210 can process the input image 1102 and the input saturation map 1104 to generate a modified image 1115. As shown, the pixels in the bottom left portion of the modified image 1115 are the same as the corresponding pixels in the input image 1102 due to no saturation effect being applied to those pixels based on the saturation map value of 0.5. The remaining pixels in the modified image 1115 have a darker green appearance (approximately grayscale) due to the desaturation being applied to those pixels based on the saturation map value of 0.1.

[0079] Another example of a spatial adjustment map is a hue map. Hue is the color in an image, and saturation (in HSV color space) is the intensity (or depth) of that color. A hue map indicates the amount of color change to apply to pixels in an input image. Hue adjustments may be applied using different techniques. For example, the Hue, Saturation, Value (HSV) color space is a representation of the RGB color space. The HSV representation uses the saturation and hue dimensions to model how different colors blend together.

[0080] Figure 12A is a conceptual diagram 1215 illustrating the HSV color space, with saturation represented on the y-axis and hue represented on the x-axis, with a value of 255. Hue (color) may be modified as shown on the x-axis. As shown, hue wraps around, in this case, hues of 0 and 180 represent the same red color (according to the OpenCV convention, Hue:0 = Hue:180). Because the diagram is shown in black and white, colors are written in letters where they appear in the HSV color space.

[0081] Note that saturation in HSV space is not equivalent to the saturation defined by the saturation map discussed above. For example, when referring to saturation in HSV color space, saturation refers to the standard HSV color space definition, where saturation is a particular color intensity. However, depending on the HSV value, a saturation value of 0 can make a pixel appear black (value = 0), gray (value = 128), or white (value = 255). This saturation may sometimes be referred to as "HSV saturation." Because color saturation and lightness are linked in HSV space, this definition of saturation (HSV saturation) is not used herein. Instead, as used herein, saturation refers to how colorful or gray a pixel appears, which can be captured by the alpha (α) value described above.

[0082] FIG. 12B is a conceptual diagram 1216 illustrating the YUV color space. An original UV vector 1217 is shown relative to a center point 1219 (U=128, V=128). The original UV vector 1217 can be modified by adjusting the angle (θ), resulting in a modified hue (color) shown by a hue-modified vector 1218. The hue-modified vector 1218 is shown relative to the center point 1219. In the YUV color space, the orientation of the UV vector may be modified with respect to the center point 1219 (U=128, V=128, as shown in FIG. 9). The modification of the UV vector shows an example of how hue adjustment may work, where the strength of the hue adjustment corresponds to the angle (θ) and the direction of the hue adjustment corresponds to whether the angle (θ) is negative or positive.

[0083] FIG. 13 is a conceptual diagram illustrating an example of an input hue map 1304 applied to the strength of a hue adjustment in an input image 1302 to generate a modified image 1315, according to some examples. The hue map key 1305 specifies that a value of 0.5 in the hue map 1304 indicates no effect (no change in hue or color) (a hue adjustment of 0 strength). Any value other than 0.5 indicates a change in hue relative to the current color (a hue adjustment of non-zero strength). For example, referring to FIG. 12A as an illustrative example, if the current color is green (i.e., hue = 60), a hue value of 60 may be added to change the color to blue (i.e., hue = 60 + 60 = 120). In another example, to change green to red, a hue value of 60 may be subtracted from the current hue value of 60 (i.e., hue = 60 - 60 = 0). The input hue map 1304 includes a value of 0.5 for a group of locations in the upper left portion of the hue map 1304 (the area of ​​medium shades of gray), indicating that no hue change is applied to the corresponding pixel in the input image 1302 (a hue adjustment of 0 strength). The hue map 1304 includes a value of 0.4 for a group of locations outside the upper left portion of the hue map 1304 (the area of ​​dark shades of gray). The value of 0.4 changes the pixel in the input image 1302 from its current hue to a modified hue. Because the value of 0.4 is closer to 0.5 than to 0.0, the change in hue compared to the input image 1302 is small. The change in hue indicated by the input hue map 1304 is the relative change to the hue of each pixel in the input image 1302. The mapping for each pixel may be determined using the color space shown in FIG. 12A. Like the other adjustment maps described herein, hue map values ​​can range from 0 to 1, which are remapped to an actual range of values ​​from 0 to 180 during application of the hue map.

[0084] Input image 1302 shows a portion of a yellow wall with the word "Griffith-" written in purple letters. Image processing ML system 210 can process input image 1302 and input hue map 1304 to generate modified image 1315. As shown, the pixels in the upper left portion of modified image 1315 are the same hue as the corresponding pixels in input image 1302 due to no hue change applied to those pixels (a hue adjustment of 0 strength) based on a hue map value of 0.5 in that area. Thus, the pixels in the upper left portion of modified image 1315 retain the yellow hue of the yellow wall in input image 1302. The remaining pixels of modified image 1315 have an altered hue compared to the pixels in input image 1302 due to the hue change applied to those pixels based on a hue map value of 0.4. Specifically, in the remainder of the modified image 1315, areas of the wall that appear yellow in the input image 1302 appear orange in the modified image 1315. In the remainder of the modified image 1315, the word "Griffith-" that appears purple in the input image 1302 appears a different shade of purple in the modified image 1315.

[0085] In some examples, multiple tuning maps may be input to image processing ML system 210 along with input image 1402. In some examples, the multiple tuning maps may include different tuning maps corresponding to different image processing functions, such as noise removal, noise addition, sharpening, blurring (e.g., blurring), tone adjustment, saturation adjustment, detail adjustment, hue adjustment, or a combination thereof. In some examples, the multiple tuning maps may include different tuning maps corresponding to the same image processing function, but may modify different areas / locations within input image 1402 in different ways, for example.

[0086] 14A is a conceptual diagram illustrating an example of a system including an image processing system 1406 that receives an input image 1402 and multiple spatially-varying tuning maps 1404. In some examples, the image processing system 1406 may be the image processing ML system 210 and / or may include a machine learning (ML) system, such as the ML system of the image processing ML system 210. The ML system may apply image processing functions using ML. In some examples, the image processing system 1406 may apply image processing functions without an ML system. The image processing system 1406 may be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof.

[0087] The spatially-varying tuning maps 1404 may include one or more of a noise map, a sharpness map, a tuning map, a saturation map, a hue map, and / or other maps associated with image processing functions. Using the input image 1402 and the plurality of tuning maps 1404 as inputs, the image processing system 1406 generates a modified image 1415. The image processing system 1406 may modify pixels of the input image 1402 based on values ​​included in the spatially-varying tuning maps 1404, resulting in the modified image 1415. For example, the image processing system 1406 may modify pixels of the input image 1402 by applying an image processing function to each pixel of the input image 1402 with a strength indicated by a value included in one of the spatially-varying tuning maps 1404 corresponding to that image processing function, and may perform such modification for each image processing function and corresponding one of the spatially-varying tuning maps 1404, resulting in the modified image 1415. The spatially varying tuning map 1404 may be generated using an ML system as further discussed with respect to at least Figures 14B, 14C, 14D, 16B, 17, 18, 19, 20, 21A, 21B, 21C, and 21D.

[0088] 14B is a conceptual diagram illustrating an example system including an image processing system 1406 that receives an input image 1402 and multiple spatially-varying tuning maps 1404, and an auto-adjusting machine learning (ML) system 1405 that receives the input image 1402 and generates the multiple spatially-varying tuning maps 1404, according to some examples. The auto-adjusting ML system 1405 can process the input image 1402 to generate the spatially-varying tuning maps 1404. The tuning maps 1404 and the input image 1402 can be provided as inputs to the image processing system 1406. The image processing system 1406 can process the input image 1402 based on the tuning maps 1404 to generate a modified image 1415 similar to that described above with respect to FIG. 14A. The auto-tuning ML system 1405 may be implemented using one or more convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or a combination thereof. The auto-tuning ML system 1405 and / or any of the above-listed machine learning elements (which may be part of the auto-tuning ML system 1405) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or a combination thereof.

[0089] 14C is a conceptual diagram illustrating an example of an auto-tuning machine learning (ML) system 1405 that receives an input image 1402 and generates multiple spatially-varying tuning maps 1404, according to some examples. In the example shown in FIG. 14C, the spatially-varying tuning maps 1404 include a noise map, a tone map, a saturation map, a hue map, and a sharpness map. Thus, although not shown in FIG. 14C, the image processing system 1406 may apply noise reduction with an intensity indicated in the noise map, apply a tone adjustment with an intensity and / or orientation indicated in the tone map, apply a saturation adjustment with an intensity and / or orientation indicated in the saturation map, apply a hue adjustment with an intensity and / or orientation indicated in the hue map, and / or apply sharpening with an intensity indicated in the sharpness map.

[0090] 14D is a conceptual diagram illustrating an example system including an image processing system 1406 that receives an input image 1402 and multiple spatially-varying tuning maps 1404, and an auto-tuning machine learning (ML) system 1405 that receives a downscaled version of the input image 1422 and generates smaller spatially-varying tuning maps 1424 that are upscaled to the multiple spatially-varying tuning maps 1404, according to some examples. The system of FIG. 14D is similar to the system of FIG. 14B but includes a downsampler 1418 and an upsampler 1426. The downsampler 1418 downsamples, downscales, or reduces the input image 1402 to generate a downsampled input image 1422. Rather than receiving the input image 1402 as input as in FIG. 14B, the auto-tuning machine learning (ML) system 1405 receives the downsampled input image 1422 as input in FIG. 14D. The self-tuning machine learning (ML) system 1405 generates small spatially-varying tuning maps 1424, which may each share the same size and / or dimensions as the downsampled input image 1422. The upsampler 1426 may upsample, upscale, and / or upscale the small spatially-varying tuning maps 1424 to generate the spatially-varying tuning maps 1404. In some examples, the upsampler 1426 may perform bilinear upsampling. The spatially-varying tuning maps 1404 may then be received as input by the image processing system 1406, along with the input image 1402.

[0091] Because the downsampled input image 1422 has a smaller size and / or resolution compared to the input image 1402, it may be faster and more efficient (in terms of computational resources, bandwidth, and / or battery consumption) for the self-tuning machine learning (ML) system 1405 to generate a small spatially-varying tuning map 1424 from the downsampled input image 1422 with downsampling performed by downsampler 1418 and / or upsampling performed by upsampler 1426 as in FIG. 14D than to generate the spatially-varying tuning map 1404 directly from the input image 1402 as in FIG. 14B.

[0092] As described above, the image processing system 1406, the image processing ML system 210, and / or the self-tuning machine learning (ML) system 1405 may include one or more neural networks that may be trained using supervised learning techniques.

[0093] 15 is a block diagram 1600A illustrating an example of a neural network 1500 that may be used by the image processing system 1406 and / or the self-tuning machine learning (ML) system 1405, according to some examples. The neural network 1500 may include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief net (DBN), a recurrent neural network (RNN), a generative adversarial network (GAN), and / or other types of neural networks.

[0094] The input layer 1510 of the neural network 1500 includes input data. The input data of the input layer 1510 may include data representing pixels of an input image frame. In an illustrative example, the input data of the input layer 1510 may include data representing pixels of the input image 1402 and / or the downsampled input image 1422 of FIGS. 14A-14D (e.g., for the NN 1500 of the auto-tuning ML system 1405 and / or the image processing system 1406). In an illustrative example, the input data of the input layer 1510 may include data representing pixels of the spatially-varying tuning map 1404 of FIGS. 14A-14D (e.g., for the NN 1500 of the auto-tuning ML system 1405). The image may include image data from an image sensor, including raw pixel data (e.g., including a single color per pixel based on a Bayer filter), or processed pixel values ​​(e.g., RGB pixels in an RGB image). The neural network 1500 includes multiple hidden layers 1512a, 1512b, through 1512n. Hidden layers 1512a, 1512b, through 1512n include n hidden layers, where "n" is an integer greater than or equal to 1. The number of hidden layers may be as many as needed for a given application. Neural network 1500 further includes an output layer 1514 that provides output resulting from the processing performed by hidden layers 1512a, 1512b, through 1512n. In an illustrative example, output layer 1514 may provide a rectified image, such as rectified image 215 of FIG. 2 or rectified image 1415 of FIGS. 14A-14D (e.g., for NN 1500 of image processing system 1406 and / or image processing ML system 210). In an illustrative example, output layer 1514 may provide spatially-varying tuning map 1404 of FIGS. 14A-14D and / or small spatially-varying tuning map 1424 of FIG. 14D (e.g., for NN 1500 of auto-tuning ML system 1405).

[0095] Neural network 1500 is a multi-layer neural network of interconnected filters. Each filter may be trained to learn features that represent input data. Information related to the filters is shared between different layers, and each layer retains information as it is processed. In some cases, neural network 1500 may include a feedforward network, in which there are no feedback connections such that the output of the network is fed back to itself. In some cases, network 1500 may include a recurrent neural network, which may have loops that allow information to be carried across nodes while reading the input.

[0096] In some cases, information can be exchanged between layers through node-to-node interconnections between various layers. In some cases, the network can include a convolutional neural network, which may not connect every node in one layer to every other node in the next layer. In a network in which information is exchanged between layers, nodes in the input layer 1510 can activate a set of nodes in the first hidden layer 1512a. For example, as shown, each of the input nodes in the input layer 1510 may be connected to each of the nodes in the first hidden layer 1512a. The nodes in the hidden layer can transform the information of each input node by applying an activation function (e.g., a filter) to that information. The information derived from the transformation can then be passed to and activated by nodes in the next hidden layer 1512b, which can perform a specific, specified function. Exemplary functions include convolutional functions, downsampling, upscaling, data transformation, and / or any other suitable function. The output of the hidden layer 1512b can then activate nodes in the next hidden layer, and so on. The output of the final hidden layer 1512n can activate one or more nodes in the output layer 1514, which results in a processed output image. In some cases, a node in the neural network 1500 (e.g., node 1516) is shown as having multiple output lines, whereas the node has a single output, and all lines shown as outputting from the node represent the same output value.

[0097] In some cases, each node or interconnection between nodes can have a weight, which is a set of parameters derived from training of neural network 1500. For example, the interconnections between nodes can represent information learned about the interconnected nodes. The interconnections can have adjustable numerical weights that can be adjusted (e.g., based on a training data set), allowing neural network 1500 to be adaptive to inputs and to learn as more data is processed.

[0098] The neural network 1500 is pre-trained to process features from the data in the input layer 1510 using different hidden layers 1512 a, 1512 b through 1512 n to provide output through the output layer 1514 .

[0099] FIG. 16A is a block diagram 1600B illustrating an example of training an image processing system 1406 (e.g., image processing ML system 210) according to some examples. Referring to FIG. 16A, a neural network (e.g., neural network 1500) implemented by image processing system 1406 (e.g., image processing ML system 210) may be pre-trained to process input images and tuning maps. As shown in FIG. 16A, the training data includes input image 1606 and input tuning map 1607. Input image 1402 may be an example of input image 1606. The spatially varying tuning map may be an example of input tuning map 1607. Input image 1606 and input tuning map 1607 may be input into a neural network (e.g., neural network 1500) of image processing ML system 210, and the neural network can generate output image 1608. Corrected image 1415 may be an example of output image 1608. For example, a single input image and several tuning maps (e.g., the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above, and / or other maps associated with an image processing function) can be input to a neural network, and the neural network can output an output image. In another example, a batch of input images and several corresponding tuning maps can be input to a neural network, and the neural network can then generate several output images.

[0100] A set of reference output images 1609 may also be provided for comparison with output images 1608 of image processing system 1406 (e.g., image processing ML system 210) to determine loss (described below). A reference output image may be provided for each input image in input images 1606. For example, output images from reference output images 1609 may include final output images previously generated by the camera system that have desired characteristics for the corresponding input images based on some tuning maps from input tuning maps 1607.

[0101] The neural network parameters may be adjusted based on a comparison of the output image 1608 and the reference output image 1609 by the backpropagation engine 1612. The parameters may include the neural network's weights, biases, and / or other parameters. In some cases, a neural network (e.g., neural network 1500) may adjust node weights using a training process called backpropagation. Backpropagation may include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. This process may be repeated for a number of iterations for each set of training images until the neural network is trained sufficiently that the layer weights are accurately adjusted. Once the neural network is properly trained, the image processing system 1406 (e.g., image processing ML system 210) can process any input image and several tuning maps to generate a modified version of the input image based on the tuning maps.

[0102] The forward pass may involve passing an input image (or a batch of input images) and several tuning maps (e.g., one or more of the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above) through a neural network. The weights of the various filters in the hidden layer may be initially randomized before the neural network is trained. The input image may include a multidimensional array of numbers representing the pixels of the image. In one example, the array may include a 128 x 128 x 11 array of numbers, with 128 rows and 128 columns of pixel locations and 11 input values ​​per pixel location.

[0103] In the first training iteration for a neural network, due to the random selection of weights at initialization, the output may contain values ​​that do not provide a preference for any feature or node. For example, if the output is an array with multiple color components per pixel location, the output image may depict an inaccurate color representation of the input. With the initial weights, the neural network is unable to determine low-level features and therefore cannot make an accurate determination of what the color intensity is likely to be. A loss function can be used to analyze the error in the output. Any suitable loss function definition can be used. An example of a loss function includes the mean squared error (MSE). MSE is:

number

[0104] The loss (or error) is large for the first training data (image data and corresponding tuning map) because the actual values ​​differ significantly from the predicted outputs. The goal of training is to minimize the amount of loss so that the predicted outputs are the same as the training labels. A neural network can perform a backward pass by determining which inputs (weights) contributed most to the network's loss, and then adjust the weights so that the loss is reduced and eventually minimized. In some cases, to determine which weights contributed most to the network's loss, the derivative (or other appropriate function) of the loss with respect to the weights (denoted as dL / dW, where W is the weight in a particular layer) can be calculated. After the derivatives are calculated, a weight update can be performed by updating all the weights in the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient.

number

[0105] FIG. 16B is a block diagram illustrating an example of training of a self-adjusting machine learning (ML) system 1405, according to some examples. Referring to FIG. 16B, a neural network (e.g., neural network 1500) implemented by the self-adjusting machine learning (ML) system 1405 may be pre-trained to process an input image 1606 and / or a downsampled input image 1616. As shown in FIG. 16B, the training data includes the input image 1606 and / or the downsampled input image 1616. The input image 1402 may be an example of the input image 1606. The downsampled input image 1422 may be an example of the downsampled input image 1616. The input image 1606 and / or the downsampled input image 1616 may be input to the neural network (e.g., neural network 1500) of the self-adjusting machine learning (ML) system 1405, and the neural network may generate an output tuning map 1618. The spatially-varying tuning map 1404 may be an example of the output tuning map 1618. The small spatially varying tuning map 1424 may be an example of an output tuning map 1618. For example, the input image 1606 may be an input to a neural network, and the neural network may output one or more output tuning maps 1618 (e.g., one or more of the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above, and / or another other map related to the image processing function). For example, the downsampled input image 1616 may be an input to a neural network, and the neural network may output one or more (small) output tuning maps 1618 (e.g., one or more small versions of the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above, and / or another other map related to the image processing function).In another example, the input images 1606 and / or the batch of downsampled input images 1616 may be input to a neural network, which can then generate several output tuning maps 1618.

[0106] 16B, a neural network (e.g., neural network 1500) implemented by self-tuning machine learning (ML) system 1405 may include a backpropagation engine 1622 similar to backpropagation engine 1612 of FIG. 16A. Backpropagation engine 1622 of FIG. 16B may receive and use reference output tuning map 1619 in a manner similar to how backpropagation engine 1612 of FIG. 16A receives and uses reference output image 1609.

[0107] As mentioned above, in some implementations, the tuning map may be generated automatically using a machine learning system separate from the image processing ML system 210.

[0108] FIG. 17 illustrates the process of filtering luminance channel data (I) according to local linearity constraints 1720, according to some examples. y ) by generating affine coefficients (a,b) that modify the luminance channel data (I y 17 is a block diagram illustrating an example of a system including an image processing ML system and an auto-tuning machine learning (ML) system 1705 that generates a spatially varying tuning map omega (Ω) from the image. The auto-tuning machine learning (ML) system 1705 may be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof. The auto-tuning machine learning (ML) system 1705 is illustrated in FIG. 17 as I RGB The input image data I can be received (e.g., input image 1402), which is shown as RGB The luminance channel data from I yIn some examples, the luminance channel data I y is essentially the input image data I RGB The auto-tuning ML system 1705 uses the input image data I RGB as input and outputs one or more affine coefficients, shown here as a and b. The one or more affine coefficients include a multiplier a. The one or more affine coefficients include an offset b. The auto-tuning ML system 1705, or another imaging system, may use the equation Ω=a*I y +b according to the luminance channel data I y The tuning map Ω can be generated by applying affine coefficients to the equation ∇Ω=a*∇I. y A local linearity constraint 1720 may be imposed according to: a. Use of the local linearity constraint 1720 can ensure that one or more gradients in the map align with one or more gradients in the image data, which can reduce halo effects when applying image processing functions. Using affine coefficients and / or local linearity constraints 1720 to generate the map can produce higher quality spatially-varying image modifications than systems that do not use affine coefficients and / or local linearity constraints to generate the map, e.g., due to better alignment between the image data and the map and reduced halo effects at the boundaries of the depicted object. Each tuning map generated may include a unique set of one or more affine coefficients (e.g., a, b). For example, in the example of FIG. 14C , the auto-tuning ML system 1405 generates five different spatially-varying tuning maps 1404 from a single input image 1402. Thus, the auto-tuning ML system 1405 of FIG. 14C can generate five sets of one or more affine coefficients (e.g., a, b), one set for each of the five different spatially-varying tuning maps 1404. One or more of the affine coefficients (eg, a and / or b) may also vary spatially.

[0109] The auto-tuning ML system 1705 may be implemented using one or more convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or a combination thereof. The auto-tuning ML system 1705 and / or any of the above-listed machine learning elements (which may be part of the auto-tuning ML system 1705) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or a combination thereof.

[0110] 18 is a block diagram illustrating details of the auto-tuning ML system 1705. As shown, the auto-tuning ML system 1705 includes a local neural network 1806 used for patch-by-patch processing (e.g., processing an image patch of an image) and a global neural network 1807 used to process the complete image. The local neural network 1806 can receive and process an image patch 1825 (which is part of a complete image input 1826) as input. The global neural network 1807 can receive and process the complete image input 1826 as input. Output from the global neural network 1807 may be provided to the local neural network 1806. In some examples, the auto-tuning ML system 1705 can output a spatially-varying tuning map 1804. In some examples, the auto-tuning ML system 1705 can output affine coefficients 1822 (e.g., a, b as in FIG. 17 ), which the auto-tuning ML system 1705 can apply to the luminance (Y) channel 1820 to generate the spatially-varying tuning map 1804.

[0111] In some examples, for computational efficiency, the global neural network 1807 may be fed one or more downsized, low-resolution images (downscaled, downsampled, and / or downsized from a full-resolution image), and the local neural network 1806 may be fed one or more full-resolution or high-resolution image patches to obtain a corresponding full-resolution spatially-varying tuning map. The high-resolution image patch (e.g., image patch 1825) and the low-resolution full image input (e.g., image input 1826) may be based on the same image (as shown in FIG. 18 ), but the high-resolution image patch 1825 is based on a high-resolution version of the low-resolution full image input 1826. For example, the low-resolution full image input 1826 may be a downscaled version of a higher-resolution image from which the high-resolution image patch 1825 was extracted. The local NN 1806 may be referred to as the high-resolution NN 1806. The global NN 1807 may be referred to as the low-resolution NN 1807.

[0112] The luminance (Y) channel 1820 is shown as originating from an image patch 1825, but may originate from a complete image input 1826. The luminance (Y) channel 1820 may be high resolution or low resolution. The luminance (Y) channel 1820 may contain luminance data for the complete image input 1826 in its original high resolution or at a downscaled low resolution. The luminance (Y) channel 1820 may contain luminance data for the image patch 1825 in its original high resolution or at a downscaled low resolution.

[0113] FIG. 19 is a block diagram illustrating an example of a neural network architecture 1900 of a local neural network 1806 of an auto-tuning ML system 1705. The local neural network 1806 may be referred to as a high-resolution local neural network 1806 or a high-resolution neural network 1806. The neural network architecture 1900 receives as input an image patch 1825, which may be high-resolution. The neural network architecture 1900 outputs affine coefficients 1822 (e.g., a and / or b, as in FIG. 17 ), which may be high-resolution. The affine coefficients 1822 may be used to generate tuning maps 1804, such as those shown in FIGS. 17 , 18 , 20 , 21A , and / or 21B . A key 1920 identifies how the operation of various NNs is depicted in FIG. 19 . For example, a convolution with a 3×3 filter and a stride of 1 is indicated by a thick white arrow with a black outline pointing to the right. Convolution with a 2x2 filter and a stride of 2 is indicated by a thick black arrow pointing down. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing up.

[0114] FIG. 20 is a block diagram illustrating an example neural network architecture 2000 of the global neural network 1807 of the self-tuning ML system 1705. The global neural network 1807 may be referred to as a low-resolution global neural network 1807 or a low-resolution neural network 1807. The neural network architecture 2000 receives as input a full image input 1826, which may be downsampled and / or low-resolution. The neural network architecture 2000 outputs global features 2010, which may be low-resolution. A key 2020 identifies how various NN operations are depicted in FIG. 20. For example, a convolution with a 3×3 filter and a stride of 1 is depicted by a thick white arrow with a black outline pointing to the right. A convolution with a 2×2 filter and a stride of 2 is depicted by a thick black arrow pointing down. Average pooling is depicted by a thick white arrow with a diagonal black stripe and a black outline pointing down. A fully connected layer is depicted by a thin black arrow pointing to the right.

[0115] FIG. 21A is a block diagram illustrating a neural network architecture 2100A of an auto-tuning ML system 1705, according to some examples. A low-resolution full image 1826 enables a global NN 1807 of the auto-tuning ML system 1705 to generate global information for the image. The low-resolution global neural network 1807 processes the full image and determines global features that are important to incorporate when going from input to output. By using the reduced-resolution image as input for the global neural network 1807, the computation of the global information is significantly lighter than if a full-resolution image were used. The high-resolution patch 1825 enables the auto-tuning network to be locally consistent, ensuring no discontinuities in the local features of the image. Global features may be incorporated into layers (e.g., via channel attention with additive bias) after one or more convolution operations so that the local NN can take global features into account while generating affine coefficients 1822 based on the image patch 1825 and / or while generating a tuning map 1804 based on the image patch 1825. As shown in FIG. 21A, the affine coefficients 1822 may be combined with the Y channel 1820 to generate the tuning map 1804.

[0116] Key 2120A identifies how the operations of various NNs are depicted in FIG. 21A. For example, a convolution with a 3x3 filter and a stride of 1 is indicated by a thick white arrow with a black outline pointing to the right. A convolution with a 2x2 filter and a stride of 2 is indicated by a thick black arrow pointing down. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing up. Channel attention with additional bias is indicated by a thin black arrow pointing up and / or left (e.g., upward from the global features determined by the global NN 1807).

[0117] FIG. 21B is a block diagram illustrating another example of a neural network architecture 2100B of the auto-tuning ML system 1705, according to some examples. The input to the neural network architecture 2100B may be an input image 2130. The input image 2130 may be a high-resolution full image. The input image 2130 may be an original high-resolution or downscaled low-resolution full image input 1826. The input image 2130 may be a high-resolution or downscaled low-resolution image patch 1825. The neural network architecture 2100B may process the input image 2130 to generate affine coefficients 1822 based on the input image 2130. The neural network architecture 2100B may process the input image 2130 to generate a tuning map 1804 based on the input image 2130. The affine coefficients 1822 may be combined with the Y channel 1820 to generate the tuning map 1804. The spatial attention engine 2110 and the channel attention engine 2115 are part of the neural network architecture 2100B. The spatial attention engine 2110 is shown in more detail in Figure 21C. The channel attention engine 2115 is shown in more detail in Figure 21D.

[0118] Key 2120B identifies how the operations of various NNs are depicted in Figures 21A, 21B, and 21C. For example, a convolution with a 3x3 filter and a stride of 1 is indicated by a thick white arrow with a black outline pointing to the right. A convolution with a 1x1 filter and a stride of 1 is indicated by a thin black arrow pointing down. A convolution with a 2x2 filter and a stride of 2 is indicated by a thick black arrow pointing down. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing up. Operations involving attention, upsampling, and multiplication are indicated by thin black arrows pointing up and / or to the side. A circled "X" symbol indicates application of affine coefficients 1822 to the Y channel 1820. A double circled "X" symbol indicates element-wise multiplication after magnification. A thin dashed black arrow extending horizontally (in Figure 21D) indicates shared parameters.

[0119] In some examples, the local NN 1808 can generate the tuning map 1804 directly without first generating the affine coefficients 1822 .

[0120] 21C is a block diagram illustrating an example neural network architecture 2100C of a spatial attention engine 2110, according to some examples. The spatial attention engine 2110 includes max pooling, average pooling, and concatenation.

[0121] 21D is a block diagram illustrating an example of a neural network architecture 2100D of a channel attention engine 2115, according to some examples. The spatial attention engine 2110 includes max pooling, average pooling, shared parameters, and summation.

[0122] By using machine learning to generate the affine coefficients 1822 and / or tuning map 1804, the imaging system achieves increased customizability and context sensitivity. The imaging system can provide adjustment maps that perform well with the content of the input image without producing visual artifacts such as halos. Generating the affine coefficients 1822 and / or tuning map 1804 based on machine learning may also be more efficient than traditional manual adjustment of image processing parameters. Imaging innovation can also be accelerated based on generating the affine coefficients 1822 and / or tuning map 1804 based on machine learning. For example, generating the affine coefficients 1822 and / or tuning map 1804 using machine learning may allow the imaging system to quickly and easily adapt to work with data from additional sensors, different types of lenses, different types of camera arrays, and other modifications.

[0123] The neural network of the auto-tuning ML system 1705 may be trained using supervised learning techniques, similar to those described above with respect to the image processing ML system 210. For example, backpropagation may be used to adjust the parameters of the neural network of the auto-tuning ML system 1705. The training data may include input images and known output images with desired characteristics by applying different tuning maps. For example, based on the input images and output images, the neural network is trained to generate a set of masks that, when applied to the input image, give a corresponding output. Using saturation and tone as an illustrative example, the neural network attempts to determine how much saturation and tone need to be applied to various pixels of the input image to achieve the characteristics of the pixel in the output image. In such an example, the neural network may generate a saturation map and a tone map based on the training.

[0124] The inference procedure of the self-tuning ML system 1705 (once the neural network is trained) may be performed by processing the input image to produce the tuning map 1804. For example, global features may be extracted from a low-resolution version of the input image by a global neural network 1807. A single inference may be performed on the full low-resolution image input. A local neural network 1806 may then be used to perform patch-based inference, with the global features from the global neural network 1807 being fed to the local network 1806 during inference. A high-resolution patch-based inference is then performed on the patches of the input image to produce the tuning map 1804.

[0125] In some examples, the image processing ML system 210 (and in some cases the auto-tuning ML system 1705) described above may be used when capturing an image or when processing a previously captured image. For example, when capturing an image, the image processing ML system 210 may process the image to generate an output image with optimal characteristics (e.g., with respect to noise, sharpness, saturation, hue, etc.) based on the use of a tuning map. In another example, a previously generated, stored image may be retrieved and processed by the image processing ML system 210 to generate an enhanced output image with optimal characteristics based on the use of a tuning map.

[0126] In some examples, the image processing ML system 210 (and in some cases the auto-tuning ML system 1705) described above may be used when adjusting the ISP. For example, parameters for an ISP are traditionally manually adjusted by an expert with experience in how to process input images for a desired output image. As a result of the correlation between the ISP modules (e.g., filters) and the true number of adjustable parameters, the expert may need several weeks (e.g., 3-8 weeks) to determine, test, and / or adjust device settings for parameters based on a particular camera sensor and ISP combination. Each camera sensor and ISP combination may be adjusted by an expert because the camera sensor or other camera characteristics (e.g., lens characteristics or imperfections, aperture size, shutter speed and movement, flash brightness and color, and / or other characteristics) may affect the captured image and therefore at least some of the adjustable parameters for the ISP.

[0127] FIG. 22 is a block diagram illustrating an example of a preconditioned image signal processor (ISP) 2208. As shown, an image sensor 2202 captures raw image data. Photodiodes in the image sensor 2202 capture varying shades of gray (or monochrome). Color filters can be applied to the image sensor to provide color-filtered raw input data 2204 (e.g., having a Bayer pattern). The ISP 2208 has individual function blocks that each apply a specific operation to the raw camera sensor data to produce a final output image. For example, the function blocks may include blocks dedicated to demosaicing, gain, white balance, color correction, gamma compression (or gamma correction), tone mapping, noise reduction (denoising), among others. For example, the demosaicing function block of the ISP 2208 can help generate an output color image 2209 using the color-filtered raw input data 2204 by interpolating the pixel's color and brightness using neighboring pixels. This demosaicing process may be used by the ISP 2208 to evaluate the color and brightness data of a given pixel and compare those values ​​with data from neighboring pixels. The ISP 2208 can then use demosaicing algorithms to produce appropriate color and brightness values ​​for the pixels. The ISP 2208 can perform various other image processing functions such as noise reduction, sharpening, tone mapping and / or conversion between color spaces, autofocus, gamma, exposure, white balance, among other possible image processing functions, before providing the final output color image 2209.

[0128] The function blocks of the ISP 2208 require a large number of manually adjusted tuning parameters 2206 to meet certain specifications. In some cases, more than 10,000 parameters need to be adjusted and controlled for a given ISP. For example, to optimize the output color image 2209 according to certain specifications, the algorithm for each function block must be optimized by adjusting the algorithm's tuning parameters 2206. New function blocks must also be continually added to handle the various cases that arise in the space. The large number of manually adjusted parameters leads to very time-consuming and expensive support requirements for the ISP.

[0129] In some cases, an ISP may be implemented using a machine learning system (referred to as a machine learning ISP) to jointly perform multiple ISP functions.

[0130] FIG. 23 is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) 2300. The machine learning ISP 2300 may include an input interface 2301 that can receive raw image data from an image sensor 2302. In some cases, the image sensor 2302 includes an array of photodiodes that capture frames 2304 of raw image data. Each photodiode can represent a pixel location and generate a pixel value for that pixel location. The raw image data from the photodiodes may include a single brightness or grayscale value for each pixel location in the frame 2304. For example, a color filter array may be integrated with the image sensor 2302 or used in conjunction with the image sensor 2302 (placed over the photodiodes) to convert monochrome information to brightness.

[0131] One illustrative example of a color filter array is a Bayer pattern color filter array (or Bayer color filter array), which enables image sensor 2302 to capture frames of pixels having a Bayer pattern in which each pixel location has one of a red, green, or blue filter. For example, raw image patch 2306 from frame 2304 of raw image data has a Bayer pattern based on the Bayer color filter used with image sensor 2302. As shown in the pattern of raw image patch 2306 shown in FIG. 23, the Bayer pattern includes a red filter, a blue filter, and a green filter. Bayer color filters operate by filtering out incoming light. For example, a photodiode with a green portion of the pattern passes green information (half of the pixel), a photodiode with a red portion of the pattern passes red information (one-quarter of the pixel), and a photodiode with a blue portion of the pattern passes blue information (one-quarter of the pixel).

[0132] In some cases, a device may include multiple image sensors (which may be similar to image sensor 2302), in which case the machine learning ISP operations described herein may be applied to raw image data acquired by the multiple image sensors. For example, a multiple-camera device may capture image data using the multiple cameras, and machine learning ISP 2300 may apply ISP operations to the raw image data from the multiple cameras. In one illustrative example, a dual-camera mobile phone, tablet, or other device may be used to capture larger images with a wider angle (e.g., a larger field of view (FOV)), capture more light (resulting in greater sharpness and clarity, among other benefits), generate 360-degree (e.g., virtual reality) video, and / or perform other functions that are improved over those achieved by a single-camera device.

[0133] Raw image patches 2306 are provided to and received by input interface 2301 for processing by machine learning ISP 2300. Machine learning ISP 2300 can use a neural network system 2303 for the ISP's tasks. For example, the neural network of neural network system 2303 can be trained to directly derive a mapping from raw image training data captured by an image sensor to a final output image. For example, the neural network can be trained using numerous examples of raw data inputs (e.g., with color-filtered patterns) and also using examples of the corresponding desired output images. Using the training data, neural network system 2303 can learn the mapping from the raw inputs required to obtain the output image, after which ISP 2300 can produce output images similar to those produced by a traditional ISP.

[0134] The neural network of ISP2300 may include an input layer, multiple hidden layers, and an output layer. The input layer includes raw image data acquired by the image sensor 2302 (e.g., a raw image patch 2306 or a complete frame of raw image data). The hidden layer may include filters that may be applied to the raw image data and / or to the output from the previous hidden layer. Each of the filters in the hidden layer may include weights used to indicate the importance of the filter's nodes. In one illustrative example, the filter may include a 3x3 convolution filter convolved around the input array, with each entry of the 3x3 filter having a unique weight value. Each convolution iteration (or stride) of the 3x3 filter applied to the input array may produce a single weighted output feature value. The neural network may have a series of multiple hidden layers, with earlier layers determining low-level characteristics of the input and later layers building a hierarchy of more complex characteristics. The hidden layers of the ISP2300 neural network are connected to a high-dimensional representation of the data. For example, a layer may include several blocks of recursive convolutions with a large number of channels (dimensions). In some cases, the number of channels may be an order of magnitude larger than the number of channels in an RGB or YCbCr image. The illustrative example provided below includes recursive convolutions with 64 channels each, resulting in a nonlinear hierarchical network structure that produces high-quality image detail. For example, as described in more detail herein, the number of channels n (e.g., 64 channels) refers to having an n-dimensional (e.g., 64-dimensional) representation of data at each pixel location. Conceptually, the number of channels n represents "n features" (e.g., 64 features) at a pixel location.

[0135] Neural network system 2303 collaboratively accomplishes various ISP functions. Certain parameters of the neural network applied by neural network system 2303 have no clear analog in a traditional ISP, and conversely, certain functional blocks of a traditional ISP system have no clear counterpart in a machine learning ISP. For example, a machine learning ISP performs signal processing functions as a single unit, rather than having individual functional blocks that a typical ISP might include to perform various functions. Further details of the neural network applied by neural network system 2303 are described below.

[0136] In some examples, the machine learning ISP 2300 may also include an optional pre-processing engine 2307 that can process additional image tuning parameters to augment the input data. Such additional image tuning parameters (or augmentation data) may include, for example, tone data, radial distance data, auto white balance (AWB) gain data, any combination thereof, and / or any other additional data that can augment the pixels of the input data. By augmenting the raw input pixels, the input becomes a multidimensional set of values ​​for each pixel location of the raw image data.

[0137] Based on the determined high-level features, the neural network system 2303 can generate an RGB output 2308 based on the raw image patch 2306. The RGB output 2308 includes red, green, and blue components for each pixel. In this application, the RGB color space is used as an example. Those skilled in the art will understand that other color spaces, such as luma and chroma (YCbCr or YCV) color components, or other suitable color components, may also be used. The RGB output 2308 may be output from the output interface 2305 of the machine learning ISP 2300 and used to generate image patches in a final output image 2309 (constituting the output layer). In some cases, the array of pixels in the RGB output 2308 may include fewer dimensions than the dimensions of the input raw image patch 2306. In one illustrative example, raw image patch 2306 may include a 128x128 array of raw pixels (e.g., in a Bayer pattern), but application of each convolution filter in neural network system 2303 results in RGB output 2308 including an 8x8 array of pixels. The smaller output size of RGB output 2308 than raw image patch 2306 is a by-product of applying the convolution filters and designing neural network system 2303 to avoid padding the data processed through each of the convolution filters. Multiple convolution layers further reduce the output size. In such cases, patches from the input frame of raw image data 2304 may overlap, so that the final output image 2309 includes a complete picture. The resulting final output image 2309 includes processed image data derived from the raw input data by neural network system 2303. The final output image 2309 may be rendered for display, used for compression (or coding), stored, or used for any other image-based purpose.

[0138] 24 is a block diagram illustrating an example of a neural network architecture 2400 of a machine learning (ML) image signal processor (ISP) 2300. Pixel shuffle upsampling is an upsampling method in which channel dimensions are reshaped along the spatial dimension. In one example using 2x upsampling for illustrative purposes, xxxx (referring to 4 channels x 1 spatial location) along four channels is used to generate a single channel with twice the spatial dimension: xx xx (referring to 1 channel x 4 spatial location). An example of pixel shuffle upsampling is described in "Checkerboard artifact-free sub-pixel convolution" by Andrew Aitken et al., which is incorporated herein by reference in its entirety for all purposes.

[0139] Using machine learning to implement ISP functions allows the ISP to be customizable. For example, various functions can be developed and applied by presenting target data examples and modifying the network weights through training. Machine learning-based ISPs can also achieve a faster turnaround for updates compared to hardwired or heuristic-based ISPs. Furthermore, machine learning-based ISPs eliminate the time-consuming task of adjusting tuning parameters required for pre-tuned ISPs, which require significant effort and manpower to manage ISP infrastructure. Holistic development can be used for machine learning ISPs, during which the end-to-end system is directly optimized and created. This holistic development contrasts with the individual development of pre-tuned ISP functional blocks. Machine learning ISPs can also accelerate innovation in imaging. For example, customizable machine learning ISPs unlock numerous innovation possibilities, enabling developers and engineers to more quickly derive, develop, and adapt solutions to work with novel sensors, lenses, and camera arrays, among other advancements.

[0140] As mentioned above, the image processing system 1406 (e.g., image processing ML system 210) and / or auto-adjustment ML system 1405 described above may be used when adjusting the ML ISP and / or the conventional ISP. The various tuning maps described above (e.g., the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above, and / or other maps related to image processing functions) may be used to adjust the ISP. When used to adjust ISO, the maps may be referred to as adjustable knobs. Examples of tuning maps (or adjustable knobs) that may be used to adjust the ISP include local tone manipulation (e.g., using a tone map), detail enhancement, and color saturation, among others. In some cases, local tone manipulation may include contrast-limited adaptive histogram equalization (e.g., using OpenCV). In some cases, detail enhancement may be performed using a domain transform for edge-aware filtering (e.g., using OpenCV). In some examples, color saturation may be performed using the Pillow library. In some implementations, the auto-tuning ML system 1705 may be used to generate tuning maps or knobs used to adjust the ISP.

[0141] Key 2420 identifies how various NN operations are depicted in Figure 21A. For example, a convolution with a 3x3 filter and a stride of 1 is depicted by a thick white arrow with a black outline pointing to the right. A convolution with a 2x2 filter and a stride of 2 is depicted by a thick black arrow pointing down. Upsampling (e.g., bilinear upsampling) is depicted by a thick black arrow pointing up.

[0142] 25A is a conceptual diagram illustrating an example of a first tone adjustment strength applied to an exemplary input image to generate a modified image, according to some examples. The tone adjustment strength for the modified image in FIG. 25A is 0.0.

[0143] 25B is a conceptual diagram illustrating an example of the strength of a second tone adjustment applied to the exemplary input image of FIG. 25A to generate a modified image, according to some examples. The strength of the tone adjustment for the modified image of FIG. 25A is 0.5.

[0144] 25A and 25B are images showing an example of applying various tone levels by performing a CLAHE using an OpenCV implementation. For example, an image can be partitioned into a fixed grid, and the histogram can be constrained to predetermined values ​​before calculating the cumulative distribution function (CDF). The constraining value can determine the strength of the local tone manipulation. In the example of FIGS. 25A and 25B, the application includes a maximum constraining value of 0.5. To obtain a final result, the grid transformation may be interpolated. As shown, changes in tone result in different amounts of brightness adjustment applied to the pixels of the image processed by the ISP, which can result in a darker or lighter-toned image.

[0145] Figure 26A is a conceptual diagram illustrating an example of a first detail adjustment strength applied to the example input image of Figure 25A to generate a modified image, according to some examples. The detail adjustment (detail enhancement) strength for the modified image of Figure 26A is 0.0.

[0146] Figure 26B is a conceptual diagram illustrating an example of a second detail adjustment strength applied to the exemplary input image of Figure 25A to generate a modified image, according to some examples. The detail adjustment (detail enhancement) strength for the modified image of Figure 26A is 0.5.

[0147] 26A and 26B are images showing examples of various detail enhancement applications using an OpenCV implementation. For example, detail enhancement may be used for edge-preserving smoothing of the original image. Detail may be represented as described above with respect to a saturation map. For example, image detail may be obtained by subtracting a filtered image (the smoothed image resulting from edge-preserving filtering) from the input image. Detail may be enhanced at multiple scales. Range sigma may be equal to 0.05 (Range sigma = 0.05). As mentioned above, spatial sigma is a hyperparameter of the bilateral filter. Spatial sigma may be used to control the strength of detail enhancement. In some examples, the maximum spatial sigma may be set to 5.0 (Max Spatial Sigma = 5.0). While FIG. 26A provides an example where spatial sigma is equal to 0, FIG. 26B shows an example where spatial sigma is equal to the maximum value of 5.0. As can be seen, the image in FIG. 26B has sharper details and more noise compared to the image in FIG. 26A.

[0148] 27A is a conceptual diagram illustrating an example of the strength of a first color saturation adjustment applied to the exemplary input image of FIG. 25A to generate a modified image, according to some examples. The first color saturation value for the modified image of FIG. 27A is 0.0, indicating maximum color desaturation. The modified image of FIG. 27A represents a grayscale version of the exemplary input image of FIG. 25A.

[0149] 27B is a conceptual diagram illustrating an example of the strength of a second color saturation adjustment applied to the exemplary input image of FIG. 25A to generate a modified image, according to some examples. The strength of the second color saturation adjustment for the modified image of FIG. 27A is 1.0, indicating no change in saturation (a saturation adjustment strength of 0). The modified image of FIG. 27B matches the exemplary input image of FIG. 25A in color saturation.

[0150] FIG. 27C is a conceptual diagram illustrating an example strength of a third color saturation adjustment applied to the exemplary input image of FIG. 25A to generate a modified image, according to some examples. The third color saturation value for the modified image of FIG. 27C is 2.0, indicating the maximum strength of color saturation increase. The modified image of FIG. 27C represents an over-saturated version of the exemplary input image of FIG. 25A. In an attempt to illustrate the over-saturated nature of the modified image of FIG. 27C, some elements are depicted as brighter in FIG. 27C compared to FIGS. 27A-27B or FIG. 25A.

[0151] Figures 27A, 27B, and 27C are images showing examples of the application of various color saturation adjustments. The image in Figure 27A shows the image when a saturation of 0 is applied, resulting in a grayscale (desaturated) image similar to that described above with respect to Figures 9 through 11. The image in Figure 27B shows the image when a saturation of 1.0 is applied, resulting in no change in saturation. The image in Figure 27C shows the image when a saturation of 2.0 is applied, resulting in a highly saturated image.

[0152] FIG. 28 is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) receiving as input various tuning parameters (similar to the tuning maps described above) used to adjust the ML ISP 2300. The ML ISP 2300 may be similar to and perform similar operations as the ML ISP 2300 discussed above with respect to FIG. 23. The ML ISP 2300 includes a trained machine learning (ML) model 2802. Based on processing of the parameters, the trained ML model 2802 of the ML ISP 2300 outputs an enhanced image based on the tuning parameters. Image data 2804 including a raw image with simulated ISO noise and radial distance from the center of the raw image is provided as input. Tuning parameters 2806 used to capture the raw image include red channel gain, blue channel gain, ISO speed, and exposure time. The tuning parameters 2806 include tone enhancement strength, detail enhancement strength, and color saturation strength. The additional tuning parameters 2806 are similar to the tuning maps discussed above and may be applied at the image level (a single value applied to the entire image) or at the pixel level (a different value for each pixel in the image).

[0153] 29 is a block diagram illustrating example tuning parameter values ​​that may be provided to the machine learning (ML) image signal processor (ISP) ML ISP 2300. An exemplary raw input image is shown along with an image representing the radial distance from the center of the raw image. The raw image data shown may represent image data for only one color channel (e.g., green), for example. For the tuning parameters 2816 used to capture the raw input image, the red gain value is 1.957935, the blue gain value is 1.703827, the ISO speed is 500, and the exposure time is 3.63e -04 The tuning parameters include 0.0 for tone strength, 0.0 for detail strength, and 1.0 for saturation strength.

[0154] FIG. 30 is a block diagram illustrating additional examples of specific tuning parameter values ​​that may be provided to the machine learning (ML) image signal processor (ISP) ML ISP 2300. The same raw input image, radial distance information, and image capture parameters (red gain, blue gain, ISO speed, and exposure time) as shown in FIG. 29 are shown in FIG. 30. Different tuning parameters 2826 are provided compared to tuning parameters 2816 of FIG. 29. Tuning parameters 2826 include a tone intensity of 0.2, a detail intensity of 1.9, and a saturation intensity of 1.9. The output enhanced image 2828 of FIG. 30 has a different appearance compared to the output enhanced image 2808 of FIG. 29 based on the different tuning parameters. For example, the output enhanced image 2828 of FIG. 30 is highly saturated and has finer detail, while the enhanced image of FIG. 29 does not have the effect of saturation or detail enhancement. Because Figures 28-30 are shown in grayscale rather than color, the red channels of output enhanced image 2828 and output enhanced image 2808 are shown in Figures 28-30 to show the change in saturation of the red color. Increased saturation of red appears as lighter areas, while decreased saturation of red appears as darker areas. Because Figure 30 shows the red channel of output enhanced image 2828, the brighter flowers in output enhanced image 2828 (compared to output enhanced image 2808 and the raw image data) indicate that the red color of the flowers in output enhanced image 2828 is more saturated.

[0155] The trained ML model 2802 of the ML ISP 2300 is trained to process tuning parameters (e.g., tuning maps or knobs) and generate enhanced or modified output images. The training setup may include a PyTorch implementation. The network receptive field may be set to 160x160. The training images may include thousands of images captured by one or more devices. Patch-based training may be performed, where patches of input images (rather than entire images) are provided to the ML ISP 2300 during training. In some examples, the input patches have a size of 320x320, and the output patches generated by the ML ISP 2300 have a size of 160x160. The size reduction may be based on the convolutional nature of the neural network of the ML ISP 2300 (e.g., the neural network is implemented as a CNN or other network using convolutional filters). In some cases, a batch of image patches may be provided to the ML ISP 2300 at each training iteration. In one illustrative example, the batch size may include 128 images. The neural network may be trained until the validation loss stabilizes. In some cases, stochastic gradient descent (e.g., an Adam optimizer) may be used to train the network. In some cases, a learning rate of 0.0015 may be used. In some examples, the trained ML model 2802 may be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof.

[0156] FIG. 31 is a block diagram illustrating examples of objective functions and various losses that may be used during training of the machine learning (ML) image signal processor (ISP) ML ISP2300. Pixel-by-pixel losses include L1 loss and L2 loss. Pixel-by-pixel losses can result in better color reproduction for the output image generated by the ML ISP2300. Structural losses may also be used to train the ML ISP2300. Structural losses include a structural similarity index (SSIM). For SSIM, a window size of 7×7 and a Gaussian sigma of 1.5 may be used. Another structural loss that can be used is multi-scale SSIM (MS-SSIM). For MS-SSIM, a window size of 7×7, a Gaussian sigma of 1.5, and scale weights of [0.9, 0.1] may be used. Structural losses may be used to better preserve high-frequency information. Pixel-by-pixel losses and structural losses, such as the L1 loss and MS-SSIM shown in FIG. 31, may be used together.

[0157] In some examples, the ML ISP2300 can perform patch-by-patch model inference after its neural network is trained. For example, the input to the neural network can include one or more raw image patches (e.g., having a Bayer pattern) from a frame of raw image data, and the output can include an output RGB patch (or a patch having another color component representation, such as YUV). In one illustrative example, the neural network takes a 128x128 pixel raw image patch as input and produces an 8x8x3 RGB patch as the final output. Based on the convolutional nature of the various convolution filters applied by the neural network, many of the pixel locations outside the 8x8 array from the raw image patch are consumed by the network to generate the final 8x8 output patch. This reduction in data from input to output is due to the amount of context required to understand neighboring information for processing pixels. Having a larger input raw image patch with all of the neighboring information and context aids in the processing and production of smaller output RGB patches.

[0158] In some examples, based on decreasing pixel locations from input to output, 128x128 raw image patches are designed so that they overlap in the raw input image. In such examples, the 8x8 outputs are non-overlapping. For example, for the first 128x128 raw image patch in the upper left corner of the raw image frame, the first 8x8 RGB output patch is generated. The next 128x128 patch in the raw image frame is eight pixels to the right of the last 128x128 patch, and therefore overlaps with the last 128x128 pixel patch. The next 128x128 patch is processed by the neural network to generate a second 8x8 RGB output patch. The second 8x8 RGB patch is placed next to the first 8x8 RGB output patch (generated using the previous 128x128 raw image patch) in the complete final output image. Such processing may be performed until the 8x8 patches that make up the complete output image are generated.

[0159] FIG. 32 is a conceptual diagram illustrating an example of patch-by-patch model estimation that results in non-overlapping output patches at a first image location, according to some examples.

[0160] FIG. 33 is a conceptual diagram illustrating an example of patch-by-patch model estimation of FIG. 32 at a second image location, according to some examples.

[0161] FIG. 34 is a conceptual diagram illustrating an example of patch-by-patch model estimation of FIG. 32 at a third image location, according to some examples.

[0162] FIG. 35 is a conceptual diagram illustrating an example of patch-by-patch model estimation of FIG. 32 at a fourth image location, according to some examples.

[0163] Figures 32, 33, 34, and 35 show examples of patch-by-patch model inference that result in non-overlapping output patches. As shown in Figures 32 to 35, the input patch may have a size of ko+160, where ko is the output patch size (output patch size equals ko). Diagonally striped pixels refer to padding pixels (e.g., reflective). The input patches overlap, as indicated by the boxes with dashed outlines in Figures 32 to 35. The output patches do not overlap, as indicated by the white boxes.

[0164] As described above, additional tuning parameters (or tuning maps or adjustment knobs) may be applied at the image level (a single value applied to the entire image) or at the pixel level (a different value for each pixel in the image).

[0165] Figure 36 is a conceptual diagram showing an example of a spatially fixed tone map (or mask) applied at the image level, where a single value of t is applied to all pixels of the image. As described above, tone maps may be provided that include different values ​​for various pixels of an image (e.g., as shown in Figure 8), and thus may vary spatially.

[0166] FIG. 37 is a conceptual diagram illustrating an example of application of a spatially-varying map 3703 (or mask) to process input image data 3702 to generate an output image 3715 in which the strength of the saturation adjustment spatially varies, according to some examples. The spatially-varying map 3703 includes a saturation map. The spatially-varying map 3703 includes a value for each location corresponding to a pixel in a raw image 3702 (also called a Bayer image) produced by an image sensor. In contrast to a spatially fixed map, the values ​​at different locations in the spatially-varying map 3703 may be different. The spatially-varying map 3703 includes a value of 0.5 corresponding to an area of ​​pixels in the raw image 3702 depicting a flower in the foreground. This area with a value of 0.5 is shown in gray in the spatially-varying map 3703. A value of 0.5 in the spatially-varying map 3703 indicates that the saturation of that area should remain the same, neither increasing nor decreasing. Spatially-varying map 3703 contains a value of 0.0 corresponding to the area of ​​pixels in raw image 3702 that depict the background behind the flower. This area with a value of 0.0 is shown as black in spatially-varying map 3703. A value of 0.0 in input saturation map 203 indicates that the area should be fully desaturated. Corrected image 3715 shows an example of an image produced (e.g., by image processing ML system 210) by applying saturation to an intensity and orientation based on spatially-varying map 3703. Corrected image 3705 represents an image in which the foreground flower is still saturated to the same extent as reference image 3716 (the saturation intensity is unchanged), but the background is fully desaturated and therefore depicted in grayscale. Reference image 3716 is used as the base image that the system is expected to produce using a neutral saturation (0.5) applied to the entire raw image 3702. Because Figure 37 is shown in grayscale rather than color, the green channels of corrected image 3705 and reference image 3716 are shown in Figure 37 to show the changes in green saturation. Increased green saturation appears as lighter areas, while decreased green saturation appears as darker areas.FIG. 37 shows the green channel of the corrected image 3705, and because the background behind the flower is primarily green, the background appears much darker in the corrected image 3705 (compared to the reference image 3716), indicating that the green color of the background in the corrected image 3705 is desaturated (compared to the reference image 3716).

[0167] FIG. 38 is a conceptual diagram illustrating an exemplary application of a spatially-variant map 3803 to process input image data 3802 to generate an output image 3815 with spatially-varying intensity of tone adjustment and spatially-varying intensity of detail adjustment, according to some examples. The spatially-variant map 3703 includes a tone map. The spatially-variant map 3803 includes a value for each location corresponding to a pixel in a raw image 3802 (also called a Bayer image) produced by an image sensor. The values ​​at different locations in the spatially-variant map 3803 may be different. The spatially-variant map 3803 includes a tone value of 0.5 and a detail value of 5.0, which corresponds to an area of ​​pixels in the raw image 3802 depicting a flower in the foreground. This area with a tone value of 0.5 and a detail value of 5.0 is shown as white in the spatially-variant map 3803. A tone value of 0.5 in the spatially-variant map 3803 corresponds to a gamma (γ) value of 1.0 (10 0.0 ) which means that the tone in that area should remain the same without change. A detail value of 5.0 in spatially-varying map 3803 indicates that the detail in that area should be increased. Spatially-varying map 3803 contains a tone value of 0.0 and a detail value of 0.0 which corresponds to an area of ​​pixels in raw image 3802 that depict the background behind the flower. This area with a tone value of 0.0 and a detail value of 0.0 is shown as black in spatially-varying map 3803. A tone value of 0.0 in input tone map 203 corresponds to a gamma (γ) value of 2.5 (10 0.43803), meaning that the tone in that area should be darkened. The orientation of the tone value to gamma value mapping in FIG. 38 is opposite to that of the tone value to gamma value mapping in FIG. 7, with 0 being lighter and 1 being darker. A detail value of 5.0 in spatially-varying map 3803 indicates that the detail in that area should be reduced. Corrected image 3805 shows an example of an image produced (e.g., by image processing ML system 210) by correcting the tone and detail to an intensity and orientation based on spatially-varying map 3803. Corrected image 3805 represents an image in which the flowers in the foreground have the same tone (compared to reference image 3816) but with increased detail (compared to reference image 3816), and the background has darkened tone (compared to reference image 3816) and reduced detail (compared to reference image 3816). The reference image 3816 is used as a base image that the system is expected to produce with a neutral tone (0.5) and neutral detail applied to the entire raw image 3802 .

[0168] FIG. 39 is a conceptual diagram illustrating an automatically adjusted image generated by an image processing system using one or more tuning maps to adjust an input image.

[0169] 40A is a conceptual diagram illustrating an output image generated by an image processing system using one or more spatially-varying tuning maps generated using a self-tuning machine learning (ML) system, according to some examples. An input image and an output image are shown in FIG. 40A. The output image is a modified version of the input image that is modified based on the one or more spatially-varying tuning maps.

[0170] Figure 40B is a conceptual diagram illustrating an example of a spatially-varying tuning map that may be used to generate the output image shown in Figure 40A from the input image shown in Figure 40A, according to some examples. In addition to again showing the input and output images of Figure 40A, Figure 40B shows a detail map, a noise map, a tone map, a saturation map, and a hue map.

[0171] 41A is a conceptual diagram illustrating an output image generated by an image processing system using one or more spatially-varying tuning maps generated using a self-tuning machine learning (ML) system, according to some examples. An input image and an output image are shown in FIG. 41A. The output image is a modified version of the input image that is modified based on the one or more spatially-varying tuning maps.

[0172] Figure 41B is a conceptual diagram illustrating examples of spatially-varying tuning maps that may be used to generate the output image shown in Figure 41A from the input image shown in Figure 41A, according to some examples. In addition to again showing the input image of Figure 41A, Figure 41B shows a sharpness map, a noise map, a tone map, a saturation map, and a hue map.

[0173] 42A is a flowchart illustrating an example of a process 4200 for processing image data using one or more neural networks using the techniques described herein. The process 4200 may be performed by an imaging system, such as the image capture and processing system 100, the image capture device 105A, the image processing device 105B, the image processor 150, the ISP 154, the host processor 152, the image processing ML system 210, the system of FIG. 14A, the system of FIG. 14B, the system of FIG. 14C, the system of FIG. 14D, the image processing system 1406, the auto-tuning ML system 1405, the downsampler 1418, the upsampler 1426, the neural network 1500, the neural network architecture 1900, the neural network architecture 2000, the neural network architecture 2100A, the neural network architecture 2100B, the spatial attention engine 2110, the channel The system may include an attention engine 2115, an image sensor 2202, a pre-conditioned ISP 2208, a machine learning (ML) ISP 2300, a neural network system 2303, a pre-processing engine 2307, an input interface 2301, an output interface 2305, a neural network architecture 2400, a trained machine learning model 2802, an imaging system that performs the process 4200, a computing system 4300, or a combination thereof.

[0174] At block 4202, process 4200 includes an imaging system acquiring image data. In some implementations, the image data includes a processed image having multiple color components for each pixel of the image data. For example, the image data may include one or more RGB images, one or more YUV images, or other color images previously captured and processed by a camera system, such as the image of scene 110 generated by image capture and processing system 100 shown in FIG. 1. In some implementations, the image data includes raw image data from one or more image sensors (e.g., image sensor 130). The raw image data includes a single color component for each pixel of the image data. In some cases, the raw image data is acquired from one or more image sensors that are filtered by a color filter array, such as a Bayer color filter array. In some implementations, the image data includes one or more patches of image data. A patch of image data includes a subset of a frame of image data. In some cases, generating the modified image includes generating multiple patches of output image data. Each patch of output image data may include a subset of pixels of the output image.

[0175] Examples of raw image data include image data captured using image capture and processing system 100, image data captured using image capture device 105A and / or image processing device 105B, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, input image 1606, downsampled input image 1616, and I of FIG. RGB ,Luminance channel I in Fig. 17 y, luminance channel 1820, image patch 1825, complete image input 1826, input image 2130, color filtered raw input data 2204, output color image 2209, frame of raw image data 2304, raw image patch 2306, RGB output 2308, final output image 2309, the raw input image of FIG. 24, the output RGB image of FIG. 24, the input image of FIGS. 25A-25B, the input image of FIGS. 26A-26B, the input image of FIGS. 27A-27C, image data 2804, image data 2814, the input patch of FIGS. 32-35, raw image 3702, reference image 3716, raw image 3802, reference image 3816, the input image of FIG. 39, the input image of FIGS. 40A-40B, the input image of FIGS. 41A-41B, the image data of block 4252, other image data described herein, other images described herein, or combinations thereof. In some examples, block 4202 of process 4200 may correspond to block 4252 of process 4250.

[0176] At block 4204, process 4200 includes the imaging system acquiring one or more maps. The one or more maps are also referred to herein as one or more tuning maps. Each map of the one or more maps is associated with a respective image processing function. Each map also includes a value indicating an amount of the image processing function to be applied to a corresponding pixel of the image data. For example, a map of the one or more maps includes multiple values ​​and is associated with an image processing function, with each value of the multiple values ​​in the map indicating an amount of the image processing function to be applied to a corresponding pixel of the image data. Figures 3A and 3B, discussed above, show values ​​of an example tuning map (tuning map 303 of Figure 3B) and corresponding pixels of an example input image (input image 302 of Figure 3A).

[0177] In some cases, the one or more image processing functions associated with the one or more maps include a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, a hue adjustment function, or any combination thereof. In some examples, multiple tuning maps may be obtained, each associated with an image processing function. For example, the one or more maps may include multiple maps, where a first map of the multiple maps is associated with a first image processing function and a second map of the multiple maps is associated with a second image processing function. In some cases, the first map may include a first plurality of values, where each value of the first plurality of values ​​of the first map indicates an amount of the first image processing function to be applied to a corresponding pixel of the image data. The second map may include a second plurality of values, where each value of the second plurality of values ​​of the second map indicates an amount of the second image processing function to be applied to a corresponding pixel of the image data. In some cases, a first image processing function associated with a first map of the plurality of maps includes one of a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function, and a second image processing function associated with a second map of the plurality of maps includes a different one of a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function. In one illustrative example, a noise map associated with the noise reduction function can be obtained, a sharpness map associated with the sharpness adjustment function can be obtained, a tone map associated with the tone adjustment function can be obtained, a saturation map associated with the saturation adjustment function can be obtained, and a hue map associated with the hue adjustment function can be obtained.

[0178] In some examples, the imaging system can generate one or more maps using a machine learning system, as in block 4254. The machine learning system may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or a combination thereof. The machine learning system and / or any of the above-listed machine learning elements (which may be part of the machine learning system) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or a combination thereof. The machine learning system that generates the one or more maps in block 4204 may be the same machine learning system as the machine learning system that generates the modified image in block 4206. The machine learning system that generates the one or more maps in block 4204 may be a different machine learning system than the machine learning system that generates the modified image in block 4206.

[0179] Examples of one or more maps include input saturation map 203, spatial tuning map 303, input noise map 404, input sharpening map 604, input tone map 804, input saturation map 1104, input hue map 1304, spatially varying tuning map 1404, small spatially varying tuning map 1424, output layer 1514, input tuning map 1607, output tuning map 1618, reference output tuning map 1619, spatially varying tuning map omega (Ω) of FIG. 17, tuning map 1804, tuning parameters 2206, tuning parameters 2806, tuning parameters 2616, tuning parameters 36, spatially varying map 3703, spatially varying map 3803, detail map 40B, noise map 40B, tone map 40B, saturation map 40B, hue map 40B, sharpness map 41B, noise map 41B, tone map 41B, saturation map 41B, hue map 41B, one or more maps in block 4254, other spatially varying maps described herein, other spatially fixed maps described herein, other maps described herein, other masks described herein, or any combination thereof. Examples of image processing functions include noise reduction, noise addition, sharpness adjustment, detail adjustment, tone adjustment, color saturation adjustment, hue adjustment, any other image processing function described herein, any other image processing parameter described herein, or any combination thereof. In some examples, block 4204 of process 4200 may correspond to block 4254 of process 4250.

[0180] At block 4206, process 4200 includes the imaging system generating a modified image using the image data and one or more maps as input to a machine learning system. The modified image may be referred to as an output image. The modified image includes characteristics based on respective image processing functions associated with each of the one or more maps. In some cases, the machine learning system includes at least one neural network. In some examples, process 4200 includes generating the one or more maps using an additional machine learning system that is different from the machine learning system used to generate the modified image. For example, the auto-tuning ML system 1705 of FIG. 17 can be used to generate the one or more maps, and the image processing ML system 210 can be used to generate the modified image based on the input image and the one or more maps.

[0181] Examples of rectified images are rectified image 215, rectified image 415, image 510, image 511, image 512, rectified image 615, output image 710, output image 711, and output image 712, rectified image 815, output image 1020, output image 1021, output image 1022, rectified image 1115, rectified image 1315, rectified image 1415, output layer 1514, output image 1608, reference output image 1609, output color image 2209, RGB output image 2308, final output image 2309, output RGB image 2409 in Figure 24. 25A-25B, the rectified image of FIGS. 26A-26B, the rectified image of FIGS. 27A-27C, output enhanced image 2808, output enhanced image 2828, the output patch of FIGS. 32-35, output image 3715, reference image 3716, output image 3815, reference image 3816, the automatically adjusted image of FIG. 39, the output image of FIGS. 40A-40B, the output image of FIG. 41A, the rectified image of block 4206, other image data described herein, other images described herein, or a combination thereof. In some examples, block 4206 of process 4200 may correspond to block 4256 of process 4250.

[0182] The machine learning system of block 4206 may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The machine learning system of block 4206 and / or any of the above-listed machine learning elements (which may be part of the machine learning system) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0183] 42B is a flowchart illustrating an example of a process 4250 for processing image data using one or more neural networks using the techniques described herein. Process 4250 may be performed by an imaging system, such as image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, image processing ML system 210, the system of FIG. 14A, the system of FIG. 14B, the system of FIG. 14C, the system of FIG. 14D, image processing system 1406, auto-tuning ML system 1405, downsampler 1418, upsampler 1426, neural network 1500, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, spatial attention engine 2110, channel The system may include an attention engine 2115, an image sensor 2202, a pre-conditioned ISP 2208, a machine learning (ML) ISP 2300, a neural network system 2303, a pre-processing engine 2307, an input interface 2301, an output interface 2305, a neural network architecture 2400, a trained machine learning model 2802, an imaging system that performs the process 4200, a computing system 4300, or a combination thereof.

[0184] At block 4252, process 4250 includes an imaging system acquiring image data. In some implementations, the image data includes a processed image having multiple color components for each pixel of the image data. For example, the image data may include one or more RGB images, one or more YUV images, or other color images previously captured and processed by a camera system, such as the image of scene 110 generated by image capture and processing system 100 shown in FIG. 1. Examples of image data include image data captured using image capture and processing system 100, image data captured using image capture device 105A and / or image processing device 105B, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, downsampled input image 1616, and the I image data of FIG. RGB ,Luminance channel I in Fig. 17 y, luminance channel 1820, image patch 1825, complete image input 1826, input image 2130, color filtered raw input data 2204, output color image 2209, frame of raw image data 2304, raw image patch 2306, RGB output 2308, final output image 2309, the raw input image of FIG. 24, the output RGB image of FIG. 24, the input image of FIGS. 25A-25B, the input image of FIGS. 26A-26B, the input image of FIGS. 27A-27C, image data 2804, image data 2814, the input patch of FIGS. 32-35, raw image 3702, reference image 3716, raw image 3802, reference image 3816, the input image of FIG. 39, the input image of FIGS. 40A-40B, the input image of FIGS. 41A-41B, the raw image data of block 4202, other image data described herein, other images described herein, or combinations thereof. The imaging system may include an image sensor that captures image data. The imaging system may include an image sensor connector coupled to the image sensor. Acquiring the image data in block 4252 may include acquiring the image data from the image sensor and / or through the image sensor connector.

[0185] In some examples, the image data includes an input image having multiple color components for each pixel of the multiple pixels of the image data. The input image may already be at least partially processed, for example, through demosaicing in ISP 154. The input image may be captured and / or processed using image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, image sensor 130, ISP 154, host processor 152, image sensor 2202, or a combination thereof. In some examples, the image data includes raw image data from one or more image sensors, the raw image data including at least one color component for each pixel of the multiple pixels of the image data. Examples of the one or more image sensors may include image sensor 130 and image sensor 2202. Examples of raw image data include image data captured using image capture and processing system 100, image data captured using image capture device 105A, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, downsampled input image 1616, and I of FIG. RGB ,Luminance channel I in Fig. 17 y, luminance channel 1820, image patch 1825, complete image input 1826, input image 2130, color filtered raw input data 2204, frame of raw image data 2304, raw image patch 2306, the raw input image of FIG. 24, the input image of FIGS. 25A-25B, the input image of FIGS. 26A-26B, the input image of FIGS. 27A-27C, image data 2804, image data 2814, the input patch of FIGS. 32-35, raw image 3702, reference image 3716, raw image 3802, reference image 3816, the input image of FIG. 39, the input image of FIGS. 40A-40B, the input image of FIGS. 41A-41B, the raw image data of block 4202, other raw image data described herein, other image data described herein, other images described herein, or combinations thereof. The image data may include a single color component (e.g., red, green, or blue) for each pixel of the image data. The image data may include multiple color components (e.g., red, green, and / or blue) for each pixel of the image data. In some cases, the image data is acquired from one or more image sensors filtered by a color filter array, such as a Bayer color filter array. In some examples, the image data includes one or more patches of image data. A patch of image data includes a subset of a frame of image data, for example, corresponding to one or more contiguous areas or regions. In some cases, generating the modified image includes generating multiple patches of output image data. Each patch of output image data may include a subset of pixels of the output image. In some examples, block 4252 of process 4250 may correspond to block 4202 of process 4200.

[0186] At block 4254, process 4250 includes the imaging system using the image data as input to one or more trained neural networks to generate one or more maps. Each map of the one or more maps is associated with a respective image processing function. In some examples, block 4254 of process 4250 may correspond to block 4204 of process 4200.Examples of one or more trained neural networks include image processing ML system 210, auto-tuning ML system 1405, neural network 1500, auto-tuning ML system 1705, local neural network 1806, global neural network 1807, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, neural network system 2303, neural network architecture 2400, machine learning ISP 2300, trained machine learning model 2802, one or more trained convolutional neural networks (CNNs), one or more trained neural networks (NNs), one or more CNNs, one or more trained neural networks (NNs), one or more neural network architectures. The neural network may include a plurality of NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, any other trained neural network described herein, any other neural network described herein, any other trained machine learning model described herein, any other machine learning model described herein, any other trained machine learning system described herein, any other machine learning system described herein, or any combination thereof.The one or more trained neural networks of block 4254 and / or any of the machine learning elements listed above (which may be part of the one or more trained neural networks of block 4254) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or a combination thereof.

[0187] Examples of one or more maps include input saturation map 203, spatial tuning map 303, input noise map 404, input sharpening map 604, input tone map 804, input saturation map 1104, input hue map 1304, spatially varying tuning map 1404, small spatially varying tuning map 1424, output layer 1514, input tuning map 1607, output tuning map 1618, reference output tuning map 1619, spatially varying tuning map omega (Ω) of FIG. 17, tuning map 1804, tuning parameters 2206, tuning parameters 2806, tuning parameters 2616, tuning parameters 36, spatially varying map 3703, spatially varying map 3803, detail map 40B, noise map 40B, tone map 40B, saturation map 40B, hue map 40B, sharpness map 41B, noise map 41B, tone map 41B, saturation map 41B, hue map 41B, one or more maps in block 4204, other spatially varying maps described herein, other spatially fixed maps described herein, other maps described herein, other masks described herein, or any combination thereof. Examples of image processing functions include noise reduction, noise addition, sharpness adjustment, detail adjustment, tone adjustment, color saturation adjustment, hue adjustment, any other image processing function described herein, any other image processing parameter described herein, or any combination thereof.

[0188] Each map of the one or more maps may include multiple values ​​and may be associated with an image processing function. Example values ​​may include values ​​V0 through V56 of spatial tuning map 303. Example values ​​may include input saturation map 203, input noise map 404, input sharpening map 604, input tone map 804, input saturation map 1104, input hue map 1304, spatially varying tuning map 1404, small spatially varying tuning map 1424, spatially varying map 3703, spatially varying map 3803, the detail map of FIG. 40B, the noise map of FIG. 40B, the tone map of FIG. 40B, the saturation map of FIG. 40B, the hue map of FIG. 40B, the sharpness map of FIG. 41B, the noise map of FIG. 41B, the tone map of FIG. 41B, the saturation map of FIG. 41B, the hue map of FIG. 41B, and one or more maps of block 4204. Each value of the map may indicate a strength to be used when applying the image processing function to the corresponding region of the image data. In some examples, a higher numerical value (e.g., 1) may indicate a stronger strength to be used when applying the image processing function to the corresponding region of the image data, while a lower numerical value (e.g., 0) may indicate a weaker strength to be used when applying the image processing function to the corresponding region of the image data. In some examples, a lower numerical value (e.g., 0) may indicate a stronger strength to be used when applying the image processing function to the corresponding region of the image data, while a higher numerical value (e.g., 1) may indicate a weaker strength to be used when applying the image processing function to the corresponding region of the image data. In some examples, a value (e.g., 0.5) may indicate applying the image processing function to the corresponding region of the image data with a strength of 0, which may result in no application of the image processing function and therefore no change to the image data using the image processing function. The corresponding region of the image data may correspond to pixels of the image data of block 4202 and / or the image of block 4256.The corresponding regions of image data may correspond to multiple pixels of the image data of block 4202 and / or the image of block 4256, for example, where the multiple pixels are binned for resampling, resizing, and / or rescaling, e.g., to form superpixels. The corresponding regions of image data may be contiguous regions. The multiple pixels may be in contiguous regions. The multiple pixels may be adjacent to each other.

[0189] Each value of the plurality of values ​​in the map may indicate an orientation to be used when applying the image processing function to the corresponding region of the image data. In some examples, a numerical value above a threshold (e.g., above 0.5) may indicate a positive orientation to be used when applying the image processing function to the corresponding region of the image data, while a numerical value below the threshold (e.g., below 0.5) may indicate a negative orientation to be used when applying the image processing function to the corresponding region of the image data. In some examples, a numerical value above the threshold (e.g., above 0.5) may indicate a negative orientation to be used when applying the image processing function to the corresponding region of the image data, while a numerical value below the threshold (e.g., below 0.5) may indicate a positive orientation to be used when applying the image processing function to the corresponding region of the image data. The corresponding region of the image data may correspond to a pixel of the image data of block 4202 and / or the image of block 4256. The corresponding regions of image data may correspond to multiple pixels of the image data of block 4202 and / or the image of block 4256, for example, where the multiple pixels are binned for resampling, resizing, and / or rescaling, e.g., to form superpixels. The corresponding regions of image data may be contiguous regions. The multiple pixels may be in contiguous regions. The multiple pixels may be adjacent to each other.

[0190] In some examples, a positive orientation for the saturation adjustment may result in over-saturation, while a negative orientation for the saturation adjustment may result in under-saturation or de-saturation. In some examples, a positive orientation for the tone adjustment may result in increased brightness and / or luminosity, while a negative orientation for the tone adjustment may result in decreased brightness and / or luminosity (or vice versa). In some examples, a positive orientation for the sharpness adjustment may result in increased sharpness (e.g., decreased blur), while a negative orientation for the sharpness adjustment may result in decreased sharpness (e.g., increased blur) (or vice versa). In some examples, a positive orientation for the noise adjustment may result in noise reduction, while a negative orientation for the tone adjustment may result in noise generation and / or addition (or vice versa). In some examples, positive and negative orientations for the hue adjustment may correspond to different angles theta (θ) of the hue-adjusted vector 1218, such as the conceptual diagram 1216 of FIG. 12B . For example, the positive and negative orientations for the hue adjustment may correspond to the positive and negative angles theta (θ) of the hue-adjusted vector 1218, such as the conceptual diagram 1216 of Figure 12B. In some examples, one or more of the image processing functions may be limited to processing in the positive orientation as discussed herein, or may be limited to processing in the negative orientation as discussed herein.

[0191] The one or more maps may include multiple maps. A first map of the multiple maps may be associated with a first image processing function, and a second map of the multiple maps may be associated with a second image processing function. The second image processing function may be different from the first image processing function. In some examples, the one or more image processing functions associated with at least one of the multiple maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, a hue adjustment function, or a combination thereof. In one illustrative example, a noise map associated with the noise reduction function may be obtained, a sharpness map associated with the sharpness adjustment function may be obtained, a tone map associated with the tone adjustment function may be obtained, a saturation map associated with the saturation adjustment function may be obtained, and a hue map associated with the hue adjustment function may be obtained.

[0192] The first map can include a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength and / or direction to use when applying the first image processing function to a corresponding region of the image data. The second map can include a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength and / or direction to use when applying the second image processing function to a corresponding region of the image data. The corresponding region of the image data may correspond to a pixel of the image data of block 4202 and / or the image of block 4256. The corresponding region of the image data may correspond to a plurality of pixels of the image data of block 4202 and / or the image of block 4256, for example, where the plurality of pixels are binned for resampling, resizing, and / or rescaling, for example, to form a superpixel. The corresponding region of the image data may be a contiguous region. The plurality of pixels may be in a contiguous region. The plurality of pixels may be adjacent to each other.

[0193] In some examples, the image data includes luminance channel data corresponding to the image. An example of luminance channel data is luminance channel I in FIG.y and / or luminance channel 1820 data. Using the image data as input to the one or more trained neural networks may include using a luminance channel corresponding to the image as input to the one or more trained neural networks. In some examples, generating an image based on the image data includes generating the image based on luminance channel data as well as color difference data corresponding to the image.

[0194] In some examples, the one or more trained neural networks output one or more affine coefficients based on using image data as input to the one or more trained neural networks. Examples of affine coefficients include affine coefficients a and b in FIG. 17. Examples of the one or more trained neural networks that output one or more affine coefficients include auto-tuning ML system 1705, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, or combinations thereof. In some examples, auto-tuning ML system 1405 can generate affine coefficients (e.g., a and / or b) and can generate spatially-varying tuning map 1404 and / or small spatially-varying tuning map 1424 based on the affine coefficients. Generating the one or more maps may include generating a first map by transforming the image data using at least one affine coefficient. The image data may include luminance channel data corresponding to the image, so that transforming the image data using one or more affine coefficients includes transforming the luminance channel data using one or more affine coefficients. For example, FIG. 17 illustrates the equation Ω=a*I y+ b to generate the tuning map Ω according to the luminance channel I using the affine coefficients a and b in Figure 17. y 17 illustrates transforming the luminance channel data from ∇Ω to ∇Ω. The one or more affine coefficients may include a multiplier, such as multiplier a in FIG. 17. Transforming the image data using the one or more affine coefficients may include multiplying luminance values ​​of at least a subset of the image data by the multiplier. The one or more affine coefficients may include an offset, such as offset a in FIG. 17. Transforming the image data using the one or more affine coefficients may include offsetting luminance values ​​of at least a subset of the image data by the offset. In some examples, the one or more trained neural networks output one or more affine coefficients that are also based on a local linearity constraint that aligns one or more gradients in the first map with one or more gradients in the image data. An example of a local linearity constraint may include a local linearity constraint 1720, which is represented by the formula ∇Ω=a*∇I y In some examples, applying one or more affine coefficients to at least a subset of the image data (e.g., to the luminance data) to transform the at least a subset of the image data (e.g., the luminance data) may be controlled by a local linearity constraint.

[0195] In some examples, one or more trained neural networks of block 4254 directly generate and / or output one or more maps in response to receiving image data as input to the one or more trained neural networks. For example, one or more trained neural networks may generate and / or output one or more maps directly based on an image without generating and / or outputting affine coefficients. Auto-tuning ML system 1405 may directly generate and / or output, for example, spatially-varying tuning map 1404 and / or small spatially-varying tuning map 1424. In some examples, various neural networks and / or machine learning systems shown and / or described herein as generating and using affine coefficients may instead be modified to directly output one or more maps without first generating affine coefficients. For example, auto-tuning ML system 1705 of FIG. 17 and / or FIG. 18 may be modified to directly generate and / or output tuning map 1804 as an output layer instead of, or in addition to, generating and / or outputting affine coefficients 1822. Neural network architecture 1900 may be modified to directly generate and / or output tuning map 1804 as an output layer instead of, or in addition to, generating and / or outputting affine coefficients 1822. Neural network architecture 2100A may be modified to directly generate and / or output tuning map 1804 as an output layer instead of, or in addition to, generating and / or outputting affine coefficients 1822. Neural network architecture 2100B may be modified to directly generate and / or output tuning map 1804 as an output layer instead of, or in addition to, generating and / or outputting affine coefficients 1822.

[0196] Each map of the one or more maps may be spatially varied. Each map of the one or more maps may be spatially varied based on different types or categories of objects depicted in different regions of the image data. Examples of such maps include input saturation map 203, spatially varying map 3703, and spatially varying map 3803, where regions of the image data depicting foreground flowers and regions of the image data depicting background map to different strengths and / or orientations to be used in applying one or more corresponding image processing functions. In other examples, a map may indicate that image processing functions may be applied with different strengths and / or orientations in regions depicting different types or categories of objects, such as people, faces, clothing, plants, sky, water, clouds, buildings, display screens, metal, plastic, concrete, brick, hair, wood, textured surfaces, any other type or category of object discussed herein, or combinations thereof. Each map of the one or more maps may be spatially varied based on different colors in different regions of the image data. Each map of the one or more maps may be spatially varied based on different image attributes in different regions of the image data. For example, input noise map 404 indicates that no denoising should be performed in areas of input image 402 that are already smooth, but that relatively strong denoising should be performed in areas of input image 402 that are noisy.

[0197] At block 4256, process 4250 includes the imaging system generating an image based on the image data and the one or more maps. The image includes characteristics based on the respective image processing functions associated with each map of the one or more maps. The image may be referred to as an output image, a rectified image, or a combination thereof. Examples of images include rectified image 215, rectified image 415, image 510, image 511, image 512, rectified image 615, output image 710, output image 711, and output image 712, rectified image 815, output image 1020, output image 1021, output image 1022, rectified image 1115, rectified image 1315, rectified image 1415, output image 1514, output image 1608, reference output image 1609, output color image 2209, RGB output 2308, final output image 2309, output RGB image of FIG. 24, and the output RGB image of FIG. 25. 5A-25B, the rectified image of Figures 26A-26B, the rectified image of Figures 27A-27C, output enhanced image 2808, output enhanced image 2828, the output patch of Figures 32-35, output image 3715, reference image 3716, output image 3815, reference image 3816, the automatically adjusted image of Figure 39, the output image of Figures 40A-40B, the output image of Figure 41A, the rectified image of block 4206, other image data described herein, other images described herein, or a combination thereof. In some examples, block 4256 of process 4250 may correspond to block 4206 of process 4200. The image may be generated using an image processor such as image capture and processing system 100, image processing device 105B, image processor 150, ISP 154, host processor 152, image processing ML system 210, image processing system 1406, neural network 1500, machine learning (ML) ISP 2300, neural network system 2303, pre-processing engine 2307, input interface 2301, output interface 2305, neural network architecture 2400, trained machine learning model 2802, the machine learning system of block 4206, or a combination thereof.

[0198] In some examples, generating an image based on the image data and the one or more maps includes using the image data and the one or more maps as inputs to a second set of one or more trained neural networks that are separate from the one or more trained neural networks. Examples of the second set of one or more trained neural networks may include image processing ML system 210, image processing system 1406, neural network 1500, machine learning (ML) ISP 2300, neural network system 2303, neural network architecture 2400, trained machine learning model 2802, the machine learning system of block 4206, or combinations thereof. In some examples, generating an image based on the image data and the one or more maps includes demosaicing the image data using the second set of one or more trained neural networks.

[0199] The second set of one or more trained neural networks may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The second set of one or more trained neural networks and / or any of the above-listed machine learning elements (which may be part of the second set of one or more trained neural networks) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof. In some examples, the second set of one or more trained neural networks may be the one or more trained neural networks of block 4254. In some examples, the second set of one or more trained neural networks may share at least one trained neural network with the one or more trained neural networks of block 4254. In some examples, the second set of one or more trained neural networks may be separate from the one or more trained neural networks of block 4254.

[0200] The imaging system may include a display screen. The imaging system may include a display screen connector coupled to the display screen. The imaging system can display images on the display screen, for example, by transmitting images to the display screen (and / or to a display controller) using the display screen connector. The imaging system may include a communication transceiver, which may be wired and / or wireless. The imaging system can transmit images to a recipient device using the communication transceiver. The recipient device may be any type of computing system 4300 or any component thereof. The recipient device may be an output device, such as a device with a display screen and / or a projector, capable of displaying images.

[0201] In some embodiments, the imaging system may include: a means for acquiring image data; a means for using the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps being associated with a respective image processing function; and a means for generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps. In some examples, the means for acquiring image data may include image sensor 130, image capture device 105A, image capture and processing system 100, image sensor 2202, or a combination thereof. In some examples, the means for generating the one or more maps may include image processing device 105B, image capture and processing system 100, image processor 150, ISP 154, host processor 152, image processing ML system 210, auto-tuning ML system 1405, neural network 1500, auto-tuning ML system 1705, local neural network 1806, global neural network 1807, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, neural network system 2303, neural network architecture 2400, machine learning ISP 2300, trained machine learning model 2802, or a combination thereof.In some examples, the means for generating the modified image may include image processing device 105B, image capture and processing system 100, image processor 150, ISP 154, host processor 152, image processing ML system 210, image processing system 1406, neural network 1500, machine learning (ML) ISP 2300, neural network system 2303, pre-processing engine 2307, input interface 2301, output interface 2305, neural network architecture 2400, trained machine learning model 2802, the machine learning system of block 4206, or a combination thereof.

[0202] In some examples, the processes described herein (e.g., process 4200, process 4250, and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 4200 and / or process 4250 may be performed by image processing ML system 210 of FIG. 2. In another example, process 4200 and / or process 4250 may be performed by a computing device with computing system 4300 shown in FIG. 43. For example, a computing device with computing system 4300 shown in FIG. 43 may include components of image processing ML system 210 and may perform the operations of FIG. 43.

[0203] The computing device may include any appropriate device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or autonomous vehicle computing device, a robotic device, a television, and / or any other computing device with the resource capabilities to perform the processes described herein, including process 4200 and / or process 4250. In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0204] Components of a computing device may be implemented with circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0205] Process 4200 and process 4250 are illustrated as logical flow diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement a process.

[0206] Additionally, process 4200, process 4250, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored in a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0207] Figure 43 is a diagram illustrating an example of a system for implementing some aspects of the present technology. Specifically, Figure 43 illustrates an example of a computing system 4300, which may be, for example, an internal computing system, a remote computing system, a camera, or any computing device comprising any of the components of the system that communicate with each other using a connection 4305. The connection 4305 may be a physical connection using a bus or a direct connection to a processor 4310, such as in a chipset architecture. The connection 4305 may also be a virtual connection, a networked connection, or a logical connection.

[0208] In some embodiments, computing system 4300 is a distributed system such that the functions described in this disclosure may be distributed across a data center, multiple data centers, a peer network, etc. In some embodiments, one or more of the described system components represent many components, each performing some or all of the functions for which the component is described. In some embodiments, a component may be a physical device or a virtual device.

[0209] The exemplary system 4300 includes at least one processing unit (CPU or processor) 4310 and connections 4305 coupling various system components to the processor 4310, including system memory 4315, such as read-only memory (ROM) 4320 and random access memory (RAM) 4325. The computing system 4300 may also include a cache 4312 of high-speed memory directly connected to, near, or integrated as part of the processor 4310.

[0210] Processor 4310 may include any general-purpose processor and hardware or software services, such as services 4332, 4334, and 4336 stored in storage device 4330, configured to control processor 4310, as well as special-purpose processors where software instructions are incorporated into the actual processor design. Processor 4310 may essentially be a completely self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0211] To enable user interaction, the computing system 4300 includes input devices 4345, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. The computing system 4300 can also include an output device 4335, which may be one or more of several output mechanisms. In some cases, a multimodal system may allow a user to provide multiple types of input / output for communicating with the computing system 4300. The computing system 4300 can include a communication interface 4340, which can generally govern and manage user input and system output.Communication interfaces include audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple® Lightning® ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, Radio Frequency Identification (RFID) wireless signal transmission, Near Field Communication (NFC) wireless signal transmission, Dedicated Short Range Communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, Wireless Local Area Network (WLAN) signal transmission, Visible Light Communication (VLC), and Worldwide Interoperability for Microwave Access The communication interface 4340 may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Maximum Access Point), infrared (IR) communications wireless signal transmissions, public switched telephone network (PSTN) signal transmissions, integrated services digital network (ISDN) signal transmissions, 3G / 4G / 5G / LTE cellular data network wireless signal transmissions, ad hoc network signal transmissions, radio wave signal transmissions, microwave signal transmissions, infrared signal transmissions, visible light signal transmissions, ultraviolet light signal transmissions, wireless signal transmissions along electromagnetic media, or any combination thereof. The communication interface 4340 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the location of the computing system 4300 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo GNSS.Because there are no constraints that apply to any particular hardware configuration, this basic feature may be easily replaced by improved hardware or firmware configurations as they are developed.

[0212] The storage device 4330 may be a non-volatile and / or non-transitory and / or computer readable memory device, and may be a hard disk, or may be a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid state memory, a compact disk read only memory (CD-ROM) optical disk, a rewritable compact disk (CD) optical disk, a digital video disk (DVD) optical disk, a Blu-ray disk (BDD) optical disk, a holographic optical disk, another optical media, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card, The memory card may also be a card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0213] The storage device(s) 4330 may include software services, servers, services, etc. that cause the system to perform functions when code defining the same is executed by the processor 4310. In some aspects, hardware services that perform particular functions may include software components stored on a computer-readable medium that interface with the necessary hardware components, such as the processor 4310, connections 4305, output devices 4335, etc., to perform the functions.

[0214] As used herein, the term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media that can store, contain, or transport instructions and / or data. Computer-readable media may also include non-transitory media, on which data may be stored and which do not include carrier waves and / or transitory electronic signals propagating wirelessly or via wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may store code and / or machine-executable instructions, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0215] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bitstreams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0216] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the present embodiments may be practiced without these specific details. For clarity of explanation, in some instances, the present technology may be presented as including individual functional blocks, including functional blocks including devices, device components, steps or routines in methods embodied in software, or a combination of hardware and software. Additional components other than those shown in the drawings and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0217] Also, particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0218] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored on or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to perform, or otherwise configure the general-purpose computer, special-purpose computer, or processing device to perform, a particular function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, etc.

[0219] Devices implementing processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., a computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in a peripheral device or add-in card. Such functionality may also be implemented on a circuit board, in different chips, or different processes running within a single device, as further examples.

[0220] The instructions, media for carrying such instructions, computational resources for executing the instructions, and other structures for supporting such computational resources are exemplary means for providing the functionality described in this disclosure.

[0221] In the foregoing description, aspects of the present application have been described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Accordingly, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the concepts of the present invention may be embodied or utilized in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the applications described above may be used individually or together. Moreover, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the present specification. Accordingly, the specification and drawings should be considered illustrative and not limiting. For illustrative purposes, methods have been described in a particular order. It should be understood that in alternative embodiments, methods may be performed in an order different from that described.

[0222] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of this description.

[0223] When a component is described as being "configured to" perform some operation, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operation, by programming programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operation, or any combination thereof.

[0224] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to another component via a wired or wireless connection, and / or other suitable communication interface).

[0225] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one element of the set or multiple elements of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" means A, B, or A and B. As another example, claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0226] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0227] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device handset, or an integrated circuit device with multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium such as a random access memory (RAM), e.g., a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a nonvolatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. The techniques may additionally or alternatively be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures, such as a propagated signal or wave, that may be accessed, read, and / or executed by a computer.

[0228] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor” as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or incorporated into a composite video encoder / decoder (codec).

[0229] Illustrative aspects of the present disclosure include the following.

[0230] Aspect 1. A method for processing image data, comprising: acquiring image data; acquiring one or more maps, each map of the one or more maps associated with a respective image processing function; and using the image data and the one or more maps as inputs to a machine learning system to generate a modified image, the modified image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0231] Aspect 2. The method of aspect 1, wherein a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, and each value of the plurality of values ​​in the map indicates an amount of the image processing function to be applied to a corresponding pixel of the image data.

[0232] Embodiment 3. The method of any of embodiments 1 to 2, wherein the one or more image processing functions associated with the one or more maps include at least one of a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0233] Embodiment 4. The method of any of embodiments 1 to 3, wherein the one or more maps include a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

[0234] Embodiment 5. The method of embodiment 4, wherein the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating an amount of a first image processing function to be applied to a corresponding pixel of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating an amount of a second image processing function to be applied to a corresponding pixel of the image data.

[0235] Embodiment 6. The method of any one of embodiments 4 or 5, wherein a first image processing function associated with a first map of the plurality of maps includes one of a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function, and a second image processing function associated with a second map of the plurality of maps includes a different one of a noise reduction function, a sharpness adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0236] Embodiment 7. The method of any one of embodiments 1 to 6, further comprising generating one or more maps using an additional machine learning system that is different from the machine learning system used to generate the modified image.

[0237] Embodiment 8. The method of any one of embodiments 1 to 7, wherein the image data includes a processed image having multiple color components for each pixel of the image data.

[0238] Embodiment 9. The method of any one of embodiments 1 to 7, wherein the image data includes raw image data from one or more image sensors, and the raw image data includes a single color component for each pixel of the image data.

[0239] Embodiment 10. The method of embodiment 9, wherein the raw image data is obtained from one or more image sensors filtered by a color filter array.

[0240] Embodiment 11. The method of embodiment 10, wherein the color filter array comprises a Bayer color filter array.

[0241] Embodiment 12. The method of any one of embodiments 1 to 11, wherein the image data comprises a patch of image data, and the patch of image data comprises a subset of the frame of image data.

[0242] Aspect 13. The method of aspect 12, wherein generating the modified image includes generating a plurality of patches of output image data, each patch of output image data including a subset of pixels of the output image.

[0243] Embodiment 14. The method of any one of embodiments 1 to 13, wherein the machine learning system includes at least one neural network.

[0244] Aspect 15. An apparatus comprising: a memory configured to store image data; and a processor implemented in circuitry and configured to perform operations according to any of aspects 1 to 14.

[0245] Embodiment 16. The device of embodiment 15, wherein the device is a camera.

[0246] Embodiment 17. The apparatus of embodiment 15, wherein the apparatus is a mobile device including a camera.

[0247] Embodiment 18: The device of any one of embodiments 15 to 17, further comprising a display configured to display one or more images.

[0248] Embodiment 19. The apparatus of any one of embodiments 15 to 18, further comprising a camera configured to capture one or more images.

[0249] Aspect 20. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any of aspects 1-14.

[0250] Embodiment 21. An apparatus comprising one or more means for performing the operations according to any of embodiments 1 to 14.

[0251] Aspect 22. An apparatus for processing image data, comprising means for performing the operations according to any of aspects 1 to 14.

[0252] Aspect 23. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operations of any of aspects 1-14.

[0253] Aspect 24. An apparatus for processing image data, comprising: a memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to: acquire the image data; use the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps associated with a respective image processing function; and generate an image based on the image data and the one or more maps, wherein the image includes characteristics based on the respective image processing function associated with each map of the one or more maps.

[0254] Embodiment 25. The device of embodiment 24, wherein a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, and each value of the plurality of values ​​in the map indicates a strength to be used when applying the image processing function to a corresponding region of the image data.

[0255] Embodiment 26. The apparatus of embodiment 25, wherein the corresponding regions of the image data correspond to pixels of the image.

[0256] Embodiment 27. The apparatus of any of embodiments 24 to 26, wherein the one or more maps include a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

[0257] Embodiment 28. The device of embodiment 27, wherein the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0258] Embodiment 29. The device of any one of embodiments 27 to 28, wherein the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used when applying a first image processing function to a corresponding region of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used when applying a second image processing function to a corresponding region of the image data.

[0259] Embodiment 30. The apparatus of any of embodiments 24 to 29, wherein the image data includes luminance channel data corresponding to an image, and wherein using the image data as input to the one or more trained neural networks includes using the luminance channel data corresponding to the image as input to the one or more trained neural networks.

[0260] Embodiment 31. The apparatus of embodiment 30, wherein generating the image based on the image data includes generating the image based on luminance channel data and color difference data corresponding to the image.

[0261] Embodiment 32. The apparatus of any of embodiments 30 to 31, wherein the one or more trained neural networks output one or more affine coefficients based on using the image data as input to the one or more trained neural networks, and generating the one or more maps includes at least generating a first map by transforming the image data using the one or more affine coefficients.

[0262] Embodiment 33. The apparatus of any of embodiments 30 to 32, wherein the image data includes luminance channel data corresponding to an image, and wherein transforming the image data using one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0263] Embodiment 34. The apparatus of any of embodiments 30 to 33, wherein the one or more affine coefficients include a multiplier, and wherein transforming the image data using the one or more affine coefficients includes multiplying luminance values ​​of at least a subset of the image data by the multiplier.

[0264] Embodiment 35. The apparatus of any of embodiments 30 to 34, wherein the one or more affine coefficients include an offset, and wherein transforming the image data using the one or more affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset.

[0265] Embodiment 36. The apparatus of any of embodiments 30 to 35, wherein the one or more trained neural networks output one or more affine coefficients that are also based on local linearity constraints that align one or more gradients in the first map with one or more gradients in the image data.

[0266] Embodiment 37. The apparatus of any of embodiments 24 to 36, wherein the one or more processors are configured to use the image data and the one or more maps as inputs to a second set of one or more trained neural networks that are separate from the one or more trained neural networks to generate an image based on the image data and the one or more maps.

[0267] Embodiment 38. The apparatus of embodiment 37, wherein the one or more processors are configured to demosaic the image data using a second set of one or more trained neural networks to generate an image based on the image data and the one or more maps.

[0268] Embodiment 39. The apparatus of any of embodiments 24 to 38, wherein each map of the one or more maps is spatially varied based on different types of objects depicted in the image data.

[0269] Embodiment 40. The apparatus of any of embodiments 24 to 39, wherein the image data includes an input image having a plurality of color components for each pixel of the plurality of pixels of the image data.

[0270] Embodiment 41. The device of any of embodiments 24 to 40, wherein the image data includes raw image data from one or more image sensors, the raw image data including at least one color component for each pixel of the plurality of pixels of the image data.

[0271] Embodiment 42. The apparatus of any of embodiments 24 to 41, further comprising an image sensor that captures image data, and wherein acquiring the image data includes acquiring the image data from the image sensor.

[0272] Embodiment 43. The apparatus of any of embodiments 24 to 42, further comprising a display screen, wherein the one or more processors are configured to display the image on the display screen.

[0273] Embodiment 44. The apparatus of any of embodiments 24 to 43, further comprising a communications transceiver, wherein the one or more processors are configured to transmit the image to a recipient device using the communications transceiver.

[0274] Aspect 45. A method for processing image data, comprising the steps of obtaining the image data; using the image data as input to one or more trained neural networks to generate one or more maps, each map of the one or more maps being associated with a respective image processing function; and generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0275] Aspect 46. The method of aspect 45, wherein a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, and each value of the plurality of values ​​in the map indicates a strength to be used when applying the image processing function to a corresponding region of the image data.

[0276] Embodiment 47. The method of embodiment 46, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0277] Embodiment 48. The method of any of embodiments 45 to 47, wherein the one or more maps include a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

[0278] Embodiment 49. The method of embodiment 48, wherein the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0279] Embodiment 50. The method of any of embodiments 48 to 49, wherein the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used when applying a first image processing function to a corresponding region of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used when applying a second image processing function to a corresponding region of the image data.

[0280] Aspect 51. The method of any of aspects 45 to 50, wherein the image data includes luminance channel data corresponding to an image, and wherein using the image data as input to the one or more trained neural networks includes using the luminance channel data corresponding to the image as input to the one or more trained neural networks.

[0281] Aspect 52. The method of aspect 51, wherein generating the image based on the image data includes generating the image based on luminance channel data and color difference data corresponding to the image.

[0282] Embodiment 53. The method of any of embodiments 45 to 52, wherein the one or more trained neural networks output one or more affine coefficients based on using image data as input to the one or more trained neural networks, and wherein generating one or more maps includes at least generating a first map by transforming the image data using the one or more affine coefficients.

[0283] Aspect 54. The method of aspect 53, wherein the image data includes luminance channel data corresponding to an image, and wherein transforming the image data using one or more affine coefficients includes transforming the luminance channel data using one or more affine coefficients.

[0284] Aspect 55. The method of any of aspects 53 to 54, wherein the one or more affine coefficients include a multiplier, and wherein transforming the image data using the one or more affine coefficients includes multiplying luminance values ​​of at least a subset of the image data by the multiplier.

[0285] Aspect 56. The method of any of aspects 53 to 55, wherein the one or more affine coefficients include an offset, and wherein transforming the image data using the one or more affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset.

[0286] Embodiment 57. The method of any of embodiments 53 to 56, wherein the one or more trained neural networks output one or more affine coefficients that are also based on local linearity constraints that align one or more gradients in the first map with one or more gradients in the image data.

[0287] Embodiment 58. The method of any of embodiments 45 to 57, wherein generating the image based on the image data and the one or more maps includes using the image data and the one or more maps as input to a second set of one or more trained neural networks that is separate from the one or more trained neural networks.

[0288] Embodiment 59. The method of embodiment 58, wherein the step of generating an image based on the image data and the one or more maps includes a step of demosaicing the image data using a second set of one or more trained neural networks.

[0289] Embodiment 60. The method of any of embodiments 45 to 59, wherein each map of the one or more maps is spatially varied based on different types of objects depicted in the image data.

[0290] Embodiment 61. The method of any of embodiments 45 to 60, wherein the image data includes an input image having a plurality of color components for each pixel of the plurality of pixels of the image data.

[0291] Embodiment 62. The method of any of embodiments 45 to 61, wherein the image data includes raw image data from one or more image sensors, and the raw image data includes at least one color component for each pixel of the plurality of pixels of the image data.

[0292] Embodiment 63. The method of any of embodiments 45 to 62, wherein acquiring the image data includes acquiring the image data from an image sensor.

[0293] Embodiment 64. The method of any of embodiments 45 to 63, further comprising displaying the image on a display screen.

[0294] Embodiment 65. The method of any of embodiments 45 to 64, further comprising transmitting the image to a recipient device using a communications transceiver.

[0295] Aspect 66. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any of aspects 45 to 64.

[0296] Embodiment 67. An apparatus comprising one or more means for performing the operations according to any of embodiments 45 to 64.

[0297] Embodiment 68. An apparatus for processing image data, comprising means for performing the operations according to any of embodiments 45 to 64.

[0298] Aspect 69. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operations of any of aspects 45 to 64.

[0299] Aspect 70. An apparatus for processing image data, comprising: a memory; and one or more processors coupled to the memory, wherein the one or more processors are configured to: acquire the image data; use the image data as input to a machine learning system to generate one or more maps, each map of the one or more maps associated with a respective image processing function; and generate an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0300] Embodiment 71. The apparatus of embodiment 70, wherein a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, and each value of the plurality of values ​​in the map indicates a strength to be used when applying the image processing function to a corresponding region of the image data.

[0301] Embodiment 72. The apparatus of embodiment 71, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0302] Embodiment 73. The apparatus of any of embodiments 70 to 72, wherein the one or more maps include a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

[0303] Embodiment 74. The device of embodiment 73, wherein the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0304] Embodiment 75. The device of any one of embodiments 73 to 74, wherein the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used in applying the first image processing function to a corresponding region of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used in applying the second image processing function to a corresponding region of the image data.

[0305] Embodiment 76. The apparatus of any of embodiments 70 to 75, wherein the image data includes luminance channel data corresponding to the image, and wherein using the image data as input to the machine learning system includes using the luminance channel data corresponding to the image as input to the machine learning system.

[0306] Embodiment 77. The apparatus of embodiment 76, wherein generating the image based on the image data includes generating the image based on luminance channel data and color difference data corresponding to the image.

[0307] Embodiment 78. The apparatus of any of embodiments 76 to 77, wherein the machine learning system outputs one or more affine coefficients based on using the image data as input to the machine learning system, and generating the one or more maps includes at least generating a first map by transforming the image data using the one or more affine coefficients.

[0308] Embodiment 79. The apparatus of any of embodiments 76 to 78, wherein the image data includes luminance channel data corresponding to an image, and wherein transforming the image data using one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0309] Embodiment 80. The apparatus of any of embodiments 76 to 79, wherein the one or more affine coefficients include a multiplier, and wherein transforming the image data using the one or more affine coefficients includes multiplying luminance values ​​of at least a subset of the image data by the multiplier.

[0310] Embodiment 81. The apparatus of any of embodiments 76 to 80, wherein the one or more affine coefficients include an offset, and wherein transforming the image data using the one or more affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset.

[0311] Embodiment 82. The apparatus of any of embodiments 76 to 81, wherein the machine learning system outputs one or more affine coefficients that are also based on local linearity constraints that align one or more gradients in the first map with one or more gradients in the image data.

[0312] Embodiment 83. The apparatus of any of embodiments 70 to 82, wherein the one or more processors are configured to use the image data and the one or more maps as inputs to a second set of machine learning systems, separate from the machine learning system, to generate an image based on the image data and the one or more maps.

[0313] Embodiment 84. The apparatus of embodiment 83, wherein the one or more processors are configured to demosaic the image data using a second set of machine learning systems to generate an image based on the image data and the one or more maps.

[0314] Embodiment 85. The apparatus of any of embodiments 70 to 84, wherein each map of the one or more maps is spatially varied based on different types of objects depicted in the image data.

[0315] Embodiment 86. The apparatus of any of embodiments 70 to 85, wherein the image data includes an input image having a plurality of color components for each pixel of the plurality of pixels of the image data.

[0316] Embodiment 87. The apparatus of any of embodiments 70 to 86, wherein the image data includes raw image data from one or more image sensors, and the raw image data includes at least one color component for each pixel of the plurality of pixels of the image data.

[0317] Embodiment 88. The apparatus of any of embodiments 70 to 87, further comprising an image sensor that captures image data, and wherein acquiring the image data includes acquiring the image data from the image sensor.

[0318] Embodiment 89. The apparatus of any of embodiments 70 to 88, further comprising a display screen, wherein the one or more processors are configured to display the image on the display screen.

[0319] Embodiment 90. The apparatus of any of embodiments 70 to 89, further comprising a communications transceiver, wherein the one or more processors are configured to transmit the image to a recipient device using the communications transceiver.

[0320] Aspect 91. A method for processing image data, comprising the steps of obtaining the image data; using the image data as input to a machine learning system to generate one or more maps, each map of the one or more maps being associated with a respective image processing function; and generating an image based on the image data and the one or more maps, the image including characteristics based on the respective image processing function associated with each map of the one or more maps.

[0321] Aspect 92. The method of aspect 91, wherein a map of the one or more maps includes a plurality of values ​​and is associated with an image processing function, and each value of the plurality of values ​​in the map indicates a strength to be used when applying the image processing function to a corresponding region of the image data.

[0322] Embodiment 93. The method of embodiment 92, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0323] Embodiment 94. The method of any of embodiments 91 to 93, wherein the one or more maps include a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

[0324] Embodiment 95. The method of embodiment 94, wherein the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

[0325] Embodiment 96. The method of any of embodiments 94 to 95, wherein the first map includes a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used when applying a first image processing function to a corresponding region of the image data, and the second map includes a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used when applying a second image processing function to a corresponding region of the image data.

[0326] Aspect 97. The method of any of aspects 91 to 96, wherein the image data includes luminance channel data corresponding to the image, and wherein using the image data as input to the machine learning system includes using the luminance channel data corresponding to the image as input to the machine learning system.

[0327] Aspect 98. The method of aspect 97, wherein generating the image based on the image data includes generating the image based on luminance channel data and color difference data corresponding to the image.

[0328] Aspect 99. The method of any of aspects 91 to 98, wherein the machine learning system outputs one or more affine coefficients based on using image data as input to the machine learning system, and wherein generating one or more maps includes at least generating a first map by transforming the image data using the one or more affine coefficients.

[0329] Aspect 100. The method of aspect 99, wherein the image data includes luminance channel data corresponding to an image, and wherein transforming the image data using one or more affine coefficients includes transforming the luminance channel data using one or more affine coefficients.

[0330] Aspect 101. The method of any of aspects 99 to 100, wherein the one or more affine coefficients include a multiplier, and wherein transforming the image data using the one or more affine coefficients includes multiplying luminance values ​​of at least a subset of the image data by the multiplier.

[0331] Aspect 102. The method of any of aspects 99 to 101, wherein the one or more affine coefficients include an offset, and wherein transforming the image data using the one or more affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset.

[0332] Embodiment 103. The method of any of embodiments 99 to 102, wherein the machine learning system outputs one or more affine coefficients that are also based on local linearity constraints that align one or more gradients in the first map with one or more gradients in the image data.

[0333] Embodiment 104. The method of any of embodiments 91 to 103, wherein generating an image based on the image data and the one or more maps includes using the image data and the one or more maps as input to a second set of machine learning systems separate from the machine learning system.

[0334] Aspect 105. The method of aspect 104, wherein generating an image based on the image data and the one or more maps includes demosaicing the image data using a second set of machine learning systems.

[0335] Embodiment 106. The method of any of embodiments 91 to 105, wherein each map of the one or more maps is spatially varied based on different types of objects depicted in the image data.

[0336] Embodiment 107. The method of any of embodiments 91 to 106, wherein the image data includes an input image having a plurality of color components for each pixel of the plurality of pixels of the image data.

[0337] Embodiment 108. The method of any of embodiments 91 to 107, wherein the image data includes raw image data from one or more image sensors, and the raw image data includes at least one color component for each pixel of the plurality of pixels of the image data.

[0338] Embodiment 109. The method of any of embodiments 91 to 108, wherein acquiring the image data includes acquiring the image data from an image sensor.

[0339] Embodiment 110. The method of any of embodiments 91 to 109, further comprising the step of displaying the image on a display screen.

[0340] Embodiment 111. The method of any of embodiments 91 to 110, further comprising transmitting the image to a recipient device using a communications transceiver.

[0341] Embodiment 112. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any of embodiments 91 to 111.

[0342] Embodiment 113. An apparatus comprising one or more means for performing the operations according to any of embodiments 91 to 111.

[0343] Embodiment 114. An apparatus for processing image data, comprising means for performing the operations according to any of embodiments 91 to 111.

[0344] Aspect 115. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operations of any of aspects 91 to 111. [Explanation of symbols]

[0345] 100 Image Capture and Processing System 105A Image Capture Device 105B Image Processing Device 110 scenes 115 Lens 120 Control Mechanism 125A Exposure Control Mechanism 125B Focus control mechanism 125C Zoom Control Mechanism 130 Image Sensor 140 RAM 145 ROM 150 Image Processor 152 Host Processor 154 ISP 156 I / O 160 I / O 201 Input image and spatial tuning map 202 input images 203 Input Saturation Map 210 Image Processing ML System 215 modified images 216 modified images 302 input images 303 Tuning Map 402 input images 404 Input Noise Map 405 Noise Map Key 415 Modified Images 510 images 511 images 512 images 602 input images 604 Input Sharpening Map 605 Sharpness Map Key 615 Modified Images 710 output image 711 output image 712 output images 715 Gamma Curve 802 input images 804 Input Tone Map 805 Tone Map Key 815 modified images 1020 output image 1021 output images 1022 output images 1102 input images 1104 Input Saturation Map 1105 Saturation Map Key 1115 Modified Images 1217 original UV vectors 1218 hue corrected vectors 1219 Center point 1302 input images 1304 Input Hue Map 1305 Hue Map Key 1315 modified images 1402 input images 1404 Spatially Varying Tuning Map 1405 Auto-tuning ML System 1406 Image Processing System 1415 modified images 1418 Down Sampler 1422 Downsampled input image 1424 Small spatially varying tuning maps 1426 Upsampler 1500 Neural Networks 1510 Input Layer 1512 hidden layer 1514 output layer 1516 nodes 1606 input images 1607 Input Tuning Map 1608 output images 1609 Reference Output Image 1612 Backpropagation Engine 1616 Downsampled input image 1618 Output Tuning Map 1619 Reference Output Tuning Map 1622 Backpropagation Engine 1705 Auto-tuning ML System 1720 Local Linearity Constraints 1804 Spatially Varying Tuning Map 1806 Local Neural Network 1807 Global Neural Network 1820 Y Channel 1822 Affine Coefficients 1825 image patches 1826 Complete Image Input 1920 keys 2010 Global Features 2020 Key 2110 Spatial Attention Engine 2120 Channel Attention Engine 2120 Key 2130 input images 2202 Image Sensor 2204 color filtered raw input data 2206 Tuning Parameters 2208 ISP 2209 Output color image 2300 Machine Learning ISP 2301 Input Interface 2302 Image Sensor 2303 Neural Network System 2304 frames 2305 Output Interface 2306 raw image patches 2307 Pre-processing Engine 2308 RGB output 2309 Final output image 2420 Key 2802 trained machine learning models 2804 Image data 2806 Tuning Parameters 2808 Enhanced Image 2814 Image data 2816 Tuning Parameters 2826 Tuning Parameters 2828 Enhanced Image 3702 Input image data 3703 Spatially Varying Maps 3715 output images 3716 Reference Image 3802 Input image data 3803 Spatially Varying Maps 3815 modified images 3816 Reference Image 4305 Connection 4310 processor 4312 Cache 4315 memory 4320 ROM 4325 RAM 4330 storage device 4332 Service 1 4334 Service 2 4335 output device 4336 Service 3 4340 communication interface 4345 Input Devices

Claims

1. 1. An apparatus for processing image data, comprising: Memory and at least one processor coupled to the memory, acquiring image data; using the image data as an input to at least one trained neural network to generate a plurality of affine coefficients based on the image data as an input to the at least one trained neural network; transforming the image data using the plurality of affine coefficients to generate at least one map, the at least one map associated with an image processing function; generating a modified image using the image data and the at least one map, the modified image including characteristics based on the image processing function; a processor configured to: An apparatus comprising:

2. 2. The apparatus of claim 1, wherein a map of the at least one map includes a plurality of values ​​and is associated with an image processing function, each value of the plurality of values ​​indicating a strength to be used in applying the image processing function to a corresponding region of the image data.

3. The apparatus of claim 2 , wherein the corresponding regions of the image data correspond to pixels of the image.

4. 2. The apparatus of claim 1, wherein the at least one map comprises a plurality of maps, a first map of the plurality of maps being associated with a first image processing function, and a second map of the plurality of maps being associated with a second image processing function.

5. 5. The apparatus of claim 4, wherein the one or more image processing functions associated with at least one of the plurality of maps include at least one of a noise reduction function, a sharpness adjustment function, a detail adjustment function, a tone adjustment function, a saturation adjustment function, and a hue adjustment function.

6. 5. The apparatus of claim 4, wherein the first map comprises a first plurality of values, each value of the first plurality of values ​​in the first map indicating a strength to be used in applying the first image processing function to a corresponding region of the image data, and the second map comprises a second plurality of values, each value of the second plurality of values ​​in the second map indicating a strength to be used in applying the second image processing function to a corresponding region of the image data.

7. 2. The apparatus of claim 1, wherein the image data includes luminance channel data corresponding to an image, and wherein using the image data as input to the at least one trained neural network includes using the luminance channel data corresponding to the image as input to the at least one trained neural network.

8. The apparatus of claim 7 , wherein generating the image based on the image data includes generating the image based on the luminance channel data and color difference data corresponding to the image.

9. 2. The apparatus of claim 1, wherein the image data includes luminance channel data corresponding to an image, and wherein transforming the image data using the plurality of affine coefficients includes transforming the luminance channel data using the plurality of affine coefficients.

10. 2. The apparatus of claim 1, wherein the plurality of affine coefficients comprises a multiplier, and wherein transforming the image data using the plurality of affine coefficients comprises multiplying luminance values ​​of at least a subset of the image data by the multiplier.

11. the plurality of affine coefficients include an offset, and transforming the image data using the plurality of affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset; and / or the at least one trained neural network is configured to output the plurality of affine coefficients that are also based on a local linearity constraint that aligns at least one gradient in the at least one map with at least one gradient in the image data; and / or the at least one processor is configured to use the image data and the at least one map as input to at least a second trained neural network separate from the at least one trained neural network to generate the image based on the image data and the at least one map, and the at least one processor is configured to demosaic the image data using the second set of trained neural networks to generate the image based on the image data and the at least one map.

10. The apparatus of claim 1.

12. each map of the at least one map is spatially varied based on different types of objects depicted in the image data; and / or the image data includes an input image having a plurality of color components for each pixel of a plurality of pixels of the image data; and / or the image data includes raw image data from one or more image sensors, the raw image data including at least one color component for each pixel of a plurality of pixels of the image data; and / or further comprising an image sensor that captures the image data, and obtaining the image data includes obtaining the image data from the image sensor; and / or further comprising a display screen, wherein the at least one processor is configured to display the image on the display screen; and / or further comprising a communications transceiver, wherein the at least one processor is configured to transmit the image to a recipient device using the communications transceiver.

10. The apparatus of claim 1.

13. 1. A method for processing image data, comprising: acquiring image data; using the image data as an input to at least one trained neural network to generate a plurality of affine coefficients based on using the image data as an input to the at least one trained neural network; transforming the image data using the plurality of affine coefficients to generate at least one map, the at least one map being associated with an image processing function; generating a modified image using the image data and the at least one map, the modified image including characteristics based on the image processing function; A method comprising:

14. a map of the at least one map includes a plurality of values ​​and is associated with an image processing function, each value of the plurality of values ​​indicating the strength with which the image processing function should be applied to a corresponding region of the image data; and / or the at least one map comprises a plurality of maps, a first map of the plurality of maps being associated with a first image processing function and a second map of the plurality of maps being associated with a second image processing function; and / or the image data includes luminance channel data corresponding to an image, and using the image data as input to the at least one trained neural network includes using the luminance channel data corresponding to the image as input to the at least one trained neural network; and / or the plurality of affine coefficients comprise multipliers, and transforming the image data using the plurality of affine coefficients comprises multiplying luminance values ​​of at least a subset of the image data by the multipliers; and / or the plurality of affine coefficients include an offset, and transforming the image data using the plurality of affine coefficients includes offsetting luminance values ​​of at least a subset of the image data by the offset; and / or generating the image based on the image data and the at least one map includes using the image data and the at least one map as inputs to at least a second trained neural network that is separate from the at least one trained neural network; and / or each map of the at least one map is spatially varied based on different types of objects depicted in the image data; and / or the image data includes luminance channel data corresponding to an image, and transforming the image data using the plurality of affine coefficients includes transforming the luminance channel data using the plurality of affine coefficients. The method of claim 13.

15. The apparatus of claim 1 , wherein at least one edge of the image data is stored in the at least one map.

Citation Information

Patent Citations

  • Image signal processor for processing images

    US20190108618A1

  • Image lighting methods and apparatuses, electronic devices, and storage media

    US20200143230A1