Machine learning based image adjustment

By generating image processing function strength maps through machine learning systems, the problem of time-consuming and costly manual parameter adjustment in traditional ISPs is solved, realizing the automation and personalization of image processing and improving image quality and efficiency.

CN115461780BActive Publication Date: 2026-02-10QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180031660.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-10
Filing Date
2021-05-13
Publication Date
2026-02-10
Estimated Expiration
2041-05-13

AI Technical Summary

Technical Problem

Traditional image signal processors (ISPs) require manual adjustment of a large number of parameters, which makes the adjustment process time-consuming and expensive, and they cannot dynamically adjust image processing parameters according to image content.

Method used

A machine learning system is used to generate an image processing function intensity map. The system generates the image processing function intensity map and dynamically adjusts the image processing parameters, including noise reduction, sharpening, and color saturation. The accuracy and efficiency of image processing are improved by using affine coefficients and local linearity constraints.

Benefits of technology

It automates and personalizes image processing, improves image quality, and reduces adjustment time and cost. The efficiency and effectiveness of image processing are superior to traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115461780B_ABST
    Figure CN115461780B_ABST
Patent Text Reader

Abstract

An imaging system can obtain image data, for example, from an image sensor. The imaging system can provide the image data as input data to a machine learning system, which can generate one or more maps based on the image data. Each map can identify an intensity to which a certain image processing function is to be applied to each pixel of the image data. Different maps can be generated for different image processing functions, for example, noise reduction, sharpening, or color saturation. The imaging system can generate a modified image based on the image data and the one or more maps, for example, by applying each of one or more image processing functions according to each of the one or more maps. The imaging system can provide the image data and the one or more maps to a second machine learning system to generate the modified image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to image processing, and more specifically to systems and techniques for performing machine learning-based image adjustments. Background Technology

[0002] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). A camera may include a processor, such as an image signal processor (ISP), which can receive and process one or more image frames. For example, raw image frames captured by the camera sensor can be processed by the ISP to generate a final image. Cameras can be configured with various image capture and image processing settings to alter the appearance of images. Some camera settings are determined and applied before or during photo capture, such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain. Other camera settings can configure post-processing of the photo, such as changing contrast, brightness, saturation, sharpness, levels, curves, or color.

[0003] Traditional image signal processors (ISPs) have individual discrete blocks that address various partitions of the image-based problem space. For example, a typical ISP has discrete function blocks, each applying a specific operation to the raw camera sensor data to create the final output image. Such function blocks can include blocks for de-mosaicing, noise reduction (denoising), color processing, tone mapping, and many other image processing functions. Each of these function blocks contains numerous manually tuned parameters, resulting in an ISP with a large number of manually tuned parameters (e.g., over 10,000), which must be readjusted according to each client's tuning preferences. This manual tuning of parameters is very time-consuming and expensive, and therefore is typically performed only once. Once tuned, traditional ISPs usually use the same parameters for every image. Summary of the Invention

[0004] In some examples, systems and techniques for performing machine learning-based image adjustments using one or more machine learning systems are described. An imaging system may, for example, acquire image data from an image sensor. The imaging system may feed the image data as input to a machine learning system, which may generate one or more maps based on the image data. Each map can identify the intensity of an image processing function applied to each pixel of the image data. Different maps may be generated for different image processing functions such as noise reduction, sharpening, or color saturation. The imaging system may generate a modified image based on the image data and one or more maps (e.g., by applying each of one or more image processing functions according to each of the one or more maps). The imaging system may also feed the image data and one or more maps to a second machine learning system to generate the modified image.

[0005] In one example, an apparatus for image processing is provided. The apparatus includes a memory and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured and capable of: acquiring image data; using the image data as input to one or more trained neural networks to generate one or more graphs, each of the one or more graphs being associated with a corresponding image processing function; and generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0006] In another example, an image processing method is provided. The method includes: acquiring image data; using the image data as input to one or more trained neural networks to generate one or more graphs, each of the one or more graphs being associated with a corresponding image processing function; and generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0007] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause one or more processors to: acquire image data; use the image data as input to one or more trained neural networks to generate one or more graphs, each of the one or more graphs being associated with a corresponding image processing function; and generate an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0008] In another example, an apparatus for image processing is provided. The apparatus includes: a unit for acquiring image data; a unit for generating one or more graphs using the image data as input to one or more trained neural networks, wherein each graph in the one or more graphs is associated with a corresponding image processing function; and a unit for generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each graph in the one or more graphs.

[0009] In some aspects, one or more graphs include multiple values ​​and are associated with an image processing function, each of the multiple values ​​of the graph indicating the intensity at which the image processing function is applied to a corresponding region of the image data. In some aspects, the corresponding region of the image data corresponds to a pixel of the image.

[0010] In some aspects, the one or more images include a plurality of images, wherein a first image of the plurality of images is associated with a first image processing function, and a second image of the plurality of images is associated with a second image processing function. In some aspects, the one or more image processing functions associated with at least one of the plurality of images include at least one of noise reduction, sharpness adjustment, detail adjustment, tone adjustment, saturation adjustment, and hue adjustment. In some aspects, the first image includes a first plurality of values, each of the first plurality of values ​​of the first image indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second image includes a second plurality of values, each of the second plurality of values ​​of the second image indicating the intensity of applying the second image processing function to a corresponding region of the image data.

[0011] In some aspects, the image data includes luminance channel data corresponding to the image, wherein using the image data as input to one or more trained neural networks includes using luminance channel data corresponding to the image as input to one or more trained neural networks. In some aspects, generating an image based on the image data includes generating an image based on the luminance channel data and chrominance data corresponding to the image.

[0012] In some aspects, one or more trained neural networks output one or more affine coefficients based on using image data as input to the one or more trained neural networks, wherein generating one or more maps includes generating a first map by transforming the image data using at least one or more affine coefficients. In some aspects, the image data includes luminance channel data corresponding to the image, wherein transforming the image data using one or more affine coefficients includes transforming the luminance channel data using one or more affine coefficients. In some aspects, the one or more affine coefficients include a multiplier, wherein transforming the image data using the one or more affine coefficients includes multiplying the luminance values ​​of at least a subset of the image data by the multiplier. In some aspects, the one or more affine coefficients include an offset, wherein transforming the image data using one or more affine coefficients includes offsetting the luminance values ​​of at least a subset of the image data by the offset. In some aspects, the one or more trained neural networks further output one or more affine coefficients based on local linearity constraints that align one or more gradients in the first map with one or more gradients in the image data.

[0013] In some aspects, generating images based on image data and one or more graphs involves using the image data and one or more graphs as input to a second set of one or more trained neural networks, distinct from the original set of trained neural networks. In other aspects, generating images based on image data and one or more graphs involves de-mosaicing the image data using the second set of one or more trained neural networks.

[0014] In some aspects, each of the one or more images varies spatially based on different types of objects depicted in the image data. In some aspects, the image data includes an input image having multiple color components for each of the multiple pixels of the image data. In some aspects, the image data includes raw image data from one or more image sensors, which includes at least one color component for each of the multiple pixels of the image data.

[0015] In some aspects, the above-described methods, apparatus, and computer-readable media further include: an image sensor that captures image data, wherein acquiring the image data includes acquiring image data from the image sensor. In some aspects, the above-described methods, apparatus, and computer-readable media further include: a display screen, wherein one or more processors are configured to display an image on the display screen. In some aspects, the above-described methods, apparatus, and computer-readable media further include: a communication transceiver, wherein one or more processors are configured to use the communication transceiver to transmit an image to a receiving device.

[0016] In some aspects, the device includes a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data.

[0017] This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. The subject matter should be understood by referring to appropriate portions of the entire specification, any or all of the drawings, and each claim.

[0018] The foregoing and other features and embodiments will become more apparent when referenced to the following description, claims and drawings. Attached Figure Description

[0019] The exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings:

[0020] Figure 1 This is a block diagram illustrating an example architecture of an image capture and processing system based on some examples;

[0021] Figure 2 This is a block diagram illustrating examples of systems including image processing machine learning (ML) systems, based on some examples;

[0022] Figure 3A This is a conceptual diagram illustrating an example of an input image, including multiple pixels labeled P0 to P63, based on some examples;

[0023] Figure 3B This is a conceptual diagram illustrating an example of a spatial tuning diagram based on some examples;

[0024] Figure 4 This is a conceptual diagram illustrating examples of input noise maps applied to adjust the noise reduction intensity in an input image to generate a modified image, based on some examples;

[0025] Figure 5 This is a conceptual diagram illustrating the effects of applying different intensities of sharpness adjustment image processing to example images based on several examples;

[0026] Figure 6 This is a conceptual diagram illustrating examples of how the sharpening intensity is applied to an input image to generate a modified image, based on several examples.

[0027] Figure 7 It is a conceptual diagram illustrating various examples of how different values ​​in example and tone maps of gamma curves, based on some examples, are converted into different intensities of tone adjustment applied to example images;

[0028] Figure 8 This is a conceptual diagram illustrating an example of how to adjust the intensity of a tone in an input image to generate a modified image, based on some examples.

[0029] Figure 9 This is a conceptual diagram illustrating saturation levels in a lightness-chromaticity (YUV) color space based on some examples;

[0030] Figure 10 This is a conceptual diagram showing processed variants of example images based on some examples, each variant being processed using different alpha (α) values ​​to adjust color saturation;

[0031] Figure 11 This is a conceptual diagram illustrating examples of how the intensity and direction of saturation are adjusted based on some examples applied to the input image to generate a modified image of the input saturation map;

[0032] Figure 12A This is a conceptual diagram illustrating the hue, saturation, and value (HSV) color space based on some examples;

[0033] Figure 12B This is a conceptual diagram illustrating modifications to the hue vectors in a lightness-chromaticity (YUV) color space based on some examples;

[0034] Figure 13 This is a conceptual diagram illustrating examples of input hue maps applied to the intensity of hue adjustments in an input image to generate a modified image, based on some examples.

[0035] Figure 14A This is a conceptual diagram illustrating an example system based on some examples, which includes an image processing system that receives an input image and multiple spatially varying tuning maps;

[0036] Figure 14B The diagram illustrates an example of a system based on several examples, including an image processing system that receives an input image and multiple spatially varying tuned maps, and an automatic tuning machine learning (ML) system that receives an input image and generates multiple spatially varying tuned maps.

[0037] Figure 14C This is a conceptual diagram illustrating an example system based on some examples, which includes an automatically adjusting machine learning (ML) system that receives an input image and generates multiple spatially varying tuned maps;

[0038] Figure 14D The following is a conceptual diagram illustrating an example system based on some examples, which includes an image processing system that receives an input image and multiple spatially varied tuned maps, and an automatic adjustment machine learning (ML) system that receives a downsized variant of the input image and generates a small spatially varied tuned map that is enlarged into multiple spatially varied tuned maps.

[0039] Figure 15 This is a block diagram illustrating examples of neural networks that can be used by image processing systems and / or automatically adapted machine learning (ML) systems, based on some examples;

[0040] Figure 16A This is a block diagram illustrating an example of training based on some example image processing systems (e.g., image processing ML systems);

[0041] Figure 16B This is a block diagram illustrating an example of automatically adjusting the training of a machine learning (ML) system based on some examples;

[0042] Figure 17 This is a block diagram illustrating an example of a system based on some examples, which includes an automatically adjusting machine learning (ML) system that generates a spatially varying tuning map from luminance channel data by generating affine coefficients that modify luminance channel data according to local linearity constraints.

[0043] Figure 18 This is a block diagram illustrating detailed examples of an automatically adjusting ML system based on some examples;

[0044] Figure 19 This is a block diagram illustrating an example of a neural network architecture for an automatically adjusting ML system's local neural network, based on some examples.

[0045] Figure 20 This is a block diagram illustrating an example of a neural network architecture for an automatically adjusting global neural network in an ML system, based on some examples.

[0046] Figure 21A This is a block diagram illustrating an example of an automatically adaptable neural network architecture for an ML system, based on several examples.

[0047] Figure 21B This is a block diagram illustrating another example of a neural network architecture for automatically adjusting an ML system based on some examples;

[0048] Figure 21C This is a block diagram illustrating an example neural network architecture for a spatial attention engine based on some examples;

[0049] Figure 21DThis is a block diagram illustrating an example neural network architecture for a channel attention engine based on some examples;

[0050] Figure 22 This is a block diagram illustrating an example of a pre-tuned image signal processor (ISP) according to some examples;

[0051] Figure 23 This is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) based on some examples;

[0052] Figure 24 This is a block diagram illustrating an example neural network architecture for a machine learning (ML) image signal processor (ISP) based on some examples;

[0053] Figure 25A This is a conceptual diagram illustrating an example of applying intensity adjustments to the first hue of an example input image to generate a modified image, based on some examples.

[0054] Figure 25B This shows the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a modified image with second-tone intensity adjustment.

[0055] Figure 26A This shows the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a modified image with first-level detail adjustment intensity.

[0056] Figure 26B This shows the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a second detail adjustment intensity for a modified image;

[0057] Figure 27A This shows the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a modified image with the first color saturation adjustment intensity.

[0058] Figure 27B This shows the application based on some examples. Figure 25A A conceptual diagram illustrating an example input image used to generate a modified image with a second color saturation adjustment intensity.

[0059] Figure 27C This shows the application based on some examples. Figure 25A A conceptual diagram illustrating an example of using a sample input image to generate a modified image with third-color saturation adjustment intensity;

[0060] Figure 28This is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) that receives various tuning parameters for tuning an MLISP as input, according to some examples.

[0061] Figure 29 This is a block diagram illustrating examples of specific tuning parameter values ​​that can be provided to a machine learning (ML) image signal processor (ISP) based on some examples;

[0062] Figure 30 This is a block diagram illustrating additional examples of specific tuning parameter values ​​that can be provided to a machine learning (ML) image signal processor (ISP) based on some examples;

[0063] Figure 31 This is a block diagram illustrating examples of objective functions and different losses that can be used during the training of a machine learning (ML) image signal processor (ISP);

[0064] Figure 32 This is a conceptual diagram illustrating examples of patch-by-patch model inference leading to non-overlapping output patches at the first image location, based on some examples.

[0065] Figure 33 This is shown at the second image position according to some examples. Figure 32 Conceptual diagram of an example of patch-by-patch model inference;

[0066] Figure 34 This is shown at the position of the third image according to some examples. Figure 32 Conceptual diagram of an example of patch-by-patch model inference;

[0067] Figure 35 This is shown at the fourth image position according to some examples. Figure 32 Conceptual diagram of an example of patch-by-patch model inference;

[0068] Figure 36 This is a conceptual diagram illustrating examples of spatially fixed tone maps applied at the image level, based on some examples;

[0069] Figure 37 This is a conceptual diagram illustrating an example application of spatial variation maps to process input image data to generate an output image with spatial variation saturation adjustment intensity;

[0070] Figure 38 This is a conceptual diagram illustrating an example application of spatial variation maps to process input image data to generate an output image with spatial variation tone adjustment intensity and spatial variation detail adjustment intensity.

[0071] Figure 39This is a conceptual diagram illustrating an automatically adjusted image generated by an image processing system using one or more tuning maps to adjust an input image, based on some examples.

[0072] Figure 40A This is a conceptual diagram illustrating the output image generated by an image processing system using one or more spatial variation tuning maps generated by an automatic adjustment machine learning (ML) system, based on some examples.

[0073] Figure 40B This shows examples that can be used based on... Figure 40A The input image shown generates Figure 40A A conceptual diagram illustrating an example of the spatial variation tuning plot of the output image;

[0074] Figure 41A This is a conceptual diagram illustrating the output image generated by an image processing system using one or more spatial variation tuning maps generated by an automatic adjustment machine learning (ML) system, based on some examples.

[0075] Figure 41B This shows examples that can be used based on... Figure 41A The input image shown generates Figure 41A A conceptual diagram illustrating an example of the spatial variation tuning plot of the output image;

[0076] Figure 42A This is a flowchart illustrating an example of a process for processing image data, based on some examples;

[0077] Figure 42B This is a flowchart illustrating an example of a process for processing image data, based on several examples; and

[0078] Figure 43 This is a diagram illustrating an example of a computing system used to implement some of the aspects described in this article. Detailed Implementation

[0079] Certain aspects and embodiments of this disclosure are provided below. It will be apparent to those skilled in the art that some of these aspects and embodiments can be applied independently and some can be combined. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.

[0080] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments will provide those skilled in the art with enabling descriptions for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0081] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms "image," "image frame," and "frame" are used interchangeably herein. Cameras can be configured with a variety of image capture and image processing settings. Different settings result in images with different appearances. Some camera settings are determined and applied before or during the capture of one or more image frames, such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain. For example, settings or parameters can be applied to the image sensor to capture one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changing contrast, brightness, saturation, sharpness, color levels, curves, or colors. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) to process one or more image frames captured by the image sensor.

[0082] A camera may include a processor, such as an ISP, which can receive and process one or more image frames from an image sensor. For example, raw image frames captured by the camera sensor can be processed by the ISP to generate a final image. In some examples, the ISP may process the image frames using multiple filters or processing blocks applied to them, such as demosaicing, gain adjustment, white balance adjustment, color balance or correction, gamma compression, tone mapping or adjustment, denoising or noise filtering, edge enhancement, contrast adjustment, intensity adjustment (e.g., darkening or brightening), etc. In some examples, the ISP may include a machine learning system (e.g., one or more trained neural networks, one or more trained machine learning models, one or more artificial intelligence algorithms, and / or one or more other machine learning components) that can process the image frames and output processed image frames.

[0083] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively, “Systems and Techniques”) are described for performing machine learning-based automatic image adjustment using one or more machine learning systems. An imaging system may acquire image data, for example, from an image sensor. The imaging system may provide the image data as input to a machine learning system. The machine learning system may generate one or more graphs based on the image data. Each of the one or more graphs may identify the intensity at which an image processing function is to be applied to each pixel of the image data. Different graphs may be generated for different image processing functions, such as noise reduction, noise addition, sharpening, desharpening, detail enhancement, detail reduction, tone adjustment, color desaturation, color saturation enhancement, or combinations thereof. The one or more graphs may be spatially varied such that different pixels of the image correspond to different intensities of the image processing function applied in each graph. The imaging system may generate a modified image based on the image data and one or more graphs, for example, by applying each of one or more image processing functions according to each of the one or more graphs. In some examples, the imaging system may provide the image data and one or more graphs to a second machine learning system to generate a modified image. The second machine learning system may be different from the machine learning system.

[0084] In some camera systems, the host processor (HP) (also sometimes called the application processor (AP)) is used to dynamically configure the image sensor with new parameter settings. The HP can be used to dynamically configure the parameter settings of the ISP pipeline to match the precise settings of the input image sensor frames, ensuring correct processing of image data. In some examples, the HP can be used to dynamically configure parameter settings based on one or more graphs. For example, parameter settings can be determined based on values ​​in a graph.

[0085] Generating and using graphs for image processing can offer various technical advantages to image processing systems. Spatially variable image processing graphs allow imaging systems to apply image processing functions with varying intensities to different regions of an image. For example, regarding one or more image processing functions, an image region depicting the sky can be processed differently from another region of the same image depicting grass. Graphs can be generated in parallel, thereby improving the efficiency of image processing.

[0086] In some examples, machine learning systems can generate one or more affine coefficients, such as multipliers and offsets, which imaging systems can use to modify components of the image data (e.g., the brightness channel of the image data) to generate each map. Local linearity constraints ensure that one or more gradients in the map are aligned with one or more gradients in the image data, which can reduce halo effects when applying image processing functions. Using affine coefficients and / or local linearity constraints to generate maps can produce higher quality spatially varied image modifications compared to systems that do not use affine coefficients and / or local linearity constraints, for example, due to better alignment between the image data and the map, and reduced halo effects at the boundaries of depicted objects.

[0087] Figure 1 This is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing scene images (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photographs) and / or can capture video including multiple images (or video frames) in a specific sequence. A lens 115 of the system 100 faces scene 110 and receives light from scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0088] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from image sensor 130 and / or image processor 150. One or more control mechanisms 120 may include multiple mechanisms and components; for example, control mechanism 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 may also include additional control mechanisms besides those shown, such as controls for analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0089] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can adjust the focus by moving the lens 115 closer to or away from the image sensor 130 via an actuating motor or servo. In some cases, the system 100 may include additional lenses, such as one or more microlenses on each photodiode of the image sensor 130, each microlens bending light received from the lens 115 toward the corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus setting can be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus setting may be referred to as the image capture setting and / or the image processing setting.

[0090] The exposure control mechanism 125A of the control mechanism 120 can obtain the exposure settings. In some cases, the exposure control mechanism 125A stores the exposure settings in a memory register. Based on these exposure settings, the exposure control mechanism 125A can control the aperture size (e.g., aperture size or f / stop), the duration of the aperture opening (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied to the image sensor 130, or any combination thereof. The exposure settings may be referred to as image capture settings and / or image processing settings.

[0091] The zoom control mechanism 125C of the control mechanism 120 can obtain zoom settings. In some examples, the zoom control mechanism 125C stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 125C can control the focal length of a lens element assembly (lens assembly) including lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a zoom lens. In some examples, the lens assembly may include a focusing lens (which in some cases may be lens 115) that first receives light from scene 110, then the light passes through a focusless zoom system between the focusing lens (e.g., lens 115) and image sensor 130, and then the light reaches image sensor 130. In some cases, a focusless zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference), with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control mechanism 125C moves one or more lenses in the focusless zoom system, such as a negative lens and one or two positive lenses.

[0092] Image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus light matching the color of the filter covering the photodiode can be measured. For example, Bayer color filters include red, blue, and green filters, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red filter, blue light data from at least one photodiode covered by the blue filter, and green light data from at least one photodiode covered by the green filter. Other types of color filters may be used in place of red, blue, and / or green filters or as an addition to red, blue, and / or green filters using yellow, magenta, and / or cyan (also known as "emerald") color filters. Some image sensors may not have color filters at all, but may instead use different photodiodes throughout the pixel array (in some cases, vertically stacked). Different photodiodes throughout the pixel array may have different spectral sensitivity profiles, and thus respond to light of different wavelengths. Monochrome image sensors may also lack color filters, thus lacking color depth.

[0093] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier to amplify the analog signal output from the photodiodes and / or include an analog-to-digital converter (ADC) to convert the analog signal output from the photodiodes (and / or the signal amplified by the analog gain amplifier) ​​into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms 120 may alternatively or additionally be included in image sensor 130. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide-semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0094] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other type of processor 5010 discussed with respect to computing device 5000. Host processor 152 may be a digital signal processor (DSP) and / or other types of processor. In some embodiments, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) including host processor 152 and ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). TMThis includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an Internal Integrated Circuit 2 (I2C) interface, an Internal Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a Serial General Purpose Input / Output (GPIO) interface, a Mobile Industrial Processor Interface (MIPI) (e.g., a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports). In an illustrative example, host processor 152 may communicate with image sensor 130 using an I2C port, and ISP 154 may communicate with image sensor 130 using a MIPI port.

[0095] Image processor 150 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, input reception, output management, memory management, or some combination thereof. Image processor 150 can store image frames and / or processed images in random access memory (RAM) 140 / 5020, read-only memory (ROM) 145 / 5025, cache, memory unit, another storage device, or some combination thereof.

[0096] Various input / output (I / O) devices 160 can be connected to the image processor 150. I / O devices 160 may include a display screen, keyboard, keypad, touchscreen, touchpad, touch-sensitive surface, printer, any other output device 5035, any other input device 5045, or some combination thereof. In some cases, text can be entered into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via a virtual keyboard or keypad on the touchscreen of the I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors that enable wired connections between the system 100 and one or more peripheral devices through which the system 100 can receive data from and / or send data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable wireless connections between the system 100 and one or more peripheral devices through which the system 100 can receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 160 previously discussed, and once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors, they can be considered I / O devices 160 in themselves.

[0097] In some cases, the image capture and processing system 100 may be a single device. In other cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 105A and the image processing device 105B may be wirelessly coupled together, for example, via one or more wires, cables, or other electrical connectors and / or via one or more wireless transceivers. In some embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0098] like Figure 1 As shown, the vertical dashed line will Figure 1 The image capture and processing system 100 is divided into two parts, namely image capture device 105A and image processing device 105B. Image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. Image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, certain components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included in image capture device 105A.

[0099] Image capture and processing system 100 may include electronic devices such as mobile or landline handsets (e.g., smartphones, cellular phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some embodiments, image capture device 105A and image processing device 105B may be different devices. For example, image capture device 105A may include a camera device, and image processing device 105B may include a computing device, such as a mobile handset, desktop computer, or other computing device.

[0100] Although the image capture and processing system 100 is shown to include certain components, those skilled in the art will understand that the image capture and processing system 100 may include more than [other components]. Figure 1 The components shown are further components. Components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, components of the image capture and processing system 100 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image capture and processing system 100.

[0101] Traditional camera systems (e.g., image sensors and ISPs) are tuned with parameters and then process the image based on those parameters. ISPs are typically tuned using a fixed tuning method during production. Camera systems (e.g., image sensors and ISPs) also typically perform global image adjustments based on predefined conditions such as illumination level, color temperature, exposure time, etc. Typical camera systems also use heuristic-based coarse-precision tuning (e.g., window-based local tone mapping). Therefore, traditional camera systems cannot enhance images based on what is contained within them.

[0102] This document describes systems, apparatuses, processes, and computer-readable media for performing machine learning-based image adjustment. Machine learning-based image adjustment techniques can be applied to processed images (e.g., output from an ISP and / or image post-processing system) and / or to raw image data from an image sensor. Machine learning-based image adjustment techniques can provide dynamic tuning for each image based on the scene content contained within it (instead of fixed tuning in traditional camera systems). Machine learning-based image adjustment techniques can also provide the ability to incorporate additional semantic context-based tuning (e.g., segmentation information) rather than simply providing heuristic-based tuning.

[0103] In some examples, one or more tuning maps may also be used to perform machine learning-based image adjustment techniques. In some examples, the tuning map may have the same resolution as the input image, where each value within the tuning map corresponds to a pixel in the input image. In some examples, the tuning map may be based on a downsampled variant of the input image and therefore may have a lower resolution than the input image, where each value within the tuning map corresponds to more than one pixel in the input image (e.g., four or more adjacent pixels in a square or rectangle if the downsampled variant of the input image has half the length and half the width of the input image). A tuning map may also be called a spatial tuning map or a spatially variable tuning map to refer to the ability of the values ​​in the tuning map to vary spatially for each location in the tuning map. In some examples, each value within the tuning map may correspond to a predetermined subset of the input image. In some examples, each value within the tuning map may correspond to the entire input image, in which case the tuning map may be called a spatially fixed tuning map.

[0104] A tuning map provides an image processing machine learning system with the ability to perform pixel-level adjustments to each image, allowing for high-precision adjustments to the image, not just global image adjustments. In some examples, the tuning map may be automatically generated by image processor 150, host processor 152, ISP 154, or a combination thereof. In some examples, the tuning map can be automatically generated using a machine learning system. In some examples, the tuning map can be generated with per-pixel precision using a machine learning system. For example, a first machine learning system can be used to generate a pixel-level tuning map, and a second machine learning system can process the input image and one or more tuning maps to generate a modified image (also referred to as an adjusted image or processed image). In some examples, the first machine learning system may be operated at least partially by and / or engaged with image processor 150, host processor 152, ISP 154, or a combination thereof. In some examples, the second machine learning system may be operated at least partially by and / or interfaced with image processor 150, host processor 152, ISP 154, or a combination thereof.

[0105] Figure 2 This is a block diagram illustrating an example of a system including an image processing machine learning (ML) system 210. The image processing ML system 210 receives one or more input images and one or more spatially tuned maps 201 as input. Example input image 202 is shown below. Figure 2 As shown. The image processing ML system 210 can process any type of image data. For example, in some examples, the input image provided to the image processing ML system 210 may include images that have already been processed by... Figure 1 Image capture device 105A image sensor 130 captures and / or is generated by ISP (e.g., Figure 1 The image processing ML system 210 can be implemented using one or more convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The image processing ML system 210 and / or any machine learning elements listed above can be trained using supervised learning (which may be part of the image processing ML system 210), unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0106] In some examples, the input image 202 may include multiple chromaticity components and / or color components (e.g., red (R) color components or samples, green (G) color components or samples, and blue (B) color components or samples) for each pixel of the image data. In some cases, the device may include multiple cameras, and the image processing ML system 210 may process images acquired by one or more of the multiple cameras. In an illustrative example, a dual-camera mobile phone, tablet, or other device may be used to capture larger images at a wider angle (e.g., with a wider field of view (FOV)), capture more light (resulting in sharper, clearer, and other advantages), generate 360-degree (e.g., virtual reality) video, and / or perform other enhanced functions (compared to those achieved by a single-camera device).

[0107] In some examples, the image processing ML system 210 can process raw image data provided by one or more image sensors. For example, the input image provided to the image processing ML system 210 may include data from an image sensor (e.g., Figure 1The image data is generated by the image sensor 130. In some cases, the image sensor may include an array of photodiodes capable of capturing frames of raw image data. Each photodiode may represent a pixel location and generate a pixel value for that pixel location. The raw image data from the photodiodes may include a single color or grayscale value for each pixel location in the frame. For example, a color filter array may be integrated with the image sensor or may be used in conjunction with the image sensor (e.g., placed on the photodiodes) to convert monochrome information into color values. An illustrative example of a color filter array includes a Bayer pattern color filter array (or Bayer color filter array) to allow the image sensor to capture pixel frames with a Bayer pattern, having one of a red, green, or blue color filter at each pixel location. In some cases, the device may include multiple image sensors, in which case the image processing ML system 210 can process the raw image data obtained from multiple image sensors. For example, a device with multiple cameras may use multiple cameras to capture image data, and the image processing ML system 210 can process the raw image data from multiple cameras.

[0108] Various types of spatial tuning maps can be provided as input to the image processing ML system 210. Each type of spatial tuning map can be associated with a corresponding image processing function. Spatial tuning maps can include, for example, noise tuning maps, sharpness tuning maps, hue tuning maps, saturation tuning maps, hue tuning maps, any combination thereof, and / or other tuning maps used for other image processing functions. For example, a noise tuning map can include values ​​corresponding to the intensity at which noise enhancement (e.g., noise reduction, noise addition) will be applied to different parts of the input image (e.g., different pixels) mapped to locations (e.g., coordinates) of different parts of the input image. A sharpness tuning map can include values ​​corresponding to the intensity at which sharpness enhancement (e.g., sharpening) will be applied to different parts of the input image (e.g., different pixels) mapped to locations (e.g., coordinates) of different parts of the input image. A hue tuning map can include values ​​corresponding to the intensity at which hue enhancement (e.g., hue mapping) will be applied to different parts of the input image (e.g., different pixels) mapped to locations (e.g., coordinates) of different parts of the input image. A saturation tuning map may include values ​​corresponding to intensity at which saturation enhancement (e.g., increasing or decreasing saturation) will be applied to different parts of the input image (e.g., different pixels) mapped to different parts of the input image at locations (e.g., coordinates). A hue tuning map may include values ​​corresponding to intensity at which hue enhancement (e.g., hue shift) will be applied to different parts of the input image (e.g., different pixels) mapped to different parts of the input image at locations (e.g., coordinates).

[0109] In some examples, the range of values ​​in the tuning graph can be from value 0 to value 1. This range can be inclusive of both ends, and therefore include 0 and / or 1 as possible values. The range can also be exclusive of both ends, and therefore exclude 0 and / or 1 as possible values. For example, each position in the tuning graph can include any value between 0 and 1. In some examples, one or more tuning graphs can include values ​​other than those between 0 and 1. Various tuning graphs are described in more detail below. Input saturation graph 203 in Figure 2 The image shown is an example of a spatial tuning diagram.

[0110] Although all figures in this document are shown in black and white, input image 202 represents a color image depicting pink and red flowers in the foreground, green leaves as part of the background, and a white wall as another part of the background. Input saturation image 203 includes a value of 0.5 corresponding to a pixel region in input image 202 depicting the pink and red flowers in the foreground. This region with a value of 0.5 is shown in gray in input saturation image 203. A value of 0.5 in input saturation image 203 indicates that the saturation in this region will remain unchanged, without increase or decrease. Input saturation image 203 includes a value of 0.0 corresponding to a pixel region in input image 202 depicting both the green leaves and the white wall. This region with a value of 0.0 is shown in black in input saturation image 203. A value of 0.0 in input saturation image 203 indicates that the saturation in this region will be completely desaturated. The modified image 216 shows that the flowers in the foreground are still saturated (still pink and red) to the same extent as in the input image 202, but the background (both the green leaves and the white wall) is completely desaturated and therefore depicted in grayscale.

[0111] Each tuning map (e.g., input saturation map 203) may have the same size and / or the same resolution as one or more input images (e.g., input image 202) provided as input to the image processing ML system 210 to be processed to generate a modified image (e.g., modified image 216). Each tuning map (e.g., input saturation map 203) may have the same size and / or the same resolution as the modified image (e.g., modified image 216) generated based on the tuning map. In some examples, the tuning map (e.g., input saturation map 203) may be smaller or larger than the input image (e.g., input image 202), in which case the tuning map or the input image (or both) may be downsized or upsized and / or upsampled before generating the modified image 216.

[0112] In some examples, tuning map 303 can be stored as an image, such as a grayscale image. In tuning map 303, the gray shading at a pixel can indicate the value of that pixel. In illustrative examples, black can encode a value of 0.0, medium gray or neutral gray can encode a value of 0.5, and white can encode a value of 1.0. Other grayscale shading can encode values ​​between those mentioned above. For example, light gray can encode values ​​between 0.5 and 1.0 (e.g., 0.6, 0.7, 0.8, 0.9). Dark gray can encode values ​​between 0.5 and 0.0 (e.g., 0.1, 0.2, 0.3, 0.4). In some examples, the opposite scheme can be used, where white encodes a value of 0.0 and black encodes a value of 1.0.

[0113] Figure 3A This is a conceptual diagram illustrating an example of an input image 302 comprising multiple pixels labeled P0 to P63, based on some examples. The input image is 7 pixels wide and 7 pixels high. The pixels are numbered from left to right in each row in the order of P0 to P63, starting from the top row and counting upwards towards the bottom row.

[0114] Figure 3B This is a conceptual diagram illustrating an example of a spatial tuning diagram 303 based on some examples. Tuning diagram 303 includes multiple values ​​labeled V0 to V63. The tuning diagram is 7 pixels wide and 7 pixels high. Pixels are numbered sequentially from left to right from V0 to V63 within each row, starting from the top row and counting upwards towards the bottom row.

[0115] Each value within the tuning graph 303 corresponds to a pixel in the input image 302. For example, the value V0 in the tuning graph 303 corresponds to pixel P0 in the input image 302. The values ​​in the tuning graph 303 are used to adjust or modify their corresponding pixels in the input image 302. In some examples, each value in the tuning graph 303 indicates the intensity and / or direction of the image processing function applied to the corresponding pixel of the image data. In some examples, each value in the tuning graph 303 indicates the amount of image processing function to be applied to the corresponding pixel. For example, a first value of V0 in the tuning graph 303 (e.g., value 0, value 1, or other value) may indicate that the image processing function of the tuning graph will be applied with zero intensity (it will not be applied at all) to the corresponding pixel P0 in the input image 302. In another example, a second value of V15 in the tuning graph 303 (e.g., value 0, value 1, or other value) indicates that the image processing function is applied to the corresponding pixel P15 in the input image 302 with maximum intensity (the maximum amount of image processing function). Values ​​in different types of tuning graphs may represent different levels of applicability of the corresponding image processing function. In an illustrative example, tuning map 303 can be a saturation map. A value of 0 in the saturation map can represent complete desaturation of the corresponding pixel, causing the pixel to be grayscaled or turned into monochrome (maximum intensity in the desaturation or negative saturation direction). A value of 0.5 in the saturation map can represent no saturation or desaturation effect applied to the corresponding pixel (saturation intensity is zero), and a value of 1 in the saturation map can represent maximum saturation (maximum intensity in the saturation or positive saturation direction). Values ​​between 0 and 1 in the saturation map will indicate different saturation levels between the above levels. For example, a value of 0.2 represents slight desaturation (low intensity in the desaturation or negative saturation direction), while a value of 0.8 represents slight saturation increase (low intensity in the saturation or positive saturation direction).

[0116] Image processing ML system 210 processes one or more input images and one or more spatial tuning maps 201 to generate a modified image 215. Example modified image 216 is shown. Figure 2 As shown in the figure. The modified image 216 is a modified version of the input image 202. As described above, the saturation map 203 includes a value of 0.5 for the regions in the saturation map 203 corresponding to the pixels depicting the pink and red flowers shown in the foreground of the input image 202. The saturation map 203 includes a value of 0 for the locations in image 202 corresponding to the background (which does not depict the flowers in the foreground of image 202). The 0 values ​​in the saturation map 203 indicate that the pixels corresponding to these values ​​will be completely desaturated (desaturation or maximum intensity in the negative saturation direction). As a result, in the modified image 215, all pixels except those representing flowers are depicted in grayscale (black and white). The value of 0.5 in the saturation map 203 indicates that the pixels corresponding to these values ​​will have no saturation change (saturation intensity is zero).

[0117] Image processing ML system 210 may include one or more neural networks trained to process one or more input images and one or more spatially tuned maps 201 to generate a modified image 215. In an illustrative example, supervised learning techniques may be used to train image processing ML system 210. For example, a backpropagation training process may be used to adjust the weights (and in some cases other parameters, such as biases) of the nodes of the neural network of image processing ML system 210. Backpropagation includes forward pass, loss function, backpropagation, and weight update. Forward pass, loss function, backpropagation, and parameter update are performed in one training iteration. The process is repeated a certain number of times for each set of training data until the weights of the neural network parameters are accurately tuned. Further details regarding image processing ML system 210 are provided in this paper.

[0118] As described above, various types of spatial tuning maps can be provided for use by the image processing ML system 210. An example of a spatial tuning map is a noise map that indicates the amount of denoising (noise reduction) to be applied to the pixels of an input image.

[0119] Figure 4 This is a conceptual diagram illustrating an example of an input noise map 404, which is applied to adjust the noise reduction intensity in an input image 402 to generate a modified image 415. The lower left portion of the input image 402 is marked in parentheses as "smooth" and has almost no visible noise. The remaining portion of the input image 402 is marked in parentheses as "noisy" and includes visible noise (with a grainy appearance).

[0120] Figure 4 The noise map key 405 is shown. Noise map key 405 indicates that a value of 0 in noise map 404 means no denoising will be applied, and a value of 1 means maximum denoising will be applied. As shown, the input noise map 404 includes a value of 0 for the black areas of noise map 404, which correspond to the "smooth" areas (pixels in the lower left region of image 402) with few or no noise. The input noise map 404 includes a value of 0.8 for the remaining light gray shaded areas, which correspond to the "noisy" areas (pixels in the lower left region of image 402) with noise. The image processing ML system 210 can process the input image 402 and the input noise map 404 to produce a modified image 415. As shown, based on the denoising performed by the image processing ML system 210 on the pixels in the input image 402 corresponding to the positions in noise map 404 with a value of 0.8, the modified image 415 does not contain noise.

[0121] Another example of a spatial tuning map is a sharpness map, which indicates the amount of sharpening to be applied to the pixels of an input image. For example, a blurred image can be focused by sharpening the pixels of the image.

[0122] Figure 5 This is a conceptual diagram illustrating the effects of applying different intensities of sharpness tuning to example images based on several examples. In some examples, a value of 0 in the sharpness map indicates that no sharpening will be applied to the corresponding pixels in the input image to generate the corresponding output image, and a value of 1 in the sharpness map indicates that maximum sharpening will be applied to the corresponding pixels in the input image to generate the corresponding output image. Image 510 was generated by the image processing ML system 210 with the sharpness tuning map value set to 0, where no sharpening was applied to image 510 (or a sharpening intensity of zero was applied). Image 511 was generated by the image processing ML system 210 with the sharpness map value set to 0.4, in which case a moderate amount of sharpening (sharpening at a medium intensity) was applied to the pixels of image 511. Image 512 was generated with the sharpness map value set to 1.0, where maximum sharpening (sharpening at maximum intensity) was applied to the pixels of image 512.

[0123] Image processing ML system 210 can generate an enhanced image (or sharpened image) based on an input image and a sharpness map. For example, image processing ML system 210 can use the sharpness map to sharpen the input image and produce a sharpened image (also called an enhanced image). The sharpened (or enhanced) image depends on the image detail and the alpha (α) parameter. For example, the following equation can be used to generate a sharpened (or enhanced) image: Enhanced Image = Original Input Image + α * Detail. The values ​​in the sharpness map allow image processing ML system 210 to modify both the alpha (α) and image detail parameters to obtain a sharpened image. As an illustrative example of generating a sharpened (or enhanced) image, image processing ML system 210 can apply an edge-preserving filter to the input image, for example, by applying an edge-preserving nonlinear filter. Examples of edge-preserving nonlinear filters include bilateral filters. In some cases, bilateral filters have two hyperparameters, including spatial sigma and range sigma. Spatial sigma controls the filter window size (where a larger window size results in smoother filtering). The range sigma controls the filter size along the intensity dimension (where larger values ​​cause different intensities to blur together). In some examples, image detail can be obtained by subtracting the filtered image (a smoothed image obtained from edge-preserving filtering) from the input image. A sharpened (or enhanced) image can then be obtained by adding a fraction of the image detail (represented by the alpha (α) parameter) to the original input image. Using image detail and alpha (α), a sharpened (enhanced) image can be obtained as follows: Enhanced Image = Original Input Image + α * Detail.

[0124] Figure 6This is a conceptual diagram illustrating an example of an input sharpness map 604 applied to adjust the sharpening intensity in an input image 602 to generate a modified image 615, based on some examples. It can be seen that the input image 602 has a blurred appearance. The sharpness map key 605 specifies that a value of 0 in the noise map 604 indicates that no sharpening will be applied to the corresponding pixels of the input image 602 (sharpening will be applied with zero intensity), and a value of 1 indicates that maximum sharpening will be applied to the corresponding pixels (sharpening will be applied with maximum intensity). As shown, the input sharpness map 604 includes a value of 0 for the black shaded area in the lower left portion of the sharpness map 604. The input sharpness map 604 includes a value of 0.9 for the remaining area (light gray shade) in the lower left portion of the sharpness map 604. The image processing ML system 210 can process the input image 602 and the input sharpness map 604 to produce the modified image 615. As shown in the figure, the pixels in the modified image 615 corresponding to the positions with a value of 0 in the sharpness map 604 are blurred (because sharpening was applied with zero intensity), while the pixels in the modified image 615 corresponding to the positions with a value of 0.9 in the sharpness map 604 are sharpened (not blurred) because the image processing ML system 210 performs sharpening at a high intensity in that area.

[0125] A tone map is another example of a spatial tone map. A tone map represents the amount of brightness adjustment applied to pixels in an input image, resulting in an image with a darker or brighter tone. The gamma value controls the amount of brightness of the pixels; this is called gamma correction. For example, gamma correction can be defined as... The transformation of I, where I in γ is the input value (e.g., a pixel or the entire image), γ is the gamma value, and I out This refers to the output value (e.g., a single pixel or the entire output image). The gamma transform function can be applied in different ways to different types of images. For example, for the red-green-blue (RGB) color space, the same gamma transform can be applied to the R, G, and B channels of an image. In another example, for the YUV color space (which has one luminance component (Y) and two chrominance components U (blue projection) and V (red projection)), the gamma transform function can be applied only to the luminance (Y) channel of the image (e.g., only to the Y component of the image pixels or samples).

[0126] Figure 7 This is a conceptual diagram illustrating various examples of gamma curves 715 and the effects of different values ​​in a tone map applied to different intensities of the example image, based on some examples. In some examples, the gamma (γ) range is 10. -0.4 Up to 10 0.4This corresponds to values ​​from 0.4 to 2.5. The gamma index -0.4 to 0.4 can be mapped to a range from 0 to 1 to calculate the tone tone map. For example, a value of 0 in the tone map indicates that a gamma (γ) value of 0.4 (corresponding to 10) will be applied to the corresponding pixel in the input image. -0.4 In the tone map, a value of 0.5 indicates that a gamma (γ) value of 1.0 (corresponding to 10) will be applied to the corresponding pixel in the input image. 0.0 The value 1 in the tone map indicates that a gamma (γ) value of 2.5 (corresponding to 10) is applied to the corresponding pixel in the input image. 0.4 Output images 710, 711, and 712 are all generated by applying different intensities of hue adjustment to the same input image. When a gamma value of 0.4 is applied, image 710, generated by the image processing ML system 210, results in increased brightness and is therefore brighter than the input image. When a gamma value of 1.0 is applied, image 711 is generated, resulting in no change in brightness. Therefore, output image 711 is hue-wise identical to the input image. When a gamma value of 2.5 is applied, image 712 is generated, resulting in a darker image that is less bright than the input image.

[0127] Figure 8 This is a conceptual diagram illustrating an example of an input tone map 804 for generating a modified image 815, based on some examples of tone adjustment intensities applied to an input image 802. The tone map key 805 specifies that a value of 0 in tone map 804 indicates that a gamma (γ) value of 0.4 (a tone adjustment with an intensity of 0.4) will be applied to the corresponding pixel in input image 802, a value of 0.5 in tone map 804 indicates that a gamma (γ) value of 1.0 (a tone adjustment with an intensity of 1.0, which indicates no change in tone) will be applied to the corresponding pixel in input image 802, and a value of 1 in tone map 804 indicates that a gamma (γ) value of 2.5 (a tone adjustment with a maximum intensity of 2.5) will be applied to the corresponding pixel in input image 802. The input tone map 804 includes a value of 0.12 for a set of locations (dark gray shaded areas) in the left half of tone map 804. The value of 0.12 corresponds to an approximate gamma (γ) value of 0.497. For example, the exponent can be determined as exponent = ((0.12-0.5)*0.4) / 0.5 = -0.304, and the gamma (γ) value can be determined as γ = 10. 指数 =10 (-0.304) =0.497. For the remaining set of locations (the area of ​​medium gray shading) in the right half of tone map 804, the value 0.5 is included. The value 0.5 corresponds to a gamma (γ) value of 1.0.

[0128] Image processing ML system 210 can process input image 802 and input tone map 804 to generate modified image 815. As shown, due to the use of a gamma (γ) value of 0.497 in the gamma transform based on tone map value of 0.12, pixels in the left half of the modified image 815 have a lighter tone (relative to the corresponding pixels in the left half of the input image 802 and / or relative to the pixels in the right half of the modified image 815). Due to the use of a gamma (γ) value of 1.0 in the gamma transform based on tone map value of 0.5, pixels in the right half of the modified image 815 have the same tone as the corresponding pixels in the right half of the input image 802 (and / or a darker tone relative to the pixels in the left half of the modified image 815).

[0129] A saturation map is another example of a spatial tuning map. A saturation map indicates the amount of saturation applied to pixels in an input image, resulting in an image with modified color values. For a saturation map, saturation refers to the difference between an RGB pixel (or a pixel defined using another color space, such as YUV) and its corresponding grayscale image (e.g., as shown in Equations 1-3 below). In some cases, this saturation can be referred to as a "saturation difference." For example, increasing the saturation of an image may cause colors in the image to become more intense, while decreasing the saturation may cause colors to become paler. If the saturation is reduced enough, the image may become desaturated, potentially resulting in a grayscale image. As described below, alpha mixing techniques can be performed in certain situations, where the alpha (α) value determines the saturation of a color.

[0130] Different techniques can be used to apply saturation adjustments. In the YUV color space, the values ​​of the U chromaticity component (blue projection) and V chromaticity component (red projection) of image pixels can be adjusted to set different saturation levels for the image.

[0131] Figure 9 This is a conceptual diagram (915) illustrating the saturation levels in the YUV color space, where the x-axis represents the U chromaticity component (blue projection) and the y-axis represents the V chromaticity component (red projection). Luminance (Y) in... Figure 9 The value is constant at 128 in the graph. Desaturation is represented at the center of the graph, with values ​​of U = 128 and V = 128, where the image is a grayscale (desaturated) image. As the distance increases (higher or lower U and / or V values), the saturation becomes higher. For example, the greater the distance from the center, the higher the saturation.

[0132] In another example, alpha blending can be used to adjust saturation. The mapping between luminance (Y) and the R, G, B channels can be expressed using the following equation:

[0133] R'=α*R+(1-α)*Y Equation (1)

[0134] G'=α*G+(1-α)*Y Equation (2)

[0135] B'=α*B+(1-α)*Y Equation (3)

[0136] Figure 10 This is a conceptual diagram illustrating processed variants of example images, each processed with a different alpha (α) value to adjust color saturation based on several examples. For α values ​​less than 1 (α < 1), the image becomes more desaturated. For example, when α = 0, all R', G', and B' values ​​are equal to the luminance (Y) value (α = 0: R' = G' = B' = Y), resulting in grayscale (desaturation). Output images 1020, 1021, and 1022 are all generated by applying saturation adjustments of different intensities to the same input image. The input image shows the faces of three women adjacent to each other. Because... Figure 10 Displayed in grayscale instead of color, therefore the red channels of output images 1020, 1021, and 1022 are in grayscale. Figure 10 The diagram illustrates the changes in red saturation. An increase in red saturation is represented by brighter areas, while a decrease in red saturation is represented by darker areas. When α = 1, there is no effect on saturation. For example, when α = 1, equations (1)-(3) result in R' = R, G' = G, and B' = B. Figure 10 The output image 1021 is an example of an image with the alpha (α) value set to 1.0 and no saturation effect. Therefore, the output image 1021 is identical to the input image in terms of color saturation. Figure 10 The output image 1020 is an example of a grayscale (desaturated) image generated by an alpha (α) value of 0.0. Because... Figure 10 The red channel of output image 1020 is shown, so the darker faces in output image 1020 (compared to output image 1021, and therefore compared to the input image) indicate that the reds in the faces are less saturated. For α values ​​greater than 1 (α>1), the image will become more saturated. R, G, and B values ​​can be cropped at the highest intensity. Figure 10 The output image 1022 is an example of a highly saturated image generated by an alpha (α) value of 2.0. Because... Figure 10 The red channel of the output image 1020 is displayed, so the brighter face in the output image 1022 (compared to the output image 1021, and therefore compared to the input image) indicates that the red in the face is more saturated.

[0137] Figure 11This is a conceptual diagram illustrating an example of an input image 1102 and an input saturation map 1104. The saturation map key 1105 specifies that a value of 0 in the saturation map 1104 represents desaturation, a value of 0.5 represents no saturation effect, and a value of 1 represents saturation. The input saturation map 1104 includes a value of 0.5 for a set of locations in the lower left portion of the saturation map 1104 to indicate that no saturation effect will be applied to the corresponding pixels of the input image 1102. The saturation map 1104 includes a value of 0.1 for a set of locations outside the lower left portion of the saturation map 1104. A value of 0.1 will result in pixels having reduced saturation (almost desaturated).

[0138] Image processing ML system 210 can process input image 1102 and input saturation map 1104 to generate modified image 1115. As shown, because no saturation effect is applied to those pixels based on the saturation map value of 0.5, the pixels in the lower left part of the modified image 1115 are the same as the corresponding pixels in the input image 1102. Because the saturation is reduced based on the saturation map value of 0.1 applied to those pixels, the pixels in the remaining part of the modified image 1115 have a deeper green appearance (almost grayscale).

[0139] Another example of a color space tuning diagram is a hue diagram. Hue is the color in an image, and saturation (in the HSV color space) is the intensity (or richness) of that color. A hue diagram indicates the amount of color variation to be applied to the pixels of an input image. Hue adjustments can be applied using different techniques. For example, the Hue, Saturation, Value (HSV) color space is a representation of the RGB color space. HSV representation uses both saturation and hue dimensions to model how different colors blend together.

[0140] Figure 12A This is diagram 1215 illustrating the concept of the HSV color space, where saturation is represented on the y-axis and hue on the x-axis, with values ​​of 255. Hue (color) can be modified, as shown on the x-axis. As shown, hue is wrapped around the image; in this case, hue 0 and 180 represent the same red (according to OpenCV convention, hue: 0 = hue: 180). Since these diagrams are displayed in black and white, the colors are written as text at their positions in the HSV color space.

[0141] Note that saturation in the HSV color space is not the same as the saturation defined by the saturation diagram above. For example, when referring to saturation in the HSV color space, saturation refers to the standard HSV color space definition, where saturation is the richness of a particular color. However, depending on the HSV value, a saturation value of 0 can make a pixel appear as black (value = 0), gray (value = 128), or white (value = 255). In some cases, this saturation can be called "HSV saturation." Because color saturation and color value are coupled in the HSV space, this definition of saturation (HSV saturation) is not used here. Instead, as used herein, saturation refers to how colorful or gray a pixel is, which can be captured by the alpha (α) value described above.

[0142] Figure 12B This is a conceptual diagram 1216 illustrating the YUV color space. The original UV vector 1217 is shown relative to the center point 1219 (U = 128, V = 128). The original UV vector 1217 can be modified by adjusting the angle (θ), thereby producing the modified hue (color) indicated by the hue modification vector 1218. The hue modification vector 1218 is shown relative to the center point 1219. In the YUV color space, the UV vector direction (U = 128, V = 128, etc.) can be modified relative to the center point 1219. Figure 9 (As shown). UV vector modification illustrates an example of how hue adjustment can work, where the intensity of the hue adjustment corresponds to the angle (θ), and the direction of the hue adjustment corresponds to whether the angle (θ) is negative or positive.

[0143] Figure 13 This is a conceptual diagram illustrating an example of an input hue map 1304 applied to an input image 1302 to generate a modified image 1315, based on some examples. The hue map key 1305 specifies that a value of 0.5 in the hue map 1304 indicates no effect (no change in hue or color) (hue adjustment with zero intensity). Any value other than 0.5 indicates a hue change relative to the current color (hue adjustment using a non-zero intensity). For example, see reference... Figure 12AAs an illustrative example, if the current color is green (i.e., hue = 60), a hue value of 60 can be added to convert the color to blue (i.e., hue = 60 + 60 = 120). In another example, a hue value of 60 can be subtracted from the existing hue value of 60 to convert green to red (i.e., hue = 60 - 60 = 0). The input hue map 1304 includes a value of 0.5 for a set of locations (medium gray shaded areas) in the upper left portion of the hue map 1304, indicating that no hue change (a hue adjustment with zero intensity) will be applied to the corresponding pixels of the input image 1302. The hue map 1304 includes a value of 0.4 for a set of locations (dark gray shaded areas) outside the upper left portion of the hue map 1304. A value of 0.4 will cause the pixels in the input image 1302 to change from the existing hue to the modified hue. Because the value of 0.4 is closer to 0.5 than 0.0, the hue change is slight compared to the input image 1302. The hue changes indicated by the input hue map 1304 are relative to the hue changes of each pixel in the input image 1302. The mapping of each pixel can be used... Figure 12A The color space shown is used to determine this. This is in contrast to other tunings described here. Figure 1 Similarly, the hue map values ​​can range from 0 to 1, and are remapped to the actual 0-180 value range during the application of the hue map.

[0144] Input image 1302 shows a portion of a yellow wall with the purple text "Griffith-". Image processing ML system 210 can process input image 1302 and input hue map 1304 to generate a modified image 1315. As shown, pixels in the upper left portion of the modified image 1315 have the same hue as their corresponding pixels in input image 1302 (this is because no hue change (hue adjustment with zero intensity) is applied to those pixels based on a hue map value of 0.5 in that region). Therefore, the pixels in the upper left portion of the modified image 1315 reproduce the yellow hue of the yellow wall in input image 1302. Pixels in the remaining portion of the modified image 1315 have a changed hue compared to the pixels in input image 1302 (this is because a hue change is applied to those pixels based on a hue map value of 0.4). Specifically, in the remaining portion of the modified image 1315, the wall area that appears yellow in input image 1302 appears orange in the modified image 1315. In the remainder of the modified image 1315, the text “Griffith-”, which appears as purple in the input image 1302, appears as a different shade of purple in the modified image 1315.

[0145] In some examples, multiple tuning maps may be input to the image processing ML system 210 along with the input image 1402. In some examples, the multiple tuning maps may include different tuning maps corresponding to different image processing functions, such as denoising, noise addition, sharpening, desharpening (e.g., blurring), hue adjustment, saturation adjustment, detail adjustment, hue adjustment, or combinations thereof. In some examples, the multiple tuning maps may include different tuning maps corresponding to the same image processing function, but may modify different regions / positions within the input image 1402 in different ways, for example.

[0146] Figure 14A This is a conceptual diagram illustrating an example of an image processing system 1406 that includes receiving an input image 1402 and multiple spatially varied tuned images 1404. In some examples, the image processing system 1406 may be an image processing ML system 210 and / or may include a machine learning (ML) system, such as the image processing ML system 210. The ML system can apply image processing functions using ML. In some examples, the image processing system 1406 can apply image processing functions without an ML system. The image processing system 1406 can be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof.

[0147] Spatial variation tuning map 1404 may include one or more of a noise map, sharpness map, hue map, saturation map, hue map, and / or another map associated with an image processing function. Using an input image 1402 and multiple tuning maps 1404 as input, image processing system 1406 generates a modified image 1415. Image processing system 1406 may modify the pixels of input image 1402 based on values ​​included in the spatial variation tuning maps 1404 to produce modified image 1415. For example, image processing system 1406 may modify the pixels of input image 1402 by applying an image processing function to each pixel of input image 1402 with an intensity indicated by a value included in one of the spatial variation tuning maps 1404 corresponding to the image processing function, and may do so for each image processing function and its corresponding one of the spatial variation tuning maps 1404, thereby producing modified image 1415. Spatial variation tuning maps 1404 may be generated using an ML system, such as at least regarding Figure 14B , Figure 14C , Figure 14D , Figure 16B , Figure 17 , Figure 18 , Figure 19 , Figure 20 , Figure 21A , Figure 21B , Figure 21C and Figure 21DAs discussed further.

[0148] Figure 14B This is a conceptual diagram illustrating an example system based on some examples. The system includes an image processing system 1406 that receives an input image 1402 and multiple spatially varied tuning maps 1404, and an automatic adjustment machine learning (ML) system 1405 that receives the input image 1402 and generates multiple spatially varied tuning maps 1404. The automatic adjustment ML system 1405 can process the input image 1402 to generate the spatially varied tuning maps 1404. The tuning maps 1404 and the input image 1402 can be provided as input to the image processing system 1406. The image processing system 1406 can process the input image 1402 based on the tuning maps 1404 to generate a modified image 1415, similar to the above description. Figure 14A As described, the Auto-Tuned ML System 1405 can be implemented using one or more Convolutional Neural Networks (CNNs), one or more CNNs, one or more Trained Neural Networks (NNs), one or more NNs, one or more Trained Support Vector Machines (SVMs), one or more SVMs, one or more Trained Random Forests, one or more Random Forests, one or more Trained Decision Trees, one or more Decision Trees, one or more Trained Gradient Boosting Algorithms, one or more Gradient Boosting Algorithms, one or more Trained Regression Algorithms, one or more Regression Algorithms, or combinations thereof. The Auto-Tuned ML System 1405 and / or any of the machine learning elements listed above (which may be part of the Auto-Tuned ML System 1405) can be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0149] Figure 14C This is a conceptual diagram illustrating an example system based on some examples, including an automatically adjusting machine learning (ML) system 1405 that receives an input image 1402 and generates multiple spatial variation tuning maps 1404. Figure 14C In the example shown, the spatial variation tuning map 1404 includes a noise map, a tone map, a saturation map, a hue map, and a sharpness map. Image processing system 1406 (although not included in...) Figure 14C (As shown in the diagram) Therefore, noise reduction can be applied at the intensity indicated in the noise graph, hue adjustment can be applied at the intensity and / or direction indicated in the hue graph, saturation adjustment can be applied at the intensity and / or direction indicated in the saturation graph, hue adjustment can be applied at the intensity and / or direction indicated in the hue graph, and / or sharpening can be applied at the intensity indicated in the sharpness graph.

[0150] Figure 14DThis is a conceptual diagram illustrating an example system based on some examples, including an image processing system 1406 that receives an input image 1402 and multiple spatial variation tuning maps 1404, and an automatic adjustment machine learning (ML) system 1405 that receives a downsized variant of the input image 1422 and generates a smaller spatial variation tuning map 1424 that is enlarged into multiple spatial variation tuning maps 1404. Figure 14D The system is similar to Figure 14B The system includes a downsampler 1418 and an upsampler 1426. The downsampler 1418 downsamples, reduces, or shrinks the input image 1402 to generate a downsampled input image 1422. It is not like... Figure 14B The machine learning (ML) system 1405 receives the downsampled input image 1422 as input, and automatically adjusts the input image 1422 as input. Figure 14D The input is the image. An automatic adjustment machine learning (ML) system 1405 generates small spatial variation tuning maps 1424, each tuning map potentially sharing the same size and / or dimensions as the downsampled input image 1422. An upsampler 1426 can upsample, enlarge, and / or expand the small spatial variation tuning maps 1424 to generate a spatial variation tuning map 1404. In some examples, the upsampler 1426 may perform bilinear upsampling. The spatial variation tuning map 1404 can then be received as input by an image processing system 1406 along with the input image 1402.

[0151] Because the downsampled input image 1422 has a smaller size and / or resolution compared to the input image 1402, it is more efficient than the automatic adjustment machine learning (ML) system 1405 to directly generate the spatial variation tuning map 1404 from the input image 1402. Figure 14B As shown, the automatically adjusted machine learning (ML) system 1405 may generate a small spatial variation tuning map 1424 from the downsampled input image 1422 faster and more efficiently (in terms of computing resources, bandwidth, and / or battery power), along with downsampling performed by downsampler 1418 and / or upsampling performed by upsampler 1426, as Figure 14D As shown.

[0152] As described above, the image processing system 1406, the image processing ML system 210, and / or the automatically adjusted machine learning (ML) system 1405 may include one or more neural networks that can be trained using supervised learning techniques.

[0153] Figure 15This is a block diagram 1600A illustrating examples of a neural network 1500 that can be used by an image processing system 1406 and / or an automatically tuned machine learning (ML) system 1405, according to some examples. The neural network 1500 can include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief network (DBN), a recurrent neural network (RNN), a generative adversarial network (GAN), and / or other types of neural networks.

[0154] The input layer 1510 of the neural network 1500 includes input data. The input data of the input layer 1510 may include data representing pixels of an input image frame. In an illustrative example, the input data of the input layer 1510 may include data representing... Figures 14A-14D The input data for the input image 1402 and / or the downsampled input image 1422 is pixel data (e.g., for an NN 1500 of an auto-adjustment ML system 1405 and / or an image processing system 1406). In an illustrative example, the input data for the input layer 1510 may include data representing... Figures 14A-14D The spatial variation tunes the pixel data of image 1404 (e.g., for NN 1500 in an auto-tuning ML system 1405). The image may include image data from an image sensor, including raw pixel data (including a single color per pixel based on, for example, a Bayer filter) or processed pixel values ​​(e.g., RGB pixels of an RGB image). The neural network 1500 includes multiple hidden layers 1512a, 1512b to 1512n. Hidden layers 1512a, 1512b to 1512n include a number of hidden layers of “n”, where “n” is an integer greater than or equal to 1. The number of hidden layers can be set to include as many layers as needed for a given application. The neural network 1500 also includes an output layer 1514, which provides the output of the processing performed by the hidden layers 1512a, 1512b to 1512n. In an illustrative example, the output layer 1514 may provide a modified image, such as Figure 2 The modified image 215 or Figures 14A-14D The modified image 1415 (e.g., for the NN 1500 of image processing system 1406 and / or image processing ML system 210). In the illustrative example, output layer 1514 can provide Figures 14A-14D Spatial variation tuning diagram 1404 and / or Figure 14D Small spatial variation tuning diagram 1424 (e.g., for NN 1500 of automatic tuning ML system 1405).

[0155] Neural Network 1500 is a multi-layer neural network of interconnected filters. Each filter can be trained to learn features representing the input data. Information associated with the filters is shared between different layers, and each layer retains information while processing it. In some cases, Neural Network 1500 may include a feedforward network, in which case there are no feedback connections, and the network's output is fed back to itself. In some cases, Network 1500 may include a recurrent neural network, which may have loops that allow information to be carried across nodes while the input is being read in.

[0156] In some cases, information can be exchanged between layers through node-to-node interconnects. In some cases, the network may include a convolutional neural network, which may not link every node in one layer to every other node in the next layer. In a network that exchanges information between layers, nodes in input layer 1510 can activate a set of nodes in the first hidden layer 1512a. For example, as shown, each input node in input layer 1510 can be connected to each node in the first hidden layer 1512a. The nodes in the hidden layers can transform the information of each input node by applying activation functions (e.g., filters) to this information. The information derived from the transformation can then be passed to and can activate nodes in the next hidden layer 1512b, which can perform their own specified functions. Example functions include convolution functions, scaling, scaling, data transformations, and / or any other suitable functions. The output of hidden layer 1512b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 1512n can activate one or more nodes in output layer 1514, which provides the processed output image. In some cases, although a node in neural network 1500 (e.g., node 1516) is shown as having multiple output lines, the node has a single output and all the lines shown as outputs from the node represent the same output value.

[0157] In some cases, each node or the interconnection between nodes can have weights derived from a set of parameters trained on the neural network 1500. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnections can have adjustable numerical weights that can be tuned (e.g., based on the training dataset) to allow the neural network 1500 to adapt to the input and learn as more and more data is processed.

[0158] The neural network 1500 is pre-trained to process features from the data in the input layer 1510 using different hidden layers 1512a, 1512b to 1512n in order to provide output through the output layer 1514.

[0159] Figure 16AThis is a block diagram 1600B illustrating a training example based on some example image processing system 1406 (e.g., image processing ML system 210). Reference Figure 16A A neural network (e.g., neural network 1500) implemented by image processing system 1406 (e.g., image processing ML system 210) can be pre-trained to process the input image and the tuning map. Figure 16A As shown, the training data includes input image 1606 and input tuning map 1607. Input image 1402 can be an example of input image 1606. Spatially varied tuning map can be an example of input tuning map 1607. Input image 1606 and input tuning map 1607 can be input into a neural network (e.g., neural network 1500) of image processing ML system 210, and the neural network can generate output image 1608. Modified image 1415 can be an example of output image 1608. For example, a single input image and multiple tuning maps (e.g., one or more of the aforementioned noise map, sharpness map, hue map, saturation map, and / or hue map, and / or another map associated with an image processing function) can be input into the neural network, and the neural network can output an output image. In another example, a batch of input images and multiple corresponding tuning maps can be input into the neural network, and the neural network can then generate multiple output images.

[0160] A set of reference output images 1609 may also be provided for comparison with the output image 1608 of the image processing system 1406 (e.g., image processing ML system 210) to determine loss (as described below). Reference output images may be provided for each input image of the input image 1606. For example, the output images from the reference output images 1609 may include a final output image previously generated by the camera system and having the desired characteristics for the corresponding input image based on multiple tuning maps from the input tuning map 1607.

[0161] The parameters of a neural network can be tuned based on a comparison of the output image 1608 and the reference output image 1609 by the backpropagation engine 1612. Parameters may include the weights, biases, and / or other parameters of the neural network. In some cases, the neural network (e.g., neural network 1500) may use a training process called backpropagation to adjust the weights of its nodes. Backpropagation may include forward pass, loss function, backpropagation, and weight update. Forward pass, loss function, backpropagation, and parameter update are performed in one training iteration. For each set of training images, this process may be repeated a certain number of iterations until the neural network is trained well enough to accurately tune the layer weights. Once the neural network is properly trained, the image processing system 1406 (e.g., image processing ML system 210) can process any input image and any number of tuned images to generate modified versions of the input image based on the tuned images.

[0162] Forward propagation may involve passing an input image (or a batch of input images) and multiple tuning maps (e.g., one or more of the noise map, sharpness map, tone map, saturation map, and / or hue map discussed above) through a neural network. The weights of the various filters in the hidden layers may be initially randomized before training the neural network. The input image may include a multidimensional digital array representing the image's pixels. In one example, this array may include a 128x128x11 digital array with 128 rows and 128 columns of pixel locations and 11 input values ​​per pixel location.

[0163] For the first training iteration of a neural network, the output may include values ​​that do not prioritize any particular feature or node because the weights are randomly selected during initialization. For example, if the output is an array with many color components at each pixel location, the output image may depict an inaccurate color representation of the input. Using the initial weights, the neural network cannot determine low-level features and therefore cannot accurately determine what the color values ​​might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined. An example of a loss function is Mean Squared Error (MSE). MSE is defined as... It calculates the mean or average of the squared differences (the actual answer minus the predicted (output) answer, squared). The term n is the number of values ​​in the sum. The loss can be set to equal E. total The value of .

[0164] For the first training data (image data and the corresponding tuned map), the loss (or error) will be high because the actual values ​​will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training labels. Neural networks can perform backpropagation by determining which inputs (weights) contribute most to the network loss, and the weights can be adjusted to reduce and eventually minimize the loss. In some cases, the derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) (or other suitable function) can be calculated to determine the weights that contribute most to the network loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be expressed as... Where w represents the weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate involves larger weight updates, while a lower value indicates smaller weight updates.

[0165] Figure 16B This is a block diagram illustrating training examples of the automatically adjusted machine learning (ML) system 1405 based on some examples. (Reference) Figure 16B A neural network (e.g., neural network 1500) implemented by an automatically tuned machine learning (ML) system 1405 can be pre-trained to process input image 1606 and / or downsampled input image 1616. Figure 16BAs shown, the training data includes input image 1606 and / or downsampled input image 1616. Input image 1402 can be an example of input image 1606. Downsampled input image 1422 can be an example of downsampled input image 1616. Input image 1606 and / or downsampled input image 1616 can be fed into a neural network (e.g., neural network 1500) of automatic tuning machine learning (ML) system 1405, and the neural network can generate output tuning map 1618. Spatial variation tuning map 1404 can be an example of output tuning map 1618. Small spatial variation tuning map 1424 can be an example of output tuning map 1618. For example, input image 1606 can be fed into a neural network, and the neural network can output one or more output tuning maps 1618 (e.g., one or more of the aforementioned noise map, sharpness map, hue map, saturation map, and / or hue map, and / or another map associated with image processing functions). For example, the downsampled input image 1616 can be fed into a neural network, and the neural network can output one or more (small) output tuning maps 1618 (e.g., small variations of one or more of the aforementioned noise map, sharpness map, hue map, saturation map, and / or hue map, and / or another other map associated with the image processing function). In another example, a batch of input images 1606 and / or the downsampled input image 1616 can be fed into a neural network, and the neural network can then generate multiple output tuning maps 1618.

[0166] refer to Figure 16B The neural network implemented by the automatically tuned machine learning (ML) system 1405 (e.g., neural network 1500) can include similar... Figure 16A The backpropagation engine 1612 and the backpropagation engine 1622. Figure 16B The backpropagation engine 1622 can be similar to Figure 16A The backpropagation engine 1612 receives and uses the reference output image 1609 and the reference output tuning image 1619.

[0167] As described above, in some embodiments, a machine learning system separate from the image processing ML system 210 can be used to automatically generate tuning maps.

[0168] Figure 17 This is a block diagram illustrating an example of a system including an image processing ML system and an automatic adjustment machine learning (ML) system 1705, which generates affine coefficients (a,b) from luminance channel data (I...). y Generate a spatially varying tuning map Ω, and modify the luminance channel data (I) according to the local linearity constraint 1720. yThe automatically tuned machine learning (ML) system 1705 can be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof. The automatically tuned machine learning (ML) system 1705 can receive input image data (e.g., input image 1402), such as... Figure 17 The image shows I. RGB From input image data I RGB The luminance channel data is called I. y In some examples, the luminance channel data I y Actually, it's the input image data I RGB The grayscale version. The ML system 1705 automatically adjusts the input image data received by I. RGB As input, and outputting one or more affine coefficients, denoted here as a and b. The one or more affine coefficients include a multiplier. The one or more affine coefficients include an offset b. The automatic adjustment of the ML system 1705 or another imaging system can be achieved according to the equation Ω = a * I. y +b applies the affine coefficients to the luminance channel data I. y To generate the tuning diagram Ω. The ML system 1705 or another imaging system can be automatically adjusted according to the equation... Implement local linearity constraints 1720. Using local linearity constraints 1720 ensures that one or more gradients in the graph are aligned with one or more gradients in the image data, which reduces halo effects in image processing applications. Generating graphs using affine coefficients and / or local linearity constraints 1720 can produce higher quality spatially varied image modifications compared to systems that do not use affine coefficients and / or local linearity constraints, for example, due to better alignment between the image data and the graph, and reduced halo effects at the boundaries of depicted objects. Each generated tuning graph may include its own set of one or more affine coefficients (e.g., a, b). For example, in Figure 14C In the example, the automatic adjustment ML system 1405 generates five different spatial variation tuning maps 1404 from an input image 1402. Therefore, Figure 14C The automatic adjustment ML system 1405 can generate one or more sets of affine coefficients (e.g., a, b), one set for each of the five different spatially varied tuning diagrams 1404. One or more affine coefficients (e.g., a and / or b) can also be spatially varied.

[0169] The Automatically Tuned ML System 1705 can be implemented using one or more Convolutional Neural Networks (CNNs), one or more CNNs, one or more Trained Neural Networks (NNs), one or more NNs, one or more Trained Support Vector Machines (SVMs), one or more SVMs, one or more Trained Random Forests, one or more Random Forests, one or more Trained Decision Trees, one or more Decision Trees, one or more Trained Gradient Boosting Algorithms, one or more Gradient Boosting Algorithms, one or more Trained Regression Algorithms, one or more Regression Algorithms, or combinations thereof. The Automatically Tuned ML System 1705 and / or any of the machine learning elements listed above (which can be part of the Automatically Tuned ML System 1705) can be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0170] Figure 18 This is a block diagram illustrating the details of the automatic adjustment ML system 1705. As shown, the automatic adjustment ML system 1705 includes a local neural network 1806 for patch-by-patch processing (e.g., processing of image patches of an image) and a global neural network 1807 for processing the complete image. The local neural network 1806 can receive an image patch 1825 (which is part of the complete image input 1826) as input and can process the image patch 1825. The global neural network 1807 can receive the complete image input 1826 as input and can process the complete image input 1826. The output from the global neural network 1807 can be provided to the local neural network 1806. In some examples, the automatic adjustment ML system 1705 can output a spatially varying tuning map 1804. In some examples, the automatic adjustment ML system 1705 can output affine coefficients 1822 (e.g., Figure 17 In (a) and (b), the automatic adjustment ML system 1705 can apply the affine coefficients to the luminance (Y) channel 1820 to generate a spatially varied tuning map 1804.

[0171] In some examples, for computational efficiency, one or more downsized low-resolution images (downsized, downsampled, and / or reduced from full-resolution images) can be fed to the global neural network 1807, and one or more full-resolution or high-resolution image patches can be fed to the local neural network 1806 to obtain a corresponding full-resolution spatial variation tuning map. The high-resolution image patch (e.g., image patch 1825) and the low-resolution full-image input (e.g., image input 1826) can be based on the same image (e.g., ...). Figure 18(As shown), however, the high-resolution image patch 1825 is a higher-resolution version of the low-resolution full image input 1826. For example, the low-resolution full image input 1826 can be a scaled-down version of a higher-resolution image from which the high-resolution image patch 1825 is extracted. The local NN 1806 can be referred to as the high-resolution NN 1806. The global NN 1807 can be referred to as the low-resolution NN 1807.

[0172] The luminance (Y) channel 1820 is shown as originating from image patch 1825, but may originate from the full image input 1826. The luminance (Y) channel 1820 may be high-resolution or low-resolution. The luminance (Y) channel 1820 may include luminance data from the full image input 1826, either at its original high resolution or a scaled-down low resolution. The luminance (Y) channel 1820 may also include luminance data from image patch 1825, either at its original high resolution or a scaled-down low resolution.

[0173] Figure 19 This is a block diagram illustrating an example of the neural network architecture 1900 of the local neural network 1806 in the automatically adjusting ML system 1705. The local neural network 1806 may be referred to as a high-resolution local neural network 1806 or a high-resolution neural network 1806. The neural network architecture 1900 receives, can be a high-resolution image patch 1825, as its input. The neural network architecture 1900 outputs affine coefficients 1822 (e.g., as shown in the image). Figure 17 (a) and / or (b) in the figure, which can be high resolution. Affine coefficient 1822 can be used to generate tuning diagram 1804, such as... Figure 17 , Figure 18 , Figure 20 , Figure 21A and / or Figure 21B As shown. Key 1920 identifier. Figure 19 The diagram illustrates how different neural network operations are described. For example, a convolution with a 3x3 filter and a stride of 1 is represented by a thick white arrow outlined in black and pointing to the right. A convolution with a 2x2 filter and a stride of 2 is represented by a thick black arrow pointing downwards. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing upwards.

[0174] Figure 20 This is a block diagram illustrating an example of the neural network architecture 2000 of the global neural network 1807 in the automatically adjusting ML system 1705. The global neural network 1807 may be referred to as a low-resolution global neural network 1807 or a low-resolution neural network 1807. The neural network architecture 2000 receives a complete image input 1826 as its input, which may be downsampled and / or may be low-resolution. The output of the neural network architecture 2000 may be low-resolution global features 2010. Key 2020 identifies... Figure 20The diagram illustrates how different neural network operations are represented. For example, a convolution with a 3x3 filter and a stride of 1 is indicated by a thick white arrow outlined in black and pointing to the right. A convolution with a 2x2 filter and a stride of 2 is indicated by a thick black arrow pointing downwards. Average pooling is indicated by a thick white arrow with black diagonal shading, outlined in black, and pointing downwards. A fully connected layer is indicated by a thin black arrow pointing to the right.

[0175] Figure 21A This is a block diagram illustrating the neural network architecture 2100A of an auto-tuning ML system 1705 according to some examples. A low-resolution full image 1826 allows the global neural network 1807 of the auto-tuning ML system 1705 to generate global information about the image. The low-resolution global neural network 1807 processes the entire image and determines global features that are important as it progresses from input to output. By using a reduced-resolution image as input to the global neural network 1807, the computation for global information is considerably lighter compared to using a full-resolution image. High-resolution patches 1825 allow the auto-tuning network to be locally consistent, thus eliminating discontinuities in the local features of the image. Global features can be incorporated into layers after one or more convolutional operations (e.g., via channel attention with additive bias), allowing the local neural network to take global features into account while also generating affine coefficients 1822 based on the image patch 1825 and / or while generating a tuning map 1804 based on the image patch 1825. Figure 21A As shown, affine coefficient 1822 can be combined with Y channel 1820 to generate tuning pattern 1804.

[0176] Key 2120A recognition Figure 21A The diagram illustrates how different NN operations are described. For example, a convolution with a 3x3 filter and a stride of 1 is represented by a thick white arrow outlined in black and pointing to the right. A convolution with a 2x2 filter and a stride of 2 is represented by a thick black arrow pointing downwards. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing upwards. Channel attention with additive bias is indicated by thin black arrows pointing upwards and / or to the left (e.g., upwards from global features determined by the global NN 1807).

[0177] Figure 21BThis is a block diagram illustrating another example of the neural network architecture 2100B of an automatically tuned ML system 1705, based on some examples. The input to the neural network architecture 2100B can be an input image 2130. The input image 2130 can be a high-resolution full image. The input image 2130 can be a raw high-resolution or downsized low-resolution full image input 1826. The input image 2130 can be a high-resolution or downsized low-resolution image patch 1825. The neural network architecture 2100B can process the input image 2130 to generate affine coefficients 1822 based on the input image 2130. The neural network architecture 2100B can process the input image 2130 to generate a tuning map 1804 based on the input image 2130. The affine coefficients 1822 can be combined with the Y channel 1820 to generate the tuning map 1804. Spatial attention engine 2110 and channel attention engine 2115 are part of the neural network architecture 2100B. Spatial attention engine 2110 in... Figure 21C This is shown in more detail below. Channel attention engine 2115 in... Figure 21D It is shown in more detail below.

[0178] Key 2120B identifier Figure 21A , Figure 21B and Figure 21C How are different NN operations described? For example, a convolution with a 3x3 filter and a stride of 1 is indicated by a thick white arrow outlined in black and pointing to the right. A convolution with a 1x1 filter and a stride of 1 is indicated by a thin black arrow pointing downwards. A convolution with a 2x2 filter and a stride of 2 is indicated by a thick black arrow pointing downwards. Upsampling (e.g., bilinear upsampling) is indicated by a thick black arrow pointing upwards. Operations including attention, upsampling, and multiplication are indicated by thin black arrows pointing upwards and / or to the side. A circled "X" symbol indicates that affine coefficient 1822 is applied to the Y channel 1820. A double-circled "X" symbol indicates element-wise multiplication after unrolling. Thin black dashed arrows extending from one side to the other (in...) Figure 21D (in Chinese) indicates shared parameters.

[0179] In some examples, the local NN 1808 can directly generate the tuned graph 1804 without first generating the affine coefficients 1822.

[0180] Figure 21C This is a block diagram illustrating an example neural network architecture 2100C of a spatial attention engine 2110 based on some examples. The spatial attention engine 2110 includes max pooling, average pooling, and cascading.

[0181] Figure 21DThis is a block diagram illustrating an example of a neural network architecture 2100D for a channel attention engine 2115, based on some examples. The spatial attention engine 2110 includes max pooling, average pooling, shared parameters, and summation.

[0182] By using machine learning to generate affine coefficients 1822 and / or tuning maps 1804, the imaging system offers improved customization and context sensitivity. The imaging system is able to provide tuning maps suitable for the content of the input image without producing visual artifacts such as halos. Machine learning-based generation of affine coefficients 1822 and / or tuning maps 1804 can also be more efficient than traditional manual adjustment of image processing parameters. Machine learning-based generation of affine coefficients 1822 and / or tuning maps 1804 can also accelerate imaging innovation. For example, using machine learning to generate affine coefficients 1822 and / or tuning maps 1804 allows for faster and easier adaptation to data and other variations from additional sensors, different types of lenses, different types of camera arrays, and other changes.

[0183] Supervised learning techniques, similar to those described above for the image processing ML system 210, can be used to train the neural network of the automatically tuned ML system 1705. For example, backpropagation can be used to tune the parameters of the neural network of the automatically tuned ML system 1705. Training data can include input images and known output images with the desired characteristics obtained by applying different tuning maps. For example, based on the input and output images, the neural network will be trained to generate a set of masks that, when applied to the input image, produce the corresponding output. Using saturation and hue as illustrative examples, the neural network will attempt to determine how much saturation and hue needs to be applied to different pixels in the input image to achieve the characteristics of the pixels in the output image. In such an example, the neural network can generate saturation and hue maps based on training.

[0184] The inference process of the automatically tuned ML system 1705 can be performed by processing the input image to generate a tuning map 1804 (once the neural network has been trained). For example, global features can be extracted from a low-resolution version of the input image by a global neural network 1807. Single-pass inference can be performed on the low-resolution full-image input. Patch-based inference can then be performed using a local neural network 1806, where global features from the global neural network 1807 are fed into the local network 1806 during inference. High-resolution patch-based inference is then performed on patches of the input image to generate the tuning map 1804.

[0185] In some examples, the image processing ML system 210 described above (and in some cases, the automatic adjustment ML system 1705) can be used when capturing an image or when processing a previously captured image. For example, when capturing an image, the image processing ML system 210 can process the image based on using a tuning map to generate an output image with optimal characteristics (e.g., regarding noise, sharpness, hue, saturation, and hue). In another example, the image processing ML system 210 can retrieve and process a previously generated, stored image to generate an enhanced output image with optimal characteristics based on using a tuning map.

[0186] In some examples, the aforementioned image processing ML system 210 (and in some cases, the automatically tuned ML system 1705) can be used when tuning the ISP. For example, the parameters of the ISP are routinely adjusted manually by an expert with experience in how to process input images to obtain the desired output image. Due to the correlation between the ISP module (e.g., filters) and a large number of adjustable parameters, the expert may need several weeks (e.g., 3-8 weeks) to determine, test, and / or adjust the device settings based on a specific combination of camera sensor and ISP. Because the camera sensor or other camera characteristics (e.g., lens characteristics or defects, aperture size, shutter speed and movement, flash brightness and color, and / or other characteristics) can affect the captured image and therefore at least some of the adjustable parameters of the ISP, each combination of camera sensor and ISP can be adjusted by an expert.

[0187] Figure 22This is a block diagram illustrating an example of a pre-tuned image signal processor (ISP) 2208. As shown, an image sensor 2202 captures raw image data. Photodiodes in the image sensor 2202 capture shading variations in grayscale (or monochrome). Color filters can be applied to the image sensor to provide color-filtered raw input data 2204 (e.g., with a Bayer pattern). The ISP 2208 has discrete function blocks, each applying a specific operation to the raw camera sensor data to create a final output image. For example, function blocks may include blocks dedicated to demosaicing, gain adjustment, white balance, color correction, gamma compression (or gamma correction), tone mapping, noise reduction (denoising), etc. For example, the demosaicing function block of the ISP 2208 can help generate an output color image 2209 from the color-filtered raw input data 2204 by interpolating the color and brightness of pixels using neighboring pixels. The ISP 2208 can use this demosaicing process to evaluate the color and brightness data of a given pixel and compare these values ​​with data from neighboring pixels. The ISP 2208 can then use a demosaicing algorithm to generate appropriate color and brightness values ​​for the pixels. The ISP 2208 can perform various other image processing functions before providing the final output color image 2209, such as noise reduction, sharpening, tone mapping and / or color space conversion, autofocus, gamma adjustment, exposure adjustment, white balance, and many other possible image processing functions.

[0188] The function blocks of the ISP 2208 need to be manually tuned to meet numerous tuning parameters 2206 of a specific specification. In some cases, more than 10,000 parameters need to be tuned and controlled for a given ISP. For example, to optimize the output color image 2209 according to a specific specification, the algorithm for each function block must be optimized using the tuning parameters 2206 of the tuning algorithm. New function blocks must also be continuously added to handle different situations that arise in space. The large number of manually tuned parameters makes supporting the ISP very time-consuming and expensive.

[0189] In some cases, machine learning systems (called machine learning ISPs) can be used to implement ISPs, thereby performing multiple ISP functions in a federated manner.

[0190] Figure 23This is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) 2300. The machine learning ISP 2300 may include an input interface 2301 that can receive raw image data from an image sensor 2302. In some cases, the image sensor 2302 includes an array of photodiodes that capture frames 2304 of raw image data. Each photodiode may represent a pixel location and may generate a pixel value for that pixel location. The raw image data from the photodiodes may include a single color or grayscale value for each pixel location in frame 2304. For example, a color filter array may be integrated with or used in conjunction with the image sensor 2302 (e.g., covering the photodiodes) to convert monochrome information into color values.

[0191] An illustrative example of a color filter array includes a Bayer patterned color filter array (or Bayer color filter array) to allow image sensor 2302 to capture pixel frames with a Bayer pattern, wherein one of the red, green, or blue color filters is located at each pixel location. For example, the raw image patch 2306 of frame 2304 from raw image data has a Bayer pattern based on a Bayer color filter array used with image sensor 2302. The Bayer pattern includes red, blue, and green filters, such as... Figure 23 The pattern of the original image patch 2306 is shown. The Bayer color filter operates by filtering out incident light. For example, the photodiodes in the green portion of the pattern pass through green information (half a pixel), the photodiodes in the red portion of the pattern pass through red information (quarter a pixel), and the photodiodes in the blue portion of the pattern pass through blue information (quarter a pixel).

[0192] In some cases, a device may include multiple image sensors (which may be similar to image sensor 2302), in which case the machine learning ISP operations described herein can be applied to the raw image data obtained from multiple image sensors. For example, a device with multiple cameras can use multiple cameras to capture image data, and the machine learning ISP 2300 can apply ISP operations to the raw image data from multiple cameras. In an illustrative example, a dual-camera mobile phone, tablet, or other device can be used to capture a larger image at a wider angle (e.g., with a wider field of view (FOV)), capture more light (resulting in sharper, clearer, and other advantages), to generate 360-degree (e.g., virtual reality) video, and / or perform other enhancements (compared to those achieved by a single-camera device).

[0193] The raw image patch 2306 is provided to and received by the input interface 2301 for processing by the machine learning ISP 2300. The machine learning ISP 2300 can use the neural network system 2303 to perform the ISP task. For example, the neural network of the neural network system 2303 can be trained to directly derive a mapping from raw image training data captured by the image sensor to the final output image. For example, the neural network can be trained using examples of a large amount of raw data input (e.g., with color filtering patterns) and also using examples of the desired corresponding output images. Using the training data, the neural network system 2303 can learn the mapping from the raw inputs required to obtain the output image, after which the ISP 2300 can produce an output image similar to that produced by a conventional ISP.

[0194] The neural network of the ISP 2300 may include an input layer, multiple hidden layers, and an output layer. The input layer includes raw image data obtained by the image sensor 2302 (e.g., a raw image patch 2306 or a complete frame of raw image data). The hidden layers may include filters that can be applied to the raw image data, and / or outputs from previous hidden layers. Each filter in the hidden layer may include weights used to indicate the importance of the filter node. In an illustrative example, the filters may include 3×3 convolutional filters that convolve around the input array, where each entry in the 3×3 filter has a unique weight value. In each convolutional iteration (or stride) of the 3×3 filter applied to the input array, a single weighted output feature value may be produced. The neural network may have a series of many hidden layers, where earlier layers determine low-level characteristics of the input, and later layers build a hierarchical structure of more complex characteristics. The hidden layers of the ISP 2300 neural network are connected to a high-dimensional representation of the data. For example, these layers may include multiple repeating convolutional blocks with a large number of channels (dimensions). In some cases, the number of channels may be an order of magnitude larger than the number of channels in an RGB or YCbCr image. The illustrative examples provided below include repeated convolutions, each with 64 channels, to provide a non-linear and hierarchical network structure that produces high-quality image detail. For example, as described in more detail herein, n channels (e.g., 64 channels) refers to an n-dimensional (e.g., 64-dimensional) representation of data at each pixel location. Conceptually, n channels represent “n features” (e.g., 64 features) at a pixel location.

[0195] The neural network system 2303 implements various signal processing ISP functions in a joint manner. The specific parameters of the neural network used by the neural network system 2303 do not have explicit analogies in conventional ISPs, and conversely, specific functional blocks of conventional ISP systems do not have explicit correspondences in machine learning ISPs. For example, the machine learning ISP performs signal processing functions as a single unit, rather than having individual functional blocks that a typical ISP might contain for performing various functions. Further details of the neural network applied by the neural network system 2303 are described below.

[0196] In some examples, the machine learning ISP 2300 may also include an optional preprocessing engine 2307, which can process additional image tuning parameters to enhance the input data. Such additional image tuning parameters (or enhancement data) may include, for example, hue data, radial distance data, automatic white balance (AWB) gain data, any combination thereof, and / or any other additional data of the pixels that can enhance the input data. By supplementing the original input pixels, the input for each pixel location of the original image data becomes a multidimensional set of values.

[0197] Based on the determined high-level features, the neural network system 2303 can generate an RGB output 2308 based on the original image patch 2306. The RGB output 2308 includes red, green, and blue components for each pixel. The RGB color space is used as an example in this application. Those skilled in the art will understand that other color spaces, such as luminance and chromaticity (YCbCr or YUV) color components, or other suitable color components, can also be used. The RGB output 2308 can be output from the output interface 2305 of the machine learning ISP 2300 and used to generate image patches in the final output image 2309 (constituting the output layer). In some cases, the pixel array in the RGB output 2308 may include dimensions smaller than the input original image patch 2306. In an illustrative example, the original image patch 2306 may contain a 128x128 original image pixel array (e.g., in Bayer mode), while the application of the repetitive convolutional filter of the neural network system 2303 results in the RGB output 2308 comprising an 8x8 pixel array. The output size of the RGB output 2308, smaller than the original image patch 2306, is a byproduct of applying convolutional filters and designing the neural network system 2303 to avoid padding the data processed by each convolutional filter. By having multiple convolutional layers, the output size is reduced further. In this case, patches from the frames 2304 of the input original image data can be overlapped, such that the final output image 2309 contains the complete picture. The resulting final output image 2309 contains the processed image data derived from the original input data by the neural network system 2303. The final output image 2309 can be rendered for display, used for compression (or decoding), stored, or used for any other image-based purpose.

[0198] Figure 24 This is a block diagram illustrating an example of the neural network architecture 2400 of a machine learning (ML) image signal processor (ISP) 2300. Pixel shuffle upsampling is an upsampling method in which the channel dimension is reshaped along the spatial dimension. In one example (which uses two (2x) upsamplings for illustrative purposes), xxxx (meaning 4 channels x 1 spatial location) along 4 channels will be used to generate a single channel with twice the spatial dimension: xxxx (meaning 1 channel x 4 spatial locations). An example of pixel shuffle upsampling is described in “Checkerboard artifact freesub-pixel convolution” by Andrew Aitken et al., which is incorporated herein by reference in its entirety for all purposes.

[0199] By using machine learning to perform ISP functions, ISPs become customizable. For example, different functions can be developed and applied by presenting target data examples and by changing network weights through training. Compared to hardwired or heuristic-based ISPs, machine learning-based ISPs also enable faster turnaround times for updates. Furthermore, machine learning-based ISPs eliminate the time-consuming task of tuning the parameters required for pre-tuned ISPs. For example, there is a significant amount of effort and personnel required to manage ISP infrastructure. Holistic development can be applied to machine learning ISPs, during which end-to-end systems are directly optimized and created. This holistic development contrasts with the piecemeal development of pre-tuned ISP functional blocks. Machine learning-based ISPs can also accelerate imaging innovation. For example, customizable machine learning ISPs unlock numerous innovative possibilities, allowing developers and engineers to drive, develop, and tune solutions more quickly to work with new sensors, lenses, camera arrays, and other advancements.

[0200] As described above, the image processing system 1406 (e.g., image processing ML system 210) and / or the aforementioned automatic adjustment ML system 1405 can be used when tuning the ML ISP and / or a conventional ISP. The various tuning maps described above (e.g., one or more of the aforementioned noise map, sharpness map, hue map, saturation map, and / or hue map, and / or another map associated with an image processing function) can be used to tune the ISP. When used to tune the ISP, these maps can be referred to as tunable knobs. Examples of tuning maps (or tunable knobs) that can be used to tune the ISP include local tone manipulation (e.g., using a hue map), detail enhancement, color saturation, etc. In some cases, local tone manipulation may include contrast-limited adaptive histogram equalization (CLAHE) (e.g., using OpenCV). In some cases, detail enhancement may be performed using edge-aware filtering with domain transformation (e.g., using OpenCV). In some examples, color saturation may be performed using the Pillow library. In some implementations, the automatic adjustment ML system 1705 can be used to generate tuning maps or knobs for tuning the ISP.

[0201] Key 2420 identifier Figure 21A The diagram illustrates how different neural network operations are described. For example, a convolution with a 3x3 filter and a stride of 1 is represented by a thick white arrow drawn in black and pointing to the right. A convolution with a 2x2 filter and a stride of 2 is represented by a thick black arrow pointing downwards. Upsampling (e.g., bilinear upsampling) is represented by a thick black arrow pointing upwards.

[0202] Figure 25AThis is a conceptual diagram illustrating an example of applying intensity adjustments to the first hue of an example input image to generate a modified image, based on some examples. Figure 25A The hue adjustment intensity of the modified image is 0.0.

[0203] Figure 25B This shows the application based on some examples. Figure 25A A conceptual diagram illustrating an example input image used to generate a modified image with a second-tone intensity adjustment. Figure 25A The hue adjustment intensity of the modified image is 0.5.

[0204] Figure 25A and Figure 25B This image illustrates an example of applying different tone levels using CLAHE implemented via OpenCV. For instance, the image can be divided into a fixed grid, and the histogram can be cropped with predefined values ​​before calculating the cumulative distribution function (CDF). The cropping constraints can determine the intensity of the local tone manipulation. Figure 25A and Figure 25B In the example, the application includes a maximum cropping limit of 0.5. Interpolation can be performed on the mesh transformation to produce the final result. As shown, changes in hue result in varying amounts of brightness adjustment, which will be applied to the pixels of the image processed by the ISP, potentially leading to an image with a darker or brighter hue.

[0205] Figure 26A This illustrates the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a modified image with first-level detail intensity adjustment. Figure 26A The intensity of detail adjustment (detail enhancement) for the modified image is 0.0.

[0206] Figure 26B This illustrates the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a second detail adjustment intensity for the modified image. Figure 26A The modified image's detail adjustment (detail enhancement) intensity is 0.5.

[0207] Figure 26A and Figure 26BThese images illustrate application examples of different detail enhancement methods implemented using OpenCV. For example, detail enhancement can be used to smooth the edges of the original image. Details can be expressed as described above regarding saturation maps. For example, image details can be obtained by subtracting a filtered image (e.g., a smoothed image produced by edge-preserving filtering) from the input image. Details can be enhanced at multiple scales. The range sigma can be equal to 0.05 (range sigma = 0.05). As mentioned above, spatial sigma is a hyperparameter of the bilateral filter. Spatial sigma can be used to control the intensity of detail enhancement. In some examples, the maximum spatial sigma can be set to 5.0 (maximum spatial sigma = 5.0). Figure 26A An example is provided where the space sigma is equal to 0, while Figure 26B An example is shown where the space sigma is equal to the maximum value of 5.0. As shown in the figure, with... Figure 26A Compared to the images in the middle, Figure 26B It features sharper details and more noise.

[0208] Figure 27A This illustrates the application based on some examples. Figure 25A A conceptual diagram illustrating the example input image used to generate a modified image with adjusted first color saturation intensity. Figure 27A The first color saturation value of the modified image is 0.0, which represents the maximum color desaturation intensity. Figure 27A Modified image representation in Figure 25A The grayscale version of the example input image.

[0209] Figure 27B This illustrates the application based on some examples. Figure 25A A conceptual diagram illustrating an example input image used to generate a modified image with a second color saturation adjustment intensity. Figure 27A The intensity of the second color saturation adjustment in the modified image is 1.0, indicating that the saturation has not changed (zero intensity saturation adjustment). Figure 27B The modified image in the image matches in terms of color saturation. Figure 25A Example input image.

[0210] Figure 27C This illustrates the application based on some examples. Figure 25A A conceptual diagram illustrating an example input image used to generate a modified image with third-color saturation adjustment intensity. Figure 27C The modified image has a third color saturation value of 2.0, which indicates the intensity of the maximum color saturation increase. Figure 27C Modified image representation in Figure 25A An oversaturated version of the example input image. In an attempt to demonstrate... Figure 27C The oversaturation properties of the modified image in the image, and Figures 27A-27B or Figure 25A In comparison, some elements are Figure 27C The center is highlighted.

[0211] Figure 27A , Figure 27B and Figure 27C This is an image showing application examples of different color saturation adjustments. Figure 27A The image in the image shows the result when a saturation of 0 is applied, leading to a grayscale (desaturated) image, similar to the one above. Figures 9-11 The image described. Figure 27B The image shown is the result of applying a saturation level of 1.0, which resulted in no change in saturation. Figure 27C The image shown is an example of a highly saturated image obtained when a saturation level of 2.0 is applied.

[0212] Figure 28 This is a block diagram illustrating an example of a machine learning (ML) image signal processor (ISP) that receives various tuning parameters (similar to the tuning diagram above) as input for tuning the ML ISP 2300. The ML ISP 2300 can be similar to the one described above. Figure 23 The ML ISP 2300 is discussed and performs similar operations. The ML ISP 2300 includes a trained machine learning (ML) model 2802. Based on parameter processing, the trained ML model 2802 of the ML ISP 2300 outputs an enhanced image based on tuning parameters. Image data 2804 (including the original image with simulated ISO noise and the radial distance from the center of the original image) is provided as input. Tuning parameters 2806 used to capture the original image include red channel gain, blue channel gain, ISO speed, and exposure time. Tuning parameters 2806 include tone enhancement intensity, detail enhancement intensity, and color saturation intensity. Additional tuning parameters 2806 are similar to the tuning diagram discussed above and can be applied at the image level (a single value is applied to the entire image) or at the pixel level (different values ​​are applied to each pixel in the image).

[0213] Figure 29 This is a block diagram illustrating an example of specific tuning parameter values ​​that can be provided to the ML ISP 2300, a machine learning (ML) image signal processor (ISP). The example raw input image is shown alongside an image representing the radial distance from the center of the raw image. The raw image data shown may, for example, represent image data for only one color channel (e.g., green). For tuning parameters 2816 used to capture the raw input image, the red gain is 1.957935, the blue gain is 1.703827, the ISO speed is 500, and the exposure time is 3.63e. -04The tuning parameters include a hue intensity of 0.0, a detail intensity of 0.0, and a saturation intensity of 1.0.

[0214] Figure 30 This is an additional block diagram illustrating specific tuning parameter values ​​that can be provided to the ML ISP 2300 image signal processor (ISP) for machine learning (ML). Figure 29 The same original input image, radial distance information, and image capture parameters (red gain, blue gain, ISO sensitivity, and exposure time) shown are displayed in [the image / data]. Figure 30 In the middle. And Figure 29 Compared to tuning parameter 2816, a different tuning parameter 2826 is provided. Tuning parameter 2826 includes a hue intensity of 0.2, a detail intensity of 1.9, and a saturation intensity of 1.9. Based on the differences in adjustment parameters, Figure 30 Output enhancement image 2828 and Figure 29 The output enhancement image 2808 has a different appearance compared to the previous one. For example, Figure 30 The output enhanced image 2828 is highly saturated and has finer details, while Figure 29 The enhanced image lacks saturation or detail enhancement. Because... Figures 28-30 Shown in grayscale instead of color, the red channels of output enhanced image 2828 and output enhanced image 2808 are... Figures 28-30 The diagram illustrates the changes in red saturation. Increased red saturation is represented by brighter areas, while decreased red saturation is represented by darker areas. Because... Figure 30 The red channel of the output enhanced image 2828 is shown, so the brighter flowers in the output enhanced image 2828 (compared to the output enhanced image 2808 and the original image data) indicate that the red in the flowers in the output enhanced image 2828 is more saturated.

[0215] The trained ML model 2802 of the ML ISP 2300 is trained to process tuning parameters (e.g., tuning maps or knobs) and generate enhanced or modified output images. Training settings may include a PyTorch implementation. The network receptive field may be set to 160x160. Training images may include thousands of images captured by one or more devices. Patch-based training can be performed, where patches of the input images (relative to the entire image) can be provided to the ML ISP 2300 during training. In some examples, the input patch size is 320x320, while the output patch size generated by the ML ISP 2300 is 160x160. Size reduction may be based on the convolutional nature of the ML ISP 2300's neural network (e.g., where the neural network is implemented as a CNN or other network using convolutional filters). In some cases, batches of image patches can be provided to the ML ISP 2300 at each training iteration. In one illustrative example, the batch size may include 128 images. The neural network can be trained until the loss is verified to be stable. In some cases, stochastic gradient descent (e.g., the Adam optimizer) can be used to train the network. In some cases, a learning rate of 0.0015 can be used. In some examples, the trained ML model 2802 can be implemented using one or more trained support vector machines, one or more trained neural networks, or a combination thereof.

[0216] Figure 31 This is a block diagram illustrating examples of objective functions and different losses that can be used during the training of the ML ISP 2300, a machine learning (ML) image signal processor (ISP). Pixel-wise losses include L1 and L2 losses. Pixel-wise losses can provide better color reproduction for the output images generated by the ML ISP 2300. Structural losses can also be used to train the MLISP 2300. Structural losses include the Structural Similarity Index (SSIM). For SSIM, a 7x7 window size and a Gaussian sigma of 1.5 can be used. Another structural loss that can be used is Multi-Scale SSIM (MS-SSIM). For MS-SSIM, a 7x7 window size, a Gaussian sigma of 1.5, and scale weights of [0.9, 0.1] can be used. Structural losses can be used to better preserve high-frequency information. Pixel-wise losses and structural losses can be used together, for example... Figure 31 The L1 loss and MS-SSIM are shown in the figure.

[0217] In some examples, the ML ISP 2300 can perform patch-by-patch model inference after the neural network of the ML ISP 2300 has been trained. For example, the input to the neural network may include one or more raw image patches (e.g., with Bayer patterns) from a raw image data frame, and the output may include an output RGB patch (or a patch with other color component representations, such as YUV). In an illustrative example, the neural network takes a 128x128 pixel raw image patch as input and produces an 8x8x3 RGB patch as the final output. Based on the convolutional properties of the various convolutional filters applied by the neural network, many pixel locations outside the 8x8 array of raw image patches are consumed by the network to generate the final 8x8 output patch. This reduction in data from input to output is due to the amount of context required to understand neighboring information to process pixels. Having a larger input raw image patch with all neighboring information and context helps in processing and producing a smaller output RGB patch.

[0218] In some examples, based on the reduction of pixel positions from input to output, the 128x128 original image patches are designed so that they overlap in the original input image. In such examples, the 8x8 outputs do not overlap. For example, for the first 128x128 original image patch in the top left corner of the original image frame, a first 8x8 RGB output patch is produced. The next 128x128 patch in the original image frame will be located 8 pixels to the right of the last 128x128 patch, and will therefore overlap with the last 128x128 pixel patch. The next 128x128 patch will be processed by a neural network to produce a second 8x8 RGB output patch. In the complete final output image, the second 8x8 RGB patch will be placed next to the first 8x8 RGB output patch (generated using the previous 128x128 original image patch). This process can be performed until the 8x8 patches that make up the complete output image are produced.

[0219] Figure 32 This is a conceptual diagram illustrating an example of patch-by-patch model inference based on some examples, leading to non-overlapping output patches at the first image location.

[0220] Figure 33 This is shown at the second image position according to some examples. Figure 32 A concept diagram of an example of patch-by-patch model inference.

[0221] Figure 34 This is shown at the position of the third image according to some examples. Figure 32 A concept diagram of an example of patch-by-patch model inference.

[0222] Figure 35This is shown at the fourth image position according to some examples. Figure 32 A concept diagram of an example of patch-by-patch model inference.

[0223] Figure 32 , Figure 33 , Figure 34 and Figure 35 This is a diagram illustrating an example of patch-by-patch model inference, which results in non-overlapping output patches. For example... Figures 32 to 35 As shown, the input patch size can be ko + 160, where ko is the output patch size (the output patch size is equal to ko). Pixels with diagonal stripes refer to fill pixels (e.g., reflective). Figures 32 to 35 As shown in the dashed box, the input tiles overlap. As shown in the white box, the output tiles do not overlap.

[0224] As mentioned above, tuning parameters (or tuning maps or tuning knobs) can be applied at the image level (a single value is applied to the entire image) or at the pixel level (different values ​​are applied to each pixel in the image).

[0225] Figure 36 This is a conceptual diagram illustrating an example of a spatially fixed tone map (or mask) applied at the image level, where a single value of t is applied to all pixels of the image. As mentioned above, tone maps that include different values ​​for each pixel of the image can be provided (e.g., such as...). Figure 8 (as shown), therefore it can vary in space.

[0226] Figure 37This is a conceptual diagram illustrating an example of processing input image data 3702 using a spatial variation diagram 3703 (or mask) to generate an output image 3715 with spatially varied saturation adjustment intensity. The spatial variation diagram 3703 includes a saturation diagram. The spatial variation diagram 3703 includes values ​​for each location corresponding to a pixel in the original image 3702 (also called a Bayer image) generated by the image sensor. In contrast to a spatially fixed diagram, the values ​​at different locations in the spatial variation diagram 3703 can be different. The spatial variation diagram 3703 includes a value of 0.5 corresponding to a pixel region in the original image 3702 depicting a flower in the foreground. This region with a value of 0.5 is shown in gray in the spatial variation diagram 3703. A value of 0.5 in the spatial variation diagram 3703 indicates that the saturation in this region remains unchanged, without increase or decrease. The spatial variation diagram 3703 includes a value of 0.0 corresponding to a pixel region in the original image 3702 depicting the background behind the flower. The region with a value of 0.0 is shown in black in spatial variation map 3703. A value of 0.0 in input saturation map 203 indicates that the saturation in this region will be completely desaturated. Modified image 3715 illustrates an example of an image generated (e.g., by image processing ML system 210) by applying saturation based on the intensity and orientation of spatial variation map 3703. Modified image 3705 shows an image where the flowers in the foreground remain saturated (at an unchanged saturation intensity) to the same degree as reference image 3716, but the background is completely desaturated and therefore depicted in grayscale. Reference image 3716 is used as the baseline image that the system expects to generate, where a neutral saturation (0.5) is applied across the entire original image 3702. Because Figure 37 Shown in grayscale instead of color, the green channel of the modified image 3705 and the reference image 3716 is... Figure 37 The diagram illustrates the changes in green saturation. Increased green saturation is represented by brighter areas, while decreased green saturation is represented by darker areas. Because... Figure 37 The green channel of the modified image 3705 is shown, and the background behind the flower is predominantly green. Therefore, the background appears dark in the modified image 3705 (compared to the reference image 3716), indicating that the green in the background of the modified image 3705 is desaturated (compared to the reference image 3716).

[0227] Figure 38This is a conceptual diagram illustrating an example application of spatial variation map 3803 to process input image data 3802 to generate an output image 3815 with spatial variation hue adjustment intensity and spatial variation detail adjustment intensity, based on some examples. Spatial variation map 3703 includes a hue map. Spatial variation map 3803 includes values ​​for each location corresponding to a pixel in the original image 3802 (also called a Bayer image) generated by the image sensor. The values ​​at different locations in spatial variation map 3803 can be different. Spatial variation map 3803 includes a hue value of 0.5 and a detail value of 5.0, corresponding to a pixel region in the original image 3802 depicting a flower in the foreground. This region with a hue value of 0.5 and a detail value of 5.0 is represented in white in spatial variation map 3803. A hue value of 0.5 in spatial variation map 3803 indicates a gamma (γ) value of 1.0 (corresponding to 10). 0.0 This means that the hue of the area remains the same and unchanged. A detail value of 5.0 in spatial variation graph 3803 indicates that the detail in this area will increase. Spatial variation graph 3803 includes a hue value of 0.0 and a detail value of 0.0, which corresponds to the pixel area depicting the background behind the flower in the original image 3802. This area with a hue value of 0.0 and a detail value of 0.0 is shown in black in spatial variation graph 3803. A hue value of 0.0 in input hue graph 203 indicates a gamma (γ) value of 2.5 (corresponding to 10). 0.4 This means that the tone of the area will darken. Figure 38 The mapping direction from hue value to gamma value in the middle Figure 7 The mapping from hue values ​​to gamma values ​​is in the opposite direction, where 0 brightens and 1 darkens. A detail value of 5.0 in spatial variation diagram 3803 indicates that detail in that area will be reduced. Modified image 3815 illustrates an example of an image generated by modifying hue and detail based on the intensity and direction of spatial variation diagram 3803 (e.g., by image processing ML system 210). Modified image 3805 represents an image in which the flowers in the foreground have the same hue (compared to reference image 3816) but increased detail (compared to reference image 3816), and where the background has a darker hue (compared to reference image 3816) and reduced detail (compared to reference image 3816). Reference image 3816 is used as the baseline image that the system expects to produce, where a neutral hue (0.5) and neutral detail are applied throughout the original image 3802.

[0228] Figure 39 This is a conceptual diagram illustrating the automatically adjusted image generated by an image processing system using one or more tuning diagrams to adjust the input image.

[0229] Figure 40AThis is a conceptual diagram illustrating output images generated by an image processing system using one or more spatially varied tuning maps (which are generated using an automated tuning machine learning (ML) system) based on some examples. The input and output images are in... Figure 40A As shown in the figure, the output image is a modified variant of the input image based on one or more spatial variation tuning maps.

[0230] Figure 40B This shows examples that can be used based on... Figure 40A The input image shown generates Figure 40A The example concept diagram of the spatial variation tuning plot of the output image is shown. In addition to showing again... Figure 40A In addition to the input and output images, Figure 40B It also shows detail images, noise images, tone images, saturation images, and hue images.

[0231] Figure 41A This is a conceptual diagram illustrating output images generated by an image processing system using one or more spatially varied tuning maps (which are generated using an automated tuning machine learning (ML) system) based on some examples. The input and output images are in... Figure 41A As shown in the figure, the output image is a modified variant of the input image based on one or more spatial variation tuning maps.

[0232] Figure 41B This shows examples that can be used based on... Figure 41A The input image shown generates Figure 41A The diagram shows an example of a spatial variation tuning plot of the output image. (Further explanation follows.) Figure 41A The input image, Figure 41B It also shows sharpness, noise, hue, saturation, and hue graphs.

[0233] Figure 42A This is a flowchart illustrating an example of a process 4200 using one or more neural networks to process image data using the techniques described herein. Process 4200 can be performed by an imaging system. The imaging system may include, for example, an image capture and processing system 100, an image capture device 105A, an image processing device 105B, an image processor 150, an ISP 154, a host processor 152, and an image processing ML system 210. Figure 14A The system Figure 14B The system Figure 14C The system Figure 14DThe system includes an image processing system 1406, an automatic adjustment ML system 1405, a downsampler 1418, an upsampler 1426, a neural network 1500, a neural network architecture 1900, a neural network architecture 2000, a neural network architecture 2100A, a neural network architecture 2100B, a spatial attention engine 2110, a channel attention engine 2115, an image sensor 2202, a pre-tuned ISP 2208, a machine learning (ML) ISP 2300, a neural network system 2303, a preprocessing engine 2307, an input interface 2301, an output interface 2305, a neural network architecture 2400, a trained machine learning model 2802, an imaging system 4200, a computing system 4300, or a combination thereof.

[0234] In block 4202, process 4200 includes an imaging system acquiring image data. In some embodiments, the image data includes a processed image having multiple color components for each pixel of the image data. For example, the image data may include one or more RGB images, one or more YUV images, or other color images previously captured and processed by a camera system, such as those obtained by... Figure 1 The image capture and processing system 100 shown generates an image of scene 110. In some embodiments, the image data includes raw image data from one or more image sensors (e.g., image sensor 130). The raw image data includes a single color component for each pixel of the image data. In some cases, the raw image data is obtained from one or more image sensors and filtered by a color filter array such as a Bayer color filter array. In some embodiments, the image data includes one or more image data patches. The image data patches include a subset of image data frames. In some cases, generating a modified image includes generating multiple output image data patches. Each output image data patch may include a subset of pixels of the output image.

[0235] Examples of raw image data may include image data captured using image capture and processing system 100, image data captured using image capture device 105A and / or image processing device 105B, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, and downsampled input image 1616. Figure 17 I RGB , Figure 17 Brightness channel I y1820 Luminosity Channel, 1825 Image Patch, 1826 Complete Image Input, 2130 Input Image, 2204 Raw Input Data After Color Filtering, 2209 Output Color Image, 2304 Frame of Raw Image Data, 2306 Raw Image Patch, 2308 RGB Output, 2309 Final Output Image Figure 24 The original input image, Figure 24 Output RGB image Figures 25A-25B Input image, Figures 26A-26B Input image, Figures 27A-27C Input image, image data 2804, image data 2814 Figure 32-35 Input patch, original image 3702, reference image 3716, original image 3802, reference image 3816 Figure 39 Input image, Figures 40A-40B Input image, Figures 41A-41B The input image, image data of box 4252, other image data described herein, other images described herein, or combinations thereof. In some examples, box 4202 of process 4200 may correspond to box 4252 of process 4250.

[0236] In box 4204, process 4200 includes an imaging system acquiring one or more maps. These one or more maps are also referred to herein as one or more tuning maps. Each of the one or more maps is associated with a corresponding image processing function. Each map also includes a value indicating an amount of image processing function to be applied to the corresponding pixel of the image data. For example, a map in one or more maps includes multiple values ​​and is associated with an image processing function, wherein each of the multiple values ​​of the map indicates an amount of image processing function to be applied to the corresponding pixel of the image data. The above discussion... Figure 3A and Figure 3B Example tuning diagram is shown. Figure 3B The values ​​of the tuning diagram 303 and the example input image ( Figure 3A The corresponding pixels of the input image 302).

[0237] In some cases, one or more image processing functions associated with one or more images include noise reduction, sharpness adjustment, hue adjustment, saturation adjustment, hue adjustment, or any combination thereof. In some examples, multiple tuned images, each associated with an image processing function, may be obtained. For example, one or more images may include multiple images, wherein a first image among the multiple images is associated with a first image processing function, and a second image among the multiple images is associated with a second image processing function. In some cases, the first image may include a first plurality of values, wherein each of the first plurality of values ​​in the first image indicates an amount of the first image processing function to be applied to a corresponding pixel of the image data. The second image may include a second plurality of values, wherein each of the second plurality of values ​​in the second image indicates an amount of the second image processing function to be applied to a corresponding pixel of the image data. In some cases, the first image processing function associated with the first image among the multiple images includes one of noise reduction, sharpness adjustment, hue adjustment, saturation adjustment, and hue adjustment, and the second image processing function associated with the second image among the multiple images includes different functions among noise reduction, sharpness adjustment, hue adjustment, saturation adjustment, and hue adjustment. In an illustrative example, a noise map associated with the noise reduction function can be obtained, a sharpness map associated with the sharpness adjustment function can be obtained, a tone map associated with the tone adjustment function can be obtained, a saturation map associated with the saturation adjustment function can be obtained, and a hue map associated with the hue adjustment function can be obtained.

[0238] In some examples, the imaging system may use a machine learning system to generate one or more graphs, as shown in box 4254. The machine learning system may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The machine learning system and / or any machine learning elements listed above (which may be part of the machine learning system) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof. The machine learning system that generates one or more graphs in box 4204 may be the same machine learning system that generates the modified image in box 4206. The machine learning system that generates one or more graphs in box 4204 may be a different machine learning system from the machine learning system that generates the modified image in box 4206.

[0239] Examples of one or more graphs include input saturation graph 203, spatial tuning graph 303, input noise graph 404, input sharpening graph 604, input hue graph 804, input saturation graph 1104, input hue graph 1304, spatial variation tuning graph 1404, small spatial variation tuning graph 1424, output layer 1514, input tuning graph 1607, output tuning graph 1618, and reference output tuning graph 1619. Figure 17 Spatial variation tuning diagrams for Omega (Ω), tuning diagram 1804, tuning parameters 2206, 2806, 2616, and 2826. Figure 36 Color tone diagram, spatial variation diagram 3703, spatial variation diagram 3803. Figure 40B Detailed images Figure 40B noise graph Figure 40B color palette, Figure 40B Saturation map Figure 40B Hue diagram Figure 41B Sharpness map Figure 41B noise graph Figure 41B color palette, Figure 41B Saturation map Figure 41B The hue map, one or more maps of box 4254, other spatial variation maps described herein, other spatial fixed maps described herein, other maps described herein, other masks described herein, or any combination thereof. Examples of image processing functions include noise reduction, noise addition, sharpness adjustment, detail adjustment, hue adjustment, color saturation adjustment, hue adjustment, any other image processing function described herein, any other image processing parameter described herein, or any combination thereof. In some examples, box 4204 of process 4200 may correspond to box 4254 of process 4250.

[0240] In box 4206, process 4200 includes an imaging system using image data and one or more graphs as input to a machine learning system to generate a modified image. The modified image may be referred to as the output image. The modified image includes features based on a corresponding image processing function associated with each of the one or more graphs. In some cases, the machine learning system includes at least one neural network. In some examples, process 4200 includes using an additional machine learning system, different from the machine learning system used to generate the modified image, to generate one or more graphs. For example, Figure 17 The automatic adjustment ML system 1705 can be used to generate one or more graphs, and the image processing ML system 210 can be used to generate a modified image based on an input image and one or more graphs.

[0241] Examples of modified images include modified image 215, modified image 415, image 510, image 511, image 512, modified image 615, output image 710, output image 711 and output image 712, modified image 815, output image 1020, output image 1021, output image 1022, modified image 1115, modified image 1315, modified image 1415, output layer 1514, output image 1608, reference output image 1609, output color image 2209, RGB output 2308, and final output image 2309. Figure 24 Output RGB image Figures 25A-25B The modified image Figures 26A-26B The modified image Figures 27A-27C Modified image, output enhanced image 2808, output enhanced image 2828 Figure 32-35 Output patch, output image 3715, reference image 3716, output image 3815, reference image 3816. Figure 39 Automatic image adjustment Figures 40A-40B Output image Figure 41A The output image, the modified image of box 4206, other image data described herein, other images described herein, or combinations thereof. In some examples, box 4206 of process 4200 may correspond to box 4256 of process 4250.

[0242] The machine learning system of Box 4206 may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The machine learning system of Box 4206 and / or any machine learning elements listed above (which may be part of the machine learning system) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0243] Figure 42B This is a flowchart illustrating an example of a process 4250 using one or more neural networks to process image data using the techniques described herein. Process 4250 can be performed by an imaging system. The imaging system may include, for example, an image capture and processing system 100, an image capture device 105A, an image processing device 105B, an image processor 150, an ISP 154, a host processor 152, and an image processing ML system 210. Figure 14AThe system Figure 14B The system Figure 14C The system Figure 14D The system includes an image processing system 1406, an automatic adjustment ML system 1405, a downsampler 1418, an upsampler 1426, a neural network 1500, a neural network architecture 1900, a neural network architecture 2000, a neural network architecture 2100A, a neural network architecture 2100B, a neural network architecture 2100C, a neural network architecture 2100D, a spatial attention engine 2110, a channel attention engine 2115, an image sensor 2202, a pre-tuned ISP 2208, a machine learning (ML) ISP 2300, a neural network system 2303, a preprocessing engine 2307, an input interface 2301, an output interface 2305, a neural network architecture 2400, a trained machine learning model 2802, an imaging system 4200, a computing system 4300, or a combination thereof.

[0244] In block 4252, process 4250 includes the imaging system acquiring image data. In some embodiments, the image data includes a processed image having multiple color components for each pixel of the image data. For example, the image data may include one or more RGB images, one or more YUV images, or other color images previously captured and processed by the camera system, such as those obtained by... Figure 1 The image of scene 110 generated by the image capture and processing system 100 shown is illustrated. Examples of image data may include image data captured using the image capture and processing system 100, image data captured using image capture device 105A and / or image processing device 105B, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, and downsampled input image 1616. Figure 17 I RGB , Figure 17 Brightness channel I y 1820 Luminosity Channel, 1825 Image Patch, 1826 Complete Image Input, 2130 Input Image, 2204 Raw Input Data for Color Filtering, 2209 Output Color Image, 2304 Frame of Raw Image Data, 2306 Raw Image Patch, 2308 RGB Output, 2309 Final Output Image Figure 24 The original input image, Figure 24 Output RGB image Figures 25A-25B Input image, Figures 26A-26B Input image, Figures 27A-27C Input image, image data 2804, image data 2814 Figure 32-35 Input patch, original image 3702, reference image 3716, original image 3802, reference image 3816 Figure 39 Input image, Figures 40A-40B Input image, Figures 41A-41B The input image, raw image data of box 4202, other image data described herein, other images described herein, or combinations thereof. The imaging system may include an image sensor that captures image data. The imaging system may include an image sensor connector coupled to the image sensor. Obtaining image data in box 4252 may include obtaining image data from the image sensor and / or through the image sensor connector.

[0245] In some examples, the image data has an input image that has multiple color components for each of the multiple pixels of the image data. The input image may have been at least partially processed, for example, by de-mosaicing at ISP 154. The input image can be captured and / or processed using image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, image sensor 130, ISP 154, host processor 152, image sensor 2202, or combinations thereof. In some examples, the image data includes raw image data from one or more image sensors that includes at least one color component for each of the multiple pixels of the image data. Examples of one or more image sensors may include image sensor 130 and image sensor 2202. Examples of raw image data may include image data captured using image capture and processing system 100, image data captured using image capture device 105A, input image 202, input image 302, input image 402, input image 602, input image 802, input image 1102, input image 1302, input image 1402, downsampled input image 1422, input layer 1510, input image 1606, and downsampled input image 1616. Figure 17 I RGB , Figure 17 Brightness channel I y 1820 Luminosity Channel, 1825 Image Patch, 1826 Complete Image Input, 2130 Input Image, 2204 Raw Input Data for Color Filtering, 2304 Frame of Raw Image Data, 2306 Raw Image Patch Figure 24 The original input image, Figures 25A-25B Input image, Figures 26A-26B Input image, Figures 27A-27C Input image, image data 2804, image data 2814 Figure 32-35Input patch, original image 3702, reference image 3716, original image 3802, reference image 3816 Figure 39 Input image, Figures 40A-40B Input image, Figures 41A-41B The input image, the raw image data of box 4202, other raw image data described herein, other image data described herein, other images described herein, or combinations thereof. The image data may include a single color component (e.g., red, green, or blue) for each pixel of the image data. The image data may include multiple color components (e.g., red, green, and / or blue) for each pixel of the image data. In some cases, the image data is obtained from one or more image sensors filtered by a color filter array, such as a Bayer color filter array. In some examples, the image data includes one or more image data patches. Image data patches include subsets of image data frames, such as corresponding to one or more contiguous regions or areas. In some cases, generating a modified image includes generating multiple output image data patches. Each output image data patch may include a subset of pixels of the output image. In some examples, box 4252 of process 4250 may correspond to box 4202 of process 4200.

[0246] At box 4254, process 4250 includes an imaging system using image data as input to one or more trained neural networks to generate one or more graphs. Each of the one or more graphs is associated with a corresponding image processing function. In some examples, box 4254 of process 4250 may correspond to box 4204 of process 4200. Examples of one or more trained neural networks include image processing ML system 210, auto-tuning ML system 1405, neural network 1500, auto-tuning ML system 1705, local neural network 1806, global neural network 1807, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, neural network system 2303, neural network architecture 2400, and machine learning ISP. 2300. Trained machine learning model; 2802. One or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, any other trained neural network described herein, any other neural network described herein, any other trained machine learning model described herein, any other machine learning model described herein, any other trained machine learning system described herein, any other machine learning system described herein, or any combination thereof. One or more trained neural networks in box 4254, and / or any machine learning elements listed above (which may be part of one or more trained neural networks in box 4254) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof.

[0247] Examples of one or more graphs include input saturation graph 203, spatial tuning graph 303, input noise graph 404, input sharpening graph 604, input hue graph 804, input saturation graph 1104, input hue graph 1304, spatial variation tuning graph 1404, small spatial variation tuning graph 1424, output layer 1514, input tuning graph 1607, output tuning graph 1618, and reference output tuning graph 1619. Figure 17 Spatial variation tuning diagrams for Omega (Ω), tuning diagram 1804, tuning parameters 2206, 2806, 2616, and 2826. Figure 36 Color tone diagram, spatial variation diagram 3703, spatial variation diagram 3803. Figure 40B Detailed images Figure 40Bnoise graph Figure 40B color palette, Figure 40B Saturation map Figure 40B Hue diagram Figure 41B Sharpness map Figure 41B noise graph Figure 41B color palette, Figure 41B Saturation map Figure 41B The hue map, one or more maps of box 4204, other spatial variation maps described herein, other spatial fixed maps described herein, other maps described herein, other masks described herein, or any combination thereof. Examples of image processing functions include noise reduction, noise addition, sharpness adjustment, detail adjustment, hue adjustment, color saturation adjustment, hue adjustment, any other image processing function described herein, any other image processing parameter described herein, or any combination thereof.

[0248] Each of the one or more graphs may include multiple values ​​and may be associated with image processing functions. Examples of multiple values ​​may include values ​​V0-V56 of spatial tuning graph 303. Examples of multiple values ​​may include different values ​​corresponding to different regions of the following: input saturation graph 203, input noise graph 404, input sharpening graph 604, input hue graph 804, input saturation graph 1104, input hue graph 1304, spatial variation tuning graph 1404, small spatial variation tuning graph 1424, spatial variation graph 3703, and spatial variation graph 3803. Figure 40B Detailed images Figure 40B noise graph Figure 40B color palette, Figure 40B Saturation map Figure 40B Hue diagram. Figure 41B Sharpness map Figure 41B noise graph Figure 41B color palette, Figure 41B Saturation map Figure 41BThe image includes a hue diagram and one or more diagrams in box 4204. Each of the multiple values ​​in the diagram can indicate the intensity of the image processing function applied to the corresponding region of the image data. In some examples, a higher value (e.g., 1) can indicate a stronger intensity of the image processing function applied to the corresponding region of the image data, while a lower value (e.g., 0) can indicate a weaker intensity of the image processing function applied to the corresponding region of the image data. In some examples, a value (e.g., 0.5) can indicate that the image processing function is applied to the corresponding region of the image data with zero intensity, which may result in no application of the image processing function and therefore no modification to the image data using the image processing function. The corresponding region of the image data may correspond to the pixels of the image data in box 4202 and / or the pixels of the image in box 4256. The corresponding region of the image data can correspond to multiple pixels of the image data in box 4202 and / or multiple pixels of the image in box 4256, for example, when multiple pixels are binned, such as to form superpixels, resample, resize, and / or rescale. The corresponding region of the image data can be a contiguous region. Multiple pixels can be in a contiguous region. Multiple pixels can be adjacent to each other.

[0249] Each of the multiple values ​​in the graph can indicate the direction in which image processing functions are applied to a corresponding region of the image data. In some examples, values ​​above a threshold (e.g., above 0.5) can indicate a positive direction for applying image processing functions to a corresponding region of the image data, while values ​​below a threshold (e.g., below 0.5) can indicate a negative direction. The corresponding region of the image data can correspond to pixels of the image data in box 4202 and / or pixels of the image in box 4256. The corresponding region of the image data can correspond to multiple pixels of the image data in box 4202 and / or multiple pixels of the image in box 4256, for example, in the case where multiple pixels are merged, such as to form superpixels, resample, resize, and / or rescale. The corresponding region of the image data can be a contiguous region. Multiple pixels can be in a contiguous region. Multiple pixels can be adjacent to each other.

[0250] In some examples, a positive direction for saturation adjustment may result in oversaturation, while a negative direction may result in undersaturation or desaturation. In some examples, a positive direction for hue adjustment may result in increased luminance and / or brightness, while a negative direction may result in decreased luminance and / or brightness (or vice versa). In some examples, a positive direction for sharpness adjustment may result in increased sharpness (e.g., reduced blur), while a negative direction may result in decreased sharpness (e.g., increased blur) (or vice versa). In some examples, a positive direction for noise adjustment may result in reduced noise, while a negative direction for hue adjustment may result in noise generation and / or addition (or vice versa). In some examples, the positive and negative directions for hue adjustment may correspond to different angles of the hue modification vector 1218, such as... Figure 12B As shown in the conceptual diagram 1216. For example, the positive and negative directions used for hue adjustment can correspond to the positive and negative angles θ of the hue modification vector 1218, as... Figure 12B The concept is illustrated in Figure 1216. In some examples, one or more of the image processing functions may be limited to processing in the positive direction as discussed herein, or may be limited to processing in the negative direction as discussed herein.

[0251] One or more images may include multiple images. A first image among the multiple images may be associated with a first image processing function, while a second image among the multiple images may be associated with a second image processing function. The second image processing function may be different from the first image processing function. In some examples, one or more image processing functions associated with at least one image among the multiple images include at least one of noise reduction, sharpness adjustment, detail adjustment, tone adjustment, saturation adjustment, hue adjustment, or combinations thereof. In an illustrative example, a noise image associated with a noise reduction function, a sharpness image associated with a sharpness adjustment function, a tone image associated with a tone adjustment function, a saturation image associated with a saturation adjustment function, and a hue image associated with a hue adjustment function can be obtained.

[0252] The first image may include a first plurality of values, wherein each of the first plurality of values ​​in the first image indicates the intensity and / or direction of applying a first image processing function to a corresponding region of the image data. The second image may include a second plurality of values, wherein each of the second plurality of values ​​in the second image indicates the intensity and / or direction of applying a second image processing function to a corresponding region of the image data. The corresponding region of the image data may correspond to pixels of the image data in box 4202 and / or pixels of the image in box 4256. The corresponding region of the image data may correspond to multiple pixels of the image data in box 4202 and / or multiple pixels of the image in box 4256, for example, in the case of multiple pixels being merged, such as to form superpixels, resampling, resizing, and / or rescaling. The corresponding region of the image data may be a contiguous region. Multiple pixels may be in a contiguous region. Multiple pixels may be adjacent to each other.

[0253] In some examples, the image data includes luminance channel data corresponding to the image. Examples of luminance channel data may include... Figure 17 Brightness channel I y And / or luminance channel 1820. Using image data as input to one or more trained neural networks may include using luminance channel data corresponding to the image as input to one or more trained neural networks. In some examples, generating an image based on image data includes generating an image based on luminance channel data and chrominance data corresponding to the image.

[0254] In some examples, one or more trained neural networks output one or more affine coefficients based on using image data as input to the one or more trained neural networks. Examples of affine coefficients include... Figure 17 The affine coefficients a and b. Examples of one or more trained neural networks that output one or more affine coefficients include autotuning ML system 1705, neural network architecture 1900, neural network architecture 2000, neural network architecture 2100A, neural network architecture 2100B, neural network architecture 2100C, neural network architecture 2100D, or combinations thereof. In some examples, autotuning ML system 1405 may generate affine coefficients (e.g., a and / or b) and may generate a spatially varied tuning map 1404 and / or a small spatially varied tuning map 1424 based on the affine coefficients. Generating one or more maps may include generating a first map by transforming image data using at least one or more affine coefficients. The image data may include luminance channel data corresponding to the image, thereby transforming the image data using one or more affine coefficients includes transforming the luminance channel data using one or more affine coefficients. For example, Figure 17 The diagram illustrates the use of Figure 17 The affine coefficients a and b from the luminance channel I y Transform the luminance channel data according to the equation Ω = a * Iy +b generates the tuning diagram Ω. One or more affine coefficients can include multipliers, for example... Figure 17 The multiplier a. Transforming image data using one or more affine coefficients may include multiplying the brightness values ​​of at least a subset of the image data by a multiplier. One or more affine coefficients may include offsets, such as... Figure 17 The offset 'a'. Transforming image data using one or more affine coefficients may include offsetting the brightness values ​​of at least a subset of the image data by this offset. In some examples, one or more trained neural networks also output one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first image with one or more gradients in the image data. Examples of local linearity constraints may include local linearity constraint 1720, which is based on the equation In some examples, applying one or more affine coefficients to at least one subset of image data (e.g., brightness data) to transform at least that subset of image data (e.g., brightness data) can be controlled by local linearity constraints.

[0255] In some examples, one or more trained neural networks in box 4254 directly generate and / or output one or more graphs in response to receiving image data as input to the one or more trained neural networks. For example, one or more trained neural networks may directly generate and / or output one or more graphs based on an image without generating and / or outputting affine coefficients. For example, an autotuning ML system 1405 may directly generate and / or output a spatially varied tuned graph 1404 and / or a small spatially varied tuned graph 1424. In some examples, various neural network and / or machine learning systems shown and / or described herein for generating and using affine coefficients may be modified to instead directly output one or more graphs without first generating affine coefficients. For example, Figure 17 and / or Figure 18 The automatic adjustment ML system 1705 can be modified to directly generate and / or output tuning graph 1804 as its output layer, instead of or appended to the generation and / or output affine coefficients 1822. The neural network architecture 1900 can be modified to directly generate and / or output tuning graph 1804 as its output layer, instead of and / or appended to the generation and / or output affine coefficients 1822. The neural network architecture 2100A can be modified to directly generate and / or output tuning graph 1804 as its output layer, instead of and / or appended to the generation and / or output affine coefficients 1822. The neural network architecture 2100B can be modified to directly generate and / or output tuning graph 1804 as its output layer, instead of or appended to the generation and / or output affine coefficients 1822.

[0256] Each graph in one or more graphs can vary spatially. Each graph in one or more graphs can vary spatially based on different types or categories of objects depicted in different regions of the image data. Examples of such graphs include input saturation graph 203, spatial variation graph 3703, and spatial variation graph 3803, where regions of image data depicting foreground flowers and regions of image data depicting the background are mapped to different intensities and / or directions of the application of one or more corresponding image processing functions. In other examples, graphs can indicate that image processing functions can be applied with different intensities and / or directions in regions depicting different types or categories of objects, such as regions depicting people, faces, clothing, plants, sky, water, clouds, buildings, displays, metal, plastic, concrete, bricks, hair, trees, textured surfaces, any other type or category of objects discussed herein, or combinations thereof. Each graph in one or more graphs can vary spatially based on different colors in different regions of the image data. Each graph in one or more graphs can vary spatially based on different image attributes in different regions of the image data. For example, input noise graph 404 indicates that no denoising is performed in the already smoothed areas of input image 402, but indicates that relatively strong denoising is performed in the noisy areas of input image 402.

[0257] In box 4256, process 4250 includes an imaging system generating an image based on image data and one or more graphs. The image includes features based on corresponding image processing functions associated with each of the one or more graphs. The image may be referred to as an output image, a modified image, or a combination thereof. Examples of images include modified image 215, modified image 415, image 510, image 511, image 512, modified image 615, output image 710, output image 711 and output image 712, modified image 815, output image 1020, output image 1021, output image 1022, modified image 1115, modified image 1315, modified image 1415, output layer 1514, output image 1608, reference output image 1609, output color image 2209, RGB output 2308, and final output image 2309. Figure 24 Output RGB image Figures 25A-25B The modified image Figures 26A-26B The modified image Figures 27A-27C Modified image, output enhanced image 2808, output enhanced image 2828 Figure 32-35 Output patch, output image 3715, reference image 3716, output image 3815, reference image 3816. Figure 39 Automatic image adjustment Figures 40A-40B Output image Figure 41AThe output image, the modified image of box 4206, other image data described herein, other images described herein, or combinations thereof. In some examples, box 4256 of process 4250 may correspond to box 4206 of process 4200. Images can be generated using image processors such as image capture and processing system 100, image processing device 105B, image processor 150, ISP 154, host processor 152, image processing ML system 210, image processing system 1406, neural network 1500, machine learning (ML) ISP 2300, neural network system 2303, preprocessing engine 2307, input interface 2301, output interface 2305, neural network architecture 2400, trained machine learning model 2802, machine learning system of box 4206, or combinations thereof.

[0258] In some examples, generating an image based on image data and one or more graphs includes using the image data and one or more graphs as input to a second set of one or more trained neural networks, distinct from the first set of trained neural networks. Examples of the second set of one or more trained neural networks may include an image processing ML system 210, an image processing system 1406, a neural network 1500, a machine learning (ML) ISP 2300, a neural network system 2303, a neural network architecture 2400, a trained machine learning model 2802, a machine learning system of box 4206, or a combination thereof. In some examples, generating an image based on image data and one or more graphs includes de-mosaicing the image data using the second set of one or more trained neural networks.

[0259] The second group of one or more trained neural networks may include one or more trained convolutional neural networks (CNNs), one or more CNNs, one or more trained neural networks (NNs), one or more NNs, one or more trained support vector machines (SVMs), one or more SVMs, one or more trained random forests, one or more random forests, one or more trained decision trees, one or more decision trees, one or more trained gradient boosting algorithms, one or more gradient boosting algorithms, one or more trained regression algorithms, one or more regression algorithms, or combinations thereof. The second group of one or more trained neural networks and / or any machine learning elements listed above (which may be part of the second group of one or more trained neural networks) may be trained using supervised learning, unsupervised learning, reinforcement learning, deep learning, or combinations thereof. In some examples, the second group of one or more trained neural networks may be one or more trained neural networks of box 4254. In some examples, the second group of one or more trained neural networks may share at least one trained neural network common to one or more trained neural networks of box 4254. In some examples, the second group of one or more trained neural networks may be different from one or more trained neural networks of box 4254.

[0260] The imaging system may include a display screen. The imaging system may include a display screen connector coupled to the display screen. The imaging system may display images on the display screen, for example, by sending images to the display screen (and / or display controller) using the display screen connector. The imaging system may include a communication transceiver, which may be wired and / or wireless. The imaging system may use the communication transceiver to transmit images to a receiving device. The receiving device may be any type of computing system 4300 or any component thereof. The receiving device may be an output device, such as a device having a display screen and / or projector capable of displaying images.

[0261] In some aspects, the imaging system may include: a unit for acquiring image data; a unit for generating one or more graphs using the image data as input to one or more trained neural networks, wherein each of the one or more graphs is associated with a corresponding image processing function; and a unit for generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs. In some examples, the unit for acquiring image data may include an image sensor 130, an image capture device 105A, an image capture and processing system 100, an image sensor 2202, or a combination thereof. In some examples, the unit for generating one or more graphs may include an image processing device 105B, an image capture and processing system 100, an image processor 150, an ISP 154, a host processor 152, an image processing ML system 210, an auto-adjustment ML system 1405, a neural network 1500, an auto-adjustment ML system 1705, a local neural network 1806, a global neural network 1807, a neural network architecture 1900, a neural network architecture 2000, a neural network architecture 2100A, a neural network architecture 2100B, a neural network architecture 2100C, a neural network architecture 2100D, a neural network system 2303, a neural network architecture 2400, a machine learning ISP 2300, a trained machine learning model 2802, or a combination thereof. In some examples, the unit for generating the modified image may include an image processing device 105B, an image capture and processing system 100, an image processor 150, an ISP 154, a host processor 152, an image processing ML system 210, an image processing system 1406, a neural network 1500, a machine learning (ML) ISP 2300, a neural network system 2303, a preprocessing engine 2307, an input interface 2301, an output interface 2305, a neural network architecture 2400, a trained machine learning model 2802, a machine learning system of box 4206, or a combination thereof.

[0262] In some examples, the processes described herein (e.g., processes 4200, 4250, and / or other processes described herein) may be performed by a computing device or apparatus. In one example, process 4200 and / or process 4250 may be performed by... Figure 2 The image processing ML system 210 executes the process. In another example, process 4200 and / or process 4250 can be performed by a system with... Figure 43 The computing device of the computing system 4300 shown performs the operation. For example, it has... Figure 43 The computing device of the computing system 4300 shown may include components of the image processing ML system 210 and can implement... Figure 43 The operation.

[0263] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, wearable devices (e.g., VR headsets, AR headsets, AR glasses, connected watches or smartwatches or other wearable devices), server computers, computing devices of autonomous vehicles or autonomous vehicles, robotic devices, televisions and / or any other computing device with the resources to perform processes (including processes 4200 and / or 4250 described herein). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors and / or other components configured to perform the process steps described herein. In some examples, a computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.

[0264] Components of a computing device can be implemented in circuitry. For example, components may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof and / or be implemented using them to perform the various operations described herein.

[0265] Processes 4200 and 4250 are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media that performs the operation when executed by one or more processors. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement a process.

[0266] Furthermore, the processes 4200, 4250, and / or other processes described herein can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As described above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0267] Figure 43 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. In particular, Figure 43 An example of a computing system 4300 is illustrated, which can be any computing device constituting, for example, an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 4305. Connection 4305 can be a physical connection using a bus, or a direct connection to processor 4310, such as in a chipset architecture. Connection 4305 can also be a virtual connection, a network connection, or a logical connection.

[0268] In some embodiments, the computing system 4300 is a distributed system, wherein the functions described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functions described for that component. In some embodiments, a component may be a physical or virtual device.

[0269] Example system 4300 includes at least one processing unit (CPU or processor) 4310 and connections 4305 that couple various system components, including system memory 4315 (e.g., read-only memory (ROM) 4320 and random access memory (RAM) 4325), to processor 4310. Computing system 4300 may include a cache 4312 of high-speed memory, which is directly connected to, adjacent to, or integrated into processor 4310.

[0270] Processor 4310 may include any general-purpose processor and hardware or software services, such as services 4332, 4334, and 4336 stored in storage device 4330, which are configured to control processor 4310 and dedicated processors in which software instructions are incorporated into the actual processor design. Processor 4310 may essentially be a completely self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0271] To enable user interaction, the computing system 4300 includes an input device 4345, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice input, etc. The computing system 4300 may also include an output device 4335, which can be one or more of a plurality of output mechanisms. In some cases, a multi-mode system allows the user to provide multiple types of input / output to communicate with the computing system 4300. The computing system 4300 may include a communication interface 4340, which typically governs and manages user input and system output. The communication interface may use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communications, including transceivers using: audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, etc. Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs Wireless signal transmission Low Energy (BLE) wireless signal transmission Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short-range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), global microwave access interoperability (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 4340 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 4300 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo Global Navigation Satellite System. There are no limitations on operation on any particular hardware configuration; therefore, the basic features described here can be easily replaced with improved hardware or firmware configurations as they are under development.

[0272] Storage device 4330 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape cassettes, flash memory cards, solid-state storage devices, digital versatile disks, magnetic tape cassettes, floppy disks, flexible disks, hard disks, magnetic tapes, magnetic stripes / strips, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, optical disc read-only memory (CD-ROM), rewritable optical disc (CD), digital video optical disc (DVD), Blu-ray disc (BDD), holographic disc, another optical medium, secure digital storage (SD) cards, micro secure digital storage (microSD) cards, memory. Cards, smart card chips, EMV chips, SIM cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette tape and / or combinations thereof.

[0273] Storage device 4330 may include software services, servers, services, etc., which enable the system to perform functions when processor 4310 executes code defining such software. In some embodiments, hardware services that perform specific functions may include software components stored in a computer-readable medium that are combined with necessary hardware components (e.g., processor 4310, connection 4305, output device 4335, etc.) to perform functions.

[0274] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include, but does not include, non-transitory media in which data can be stored, carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions that may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0275] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0276] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that embodiments can be practiced without these specific details. For clarity of explanation, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods implemented in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0277] The various embodiments can be described above as processes or methods depicted as flowcharts, diagrams, data flow diagrams, structural diagrams, or block diagrams. Although a flowchart can describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but there may be other steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.

[0278] The processes and methods described in the examples above can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a particular function or group of functions. Some of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information, and / or information created during the methods according to the examples include hard disks or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.

[0279] Devices implementing the processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablets or other small personal computers, personal digital assistants, rack-mount devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. As a further example, such functionality may also be implemented on a circuit board between different chips or different processes executed in a single device.

[0280] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functionality described in this disclosure.

[0281] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, while illustrative embodiments of this application have been described in detail herein, it should be understood that the inventive concept can be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations unless limited by prior art. Various features and aspects of the above applications can be used alone or in combination. Furthermore, embodiments can be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be appreciated that, in alternative embodiments, these methods may be performed in a different order than that described.

[0282] Those skilled in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of this specification.

[0283] When a component is described as being "configured" to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits), or any combination thereof.

[0284] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection, and / or other suitable communication interface).

[0285] The language of the claims, or other languages ​​that refer to “at least one” and / or “one or more” in a group, indicates that one or more members of the group (in any combination) satisfy the claims. For example, the claim language stating “at least one of A and B” means A, B, or A and B. In another example, the claim language stating “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one” in a group and / or “one or more” in a group does not limit the group to the items listed in the group. For example, the claim language stating “at least one of A and B” can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0286] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally according to their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0287] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or multi-purpose integrated circuit devices, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, the technology can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging materials. The computer-readable medium can include memory or data storage media, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, these technologies may be implemented at least in part through a computer-readable communication medium that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer, such as a propagating signal or wave.

[0288] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described herein. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

[0289] The illustrative aspects of this disclosure include:

[0290] Aspect 1. A method for processing image data, the method comprising: acquiring image data; obtaining one or more graphs, wherein each of the one or more graphs is associated with a corresponding image processing function; and using the image data and the one or more graphs as input to a machine learning system to generate a modified image, the modified image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0291] Aspect 2. The method according to Aspect 1, wherein the graph in the one or more graphs includes a plurality of values ​​and is associated with an image processing function, each of the plurality of values ​​of the graph indicating the amount by which the image processing function is applied to a corresponding pixel of the image data.

[0292] Aspect 3. The method according to any one of Aspects 1 to 2, wherein one or more image processing functions associated with the one or more images include at least one of noise reduction function, sharpness adjustment function, tone adjustment function, saturation adjustment function, and hue adjustment function.

[0293] Aspect 4. The method according to any one of Aspects 1 to 3, wherein the one or more graphs comprise a plurality of graphs, a first graph of the plurality of graphs is associated with a first image processing function, and a second graph of the plurality of graphs is associated with a second image processing function.

[0294] Aspect 5. The method according to aspect 4, wherein the first graph includes a first plurality of values, each of the first plurality of values ​​of the first graph indicating an amount of the first image processing function to be applied to a corresponding pixel of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​of the second graph indicating an amount of the second image processing function to be applied to a corresponding pixel of the image data.

[0295] Aspect 6. The method according to any one of Aspects 4 to 5, wherein the first image processing function associated with the first image in the plurality of images includes one of a noise reduction function, a sharpness adjustment function, a hue adjustment function, a saturation adjustment function, and a hue adjustment function, and wherein the second image processing function associated with the second image in the plurality of images includes one of the noise reduction function, the sharpness adjustment function, the hue adjustment function, the saturation adjustment function, and the hue adjustment function.

[0296] Aspect 7. The method according to any one of Aspects 1 to 6 further includes: generating one or more graphs using an additional machine learning system different from the machine learning system used to generate the modified image.

[0297] Aspect 8. The method according to any one of Aspects 1 to 7, wherein the image data includes a processed image having a plurality of color components for each pixel of the image data.

[0298] Aspect 9. The method according to any one of Aspects 1 to 7, wherein the image data comprises raw image data from one or more image sensors, the raw image data comprising a single color component for each pixel of the image data.

[0299] Aspect 10. The method according to aspect 9, wherein the raw image data is obtained from one or more image sensors filtered by a color filter array.

[0300] Aspect 11. The method according to aspect 10, wherein the color filter array comprises a Bayer color filter array.

[0301] Aspect 12. The method according to any one of Aspects 1 to 11, wherein the image data includes an image data patch, the image data patch comprising a subset of image data frames.

[0302] Aspect 13. The method according to aspect 12, wherein generating the modified image includes generating a plurality of output image data patches, each output image data patch including a subset of pixels of the output image.

[0303] Aspect 14. The method according to any one of Aspects 1 to 13, wherein the machine learning system comprises at least one neural network.

[0304] Aspect 15. An apparatus comprising a memory configured to store image data and a processor implemented in a circuit and configured to perform operations according to any one of aspects 1 to 14.

[0305] Aspect 16. The apparatus according to aspect 15, wherein the apparatus is a camera.

[0306] Aspect 17. The apparatus according to aspect 15, wherein the apparatus is a mobile device including a camera.

[0307] Aspect 18. The apparatus according to any one of aspects 15 to 17 further includes a display configured to display one or more images.

[0308] Aspect 19. The apparatus according to any one of aspects 15 to 18 further includes a camera configured to capture one or more images.

[0309] Aspect 20. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any one of Aspects 1 to 14.

[0310] Aspect 21. An apparatus comprising one or more units for performing operations according to any one of aspects 1 to 14.

[0311] Aspect 22. An apparatus for processing image data, the apparatus comprising a unit for performing an operation according to any one of aspects 1 to 14.

[0312] Aspect 23. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform operations according to any one of aspects 1 to 14.

[0313] Aspect 24. An apparatus for processing image data, the apparatus comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: acquire image data; use the image data as input to one or more trained neural networks to generate one or more graphs, each of the one or more graphs being associated with a corresponding image processing function; and generate an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0314] Aspect 25. The apparatus of aspect 24, wherein the graph in one or more graphs includes a plurality of values ​​and is associated with an image processing function, each of the plurality of values ​​of the graph indicating the intensity of applying the image processing function to a corresponding region of the image data.

[0315] Aspect 26. The apparatus according to aspect 25, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0316] Aspect 27. The apparatus according to any one of aspects 24 to 26, wherein the one or more figures comprise a plurality of figures, a first figure in the plurality of figures being associated with a first image processing function, and a second figure in the plurality of figures being associated with a second image processing function.

[0317] Aspect 28. The apparatus according to aspect 27, wherein one or more image processing functions associated with at least one of the plurality of figures include at least one of noise reduction function, sharpness adjustment function, detail adjustment function, tone adjustment function, saturation adjustment function, and hue adjustment function.

[0318] Aspect 29. The apparatus according to any one of Aspects 27 to 28, wherein the first graph includes a first plurality of values, each of the first plurality of values ​​of the first graph indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​of the second graph indicating the intensity of applying the second image processing function to a corresponding region of the image data.

[0319] Aspect 30. The apparatus according to any one of aspects 24 to 29, wherein the image data includes luminance channel data corresponding to an image, wherein using the image data as input to the one or more trained neural networks includes using the luminance channel data corresponding to the image as input to the one or more trained neural networks.

[0320] Aspect 31. The apparatus according to aspect 30, wherein generating the image based on the image data includes generating the image based on the luminance channel data and chrominance data corresponding to the image.

[0321] Aspect 32. The apparatus according to any one of aspects 30 to 31, wherein the one or more trained neural networks output one or more affine coefficients based on using the image data as input to the one or more trained neural networks, wherein generating the one or more graphs includes generating a first graph by at least transforming the image data using the one or more affine coefficients.

[0322] Aspect 33. The apparatus according to any one of aspects 30 to 32, wherein the image data includes luminance channel data corresponding to an image, wherein transforming the image data using the one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0323] Aspect 34. The apparatus according to any one of aspects 30 to 33, wherein the one or more affine coefficients include multipliers, wherein transforming the image data using the one or more affine coefficients includes multiplying the luminance values ​​of at least a subset of the image data by the multipliers.

[0324] Aspect 35. The apparatus according to any one of aspects 30 to 34, wherein the one or more affine coefficients include an offset, wherein transforming the image data using the one or more affine coefficients includes offsetting the luminance values ​​of at least a subset of the image data by the offset.

[0325] Aspect 36. The apparatus of any one of aspects 30 to 35, wherein the one or more trained neural networks further output the one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first image with one or more gradients in the image data.

[0326] Aspect 37. The apparatus of any one of Aspects 24 to 36, wherein, in order to generate the image based on the image data and the one or more graphs, the one or more processors are configured to use the image data and the one or more graphs as input to a second set of one or more trained neural networks, different from the one or more trained neural networks.

[0327] Aspect 38. The apparatus according to aspect 37, wherein, in order to generate the image based on the image data and the one or more graphs, the one or more processors are configured to use the second set of one or more trained neural networks to demosaic the image data.

[0328] Aspect 39. The apparatus according to any one of aspects 24 to 38, wherein each of the one or more figures varies spatially based on different types of objects depicted in the image data.

[0329] Aspect 40. The apparatus according to any one of aspects 24 to 39, wherein the image data includes an input image having a plurality of color components for each of a plurality of pixels of the image data.

[0330] Aspect 41. The apparatus according to any one of Aspects 24 to 40, wherein the image data comprises raw image data from one or more image sensors, the raw image data comprising at least one color component of each of a plurality of pixels of the image data.

[0331] Aspect 42. The apparatus according to any one of aspects 24 to 41 further includes: an image sensor that captures image data, wherein obtaining the image data includes obtaining the image data from the image sensor.

[0332] Aspect 43. The apparatus according to any one of aspects 24 to 42, further comprising: a display screen, wherein the one or more processors are configured to display the image on the display screen.

[0333] Aspect 44. The apparatus according to any one of aspects 24 to 43 further includes: a communication transceiver, wherein the one or more processors are configured to use the communication transceiver to transmit the image to a receiving device.

[0334] Aspect 45. A method for processing image data, the method comprising: acquiring image data; using the image data as input to one or more trained neural networks to generate one or more graphs, wherein each graph in the one or more graphs is associated with a corresponding image processing function; and generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each graph in the one or more graphs.

[0335] Aspect 46. The method according to aspect 45, wherein the graph in the one or more graphs includes a plurality of values ​​and is associated with an image processing function, each of the plurality of values ​​of the graph indicating the intensity of applying the image processing function to a corresponding region of the image data.

[0336] Aspect 47. The method according to aspect 46, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0337] Aspect 48. The method according to any one of aspects 45 to 47, wherein the one or more graphs comprise a plurality of graphs, a first graph of the plurality of graphs is associated with a first image processing function, and a second graph of the plurality of graphs is associated with a second image processing function.

[0338] Aspect 49. The method according to aspect 48, wherein one or more image processing functions associated with at least one of the plurality of figures include at least one of noise reduction function, sharpness adjustment function, detail adjustment function, tone adjustment function, saturation adjustment function, and hue adjustment function.

[0339] Aspect 50. The method according to any one of Aspects 48 to 49, wherein the first graph includes a first plurality of values, each of the first plurality of values ​​of the first graph indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​of the second graph indicating the intensity of applying the second image processing function to a corresponding region of the image data.

[0340] Aspect 51. The method according to any one of Aspects 45 to 50, wherein the image data includes luminance channel data corresponding to the image, wherein using the image data as input to the one or more trained neural networks includes using luminance channel data corresponding to the image as input to the one or more trained neural networks.

[0341] Aspect 52. The method according to aspect 51, wherein generating the image based on the image data includes generating the image based on the luminance channel data and chrominance data corresponding to the image.

[0342] Aspect 53. The method according to any one of aspects 45 to 52, wherein the one or more trained neural networks output one or more affine coefficients based on using the image data as input to the one or more trained neural networks, wherein generating the one or more graphs includes generating a first graph by at least transforming the image data using the one or more affine coefficients.

[0343] Aspect 54. The method according to aspect 53, wherein the image data includes luminance channel data corresponding to the image, wherein transforming the image data using the one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0344] Aspect 55. The method according to any one of Aspects 53 to 54, wherein the one or more affine coefficients include multipliers, wherein transforming the image data using the one or more affine coefficients includes multiplying the luminance values ​​of at least a subset of the image data by the multipliers.

[0345] Aspect 56. The method according to any one of aspects 53 to 55, wherein the one or more affine coefficients include an offset, wherein transforming the image data using the one or more affine coefficients includes offsetting the brightness values ​​of at least a subset of the image data by the offset.

[0346] Aspect 57. The method of any one of Aspects 53 to 56, wherein the one or more trained neural networks further output one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first figure with one or more gradients in the image data.

[0347] Aspect 58. The method according to any one of Aspects 45 to 57, wherein generating the image based on the image data and the one or more graphs includes using the image data and the one or more graphs as input to a second set of one or more trained neural networks different from the one or more trained neural networks.

[0348] Aspect 59. The method according to aspect 58, wherein generating the image based on the image data and the one or more graphs includes de-mosaicing the image data using a second set of one or more trained neural networks.

[0349] Aspect 60. The method according to any one of aspects 45 to 59, wherein each of the one or more figures varies spatially based on different types of objects depicted in the image data.

[0350] Aspect 61. The method according to any one of Aspects 45 to 60, wherein the image data includes an input image having a plurality of color components for each of a plurality of pixels of the image data.

[0351] Aspect 62. The method according to any one of Aspects 45 to 61, wherein the image data comprises raw image data from one or more image sensors, the raw image data comprising at least one color component for each of a plurality of pixels of the image data.

[0352] Aspect 63. The method according to any one of aspects 45 to 62, wherein obtaining the image data comprises obtaining the image data from an image sensor.

[0353] Aspect 64. The method according to any one of aspects 45 to 63 further includes: displaying the image on a display screen.

[0354] Aspect 65. The method according to any one of aspects 45 to 64 further includes: transmitting the image to a receiving device using a communication transceiver.

[0355] Aspect 66. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any one of Aspects 45 to 64.

[0356] Aspect 67. An apparatus comprising one or more units for performing operations according to any one of aspects 45 to 64.

[0357] Aspect 68. An apparatus for processing image data, the apparatus comprising a unit for performing an operation according to any one of aspects 45 to 64.

[0358] Aspect 69: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform operations according to any one of aspects 45 to 64.

[0359] Aspect 70. An apparatus for processing image data, the apparatus comprising: a memory; one or more processors coupled to the memory, the one or more processors being configured to: acquire image data; use the image data as input to a machine learning system to generate one or more graphs, each of the one or more graphs being associated with a corresponding image processing function; and generate an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0360] Aspect 71. The apparatus according to aspect 70, wherein the graph in the one or more graphs includes a plurality of values ​​and is associated with an image processing function, each of the plurality of values ​​of the graph indicating the intensity of applying the image processing function to a corresponding region of the image data.

[0361] Aspect 72. The apparatus according to aspect 71, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0362] Aspect 73. The apparatus according to any one of aspects 70 to 72, wherein the one or more figures comprise a plurality of figures, a first figure in the plurality of figures being associated with a first image processing function, and a second figure in the plurality of figures being associated with a second image processing function.

[0363] Aspect 74. The apparatus according to aspect 73, wherein one or more image processing functions associated with at least one of the plurality of figures include at least one of noise reduction function, sharpness adjustment function, detail adjustment function, tone adjustment function, saturation adjustment function, and hue adjustment function.

[0364] Aspect 75. The apparatus according to any one of aspects 73 to 74, wherein the first graph includes a first plurality of values, each of the first plurality of values ​​of the first graph indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​of the second graph indicating the intensity of applying the second image processing function to a corresponding region of the image data.

[0365] Aspect 76. The apparatus according to any one of Aspects 70 to 29, wherein the image data includes luminance channel data corresponding to the image, wherein using the image data as input to a machine learning system includes using the luminance channel data corresponding to the image as input to the machine learning system.

[0366] Aspect 77. The apparatus according to aspect 76, wherein generating the image based on the image data includes generating the image based on the luminance channel data and chrominance data corresponding to the image.

[0367] Aspect 78. The apparatus according to any one of aspects 76 to 77, wherein the machine learning system outputs one or more affine coefficients based on using the image data as input to the machine learning system, wherein generating the one or more graphs includes generating a first graph by at least transforming the image data using the one or more affine coefficients.

[0368] Aspect 79. The apparatus according to any one of aspects 76 to 78, wherein the image data includes luminance channel data corresponding to an image, wherein transforming the image data using the one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0369] Aspect 80. The apparatus according to any one of aspects 76 to 79, wherein the one or more affine coefficients include multipliers, wherein transforming the image data using the one or more affine coefficients includes multiplying the luminance values ​​of at least a subset of the image data by the multipliers.

[0370] Aspect 81. The apparatus according to any one of aspects 76 to 80, wherein the one or more affine coefficients include an offset, wherein transforming the image data using the one or more affine coefficients includes offsetting the luminance values ​​of at least a subset of the image data by the offset.

[0371] Aspect 82. The apparatus according to any one of aspects 76 to 81, wherein the machine learning system further outputs the one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first image with one or more gradients in the image data.

[0372] Aspect 83. The apparatus according to any one of aspects 70 to 82, wherein, in order to generate the image based on the image data and the one or more graphs, the one or more processors are configured to use the image data and the one or more graphs as input to a second set of machine learning systems different from the machine learning system.

[0373] Aspect 84. The apparatus according to aspect 83, wherein, in order to generate the image based on the image data and the one or more images, the one or more processors are configured to use the second set of machine learning systems to demosaic the image data.

[0374] Aspect 85. The apparatus according to any one of aspects 70 to 84, wherein each of the one or more figures varies spatially based on different types of objects depicted in the image data.

[0375] Aspect 86. The apparatus according to any one of Aspects 70 to 85, wherein the image data includes an input image having a plurality of color components for each of a plurality of pixels of the image data.

[0376] Aspect 87. The apparatus according to any one of aspects 70 to 86, wherein the image data comprises raw image data from one or more image sensors, the raw image data comprising at least one color component of each of a plurality of pixels of the image data.

[0377] Aspect 88. The apparatus according to any one of aspects 70 to 87 further includes: an image sensor that captures image data, wherein obtaining the image data includes obtaining the image data from the image sensor.

[0378] Aspect 89. The apparatus according to any one of aspects 70 to 88, further comprising: a display screen, wherein the one or more processors are configured to display the image on the display screen.

[0379] Aspect 90. The apparatus according to any one of aspects 70 to 89 further includes: a communication transceiver, wherein the one or more processors are configured to use the communication transceiver to transmit the image to a receiving device.

[0380] Aspect 91. A method for processing image data, the method comprising: acquiring image data; using the image data as input to a machine learning system to generate one or more graphs, wherein each of the one or more graphs is associated with a corresponding image processing function; and generating an image based on the image data and the one or more graphs, the image including features based on the corresponding image processing function associated with each of the one or more graphs.

[0381] Aspect 92. The method according to aspect 91, wherein the graph in the one or more graphs includes a plurality of values ​​and is associated with an image processing function, each of the plurality of values ​​of the graph indicating the intensity of applying the image processing function to a corresponding region of the image data.

[0382] Aspect 93. The method according to aspect 92, wherein the corresponding region of the image data corresponds to a pixel of the image.

[0383] Aspect 94. The method according to any one of aspects 91 to 93, wherein the one or more graphs comprise a plurality of graphs, a first graph of the plurality of graphs is associated with a first image processing function, and a second graph of the plurality of graphs is associated with a second image processing function.

[0384] Aspect 95. The method according to aspect 94, wherein one or more image processing functions associated with at least one of the plurality of figures include at least one of noise reduction function, sharpness adjustment function, detail adjustment function, tone adjustment function, saturation adjustment function, and hue adjustment function.

[0385] Aspect 96. The method according to any one of Aspects 94 to 95, wherein the first graph includes a first plurality of values, each of the first plurality of values ​​of the first graph indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​of the second graph indicating the intensity of applying the second image processing function to a corresponding region of the image data.

[0386] Aspect 97. The method according to any one of Aspects 91 to 96, wherein the image data includes luminance channel data corresponding to the image, wherein using the image data as input to the machine learning system includes using the luminance channel data corresponding to the image as input to the machine learning system.

[0387] Aspect 98. The method according to aspect 97, wherein generating the image based on the image data includes generating the image based on the luminance channel data and chrominance data corresponding to the image.

[0388] Aspect 99. The method according to any one of Aspects 91 to 98, wherein the machine learning system outputs one or more affine coefficients based on using the image data as input to the machine learning system, wherein generating the one or more graphs includes generating a first graph by at least transforming the image data using the one or more affine coefficients.

[0389] Aspect 100. The method according to aspect 99, wherein the image data includes luminance channel data corresponding to the image, wherein transforming the image data using the one or more affine coefficients includes transforming the luminance channel data using the one or more affine coefficients.

[0390] Aspect 101. The method according to any one of Aspects 99 to 100, wherein the one or more affine coefficients include multipliers, wherein transforming the image data using the one or more affine coefficients includes multiplying the luminance values ​​of at least a subset of the image data by the multipliers.

[0391] Aspect 102. The method according to any one of aspects 99 to 101, wherein the one or more affine coefficients include an offset, wherein transforming the image data using the one or more affine coefficients includes offsetting the luminance values ​​of at least a subset of the image data by the offset.

[0392] Aspect 103. The method according to any one of aspects 99 to 102, wherein the machine learning system further outputs the one or more affine coefficients based on a local linearity constraint that aligns one or more gradients in the first figure with one or more gradients in the image data.

[0393] Aspect 104. The method according to any one of aspects 91 to 103, wherein generating the image based on the image data and the one or more graphs includes using the image data and the one or more graphs as input to a second set of machine learning systems different from the machine learning system.

[0394] Aspect 105. The method according to aspect 104, wherein generating an image based on the image data and the one or more graphs includes de-mosaicing the image data using a second set of machine learning systems.

[0395] Aspect 106. The method according to any one of aspects 91 to 105, wherein each of the one or more figures varies spatially based on different types of objects depicted in the image data.

[0396] Aspect 107. The method according to any one of aspects 91 to 106, wherein the image data includes an input image having a plurality of color components for each of a plurality of pixels of the image data.

[0397] Aspect 108. The method according to any one of Aspects 91 to 107, wherein the image data comprises raw image data from one or more image sensors, the raw image data comprising at least one color component for each of a plurality of pixels of the image data.

[0398] Aspect 109. The method according to any one of aspects 91 to 108, wherein obtaining the image data includes obtaining the image data from an image sensor.

[0399] Aspect 110. The method according to any one of aspects 91 to 109 further includes: displaying the image on a display screen.

[0400] Aspect 111. The method according to any one of aspects 91 to 110, further comprising: transmitting the image to a receiving device using a communication transceiver.

[0401] Aspect 112. A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device to perform the method of any one of Aspects 91 to 111.

[0402] Aspect 113. An apparatus comprising one or more units for performing operations according to any one of aspects 91 to 111.

[0403] Aspect 114. An apparatus for processing image data, the apparatus comprising a unit for performing operations according to any one of aspects 91 to 111.

[0404] Aspect 115. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform operations according to any one of aspects 91 to 111.

Claims

1. An apparatus for processing image data, the apparatus comprising: At least one memory; as well as At least one processor coupled to the at least one memory, the at least one processor being configured to: Acquire image data; Using the image data as input to at least one trained neural network, a plurality of affine coefficients are generated based on using the image data as input to the at least one trained neural network. At least one graph is generated by transforming the image data using the plurality of affine coefficients and implementing at least one local linearity constraint, wherein the at least one graph is associated with an image processing function, and wherein the at least one local linearity constraint is configured to set the pixel value difference of the graph gradient in the at least one graph based on the pixel value difference between pixels of the image gradient in the image data to increase the alignment between the graph gradient and the image gradient; and The image processing function is used to modify the characteristics of the image data to generate an image based on at least one of the graphs.

2. The apparatus according to claim 1, wherein, The graph in the at least one graph includes multiple values ​​and is associated with an image processing function, each of the multiple values ​​indicating the intensity of applying the image processing function to a corresponding region of the image data.

3. The apparatus according to claim 2, wherein, The corresponding region of the image data corresponds to a pixel in the image.

4. The apparatus according to claim 1, wherein, The at least one graph includes a plurality of graphs, wherein a first graph in the plurality of graphs is associated with a first image processing function, and a second graph in the plurality of graphs is associated with a second image processing function.

5. The apparatus according to claim 4, wherein, At least one image processing function associated with at least one of the plurality of images includes at least one of the following: noise reduction function, sharpness adjustment function, detail adjustment function, tone adjustment function, saturation adjustment function, or hue adjustment function.

6. The apparatus according to claim 4, wherein, The first graph includes a first plurality of values, each of the first plurality of values ​​in the first graph indicating the intensity of applying the first image processing function to a corresponding region of the image data, and wherein the second graph includes a second plurality of values, each of the second plurality of values ​​in the second graph indicating the intensity of applying the second image processing function to a corresponding region of the image data.

7. The apparatus according to claim 1, wherein, The image data includes luminance channel data corresponding to the image, wherein using the image data as input to the at least one trained neural network includes using the luminance channel data corresponding to the image as input to the at least one trained neural network.

8. The apparatus according to claim 7, wherein, Generating the image based on the image data includes: generating the image based on the luminance channel data and chrominance data corresponding to the image.

9. The apparatus according to claim 1, wherein, The image data includes luminance channel data corresponding to the image, wherein transforming the image data using the plurality of affine coefficients includes transforming the luminance channel data using the plurality of affine coefficients.

10. The apparatus according to claim 1, wherein, The plurality of affine coefficients includes a multiplier, wherein transforming the image data using the plurality of affine coefficients includes multiplying the brightness values ​​of at least one subset of the image data by the multiplier.

11. The apparatus according to claim 1, wherein, The plurality of affine coefficients includes an offset, wherein using the plurality of affine coefficients to transform the image data includes offsetting the brightness values ​​of at least one subset of the image data by the offset.

12. The apparatus according to claim 1, wherein, The at least one local linearity constraint is configured to reduce the halo effect in the at least one graph.

13. The apparatus according to claim 1, wherein, In order to generate the image based on the image data and the at least one graph, the at least one processor is configured to use the image data and the at least one graph as input to at least a second trained neural network that is different from the at least one trained neural network.

14. The apparatus according to claim 13, wherein, In order to generate the image based on the image data and the at least one graph, the at least one processor is configured to demosaic the image data using at least the second trained neural network.

15. The apparatus according to claim 1, wherein, Each of the at least one graph varies spatially based on the different types of objects depicted in the image data.

16. The apparatus according to claim 1, wherein, The image data includes an input image, which has multiple color components for each of the multiple pixels in the image data.

17. The apparatus according to claim 1, wherein, The image data includes raw image data from at least one image sensor, wherein the raw image data includes at least one color component for each of the plurality of pixels in the image data.

18. The apparatus according to claim 1, further comprising: An image sensor that captures the image data, wherein obtaining the image data includes obtaining the image data from the image sensor.

19. The apparatus according to claim 1, further comprising: A display screen, wherein the at least one processor is configured to display the image on the display screen.

20. The apparatus according to claim 1, further comprising: A communication transceiver, wherein the at least one processor is configured to use the communication transceiver to transmit the image to a receiving device.

21. The apparatus according to claim 1, wherein, At least one edge from the image data is preserved in the at least one image.

22. A method for processing image data, the method comprising: Acquire image data; Using the image data as input to at least one trained neural network, a plurality of affine coefficients are generated based on using the image data as input to the at least one trained neural network. At least one graph is generated by transforming the image data using the plurality of affine coefficients and implementing at least one local linearity constraint, wherein the at least one graph is associated with an image processing function, and wherein the at least one local linearity constraint is configured to set the pixel value difference of the graph gradient in the at least one graph based on the pixel value difference between pixels of the image gradient in the image data to increase the alignment between the graph gradient and the image gradient; and The image processing function is used to modify the characteristics of the image data to generate an image based on at least one of the graphs.

23. The method according to claim 22, wherein, The graph in the at least one graph includes multiple values ​​and is associated with an image processing function, each of the multiple values ​​indicating the intensity of applying the image processing function to a corresponding region of the image data.

24. The method according to claim 22, wherein, The at least one graph includes a plurality of graphs, wherein a first graph in the plurality of graphs is associated with a first image processing function, and a second graph in the plurality of graphs is associated with a second image processing function.

25. The method according to claim 22, wherein, The image data includes luminance channel data corresponding to the image, wherein using the image data as input to the at least one trained neural network includes using the luminance channel data corresponding to the image as input to the at least one trained neural network.

26. The method according to claim 22, wherein, The plurality of affine coefficients includes a multiplier, wherein transforming the image data using the plurality of affine coefficients includes multiplying the brightness values ​​of at least one subset of the image data by the multiplier.

27. The method according to claim 22, wherein, The plurality of affine coefficients includes an offset, wherein using the plurality of affine coefficients to transform the image data includes offsetting the brightness values ​​of at least one subset of the image data by the offset.

28. The method according to claim 22, wherein, Generating the image based on the image data and the at least one graph includes: using the image data and the at least one graph as input to at least a second trained neural network that is different from the at least one trained neural network.

29. The method according to claim 22, wherein, Each of the at least one graph varies spatially based on the different types of objects depicted in the image data.

30. The method according to claim 22, wherein, The image data includes luminance channel data corresponding to the image, wherein transforming the image data using the plurality of affine coefficients includes transforming the luminance channel data using the plurality of affine coefficients.

Citation Information

Patent Citations

  • System and method for demosaicing image data using weighted gradients

    CN102640499A

  • Image lighting methods and apparatuses, electronic devices, and storage media

    US20200143230A1