Method and apparatus for processing image by using neural network model
A lightweight neural network model infers parameters for image signal processing operations, addressing complexity and retraining issues in existing models, ensuring efficient and customizable image processing.
Patent Information
- Application Number
- PCT/KR2025/095112
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-09
- Filing Date
- 2025-03-26
- Publication Date
- 2025-11-06
AI Technical Summary
Existing neural network models for processing raw images require high computational complexity, large memory usage, and are difficult to streamline, necessitating retraining or additional signal processing to handle diverse shooting conditions, and often rely on ground truth data for training.
A lightweight neural network model is trained to infer parameters for image signal processing operations, allowing it to be integrated into an image signal processing pipeline without retraining, using inference-based parameter adjustment to customize image characteristics.
The solution reduces computational complexity and memory requirements while enabling efficient image processing with customizable output quality, eliminating the need for extensive retraining and ground truth data.
Smart Images

Figure KR2025095112_06112025_PF_FP_ABST
Abstract
Description
Method and device for processing images using a neural network model
[0001] The present disclosure relates to a method and apparatus for processing an image using a neural network model, and more particularly, to a method and apparatus for processing an image using parameters obtained based on a neural network model.
[0002] Recent advancements in neural network technology have led to the utilization of neural networks in a variety of fields. Neural network models are also widely used in image signal processing, such as for image correction and restoration. For example, convolutional neural networks (CNNs) can be used for image processing. CNNs can effectively extract spatial characteristics of images.
[0003] A raw image can contain all data captured by an image sensor. For example, a raw image may correspond to an unprocessed image captured by a camera sensor. Because raw images contain original data without any compression or correction, they can offer a wider range of adjustments for editing, including richer gradations and extreme shadows and highlights. Therefore, neural network models are being utilized to process raw images. There is a growing demand for neural network models that can effectively process a variety of raw images acquired under diverse shooting conditions while maintaining low complexity.
[0004] The present disclosure can be implemented in various ways, including as a method, system, device, or computer program stored on a computer-readable storage medium.
[0005] A method according to one embodiment of the present disclosure may include: obtaining a first image from a raw image; obtaining one or more parameters based on the first image using a neural network model; and obtaining a second image by performing one or more signal processing operations on the first image using the one or more parameters. The neural network model may be learned by obtaining parameters from a training image using the neural network model, performing the one or more signal processing operations on the training image based on the parameters to obtain an output image, and updating the neural network model based on the training image and the output image.
[0006] An electronic device according to one embodiment of the present disclosure may include: at least one processor including a processing circuit; and a memory storing one or more instructions including one or more recording media. The one or more instructions may be individually or in combination executed by the at least one processor to cause the electronic device to: obtain a first image from a raw image; obtain one or more parameters based on the first image using a neural network model; and obtain a second image by performing one or more signal processing operations on the first image using the one or more parameters. The neural network model may be learned by obtaining parameters from a training image using the neural network model, performing the one or more signal processing operations on the training image based on the parameters to obtain an output image, and updating the neural network model based on the training image and the output image.
[0007] A recording medium according to one embodiment of the present disclosure may be a computer-readable recording medium having recorded thereon a program for performing at least some of the methods described in the present disclosure on a computer.
[0008] FIG. 1 illustrates an exemplary block diagram of an electronic device according to one embodiment of the present disclosure.
[0009] FIG. 2 illustrates a block diagram of image signal processing for a raw image according to one embodiment of the present disclosure.
[0010] FIG. 3 illustrates a block diagram of image signal processing for a linear image according to one embodiment of the present disclosure.
[0011] FIG. 4 illustrates images generated by the first exposure fusion of FIG. 3 according to one embodiment of the present disclosure.
[0012] FIG. 5 illustrates images generated by the second exposure fusion of FIG. 3 according to one embodiment of the present disclosure.
[0013] FIG. 6 illustrates an output image and its histogram generated by the dynamic range scaling of FIG. 3 according to one embodiment of the present disclosure.
[0014] FIG. 7 exemplarily illustrates the structure of a neural network model according to one embodiment of the present disclosure.
[0015] FIG. 8 illustrates an example of an image output as parameters change according to one embodiment of the present disclosure.
[0016] FIG. 9 exemplarily illustrates a flowchart of a method according to one embodiment of the present disclosure.
[0017] FIG. 10 illustrates an exemplary block diagram of an electronic device according to one embodiment of the present disclosure.
[0018] FIG. 11 illustrates an example of a user interface displayed on a display of an electronic device according to one embodiment of the present disclosure.
[0019] This disclosure may be subject to various modifications and various embodiments. Specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the disclosure.
[0020] When describing embodiments, detailed descriptions of related known technologies are omitted if they are deemed to unnecessarily obscure the main point. Furthermore, numbers (e.g., "first," "second," etc.) used in the description of embodiments are merely identifiers used to distinguish one component from another.
[0021] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this specification should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0022] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.
[0023] Throughout this disclosure, when a part is said to "include" a certain component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part," "module," and "block" used herein mean a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0024] The expression “configured to” as used herein can be used interchangeably with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” does not necessarily mean something that is “specifically designed to” in terms of hardware. Instead, in some contexts, the expression “a system configured to” can mean that the system is “capable of” doing something in conjunction with other devices or components. For example, the phrase “a processor configured to perform A, B, and C” can mean a dedicated processor (e.g., an embedded processor) for performing the operations, or a generic-purpose processor that can perform the operations by executing one or more software programs stored in memory.
[0025] Additionally, when a component is referred to as being “connected” or “connected” to another component in the present disclosure, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless otherwise specifically stated.
[0026] All functions or operations described in the present disclosure may be individually processed by a single processor and / or collectively processed by multiple processors. A single processor or a combination of multiple processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), an Integrated Chip (IC), etc.
[0027] It should be understood that the blocks and combinations of flowcharts illustrated in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.
[0028] In this disclosure, components expressed as "units", "modules", etc. may be two or more components combined into one component, or one component may be divided into two or more components with more detailed functions. In addition, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and of course, some of the main functions performed by each component may be exclusively performed by other components.
[0029] Before going into a specific description of the invention, the terms used in the present disclosure may be defined or understood as follows.
[0030] In the present disclosure, a 'raw image' may be understood as an image representing the intensity of light detected by a light receiving sensor after external light passes through a color filter corresponding to each pixel for a plurality of pixels. The raw image may have a designated pattern based on the pattern of a color filter array composed of a plurality of types of color filters and may be understood as data that has not been demosaiced. In various embodiments, the raw image having the designated pattern may have only a color value corresponding to a specific color filter for each of the plurality of pixels. In one embodiment, the raw image may be data on which some processing such as Lens Shading Correction (LSC) and Bad Pixel Correction (BPC) has been performed.
[0031] In the present disclosure, 'demosaicing' can be understood as a process of calculating an actual color value from a color value corresponding to any one of a plurality of color filters for each of a plurality of pixels included in a raw image. In one embodiment, demosaicing can mean a process of calculating an actual color value of an object corresponding to each of the plurality of pixels by predicting or calculating a color value for a color filter other than the specific color filter for each of the plurality of pixels.
[0032] In the present disclosure, the term 'pixel value' may be understood as a color value corresponding to each pixel in data (e.g., image data) for a plurality of pixels consisting of N columns and M rows (N and M are natural numbers). The pixel value may include color values for one or more colors determined according to a color filter array used to obtain the corresponding data. For example, when a Bayer pattern is used as the color filter array, the image may be understood as an image having an RGB (Red, Green, Blue) color space, and three pixel values for red, green, and blue may exist for one pixel of the image.
[0033] In the present disclosure, the term 'RGB image' may be understood as an image in the RGB color space. For example, RGB image data may include pixel values for red, green, and blue.
[0034] In the present disclosure, a set of pixel values for the red component (or associated with red) in RGB image data may be referred to as a red channel (or R channel). A set of pixel values for the green component (or associated with green) in RGB image data may be referred to as a green channel (or G channel). A set of pixel values for the blue component (or associated with blue) in RGB image data may be referred to as a blue channel (or B channel).
[0035] In the present disclosure, a 'linear RGB image' may represent an RGB image converted into three channels (R channel, G channel, B channel) by applying demosaicing to a Bayer image. In one embodiment, the 'linear RGB image' may correspond to an image to which demosaicing and WBG (white balance gain) are applied to a raw image.
[0036] In some embodiments, an electronic device for processing images may include a neural network model that directly processes raw images. This neural network model may be trained to output full-size images. For example, the electronic device may use an end-to-end neural network model or an artificial intelligence model. Inputting and outputting images using an end-to-end neural network model may require a large amount of memory, high computational complexity, and long processing times. Consequently, it is difficult to streamline the neural network, and the memory usage of the neural network model may be large. Furthermore, retraining the neural network model or introducing additional signal processing may be necessary to process images with different characteristics. Additionally, to output appropriate images in response to various requirements or issues, the end-to-end neural network model may need to be retrained or pre-trained using various techniques. Furthermore, training an end-to-end neural network model requires a training data set containing ground truth or labeled data acquired by experts or other individuals.
[0037] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.
[0038] FIG. 1 illustrates a block diagram of an electronic device (100) according to one embodiment of the present disclosure.
[0039] Referring to FIG. 1, an electronic device (100) may include an image sensor (102), a neural network model (104), and an ISP pipeline (106). In one embodiment, the electronic device (100) may be an ISP device including a signal processing unit (e.g., an ISP pipeline (106)) for processing a raw image acquired from a sensor of a camera module (e.g., an image sensor (102)) and an artificial intelligence (AI) unit (e.g., a neural network model (104)) for determining parameters used in the signal processing unit.
[0040] Light entering the electronic device through the camera lens can be converted into an electrical image signal using an image sensor (102). In one embodiment, the image sensor (102) may include a charged coupled device (CCD) sensor or a complementary metal-oxide semiconductor (CMOS) sensor. The image sensor (102) may include a color filter array (CFA) composed of a plurality of types of color filters. The image sensor (102) can obtain a raw image based on the converted electrical image signal. The image sensor (102) can provide the raw image to a neural network model (104) and an ISP pipeline (106).
[0041] The electronic device (100) can apply a series of signal processing operations (or signal processing algorithms) to a raw image acquired through a sensor of a camera module, a raw image loaded from a storage device within the electronic device (100), or a raw image provided from outside the electronic device (100). Accordingly, a final image having optimal contrast ratio, brightness, color, and sharpness can be generated from the electronic device (100). The setting parameters used in each signal processing operation can be acquired using an AI model. When the parameters acquired based on the inference of the AI model are applied to each signal processing operation, the AI model can be trained so that the image generated from the electronic device (100) satisfies the intended characteristics.
[0042] The neural network model (104) can perform inference based on a raw image or an image corresponding to the raw image. Based on the inference of the neural network model (104), one or more parameters can be obtained. For example, using the neural network model (104), the electronic device (100) can obtain one or more parameters to be applied to the ISP pipeline (106). The electronic device (100) can obtain a raw image from the image sensor (102), generate an image corresponding to the raw image, input the generated image to the neural network model (104), and calculate one or more parameters based on the inference from the neural network model (104).
[0043] In one embodiment, parameters obtained using the neural network model (104) may be mapped to specific signal processing operations of the ISP pipeline (106). For example, one parameter obtained using the neural network model (104) may be mapped to one of the parameters used for one function among the signal processing functions provided by an image editing application provided to a user of the electronic device (100). Accordingly, the electronic device (100) may be easily integrated into an existing image processing system of the electronic device (100).
[0044] In one embodiment, the neural network model (104) may be trained by obtaining parameters from an input image (e.g., a training image from a training dataset) using the neural network model (104), applying the obtained parameters to the ISP pipeline (106), obtaining an output image from the ISP pipeline (106), and updating the neural network model (104) based on the input image and the output image. For example, the neural network model may be trained by optimizing (or updating) the neural network model (104) based on a training loss calculated based on the input image and the output image. The characteristics of the input image or the output image used for calculating the loss may include local variance between adjacent pixels, an average of all pixel values, or entropy. Accordingly, ground truth data for training the neural network model (104) may be unnecessary, and as a result, it may be easy to obtain training data for the neural network model (104).
[0045] The ISP pipeline (106) may include one or more signal processing operations (108). For example, the ISP pipeline (106) may be a pipeline composed of one or more signal processing operations connected in series, one or more signal processing operations connected in parallel, or a combination thereof. Using the ISP pipeline (106), the electronic device (100) may process a raw image. For example, by performing one or more signal processing operations (108) within the ISP pipeline (106), the electronic device (100) may process the raw image.
[0046] The electronic device (100) can individually or in combination perform one or more signal processing operations (108) within the ISP pipeline (106) based on one or more parameters obtained using the neural network model (104). The electronic device (100) can apply one or more parameters obtained using the neural network model (104) to a corresponding signal processing operation(s). The one or more parameters obtained using the neural network model (104) can be used to set (or control, configure) the corresponding signal processing operation. In one embodiment, the electronic device (100) can infer parameters for some or all of the signal processing operations (or algorithms) included in the electronic device (100) using a single neural network model. For example, a first parameter obtained using a neural network model (104) may be used to perform a first signal processing operation within an ISP pipeline (106), and a second parameter obtained using a neural network model (104) may be used to perform a second signal processing operation within the ISP pipeline (106).
[0047] In one embodiment, the electronic device (100) may include a lightweight neural network model for each signal processing operation (or algorithm). To appropriately perform a specific signal processing, a dedicated neural network model may be independently trained and updated. For example, the electronic device (100) may include multiple neural network models. A first neural network model among the multiple neural network models may be used to obtain parameters for performing a first signal processing operation within the ISP pipeline (106). A second neural network model among the multiple neural network models may be used to obtain parameters for performing a second signal processing operation within the ISP pipeline (106).
[0048] In one embodiment, one or more signal processing operations (108) within the ISP pipeline (106) may include at least one of exposure fusion, brightness adjustment, contrast adjustment, histogram equalization, tone adjustment, and dynamic range scaling. The at least one signal processing operation within the ISP pipeline (106) may be executed using one or more parameters inferred using the neural network model (104).
[0049] The neural network model (104) of the electronic device (100) can be trained to output optimized parameter(s) for at least some of one or more signal processing operations (108) of the ISP pipeline (106). In one embodiment, when a scenario for signal processing by the electronic device (100) changes (e.g., when the settings of the image sensor (102) change, or when a user of the electronic device (100) requests something), the parameters obtained using the neural network model (104) can be modified while maintaining the neural network model (104) (e.g., without retraining or modifying the neural network model (104). For example, the electronic device (100) can modify the parameters obtained using the inference results of the neural network model (104) and apply the modified parameters to the ISP pipeline (106). By adjusting the parameters inferred based on the neural network model (104), the characteristics of the output image can be customized. Furthermore, training a new model or retraining an existing neural network model to generate the desired output may be unnecessary. Instead of spending a long time retraining the neural network model (104), issues that arise when actually operating the ISP pipeline (106) (or when verifying the ISP pipeline (106)) can be quickly resolved by adjusting the parameters obtained using the neural network model (104).
[0050] The neural network model (104) of the electronic device (100) can be trained to infer one or more values for calculating parameters for the input image, instead of the image itself, from the input image. A neural network model trained to infer the image itself can be large, can infer unexpected image correction results, and can be difficult to control. By being trained to infer values for calculating parameters instead of the image, the neural network model (104) can be lightweight. Furthermore, the image signal processing of the electronic device (100) can be performed using the ISP pipeline (106) rather than the neural network model (104). Therefore, even if the accuracy of the inference of the neural network model (104) is low (e.g., an execution error of the neural network model (104)), it is unlikely to significantly degrade the performance of the ISP pipeline (106), and thus a certain level of image signal processing can be guaranteed.
[0051] In one embodiment, an electronic device (100) including an image sensor (102), a neural network model (104), and an ISP pipeline (106) may be implemented in the form of a SoC (System on Chip) mounted on a single chip. For example, the electronic device (100) may include a camera module including a lens module (not shown), an image sensor (102), and an image signal processor, wherein the image signal processor may include a neural network model (104) and an ISP pipeline (106). The image sensor (102) receives light transmitted through the lens module and outputs an image, and the image signal processor may correct the image from the image sensor (102) using the neural network (104) and the ISP pipeline (106). For example, when a user takes an image using an electronic device (100), a neural network (104) and an ISP pipeline (106) may be executed on the image taken by the image signal processor, and an image that has passed through the ISP pipeline (106) may be provided to the user.
[0052] In one embodiment, the electronic device (100) may include a main processor (not shown) capable of executing a neural network model (104) and an ISP pipeline (106). The main processor may correct a raw image acquired from an image sensor (102) or a raw image received from an external source using the neural network model (104) and the ISP pipeline (106). For example, the main processor may receive a request from a user to correct a raw image, and in response to the request, may correct the raw image using the neural network model (104) and the ISP pipeline (106).
[0053] FIG. 2 illustrates a block diagram of image signal processing (200) for a raw image according to one embodiment of the present disclosure.
[0054] Referring to FIG. 2, image signal processing (200) may include blocks (202, 204, 206, 208, 210). In one embodiment, image signal processing (200) may be performed by the electronic device (100) of FIG. 1.
[0055] Prior to being input to the ISP pipeline of block (210), the raw image may be preprocessed in block (202). For example, the electronic device (100) may perform various signal processing operations on the raw image, such as Bayer demosaicing, white balance gain and color correction matrices, pixel value normalization, or exposure normalization. The preprocessing performed in block (202) may be understood as refining the raw image. As a result of preprocessing the raw image, an input image may be output from block (202).
[0056] An input image from block (202) may be input to block (210) and block (204). Block (210) may be an ISP pipeline including one or more signal processing operations (211, …, 21n; n is a natural number). To obtain one or more parameters for the signal processing operations (211, …, 21n), the input image may be input to block (204). In block (204), the input image may be downsampled. For example, in block (204), the electronic device (100) may reduce the size of the input image. By using the downsampled input image instead of the input image, the input size of the subsequent neural network model (206) may be reduced, and thus the amount of computation of the neural network model (206) may be reduced.
[0057] In block (204), the downsampled input image can be input to a neural network model (206). The neural network model (206) can perform inference based on the downsampled input image. Based on the inference of the neural network model (206), in block (208), one or more parameters for signal processing operations (211, ..., 21n) can be calculated. For example, in block (208), the electronic device (100) can perform various operations based on the inference of the neural network model (206), thereby calculating one or more parameters. The parameters calculated in block (208) can be provided to the ISP pipeline of block (210).
[0058] The signal processing operations (211, …, 21n) may be performed based on one or more corresponding parameters among the parameters calculated in block (208). For example, the electronic device (100) may perform the signal processing operation (211) on the input image based on a first parameter among the parameters calculated in block (208). The electronic device (100) may perform the signal processing operation (21n) based on a second parameter and a third parameter among the parameters calculated in block (208). The result of the signal processing operation (21n) may correspond to an output image of the image signal processing (200).
[0059] FIG. 3 illustrates a block diagram of image signal processing for a linear image according to one embodiment of the present disclosure.
[0060] Referring to FIG. 3, image signal processing (300) may include blocks (302, 306, 308, 310). In one embodiment, image signal processing (300) illustrated in FIG. 3 may be performed by the electronic device (100) of FIG. 1. By performing image signal processing (300), an output image may be obtained from an input image.
[0061] In one embodiment, the input image of the image signal processing (300) may be an image in the linear RGB domain. For example, before being input to the ISP pipeline of block (302), Bayer demosaicing and application of white balance gains and color correction matrices may be completed on the raw image.
[0062] The ISP pipeline (302) may include blocks (304, 312, 314, 316, and 318). Each block of the ISP pipeline (302) may correspond to a separate signal processing operation. In block (304), exposure normalization may be performed on the input image. Depending on the image capturing conditions (e.g., the settings of the camera module when capturing the image), the exposure of the image may be set relatively high or low. For example, under some conditions, the raw image may be acquired under low exposure to prevent clipping in the highlight area. Consequently, the brightness between the images may be inconsistent. To address this inconsistency, the input image may first be adjusted as if it were captured under normal exposure. In one embodiment, the exposure normalization in block (304) may normalize the pixel values of the input image to the range [0, 1].
[0063] To perform signal processing operations in blocks (312, 314) of the ISP pipeline (302), one or more parameters may be calculated from a normalized input image based on a neural network model (308). Prior to being input to the neural network model (308), the normalized input image may be downsampled in block (306). In one embodiment, the normalized input image may be downsampled by a ratio of 16 (e.g., a ratio of 1:16 or a ratio of 1 / 16). For example, downsampling may reduce the size of the normalized input image to 1 / 16.
[0064] In block (306), the downsampled input image may be input into the neural network model (308). For example, the electronic device (100) may perform inference by inputting the downsampled input image into the neural network model (308). Based on the inference from the neural network model (308), in block (310), one or more parameters may be calculated. The parameters calculated in block (310) may be provided to corresponding signal processing operations of the ISP pipeline (302). For example, among the parameters calculated in block (310), a first parameter and a second parameter may be used for the first exposure fusion in block (312). A third parameter and a fourth parameter among the parameters calculated in block (310) may be used for the second exposure fusion in block (314).
[0065] In the embodiment illustrated in FIG. 3, the parameters calculated in block (310) may be used in the first exposure fusion of block (312) and the second exposure fusion of block (314). However, embodiments of the present disclosure are not limited thereto. For example, in block (310), one or more parameters for gamma correction of block (316) and / or one or more parameters for dynamic range scaling of block (318) may be additionally calculated. The additionally calculated parameters may be used for corresponding signal processing operations.
[0066] In blocks (312) and (314), first exposure fusion and second exposure fusion may be performed on the normalized input image, respectively, based on one or more parameters calculated using a neural network model. Through the first exposure fusion and the second exposure fusion, local brightness and contrast of all regions within the image may be optimized. In one embodiment, the first exposure fusion of block (312) may be understood as exposure fusion darken, and the second exposure fusion of block (314) may be understood as exposure fusion brighten.
[0067] Due to the limited dynamic range of camera sensors, captured images of high dynamic scenes may have overexposed or underexposed areas. In block (314), the second exposure fusion may enhance the tones of underexposed areas. At the same time, the second exposure fusion may cause undesirable brightening of well-exposed areas. Therefore, prior to block (314), the first exposure fusion in block (312) may compensate for the undesired brightening. For example, prior to the second exposure fusion in block (314), the first exposure fusion in block (312) may selectively darken well-exposed areas or overexposed areas. The first exposure fusion in block (312) and the second exposure fusion in block (314) will be described in more detail below with reference to FIGS. 4 and 5, respectively.
[0068] In block (316), gamma correction may be performed on the exposure-fused image. For example, the electronic device (100) may non-linearly adjust the brightness of pixel values of the exposure-fused image by applying a specific value. Accordingly, linear RGB values of the exposure-fused image may be converted into non-linear RGB values. In one embodiment, the electronic device (100) may perform gamma correction by taking the square of a predetermined gamma value (e.g., 2.2) for the exposure-fused image.
[0069] At block (318), dynamic range scaling may be performed on the gamma-corrected image. For example, the electronic device (100) may adjust the brightness values of the gamma-corrected image to expand or contract the contrast ratio. The electronic device (100) may linearly scale the gamma-corrected image to fit the image (or the range of pixel values of the image) to the new dynamic range. Dynamic range scaling may allow for better representation of details in very bright or very dark areas. The dynamic range scaling of block (318) will be described in detail below with reference to FIG. 6.
[0070] The image signal processing (300) illustrated in FIG. 3 can be understood as a hybrid AI ISP. The image signal processing (300) may comprise an AI portion including a neural network model and a signal processing portion. The neural network model may be trained to infer parameters used in signal processing algorithms, thereby improving local contrast in areas with different exposures. In one embodiment, the image signal processing (300) may comprise a small neural network with fast image output suitable for mobile environments. Furthermore, the parameters inferred using the neural network model may be customized to output images suited to the user's intent or the application's purpose. Accordingly, retraining the neural network model or including multiple pre-trained neural network models may be unnecessary.
[0071] In some embodiments, the ISP pipeline (302) may additionally include other signal processing operations. For example, various image signal processing operations such as color temperature adjustment, image denoising, tint adjustment, or saturation adjustment may be additionally introduced into the ISP pipeline (302).
[0072] FIG. 4 illustrates images generated by the first exposure fusion of FIG. 3 according to one embodiment of the present disclosure.
[0073] Referring to Figure 4, the image (402) may be a normalized input image input to the first exposure fusion of block (312) of Fig. 3. In block (312), the image (402) is a function by his dark version (i.e. image (404)) can be used to generate a function can be defined as in mathematical expression 1.
[0074]
[0075] In mathematical expression 1, can be a real number greater than 1. Function By, image The dark area of (402) can be further darkened. In one embodiment, the image is normalized by exposure of block (304). The pixel values of (402) can be normalized in the range [0, 1]. Accordingly, in the first exposure fusion of block (312), the image As (402) is squared, the pixel values in the dark areas may decrease. As a result, the image (402) The pixels in the dark area are the function It can be darker by .
[0076] illumination map (406) can be derived from image P(402) using mathematical expression 2.
[0077]
[0078] In mathematical expression 2, may refer to a guided filter. A guided filter can be used to blur out textures and leave only structural surfaces. In one embodiment, to apply a guided filter, an image is used as a guide image. (402) may be used. In one embodiment, a separate guide image may be provided as a block (312) to apply a guided filter. Illumination map (406) is an image (402) may be a weight map for representing weights for the edges.
[0079] In mathematical expressions 1 and 2, the parameter and may be parameters calculated in block (310) based on the inference of the neural network model (308). Parameters is an image by first exposure fusion of block (312). (402) may be related to the degree of darkening. Parameter is an image by a guided filter (402) may be related to the degree of blurring. Based on the neural network model (308), the image (402) Optimized parameters can be determined.
[0080] image (402) is image 1- (408) can be multiplied by Image 1- (408) is a lighting map (406) can correspond to an image with pixel values subtracted from 1. Image 1- (408), pixels whose pixel value is close to 1 (e.g., image 1- (408) is an image (402) may be pixels in the dark area. In one embodiment, the exposure of the image may be normalized to the range [0, 1] by block (304) (e.g., the pixel values of the image may be normalized to the range [0, 1]), and thus the image 1- (408) is a lighting map (406) This can correspond to an inverted image.
[0081] image (404) is a lighting map (406) can be multiplied. In block (410), the image (402) and Image 1- Product and image of (408) (404) and lighting map (406) The product can be merged (or added). As a result of the merge, the image (412) can be generated. Referring to mathematical expression 1, The larger the value, the better the image (412) is an image (402) It may darken in contrast.
[0082] Histogram (414) is an image (402) represents the distribution of pixel values, and the histogram (416) is an image (412) can express the distribution of pixel values. When comparing the histogram (414) and the histogram (416), the pixel values of the bright area (e.g., pixels to the right of the center of the X-axis of the histogram) and the middle area (e.g., pixels near the center of the X-axis of the histogram) can be moved to the dark area (e.g., pixels to the left of the center of the X-axis of the histogram) by the first exposure fusion. Accordingly, the image (402), compared to the image generated by the first exposure fusion (412) can be darker.
[0083] FIG. 5 illustrates images generated by the second exposure fusion of FIG. 3 according to one embodiment of the present disclosure.
[0084] Referring to Figure 5, the image (502) may be an image input from the first exposure fusion of block (312) to the second exposure fusion of block (314) in FIG. 3. For example, the image (502) is an image resulting from the first exposure fusion of block (312), as shown in FIG. 4. (412) may be. In block (314), the image (502) is a function by his dark version (i.e. image (504)) can be used to generate a function can be defined as in mathematical formula 3.
[0085]
[0086] In mathematical expression 3, It can be a real number greater than 1. Function By, image The bright area of (502) can be further brightened. As the image increases, (502) The pixel values may increase, resulting in (504) can be brighter.
[0087] Lighting map (508) is an image using mathematical expression 4. (502) can be derived from.
[0088]
[0089] In mathematical equation 4, may mean a guided filter. In one embodiment, to apply the guided filter, an image as a guide image (502) may be used. In one embodiment, a separate guide image may be provided as a block (312) to apply a guided filter. Illumination map (508) may be a weight map for representing weights for the boundaries of image R (502).
[0090] In Equations 3 and 4, the parameter and may be parameters calculated in block (310) based on the inference of the neural network model (308). Parameters is an image by first exposure fusion of block (314). (502) may be related to the degree of brightness. Parameter is an image by a guided filter (502) may be related to the degree of blurring. Based on the neural network model (308), the image (502) Optimized parameters can be determined.
[0091] image (502) is a lighting map (508) can be multiplied by the image (504) is image 1- (506) can be multiplied by Image 1- (506) is a lighting map (508) can correspond to an image in which the pixel values are subtracted from 1. In block (510), the image (502) and lighting map Product and image of (508) (504) and Image 1- (506) The product can be merged (or added). As a result of the merge, the image (512) can be generated. Referring to mathematical expression 3, The larger the value, the better the image (512) is an image (502) It can be brighter in contrast.
[0092] Histogram (514) is an image (502) represents the distribution of pixel values, and the histogram (516) represents the image (504) represents the distribution of pixel values, and the histogram (518) is an image (512) can express the distribution of pixel values. Comparing the histogram (514) and histogram (516), the image (502) Image (504) may include pixels with high pixel values (e.g., relatively bright pixels). Comparing the histogram (514) and histogram (516), the image (504) without affecting the dark areas (e.g., pixels to the left of the center of the X-axis of the histogram) and mid-tone areas (e.g., pixels near the center of the X-axis of the histogram) of the image. The brightest areas of (504) (e.g., pixels near the rightmost side of the X-axis of the histogram) can be darkened. For example, through second exposure fusion, the rightmost peak of the histogram (518) (e.g., image (512) sky area) can be moved to a darker part (e.g., to the left) than the rightmost peak of the histogram (516).
[0093] FIG. 6 illustrates an output image and its histogram generated by the dynamic range scaling of FIG. 3 according to one embodiment of the present disclosure.
[0094] Referring back to FIG. 3, for the image that has been exposed, fused, and gamma-corrected by blocks (312, 314, and 316) of FIG. 3, dynamic range scaling can be performed in block (318). By signal processing in block (318), an image (602) can be finally output. The image input to block (318) can be linearly scaled to fit the new dynamic range by Equation 5.
[0095]
[0096] In mathematical equation 5, and can be the maximum and minimum pixel values of image I, respectively. and may be the upper and lower boundaries of the new dynamic range, respectively. and can be determined based on mathematical formula 6.
[0097]
[0098] In mathematical expression 6, and may be the maximum pixel value and the minimum pixel value of the image (e.g., the normalized input image) output by exposure normalization of block (304) of FIG. 3, respectively. In one embodiment, and Instead of mathematical expression 6, it can be calculated in block (310) based on the inference of the neural network model (308).
[0099] After dynamic range scaling is performed, an image (602) can be output from block (318). Image of Fig. 5 Compared to the histogram (518) of (512), the histogram (604) of the image (602) of FIG. 6 may be stretched to the right. For example, the rightmost peak of the histogram (518) of FIG. 5 may be observed to have shifted to the right in the histogram (604). Consequently, through dynamic range scaling, the final image may have enhanced tones in the mid-tones and shadows.
[0100] FIG. 7 exemplarily illustrates the structure of a neural network model (700) according to one embodiment of the present disclosure.
[0101] Referring to FIG. 7, a neural network model (700) may output an inference based on a downsampled input image (702). In one embodiment, the neural network model (700) may correspond to the neural network model (308) of FIG. 3. Referring to FIG. 3, an input image to the ISP pipeline (302) may first be downsampled in block (306) and then input to the neural network model (308). In one embodiment, the input image may be downsampled at a ratio of 1 / 16.
[0102] In the illustrated embodiment, the size of the downsampled input image (702) is It could be. Here, is the number of images, is the height of the image, and may be a natural number corresponding to the size of the image. In one embodiment, the downsampled input image (702) may include four channels: three color channels (e.g., RGB channels) and an additional channel corresponding to additional information of the image.
[0103] In one embodiment, the additional information of the image may include the average pixel value of the input image and / or the variance of the input image. In one embodiment, the additional channel for the average pixel value of the input image is filled with the value 1. A channel of size may be created by multiplying that channel by the average pixel value of the input image. In one embodiment, the additional channel may be created by converting the input image to grayscale and multiplying the converted grayscale image by the average pixel value of the input image. Alternatively or additionally, the downsampled input image (702) may include an additional channel for the variance of the input image. The additional channel for the variance may be created in a manner similar to the manner in which the channel for the average pixel value of the input image is created as described above.
[0104] In one embodiment, the additional information of an image may include metadata contained in the corresponding raw image. For example, the metadata may include information related to the image sensor used to capture the image, the format of the image file, the size of the image, and the resolution of the image. In one embodiment, the metadata may include bit depth, white level, baseline exposure, ISO, orientation, aperture, and / or shutter speed.
[0105] The neural network model (700) may include four CNN (Convolution Neural Network) layers (704, 706, 708, 710). Each CNN layer has eight channels, kernels, and It can have a stride. CNN layers (704, 706, 708) can have Leaky ReLU (Rectified Linear Unit) as an activation function. The downsampled input image (702) can be sequentially passed through layers (704, 705, 708, 710) and then pooled in block (712).
[0106] In block (712), from the last layer (710) An image of the size may be output. Due to the stride of 22 in each layer, the height and width of the input image (702) may be 1 / 16 times. In block (712), the output of the last layer (710) may be pooled after being converted to an absolute value. In one embodiment, max pooling or average pooling may be performed on the output converted to an absolute value. For example, in block (712), for each channel, the maximum value among the absolute values of the values in the output of the last layer (710) may be output. For example, in block (712), the average of the absolute values of the values in the output of the last layer (710) may be output.
[0107] By pooling in block (712), For the downsampled input image of the dog, The values of the dog are output and can be used in blocks (714, 716, 718, 720). In one embodiment, blocks (714, 716, 718, 720) can be included in block (310) of FIG. 3. Blocks (714, 716, 718, 720) can be performed by the electronic device (100) of FIG. 1.
[0108] In block (714), The maximum of four values for a specific image among the dog images can be calculated. The maximum value calculated in block (714) is the parameter can be used as. In block (716), (the average of four values for a specific image) Х5+1 can be calculated. The value calculated in block (716) is a parameter can be used as. In block (718), the maximum of four values for a specific image can be calculated. The maximum value calculated in block (718) is a parameter can be used as. In block (720), (the maximum of four values for a specific image)+1 can be calculated. The maximum value calculated in block (720) is a parameter can be used as. About the dog image, {set of parameters , , , } can be obtained.
[0109] In one embodiment, the electronic device (100) can modify the method for deriving parameters of each of the blocks (714, 716, 718, 720). For example, the electronic device (100) can modify the calculation formulas of at least some of the blocks (714, 716, 718, 720) based on a request from a user of the electronic device (100) or a changed shooting condition of the image sensor (102). Accordingly, the parameters derived using the neural network model (700) can be adaptively changed without retraining the neural network model (700).
[0110] To derive optimized parameters corresponding to input images, a neural network model (700) can be trained. The neural network model (700) can be trained using an input image of an ISP pipeline associated with the neural network model (700) and an output image of a signal processing operation of the ISP pipeline. In one embodiment, referring to FIG. 3, to train the neural network model (700), an input image of the ISP pipeline (302) and the output image of the second exposure fusion of block (314) Loss using It can be calculated as in mathematical formula 7.
[0111]
[0112] In mathematical expression 7, is the input image is the width of, is the input image is the height of, is the input image can correspond to each channel. In mathematical expression 7, the operator can be defined as in mathematical formula 8.
[0113]
[0114] Referring to mathematical expression 8, the operator By the difference between a pixel and its four neighboring pixels (e.g., four pixels above, below, left, and right of the pixel) can be obtained. The neural network model (700) has a loss can be learned with the goal of minimizing the loss. Referring to Equations 7 and 8, the loss Minimizing the difference between adjacent pixel values can maximize the loss. The more this is minimized, the better the contrast can be.
[0115] In one embodiment, the loss L may be determined based on the characteristics of the camera module of the electronic device in which the neural network model (700) is mounted. For example, the loss L may be determined in a different manner from Equation 7, taking into account the characteristics of the image sensor of the camera module. Accordingly, the neural network model (700) may be trained so that the parameters obtained according to the inference of the neural network model (700) satisfy the characteristics intended by the manufacturer of the camera module. The neural network model (700) may be trained to cover corner cases associated with image capturing. Accordingly, when the neural network model (700) is used to obtain parameters for signal processing operations, it may not be necessary to preset various values for the parameters of the signal processing operations by taking corner cases into account.
[0116] FIG. 8 illustrates an example of an image output as parameters change according to one embodiment of the present disclosure.
[0117] More specifically, Fig. 8 shows the parameters and The images (802, 804, 806, 808, 810, 812, 814, 816, 818) output from the ISP pipeline (302) of FIG. 3 are illustrated when various ratios are multiplied. For example, the image (802) is calculated using parameters calculated in block (310). and It may be an image output from the ISP pipeline (302) when it is provided to the ISP pipeline (302) after being multiplied by 0.5. The image (810) is the original parameters calculated in the block (310) using the inference of the neural network model (308). and It could be.
[0118] The image output by the image signal processing of FIG. 3 may be customized by simply tuning the inferred parameters. For example, instead of retraining the neural network model (308) to output a brighter image than the image (810), the electronic device (100) may use the inference of the neural network model (308) to tune the parameters calculated in block (310). Multiply by 2, The second exposure fusion of block (314) can be performed using .
[0119] Comparing images (802, 804, 806, 808, 810, 812, 814, 816, 818), the parameters As the parameter increases, the shadow area becomes brighter, and the As this increases, the sky area becomes darker and clouds become more prominent. To output the desired image, the calculated parameters and This can be adjusted. For example, to output a brighter image, the parameter can be increased. The image signal processing of FIG. 3 may not require retraining of the neural network model (308) to customize the characteristics of the output image.
[0120] FIG. 9 illustrates an exemplary flowchart of a method (900) according to one embodiment of the present disclosure.
[0121] Referring to FIG. 9, the method (900) may include steps (902, 904, 906). In one embodiment, the method (900) may be performed by the electronic device (100) of FIG. 1. However, the present disclosure is not limited thereto, and steps (902, 904, 906) may be performed individually or in combination by any electronic device. The method according to one embodiment of the present disclosure is not limited to that illustrated in FIG. 9, and any one of the steps illustrated in FIG. 9 may be omitted, or steps not illustrated in FIG. 9 may be further included. In some embodiments, the order of at least some of the steps (902, 904, 906) may be changed.
[0122] In step (902), the electronic device (100) can obtain a first image from a raw image. In step (904), the electronic device (100) can obtain one or more parameters based on the first image using a neural network model. In step (906), the electronic device (100) can obtain a second image by performing one or more signal processing operations on the first image using the one or more parameters. In some embodiments, the one or more signal processing operations can include at least one of: exposure fusion, brightness adjustment, contrast adjustment, histogram equalization, tone adjustment, and dynamic range scaling.
[0123] In some embodiments, the step of obtaining one or more parameters from the first image may include the steps of: downsampling the first image and inputting the downsampled first image into a neural network model.
[0124] In some embodiments, the step of obtaining the second image may include: obtaining a first intermediate image darker than the first image using a first parameter of the one or more parameters; obtaining a second intermediate image brighter than the first image using the first intermediate image and a second parameter of the one or more parameters; and obtaining the second image by correcting the second intermediate image.
[0125] In some embodiments, the step of obtaining the first intermediate image may include: obtaining a third intermediate image darker than the first image from the first image using a first parameter; obtaining a first illumination map from the first image using a third parameter of the one or more parameters; and obtaining the first intermediate image from the first image, the third intermediate image, and the first illumination map. In some embodiments, the step of obtaining the first illumination map may include applying a guided filter to the first image using the third parameter.
[0126] In some embodiments, the step of obtaining the second intermediate image may include: obtaining a fourth intermediate image from the first intermediate image using the second parameter, the fourth intermediate image being brighter than the first intermediate image; obtaining a second illumination map from the first intermediate image using the fourth parameter of the one or more parameters; and obtaining the second intermediate image from the first intermediate image, the fourth intermediate image, and the second illumination map. In some embodiments, the step of obtaining the second illumination map may include applying a guided filter to the first intermediate image using the fourth parameter.
[0127] In some embodiments, the step of obtaining the second image may include: obtaining the second image by performing gamma correction and dynamic range scaling on the second intermediate image.
[0128] In some embodiments, the neural network model can be trained by: obtaining parameters from a third image using the neural network model, performing signal processing operations on the third image using the parameters to obtain a fourth image, calculating a loss function based on the third image and the fourth image, and updating the neural network model based on the calculated loss function.
[0129] FIG. 10 illustrates an exemplary block diagram of an electronic device (1000) according to one embodiment of the present disclosure.
[0130] Referring to FIG. 10, the electronic device (1000) may include a processor (1002), a memory (1004), a display (1006), a camera module (1008), and a communication interface (1010). In one embodiment, the electronic device (100) of FIG. 1 may be implemented in a similar manner to the electronic device (1000) of FIG. 10. For example, the image sensor (102) of the electronic device (100) may be included in the camera module (1008) of the electronic device (1000) (or may correspond to the camera module (1008)). The neural network model (104) of the electronic device (100) may correspond to the neural network model (1012) stored in the memory (1004) of the electronic device (1000). The ISP pipeline (108) of the electronic device (100) may be included in (or processed by) the processor (1002). In one embodiment, the camera module (1008) may include a separate processor for processing images, and in such an embodiment, the ISP pipeline (108) of the electronic device (100) may be included in (or processed by) the processor of the camera module (1008).
[0131] In FIG. 10, only essential components for explaining the function and / or operation of the electronic device (1000) are illustrated, and the components included in the electronic device (1000) are not limited as illustrated in FIG. 10. In one embodiment, the electronic device (1000) may be a portable device or a mobile device, and in such embodiments, the electronic device (1000) may further include a battery that supplies driving power to the processor (1002), the memory (1004), the display (1006), the camera module (1008), and the communication interface (1010).
[0132] The processor (1002) may execute one or more instructions of a program stored in the memory (1004). For example, the processor (1002) may execute a neural network model (1012) stored in the memory (1004). The processor (1002) may be configured with hardware components that perform arithmetic, logic, and input / output operations and image processing. Although the processor (1002) is illustrated as a single element in FIG. 10 , it is not limited thereto. In one embodiment of the present disclosure, the processor (1002) may be configured with one or more elements.
[0133] The processor (1002) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include at least one processor, and additionally or alternatively, may include various processing circuits. One or more processors may be individually and / or collectively configured to perform various functions described herein in a distributed fashion. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform multiple functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, the at least one processor may include a combination of processors that perform various functions of the disclosed functions in a distributed manner. The at least one processor may individually or in combination execute program instructions to achieve or perform various functions.
[0134] The processor (1002) may be implemented as a general-purpose processor such as a CPU (Central Processing Unit), an AP (Application Processor), a DSP (Digital Signal Processor), a graphics-only processor such as a GPU (Graphics Processing Unit), a VPU (Vision Processing Unit), or an AI-only processor such as an NPU (Neural Processing Unit), for example. The processor (1002) may be controlled to process input data according to predefined operating rules or an AI model. Alternatively, if the processor (1002) is an AI-only processor, the AI-only processor may be designed with a hardware structure specialized for processing a specific AI model.
[0135] The memory (1004) can store instructions, data structures, and program codes that can be read by the processor (1002). In one embodiment, the memory (1004) can store instructions that, when executed individually or in combination by the processor (1002), can cause the electronic device (1000) to perform at least some of the operations of the electronic device (100) described above with reference to FIGS. 1 to 9. For example, the processor (1002) can perform at least some of the operations described in FIGS. 1 to 9 by executing one or more instructions or codes stored in the memory (1004).
[0136] The memory (1004) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory, and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk, and / or a volatile memory such as a DRAM (Dynamic Random Access Memory) or an SRAM (Static Random Access Memory).
[0137] In one embodiment, the memory (1004) may store one or more instructions and / or program codes that cause the electronic device (1000) to process a raw image. For example, the memory (1004) may store instructions and / or program codes for implementing the functions of the neural network model (104) and the ISP pipeline (106) of FIG. 1. The memory (1004) may at least temporarily store various formulas for deriving one or more parameters for the ISP pipeline (106). Meanwhile, the elements stored in the memory (1004) described above are for convenience of explanation and are not necessarily limited thereto.
[0138] The display (1006) can output an image signal to the screen of the electronic device (1000) under the control of the processor (1002). For example, the display (1006) can output at least some of the images acquired in the process of the electronic device (1000) performing signal processing on the images, such as the raw image of FIG. 1 or the image output from the ISP pipeline (106) (e.g., the output image of FIG. 2 or FIG. 3).
[0139] In one embodiment, the display (1006) may include a touch panel. The touch panel may include one or more touch sensors that detect touch input. In one embodiment, a signal may be input through the touch panel to adjust the degree of image signal processing performed by the electronic device (1000). In one embodiment, a signal may be input through the touch panel to adjust at least some of one or more parameters acquired using the neural network model (104).
[0140] The camera module (1008) can capture an object to obtain a still image. For example, the camera module (1008) can obtain the raw image of FIG. 1. The image obtained by the camera module (1008) can be stored in the memory (1004). In one embodiment, the camera module (1008) can include a lens module, an image sensor, and / or an image processing module. In one embodiment, the electronic device (1000) can include a plurality of camera modules.
[0141] The communication interface (1010) can perform data communication with other external devices under the control of the processor (1002). In one embodiment, the communication interface (1010) can include communication circuit(s) that can perform data communication between the electronic device (1000) and other electronic devices by using at least one of data communication methods including a Local Area Network (LAN), a Wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication. In one embodiment, the electronic device (1000) can obtain a raw image captured by an external device through a communication interface (1010).
[0142] FIG. 11 illustrates an example of a user interface (UI) displayed on a display of an electronic device according to one embodiment of the present disclosure.
[0143] Referring to FIG. 11, the electronic device (1000) of FIG. 10 may display a graphic UI (1102) related to image signal processing on an area of the display (1006). The graphic UI (1102) may include an image (1104), an icon (1106) indicating that an image signal processing operation for adjusting a highlight of the image is currently being performed, an adjustment degree UI (1108) for receiving a user's input for the degree of highlight adjustment, a cancel icon (1112) for receiving a user's input for canceling the highlight adjustment, and a completion icon (1114) for receiving a user's input for completing the highlight adjustment.
[0144] In one embodiment, the brightness of a highlight region of an image may be adjusted by a highlight adjustment operation. The highlight adjustment operation may be associated with the first exposure fusion of block (312) and / or the second exposure fusion of block (314) of FIG. 3. For example, parameters for the first exposure fusion and / or parameters for the second exposure fusion may be adjusted to change the brightness of a highlight region of an image.
[0145] In response to a user input (e.g., a user input on a touch panel of the electronic device (1000)) to adjust a highlight of an image (1104), the electronic device (1000) may display at least a portion of a graphical UI (1102) on an area of the display (1006). In one embodiment, at least some of the UIs or icons described above may be omitted from the graphical UI, may be positioned at a different location on the graphical UI (1102) than the location illustrated in FIG. 11 , and / or UIs or icons associated with one or more functions provided to the user by the electronic device (1000) may be additionally positioned on the graphical UI (1102).
[0146] In response to receiving a user input for the auto highlight icon (1110), the electronic device (1000) may set the highlight level to an automatic preset value. For example, the electronic device (1000) may set parameters used in the image signal processing (300) of FIG. 3. , , , The raw image corresponding to the image (1104) can be set to the original values calculated using the neural network model (308) (e.g., values calculated in the manner described with reference to FIG. 7). Based on the set parameters, the electronic device (1000) can input the linear image generated from the raw image corresponding to the image (1104) into the ISP pipeline (302) of FIG. 3. The electronic device (1000) can display the output of the ISP pipeline (302). Accordingly, the result of the automatic highlighting can be provided to the user.
[0147] The adjustment degree UI (1108) may be a graphical interface that visualizes at least some of the scales between the minimum value (e.g., 0 [%]) and the maximum value (e.g., 100 [%]) of the highlight adjustment intensity. The adjustment degree UI (1108) may be a graphical interface that indicates at least some of a sequence of scales that divide the range of the highlight adjustment degree into equal intervals. The central scale of the adjustment degree UI (1108) may correspond to the value of the current highlight adjustment degree. For example, in the embodiment of FIG. 11, the current highlight adjustment degree may be 50 [%]. The current highlight adjustment degree may be displayed on the icon (1106). The image (1104) may be an output of the image signal processing (300) corresponding to the highlight adjustment degree.
[0148] In one embodiment, the electronic device (1000) may obtain a user's scroll input (or swipe input) for the adjustment degree UI (1108). In response to the user's scroll input for the intensity adjustment UI (1120), a portion of a sequence of scales displayed on the adjustment degree UI (1108) may change. For example, the user may move the scales of the adjustment degree UI (1108) left and right through a touch input, thereby changing the highlight adjustment degree. The electronic device (1000) may obtain a new highlight adjustment degree corresponding to the user's touch input and display the obtained new highlight degree on the icon (1106). The electronic device (1000) may modify parameters for the image signal processing (300) in response to the new highlight degree. In one embodiment, the new highlight degree may be determined based on the intensity of the user's touch input. For example, the electronic device (1000) may determine an increment or decrement of the highlight level based on the intensity of the user's touch input. The intensity of the user's touch input and the amount of change in the highlight level may be proportional.
[0149] In one embodiment, in response to receiving input from a user instructing to increase the degree of highlighting, the electronic device (1000) sets a parameter and / or , and a linear image generated from a raw image corresponding to the image (1104) based on the modified parameters can be input back into the ISP pipeline (302) of FIG. 3. The electronic device (1000) can display the image output from the ISP pipeline (302) based on the modified parameters on the display (1006).
[0150] In response to receiving a user input for the cancel icon (1112), the electronic device (1000) may cancel highlight adjustment for the image (1104). For example, the electronic device (1000) may stop the highlight adjustment operation without saving the highlight adjustment results for the image (1104).
[0151] In response to receiving a user input for the completion icon (1114), the electronic device (1000) may store a highlight adjustment result for the image (1104) and stop the highlight adjustment operation. For example, the electronic device (1000) may store the highlight adjustment result displayed on the display (1006) in the memory (1020) at the time of receiving the user input for the completion icon (1114). Thereafter, the electronic device (1000) may stop the highlight adjustment operation.
[0152] In one embodiment, in a manner similar to the highlight adjustment operation, the electronic device (1000) may receive a request from a user to adjust the degree of shadows in an image. In response to the request from the user, the electronic device (1000) may adjust at least some of the parameters obtained in the image signal processing operation (300) of FIG. 3 and may correct the input image based on the ISP pipeline (302) using the adjusted parameters. In response to a request from the user to darken the shadows, the electronic device (1000) may adjust the parameters and / or can be adjusted. For example, the electronic device (1000) may obtain parameters using a neural network model (308). can increase. For example, in response to a user request to brighten the shadow, the electronic device (1000) uses a neural network model (308) to obtain parameters can reduce.
[0153] In some embodiments, the electronic device (1000) may perform various image signal processing operations on a raw image using parameters obtained using a neural network model. For example, the ISP pipeline of the electronic device (1000) may include at least one of exposure fusion, brightness adjustment, contrast adjustment, histogram equalization, tone adjustment, and dynamic range scaling. At least one signal processing operation of the ISP pipeline may be performed using one or more parameters obtained using a neural network. The electronic device (1000) may provide a UI for adjusting the degree of the at least one signal processing operation to a user of the electronic device (1000) in a manner similar to the embodiment illustrated in FIG. 11. For example, the electronic device (1000) may provide a UI for adjusting the degree of the at least one signal processing operation to a user in a manner similar to the adjustment degree UI (1108). The electronic device (1000) may receive a request from a user to adjust the degree of the at least one signal processing operation. In response to a request from a user, the electronic device (1000) may adjust at least some of the parameters acquired using the neural network. The electronic device (1000) may re-run the ISP pipeline using the adjusted parameters and provide the output image back to the user.
[0154] In some embodiments of the present disclosure, an electronic device may include an ISP pipeline and one or more neural network models for obtaining parameters used in the ISP pipeline from an input image. The neural network models may infer values (e.g., values corresponding to characteristics of the input image) for calculating the parameters, rather than inferring the output image itself. Accordingly, the size of the neural network models may be reduced. In some embodiments, the electronic device may include a plurality of lightweight neural network models, and in such embodiments, the electronic device may be implemented with low complexity through various combinations of lightweight neural network models.
[0155] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0156] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0157] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various changes and modifications may be made based on the above description. For example, appropriate results may still be achieved if the described techniques are performed in a different order than described, and / or if components such as the described computer system or modules are combined or combined in a different manner than described, or if they are replaced or substituted with other components or equivalents.
Claims
1. A step of obtaining a first image (402) from a raw image; A step of obtaining one or more parameters based on the first image (402) using a neural network model; and A step of obtaining a second image by performing one or more signal processing operations on the first image (402) using the one or more parameters, A method wherein the neural network model is learned by obtaining parameters based on a learning image using the neural network model, performing one or more signal processing operations on the learning image using the parameters to obtain an output image, and updating the neural network model based on the learning image and the output image.
2. In paragraph 1, The step of obtaining one or more of the above parameters comprises: A step of downsampling the first image (402); and A method comprising the step of inputting the downsampled first image into the neural network model.
3. In paragraph 1 or 2, The steps of obtaining the second image are: A step of obtaining a first intermediate image (412) darker than the first image (402) by using a first parameter among the one or more parameters; A step of obtaining a second intermediate image (512) brighter than the first image (402) by using the first intermediate image (412) and a second parameter among the one or more parameters; and A method comprising the step of obtaining the second image by correcting the second intermediate image (512).
4. In paragraph 3, The steps of obtaining the above first intermediate image (412) are: A step of obtaining a third intermediate image (404) darker than the first image (402) from the first image (402) using the first parameter; A step of obtaining a first illumination map (406) from the first image (402) using a third parameter among the one or more parameters; and A method comprising the step of obtaining the first intermediate image (412) from the first image (P), the third intermediate image (404), and the first illumination map (406).
5. In paragraph 4, A method, wherein the step of obtaining the first illumination map (406) includes the step of applying a guided filter to the first image (402) using the third parameter.
6. In any one of paragraphs 3 to 5, The steps of obtaining the above second intermediate image (512) are: A step of obtaining a fourth intermediate image (504) brighter than the first intermediate image (412) from the first intermediate image (412) using the second parameter; A step of obtaining a second illumination map (508) from the first intermediate image (412) using a fourth parameter among the one or more parameters; and A method comprising the step of obtaining the second intermediate image (512) from the first intermediate image (412), the fourth intermediate image (504), and the second illumination map (508).
7. In paragraph 6, A method wherein the step of obtaining the second illumination map (508) includes the step of applying a guided filter to the first intermediate image (412) using the fourth parameter.
8. In any one of paragraphs 3 to 7, The steps of obtaining the second image are: A method comprising the step of obtaining the second image by performing gamma correction and dynamic range scaling on the second intermediate image (512).
9. In any one of paragraphs 1 to 8, A method in which the neural network model is learned by calculating a loss based on pixel values of the training image and pixel values of the output image, and updating the neural network model based on the calculated loss.
10. In paragraph 9, A method wherein the loss is calculated based on the variance of one or more adjacent pixels of a pixel of the output image, the average of the entire pixel values of the output image, or the entropy of the output image.
11. In any one of paragraphs 1 to 10, A method wherein said one or more signal processing operations comprise at least one of: exposure fusion, brightness adjustment, contrast adjustment, histogram equalization, tone adjustment, and dynamic range scaling.
12. A computer-readable recording medium having recorded thereon a program for performing the method of any one of claims 1 to 11 on a computer.
13. At least one processor (1002) comprising a processing circuit; and An electronic device (1000) comprising a memory (1004) storing one or more commands, the memory including one or more recording media, The one or more instructions are individually or in combination executed by the at least one processor (1002) to cause the electronic device (1000) to: Obtain a first image from a raw image; Obtaining one or more parameters based on the first image using a neural network model; and Acquire a second image by performing one or more signal processing operations on the first image using the one or more parameters, An electronic device wherein the neural network model is learned by obtaining parameters from a learning image using the neural network model, performing one or more signal processing operations on the learning image based on the parameters to obtain an output image, and updating the neural network model based on the learning image and the output image.
14. In paragraph 13, The one or more instructions are executed individually or in combination by the at least one processor (1002) to cause the electronic device (1000) to further: Obtaining a first intermediate image darker than the first image using a first parameter among the one or more parameters; Using the first intermediate image and a second parameter among the one or more parameters, a second intermediate image brighter than the first image is obtained; and An electronic device that obtains the second image by correcting the second intermediate image.
15. In paragraph 14, The one or more instructions are executed individually or in combination by the at least one processor (1002) to cause the electronic device (1000) to further: Using the first parameter, a third intermediate image darker than the first image is obtained from the first image; Obtaining a first illumination map from the first image using a third parameter among the one or more parameters; and An electronic device that obtains the first intermediate image from the first image, the third intermediate image, and the first illumination map.
Citation Information
Patent Citations
Sweet potato paste and manufacturing method thereof
KR1020230033931A
Roof Rack for Automobile
KR1020250140221A
Method, apparatus and system for automating the creation of online shopping malls that can be customized based on database linkage
KR102551399B1
Visibility determinations in physical space using evidential illumination values
US20240046605A1
KR20240022265A