Method and apparatus for image processing

The neural network-based image processing method addresses the challenge of separating illumination and reflection components in images under AC lighting by using a Retinex theory-inspired model, achieving improved accuracy and robustness in image classification and recognition.

CN115834996BActive Publication Date: 2025-07-15SAMSUNG ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210585609.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-15
Filing Date
2022-05-26
Publication Date
2025-07-15
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively decompose the illuminance and reflection components in the image, resulting in changes in the color and brightness of the light source in image processing that affect the recognition accuracy.

Method used

Using Retinex theory and deep learning method, the illumination map, reflection map and light source color are extracted from multiple images through illumination extraction model and reflection extraction model, and the image decomposition is used to generate white balanced images and high dynamic range images.

Benefits of technology

It improves the accuracy and robustness of image processing, reduces the impact of light source color and brightness changes on image recognition, and enhances the applicability and decomposition performance of image decomposition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115834996B_ABST
    Figure CN115834996B_ABST
Patent Text Reader

Abstract

A method and apparatus for image processing are provided. The apparatus includes: an image acquirer configured to acquire a plurality of images each having a different brightness; and one or more processors configured to extract an illuminance map of an input image and a light source color of the input image from the input image of the plurality of images and time-related information of the plurality of images based on an illuminance extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of Korean Patent Application No. 10-2021-0123322, filed on Sep. 15, 2021, in the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field

[0002] The following description relates to a method and apparatus for image processing. Background Art

[0003] As a means for solving the problem of classifying an input pattern into a specific group, an effective pattern recognition method may be applied to an actual computer. To solve the problem of classifying an input pattern into a specific group, a neural network may use a method having a learning ability. Through this method, the neural network may generate a mapping between an input pattern and an output pattern, which may be represented as a neural network having a learning ability. In addition, the neural network may have a generalization ability to generate a relatively correct output for an input pattern not used for learning based on the learning result. Summary of the Invention

[0004] The present invention is provided to introduce, in a simplified form, a selection of concepts that are further described below in the detailed description. The present invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0005] In one general aspect, an apparatus having image processing includes: an image acquirer configured to acquire a plurality of images each having a different brightness; and one or more processors configured to extract an illumination map of an input image and a light source color of the input image from the input image of the plurality of images and temporal correlation information of the plurality of images based on an illumination extraction model.

[0006] The image acquirer may be configured to acquire the plurality of images including one or more images having a brightness different from that of the input image.

[0007] To extract the illumination map and the light source color, the one or more processors may be configured to reshape data generated from the plurality of images based on a time frame by compressing color channels of the plurality of images into a single channel; and determine temporal correlation information based on the plurality of images and attention data generated from the reshaped data.

[0008] To extract the illumination map and the light source color, the one or more processors may be configured to extract a color map of each color channel from the input image using one or more convolutional layers of the illumination extraction model; and determine a light source color vector indicating the light source color based on the extracted color map of each color channel and a light source confidence map.

[0009] The one or more processors may be configured to: generate a reflection map from an input image using an illuminance map and a light source color vector.

[0010] The one or more processors may be configured to: generate a temporal gradient map as a light source confidence map by accumulating the differences of each time frame pixel-by-pixel from the plurality of images.

[0011] The one or more processors may be configured to: generate another illuminance map using the input image and the reflection map; and generate light source related information between the illuminance map and the another illuminance map as a light source confidence map.

[0012] The illuminance extraction model may include: a pyramid pooling layer that propagates output data to a subsequent layer, in which results of performing separate convolution operations on data pooled from the input data in different sizes are spliced to the input data.

[0013] The one or more processors may be configured to: generate a white balance image from the input image using the extracted light source color.

[0014] The one or more processors may be configured to: extract a reflection map of the same time frame as that of the illuminance map from the input image based on a reflection extraction model.

[0015] The one or more processors may be configured to: share feature data extracted from at least a part of layers of the reflection extraction model with the illuminance extraction model.

[0016] The image acquirer may be configured to: acquire the plurality of images captured under an alternating current (AC) light source.

[0017] The image acquirer may be configured to: acquire each of the plurality of images with different exposure times.

[0018] The processor may be configured to: generate a plurality of illuminance maps corresponding to respective time frames from the plurality of images using an illuminance extraction model; generate a plurality of reflection maps corresponding to respective time frames from the plurality of images using a reflection extraction model; and generate a synthesized image from the plurality of illuminance maps and the plurality of reflection maps.

[0019] The processor may be configured to: reconstruct a high dynamic range (HDR) image from the plurality of illuminance maps and the plurality of reflection maps based on an image fusion model.

[0020] In another general aspect, a method having image processing includes: acquiring a plurality of images each having a different brightness; extracting an illuminance map of an input image and a light source color of the input image from the input image of the plurality of images and temporal related information of the plurality of images based on an illuminance extraction model.

[0021] The steps of extracting the illuminance map and the light source color may include: shaping the data generated from the plurality of images by compressing the color channels of the plurality of images into a single channel based on time frames; and determining time-related information based on the plurality of images and the attention data generated from the shaped data.

[0022] The steps of extracting the illuminance map and the light source color may include: using one or more convolutional layers of an illuminance extraction model to extract a color map of each color channel from an input image; and determining a light source color vector indicating the light source color based on the color map of each color channel extracted and a light source confidence map.

[0023] The steps of extracting the illuminance map and the light source color may include: propagating output data from a pyramid pooling layer of an illuminance extraction model to subsequent layers, in which the results of performing separate convolution operations on data pooled from the input data in different sizes are concatenated to the input data.

[0024] The method may include: training an illuminance extraction model based on a loss, the loss being determined based on either or both of the illuminance map and the light source color.

[0025] In another general aspect, one or more embodiments include: a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform any one, any combination, or all of the operations and methods described herein.

[0026] In another general aspect, a method with image processing includes: using an illuminance extraction model to extract an illuminance map and a light source color of an input image among a plurality of images each having a different brightness; determining a light source color vector of the light source color based on color maps extracted for each color channel from the input image using a part of the illuminance extraction model; extracting a reflection map of the input image based on the illuminance map and the light source color vector; and generating a white balance image of the input image based on the illuminance map and the reflection map.

[0027] The step of extracting the reflection map may include: performing element-wise division on the input image using the illuminance map and the light source color vector.

[0028] The illuminance extraction model may include: an encoder part and a decoder part, and color maps extracted from the input image for each color channel are output from convolutional layers of the encoder part.

[0029] Other features and aspects will be apparent from the following detailed description, the drawings, and the claims. Description of the Drawings

[0030] Figure 1 An example of decomposing an image captured under illumination is shown.

[0031] Figure 2 An example of a computing device configured to perform image decomposition using an illuminance extraction model and a reflection extraction model is shown.

[0032] Figure 3 An example of an operation of extracting time-related information using an illuminance extraction model is shown.

[0033] Figure 4 An example of the structures of an illuminance extraction model and a reflection extraction model is shown.

[0034] Figure 5 An example of a pyramid pooling layer in an illuminance extraction model is shown.

[0035] Figure 6 An example of calculating a color vector is shown.

[0036] Figure 7 An example of calculating a confidence map is shown.

[0037] Figure 8 An example of extracting a reflection map using an illuminance extraction model is shown.

[0038] Figure 9 An example of an operation of generating a synthetic image using an image decomposition model is shown.

[0039] Figures 10 to 13 An example of training an image decomposition model is shown.

[0040] Figure 14 An example of an image processing device is shown.

[0041] Throughout the drawings and the detailed description, unless otherwise described or provided, the same reference numerals will be understood to refer to the same elements, features, and structures. For clarity, illustration, and convenience, the drawings may not be to scale, and the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated. Detailed Description

[0042] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent after understanding the disclosure of this application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. Additionally, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0043] Although terms such as "first" and "second" may be used to explain various components, elements, regions, layers or parts, these components, elements, regions, layers or parts are not limited by these terms. Instead, these terms are only used to separate one component, element, region, layer or part from another component, element, region, layer or part. For example, in the examples described herein, the first component, first element, first region, first layer or first part referred to may also be referred to as the second component, second element, second region, second layer or second part without departing from the teachings of the examples.

[0044] Throughout the specification, when an element (such as a layer, region or substrate) is described as being "on", "connected to" or "coupled to" another element, the element may be directly "on" the other element, directly "connected to" or "coupled to" the other element, or there may be one or more other elements therebetween. In contrast, when an element is described as being "directly on" another element, "directly connected to" or "directly coupled to" another element, there are no other elements therebetween. Similarly, for example, the expressions "between" and "directly between" and "adjacent" and "directly adjacent" may also be interpreted as described previously.

[0045] The terms used herein are for the purpose of describing particular examples only and are not intended to limit the disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. As used herein, the term "and / or" includes any one and any combination of any two or more of the associated listed items. It will also be understood that when used in this specification, the terms "comprises", "comprising" and "has" specify the presence of stated features, integers, steps, operations, elements, components and / or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. The use of the term "may" herein with respect to an example or embodiment (e.g., what an example or embodiment may include or achieve) means that there is at least one example or embodiment that includes or achieves such a feature, while all examples are not limited thereto.

[0046] Unless otherwise defined herein, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless otherwise defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with the context of the relevant art and this disclosure, and shall not be interpreted in an idealized or overly formal sense.

[0047] In the following, examples will be described in detail with reference to the accompanying drawings. Similar reference numerals shown in the respective drawings denote similar elements, and further descriptions related thereto will be omitted.

[0048] Figure 1 An example of an image captured under illumination and decomposed is shown.

[0049] In one example, an image (e.g., a high-speed image) 130 may be decomposed into an illumination map 141 and a reflection map 143 according to the Retinex theory. The illumination map 141 may refer to a map representing the level of light incident from the light source 110 onto the object 190 and / or the background. For example, the element value corresponding to each pixel of the image 130 in the illumination map 141 may represent the intensity of light incident on the point corresponding to the pixel in the scene captured in the image 130. The reflection map 143 refers to a map representing the level at which the object 190 and / or the background reflect the incident light. For example, the element value corresponding to each pixel of the image 130 in the reflection map 143 may represent the reflection coefficient of the point corresponding to the pixel in the scene captured in the image 130. Each element value of the reflection map 143 may represent the reflection coefficient according to each color channel of the color space, and may represent the reflection level of light having a wavelength corresponding to the color of the corresponding color channel. For example, in the case of the RGB color space, for each pixel, the reflection map 143 may include an element value indicating the reflection coefficient in the red channel, an element value indicating the reflection coefficient in the green channel, and an element value indicating the reflection coefficient in the blue channel. In addition, in addition to the illumination map 141 and the reflection map 143, the image processing device may extract a light source color 142 indicating the color 112 of the light source 110 from the image 130.

[0050] When only the brightness changes when capturing a plurality of images of the same scene, the illumination map 141 may change, and the reflection map 143 may be invariant (e.g., for a plurality of time frames corresponding to the plurality of images). For example, each of a series of images captured in consecutive frames may have a different brightness. For example, the illumination intensity of a light source 110 operated by an AC power supply (hereinafter, also referred to as an alternating current (AC) light source 110) may change according to the frequency of the AC power supply. When the AC power supply supplies sinusoidal power having a frequency of 60 hertz (Hz), although the AC power supply has a negative value, the illumination intensity of the AC light source 110 may increase and the brightness may increase. That is, the AC light source 110 having an AC power supply of 60 Hz may change the brightness to 120 Hz. The brightness change of 120 Hz may be captured by the high-speed camera 120.

[0051] According to the Retinex theory, for example, each of multiple images in a sequence of frames captured using a high-speed camera 120 under an AC light source 110 can be decomposed into a consistent (or invariant) reflectance map 143 and an illuminance map 141 with varying brightness for each time frame. That is, the image 130 captured by the high-speed camera 120 can show a sinusoidal brightness change 129 along the time axis, the reflectance map 143 can show a constant invariance 193 along the time axis, and the illuminance map 141 can show a sinusoidal brightness change 119 along the time axis. The captured image 130 can be represented as the multiplication of the illuminance map 141 and the reflectance map 143 for each time frame. As described above, since the reflectance map 143 shows a constant invariance 193 along the time axis, the brightness change 129 of the image 130 along the time axis can depend on the brightness change 119 of the illuminance map 141. Since the brightness change of the AC light source 110 is sinusoidal, this constraint of showing a sinusoidal brightness change even in the captured image 130 can be used to more accurately predict the illuminance map 141.

[0052] An image processing apparatus according to one or more embodiments can estimate (e.g., determine) the illuminance map 141, the reflectance map 143, and the light source color 142 from the image 130 based on an image decomposition model that is accurately trained or learned using a loss function designed considering the above constraint (e.g., showing a sinusoidal brightness change in the captured image 130). For example, the image processing apparatus can generate time-related information representing the brightness change of the AC light source 110 of the multiple images. The image processing apparatus can perform image decomposition using the generated time-related information and a single input image among the multiple images. The following refers to Figure 2 A non-limiting example of image decomposition is further described, and the following refers to Figure 10 A non-limiting example of the training of the image decomposition model and the loss function is further described.

[0053] The illuminance map 141, the reflectance map 143, and the light source color 142 can be estimated in a complex manner from an image decomposition model trained using a loss function designed based on the brightness change 119 of the illuminance map 141 and the invariance 193 of the reflectance map 143 for each time frame in the above brightness change of the AC light source 110. When using the time-related information and the consistency of the reflectance map 143, an image processing apparatus according to one or more embodiments can perform more accurate image decomposition. For example, the image decomposition model can use the brightness change characteristics of the AC light source 110 to learn the time characteristics and gradients of the AC light source 110. When estimating the light source color 142 without assuming white illumination, an image processing apparatus according to one or more embodiments can more accurately estimate the illuminance map 141 and the reflectance map 143. An image processing apparatus according to one or more embodiments can prevent deterioration of complex decomposition performance caused by the illumination brightness (e.g., low illuminance) and color distortion of the light source color.

[0054] In an image processing apparatus according to one or more embodiments, an image decomposition model may include an illuminance extraction model and a reflection extraction model, and an output result of the illuminance extraction model may be used as training data for training the reflection extraction model. Conversely, an output result of the reflection extraction model may be used as data for training the illuminance extraction model. Thus, although manual marking by humans is limited, a large amount of training data can be obtained.

[0055] As a reference, the image processing apparatus according to one or more embodiments is not limited to decomposing only the image 130 captured by the high-speed camera 120. Since an ordinary camera with an adjusted capture period and / or exposure time can capture a part of the brightness change, the above method can also be applied. For example, even the image 130 captured at a frames per second (FPS) different from the illumination period of the AC light source 110 may include a brightness change similar to the brightness change of the image 130 captured by the high-speed camera 120. When an ordinary camera with a fixed exposure time acquires multiple frames of images, although the period of the AC light source 110 is not given, the brightness may vary in the captured image 130. In addition, for the AC light source 110, the brightness change of the AC waveform can be predicted, and the brightness change according to multiple exposure times can be predicted to be a linear change. Thus, similarly, even for images captured at multiple exposure times, time-related information can be used for image decomposition based on the Retinex theory.

[0056] Figure 2 An example of a computing device configured to perform image decomposition using an illuminance extraction model and a reflection extraction model is shown.

[0057] Referring to Figure 2 , the image processing apparatus 200 according to one or more embodiments may include an image acquirer 201 (e.g., one or more sensors (such as one or more cameras)), a processor 202 (e.g., one or more processors), and a memory 203 (e.g., one or more memories).

[0058] The image acquirer 201 may acquire a plurality of images 210 each having a different brightness. According to one example, the image acquirer 201 may acquire a plurality of images 210, the plurality of images 210 including at least one image having a brightness different from the brightness of the input image 211 among the plurality of images 210. For example, the image acquirer 201 may acquire a plurality of images 210 captured under an AC light source. In Figure 2 which, the plurality of images 210 may include a first image I1 to an Nth image I N(e.g., corresponding to the first time frame to the Nth time frame). N represents an integer of 2 or greater, and i represents an integer greater than or equal to 1 and less than or equal to N. Here, an example of capturing multiple images 210 mainly under an AC light source is described. However, this is only provided as an example, and as another example, referring to Figure 9 , the image acquirer 201 can acquire multiple images 210 each captured with a different exposure time. The image acquirer 201 can be or include, for example, a camera sensor, but is not limited thereto. The image acquirer 201 can be a communicator configured to receive multiple images 210 from a device including another camera sensor through wired communication and / or wireless communication.

[0059] The memory 203 can store an image decomposition model. The image decomposition model can be in a machine learning structure trained to output an image decomposition result from an input image 211. The input image 211 can be the nth image among the N images. Here, n represents an integer greater than or equal to 1 and less than or equal to N. The image decomposition model can include an illuminance extraction model 230 and a reflection extraction model 240. The illuminance extraction model 230 can be in a machine learning structure trained to extract an illuminance map 231 from the input image 211 and time-related information 212. The time-related information 212 is shown as F in Figure 2 . The reflection extraction model 240 can be in a machine learning structure trained to extract a reflection map 241 from the input image 211. For example, the image decomposition model, the illuminance extraction model 230, and the reflection extraction model 240 can include a neural network 239. The neural network 239 can be, for example, a deep neural network (DNN). The DNN can include a fully connected network, a deep convolutional network, and a recurrent neural network. The neural network 239 can perform image decomposition by non-linearly mapping input data and output data based on deep learning, and the image decomposition includes the extraction of the reflection map 241, the extraction of the illuminance map 231, and the extraction of the light source color. In Figure 2 , for example, the reflection map 241 can represent the nth reflection map R of the nth time frame n , and the illuminance map 231 can represent the nth illuminance map L of the nth time frame n . Deep learning refers to a machine learning scheme for solving image processing problems from large datasets. The input data and the output data can be mapped through supervised or unsupervised learning of deep learning.

[0060] Referring to Figure 2 , the neural network 239 can include an input layer, a hidden layer, and an output layer. Each of the input layer, the hidden layer, and the output layer can include multiple nodes. Although it is shown in Figure 2 that the hidden layer includes three layers for ease of description, the hidden layer can include various numbers of layers (e.g., four or more layers). In addition, although it is shown in Figure 2The neural network 239 is shown to include a separate input layer for receiving input data, but the input data (e.g., the input image 211 and the time-related information 212) can be directly input into the hidden layer. In the neural network 239, the nodes of the layers other than the output layer can be connected to the nodes of the subsequent layer through links (e.g., connections) for sending output signals. The number of links can correspond to the number of nodes included in the subsequent layer. Such links can be referred to as connection lines. The output of the activation function including the weighted input of the nodes included in the previous layer can be input into each node included in the hidden layer. The weighted input is obtained by multiplying the input of the nodes included in the previous layer by the connection weights. The connection weights can be referred to as the parameters of the neural network 239. The activation function can include sigmoid, hyperbolic tangent (tanh), and rectified linear unit (ReLU), and non-linearity can be formed in the neural network 239 through the activation function. The weighted input of the nodes included in the previous layer can be input into each node included in the output layer.

[0061] For reference, the following refers to Figure 4 a non-limiting example of the structures of the image decomposition model, the illuminance extraction model 230, and the reflection extraction model 240.

[0062] The processor 202 can extract the illuminance map 231 of the input image 211 and the light source color of the input image 211 from at least one input image 211 among the multiple images 210 and the time-related information based on the illuminance extraction model 230. In addition, the processor 202 can extract the reflection map 241 of the corresponding input image 211 from at least one input image 211 among the multiple images 210 based on the reflection extraction model 240.

[0063] For example, the processor 202 can obtain multiple images 210 from consecutive frames through the image acquirer 201. For example, each of the multiple images 210 can have a different brightness. The processor 202 can select a single input image 211 from among the multiple images 210. The selected input image 211 can correspond to a single time frame among the multiple time frames. The processor 202 can also generate the time-related information 212 of the multiple images 210. The following refers to Figure 3Describe a non - limiting example of generating time - related information 212. The processor 202 can calculate (e.g., determine) an illuminance map 231 and a local color map 232 by applying an illuminance extraction model 230 to the input image 211 and the time - related information 212. The processor 202 can calculate a reflection map 241 by applying a reflection extraction model 240 to the input image 211. The processor 202 can extract a reflection map 241 of the same time frame as that of the illuminance map 231 from the input image 211 based on the reflection extraction model 240. In addition, the processor 202 can calculate a light source confidence map 220, and can determine a light source color vector 233 indicating the light source color based on the light source confidence map 220 and the local color map 232. Refer to the following with respect to Figure 6 Describe a non - limiting example of calculating the light source confidence map 220 and the light source color vector 233.

[0064] The image processing apparatus 200 of one or more embodiments can predict light source color information by using the illuminance map 231 and the reflection map 241 based on the Retinex theory. For example, the image processing apparatus 200 can supplementarily improve the estimation accuracy of each piece of information by predicting the light source color and the illuminance map 231 with high correlation as described above. Since the image decomposition model does not depend on a priori, the applicability and the decomposition performance can be improved. The image processing apparatus 200 of one or more embodiments can synthesize a further improved white - balance image 290 (e.g., ) and a multi - exposure fusion image. For example, the image processing apparatus 200 can generate the white - balance image 290 by performing an element - by - element operation (e.g., element - by - element multiplication) between the respective element values of the illuminance map 231 and the reflection map 241. However, this is provided only as an example, and the image processing apparatus 200 can alternatively or additionally generate the white - balance image 290 from the input image 211 using the extracted light source color. For example, the image processing apparatus 200 can generate the white - balance image 290 by dividing each pixel value of the input image 211 by the light source color vector 233.

[0065] The image processing apparatus 200 can be applied to a camera (e.g., an image acquirer 201) that uses artificial intelligence computing and / or a neural processor and server - oriented image processing in the field of depth image processing. In addition, the image processing apparatus 200 of one or more embodiments can generate an image with standardized illuminance by white - balancing as a pre - processing to reduce object recognition confusion caused by illuminance in image processing tasks including image classification, object tracking, optical flow calculation, and / or depth estimation.

[0066] Figure 3 Show an example of an operation of extracting time - related information (e.g., Figure 2 the time - related information 212) using the illuminance extraction model.

[0067] The image processing apparatus of one or more embodiments may generate information representing the temporal correlation between a plurality of images 310 (e.g., temporal correlation information 312). For example, when changes in the images of the same scene typically occur due to changes in the illumination intensity of the light source, it may be assumed that regions that change more along the time axis have more illuminance information. The image processing apparatus may extract the temporal correlation information 312 targeting the above regions. The temporal correlation information 312 may be a graph in which pixels of portions related to illuminance changes in a plurality of input images are emphasized.

[0068] For example, the image processing apparatus may shape the data 301 generated from a plurality of images 310 based on time frames by compressing the color channels of the plurality of images 310 into a single channel. For example, each of the plurality of images 310 may include pixels of height H and width W, and the image of each of the T time frames may include C-channel images. Here, each of H, W, T, and C represents an integer of 1 or greater. The data 301 compressed into a single channel may be H×W×T-dimensional data. The image processing apparatus may generate data 302a having HW×T dimensions and data 302b having T×HW dimensions by shaping the data 301 compressed into a single channel.

[0069] The image processing apparatus may calculate the temporal correlation information 312 based on the plurality of images 310 and attention data generated from the shaped data 302a and 302b. For example, the image processing apparatus may calculate HW×HW-dimensional data by multiplying the shaped data 302a and 302b. The image processing apparatus may generate attention data of H×W×C×T dimensions by performing a SoftMax operation 305 on the HW×HW-dimensional data. The image processing apparatus may generate temporal correlation information 312 of H×W×C×T dimensions by multiplying the attention data by the plurality of images 310.

[0070] The above temporal correlation information 312 may be a graph in which pixel values of temporal attention regions are emphasized using a non-local network scheme. The temporal correlation information 312 may be calculated from the input images themselves before training, so the computational amount may be minimized and the training time may be minimized.

[0071] Figure 4 An example of the structure of an illuminance extraction model (e.g., Figure 2 the illuminance extraction model 230) and a reflection extraction model (e.g., Figure 2 the reflection extraction model 240) is shown.

[0072] The image decomposition model according to the example may include an illuminance extraction model 430 and a reflection extraction model 440. As described above, the image processing apparatus of one or more embodiments may use the reflection extraction model 440 to extract a reflection map 449 from the input image 411. The image processing apparatus may use the illuminance extraction model 430 to extract a local color map 432 and an illuminance map 439 from the input image 411 and the time-related information 412.

[0073] For example, the reflection extraction model 440 may include a neural network including at least one convolutional layer and may be in an autoencoder structure based on VGG16. The reflection extraction model 440 may include, for example, an encoder part and a decoder part, and may include a shortcut connection that propagates data from a layer belonging to the encoder part to a layer corresponding to the decoder part. The encoder part may include one or more layers that extract the input image 411 and compress the extracted input image 411 into a representation vector (e.g., the representation vector 441), and the decoder part may include one or more layers that estimate the reflection map 449 from the compressed representation vector. Here, the structure of the reflection extraction model 440 is not limited thereto. For reference, the term "layer" used herein may also be represented as a block. For example, a convolutional layer may also be referred to as a convolutional block.

[0074] In addition, for example, the illuminance extraction model 430 may include a neural network including at least one convolutional layer and may be in an autoencoder structure based on VGG16. The illuminance extraction model 430 may include, for example, an encoder part and a decoder part. The encoder part of the illuminance extraction model 430 may include at least one pyramid pooling layer 431a. The pyramid pooling layer 431a may be disposed between convolutional layers in the encoder part. Refer to the following Figure 5 for a non-limiting example of the pyramid pooling layer 431a. The pyramid pooling layer 431a may also be referred to as a pyramid pooling block. The illuminance extraction model 430 may include a convolutional layer connected to the pyramid pooling layer 431a and outputting the local color map 432.

[0075] In addition, the image processing apparatus may share feature data extracted from at least a part of the layers of the reflection extraction model 440 with the illuminance extraction model 430. For example, the image processing apparatus may transmit the representation vector 441 compressed by the encoder part of the reflection extraction model 440 to the illuminance extraction model 430 (e.g., to the decoder part of the illuminance extraction model 430).

[0076] In one example, refer to Figure 4, the image processing apparatus can extract a local color map 432 for estimating the color of a light source from some layers 431 of the encoder part of the illuminance extraction model 430. The local color map 432 can include, for example, the element values of the red channel, the green channel, and the blue channel in the RGB color space as a color map, and the input image 411 is abstracted in the form of this color map. The local color map 432 can be extracted from an intermediate layer of the encoder part by the pyramid pooling layer 431a. Therefore, the global features can be applied to the local color map 432.

[0077] Figure 5 An example of a pyramid pooling layer in the illuminance extraction model is shown (e.g., Figure 4 the pyramid pooling layer 431a).

[0078] As described above with reference to Figure 4 the illuminance extraction model of one or more embodiments may include a pyramid pooling layer 530. The pyramid pooling layer 530 can be a residual block including a plurality of convolutional layers. For example, the pyramid pooling layer 530 can propagate the output data 539 to a subsequent layer by concatenating or combining the result of performing a separate convolution operation on data having different sizes (or referred to as dimensions) obtained by performing pooling on the input data 531 with the input data 531. Referring to Figure 5 , the image processing apparatus of one or more embodiments can obtain a plurality of pooled data 533a, 533b, 533c, and 533d from the input data 531 input to the pyramid pooling layer 530 through pooling 532. The plurality of pooled data 533a, 533b, 533c, and 533d can be pooled in different sizes. The image processing apparatus can perform a convolution operation on each of the plurality of pooled data 533a, 533b, 533c, and 533d. The image processing apparatus can generate data having the same size as the size of the input data 531 by performing upsampling 535 on each convolution data. The image processing apparatus can generate the output data 539 to be propagated to a subsequent layer by concatenating or combining the input data 531 and the upsampled data.

[0079] Figure 6 An example of calculating a color vector (e.g., Figure 2 the light source color vector 233) is shown.

[0080] The image processing apparatus of one or more embodiments can extract a color map of each color channel from the input image using at least one convolutional layer of the illuminance extraction model. For example, the color map of each color channel can be a local color map 632 (such as, for example, the local color map 432 described above with reference to Figure 4 )

[0081] Rather than using a neural network to extract the light source color, the image processing device can estimate a local color map 632 indicating the light source color of a local area, and can calculate a light source color vector 633 indicating the light source color through a weighted sum between the local color map 632 and the light source confidence map.

[0082] The image processing device can calculate a light source color vector 633 indicating the light source color based on the color map and the light source confidence map of each extracted color channel. For example, referring to Figure 6 , the image processing device can determine the weighted sum between the element value of the light source confidence map and the element value of the color map of each color channel as the color value of the corresponding color channel. That is, the light source color vector 633 can include a red value, a green value, and a blue value as a vector with a 3×1 dimension.

[0083] The processor of the image processing device can generate a temporal gradient map 620 as the light source confidence map by accumulating the differences of each time frame pixel by pixel from the plurality of images. In the input image, it can be assumed or determined that an area with a relatively high temporal gradient of the image (e.g., greater than or equal to a predetermined threshold) has a relatively high confidence in the illuminance value.

[0084] Figure 7 An example of calculating the confidence map is shown.

[0085] The image processing device according to one or more embodiments can generate another illuminance map 791 using the input image 711 and the reflection map 741. For example, in addition to the illuminance map 731 extracted from the input image 711 and the time-related information 712 through the illuminance extraction model 730, the image processing device can also generate other illuminance maps 791. The image processing device can generate another illuminance map 791 using the input image 711 and the reflection map 741 extracted through the reflection extraction model 740. The image processing device can generate another illuminance map 791 by dividing the input image 711 by the reflection map 741 and the light source color vector 733.

[0086] The image processing device can generate illuminance-related information 720 between the illuminance map 731 and another illuminance map 791 (e.g., ) as the light source confidence map. For example, the image processing device can generate light source-related information 720 by calculating the correlation of each position and / or each region between the illuminance map 731 and another illuminance map 791. The light source-related information 720 can be a map representing the correlation of each element between the illuminance map 731 and another illuminance map 791. The image processing device can use the light source-related information 720 to estimate the light source color vector 733 from the local color map 732, rather than using Figure 6 the temporal gradient map 620.

[0087] According to one example, the image processing apparatus may generate a white balance image by dividing the input image 711 by the light source color vector 733.

[0088] Figure 8 An example of extracting a reflection map using an illuminance extraction model is shown.

[0089] The image processing apparatus according to one or more embodiments may generate a reflection map 841 from the input image 811 using the illuminance map 831 and the light source color vector 833. Although mainly with reference to Figures 1 to 8 An example in which the image decomposition model includes a reflection extraction model and an illuminance extraction model 830 is described, but in another non-limiting example, the image decomposition model may include the illuminance extraction model 830 without including the reflection extraction model.

[0090] The image processing apparatus may extract the illuminance map 831 from the input image 811 and the time-related information 812 using the illuminance extraction model 830. Similar to the above description, the local color map 832 and the light source color vector 833 may be extracted from the front end of the illuminance extraction model 830 based on the light source confidence map 820. The image processing apparatus may extract the reflection map 841 by performing an element-wise division on the input image 811 using the illuminance map 831 and the light source color vector 833.

[0091] The image processing apparatus may generate a new image (e.g., a white balance image and a high dynamic range (HDR) image) using one or a combination of at least two of the extracted illuminance map 831, reflection map 841, and light source color vector 833.

[0092] Figure 9 An example of an operation of generating a composite image using an image decomposition model is shown.

[0093] The image acquirer according to one or more embodiments may acquire a plurality of images 910 each captured at a different exposure time (e.g., Figure 9 Exp1 to Exp4 shown in

[0094] The image processing apparatus according to one or more embodiments may generate a reflection map 941 and an illuminance map 931 by selecting the input image 911 from among the plurality of images 910 and by applying the reflection extraction model 940 and the illuminance extraction model 930 to the selected input image 911. Even for another image among the plurality of images 910, the image processing apparatus may repeat the above-described image decomposition operation. For reference, although not shown in Figure 9Although not shown, time-related information, a light source confidence map, and a light source color vector can be extracted. As a reference, a luminance change pattern different from the luminance change pattern of the AC light source described above (e.g., a linear increase) is presented in the images captured each time at multiple exposure times. Therefore, a loss function designed for multiple exposure times can be used to train the image decomposition model.

[0095] As described above, the processor of the image processing apparatus according to one or more embodiments can generate a plurality of illuminance maps 939 corresponding to respective time frames from a plurality of images 910 using the illuminance extraction model 930. The processor can generate a plurality of reflection maps 949 corresponding to respective time frames from the plurality of images 910 using the reflection extraction model 940. The processor can generate a synthesized image 951 from the plurality of illuminance maps 939 and the plurality of reflection maps 949. For example, the processor can reconstruct an HDR image from the plurality of illuminance maps 939 and the plurality of reflection maps 949 based on the image fusion model 950. The image fusion model 950 can be a machine learning model designed and trained to output an HDR image from the plurality of illuminance maps 939 and the plurality of reflection maps 949. However, it is provided only as an example, and the generation of the synthesized image 951 is not limited thereto. The image processing apparatus can reconstruct the synthesized image 951 by multiplying the average illuminance image of the illuminance map 939 by the average reflection image of the reflection map 949. In addition, the synthesized image 951 is not limited to an HDR image. An image having a white balance, illuminance, color, exposure time, and dynamic range specified by a user can be generated by fusing the illuminance map 939 and the reflection map 949.

[0096] Figures 10 to 13 An example of training an image decomposition model is shown.

[0097] Figure 10 The overall training and angular loss of the image decomposition model are shown. As a reference, Figure 10 has the same structure as Figure 2 and Figure 10 the loss described in Figure 2 is not limited to being applied only to Figures 7 to 9 For example, even in the structure of Figure 10 the loss described in

[0098] An image model construction device (e.g., an image processing device) of one or more embodiments may represent a device that constructs an image decomposition model. For example, the image model construction device may generate and train an image decomposition model (e.g., an illuminance extraction model 1030 and a reflection extraction model 1040). The operation of constructing the image decomposition model may include the operations of generating and training the image decomposition model. The image processing device may decompose an input image into an illuminance map, a reflection map, and a light source color based on the image decomposition model. However, it is provided only as an example, and the image model construction device may be integrated with the image processing device.

[0099] Referring to Figure 11 , the image model construction device may perform Siamese training. For example, the image model construction device may construct a first image decomposition model 1110a and a second image decomposition model 1110b (e.g., a neural network or a deep learning network) that share the same weight parameters for a first training image 1101 and a second training image 1102. The image model construction device may backpropagate the loss calculated using a first temporary output 1108 (Fx1) output from the first image decomposition model 1110a and a second temporary output 1109 (Fx2) output from the second image decomposition model 1110b. For example, the image model construction device may calculate the loss by applying a Siamese network to the images of each time frame, and may backpropagate the loss each time. Since testing can be performed without considering the number of frames of the image, training can be performed in a direct current (DC) lighting environment, a natural light environment, a single image, and an alternating current (AC) lighting environment.

[0100] According to one example, the image model construction device may perform training using a total loss that includes multiple losses. The image model construction device may repeat the parameter update of the image decomposition model until the total loss converges and / or until the total loss becomes less than a critical loss. The total loss may be expressed as, for example, Equation 1 below.

[0101] Equation 1:

[0102] L tot = L recon + L invar + L smooth + L color + L AC

[0103] In Equation 1, L tot represents the total loss, L recon represents the reconstruction loss 1091 between the input image and the reconstructed image, L invar represents the invariant loss 1092 between the reflection maps, L smooth represents the smooth loss 1093 for the illumination form, Lcolor Indicates a color loss of 1094, and L AC Indicates a luminance fitting loss of 1095.

[0104] The reconstruction loss 1091 refers to a loss function representing the satisfaction level of the illumination map, reflection map, and light source color obtained through the entire network for the Retinex theory, and can be designed as, for example, Equation 2 below.

[0105] Equation 2:

[0106]

[0107] In Equation 2, M j Indicates the intensity mask 1080 corresponding to the j-th time frame among N time frames, I j Indicates the input image corresponding to the j-th time frame, R i Indicates the reflection map corresponding to the i-th time frame, L j Indicates the illumination map corresponding to the j-th time frame, c j Indicates the light source color vector corresponding to the j-th time frame, and α ij Indicates an arbitrary coefficient. In Figure 10 can be calculated according to the convolution operation of the reflection map R i , the illumination map L j and the light source color vector c j * represents the convolution operation. The above Equation 2 can use the L1 function to represent the constraint loss, making the convolutional multiplication of the reflection map, illumination map, and light source color vector the same as the input image. Here, to prevent a reduction in accuracy in the light saturation region, for the remaining region determined based on the intensity mask 1080 that does not include the saturation region, a loss including the reconstruction loss 1091 calculated between the temporary output image and the input image based on the illumination map and reflection map can be used to train the illumination extraction model 1030. That is, in Equation 2, the intensity mask 1080 can have a value of 0 for the saturation region, and the image model construction device can exclude the calculation of the reconstruction loss 1091 for the saturation region.

[0108] For example, the invariant loss 1092 can be designed as Equation 3 below.

[0109] Equation 3:

[0110]

[0111] Equation 3 can represent the mean squared error (MSE) loss, which is designed such that the reflection maps (R i , R j ) of all time frames of multiple images can be the same.

[0112] For example, the smooth loss 1093 can be designed as Equation 4 below.

[0113] Equation 4:

[0114]

[0115] Equation 4 can be designed as a total variation L2 regularizer to reflect the smooth nature of the illuminance.

[0116] Figure 12 The color loss 1094 is described.

[0117] For example, the color loss 1094 can be designed as Equation 5 below.

[0118] Equation 5:

[0119] L color = L CC + L HS

[0120] In Equation 5, L cc represents the color constancy loss function, and L HS represents the hue saturation loss function.

[0121] For example, the color constancy loss function can be designed as Equation 6 and Equation 7 below.

[0122] Equation 6:

[0123]

[0124] Equation 7:

[0125] Γ t = mean(M g * C local )

[0126] Equation 6 can represent the error between the illuminance estimated using the illuminance extraction model 1030 and the actual illuminance. Equation 7 can represent the light source color vector Γ local at the t-th time frame calculated by the weighted sum between the estimated local color map C g and the temporal gradient map M t . Here, the function mean() represents the function for calculating the average value. The above Equation 6 can represent the angular error between the color vector Γ t and the ground truth illuminance vector Γ g .

[0127] In addition, according to an example, the hue saturation loss L HSTo train the illuminance extraction model 1030. For example, the hue-saturation loss LHS can be expressed as Equation 8 below.

[0128] Equation 8:

[0129]

[0130] Equation 8 can be a loss function representing the difference in hue and the difference in saturation between the reflectance map and the input image. In the input image 1210 from which the illuminance component has been removed and the reflectance map 1220 based on the HSV (hue (H), saturation (S), and value (V)) color space 1200, the hue and saturation can appear the same regardless of the illuminance. The input image (e.g., white balance image) from which the illuminance has been removed using the ground truth illuminance value 1210 and the reflectance map (R i ) 1220 can be respectively converted into a hue value H(x) and a saturation value S(x). The image model construction device can calculate the L1 loss so that the hue component and the saturation component between the two images can be the same.

[0131] Refer to Figure 13 to describe the luminance fitting loss 1095.

[0132] According to one example, the illuminance extraction model 1030 can be trained using a loss including the luminance fitting loss 1095 between the luminance of the illuminance map and the luminance of the light source extracted for multiple time frames. For example, the luminance fitting loss 1095 can be designed as Equation 9 and Equation 10 below.

[0133] Equation 9:

[0134]

[0135] Equation 10:

[0136]

[0137] In Equation 9, represents the average value of the illuminance values 1310 in the illuminance map for the t-th time frame. Equation 10 refers to a function representing the illumination intensity curve 1320 according to the Gauss-Newton method for sinusoidal regression modeling. f ac represents the AC frequency of the light source, and f cam represents the frequency of the camera. and off represent the offset variables updated during the regression. Equation 9 can represent the difference between the average value of the illuminance values 1310 and the illumination intensity curve 1320 according to the illuminance phase of the light source according to the Gauss-Newton method.

[0138] Figure 14 is a flowchart showing an example of the image processing method.

[0139] In operation 1410, an image processing apparatus according to one or more embodiments may acquire a plurality of images each having a different luminance.

[0140] In operation 1420, the image processing apparatus may extract an illuminance map of the input image and a light source color of the input image from at least one input image among the plurality of images and time-related information of the plurality of images based on an illuminance extraction model.

[0141] However, in a case of operations other than Figure 14 the image processing apparatus may perform, in time series and / or in parallel, one or a combination of at least two of the operations described above with reference to Figures 1 to 13 An image processing apparatus according to one or more embodiments may also perform white balance, illuminance component estimation, and reflection component estimation. In addition, the image processing apparatus may exhibit further improved white balance performance and accurate estimation performance. The image processing apparatus may exhibit performance that is robust to low illuminance noise and artifacts that may occur around strong lighting components.

[0142]

[0143] Here, regarding Figures 1 to 14 ​The described light source, high-speed camera, image processing device, image acquirer, processor, memory, light source 110, high-speed camera 120, image processing device 200, image acquirer 201, processor 202, memory 203, and other devices, apparatuses, units, modules, and components are implemented by or represent hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate, controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). The processor or computer can be implemented by one or more processing elements (such as logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve the desired result). In one example, the processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) to perform the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For simplicity, the singular terms "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or the processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component, or two or more hardware components, can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can implement a single hardware component, or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of different processing configurations including: single processor, independent processor, parallel processor, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.

[0144] performing the operations described in this application Figure 1 - 14The method shown is performed by computing hardware (e.g., by one or more processors or computers), which is implemented to execute instructions or software as described above to perform the operations performed by the method in this application. For example, a single operation, or two or more operations, can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors, or a processor and a controller, and one or more other operations can be performed by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, can perform a single operation, or two or more operations.

[0145] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and perform the method as described above can be written as a computer program, code segment, instruction, or any combination thereof to individually or jointly direct or configure one or more processors or computers to operate as a machine or special-purpose computer to perform the operations performed by the hardware components and method as described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by one or more processors or computers. In another example, the instructions or software include high-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the figures and the corresponding descriptions in the specification, which disclose algorithms for performing the operations performed by the hardware components and method as described above.

[0146] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement the hardware components and execute the methods described above, along with any associated data, data files, and data structures, can be recorded, stored, or fixed on one or more non-transitory computer-readable storage media, or can be recorded, stored, or fixed on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage devices, hard disk drives (HDD), solid state drives (SSD), flash memory, cartridge memory (such as, multimedia card or micro card (e.g., Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and any other device that is configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to one or more processors or computers such that the one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a networked computer system such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0147] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of this application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered only as descriptive and not for purposes of limitation. The description of a feature or aspect in each example will be considered applicable to similar features or aspects in other examples. Appropriate results can be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or are replaced or supplemented by other components or their equivalents.

Claims

1. An apparatus for image processing, the apparatus comprising: an image acquirer configured to: acquire a plurality of images each having a different brightness; and one or more processors configured to: extract an illumination map of an input image and a light source color of the input image from the input image of the plurality of images and time-related information of the plurality of images based on an illumination extraction model, wherein, in order to extract the illumination map and the light source color, the one or more processors are configured to: reshape data generated from the plurality of images based on time frames by compressing color channels of the plurality of images into a single channel; and determine time-related information according to the plurality of images and attention data generated from the reshaped data.

2. The device according to claim 1, wherein, The image acquirer is configured to: acquire the plurality of images including one or more images having a brightness different from the brightness of the input image.

3. The device according to claim 1, wherein In order to extract the illumination map and the light source color, the one or more processors are configured to: extract a color map of each color channel from the input image using one or more convolutional layers of the illumination extraction model; and determine a light source color vector indicating the light source color based on the extracted color map of each color channel and a light source confidence map.

4. The apparatus according to claim 3, wherein, The one or more processors are configured to: generate a reflection map from the input image using the illumination map and the light source color vector.

5. The apparatus according to claim 3, wherein, The one or more processors are configured to: generate a time gradient map as a light source confidence map by accumulating differences of each time frame pixel by pixel from the plurality of images.

6. The apparatus according to claim 3, wherein, The one or more processors are configured to: generate another illumination map using the input image and the reflection map; and generate light source-related information between the illumination map and the another illumination map as a light source confidence map.

7. The device according to claim 1, wherein, The illumination extraction model includes: a pyramid pooling layer that propagates output data to a subsequent layer, and obtains the output data by concatenating results of performing separate convolutional operations on data of different sizes obtained by performing pooling on input data with the input data.

8. The device according to any one of claims 1 to 6, wherein, The one or more processors are configured to: generate a white balance image from the input image using the extracted light source color.

9. The device according to any one of claims 1 to 6, wherein, The one or more processors are configured to: extract a reflection map of a time frame same as the time frame of the illumination map from the input image based on a reflection extraction model.

10. The apparatus according to claim 9, wherein, The one or more processors are configured to: share feature data extracted from at least a part of layers of the reflection extraction model with the illumination extraction model.

11. The device according to claim 1, wherein, The image acquirer is configured to: acquire the plurality of images captured under an AC light source.

12. The device according to claim 1, wherein, The image acquirer is configured to: acquire each of the plurality of images with different exposure times.

13. The apparatus according to claim 1, wherein, The one or more processors are configured to: generate a plurality of illumination maps corresponding to respective time frames from the plurality of images using the illumination extraction model; generate a plurality of reflection maps corresponding to respective time frames from the plurality of images using the reflection extraction model; and generate a synthetic image from the plurality of illumination maps and the plurality of reflection maps.

14. The apparatus according to claim 13, wherein, The one or more processors are configured to: reconstruct a high dynamic range image from the plurality of illumination maps and the plurality of reflection maps based on an image fusion model.

15. A method for image processing, the method comprising: acquiring a plurality of images each having a different brightness; and Based on the illuminance extraction model, extract the illuminance map of the input image and the light source color of the input image from the input image of the plurality of images and the time-related information of the plurality of images. Wherein, the steps of extracting the illuminance map and the light source color include: Shaping the data generated from the plurality of images by compressing the color channels of the plurality of images into a single channel based on time frames; and Determining the time-related information according to the plurality of images and the attention data generated from the shaped data.

16. The method according to claim 15, wherein, The steps of extracting the illuminance map and the light source color include: Extracting a color map of each color channel from the input image using one or more convolutional layers of the illuminance extraction model; and Based on the color map of each extracted color channel and the light source confidence map, determining a light source color vector indicating the light source color.

17. The method according to claim 15, wherein, The steps of extracting the illuminance map and the light source color include: Propagating the output data from the pyramid pooling layer of the illuminance extraction model to subsequent layers, and obtaining the output data by concatenating the results of performing separate convolution operations on data of different sizes obtained by performing pooling on the input data with the input data.

18. The method according to claim 15 further comprises: Training the illuminance extraction model based on a loss, the loss being determined based on either or both of the illuminance map and the light source color.

19. A method for image processing, the method comprising: Using an illuminance extraction model to extract the illuminance map of an input image and the light source color among a plurality of images each having a different brightness; Determining a light source color vector of the light source color based on the color map extracted for each color channel from the input image using a part of the illuminance extraction model; Extracting a reflection map of the input image based on the illuminance map and the light source color vector; and And Generating a white balance image of the input image based on the illuminance map and the reflection map.

20. The method according to claim 19, wherein The step of extracting the reflection map includes: Applying element-wise division to the input image using the illuminance map and the light source color vector.

21. The method according to claim 19, wherein, The illuminance extraction model includes: an encoder part and a decoder part, and outputs a color map extracted from the input image for each color channel from the convolutional layer of the encoder part.

22. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method according to any one of claims 15 to 21.

Citation Information

Patent Citations

  • High-reliability, high-power, high-brightness blue laser diode system and manufacturing method thereof

    KR1020210123322A

  • Low-illumination image enhancement method based on Retinex and deep learning

    CN111968044A