Image processing apparatus, image processing method, and computer program
The image processing device and method address color constancy by dividing color correction into two steps with inverse transformations, enhancing color consistency and task performance under varying lighting conditions.
Patent Information
- Application Number
- JP2025071229
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-30
- Filing Date
- 2025-04-23
- Publication Date
- 2025-11-12
AI Technical Summary
Existing image processing technologies struggle to achieve color constancy under varying lighting conditions, affecting tasks that require color consistency and comparability such as image stitching and object recognition.
An image processing device and method that divides color correction into two steps: transforming an input image to an intermediate image in the original color space and then adjusting it to a predetermined lighting condition using inverse transformations, with a feature extraction unit calculating the necessary parameters.
The method effectively adjusts images to specified lighting conditions, improving color consistency and enabling better performance in tasks like image stitching and object recognition.
Smart Images

Figure 2025169194000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the technical field of image processing, and more particularly to an image processing device, an image processing method, and a computer program for adjusting image colors. [Background technology]
[0002] In the fields of computer vision and image processing, color constancy is one of the important issues. Achieving color constancy usually involves algorithms that adjust image colors so that they match the colors that can be perceived under given lighting conditions (e.g., standard lighting conditions). This is very important for tasks that require color consistency and comparability, such as image stitching, object recognition, and any application where a machine needs to make decisions based on color. Summary of the Invention [Problem to be solved by the invention]
[0003] An object of the present invention is to provide an image processing device, an image processing method, and a computer program for adjusting the color of an image under predetermined lighting conditions. [Means for solving the problem]
[0004] According to one aspect of the present invention, there is provided an image processing device, comprising: a first transformation unit that receives an image output by a photographing device as an input image and performs a first set of transformations on the input image to obtain intermediate images, the intermediate images being images in an original color space of the photographing device; a second transformation unit that performs a second set of transformations on the intermediate image according to an inverse order of transformations in the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; and A feature extraction unit is included to calculate the parameters required for the first set of transforms and the second set of transforms.
[0005] According to another aspect of the present invention, there is provided an image processing method, comprising: receiving an image output by a photographing device as an input image, and performing a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in the original color space of the photographing device; and performing a second set of transformations on the intermediate image in an order reverse to that of the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; A feature extraction unit calculates the parameters required for the first set of transformations and the second set of transformations.
[0006] According to another aspect of the present invention, there is provided a computer-readable storage medium carrying a program product including computer-readable instruction code which, when readable and executed by a computer, causes the computer to perform an image processing method according to the present invention.
[0007] According to another aspect of the present invention, there is further provided a computer program and a computer program product for implementing the image processing method according to the present invention. [Effects of the Invention]
[0008] According to the image processing device, image processing method and computer program of the present invention, the image color correction step is divided into two steps: from the input image to the electronic negative film (intermediate image) and from the electronic negative film to the corrected image, and the calculation method from the electronic negative film to the input image is made to match the calculation method from the electronic negative film to the corrected image, so that the input image can be adjusted to a color under a specified lighting condition or a specified lighting condition. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a configuration of an image processing apparatus according to an embodiment of the present invention; [Figure 2] 1 is a block diagram showing a configuration of an image processing device according to a first embodiment of the present invention. [Figure 3] 1 is a simplified configuration diagram of an image processing device according to a first embodiment of the present invention; [Figure 4] 3 is an image correction flowchart of the image processing apparatus according to the first embodiment of the present invention. [Figure 5] FIG. 2 is a block diagram showing the configuration of a global feature extraction unit in the image processing device according to the first embodiment of the present invention. [Figure 6] FIG. 2 is a block diagram showing the configuration of a local feature extraction unit in the image processing device according to the first embodiment of the present invention. [Figure 7] FIG. 10 is a block diagram showing the configuration of an image processing device according to a second embodiment of the present invention. [Figure 8] FIG. 10 is a simplified configuration diagram of an image processing device according to a second embodiment of the present invention. [Figure 9] 10 is a flowchart of image correction performed by an image processing apparatus according to a second embodiment of the present invention. [Figure 10] FIG. 10 is a block diagram showing the configuration of an image processing device according to a third embodiment of the present invention. [Figure 11] FIG. 10 is a simplified configuration diagram of an image processing device according to a third embodiment of the present invention. [Figure 12] 10 is a flowchart of image correction performed by an image processing apparatus according to a third embodiment of the present invention. [Figure 13] 1 is a flowchart of an image processing method according to an embodiment of the present invention. [Figure 14] 1 is a block diagram showing an exemplary configuration of a general-purpose personal computer capable of implementing an image processing apparatus and method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. Note that the following embodiments are merely illustrative and are not intended to limit the scope of the present invention.
[0011] 1 is a diagram illustrating an image processing apparatus according to an embodiment of the present invention. As shown in FIG. 1, an image processing apparatus 100 according to an embodiment of the present invention may include a first conversion unit 110, a second conversion unit 120, and a feature extraction unit .
[0012] The first transformation unit 110 receives an image output by a photographing device as an input image and performs a first set of transformations on the input image to obtain an intermediate image, where the intermediate image is an image in the original color space of the photographing device. For example, the intermediate image is an electronic negative film of the photographing device obtained by the first set of transformations. Note that the electronic negative film here is not completely the same as the real (actual) electronic negative film of the photographing device, but is an image that is close to or equivalent to the real electronic negative film and is in the same color space (original color space) as the real electronic negative film. In addition, the image output by the photographing device is an image in the output color space of the photographing device.
[0013] The second transformation unit 120 can perform a second set of transformations on the intermediate image in a reverse order to that of the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is the inverse process of the corresponding transformation in the first set of transformations. In the present invention, the output image may be referred to as a corrected image.
[0014] As mentioned above, the order of the transformations in the second set of transformations is the reverse of the order of the transformations in the first set of transformations, and each transformation in the second set of transformations is the inverse of the corresponding transformation in the first set of transformations. For example, the first set of transformations may include a first transformation, a second transformation, and a third transformation, and the second set of transformations may include a fourth transformation, a fifth transformation, and a sixth transformation. In such a case, the fourth transformation is the inverse of the third transformation, the fifth transformation is the inverse of the second transformation, and the sixth transformation is the inverse of the first transformation.
[0015] For example, the first set of transformations and the second set of transformations in the present invention may each be a set of arithmetic operations, and the arithmetic operation process of each transformation in the second set of transformations may be the inverse process of the corresponding transformation in the first set of transformations. In this case, the first transformation unit 110 and the second transformation unit 120 may be configured as one or more calculation units corresponding to different operations in the corresponding set of arithmetic operations. Each of these calculation units performs a corresponding operation corresponding to a certain transformation in the set of arithmetic operations. Note that the first set of transformations simultaneously converts the input image into its original color space while obtaining an electronic negative of the input image, while the second set of transformations simultaneously adjusts the input image to a predetermined lighting condition while restoring the intermediate image to the color space of the input image. Therefore, although the corresponding transformation processes in both sets are reversible, the parameters used in the corresponding transformations are not necessarily completely identical. That is, the parameters of some transformations in the second set of transformations are related to the requirements of the predetermined lighting condition. The predetermined lighting condition may also be a standard lighting condition.
[0016] The feature extraction unit 130 can calculate parameters necessary for the first set of transformation and the second set of transformation. For example, the feature extraction unit 130 can be realized by a neural network. By training the feature extraction unit 130, the first transformation unit 110, and the second transformation unit 120, the overall model parameters of the image processing device 100 can be fixed. As a result, when the image processing device 100 is used, the feature extraction unit 130 and the first transformation unit 110 can obtain the necessary parameters suitable for transforming an input image into an electronic negative film and adjusting the electronic negative film to a predetermined lighting condition. Furthermore, the first transformation unit 110 and the second transformation unit 120 can each perform a corresponding set of transformation by performing a respective set of arithmetic operations using the obtained parameters necessary for the transformation.
[0017] Typically, when a photographic device captures a photograph / image (picture), it first acquires an electronic negative film, then performs RGB demosaicing, noise reduction, white balance and color space conversion, style conversion, image scaling, mapping to an output color space (e.g., sRGB (standard Red Green Blue) color space, P3 color space, etc.), JPEG (Joint Photographic Experts Group) / HEIC (High Efficiency Image Format) compression, etc., before outputting the image. However, all of these image signal processing algorithms are performed within the photographic device and cannot be applied to images in, for example, sRGB color space obtained from the photographic device (e.g., final video).
[0018] In the present invention, the first conversion unit 110 first converts the image output by the photographing device into an electronic negative film to obtain an image in the original color space of the photographing device. That is, the first conversion unit 110 converts the image from the color space of the image output by the photographing device to the original color space of the photographing device. The first conversion unit 110 may also perform other conversions, such as an inverse white balance conversion. The second conversion unit 120 can convert the intermediate image (electronic negative film) back to the color space of the input image that needs to be processed. Similarly, the second conversion unit 120 may also perform other conversions, such as a white balance conversion. The conversion performed by the second conversion unit 120 can adjust the lighting conditions to the desired predetermined lighting conditions.
[0019] As a result, the image processing device 100 in an embodiment of the present invention divides the image color correction step into two steps: from the input image to the electronic negative film (intermediate image) and from the electronic negative film to the corrected image, and makes the calculation method from the electronic negative film to the input image consistent with the calculation method from the electronic negative film to the corrected image, thereby adjusting the input image to a specified lighting condition or a color under a specified lighting condition.
[0020] <Applicable scenes> The present invention can be applied to tasks requiring color consistency and comparability, such as image stitching, object recognition, and any application where a machine needs to make decisions based on color. For example, the present invention can better recognize or classify objects based on their color characteristics. The objects may be, for example, images obtained by operations such as photography. For example, specific application scenarios may include identifying or classifying products in a supermarket, or tracking the disappearance of a specific object (e.g., a thief, a person / object that needs protection, etc.) in an image or video.
[0021] This allows the present invention to improve the performance of downstream color-related tasks.
[0022] An image processing device 200 according to a first embodiment of the present invention will now be described in conjunction with FIGS. 2 to 6. FIG. 2 is a block diagram of the image processing device 200 according to the first embodiment of the present invention. As shown in FIG. 2, the image processing device 200 may include a first transformation unit 210, a second transformation unit 220, a feature extraction unit 230, a first adjustment unit 240, and a second adjustment unit 250. The first transformation unit 210 may include a first transformation subunit 211, a second transformation subunit 212, a third transformation subunit 213, and a fourth transformation subunit 214. The second transformation unit 220 may include a first inverse transformation subunit 221, a second inverse transformation subunit 222, a third inverse transformation subunit 223, and a fourth inverse transformation subunit 224. The feature extraction unit 230 may include a first global feature extraction module 231, a local feature extraction module 232, and a second global feature extraction module 233. Note that the term "inverse transformation" in this description refers to the transformation process in the first set of transformations and does not limit the type of transformation performed by the second transformation unit 220. 1. Furthermore, the first conversion unit 210, the second conversion unit 220, and the feature extraction unit 230 correspond to the first conversion unit 110, the second conversion unit 120, and the feature extraction unit 130 in FIG. 1, and therefore the description regarding FIG. 1 can be similarly applied to the image processing device 200.
[0023] In this embodiment, the first set of transformations may sequentially include a first spatial transformation, a de-style transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations may sequentially include a white balance transformation, a second inverse spatial transformation, a style transformation, and a first inverse spatial transformation. Correspondingly, the first transformation subunit 211 may perform the first set of spatial transformation, the second transformation subunit 212 may perform the de-style transformation, the third transformation subunit 213 may perform the second spatial transformation, and the fourth transformation subunit 214 may perform the inverse white balance transformation. Correspondingly, the first inverse transformation subunit 221 may perform the white balance transformation, the second inverse transformation subunit 222 may perform the second inverse spatial transformation, the third inverse transformation subunit 223 may perform the style transformation, and the fourth inverse transformation subunit 224 may perform the first inverse spatial transformation.
[0024] For example, if the input image is in the sRGB color space, the first space transformation transforms the input image from the sRGB color space to a linear RGB color space, and the second space transformation transforms the image from the standard color space CIE XYZ to the original color space of the image capture device. The first inverse space transformation transforms the image from the linear RGB color space to the sRGB color space, and the second inverse space transformation transforms the image from the original color space of the image capture device to the standard color space CIE XYZ. Of course, the present invention is not limited to this. If the input image is in the P3 color space, the first space transformation can transform from the P3 color space to a linear RGB color space, and the first inverse space transformation can transform from the linear RGB color space to the P3 color space. Other transformations remain the same.
[0025] Fig. 3 is a simplified configuration diagram of an image processing device according to a first embodiment of the present invention, and Fig. 4 is an image correction flowchart of the image processing device according to the first embodiment of the present invention. As shown in Figs. 3 and 4, the local feature extraction module 232 may include two local feature extraction modules (a first local feature extraction module and a second local feature extraction module) connected in parallel to extract different parameters used for both destyling and style transfer. Of course, the local feature extraction module 232 in the present invention may be implemented using a single local feature extraction module, and in such a case, the local feature extraction module 232 may output multiple parameters used for destyling and style transfer.
[0026] 2 to 4, the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 are connected in series and are also connected in series with the subunits of the first transformation unit 210, so that the feature extraction unit 230 and the first transformation unit 210 can provide parameters for the first set of transformation and the second set of transformation. The principle of the image processing device 200 will be described in detail below in conjunction with FIGS. 3 and 4. For the sake of simplicity and ease of explanation, only the first global feature extraction module, the first local feature extraction module, the second local feature extraction module, and the second global feature extraction module are shown in FIGS. 3 and 4, and the mathematical operations performed by different transformation subunits / inverse transformation subunits are replaced by operators of the corresponding transformation subunits / inverse transformation subunits.
[0027] As shown in FIGS. 2 to 4, first, an input image (original picture captured by a photographing device) I in sRGB is input to the first global feature extraction module 231, which calculates and obtains the parameter γ (and thereby obtains 1 / γ). in sRGBis also input to the first transformation subunit 211, and the first transformation subunit 211 performs a power transformation on each pixel point of the input image based on the parameter γ provided by the first global feature extraction module 231, that is, performs a 1 / γ power operation to obtain I s RGB =I in 1 / γ where I in I in sRGB This step is used to remove the nonlinear power transformation that simulates the perception of the human optic nerve and convert from the nonlinear sRGB color space to the linear RGB color space. The parameter γ is used to perform the gamma transformation, also known as perceptual encoding, which aims to remap linear colors to better accommodate the nonlinear response of the human visual system to radiant power.
[0028] The first adjustment unit 240 (corresponding to the two operators in the dotted box in FIG. 3) adjusts the image I obtained after the inverse operation of the gamma transformation by the first transformation sub-unit 211. s RGB and Original Picture I in sRGB After performing the addition with , the result can be sent to two parallel local feature extraction modules to calculate the nonlinear transformation fitting parameter matrices M and A. This process is for fitting steps such as style transfer and color enhancement in image processing. After the first local feature extraction module and the second local feature extraction module respectively calculate and obtain A (and -A) and M (and 1 / M), the second transformation subunit 212 (represented by two operators in the dotted line box in Figure 3) calculates I obtained by the first transformation subunit 211. s RGB By using , dot product with 1 / M and then add with -A, the result after stylization removal is
[0029]
number
[0030] Preferably, the first adjustment unit 240 adjusts the first transformed image I input to the local feature extraction module. s RGB For example, the first adjusting unit 240 is configured to adjust a first proportional coefficient of the image I s RGB is multiplied by a proportional coefficient α1 and then the original input image I in sRGB This proportionality coefficient is a channel-independent learnable 3-dimensional vector. For example, during training, α1 is set to 10 -4 Initializing to is used to accelerate convergence and keep the network structure stable. Taking the original image as input and multiplying the processed image by a learnable factor helps the model learn accurately from the beginning, and adding additional learnable proportions to the processed image can improve performance.
[0031] The second adjustment unit 250 (in FIG. 3 for two operators in a box of points) adjusts the calculation result I of the second transformation subunit 212. ct XYZ and Original Picture I in sRGB After performing the addition, the result is passed to the second global feature extraction module 233, which then calculates and outputs the parameter T ideal ,T ct -1 ,W ct -1 ,W ideal Among them, the parameter W ct -1 and W ideal represents the white balance transformation matrix, which may be, for example, a single 3x3 diagonal matrix. ct -1represents the white balance matrix from the current color temperature (with respect to the input image that needs to be processed) to the original white balance, and the parameter W ideal may represent the white balance matrix from the original white balance to the ideal color temperature (under a given desired lighting condition). ct -1 represents the transformation matrix from the standard color space CIE XYZ to the original color space (raw-RGB) of the image capture device under the current lighting conditions, and the parameter T ideal represents the transformation matrix from the original color space (raw-RGB) of the image capture device to the standard color space CIE XYZ under the desired lighting conditions. For example, the parameter T ct -1 and T ideal is a square matrix of order 3. These two transformation matrices may have different values as the color temperature changes, so they need to be estimated separately. Therefore, introducing a different T matrix can make the transformation more accurate.
[0032] Preferably, the second adjustment unit 250 adjusts the second transformed image I input to the second global feature extraction module 233. ct XYZ The second adjusting unit 250 is configured to adjust a second proportional coefficient of the image I ct XYZ is multiplied by a proportional coefficient α2 and then the original input image I in sRGB This proportionality coefficient is a channel-independent, learnable three-dimensional vector. For example, during training, α2 is set to 10 -4 Initializing to is used to accelerate convergence and maintain the stability of the network structure. The proportional coefficients α1 and α2 can be independent coefficients. Using the original image as input and multiplying the processed image by a learnable factor helps the model learn accurately from the beginning, and adding an additional learnable proportionality to the processed image can improve performance.
[0033] The second global feature extraction module 233 calculates the parameter Tct -1 to the third conversion subunit 213, and the parameter W ct -1 The third transformation subunit 213 can provide the image after the second spatial transformation to the fourth transformation subunit 214.
[0034]
number
[0035]
number
[0036] The fourth transformation sub-unit 214 provides the calculated intermediate image to the second transformation unit 220. The second global feature extraction module 233 also provides the calculated parameter T ideal and W ideal to the second transformation unit 220. Then, the first inverse transformation sub-unit 221 performs a white balance transformation to convert the image W ideal I raw and a second inverse transform subunit 222 performs a second inverse spatial transform to obtain the image T ideal W ideal I raw and a third inverse transformation subunit 223 performs style transformation to obtain the image
[0037]
number
[0038]
number
[0039] This results in an output image I out is the input original image I insRGB This corresponds to the result of a series of transformations on the , and the process is as shown in equation (1) below.
[0040]
number
[0041] As mentioned above, the above description has been given using an example in which the image output by the image capture device is an image in the sRGB color space, but the image capture device can also output images in other color spaces, such as the P3 color space. The present invention can also be applied to input images in other color spaces. In this case, it is sufficient to change the mapping matrix for the different space.
[0042] In the above calculation process, the feature extraction unit 230 and the first transformation unit 210 provide parameters for the first set of transformation and the second set of transformation. That is, only the feature extraction unit 230 and the first transformation unit 210 participate in the parameter calculation, and the second transformation unit 220 only uses the parameters obtained by the calculation. The entire image correction process can be seen in conjunction with FIG. 4. In FIG. 4, the forward process (the calculation process of the feature extraction unit and the first transformation unit) is represented by a solid line, and the backward process (the calculation process of the second transformation unit) is represented by a dotted line. As shown in FIG. 4, the parameters are calculated only in the forward process, and the parameters themselves are not calculated in the backward process but are directly used.
[0043] In the present invention, the corresponding parameters and proportional coefficients α1 and α2 of the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 are determined by training the image processing device 200. When using the trained image processing device 200, the parameters and proportional coefficients α1 and α2 of the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 determined by training are used to calculate the parameters required for the first set of transformations and the second set of transformations. Furthermore, when training the image processing device, the inputs are different images of different scenes and under different lighting conditions, and the outputs are images of the corresponding scenes under predetermined lighting conditions. The training process does not require the true values of the electronic negative film, but only requires the above-mentioned input images and output images under the desired lighting conditions. After the model parameters are fixed by training, the second transformation unit directly uses the transformation parameters calculated by the feature extraction unit and the first transformation unit.
[0044] When training the image processing device 200, the same scene was photographed with the same photographing device under different color temperature and color style settings using the same electronic negative film I. raw Therefore, in the training process, we increase the L1 loss function for electronic negative film. Therefore, the total loss function formula is L=L1(Iout sRGB ,I gt sRGB )+λL1(I1 raw ,I2 raw ), where λ is the proportionality coefficient and L1(I out sRGB ,I gt sRGB ) represents the difference between the output image and the true value (e.g., expressed in L1 norm), and L1(I1 raw ,I2 raw ) represents the difference between intermediate images calculated from different images of the same scene by the same image capture device (e.g., expressed in L1 norm). By way of example, λ may take the value 0.1.
[0045] In the image processing device 200 according to an embodiment of the present invention, the parameters required for the first set of transformation and the second set of transformation are extracted using a global feature extraction module, a local feature extraction module, and a first transformation unit, which are connected in series. Compared to a case in which the global feature extraction module and the local feature extraction module are connected in parallel and output the transformation parameters simultaneously, the network structure transmits the calculation results of the above steps in series and allows more useful information to be incorporated into the transformation parameters, thereby achieving better color adjustment results.
[0046] In addition, by dividing the image color correction step into two steps, from the input image to the electronic negative film (intermediate image) and from the electronic negative film to the corrected image, and by making the calculation method from the electronic negative film to the input image consistent with the calculation method from the electronic negative film to the corrected image, the input image can be effectively adjusted to a specified lighting condition.
[0047] The structures of the global feature extraction modules (first global feature extraction module and second global feature extraction module) and the local feature extraction modules (first local feature extraction module and second local feature extraction module) will be described below with reference to Figures 5 and 6. Figure 5 is a block diagram showing the configuration of a global feature extraction unit in an image processing device according to a first embodiment of the present invention, and Figure 6 is a block diagram showing the configuration of a local feature extraction unit in an image processing device according to the first embodiment of the present invention.
[0048] As shown in Figure 5, the global feature extraction module is composed of a first convolution module, a cross-attention module, a first normalization module, a forward propagation network module, and a linear fully connected layer. The difference between the first global feature extraction module 231 and the second global feature extraction module 233 lies in the length of the query vector. The query vector is a trainable vector of a predetermined length and fixed dimensions. Its length varies depending on the number of dimensions that need to be queried. For example, when querying a 3D diagonal matrix, the length is 3, and when querying a 3D matrix, the length is 9.
[0049] As shown in Figure 6, the local feature extraction module consists of a first convolution module, a first normalization module, a first fully connected layer, a second convolution module, a second fully connected layer, a second normalization module, a forward propagation network module, and a linear fully connected layer. β1 and β2 are learnable proportional coefficients. The two local feature extraction modules have the same structure.
[0050] The sub-modules in the global feature extraction module and the local feature extraction module may be implemented in any available manner in the prior art, and therefore detailed description thereof will be omitted here.
[0051] An image processing device 300 according to a second embodiment of the present invention will now be described with reference to Figures 7 to 9. Figure 7 is a block diagram showing the configuration of the image processing device according to the second embodiment of the present invention, Figure 8 is a simplified configuration diagram of the image processing device according to the second embodiment of the present invention, and Figure 9 is an image correction flowchart of the image processing device according to the second embodiment of the present invention.
[0052] 7 , the image processing device 300 may include a first transformation unit 310, a second transformation unit 320, a feature extraction unit 330, a first adjustment unit 340, and a second adjustment unit 350. The first transformation unit 310 may include a first transformation subunit 311, a second transformation subunit 312, a third transformation subunit 313, a fourth transformation subunit 314, and a fifth transformation subunit 315. The second transformation unit 320 may include a first inverse transformation subunit 321, a second inverse transformation subunit 322, a third inverse transformation subunit 323, a fourth inverse transformation subunit 324, and a fifth inverse transformation subunit 325. In addition, the feature extraction module 330 may include a first global feature extraction module 331, a local feature extraction module 332, and a second global feature extraction module 333.
[0053] The difference between the image processing device 300 and the image processing device 200 is the addition of a fifth transformation subunit 315 and a fifth inverse transformation subunit 325. The other subunits and feature extraction modules are the same as those of the image processing device 200, and detailed descriptions thereof will be omitted here.
[0054] As mentioned above, the image after undergoing the inverse operation of the gamma conversion of the first conversion sub-unit 311 is I s RGB and it is in the linear RGB color space. The linear RGB color space and the standard color space CIE XYZ have a fixed mapping matrix P -1 Conversely, there is a mapping matrix P from the standard color space CIE XYZ to the linear RGB color space. This transformation matrix is independent of external factors such as lighting, and is specified by the CIE Institute under standard D65 lighting, i.e.,
[0055]
number
[0056] Therefore, in this embodiment, after the first transformation sub-unit 311 performs the first spatial transformation, the fifth transformation sub-unit 315 performs another spatial transformation (by the mapping matrix) to obtain I s XYZ Thereafter, the first adjustment unit 340 provides input to the local feature extraction module 332 based on the input image and the output of the fifth transformation sub-unit 315.
[0057] Note that in the first embodiment, the mapping matrix from the linear RGB color space to the standard color space CIE XYZ is fitted to the parameters M and A during the training of the image processing device 200, so that after undergoing the destyled transformation process, the image can be directly transformed into the standard color space CIE XYZ.
[0058] In the second embodiment, the first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a de-style transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, a second inverse spatial transformation, a style transformation, a third inverse spatial transformation, and a first inverse spatial transformation. Correspondingly, the first transformation sub-unit 311 performs the first spatial transformation, i.e., transforms the input image from the sRGB color space to the linear RGB color space, the fifth transformation sub-unit 315 performs the third spatial transformation, i.e., transforms the image from the linear RGB color space to the standard color space CIE XYZ, the second transformation sub-unit 312 performs the de-style transformation, the third transformation sub-unit 313 performs the second spatial transformation, i.e., transforms the image from the standard color space CIE XYZ to the original color space of the image capture device, and the fourth transformation sub-unit 314 performs the inverse white balance transformation.
[0059] Accordingly, the first inverse transformation subunit 321 can perform white balance transformation, the second inverse transformation subunit 322 can perform second inverse space transformation, i.e., transformation from the original color space of the image capture device to the standard color space CIE XYZ, the third inverse transformation subunit 323 can perform style transformation, the fifth inverse transformation subunit 325 can perform third inverse space transformation, i.e., transformation of the image from the standard color space CIE XYZ to the linear RGB color space, and the fourth inverse transformation subunit 324 can perform first inverse space transformation, i.e., transformation of the input image from the linear RGB color space to the sRGB color space.
[0060] As shown in FIGS. 7 to 9, the first global feature extraction module 331 extracts the input image I in sRGB The parameter γ can be obtained based on the above equation and provided to the first transformation subunit 311 and the fourth inverse transformation subunit 324.
[0061] The first transformation subunit 311 converts the input image I in sRGB , and the parameter γ provided by the first global feature extraction module 331, I s RGB =I in 1 / γ The fifth conversion subunit 315 can obtain I s RGB and based on the transformation matrix P, I s XYZ =P -1 I in 1 / γ can be obtained and provided to the first adjusting unit 340 and the second converting sub-unit 312.
[0062] The first adjustment unit 340 adjusts the image I s XYZ is multiplied by a proportional coefficient α1 and then the original input image I in sRGBand provide the result to the local feature extraction module 332. The local feature extraction module 332 can obtain the parameters M and A and provide them to the second transformation subunit 312 and the third inverse transformation subunit 323. The second transformation subunit 312 can obtain the parameters M and A and provide them to the second transformation subunit 312 and the third inverse transformation subunit 323.
[0063]
number
[0064] The second global feature extraction module 333 calculates the parameter T ct -1 to the third conversion subunit 313, and the parameter W ct -1 The third transformation subunit 313 can provide the image after the second spatial transformation to the fourth transformation subunit 314.
[0065]
number
[0066]
number
[0067] The fourth transformation sub-unit 314 provides the calculated intermediate image to the second transformation unit 320. The second global feature extraction module 333 also provides the calculated parameter T ideal and Wideal to the second transformation unit 320. Then, the first inverse transformation sub-unit 321 performs a white balance transformation to convert the image W ideal I raw and a second inverse transform subunit 322 performs a second inverse spatial transform to obtain the image T ideal W ideal I raw and a third inverse transformation subunit 323 performs style transformation to obtain the image
[0068]
number
[0069]
number
[0070]
number
[0071] So the output image is
[0072]
number
[0073] In the image processing device 300 according to the embodiment of the present invention, the parameters required for the first and second set of transformations are extracted using a serially connected global feature extraction module, a local feature extraction module, and a first transformation unit, and compared to a case in which the global feature extraction module and the local feature extraction module are connected in parallel and simultaneously output the transformation parameters, the network structure serially transmits the calculation results of the above steps, allowing more useful information to be incorporated into the transformation parameters, thereby achieving better color adjustment results. Furthermore, the image color correction step is divided into two steps: from the input image to an electronic negative film (intermediate image) and from the electronic negative film to the corrected image, and the calculation method from the electronic negative film to the input image is consistent with the calculation method from the electronic negative film to the corrected image, thereby effectively adjusting the input image to a predetermined lighting condition.
[0074] An image processing device 400 according to the third embodiment of the present invention will now be described with reference to Figures 10 to 12. Figure 10 is a block diagram showing the configuration of the image processing device according to the third embodiment of the present invention, Figure 11 is a simplified configuration diagram of the image processing device according to the third embodiment of the present invention, and Figure 12 is a flowchart showing image correction in the image processing device according to the third embodiment of the present invention.
[0075] 10 , the image processing device 400 may include a first transformation unit 410, a second transformation unit 420, a feature extraction unit 430, a first adjustment unit 440, and a second adjustment unit 450. The first transformation unit 410 may include a first transformation subunit 411, a second transformation subunit 412, a third transformation subunit 413, a fourth transformation subunit 414, a fifth transformation subunit 415, and a sixth transformation subunit 416. The second transformation unit 420 may include a first inverse transformation subunit 421, a second inverse transformation subunit 422, a third inverse transformation subunit 423, a fourth inverse transformation subunit 424, a fifth inverse transformation subunit 425, and a sixth inverse transformation subunit 426. The feature extraction module 430 may include a first global feature extraction module 431, a local feature extraction module 432, and a second global feature extraction module 433.
[0076] The difference between the image processing device 400 and the image processing device 200 is the addition of a fifth transformation subunit 415, a sixth transformation subunit 416, a fifth inverse transformation subunit 425, and a sixth inverse transformation subunit 426. The other subunits and feature extraction modules are similar to those of the image processing device 200, and detailed descriptions thereof will be omitted here. Furthermore, the fifth transformation subunit 415 and the fifth inverse transformation subunit 425 in the image processing device 400 perform the same processing as the fifth transformation subunit 315 and the fifth inverse transformation subunit 325 in the image processing device 300, and therefore the above description of the fifth transformation subunit 315 and the fifth inverse transformation subunit 325 can be similarly applied to the fifth transformation subunit 415 and the fifth inverse transformation subunit 425.
[0077] In this embodiment, compared to the image processing devices 200 and 300, an illumination conversion ratio k is introduced, which is used to represent the ratio between the exposure amount under the current color temperature and the exposure amount under the ideal color temperature (desired predetermined illumination condition). This proportionality coefficient is k, which is the ratio of the exposure amount under the current color temperature to the exposure amount under the original white balance. ct -1 and k from the exposure at the original white balance to the exposure at the ideal color temperature ideal It may be divided into
[0078] In this embodiment, the first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, an inverse exposure compensation transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, an exposure compensation transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform.
[0079] Accordingly, the first conversion sub-unit 411 performs a first space conversion, i.e., converts the input image from the sRGB color space to the linear RGB color space; the fifth conversion sub-unit 415 performs a third space conversion, i.e., converts the image from the linear RGB color space to the standard color space CIE XYZ; the second conversion sub-unit 412 performs a destyle conversion; the third conversion sub-unit 413 performs a second space conversion, i.e., converts the image from the standard color space CIE XYZ to the original color space of the photographing device; the sixth conversion sub-unit 416 performs an inverse exposure compensation conversion; and the fourth conversion sub-unit 414 performs an inverse white balance conversion.
[0080] Accordingly, the first inverse transformation subunit 421 performs white balance transformation, the sixth inverse transformation subunit 426 performs exposure compensation transformation, the second inverse transformation subunit 422 performs second inverse space transformation, i.e., transforming from the original color space of the photographing device to the standard color space CIE XYZ, the third inverse transformation subunit 423 performs style transformation, the fifth inverse transformation subunit 425 performs third inverse space transformation, i.e., transforming the image from the standard color space CIE XYZ to the linear RGB color space, and the fourth inverse transformation subunit 424 performs first inverse space transformation, i.e., transforming the input image from the linear RGB color space to the sRGB color space.
[0081] As shown in FIGS. 10 to 12, the first global feature extraction module 431 extracts the input image I in sRGB The parameter γ can be obtained based on the above equation and provided to the first transformation subunit 411 and the fourth inverse transformation subunit 424.
[0082] The first transformation subunit 411 converts the input image I in sRGB , and based on the parameter γ provided by the first global feature extraction module 431, I s RGB =I in 1 / γ The fifth conversion subunit 415 can obtain I s RGB and based on the transformation matrix P, I sXYZ =P -1 I in 1 / γ can be obtained and provided to the first adjusting unit 440 and the second converting sub-unit 412.
[0083] The first adjustment unit 440 adjusts the image I s XYZ is multiplied by a proportional coefficient α1 and then the original input image I in sRGB and provide the result to the local feature extraction module 432. The local feature extraction module 432 obtains the parameters M and A and provides them to the second transformation subunit 412 and the third inverse transformation subunit 423. The second transformation subunit 412 obtains the parameters M and A and provides them to the second transformation subunit 412 and the third inverse transformation subunit 423.
[0084]
number
[0085] The second global feature extraction module 433 calculates the parameter T ct -1 to the third conversion subunit 413, and the parameter k ct -1 to the sixth conversion subunit 416, and the parameter W ct -1 The third transformation subunit 413 can provide the image after the second spatial transformation to the fourth transformation subunit 414.
[0086]
number
[0087]
number
[0088]
number
[0089] The fourth transformation sub-unit 414 provides the calculated intermediate image to the second transformation unit 420. Also, the second global feature extraction module 433 provides the calculated T ideal , k ideal and W ideal to the second transformation unit 420. Then, the first inverse transformation sub-unit 421 performs a white balance transformation to convert the image W ideal I raw and the sixth inverse transformation subunit 426 obtains the image k after the exposure compensation transformation. ideal W ideal I raw and a second inverse transform subunit 422 performs a second inverse spatial transform to obtain the image T ideal k ideal W ideal I raw and a third inverse transform subunit 423 performs style transformation to obtain the image
[0090]
number
[0091]
number
[0092]
number
[0093] Thus, the output image may be expressed as:
[0094]
number
[0095] Typically, color constancy can be affected by both white balance and exposure compensation. White balance adjusts the color balance of an image to reflect how the human eye perceives it under natural light, compensates for the color temperature of the light source to ensure that white objects appear white, and adjusts the red, green, and blue channels to remove color casts caused by lighting conditions. Exposure compensation adjusts the overall brightness and contrast of the image to correct for underexposure (too dark) or overexposure (too bright) that may occur during shooting. White balance and exposure compensation are closely related, and changes in one can affect the other. In this embodiment, simultaneous exposure compensation and white balance adjustments can improve color correction performance.
[0096] <Other Examples> The sixth transformation sub-unit 416 and the sixth inverse transformation sub-unit 426 of the image processing device 400 can be added to the image processing device 200, and in this way, simultaneous processing of exposure compensation and white balance can be similarly realized.
[0097] In the above-described embodiments of the image processing device, The first conversion unit converts the input image from the sRGB color space to the original color space of the image capture device through space conversion before inverse white balance conversion; the first global feature extraction module is configured to obtain parameters for a first spatial transformation and a first inverse spatial transformation based on the input image; The local feature extraction module is configured to obtain parameters for the de-style transformation and the style transformation based on the input image and a first transformed image obtained by a transformation prior to the de-style transformation; and The second global feature extraction module is configured to obtain, based on the input image and the second transformed image that has undergone the destyled transformation, parameters for transformations corresponding to subsequent transformations after the destyled transformation in the first set of transformations and subsequent transformations in the second set of transformations.
[0098] Tests on the image processing device according to the present invention will now be described.
[0099] The basic model is a lightweight exposure compensation neural network presented at the 2022 BMVC (British Machine Vision Conference), which adopts parallel local feature extraction modules and global feature extraction modules to calculate parameters and does not have the reversible transformation process in this invention.
[0100] The test model is an image processing device according to the third embodiment.
[0101] Table 1 shows the test results using the white balance correction dataset.
[0102] [Table 1] MSE (Mean Square Error) is the average value of the Euclidean distance between the value of each channel of the output image and the value of each channel of the true image in the RGB color space. MAE (Mean Angle Error) is the average value of the included angle between each pixel vector of the output image and each pixel vector of the true image when the three RGB channels are the three axes of the coordinate system in the RGB color space.
[0103] ΔE 2000 (Delta E 2000) is a color difference index used to quantify the visual difference between two colors. It is an improved version of DeltaE, and has better consistency and adaptability to the characteristics of the human visual system, especially when dealing with subtle differences between colors.
[0104] Delta E itself is a measure of the difference between two colors, and the original Delta E (or ΔE * Delta E 2000 (ab) is the Euclidean distance between two points based on the CIELAB color space. However, this original version does not fully match human visual perception of color differences, especially sensitivity to certain color ranges. To solve this problem, Delta E 2000 was proposed, which takes into account the influence of color attributes such as hue, saturation, and brightness on visual perception and adjusts the calculation formula to more accurately reflect these differences. Currently, Delta E 2000 is one of the most accurate color difference evaluation methods and is widely used in color-related industries that require high color accuracy, such as printing, paints, plastics, lighting, and display calibration.
[0105] The calculation formula for Delta E 2000 is relatively complex and includes multiple correction terms to account for interactions between different color parameters and non-linear changes in color space. In practice, the calculation of these terms is usually performed by dedicated software or customized algorithms.
[0106] Numerically, the Delta E 2000 value is usually interpreted as follows: ΔE<1.0: The difference is not noticeable to the human eye; 1.0<ΔE<2.0: A difference that can be detected only by a trained observer under certain conditions; 2.0<ΔE<10.0: a difference that can be detected by a non-expert observer; and ΔE>10.0: There is a clear color difference.
[0107] As can be seen from Table 1, compared with the basic model, all three test values are reduced, so the present invention can achieve a better white balance correction effect.
[0108] Table 2 shows the test results using an exposure correction dataset (e.g., from MIT-Adobe FiveK).
[0109] [Table 2] SSIM (Structural Similarity Index Measure) and PSNR (Peak Signal-to-Noise Ratio) are both metrics for quantifying image quality, and are used in the fields of image processing and computer vision to evaluate the impact of operations such as image compression and image restoration on image quality. They use different methods to evaluate the similarity or difference between images.
[0110] SSIM (Structural Similarity Index) is a method to compare the difference in visual effect between two images. Unlike indices based on error accumulation (e.g., MSE), SSIM takes into account the structural information of the image. The SSIM value ranges from -1 to 1, with 1 representing that the two images are completely identical. SSIM takes into account three aspects: brightness, contrast, and structure. That is, Brightness: compares the average brightness of two images; Contrast: Compare the standard deviation of the contrast of the two images; and Structural: Compares the structural information of two images, which is essentially the correlation between their pixel values.
[0111] SSIM is a perceptual model that aims to approximate the characteristics of human visual perception. In practical applications, SSIM is widely used to evaluate the effectiveness of image processing tasks such as image compression, image transmission, watermarking, and noise reduction.
[0112] PSNR (Peak Signal to Noise Ratio) is an index used to evaluate image restoration quality, especially in the field of image compression. PSNR represents the pixel-level difference between the original image and the image after some processing. It is a logarithmic scale expression based on the mean square error (MSE). That is, Calculating MSE: measuring the mean squared difference between the pixels corresponding to the two images; and Calculating PSNR: Place the MSE results on a logarithmic scale, usually based on the maximum possible pixel value of the image (e.g., 255 for an 8-bit image).
[0113] The unit of PSNR is decibels (dB), and the higher the value, the better the image quality. However, because PSNR is an error-based evaluation, it cannot always well reflect the visual difference in image quality.
[0114] Both SSIM and PSNR are important tools for assessing image quality, but they focus on different aspects: SSIM focuses on simulating the sensitivity of the human eye to structural changes, while PSNR focuses on global pixel-level errors.
[0115] As shown in Table 2, both of these test indices are increased, so the present invention can achieve a better exposure compensation effect.
[0116] In summary, compared with the basic model in the prior art, the method according to the present invention can automatically correct color by simultaneously changing white balance and exposure, and obtain better performance.
[0117] An image processing method according to an embodiment of the present invention will be described below in conjunction with FIG.
[0118] 13, the image processing method according to the embodiment of the present invention starts in step S110. In step S110, an image output by a photographing device is received as an input image, and a first set of transformations is performed on the input image to obtain an intermediate image, which is an image in the original color space of the photographing device.
[0119] Next, in step S120, a second set of transformations is performed on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image, where the input image is adjusted to a predetermined lighting condition. Also, each transformation in the second set of transformations is the inverse process of the corresponding transformation in the first set of transformations. Then, the process ends.
[0120] In the above process, the feature extraction unit calculates the parameters required for the first set of transformations and the second set of transformations.
[0121] As a result, in the image processing method according to an embodiment of the present invention, the image color correction step is divided into two steps: from the input image to the electronic negative film (intermediate image) and from the electronic negative film to the corrected image, and the calculation method from the electronic negative film to the input image is made to match the calculation method from the electronic negative film to the corrected image, so that the input image can be adjusted to a specified lighting condition or to a color under a specified lighting condition.
[0122] Various specific implementation manners of the above steps of the image processing method according to the embodiment of the present invention have been described in detail, so detailed description thereof will be omitted here.
[0123] Of course, the process of each operation in the image processing method according to the present invention may be realized by a computer-executable program stored in various computer-readable storage media.
[0124] The object of the present invention may also be realized in the following manner: a storage medium storing the above-mentioned executable program code is provided directly or indirectly to a system or device, and the program code is read and executed by a computer or central processing unit (CPU) in the system or device. In this case, as long as the system or device has the function of executing a program, the embodiment of the present invention is not limited to a program, and the program may be in any form, such as an object-oriented program, an interpreted program, a script program provided to an operating system, etc.
[0125] The computer-readable storage media mentioned above may be various storage devices and units, semiconductor devices, magnetic units such as optical, magnetic and magneto-optical disks, and any other medium suitable for storing information.
[0126] In addition, the technical solution of the present invention can also be realized by connecting a computer to a corresponding website on the Internet, downloading and installing the computer program code of the present invention into the computer, and then running the program.
[0127] FIG. 14 is a block diagram showing an exemplary configuration of a general-purpose personal computer that can implement an image processing apparatus and method according to an embodiment of the present invention.
[0128] 14, the computer 1400 may be, for example, a computer system. Note that the computer 1400 is merely an example and does not limit the scope or functionality of the method and apparatus according to the present invention. Furthermore, the computer 1400 does not depend on any module, assembly, or combination thereof in the above-described method and apparatus.
[0129] 14, a central processing unit (CPU) 1401 performs various processes based on programs stored in a ROM 1402 or programs loaded from a storage unit 1408 into a RAM 1403. The RAM 1403 can also store data required when the CPU 1401 performs various processes, depending on the needs. The CPU 1401, ROM 1402, and RAM 1403 are connected to one another via a bus 1404. An input / output interface 1405 is also connected to the bus 1404.
[0130] The input / output interface 1405 is further connected to the following components: an input device 1406 including a keyboard; an output device 1407 including a display such as a liquid crystal display (LCD) and a speaker; a storage device 1408 including a hard disk; and a communication device 1409 including a network interface card such as a LAN card or a modem. The communication device 1409 performs communication processing via a network such as the Internet or a LAN. A drive 1410 may be connected to the input / output interface 1405 as needed. A removable medium 1411, such as a semiconductor memory, can be inserted into the drive 1410 as needed, allowing a computer program read from the medium to be installed in the storage device 1408.
[0131] The present invention also provides a program product including machine-readable instruction codes, which, when read and executed by a machine, can perform the methods of the above-described embodiments of the present invention. Accordingly, various storage media for carrying such program products, such as magnetic disks (including floppy disks (registered trademark)), optical disks (including CD-ROMs and DVDs), magneto-optical disks (including MDs (registered trademark)), and semiconductor storage devices, are also included in the present invention.
[0132] The storage medium may include, for example, a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory device, etc., but is not limited to these.
[0133] Furthermore, each operation (process) in the above-described method can also be realized in the form of a computer-executable program stored in various machine-readable storage media.
[0134] Furthermore, the following supplementary notes are disclosed regarding the above-mentioned embodiments.
[0135] (Appendix 1) An image processing device, a first transformation unit that receives an image output by a photographing device as an input image and performs a first set of transformations on the input image to obtain intermediate images, the intermediate images being images in an original color space of the photographing device; a second transformation unit that performs a second set of transformations on the intermediate image according to an inverse order of transformations in the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; and a feature extraction unit for calculating parameters necessary for the first set of transforms and the second set of transforms.
[0136] (Appendix 2) 10. The image processing device according to claim 1, The first set of transformations and the second set of transformations each include one of the following cases: the first set of transforms includes, in order, a first spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, and the first inverse spatial transform; the first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform; and The first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, an inverse exposure compensation transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, an exposure compensation transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform.
[0137] (Appendix 3) 3. The image processing device according to claim 2, The first conversion unit converts the input image from the sRGB color space to the original color space of the image capture device through a space conversion before the inverse white balance conversion.
[0138] (Appendix 4) 3. The image processing device according to claim 2, The feature extraction unit includes a first global feature extraction module, a local feature extraction module, and a second global feature extraction module connected in series, and the feature extraction unit and the first transformation unit provide parameters related to the first set of transformation and the second set of transformation.
[0139] (Appendix 5) 5. The image processing device according to claim 4, The first global feature extraction module is configured to obtain parameters for the first spatial transformation and the first inverse spatial transformation based on the input image.
[0140] (Appendix 6) 6. The image processing device according to claim 5, The local feature extraction module is configured to obtain parameters for the de-styling transformation and the style transfer based on the input image and a first transformed image obtained by a transformation prior to the de-styling transformation.
[0141] (Appendix 7) 7. The image processing device according to claim 6, The second global feature extraction module is configured to obtain, based on the input image and a second transformed image that has undergone the destyled transformation, parameters for a subsequent transformation after the destyled transformation in the first set of transformations and a transformation corresponding to the subsequent transformation in the second set of transformations.
[0142] (Appendix 8) 5. The image processing device according to claim 4, Corresponding parameters of the first global feature extraction module, the local feature extraction module and the second global feature extraction module are determined by training the image processing device, and during the training process for the image processing device, a loss function includes a term representing the difference between the intermediate images calculated from different images of the same scene by the same image capturing device.
[0143] (Appendix 9) 3. The image processing device according to claim 2, the first spatial transformation transforms the input image from an sRGB color space to a linear RGB color space, and the first inverse spatial transformation transforms the image from the linear RGB color space to the sRGB color space; The second space transformation transforms the image from the standard color space CIE XYZ to the original color space of the image capture device, and the second inverse space transformation transforms the image from the original color space of the image capture device to the standard color space CIE XYZ; and The third space transformation transforms the image from a linear RGB color space to a standard color space CIE XYZ, and the third inverse space transformation transforms the image from the standard color space CIE XYZ to a linear RGB color space.
[0144] (Appendix 10) 5. The image processing device according to claim 4, The local feature extraction module is realized by two local feature extraction units connected in parallel, and the two local feature extraction units connected in parallel are used to extract different parameters used for both the destyling and the style conversion.
[0145] (Appendix 11) 7. The image processing device according to claim 6, The image processing device further includes a first adjustment unit, which is used to adjust a first proportionality coefficient of the first transformed image input to the local feature extraction module.
[0146] (Appendix 12) 8. The image processing device according to claim 7, The image processing device further includes a second adjustment unit, which is used to adjust a second proportionality coefficient of the second transformed image input to the second global feature extraction module.
[0147] (Appendix 13) 13. The image processing device according to claim 11, The first proportional coefficient and the second proportional coefficient are learnable vectors.
[0148] (Appendix 14) 1. An image processing method, comprising: receiving an image output by a photographing device as an input image, and performing a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in the original color space of the photographing device; and performing a second set of transformations on the intermediate image in an order opposite to that of the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; Calculating parameters necessary for the first set of transformations and the second set of transformations by a feature extraction unit.
[0149] (Appendix 15) 15. The method of claim 14, The first set of transformations and the second set of transformations each include one of the following cases: the first set of transforms includes, in order, a first spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, and the first inverse spatial transform; the first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform; and The first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, an inverse exposure compensation transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, an exposure compensation transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform.
[0150] (Appendix 16) 16. The method of claim 15, The feature extraction unit includes a first global feature extraction module, a local feature extraction module and a second global feature extraction module connected in series, and the calculation of the parameters includes: providing parameters for the first set of transforms and the second set of transforms by the feature extraction unit and a first transform unit that performs the first set of transforms.
[0151] (Appendix 17) 17. The method of claim 16, The calculation of the parameters is obtaining, by the first global feature extraction module, parameters for the first spatial transformation and the first inverse spatial transformation based on the input image.
[0152] (Appendix 18) 18. The method of claim 17, The calculation of the parameters is The method includes obtaining parameters for the de-styling and the style transfer based on the input image and a first transformed image obtained by a transformation prior to the de-styling, by the local feature extraction module.
[0153] (Appendix 19) 19. The method of claim 18, The calculation of the parameters is The second global feature extraction module obtains, based on the input image and a second transformed image that has undergone the destyled transformation, parameters for a subsequent transformation after the destyled transformation in the first set of transformations and a transformation corresponding to the subsequent transformation in the second set of transformations.
[0154] (Appendix 20) 1. A computer-readable storage medium, comprising: A program product is carried that includes computer-readable instruction code, which, when read and executed by a computer, causes the computer to perform the image processing method described in Appendix 14-19.
[0155] Although the preferred embodiment of the present invention has been described above, the present invention is not limited to this embodiment, and any modification to the present invention falls within the technical scope of the present invention as long as it does not depart from the spirit of the present invention.
Claims
1. An apparatus for processing images, comprising: a first transformation unit that receives an image output by a photographing device as an input image and performs a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in an original color space of the photographing device; a second transformation unit that performs a second set of transformations on the intermediate image according to an order of transformations that is reverse to the order of transformations in the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; and an apparatus comprising a feature extraction unit for calculating parameters necessary for said first set of transforms and said second set of transforms.
2. 10. The apparatus of claim 1, The first set transformation and the second set transformation include one of the following cases: the first set of transforms includes, in order, a first spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, and the first inverse spatial transform; the first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform; and the first set of transforms includes, in order, a first spatial transform, a third spatial transform, a destyle transform, a second spatial transform, an inverse exposure compensation transform, and an inverse white balance transform, and the second set of transforms includes, in order, a white balance transform, an exposure compensation transform, a second inverse spatial transform, a style transform, a third inverse spatial transform, and the first inverse spatial transform.
3. 3. The apparatus of claim 2, The first conversion unit converts the input image from the sRGB color space to the original color space of the image capture device through a space conversion before the inverse white balance conversion.
4. 3. The apparatus of claim 2, The feature extraction unit includes a first global feature extraction module, a local feature extraction module, and a second global feature extraction module connected in series, and the feature extraction unit and the first transformation unit provide parameters for the first set of transformations and the second set of transformations.
5. 5. The apparatus of claim 4, The apparatus, wherein the first global feature extraction module is configured to obtain parameters for the first spatial transform and the first inverse spatial transform based on the input image.
6. 6. The apparatus of claim 5, The apparatus, wherein the local feature extraction module is configured to derive parameters for the de-styling transformation and the style transfer based on the input image and a first transformed image obtained by a transformation prior to the de-styling transformation.
7. 7. The apparatus of claim 6, The apparatus, wherein the second global feature extraction module is configured to obtain, based on the input image and a second transformed image that has undergone the destyled transformation, parameters for a subsequent transformation after the destyled transformation in the first set of transformations and a transformation corresponding to the subsequent transformation in the second set of transformations.
8. 5. The apparatus of claim 4, An apparatus in which corresponding parameters of the first global feature extraction module, the local feature extraction module and the second global feature extraction module are determined by training the apparatus, and during the training process for the apparatus, a loss function includes a term representing the difference between the intermediate images calculated from different images of the same scene by the same imaging device.
9. 1. A method for processing an image, comprising: receiving an image output by a photographing device as an input image, and performing a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in the original color space of the photographing device; and performing a second set of transformations on the intermediate image in an order reverse to that of the first set of transformations to obtain an output image, in which the input image is adjusted to a predetermined lighting condition, and each transformation in the second set of transformations is an inverse process of a corresponding transformation in the first set of transformations; A method for calculating parameters required for the first set of transformations and the second set of transformations by a feature extraction unit.
10. A program for causing a computer to execute the method according to claim 9.