Image processing apparatus, image processing method, and machine-readable storage medium
By using an image processing device and method, the image is transformed from the color space of the camera device to the original color space, and then adjusted to the predetermined lighting conditions in reverse process, thus solving the image color matching problem and improving the image stitching and object recognition effects.
Patent Information
- Application Number
- CN202410543833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies struggle to effectively adjust image colors to match predetermined lighting conditions, impacting the color consistency and comparability of image stitching and object recognition.
An image processing device and method are used to transform the image from the color space of the camera device to the original color space through a first set of transformations. Then, the image is adjusted to a predetermined lighting condition through a second set of transformations in the reverse process. The transformation parameters are calculated using a feature extraction unit. The process is divided into two steps: from the input image to the electronic film and from the electronic film to the corrected image.
It achieves accurate adjustment of image colors to predetermined lighting conditions, improving the performance of image stitching and object recognition.
Smart Images

Figure CN120881402A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing, and more particularly to image processing apparatus, image processing method, and machine-readable storage medium for adjusting image colors. Background Technology
[0002] This section provides background information relating to this disclosure, which is not necessarily prior art.
[0003] Achieving color constancy is a key challenge in computer vision and image processing. It typically involves algorithms that adjust the colors of an image to match those perceptible under predetermined lighting conditions (e.g., standard lighting). This is crucial for tasks requiring color consistency and comparability, such as image stitching, object recognition, and any application where machines need to make decisions based on color. Summary of the Invention
[0004] This section provides a general overview of this disclosure, rather than a full disclosure of its entire scope or all its features.
[0005] The purpose of this disclosure is to provide an image processing apparatus, an image processing method, and a machine-readable storage medium for adjusting the color of an image to a predetermined lighting condition.
[0006] According to one aspect of this disclosure, an image processing apparatus is provided, comprising: a first transformation unit that receives an image output by a camera as an input image and performs a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in the original color space of the camera; a second transformation unit that performs a second set of transformations on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image, wherein the input image is adjusted to predetermined lighting conditions in the output image, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations; and a feature extraction unit that calculates parameters required in the first set of transformations and the second set of transformations.
[0007] According to another aspect of this disclosure, an image processing method is provided, comprising: receiving an image output by a camera device as an input image, and performing a first set of transformations on the input image to obtain an intermediate image, the intermediate image being an image in the original color space of the camera device; and performing a second set of transformations on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image, wherein the input image is adjusted to predetermined lighting conditions in the output image, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations, wherein parameters required in the first set of transformations and the second set of transformations are calculated by a feature extraction unit.
[0008] According to another aspect of this disclosure, a machine-readable storage medium is provided that carries a program product including machine-readable instruction code stored thereon, wherein the instruction code, when read and executed by a computer, enables the computer to perform an image processing method according to this disclosure.
[0009] In accordance with other aspects of this disclosure, computer program code and computer program products for implementing the image processing methods according to this disclosure are also provided.
[0010] According to the image processing apparatus, image processing method, and machine-readable storage medium disclosed herein, the image color correction step is divided into two steps: from the input image to the electronic film (intermediate image) and from the electronic film to the corrected image. The calculation method from the electronic film to the input image is consistent with the calculation method from the electronic film to the corrected image, thereby adjusting the input image to the color under predetermined lighting conditions or under predetermined lighting conditions.
[0011] Further areas of applicability will become apparent from the description provided herein. The descriptions and specific examples in this summary are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description
[0012] The accompanying drawings described herein are for illustrative purposes only and not for all possible implementations, and are not intended to limit the scope of this disclosure. In the drawings:
[0013] Figure 1 This is a block diagram illustrating the structure of an image processing apparatus according to an embodiment of the present disclosure;
[0014] Figure 2 This is a block diagram illustrating the structure of an image processing apparatus according to a first embodiment of the present disclosure;
[0015] Figure 3 This is a simplified structural diagram illustrating an image processing apparatus according to a first embodiment of the present disclosure;
[0016] Figure 4 This is a flowchart illustrating the image correction process of an image processing apparatus according to a first embodiment of the present disclosure;
[0017] Figure 5 This is a block diagram illustrating the structure of a global feature extraction unit in an image processing apparatus according to a first embodiment of the present disclosure;
[0018] Figure 6 This is a block diagram illustrating the structure of a local feature extraction unit in an image processing apparatus according to a first embodiment of the present disclosure;
[0019] Figure 7 This is a block diagram illustrating the structure of an image processing apparatus according to a second embodiment of the present disclosure;
[0020] Figure 8 This is a simplified structural diagram illustrating an image processing apparatus according to a second embodiment of the present disclosure;
[0021] Figure 9 This is a flowchart illustrating the image correction process of an image processing apparatus according to a second embodiment of the present disclosure;
[0022] Figure 10 This is a block diagram illustrating the structure of an image processing apparatus according to a third embodiment of the present disclosure;
[0023] Figure 11 This is a simplified structural diagram illustrating an image processing apparatus according to a third embodiment of the present disclosure;
[0024] Figure 12 This is a flowchart illustrating the image correction process of an image processing apparatus according to a third embodiment of the present disclosure;
[0025] Figure 13 A flowchart illustrating an image processing method according to an embodiment of the present disclosure; and
[0026] Figure 14 This is a block diagram of an exemplary structure of a general-purpose personal computer in which image processing apparatus and methods according to embodiments of the present disclosure can be implemented.
[0027] While this disclosure is readily subject to various modifications and substitutions, specific embodiments thereof have been shown by way of example in the accompanying drawings and are described in detail herein. However, it should be understood that the description of specific embodiments herein is not intended to limit this disclosure to the specific forms disclosed, but rather, this disclosure is intended to cover all modifications, equivalents, and substitutions falling within the spirit and scope of this disclosure. It should be noted that throughout the drawings, corresponding reference numerals indicate corresponding parts. Detailed Implementation
[0028] Examples of this disclosure will now be described more fully with reference to the accompanying drawings. The following description is merely exemplary and is not intended to limit the disclosure, its application, or its uses.
[0029] Example embodiments are provided so that this disclosure will become exhaustive and will fully convey its scope to those skilled in the art. Numerous specific details, such as examples of particular components, apparatus, and methods, are set forth to provide a detailed understanding of embodiments of this disclosure. It will be apparent to those skilled in the art that the specific details are not required, and that the example embodiments may be implemented in many different forms, none of which should be construed as limiting the scope of this disclosure. In some example embodiments, well-known processes, well-known structures, and well-known techniques are not described in detail.
[0030] Figure 1 An image processing apparatus according to an embodiment of the present disclosure is illustrated. For example... Figure 1 As shown, the image processing apparatus 100 according to an embodiment of the present disclosure may include a first transformation unit 110, a second transformation unit 120, and a feature extraction unit 130.
[0031] The first transformation unit 110 can receive an image output by the camera device as an input image and perform a first set of transformations on the input image to obtain an intermediate image, which is an image in the original color space of the camera device. For example, the intermediate image is the electronic film of the camera device obtained through the first set of transformations. It should be noted that the electronic film here is not exactly the same as the actual electronic film of the camera device, but is close to or equivalent to the actual electronic film and is an image in the same color space (original color space) as the actual electronic film. Furthermore, the image output by the camera device is an image in the output color space of the camera device.
[0032] The second transformation unit 120 can perform a second set of transformations on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image. In the output image, the input image is adjusted to predetermined lighting conditions, and the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations. In this application, the output image can also be referred to as a corrected image.
[0033] As described above, the order of transformations in the second set of transformations is the reverse of the order of transformations in the first set of transformations, and the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations. For example, the first set of transformations may include the first, second, and third transformations, while the second set of transformations may include the fourth, fifth, and sixth transformations. In this case, the fourth transformation is the inverse process of the third transformation, the fifth transformation is the inverse process of the second transformation, and the sixth transformation is the inverse process of the first transformation.
[0034] Furthermore, as an example, the first set of transformations and the second set of transformations in this application can each be a set of arithmetic operations, and the arithmetic operation process of each transformation in the second set of transformations is the inverse process of the corresponding transformation in the first set of transformations. In this case, the first transformation unit 110 and the second transformation unit 120 can be configured as one or more computing units corresponding to different operations in the corresponding set of arithmetic operations. These computing units respectively execute the corresponding operation in the set of arithmetic operations corresponding to a certain transformation. However, it should be noted that since the electronic negative of the input image is obtained through the first set of transformations while transforming the input image to the original color space, and the input image is adjusted to the predetermined lighting conditions while restoring the intermediate image to the color space of the input image through the second set of transformations, although the corresponding transformation processes in the two are reversible, the parameters used in the corresponding transformations are not necessarily exactly the same. That is, some transformation parameters in the second set of transformations are related to the requirements of the predetermined lighting conditions. In addition, the predetermined lighting conditions can be standard lighting conditions.
[0035] The feature extraction unit 130 can calculate the parameters required in the first set of transformations and the second set of transformations. For example, the feature extraction unit 130 can be implemented using a neural network. The model parameters of the entire image processing apparatus 100 can be fixed by training the feature extraction unit 130, the first transformation unit 110, and the second transformation unit 120. Thus, when using the image processing apparatus 100, the feature extraction unit 130 and the first transformation unit 110 obtain the parameters required for transformations suitable for converting the input image into an electronic film and adjusting the electronic film to predetermined lighting conditions. Furthermore, the first transformation unit 110 and the second transformation unit 120 can each perform a set of arithmetic operations using the obtained transformation parameters to perform the corresponding set of transformations.
[0036] Typically, when a camera captures a photograph / image, it first obtains a digital negative, then performs RGB desacrifice, noise reduction, white balance and color space transformation, style transformation, image scaling, mapping to an output color space (e.g., sRGB (standard Red Green Blue) color space, P3 color space), JPEG (Joint Photographic Experts Group) / HEIC (High Efficiency Image Format) compression, and other processing before outputting the image. However, such image signal processing algorithms are all performed within the camera device and cannot be applied to images obtained from the camera device (e.g., the final video) in a color space such as sRGB.
[0037] In this application, the image output by the imaging device is first converted into an electronic film by the first conversion unit 110 to obtain an image in the original color space of the imaging device. That is, the first conversion unit 110 converts the image from the color space of the image output by the imaging device to the original color space of the imaging device. Furthermore, the first conversion unit 110 can also perform other transformations such as inverse white balance transformation. The second conversion unit 120 can convert the intermediate image (electronic film) back to the color space of the input image to be processed. Similarly, the second conversion unit 120 can also perform other transformations such as white balance transformation. The transformations performed by the second conversion unit 120 can adjust the lighting conditions to the desired predetermined lighting conditions.
[0038] Therefore, the image processing apparatus 100 according to the embodiments of the present disclosure can divide the image color correction step into two steps: from the input image to the electronic film (intermediate image) and from the electronic film to the corrected image. The calculation method from the electronic film to the input image is consistent with the calculation method from the electronic film to the corrected image, thereby adjusting the input image to the color under predetermined lighting conditions or under predetermined lighting conditions.
[0039] Application scenarios:
[0040] This application can be applied to tasks requiring color consistency and comparability, such as image stitching, object recognition, and any application where machines need to make decisions based on color. For example, this application can better identify or classify objects based on their color characteristics. Objects can be images obtained through operations such as taking a picture. Specific applications include supermarket goods recognition or classification, and tracking specific objects (such as thieves, people / objects requiring protection, etc.) in images or videos.
[0041] Therefore, this application can improve the performance of color-related downstream tasks.
[0042] The following is combined Figures 2 to 6 The image processing apparatus 200 according to a first embodiment of the present disclosure will be described. Figure 2 A structural block diagram of an image processing apparatus 200 according to a first embodiment of the present disclosure is shown. Figure 2As shown, the image processing device 200 may include a first transformation unit 210, a second transformation unit 220, a feature extraction unit 230, a first adjustment unit 240, and a second adjustment unit 250. The first transformation unit 210 may include a first transformation subunit 211, a second transformation subunit 212, a third transformation subunit 213, and a fourth transformation subunit 214. The second transformation unit 220 may include a first inverse transformation subunit 221, a second inverse transformation subunit 222, a third inverse transformation subunit 223, and a fourth inverse transformation subunit 224. The feature extraction unit 230 may include a first global feature extraction module 231, a local feature extraction module 232, and a second global feature extraction module 233. It should be noted that the inverse transformation in this naming refers to the transformation process relative to the first set of transformations and does not limit the type of transformation performed by the second transformation unit 220. Furthermore, the first transformation unit 210, the second transformation unit 220, and the feature extraction unit 230 correspond to... Figure 1 The first transformation unit 110, the second transformation unit 120, and the feature extraction unit 130, therefore regarding Figure 1 The same description applies to the image processing device 200.
[0043] In this embodiment, the first set of transformations may sequentially include a first spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations may sequentially include a white balance transformation, an inverse second spatial transformation, a style transformation, and an inverse first spatial transformation. Accordingly, the first transformation subunit 211 may perform the first set of spatial transformations, the second transformation subunit 212 may perform the destyling transformation, the third transformation subunit 213 may perform the second spatial transformation, and the fourth transformation subunit 214 may perform the inverse white balance transformation. Correspondingly, the first inverse transformation subunit 221 may perform the white balance transformation, the second inverse transformation subunit 222 may perform the second inverse spatial transformation, the third inverse transformation subunit 223 may perform the style transformation, and the fourth inverse transformation subunit 224 may perform the first inverse spatial transformation.
[0044] Taking an input image in the sRGB color space as an example, the first spatial transformation transforms the input image from the sRGB color space to the linear RGB color space, and the second spatial transformation transforms the image from the standard CIE XYZ color space to the original color space of the camera device. The inverse first spatial transformation transforms the image from the linear RGB color space to the sRGB color space, and the inverse second spatial transformation transforms the image from the original color space of the camera device back to the standard CIE XYZ color space. However, this disclosure is not limited to this. When the input image is in the P3 color space, the first spatial transformation can transform the P3 color space to the linear RGB color space, and the inverse first spatial transformation can transform the linear RGB color space to the P3 color space. Other than this, the other transformations are the same.
[0045] Figure 3 The illustration is a simplified structural diagram of an image processing apparatus according to a first embodiment of the present disclosure, and Figure 4 This is a flowchart illustrating the image correction process of an image processing apparatus according to a first embodiment of the present disclosure. Figure 3 and Figure 4 As shown, the local feature extraction module 232 may include two parallel local feature extraction modules (a first local feature extraction module and a second local feature extraction module) to extract the different parameters used in both the destyling transformation and the style transformation. Of course, the local feature extraction module 232 of this application can also be implemented using a single local feature extraction module. In this case, the local feature extraction module 232 can output multiple parameters used in both the destyling transformation and the style transformation.
[0046] like Figures 2 to 4 As shown, the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 are connected in series and can be connected in series with the sub-units of the first transformation unit 210, thereby providing parameters related to the first set of transformations and the second set of transformations through the feature extraction unit 230 and the first transformation unit 210. The following is a combination of... Figure 3 and Figure 4 To describe the principle of the image processing device 200 in more detail. It should be noted that... Figure 3 and Figure 4 In order to simplify the structure and facilitate description, only the first global feature extraction module, the first local feature extraction module, the second local feature extraction module, and the second global feature extraction module are shown, and each transformation subunit / inverse transformation subunit is replaced by an operator of mathematical operations performed by different transformation subunits / inverse transformation subunits.
[0047] like Figures 2 to 4 As shown, first, input the image (the original picture taken by the camera device). The input is given to the first global feature extraction module 231, and the first global feature extraction module 231 calculates the parameter γ (and therefore...). Input image The input image is also fed into the first transformation subunit 211, and the first transformation subunit 211 performs a power transformation on each pixel of the input image based on the parameter γ provided by the first global feature extraction module 231, i.e., performs... The exponentiation operation yields Among them, I in yes This step aims to remove the nonlinear power transformation that simulates human visual perception and convert the nonlinear sRGB color space to a linear RGB color space. The parameter γ is used for the gamma transformation, which, also known as perceptual encoding, remaps linear colors to better suit the nonlinear response of our visual system to radiant power.
[0048] First adjustment unit 240 (in) Figure 3 The two operators (corresponding to the two operators in the dashed box) can transform the image obtained after the inverse gamma transform operation is performed by the first transform subunit 211. With the original image The sums are then fed into two parallel local feature extraction modules to calculate the nonlinear transformation fitting parameter matrices M and A. This process is used to fit steps such as style transformation and color enhancement in image processing. A (and therefore -A) and M (and therefore...) are calculated in the first and second local feature extraction modules, respectively. After that, the second transformation subunit 212 (in) Figure 3 (The two operators in the dotted line box are used to represent Is, which can be obtained using the first transformation subunit 211) RGB Perform element-wise multiplication Adding -A to the element yields the destylated result.
[0049] Preferably, the first adjustment unit 240 is configured to adjust the first transformed image input to the local feature extraction module. The first scaling factor. For example, the first adjustment unit 240 can adjust the image. Multiply by the scaling factor α1 and then combine with the original input image Add them together. This scaling factor is a channel-independent, learnable three-dimensional vector. For example, α1 is initialized to 10 during training. -4 This is used to accelerate convergence and maintain network stability. By taking the original image as input and multiplying the processed image by a learnable factor, the model can learn correctly from the start. Assigning a learnable additional scale to the processed image improves performance.
[0050] Second adjustment unit 250 (in) Figure 3 The two operators in the box corresponding to the points can calculate the result of the second transformation subunit 212. With the original image The sums are then fed into the second global feature extraction module 233, which calculates the parameter T. ideal , W ideal Among them, parameters and W ideal This represents the white balance transformation matrix, which can be, for example, a 3x3 diagonal matrix. Parameters It can represent the white balance matrix from the current color temperature (with respect to the input image to be processed) to the original white balance, and the parameter W ideal This can represent a white balance matrix from the original white balance to the ideal color temperature (under desired predetermined lighting conditions). Parameters This represents the transformation matrix from the standard color space CIE XYZ to the camera device's raw color space (raw-RGB) under the current lighting conditions, with parameter T. ideal This represents the transformation matrix from the camera's native color space (raw-RGB) to the standard color space CIE XYZ under desired, predetermined lighting conditions. For example, the parameters... and T ideal It is a 3x3 square matrix. These two transformation matrices will have different values depending on the color temperature, therefore they need to be estimated separately. Therefore, introducing different T matrices can make the transformation more accurate.
[0051] Preferably, the second adjustment unit 250 is configured to adjust the second transformed image input to the second global feature extraction module 233. The second scaling factor. The second adjustment unit 250 can adjust the image... Multiply by the scaling factor α2 and then combine with the original input image Add them together. This scaling factor is a channel-independent, learnable three-dimensional vector. For example, α2 can be initialized to 10 during training. -4 These are used to accelerate convergence and maintain network stability. The scaling coefficients α1 and α2 can be independent coefficients. By using the original image as input and multiplying the processed image by a learnable factor, the model can learn correctly from the start. Assigning a learnable scaling factor to the processed image improves performance.
[0052] The second global feature extraction module 233 can extract the calculated parameters. Provided to the third transformation subunit 213, and the parameters are... Provided to the fourth transformation subunit 214. The third transformation subunit 213 calculates the image after the second spatial transformation. Furthermore, the fourth transformation subunit 214 calculates the intermediate image after inverse white balance transformation.
[0053] The fourth transformation subunit 214 provides the calculated intermediate image to the second transformation unit 220. Furthermore, the second global feature extraction module 233 calculates the parameter T... ideal and W idealThe image is provided to the second transformation unit 220. Next, the first inverse transformation subunit 221 can perform a white balance transformation to obtain the image W. ideal I raw The second inverse transform subunit 222 can perform a second spatial inverse transform to obtain image T. ideal W ideal I raw The third inverse transform subunit 223 can perform style transformation to obtain the image (T). ideal W ideal I raw +A)⊙M, and the fourth inverse transform subunit 224 can perform the first spatial inverse transform, i.e., the gamma transform, to obtain the output image I. out =((T) ideal W ideal I raw +A)⊙M) γ Using the same parameters M and A in style transformation and destyling transformation can preserve more information about the camera's style and the image.
[0054] Therefore, the output image I out Essentially, it's the original input image. The result of a series of transformations is shown in equation (1) below.
[0055]
[0056] The purpose is to obtain the electronic negative I by sequentially removing gamma transform, nonlinear style transform, color space conversion, and current color temperature white balance. raw After white balance correction is performed in the raw-RGB color space of the camera device, color space conversion, nonlinear style transformation, and gamma transformation are then applied sequentially to the white balance-corrected image. The input image is corrected using the aforementioned reversible transformation formula. This reversible transformation formula allows the image transformation to more closely resemble the actual physical process, improving the performance of image adjustment.
[0057] It should be noted that, as described above, although the image output by the camera device is in the sRGB color space, the camera device can also output images in other color spaces, such as the P3 color space. This application can be applied to input images in other color spaces. Only the mapping matrix for different spaces needs to be changed.
[0058] In the above calculation process, the feature extraction unit 230 and the first transformation unit 210 provide parameters related to the first set of transformations and the second set of transformations. That is, only the feature extraction unit 230 and the first transformation unit 210 participate in the parameter calculation; the second transformation unit 220 only uses the calculated parameters. Combined with... Figure 4The entire image correction process can be seen more clearly. Figure 4 The forward process (the calculation process of the feature extraction unit and the first transformation unit) is represented by a solid line, and the reverse process (the calculation process of the second transformation unit) is represented by a dashed line. For example... Figure 4 As shown, parameters are calculated only in the forward process, and the parameters are used directly in the reverse process without calculating the parameters themselves.
[0059] This application determines the corresponding parameters and scaling factors α1 and α2 of the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 by training the image processing device 200. When using the trained image processing device 200, the parameters and scaling factors α1 and α2 of the first global feature extraction module 231, the local feature extraction module 232, and the second global feature extraction module 233 determined through training are used to calculate the parameters required for the first set of transformations and the second set of transformations. Furthermore, during the training of the image processing device, the input consists of different images under different scenes and lighting conditions, and the output is an image of the corresponding scene under predetermined lighting conditions. During the training process, only the above-mentioned input images and the output images under the desired lighting conditions are needed, and the ground truth of the electronic film is not required. After the model parameters are fixed through training, the second transformation unit directly uses the transformation parameters calculated by the feature extraction unit and the first transformation unit.
[0060] When training the image processing device 200, it is considered that the same scene captured by the same camera device under different color temperature and color style settings should have the same electronic negative I. raw Based on this premise, an L1 loss function for the electronic film is added during the training process. Therefore, the total loss function formula is L = Where λ is the proportionality coefficient. This represents the difference between the output image and the true value (e.g., expressed using the L1 norm), and This represents the difference between calculated intermediate images (e.g., represented by the L1 norm) of different images taken by the same camera device in the same scene. As an example, λ can be 0.1.
[0061] In the image processing apparatus 200 according to an embodiment of the present disclosure, a global feature extraction module, a local feature extraction module, and a first transformation unit connected in series are used to extract the parameters required for the first set of transformations and the second set of transformations. Compared with the case where the global feature extraction module and the local feature extraction module are connected in parallel and output transformation parameters simultaneously, the network structure serially transmits the calculation results of the previous steps, incorporates more useful information into the transformation parameters, and thus obtains better color adjustment results.
[0062] Furthermore, the image color correction process is divided into two steps: from the input image to the electronic negative (intermediate image) and from the electronic negative to the corrected image. The calculation method from the electronic negative to the input image is made consistent with the calculation method from the electronic negative to the corrected image, thereby effectively adjusting the input image to the predetermined lighting conditions.
[0063] The following is combined Figure 5 and Figure 6 This section explains the structure of the global feature extraction module (first global feature extraction module and second global feature extraction module) and the local feature extraction module (first local feature extraction module and second local feature extraction module). Figure 5 This is a block diagram illustrating the structure of a global feature extraction unit in an image processing apparatus according to a first embodiment of the present disclosure. Figure 6 This is a block diagram illustrating the structure of a local feature extraction unit in an image processing apparatus according to a first embodiment of the present disclosure.
[0064] like Figure 5 As shown, the global feature extraction module consists of a first convolutional module, a cross-force attention module, a first normalization module, a forward propagation network module, and a linear fully connected layer. The difference between the first global feature extraction module 231 and the second global feature extraction module 233 lies in the length of the query vector. The query vector is a learnable vector of a certain length and fixed dimensions. Its length varies with the number of dimensions to be queried. For example, when querying a three-dimensional diagonal matrix, the length is 3; when querying a three-dimensional matrix, the length is 9.
[0065] like Figure 6 As shown, the local feature extraction module consists of a first convolutional module, a first normalization module, a first fully connected layer, a second convolutional module, a second fully connected layer, a second normalization module, a forward propagation network module, and a linear fully connected layer. β1 and β2 are learnable scaling coefficients. The two local feature extraction modules have the same structure.
[0066] The sub-modules in the global feature extraction module and the local feature extraction module can be implemented using any available method in the prior art, and therefore will not be described in detail.
[0067] The following is combined Figures 7 to 9 The image processing apparatus 300 according to the second embodiment of the present disclosure will be described. Figure 7 This is a block diagram illustrating the structure of an image processing apparatus according to a second embodiment of the present disclosure. Figure 8 This is a simplified structural diagram illustrating an image processing apparatus according to a second embodiment of the present disclosure. Figure 9 This is a flowchart illustrating the image correction process of an image processing apparatus according to a second embodiment of the present disclosure.
[0068] like Figure 7As shown, the image processing device 300 may include a first transformation unit 310, a second transformation unit 320, a feature extraction unit 330, a first adjustment unit 340, and a second adjustment unit 350. The first transformation unit 310 may include a first transformation subunit 311, a second transformation subunit 312, a third transformation subunit 313, a fourth transformation subunit 314, and a fifth transformation subunit 315. The second transformation unit 320 may include a first inverse transformation subunit 321, a second inverse transformation subunit 322, a third inverse transformation subunit 323, a fourth inverse transformation subunit 324, and a fifth inverse transformation subunit 325. The feature extraction module 330 may include a first global feature extraction module 331, a local feature extraction module 332, and a second global feature extraction module 333.
[0069] The only difference between the image processing apparatus 300 and the image processing apparatus 200 is the addition of a fifth transformation subunit 315 and a fifth inverse transformation subunit 325. Other subunits and the feature extraction module are similar to those of the image processing apparatus 200, and their descriptions will be omitted as appropriate.
[0070] As mentioned earlier, the image after the inverse gamma transform operation of the first transform subunit 311 is: It resides in the linear RGB color space. The linear RGB color space has a fixed mapping matrix P to the standard color space CIE XYZ. -1 Conversely, there is a mapping matrix P from the standard CIE XYZ color space to the linear RGB color space. This transformation matrix is independent of external factors such as lighting and is calibrated by the CIE association under standard D65 lighting conditions, as shown below.
[0071]
[0072] Therefore, in this embodiment, after the first spatial transformation is performed by the first transformation subunit 311, the fifth transformation subunit 315 can perform another spatial transformation (through the mapping matrix) to obtain I. s XYZ Then, the first adjustment unit 340 provides input to the local feature extraction module 332 based on the input image and the output of the fifth transformation subunit 315.
[0073] It should be noted that in the first embodiment, the mapping matrix from the linear RGB color space to the standard color space CIE XYZ is fitted into parameters M and A during the training of the image processing device 200. Therefore, after the destyle transformation process, the image can be directly transformed to the standard color space CIE XYZ.
[0074] In the second embodiment, the first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a stylistic transformation, an inverse third spatial transformation, and an inverse first spatial transformation. Accordingly, the first transformation subunit 311 can perform the first spatial transformation, that is, transform the input image from the sRGB color space to the linear RGB color space; the fifth transformation subunit 315 can perform the third spatial transformation, that is, transform the image from the linear RGB color space to the standard color space CIE XYZ; the second transformation subunit 312 can perform the destyling transformation; the third transformation subunit 313 can perform the second spatial transformation, that is, transform the image from the standard color space CIE XYZ to the original color space of the camera device; and the fourth transformation subunit 314 can perform the inverse white balance transformation.
[0075] Accordingly, the first inverse transform subunit 321 can perform white balance transformation, the second inverse transform subunit 322 can perform second inverse spatial transformation, that is, transform from the original color space of the camera device to the standard color space CIE XYZ, the third inverse transform subunit 323 can perform style transformation, the fifth inverse transform subunit 325 can perform third inverse spatial transformation, that is, transform the image from the standard color space CIE XYZ to the linear RGB color space, and the fourth inverse transform subunit 324 can perform first inverse spatial transformation, that is, transform the input image from the linear RGB color space to the sRGB color space.
[0076] like Figures 7 to 9 As shown, the first global feature extraction module 331 can extract features based on the input image. The parameter γ is obtained and provided to the first transformation subunit 311 and the fourth inverse transformation subunit 324.
[0077] The first transformation subunit 311 can be based on the input image. The parameters γ provided by the first global feature extraction module 331 are obtained. The fifth transformation subunit 315 is based on And the transformation matrix P yields It is also provided to the first adjustment unit 340 and the second transformation subunit 312.
[0078] The first adjustment unit 340 can adjust the image Multiply by the scaling factor α1 and then combine with the original input image The parameters are added together and provided to the local feature extraction module 332. The local feature extraction module 332 can obtain parameters M and A, and provide them to the second transformation subunit 312 and the third inverse transformation subunit 323. The second transformation subunit 312 can obtain... It is also provided to the second adjustment unit 350 and the third transformation subunit 313. The second adjustment unit 350 can adjust the image... Multiply by the scaling factor α2 and then combine with the original input image The sums are then provided to the second global feature extraction module 333.
[0079] The second global feature extraction module 333 can extract the calculated parameters. Provided to the third transformation subunit 313, and the parameters are... Provided to the fourth transformation subunit 314. The third transformation subunit 313 calculates the image after the second spatial transformation. Furthermore, the fourth transformation subunit 314 calculates the intermediate image after inverse white balance transformation.
[0080] The fourth transformation subunit 314 provides the calculated intermediate image to the second transformation unit 320. Furthermore, the second global feature extraction module 333 calculates the parameter T... ideal and W ideal The image is provided to the second transformation unit 320. Next, the first inverse transformation subunit 321 can perform a white balance transformation to obtain image W. ideal I raw The second inverse transform subunit 322 can perform a second spatial inverse transform to obtain image T. ideal W ideal I raw The third inverse transform subunit 323 can perform style transformation to obtain the image (T). ideal W ideal I raw +A)⊙M, the fifth inverse transform subunit 325 can perform the third spatial inverse transform to obtain the image P((T) ideal W ideal I raw +A)⊙M), and the fourth inverse transform subunit 324 can perform the first spatial inverse transform, i.e., the gamma transform, to obtain the output image I. out =(P((T) ideal W ideal I raw +A)⊙<)) γ .
[0081] Therefore, the output image can also be represented as:
[0082]
[0083] In the image processing apparatus 300 according to an embodiment of the present disclosure, a global feature extraction module, a local feature extraction module, and a first transformation unit connected in series are used to extract the parameters required for the first set of transformations and the second set of transformations. Compared with the case where the global feature extraction module and the local feature extraction module are connected in parallel and output transformation parameters simultaneously, serially transmitting the calculation results of the preceding steps through the network structure allows more useful information to be incorporated into the transformation parameters, thereby obtaining better color adjustment results. Furthermore, the image color correction step is divided into two steps: from the input image to the electronic film (intermediate image) and from the electronic film to the corrected image. The calculation method from the electronic film to the input image is made consistent with the calculation method from the electronic film to the corrected image, thereby effectively adjusting the input image to the predetermined lighting conditions.
[0084] The following is combined Figures 10 to 12 To illustrate the image processing apparatus 400 according to the third embodiment of the present disclosure. Figure 10 This is a block diagram illustrating the structure of an image processing apparatus according to a third embodiment of the present disclosure. Figure 11 This is a simplified structural diagram illustrating an image processing apparatus according to a third embodiment of the present disclosure. Figure 12 This is a flowchart illustrating the image correction process of an image processing apparatus according to a third embodiment of the present disclosure.
[0085] like Figure 10 As shown, the image processing device 400 may include a first transformation unit 410, a second transformation unit 420, a feature extraction unit 430, a first adjustment unit 440, and a second adjustment unit 450. The first transformation unit 410 may include a first transformation subunit 411, a second transformation subunit 412, a third transformation subunit 413, a fourth transformation subunit 414, a fifth transformation subunit 415, and a sixth transformation subunit 416. The second transformation unit 420 may include a first inverse transformation subunit 421, a second inverse transformation subunit 422, a third inverse transformation subunit 423, a fourth inverse transformation subunit 424, a fifth inverse transformation subunit 425, and a sixth inverse transformation subunit 426. The feature extraction module 430 may include a first global feature extraction module 431, a local feature extraction module 432, and a second global feature extraction module 433.
[0086] The image processing apparatus 400 differs from the image processing apparatus 200 in that it adds a fifth transform subunit 415, a sixth transform subunit 416, a fifth inverse transform subunit 425, and a sixth inverse transform subunit 426. Other subunits and feature extraction modules are similar to those in the image processing apparatus 200 and will be described appropriately. Furthermore, the fifth transform subunit 415 and the fifth inverse transform subunit 425 in the image processing apparatus 400 perform the same processing as the fifth transform subunit 315 and the fifth inverse transform subunit 325 in the image processing apparatus 300. Therefore, the preceding descriptions of the fifth transform subunit 315 and the fifth inverse transform subunit 325 also apply to the fifth transform subunit 415 and the fifth inverse transform subunit 425.
[0087] In this embodiment, compared to image processing devices 200 and 300, a light conversion ratio k is introduced to represent the ratio of the exposure at the current color temperature to the exposure at the ideal color temperature (desired predetermined lighting conditions). This ratio can also be broken down into the exposure from the current color temperature to the original white balance exposure. And from the original white balance exposure to the ideal color temperature exposure (k). ideal .
[0088] In this embodiment, the first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a style transformation, a second spatial transformation, an inverse exposure correction transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an exposure correction transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation.
[0089] Accordingly, the first transformation subunit 411 can perform a first spatial transformation, that is, transform the input image from the sRGB color space to the linear RGB color space; the fifth transformation subunit 415 can perform a third spatial transformation, that is, transform the image from the linear RGB color space to the standard color space CIE XYZ; the second transformation subunit 412 can perform a destyling transformation; the third transformation subunit 413 can perform a second spatial transformation, that is, transform the image from the standard color space CIE XYZ to the original color space of the camera device; the sixth transformation subunit 416 can perform an inverse exposure correction transformation; and the fourth transformation subunit 414 can perform an inverse white balance transformation.
[0090] Accordingly, the first inverse transform subunit 421 can perform white balance transformation, the sixth inverse transform subunit 426 can perform exposure correction transformation, the second inverse transform subunit 422 can perform second inverse space transformation, that is, transform from the original color space of the camera device to the standard color space CIE XYZ, the third inverse transform subunit 423 can perform style transformation, the fifth inverse transform subunit 425 can perform third inverse space transformation, that is, transform the image from the standard color space CIE XYZ to the linear RGB color space, and the fourth inverse transform subunit 424 can perform first inverse space transformation, that is, transform the input image from the linear RGB color space to the sRGB color space.
[0091] like Figures 10 to 12 As shown, the first global feature extraction module 431 can extract features based on the input image. The parameter γ is obtained and provided to the first transformation subunit 411 and the fourth inverse transformation subunit 424.
[0092] The first transformation subunit 411 can be based on the input image. The parameters γ provided by the first global feature extraction module 431 are obtained. The fifth transformation subunit 415 is based on And the transformation matrix P yields It is also provided to the first adjustment unit 440 and the second transformation subunit 412.
[0093] The first adjustment unit 440 can adjust the image Multiply by the scaling factor α1 and then combine with the original input image The parameters are added together and provided to the local feature extraction module 432. The local feature extraction module 432 can obtain parameters M and A, and provide them to the second transformation subunit 412 and the third inverse transformation subunit 423. The second transformation subunit 412 can obtain... It is also provided to the second adjustment unit 450 and the third transformation subunit 413. The second adjustment unit 450 can adjust the image... Multiply by the scaling factor α2 and then combine with the original input image The sums are then provided to the second global feature extraction module 433.
[0094] The second global feature extraction module 433 can extract the calculated parameters. Provided to the third transformation subunit 413, the parameters Provided to the sixth transformation subunit 416, and the parameters are... The image is provided to the fourth transformation subunit 414. The third transformation subunit 413 obtains the image after the second spatial transformation. The sixth transformation subunit 416 obtains the image after exposure correction inverse transformation. Furthermore, the fourth transformation subunit 414 obtains the intermediate image after inverse white balance transformation.
[0095] The fourth transformation subunit 414 provides the calculated intermediate image to the second transformation unit 420. Furthermore, the second global feature extraction module 433 calculates the T... ideal k ideal and W ideal The image is provided to the second transformation unit 420. Next, the first inverse transformation subunit 421 can perform a white balance transformation to obtain image W. ideal I raw The sixth inverse transform subunit 426 can obtain the image k after exposure correction transform. ideal W ideal I raw The second inverse transform subunit 422 can perform a second spatial inverse transform to obtain image T. ideal k ideal W ideal I raw The third inverse transform subunit 423 can perform style transformation to obtain the image (T). ideal k ideal W ideal I raw +A)⊙M, the fifth inverse transform subunit 425 can perform the third spatial inverse transform to obtain the image P[(T ideal k ideal W ideal I raw +A)⊙M], and the fourth inverse transform subunit 424 can perform the first spatial inverse transform, i.e., the gamma transform, to obtain the output image I. out ={P[(T ideal k ideal W ideal I raw +A)⊙M]} γ .
[0096] Therefore, the output image can also be represented as:
[0097]
[0098] Typically, color constancy is affected by both white balance and exposure correction. White balance adjusts the color balance of an image to reflect human perception under natural light; it compensates for the color temperature of the light source to ensure that white objects appear white; and it includes adjusting the red, green, and blue channels to eliminate color casts caused by lighting conditions. Exposure correction adjusts the overall brightness and contrast of an image; it corrects underexposure (too dark) or overexposure (too bright) that may have occurred during shooting. White balance and exposure correction are closely related; a change in one will affect the other. In this embodiment, exposure correction and white balance are processed simultaneously, improving the performance of color correction.
[0099] Other embodiments:
[0100] The sixth transformation subunit 416 and the sixth inverse transformation subunit 426 of the image processing device 400 can be added to the image processing device 200, so that the exposure correction and white balance of the image can be processed simultaneously.
[0101] In the above-described embodiments of the image processing apparatus:
[0102] The first transformation unit transforms the input image from the sRGB color space to the original color space of the camera device through spatial transformation prior to the white balance inverse transformation.
[0103] The first global feature extraction module is configured to obtain parameters related to the first spatial transformation and the first spatial inverse transformation based on the input image.
[0104] The local feature extraction module is configured to obtain parameters related to destyling and style transformation based on the input image and a first transformed image obtained after the transformation prior to destyling.
[0105] The second global feature extraction module obtains parameters related to subsequent transformations after destyling in the first set of transformations and transformations corresponding to subsequent transformations in the second set of transformations, based on the input image and the second transformed image after destyling transformation.
[0106] The following describes a test of an image processing apparatus according to the present disclosure.
[0107] The basic model is a lightweight exposure correction neural network published at the British Machine Vision Conference (BMVC) in 2022. It uses parallel local feature extraction modules and global feature extraction modules to calculate parameters and does not have the reversible transformation process in this application.
[0108] The test model is: the image processing apparatus according to the third embodiment.
[0109] Table 1 shows the test results using the white balance correction dataset.
[0110] Table 1
[0111]
[0112] MSE (Mean Square Error) is the average Euclidean distance between the values in each channel of the output image and the values in each channel of the ground truth image in the RGB color space. MAE (Mean Angle Error) is the average angle between the pixel vector of the output image and the pixel vector of the ground truth image, calculated using the RGB three channels as the three axes of a coordinate system in the RGB color space.
[0113] ΔE 2000 (Delta E 2000) is a color difference metric used to quantify the visual difference between two colors. It is an improved version of Delta E, offering better consistency and adaptability to the characteristics of the human visual system, especially when dealing with subtle differences between colors.
[0114] Delta E is a measure of the difference between two colors. The original Delta E (or ΔE*ab) is based on the CIELAB color space and calculates the Euclidean distance between two points. However, this original version does not perfectly match human visual perception of color differences, especially sensitivity to certain color ranges. To address this issue, Delta E 2000 was proposed. It considers the influence of color attributes such as hue, saturation, and brightness on visual perception and adjusts the calculation formula to more accurately reflect these differences. Delta E 2000 is currently one of the most accurate color difference assessment methods and is widely used in color-related industries such as printing, coatings, plastics, lighting, and display calibration, where high color accuracy is required.
[0115] The calculation formula for Delta E 2000 is quite complex, including multiple correction terms to account for the interactions between different color parameters and nonlinear variations in the color space. In practice, this calculation is typically performed by specialized software or a custom algorithm.
[0116] Numerically, the value of Delta E 2000 is usually interpreted as follows:
[0117] ΔE<1.0: Differences not perceptible to the human eye.
[0118] 1.0 < ΔE < 2.0: Differences that can only be perceived by trained observers under specific conditions.
[0119] 2.0 < ΔE < 10.0: Differences that can be perceived by non-professional observers.
[0120] ΔE>10.0: Significant color difference.
[0121] As can be seen from Table 1, all three test values are reduced compared to the base model, thus this application can achieve better white balance correction.
[0122] Table 2 shows the test results using an exposure correction dataset (e.g., from MIT-Adobe FiveK).
[0123] Table 2
[0124]
[0125] SSIM (Structural Similarity Index Measure) and PSNR (Peak Signal-to-Noise Ratio) are both metrics used to quantify image quality, especially in image processing and computer vision, to assess the impact of operations such as image compression or image restoration on image quality. They measure the similarity or difference between images using different methods.
[0126] SSIM (Structural Similarity Index) is a method for comparing the visual differences between two images. Unlike error-accumulation-based metrics (such as MSE), SSIM considers the structural information of the images. SSIM values range from -1 to 1, with 1 indicating that the two images are identical. SSIM considers three aspects: brightness, contrast, and structure.
[0127] Brightness: Compare the average brightness of the two images.
[0128] Contrast ratio: The standard deviation of the contrast ratio of two images.
[0129] Structure: Comparing the structural information of two images essentially involves the correlation of their pixel values.
[0130] SSIM is a perceptual model designed to more closely resemble human visual perception. In practical applications, SSIM is widely used to evaluate the effectiveness of image processing tasks such as image compression, image transmission, watermarking, and denoising.
[0131] PSNR (Peak Signal-to-Noise Ratio) is a metric for evaluating the quality of image restoration, especially in the field of image compression. PSNR calculates the pixel-level difference between the original image and the processed image. It is a logarithmic representation based on the mean squared error (MSE).
[0132] Calculate MSE: Measure the average of the squared differences between corresponding pixels in two images.
[0133] Calculate PSNR: Put the result of MSE into a logarithmic scale, usually based on the maximum possible pixel value of the image (e.g., 255 for an 8-bit image).
[0134] PSNR is measured in decibels (dB), and a higher value indicates better image quality. However, because PSNR is an error-based measure, it does not always accurately reflect visual differences in image quality.
[0135] SSIM and PSNR are both important tools for evaluating image quality, but they focus on different aspects. SSIM focuses more on simulating the human eye's sensitivity to structural changes, while PSNR focuses on global pixel-level errors.
[0136] As shown in Table 2, both of these test indicators have increased, thus this application can achieve better exposure correction results.
[0137] In summary, compared with the basic model in the prior art, the method proposed in this application can automatically correct colors by simultaneously changing white balance and exposure, and achieve better performance.
[0138] The following is combined Figure 13 To describe an image processing method according to embodiments of the present disclosure.
[0139] like Figure 13 As shown, the image processing method according to an embodiment of the present disclosure begins at step S110. In step S110, an image output by a camera device is received as an input image, and a first set of transformations is performed on the input image to obtain an intermediate image. The intermediate image is an image in the original color space of the camera device.
[0140] Next, in step S120, a second set of transformations is performed on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image. In the output image, the input image is adjusted to predetermined lighting conditions. Furthermore, the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations. After this, the process ends.
[0141] In the above process, the parameters required for the first and second transformations are calculated by the feature extraction unit.
[0142] Therefore, the image processing method according to the embodiments of this disclosure can divide the image color correction step into two steps: from the input image to the electronic film (intermediate image) and from the electronic film to the corrected image. The calculation method from the electronic film to the input image is consistent with the calculation method from the electronic film to the corrected image, thereby adjusting the input image to the color under predetermined lighting conditions or under predetermined lighting conditions.
[0143] Various specific embodiments of the above-described steps of the image processing method according to the embodiments of this disclosure have been described in detail above and will not be repeated here.
[0144] Obviously, the various operational processes of the image processing method according to this disclosure can be implemented as a computer executable program stored in various machine-readable storage media.
[0145] Furthermore, the objective of this disclosure can also be achieved by providing a storage medium storing the aforementioned executable program code directly or indirectly to a system or device, and having a computer or central processing unit (CPU) in the system or device read and execute the aforementioned program code. In this case, as long as the system or device has the function of executing a program, the implementation of this disclosure is not limited to a program, and the program can be in any form, such as an object program, a program executed by an interpreter, or a script program provided to an operating system.
[0146] The aforementioned machine-readable storage media include, but are not limited to: various memories and storage units, semiconductor devices, disk units such as optical, magnetic and magneto-optical disks, and other media suitable for storing information.
[0147] Alternatively, the technical solution of this disclosure can also be implemented by connecting to a corresponding website on the Internet, downloading and installing the computer program code according to this disclosure onto the computer, and then executing the program.
[0148] Figure 14 This is a block diagram of an exemplary structure of a general-purpose personal computer in which image processing apparatus and methods according to embodiments of the present disclosure can be implemented.
[0149] like Figure 14 As shown, CPU 1401 executes various processes based on programs stored in read-only memory (ROM) 1402 or programs loaded into random access memory (RAM) 1403 from storage device 1408. RAM 1403 also stores data required as needed when CPU 1401 executes various processes, etc. CPU 1401, ROM 1402, and RAM 1403 are interconnected via bus 1404. Input / output interface 1405 is also connected to bus 1404.
[0150] The following components are connected to input / output interface 1405: input device 1406 (including keyboard, mouse, etc.), output device 1407 (including display, such as cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.), storage device 1408 (including hard disk, etc.), and communication device 1409 (including network interface card, such as LAN card, modem, etc.). Communication device 1409 performs communication processing via a network, such as the Internet. Drive 1410 may also be connected to input / output interface 1405 as needed. Removable media 1411, such as disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1410 as needed, so that computer programs read from them can be installed into storage device 1408 as needed.
[0151] When the above series of processes are implemented by software, the program constituting the software is installed from a network such as the Internet or a storage medium such as removable medium 1411.
[0152] Those skilled in the art will understand that such storage media are not limited to Figure 14 The illustration shows a removable medium 1411 containing a program, distributed separately from the device to provide the program to the user. Examples of removable media 1411 include magnetic disks (including floppy disks (registered trademark)), optical disks (including optical disc read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including mini-discs (MD) (registered trademark)), and semiconductor memory. Alternatively, the storage medium may be ROM 1402, a hard disk included in storage device 1408, etc., containing programs and distributed to the user along with the device containing them.
[0153] In the systems and methods of this disclosure, it is apparent that the components or steps can be decomposed and / or recombined. Such decomposition and / or recombination should be considered equivalent to those disclosed. Furthermore, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order. Some steps can be performed in parallel or independently of each other.
[0154] While embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative and do not constitute a limitation thereof. Those skilled in the art can make various modifications and alterations to the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is defined only by the appended claims and their equivalents.
[0155] The various techniques described in this specification can be performed independently of each other unless there is a contradiction. Of course, any of the various techniques can be performed in combination. In one example, some or all of the techniques described in another embodiment can be combined to perform some or all of the techniques described in any embodiment. Additionally, any part or all of the techniques described above can be combined with another technique not described above.
[0156] Regarding the implementation methods including the above embodiments, the following notes are also disclosed:
[0157] Appendix 1. An image processing apparatus, comprising:
[0158] The first transformation unit receives the image output by the camera device as the input image and performs a first set of transformations on the input image to obtain an intermediate image, wherein the intermediate image is the image in the original color space of the camera device.
[0159] A second transformation unit performs a second set of transformations on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image in which the input image is adjusted to predetermined lighting conditions, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations; and
[0160] The feature extraction unit calculates the parameters required in the first set of transformations and the second set of transformations.
[0161] Appendix 2. The image processing apparatus according to Appendix 1, wherein the first set of transformations and the second set of transformations each include one of the following:
[0162] The first set of transformations sequentially includes a first spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, and an inverse first spatial transformation;
[0163] The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation; and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation; and
[0164] The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a style transformation, a second spatial transformation, an inverse exposure correction transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an exposure correction transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation.
[0165] Note 3. According to the image processing apparatus described in Note 2, the first transformation unit transforms the input image from the sRGB color space to the original color space of the camera device through a spatial transformation prior to the white balance inverse transformation.
[0166] Appendix 4. The image processing apparatus according to Appendix 2, wherein the feature extraction unit includes a first global feature extraction module, a local feature extraction module and a second global feature extraction module connected in series, and provides parameters related to the first set of transformations and the second set of transformations through the feature extraction unit and the first transformation unit.
[0167] Note 5. The image processing apparatus according to Note 4, wherein the first global feature extraction module is configured to obtain parameters related to the first spatial transformation and the first spatial inverse transformation based on the input image.
[0168] Note 6. The image processing apparatus according to Note 5, wherein the local feature extraction module is configured to obtain parameters related to the destyling transformation and the style transformation based on the input image and a first transformed image obtained after the transformation prior to the destyling transformation.
[0169] Note 7. The image processing apparatus according to Note 6, wherein the second global feature extraction module obtains parameters related to subsequent transformations in the first set of transformations after the destyling transformation and transformations in the second set of transformations corresponding to the subsequent transformations, based on the input image and the second transformed image after the destyling transformation.
[0170] Note 8. The image processing apparatus according to Note 4, wherein the corresponding parameters of the first global feature extraction module, the local feature extraction module and the second global feature extraction module are determined by training the image processing apparatus, and during the training of the image processing apparatus, the loss function includes a term representing the difference between the calculated intermediate images of different images from the same camera device in the same scene.
[0171] Appendix 9. The image processing apparatus according to Appendix 2, wherein,
[0172] The first spatial transformation transforms the input image from the sRGB color space to the linear RGB color space, and the first inverse spatial transformation transforms the image from the linear RGB color space back to the sRGB color space.
[0173] The second spatial transformation transforms the image from the standard color space CIE XYZ to the original color space of the camera device, and the second inverse spatial transformation transforms the image from the original color space of the camera device to the standard color space CIE XYZ.
[0174] The third space transformation transforms the image from the linear RGB color space to the standard color space CIE XYZ, and the third space inverse transformation transforms the image from the standard color space CIE XYZ back to the linear RGB color space.
[0175] Note 10. The image processing apparatus according to Note 4, wherein the local feature extraction module is implemented by two parallel local feature extraction units, the two parallel local feature extraction units being used to extract different parameters used in both the destyling transformation and the style transformation.
[0176] Note 11. The image processing apparatus according to Note 6, wherein the image processing apparatus further comprises: a first adjustment unit configured to adjust a first scaling factor of the first transformed image input to the local feature extraction module.
[0177] Note 12. The image processing apparatus according to Note 7, wherein the image processing apparatus further comprises: a second adjustment unit configured to adjust a second scaling factor of the second transformed image input to the second global feature extraction module.
[0178] Note 13. The image processing apparatus according to Note 11 or 12, wherein the first scaling factor and the second scaling factor are learnable vectors.
[0179] Appendix 14. An image processing method, comprising:
[0180] The image output by the camera device is received as an input image, and a first set of transformations is performed on the input image to obtain an intermediate image, wherein the intermediate image is an image in the original color space of the camera device; and
[0181] A second set of transformations is performed on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image in which the input image is adjusted to predetermined lighting conditions, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations.
[0182] Specifically, the feature extraction unit calculates the parameters required for the first set of transformations and the second set of transformations.
[0183] Note 15. According to the method described in Note 14, the first set of transformations and the second set of transformations each include one of the following cases:
[0184] The first set of transformations sequentially includes a first spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, and an inverse first spatial transformation;
[0185] The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation; and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation; and
[0186] The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a style transformation, a second spatial transformation, an inverse exposure correction transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an exposure correction transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation.
[0187] Appendix 16. According to the method described in Appendix 15, the feature extraction unit comprises a first global feature extraction module, a local feature extraction module, and a second global feature extraction module connected in series, and the calculation of the parameters includes:
[0188] The feature extraction unit and the first transformation unit that performs the first set of transformations provide parameters related to the first set of transformations and the second set of transformations.
[0189] Note 17. According to the method described in Note 16, calculating the parameter includes:
[0190] The first global feature extraction module obtains parameters related to the first spatial transformation and the first spatial inverse transformation based on the input image.
[0191] Note 18. According to the method described in Note 17, calculating the parameter includes:
[0192] The local feature extraction module obtains parameters related to the destyling transformation and the style transformation based on the input image and the first transformed image obtained after the transformation before the destyling transformation.
[0193] Note 19. The method according to Note 18, wherein calculating the parameter includes:
[0194] The second global feature extraction module obtains parameters related to subsequent transformations in the first set of transformations after the destyling transformation and transformations corresponding to the subsequent transformations in the second set of transformations, based on the input image and the second transformed image after the destyling transformation.
[0195] Appendix 20. A machine-readable storage medium carrying a program product including machine-readable instruction code stored thereon, wherein, when read and executed by a computer, the instruction code enables the computer to perform the image processing method according to Appendices 14-19.
Claims
1. An image processing apparatus, comprising: The first transformation unit receives the image output by the camera device as the input image and performs a first set of transformations on the input image to obtain an intermediate image, wherein the intermediate image is the image in the original color space of the camera device. The second transformation unit performs a second set of transformations on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image in which the input image is adjusted to predetermined lighting conditions, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations. as well as The feature extraction unit calculates the parameters required in the first set of transformations and the second set of transformations.
2. The image processing apparatus according to claim 1, wherein, The first set of transformations and the second set of transformations each include one of the following cases: The first set of transformations sequentially includes a first spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, and an inverse first spatial transformation; The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a destyling transformation, a second spatial transformation, and an inverse white balance transformation; and the second set of transformations sequentially includes a white balance transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation; and The first set of transformations sequentially includes a first spatial transformation, a third spatial transformation, a style transformation, a second spatial transformation, an inverse exposure correction transformation, and an inverse white balance transformation, and the second set of transformations sequentially includes a white balance transformation, an exposure correction transformation, an inverse second spatial transformation, a style transformation, an inverse third spatial transformation, and an inverse first spatial transformation.
3. The image processing apparatus according to claim 2, wherein, The first transformation unit transforms the input image from the sRGB color space to the original color space of the camera device through spatial transformation prior to the white balance inverse transformation.
4. The image processing apparatus according to claim 2, wherein, The feature extraction unit includes a first global feature extraction module, a local feature extraction module, and a second global feature extraction module connected in series, and provides parameters related to the first set of transformations and the second set of transformations through the feature extraction unit and the first transformation unit.
5. The image processing apparatus according to claim 4, wherein, The first global feature extraction module is configured to obtain parameters related to the first spatial transformation and the first spatial inverse transformation based on the input image.
6. The image processing apparatus according to claim 5, wherein, The local feature extraction module is configured to obtain parameters related to the destyling transformation and the style transformation based on the input image and a first transformed image obtained after the transformation prior to the destyling transformation.
7. The image processing apparatus according to claim 6, wherein, The second global feature extraction module obtains parameters related to subsequent transformations in the first set of transformations after the destyling transformation and transformations corresponding to the subsequent transformations in the second set of transformations based on the input image and the second transformed image after the destyling transformation.
8. The image processing apparatus according to claim 4, wherein, The parameters of the first global feature extraction module, the local feature extraction module, and the second global feature extraction module are determined by training the image processing device. During the training process of the image processing device, the loss function includes a term representing the difference between the calculated intermediate images of different images from the same camera device in the same scene.
9. An image processing method, comprising: The image output by the camera device is received as the input image, and a first set of transformations is performed on the input image to obtain an intermediate image, which is an image in the original color space of the camera device; as well as A second set of transformations is performed on the intermediate image in the reverse order of the transformations in the first set of transformations to obtain an output image in which the input image is adjusted to predetermined lighting conditions, and wherein the transformations in the second set of transformations are the inverse processes of the corresponding transformations in the first set of transformations. Specifically, the feature extraction unit calculates the parameters required for the first set of transformations and the second set of transformations.
10. A machine-readable storage medium carrying a program product including machine-readable instruction code stored thereon, wherein, When the instruction code is read and executed by a computer, it enables the computer to perform the method according to claim 9.