A method for harmonizing color deviation in image fusion

Through the conditional probability diffusion model and the mask deep supervision loss function, the problem that real-world image harmony algorithm in the existing technology cannot handle complex degradation is solved, and high-quality image harmony effect is achieved.

CN116385279BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310024355.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-06
Publication Date
2025-07-22
Estimated Expiration
2043-01-06

AI Technical Summary

Technical Problem

Existing image harmony algorithms cannot effectively deal with complex degraded scenes in the real world, ignoring factors such as noise and resolution, resulting in the generated images being unrealistic.

Method used

Using a method based on conditional probability diffusion model, the degradation estimation network, U-Net and Gaussian distribution sampling is gradually iteratively generated harmonious images, and the mask deep supervision loss function is used for training to achieve flexible modeling and high-quality harmony of different degrees of degradation.

Benefits of technology

The generated images have high fidelity and can effectively process images containing noise and other degraded in the real world, improving the stability and generation quality of the harmonious model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385279B_ABST
    Figure CN116385279B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for harmonizing color deviation in image fusion, which solves the problem that the existing image harmonization algorithms lack the ability to model and process the harmonization of complex real scenes. The method includes the following steps: Step 1, foreground degradation estimation, which includes inputting a discordant synthetic image and a foreground mask image into a degradation estimation network to estimate a degradation map, outputting a degraded image, and further encoding the degraded image to obtain an encoding of the degradation degree; Step 2, noise mean and variance estimation, which includes obtaining estimates of the noise mean and variance according to the discordant synthetic image, the foreground mask image, and the obtained encoding of the degradation degree; Step 3, harmonized image sampling, which includes uniquely determining a Gaussian distribution from the obtained noise mean and variance, and sampling from this Gaussian distribution to obtain the harmonized image in this iteration; Step 4, execute Steps 1 to 3 iteratively to generate a harmonized image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and image processing, and particularly relates to a method for harmonizing color deviation of synthetic images in the real world. Background Art

[0002] Image harmonization is a task of adjusting the visual appearance (brightness, contrast, texture, etc.) of the foreground in a synthetic image to make it consistent with the background. Image harmonization technology is widely applied in scenarios such as virtual reality, art creation, video conferencing, and e-commerce. Although existing deep learning-based methods have made great progress, they all implicitly assume that only the illumination distribution is different between the foreground and background images. This assumption ignores the fact that images captured by different acquisition devices in different real-world scenarios inevitably have some degradations, such as noise and resolution. Existing methods cannot model such complex degradation scenarios and are therefore not suitable for dealing with the problem of real-world image harmonization.

[0003] Traditional image transfer and generation are achieved based on the mapping between the inherent statistical characteristics of images, such as color distribution mapping, gradient matching, and multi-scale mapping. However, image statistical features cannot model the distribution of complex images, resulting in their ability to handle only some simple situations. With the rise of deep learning technology, deep learning-based image harmonization technology has received increasing attention. According to whether paired data is used for training, deep learning-based image harmonization methods can be divided into supervised methods and unsupervised methods.

[0004] Supervised methods. Tsai et al. proposed an end-to-end CNN image harmonization network in the document "Y. Tsai, X. Shen, Z. Lin, K. Sunkavalli, X. Lu, and M. Yang. Deep Image Harmonization. in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3789-3797." and used an auxiliary scene parsing branch to improve the basic image harmonization network. Cun et al. designed an additional spatial separation attention module for image harmonization in the document "X. Cun and C. Pun. Improving the Harmony of the Composite Image by Spatial-Separated Attention Module. IEEE Transactions on Image Processing, vol. 29, no. 1, pp. 4759-4771, 2020." to learn regional appearance changes in low-level features. Cong et al. introduced the concept of domain and proposed a domain verification discriminator to bring the foreground domain and background domain closer in the document "W. Cong, J. Zhang, L. Niu, L. Liu, Z. Ling, W. Li, and L. Zhang. DoveNet: Deep Image Harmonization via Domain Verification. in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 8394-8403." Guo et al. in the paper "Z. Guo, H. Zheng, Y. Jiang, Z. Gu, and B. Zheng. Intrinsic Image Harmonization. in Proc. IEEE Conference on Computer Vision and Pattem Recognition, 2021, pp. 16362-16371." realized that the disharmony stems from the intrinsic reflectivity and illumination differences between the foreground and background, and proposed an autoencoder to decompose the synthetic image into reflectivity and illumination for harmonization separately, where the reflectivity is harmonized through material consistency penalty, and the illumination is harmonized by learning and transferring light from the background to the foreground.In the literature "J. Ling, H. Xue, L. Song, R. Xie, and X. Gu. Region-Aware Adaptive Instance Normalization for Image Harmonization. in Proc. IEEE Conference on Computer Vision and Pattem Recognition, 2021, pp. 9357-9366.", Ling et al. reformulated image harmonization as a background-to-foreground style transfer problem and proposed a region-aware adaptive instance normalization module to explicitly represent the visual styles from the background and adaptively apply them to the foreground. Guo et al. proposed using the ability of Transformer to model long-range dependencies to solve the problem of image harmonization in the literature "Z. Guo, D. Guo, H. Zheng, Z. Gu, B. Zheng, and J. Dong. Image Harmonization with Transformer. in Proc. IEEE International Conference on Computer Vision, 2021, pp. 14870-14879.".

[0005] Unsupervised methods. Chen and Kae proposed a geometric and color-consistent GAN (FCC-GAN) for realistic image synthesis in the literature "B. Chen and A. Kae. Toward Realistic Image Compositing with Adversarial Learning. in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 8407-8416." GCC-GAN considers the consistency of geometry, color, and boundaries simultaneously, making the synthesized image indistinguishable from the real image. Jiang et al. proposed a self-supervised framework for image harmonization in the literature "Y. Jiang, H. Zhang, J. Zhang, Y. Wang, Z. Lin, K. Sunkavalli, S. Chen, S. Amirghodsi, S. Kong, and Z. Wang. SSH: A Self-Supervised Framework for Image Harmonization. in Proc. IEEE International Conference on Computer Vision, 2021, pp. 4832-4841." The framework extracts content and appearance representations from the foreground and background respectively, and then aggregates these representations to generate the harmonized output. Based on this form, the authors proposed a dual data augmentation engine to generate various synthetic data triples to support self-supervised training. At the same time, the authors proposed to use a three-dimensional color lookup table to replace the traditional color transfer enhancement, which can instantly generate more diverse samples. Summary of the Invention

[0006] The present invention discloses a method for harmonizing real-world images, which mainly solves the problem that existing image harmonization algorithms lack the ability to model and process the harmonization of complex real scenes. Specifically, the object of the present invention is to improve the following aspects.

[0007] 1. Existing image harmonization methods only consider the difference in illumination distribution between the foreground and background regions, lacking the consideration of other distribution differences, such as noise, resolution, etc.

[0008] 2. Existing image harmonization methods only model the illumination distribution, lacking the ability to model different degrees of degradation distributions.

[0009] 3. Existing methods for image harmonization use a pixel-based loss function to train a neural network to complete the mapping from the synthetic image to the real image, making the generated image easily regress to the mean of multiple images.

[0010] The present invention proposes a method for harmonizing real-world images based on a conditional probability diffusion model to improve the quality of image harmonization in real-world scenarios. The technical solution for implementing the present invention includes the following steps:

[0011] 1. Foreground degradation estimation. First, a degradation estimation network is used to estimate the degradation map. The degradation estimation network consists of multiple cascaded convolutional blocks, and each convolutional block is composed of a convolutional layer and a ReLU activation layer. This network takes the disharmonious synthetic image and the foreground mask image as inputs. The output is the degraded image. Then, the mean values are calculated in both the vertical and horizontal dimensions, and the mean values are input into an nn.Embedding layer to obtain the encoding regarding the degradation degree.

[0012] 2. Noise mean and variance estimation. The estimation of the noise mean and variance is completed by a U-Net. Its input consists of the following three parts: the harmonized image sampled in the previous step, the original synthetic image, and the foreground mask. These three parts are concatenated as the input of the U-Net. Additionally, the current step needs to be encoded as and then added to and inserted into the residual blocks at each level of the U-Net so that the current step number and degradation degree can be perceived during the process of estimating the noise mean and variance. The output of the U-Net is the mean and variance of the noise distribution.

[0013] 3. Sampling of the harmonized image. The mean and variance obtained can uniquely determine a Gaussian distribution, and sampling from this Gaussian distribution can obtain the harmonized image at this step.

[0014] 4. Iterative generation of the harmonized image. The inference process of the diffusion model is a step-by-step iterative generation process. The more sampling steps, the higher the quality of the finally generated harmonized image. Weighing the generation quality and time cost, we set the number of sampling steps to 1000. At each step, the above four steps 1-4 are sequentially executed, and finally a complete harmonized image is obtained.

[0015] Generally speaking, the present invention makes full use of the inherent attributes of real-world synthetic images and uses 1000 steps to generate the final harmonized image. At each step, first, the degradation estimation network is used to estimate the degradation level of the foreground area of the synthetic image. Then, a U-Net is used to estimate the noise mean and variance of the intermediate result of the harmonized image generated in the previous step. Finally, the intermediate result of the harmonized image at this step is sampled from the Gaussian distribution. Our method is stable in training, and the generated harmonized images have high fidelity. Experiments prove that our method is significantly better than other methods.

[0016] The present invention has the following advantages over traditional hyperspectral image restoration methods:

[0017] 1. It can not only harmonize simple illumination, but also harmonize synthetic images in the real world with other degradations.

[0018] 2. According to different types of degradations existing in the foreground region of the synthetic image, different types and degrees of degradations can be flexibly modeled.

[0019] 3. By using a diffusion model to model the distribution migration, the stability of the distribution migration and the harmonization model training can be improved.

[0020] 4. By using a masked depth supervision loss function to train the model, the fidelity of the generated harmonized image can be greatly improved.

[0021] A method for harmonizing color deviation in image fusion disclosed by the present invention is characterized in that the method includes:

[0022] Step 1, foreground degradation estimation, including inputting a discordant synthetic image and a foreground mask image into a degradation estimation network to estimate a degradation map, outputting a degraded image, and further encoding the degraded image to obtain an encoding of the degradation degree, wherein the degradation estimation network is composed of multiple cascaded convolutional blocks;

[0023] Step 2, noise mean and variance estimation, including obtaining an estimation of the noise mean and variance according to the discordant synthetic image, the foreground mask image, and the obtained encoding of the degradation degree;

[0024] Step 3, harmonized image sampling, including uniquely determining a Gaussian distribution from the obtained noise mean and variance, and sampling from the Gaussian distribution to obtain the harmonized image in this iteration;

[0025] Step 4, execute Steps 1 to 3 iteratively to generate a harmonized image.

[0026] In one embodiment, in the foreground degradation estimation, the degradation estimation network takes the discordant synthetic image I cond and the foreground mask image I mask as inputs, and the output is the degraded image I D , and the calculation formula is expressed as:

[0027] I D = DEN(I cond , I mask ), (1)

[0028] E D = nn.Embedding(I D ), (2)

[0029] where DEN(.) represents the degradation estimation network composed of multiple cascaded convolutional blocks, each convolutional block consists of a convolutional layer and a ReLU activation layer, and nn.Embedding(.) directly represents the embedding layer in the Pytorch architecture to obtain the encoding E of the degradation degree D 。

[0030] In one embodiment, the noise mean and method estimation includes using a segmentation network U-Net to complete the estimation of the mean and variance of the Gaussian distribution noise, and its input consists of the following three parts: the harmonized image I sampled in the previous iteration t-1 , where I0 represents the image of pure isotropic Gaussian noise, where t is the current iteration step number, and the unharmonious synthetic image I cond and the foreground mask image I mask are input into the segmentation network U-Net, and the current iteration step is encoded as E T Then it is added to E D and inserted into the residual blocks of each level of the segmentation network U-Net, so that the current iteration step number and degradation degree can be perceived during the process of estimating the noise mean and variance. This process can be expressed as follows:

[0031] E = E T + E D , (3)

[0032] u t , σ t 2 = UNet(I t-1 , I cond , I mask , E), (4)

[0033] where, u t , σ t 2 respectively represent the mean and variance of the Gaussian noise distribution, and E represents the sum of the current iteration step number encoding E T and the degradation encoding E D .

[0034] In one embodiment, for the harmonized image sampling, the harmonized image sample I at the t-th iteration step t can be sampled from the Gaussian distribution N(u t , σ t 2 ) as follows:

[0035] I t = u t + σε, σ ∼ N(ε, 0, I), (5)

[0036] Among them, u t , σ t 2 are the mean and variance obtained from the previous step prediction. ò is sampled from the standard normal distribution to obtain the foreground region of I t as the foreground for generating the harmonized result, and its background region is directly copied from I cond , which is expressed by the formula:

[0037]

[0038] where * represents pixel-wise multiplication, and I t is the intermediate harmonized image generated in the t-th iteration.

[0039] In one embodiment, the iterative generation of the harmonized image includes sequentially executing the above steps 1 to 3. After executing T = 1000 times, the final harmonized image I T is obtained.

[0040] In one embodiment, the following loss function L total is used to optimize the entire model:

[0041] L total = L est + L deep , (7)

[0042] where L est is the degradation estimation loss function, which is expressed as:

[0043] L est = E t~[1,T] ||D gt - I D || 2 , (8)

[0044] where D gt represents the true degradation map, and I D is the degradation map output by the degradation estimation network.

[0045] L deep represents the multi-scale deep supervision loss function, which can be expressed as follows,

[0046]

[0047] It is characterized in that two scales of outputs are added to the segmentation network U-Net. The prediction result of the original size is output at the last level of the decoder of the segmentation network U-Net. A lightweight convolutional layer is connected to the second last and the third last levels respectively to output prediction results of different scales, which are the noise means and variances of 1 / 2 and 1 / 4 times of the original size respectively. Here, s represents the level, and the maximum value is 3, indicating that the generated harmonized images are supervised at three different scales. When the value of s is 1, it represents the original size. When the values of s are 2 and 3, they represent 1 / 2 and 1 / 4 times of the original size respectively. denotes the true harmonized image at level s. denotes the generated harmonized image at level s. MSE represents the mean squared error loss. * represents the pixel-wise product. denotes the mask at level s, multiplied by denotes that the supervision is only performed in the mask region. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the specific flowchart for the implementation of the present invention (RIHD). DETAILED DESCRIPTION OF THE INVENTION

[0049] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. The preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will further clarify the present invention by combining the drawings in the embodiments of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0051] Refer to Figure 1 , the specific implementation steps of the present invention are as follows:

[0052] Step 1, foreground degradation estimation. First, a degradation estimation network is used to estimate the degradation map. This network takes the unharmonious synthetic image I cond and the foreground mask image I mask as inputs. The output is the degraded image I D . It is expressed by the formula:

[0053] ID = DEN(I cond , I mask ), (10)

[0054] E D = nn.Embedding(I D ), (11)

[0055] where DEN(.) represents a degradation estimation network composed of multiple cascaded convolutional blocks, and each convolutional block consists of a convolutional layer and a ReLU activation layer. nn.Embedding(.) directly represents the embedding layer in the Pytorch architecture.

[0056] Step 2, noise mean and method estimation. The present invention uses a U-Net to estimate the mean and variance of Gaussian distributed noise. Its input consists of the following three parts: the harmonized image I t-1 sampled in the previous step, the original synthetic image Ic ond and the foreground mask I mask . These three parts are concatenated as the input of the U-Net. Additionally, the current step is encoded as E T and then added to E D and inserted into the residual blocks at each level of the U-Net, so that the current step number and degradation degree can be perceived during the process of estimating the noise mean and variance. This process can be expressed as follows:

[0057] E = E T + E D , (12)

[0058] u t , σ t 2 = UNet(I t+1 , I cond , I mask , E), (13)

[0059] where u t , σ t 2 respectively represent the mean and variance of the Gaussian noise distribution, and E represents the sum of the current iteration step encoding E T and the degradation encoding E D .

[0060] Step 3, harmonized image sampling. The harmonized image sample at the t-th step can be sampled from the Gaussian distribution N(u t , σ t 2 ), as follows:

[0061]

[0062] Among them, u t , σ t 2 are the mean and variance obtained from the previous step prediction, and ò is sampled from the standard normal distribution. We obtain the foreground region of I t as the foreground for generating the harmonization result, and its background region is directly copied from I cond , which is expressed by the formula:

[0063]

[0064] where * represents pixel-by-pixel multiplication. I t is the intermediate harmonization result generated at the t-th step (where I0 represents the image of pure isotropic Gaussian noise).

[0065] Step 4, iterative generation of the harmonized image. After sequentially executing the above three steps 1-3 and performing T = 1000 steps, the final harmonized image I T is obtained.

[0066] The present invention uses the following loss function L total to optimize the entire model:

[0067] L total = L est + L deep , (16)

[0068] where L est represents the degradation estimation loss function as:

[0069] L est = E t~[1,T] ||D gt - I D || 2 , (17)

[0070] where D gt represents the true degradation map. I D is the degradation map output by the degradation estimation network.

[0071] L deep represents the multi-scale deep supervision loss function, which can be expressed as follows,

[0072]

[0073] We added two scales of outputs to the U-Net. The prediction results of the original size are output at the last level of the U-Net decoder. A lightweight convolutional layer is connected to the second last and the third last levels respectively to output the prediction results of different scales, which are 1 / 2 and 1 / 4 times the original size, the noise mean and variance. Here, s represents the level, and the maximum value is 3, indicating that the generated harmonized images are supervised at three different scales. When the s value is 1, it represents the original size. When the s values are 2 and 3, they represent 1 / 2 and 1 / 4 times the original size respectively. represents the real harmonized image at level s. represents the generated harmonized image at level s. MSE represents the MSE Loss, and * represents the pixel-by-pixel product. represents the mask at level s, multiplied by indicates that the supervision is only carried out in the mask area. Experiments have proved that this method can greatly improve the fidelity and sampling efficiency of the generated harmonized images.

[0074] The effects of the present invention can be further illustrated by the following simulation experiments.

[0075] 1. Simulation conditions

[0076] The present invention is based on the Pytorch framework for experimental simulation on a central processing unit of i7-6800K@3.4GHz CPU, 64G memory, NVIDIA GeForce RTX 3090 GPU, and Ubuntu 18.04 operating system.

[0077] The data used in the experiment is the D-iHarmony4 dataset newly created by us based on the iHarmony4 image harmonization dataset proposed by W. Cong et al. in the literature "W. Cong, J. Zhang, L. Niu, L. Liu, Z. Ling, W. Li, and L. Zhang. DoveNet: Deep Image Harmonization via Domain Verification. in Proc. IEEE Conference on Computer Vision and Pattem Recognition, 2020, pp. 8394-8403." to adapt to real-world scenarios. The newly created dataset is called D-iHarmony4. D-iHarmony4 also consists of four sub-datasets, namely D-HCOCO, D-HAdobe5k, D-HFlickr, and D-Hday2night. Each sample in the D-iHarmony4 dataset consists of a disharmonious synthetic image and its paired harmonized image. Compared with the original iHarmony4 dataset, various degrees of degradation, such as noise, are added to the foreground of the synthetic images in D-iHarmony4 to make the synthetic images more in line with real-world scenarios.

[0078] Simulation content

[0079] In the simulation, the algorithm model is first trained using the training set, and then the disharmonious synthetic images of each sample in the test set are input into the trained model. Finally, the harmonized images are obtained, and the harmonized images are compared and evaluated with the real images.

[0080] To prove the effectiveness of the algorithm, the deep learning-based harmonization method DoveNet, S 2 AM and IntrinsicNet are selected as comparison algorithms on the D-iHarmony4 dataset. Among them, DoveNet was proposed in the literature "W. Cong, J. Zhang, L. Niu, L. Liu, Z. Ling, W. Li, and L. Zhang. DoveNet: Deep Image Harmonization via Domain Verification. in Proc. IEEE Conference on Computer Vision and Pattem Recognition, 2020, pp. 8394-8403."; S 2AM was proposed in the literature "X. Cun and C. Pun. Improving the Harmony of the Composite Image by Spatial-Separated Attention Module. IEEE Transactions on Image Processing, vol. 29, no. 1, pp. 4759-4771, 2020."; IntrinsicNet was proposed in the literature "Z. Guo, H. Zheng, Y. Jiang, Z. Gu, and B. Zheng. Intrinsic Image Harmonization. in Proc. IEEE Conference on Computer Vision and Pattem Recognition, 2021, pp. 16362-16371."; RIHD represents our method. MSE and PSNR are evaluation metrics for the harmonized images, and Table 1 shows the evaluation results on the test set of the D-iHarmony4 dataset.

[0081] Table 1 Comparison Experiment Results

[0082]

[0083] As can be seen from Table 1, on the D-iHarmony4 test set, the quality of the harmonized images of the present invention (RIHD) is significantly better than that of other algorithms. Among them, the PSNR value is on average 1.11 dB higher than that of the sub-optimal IntrinsicNet method on the four sub-datasets.

[0084] The above embodiments are the preferred embodiments of this patent, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for harmonizing color deviation in image fusion, characterized in that, The method includes: Step 1, foreground degradation estimation, which includes inputting a discordant synthetic image and a foreground mask image into a degradation estimation network to estimate a degradation map, outputting a degraded image, and further encoding the degraded image to obtain an encoding of the degradation degree, where the degradation estimation network consists of multiple cascaded convolutional blocks; Step 2, noise mean and variance estimation, which includes obtaining estimates of the noise mean and variance based on the discordant synthetic image, the foreground mask image, and the obtained encoding of the degradation degree; Step 3, harmonious image sampling, which includes uniquely determining a Gaussian distribution from the obtained noise mean and variance, and sampling from this Gaussian distribution to obtain the harmonious image in this iteration; Step 4, performing Steps 1 to 3 iteratively to generate a harmonious image; It is characterized in that the noise mean and method estimation include using a segmentation network U-Net to complete the estimation of the mean and variance of Gaussian distribution noise, and its input consists of the following three parts: the harmonized image obtained by the previous iteration sampling , where represents an image of pure isotropic Gaussian noise, where t is the current iteration step number, and the unharmonious synthetic image and the foreground mask image are input into the segmentation network U-Net, and the current iteration step is encoded as Then it is added to and inserted into the residual blocks of each level of the segmentation network U-Net, so that the current iteration step number and the degradation degree can be perceived during the process of estimating the noise mean and variance. This process can be expressed as follows: , (3) ,(4) Among them, respectively represent the mean and variance of the Gaussian noise distribution, E represents the encoding of the current iteration step and the degraded encoding The sum; Among them, for the harmonized image sampling, the harmonized image sample of the t th iteration step can be sampled from a Gaussian distribution as follows: ,(5) Among them, is the mean and variance obtained from the previous step of prediction, is sampled from the standard normal distribution to obtain The foreground region of is used as the foreground for generating the harmonized result, and its background region is directly copied from , which is expressed by the formula: , (6) Among them represents the pixel-by-pixel product, is the intermediate harmonized image generated in the t th iteration.

2. The method for harmonizing color deviation in image fusion according to claim 1, wherein In the foreground degradation estimation, the degradation estimation network uses the discordant composite image and the foreground mask image as inputs, and the output is the degraded image , and the calculation formula is expressed as: ,(1) ,(2) Among them represents the degradation estimation network composed of multiple cascaded convolutional blocks, and each convolutional block consists of a convolutional layer and a ReLU activation layer directly represents the encoding of the degradation degree obtained by the embedding layer in the Pytorch architecture .

3. The method for harmonizing color deviation in image fusion according to claim 1, characterized in that, The harmonious image iterative generation includes sequentially executing the above-mentioned Step 1 to Step 3. After executing times, the final harmonious image is obtained. .

4. The method for harmonizing color deviation in image fusion according to any one of claims 1 to 3, characterized in that Use the following loss function to optimize the entire model: (7) Among them is the degradation estimation loss function, expressed as: (8) Among them represents the true degradation map, which is the degradation map output by the degradation estimation network; Represents a multi-scale deep supervision loss function, which can be expressed as follows: (9) It is characterized in that two scales of outputs are added to the segmentation network U-Net. The prediction result of the original size is output at the last level of the decoder of the segmentation network U-Net. A lightweight convolutional layer is connected to the second last and the third last levels respectively to output prediction results of different scales, which are the noise means and variances of 1 / 2 and 1 / 4 times of the original size respectively. Among them represents the level, and the maximum value is 3, indicating that the generated harmonized images are supervised at 3 different scales. When the value is 1, it represents the original size. When the s value is 2 and 3, they represent 1 / 2 and 1 / 4 times of the original size respectively. represents that the level is the true harmonized image at represents that the level is the harmonized image generated at , and MSE represents the mean square error loss. represents the per-pixel product. represents that the level is the mask at , and multiplying by means that the supervision is only carried out in the mask area.

Citation Information

Patent Citations

  • Real image denoising method based on multi-scale selection feedback network

    CN112927159A

  • Method and system for harmonizing synthetic image based on foreground reference image

    CN115205544A