A Pixel-Level Color Consistency Method and Device Based on Deep Learning
Through the pixel-level color consistency method based on deep learning, the image cutout and color correction are used to use the Unet network, attention mechanism and loop cutout mechanism to perform image cutout and color correction in the existing technology, and a more efficient image processing effect is achieved.
Patent Information
- Application Number
- CN202210589156.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-27
AI Technical Summary
The existing technology has shortcomings in image cutout and color correction. The cutout algorithm requires a lot of prior knowledge, the process is cumbersome, the operation efficiency is low, and it is difficult to balance the cutout accuracy and receptive field; manual color correction requires learning complex usage techniques, and the color correction algorithm is limited in practicality, slow processing speed and unsatisfactory effect.
A pixel-level color consistency method based on deep learning is proposed. By acquiring the image dataset, using the Unet network, attention mechanism and circular cutout mechanism for cutting the image, the first cutout result is obtained, and local and overall color correction is performed based on this result to obtain the result image.
It reduces image feature loss, improves image edge detail information extraction ability, enhances color consistency of synthetic images, and improves cutout accuracy and color correction efficiency.
Smart Images

Figure CN114897739B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular, to a pixel-level color consistency method and device based on deep learning. Background Art
[0002] With the development of technologies such as virtual reality and augmented reality, the combination of virtual and real images has gradually come into the public eye. As one of the main problems in the technology of combining virtual and real, the realism of the synthesized image is a hot topic of research, and the color consistency between the inserted object and the original image is one of the important factors affecting the realism. Existing image color processing methods are mainly divided into two categories:
[0003] 1. Photo editing software: Photo editing software such as PS (Photoshop, photo shop) has powerful image processing functions, but for non-professionals, learning how to use it is a cumbersome process;
[0004] 2. Color correction algorithms:
[0005] a) Some people propose to extract the areas with normal colors from the image sequence and blend them into a high-quality image. Although this method has achieved certain success to a certain extent and can obtain a synthetic image with relatively ideal effects, it requires the image sequence to be input multiple times, requires a large amount of computation, and cannot be directly applied to a single image;
[0006] b) For the method of enhancing the darker parts of video colors, some people propose that first, use the sampled tone mapping curve to construct multiple image sequences for each video frame, and then gradually fuse the image sequences in a spatio-temporal perception manner to obtain an enhanced video, but this method is only applicable to the darker areas.
[0007] In particular, the steps of the existing two-way light estimation method in the prior art are mainly as follows:
[0008] 1. Given an input image, perform two-way light estimation to obtain underexposed areas;
[0009] 2. Invert the image to obtain overexposed areas;
[0010] 3. Perform exposure correction on the obtained overexposed areas and underexposed areas;
[0011] 4. Fuse the underexposure-corrected image, the normal-exposure image, and the overexposure-corrected image to obtain the final output.
[0012] Generally speaking, on the one hand, the existing technologies require mastering the usage skills of image editing software such as PS. On the other hand, most of the color correction algorithms in the existing technologies are used to correct the problems of excessive or insufficient colors. Therefore, their practicality has certain limitations. In addition, although there are some color correction algorithms that take both into account, these algorithms usually have slow processing speed and unsatisfactory effects.
[0013] In addition, for synthetic images, a matte extraction operation may be required first. The existing image matte extraction algorithms are mainly divided into traditional image matte extraction methods and deep learning-based methods:
[0014] a) Traditional image matte extraction methods, such as mathematical methods based on Bayes' formula, simplify the model by assuming that adjacent pixels and pixels with similar colors have similar opacities. However, this is an ill-posed problem with too many unknowns. Although a large number of variables can be reduced through simple color assumptions, more prior knowledge is required and the matte extraction process is relatively cumbersome.
[0015] b) Deep learning-based methods learn the internal laws and representations of samples through a large amount of data to obtain an end-to-end mapping. Most of the existing methods are based on convolutional neural networks and have the following disadvantages when facing the matte extraction task: First, the overlapping parts of the image patches obtained at adjacent pixel points of the image are too many, resulting in redundant network parameters, wasting computing resources and affecting the operation efficiency. Second, in order to expand the receptive field, the accuracy has to be sacrificed. The downsampling operation will ignore some pixels while expanding the receptive field, resulting in the loss of some image detail information, ultimately leading to a decrease in matte extraction accuracy. It has become an insoluble problem to improve both accuracy and expand the receptive field at the same time.
[0016] In particular, for the matte extraction network model based on multi-scale channel attention, the main steps of its model are as follows:
[0017] 1. The input is a joint input of an RGB (Red Green Blue) image and a single-channel trimap, which is then convolved and normalized.
[0018] 2. The obtained image is downsampled to obtain a feature map, which is then fused with the low-level spatial features extracted by the ASPP (atrous spatial pyramid pooling) module.
[0019] 3. The obtained fusion result passes through the channel domain attention mechanism.
[0020] 4. The input encoding of the channel domain attention mechanism is input into the decoder to obtain the final result.
[0021] Generally speaking, the evaluation of matte extraction algorithms usually includes four criteria: mean square error, sum of absolute differences, connectivity error, and gradient error. However, as can be seen from the above methods, the existing technologies are difficult to meet the requirements. On the one hand, the existing technologies require a large amount of prior knowledge and the matte extraction process is cumbersome. On the other hand, the existing technologies have low operating efficiency, and it is difficult to balance the matte extraction accuracy and the receptive field.
[0022] In summary, in the existing technologies, there are deficiencies in both matte extraction algorithms and color correction in terms of image color consistency: Matte extraction algorithms require a large amount of prior knowledge, the process is cumbersome, and the operating efficiency is low. It is difficult to balance the matte extraction accuracy and the receptive field; Manual color correction requires learning complex usage skills of image editing software, the practicality of color correction algorithms is limited, the processing speed is slow, and the effect is not ideal. The deficiencies in matte extraction algorithms and color correction indicate that the problem of image color consistency still needs to be further studied.
[0023] Explanation of related terms:
[0024] (1) Color consistency: The problem of color consistency refers to the situation in an image where when a new object that originally does not belong to the image scene is incorporated, there will be a difference in color between the surface of the incorporated object and the scene color. In order to improve the realism of the synthesized image, the color difference between the newly added image and the original scene is corrected to make the final image more harmonious and more in line with the perception of the human eye.
[0025] (2) Convolutional Neural Network (CNN): A convolutional neural network is a type of feedforward neural network that contains convolutional computations and has a deep structure. It is one of the representative algorithms of deep learning. A convolutional neural network has the ability of representation learning and can perform shift-invariant classification on input information according to its hierarchical structure. Therefore, it is also called "Shift-Invariant Artificial Neural Networks (SIANN)".
[0026] (3) Long Short-Term Memory (LSTM): A long short-term memory network is a type of recurrent neural network designed specifically to solve the problem of long-term dependencies existing in general RNNs (Recurrent Neural Networks). All RNNs have a chain-like form of repeating neural network modules.
[0027] (4) Attention mechanism: The attention mechanism originates from the study of human vision. In cognitive science, due to the bottleneck of information processing, humans selectively focus on part of all information while ignoring other visible information. The above mechanism is usually called the attention mechanism. The deep learning attention mechanism is a bionic imitation of the human visual attention mechanism, which is essentially a resource allocation mechanism. The physiological principle is that human visual attention can be received at a certain area on the image with high resolution, and its surrounding area can be perceived at low resolution, and the viewpoint can change over time.
[0028] (5) Unet: The shallower convolution blocks in the Unet network structure model are used to extract low-level features, such as color and position information within the image. The deeper convolution blocks can extract higher-level semantic features, such as pixel category information. It is a typical Encoder-Decoder structure, and the encoding structure and decoding structure are symmetrical, shaped like the English letter U, so it is called the U-Net network structure.
[0029] (6) Encoder-Decoder Model: Encoder-Decoder is not a specific model, but a general framework. Encoding is to transform the input sequence into a fixed-length vector through a certain model, and decoding is to transform the previously generated fixed vector into an output sequence.
[0030] Application Contents
[0031] (Technical Problems + Technical Solutions + Technical Effects)
[0032] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0033] To this end, the purpose of this application is to address the shortcomings in graphic color consistency and propose a pixel-level color consistency method based on deep learning.
[0034] Another object of the present application is to propose a pixel-level color consistency device based on deep learning.
[0035] To achieve the above objectives, the present application proposes a pixel-level color consistency method based on deep learning, comprising the following steps:
[0036] Acquire an image data set, wherein the image data set includes an RGB three-channel image and a three-dimensional image;
[0037] Perform matte extraction operations on the RGB three-channel image and the trimap using a matte extraction model to obtain a first matte extraction result, where the matte extraction model includes a Unet (image segmentation network) network, an attention mechanism, and a cyclic matte extraction mechanism;
[0038] Perform color correction based on the first matte extraction result to obtain a result image.
[0039] In some possible embodiments, the performing matte extraction operations on the RGB three-channel image and the trimap using a matte extraction model to obtain a first matte extraction result includes:
[0040] Input the RGB three-channel image and the trimap into the encoder of the Unet network to output a first feature image;
[0041] Perform atrous pooling on the RGB three-channel image and the trimap to obtain an atrous pooling image;
[0042] Synthesize the first feature image and the atrous pooling image to obtain a synthesized image;
[0043] Perform a convolution operation on the synthesized image through the attention mechanism to obtain a second feature image;
[0044] Input the second feature image into the decoder of the Unet network to output a second matte extraction result;
[0045] Perform cyclic matte extraction on the second matte extraction result through the cyclic matte extraction mechanism to obtain the first matte extraction result.
[0046] In some possible embodiments, the attention mechanism includes a channel domain attention mechanism and a spatial domain attention mechanism. The performing a convolution operation on the synthesized image through the attention mechanism to obtain a second feature image includes:
[0047] Obtain a channel domain attention mechanism weight matrix and a spatial domain attention weight matrix;
[0048] Perform a convolution operation on the synthesized image based on the channel domain attention mechanism weight matrix and the spatial domain attention weight matrix to obtain the second feature image.
[0049] In some possible embodiments, the performing cyclic matte extraction on the second matte extraction result through the cyclic matte extraction mechanism to obtain the first matte extraction result includes:
[0050] Calculate the matte extraction score of the second matte extraction result;
[0051] In the case where the matte extraction score is less than the preset matte extraction score threshold, repeat the matte extraction operation based on the second matte extraction result;
[0052] When the number of repeated executions reaches a preset quantity, the first matte result is obtained.
[0053] In some possible embodiments, the deep learning-based pixel-level color consistency method further includes:
[0054] When the matte score is greater than or equal to the preset matte score threshold, the first matte result is obtained.
[0055] In some possible embodiments, the color correction includes local color correction and overall color correction. Performing color correction based on the first matte result to obtain a result image includes:
[0056] Performing the local color correction on the first matte result to obtain a restored image;
[0057] Performing the overall color correction based on the restored image to obtain the result image.
[0058] In some possible embodiments, performing the local color correction on the first matte result to obtain a restored image includes:
[0059] Separating the first matte result to obtain a target image and a background image;
[0060] Calculating the average pixel intensity values of the target image and the background image respectively;
[0061] Inputting the average pixel intensity values into a first preset function to output relative pixel intensity values;
[0062] Inputting the relative pixel intensity values into a second preset function to output a first relative pixel intensity weight;
[0063] Synthesizing and restoring the target image and the background image based on the first relative pixel intensity weight to obtain the restored image.
[0064] In some possible embodiments, performing the overall color correction based on the restored image to obtain the result image includes:
[0065] Scanning the restored image to obtain the pixel intensity value of each pixel;
[0066] Comparing the pixel intensity value of each pixel with a preset pixel intensity threshold to obtain a comparison result;
[0067] Dividing the restored image into a color transition region, a color normal region, and a color deficiency region according to the comparison result;
[0068] Calculate the second relative pixel intensity weights of the color over - area, the color normal - area, and the color lacking - area;
[0069] Perform color correction on the restored image based on the second relative pixel intensity weights and a preset formula to obtain the result image.
[0070] In some possible embodiments, calculating the matte score of the second matte result includes:
[0071] Obtain the weights of the intersection - over - union and the pixel intensity values;
[0072] Calculate the intersection - over - union score and the pixel intensity value score of the second matte result;
[0073] Calculate the matte score of the second matte result according to the weights of the intersection - over - union and the pixel intensity values, the intersection - over - union score and the pixel intensity value score of the second matte result.
[0074] To achieve the above object, another aspect of the present application provides a pixel - level color consistency device based on deep learning, including:
[0075] An acquisition module, configured to acquire an image data set, where the image data set includes RGB three - channel images and trimaps;
[0076] A matte module, configured to perform matte operations on the RGB three - channel images and the trimaps in the image data set through a matte model, to obtain a first matte result, where the matte model includes a Unet network, an attention mechanism, and a cyclic matte mechanism;
[0077] A correction module, configured to perform color correction based on the first matte result to obtain a result image.
[0078] Advantages of the present application:
[0079] According to a pixel - level color consistency method based on deep learning in an embodiment of the present application, by acquiring an image data set, performing matte operations on the RGB three - channel images and the trimaps in the image data set through a matte model to obtain a first matte result, and performing color correction based on the first matte result to obtain a result image. The present application can reduce the loss of image features, improve the ability to extract image edge detail information, and enhance the color consistency of the synthesized image.
[0080] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0081] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0082] Figure 1 FIG. is a flowchart of a pixel-level color consistency method based on deep learning according to an embodiment of the present application;
[0083] Figure 2 FIG. is an illustrative diagram of a pixel-level color consistency method based on deep learning according to an embodiment of the present application;
[0084] Figure 3 FIG. is a structural diagram of a pixel-level color consistency device based on deep learning according to an embodiment of the present application. Detailed Embodiments
[0085] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0086] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0087] A pixel-level color consistency method and device based on deep learning according to an embodiment of the present application will be described below with reference to the accompanying drawings. First, a pixel-level color consistency method based on deep learning according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0088] Figure 1 FIG. is a flowchart of a pixel-level color consistency method based on deep learning according to an embodiment of the present application.
[0089] As Figure 1 shown, the pixel-level color consistency method based on deep learning includes the following steps:
[0090] Step S110, obtaining an image data set, where the image data set includes RGB three-channel images and tripartite graphs.
[0091] In the embodiments of the present application, the image dataset can be obtained based on the Adobe image matting dataset. The Adobe image matting dataset includes many details such as hair, and can better reflect the characteristics of the images obtained in daily life. By synthesizing the foreground images in the Adobe image matting dataset, an RGB three-channel image and a trimap are obtained. Data augmentation can also be performed on the RGB three-channel image and the trimap, including rotation, scaling, and cropping.
[0092] Step S120: Perform a matting operation on the RGB three-channel image and the trimap through a matting model to obtain a first matting result.
[0093] Among them, the matting model includes a Unet network, an attention mechanism, and a cyclic matting mechanism. The Unet network is used to extract low-level features of the image. The attention mechanism is used to receive a certain area on the image at a high resolution and its surrounding area at a low resolution. The high-resolution receiving area can change over time. The cyclic matting mechanism is used to perform repeated matting on the matting result to improve the matting accuracy. The first matting result can be the matting result obtained by performing a matting operation on the RGB three-channel image and the trimap through the matting model.
[0094] In the embodiments of the present application, the matting model includes a Unet network, an attention mechanism, and a cyclic matting mechanism. The RGB three-channel image and the trimap are input into the matting model for matting operation to obtain a first matting result. The Unet network is the backbone network of the matting model and is a typical Encoder-Decoder structure with skip connection ability, which can reduce the loss of image features during the matting process. There are many other types of networks that can be used as the backbone network of the matting model, such as image segmentation networks like VGG and AlexNet. Adding an attention mechanism on the basis of the Unet network can enhance the attention of the matting model to regions such as edges and hair. The cyclic matting mechanism can obtain a higher-precision matting effect by connecting multiple Unet networks in series, and an attention mechanism is added to each Unet network.
[0095] Step S130: Perform color correction based on the first matting result to obtain a result image.
[0096] In the embodiments of the present application, after obtaining the first matting result, color correction can be performed based on the first matting result to improve the color consistency of the image and obtain a result image with improved color consistency.
[0097] A pixel-level color consistency method based on deep learning according to an embodiment of the present application obtains an image data set, performs matte extraction operations on the RGB three-channel image and the trimap in the image data set through a matte extraction model to obtain a first matte extraction result, and performs color correction based on the first matte extraction result to obtain a result image. The present application can reduce the loss of image features, improve the ability to extract edge detail information of the image, and enhance the color consistency of the synthesized image.
[0098] In some possible embodiments, performing matte extraction operations on the RGB three-channel image and the trimap through a matte extraction model to obtain a first matte extraction result includes:
[0099] Input the RGB three-channel image and the trimap into the encoder of the Unet network to output a first feature image;
[0100] Perform atrous pooling on the RGB three-channel image and the trimap to obtain an atrous pooling image;
[0101] Synthesize the first feature image and the atrous pooling image to obtain a synthesized image;
[0102] Perform convolutional operations on the synthesized image through an attention mechanism to obtain a second feature image;
[0103] Input the second feature image into the decoder of the Unet network to output a second matte extraction result;
[0104] Perform iterative matte extraction on the second matte extraction result through an iterative matte extraction mechanism to obtain a first matte extraction result.
[0105] Among them, the first feature image is a feature image obtained by encoding the RGB three-channel image and the trimap through the encoder of the Unet network, the second feature image is a feature image obtained by performing convolutional operations on the synthesized image of the first feature image and the atrous pooling image through an attention mechanism, and the second matte extraction result is a matte extraction result obtained by decoding the second feature image through the decoder of the Unet network.
[0106] In the embodiment of the present application, the number of input channels of the matte extraction model is set to four, and the RGB three-channel image and the trimap can be input into the encoder of the Unet network for encoding according to four input channels to obtain a first feature image.
[0107] At the same time, perform atrous pooling on the RGB three-channel image and the trimap to obtain an atrous pooling image. On the basis of atrous pooling, the RGB three-channel image and the trimap can be scaled according to any ratio. The purpose of this is to expand the receptive field of the network neurons. The process of atrous pooling can be expressed by the following formula:
[0108] Y = C 3,6 (P) + C 3,12(P)+C 3,18 (P)+C 3,24 (P)
[0109] Among them, C k,d (X), k represents the convolution kernel size, d represents the dilation rate. The convolution kernel size can uniformly adopt 3*3, and the dilation rate can adopt four values d = {6, 12, 18, 24}.
[0110] Then, the first feature image and the dilated pooling image are combined to obtain a combined image. Then, a convolution operation can be performed on the combined image through an attention mechanism to obtain a second feature image. Next, the second feature image can be input into the decoder of the Unet network to output a second matte result. Finally, a cyclic matte mechanism can perform cyclic matting on the second matte result to obtain a first matte result.
[0111] In some possible embodiments, the attention mechanism includes a channel-domain attention mechanism and a spatial-domain attention mechanism. Performing a convolution operation on the combined image through the attention mechanism to obtain a second feature image includes:
[0112] Obtaining a channel-domain attention mechanism weight matrix and a spatial-domain attention weight matrix;
[0113] Performing a convolution operation on the combined image based on the channel-domain attention mechanism weight matrix and the spatial-domain attention weight matrix to obtain a second feature image.
[0114] In the embodiments of the present application, the attention mechanism includes a channel-domain attention mechanism and a spatial-domain attention mechanism. Considering that different channels have different importance for the result image, different weights can be set for different channels. At the same time, in the spatial-domain attention mechanism, the weight of the matte target edge can be increased, so that the attention mechanism can be represented as a weight matrix. By obtaining the channel-domain attention mechanism weight matrix and the spatial-domain attention weight matrix, and then performing a convolution operation on the combined image based on the channel-domain attention mechanism weight matrix and the spatial-domain attention weight matrix, a second feature image can be obtained. The above process can be simply described by the following formula:
[0115] output = W CA ·W SA ·input
[0116] Among them, W CA represents the channel-domain attention mechanism weight matrix, W SA represents the spatial-domain channel attention mechanism weight matrix, input represents the input combined image, and output represents the output second feature image.
[0117] In some possible embodiments, performing cyclic matting on the second matting result through a cyclic matting mechanism to obtain a first matting result includes:
[0118] Calculating the matting score of the second matting result;
[0119] In the case where the matting score is less than a preset matting score threshold, repeatedly performing the matting operation based on the second matting result;
[0120] In the case where the number of repeated executions reaches a preset number, obtaining the first matting result.
[0121] Among them, the preset matting score threshold can be a preset matting score, such as 0.9, and the preset number can be the upper limit of the preset number of repeated executions. The preset number can be limited by the number of cascaded Unet networks. For example, 3 cascaded Unet networks are used to perform the cyclic matting mechanism.
[0122] In the embodiments of the present application, after inputting the second feature image into the decoder of the Unet network and outputting the second matting result, the matting score of the second matting result can be calculated. In the case where the matting score is less than the preset matting score threshold, the matting operation can be repeatedly performed based on the second matting result, that is, inputting the second matting result into the next cascaded Unet network to continue the matting operation, obtaining the matting result of the next round, updating it as the second matting result. In the case where the number of repeated executions reaches the preset number, that is, the second matting result completes the matting operation in all cascaded Unet networks, the matting result of the last cascaded Unet network is output as the first matting result to the next step.
[0123] In some possible embodiments, the pixel-level color consistency method based on deep learning further includes:
[0124] In the case where the matting score is greater than or equal to the preset matting score threshold, obtaining the first matting result.
[0125] In the embodiments of the present application, before the number of repeated executions reaches the preset number, when the matting score is greater than or equal to the preset matting score threshold, the matting result of the current situation can also be output as the first matting result to the next step and the current cyclic matting operation is terminated.
[0126] In some possible embodiments, color correction includes local color correction and overall color correction. Performing color correction based on the first matting result to obtain a result image includes:
[0127] Performing local color correction on the first matting result to obtain a restored image;
[0128] Performing overall color correction based on the restored image to obtain a result image.
[0129] In the embodiments of the present application, color correction includes local color correction and overall color correction. After obtaining the first matte result, local color correction can be performed on the first matte result to obtain a restored image, and then overall color correction can be performed based on the restored image to obtain a result image, which can be output as the final color consistency processing image.
[0130] In some possible embodiments, performing local color correction on the first matte result to obtain a restored image includes:
[0131] Separating the first matte result to obtain a target image and a background image;
[0132] Calculating the average pixel intensity values of the target image and the background image respectively;
[0133] Inputting the average pixel intensity values into a first preset function to output relative pixel intensity values;
[0134] Inputting the relative pixel intensity values into a second preset function to output a first relative pixel intensity weight;
[0135] Synthesizing the restored target image and the background image based on the first relative pixel intensity weight to obtain a restored image.
[0136] Wherein, the first preset function is used to convert the average pixel intensity values into relative pixel intensity values, the second preset function is used to convert the relative pixel intensity values into the first relative pixel intensity weight, and the first relative pixel intensity weight is used to represent the pixel intensity weights occupied by the target image and the background image respectively.
[0137] In the embodiments of the present application, after obtaining the first matte effect, the first matte result can be separated to obtain a target image and a background image, and the average pixel intensity values of the target image and the background image are calculated respectively. The average pixel intensity value can be obtained by respectively accumulating the pixel intensity values of the target image and the background image and then dividing by the total number of pixels.
[0138] After calculating the average pixel intensity values of the target image and the background image, the average pixel intensity values can be input into the first preset function to output relative pixel intensity values. The first preset function can adopt the Sigmoid function (mapping curve function), and the function form is as follows:
[0139]
[0140] Wherein, x represents the input average pixel intensity value, and S(x) represents the output relative pixel intensity value.
[0141] After obtaining the relative pixel intensity value, the relative pixel intensity value can be input into a second preset function to output a first relative pixel intensity weight. The second preset function can adopt a subtraction form, specifically as follows:
[0142] w = 1 - S(x)
[0143] Among them, S(x) represents the output relative pixel intensity value, and w represents the output first relative pixel intensity weight.
[0144] After obtaining the first relative pixel intensity weight, the target image and the background image can be synthesized based on the first relative pixel intensity weight to obtain a restored image. That is to say, the restored image is obtained by synthesizing the target image and the background image based on the first relative pixel intensity weight.
[0145] In some possible embodiments, overall color correction is performed on the restored image to obtain a result image, including:
[0146] Scanning the restored image to obtain the pixel intensity value of each pixel;
[0147] Comparing the pixel intensity value of each pixel with a preset pixel intensity threshold to obtain a comparison result;
[0148] According to the comparison result, the restored image is divided into a color transition region, a color normal region, and a color deficiency region;
[0149] Calculating the second relative pixel intensity weights of the color transition region, the color normal region, and the color deficiency region;
[0150] Performing color correction on the restored image based on the second relative pixel intensity weight and a preset formula to obtain a result image.
[0151] Among them, the preset pixel intensity threshold can be a preset pixel intensity value used to measure the size of the image pixel intensity, and thus divide the image color strength region. The preset formula can be a formula preselected for performing color correction on the obtained restored image.
[0152] In the embodiments of the present application, after obtaining the restored image, the restored image can be scanned to obtain the pixel intensity value of each pixel. The pixel intensity value of each pixel is compared with a preset pixel intensity threshold to obtain a comparison result. The comparison result may be that the pixel intensity value is greater than the preset pixel intensity threshold, the pixel intensity value is equal to the preset pixel intensity threshold, or the pixel intensity value is less than the preset pixel intensity threshold. According to the comparison result, the restored image can be divided into a color over - region, a color normal region, and a color lacking region. The color over - region can be an image region where the pixel intensity value is greater than the preset pixel intensity threshold, the color normal region can be an image region where the pixel intensity value is equal to the preset pixel intensity threshold, and the color lacking region can be an image region where the pixel intensity value is less than the preset pixel intensity threshold. After dividing the color over - region, the color normal region, and the color lacking region, the second relative pixel intensity weights of the color over - region, the color normal region, and the color lacking region can be calculated. The calculation of the second relative pixel intensity weights can refer to the calculation of the first relative pixel intensity weights, which will not be elaborated here. Based on the second relative pixel intensity weights and a preset formula, the restored image can be color - corrected, and the color over - region, the color normal region, and the color lacking region can be added and fused according to the second relative pixel intensity weights of the color over - region, the color normal region, and the color lacking region to obtain a result image. Among them, the form of the preset formula can be as follows:
[0153]
[0154] Among them, (L - L′ p ) 2 means to make the result image as similar as possible to the initialization image. λ is a hyperparameter that balances the weights of the two before and after. The initialization image can be selected as the image of the maximum value of each RGB channel as the initialization image. is used to minimize the redundancy of the detailed parts of the image. W p represents the second relative pixel intensity weight. represents taking the partial derivative of L.
[0155] In some possible embodiments, calculating the matte score of the second matte result includes:
[0156] Obtaining the weight of the intersection - over - union and the pixel intensity value;
[0157] Calculating the intersection - over - union score and the pixel intensity value score of the second matte result;
[0158] Calculating the matte score of the second matte result according to the weight of the intersection - over - union and the pixel intensity value, the intersection - over - union score and the pixel intensity value score of the second matte result.
[0159] In the embodiment of the present application, after obtaining the second matte result, the weights of the intersection over union (IoU) and the pixel intensity value can be obtained, and the weights of the IoU and the pixel intensity value can be set in advance. It is also possible to calculate the IoU score and the pixel intensity value score of the second matte result, and the calculation methods of the IoU score and the pixel intensity value score can be as follows:
[0160]
[0161]
[0162] Among them, score IoU represents the IoU score, score p represents the pixel intensity value score, S i represents the area, which is used to determine the position of the matte object, P i represents the pixel intensity value of the region, which is used to optimize the pixels in the detailed region. Specifically, S object represents the area occupied by the target object prediction, S gt represents the actual area occupied by the target object, P object represents the predicted intensity of the target object pixels, P gt represents the actual pixel intensity of the target object.
[0163] After obtaining the weights of the IoU and the pixel intensity value and the IoU score and the pixel intensity value score of the second matte result, the matte score of the second matte result can be calculated according to the weights of the IoU and the pixel intensity value, the IoU score and the pixel intensity value score of the second matte result, and the calculation method can be as follows:
[0164] score = W IoU ·score IoU + W p ·sxore p
[0165] Among them, score IoU represents the IoU score, score p represents the pixel intensity value score,, w IoU represents the IoU weight, w p represents the pixel intensity value weight, and score represents the matte score of the second matte result.
[0166] It should be noted that Figure 2 is an illustrative diagram of a pixel-level color consistency method based on deep learning according to an embodiment of the present application, as Figure 2As shown in the figure, first, image input is performed. Preprocessing operations can be carried out on the foreground images in the image dataset, and then image matting can be performed. Based on the Unet network model, a channel domain attention mechanism and a spatial domain attention mechanism are added, and a step-by-step matting strategy is adopted. By calculating and comparing the relationship between the matting score and the threshold, the matting result is output. Then, local color correction is performed. By calculating the average pixel intensity of the matting result (the target object and the environment) respectively, it is converted into the corresponding weights for weighted fusion. After that, overall color correction is performed. The two-way color estimation method is adopted, considering both cases of color overabundance and color deficiency. By dividing the image into three regions with different situations, the average pixel intensity of the corresponding regions is calculated and weighted fusion is performed to obtain the final correction result. Finally, image output is performed, that is, the correction result is output.
[0167] It should be noted that in the specific experimental operation, the matting model needs to be pre-trained to achieve a better matting effect. The training environment can use NVIDIA GTX2080Ti (graphics card model) for training, and the training parameters can be as shown in the following table:
[0168] Training parameters Parameter values Learning rate 1e-3 Number of training epochs 150 Number of samples per training 64 Matting score threshold 0.9 Number of repetitions of matting operation 3
[0169] The experimental results show that after local color correction, the obtained effect diagram has been improved to some extent. However, if overall color correction is added, compared with not adding overall color correction, the final result is closer to the human eye's habit, which reflects the necessity of overall color correction.
[0170] To implement the above embodiments, as Figure 3 shown, in this embodiment, a pixel-level color consistency device 300 based on deep learning is also provided. The device 300 includes: an acquisition module 310, a matting module 320, and a correction module 330.
[0171] The acquisition module 310 is used to acquire an image dataset, where the image dataset includes RGB three-channel images and trimaps;
[0172] The matting module 320 is used to perform matting operations on the RGB three-channel images and trimaps through a matting model to obtain a first matting result, where the matting model includes a Unet network, an attention mechanism, and a cyclic matting mechanism;
[0173] The correction module 330 is used to perform color correction based on the first matting result to obtain a result image.
[0174] A pixel-level color consistency device based on deep learning according to an embodiment of the present application obtains an image data set, performs matte extraction operations on the RGB three-channel images and trimaps in the image data set through a matte extraction model to obtain a first matte extraction result, and performs color correction based on the first matte extraction result to obtain a result image. The present application can reduce the loss of image features, improve the ability to extract image edge detail information, and enhance the color consistency of the synthesized image.
[0175] It should be noted that the foregoing explanations of the embodiments of the pixel-level color consistency method based on deep learning also apply to the pixel-level color consistency device based on deep learning in this embodiment, and will not be repeated here.
[0176] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0177] In the description of this specification, the descriptions with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0178] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A pixel-level color consistency method based on deep learning, characterized in that, it includes: Obtain an image dataset, where the image dataset includes RGB three-channel images and trimaps; Perform matte extraction operations on the RGB three-channel images and the trimaps through a matte extraction model to obtain a first matte extraction result, where the matte extraction model includes a Unet network, an attention mechanism, and a cyclic matte extraction mechanism; Perform color correction based on the first matte extraction result to obtain a result image, and the color correction refers to correcting the color difference between the first matte extraction result and the original scene; The step of performing matte extraction operations on the RGB three-channel images and the trimaps through the matte extraction model to obtain a first matte extraction result includes: Input the RGB three-channel images and the trimaps into the encoder of the Unet network to output a first feature image; Perform atrous pooling on the RGB three-channel images and the trimaps to obtain an atrous pooling image; Synthesize the first feature image and the atrous pooling image to obtain a synthesized image; Perform a convolution operation on the synthesized image through the attention mechanism to obtain a second feature image; Input the second feature image into the decoder of the Unet network to output a second matte extraction result; Perform cyclic matte extraction on the second matte extraction result through the cyclic matte extraction mechanism to obtain the first matte extraction result; The color correction includes local color correction and overall color correction, and the step of performing color correction based on the first matte extraction result to obtain a result image includes: Perform the local color correction on the first matte extraction result to obtain a restored image; Perform the overall color correction based on the restored image to obtain the result image.
2. The method according to claim 1, characterized in that, the attention mechanism includes a channel domain attention mechanism and a spatial domain attention mechanism, and the step of performing a convolution operation on the synthesized image through the attention mechanism to obtain a second feature image includes: Obtain a channel domain attention mechanism weight matrix and a spatial domain attention weight matrix; Perform a convolution operation on the synthesized image based on the channel domain attention mechanism weight matrix and the spatial domain attention weight matrix to obtain the second feature image.
3. The method according to claim 1, characterized in that, the step of performing cyclic matte extraction on the second matte extraction result through the cyclic matte extraction mechanism to obtain the first matte extraction result includes: Calculate the matte extraction score of the second matte extraction result; In the case where the matte extraction score is less than a preset matte extraction score threshold, repeatedly perform the matte extraction operation based on the second matte extraction result; In the case where the number of repeated executions reaches a preset number, obtain the first matte extraction result.
4. The method according to claim 3, characterized in that, it further includes: In the case where the matte extraction score is greater than or equal to the preset matte extraction score threshold, obtain the first matte extraction result.
5. The method according to claim 4, characterized in that, the step of performing the local color correction on the first matte extraction result to obtain a restored image includes: Separate the first matte extraction result to obtain a target image and a background image; Calculate the average pixel intensity values of the target image and the background image respectively; Input the average pixel intensity values into a first preset function to output relative pixel intensity values; Input the relative pixel intensity values into a second preset function to output a first relative pixel intensity weight; Based on the first relative pixel intensity weight, synthesize and restore the target image and the background image to obtain the restored image.
6. The method according to claim 5, wherein, the performing the overall color correction based on the restored image to obtain the result image includes: Scan the restored image to obtain the pixel intensity value of each pixel; Compare the pixel intensity value of each pixel with a preset pixel intensity threshold to obtain a comparison result; According to the comparison result, divide the restored image into a color over - region, a color normal region, and a color lack region; Calculate the second relative pixel intensity weights of the color over - region, the color normal region, and the color lack region; Based on the second relative pixel intensity weights and a preset formula, perform color correction on the restored image to obtain the result image.
7. The method according to claim 3, wherein, the calculating the matte score of the second matte result includes: Obtain the weights of the intersection - over - union and the pixel intensity value; Calculate the intersection - over - union score and the pixel intensity value score of the second matte result; Calculate the matte score of the second matte result according to the intersection - over - union and the weights of the pixel intensity value, the intersection - over - union score and the pixel intensity value score of the second matte result.
8. A pixel - level color consistency device based on deep learning, wherein, it includes: An acquisition module, configured to acquire an image data set, where the image data set includes RGB three - channel images and trimaps; A matte module, configured to perform a matte operation on the RGB three - channel images and the trimaps through a matte model to obtain a first matte result, where the matte model includes a Unet network, an attention mechanism, and a cyclic matte mechanism; A correction module, configured to perform color correction based on the first matte result to obtain a result image, and the color correction refers to correcting the color difference between the first matte result and the original scene; The performing a matte operation on the RGB three - channel images and the trimaps through a matte model to obtain a first matte result includes: Input the RGB three - channel images and the trimaps into the encoder of the Unet network to output a first feature image; Perform atrous pooling on the RGB three - channel images and the trimaps to obtain an atrous pooling image; Synthesize the first feature image and the atrous pooling image to obtain a synthesized image; Perform a convolution operation on the synthesized image through the attention mechanism to obtain a second feature image; Input the second feature image into the decoder of the Unet network to output a second matte result; Perform cyclic matting on the second matte result through the cyclic matte mechanism to obtain the first matte result; The color correction includes local color correction and overall color correction. Performing color correction based on the first matte result to obtain a result image includes: Performing the local color correction on the first matte result to obtain a restored image; Performing the overall color correction based on the restored image to obtain the result image.
Citation Information
Patent Citations
Automatic matting system, method and device
CN110400323A
Single image input self-supervision matting model training method, matting method and device
CN114119639A