Coloring model training method based on color transformation
Through the color transformation-based shading model training method, the clustering algorithm and weight attention map are used to solve the problem of color transformation destroying natural coloring in unconditional areas, realizing the accuracy of local conditional coloring and automatic generation of natural laws, and improving the efficiency of image coloring repair.
Patent Information
- Application Number
- CN202510190175.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-18
AI Technical Summary
During the training process of existing shading models, color transformation will destroy the natural shading ability of the unconditional area of the image, resulting in large errors in the model prediction results and affecting the training effect.
The color transformation-based shading model training method is adopted, and subpixel segmentation is performed through the clustering algorithm, prompt points are randomly generated, distance diffusion map and color similarity matrix are calculated, and model training is performed in combination with the weight attention map, which is divided into two-stage training process. The first stage does not perform color transformation, and the second stage adjusts the target loss function with the weight attention map.
It improves the accuracy of the model's coloring of conditional areas, while retaining the natural coloring ability of unconditional areas, improving the efficiency and effect of image coloring and repair.
Smart Images

Figure CN120339424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of coloring model training, and specifically to a coloring model training method based on color transformation. Background Technique
[0002] Image conditional coloring based on point prompts deterministically colors local regions of an image based on the color information of the prompt points, and colors other object regions without prompt points according to the learned natural color distribution law. This technology has wide applications and important significance, such as coloring old films. In the actual practical process, we may hope that the upper garment of a person in the image is red and the pants are green. For other background regions or other objects, we hope that they can be automatically colored and their colors conform to the natural law, avoiding the need for manual prompts for all regions.
[0003] Currently, most coloring models still focus on how to optimize the model through a stronger network structure, such as the recently popular transformers and diffusion models, and how to make the model support multi-modal input.
[0004] Method 1 hopes to rely on the good receptive field of the transformer to improve the diffusion ability of local region colors. Refer to the reference "Yun J, Lee S, Park M, et al. iColoriT: Towards propagating local hints to the right region in interactive colorization by leveraging vision transformer[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2023: 1787-1796".
[0005] Method 2 realizes multi-modal input by converting multi-modal inputs into prompt points. Refer to the reference "Huang Z, Zhao N, Liao J. Unicolor: A unified framework for multi-modal colorization with transformer[J]. ACM Transactions on Graphics (TOG), 2022, 41(6): 1-16".
[0006] Method 3 relies on the diffusion model and embeds conditional information into the latent space to improve the model's capabilities. It also mentions adding color jump enhancement to indicate the model's dependence on conditions. Reference "Liang Z, Li Z, Zhou S, etal. Control Color: Multimodal Diffusion-based Interactive Image Colorization[J]. arXiv preprint arXiv:2402.10855, 2024".
[0007] Among them, the model's trust in the conditions is enhanced by adding image color transformation. Without color transformation, the color of the prompt point is originally a natural law. Only when the law is destroyed will the model tend to rely on the conditions. However, there is a problem here. When the image color is transformed, the unconditional coloring learning of the area without prompt points will be destroyed.
[0008] Method 4 performs local image coloring based on traditional machine learning methods. When training the conditional coloring model, these deep learning-based methods randomly select pixel points in some areas of the image color channel as prompt points, and then use the color information of the prompt points and the brightness channel of the image as model inputs, hoping that the model outputs the corresponding color channel, and finally performs network feedback learning by calculating the error between the color channel information predicted by the model and the color channel of the original image. It can be found that there are some areas in the input brightness channel without prompt points, that is, this is equivalent to unconditional coloring, relying on the model's learning of natural color laws, and for areas with prompt points, it is hoped that the color that can be predicted is consistent with the prompt points, which relies on the model's trust in the prompt points. The Chinese patent with publication number CN104851114B disclosed a method and terminal for realizing local image color change on May 1, 2018.
[0009] When training a model, current methods randomly change the color of the image to improve the generalization of the model or to increase the model's dependence on conditions. However, they do not consider that color change will destroy the image's unconditional natural coloring ability, resulting in the tendency of areas without conditions to randomly generate colors during training, rather than generating colors that conform to the natural distribution, that is, the colors corresponding to the labels. This will cause the model's prediction results to have large errors, affecting the final model training effect. Summary of the invention
[0010] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a coloring model training method based on color transformation, which can enable the model to focus on learning coloring conditions in areas with prompt points, and focus on learning natural color rules in areas without prompt points, without affecting each other.
[0011] According to one aspect of the specification of the present invention, a method for training a coloring model based on color transformation is provided, including:
[0012] Based on a clustering algorithm, perform sub-pixel segmentation on the original image to obtain an image segmentation map;
[0013] Randomly change the color of the original image to obtain a transformed image;
[0014] Randomly generate the number of hint points and their corresponding coordinates;
[0015] Based on the image segmentation map, calculate the places with the same pixels as the hint points to obtain a hint point region binary map, and combine it with the transformed image to obtain the hint point color;
[0016] Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change;
[0017] Convert the color space of the transformed image to obtain a Lab image, and separate the channels of the Lab image to obtain an L-channel image;
[0018] Based on the color of the hint points, calculate the similarity between other pixels to obtain a color similarity matrix, and obtain a weight attention map by multiplying the color similarity matrix by the distance diffusion map;
[0019] The model inputs the hint point coordinates, the corresponding hint point colors, and the L-channel image, and based on the target loss function, combines the weight attention map to perform model training.
[0020] As a further technical solution, randomly changing the color of the original image to obtain a transformed image includes: randomly transforming the color by arbitrarily adjusting the RGB values to obtain the transformed image.
[0021] As a further technical solution, based on the image segmentation map, calculating the places with the same pixel values as the hint points to obtain a hint point region binary map, and combining it with the transformed image to obtain the hint point color includes:
[0022] Extract the color value hint_value corresponding to the hint points on the image segmentation map, and then traverse all pixels on the image segmentation map. If the pixel value is equal to hint_value, it is the hint point region, and a hint point region binary map is obtained;
[0023] In the hint point region of the binary map, take the average value of the corresponding pixels on the color-transformed image to obtain the hint point color.
[0024] As a further technical solution, calculating the distance between the pixel coordinates of the image segmentation map and the hint points, and obtaining a distance diffusion map based on the distance change includes:
[0025] Traverse all pixel coordinates on the image segmentation map, calculate the Euclidean distance between all pixel coordinates and the hint points, perform Gaussian smoothing and normalization to obtain a distance diffusion map.
[0026] As a further technical solution, based on the color of the hint points, calculate the similarity with other pixels to obtain a color similarity matrix, including:
[0027] Convert the hint color and the transformed image to the Lab color space, separate the channels to obtain the L-channel image, compare the similarity of the ab channels, and obtain a similarity matrix map.
[0028] As a further technical solution, the model training is divided into two stages:
[0029] The first stage: During training, no color transformation is performed. The objective loss function during training is the robust regression loss. When the error predicted by the model is less than the set threshold, the squared loss function is used. If the error is greater than or equal to the set delta value, the absolute value loss function is used. When the training loss converges, the first stage ends;
[0030] The second stage: During training, randomly perform color transformation on some samples and do not perform color transformation on some samples. The unchanged samples still use the objective loss function of the first stage, and the changed samples will adjust the objective loss function in combination with the weighted attention map.
[0031] According to one aspect of the specification of the present invention, there is provided a method for training a coloring model based on color transformation, including:
[0032] The first main module is used to perform sub-pixel segmentation on the original image based on the clustering algorithm to obtain an image segmentation map;
[0033] The second main module is used to randomly perform color changes on the original image to obtain a transformed image;
[0034] The third main module is used to randomly generate the number of hint points and the corresponding coordinates;
[0035] The fourth main module is used to calculate the places where the pixels are the same as the hint points based on the image segmentation map to obtain a binary map of the hint point area, and combine it with the transformed image to obtain the hint point color;
[0036] The fifth main module calculates the distance between the pixel coordinates of the image segmentation map and the hint points, and obtains a distance diffusion map based on the distance change;
[0037] The sixth main module converts the transformed image to the Lab image, separates the channels of the Lab image to obtain the L-channel image;
[0038] The seventh main module calculates the similarity with other pixels based on the color of the hint points to obtain a color similarity matrix, and multiplies the color similarity matrix by the distance diffusion map to obtain a weight attention map.
[0039] The eighth main module combines the weight attention map with the input of the model, which includes the hint point coordinates, the corresponding hint point colors, and the L-channel image, for model training.
[0040] According to one aspect of the specification of the present invention, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the coloring model training method based on color transformation are implemented.
[0041] According to one aspect of the specification of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the coloring model training method based on color transformation.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] The local conditional coloring technology designed by the present invention can help users quickly color old photos or black-and-white photos. When coloring, colors can be specified for some local areas, and for areas where no color is specified, colors that conform to natural laws will also be automatically generated, avoiding the need for users to interactively color each local area. At the same time, due to the enhancement of color transformation, the accuracy of the conditional coloring area is improved, greatly improving the efficiency of users in coloring and restoring images. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0045] Figure 1 It is a schematic flowchart of a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0046] Figure 2 It is a schematic diagram of obtaining sub-pixel segmentation of a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0047] Figure 3 It is a schematic diagram of color transformation of a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0048] Figure 4 Schematic diagram of obtaining a binary map of a hint point region based on a sub-pixel segmentation map for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0049] Figure 5 Schematic diagram of obtaining a hint color based on a hint point region for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0050] Figure 6 Schematic diagram of obtaining a distance diffusion map based on distance change for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0051] Figure 7 Schematic diagram of RGB to LAB conversion in the color space for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0052] Figure 8 Schematic diagram of obtaining a similarity matrix map based on color similarity for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0053] Figure 9 Schematic diagram of obtaining a weight attention map for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0054] Figure 10 Schematic diagram of the model structure for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0055] Figure 11 Schematic diagram of a sample example during the first-stage training process for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0056] Figure 12 Schematic diagram of a sample example during the second-stage training process for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0057] Figure 13 Schematic diagram of Embodiment 1 for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0058] Figure 14 Schematic diagram of Embodiment 2 for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0059] Figure 15 Schematic diagram of Embodiment 3 for a coloring model training method based on color transformation provided by an embodiment of the present invention.
[0060] Figure 16Schematic diagram of Embodiment 4 of a colorization model training method based on color transformation provided by an embodiment of the present invention.
[0061] Figure 17 Schematic diagram of Embodiment 5 of a colorization model training method based on color transformation provided by an embodiment of the present invention.
[0062] Figure 18 Schematic diagram of a system for a colorization model training method based on color transformation provided by an embodiment of the present invention. Detailed implementation manners
[0063] Image conditional colorization is a classic and powerful technique in the field of image processing. It generally refers to coloring or color adjustment of images under specific conditions. This technique can be used for various tasks, such as image enhancement, style transfer, image editing, etc. By utilizing certain specific attributes of the input image (such as brightness, texture, structure, etc.), targeted color processing can be carried out, so that the final image presents a more satisfactory effect. For example, in medical image processing such as CT, MRI, and X-ray, conditional colorization can color different regions according to different pathological features (such as tumor regions, normal tissues, etc.) to help doctors better identify and analyze abnormalities. When a local area of a picture is damaged, conditional colorization is performed based on the information of the surrounding area to restore the original color or style of the image. For example, it is applied in scenarios such as old photo restoration and missing data restoration. In old photo restoration or digitalization of historical relics, conditional colorization technology is used to re-color black-and-white photos or damaged photos to restore their original colors. In advertising production, conditional colorization technology is used to color and adjust product pictures to make them more in line with the aesthetics of the target audience or the color standards of the brand. Image conditional colorization technology has penetrated into multiple industries and application fields, especially showing great potential in computer vision, medical imaging, autonomous driving, art creation, etc. With the development of deep learning and artificial intelligence technologies, the accuracy and efficiency of image conditional colorization are expected to be continuously improved, further expanding its application scope.
[0064] The terms "including" and "having" in the specification, claims and above-mentioned drawings of the present invention, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a new technical solution. This combination is not restricted by the order of steps and / or the structural composition mode, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0066] An embodiment of the present invention provides a coloring model training method based on color transformation, as Figure 1 shown, including:
[0067] Step 1: Based on the clustering algorithm, perform sub-pixel segmentation on the original image to obtain an image segmentation map;
[0068] Step 2: Randomly change the color of the original image to obtain a transformed image;
[0069] Step 3: Randomly generate the number of hint points and the corresponding coordinates;
[0070] Step 4: Based on the image segmentation map, calculate the places with the same pixels as the hint points to obtain a hint point region binary map, and combine it with the transformed image to obtain the hint point color;
[0071] Step 5: Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change;
[0072] Step 6: Convert the color space of the transformed image to obtain an Lab image, and perform channel separation on the Lab image to obtain an L-channel image;
[0073] Step 7: Based on the color of the hint points, calculate the similarity with other pixels to obtain a color similarity matrix, and multiply the color similarity matrix by the distance diffusion map to obtain a weight attention map;
[0074] Step 8: Input the hint point coordinates, the corresponding hint point colors, and the L-channel image into the model, and perform model training based on the target loss function in combination with the weight attention map.
[0075] As Figure 3As shown, in step 2, the original image is randomly color-changed to obtain the transformed image, including: randomly changing the color by arbitrarily adjusting the RGB (red, green, blue) values to obtain the transformed image.
[0076] In step 3: randomly generate the number of hint points n and the corresponding coordinates hints, and the functional relationship of hints is as follows;
[0077] hints = [(x1, y1), (x2, y2),... (xn, yn)]
[0078] Without special instructions in this article, it is assumed that n = 1 and only one hint point is taken.
[0079] As Figures 4 - 5 shown, in step 4, based on the image segmentation map, calculate the places with the same pixels as the hint points to obtain the hint point region binary map, and combine it with the transformed image to obtain the hint point color, including:
[0080] Extract the color value hint_value corresponding to the hint points on the image segmentation map, and then traverse all the pixels on the image segmentation map. If the pixel value is equal to hint_value, it is the hint point region, and the hint point region binary map is obtained;
[0081] In the hint point region of the binary map, take the average value of the corresponding pixels on the color-transformed image to obtain the hint point color.
[0082] Specifically, first extract the color value hint_value corresponding to the hint points on the sub-pixel segmentation image, then traverse all the pixels on the sub-pixel segmentation image. If the pixel value is equal to hint_value, it is the hint point region, and finally a binary map is obtained. The pixel value of the white region is 255, representing the hint point region.
[0083] Then in the hint point region of the previous binary map, take the average value of the corresponding pixels on the color-transformed image, which is the color information of the hint points. The color information and coordinates of the hint points construct the color condition.
[0084] As Figure 6 shown, in step 5, calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain the distance diffusion map based on the distance change, including:
[0085] Traverse all the pixel coordinates on the image segmentation map, calculate the Euclidean distance between all the pixel coordinates and the hint points, perform Gaussian smoothing and normalization to obtain the distance diffusion map.
[0086] Specifically, traverse all the pixel coordinates on the image, calculate their Euclidean distance from the hint points (x1, x2), and then perform Gaussian smoothing and normalization to obtain a distance diffusion map.
[0087] As Figures 7 - 9 shown, in step 6, based on the color of the hint point, calculate the similarity with other pixels to obtain a color similarity matrix, including:
[0088] Convert the hint color and the transformed image to the Lab color space, separate the channels to obtain the L-channel image, compare the similarity of the ab channels, and obtain the similarity matrix diagram.
[0089] Specifically, first convert the hint color and the input image to the Lab color space, and then calculate the weighted Euclidean distance between the pixel values at all positions of the image and the pixel values of the hint point in the Lab color space. The formula is as follows, where the range of m is from 0 to 1, and in this paper, m is taken as 0.7.
[0090] distance = (m * (L1 - Lk)^2 + (a1 - ak)^2 + (b1 - bk)^2)^0.5
[0091] Because the LAB color space is based on the human eye's perception of color and is a more advanced color representation method. The L channel represents the brightness value of the color, the a channel represents the color range from green to red, and the b channel represents the color range from blue to yellow. The image details are mainly in the L channel. Therefore, in order to ensure that the details of the original image are not affected, the coloring model generally only predicts the ab channels representing the color. So when calculating the color similarity in this paper, the similarity of the ab channels is considered first.
[0092] At the same time, by multiplying the color similarity matrix with the distance diffusion map, a weight attention map can be obtained, which takes into account both the color similarity and the influence of distance. This avoids only the object regions with color similarity being regarded as the target coloring regions. As shown in the figure, the rightmost image is the weight attention map.
[0093] Without additional conditions, the model will actually learn this internal weight attention by itself through a large amount of data, that is, it will find the regions with the same semantics as the hint point through the similarity of features, such as the chair region shown in this paper's example. However, this requires the network to have very high learning ability to ensure that the conditional color fully spreads in the target local region while avoiding color spillover, similar to image semantic segmentation. And the weight attention map designed in this paper is used as an additional condition to help the network learn better without increasing the time consumption in the inference stage. Most importantly, for the transformed image, the non-conditional coloring regions can be suppressed through this weight attention map to avoid the natural conditional coloring ability of the non-conditional coloring regions being damaged.
[0094] As Figure 10As shown, in step 7, during model training, the model input is the prompt point coordinates and the corresponding color information and the L channel image, and the model output is the predicted a and b channels. The network structure is constructed based on the transformer.
[0095] Model training is divided into two stages: No color transformation enhancement is performed, and all data is taken in natural scenes. This is because the model's ability focuses on natural color laws rather than conditions.
[0096] The first stage: During training, no color transformation is performed, and the target loss function during training is Huber Loss. In addition to being called "Huber loss", Huber loss can also be called Huber regression loss or robust regression loss. These names reflect its application in regression tasks and its robustness to outliers.
[0097] The formula is as follows: if the error predicted by the model is lower than the threshold k, the square loss function is used; if the error is greater than or equal to delta, delta=0.01, the absolute value loss function is used. This avoids the feedback of the absolute value loss function being too large and missing the local optimal solution, while also avoiding the sensitivity of the square loss function to outliers.
[0098]
[0099] like Figure 11 As shown in the figure, the input image is a blue sky and sea, the cue point is located in the chair area, and the cue color is light yellow. When the model predicts, the chair is still light yellow and the sky and sea are still blue.
[0100] The second stage: During training, some samples will be randomly transformed in color, while others will not. The samples that do not change still use the target loss function of the first stage, and the samples that change will adjust the target loss function in combination with the weighted attention map. That is, when calculating the error of each pixel, the value corresponding to the weighted attention map will be used as the weight, where w is the corresponding weight.
[0101]
[0102] like Figure 12 As shown in the figure, after the input image is transformed, the sky and the sea are green, the cue point is located in the chair area, and the cue color is yellow. When the model predicts, the chair remains yellow, while the sky and the sea are unconditionally colored areas. This is because during training, the learning of these areas is suppressed by the weighted attention map, so these areas will still retain the unconditional natural coloring ability instead of being predicted as green.
[0103] In summary of the above embodiments, the present invention optimizes the objective loss function during model training through a weight attention map, where the weight attention map is obtained by combining color similarity and distance diffusion. First, the sub-pixel segmentation maps of all images are obtained offline, and then the color information of the cue points is obtained according to the random cue point coordinates, the sub-pixel segmentation maps, and the input images. Next, a color similarity matrix is obtained based on the similarity between the colors of the cue points and other pixels. Additionally, a distance diffusion matrix is constructed based on the coordinate distances between the cue point coordinates and other pixels. Finally, the color similarity matrix and the distance diffusion matrix are multiplied in a certain proportion to obtain the final weight attention map. Through the weight attention map, the conditional regions can learn with greater weights, while the unconditional regions learn with smaller weights. This can not only enhance the generalization ability of the model through color transformation and improve the model's dependence on conditions, but also ensure that the coloring ability of the unconditional regions is not damaged. And a two-stage training method is adopted, that is, color enhancement is not performed in the first stage, and color enhancement is performed in the second stage.
[0104] As Figure 13 shown, in Embodiment 1, the cue conditions are: the cue point position is in the area of the first dog, and the color is blue;
[0105] As Figure 14 shown, in Embodiment 2, the cue conditions are: the cue point position is the second apple from the left, and the color is red;
[0106] As Figure 15 shown, in Embodiment 3, the cue conditions are: the cue point position is the upper garment of the person, and the color is light green;
[0107] As Figure 16 shown, in Embodiment 4, the cue conditions are: the cue point position is the body of the leopard, and the color is purple;
[0108] As Figure 17 shown, in Embodiment 5, the cue conditions are: the cue point position is the upper garment area, and the color is dark green.
[0109] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and their functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides a coloring model training method based on color transformation, and the system is used to execute a coloring model training method based on color transformation in the above method embodiments.
[0110] See Figure 18, the system includes: a first main module for sub-pixel segmentation of the original image based on a clustering algorithm to obtain an image segmentation map; a second main module for randomly changing the color of the original image to obtain a transformed image; a third main module for randomly generating the number of hint points and their corresponding coordinates; a fourth main module for calculating, based on the image segmentation map, the areas with the same pixels as the hint points to obtain a binary map of the hint point areas, and combining with the transformed image to obtain the hint point colors; a fifth main module for calculating the distance between the pixel coordinates of the image segmentation map and the hint points, and obtaining a distance diffusion map based on the distance change; a sixth main module for performing color space conversion on the transformed image to obtain an Lab image, and separating the channels of the Lab image to obtain an L-channel image; a seventh main module for calculating the similarity between other pixels based on the colors of the hint points to obtain a color similarity matrix, and obtaining a weighted attention map by multiplying the color similarity matrix with the distance diffusion map; an eighth main module for inputting the hint point coordinates, the corresponding hint point colors, and the L-channel image into the model, and combining with the weighted attention map to perform model training.
[0111] A coloring model training method based on color transformation provided by an embodiment of the present invention adopts Figure 18 several modules therein. When coloring, it can specify colors for some local areas. For areas without specification, it will also automatically generate colors that conform to natural laws, avoiding the need for users to interactively color each local area. At the same time, due to the enhancement of color transformation, the accuracy of the conditional coloring area is improved. The efficiency of users in coloring and repairing images is greatly improved.
[0112] It should be noted that the system embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting the corresponding functional modules. Its principle is basically the same as the principle of the above system embodiment provided by the present invention. As long as those skilled in the art, based on the above system embodiment, refer to the specific technical solutions in other method embodiments, obtain the corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicality of the technical solutions, improve the modules in the above system embodiment to obtain the corresponding system-like embodiments for implementing the methods in other method-like embodiments. For example:
[0113] Based on the content of the above system embodiment, as a preferred embodiment, a coloring model training method based on color transformation provided by an embodiment of the present invention randomly changes the color of the original image to obtain a transformed image, including: randomly transforming the color by arbitrarily adjusting the RGB (red, green, blue) values to obtain the transformed image, that is, the input image of the model during training.
[0114] Based on the content of the above system embodiments, as a preferred embodiment, a coloring model training method based on color transformation provided in an embodiment of the present invention, based on an image segmentation map, calculates the same places as the pixels of the hint points to obtain a binary map of the hint point region, and combines the transformed image to obtain the hint point color, including:
[0115] Extract the color value hint_value corresponding to the hint points on the image segmentation map, and then traverse all the pixels on the image segmentation map. If the pixel value is equal to hint_value, it is the hint point region, and a binary map of the hint point region is obtained;
[0116] In the hint point region of the binary map, take the average value of the corresponding pixels on the color-transformed image to obtain the hint point color.
[0117] Based on the content of the above system embodiments, as a preferred embodiment, in a coloring model training method based on color transformation provided in an embodiment of the present invention, calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change, including:
[0118] Traverse all the pixel coordinates on the image segmentation map, calculate the Euclidean distance between all the pixel coordinates and the hint points, perform Gaussian smoothing and normalization to obtain a distance diffusion map.
[0119] Based on the content of the above system embodiments, as a preferred embodiment, in a coloring model training method based on color transformation provided in an embodiment of the present invention, calculate the similarity between other pixels based on the color of the hint points to obtain a color similarity matrix, including:
[0120] Perform color space conversion on the hint color and the input image, convert it to the Lab space, separate the channels to obtain the L-channel image, compare the similarity of the ab channels, and obtain a similarity matrix map.
[0121] Based on the content of the above system embodiments, as a preferred embodiment, in a coloring model training method based on color transformation provided in an embodiment of the present invention, the model training is divided into two stages:
[0122] The first stage: During training, no color transformation is performed. The target loss function during training is HuberLoss. If the error predicted by the model is less than the set threshold, the square loss function is used. If the error is greater than or equal to the set delta value, the absolute value loss function is used. When the training loss converges, the first stage ends;
[0123] The first stage: During training, randomly perform color transformation on some samples and do not perform color transformation on some samples. The unchanged samples still use the target loss function of the first stage, and the changed samples will adjust the target loss function in combination with the weight attention map.
[0124] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a color transformation-based coloring model training method are implemented as follows:
[0125] Based on a clustering algorithm, perform sub-pixel segmentation on the original image to obtain an image segmentation map;
[0126] Randomly change the color of the original image to obtain a transformed image;
[0127] Randomly generate the number of hint points and the corresponding coordinates;
[0128] Based on the image segmentation map, calculate the places where the pixels are the same as the hint points to obtain a hint point region binary map, and combine it with the transformed image to obtain the hint point color;
[0129] Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change;
[0130] Convert the color space of the transformed image to obtain an Lab image, and perform channel separation on the Lab image to obtain an L-channel image;
[0131] Based on the color of the hint points, calculate the similarity between other pixels to obtain a color similarity matrix, and multiply the color similarity matrix by the distance diffusion map to obtain a weight attention map;
[0132] The model inputs the hint point coordinates, the corresponding hint point colors, and the L-channel image, and based on the target loss function, combines the weight attention map to perform model training.
[0133] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the steps of a color transformation-based coloring model training method executed by the computer are as follows:
[0134] Based on a clustering algorithm, perform sub-pixel segmentation on the original image to obtain an image segmentation map;
[0135] Randomly change the color of the original image to obtain a transformed image;
[0136] Randomly generate the number of hint points and the corresponding coordinates;
[0137] Based on the image segmentation map, calculate the places where the pixels are the same as the hint points to obtain a hint point region binary map, and combine it with the transformed image to obtain the hint point color;
[0138] Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain the distance diffusion map based on the distance change;
[0139] Convert the transformed image to a color space to obtain an Lab image, and separate the channels of the Lab image to obtain an L-channel image;
[0140] Calculate the similarity between other pixels based on the color of the hint points to obtain a color similarity matrix, and obtain a weighted attention map by multiplying the color similarity matrix by the distance diffusion map;
[0141] The model inputs the hint point coordinates, the corresponding hint point colors, and the L-channel image, and based on the target loss function, combines the weighted attention map to perform model training.
[0142] In summary, the local conditional coloring technology designed by the present invention can help users quickly color old photos or black-and-white photos. When coloring, users can specify colors for some local areas. For areas without specification, colors that conform to natural laws will also be automatically generated, avoiding the need for users to interactively color each local area. At the same time, due to the enhancement of color transformation, the accuracy of the conditional coloring area is improved. The efficiency of users in coloring and restoring images is greatly improved.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for training a coloring model based on color transformation, characterized in that Including: Based on the clustering algorithm, perform sub-pixel segmentation on the original image to obtain an image segmentation map; Randomly change the color of the original image to obtain a transformed image; Randomly generate the number of hint points and their corresponding coordinates; Based on the image segmentation map, calculate the areas with the same pixel values as the hint points to obtain a binary map of the hint point regions, and combine it with the transformed image to obtain the hint point colors; Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change; Convert the color space of the transformed image to obtain an Lab image, and separate the channels of the Lab image to obtain an L-channel image; Based on the colors of the hint points, calculate the similarity between other pixels to obtain a color similarity matrix, and multiply the color similarity matrix by the distance diffusion map to obtain a weighted attention map; The model inputs the hint point coordinates, the corresponding hint point colors, and the L-channel image, and based on the target loss function, combines the weighted attention map to perform model training.
2. The coloring model training method based on color transformation according to claim 1, wherein Randomly change the color of the original image to obtain a transformed image, including: randomly changing the color by arbitrarily adjusting the RGB values to obtain the transformed image.
3. The coloring model training method based on color transformation according to claim 2, wherein Based on the image segmentation map, calculate the areas with the same pixel values as the hint points to obtain a binary map of the hint point regions, and combine it with the transformed image to obtain the hint point colors, including: Extract the color value hint_value corresponding to the hint points on the image segmentation map, and then traverse all the pixels on the image segmentation map. If the pixel value is equal to hint_value, it is the hint point region, and a binary map of the hint point regions is obtained; In the hint point region of the binary map, take the average value of the corresponding pixels on the color-transformed image to obtain the hint point colors.
4. The coloring model training method based on color transformation according to claim 3, characterized in that, Calculate the distance between the pixel coordinates of the image segmentation map and the hint points, and obtain a distance diffusion map based on the distance change, including: Traverse all the pixel coordinates on the image segmentation map, calculate the Euclidean distance between all the pixel coordinates and the hint points, perform Gaussian smoothing and normalization to obtain the distance diffusion map.
5. The coloring model training method based on color transformation according to claim 3, wherein Based on the colors of the hint points, calculate the similarity between other pixels to obtain a color similarity matrix, including: Convert the hint color and the transformed image to the Lab space for color space conversion, separate the channels to obtain the L-channel image, compare the similarity of the ab channels, and obtain the similarity matrix map.
6. The coloring model training method based on color transformation according to claim 2, wherein The model training is divided into two stages: The first stage: During training, no color transformation is performed. The target loss function during training is the robust regression loss. When the error predicted by the model is less than the set threshold, the squared loss function is used. If the error is greater than or equal to the set delta value, the absolute value loss function is used. When the training loss converges, the first stage ends; The second stage: During training, randomly perform color transformation on some samples and do not perform color transformation on some samples. The samples that do not change still use the target loss function of the first stage, and the samples that change will adjust the target loss function by combining the weighted attention map.
7. A method for training a coloring model based on color transformation, characterized in that Including: The first main module is used to perform sub-pixel segmentation on the original image based on the clustering algorithm to obtain an image segmentation map; The second main module is used to randomly change the color of the original image to obtain a transformed image; The third main module is used to randomly generate the number of hint points and their corresponding coordinates; The fourth main module is used to calculate the areas where the pixels are the same as those of the hint points based on the image segmentation map, obtain the binary map of the hint point area, and combine it with the transformed image to obtain the hint point colors; The fifth main module calculates the distance between the pixel coordinates of the image segmentation map and the hint points, and obtains the distance diffusion map based on the distance change; The sixth main module converts the color space of the transformed image to obtain the Lab image, and separates the channels of the Lab image to obtain the L-channel image; The seventh main module calculates the similarity between other pixels based on the colors of the hint points to obtain the color similarity matrix, and multiplies the color similarity matrix by the distance diffusion map to obtain the weight attention map; The eighth main module inputs the hint point coordinates, the corresponding hint point colors, and the L-channel image into the model, and combines the weight attention map to perform model training.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for training a coloring model based on color transformation according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the method for training a coloring model based on color transformation according to any one of claims 1 to 6.
Citation Information
Patent Citations
A method and terminal for realizing partial color change of an image
CN104851114B