Image coloring method and device based on iterative optimization and medium

Through the iteratively optimized image coloring method, the circular iteration process of the conditional coloring model and the evaluation model are used to identify and adjust unsatisfactory areas, solving the problem of low realism and fidelity of image coloring in the prior art, and achieving higher quality image coloring effects.

CN120543673AActive Publication Date: 2025-08-26NANJING PAIMI INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510623954.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing image shading methods lack iterative optimization, resulting in low realistic and fidelity of image shading, and the manual shading process is time-consuming and laborious.

Method used

Using an iterative optimization-based image coloring method, through the trained conditional coloring model and the cyclic iteration process of evaluating the model, unsatisfactory areas are identified and adjusted to generate higher quality color images.

Benefits of technology

It achieves higher image coloring reality and fidelity, improves image quality and consistency, and is close to the fine adjustment effect of professional colorists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543673A_ABST
    Figure CN120543673A_ABST
Patent Text Reader

Abstract

The invention discloses an image coloring method and device based on iterative optimization and a medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining a grayscale image; performing initial coloring on the grayscale image by using the trained condition coloring model to obtain an initial coloring result; performing color evaluation on the initial coloring result by using a trained evaluation model, and identifying an area which does not accord with a set standard of color confidence in the initial coloring result to obtain an unsatisfied area; re-coloring the unsatisfied area by using the trained condition coloring model to obtain an updated coloring result; and replacing the initial coloring result with the updated coloring result, repeating the re-coloring and evaluation process to obtain a final coloring result, and decoding the final coloring result to obtain a final color image, thereby realizing higher sense of reality and fidelity in the aspect of image coloring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image coloring method, device and medium based on iterative optimization. Background Art

[0002] Many black-and-white photographs from bygone eras still exist today, often requiring the skillful hand of an artist to add color to bring them to life. However, this manual colorization process is both time-consuming and laborious. Employing a single-shot colorization paradigm, numerous methods have been proposed that aim to replicate the prior knowledge and intuition of human experts. These methods have achieved significant progress, providing high-quality results. While some of these related techniques incorporate additional conditioning, they often lack the iterative optimization typically performed by human experts, resulting in lower realism and fidelity in image colorization. Summary of the Invention

[0003] The purpose of this application is to provide an image coloring method, device and medium based on iterative optimization, which can achieve higher realism and fidelity in image coloring.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides an image colorization method based on iterative optimization, comprising the following steps.

[0006] Get a grayscale image.

[0007] The grayscale image is initially colored using the trained conditional coloring model to obtain an initial coloring result.

[0008] The trained evaluation model is used to perform color evaluation on the initial coloring result, and regions in the initial coloring result that do not meet the set standard of color confidence are identified to obtain unsatisfactory regions.

[0009] The unsatisfactory area is recolored using the trained conditional coloring model to obtain an updated coloring result.

[0010] The updated coloring result replaces the initial coloring result, and the recoloring and evaluation process is repeated to obtain a final coloring result, and the final coloring result is decoded to obtain a final color image.

[0011] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned image coloring method based on iterative optimization.

[0012] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned image coloring method based on iterative optimization.

[0013] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0014] The present application provides an image colorization method, device and medium based on iterative optimization, which performs color evaluation on the initial colorization result through a trained evaluation model, identifies the area in the initial colorization result that does not meet the set standard of color confidence, obtains an unsatisfactory area, and uses the trained conditional colorization model to recolor the unsatisfactory area. In the colorization stage, the colorization model generates an initial color prediction. The evaluation stage uses the evaluation model to identify areas with low color confidence. In the recoloring stage, the unsatisfactory areas (i.e., areas with low color confidence) are recolored, and the areas that need improvement can be adjusted in a targeted manner. The adaptive feedback loop can iteratively improve the quality and consistency of color, achieve higher realism and fidelity in coloring, and improve the image coloring quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0016] Figure 1 This is an application environment diagram of an image coloring method based on iterative optimization in one embodiment of the present application.

[0017] Figure 2 A flowchart of an image colorization method based on iterative optimization provided in one embodiment of the present application.

[0018] Figure 3 A schematic diagram of the iterative coloring-evaluation-recoloring process provided in one embodiment of the present application.

[0019] Figure 4 A flowchart of the training and inference phases of image colorization provided in one embodiment of the present application.

[0020] Figure 5 This is a demonstration diagram of the intermediate results of the iterative coloring provided in one embodiment of the present application.

[0021] Figure 6 A schematic diagram of the visual comparison results of unconditional coloring provided by an embodiment of the present application.

[0022] Figure 7 A schematic diagram of the visual comparison results of conditional coloring provided in an embodiment of the present application.

[0023] Figure 8 This is an ablation study diagram of different loss settings used for training according to one embodiment of the present application.

[0024] Figure 9 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] Image colorization uses generative models to colorize grayscale images, and can be categorized into conditional and unconditional methods. Generative models are further categorized into single-shot and multi-step methods, with the latter being more suitable for the colorization, evaluation, and recolorization process. This section reviews both types of models.

[0027] Unconditional colorization aims to automatically add color to grayscale images. Early methods treated it as a regression problem or classification task, focusing on learning the relationship between semantics and color from large data sets. To improve the vividness of color, additional semantic information such as category labels, semantic segmentation maps, or instance bounding boxes are usually incorporated. Recent progress has utilized generative adversarial networks (GANs), Transformers, and diffusion models to enhance colorization by combining prior knowledge, such as DGP, MGANPrior, GCP, BigColor, ColTran, CT 2 , UniColor, DISCO, DDColor. However, the knowledge of colors is still limited to the scope of training data, which limits their ability to generate vivid and realistic colors.

[0028] Conditional colorization traditionally relies on scribbles for guidance. Example-based methods simplify the process by copying colors from a reference image. Some techniques use AdaIN to transfer global color distributions, while others transfer colors through pixel-level semantic alignment. Despite performance improvements, finding an ideal reference image is still time-consuming and unfriendly. Generative colorization provides an alternative by extending the unconditional generative model to include grayscale image conditions. Methods include cINN using conditional normalization flow, SCC-DC using conditional VAE, UniColor using Transformer, Diffusing Colors using diffusion models, and GAN-based methods such as ColorFormer, BigColor, and DDColor.

[0029] Single-shot generative models, such as GANs and diffusion models, generate the final output in a single shot. GANs have made significant progress in learning low-dimensional latent representations of natural color images, enabling their application in restoration, super-resolution, and colorization. Diffusion models, driven by DDPM and DDIM, have led to the emergence of LDMs such as Stable Diffusion, which set a benchmark in text-to-image generation. ControlNet uses LDM to control the pre-trained diffusion model through task-specific conditions, supporting multimodal inputs and broadening application scenarios. However, directly using grayscale images as conditions to train ControlNet will lead to significant texture and detail inconsistencies, which undermines the goal of colorization.

[0030] Multi-step generative models, such as autoregressive models and masked visual token models, iteratively generate outputs by repeatedly passing inputs and partially completed outputs back to the model until a final result is obtained. Early autoregressive image models focused on pixel sequences, while VQGAN converts these sequences into latent tokens for prediction of the next token, similar to BERT. Regression prediction of color tokens has shown excellent results. Traditional methods view images as 1D token sequences in raster scan order (from left to right, row by row), which is not ideal in terms of efficiency and representation. In contrast, masked generation models predict new tokens in all directions or generate multiple tokens simultaneously in a random order, while maintaining the autoregressive principle of predicting subsequent tokens based on known tokens.

[0031] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] The image coloring method based on iterative optimization provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the grayscale image to be processed to the server 104. After the server 104 receives the grayscale image to be processed, the server 104 uses the trained conditional coloring model to perform initial coloring on the grayscale image to obtain an initial coloring result. The server 104 uses the trained evaluation model to perform color evaluation on the initial coloring result, identifies areas in the initial coloring result that do not meet the set color confidence standards, and obtains unsatisfactory areas. The unsatisfactory areas are recolored using the trained conditional coloring model to obtain an updated coloring result. The updated coloring result replaces the initial coloring result. The recoloring and evaluation process is repeated to obtain a final coloring result, and the final coloring result is decoded to obtain a final color image. The server 104 can feedback the final color image obtained for the grayscale image to the terminal 102. In addition, in some embodiments, the image colorization method based on iterative optimization can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform image colorization processing on the grayscale image to be processed, or the server 104 can obtain the grayscale image to be processed from the data storage system and perform image colorization processing on the grayscale image to be processed.

[0033] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0034] In an exemplary embodiment, Figure 2 As shown, an image coloring method based on iterative optimization is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, including the following steps 201 to 205.

[0035] Step 201: Obtain a grayscale image. The grayscale image can be an uncolored image in any field.

[0036] Step 202: Perform initial coloring on the grayscale image using the trained conditional coloring model to obtain an initial coloring result.

[0037] Step 203 : Performing color evaluation on the initial colorization result using the trained evaluation model, identifying areas in the initial colorization result that do not meet the set standard of color confidence, and obtaining unsatisfactory areas.

[0038] Step 204 : recolor the unsatisfactory area using the trained conditional coloring model to obtain an updated coloring result.

[0039] Step 205 : replacing the initial coloring result with the updated coloring result, repeating the recoloring and evaluation process to obtain a final coloring result, and decoding the final coloring result to obtain a final color image.

[0040] Implementing steps 201 to 205 above provides a CAR method, an iterative colorization method involving colorization, evaluation, and recoloring. This method is similar to the iterative workflow of a professional colorist. This method combines a conditional colorization model with an evaluation model to replicate the cycle of colorization, evaluation, and recoloring, similar to the color correction stage in film post-production. Unlike traditional single-shot colorization methods, CAR employs a progressive, multi-step approach. In the colorization phase, the colorization model generates initial color predictions. In the evaluation phase, the evaluation model is used to identify areas with low color confidence, while in the recolorization phase, well-colored areas are masked so that the colorization model can be used to make targeted adjustments to areas requiring improvement. This adaptive feedback loop enables CAR to iteratively improve color quality and consistency, approaching the fine-tuning of professional colorists. Experimental results demonstrate that CAR outperforms traditional single-step methods, achieving higher realism and fidelity in colorization.

[0041] Colorists begin with a basic rough sketch and gradually enhance it by adding or adjusting details. They evaluate the initial shading, identify areas of dissatisfaction, and recolor those areas. Knowing that perfection in shading can't be achieved overnight, they achieve a high degree of visual consistency and accuracy through multiple cycles of coloring, evaluation, and recoloring. This iterative, human-involved approach differs from the single-shot shading used in earlier methods.

[0042] This application introduces an iterative coloring method through coloring, evaluation and recoloring, e.g. Figure 3As shown in the figure, this is a new method that simulates the colorization-evaluation-restaining process, transitioning from single-shot colorization to multi-step colorization. The left part outlines the iterative colorization process, the top image shows the results of the conditional colorization model during the colorization and recoloring steps, and the bottom image highlights the unsatisfactory areas identified by the evaluation model. The right part expands on the colorization iteration described in the left figure. Notably, a close-up view of the recolorization result is appended on top to highlight the quality improvement achieved by recoloring. CAR consists of two main models: an evaluation model that simulates the evaluation step of a colorist; and a conditional colorization model that incorporates expert knowledge and supports iterative color optimization. Specifically, the evaluation model identifies unsatisfactory latent codes in the colorization results, while the conditional colorization model combines a masked visual token model and a conditional diffusion model to perform initial colorization based on the grayscale image and recolor the unsatisfactory latent codes of the mask.

[0043] The key advances of CAR include: (1) using an evaluation model to identify unsatisfactory latent codes that need to be optimized; (2) retaining the satisfactory latent codes, along with the brush strokes and grayscale images, as input to the mask autoregressive model; (3) enhancing the results through a conditional diffusion model instead of directly using the autoregressive model to predict the final image; (4) iteratively optimizing the colorization through a multi-step process to achieve high-quality and consistent output. The colorization paradigm is transformed from single-shot colorization to multi-stage colorization to simulate the color correction process of a colorist. This application shows that the mask autoregressive model provides controllability and user interactivity, while the diffusion model achieves diverse colorization results. Compared with previous automatic colorization methods, the iterative colorization framework of colorization, evaluation and recoloring provided by this application achieves state-of-the-art performance and generalization ability.

[0044] In another exemplary embodiment of the present application, the above step 202 may include the following steps 301 to 304.

[0045] Step 301: Encode the grayscale image and the color hint image respectively to obtain a gray token sequence and a color hint token sequence; the color hint image is an image obtained by color-marking a portion of the grayscale image; the gray token sequence includes a plurality of gray tokens; the color hint token sequence includes a plurality of color hint tokens.

[0046] Step 302: Generate an intermediate token sequence based on the gray token sequence and the color hint token sequence. Step 302 specifically includes: concatenating the gray token sequence and the color hint token sequence to obtain a concatenated token sequence; and processing the concatenated token sequence using an MLP to obtain an intermediate token sequence.

[0047] Step 303: concatenate the intermediate token sequence and the mask sequence to obtain a concatenated mask sequence; the mask sequence is a token sequence of the same size as the intermediate token sequence but masked by all masks.

[0048] Step 304: Input the splicing mask sequence into the trained conditional coloring model for initial coloring to obtain an initial coloring result.

[0049] The above-mentioned step 203 specifically includes: inputting the initial coloring result and the gray token sequence into the trained evaluation model, evaluating the color confidence in the initial coloring result, and determining the area where the color confidence of the initial coloring result does not meet the set standard as an unsatisfactory area; masking the unsatisfactory area to obtain an updated occlusion mask sequence.

[0050] The coloring framework used in this application consists of a Tokenizer, a conditional coloring model, and an evaluation model.

[0051] (1) The Tokenizer is used in step 301 to encode the grayscale image and the color cue image, respectively, to obtain a grayscale token sequence and a color cue token sequence. The Tokenizer encodes the image. The Tokenizer is the KL-16Tokenizer, which is based on the VQ-16Tokenizer and derived from the VQ-GAN. Unlike the VQ-GAN, the KL-16Tokenizer uses the Kullback-Leibler (KL) divergence for regularization instead of vector quantization. "16" represents the stride of the Tokenizer.

[0052] The main network components of Tokenizer include: Convolutional Neural Network (CNN), Embedding Layer and Transformer Encoder. The main function of CNN is to extract local feature information of the image, the Embedding Layer is responsible for mapping the image information extracted by CNN to the vector space, and finally the Transformer enables the model to learn the relationship between different parts of the image. The function of Tokenizer is to divide the image into blocks, and then extract the image features through the network structure therein to form a token sequence of the image. The specific process of step 301 is: divide the grayscale image and the color prompt image respectively to obtain a number of first segments and a number of second segments. Each first segment or each second segment is input into the Tokenizer to obtain a gray token corresponding to each first segment or a color prompt token corresponding to each second segment; the gray tokens corresponding to all first segments constitute a gray token sequence, and the color prompt tokens corresponding to all second segments constitute a color prompt token sequence.

[0053] (2) Conditional shading model (i.e. Figure 4 The coloring model in ( ) includes a visual mask token model and a conditional diffusion model; the visual mask token model is used to process the concatenated mask sequence to obtain a conditional token sequence; the conditional diffusion model is used to perform multi-step noise iteration on the conditional token sequence to generate an intermediate coloring sequence. The conditional diffusion model includes several residual blocks; each residual block includes LayerNorm, linear transformation and SiLU activation function layers. The conditional diffusion model uses a compact MLP containing several residual blocks. Each residual block includes LayerNorm, linear transformation, SiLU activation and residual connection. We use three blocks with 1024 channels. The conditional diffusion model is conditioned on the vector z (conditional token sequence) generated by the masked visual token model and connected with the time embedding of the noise schedule at step t. This combined vector is conditioned on the LayerNorm layer in the MLP through AdaLN, thereby achieving effective refinement during the coloring process.

[0054] The conditional colorization model consists of a masked visual token model and a conditional diffusion model. The colorization framework accepts three inputs: a grayscale image for target color content, user-provided color strokes for local hints (i.e., color hint images), and the colorization results of the previous iteration. These inputs are converted from pixel level to latent representations through a Tokenizer. The evaluation model then identifies unsatisfactory latent codes in the colored image and replaces them with [MASK]Tokens. The masked visual token model combines the unmasked latent features with the stroke information in the grayscale image and the color hint image to generate conditional (conditional token sequence) input for use by the conditional diffusion model, which iteratively refines the colored image.

[0055] The 256×16 tokens (token sequence) extracted by the tokenizer are concatenated along the channel dimension with the latent codes from the grayscale image and the color cue image (the grayscale image's latent code is a gray token sequence, and the color cue image's latent code is a color cue token sequence) to form a 256×32 code, resulting in a concatenated token sequence. This concatenated token sequence is processed through an MLP with an output size of 256×768, yielding an intermediate token sequence. Unsatisfactory tokens identified by the evaluation model are masked and replaced with [MASK] tokens, which are also processed through an MLP. The two processed token sets are concatenated along the width dimension and augmented with positional embeddings, resulting in a concatenated mask sequence consisting of the intermediate token sequence and the masked mask sequence. The resulting sequence is fed into a Transformer inspired by the original Transformer, the Masked Visual Token model, and adapted in ViT. We employ a bidirectional attention mechanism similar to that used in masked autoencoders, allowing known tokens to provide information to unknown tokens, thereby promoting effective communication between tokens.

[0056] (3) An evaluation model is used to assess the quality of the colorization results generated by the conditional diffusion model relative to the original grayscale image. It uses a classifier to process tokens from the grayscale image and the colorization result, and generates a score through the softmax output. The higher the score, the higher the confidence that the colorization result is unsatisfactory. This evaluation model is based on the MLP-Mixer architecture, relying entirely on the MLP and repeatedly operating on the token and feature channels.

[0057] The evaluation model (primarily an MLP-Mixer) consists of a fully connected layer and N mixer layers (N is the number of patches into which the image is divided). The mixer layers are still composed of MLPs. Specifically, the evaluation model consists of a first fully connected layer, a mixer module, and a second fully connected layer, connected in sequence. The mixer module consists of several mixer layers. Its purpose: To evaluate which of the generated colored token sequences are well colored and which are poorly colored.

[0058] The visual mask token model (mainly the MAE architecture) mainly consists of an encoder-decoder structure. The entire processing flow is as follows: the user inputs a gray image I g , you can choose to add additional color hint points I s Alternatively, it may not be added. After the encoder process, the grayscale image becomes a 256*16 grayscale token sequence, and the color cue image becomes a corresponding 256*16 color cue token sequence with color information on some of the tokens. The two are concatenated to form a 256*32 concatenated token sequence. After the MLP process, it becomes a 256*768 intermediate token sequence.

[0059] During the first colorization, the intermediate token sequence and the mask sequence (the sequence of tokens completely masked by the mask) are concatenated. Within the visual mask token model, the mask sequence passes through the internal transformer layer to learn the first 256 tokens. This first 256 tokens are then discarded, and the remaining tokens are fed into the conditional diffusion model as conditions. After t steps of denoising, the token sequence with the colorization result is generated as the initial colorization result.

[0060] The initial colorization result (the token sequence generated by the conditional diffusion model) and the gray token sequence obtained at the beginning are repeatedly applied to the token and feature channels within the evaluation model under the different effects of multiple MLP layers. Feature information is extracted and compared. That is, the image feature information in the token of the generated color image is compared with the image feature information in the token of the original input gray image. Finally, a score for each token is output. Unsatisfactory areas with scores exceeding the threshold will be re-masked before the next iteration.

[0061] In subsequent iterations, the unsatisfactory portions that agree with the evaluation model's results are masked, while the token sequence for the satisfactory portions is output normally, resulting in an updated masked sequence. The grayscale and color cue-fused token sequence is no longer concatenated with the full masked sequence. Instead, it is concatenated with the intermediate token sequence and the updated masked sequence, which is the updated concatenated mask sequence. This updated concatenated mask sequence is passed through the transformer architecture of the visual mask token model. After the mask learns the information from the previous concatenated token sequence and the token blocks containing coloring information, all tokens except those that were masked are discarded. The updated concatenated mask sequence is then fed into the conditional diffusion model. The subsequent process is largely identical to the first iteration, except that the conditional diffusion model now generates colored tokens corresponding to the masked portions, rather than all tokens. These are then combined with the previously unmasked colored tokens to form a complete colored token sequence. This is then fed into the evaluation model for evaluation. After repeated iterations, the color image is finally decoded by the decoder.

[0062] The above-mentioned step 204 specifically includes: replacing the mask sequence with the updated mask sequence, returning to the step of "splicing the intermediate token sequence and the mask sequence" in step 303 to obtain an updated spliced ​​mask sequence; inputting the updated spliced ​​mask sequence into the trained conditional coloring model to obtain an updated coloring result.

[0063] The evaluation model and the conditional coloring model jointly perform the coloring process, such as Figure 4 As shown on the left, starting with a zero image as a prior, the conditional colorization model applies user-provided color cues to colorize the grayscale image. In subsequent iterations, the evaluation model identifies unsatisfactory regions exceeding a threshold and masks them for optimization. The input is encoded, and unsatisfactory regions identified by the evaluation model are replaced with [MASK] tokens. These tokens are processed by the masked vision token model to generate the conditional input to condition the conditional diffusion model.

[0064] In each iteration, the conditional diffusion model predicts all tokens in parallel and only retains tokens with evaluation model scores above the threshold. Tokens with lower scores will be re-predicted in the next iteration. The threshold is gradually lowered so that token generation is completed in several optimization cycles. The conditional diffusion model uses the conditional probability p(z|y) for prediction, where y includes the unmasked token, the grayscale image, and the user color prompt. From z TStarting from ~N(0,I), the back-diffusion process generates z0 aligned with p(z|y). This iteration automates the colorist's colorization-evaluation-recolorization process. By replacing the evaluation model with human input, it becomes interactive: the user marks incorrect areas and provides additional color hints. In the next iteration, the conditional colorization model optimizes the results based on the colorization prior and user feedback.

[0065] Flowchart of the Colorization-Evaluation-Recolorization Process During inference, the evaluation model checks the colorization quality of the previous iteration against the grayscale image, masks any poor markers, and fills the missing markers with [MASK] markers to create a partial result. If it is the first iteration, the previous result is considered zero. Then, the grayscale image and the color hints provided by the user are encoded to form a colorization condition, which is combined with the partial colorization and fed into the visual mask token model. This output instructs the conditional diffusion model to convert Gaussian noise into a color image in the opposite process. During training, the conditional colorization model (including the visual mask token model and the conditional diffusion model) and the evaluation model are trained independently. In order to simulate the function of the evaluation model during inference, the color markers on the ground are randomly masked, and the evaluation model is trained as a classifier.

[0066] The following describes the training process of the evaluation model and the conditional coloring model.

[0067] The evaluation model measures colorization quality, while the conditional colorization model uses feedback from the evaluation model to optimize the results. Although they collaborate to produce the final output during inference, they are trained independently. This separation of training allows for more efficient optimization of the colorization process.

[0068] Evaluation Model Training: The evaluation model is compatible with the output of the conditional diffusion model and operates on a token-by-token basis, taking as input two tokens: one from the grayscale image and the other from the colorized image. It outputs a score indicating the quality of the colorized token relative to the grayscale token, with lower values ​​indicating better results. The evaluation model is trained using cross-entropy, as logical values ​​effectively reflect the quality of the colorization results, similar to the work of a colorist in the evaluation step.

[0069] Conditional Colorization Model Training: This work uses a conditional diffusion model to model the conditional probability p(z|y), where z represents the true token and y is generated by the masked visual token model. To simulate colorist evaluation, a random mask is applied to remove tokens from the previous colorization result. For the initial iteration, all uncolored tokens are masked. During training, the random mask ratio is sampled between [0.7, 1.0].

[0070] Let ∈∈R dis the noise sampled from N(0,I), t is the noise scheduling time step, The noise schedule is defined. The loss function is formulated as the denoising objective.

[0071] L diff =E ε,t [||ε-ε θ (z t |y,t)|| 2 ] (1);

[0072] Where, L diff is the diffusion loss function of the conditional diffusion model, which is used to measure the difference between the noise predicted by the model and the actual noise; E ε,t is the expectation of the noise and time step t; ε is a noise vector sampled from a standard Gaussian distribution N(0,I); θ represents a network parameterized by θ, which contaminates the vector z with noise t is the input, the time step t and the condition vector y are the conditions; ||ε-ε θ (z t |y,t)|| 2 It represents the mean square error (MSE) between the predicted noise and the true noise.

[0073] The noise pollution vector z t Given by formula (2):

[0074]

[0075] Where z t is the noise data generated in step t, z is the original input of the model, is the noise scheduling parameter, which is consistent with the previous use of heavyweight architectures such as U-Net and DiT as noise estimators ε θ Different from the method of

[15] , we adopt a lightweight MLP network. The motivation for this shift is that the conditional diffusion model focuses on a single token instead of the entire image.

[0076] To enhance the colorization performance, this paper optimizes the color generation process by explicitly optimizing the cycle consistency between the generated image and the ground truth image. Specifically, when the noise added in Equation (2) is small (corresponding to low time steps t), it destroys the consistency, thus achieving effective reward fine-tuning. On the contrary, when t is large, z t Approximate random noise z T , directly from z t Predicting z'0 is prone to significant distortion. Therefore, this application uses the perturbation image z at a low time step t to T Single-step sampling is applied to predict the original image z'0 to maintain image fidelity. The color generation can be expressed as shown in Equation (3).

[0077]

[0078] In the formula, z0 represents the final generated image Token, that is, the coloring result output by the conditional diffusion model, and z'0 represents the original image Token predicted by single-step sampling; t is the noise image at time step t, ε θ (z t |y,t) is the noise estimator, where the symbols have the same meaning as in formula (1); Represents the noise correction factor, which is used to adjust the noise; is a scaling factor used to adjust the step size.

[0079] Then, we apply the denoised image z'0 for reward fine-tuning as shown in formula (4), where h(z) represents the Holistically-Nested Edge Detection (HED) operator, which calculates the HED image from the image.

[0080] L reward =E ε,t [||h(z0)-h(z'0)|| 2 ](4);

[0081] Where, L reward is the loss function of reward fine-tuning, whose purpose is to make the HED map of the generated image z'0 closer to the HED map of the real image z0, E ε ,t is the expectation of noise and time step t, ||h(z0)-h(z'0)|| 2 is the mean square error (MSE) between the real image and the HED image of the generated image, and h(·) represents the HED operator.

[0082] Reward fine-tuning guides the conditional diffusion model to generate images, thereby restoring cycle consistency, improving the adherence to the conditions during the generation process, and prompting the masked visual token model to generate more accurate color conditions. The total loss combines the diffusion training loss and the reward fine-tuning loss, and the total loss function of the conditional diffusion model is expressed as:

[0083]

[0084] Among them, L total represents the total loss function; λ is a hyperparameter that adjusts the weight of the reward fine-tuning loss; T threshold is the time step threshold, which is a hyperparameter that adjusts whether the loss of reward fine-tuning is included in the total loss function of the current time step, and is used to determine whether the noisy image z should be used. t Make fine-tuning of rewards.

[0085] Training Data Preparation: Training the conditional colorization model requires cue points to provide color information for grayscale images, while evaluating the model requires color-grayscale token pairs with good and bad colorization results. Because collecting large-scale annotated datasets is impractical, this application uses the SLIC algorithm for superpixel segmentation. The average color of each segment is assigned to its corresponding superpixel, creating a superpixel image. During training, these superpixels are randomly selected as cue points, generating different color cue images.

[0086] For the evaluation model, good token pairs are obtained by encoding a color image and its grayscale version using the KL-16 tokenizer. Bad token pairs are generated by colorizing the grayscale input using a trained conditional colorization model on a zero-stroke image. This application computes HED images of the ground truth and the colorized output, identifying areas where colorization failed due to issues like color overflow or omission. Tokens in these areas are classified as bad color-grayscale token pairs.

[0087] The intermediate results of CAR iterative coloring are shown as follows Figure 5 As shown in Figure 2. Given a grayscale input, CAR produces the final colorization result through multiple steps. In the first step, the conditional colorization model generates an initial colorization result based on the input. Then, the evaluation model identifies unsatisfactory latent codes with high scores. In the subsequent step, the conditional colorization model uses the input and the satisfactory latent code to recolor the unsatisfactory regions. The final colorization result is obtained after several evaluation-recolorization iterations.

[0088] CAR is a new colorization method that simulates the iterative colorization process of a colorist. The following experiments demonstrate the effectiveness of the image colorization method based on iterative optimization proposed in this application.

[0089] The experimental setup is as follows:

[0090] (1) Datasets: The proposed method is trained on ImageNet and OpenImage. Evaluation is performed on (i) ImageNet images; (ii) COCO images; and (iii) 500 natural scene photos. Evaluating the colorization performance of COCO and grayscale images from the internet is crucial because the color distribution of these images is often different from that of ImageNet and OpenImage, challenging the generalization ability of the model.

[0091] (2) Baseline: To verify the effectiveness of our method, we benchmark it against seven state-of-the-art unconditional colorization methods: DISCO, CT 2, ColorFormer, BigColor, UniColor, DDColor, and CtrlColor. This application uses the official code and pre-trained weights of all baselines. For UniColor, this application evaluates its automatic colorization capabilities without user-provided prompts. For conditional colorization, this application compares our method with UniColor and iColoriT using brush strokes.

[0092] (3) Evaluation Metrics: To evaluate perceptual realism, we use FID, which measures the distributional similarity between generated images and ground truth images. To evaluate color vividness, we use CF to quantify the vividness of generated images. However, since a high CF score does not necessarily mean higher visual quality, we also incorporate PSNR evaluation, following previous research, to provide additional insights. In addition, we conduct a user study to measure subjective preferences, ensuring a comprehensive evaluation that combines quantitative metrics and human perceptual evaluation.

[0093] (4) Training configuration: This application is implemented in PyTorch and trained on 10 3090 GPUs. This application uses the AdamW optimizer with β1 = 0.9 and β2 = 0.95 for 200 epochs. The learning rate is set to 10 -5 This application maintains an exponential moving average (EMA) of the model parameters with a momentum of 0.9999.

[0094] The visual comparison results of unconditional coloring are as follows Figure 6 As shown in the figure, qualitative visualization of image colorization is demonstrated by using various automatic colorization methods, including DISCO, CT2, ColFormer, BigColor, UniColor, DDColor, and CtrlColor. Compared with the related art methods, the method of this application is able to generate more natural and realistic colors.

[0095] CAR is the first method to implement the iterative coloring-evaluation-recoloring process, so it is impractical to directly compare the intermediate results of the iterations with the related art methods. Figure 5 Intermediate results are shown in to illustrate how the colorization results gradually improve as the evaluation model detects unsatisfactory latent codes and the conditional colorization model optimizes them.

[0096] This paper divides the final colorization results into two categories: unconditional colorization and conditional colorization. For each type, this paper conducts qualitative and quantitative comparisons.

[0097] 1) Qualitative comparison: This application visualizes the results of unconditional coloring and conditional coloring for quantitative comparison.

[0098] Unconditional colorization: This paper compares our method with seven related art methods for colorizing grayscale images without user prompts. The results are shown in Figure 2. Figure 6 As shown in Figure 2. It is worth noting that the ground truth images are omitted because the evaluation should not be based solely on color similarity due to the multimodal uncertainty of the colorization task itself. Our method generates more natural and vivid results while reducing color bleeding.

[0099] Conditional colorization: This application evaluates the comparison of our method with four stroke-based colorization methods, where the user provides a color stroke as a condition. For fairness, our method does not generate hint points from superpixel images, but follows UniColor and assigns the average color of the cell to the hint point. Figure 7 As shown, our method smoothly and consistently propagates brushstroke color across the entire object. Furthermore, for areas without assigned hint points, our method generates diverse and vibrant colors. Unlike unconditional colorization, the input includes both a grayscale image and stroke-based color hints. A qualitative comparison of iColoriT, UniColor, DDColor, and CtrlColor demonstrates that our method generates more natural and realistic images.

[0100] 2) Quantitative Comparison: We collected statistics on unconditional and conditional colorization results, covering ImageNet, COCO, and natural scene photos. We compared our method with unconditional and conditional colorization methods on the test datasets. The quantitative results are shown in Table 1. The top half of the table corresponds to unconditional colorization, while the bottom half focuses on conditional colorization. All methods were evaluated using their official code and weights. Our method achieved the lowest value on the Fréchet Inception Distance (FID) metric, demonstrating its ability to produce high-quality colorization results and exhibiting strong generalization. Although the color vividness score reflects the vividness of the image, and some methods surpass our method on this metric, a higher score does not always correlate with better visual quality. To further evaluate performance, we calculated the peak signal-to-noise ratio (PSNR) to measure the deviation between the generated images and the real images. Our method achieved the highest PSNR on all datasets, confirming its ability to produce more natural and realistic colorization results.

[0101] In order to evaluate the authenticity of the output results of this application on natural scene photos, this application conducted a "preference and non-preference" perception study. This application conducted an experiment, and a total of 25 people participated in the experiment. These 25 people were randomly selected, with no special requirements for academic qualifications and occupations, covering three age groups: old, middle-aged, and young. They will see five gray pictures and corresponding brushstroke color prompts, as well as the color results generated by the model of this application and the compared model. They will be asked to choose their favorite color picture from them. This application only collected data from 25 participants for each algorithm. The result data of unconditional coloring and conditional coloring are shown in Table 1.

[0102] This application conducts a comprehensive quantitative evaluation on multiple objective benchmarks, including ImageNet, COCO, and images in the wild (i.e., images not found in the aforementioned training sets and without labels), measuring the FID, CF, and PSNR metrics. Furthermore, we conduct a user study to assess subjective preferences. Our colorization framework achieves the best FID score, demonstrating superior perceptual realism, while the highest preference rate confirms its overall quality over existing methods.

[0103] Table 1 Quantitative comparison

[0104] ImageNet COCO UserStudy FID↓ CF↑ PSNR↑ DISCO 10.29 40.95 20.72 CT2 5.51 38.48 23.50 ColFormer 4.91 38.00 23.10 BigColor 5.36 39.74 21.24 UniColor 9.46 39.01 22.17 DDColor 4.38 37.66 23.54 CtrlColor 8.88 47.17 24.67 Ours 4.35 39.13 28.36 iColoriT 5.01 34.33 28.86 UniColor 7.04 36.16 24.06 CtrlColor 4.67 32.36 26.14 Ours 4.23 36.59 28.99

[0105] Ours in Table 1 refers to the colorization framework in this application, including Tokenizer, conditional colorization model, and evaluation model.

[0106] CAR uses two loss terms to train the coloring network: L diff , which is the diffusion loss used to train the conditional diffusion model; L reward , used to ensure that the HED (Holistically-Nested Edge Detection) edges of the colorized result are aligned with the original color image, which is the loss function for reward fine-tuning. Relying solely on diffusion loss will lead to a decrease in colorization quality, while combining it with reward loss significantly improves performance. This improvement is in Figure 8 This is clearly demonstrated in both unconditional and conditional colorization scenarios. The performance improvement comes from addressing the color loss (first row) and color overflow (second row) issues, which distort the HED edges of the colorized output. By adding this constraint, the model is guided to prioritize color edges, which is a key factor in achieving higher visual quality.

[0107] Ablation studies of different loss settings for training are shown in Figure 8As shown in Figure 2, using diffusion loss alone leads to color deficiency and outliers, reducing the final colorization effect. However, combining it with reward loss improves the colorization results for both unconditional (first row) and conditional (second row).

[0108] This application proposes a novel iterative image colorization framework that mimics the colorization process of human colorists. The main contributions of this application are the conditional colorization model and the evaluation model. The evaluation model identifies areas with suboptimal colorization, while the conditional colorization model specifically optimizes these masked areas. This application's framework surpasses the current state-of-the-art baseline methods both qualitatively and quantitatively, providing controllable and editable colorization results. This application has the potential to advance the development of a wide range of computer vision tasks.

[0109] The present application also provides an application scenario, which applies the above-mentioned image colorization method based on iterative optimization. Specifically: the image colorization method based on iterative optimization provided in this embodiment can be applied in an image coloring scenario. The image coloring scenario includes an image acquisition link and an image coloring link; the grayscale image enters the image coloring link from the image acquisition link, and the corresponding final color image is obtained through human-computer collaboration. The image colorization method based on iterative optimization provided in this embodiment belongs to the image coloring link. Specifically, in the image coloring link process for grayscale images, the grayscale image can be initially colored using a trained conditional coloring model to obtain an initial coloring result, the initial coloring result can be color evaluated using a trained evaluation model, the area in the initial coloring result that does not meet the set color confidence standard is identified, and an unsatisfactory area is obtained. The unsatisfactory area is recolored using the trained conditional coloring model to obtain an updated coloring result, the initial coloring result is replaced by the updated coloring result, the recoloring and evaluation process is repeated to obtain a final coloring result, and the final coloring result is decoded to obtain a final color image.

[0110] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 9As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image coloring processing data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image coloring method based on iterative optimization is implemented.

[0111] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0112] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0113] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0114] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0115] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0116] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0117] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0118] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. An image coloring method based on iterative optimization, characterized in that: The image coloring method based on iterative optimization includes: Get a grayscale image; Performing initial coloring on the grayscale image using the trained conditional coloring model to obtain an initial coloring result; Performing color evaluation on the initial colorization result using the trained evaluation model, identifying areas in the initial colorization result that do not meet the set color confidence standards, and obtaining unsatisfactory areas; Recoloring the unsatisfactory area using the trained conditional coloring model to obtain an updated coloring result; The updated coloring result replaces the initial coloring result, and the recoloring and evaluation process is repeated to obtain a final coloring result, and the final coloring result is decoded to obtain a final color image.

2. The image coloring method based on iterative optimization according to claim 1, characterized in that: Performing initial colorization on the grayscale image using the trained conditional colorization model to obtain an initial colorization result, specifically including: Encoding the grayscale image and the color hint image respectively to obtain a gray token sequence and a color hint token sequence; the color hint image is an image obtained by color-marking a portion of the grayscale image; the gray token sequence includes a plurality of gray tokens; the color hint token sequence includes a plurality of color hint tokens; generating an intermediate token sequence according to the gray token sequence and the color prompt token sequence; The intermediate token sequence and the masking mask sequence are concatenated to obtain a concatenated mask sequence; the masking mask sequence is a token sequence of the same size as the intermediate token sequence and masked by all masks; The splicing mask sequence is input into a trained conditional coloring model for initial coloring to obtain an initial coloring result.

3. The image coloring method based on iterative optimization according to claim 2, characterized in that: Generating an intermediate token sequence according to the gray token sequence and the color prompt token sequence specifically includes: Splicing the gray token sequence and the color prompt token sequence to obtain a spliced ​​token sequence; The concatenated token sequence is processed using MLP to obtain an intermediate token sequence.

4. The image coloring method based on iterative optimization according to claim 2, characterized in that: The conditional coloring model includes a visual mask token model and a conditional diffusion model; the visual mask token model is used to process the spliced ​​mask sequence to obtain a conditional token sequence; the conditional diffusion model is used to perform multi-step noise addition iteration on the conditional token sequence to generate an intermediate coloring sequence.

5. The image coloring method based on iterative optimization according to claim 4, characterized in that: The conditional diffusion model includes several residual blocks; each residual block includes LayerNorm, linear transformation and SiLU activation function layers.

6. The image coloring method based on iterative optimization according to claim 2, characterized in that: The trained evaluation model is used to perform color evaluation on the initial colorization result, and the areas in the initial colorization result that do not meet the set color confidence standards are identified to obtain unsatisfactory areas, specifically including: Inputting the initial colorization result and the gray token sequence into a trained evaluation model, evaluating the color confidence of the initial colorization result, and determining the area of ​​the initial colorization result where the color confidence does not meet the set standard as an unsatisfactory area; Mask the unsatisfactory area to obtain an updated occlusion mask sequence.

7. The image coloring method based on iterative optimization according to claim 6, characterized in that: Recoloring the unsatisfactory area using the trained conditional coloring model to obtain an updated coloring result, specifically including: Replace the masking mask sequence with the updated masking mask sequence, and return to the step of "joining the intermediate token sequence and the masking mask sequence" to obtain an updated joint mask sequence; The updated splicing mask sequence is input into the trained conditional colorization model to obtain the updated colorization result.

8. The image coloring method based on iterative optimization according to claim 1, characterized in that: The evaluation model includes a first fully connected layer, a Mixer module and a second fully connected layer connected in sequence; the Mixer module includes several Mixer layers.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the iterative optimization-based image coloring method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image coloring method based on iterative optimization according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Optimal model based interactive image recolorating method

    CN105893649A

  • Single image re-coloring method based on neural network

    CN110163927A

  • Cartoon image coloring method based on conditional generative adversarial network

    CN118736061A

  • Zero-sample infrared image coloring method and system based on multilevel representation fusion

    CN118823175A

  • Local image condition coloring method and system and electronic equipment

    CN119600033A