Method and apparatus for training image coloring model, method and apparatus for coloring image, and computer readable storage medium

By combining the optical flow model and the control model with the image restoration model, the inter-frame consistency and quality issues in video stream colorization are solved, and efficient and accurate video colorization effects are achieved.

CN120612255APending Publication Date: 2025-09-09FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410263990.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have difficulty maintaining temporal consistency and accuracy between frames in video stream colorization, and traditional methods often sacrifice image quality or require tedious manual intervention.

Method used

Using a pre-trained optical flow model and control model, by obtaining the correspondence between line drawings and occlusion masks, combined with an image restoration model, iterative training is performed to generate time-consistent and high-quality colorized videos.

Benefits of technology

Ensure the consistency of video backgrounds and the accuracy of character coloring, adapt to new scenes, reduce manual intervention, and improve image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612255A_ABST
    Figure CN120612255A_ABST
Patent Text Reader

Abstract

Disclosed are a method and apparatus for training a model for image coloring, an image coloring method and apparatus, and a medium. The training method comprises the following steps: obtaining a corresponding relation between pixel points of a current frame line draft and pixel points of a reference frame line draft in a color video stream and a shielding mask for indicating an area where a difference exists between pictures of a current frame and a reference frame; twisting the reference frame based on the corresponding relation to obtain a twisted image corresponding to the current frame; inputting the line draft of the current frame into a control model to extract features, and inputting the extracted features into an image restoration model; inputting the twisted image and the shielding mask into an image restoration model; through an image restoration model, on the basis of the extracted features and the shielding mask, restoring colors of regions with differences in the twisted image; and training the image restoration model and the control model by iteratively performing the above steps with respect to two or more frames in the video stream as the current frame, such that the image restoration model outputs the accurately colored current frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular to a technique for colorizing a video stream by line drawing guided image restoration. Background Art

[0002] Line art colorization is a highly sought-after research topic. One application is helping anime creators automatically color their line art, reducing manual coloring costs and facilitating the animation production process. The primary challenge in automatic line art colorization is maintaining fidelity to the creator's intent. For example, the colors filled in must match the creator's vision and satisfy the creator in terms of hue, brightness, and other aspects.

[0003] Furthermore, coloring a video stream based on line drawings is an even more challenging task. The challenge lies in maintaining temporal consistency between frames in the video stream. This includes consistent coloring of the same object and background stability. While many new methods for video generation have emerged, they have improved temporal consistency to some extent. However, these methods often sacrifice image quality while improving consistency.

[0004] Traditional line drawing coloring methods typically use a color palette, reference image, or color hints to express the artist's intent, and colorize the line drawing based on these intentions. However, the expression of these intentions is sometimes vague, and it is difficult to accurately convey the artist's original intention. For example, coloring based on a color palette or reference image results in a relatively random manner, because each part of the line drawing is assigned a color from the palette or reference image. This uncertain correspondence between line drawing colors leads to inconsistencies between video frames, and the resulting coloring naturally does not accurately correspond to the artist's intended work. Methods based on color hints require the artist to select a color for each part of the line drawing. These methods, based on the color hints provided by the artist, ensure that the coloring results more closely match the artist's intent. However, for video stream coloring, requiring the artist to provide detailed color hints for each frame is obviously a tedious task and is not conducive to practical application.

[0005] In recent years, the field of image generation has made tremendous progress with the rise of diffusion models. Diffusion models are a type of text-to-image model that outputs a corresponding image based on text input. Diffusion models have surpassed generative adversarial networks in image generation capabilities. Stable diffusion models are currently the most advanced.

[0006] Based on the stable diffusion model, the researchers further proposed a control model (Controlnet), which adds other control conditions such as line drawings, depth maps, and postures, so that images that meet both text descriptions and specific control conditions can be generated. Other similar models include ControlVideo, AnimateAnyone, and so on. Based on Controlnet, ControlVideo has expanded from the generation of single-frame images to the simultaneous generation of multiple-frame images, and can output multi-frame images with consistent timing by inputting control conditions (for example, line drawings) of multiple-frame images, where temporal consistency can be achieved through a cross-frame attention mechanism. However, the disadvantage of ControlVideo is that the quality of the generated images is not high, and the temporal consistency of longer videos cannot be ensured. Compared with ControlVideo, AnimateAnyone performs better, but AnimateAnyone generates human action videos based on posture, which makes it difficult for human posture control to generalize to animal scenes with different shapes. Summary of the Invention

[0007] A brief overview of the present disclosure is provided below to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important portions of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts in a simplified form as a prelude to a more detailed description that will be discussed later.

[0008] According to one aspect of the present disclosure, a method for training a model for image coloring is provided, comprising: obtaining a correspondence and an occlusion mask between pixel points of a line draft of a current frame in a color video stream and pixel points of a line draft of a reference frame, wherein the occlusion mask indicates an area where differences exist between the picture of the current frame and the picture of the reference frame; twisting the reference frame based on the correspondence to obtain a twisted image corresponding to the current frame; inputting the line draft of the current frame into a control model to extract features, and inputting the extracted features into an image restoration model; inputting the twisted image and the occlusion mask into the image restoration model; repairing the colors of the areas where differences exist in the twisted image based on the extracted features and the occlusion mask by the image restoration model; and training the image restoration model and the control model by iteratively performing the above steps on two or more frames in the video stream as current frames, so that the image restoration model outputs an accurately colored current frame.

[0009] According to a preferred embodiment, the twisted image and the occlusion mask are obtained by inputting the line drawing of the current frame and the line drawing of the reference frame into a pre-trained optical flow model.

[0010] According to a preferred embodiment, the reference frame is a frame in the video stream that is adjacent to the current frame, or is any frame in the video stream.

[0011] According to a preferred embodiment, the method further comprises inputting a textual prompt into the image restoration model.

[0012] According to a preferred embodiment, training the image restoration model includes training a U-shaped network encoder and a U-shaped network decoder in the image restoration model.

[0013] According to a preferred embodiment, the image restoration model is a stable diffusion XL restoration model.

[0014] According to a preferred embodiment, the loss function used for training includes a mean square error loss function.

[0015] According to another aspect of the present disclosure, a method for coloring an image is provided, comprising: obtaining a correspondence between pixel points of a line draft of each frame of a color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between the picture of each frame and the picture of the reference frame; twisting the reference frame based on the correspondence of each frame to obtain a twisted image corresponding to each frame; inputting the line draft of each frame into a pre-trained control model to extract features, and inputting the extracted features into a pre-trained image restoration model; inputting the twisted image and the occlusion mask corresponding to each frame into the image restoration model; and, through the image restoration model, based on the extracted features and the occlusion mask of each frame, repairing the color of the area where differences exist in the twisted image of each frame and outputting the repaired frame.

[0016] According to a preferred embodiment, the image restoration model and the control model are first trained by using the speed value as the prediction value, and the image restoration model and the control model are second trained by using the noise value as the prediction value.

[0017] According to a preferred embodiment, repairing the color of the area with differences in the torsion image of each frame also includes: performing a first repair on the color of the area with differences in the torsion image based on random noise by inputting the extracted features, torsion image and occlusion mask corresponding to the frame into a first trained image restoration model and a control model; performing a denoising diffusion implicit model (DDIM) inverse mapping on the repaired torsion image to obtain initial latent variables; obtaining a segmentation mask of the line drawing of each frame, the segmentation mask indicating the area containing the character in the line drawing; and performing a second repair on the color of the area with differences in the torsion image by inputting a combination of the segmentation mask, the initial latent variables and random noise as the starting latent variables into a second trained image restoration model and a control model.

[0018] According to a preferred embodiment, the segmentation mask of the line drawing of each frame is obtained by inputting the line drawing into a segmentation model.

[0019] According to a preferred embodiment, the segmentation model is a segmentation-everything model.

[0020] According to another aspect of the present disclosure, there is provided a device for training a model for image coloring, comprising: an obtaining device configured to obtain a correspondence and an occlusion mask between pixel points of a line draft of a current frame in a color video stream and pixel points of a line draft of a reference frame, wherein the occlusion mask indicates an area where there is a difference between the picture of the current frame and the picture of the reference frame; a twisting device configured to twist the reference frame based on the correspondence to obtain a twisted image corresponding to the current frame; an extraction device configured to input the line draft of the current frame into a control model to extract features; a repairing device configured to input the extracted features, the twisted image and the occlusion mask into an image repair model, and repair the color of the area where there is a difference in the twisted image based on the extracted features and the occlusion mask through the image repair model; and a training device configured to train the image repair model and the control model by iteratively performing the above steps for two or more frames in the video stream as current frames, so that the image repair model outputs an accurately colored current frame.

[0021] According to another aspect of the present disclosure, a device for coloring an image is provided, comprising: a first obtaining device configured to obtain a correspondence between pixel points of a line draft of each frame of a color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where there is a difference between the picture of each frame and the picture of the reference frame; a twisting device configured to twist the reference frame based on the correspondence of each frame to obtain a twisted image corresponding to each frame; an extraction device configured to input the line draft of each frame into a first pre-trained control model to extract features; and a first repairing device configured to input the extracted features, the twisted image and the occlusion mask corresponding to each frame into a first pre-trained image repair model, and through the first pre-trained image repair model, based on the extracted features and the occlusion mask of each frame, perform a first repair on the color of the area where there is a difference in the twisted image of each frame and output a first repaired frame.

[0022] According to a preferred embodiment, the first pre-trained control model and the first pre-trained image restoration model are pre-trained based on the speed value, and the device for coloring the image also includes: an inverse mapping device, which is configured to perform a denoising diffusion implicit model inverse mapping on the first restored frame to obtain an initial latent variable; a second obtaining device, which is configured to obtain a segmentation mask of the line drawing of each frame, the segmentation mask indicating the area in the line drawing containing the character; and a second restoration device, which is configured to perform a second restoration on the color of the area where there is a difference in the twisted image corresponding to each frame by inputting a combination of the segmentation mask, the initial latent variable and random noise as the starting latent variable into the second pre-trained image restoration model and the second pre-trained control model, and output a second restored frame, wherein the second pre-trained control model and the second pre-trained image restoration model are pre-trained based on the noise value.

[0023] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a program stored thereon, which, when executed by a processor, causes a computer to perform the steps of the above-mentioned method of training a model for image colorization.

[0024] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a program stored thereon, which, when executed by a processor, causes a computer to perform the steps of the above-mentioned method for coloring an image.

[0025] According to other aspects of the present disclosure, corresponding computer program codes and computer program products are also provided.

[0026] Through the method and apparatus for training a model for image coloring and the method and apparatus for coloring an image of the present disclosure, the background consistency of the video and good adaptability to new scenes are effectively ensured, and the coloring accuracy of the character is also ensured.

[0027] These and other advantages of the present disclosure will become more apparent from the following detailed description of preferred embodiments of the present disclosure in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to further illustrate the above and other advantages and features of the present disclosure, the following is a further detailed description of specific embodiments of the present disclosure in conjunction with the accompanying drawings. The drawings, together with the detailed description below, are included in this specification and form a part of this specification. Elements with the same function and structure are represented by the same reference numerals. It should be understood that these drawings only depict typical examples of the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the drawings:

[0029] Figure 1is a flowchart of a method for training a model for image colorization according to an embodiment of the present disclosure;

[0030] Figure 2 Schematically shows Figure 1 An exemplary implementation process of the method;

[0031] Figure 3 A flowchart illustrating a method for coloring an image according to an embodiment of the present disclosure is shown;

[0032] Figure 4 A fusion method based on local inverse mapping according to an embodiment of the present disclosure is shown;

[0033] Figure 5 The following schematically illustrates an exemplary implementation process of coloring a video stream;

[0034] Figure 6 Schematically shows a block diagram of an apparatus for training a model for image colorization according to an embodiment of the present disclosure;

[0035] Figure 7 Schematically shows a block diagram of a device for coloring an image according to an embodiment of the present disclosure;

[0036] Figure 8 is a block diagram of an exemplary structure of a general-purpose personal computer in which methods and / or apparatus according to embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0037] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. For the sake of clarity and conciseness, not all features of an actual implementation are described in this specification. However, it should be understood that in the process of developing any such actual implementation, many implementation-specific decisions must be made in order to achieve the developer's specific goals, such as meeting those constraints related to the system and business, and these constraints may vary from implementation to implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is a routine task for those skilled in the art who benefit from the content of this disclosure.

[0038] It is also necessary to explain here that, in order to avoid obscuring the present disclosure due to unnecessary details, the accompanying drawings only show the device structure and / or processing steps that are closely related to the solution according to the present disclosure, while other details that are not closely related to the present disclosure are omitted.

[0039] To address the existing technical issues, this disclosure proposes a method for colorizing a color video stream based on line drawings. Specifically, this method uses a pre-trained optical flow model, a control model such as ControlNet, and an image inpainting model as the overall model. The control model and image inpainting model must be pre-trained before colorizing the video stream.

[0040] The following combination Figure 1 and Figure 2 The method 100 for training an image colorization model according to an embodiment of the present disclosure is described below. Figure 2 As shown, the image colorization model 200 according to this embodiment includes an optical flow model 10, a restoration model 20 and a control model 30, wherein the optical flow model 10 is pre-trained, and the training of the image colorization model actually refers to the training of the restoration model 20 and the control model 30.

[0041] like Figure 1 As shown, in step 101, the correspondence between the pixel points of the line draft of the current frame in the color video stream and the pixel points of the line draft of the reference frame and the occlusion mask are obtained. Specifically, in this embodiment, by inputting the line draft 1 of the current frame and the line draft 2 of the reference frame into the optical flow model 10, the correspondence between the pixel points of the line draft 1 of the current frame and the pixel points of the line draft 2 of the reference frame and the occlusion mask 5 can be obtained.

[0042] It should be understood that the optical flow model 10 may be any known optical flow model, such as Recurrent All-Pairs Field Transforms (RAFT), Global Matching Flow (GMFlow), and the like.

[0043] It should also be noted that the reference frame and the current frame refer to any two frames in the same video stream. According to a preferred embodiment, a frame adjacent to the current frame is selected as the reference frame. However, the present disclosure is not limited thereto, and any frame in the video stream can also be selected as the reference frame as needed.

[0044] It should be understood that the occlusion mask 5 indicates the area where there are differences between the picture of the current frame and the picture of the reference frame. For example, the area corresponding to the animated character in the occlusion mask 5 is the area where there are differences between the picture of the reference frame and the picture of the current frame.

[0045] As is known, optical flow records the motion information of each pixel in the line drawing of the reference frame to the corresponding pixel in the line drawing of the current frame, thereby obtaining pixel points corresponding to most pixels in the current frame from the reference frame under the assumption that there is little change between adjacent frames or short-distance frames.

[0046] Next, in step 102, the reference frame is twisted based on the corresponding relationship to obtain a twisted image corresponding to the current frame.

[0047] Specifically, in this embodiment, reference frame 3 is twisted based on the motion information from each pixel in line art 2 of the reference frame to the corresponding pixel in line art 1 of the current frame, calculated by optical flow model 10, to obtain twisted image 4 corresponding to the current frame. In twisted image 4, only the colors of the occluded areas (i.e., the areas with differences in occlusion mask 5) do not accurately correspond to the colors of the current frame, while the colors of all other areas accurately correspond to the colors of the current frame. Therefore, only the occluded areas need to be colored.

[0048] Next, in step 103, the line drawing of the current frame is input into the control model to extract features, and the extracted features are input into the image restoration model. Specifically, in this embodiment, the line drawing 1 of the current frame is input into the control model 30 to extract features, and the extracted features are input into the restoration model 20.

[0049] It should be understood that the repair model 20 and the control model 30 may adopt any known model. For example, the repair model 20 may adopt the stable diffusion XL-inpainting model, and the control model 30 may adopt the Controlnet described above.

[0050] Next, in step 104 , the twisted image and the occlusion mask are input into the image inpainting model. Specifically, in this embodiment, the twisted image 4 and the occlusion mask 5 are input into the inpainting model 20 .

[0051] Next, in step 105, the image inpainting model inpaints the colors of the areas in the torsion image that differ based on the extracted features and the occlusion mask. Specifically, in this embodiment, the inpainting model 20 inpaints the colors of the areas in the torsion image 4 that differ from those in the occlusion mask 5 based on the features extracted by the control model 30 and the occlusion mask 5.

[0052] According to a preferred embodiment, a hint 7 may be input into the repair model 20. The hint 7 may be, for example, a text expressing the author's intention.

[0053] Next, in step 106 , it is determined whether the color of the colored current frame 6 output by the restoration model 20 is accurate. If so, the method 100 ends; if not, the method 100 proceeds to step 107 .

[0054] In step 107, another frame in the same video stream is selected as the current frame. Steps 101 to 107 are iteratively performed to train the restoration model 20 and the control model 30 until the restoration model 20 outputs a current frame with accurate coloring.

[0055] It should be noted that during the training process, the pairing of the current frame and the reference frame can be pre-set. For example, the adjacent frame of each current frame can be set as the reference frame of the current frame, or the same reference frame can be specified for all current frames, such as the first frame or any other frame in the video stream.

[0056] It should also be noted that the loss function of the image colorization model 200 can use, for example, the loss function of the stable diffusion XL inpainting model, that is, the mean squared error (MSE) is used to measure the difference between the predicted result of the image colorization model 200 and the true value. The model prediction value can be a noise value or a speed value obtained based on the noise and the image latent variable.

[0057] According to a preferred embodiment, during the training process, the line drawings of two adjacent frames in the video stream are used as input, the previous frame is used as a reference frame, and the next frame is used as the current frame, and the image coloring model 200 outputs the coloring result of the next frame.

[0058] According to a preferred embodiment, the training of the image restoration model 200 includes performing a first training using the speed value as a prediction value, and performing a second training using the noise value as a prediction value.

[0059] It should be understood that the area indicated by the occlusion mask 5 that needs to be repaired, that is, the area where there is a difference between the current frame and the reference frame mentioned above, usually exists in the part of the foreground character, because movement often occurs on the foreground character. The difficulty of coloring the foreground character is that it needs to comply with the creator's intention and maintain consistency throughout the video stream. Therefore, in the training of the image coloring model 200, specific animation video data can be used to fine-tune the repair model 20 and the control model 30 to ensure the consistency of the character coloring. That is, a part of the animation video can be used as training data to fine-tune the repair model 20 and the control model 30, so that the fine-tuned repair model 20 and the control model 30 can automatically and accurately color the rest of the animation video given a reference frame. Through fine-tuning, the repair model 20 and the control model 30 remember the line drawings and color correspondences of each character in the animation video.

[0060] It should also be understood that fine-tuning the restoration model 20 refers to fine-tuning the U-shaped encoder (Unet) in the restoration model 20 .

[0061] It should also be understood that Figure 2The image colorization model 200 shown includes not only the optical flow model 10, the restoration model 20, and the control model 30 shown, but also includes a variational autoencoder (VAE), a tokenizer, a text encoder, and other components not shown. During training, the optical flow model 10 and the components not shown above remain unchanged.

[0062] It should also be understood that although Figure 2 Not shown, the control model 30 and the restoration model 20 such as the stable diffusion XL restoration model use random noise (eg, randomly sampled Gaussian noise) as a starting latent variable to extract line drawing features and restore the twisted image, respectively.

[0063] The following combination Figures 3 to 5 A method 300 for coloring an image according to an embodiment of the present disclosure is described.

[0064] It should be pointed out that Figure 3 Method 300 is in Figure 5 The image coloring model 400 is implemented on the image coloring model 400 shown, and the image coloring model 400 is implemented on the image coloring model 400 Figure 1 The method 100 is pre-trained. Therefore, Figure 3 Steps 301 to 304 of the method 300 shown correspond to Figure 1 The details of steps 101 to 104 , and therefore steps 301 to 304 , of the illustrated method 100 are not repeated here.

[0065] In step 305 , the image restoration model is used to restore the color of the different regions in the twisted image of each frame based on the extracted features and occlusion mask of each frame, and the restored frame is output.

[0066] Figure 5 An exemplary implementation process for coloring a video stream is shown. The video stream includes N frames, where N is an integer greater than 1. During the inference process, for convenience, the first frame 22 in the video stream can be uniformly selected as the reference frame, and then the line draft 21 of the first frame to the line draft 23 of the Nth frame are input into the optical flow model 10. Subsequently, the first frame 22 is subjected to image twisting according to the optical flow information to obtain a twisted image 24 and an occlusion mask 25 of the first to Nth frames. Subsequently, the twisted image 24, the occlusion mask 25, and the features extracted from the line drafts 21 to 23 by the control model 30' are input into the restoration model 20'. Finally, the restoration model 20' outputs the colored first to Nth frames.

[0067] It should be understood that although Figure 5 The first frame 22 is shown as the reference frame, but the present disclosure is not limited thereto, and any suitable frame in the video stream may be selected as the reference frame.

[0068] According to a preferred embodiment, in order to simultaneously utilize the coloring results of the noise model with stable background and good restoration performance and the accurate character (animal) coloring results of the speed model, the noise model and the speed model can be fused through a local inverse mapping method.

[0069] The following combination Figure 4 The fusion method 300 ′ according to an embodiment of the present disclosure is described.

[0070] In execution Figure 3 After step 303 of method 300, execute Figure 4 Step 3051 of method 300'.

[0071] Specifically, in step 3051, the extracted features, twisted image and occlusion mask corresponding to each frame are input into an image restoration model and a control model pre-trained based on speed values, and the color of the area with differences in the twisted image is first restored based on random noise.

[0072] It should be understood that being pre-trained based on the speed value means using the speed value as a prediction value to train the Unet in the restoration model 20 and the control model 30 in the image colorization model 200.

[0073] It should also be understood that performing the first restoration based on random noise means using random noise as the starting latent variable of the restoration model 20 and the control model 30 .

[0074] Next, in step 3052, the first restored frame is subjected to DDIM inverse mapping to obtain initial latent variables.

[0075] It should be noted that how to obtain latent variables through DDIM inverse mapping is known in the prior art and will not be described in detail here.

[0076] Next, in step 3053 , a segmentation mask of the line drawing of each frame is obtained, where the segmentation mask indicates the area containing the character in the line drawing.

[0077] Specifically, in this embodiment, the segmentation mask of the line drawing of each frame can be obtained by inputting the line drawing into a segmentation model, which can be, for example, a segmentation everything model (SAM). The segmentation mask indicates the area in the line drawing, for example, containing an animal.

[0078] Finally, in step 3054, a second restoration is performed on the color of the different areas in the twisted image corresponding to each frame by inputting the combination of the segmentation mask, initial latent variables and random noise of each frame as the starting latent variables into the image restoration model and the control model pre-trained based on the noise value, and the restored frame is output.

[0079] It should be understood that being pre-trained based on noise values ​​means using noise values ​​as prediction values ​​to train the Unet in the restoration model 20 and the control model 30 in the image colorization model 200.

[0080] In this embodiment, the segmentation mask mask indicating the character area and the initial latent variable initial latent and random noise noise constitute the starting latent variable start of the image colorization model 400 pre-trained based on the noise value latent Specifically, when the restoration model 20 restores the twisted image, the initial latent variable is used in the character area and random noise is used in other areas, and the random noise in other areas is multiplied by an appropriate coefficient to balance the adverse effects caused by the initial latent variable. For example, the initial latent variable start can be set based on the following formula latent :

[0081] start latent =initial latent ×mask+noise×(1-mask)×scale formula (1)

[0082] The value of mask is 0 or 1, where 1 represents the character area and 0 represents other areas, and scale can be set to 0.6, for example. It should be understood that the present disclosure is not limited thereto, and the value of scale can be set differently according to the characteristics of different data.

[0083] The restoration model 20 and the control model 30 use the combination of the segmentation mask, initial latent variables and random noise of each frame according to formula (1) as the starting latent variables to perform image restoration and line drawing feature extraction respectively.

[0084] Figure 4 Steps 3052 to 3054 can be achieved by Figure 5 Specifically, the latent variable module 40 first performs DDIM inverse mapping on the first repaired frame to obtain the initial latent variables, then uses the segmentation model to obtain the segmentation mask of the line draft of each frame, and finally combines the segmentation mask, the initial latent variables and the random noise according to formula (1) as the starting latent variables, and inputs the repair model 20' and the control model 30' pre-trained based on the noise value, so that the control model 30' extracts the features of the line draft of each frame based on the combined starting latent variables, and the repair model 20' repairs the torsion image corresponding to each frame based on the combined starting latent variables.

[0085] experiment

[0086] Experimental setup

[0087] Consider an animal animation video containing multiple animals and multiple scenes, specifically 27 animals and 24 scenes (video streams). The training and test sets are divided by scene, for example, 15 scenes as the training set and 9 scenes as the test set. The training set contains all the animals.

[0088] First, for example, the training set is parsed at a frame rate of 2 FPS to obtain 435 images as training data. For example, the test set is parsed at a frame rate of 6 FPS to make the test video smoother. Next, the same line art extraction algorithm is used to extract line art from the images.

[0089] In order to improve the accuracy of coloring characters (such as animals in this scene), data enhancement was performed on the characters in the training dataset. Specifically, the animals in the training pictures were segmented, and the animals were randomly selected and pasted back into the training pictures after being scaled, rotated, flipped, and other processing. 3 to 5 animals were randomly pasted into each training picture, resulting in 435 animal-enhanced training pictures. Therefore, in this experiment, the total number of training pictures was 870. During the training process, two adjacent frames of line drawings of the same scene were input, and the previous frame was used as a reference frame to enable the image coloring model 200 to learn the coloring of the current frame.

[0090] When fine-tuning the image colorization model using the above training data, for example, AdamW can be used as the optimizer, with an initial learning rate of 1e-5, a batch size of 2, and a total number of training steps of 80,000. Two image colorization models are trained using noise prediction and speed prediction, respectively.

[0091] Experimental results

[0092] Using a noise prediction model, the image colorization model takes as input a line drawing video and the actual colorized image of the first frame. The model then outputs a colored video corresponding to the entire line drawing video. Experimental results using a single noise model for video colorization demonstrate that the model can output high-quality colored images that are consistent with the line drawing content, accurately colorize characters, and maintain a background color consistent with the first frame. Furthermore, the single noise model exhibits excellent adaptability, meaning that even in scenes not seen in the training set, the model can effectively colorize the line drawing video based on the background displayed in the reference frame, while maintaining background stability.

[0093] In scenes with many animals or where the animals' movements vary significantly, using a single noise model cannot achieve good animal colorization results. Experiments have shown that the speed prediction model has a relatively higher accuracy for animal colorization, but the images it outputs contain more noise, resulting in lower image quality. Therefore, the model fusion colorization method based on local inverse mapping described above can combine the advantages of the noise prediction model and the speed prediction model to produce better results. Experimental results show that by transferring the accurate animal colorization results of the speed model to the noise model, good video colorization results are achieved.

[0094] Therefore, the method 100 for training a model for coloring an image and the method 300 for coloring an image according to the present disclosure make it possible to ensure the timing consistency between frames in a video stream, the coloring consistency of character objects, and the stability of the background while adhering to the creator's intention.

[0095] The methods discussed above can be implemented entirely by a computer-executable program, or partially or completely by hardware and / or firmware. When implemented in hardware and / or firmware, or when a computer-executable program is loaded into a hardware device capable of running the program, the apparatus for training a model for image coloring and the apparatus for coloring an image, which will be described below, are implemented. Below, an overview of these apparatuses is given without repeating some of the details already discussed above, but it should be noted that although these apparatuses can perform the methods described above, the methods do not necessarily employ or are not necessarily performed by the components of the described apparatuses.

[0096] Figure 6A device 600 for training a model for image colorization is shown. The device 600 includes an acquisition device 601, a twisting device 602, an extraction device 603, a restoration device 604, and a training device 605. The acquisition device 601 is configured to obtain the correspondence between the pixel points of the line drawing of the current frame in the color video stream and the pixel points of the line drawing of the reference frame and an occlusion mask, wherein the occlusion mask indicates the area where there is a difference between the picture of the current frame and the picture of the reference frame. The twisting device 602 is configured to twist the reference frame based on the correspondence to obtain a twisted image corresponding to the current frame. The extraction device 603 is configured to input the line drawing of the current frame into the control model to extract features. The restoration device 604 is configured to input the extracted features, the twisted image, and the occlusion mask into the image restoration model, and through the image restoration model, based on the extracted features and the occlusion mask, the color of the area where there is a difference in the twisted image is restored. The training device 605 is configured to train the image restoration model and the control model by iteratively performing the above steps on two or more frames in the video stream as current frames, so that the image restoration model outputs a current frame with accurate coloring.

[0097] Figure 6 The apparatus 600 shown for training a model for image colorization corresponds to Figure 1 The method 100 for training a model for image colorization is shown. Therefore, the details of the various devices in the apparatus 600 for training a model for image colorization have been described in detail. Figure 1 The method 100 for training a model for image colorization is described in detail and will not be repeated here.

[0098] Figure 7 A device 700 for coloring an image is shown. Device 700 includes a first obtaining device 701, a twisting device 702, an extraction device 703, and a first restoration device 704. The first obtaining device 701 is configured to obtain a correspondence between the pixel points of the line drawing of each frame of a color video stream and the pixel points of the line drawing of a reference frame, and an occlusion mask, wherein the occlusion mask indicates the areas where there are differences between the image of each frame and the image of the reference frame. The twisting device 702 is configured to twist the reference frame based on the correspondence for each frame, thereby obtaining a twisted image corresponding to each frame. The extraction device 703 is configured to input the line drawing of each frame into a first pre-trained control model to extract features. The first restoration device 704 is configured to input the extracted features, twisted image, and occlusion mask corresponding to each frame into a first pre-trained image restoration model. The first pre-trained image restoration model then performs a first restoration on the color of the areas where there are differences in the twisted image of each frame based on the extracted features and occlusion mask for each frame, and outputs a first restored frame.

[0099] According to a preferred embodiment, the first pre-trained control model and the first pre-trained image restoration model are pre-trained based on the speed value, and the device 700 further includes: an inverse mapping device 705, a second obtaining device 706 and a second restoration device 707. The inverse mapping device 705 is configured to perform a denoising diffusion implicit model inverse mapping on the first restored frame to obtain an initial latent variable. The second obtaining device 706 is configured to obtain a segmentation mask of the line drawing of each frame, and the segmentation mask indicates the area containing the character in the line drawing. The second restoration device 707 is configured to perform a second restoration on the color of the area with differences in the twisted image corresponding to each frame by inputting a combination of the segmentation mask, the initial latent variable and random noise as the starting latent variable into the second pre-trained image restoration model and the second pre-trained control model, and output a second restored frame, wherein the second pre-trained control model and the second pre-trained image restoration model are pre-trained based on the noise value.

[0100] Figure 7 The device 700 for coloring an image shown corresponds to Figure 3 The method 300 for coloring an image is shown. Therefore, the relevant details of each device in the apparatus 700 for coloring an image have been described in detail. Figure 3 and Figure 4 The method 300 and 300 ′ for coloring an image are described in detail and will not be repeated here.

[0101] It should be understood that Figure 7 The inverse mapping means 705, the second obtaining means 706 and the second repairing means 707 in the embodiment can be, for example, Figure 5 It is implemented by the latent variable module 40.

[0102] Each component module or unit in the above-mentioned device or equipment can be configured by software, firmware, hardware or a combination thereof. The specific means or methods that can be used for configuration are well known to those skilled in the art and will not be described in detail here. In the case of implementation by software or firmware, the data is transferred from a storage medium or network to a computer with a dedicated hardware structure (e.g., Figure 8 The general-purpose computer 800 shown in FIG. 1 is installed with programs constituting the software. When various programs are installed, the computer can execute various functions, etc.

[0103] Figure 8 FIG. 1 is a block diagram of an exemplary structure of a general-purpose personal computer in which the methods and / or apparatus according to embodiments of the present invention may be implemented. Figure 8As shown, a central processing unit (CPU) 801 executes various processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 to a random access memory (RAM) 803. In the RAM 803, data required when the CPU 801 executes various processes, etc., is also stored as needed. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output interface 805 is also connected to the bus 804.

[0104] The following components are connected to the input / output interface 805: an input device 806 (including a keyboard, a mouse, etc.), an output device 807 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), a storage device 808 (including a hard disk, etc.), and a communication device 809 (including a network interface card such as a LAN card, a modem, etc.). The communication device 809 performs communication processing via a network such as the Internet. A drive 810 may also be connected to the input / output interface 805 as needed. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the drive 810 as needed, so that a computer program read therefrom is installed in the storage device 808 as needed.

[0105] In the case of realizing the above-described series of processing by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 811 .

[0106] It should be understood by those skilled in the art that such storage media is not limited to Figure 8 The removable medium 811 shown has a program stored therein and is distributed separately from the device to provide the program to the user. Examples of removable medium 811 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be ROM 802, a hard disk included in storage device 808, or the like, in which the program is stored and distributed to the user along with the device containing it.

[0107] The present disclosure also provides corresponding computer program code and a computer program product storing machine-readable instruction code. When the instruction code is read and executed by a machine, the method 100 for training a model for image colorization and the method 400 for colorizing an image according to an embodiment of the present disclosure can be performed.

[0108] Accordingly, storage media configured to carry the program product storing the machine-readable instruction code are also included in the disclosure of the present disclosure, including but not limited to floppy disks, optical disks, magneto-optical disks, memory cards, memory sticks, and the like.

[0109] Through the above description, the embodiments of the present disclosure provide the following technical solutions, but are not limited thereto.

[0110] Note 1. A method for training a model for image colorization, comprising:

[0111] Obtaining a correspondence between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between an image of the current frame and an image of the reference frame;

[0112] Twisting the reference frame based on the correspondence to obtain a twisted image corresponding to the current frame;

[0113] Inputting the line drawing of the current frame into the control model to extract features, and inputting the extracted features into the image restoration model;

[0114] Input the twisted image and occlusion mask into the image inpainting model;

[0115] Inpainting the color of the discrepant region in the twisted image based on the extracted features and the occlusion mask using an image inpainting model; and

[0116] By iteratively executing the above steps for two or more frames in the video stream as the current frame, the image restoration model and the control model are trained so that the image restoration model outputs a current frame with accurate coloring.

[0117] Note 2. The method according to Note 1, wherein the twisted image and the occlusion mask are obtained by inputting the line drawing of the current frame and the line drawing of the reference frame into a pre-trained optical flow model.

[0118] Note 3. The method according to Note 1, wherein the reference frame is a frame in the video stream that is adjacent to the current frame.

[0119] Note 4. The method according to Note 1, wherein the reference frame is any frame in the video stream.

[0120] Note 5. The method according to any one of Notes 1 to 4 also includes inputting text prompts into the image restoration model.

[0121] Note 6. A method according to any one of Notes 1 to 4, wherein training the image restoration model includes training a U-shaped network encoder and a U-shaped network decoder in the image restoration model.

[0122] Note 7. The method according to any one of Notes 1 to 4, wherein the image restoration model is a stable diffusion XL restoration model.

[0123] Note 8. The method according to Note 7, wherein the loss function used for training includes a mean square error loss function.

[0124] Note 9. The method according to Note 8, wherein training the image restoration model and the control model includes performing a first training using a speed value as a prediction value, and performing a second training using a noise value as a prediction value.

[0125] Note 10. A method for coloring an image, comprising:

[0126] Obtaining a correspondence between pixel points of a line draft of each frame of the color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between an image of each frame and an image of the reference frame;

[0127] The reference frame is twisted based on the corresponding relationship of each frame to obtain a twisted image corresponding to each frame;

[0128] Input the line drawing of each frame into the pre-trained control model to extract features, and input the extracted features into the pre-trained image restoration model;

[0129] Input the twisted image and occlusion mask corresponding to each frame into the image inpainting model;

[0130] Through the image restoration model, based on the extracted features and occlusion mask of each frame, the color of the different areas in the torsion image of each frame is restored and the restored frame is output.

[0131] Note 11. The method according to Note 10, wherein the twisted image and occlusion mask corresponding to each frame are obtained by inputting the line drawing of the frame and the line drawing of the reference frame into a pre-trained optical flow model.

[0132] Note 12. The method according to Note 10 or 11, wherein the image restoration model and the control model are first trained by using speed values ​​as prediction values, and the image restoration model and the control model are second trained by using noise values ​​as prediction values.

[0133] Supplementary note 13. The method according to Supplementary note 12, wherein the step of restoring the color of the region with a difference in the twisted image of each frame further comprises:

[0134] Performing a first restoration on the color of the different region in the torsion image based on random noise by inputting the extracted features, the torsion image, and the occlusion mask corresponding to the frame into a first trained image restoration model and a control model;

[0135] The restored torsion image is subjected to inverse mapping of the denoising diffusion implicit model (DDIM) to obtain the initial latent variables;

[0136] Obtaining a segmentation mask for the line art of each frame, the segmentation mask indicating the region of the line art containing the character; and

[0137] A second restoration is performed on the color of the different areas in the twisted image by inputting a combination of the segmentation mask, the initial latent variable and the random noise as the starting latent variable into the second trained image restoration model and the control model.

[0138] Note 14. The method according to Note 12, wherein the segmentation mask of the line drawing of each frame is obtained by inputting the line drawing into the segmentation model.

[0139] Note 15. The method according to Note 10 or 11, wherein the reference frame is a frame in the video stream that is adjacent to the frame to be colored, or the reference frame is any frame in the video stream.

[0140] Note 16. A device for training a model for image colorization, comprising:

[0141] An obtaining device is configured to obtain a correspondence between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between the picture of the current frame and the picture of the reference frame;

[0142] a twisting device configured to twist the reference frame based on the correspondence to obtain a twisted image corresponding to the current frame;

[0143] an extraction device configured to input the line drawing of the current frame into the control model to extract features;

[0144] a restoration device configured to input the extracted features, the torsion image, and the occlusion mask into an image restoration model, and restore the color of the region where the difference exists in the torsion image based on the extracted features and the occlusion mask through the image restoration model; and

[0145] A training device is configured to train an image restoration model and a control model by iteratively performing the above steps on two or more frames in a video stream as current frames, so that the image restoration model outputs a current frame with accurate coloring.

[0146] Note 17. A device for coloring an image, comprising:

[0147] A first obtaining device is configured to obtain a correspondence between pixel points of a line draft of each frame of the color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where there is a difference between the picture of each frame and the picture of the reference frame;

[0148] a twisting device configured to twist the reference frame based on the correspondence between each frame to obtain a twisted image corresponding to each frame;

[0149] An extraction device configured to input the line drawing of each frame into a first pre-trained control model to extract features; and

[0150] A first restoration device is configured to input the extracted features, torsion image and occlusion mask corresponding to each frame into a first pre-trained image restoration model, and through the first pre-trained image restoration model, based on the extracted features and occlusion mask of each frame, perform a first restoration on the color of the area where there are differences in the torsion image of each frame and output a first restored frame.

[0151] Supplement 18. The apparatus for coloring an image according to Supplement 17, wherein the first pre-trained control model and the first pre-trained image restoration model are pre-trained based on a speed value, and the apparatus further comprises:

[0152] an inverse mapping device configured to perform inverse mapping of the denoising diffusion implicit model on the first restored frame to obtain initial latent variables;

[0153] A second obtaining device is configured to obtain a segmentation mask of the line drawing of each frame, the segmentation mask indicating an area containing a character in the line drawing; and

[0154] A second restoration device is configured to perform a second restoration on the color of the different areas in the twisted image corresponding to each frame by inputting a combination of a segmentation mask, an initial latent variable and random noise as a starting latent variable into a second pre-trained image restoration model and a second pre-trained control model, and output a second restored frame, wherein the second pre-trained control model and the second pre-trained image restoration model are pre-trained based on the noise value.

[0155] Supplement 19. A computer-readable storage medium having a program stored thereon, the program causing a computer to perform the method according to any one of Supplements 1 to 9 when executed by a processor.

[0156] Note 20. A computer-readable storage medium having a program stored thereon, the program causing a computer to perform the method according to any one of Notes 10 to 16 when executed by a processor.

[0157] Finally, it should be noted that the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Furthermore, in the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0158] Although the embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative of the present disclosure and do not constitute a limitation of the present disclosure. Those skilled in the art will appreciate that various modifications and variations can be made to the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is limited solely by the appended claims and their equivalents.

Claims

1. A method for training a model for image colorization, comprising: Obtaining a correspondence between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between an image of the current frame and an image of the reference frame; twisting the reference frame based on the corresponding relationship to obtain a twisted image corresponding to the current frame; Inputting the line drawing of the current frame into a control model to extract features, and inputting the extracted features into an image restoration model; inputting the twisted image and the occlusion mask into the image inpainting model; Inpainting the color of the region with the difference in the twisted image based on the extracted features and the occlusion mask using the image inpainting model; and The image restoration model and the control model are trained by iteratively performing the above steps on two or more frames in the video stream as the current frame, so that the image restoration model outputs a current frame with accurate coloring.

2. The method according to claim 1, wherein The twisted image and the occlusion mask are obtained by inputting the line drawing of the current frame and the line drawing of the reference frame into a pre-trained optical flow model.

3. The method according to claim 1, wherein The reference frame is a frame in the video stream that is adjacent to the current frame, or is any frame in the video stream.

4. The method according to any one of claims 1 to 3, wherein Training the image restoration model includes training a U-shaped network encoder and a U-shaped network decoder in the image restoration model.

5. A method for coloring an image, comprising: Obtaining a correspondence between pixel points of a line draft of each frame of a color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where a difference exists between an image of each frame and an image of the reference frame; Twisting the reference frame based on the corresponding relationship of each frame to obtain a twisted image corresponding to each frame; Input the line drawing of each frame into the pre-trained control model to extract features, and input the extracted features into the pre-trained image restoration model; Inputting the twisted image and the occlusion mask corresponding to each frame into the image inpainting model; The image restoration model is used to restore the color of the different regions in the twisted image of each frame based on the extracted features and occlusion mask of each frame, and the restored frame is output.

6. The method according to claim 5, wherein: The twisted image and occlusion mask corresponding to each frame are obtained by inputting the line drawing of the frame and the line drawing of the reference frame into a pre-trained optical flow model.

7. The method according to claim 5 or 6, wherein: performing a first training on the image restoration model and the control model by using a speed value as a prediction value, and performing a second training on the image restoration model and the control model by using a noise value as a prediction value, and The step of repairing the color of the different regions in the twisted image of each frame further includes: Performing a first restoration on the color of the different region in the torsion image based on random noise by inputting the extracted features, the torsion image, and the occlusion mask corresponding to the frame into the first trained image restoration model; The restored torsion image is subjected to inverse mapping of the denoising diffusion implicit model to obtain the initial latent variables; Obtaining a segmentation mask for the line art of each frame, the segmentation mask indicating an area in the line art containing the character; and The second restoration is performed on the color of the different area in the twisted image by inputting the combination of the segmentation mask, the initial latent variable and the random noise as the starting latent variable into the second trained image restoration model and the control model.

8. A device for training a model for image colorization, comprising: An obtaining device configured to obtain a correspondence between pixel points of a line draft of a current frame and pixel points of a line draft of a reference frame in a color video stream and an occlusion mask, wherein the occlusion mask indicates an area where differences exist between the picture of the current frame and the picture of the reference frame; a twisting device configured to twist the reference frame based on the corresponding relationship to obtain a twisted image corresponding to the current frame; an extraction device configured to input the line drawing of the current frame into a control model to extract features; a restoration device configured to input the extracted features, the torsion image, and the occlusion mask into an image restoration model, and restore the color of the different region in the torsion image based on the extracted features and the occlusion mask through the image restoration model; and A training device is configured to train the image restoration model and the control model by iteratively performing the above steps on two or more frames in the video stream as the current frame, so that the image restoration model outputs a current frame with accurate coloring.

9. A device for coloring an image, comprising: A first obtaining device is configured to obtain a correspondence between pixel points of a line draft of each frame of a color video stream and pixel points of a line draft of a reference frame and an occlusion mask, wherein the occlusion mask indicates an area where there is a difference between the picture of each frame and the picture of the reference frame; a twisting device configured to twist the reference frame based on the correspondence between each frame to obtain a twisted image corresponding to each frame; An extraction device configured to input the line drawing of each frame into a first pre-trained control model to extract features; and A first restoration device is configured to input the extracted features, torsion image and occlusion mask corresponding to each frame into a first pre-trained image restoration model, and through the first pre-trained image restoration model, based on the extracted features and occlusion mask of each frame, perform a first restoration on the color of the area where there are differences in the torsion image of each frame and output a first restored frame.

10. The apparatus according to claim 9, wherein The first pre-trained control model and the first pre-trained image restoration model are pre-trained based on a speed value, and the apparatus further comprises: an inverse mapping device configured to perform a denoising diffusion implicit model inverse mapping on the first restored frame to obtain an initial latent variable; A second obtaining device is configured to obtain a segmentation mask of the line drawing of each frame, wherein the segmentation mask indicates an area containing a character in the line drawing; and A second restoration device is configured to perform a second restoration on the color of the different areas in the twisted image corresponding to each frame by inputting a combination of the segmentation mask, the initial latent variable and random noise as a starting latent variable into a second pre-trained image restoration model and a second pre-trained control model, and output a second restored frame, wherein the second pre-trained control model and the second pre-trained image restoration model are pre-trained based on the noise value.