Training method and apparatus for model for image coloring, image coloring method and apparatus, and computer-readable storage medium

The method uses optical flow and inpainting models to establish frame correspondences and iteratively train models for accurate coloring of line drawings in video streams, addressing consistency and quality issues in existing technologies.

JP2025137444APending Publication Date: 2025-09-19FUJITSU LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2025028560
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2025-02-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing methods for coloring line drawings in video streams struggle with maintaining timing consistency and image quality, particularly in long videos, and fail to accurately reflect the artist's intentions due to vague intention expressions and cumbersome artist input requirements.

Method used

A method involving a pre-trained optical flow model to establish correspondences between frames, a control model for feature extraction, and an image inpainting model to accurately color frames using occlusion masks and iterative training, ensuring consistency and quality through models like ControlNet and stable diffusion XL inpainting.

Benefits of technology

Ensures consistent timing and accurate coloring of video streams by maintaining background stability and adaptability to new scenarios, while accurately reflecting the artist's intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137444000001_ABST
    Figure 2025137444000001_ABST
Patent Text Reader

Abstract

To provide a training method and apparatus for a model for image coloring, an image coloring method and apparatus, and a storage medium.SOLUTION: A training method comprises the steps of: acquiring a correspondence relationship between pixel points of a line drawing of a current frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask indicating a region where there is a difference between the image of the current frame and the image of the reference frame; warping the reference frame on the basis of the correspondence relationship to acquire a warped image; inputting the line drawing of the current frame into a control model to extract features, and inputting the extracted features into an image restoration model; inputting the warped image and the occlusion mask into the image restoration model; and restoring the color of the region having a difference in the warped image. The image restoration model and the control model are trained by repeatedly executing the above steps so that the image restoration model outputs a current frame that has been accurately colored.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of image processing, and more particularly to techniques for coloring video streams by guiding image inpainting with line drawings. [Background technology]

[0002] Coloring line drawings is a hot research topic. One of its applications is to help animators automatically color line drawings, reducing the cost of manual coloring and facilitating the animation production process. The main challenge of automatic coloring of line drawings is to achieve faithfulness to the artist's intentions. For example, it is necessary to paint colors exactly as the artist intended and to satisfy the artist in terms of hue, brightness, etc.

[0003] Furthermore, coloring a line-drawing-based video stream is a more challenging task. It is difficult to maintain timing consistency between frames of the video stream. This consistency includes coloring consistency for the same object and background stability. Currently, many new methods for video generation have emerged that have improved the timing consistency problem to some extent. However, although these methods can improve consistency, many of them fail to maintain image quality.

[0004] Conventional line drawing coloring methods typically express the artist's intentions using a palette, reference image, or color prompts, and then color the line drawing according to those intentions. However, these intention expressions can be vague, making it difficult to accurately express the artist's intentions. For example, coloring results based on a palette or reference image are relatively random because each part of the line drawing is assigned a color from the palette or reference image. This uncertain color correspondence of the line drawing causes inconsistencies between video frames, and the coloring results do not accurately correspond to the artist's intended work. In methods based on color prompts, the artist must select a color for each part of the line drawing. Because such methods rely on the color prompts provided by the artist, the coloring results tend to more closely match the artist's intentions. However, coloring a video stream requires the artist to provide detailed color prompts for each frame, which is obviously cumbersome and impractical.

[0005] In recent years, the development of diffusion models has made great progress in the field of image generation. Diffusion models are character-based image generation models that output images corresponding to input characters. The image generation capabilities of diffusion models have already surpassed those of generative adversarial networks. The stable diffusion model is currently the most advanced.

[0006] Based on the stable diffusion model, researchers have further proposed a control model (ControlNet), which can add other control conditions, such as line drawings, depth maps, and posture, to generate images that satisfy both text descriptions and specific control conditions. Other similar models include ControlVideo and AnimateAnyone. ControlVideo, based on ControlNet, extends the generation of single-frame images to the simultaneous generation of multi-frame images. By inputting control conditions (e.g., line drawings) for multi-frame images, it can output multiple frame images with consistent timing. Here, timing consistency can be achieved using a cross-frame attention mechanism. However, ControlVideo has the disadvantage of generating low-quality images and being unable to ensure timing consistency for long videos. While AnimateAnyone performed better than ControlVideo, AnimateAnyone is a posture-based video generation of human movements, making it difficult to generalize human posture control to various animal scenarios. Summary of the Invention [Problem to be solved by the invention]

[0007] The following presents a simplified summary of the disclosure in order to provide a basic understanding of aspects of the disclosure. However, this summary is not an exhaustive overview of the disclosure, and it is not intended to identify key or important portions of the disclosure or to limit the scope of the disclosure. Rather, it is intended to merely introduce concepts in a simplified form as a prelude to the more detailed description that is presented later.

[0008] The present disclosure provides a method and apparatus for training a model for image coloring, an image coloring method and apparatus, and a storage medium. [Means for solving the problem]

[0009] In one aspect of the present disclosure, there is provided a method for training a model for image colorization, the method comprising the steps of: obtaining correspondences between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas of difference between the image of the current frame and the image of the reference frame; warping the reference frame based on the correspondences to obtain a warped image corresponding to the current frame; inputting the line drawing of the current frame into a control model to extract features and inputting the extracted features into an image inpainting model; inputting the warped image and the occlusion mask into the image inpainting model; inpainting the color of the areas of difference in the warped image based on the extracted features and the occlusion mask using the image inpainting model; and training the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame.

[0010] Preferably, the warped image and the occlusion mask are obtained by inputting a line drawing of the current frame and a line drawing of the reference frame into a pre-trained optical flow model.

[0011] Preferably, the reference frame is a frame adjacent to the current frame in the video stream, or any frame in the video stream.

[0012] Preferably, the method further comprises the step of inputting a text prompt into the image inpainting model.

[0013] Preferably, the step of training the image inpainting model includes training a U-Net encoder and a U-Net decoder in the image inpainting model.

[0014] Preferably, the image inpainting model is a stable diffusion XL inpainting model.

[0015] Preferably, the loss function used for training comprises a mean squared error loss function.

[0016] Another aspect of the present disclosure provides a method for coloring an image, the method including the steps of: obtaining correspondences between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, wherein the occlusion mask indicates areas where there are differences between the image of each frame and the image of the reference frame; warping the reference frame based on the correspondences of each frame to obtain warped images corresponding to each frame; inputting the line drawings of each frame into a pre-trained control model to extract features, and inputting the extracted features into a pre-trained image inpainting model; inputting the warped images and occlusion mask corresponding to each frame into the image inpainting model; and using the image inpainting model to inpaint colors in the areas where there are differences in the warped images of each frame based on the extracted features and occlusion mask of each frame, and outputting inpainted frames.

[0017] Preferably, a first training is performed on the image restoration model and the control model using velocity values ​​as predicted values, and a second training is performed on the image restoration model and the control model using noise values ​​as predicted values.

[0018] Preferably, the step of inpainting the colors of the different regions in the warped image of each frame includes: performing a first inpainting on the colors of the different regions in the warped image based on random noise by inputting the extracted features, the warped image, and the occlusion mask corresponding to the frame into the first trained image inpainting model; performing a denoising diffusion implicit model (DDIM) inverse mapping on the inpainted warped image to obtain initial latent variables; obtaining a segmentation mask of the line drawing of each frame, where the segmentation mask indicates regions containing characters in the line drawing; and performing a second inpainting on the colors of the different regions in the warped image by inputting a combination of the segmentation mask, the initial latent variables, and the random noise as initial latent variables into the second trained image inpainting model and a control model.

[0019] Preferably, the segmentation mask for each frame's line drawing is obtained by inputting the line drawing into a segmentation model.

[0020] Preferably, the segmentation model is a segment-anything model.

[0021] Another aspect of the present disclosure provides an apparatus for training a model for image coloring, the apparatus including: an acquisition unit that acquires correspondences between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas where there are differences between the image of the current frame and the image of the reference frame; a warping unit that warps the reference frame based on the correspondences and acquires a warped image corresponding to the current frame; an extraction unit that inputs the line drawing of the current frame to a control model to extract features; a repair unit that inputs the extracted features, the warped image, and the occlusion mask to an image inpainting model and inpaints colors of the areas where there are differences in the warped image based on the extracted features and the occlusion mask using the image inpainting model; and a training unit that trains the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame.

[0022] Another aspect of the present disclosure provides an apparatus for coloring an image, the apparatus including: a first acquisition unit that acquires correspondences between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, wherein the occlusion mask indicates areas where there are differences between the image of each frame and the image of the reference frame; a warping unit that warps the reference frame based on the correspondences of each frame and acquires warped images corresponding to each frame; an extraction unit that inputs the line drawings of each frame into a first pre-trained control model to extract features; and a first restoration unit that inputs the extracted features, warped images, and occlusion masks corresponding to each frame into a first pre-trained image restoration model, and performs a first restoration on the colors of the areas where there are differences in the warped images of each frame using the first pre-trained image restoration model based on the extracted features and occlusion mask of each frame, and outputs a first-restored frame.

[0023] Preferably, the first pre-trained control model and the first pre-trained image inpainting model are pre-trained based on a velocity value, and the apparatus for training a model for image coloration further includes: an inverse mapping unit that performs denoising diffusion implicit model inverse mapping on the first inpainted frames to obtain initial latent variables; a second acquisition unit that obtains a segmentation mask of the line drawing of each frame, where the segmentation mask indicates an area including a character in the line drawing; and a second inpainting unit that performs a second inpainting on the colors of the different areas in the warped image of each frame by inputting a combination of the segmentation mask, the initial latent variables, and random noise as an initial latent variable into the second pre-trained image inpainting model and the second pre-trained control model, and outputs a second inpainted frame, where the second pre-trained control model and the second pre-trained image inpainting model are pre-trained based on a noise value.

[0024] Another aspect of the present disclosure provides a computer-readable storage medium having a program stored thereon, the program, when executed by a processor, causing the computer to perform the method for training a model for coloring an image described above.

[0025] Another aspect of the present disclosure provides a computer-readable storage medium having a program stored thereon, the program, when executed by a processor, causing the computer to perform the method for coloring an image described above.

[0026] Other aspects of the present disclosure further provide corresponding computer program code and computer program products.

[0027] The method and apparatus for training a model for image coloring and the method and apparatus for image coloring disclosed herein can effectively ensure the consistency of the video background and good adaptability to new scenarios, while ensuring the accuracy of character coloring.

[0028] The above and other advantages of the present disclosure will become more apparent from the following detailed description of preferred embodiments of the present disclosure with reference to the drawings. [Brief explanation of the drawings]

[0029] In order to make the above and other advantages and features of the present disclosure comprehensible, specific embodiments of the present disclosure will be described in detail below with reference to the drawings. The drawings and the following detailed description are incorporated into and constitute a part of this specification. Elements having the same function and structure are designated by the same reference numerals. It should be noted that these drawings are merely for illustrating typical examples of the present disclosure and are not intended to limit the scope of the present disclosure. [Figure 1] 1 is a flowchart of a method for training a model for image coloring according to an embodiment of the present disclosure. [Figure 2] 2 is a schematic diagram illustrating the flow of an exemplary implementation of the method of FIG. 1. [Figure 3] 1 is a flowchart of an image coloring method according to an embodiment of the present disclosure. [Figure 4] 1 is a flowchart illustrating a fusion method based on local inverse mapping according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram illustrating the flow of an exemplary implementation of coloring a video stream. [Figure 6] FIG. 1 is a schematic block diagram illustrating an apparatus for training a model for coloring an image according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic block diagram illustrating an image coloring apparatus according to an embodiment of the present disclosure. [Figure 8] FIG. 1 is a block diagram illustrating an exemplary configuration of a general-purpose personal computer in which methods and / or apparatus according to embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0030]

[0023] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. For convenience of description, the specification does not include all features of actual embodiments. It should be noted that, in actual implementation, specific embodiments may be modified to achieve specific goals of developers, for example, according to system and business constraints. Although the development work is very complex and time-consuming, it is merely an example for those skilled in the art of this disclosure.

[0031] It should be noted that, for clarity of the present disclosure, the drawings only show device components and / or process steps closely related to the present disclosure, and omit details unrelated to the present disclosure.

[0032] To solve the technical problems of the prior art, the present disclosure provides a method for coloring a color video stream based on line drawings. Specifically, the method adopts a pre-trained optical flow model, such as a control model such as Controlnet, and an image inpainting model as a global model. Here, the control model and the image inpainting model need to be pre-trained before coloring the video stream.

[0033] The following describes a method 100 for training a model for image coloring according to an embodiment of the present disclosure with reference to Figures 1 and 2. As shown in Figure 2, the image coloring model 200 according to this embodiment includes an optical flow model 10, an inpainting model 20, and a control model 30. Here, the optical flow model 10 is pre-trained, and the training of the image coloring model actually means the training of the inpainting model 20 and the control model 30.

[0034] 1, in step 101, a correspondence relationship between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask are obtained. Specifically, in this embodiment, a line drawing 1 of the current frame and a line drawing 2 of the reference frame may be input into an optical flow model 10 to obtain a correspondence relationship between pixel points of the line drawing 1 of the current frame and pixel points of the line drawing 2 of the reference frame and an occlusion mask 5.

[0035] The optical flow model 10 may be any known optical flow model, such as Recurrent All-Pairs Field Transforms (RAFT) or Global Matching Flow (GMFlow).

[0036] Furthermore, the reference frame and the current frame refer to any two frames in the same video stream. Preferably, a frame adjacent to the current frame may be selected as the reference frame. However, the present disclosure is not limited thereto, and any frame in the video stream may be selected as the reference frame as necessary.

[0037] The occlusion mask 5 indicates an area where there is a difference between the image of the current frame and the image of the reference frame. For example, an area in the occlusion mask 5 corresponding to an animation character is an area where there is a difference between the image of the reference frame and the image of the current frame.

[0038] For example, optical flow records motion information from each pixel in a line drawing of a reference frame to the corresponding pixel in a line drawing of a current frame. Therefore, assuming that there is little change between adjacent frames or frames with small intervals, pixel points corresponding to most of the pixels in the current frame can be obtained from the reference frame.

[0039] Next, in step 102, the reference frame is warped based on the correspondence to obtain a warped image corresponding to the current frame.

[0040] Specifically, in this embodiment, the reference frame 3 is warped based on the motion information from each pixel in the line drawing 2 of the reference frame to the corresponding pixel in the line drawing 1 of the current frame calculated by the optical flow model 10, to obtain a warped image 4 corresponding to the current frame. In the warped image 4, not only the color of the occluded region (i.e., the region with differences in the occlusion mask 5) exactly corresponds to the color of the current frame, but the color of the other region exactly corresponds to the color of the current frame. Therefore, only the occluded region needs to be colored.

[0041] Next, in step 103, the line drawing of the current frame is input to the control model to extract features, and the extracted features are input to the image restoration model. Specifically, in this embodiment, the line drawing 1 of the current frame is input to the control model 30 to extract features, and the extracted features are input to the restoration model 20.

[0042] The restoration model 20 and the control model 30 may be any known model. For example, the restoration model 20 may be a stable diffusion XL-inpainting model, and the control model 30 may be the Controlnet model.

[0043] Next, the warped image and the occlusion mask are input into an image inpainting model in step 104. Specifically, in this embodiment, the warped image 4 and the occlusion mask 5 are input into the inpainting model 20.

[0044] Next, in step 105, the image inpainting model inpaints the color of the different area in the warped image based on the extracted features and the occlusion mask. Specifically, in this embodiment, the inpainting model 20 inpaints the color of the area in the warped image 4 that is different from the occlusion mask 5 based on the features extracted by the control model 30 and the occlusion mask 5.

[0045] Preferably, a prompt 7 may also be input into the repair model 20. The prompt 7 may be, for example, text expressing the author's intention.

[0046] Next, in step 106, it is determined whether the color of the colored current frame 6 output by the inpainting model 20 is accurate, and if so, the method 100 ends. Otherwise, the method 100 proceeds to step 107.

[0047] In step 107, another frame in the same video stream is selected as the current frame. Steps 101-107 are performed iteratively to train the inpainting model 20 and the control model 30 until the inpainting model 20 outputs an accurately colored current frame.

[0048] In the training process, pairs of current frames and reference frames may be preset, for example, adjacent frames of each current frame may be set as the reference frame of the current frame, or the same reference frame may be specified for all current frames, such as the first frame or any other frame in the video stream.

[0049] In addition, the loss function of the image coloring model 200 may use, for example, the loss function of the stable diffusion XL inpainting model, that is, the mean squared error (MSE) may be used to evaluate the difference between the predicted result of the image coloring model 200 and the true value. The predicted value of the model may be a noise value or a velocity value obtained based on the noise and latent variables of the image.

[0050] Preferably, in the training process, line drawings of two adjacent frames in the video stream are used as input, the previous frame is used as the reference frame, and the later frame is used as the current frame, and the image coloring model 200 outputs the coloring result of the later frame.

[0051] Preferably, training the image colorization model 200 includes performing a first training run using the velocity values ​​as predicted values ​​and a second training run using the noise values ​​as predicted values.

[0052] Since foreground characters often move, the area to be inpainted indicated by the occlusion mask 5, i.e., the area where there is a difference between the current frame and the reference frame, is usually located in the foreground character. The difficulty with coloring foreground characters is that it must comply with the creator's intention and maintain consistency throughout the entire video stream. Therefore, when training the image coloring model 200, data from a specific animation video may be used to fine-tune the inpainting model 20 and the control model 30 to ensure consistency in character coloring. In other words, a portion of the animation video may be used as training data to fine-tune the inpainting model 20 and the control model. Thus, after fine-tuning, the inpainting model 20 and the control model 30 can automatically and accurately color the remaining portion of the animation video when a reference frame is given. Through fine-tuning, the inpainting model 20 and the control model 30 can memorize the correspondence between the line drawings and colors of each character in the animation video.

[0053] Note that fine-tuning the repair model 20 means fine-tuning the U-Net encoder (Unet) in the repair model 20.

[0054] 2 may further include not only the optical flow model 10, the inpainting model 20, and the control model shown, but also parts such as a variational autoencoder (VAE), a tokenizer, a text encoder, etc. In training, both the optical flow model 10 and the above parts not shown remain unchanged.

[0055] Although not shown in FIG. 2, the control model 30 and the inpainting model 20, e.g., the stable diffusion XL inpainting model, use random noise (e.g., randomly sampled Gaussian noise) as starting latent variables to extract features from the line drawing and inpaint the warped image, respectively.

[0056] The following describes an image coloring method 300 according to an embodiment of the present disclosure with reference to FIGS.

[0057] It should be noted that the method 300 of Fig. 3 is implemented by the image coloring model 400 shown in Fig. 5, which is previously trained by the method 100 of Fig. 1. Therefore, steps 301 to 304 of the method 300 shown in Fig. 3 correspond to steps 101 to 104 of the method 100 shown in Fig. 1, and therefore, the details of steps 301 to 304 will not be described here.

[0058] In step 305, the image inpainting model inpaints the colors of the different regions in the warped image of each frame based on the extracted features and occlusion mask of each frame, and outputs the inpainted frame.

[0059] FIG. 5 is a schematic diagram illustrating an exemplary implementation flow of coloring a video stream. The video stream includes N frames, where N is an integer greater than 1. In the inference process, for convenience, the first frame 22 in the video stream may be collectively selected as a reference frame, and the line drawings 21 of the first frame through the line drawings 23 of the Nth frame may be input into the optical flow model 10. Then, image warping is performed on the first frame 22 based on the optical flow information to obtain warped images 24 and occlusion masks 25 of the first through Nth frames. The warped images 24, occlusion masks 25, and features extracted from the line drawings 21 through 23 by the control model 30′ are input together into the inpainting model 20′. Finally, the inpainting model 20′ outputs the colored frames 1 through N.

[0060] It should be noted that although FIG. 5 shows the first frame 22 as the reference frame, the present disclosure is not limited thereto and any suitable frame in the video stream may be selected as the reference frame.

[0061] Preferably, the noise model and the velocity model may be fused using a local inverse mapping method to simultaneously utilize the background stability and good restoration performance of the noise model's coloring results and the accurate character (animal) coloring results of the velocity model.

[0062] The following describes a fusion method 300' according to an embodiment of the present disclosure with reference to FIG.

[0063] After performing step 303 of method 300 of FIG. 3, step 3051 of method 300' of FIG. 4 is performed.

[0064] Specifically, in step 3051, the extracted features, warped image and occlusion mask corresponding to each frame are input into a pre-trained image inpainting model and a control model based on velocity values, thereby performing a first inpainting on the color of the different regions in the warped image based on random noise.

[0065] Note that pre-training based on velocity values ​​means using velocity values ​​as predictive values ​​to train the Unet and control model 30 in the inpainting model 20 in the image coloring model 200.

[0066] Also, performing the first repair based on random noise means using the random noise as the starting latent variable of the repair model 20 and the control model 30 .

[0067] Next, in step 3052, DDIM inverse mapping is performed on the first repaired frame to obtain the initial latent variables.

[0068] The method of obtaining latent variables by DDIM inverse mapping is known in the prior art, and therefore will not be described here.

[0069] Next, in step 3053, a segmentation mask for the line drawing of each frame is obtained. The segmentation mask indicates the area in the line drawing that contains characters.

[0070] Specifically, in this embodiment, a segmentation mask for each frame of a line drawing may be obtained by inputting the line drawing into a segmentation model. The segmentation model may be, for example, a Segment Anything Model (SAM). The segmentation mask indicates a region of the line drawing that includes, for example, an animal.

[0071] Finally, in step 3054, a second inpainting is performed on the colors of the different regions in the warped image of each frame by inputting the combination of the segmentation mask, initial latent variables, and random noise of each frame as the starting latent variables into the pre-trained image inpainting model and control model based on the noise value, and the inpainted frame is output.

[0072] Note that pre-training based on noise values ​​means using noise values ​​as predictive values ​​to train the Unet and control model 30 in the inpainting model 20 in the image colorization model 200.

[0073] In this embodiment, a segmentation mask mask indicating the character region, an initial latent variable initial latent and the random noise is the initial latent variable start of the pre-trained image coloring model 400 based on the noise value. latent Specifically, when the inpainting model 20 inpaints the warped image, it uses an initial latent variable in the character region and random noise in other regions, and the random noise in other regions is multiplied by an appropriate coefficient to balance the adverse effect of the initial latent variable. For example, the starting latent variable start is calculated according to the following formula: latent may be set.

[0074] start latent =initial latent ×mask+noise×(1-mask)×scale Equation (1) Here, the value of mask is 0 or 1, where 1 represents the character region and 0 represents other regions, and scale may be set to, for example, 0.6. Note that the present disclosure is not limited to this, and the value of scale may be set differently depending on the characteristics of different data.

[0075] The inpainting model 20 and the control model 30 perform image inpainting and line drawing feature extraction, respectively, using a combination of a segmentation mask for each frame, an initial latent variable, and random noise as the starting latent variable, as shown in Equation (1).

[0076] Steps 3052 to 3054 in Figure 4 may be implemented by the latent variable module 40 in Figure 5. Specifically, the latent variable module 40 first performs DDIM inverse mapping on the first inpainted frame to obtain initial latent variables, then uses a segmentation model to obtain a segmentation mask for each frame's line drawing, and finally inputs the combination of the segmentation mask, the initial latent variables, and random noise (Equation (1)) as the starting latent variable into the inpainting model 20' and the control model 30' pre-trained based on the noise value. The control model 30' then extracts features of the line drawing for each frame based on the combined starting latent variables, and the inpainting model 20' inpaints the warped image corresponding to each frame based on the combined starting latent variables.

[0077] (experiment) (Experimental settings) Assume an animal animation video, which includes multiple animals and multiple scenarios, specifically 27 kinds of animals and 24 scenarios (video streams). The training set and test set are classified according to the scenarios, for example, 15 scenarios are the training set and 9 scenarios are the test set. The training set includes all animals.

[0078] First, we analyze the training set at a frame rate of, say, 2 FPS, to obtain 435 images as training data. Then, we analyze the test set at a frame rate of, say, 6 FPS, to obtain smoother test video. Then, we use the same line drawing extraction algorithm to extract line drawings from the images.

[0079] To improve the accuracy of character coloring (e.g., animals in this scenario), character data was augmented in the training dataset. Specifically, each animal in the training image was segmented, and animals were randomly selected and scaled, rotated, flipped, and then re-attached to the training image. Three to five animals were randomly attached to each training image, resulting in 435 animal-augmented training images. Therefore, in this experiment, a total of 870 training images were used. During the training process, line drawings from two adjacent frames of the same scenario were input, and the previous frame was used as a reference frame to teach the image coloring model 200 how to color the current frame.

[0080] When fine-tuning the image coloring model using the above training data, for example, AdamW may be used as the optimizer, the initial learning rate may be set to 1e-5, the batch size may be set to 2, and the total number of training steps may be set to 80000. Two image coloring models were trained using noise prediction and velocity prediction, respectively.

[0081] (Experimental results) Using a noise prediction model, the line drawing video and the actual colored image of the first frame are input, and the image coloring model outputs a colored video corresponding to the entire line drawing video. Experimental results on coloring single-noise model videos show that the model can output high-quality colored images that match the content of the line drawing, accurately color the characters, and match the background color of the first frame. The single-noise model also has good adaptability. In other words, even in scenarios where the training set does not appear, the model can effectively color the line drawing video using the background indicated by the reference frame while maintaining background stability.

[0082] In scenarios with a large number of animals or where the animal's movement changes significantly, a single noise model cannot achieve good animal coloring results. Experiments have shown that the velocity prediction model has a relatively high accuracy rate for coloring animals, but the images output by the velocity prediction model are noisy and have poor image quality. Therefore, the model fusion coloring method based on the above-mentioned local inverse mapping can combine the advantages of the noise prediction model and the velocity prediction model to output good results. Experimental results show that good video coloring results can be achieved by transferring the accurate animal coloring results of the velocity model to the noise model.

[0083] Therefore, the method 100 for training a model for image coloring and the method 300 for image coloring according to the present disclosure can ensure conformance to the creator's intentions, as well as ensure consistency in timing between frames in the video stream, consistency in coloring of character objects, and stability of the background.

[0084] The above method may be implemented entirely by a computer-executable program, or partially or entirely by using hardware and / or firmware. When implemented by hardware and / or firmware, or when a computer-executable program is loaded into a hardware device capable of executing the program, a neural network training device, described below, is realized. The following will omit the above-mentioned detailed description and provide an overview of these devices. Note that these devices can execute the above method, but the method is not limited to being implemented using or by the components of the devices described below.

[0085] FIG. 6 is a schematic block diagram illustrating a training device 600 for a model for image coloring according to an embodiment of the present disclosure. The training device 600 includes an acquisition unit 601, a warping unit 602, an extraction unit 603, an inpainting unit 604, and a training unit 605. The acquisition unit 601 acquires correspondences between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream, and an occlusion mask. Here, the occlusion mask indicates a difference area between the image of the current frame and the image of the reference frame. The warping unit 602 warps the reference frame based on the correspondences to obtain a warped image corresponding to the current frame. The extraction unit 603 inputs the line drawing of the current frame into a control model to extract features. The inpainting unit 604 inputs the extracted features, the warped image, and the occlusion mask into an image inpainting model, and the image inpainting model inpaints the color of the difference area in the warped image based on the extracted features and the occlusion mask. The training unit 605 trains the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame.

[0086] The device 600 for training a model for image coloring shown in Fig. 6 corresponds to the method 100 for training a model for image coloring shown in Fig. 1. Therefore, the details of each part of the device 600 for training a model for image coloring have already been described in detail in the description of the method 100 for training a model for image coloring in Fig. 1, and therefore the description thereof will be omitted here.

[0087] FIG. 7 is a schematic block diagram illustrating an image coloring device 700 according to an embodiment of the present disclosure. The image coloring device 700 includes a first acquisition unit 701, a warping unit 702, an extraction unit 703, and a first restoration unit 704. The first acquisition unit 701 acquires a correspondence between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, as well as an occlusion mask. Here, the occlusion mask indicates areas where there are differences between the image of each frame and the image of the reference frame. The warping unit 702 warps the reference frame based on the correspondence between each frame and acquires a warped image corresponding to each frame. The extraction unit 703 inputs the line drawing of each frame into a first pre-trained control model to extract features. The first restoration unit 704 inputs the extracted features, warped image, and occlusion mask corresponding to each frame into a first pre-trained image restoration model, and performs a first restoration on the colors of the different regions in the warped image of each frame based on the extracted features and occlusion mask of each frame using the first pre-trained image restoration model, and outputs the first restored frame.

[0088] Preferably, the first pre-trained control model and the first pre-trained image inpainting model are pre-trained based on velocity values. The image coloring device 700 further includes an inverse mapping unit 705, a second acquisition unit 706, and a second inpainting unit 707. The inverse mapping unit 705 performs denoising diffusion implicit model inverse mapping on the first inpainted frame to obtain initial latent variables. The second acquisition unit 706 acquires a segmentation mask of the line drawing of each frame. The segmentation mask indicates an area containing a character in the line drawing. The second inpainting unit 707 inputs a combination of the segmentation mask, the initial latent variables, and random noise as initial latent variables into the second pre-trained image inpainting model and the second pre-trained control model, thereby performing a second inpainting on the colors of the different areas in the warped image of each frame, and outputs a second inpainted frame. Here, the second pre-trained control model and the second pre-trained image inpainting model are pre-trained based on noise values.

[0089] The image coloring device 700 shown in Fig. 7 corresponds to the image coloring method 300 shown in Fig. 3. Therefore, the details of each part of the image coloring device 700 have already been explained in detail in the explanation of the image coloring method 300 in Fig. 3, and therefore the explanation thereof will be omitted here.

[0090] The inverse mapping unit 705, the second obtaining unit 706, and the second repair unit 707 in FIG. 7 may be realized by, for example, the latent variable module 40 in FIG.

[0091] The above processes and devices may be realized by software and / or firmware. When implemented by software and / or firmware, a program constituting software for implementing the above methods may be installed from a storage medium or a network into a computer having a dedicated hardware configuration (e.g., a general-purpose personal computer 800 shown in FIG. 8), and the computer can execute various functions when various programs are installed.

[0092] 8 is a block diagram showing an exemplary configuration of a general-purpose personal computer capable of implementing a method and / or apparatus according to an embodiment of the present disclosure. As shown in FIG. 8, a central processing unit (CPU) 801 executes various processes according to programs stored in a read-only memory (ROM) 802 or programs loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 stores data necessary for the CPU 801 to execute various processes as needed. The CPU 801, the ROM 802, and the RAM 803 are connected to one another via a bus 804. An input / output interface 805 is also connected to the bus 804.

[0093] An input unit 806 (including a keyboard, a mouse, etc.), an output unit 807 (including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), a storage unit 808 (including, for example, a hard disk, etc.), and a communication unit 809 (including, for example, a network interface card, such as a LAN card, a modem, etc.) are connected to the input / output interface 805. The communication unit 809 performs communication processing via a network, such as the Internet. If necessary, a driver 810 may be connected to the input / output interface 805. A removable medium 811 is, for example, a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, and is set up in the driver 810 as necessary, and a computer program read from the removable medium 811 is installed in the storage unit 808 as necessary.

[0094] When the above processing is performed by software, a program constituting the software is installed via a network, such as the Internet, or a storage medium, such as a removable medium 811 .

[0095] 8, which stores the program and provides the program to the user separately from the device. The removable medium 811 includes, for example, a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including an optical disk-read only memory (CD-ROM) and a digital versatile disk (DVD)), a magneto-optical disk (a minidisk (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be the ROM 802, a hard disk included in the storage unit 808, or the like, which stores the program and is provided to the user together with the device containing the program.

[0096] The present disclosure further provides a computer program product having stored thereon corresponding computer program code, machine-readable instruction code, which, when readable and executed by a machine, can perform the method 100 for training a model for image coloring and the method 300 for image coloring of FIG. 3 according to embodiments of the present disclosure described above.

[0097] Accordingly, the present disclosure further includes a storage medium having recorded thereon a program product including machine-readable instruction code, including, but not limited to, a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, and the like.

[0098] Furthermore, the following supplementary notes are disclosed regarding the embodiments including the above-described examples. (Appendix 1) 1. A method for training a model for coloring an image, comprising: obtaining a correspondence between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of the current frame and an image of the reference frame; warping the reference frame based on the correspondence to obtain a warped image corresponding to the current frame; inputting the line drawing of the current frame into a control model to extract features, and inputting the extracted features into an image inpainting model; inputting the warped image and the occlusion mask into the image inpainting model; Inpainting the color of the different regions in the warped image based on the extracted features and the occlusion mask using the image inpainting model; and training the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame. (Appendix 2) 2. The method of claim 1, wherein the warped image and the occlusion mask are obtained by inputting a line drawing of the current frame and a line drawing of the reference frame into a pre-trained optical flow model. (Appendix 3) 2. The method of claim 1, wherein the reference frame is a frame adjacent to the current frame in the video stream. (Appendix 4) 2. The method of claim 1, wherein the reference frame is any frame in the video stream. (Appendix 5) 5. The method of any of claims 1 to 4, further comprising inputting a text prompt into the image inpainting model. (Appendix 6) 5. The method of any one of claims 1 to 4, wherein the step of training the image inpainting model includes training a U-Net encoder and a U-Net decoder in the image inpainting model. (Appendix 7) 5. The method of any one of claims 1 to 4, wherein the image inpainting model is a stable diffusion XL inpainting model. (Appendix 8) 8. The method of claim 7, wherein the loss function used for training comprises a mean squared error loss function. (Appendix 9) 9. The method of claim 8, wherein the step of training the image restoration model and the control model includes performing a first training using velocity values ​​as predicted values ​​and performing a second training using noise values ​​as predicted values. (Appendix 10) 1. A method for coloring an image, comprising: obtaining a correspondence between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of each frame and an image of the reference frame; warping the reference frame based on the correspondence relationship between each frame to obtain a warped image corresponding to each frame; Inputting the line drawing of each frame into a pre-trained control model to extract features, and inputting the extracted features into a pre-trained image inpainting model; inputting a warped image and an occlusion mask corresponding to each frame into the image inpainting model; and restoring colors of the different regions in the warped image of each frame based on the extracted features and occlusion mask of each frame using the image restoration model, and outputting the restored frame. (Appendix 11) 11. The method of claim 10, wherein the warped image and occlusion mask corresponding to each frame are obtained by inputting a line drawing of the frame and a line drawing of the reference frame into a pre-trained optical flow model. (Appendix 12) performing a first training run on the image inpainting model and the control model using velocity values ​​as predicted values; 12. The method of claim 10 or 11, further comprising performing a second training on the image restoration model and the control model using noise values ​​as predicted values. (Appendix 13) The step of restoring the color of the difference region in the warped image of each frame includes: performing a first inpainting of the colors of the different regions in the warped image based on random noise by inputting the extracted features, the warped image, and the occlusion mask corresponding to the frame into the first trained image inpainting model; Performing a denoising diffusion implicit model (DDIM) inverse mapping on the inpainted warped image to obtain initial latent variables; obtaining a segmentation mask for each frame of line art, the segmentation mask indicating an area of ​​the line art that includes a character; and performing a second inpainting on the colors of the different regions in the warped image by inputting a combination of the segmentation mask, the initial latent variables, and the random noise as starting latent variables into the second trained image inpainting model and a control model. (Appendix 14) 13. The method of claim 12, wherein the segmentation mask for the line drawing of each frame is obtained by inputting the line drawing into a segmentation model. (Appendix 15) 12. The method of claim 10 or 11, wherein the reference frame is a frame adjacent to the frame to be colored in the video stream, or an arbitrary frame in the video stream. (Appendix 16) 1. An apparatus for training a model for coloring an image, comprising: an acquisition unit for acquiring correspondences between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of the current frame and an image of the reference frame; a warping unit that warps the reference frame based on the correspondence relationship to obtain a warped image corresponding to the current frame; an extraction unit that inputs the line drawing of the current frame into a control model to extract features; an inpainting unit that inputs the extracted features, the warped image, and the occlusion mask into an image inpainting model, and inpaints the color of the different region in the warped image based on the extracted features and the occlusion mask using the image inpainting model; and a training unit that trains the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame. (Appendix 17) 1. An apparatus for coloring an image, comprising: a first acquisition unit that acquires a correspondence relationship between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, the occlusion mask indicating a region where there is a difference between an image of each frame and an image of the reference frame; a warping unit that warps the reference frame based on the correspondence between the frames and obtains a warped image corresponding to each frame; an extraction unit that inputs the line drawing of each frame into a first pre-trained control model to extract features; a first inpainting unit that inputs the extracted features, warped image, and occlusion mask corresponding to each frame into a first pre-trained image inpainting model, performs a first inpainting on the colors of the different regions in the warped image of each frame based on the extracted features and occlusion mask of each frame using the first pre-trained image inpainting model, and outputs the first inpainted frame. (Appendix 18) the first pre-trained control model and the first pre-trained image inpainting model are pre-trained based on velocity values; an inverse mapping unit that performs a denoising diffusion implicit model inverse mapping on the first repaired frame to obtain initial latent variables; a second acquisition unit that acquires a segmentation mask of the line drawing of each frame, the segmentation mask indicating an area including a character in the line drawing; and a second inpainting unit that performs a second inpainting on colors of different regions in the warped image of each frame by inputting a combination of the segmentation mask, the initial latent variables, and random noise as a starting latent variable into a second pre-trained image inpainting model and a second pre-trained control model, and outputs the second-inpainted frame, wherein the second pre-trained control model and the second pre-trained image inpainting model are pre-trained based on noise values. (Appendix 19) A computer-readable storage medium having a program stored thereon, the program causing a computer to perform the method of any one of claims 1 to 9 when executed by a processor. (Appendix 20) A computer-readable storage medium having a program stored thereon, the program causing a computer to perform the method of any one of claims 10 to 16 when executed by a processor.

[0099] It should be noted that the terms "comprise," "have," or any other variation thereof, are not limited to exclusive inclusion, and a process, method, article, or apparatus that includes a set of elements not only includes those elements, but also includes other elements not expressly listed or inherent elements of the process, method, article, or apparatus. Furthermore, unless further limited, a more specific term "comprises a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0100] Although the preferred embodiments of the present disclosure have been described above with reference to the drawings, the above embodiments and examples are illustrative and not limiting. Those skilled in the art may make various modifications, improvements, and equivalent changes to the present disclosure within the spirit and scope of the claims. These modifications, improvements, and equivalent changes are intended to fall within the scope of protection of the present disclosure.

Claims

1. 1. A method for training a model for coloring an image, comprising: obtaining a correspondence between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of the current frame and an image of the reference frame; warping the reference frame based on the correspondence to obtain a warped image corresponding to the current frame; inputting the line drawing of the current frame into a control model to extract features, and inputting the extracted features into an image inpainting model; inputting the warped image and the occlusion mask into the image inpainting model; Inpainting the color of the different regions in the warped image based on the extracted features and the occlusion mask using the image inpainting model; and training the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame.

2. The method of claim 1 , wherein the warped image and the occlusion mask are obtained by inputting a line drawing of the current frame and a line drawing of the reference frame into a pre-trained optical flow model.

3. The method of claim 1 , wherein the reference frame is a frame adjacent to the current frame in the video stream or any frame in the video stream.

4. The method according to claim 1 , wherein the step of training the image inpainting model comprises training a U-Net encoder and a U-Net decoder on the image inpainting model.

5. 1. A method for coloring an image, comprising: obtaining a correspondence between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of each frame and an image of the reference frame; warping the reference frame based on the correspondence relationship between each frame to obtain a warped image corresponding to each frame; Inputting the line drawing of each frame into a pre-trained control model to extract features, and inputting the extracted features into a pre-trained image inpainting model; inputting a warped image and an occlusion mask corresponding to each frame into the image inpainting model; and restoring colors of the different regions in the warped image of each frame based on the extracted features and occlusion mask of each frame using the image restoration model, and outputting the restored frame.

6. The method of claim 5 , wherein the warped image and occlusion mask corresponding to each frame are obtained by inputting a line drawing of the frame and a line drawing of the reference frame into a pre-trained optical flow model.

7. performing a first training run on the image inpainting model and the control model using velocity values ​​as predicted values; performing a second training run on the image inpainting model and the control model using noise values ​​as predicted values; The step of restoring the color of the difference region in the warped image of each frame includes: performing a first inpainting of the colors of the different regions in the warped image based on random noise by inputting the extracted features, the warped image, and the occlusion mask corresponding to the frame into the first trained image inpainting model; Performing denoising diffusion implicit model inverse mapping on the restored warped image to obtain initial latent variables; obtaining a segmentation mask for each frame of line art, the segmentation mask indicating an area of ​​the line art that includes a character; and performing a second inpainting on the colors of the different regions in the warped image by inputting a combination of the segmentation mask, the initial latent variables, and the random noise as starting latent variables into the second trained image inpainting model and a control model.

8. 1. An apparatus for training a model for coloring an image, comprising: an acquisition unit for acquiring correspondences between pixel points of a line drawing of a current frame and pixel points of a line drawing of a reference frame in a color video stream and an occlusion mask, the occlusion mask indicating areas where there are differences between an image of the current frame and an image of the reference frame; a warping unit that warps the reference frame based on the correspondence relationship to obtain a warped image corresponding to the current frame; an extraction unit that inputs the line drawing of the current frame into a control model to extract features; an inpainting unit that inputs the extracted features, the warped image, and the occlusion mask into an image inpainting model, and inpaints the color of the different region in the warped image based on the extracted features and the occlusion mask using the image inpainting model; a training unit that trains the image inpainting model and the control model by iteratively performing the above steps with two or more frames in the video stream as the current frame, so that the image inpainting model outputs an accurately colored current frame.

9. 1. An apparatus for coloring an image, comprising: a first acquisition unit that acquires a correspondence relationship between pixel points of a line drawing of each frame in a color video stream and pixel points of a line drawing of a reference frame, and an occlusion mask, the occlusion mask indicating an area where there is a difference between an image of each frame and an image of the reference frame; a warping unit that warps the reference frame based on the correspondence between the frames and obtains a warped image corresponding to each frame; an extraction unit that inputs the line drawing of each frame into a first pre-trained control model to extract features; a first restoration unit that inputs the extracted features, the warped image, and the occlusion mask corresponding to each frame into a first pre-trained image restoration model, performs a first restoration on the colors of the different regions in the warped image of each frame based on the extracted features and the occlusion mask of each frame using the first pre-trained image restoration model, and outputs the first restored frame.

10. the first pre-trained control model and the first pre-trained image inpainting model are pre-trained based on velocity values; an inverse mapping unit that performs a denoising diffusion implicit model inverse mapping on the first repaired frame to obtain initial latent variables; a second acquisition unit that acquires a segmentation mask of the line drawing of each frame, the segmentation mask indicating an area including a character in the line drawing; 10. The apparatus of claim 9, further comprising: a second inpainting unit that performs a second inpainting on colors of different regions in the warped image of each frame by inputting a combination of the segmentation mask, the initial latent variables, and random noise as starting latent variables into a second pre-trained image inpainting model and a second pre-trained control model, and outputs the second inpainted frame, wherein the second pre-trained control model and the second pre-trained image inpainting model are pre-trained based on noise values.

Citation Information

Cited By

  • Image coloring model construction method, image coloring method, equipment and medium

    CN121999090A