A Comic Image Coloring Method Based on Conditional Generative Adversarial Network
By applying the image coloring method of conditional generation adversarial network in comic production, the problems of high labor costs and low production efficiency are solved, and automated coloring and shadow processing are realized, and efficiency and quality consistency are improved.
Patent Information
- Application Number
- CN202411215068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-09-02
AI Technical Summary
The existing comic production methods have high human cost and low production efficiency, resulting in most comic products having only the cover or the first few pages in color, which consumes high time and labor costs.
The comic image coloring method based on conditional generation adversarial network is adopted, and by constructing a coloring model and shadow addition model, the coloring and shadow processing of line drafts to color drafts is automatically completed using the deep learning network.
It improves the efficiency of comic production, reduces manual intervention, ensures the consistency of the quality of the image, and maintains a good visual effect.
Smart Images

Figure CN118736061B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method for coloring comic images based on a conditional generative adversarial network. Background Art
[0002] In the process of comic production, it generally includes eight main steps: concept conception, storyboard design, pencil draft, line drawing, coloring and shading, background drawing, fonts and dialog boxes, printing and distribution, etc. Through the above methods, the production of comics can be completed relatively completely.
[0003] Currently, most of these eight steps rely on pure manual production, which is both time-consuming and laborious. Due to the time-consuming and labor-intensive nature of this work, until now, many comic products on the market only have colored pictures on the cover or the first few pages to attract the attention of buyers, which results in a relatively high consumption of time and labor costs.
[0004] In many fields of artificial intelligence, deep networks have achieved far better performance than traditional methods, including fields such as speech, natural language, vision, games, etc. Therefore, it is necessary to provide a method for coloring comic images based on a deep learning network. Summary of the Invention
[0005] The present invention provides a method for coloring comic images based on a conditional generative adversarial network, which solves the technical problems of high labor costs and low comic production efficiency in existing comic production methods.
[0006] According to the first aspect of the present invention, there is provided a method for coloring comic images based on a conditional generative adversarial network, including the following steps:
[0007] Obtain pairs of comic line drawing and color manuscript data and high-dynamic range imaging image data, and construct a training set and a test set;
[0008] Construct a coloring model and a shadow adding model, train the coloring model and the shadow adding model using the training set, and evaluate the coloring model and the shadow adding model using the test set;
[0009] Re-obtain a line drawing image, input it into the trained coloring model and shadow adding model for processing, and obtain a final image.
[0010] In the above aspect and any possible implementation, a further implementation is provided, where the training set and the test set include the pix2pix dataset and the SwitchLight dataset;
[0011] The pix2pix dataset includes multiple pairs of line drawing images and color image data;
[0012] The SwitchLight dataset includes a per-light illumination dataset and a high-dynamic range imaging dataset.
[0013] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The process of obtaining the high-dynamic range imaging dataset is as follows:
[0014] Use a light field device to construct a per-light illumination dataset;
[0015] Obtain an original high-dynamic range imaging image, combine the original high-dynamic range imaging image with the Phong reflection model, and perform convolution processing to generate a convolved high-dynamic range imaging image, thereby obtaining a high-dynamic range imaging dataset.
[0016] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The process of combining the original high-dynamic range imaging image with the Phong reflection model is as follows: Use the Phong reflection lobe as a query, and use the original high-dynamic range imaging atlas as keys and values. Integrate the original high-dynamic range imaging atlas information into the Phong reflection lobe representation through a cross-attention mechanism to complete the combination of the original high-dynamic range imaging image and the Phong reflection model.
[0017] For the aspects and any possible implementation manners described above, a further implementation manner is provided.
[0018] The coloring model includes a generator and a discriminator. The generator is used to generate a color image, and the discriminator is used to generate an image and a real image;
[0019] The generator is of a U-Net architecture, including an encoder and a decoder. The encoder reduces the spatial size of the input image and increases the depth of the feature map, and the decoder restores the image spatial size; the encoder and the decoder use skip connections;
[0020] The discriminator adopts a PatchGAN architecture, including a convolutional layer, an activation function, and an output layer connected in sequence.
[0021] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The output of the discriminator is a two-dimensional matrix;
[0022] Each element of the two-dimensional matrix is the authenticity of the corresponding image patch.
[0023] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The shadow addition model adopts a UNet architecture, including an inverse rendering network and a re-rendering network connected in sequence;
[0024] The inverse rendering network includes a convolutional layer and a transposed convolutional layer, extracts image features through a convolutional neural network, and restores the image spatial dimensions through the transposed convolutional layer;
[0025] The re-rendering network includes a convolutional layer and a transposed convolutional layer, and re-renders the image through the convolutional layer and the transposed convolutional layer.
[0026] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The processing process of the shadow addition model is as follows:
[0027] Input the training set images into the inverse rendering network for decomposition processing to obtain the inherent attributes of the images and the target illumination condition information;
[0028] Input the inherent attributes of the images and the illumination condition information into the re-rendering network for processing to obtain the rendered images under the target illumination conditions.
[0029] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The process of re-obtaining the line drawing image, inputting it into the trained model for processing, and obtaining the final image is as follows:
[0030] Re-obtain the line drawing image, input the line drawing image into the coloring model for coloring processing to obtain a high-quality colored image; input the high-quality colored image into the shadow addition model to generate an image with a shadow effect, and obtain the final image.
[0031] In the above-described aspects and any possible implementation manners, a further implementation manner is provided. The loss function of the coloring model includes a GAN loss function and an L1 loss function;
[0032] The loss function of the shadow addition model includes a reconstruction loss function, a perceptual loss function, an adversarial loss function, and a specular reflection loss function.
[0033] The present invention discloses the following technical effects:
[0034] The present invention makes technical improvements to the coloring and shadow parts in the process of comic production. By introducing advanced artificial intelligence technologies, AI can replace manual labor to complete the work from line drawing to colored drawing, thereby improving the production efficiency and ensuring quality consistency; by learning the coloring mode and shadow mode of hand-drawn images, the obtained model can be used to automatically color the input hand-drawn images, which not only requires no manual intervention but also can maintain a good visual effect.
[0035] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. Brief Description of the Drawings
[0036] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0037] Figure 1 shows a flowchart of a method for coloring comic images based on a conditional generative adversarial network according to an embodiment of the present invention;
[0038] Figure 2 shows a schematic diagram of the coloring process of a method for coloring comic images based on a conditional generative adversarial network according to an embodiment of the present invention;
[0039] Figure 3 shows a schematic diagram of the training and coloring process of a method for coloring comic images based on a conditional generative adversarial network according to an embodiment of the present invention. Detailed Description of the Embodiments
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.
[0041] To make the above objectives, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0042] Please refer to Figure 1 as shown, this embodiment provides a method for coloring comic images based on a conditional generative adversarial network, including the following steps:
[0043] S101, obtain line drawing and color drawing data pairs and high-dynamic range imaging image data, and construct a training set and a test set.
[0044] First, collect and prepare a training data set, which includes paired input images and target output images. The data set is divided into two parts, namely the pix2pix data set and the SwitchLight data set.
[0045] pix2pix Data Preparation:
[0046] Construct the AnitaDataset and comics_dataset_1024 datasets: containing 42,000 data pairs of line drawings and color pictures. At the same time, construct a self-reserved dataset: accumulated from the company's comic business, containing 92,000 data pairs of line drawings and color pictures.
[0047] SwitchLight Data Preparation:
[0048] Construct the OLAT (Optical Lighting) dataset: This dataset is constructed using a light field device, containing images of 287 objects, with each object photographed in 15 different poses, resulting in a total of 29,705 OLA sequences. Use photometric stereo to obtain normal and albedo maps.
[0049] HDRI (High Dynamic Range Imaging) Dataset: Obtain HDRI images from public archives such as Polyhaven, Noah Witchell, HDRMAPS, and iHDRI.
[0050] Perform convolution using the Phong reflection model: Combine the HDRI image with the Phong reflection model and perform convolution processing to generate a convolved HDRI image. The specific combination method is to use the Phong reflection lobe as a query, and the original image as the key and value. Integrate the image information into the Phong reflection lobe representation through a cross-attention mechanism, which can effectively simplify the HDRI reconstruction task.
[0051] S102. Construct a coloring model and a shadow addition model, train the coloring model and the shadow addition model using the training set, and evaluate the model using the test set.
[0052] This embodiment is based on the Pix2Pix architecture and uses a conditional generative adversarial network (cGAN) for the design of the coloring model. This model includes a generator (Generator) and a discriminator (Discriminator). The generator adopts a U-Net structure, and the discriminator adopts a PatchGAN structure.
[0053] Specifically, the generator adopts a U-Net architecture, including an encoder and a decoder. The encoder gradually reduces the spatial size of the input image through a series of convolutional layers while increasing the depth of the feature map; the decoder gradually restores the spatial size of the image through a series of deconvolutional layers while combining the feature maps of the encoder. A skip connection is used between the encoder and the decoder to retain high-resolution feature information.
[0054] The discriminator adopts the PatchGAN architecture, divides the image into several small patches, and independently judges whether each small patch is real. The discriminator includes several convolutional layers, and extracts local features of the image through a series of convolutional layers. Its activation function uses the LeakyReLU activation function and outputs a two-dimensional matrix, where each element represents the authenticity of the corresponding image patch.
[0055] The loss function of the colorization model includes the GAN loss function and the L1 loss function, which can effectively ensure the authenticity of the generated image and the retention of details.
[0056] Specifically, the GAN loss function is expressed as:
[0057] (1);
[0058] Among them, G is the function of the Generator, which is responsible for generating the output image G(x,z) from the input image and random noise. The goal of the generator is to generate realistic images that are difficult for the discriminator to distinguish between real and fake; D is the function of the Discriminator, which is responsible for distinguishing whether the input image is a real image or a fake image G(x,z) generated by the generator. The goal of the discriminator is to maximize its accuracy, that is, to distinguish between real images and generated images as accurately as possible; x is the input image, usually the original image or conditional image obtained from the dataset; y is the real target image, which is the real output image corresponding to the input image. During training, the discriminator uses y as the real sample for learning;
[0059] z is a random noise vector used to introduce randomness into the generator. Generally speaking, the noise vector z can help the generator generate more diverse output images; E[.] is the expected value, which represents the average operation on relevant variables over the entire training dataset.
[0060] The L1 loss function represents the pixel-level error between the output of the generator and the real image, ensuring the consistency of image details. Specifically, it is:
[0061] (2);
[0062] For the shadow addition model, in this embodiment, the SwitchLight model design is adopted, and the UNet architecture is used to implement the inverse rendering and re-rendering processes. The model design is combined with a physically-driven rendering model and a self-supervised pre-training framework. The shadow addition model simulates the interaction between light and surface microfacets by adopting the Cook-Torrance reflection model and uses the multi-mask autoencoder (MMAE) for self-supervised pre-training to extract image features.
[0063] The shadow addition model specifically includes an inverse rendering network and a relighting network. The inverse rendering network (Inverse Rendering Network) uses a convolutional neural network (CNN) to extract image features, gradually restores the image spatial dimensions through deconvolution layers, and decomposes the inherent attributes of the image, including surface normal, albedo, roughness, lighting conditions, etc. by inputting the image and performing deconvolution processing.
[0064] The relighting network (Relighting Network) uses convolutional layers and deconvolution layers to combine with a physical model for image reconstruction, and obtains an image re-rendered under the target lighting conditions by processing the attributes output by the inverse rendering network and the target lighting conditions.
[0065] The rendering of this embodiment combines diffuse reflection and specular reflection to calculate the final image brightness. The specific equation is:
[0066] (3);
[0067] where, L o (v) is the outgoing radiance, representing the light intensity seen from the viewing direction v; Ω is the hemispherical space, indicating that the integration is performed within the hemisphere around the surface normal n; f r (v, l) is the bidirectional reflectance distribution function (BRDF, Bidirectional Reflectance Distribution Function), describing the reflection characteristics of light from the incident direction l to the outgoing direction v; L i (l) is the incident radiance, representing the light intensity incident on the surface from the direction l; n is the surface normal, representing the direction of the surface at this point; l is the incident light direction; (n * l) is the cosine value of the incident angle, representing the angle between the incident light and the surface normal, describing the influence degree of the light on the surface.
[0068] The shadow addition model of this embodiment is also provided with a reconstruction loss function, a perceptual loss function, an adversarial loss function, and a specular loss function. Specifically:
[0069] The reconstruction loss function (Reconstruction Loss) can measure the pixel-level error between the generated image and the real image. Specifically:
[0070] (4);
[0071] where, I gtis the Ground Truth Image; I pred is the PredictedImage; ||*|| is the L1 norm, which is used to measure the absolute error between two images.
[0072] The Perceptual Loss extracts features based on a pre-trained VGG network and measures the differences in high-level features, specifically:
[0073] (5);
[0074] Among them, is the Feature Map extracted by a pre-trained network such as VGG; is the square of the L2 norm, which is used to measure the Euclidean distance in the feature space.
[0075] The Adversarial Loss can promote the realism of the generated image so that it cannot be distinguished by the discriminator, specifically:
[0076] (6);
[0077] Among them, D( ) is the discriminator network, which is used to distinguish between the real image and the generated image; E is the Expectation, which is used to calculate the average.
[0078] The Specular Loss can enhance the specular highlight details in the generated image, specifically:
[0079] (7);
[0080] Among them, S( ) is the operation or network for extracting the specular reflection highlight area in the image; ||*||1 is the L1 norm, which is used to measure the differences between the specular reflection areas.
[0081] During model training, the training dataset is used to iteratively optimize the model and update the model parameters until convergence.
[0082] After training is completed, model evaluation is carried out. By inputting the test line drawing image, the corresponding color image is generated through the coloring model. By inputting the generated color image, the final image with shadow effects is generated through the shadow addition model. The performance of the model is comprehensively evaluated through quantitative metrics (such as MAE, MSE, SSIM, and LPIPS) and qualitative evaluations (such as user studies and expert reviews).
[0083] S103. Re-obtain the line drawing image, input it into the trained model for processing, and obtain the final image.
[0084] Input the line drawing image and use the coloring model for processing to generate a high-quality color image. The specific steps are as follows: Preprocess the input line drawing image to ensure that the image size and resolution meet the model input requirements. Use the trained coloring model for image generation and output a color image.
[0085] Input the generated color image into the shadow addition model to generate the final image with a shadow effect. The specific steps are as follows: Input the preprocessed color image to ensure that the image size and resolution meet the model input requirements. Use the trained shadow addition model for image processing to generate the final image with a shadow effect. The obtained image is as Figure 2 shown.
[0086] As Figure 3 shown, in this embodiment, technical improvements are made to the coloring and shadow parts in the comic production process. By introducing advanced artificial intelligence technology, AI can replace manual labor to complete the work from line drawing to color drawing, thereby improving production efficiency and ensuring quality consistency; by learning the coloring mode and shadow mode of hand-drawn images, the obtained model can be used to automatically color the input hand-drawn images. Not only does it not require manual intervention, but it can also maintain a good visual effect.
[0087] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0088] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. No limitation is imposed herein.
[0089] The above specific implementation manners do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A comic image coloring method based on conditional generative adversarial network, characterized in that: The following steps are involved: Obtain comic line and color draft data pairs and high dynamic range imaging image data, and build training and test sets; Wherein, the training set and the test set include a pix2pix data set and a SwitchLight data set; The pix2pix dataset includes multiple pairs of line drawings and color images; The SwitchLight dataset includes a light-by-light illumination dataset and a high dynamic range imaging dataset; The acquisition process of the high dynamic range imaging data set is as follows: Use light field equipment to build a light-by-light illumination dataset; Acquire an original high dynamic range imaging image, combine the original high dynamic range imaging image with a Phong reflection model, and perform convolution processing to generate a convolved high dynamic range imaging image, thereby obtaining a high dynamic range imaging data set; The process of combining the original high dynamic range imaging image with the Phong reflection model is as follows: using the Phong reflection lobe as a query and the original high dynamic range imaging atlas as a key and a value, integrating the high dynamic range imaging atlas information into the Phong reflection lobe representation through a cross attention mechanism, and completing the combination of the original high dynamic range imaging image with the Phong reflection model; Constructing a coloring model and a shading model, training the coloring model and the shading model using the training set, and evaluating the coloring model and the shading model using the test set; Re-acquiring the line drawing image and inputting it into the trained coloring model and the shadow adding model for processing to obtain the final image, including: Reacquiring the line draft image, inputting the line draft image into the coloring model for coloring, and obtaining a high-quality color image; inputting the high-quality color image into the shadow adding model, generating an image with a shadow effect, and obtaining a final image; The shadow adding model adopts the UNet model architecture, including an inverse rendering network and a re-rendering network connected in sequence; The inverse rendering network includes a convolution layer and a deconvolution layer, extracts image features through a convolutional neural network, and restores the image space size through a deconvolution layer; The re-rendering network includes a convolution layer and a deconvolution layer, and the image is re-rendered through the convolution layer and the deconvolution layer; The shadow adding model simulates the interaction between light and surface microfacets by adopting the Cook-Torrance reflection model and uses a multi-mask autoencoder for self-supervised pre-training to extract image features; The shadow adding model specifically includes an inverse rendering network and a re-rendering network; the inverse rendering network uses a convolutional neural network to extract image features, gradually restores the image space size through a deconvolution layer, and decomposes the inherent properties of the image by inputting an image and performing deconvolution processing, including surface normals, albedo, roughness, and lighting conditions; The re-rendering network uses convolutional layers and deconvolutional layers combined with physical models to reconstruct images. By processing the attributes and target lighting conditions output by the inverse rendering network, the re-rendered image under the target lighting conditions is obtained. The coloring model adopts the Pix2Pix model architecture, including a generator and a discriminator, wherein the generator is used to generate a color image, and the discriminator is used to distinguish between the generated image and the real image; The generator is a U-Net architecture, including an encoder and a decoder, wherein the encoder reduces the spatial size of the input image and increases the feature map depth, and the decoder restores the image spatial size; the encoder and the decoder use a jump connection; The discriminator adopts the PatchGAN architecture, including a convolutional layer, an activation function and an output layer connected in sequence.
2. The method for coloring comic images based on conditional generative adversarial networks according to claim 1, characterized in that: The output of the discriminator is a two-dimensional matrix; each element of the two-dimensional matrix is the authenticity of the corresponding image patch.
3. The method for coloring comic images based on conditional generative adversarial networks according to claim 2, characterized in that: The processing process of the shadow adding model is: Inputting the training set images into the inverse rendering network for decomposition processing to obtain image intrinsic attributes and target lighting condition information; The inherent properties of the image and the lighting condition information are input into the re-rendering network for processing to obtain a rendered image under target lighting conditions.
Citation Information
Patent Citations
Image scene relighting network structure and method based on GAN network
CN115578497A