Comicization Model Construction Method, Device, Equipment, Storage Medium and Program Product
By constructing a comic-based model based on the generative model and using sample images to fit the initial model, the problem of insufficient diversity and flexibility of comic-based processing in the existing technology is solved, and a more efficient full-picture comic-based effect is achieved.
Patent Information
- Application Number
- CN202111356773.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-16
AI Technical Summary
In the prior art, the comic images generated by the image comic processing are single, not flexible enough, insufficient diversity, poor user characteristics similarity, and weak comic sense.
The sample real map is generated using the pre-trained first generative model, the second generative model is constructed and the sample comic map is generated, and the sample set is formed by combining the sample image pairs. The initial comic model is fitted using the weight of the second generative model to generate a comic model for converting the target image into a full-picture comic image.
It improves the robustness and generalization of comic-based models, improves the effect of comic-based whole pictures, and reduces the data volume requirement.
Smart Images

Figure CN114067052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular, to a method for constructing a cartoonization model, a device for constructing a cartoonization model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] The technology of image cartoonization is one of the common tasks in image editing in computer vision, and it is widely used in life and entertainment. For example, in scenarios such as film production, animation production, short videos, and live broadcasts, images are processed into cartoons.
[0003] In the related art, the implementation methods of image cartoonization are as follows:
[0004] One is the processing method based on basic signals. This method mainly constructs a material library, and through various relevant basic signals, such as height, weight, hair color, clothing color, etc., the most suitable material is matched in the material library, and then the matched materials are combined into an anime image. This method has disadvantages such as a single image, lack of flexibility, insufficient diversity, and poor similarity of user characteristics.
[0005] Another is the processing method of texture face pinching. This method deforms the real person's face into the shape of an anime face through deformation, and then realizes image cartoonization by pasting various materials, such as anime faces, eyes, eyebrows, etc. However, the effect of this method is single, the anime images constructed by different people are similar, with poor diversity, weak cartoon sense, and poor authenticity. Summary of the Invention
[0006] The present application provides a method, device, equipment, storage medium, and program product for constructing a cartoonization model to solve the problems in the prior art that the generated cartoon images in cartoonization processing are single in image, lack of flexibility, insufficient diversity, poor similarity of user characteristics, and weak cartoon sense.
[0007] In a first aspect, an embodiment of the present application provides a method for constructing a cartoonization model, the method comprising:
[0008] Generating a preset number of sample real images by using a pre-trained first generation model;
[0009] Constructing a second generation model based on the first generation model, and generating corresponding sample cartoon images for each sample real image by using the second generation model;
[0010] Combining the sample real images with the corresponding sample cartoon images into sample image pairs;
[0011] Based on a sample set composed of multiple pairs of the sample images, using the weights corresponding to the second generation model as the initial weights, fitting a preset initial cartoonization model to generate a cartoonization model for converting a target image into a full-image cartoonized image.
[0012] In a second aspect, an embodiment of the present application further provides a device for constructing a cartoonization model. The device includes:
[0013] A sample real-image generation module, configured to generate a preset number of sample real-images by using a pre-trained first generation model;
[0014] A sample cartoon-image generation module, configured to construct a second generation model based on the first generation model and generate sample cartoon-images corresponding to the respective sample real-images by using the second generation model;
[0015] An image pair combination module, configured to combine the sample real-images with the corresponding sample cartoon-images into sample image pairs;
[0016] A cartoonization model fitting module, configured to, based on a sample set composed of multiple pairs of the sample images, use the weights corresponding to the second generation model as the initial weights to fit a preset initial cartoonization model to generate a cartoonization model for converting a target image into a full-image cartoonized image.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device. The electronic device includes:
[0018] One or more processors;
[0019] A storage device, configured to store one or more programs,
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the first aspect or the second aspect above.
[0021] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the methods in the first aspect or the second aspect above are implemented.
[0022] In a fifth aspect, an embodiment of the present application further provides a computer program product. The computer program product includes computer-executable instructions, and the computer-executable instructions are used to implement the methods in the first aspect or the second aspect above when executed.
[0023] The technical solution provided by the present application has the following beneficial effects:
[0024] In this embodiment, when constructing a caricature model for converting a target image into a full-image caricature image, first, a preset number of sample real images are randomly generated by a pre-trained first generation model. Then, a second generation model for generating caricature images is constructed based on the first generation model, and the second generation model is used to generate sample caricature images corresponding to each sample real image. A sample set is obtained by combining the sample real images with the corresponding sample caricature images into sample image pairs. Next, using the weights corresponding to the second generation model as the initial weights, the preset initial caricature model is fitted with this sample set. The fitted model is the caricature model, which can perform full-image caricature processing. The second generation model in this embodiment is associated with the first generation model, and the weights of the second generation model are used as the initial weights of the caricature model. By using the image pair method to obtain image pairs as training data, the fitting of the caricature model is realized, so that the finally obtained caricature model has higher robustness and generalization ability, and the effect of full-image caricature is improved. In addition, this embodiment requires less data volume than other solutions. Description of the Drawings
[0025] Figure 1 is a flowchart of an embodiment of a method for constructing a caricature model provided in Embodiment 1 of the present application;
[0026] Figure 2 is a schematic diagram of the effect of full-image caricature processing of an image based on the caricature model provided in Embodiment 1 of the present application;
[0027] Figure 3 is a schematic diagram of the model architecture of a StyleGAN2 model provided in Embodiment 1 of the present application;
[0028] Figure 4 is a flowchart of an embodiment of a method for constructing a caricature model provided in Embodiment 2 of the present application;
[0029] Figure 5 is a flowchart of an embodiment of a method for constructing a caricature model provided in Embodiment 3 of the present application;
[0030] Figure 6 is a flowchart of an embodiment of a method for constructing a caricature model provided in Embodiment 4 of the present application;
[0031] Figure 7 is a schematic diagram of the architecture of an initial caricature model provided in Embodiment 4 of the present application;
[0032] Figure 8 is a flowchart of an embodiment of a method for constructing a caricature model provided in Embodiment 5 of the present application;
[0033] Figure 9It is a structural block diagram of an apparatus embodiment for constructing a caricature model provided in Embodiment 6 of the present application;
[0034] Figure 10 It is a schematic structural diagram of an electronic device provided in Embodiment 7 of the present application. Detailed implementation manners
[0035] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. In addition, it should be noted that for the convenience of description, only parts related to the present application rather than all structures are shown in the drawings.
[0036] Embodiment 1
[0037] Figure 1 It is a flowchart of a method embodiment for constructing a caricature model provided in Embodiment 1 of the present application. This method can be implemented by a caricature model construction device. Among them, the caricature model construction device can be located in a server or a client, and this embodiment does not limit this.
[0038] The caricature model constructed in this embodiment can be applied to scenarios such as image processing, short videos, movie production, live broadcasts, 3D cartoons, etc., and is used to process images in scenarios such as image processing, short videos, movies, live broadcasts, etc. into full-image caricature images. For example, as Figure 2 shown, given an image, input the image into the caricature model, and the caricature model can output the full-image caricatured image (that is, the full-image caricature image). The content of the full-image caricature image and the input image remains the same, but it becomes the style of a caricature, that is, all elements in the input image are caricatured. Another example is that given a short video, input each frame image of the short video into the caricature model, and the caricature model can output the full-image caricature images corresponding to each frame image to realize full-image caricaturing of each frame image in the input short video.
[0039] The caricature model constructed in this embodiment can be accessed to an APP or a web page according to the development document.
[0040] As Figure 1 shown, this embodiment may include the following steps:
[0041] Step 110, generate a preset number of sample real images by using a pre-trained first generation model.
[0042] A generative model is an important type of model in probability statistics and machine learning, referring to a model used to randomly generate observable data. Exemplarily, the first generative model can be a StyleGAN2 (Style Generative Adversarial Network 2) model, and the StyleGAN2 model can be used to randomly generate a preset number of genuine sample images. Among them, the genuine sample images can be images that have not been caricatured, for example, images containing real people.
[0043] GAN (Generative Adversarial Networks) is a deep learning model and a generative model capable of generating new content. StyleGAN is a type of GAN and a style-based generative model. StyleGAN is an advanced high-resolution image synthesis method that has been proven to work reliably on various datasets. In addition to realistic portraits, StyleGAN can also be used to generate other animals, cars, and even rooms. However, StyleGAN is not perfect. The most obvious defect is that the generated images sometimes contain speckle-like artifacts, and this defect has been overcome by StyleGAN2, thus further improving the quality of the generated images.
[0044] As Figure 3 shown in the model architecture of the StyleGAN2 model, StyleGAN2 consists of two parts, including Figure 3 the left part, the Mapping Network, and the right part, the synthesis network.
[0045] The Mapping Network can better disentangle the input. As Figure 3 shown, the Mapping Network is composed of 8 fully connected layers (FC for short). Its input is Gaussian noise (latent Z), and through the Mapping Network, a latent variable (W) is obtained.
[0046] The synthesis network is composed of modules such as a learnable affine transformation A, a modulation module Mod-Demod, and an upsampling Upsample. In addition, the synthesis network also includes weights (w), biases (b), and a constant input (c, that is, Const 4*4*512, representing a learnable constant). The activation function (Leaky ReLU) is always applied immediately after adding the bias.
[0047] Among them, the learnable affine transformation A can be composed of a fully connected layer; Upsample can use transposed convolution (also called deconvolution) for upsampling operations.
[0048] The processing flow of the modulation module Mod-Demod is as follows:
[0049] w′ ijk = s i · w ijk
[0050] Among them, s i is the scaling ratio of the i-th input feature map;
[0051] After scaling and convolution, demodulate the weights of the convolutional layer, and the standard deviation of the output activation is:
[0052]
[0053] Demodulate the demod weights, aiming to restore the output to the unit standard deviation, that is, the weights of the new convolutional layer are:
[0054]
[0055] In the above formula, adding ∈ is to avoid the denominator being 0.
[0056] Figure 3 The rightmost in is the injection of random noise. B is a learnable noise parameter. Introducing random noise is to make the generated image more realistic. For example, these noises can generate tiny features of the face when generating, such as freckles on the face.
[0057] Step 120, construct a second generation model based on the first generation model, and use the second generation model to generate sample comic images corresponding to each sample real image.
[0058] In this step, the second generation model can also be the StyleGAN2 model. The model architectures of the first generation model and the second generation model are the same. The difference lies in that different training objectives lead to different weights of the models. The training model of the first generation model is to generate sample real images, that is, images that are not cartoonized. And the training model of the second generation model is to generate sample comic images, that is, cartoonized images.
[0059] In one implementation, the weights of the pre-trained first generation model can be fine-tuned with the cartoon data as the training objective to obtain the cartoonized second generation model. Then use the second generation model to generate sample comic images corresponding to each sample real image.
[0060] Step 130, combine the sample real image and the corresponding sample comic image into a sample image pair.
[0061] In this step, after generating a preset number of sample real images and their corresponding sample comic images, each sample real image and its corresponding sample comic image can be combined to form a sample image pair (picture TO picture, abbreviated as P2P). All the sample image pairs form a sample set for subsequent model fitting.
[0062] It should be noted that the preset number can be an empirical value, and the specific value of the preset number can be determined according to the required accuracy of the model. For example, the preset number can be 150,000, that is, 150,000 pairs of sample image pairs are generated.
[0063] Step 140: Based on the sample set composed of multiple said sample image pairs, using the weights corresponding to the second generation model as the initial weights, fit a preset initial cartoonization model to generate a cartoonization model for converting a target image into a full-image cartoonized image.
[0064] In this step, after obtaining the sample set composed of sample image pairs, the sample real images in this sample set can be used as training data, the weights corresponding to the finally generated second generation model can be used as the initial weights, and the sample comic images corresponding to each sample real image can be used as the optimization target to fit the preset initial cartoonization model. Finally, a well-fitted cartoonization model is obtained. This cartoonization model is used to convert a target image into a full-image cartoonized image.
[0065] In this embodiment, when constructing a cartoonization model for converting a target image into a full-image cartoonized image, first, a preset number of sample real images are randomly generated by a pre-trained first generation model. Then, a second generation model for generating comic images is constructed based on this first generation model, and the second generation model is used to generate sample comic images corresponding to each sample real image. A sample set is obtained by combining the sample real images with the corresponding sample comic images into sample image pairs. Then, using the weights corresponding to the second generation model as the initial weights, the preset initial cartoonization model is fitted with this sample set. The well-fitted model is the cartoonization model, which can achieve full-image cartoonization processing. The second generation model in this embodiment is associated with the first generation model, and the weights of the second generation model are used as the initial weights of the cartoonization model. The image pairs are used as training data in the form of image pairs to fit the cartoonization model, so that the finally obtained cartoonization model has higher robustness and generalization ability, and improves the effect of full-image cartoonization. In addition, the data volume required in this embodiment is less than that of other solutions.
[0066] Embodiment 2
[0067] Figure 4The flowchart of a method embodiment for constructing a cartoonization model provided in Embodiment 2 of the present application. This embodiment further specifically describes the process of constructing the second generation model on the basis of Embodiment 1. As Figure 4 shown, this embodiment may include the following steps:
[0068] Step 410, generating a preset number of sample real images by using a pre-trained first generation model.
[0069] Step 420, adjusting the weights of the first generation model to generate an intermediate cartoon model.
[0070] In this embodiment, the training target of the first generation model is the original image without cartoonization, while the training target of the intermediate cartoon model is the cartoon image after cartoonization processing. Therefore, the weights of the first generation model can be used as the initial weights of the intermediate cartoon model, and the cartoon image is used as the training target to generate the intermediate cartoon model. In this way, the weights of the intermediate cartoon model are obtained by adjusting the weights of the first generation model.
[0071] Step 430, replacing the weights corresponding to some specified layers in the intermediate cartoon model with the weights of the first generation model corresponding to the some specified layers, and performing weight interpolation to generate a second generation model.
[0072] In order to ensure that some attributes of the finally output cartoon image are consistent with the attributes in the original image generated by the first generation model, after generating the intermediate cartoon model, the weights corresponding to some specified layers in the intermediate cartoon model can also be replaced with the weights of the first generation model corresponding to the some specified layers, and weight interpolation is performed to generate a second generation model.
[0073] For example, the some specified layers may include at least one of the following: the layer for controlling the pose of the character, the layer for controlling the skin color of the character. That is to say, in order to ensure that the pose and skin color of the character after cartoonization are consistent with the pose and skin color of the real person in the original image, after obtaining the intermediate cartoon model, the weights of the layer for controlling the pose of the character and the layer for controlling the skin color of the character in the intermediate cartoon model can be replaced with the weights of the layer for controlling the pose of the character and the layer for controlling the skin color of the character in the first generation model, and weight interpolation is performed in the intermediate cartoon model to finally obtain the weights of the second generation model.
[0074] Weight interpolation refers to calculating a new weight between two weights using an interpolation algorithm and inserting the new weight between the two weights. In this embodiment, the specific interpolation algorithm for weight interpolation is not limited. For example, it may include inverse distance weighted (IDW) interpolation for weight interpolation. Inverse distance weighted interpolation can also be called the method of multiplying the reciprocal of the distance by the grid. It is a weighted average interpolation method that can interpolate in an exact or smooth manner. The inverse distance weighted (IDW) interpolation explicitly assumes that things that are closer to each other are more similar than things that are farther apart. When predicting values for any unmeasured location, the inverse distance weighted method uses the measured values around the predicted location. Compared with the measured values that are farther from the predicted location, the measured values that are closest to the predicted location have a greater impact on the predicted value. The inverse distance weighted method assumes that each measurement point has a local influence, and this influence decreases with the increase of distance. Since this method assigns a larger weight to the point closest to the predicted location, and the weight decreases as a function of distance, it is called the inverse distance weighted method.
[0075] In addition, those skilled in the art can also use covariance-based weight interpolation algorithms, Kriging interpolation methods, etc. for weight interpolation.
[0076] Step 440: Use the second generation model to generate sample comic images corresponding to each sample real image.
[0077] Step 450: Combine the sample real image and the corresponding sample comic image into a sample image pair.
[0078] Step 460: Based on the sample set composed of multiple sample image pairs, use the weights corresponding to the second generation model as the initial weights to fit a preset initial cartoonization model, and generate a cartoonization model for converting a target image into a full-image cartoonized image.
[0079] In this embodiment, when constructing the second generation model, the pre-trained first generation model is used as the basis, and the weights of the first generation model are adjusted with the cartoon image as the training target to obtain an intermediate cartoon model. Then, partial layer weight replacement and weight interpolation are performed on the intermediate cartoon model to obtain the weights of the final second generation model, thereby completing the construction of the second generation model. Compared with simply training the second generation model with the cartoonized image as the training target, the second generation model constructed in the above manner of this embodiment has higher robustness and improves the authenticity of the cartoonized image.
[0080] Embodiment Three
[0081] Figure 5The flowchart of a method embodiment for constructing a caricature model provided in Embodiment 3 of this application. This embodiment further specifically describes the processing process of training samples on the basis of Embodiment 1 or Embodiment 2. As Figure 5 shown, this embodiment may include the following steps:
[0082] Step 510, generate a preset number of sample real images using a pre-trained first generation model.
[0083] Step 520, construct a second generation model based on the first generation model, and use the second generation model to generate sample caricature images corresponding to each sample real image.
[0084] Step 530, combine the sample real images and the corresponding sample caricature images into sample image pairs.
[0085] Step 540, perform data augmentation on a sample set composed of multiple sample image pairs, where the data augmentation includes at least one of randomly rotating the sample real images and the sample caricature images by a random angle, randomly cropping, randomly magnifying, randomly reducing, etc.
[0086] In this step, all sample image pairs can be combined into a sample set, and then data augmentation is performed on the sample real images and sample caricature images in the sample set to increase the amount of training data and improve the robustness and generalization ability of the model.
[0087] In implementation, the data augmentation may include at least one of the following methods: various noise augmentations, augmentations such as downsampling first and then upsampling, data augmentation of sample image pairs, and so on.
[0088] Exemplarily, the data augmentation of sample image pairs may include but is not limited to: randomly rotating the sample real images and / or the sample caricature images by a random angle, randomly cropping, randomly magnifying, randomly reducing, etc.
[0089] Step 550, based on the sample set, use the weights corresponding to the second generation model as the initial weights to fit a preset initial caricature model, and generate a caricature model for converting a target image into a full-image caricature image.
[0090] After obtaining the sample set through data augmentation, the sample set can be used, with the weights corresponding to the second generation model as the initial weights, to fit a preset initial caricature model to generate a caricature model.
[0091] In this embodiment, after obtaining a sample real image and the corresponding sample comic image to form a sample image pair, the sample image pair can be used as training data to form a sample set. Then, various data augmentation methods are applied to the sample set, and the sample set after data augmentation is used to train the comicization model to implement the full-image comicization technology. This can make the comicization model more robust, making the model robust to objects (such as human objects) at any angle, and having strong generalization ability for various scenes. The full-image comicization effect for various low-quality images is still relatively good.
[0092] Embodiment 4
[0093] Figure 6 The flowchart of a method embodiment for constructing a comicization model provided in Embodiment 4 of this application. Based on Embodiment 1 or Embodiment 2 or Embodiment 3, this embodiment further specifically describes the construction process of the comicization model. As Figure 6 shown, this embodiment may include the following steps:
[0094] Step 610, use a pre-trained first generation model to generate a preset number of sample real images.
[0095] Step 620, construct a second generation model based on the first generation model, and use the second generation model to generate sample comic images corresponding to each sample real image.
[0096] Step 630, combine the sample real image and the corresponding sample comic image to form a sample image pair.
[0097] Step 640, use the encoder in a preset initial comicization model to extract features from the sample real images in the sample set to obtain corresponding feature maps and style attribute information, and output the feature maps and the style attribute information to the decoder of the initial comicization model.
[0098] In one implementation, the initial comicization model may include an encoder Encoder and a decoder Decoder. As Figure 7 shown, the left dashed box part is Encoder, and the right dashed box part is Decoder. The function of Encoder is to extract information from each sample real image, and output the extracted feature maps and style attribute information to Decoder. Decoder combines the feature maps and the style attribute information to output a full-image comicized image.
[0099] The initial weights of the Encoder in this embodiment are the weights of an encoder that has previously edited various real-person images.
[0100] In one embodiment, the structure of the Encoder may include: an input layer, a number of residual layers, and a fully-connected layer. Among them, each residual layer is used to extract the feature map from the sample real image and output the feature map to the corresponding layer of the decoder, and the fully-connected layer is used to extract the style attribute information of the sample real image and output the style attribute information to each layer of the decoder.
[0101] For example, the structure of the Encoder is shown in Table 1 below. In Table 1, there are 5 residual layers (ResBlock), and the size of the feature map (Featuremap) output by each residual layer is specified, such as 512*512*3, 256*256*32, etc. in Table 1. The fully-connected layer FC outputs style attribute information of size 16*512.
[0102]
[0103]
[0104] Table 1
[0105] As Figure 7 shown, for the feature map extracted by each residual layer, on the one hand, it is output to the next layer for processing, and on the other hand, it also needs to be output to the corresponding layer of the Decoder (except for the last residual layer, which only outputs the result to the corresponding layer of the Decoder). The corresponding layer here refers to the decoding layer that matches the size of the currently output feature map. For example, if the size of the currently output feature map is 32*32*512, the corresponding layer in the Decoder refers to the decoding layer that can process the feature map of size 32*32*512.
[0106] In Figure 7 it, for the two output layers on the rightmost side of the Encoder, the upper one is the last residual layer ResBlock, which outputs a feature map of size 16*16*512; the lower one is the FC layer, which outputs style attribute information of size 16*512. The FC layer outputs the style attribute information to each layer of the Decoder so that the Decoder can perform full-image anime processing based on the style attribute information.
[0107] Step 650, using the decoder, taking the sample comic images in the sample set as the training target, using the weights of the second generation model as the initial weights, and training the feature map and the style attribute information with a preset loss function to obtain a caricaturing model.
[0108] In one implementation, the structure of the decoder Decoder is the same as the structure of the synthesis network of the second generation model StyleGAN2 model, and is trained with the weights of the second generation model as the initial weights.
[0109] As Figure 7 shown, for each decoding layer of the Decoder, after obtaining the feature map and style attribute information input by the Encoder, the feature map and the style attribute information are decoded and synthesized, and the decoding result is output to the next layer, and so on, and the fully rendered cartoon image is output by the last decoding layer.
[0110] In one embodiment, the loss function used to train the cartoonization model may include the combination of the following loss functions: the adversarial network loss function GAN loss , the perceptual loss function perceptual loss , and the regression loss function L1 loss , that is:
[0111] Loss = GAN loss + perceptual loss + L1 loss
[0112] Among them, the adversarial network loss function GAN loss is a classification loss function, which is used to judge the authenticity of the fully rendered cartoon image generated by the cartoonization model, and calculate the loss according to the judgment result, so that the cartoon sense of the fully rendered cartoon image generated by the cartoonization model is more realistic.
[0113] In one implementation, GAN can be calculated using the following formula loss :
[0114] GAN loss = E[D(G(x)-1) 2 + E[D(G(x)) 2
[0115] Among them, D represents the discriminator, E represents the mean value, and G(x) represents the fully rendered cartoon image output by the cartoonization model.
[0116] The perceptual loss function perceptual loss is used to input the fully rendered cartoon image output by the cartoonization model and the corresponding sample cartoon image in the sample set into a preset neural network model respectively, obtain the corresponding first feature map and second feature map output by the neural network model, and calculate the L2 loss between the first feature map and the second feature map.
[0117] Exemplarily, the preset neural network model can be a VGG model, such as VGG-19 or VGG-16, etc.
[0118] In one implementation, perceptual can be calculated using the following formulaloss :
[0119] perceptual loss = E((VGG(x) - VGG(G(x)) 2 )
[0120] Wherein, E represents the mean value, G(x) represents the full-image cartoonized image output by the cartoonization model, and x represents the sample cartoon image corresponding to the original sample image input to the cartoonization model.
[0121] L1 loss For calculating the L1 loss between the full-image cartoonized image output by the cartoonization model and the corresponding sample cartoon image in the sample set, it can be expressed by the following formula:
[0122] L1 loss = E(x - G(x))
[0123] It should be noted that for the design of the loss function in this embodiment, in addition to the combination of the three loss functions listed above, other loss functions can also be adopted according to the actual optimization goal, and this embodiment does not limit this.
[0124] In this embodiment, the initial cartoonization model includes an encoder and a decoder. When fitting the initial cartoonization model, the initial weights of the encoder are the weights of the encoder that has previously edited various real-person images, and the initial weights of the decoder are the weights of the second generation model. Using the above model architecture, with the paired data formed by the sample real image and the sample cartoon image as the training data, combined with the adversarial network loss function GAN loss 、perceptual loss function perceptual loss and the regression loss function L1 loss Three loss functions are used to fit the cartoonization model, so that the fitted cartoonization model can better extract the feature map and sample attribute information of the image through the encoder, and perform full-image cartoonization processing on the feature map and sample attribute information through the decoder, making the cartoon sense of the full-image cartoonized image output by the cartoonization model stronger, and the content after full-image cartoonization more consistent with the real image, better improving the robustness and generalization ability of the cartoonization model, and can be applied to low-quality images and complex scenes.
[0125] Embodiment Five
[0126] Figure 8 This is a flowchart of a method embodiment for constructing a cartoonization model provided by Embodiment Five of this application. Based on Embodiment One or Embodiment Two or Embodiment Three or Embodiment Four of this application, the inference process of the cartoonization model is described in more detail. As Figure 8 shown, this embodiment may include the following steps:
[0127] Step 810: Generate a preset number of sample real images using a pre-trained first generation model.
[0128] Step 820: Construct a second generation model based on the first generation model, and use the second generation model to generate sample comic images corresponding to each sample real image.
[0129] Step 830: Combine the sample real images with the corresponding sample comic images into sample image pairs.
[0130] Step 840: Based on a sample set composed of multiple sample image pairs, using the weights corresponding to the second generation model as the initial weights, fit a preset initial cartoonization model to generate a cartoonization model.
[0131] Step 850: Obtain a target image, and input the target image into the cartoonization model.
[0132] In one example, the target image may include: an image input via an image editing page. For example, after opening an image editing page in an image editing application or an application with image editing functions, the image imported by the user is used as the target image. When the user triggers the full-image cartoonization function on the image editing page, the full-image cartoonization technology of the present application can be used to perform full-image cartoonization processing on the image.
[0133] In another example, the target image may further include: each image frame in a target video. For example, in a live broadcast scenario, when the user triggers the full-image cartoonization function on the live broadcast interface, the full-image cartoonization technology of the present application can be used to perform full-image cartoonization processing on each image frame in the live broadcast video. Also, in a short video or video playback scenario, when the user triggers the full-image cartoonization function on the playback interface, the full-image cartoonization technology of the present application can be used to perform full-image cartoonization processing on each image frame in the video.
[0134] Step 860: In the cartoonization model, the encoder extracts features from the target image to extract the target feature map and target style attribute information of the target image, and inputs the target feature map and the target style attribute information into the decoder; the decoder generates a corresponding full-image cartoonization image based on the target feature map and the target style attribute information, and outputs the full-image cartoonization image.
[0135] In this embodiment, after the cartoonization model obtains the target image through the input layer of the encoder, the input layer inputs the target image into, such as Figure 7In the first residual layer of the encoder shown, the first residual layer extracts the feature map of the target image and inputs it into the next residual layer and the corresponding layer of the decoder. Then, the next residual layer continues to extract features, and so on, until the last residual layer and the FC layer are processed. At this point, the work of the encoder is completed. Then, it comes to the work of the decoder. In each layer of the decoder, based on the received target feature map and target style attribute information, caricaturing processing is performed, and the processing result is transmitted to the next layer for processing, and so on, until the last decoding layer outputs the full-image caricatured image to the output layer, and the output layer outputs the full-image caricatured image. Thus, the work of the decoder is completed. Then, the processing of the next target image can be carried out.
[0136] In this embodiment, the full-image caricaturing technology is realized through the encoder and decoder of the caricaturing model. While maintaining the style of the real image unchanged, the full-image caricaturing style is strong, the sense of caricature is real, and the immersion is high, which is suitable for various different caricaturing styles.
[0137] Embodiment Six
[0138] Figure 9 The structure block diagram of an apparatus embodiment for constructing a caricaturing model provided in Embodiment Six of this application may include the following modules:
[0139] The sample real-image generation module 910 is used to generate a preset number of sample real images by using a pre-trained first generation model;
[0140] The second generation module construction module 920 is used to construct a second generation model based on the first generation model;
[0141] The sample caricature-image generation module 930 is used to generate sample caricature images corresponding to each sample real image by using the second generation model;
[0142] The image pairing module 940 is used to combine the sample real image and the corresponding sample caricature image into a sample image pair;
[0143] The caricaturing model fitting module 950 is used to fit a preset initial caricaturing model based on a sample set composed of a plurality of the sample image pairs, with the weights corresponding to the second generation model as the initial weights, to generate a caricaturing model for converting a target image into a full-image caricatured image.
[0144] In one embodiment, the second generation module construction module 920 is specifically used for:
[0145] Adjust the weights of the first generation model to generate an intermediate caricature model;
[0146] Replace the weights corresponding to the specified layers in the middle comic model with the weights of the first generation model corresponding to the specified layers, and perform weight interpolation to generate a second generation model.
[0147] In one embodiment, the specified layers include at least one of the following: a layer for controlling the pose of a character, a layer for controlling the skin color of a character.
[0148] In one embodiment, the initial cartoonization model includes an encoder and a decoder;
[0149] The cartoonization model fitting module 950 may include the following sub-modules:
[0150] An encoding sub-module for extracting features from the sample real images in the sample set using the encoder to obtain corresponding feature maps and style attribute information, and outputting the feature maps and the style attribute information to the decoder;
[0151] A decoding sub-module for using the decoder with the sample cartoon images in the sample set as the training target, using the weights of the second generation model as the initial weights, and training the feature maps and the style attribute information using a preset loss function to obtain a cartoonization model.
[0152] In one embodiment, the loss function includes a combination of the following loss functions: an adversarial network loss function, a perceptual loss function, and a regression loss function L1_loss;
[0153] The adversarial network loss function is used to judge the authenticity of the full-image cartoonization image generated by the cartoonization model and calculate the loss according to the judgment result;
[0154] The perceptual loss function is used to input the full-image cartoonization image output by the cartoonization model and the corresponding sample cartoon image in the sample set into a preset neural network model respectively, obtain the corresponding first feature map and second feature map output by the neural network model, and calculate the L2 loss between the first feature map and the second feature map;
[0155] The L1_loss is used to calculate the L1 loss between the full-image cartoonization image output by the cartoonization model and the corresponding sample cartoon image in the sample set.
[0156] In one embodiment, the structure of the encoder is as follows:
[0157] An input layer, a plurality of residual layers, and a fully connected layer, wherein each residual layer is used to extract a feature map from a sample real image and output the feature map to a corresponding layer of a decoder, and the fully connected layer is used to extract style attribute information of the sample real image and output the style attribute information to each layer of the decoder.
[0158] In one embodiment, the initial weights of the encoder are the weights of an encoder that has previously edited various real human images.
[0159] In one embodiment, the second generation model is a StyleGAN2 model, and the structure of the decoder is the same as the structure of the synthesis network of the StyleGAN2 model.
[0160] In one embodiment, the apparatus may further include the following modules:
[0161] A target image acquisition module, configured to acquire a target image and input the target image into the cartoonization model;
[0162] A full-image cartoonization processing sub-module, configured to, in the cartoonization model, extract features of the target image by the encoder to extract a target feature map and target style attribute information of the target image, and input the target feature map and the target style attribute information into the decoder; and generate a corresponding full-image cartoonization image based on the target feature map and the target style attribute information by the decoder, and output the full-image cartoonization image.
[0163] In one embodiment, the target image includes at least one of the following:
[0164] An image input via an image editing page;
[0165] Each image frame in a target video.
[0166] In one embodiment, the apparatus may further include the following modules:
[0167] A data augmentation module, configured to perform data augmentation on the sample set before using the sample set for model fitting, wherein the data augmentation includes at least one of randomly rotating the sample real image and the sample cartoon image by a random angle, randomly cropping, randomly magnifying, and randomly reducing.
[0168] An apparatus for page rendering provided by an embodiment of the present application can execute a method for page rendering in Embodiment 1 or Embodiment 2 of the present application, and has corresponding functional modules and beneficial effects for executing the method.
[0169] Embodiment 7
[0170] Figure 10The following is a schematic structural diagram of an electronic device provided in the seventh embodiment of the present application. As shown in Figure 10 the figure, the electronic device includes a processor 1010, a memory 1020, an input device 1030, and an output device 1040. The number of processors 1010 in the electronic device can be one or more. Figure 10 In this example, one processor 1010 is taken as an example. The processor 1010, the memory 1020, the input device 1030, and the output device 1040 in the electronic device can be connected through a bus or other means. Figure 10 In this example, connection through a bus is taken as an example.
[0171] The memory 1020, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to any one of the above-mentioned Embodiments 1 to 5 in the present application embodiment. The processor 1010 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 1020, that is, implements the methods described in any one of the above-mentioned Method Embodiments 1 to 5.
[0172] The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the terminal, etc. In addition, the memory 1020 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 1020 can further include a memory remotely set relative to the processor 1010, and these remote memories can be connected to the device / terminal / server through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0173] The input device 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device. The output device 1040 can include a display device such as a display screen.
[0174] Embodiment 8
[0175] The present application embodiment 8 also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the methods described in any one of the above-mentioned Method Embodiments 1 to 5 when executed by a computer processor.
[0176] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present application, the computer-executable instructions are not limited to the method operations described above, and can also execute relevant operations in the methods provided by any embodiment of the present application.
[0177] Embodiment Nine
[0178] Embodiment Nine of the present application further provides a computer program product, which includes computer-executable instructions that are used to execute the method of any one of Embodiments 1 to 5 above when executed by a computer processor.
[0179] Of course, for a computer program product provided by an embodiment of the present application, the computer-executable instructions are not limited to the method operations described above, and can also execute relevant operations in the methods provided by any embodiment of the present application.
[0180] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, and includes several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0181] It should be noted that in the embodiments of the above device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present application.
[0182] Note that the above is only a preferred embodiment of the present application and the applied technical principle. Those skilled in the art will understand that the present application is not limited to the specific embodiments described here, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the appended claims.
Claims
1. A method for constructing a caricature model, characterized in that The method includes: Generating a preset number of sample real images using a pre-trained first generation model; Constructing a second generation model based on the first generation model and generating corresponding sample comic images for each sample real image using the second generation model; Combining the sample real images with the corresponding sample comic images into sample image pairs; Based on a sample set composed of multiple sample image pairs, using the weights corresponding to the second generation model as the initial weights, fitting a preset initial cartoonization model to generate a cartoonization model for converting a target image into a full-image cartoonized image; The initial cartoonization model includes an encoder and a decoder; The step of, based on a sample set composed of multiple sample image pairs, using the weights corresponding to the second generation model as the initial weights, fitting a preset initial cartoonization model to generate a cartoonization model for converting a target image into a full-image cartoonized image, includes: Using the encoder to extract features from the sample real images in the sample set to obtain corresponding feature maps and style attribute information, and outputting the feature maps and the style attribute information to the decoder; Using the decoder with the sample comic images in the sample set as the training target, using the weights of the second generation model as the initial weights, and training the feature maps and the style attribute information using a preset loss function to obtain a cartoonization model.
2. The method according to claim 1, wherein The step of training the second generation model based on the first generation model includes: Adjusting the weights of the first generation model to generate an intermediate comic model; Replacing the weights corresponding to some specified layers in the intermediate comic model with the weights of the first generation model corresponding to the some specified layers, and performing weight interpolation to generate a second generation model.
3. The method according to claim 2, wherein The some specified layers include at least one of the following: a layer for controlling the pose of a person, a layer for controlling the skin color of a person.
4. The method according to claim 1, wherein The loss function includes a combination of the following loss functions: an adversarial network loss function, a perceptual loss function, and a regression loss function L1_loss; The adversarial network loss function is used to judge the authenticity of the full-image cartoonized image generated by the cartoonization model and calculate the loss according to the judgment result; The perceptual loss function is used to input the full-image cartoonized image output by the cartoonization model and the corresponding sample comic image in the sample set into a preset neural network model respectively, obtain corresponding first feature maps and second feature maps output by the neural network model, and calculate the L2 loss between the first feature maps and the second feature maps; The L1_loss is used to calculate the L1 loss between the full-image cartoonized image output by the cartoonization model and the corresponding sample comic image in the sample set.
5. The method according to claim 1, wherein The structure of the encoder is as follows: An input layer, several residual layers, and a fully connected layer. Among them, each residual layer is used to extract the feature maps in the sample real images and output the feature maps to the corresponding layers of the decoder, and the fully connected layer is used to extract the style attribute information of the sample real images and output the style attribute information to each layer of the decoder.
6. The method according to claim 5, wherein The initial weights of the encoder are the weights of an encoder that has previously edited various real human images.
7. The method according to claim 1, characterized in that, The second generation model is a StyleGAN2 model, and the structure of the decoder is the same as the structure of the synthesis network of the StyleGAN2 model.
8. The method according to claim 1, characterized in that, It further includes: Obtain a target image and input the target image into the cartoonization model; In the cartoonization model, the encoder extracts features from the target image to extract the target feature map and target style attribute information of the target image, and inputs the target feature map and the target style attribute information into the decoder; The decoder generates a corresponding full-image cartoonized image based on the target feature map and the target style attribute information, and outputs the full-image cartoonized image.
9. The method according to claim 8, wherein The target image includes at least one of the following: An image input via an image editing page; Each image frame in a target video.
10. The method according to claim 1, wherein The method further includes: Before fitting the model using the sample set, perform data augmentation on the sample set, where the data augmentation includes at least one of randomly rotating, randomly cropping, randomly magnifying, and randomly reducing the sample real image and the sample cartoon image at random angles.
11. An apparatus for constructing a caricature model, characterized in that, The device includes: A sample real image generation module for generating a preset number of sample real images using a pre-trained first generation model; A second generation model construction module for constructing a second generation model based on the first generation model; A sample cartoon image generation module for generating sample cartoon images corresponding to each sample real image using the second generation model; An image pairing module for combining the sample real image and the corresponding sample cartoon image into a sample image pair; A cartoonization model fitting module for fitting a preset initial cartoonization model based on a sample set composed of multiple sample image pairs, using the weights corresponding to the second generation model as the initial weights, to generate a cartoonization model for converting a target image into a full-image cartoonized image; The initial cartoonization model includes an encoder and a decoder; The cartoonization model fitting module includes: An encoding sub-module for extracting features from the sample real images in the sample set using the encoder to obtain corresponding feature maps and style attribute information, and outputting the feature maps and the style attribute information to the decoder; A decoding sub-module for using the decoder to train the feature maps and the style attribute information with the sample cartoon images in the sample set as the training target, using the weights of the second generation model as the initial weights, and using a preset loss function to obtain a cartoonization model.
12. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-10.
14. A computer program product comprising computer-executable instructions which, when executed, are configured to implement the method according to any one of claims 1-10.
Citation Information
Patent Citations
Training method of image generation model, generating method and device and equipment
CN112862669A
Caricaturization model construction method and apparatus, and device, storage medium and program product
WO2023088276A1