A method, apparatus, electronic device and storage medium for generating a comic face

By generating comic faces through the generation of adversarial networks, the problem of poor comic face generation effect in complex scenes is solved, high-quality image conversion is achieved, data preparation process is simplified, and it is suitable for real-time applications.

CN115330589BActive Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210963016.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-07-11
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The existing comic face generation methods have poor effects in complex scenarios and rough image conversion. Traditional methods require professional knowledge and are cost-effective, while deep learning-based methods need to match data sets, which are difficult and can only convert face parts.

Method used

Generative adversarial network is used to generate comic faces, and the target area is obtained through face detection, and mutual game learning is used between generator and discriminator, combining the content generation module and attention mask generation module to generate high-quality comic face images.

Benefits of technology

It improves the authenticity of comic face generation effects and image conversion details in complex scenes, reduces the dependence on professional knowledge, reduces the difficulty of data set preparation, and enables the conversion of faces and other elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115330589B_ABST
    Figure CN115330589B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, electronic device and storage medium for generating a comic face, which is used to solve the technical problems that the existing comic face generation methods have poor comic face generation effects in complex scenarios and rough image conversion. The present invention includes: obtaining a face image to be converted; performing face detection on the face image to be converted to obtain a target face area; inputting the target face area into a preset generative adversarial network to generate a target comic face image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face conversion, and in particular, to a method, apparatus, electronic device, and storage medium for generating a cartoon face. Background Technique

[0002] Comics are an art form widely accessed by people. With the development of digital technology, animations have become increasingly popular, especially with the development of basic platform technologies such as live streaming, VR, and the metaverse, which have further promoted the development of animation technology. However, to obtain a cartoon-style portrait, traditional methods require manual painting of a cartoon-style portrait, which is a task with high costs and extremely high workload, and this method is not applicable to real-time live streaming application scenarios.

[0003] Since the generative adversarial network was proposed, it has been widely used in various tasks. Among them, image-to-image conversion is a very important application field. The goal of the image conversion task is to establish a mapping from the source domain of the image to the target domain, which is widely used in multiple fields such as image super-resolution reconstruction, style transfer, image coloring, and image dehazing, and has great application value. After years of development, there are a series of excellent image conversion algorithms in the field of image conversion, such as algorithms like Pix2pix, DualGAN, DiscoGAN, UNIT, MUNIT, DRIT, CycleGAN, etc. Among them, the Pix2pix algorithm belongs to a supervised image conversion algorithm, which requires a matching image set to complete the image conversion task. However, the production of the matching data set is difficult and costly. To achieve the conversion between unpaired images, network models such as DualGAN, DiscoGAN, and CycleGAN introduce a cyclic consistency constraint. Experimental results show that the above models have achieved good results in unpaired image conversion tasks. Although algorithms such as CycleGAN perform excellently in unpaired image conversion tasks, due to the generator needing to maintain the background area of the image unchanged while also converting the target foreground, the efficiency of the generator for converting the target image is low.

[0004] In existing solutions for generating cartoon faces, they can mainly be divided into two categories. One is the cartoon face generation solution based on traditional image algorithms, and the other is the cartoon face generation solution based on deep learning.

[0005] The first category: The advantage of the comic face generation scheme based on traditional image algorithms is that it can generate comic faces of corresponding styles without a large amount of training data sets. However, designing tasks based on traditional image algorithms is very complex. It requires not only algorithm designers to have a very deep reserve of professional knowledge in the field of computer science, but also professional knowledge in the field of art. Moreover, the success rate of traditional image algorithms in recognizing and segmenting human faces or other key parts is relatively low. Therefore, the generation effect of comic faces in complex scenarios is poor, the success rate is low, and the application scenarios are not extensive.

[0006] The second category: The comic face generation scheme based on deep learning first obtains a target local image including a human face from the image to be processed, then inputs the target local image into a trained image conversion generation model based on the Pix2pix (generative adversarial network) algorithm to obtain a comic face image. Finally, the content of the human face area in the target local image is replaced with the content of the human face area in the comic face image to obtain the target comic face image. Among them, the Pix2pix algorithm belongs to a supervised image conversion algorithm, and this algorithm requires a matching image group to complete the task of image conversion, that is, a large number of one-to-one real human faces and comic face training sets are required. However, the production of the matching data set is difficult and costly. Moreover, this scheme can only perform comic conversion on the human face part and cannot realize comic style conversion of other elements such as hair. Summary of the Invention

[0007] The present invention provides a comic face generation method, device, electronic device and storage medium, which are used to solve the technical problems that the existing comic face generation methods have poor comic face generation effects and rough image conversion in complex scenarios.

[0008] The present invention provides a comic face generation method, including:

[0009] Obtain the face image to be converted;

[0010] Perform face detection on the face image to be converted to obtain a target face area;

[0011] Input the target face area into a preset generative adversarial network to generate a target comic face image.

[0012] Optionally, the generation of the preset generative adversarial network includes:

[0013] Obtain a face image from a preset real face data set, and perform face detection on the face image to obtain a face area;

[0014] Obtain a comic face image from a preset comic face data set, and perform face detection on the comic face image to obtain a comic face area;

[0015] Construct an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator;

[0016] Generate a first comic face generation image of the face area through the generator;

[0017] Generate a first face generation image of the comic face area through the generator;

[0018] Use the discriminator to calculate a first probability score of the first comic face generation image, and use the discriminator to calculate a second probability score of the first face generation image;

[0019] Use the discriminator, the first comic face generation image and the first face generation image to generate a generator loss function of the generator;

[0020] Generate a first optimization parameter of the generator using the generator loss function;

[0021] Calculate the cross-entropy error using the first probability score and the second probability score to obtain a discriminator loss function of the discriminator;

[0022] Generate a second optimization parameter of the discriminator using the discriminator loss function;

[0023] Update the generator using the first optimization parameter and update the discriminator using the second optimization parameter;

[0024] Determine whether a preset iteration termination condition is satisfied;

[0025] If not, return to the step of generating a first comic face generation image of the face area through the generator and generating a first face generation image of the comic face area;

[0026] If so, generate a generative adversarial network using the updated generator and the updated discriminator.

[0027] Optionally, the generator includes a content generation module and an attention mask generation module; the step of generating a first comic face generation image of the face area through the generator includes:

[0028] Extract the face feature information of the face area;

[0029] Transform the face feature information to obtain face feature transformation information;

[0030] Perform feature encoding on the face feature transformation information to obtain face feature decoding information;

[0031] Input the face feature decoding information into the content generation module to generate a first content map;

[0032] Input the face feature decoding information into the attention mask generation module to generate a first foreground mask map and a first background mask map;

[0033] Multiply the first content map and the first foreground mask map to generate a first foreground map;

[0034] Multiply the face region and the first background mask map to generate a first background map;

[0035] Add the first foreground map and the first background map to generate a first generated cartoon face image.

[0036] Optionally, the step of generating the first generated face image of the cartoon face region by the generator includes:

[0037] Extract the cartoon face feature information of the cartoon face region;

[0038] Convert the cartoon face feature information to obtain cartoon face feature conversion information;

[0039] Perform feature encoding on the cartoon face feature conversion information to obtain cartoon face feature decoding information;

[0040] Input the cartoon face feature decoding information into the content generation module to obtain a second content map;

[0041] Input the cartoon face feature decoding information into the attention mask generation module to generate a second foreground mask map and a second background mask map;

[0042] Multiply the second content map and the second foreground mask map to generate a second foreground map;

[0043] Multiply the cartoon face region and the second background mask map to generate a second background map;

[0044] Add the second foreground map and the second background map to generate a first generated face image.

[0045] Optionally, the discriminator includes an auxiliary discriminator and a final discriminator; the step of generating the generator loss function of the generator by using the discriminator, the first generated cartoon face image and the first generated face image includes:

[0046] Calculate the first generated face adversarial loss of the auxiliary discriminator by using the first generated cartoon face image;

[0047] Calculate the first generated cartoon face adversarial loss of the auxiliary discriminator by using the first generated face image;

[0048] Use the generated image of the first comic face to calculate the second face generation adversarial loss of the final discriminator;

[0049] Use the generated image of the first face to calculate the second comic face generation adversarial loss of the final discriminator;

[0050] Use the generated image of the first face and the generated image of the first comic face to generate an overall image cycle consistency loss;

[0051] Obtain the background mask cycle consistency loss;

[0052] Use the first face generation adversarial loss, the first comic face generation adversarial loss, the second face generation adversarial loss, the second comic face generation adversarial loss, the overall image cycle consistency loss, and the background mask cycle consistency loss to generate the generator loss function of the generator.

[0053] Optionally, after the step of using the updated generator and the updated discriminator to generate the generative adversarial network, it further includes:

[0054] Input a preset test image into the generative adversarial network to obtain a generated test image;

[0055] Calculate the evaluation index of the generated test image;

[0056] Generate a measurement result of the generative adversarial network according to the evaluation index.

[0057] The present invention also provides a comic face generation device, including:

[0058] A to-be-converted face image acquisition module, configured to acquire a to-be-converted face image;

[0059] A face detection module, configured to perform face detection on the to-be-converted face image to obtain a target face region;

[0060] A target comic face image generation module, configured to input the target face region into a preset generative adversarial network to generate a target comic face image.

[0061] Optionally, the generation of the preset generative adversarial network includes:

[0062] A face region acquisition module, configured to acquire a face image from a preset real face dataset and perform face detection on the face image to obtain a face region;

[0063] A comic face region acquisition module, configured to acquire a comic face image from a preset comic face dataset and perform face detection on the comic face image to obtain a comic face region;

[0064] An initial generative adversarial network construction module for constructing an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator;

[0065] A first comic face generated image generation module for generating a first comic face generated image of the face region through the generator;

[0066] A first human face generated image generation module for generating a first human face generated image of the comic face region through the generator;

[0067] A probability score calculation module for calculating a first probability score of the first comic face generated image using the discriminator, and calculating a second probability score of the first human face generated image using the discriminator;

[0068] A generator loss function generation module for generating a generator loss function of the generator using the discriminator, the first comic face generated image, and the first human face generated image;

[0069] A first optimization parameter generation module for generating a first optimization parameter of the generator using the generator loss function;

[0070] A discriminator loss function calculation module for calculating cross-entropy error using the first probability score and the second probability score to obtain a discriminator loss function of the discriminator;

[0071] A second optimization parameter generation module for generating a second optimization parameter of the discriminator using the discriminator loss function;

[0072] A discriminator update module for updating the generator using the first optimization parameter and updating the discriminator using the second optimization parameter;

[0073] An iteration judgment module for judging whether a preset iteration termination condition is satisfied;

[0074] A return module for, if not, returning to the step of generating a first comic face generated image of the face region through the generator and generating a first human face generated image of the comic face region;

[0075] A generative adversarial network generation module for, if so, generating a generative adversarial network using the updated generator and the updated discriminator.

[0076] The present invention also provides an electronic device, the device including a processor and a memory:

[0077] The memory is used for storing program code and transmitting the program code to the processor;

[0078] The processor is used to execute the comic face generation method described in any one of the above according to the instructions in the program code.

[0079] The present invention also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the comic face generation method described in any one of the above.

[0080] As can be seen from the above technical solutions, the present invention has the following advantages: The present invention discloses a comic face generation method, including: obtaining a face image to be converted; performing face detection on the face image to be converted to obtain a target face area; inputting the target face area into a preset generative adversarial network to generate a target comic face image. Thereby improving the generation effect of comic faces in complex scenarios and having more realistic image conversion details. Description of the Drawings

[0081] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0082] Figure 1 It is a flowchart of the steps of a comic face generation method provided by an embodiment of the present invention;

[0083] Figure 2 It is a training schematic diagram of a generative adversarial network provided by an embodiment of the present invention;

[0084] Figure 3 It is a network model of a generator provided by an embodiment of the present invention;

[0085] Figure 4 It is a network structure of a discriminator provided by an embodiment of the present invention;

[0086] Figure 5 It is a block diagram of the structure of a comic face generation device provided by an embodiment of the present invention. Detailed Embodiments

[0087] Embodiments of the present invention provide a comic face generation method, device, electronic device and storage medium, which are used to solve the technical problems that the existing comic face generation methods have poor comic face generation effects in complex scenarios and rough image conversion.

[0088] To make the objectives, features, and advantages of the present invention more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0089] Please refer to Figure 1 , Figure 1 which is a flowchart of the steps of a method for generating a cartoon face provided in an embodiment of the present invention.

[0090] A method for generating a cartoon face provided by the present invention may specifically include the following steps:

[0091] Step 101, obtain a face image to be converted;

[0092] Step 102, perform face detection on the face image to be converted to obtain a target face region;

[0093] In the embodiment of the present invention, the Mtcnn face detection algorithm can be used to perform face detection to obtain the target face region and reduce irrelevant image conversion regions.

[0094] Mtcnn (Multi-task Cascaded Convolutional Networks) is a deep cascaded multi-task framework. This framework is used to solve the problem of face detection and alignment in an unconstrained environment due to various poses, illuminations, and occlusions. It can simultaneously complete the tasks of face detection and face alignment. Compared with traditional algorithms, it has better performance and faster detection speed.

[0095] Step 103, input the target face region into a preset generative adversarial network to generate a target cartoon face image.

[0096] A generative adversarial network is a deep learning model. The model generates better outputs through the mutual game learning of (at least) two modules in the framework: the generative model and the discriminative model. The discriminative model needs to input variables and predict through a certain model. The generative model is to randomly generate observed data given some implicit information.

[0097] The present invention obtains a face image to be converted; performs face detection on the face image to be converted to obtain a target face region; and inputs the target face region into a preset generative adversarial network to generate a target cartoon face image, thereby improving the generation effect of cartoon faces in complex scenarios and the authenticity of image conversion details.

[0098] Please refer toFigure 2 , Figure 2 This is a training schematic diagram of a generative adversarial network provided by an embodiment of the present invention. This embodiment is an explanation of the training process of the generative adversarial network involved in the above embodiment. Specifically, it may include the following steps:

[0099] Step 201: Obtain a face image from a preset real face dataset, and perform face detection on the face image to obtain a face region;

[0100] Step 202: Obtain a cartoon face image from a preset cartoon face dataset, and perform face detection on the cartoon face image to obtain a cartoon face region;

[0101] In an embodiment of the present invention, a face image can be obtained from a real face dataset to be converted, and a cartoon face image can be obtained from a cartoon face dataset, and face detection is respectively performed on the face image and the cartoon face image to obtain a face region and a cartoon face region.

[0102] In one example, the Selfie2Anime training set provided by the CycleGAN network can be selected, which contains 6800 images of face selfies and cartoon faces.

[0103] Step 203: Construct an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator;

[0104] Step 204: Generate a first cartoon face generated image of the face region through the generator;

[0105] In an embodiment of the present invention, after obtaining the face region, the face region can be fed into the generator to generate a first cartoon face generated image of the face region.

[0106] In one example, the generator includes a content generation module and an attention mask generation module; the step of generating a first cartoon face generated image of the face region through the generator includes:

[0107] S41: Extract the face feature information of the face region;

[0108] S42: Convert the face feature information to obtain face feature conversion information;

[0109] S43: Perform feature encoding on the face feature conversion information to obtain face feature decoding information;

[0110] S44: Input the face feature decoding information into the content generation module to generate a first content map;

[0111] S45: Input the face feature decoding information into the attention mask generation module to generate a first foreground mask map and a first background mask map;

[0112] S46, Multiply the first content image and the first foreground mask image to generate a first foreground image;

[0113] S47, Multiply the face region with the first background mask image to generate a first background image;

[0114] S48, Add the first foreground image and the first background image to generate a first generated caricature face image.

[0115] In an embodiment of the present invention, the face region can be input into a generator network. The face feature information of the image is extracted through the feature encoding part of the generator network; then, the face feature information is subjected to feature transformation to generate face feature transformation information; then, the face feature transformation information is subjected to feature encoding to generate face feature decoding information; the generated face feature decoding information is then respectively input into a content generation module and an attention mask generation module. The content generation module generates a first content image, and the attention mask generation module generates a first foreground mask image and a first background mask image. Then, multiplying the first content image by the first foreground mask image generated by the attention mask generation module can obtain a first foreground image, and multiplying the first background mask image generated by the attention mask generation module by the face region can obtain a first background image; finally, adding the first foreground image and the first background image can obtain a first generated caricature face image corresponding to the face region generated by the generator.

[0116] In one example, the network model of the generator is as Figure 3 shown, where the feature encoding, feature transformation, feature decoding, and content generation modules are consistent with the CycleGAN network. The process of generating a first generated caricature face image of the face region is as follows:

[0117] 1), Input a 3D image x (face region) in the X domain (face domain) to the generator G. After passing through three convolutional layers, an instance normalization layer, and a ReLu activation function, 256-dimensional image feature encoding, that is, face feature information, is extracted;

[0118] 2), Input the 256-dimensional face feature information in 1) into a feature transformation module composed of 9 residual blocks, and also output 256-dimensional transformed features (face feature transformation information); then input the 256-dimensional face feature transformation information into a feature decoding module composed of two transposed convolutional layers, an instance normalization layer, and a ReLu activation function, and output 128-dimensional decoded features, that is, face feature decoding information.

[0119] 3), Input the 128-dimensional decoded face feature decoding information into a content generation module composed of padding, convolution, and Tanh to generate a first content image Meanwhile, the decoded face feature decoding information of 128 dimensions is input into an attention mask generation module composed of convolution and a Softmax activation function to generate a first foreground mask map respectively. and a first background mask map The numerical ranges of the first foreground mask map and the first background mask map are between [0, 1].

[0120] 4), Multiply the generated first content map and the aforementioned foreground mask map one by one and then sum them up to obtain an output of a 3-dimensional foreground map, that is, the first foreground map; after multiplying the input image x by the first background mask map, an output of a 3-dimensional background map is obtained, that is, the first background map; adding the first foreground map and the first background map can obtain a 3-dimensional output image, that is, the first comic face generation image.

[0121] The mathematical expression of the above generator is shown in the following formula. A generator can generate a content map and an attention mask map (foreground mask map and background mask map ). The content map is multiplied by the foreground attention mask to generate a foreground map, the background attention mask is multiplied by the input image to generate a background map, and the foreground map and the background map are added to obtain the final output image.

[0122]

[0123] Among them, x represents the face region image. Set N to 10, then N - 1 = 9, which means that the 9-dimensional content image matrix is multiplied by the corresponding foreground mask image matrix and then accumulated, and b is a constant 1.

[0124] Step 205, generate a first face generation image of the comic face region through the generator;

[0125] In the embodiment of the present invention, after obtaining the comic face region, the comic face region can be sent into the generator to generate a first face generation image of the comic face region.

[0126] In one example, the step of generating a first face generation image of the comic face region through the generator may include the following sub-steps:

[0127] S51, extract the comic face feature information of the comic face region;

[0128] S52, convert the comic face feature information to obtain comic face feature conversion information;

[0129] S53, perform feature encoding on the comic face feature conversion information to obtain comic face feature decoding information;

[0130] S54. Input the decoded information of the comic face features into the content generation module to obtain a second content image;

[0131] S55. Input the decoded information of the comic face features into the attention mask generation module to generate a second foreground mask image and a second background mask image;

[0132] S56. Multiply the second content image by the second foreground mask image to generate a second foreground image;

[0133] S57. Multiply the comic face region by the second background mask image to generate a second background image;

[0134] S58. Add the second foreground image and the second background image to generate a first face generation image.

[0135] In an embodiment of the present invention, the comic face region can be sent into a generator network. The comic face feature information of the image is extracted through the feature encoding part of the generator network; then, the comic face feature information is subjected to feature conversion to generate comic face feature conversion information; then, the comic face feature conversion information is subjected to feature encoding to generate decoded comic face feature information; the generated decoded comic face feature information is respectively put into the content generation module and the attention mask generation module. The content generation module generates a second content image, and the attention mask generation module generates a second foreground mask image and a second background mask image. Then, multiplying the second content image by the second foreground mask image generated by the attention mask generation module can obtain a second foreground image, and multiplying the second background mask image generated by the attention mask generation module by the comic face region can obtain a second background image; finally, adding the second foreground image and the second background image can obtain a first face generation image corresponding to the comic face region generated by the generator.

[0136] Similarly, when the input is the comic face region, the mathematical expression of the generator can be as shown in the following formula. A generator can generate a content image and an attention mask image (foreground mask image and background mask image ). Multiplying the content image by the foreground attention mask generates a foreground image, multiplying the background attention mask by the input image generates a background image, and adding the foreground image and the background image can obtain the final output image.

[0137]

[0138] Among them, y represents the comic face region image. Set N to 10, then N - 1 = 9, indicating that the multiplication operation of the 9-dimensional content image matrix and the corresponding foreground mask image matrix is performed and then accumulated, and b is a constant 1.

[0139] Step 206, calculate the first probability score of the first generated image of the cartoon face using the discriminator, and calculate the second probability score of the first generated image of the human face using the discriminator;

[0140] Step 207, generate the generator loss function of the generator using the discriminator, the first generated image of the cartoon face, and the first generated image of the human face;

[0141] In an example, the discriminator includes an auxiliary discriminator and a final discriminator; the step of generating the generator loss function of the generator using the discriminator, the first generated image of the cartoon face, and the first generated image of the human face includes:

[0142] S71, calculate the first adversarial loss of generating a human face of the auxiliary discriminator using the first generated image of the cartoon face;

[0143] S72, calculate the first adversarial loss of generating a cartoon face of the auxiliary discriminator using the first generated image of the human face;

[0144] S73, calculate the second adversarial loss of generating a human face of the final discriminator using the first generated image of the cartoon face;

[0145] S74, calculate the second adversarial loss of generating a cartoon face of the final discriminator using the first generated image of the human face;

[0146] S75, generate the global image cycle consistency loss using the first generated image of the human face and the first generated image of the cartoon face;

[0147] S76, obtain the background mask cycle consistency loss;

[0148] S77, generate the generator loss function of the generator using the first adversarial loss of generating a human face, the first adversarial loss of generating a cartoon face, the second adversarial loss of generating a human face, the second adversarial loss of generating a cartoon face, the global image cycle consistency loss, and the background mask cycle consistency loss.

[0149] In a specific implementation, the specific network structure of the discriminator can be as Figure 4 shown. The overall discriminator consists of two discriminators, namely an auxiliary discriminator and a final discriminator, and the two highly share weight parameters; the specific process is as follows:

[0150] 1), The 3-dimensional Y domain (the first generated image of the cartoon face) or X domain (the first generated image of the human face) is fed into the discriminator. After passing through 4 layers of spacing, spectral normalization, and the LeakyReLu activation function, a feature map of 512 dimensions is output;

[0151] 2) After the 512-dimensional features are respectively subjected to adaptive average pooling, they become 1*512 dimensions, and then after passing through a linear layer and a spectral normalization layer, they become a 1-dimensional probability output, so as to judge whether the overall image is true. This step constitutes the discrimination of the auxiliary discriminator;

[0152] 3) The 512-dimensional feature map obtained in 1) is subjected to adaptive global pooling to become 1*512 dimensions, and then after passing through a linear layer and a spectral normalization layer, it becomes a 1-dimensional probability output, so as to judge whether the overall image is true. This step constitutes the discrimination of the auxiliary discriminator;

[0153] 4) Multiply the 512-dimensional feature map obtained in 1) with the weights obtained after pooling in 2) and 3) and passing through the Sigmoid activation function respectively, and perform matrix dimension splicing to obtain a 1024-dimensional attention-weighted feature matrix, so that the final discriminator further converges to judge the foreground target rather than the background element;

[0154] 5) After two layers of convolution, spectral normalization, and the LeakyReLu activation function, a 1-dimensional judgment matrix is output. This step constitutes the discrimination of the final discriminator.

[0155] It can be seen from Figure 4 that the discriminator structure guided by the attention mechanism adopted in the embodiment of the present invention has two sets of outputs. Among them, ηD X or ηD Y is the output of the auxiliary discriminator, which can judge the authenticity of the image in a global form. The generative adversarial loss function composed of the auxiliary discriminator (including the first face generative adversarial loss L CAM (G, ηD Y , X, Y), the first comic face generative adversarial loss L CAM (F, ηD X , X, Y)) is shown by the following mathematical expressions of the formula:

[0156]

[0157]

[0158] Among them, x and y are real images, P data (x) and P data (y) represent the sample distributions of real images, x~P data (x) and y~P data (y) mean that the samples x and y are randomly taken from the P data distribution, and E is to solve the mathematical expectation.

[0159] The generative adversarial loss function composed of the final discriminator output is consistent with CycleGAN, including the second face generative adversarial loss L GAN (G, D Y , X, Y) and the second comic face generative adversarial loss L GAN (F, D X , X, Y). Their mathematical expressions are as follows:

[0160]

[0161]

[0162] Among them, D X or D Y is the output of the final discriminator.

[0163] The error of the cycle generative adversarial network for the converted Y-domain and X-domain images is calculated using the loss function, which includes three parts: the adversarial loss of the overall image, the cycle consistency loss of the overall image, and the cycle consistency loss of the background mask. Among them, the cycle consistency loss of the background mask and the adversarial loss composed of the auxiliary discriminator in the adversarial loss of the overall image are unique to the present invention. Specifically as follows:

[0164] The adversarial loss of the overall image consists of two major parts. The first part is the adversarial loss composed of two auxiliary discriminators and the adversarial loss composed of the final discriminator.

[0165] The cycle consistency loss of the overall image is as shown in the following formula:

[0166]

[0167] Among them, x and y are real images, and P data (x) and P data (y) represent the sample distributions of real images. x ∼ P data (x) and y ∼ P data (y) indicate that the samples x and y are randomly taken from the P data distribution. E is to solve the mathematical expectation, and ||F(G(x)) - x||1 and ||G(F(y)) - y||1 represent the solution of the L1 norm for the image after the model cycle restoration and the original image. The smaller the value of L cycle (G, F), the better.

[0168] In the cyclic consistent mapping of x→G(x)→F(G(x))≈x, in the embodiments of the present invention, it is desired that the background mask image converted from the X domain to the Y domain is consistent with the background mask image restored from the Y domain to the X domain, so that the attention mask generation sub-module can guide the generator to more precisely modify the target to be converted. Inspired by the cyclic consistency loss function, the present invention proposes a cyclic consistency loss function for the background mask, and the smaller the loss value, the better.

[0169] The cyclic consistency loss of the background mask is shown in the following formula:

[0170]

[0171] The complete loss function equation of the model consists of six parts, specifically shown in the following formula, namely the generative adversarial losses of the final discriminators in the X domain and the Y domain; the generative adversarial losses of the auxiliary discriminators in the X domain and the Y domain; the cyclic consistency loss of the overall image and the cyclic consistency loss of the background mask. Among them, respectively represent the generator for converting the X-domain image to the Y domain, the generator for converting the Y-domain image to the X domain, the final discriminator for discriminating the relationship between the fake image generated in the X domain and the real image in the Y domain, the final discriminator for discriminating the relationship between the fake image generated in the Y domain and the real image in the X domain, the auxiliary discriminator for discriminating the relationship between the fake image generated in the X domain and the real image in the Y domain, the auxiliary discriminator for discriminating the relationship between the fake image generated in the Y domain and the real image in the X domain, the background mask image generated by the attention module in the G generator, the background mask image generated by the attention module in the F generator. Among them, α and λ are hyperparameters, and in the present invention, they are respectively set to 12.0 and 1.0.

[0172]

[0173] Step 208, generate the first optimization parameter of the generator by using the generator loss function;

[0174] Fix the model parameters of the discriminator, calculate the backward gradient of the generator according to the generator loss function calculated above, and update the model parameters of the generator according to the gradient information to obtain the first optimization parameter.

[0175] Step 209, calculate the cross-entropy error by using the first probability score and the second probability score to obtain the discriminator loss function of the discriminator;

[0176] Next, use the first probability score, the second probability score, and the preset labels of the X domain and the Y domain (such as 0 represents the X domain and 1 represents the Y domain) to calculate the cross-entropy error to obtain the discriminator loss function.

[0177] Cross entropy can be used as a loss function in neural networks (machine learning). p represents the distribution of true labels, and q represents the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q. Another advantage of using cross entropy as a loss function is that when using the sigmoid function in gradient descent, it can avoid the problem of the learning rate reduction of the mean squared error loss function, because the learning rate can be controlled by the output error.

[0178] Step 210, generating the second optimization parameter of the discriminator by using the discriminator loss function;

[0179] Step 211, updating the generator by using the first optimization parameter and updating the discriminator by using the second optimization parameter;

[0180] After calculating the discriminator loss function, the backward gradient of the discriminator can be calculated according to the discriminator loss function, and the model parameters of the discriminator can be updated according to the gradient information to obtain the second optimization parameter of the discriminator, and the discriminator is updated by using the second optimization parameter.

[0181] Step 212, determining whether a preset iteration termination condition is satisfied;

[0182] In the embodiment of the present invention, the termination condition can be the value of the overall loss function of the model less than 0.1.

[0183] Step 213, if not, return to the step of generating the first comic face generation image of the face area through the generator and generating the first human face generation image of the comic face area;

[0184] Step 214, if so, generating a generative adversarial network by using the updated generator and the updated discriminator.

[0185] Through continuous iteration and model optimization, a trained generative adversarial network can finally be obtained.

[0186] Further, in the embodiment of the present invention, after the step of generating a generative adversarial network by using the updated generator and the updated discriminator, the following steps are further included:

[0187] S215, inputting a preset test image into the generative adversarial network to obtain a generated test image;

[0188] S216, calculating the evaluation index of the generated test image;

[0189] S217, generating a measurement result of the generative adversarial network according to the evaluation index.

[0190] In an embodiment of the present invention, after the model is trained, the test set images can be input into the trained generative adversarial network to obtain generated test images. By calculating the corresponding KID and FID evaluation metrics for the generated test images, as well as calculating metrics such as the actual training time and model complexity, the measurement results of the generative adversarial network can be obtained.

[0191] Among them, KID and FID are used to evaluate the quality of the generated images. The KID metric measures the difference between real and fake samples by calculating the square of the maximum mean difference between the original representations. The lower the KID parameter, the more similar the two sets of samples are. The FID metric uses a neural network model to extract the high-level semantic information of the images, and measures the quality of the images generated by the generative adversarial network and the similarity between real and fake images by calculating the mean and covariance distance of the feature vectors extracted from real and generated images. When the generated images are more similar to the real images in terms of features, the FID value is smaller.

[0192] Model complexity evaluation: Floating Point Operations (FLOPs) and Multiply Accumulate Operations (MACs) are commonly used model complexity statistical metrics, which can count the amount of computation required for data to pass through the network model, that is, the computing power required when enabling the model. The number of model parameters (Parameters) is also one of the metrics describing model complexity. Times is the actual time consumed during model operation, and Memory is the actual video memory space occupied during model training. The smaller the values of these three, the more superior the model.

[0193] The present invention improves the generation effect of comic faces in complex scenarios and the authenticity of image conversion details by obtaining the face image to be converted, performing face detection on the face image to be converted to obtain the target face region, and inputting the target face region into a preset generative adversarial network to generate the target comic face image.

[0194] Please refer to Figure 5 , Figure 5 which is the structural block diagram of a comic face generation device provided by an embodiment of the present invention.

[0195] An embodiment of the present invention provides a comic face generation device, including:

[0196] A face image to be converted acquisition module 501, configured to acquire a face image to be converted;

[0197] A face detection module 502, configured to perform face detection on the face image to be converted to obtain a target face region;

[0198] The target comic face image generation module 503 is configured to input the target face region into a preset generative adversarial network to generate a target comic face image.

[0199] In the embodiments of the present invention, the generation of the preset generative adversarial network includes:

[0200] The face region acquisition module is configured to acquire face images from a preset real face dataset and perform face detection on the face images to obtain face regions;

[0201] The comic face region acquisition module is configured to acquire comic face images from a preset comic face dataset and perform face detection on the comic face images to obtain comic face regions;

[0202] The initial generative adversarial network construction module is configured to construct an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator;

[0203] The first comic face generated image generation module is configured to generate a first comic face generated image of the face region through the generator;

[0204] The first face generated image generation module is configured to generate a first face generated image of the comic face region through the generator;

[0205] The probability score calculation module is configured to calculate a first probability score of the first comic face generated image by using the discriminator, and calculate a second probability score of the first face generated image by using the discriminator;

[0206] The generator loss function generation module is configured to generate a generator loss function of the generator by using the discriminator, the first comic face generated image, and the first face generated image;

[0207] The first optimization parameter generation module is configured to generate a first optimization parameter of the generator by using the generator loss function;

[0208] The discriminator loss function calculation module is configured to calculate a cross-entropy error by using the first probability score and the second probability score to obtain a discriminator loss function of the discriminator;

[0209] The second optimization parameter generation module is configured to generate a second optimization parameter of the discriminator by using the discriminator loss function;

[0210] The discriminator update module is configured to update the generator by using the first optimization parameter and update the discriminator by using the second optimization parameter;

[0211] The iteration judgment module is configured to judge whether a preset iteration termination condition is satisfied;

[0212] A return module, configured to, if not, return the steps of generating a first comic face generation image of a face region through a generator and generating a first face generation image of a comic face region;

[0213] A generative adversarial network generation module, configured to, if so, generate a generative adversarial network by using the updated generator and the updated discriminator.

[0214] In an embodiment of the present invention, the generator includes a content generation module and an attention mask generation module; the first comic face generation image generation module includes:

[0215] A face feature information extraction sub-module, configured to extract face feature information of a face region;

[0216] A face feature conversion information generation sub-module, configured to convert the face feature information to obtain face feature conversion information;

[0217] A face feature decoding information generation sub-module, configured to perform feature encoding on the face feature conversion information to obtain face feature decoding information;

[0218] A first content map generation sub-module, configured to input the face feature decoding information into the content generation module to generate a first content map;

[0219] A first foreground mask map and first background mask map generation sub-module, configured to input the face feature decoding information into the attention mask generation module to generate a first foreground mask map and a first background mask map;

[0220] A first foreground map generation sub-module, configured to multiply the first content map and the first foreground mask map to generate a first foreground map;

[0221] A first background map generation sub-module, configured to multiply the face region and the first background mask map to generate a first background map;

[0222] A first comic face generation image generation sub-module, configured to add the first foreground map and the first background map to generate a first comic face generation image.

[0223] In an embodiment of the present invention, the first face generation image generation module includes:

[0224] A comic face feature information extraction sub-module, configured to extract comic face feature information of a comic face region;

[0225] A comic face feature conversion information generation sub-module, configured to convert the comic face feature information to obtain comic face feature conversion information;

[0226] A comic face feature decoding information generation sub-module, configured to perform feature encoding on the comic face feature conversion information to obtain comic face feature decoding information;

[0227] The second content image generation sub-module is configured to input the decoded information of the comic face features into the content generation module to obtain a second content image;

[0228] The second foreground mask image and second background mask image generation sub-module is configured to input the decoded information of the comic face features into the attention mask generation module to generate a second foreground mask image and a second background mask image;

[0229] The second foreground image generation sub-module is configured to multiply the second content image and the second foreground mask image to generate a second foreground image;

[0230] The second background image generation sub-module is configured to multiply the comic face region and the second background mask image to generate a second background image;

[0231] The first face generation image generation sub-module is configured to add the second foreground image and the second background image to generate a first face generation image.

[0232] In an embodiment of the present invention, the discriminator includes an auxiliary discriminator and a final discriminator; the generator loss function generation module includes:

[0233] The first face generation adversarial loss generation sub-module is configured to calculate the first face generation adversarial loss of the auxiliary discriminator by using the first comic face generation image;

[0234] The first comic face generation adversarial loss generation sub-module is configured to calculate the first comic face generation adversarial loss of the auxiliary discriminator by using the first face generation image;

[0235] The second face generation adversarial loss generation sub-module is configured to calculate the second face generation adversarial loss of the final discriminator by using the first comic face generation image;

[0236] The second comic face generation adversarial loss generation sub-module is configured to calculate the second comic face generation adversarial loss of the final discriminator by using the first face generation image;

[0237] The overall image cycle consistency loss generation sub-module is configured to generate an overall image cycle consistency loss by using the first face generation image and the first comic face generation image;

[0238] The background mask cycle consistency loss acquisition sub-module is configured to acquire a background mask cycle consistency loss;

[0239] The generator loss function generation sub-module is configured to generate a generator loss function of the generator by using the first face generation adversarial loss, the first comic face generation adversarial loss, the second face generation adversarial loss, the second comic face generation adversarial loss, the overall image cycle consistency loss, and the background mask cycle consistency loss.

[0240] In an embodiment of the present invention, it further includes:

[0241] A test image generation module, configured to input a preset test image into a generative adversarial network to obtain a generated test image;

[0242] An evaluation index calculation module, configured to calculate the evaluation index of the generated test image;

[0243] A metric result generation module, configured to generate a metric result of the generative adversarial network according to the evaluation index.

[0244] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory:

[0245] The memory is used to store program code and transmit the program code to the processor;

[0246] The processor is configured to execute the comic face generation method of the embodiment of the present invention according to the instructions in the program code.

[0247] The embodiment of the present invention further provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the comic face generation method of the embodiment of the present invention.

[0248] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0249] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same and similar parts among the various embodiments can be referred to each other.

[0250] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0251] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0252] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0253] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0254] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

[0255] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0256] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a comic face, characterized in that, Including: Obtain a face image to be converted; Perform face detection on the face image to be converted to obtain a target face region; Input the target face region into a preset generative adversarial network to generate a target cartoon face image; Among them, generating the preset generative adversarial network includes: Obtain face images from a preset real face dataset, and perform face detection on the face images to obtain face regions; Obtain cartoon face images from a preset cartoon face dataset, and perform face detection on the cartoon face images to obtain cartoon face regions; Construct an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator; Generate a first cartoon face generated image of the face region through the generator; Generate a first face generated image of the cartoon face region through the generator; Among them, the generator includes a content generation module and an attention mask generation module; the step of generating a first cartoon face generated image of the face region through the generator includes: Extract the face feature information of the face region; Convert the face feature information to obtain face feature conversion information; Perform feature decoding on the face feature conversion information to obtain face feature decoding information; Input the face feature decoding information into the content generation module to generate a first content map; Input the face feature decoding information into the attention mask generation module to generate a first foreground mask map and a first background mask map; Multiply the first content map and the first foreground mask map to generate a first foreground map; Multiply the face region by the first background mask map to generate a first background map; Add the first foreground map and the first background map to generate a first cartoon face generated image; Among them, the step of generating a first face generated image of the cartoon face region through the generator includes: Extract the cartoon face feature information of the cartoon face region; Convert the cartoon face feature information to obtain cartoon face feature conversion information; Perform feature decoding on the cartoon face feature conversion information to obtain cartoon face feature decoding information; Input the cartoon face feature decoding information into the content generation module to obtain a second content map; Input the cartoon face feature decoding information into the attention mask generation module to generate a second foreground mask map and a second background mask map; Multiply the second content map and the second foreground mask map to generate a second foreground map; Multiply the cartoon face region by the second background mask map to generate a second background map; Add the second foreground map and the second background map to generate a first face generated image.

2. The method according to claim 1, wherein The generation of the preset generative adversarial network further includes: Use the discriminator to calculate a first probability score of the first cartoon face generated image, and use the discriminator to calculate a second probability score of the first face generated image; Use the discriminator, the first cartoon face generated image and the first face generated image to generate a generator loss function of the generator; Use the generator loss function to generate a first optimization parameter of the generator; Calculate the cross-entropy error using the first probability score and the second probability score to obtain the discriminator loss function of the discriminator; Generate the second optimization parameter of the discriminator using the discriminator loss function; Update the generator using the first optimization parameter and update the discriminator using the second optimization parameter; Determine whether a preset iteration termination condition is satisfied; If not, return to the step of generating the first comic face generation image of the face area and the first face generation image of the comic face area through the generator; If so, generate a generative adversarial network using the updated generator and the updated discriminator; 3. The method according to claim 2, wherein The discriminator includes an auxiliary discriminator and a final discriminator; The step of generating the generator loss function of the generator using the discriminator, the first comic face generation image, and the first face generation image includes: Calculate the first face generation adversarial loss of the auxiliary discriminator using the first comic face generation image; Calculate the first comic face generation adversarial loss of the auxiliary discriminator using the first face generation image; Calculate the second face generation adversarial loss of the final discriminator using the first comic face generation image; Calculate the second comic face generation adversarial loss of the final discriminator using the first face generation image; Generate an overall image cycle consistency loss using the first face generation image and the first comic face generation image; Obtain the background mask cycle consistency loss; Generate the generator loss function of the generator using the first face generation adversarial loss, the first comic face generation adversarial loss, the second face generation adversarial loss, the second comic face generation adversarial loss, the overall image cycle consistency loss, and the background mask cycle consistency loss.

4. The method according to claim 2, characterized in that, After the step of generating a generative adversarial network using the updated generator and the updated discriminator, it further includes: Input a preset test image into the generative adversarial network to obtain a generated test image; Calculate the evaluation index of the generated test image; Generate a metric result of the generative adversarial network according to the evaluation index.

5. A comic face generation device, characterized in that, Includes: A module for obtaining a face image to be converted, which is used to obtain a face image to be converted; A face detection module, which is used to perform face detection on the face image to be converted to obtain a target face area; A target comic face image generation module, which is used to input the target face area into a preset generative adversarial network to generate a target comic face image; Among them, it further includes: A face area acquisition module, which is used to acquire a face image from a preset real face dataset and perform face detection on the face image to obtain a face area; A comic face area acquisition module, which is used to acquire a comic face image from a preset comic face dataset and perform face detection on the comic face image to obtain a comic face area; An initial generative adversarial network construction module, which is used to construct an initial generative adversarial network; the initial generative adversarial network includes a generator and a discriminator; A first comic face generation image generation module, which is used to generate a first comic face generation image of the face area through the generator; The first face generation image generation module is used to generate a first face generation image of the comic face area through the generator; Among them, the generator includes a content generation module and an attention mask generation module; the first comic face generation image generation module includes: The face feature information extraction sub-module is used to extract the face feature information of the face area; The face feature conversion information generation sub-module is used to convert the face feature information to obtain face feature conversion information; The face feature decoding information generation sub-module is used to perform feature decoding on the face feature conversion information to obtain face feature decoding information; The first content map generation sub-module is used to input the face feature decoding information into the content generation module to generate a first content map; The first foreground mask map and first background mask map generation sub-module is used to input the face feature decoding information into the attention mask generation module to generate a first foreground mask map and a first background mask map; The first foreground map generation sub-module is used to multiply the first content map and the first foreground mask map to generate a first foreground map; The first background map generation sub-module is used to multiply the face area by the first background mask map to generate a first background map; The first comic face generation image generation sub-module is used to add the first foreground map and the first background map to generate a first comic face generation image; Among them, the first face generation image generation module includes: The comic face feature information extraction sub-module is used to extract the comic face feature information of the comic face area; The comic face feature conversion information generation sub-module is used to convert the comic face feature information to obtain comic face feature conversion information; The comic face feature decoding information generation sub-module is used to perform feature decoding on the comic face feature conversion information to obtain comic face feature decoding information; The second content map generation sub-module is used to input the comic face feature decoding information into the content generation module to obtain a second content map; The second foreground mask map and second background mask map generation sub-module is used to input the comic face feature decoding information into the attention mask generation module to generate a second foreground mask map and a second background mask map; The second foreground map generation sub-module is used to multiply the second content map and the second foreground mask map to generate a second foreground map; The second background map generation sub-module is used to multiply the comic face area by the second background mask map to generate a second background map; The first face generation image generation sub-module is used to add the second foreground map and the second background map to generate a first face generation image.

6. The device according to claim 5, characterized in that, It also includes: The probability score calculation module is used to calculate the first probability score of the first comic face generation image by using the discriminator, and calculate the second probability score of the first face generation image by using the discriminator; The generator loss function generation module is used to generate the generator loss function of the generator by using the discriminator, the first comic face generation image and the first face generation image; The first optimization parameter generation module is used to generate the first optimization parameter of the generator by using the generator loss function; The discriminator loss function calculation module is used to calculate the cross-entropy error by using the first probability score and the second probability score to obtain the discriminator loss function of the discriminator; A second optimization parameter generation module, configured to generate second optimization parameters for the discriminator by using the discriminator loss function; A generator and discriminator update module, configured to update the generator by using the first optimization parameters and update the discriminator by using the second optimization parameters; An iteration determination module, configured to determine whether a preset iteration termination condition is satisfied; A return module, configured to, if the determination result is negative, return to the first comic face generation image generation module and the first human face generation image generation module; A generative adversarial network generation module, configured to, if the determination result is positive, generate a generative adversarial network by using the updated generator and the updated discriminator.

7. An electronic device, characterized in that, The device includes a processor and a memory: The memory is configured to store program code and transmit the program code to the processor; The processor is configured to execute the comic face generation method according to any one of claims 1-4 based on the instructions in the program code.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store program code, and the program code is used to execute the comic face generation method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for generating a cartoon head portrait generation model

    CN109800732A

  • Portrait cartooning method and device, robot and storage medium

    CN112465936A