3D cartoon digital human making method based on deep learning technology

Through the style transfer and 3D reconstruction methods of deep learning technology, the problem of low efficiency of manual modeling is solved, and personalized 3D cartoon digital humans are quickly generated, which improves the user experience.

CN120655796APending Publication Date: 2025-09-16信华信(大连)软件服务股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510926125.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing methods for producing cartoon digital humans rely on manual modeling, resulting in low production efficiency, difficulty in mass production, and the inability to generate cartoon images personalized for users, reducing the interactive effect.

Method used

Using style transfer and 3D reconstruction methods based on deep learning technology, key point detection, style transfer, feature coefficient extraction and 3D reconstruction are performed by acquiring real-life images, and rendering technology is combined to generate personalized 3D cartoon digital humans.

Benefits of technology

It enables the rapid generation of personalized 3D cartoon digital humans, improves production efficiency, and enhances the diversity and interactive effects of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655796A_ABST
    Figure CN120655796A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D cartoon digital human making method based on a deep learning technology. The 3D cartoon digital human making method comprises the following steps of S1, obtaining a real user image; s2, inputting the front face image of the user into a face detection and recognition module, and carrying out key point detection and image preprocessing on the front face image; and S3, inputting the image output by the face detection and recognition module into a style migration module, and obtaining a face cartoonized image through a style migration model. The invention relates to the technical field of computer vision, and has the beneficial effects that a deep learning-based style migration technology and a three-dimensional reconstruction technology are mainly adopted, a face needing to be reconstructed is converted through a pre-trained style migration model, and a cartoon figure image of a user is obtained; training the three-dimensional reconstruction network by using the cartoon figure image to obtain a three-dimensional reconstruction network model;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method for producing 3D cartoon digital humans based on deep learning technology. Background Art

[0002] With the continuous advancement of computer technology, various technologies related to virtual digital humans are emerging. Virtual digital humans are virtual characters with human appearance and behavior, created using artificial intelligence and computer technology. These virtual characters can simulate human appearance, behavior, and voice, thereby achieving a certain degree of interaction with humans. Furthermore, cartoon digital humans, developed based on virtual digital human technology, also show great commercial potential. Compared with real-life virtual digital humans, cartoon digital humans have a more distinctive cartoon style and appearance, making them more suitable for use in entertainment fields such as games, animation, and live broadcasts. They are more entertaining and applicable in a wider range of scenarios.

[0003] However, the current method for producing cartoon digital humans primarily relies on manual creation by 3D modelers. While this manual modeling method can produce cartoon digital humans of a certain style, it requires a significant amount of time to design and create, making it difficult to mass-produce them in a short period of time. For entertainment scenarios, the number of manually modeled virtual digital humans is limited. As the number of users increases, different users will end up with the same digital human image, making it difficult to create differentiated products for each user. Furthermore, the image of the manually prefabricated digital human is fixed, making it impossible to create an image that matches the user's own, thus reducing the interactive effect. Summary of the Invention

[0004] The purpose of this invention is to solve the above problems and design a 3D cartoon digital human production method based on deep learning technology.

[0005] The technical solution of the present invention to achieve the above-mentioned purpose is a method for producing a 3D cartoon digital human based on deep learning technology, comprising the following steps:

[0006] Step S1: Obtain a real user image;

[0007] Step S2: Input the user's front face image into the face detection and recognition module to perform key point detection and image preprocessing;

[0008] Step S3: Input the image output by the face detection and recognition module into the style transfer module, and obtain a cartoonized face image through the style transfer model;

[0009] Step S4: The image output by the style transfer module is input into the face reconstruction module, facial feature coefficients are obtained through the deep learning model, and the user's face is reconstructed in three dimensions using the feature coefficients;

[0010] Step S5: The cartoon face model output by the face reconstruction module is input into the fusion module. The cartoon face model and the head model are fused and completed through the optimization algorithm. The cartoon model is matched with relevant accessories and finally the cartoon model is rendered to obtain the final result.

[0011] The step S1 is specifically as follows:

[0012] The user is placed in a brightly lit environment and faces the camera directly;

[0013] Use the camera to capture the user's front face image.

[0014] The face detection and recognition module in step S2 includes two parts: face key point detection and image preprocessing;

[0015] The facial key point detection is to use the MTCNN deep learning model to detect the face and facial key points of the input image to obtain the position coordinate information of the five key points of the face;

[0016] The image preprocessing is to use the BFM face model as the basic model, obtain the position coordinates of the five key points corresponding to the BFM model and the two-dimensional image, perform least squares calculation on the key points of the facial image and the key points of the BFM face model to obtain the image scaling parameters and cropping parameters, and preprocess the input image according to the scaling parameters and cropping parameters to obtain an image of size 224×224.

[0017] The step S3 is specifically as follows:

[0018] The style transfer method in step S3: using a pre-trained deep learning model to infer a cartoon image of the input character image;

[0019] Before the deep learning model performs inference on the input image, the method further includes: training the style transfer model: manually collecting about 200 cartoon face images of the same style, using the series of styles as the specified generation style, using the collected cartoon images and the face dataset FFHQ as training sets, and training a styleGAN network model that can generate real face images through the FFHQ dataset. For the same Gaussian noise input, two styleGAN network models Gs and Gt are used to generate real face images Xs and Xt respectively. The face images generated by the Gt model are supervised using identity feature loss and style loss respectively to ensure that the generated images have a certain cartoon style while maintaining identity features. During the supervision process, The identity feature loss is established using the identity features of the Gt output image and the Gs output image, and the style loss is established using the style features of the Gt output image and the collected cartoon face images. Since the dataset FFHQ used to train the styleGAN network are all face images after standard face alignment, the final output of the Gt model lacks the generalization of a certain head posture. Here, the diversity of the image is enhanced by using a geometric expansion module to perform random scaling and angle rotation in a certain proportion. Finally, Xs is used as the input of the autoencoder to train the autoencoder. During the training process, the autoencoder is supervised by using style loss, feature loss and facial perception loss, so that the autoencoder can generate cartoon images that conform to the characteristics and postures of real people.

[0020] The step S4 is specifically as follows:

[0021] The feature coefficient method in step S4: using a deep learning model to derive the facial feature coefficients of the user image from the input image;

[0022] Before the deep learning model performs inference on the input image, the method further includes: deep learning model training: using a generative adversarial network to generate a certain number of Asian face images, and mixing the images with a face dataset FFHQ to form a dataset of the present method; dataset preprocessing: detecting the dataset using a 68-point face key point detector to obtain 68 face key point coordinates for each image; processing the dataset using a skin attention-based Gaussian mixture model to obtain a mask for each image, dividing the processed dataset into a training set, a validation set, and a test set in a certain proportion, using the divided training set and validation set to train a face feature coefficient model, inputting the test set images into the trained face feature coefficient model for model evaluation, and finally obtaining a face feature coefficient deep learning model used in the present method;

[0023] The three-dimensional reconstruction of the face in step S4: using facial feature coefficients, including character identity feature coefficients, expression coefficients, posture coefficients and color coefficients, matrix operations are performed on different category coefficients with the BFM model to obtain a three-dimensional face model of the user image.

[0024] Decoration matching in step S5: After the head model is completed and fused, the character's glasses, hairstyle and other accessories are classified using a classification detection algorithm, and the manually prefabricated decorations are matched with the fused 3D cartoon digital human based on the classification results;

[0025] The model rendering in step S5: using a rendering engine to render the cartoon model to obtain a final 3D cartoon digital human image.

[0026] The present invention utilizes a deep learning-based method for creating a 3D cartoon digital human. This method primarily employs deep learning-based style transfer and 3D reconstruction techniques. A pre-trained style transfer model is used to transform the face to be reconstructed, resulting in a cartoon character image of the user. The cartoon character image is then used to train a 3D reconstruction network to produce a 3D reconstruction network model. Finally, the cartoon image's accessories, such as glasses and hairstyles, are detected and classified. The classification results are then matched with the pre-designed outfits. Finally, the outfits are combined with the 3D cartoon digital human to create the user's 3D cartoon digital human. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flow chart of a method for producing a 3D cartoon digital human based on deep learning technology according to the present invention;

[0028] Figure 2 This is a flow chart of the style transfer module of the 3D cartoon digital human production method based on deep learning technology described in the present invention.

[0029] Figure 3 This is a flowchart of the face reconstruction module of the 3D cartoon digital human production method based on deep learning technology described in the present invention. DETAILED DESCRIPTION

[0030] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1-3 As shown in the figure, a 3D cartoon digital human production method based on deep learning technology is Figure 1 As shown, the present invention provides a method for producing a 3D cartoon digital human based on deep learning technology, which includes the following steps:

[0031] Step S1: Acquire a real person image;

[0032] The user is placed in a brightly lit environment and faces the camera with a straight face and neutral expression;

[0033] Capture a frontal neutral expression image of the user through a camera;

[0034] Step S2: Input the user's front face image into the face detection and recognition module to perform key point detection and image preprocessing;

[0035] The face detection and recognition module includes two functions: face key point detection and image preprocessing;

[0036] The facial key point detection: using the MTCNN key point detection neural network model to detect the face and facial key points of the input image, and obtain the position coordinate information of the five key points of the face;

[0037] Image preprocessing in step S2: Using the BFM face model as the base model, obtain the coordinates of the five key points corresponding to those in step S2 on the 3D face of the BFM model. Least squares calculations are performed on the facial key points in step S2 and the key points of the BFM face model to obtain scaling and cropping parameters for the input facial image. The image is preprocessed based on the scaling and cropping parameters to obtain a 224×224 image.

[0038] Step S3: input the image output by the face detection and recognition module into the style transfer module, and obtain a cartoonized face image through the style transfer model;

[0039] The style transfer method in S3 uses a pre-trained deep learning model to infer a cartoon-like image specific to the input character image.

[0040] Furthermore, before the deep learning model performs inference on the input image, the method also includes: training a style transfer model: manually collecting approximately 200 cartoon face images of the same style, and using this series of styles as the designated generation style. The collected cartoon images and the FFHQ face dataset are used as training sets, and a styleGAN network model capable of generating realistic face images is trained using the FFHQ dataset. Given the same Gaussian noise input, two styleGAN network models, Gs and Gt, are used to generate realistic face images Xs and Xt, respectively. The face images generated by the Gt model are supervised using identity feature loss and style loss, respectively, to ensure that the generated images retain identity features while also possessing a certain cartoon style. During the supervision process, the identity feature loss is established using the identity features of both the Gt output image and the Gs output image, and the style loss is established using the style features of both the Gt output image and the collected cartoon face images. Because the FFHQ dataset used to train the styleGAN network consists of face images that have undergone standard face alignment, the final output of the Gt model lacks generalizability across certain head poses. Here, a geometric expansion module is used to randomly scale and rotate images by a certain percentage to enhance image diversity. Finally, Xs is used as input to the autoencoder to train the model. During training, the autoencoder is supervised using style loss, feature loss, and face-aware loss, enabling it to generate cartoon-like images that match the features and poses of real people.

[0041] Step S4: The image output by the style transfer module is input into the face reconstruction module, facial feature coefficients are obtained through the deep learning model, and the user's face is 3D reconstructed using the feature coefficients;

[0042] The feature coefficient method in step S4: uses a deep learning model to derive the user's facial feature coefficients from the input image.

[0043] Furthermore, before the deep learning model infers the input image, the method also includes: deep learning model training: using a generative adversarial network to generate a certain number of Asian face images, and mixing the images with the face dataset FFHQ to form the dataset of this method; dataset preprocessing: using a 68-point face key point detector to detect the dataset to obtain 68 face key point coordinates for each image; using a skin attention-based Gaussian mixture model to process the dataset to obtain a mask for each image. The processed dataset is divided into a training set, a validation set, and a test set in a certain proportion. The facial feature coefficient model is trained using the divided training set and validation set, and the test set images are input into the trained facial feature coefficient model for model evaluation, and finally the facial feature coefficient deep learning model used in this method is obtained;

[0044] The three-dimensional face reconstruction in step S4: using facial feature coefficients, including character identity feature coefficients, expression coefficients, posture coefficients and color coefficients, performing matrix operations on different category coefficients and the BFM model respectively to obtain a three-dimensional face model of the user image;

[0045] Step S5: The cartoon face model output by the face reconstruction module is input into the fusion module. The cartoon face model and the head model are fused and completed using an optimization algorithm, and the cartoon model is matched with relevant accessories and costumes. Finally, the cartoon model is rendered to obtain the final result.

[0046] Decoration matching in step S5: After the head model is completed and fused, the character's glasses, hairstyle and other accessories are classified through a classification detection algorithm. The manually prefabricated decorations are matched with the fused 3D cartoon digital human based on the classification results.

[0047] Model rendering in step S5: using a rendering engine to render the cartoon model to obtain a final 3D cartoon digital human image.

[0048] The above technical solutions only reflect the preferred technical solutions of the technical solutions of the present invention. Any changes that may be made to certain parts thereof by those skilled in the art all reflect the principles of the present invention and fall within the scope of protection of the present invention.

Claims

1. A method for producing 3D cartoon digital humans based on deep learning technology, characterized in that: The following steps are involved: Step S1: Obtain a real user image; Step S2: Input the user's front face image into the face detection and recognition module to perform key point detection and image preprocessing; Step S3: Input the image output by the face detection and recognition module into the style transfer module, and obtain a cartoonized face image through the style transfer model; Step S4: The image output by the style transfer module is input into the face reconstruction module, facial feature coefficients are obtained through the deep learning model, and the user's face is reconstructed in three dimensions using the feature coefficients; Step S5: The cartoon face model output by the face reconstruction module is input into the fusion module. The cartoon face model and the head model are fused and completed through the optimization algorithm. The cartoon model is matched with relevant accessories and finally the cartoon model is rendered to obtain the final result.

2. The method for producing 3D cartoon digital humans based on deep learning technology according to claim 1, characterized in that: The step S1 is specifically as follows: The user is placed in a brightly lit environment and faces the camera directly; Use the camera to capture the user's front face image.

3. The method for producing 3D cartoon digital humans based on deep learning technology according to claim 1, characterized in that: The face detection and recognition module in step S2 includes two parts: face key point detection and image preprocessing; The facial key point detection is to use the MTCNN deep learning model to detect the face and facial key points of the input image to obtain the position coordinate information of the five key points of the face; The image preprocessing is to use the BFM face model as the basic model, obtain the position coordinates of the five key points corresponding to the BFM model and the two-dimensional image, perform least squares calculation on the key points of the facial image and the key points of the BFM face model to obtain the image scaling parameters and cropping parameters, and preprocess the input image according to the scaling parameters and cropping parameters to obtain an image of size 224×224.

4. The method for producing 3D cartoon digital humans based on deep learning technology according to claim 1, characterized in that: The step S3 is specifically as follows: The style transfer method in step S3: using a pre-trained deep learning model to infer a cartoon image of the input character image; Before the deep learning model performs inference on the input image, it also includes: style transfer model training: manually collect about 200 cartoon face images of the same style, use this series of styles as the specified generation style, use the collected cartoon images and the face dataset FFHQ as training sets, and use the FFHQ dataset to train a styleGAN network model that can generate real face images. For the same Gaussian noise input, two styleGAN network models Gs and Gt are used to generate real face images Xs and Xt respectively. The face images generated by the Gt model are supervised using identity feature loss and style loss respectively to ensure that the generated images have a certain cartoon style while maintaining identity features. During the supervision process, The identity feature loss is established using the identity features of the Gt output image and the Gs output image, and the style loss is established using the style features of the Gt output image and the collected cartoon face images. Since the dataset FFHQ used to train the styleGAN network are all face images after standard face alignment, the final output of the Gt model lacks generalization of a certain head posture. The image diversity is enhanced by using a geometric expansion module to perform random scaling and angle rotation in a certain proportion. Xs is used as the input of the autoencoder to train the autoencoder. During the training process, the autoencoder is supervised by using style loss, feature loss and facial perception loss, so that the autoencoder can generate cartoon images that conform to the characteristics and postures of real people.

5. The method for producing 3D cartoon digital humans based on deep learning technology according to claim 1, characterized in that: The step S4 is specifically as follows: The feature coefficient method in step S4: using a deep learning model to derive the facial feature coefficients of the user image from the input image; Before the deep learning model performs inference on the input image, the method further includes: deep learning model training: using a generative adversarial network to generate a certain number of Asian face images, and mixing the images with the face dataset FFHQ to form a dataset of the method; Dataset preprocessing: The dataset is detected using a 68-point facial landmark detector to obtain the coordinates of 68 facial landmarks for each image; the dataset is processed using a Gaussian mixture model based on skin attention to obtain a mask for each image, and the processed dataset is divided into a training set, a validation set, and a test set according to a certain ratio. The facial feature coefficient model is trained using the divided training set and validation set, and the test set images are input into the trained facial feature coefficient model for model evaluation, ultimately obtaining the facial feature coefficient deep learning model used in this method; The three-dimensional reconstruction of the face in step S4: using facial feature coefficients, including character identity feature coefficients, expression coefficients, posture coefficients and color coefficients, matrix operations are performed on different category coefficients with the BFM model to obtain a three-dimensional face model of the user image.

6. The method for producing 3D cartoon digital humans based on deep learning technology according to claim 1, characterized in that: Decoration matching in step S5: After the head model is completed and fused, the glasses and hairstyle of the character are classified using a classification detection algorithm, and the manually prefabricated decorations are matched with the fused 3D cartoon digital human based on the classification results; The model rendering in step S5: using a rendering engine to render the cartoon model to obtain a final 3D cartoon digital human image.