Mimicry digital human automatic generation system and method based on generative adversarial network
By introducing data preprocessing, interaction and feedback, and ethics and privacy modules into the digital human automatic generation system, the pattern crash and data quality problems in the generative and adversarial network are solved, the fidelity and diversity of digital human images or videos are improved, and the ethical and privacy security of technology is ensured.
Patent Information
- Application Number
- CN202411816220.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-13
AI Technical Summary
The mimicable digital human automatic generation system based on the generative adversarial network has a pattern crash, generating a large number of similar digital humans, lack of diversity, and insufficient training data or low quality affecting the fidelity and diversity of the generated digital humans.
The data preprocessing module collects and processes high-quality face data, combines the adversarial training of the generator and discriminator, and introduces the interaction and feedback module to adjust the generator output based on user feedback, and ensures that the generation process complies with ethical standards in the ethical and privacy module and protects user privacy.
It improves the fidelity and diversity of the generated digital human images or videos, solves pattern crashes and data quality problems, enhances the application value and user experience of the system, and ensures the ethics and privacy of the technology.
Smart Images

Figure CN119991890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital human technology, and in particular to a system and method for automatically generating morphable digital humans based on a generative adversarial network. Background Art
[0002] The automatic generation system of morphable digital human based on generative adversarial network (GAN) is a cutting-edge artificial intelligence technology. It uses the powerful generation ability of GAN to create realistic virtual digital human. Based on the core principle of generative adversarial network, the system gradually optimizes the output of the generator through continuous adversarial training between the generator and the discriminator, so that it can generate highly realistic virtual digital human. These digital human beings are not only similar to real humans in appearance, but also can simulate human behavior, language and expression, and realize natural interaction with users. They receive random noise or latent space vectors as input and generate realistic virtual digital human images or videos. They are usually composed of multi-layer convolutional neural networks, deconvolution layers, fully connected layers, etc., and generate high-resolution images through upsampling and other technologies. During the training process, the generator and the discriminator confront each other and continuously adjust parameters. The generator tries to generate more and more realistic virtual digital human beings to deceive the discriminator; while the discriminator strives to improve its discrimination ability to accurately distinguish between real humans and virtual digital human beings generated by the generator. Through this continuous confrontation process, the generator can eventually generate highly realistic virtual digital human beings. The system can be widely used in advertising, entertainment, games, education, online services and other fields. For example, in the field of advertising, virtual spokespersons can be generated to promote products; in the field of entertainment, virtual idols can be created to interact with fans; in the field of education, virtual teachers can be generated to provide personalized tutoring for students. With the continuous development of technology, the automatic generation system of morphable digital humans based on generative adversarial networks will demonstrate its application value in more fields. In the future, the system is expected to achieve a higher degree of intelligence and autonomy, bringing users a richer and more realistic virtual digital human experience. At the same time, we also need to pay attention to issues such as data privacy and ethics to ensure the healthy development of technology.
[0003] However, the existing system and method for automatically generating morphable digital humans based on generative adversarial networks have the following problems:
[0004] Generative adversarial networks will fall into mode collapse, generating a large number of similar digital humans and lacking diversity;
[0005] Insufficient or low-quality training data for generative adversarial networks will affect the realism and diversity of generated digital humans;
[0006] In response to the above problems, a system and method for automatically generating a morphable digital human based on a generative adversarial network is provided. Summary of the invention
[0007] The purpose of the present invention is to provide a system and method for automatically generating a morphable digital human based on a generative adversarial network to solve the problems raised in the above background technology. To achieve the above purpose, the present invention provides the following technical solutions: a system for automatically generating a morphable digital human based on a generative adversarial network, comprising a data preprocessing module for collecting and processing high-quality face data, the data preprocessing module is coupled with a generator module for receiving random noise or latent space vectors as input to generate realistic virtual digital human images or videos, the generator module is coupled with a discriminator module for judging whether the input image or video is a real human or a virtual digital human generated by the generator, the discriminator module is coupled with an interaction and feedback module for interaction between a user and the generated virtual digital human to adjust the output of the generator according to user feedback, and the interaction and feedback module is coupled with an ethics and privacy module for the generation process to comply with ethical standards and protect user privacy.
[0008] Preferably, the collecting and processing of high-quality facial data includes cleaning, labeling, and enhancement.
[0009] Preferably, the generator module includes an input processing module, the input processing module is coupled to a feature generation module, the feature generation module is coupled to an image generation module, the image generation module is coupled to a detail enhancement module, and the detail enhancement module is coupled to a loss calculation and feedback module.
[0010] Preferably, the feature generation module generates features through a deep neural network algorithm.
[0011] Preferably, the deep neural network algorithm further includes:
[0012] S1. Build a deep neural network model according to the data set requirements, randomly initialize the parameters in the deep neural network model, and divide the data set into a training set and a test set;
[0013] S2, calculate the gradient of the deep neural network model corresponding to the current training data;
[0014] S3. Calculate the update amount of AdamW and SGDM parts in the deep neural network optimization algorithm according to the gradient of the current training data;
[0015] S4, discard part of the update amount of AdamW in the deep neural network optimization algorithm;
[0016] S5. If the deep neural network stops optimizing, adjust the learning rate of the SGDM part;
[0017] S6, multiply the processed AdamW and SGDM parts in S3 and S4 by their respective learning rates and add them together to obtain the update amount of each parameter;
[0018] S7, complete the weight decay part and update the parameters in the deep neural network according to the update amount in S6, and the number of iterations increases by 1; if the training of all data in the training set has been completed, the epoch increases by 1;
[0019] S8. When the epoch reaches the set value, the current model parameters are output and the process ends; otherwise, return to S2.
[0020] Preferably, the S1 includes the following sub-steps: S11: reading preset hyperparameters and random seeds;
[0021] S12: According to the random seed, the parameters of the deep neural network model are randomly initialized according to a normal distribution with a mean of 0 and a variance of 1.
[0022] Preferably, the S3 comprises the following sub-steps: S31: calculating the accumulation of gradients and the accumulation of squared gradients;
[0023] S32: Calculate the adaptive learning rate part of AdamW;
[0024] S33: The cumulative gradient, the inverse of the cumulative square of the gradient and the adaptive learning rate are multiplied together to form the update amount of AdamW. At the same time, the cumulative gradient also constitutes the update amount of SGDM
[0025] The S4 includes the following sub-steps: S41: generating a number sequence that is consistent with the total amount of parameters included in the deep learning model and conforms to the Bernoulli distribution according to the hyperparameters;
[0026] S42: Multiply the update amount of AdamW by the numbers in the sequence in sequence to obtain the processed update amount of AdamW.
[0027] Preferably, the S5 includes the following sub-steps: S51: after the last training of each epoch is completed, the training set accuracy of the current training is recorded. If the number of recorded accuracies reaches 5 at this time, enter S52, otherwise enter S6;
[0028] S52: Delete the earliest recorded training set accuracy. If the difference between the maximum and minimum values in the record is less than 1%, the current epoch number is recorded as epst, and the last epoch after 5 epochs is recorded as eped.
[0029] S53: If epst≤epoch, increase the learning rate corresponding to the SGDM part; otherwise, restore the learning rate to the initial value of the hyperparameter.
[0030] The method for automatically generating a mimetic digital human based on a generative adversarial network comprises the following steps:
[0031] Step 1: Data preparation: collect high-quality face data and perform preprocessing and enhancement operations;
[0032] Step 2: Model initialization: Initialize the parameters of the generator and discriminator, and set the training hyperparameters;
[0033] Step 3: adversarial training to reach the preset training rounds or convergence conditions;
[0034] Step 4: Interaction and feedback: During or after training, allow users to interact with the generated virtual digital human, and further adjust the output of the generator based on user feedback;
[0035] Step 5, Ethics and Privacy Protection: During the generation process, strictly abide by ethical standards to protect user privacy and data security.
[0036] Preferably, the adversarial training process further includes:
[0037] Step 1: The generator generates a virtual digital human image or video.
[0038] In step 2, the discriminator determines whether the generated image or video is a real human or a virtual digital human.
[0039] Step three: According to the feedback from the discriminator, adjust the parameters of the generator to generate a more realistic virtual digital human.
[0040] Step 4: Repeat steps 1 to 3 until the preset training rounds or convergence conditions are reached.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] In the present invention, high-quality face data can be collected and processed through the data preprocessing module, random noise or latent space vector can be received as input through the generator module to generate realistic virtual digital human images or videos, the discriminator module can determine whether the input image or video is a real human or a virtual digital human generated by the generator, the interaction and feedback module can enable the interaction between the user and the generated virtual digital human to adjust the output of the generator according to user feedback, and the ethics and privacy module can be used to make the generation process meet ethical standards and protect user privacy, thereby solving the problems of mode collapse based on generative adversarial networks, generating a large number of similar digital humans, lack of diversity, and insufficient or low-quality training data based on generative adversarial networks that affect the fidelity and diversity of generated digital humans. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a system block diagram of the system and method for automatically generating a morphable digital human based on a generative adversarial network of the present invention;
[0044] Figure 2 The present invention is a flow chart of the system and method for automatically generating a morphable digital human based on a generative adversarial network. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technical personnel in this field without creative work are within the scope of protection of the present invention.
[0046] See also Figure 1 to Figure 2 The present invention provides a technical solution: a system for automatically generating morphable digital humans based on a generative adversarial network, including a data preprocessing module for collecting and processing high-quality face data, using data enhancement technology to expand the data set and improve the generalization ability of the model, the data preprocessing module is coupled with a generator module for receiving random noise or latent space vectors as input to generate realistic virtual digital human images or videos, using an improved GAN structure such as StyleGAN, BigGAN, etc. to improve the generation quality and diversity, introducing style transfer, feature fusion and other technologies to make the generated digital humans have richer appearance and expression changes, and the generator module is coupled with a discriminator module. The block is used to determine whether the input image or video is a real human or a virtual digital person generated by the generator. It adopts a multi-layer convolutional neural network structure to improve the discrimination ability, introduces attention mechanism, spectral normalization and other technologies to improve the stability and accuracy of the discrimination, and the discriminator module is coupled with the interaction and feedback module for the interaction between the user and the generated virtual digital person. The output of the generator is adjusted according to the user feedback, and natural language processing, speech recognition and other technologies are used to achieve multimodal interaction to improve the user experience. The interaction and feedback module is coupled with the ethics and privacy module for the generation process to comply with ethical standards and protect user privacy. Differential privacy, federated learning and other technologies are introduced to ensure the security and privacy of the data.
[0047] In this embodiment, collecting and processing high-quality face data includes cleaning, labeling, and enhancement.
[0048] In this embodiment, the generator module includes an input processing module, which is responsible for receiving random noise vectors or latent space codes as inputs. These inputs are the basis for the generator to generate virtual digital people, including noise sampling, latent space coding and other operations. The input processing module is coupled with a feature generation module, which generates features through a series of neural network layers (such as fully connected layers, convolutional layers, etc.) using structures such as deep neural networks and convolutional neural networks (CNNs), and introduces nonlinear characteristics through nonlinear activation functions to generate feature representations with complex patterns. The feature generation module is coupled with an image generation module, which converts the features output by the feature generation module into virtual digital people in image or video format, including deconvolution layers (also called transposed convolution layers), upsampling operations, etc. It is used to gradually increase the image resolution and generate high-resolution virtual digital human images. The image generation module is coupled with a detail enhancement module to enhance the details of the preliminary image output by the image generation module to improve the realism and detail expression of the virtual digital human. Style transfer, texture synthesis and other technologies are used to optimize the image's texture, lighting, shadow and other details. The detail enhancement module is coupled with a loss calculation and feedback module to calculate the difference between the virtual digital human generated by the generator and the real digital human, and feed the loss back to the generator module for optimization, including the combined use of multiple loss functions such as adversarial loss, L1 loss, L2 loss, and back propagation algorithm.
[0049] In this embodiment, the feature generation module generates features through a deep neural network algorithm.
[0050] In this embodiment, the deep neural network algorithm also includes:
[0051] S1. Build a deep neural network model according to the data set requirements, randomly initialize the parameters in the deep neural network model, and divide the data set into a training set and a test set;
[0052] S2, calculate the gradient of the deep neural network model corresponding to the current training data;
[0053] S3. Calculate the update amount of AdamW and SGDM parts in the deep neural network optimization algorithm according to the gradient of the current training data;
[0054] S4, discard part of the update amount of AdamW in the deep neural network optimization algorithm;
[0055] S5. If the deep neural network stops optimizing, adjust the learning rate of the SGDM part;
[0056] S6, multiply the processed AdamW and SGDM parts in S3 and S4 by their respective learning rates and add them together to obtain the update amount of each parameter;
[0057] S7, complete the weight decay part and update the parameters in the deep neural network according to the update amount in S6, and the number of iterations increases by 1; if the training of all data in the training set has been completed, the epoch increases by 1;
[0058] S8. When the epoch reaches the set value, output the current model parameters and end this process; otherwise, return to S2.
[0059] In this embodiment, S1 includes the following sub-steps: S11: reading preset hyperparameters and random seeds;
[0060] S12: According to the random seed, the parameters of the deep neural network model are randomly initialized according to a normal distribution with a mean of 0 and a variance of 1.
[0061] In this embodiment, S3 includes the following sub-steps: S31: calculating the accumulation of gradients and the accumulation of squared gradients;
[0062] S32: Calculate the adaptive learning rate part of AdamW;
[0063] S33: The cumulative gradient, the inverse of the cumulative square of the gradient and the adaptive learning rate are multiplied together to form the update amount of AdamW. At the same time, the cumulative gradient also constitutes the update amount of SGDM
[0064] S4 includes the following sub-steps: S41: generating a sequence that is consistent with the total amount of parameters included in the deep learning model and conforms to the Bernoulli distribution according to the hyperparameters;
[0065] S42: Multiply the update amount of AdamW by the numbers in the sequence in sequence to obtain the processed update amount of AdamW.
[0066] In this embodiment, S5 includes the following sub-steps: S51: after the last training of each epoch is completed, the training set accuracy of the current training is recorded. If the number of recorded accuracies reaches 5 at this time, enter S52, otherwise enter S6;
[0067] S52: Delete the earliest recorded training set accuracy. If the difference between the maximum and minimum values in the record is less than 1%, the current epoch number is recorded as epst, and the last epoch after 5 epochs is recorded as eped.
[0068] S53: If epst≤epoch, increase the learning rate corresponding to the SGDM part; otherwise, restore the learning rate to the initial value of the hyperparameter.
[0069] The method for automatically generating a mimetic digital human based on a generative adversarial network comprises the following steps:
[0070] Step 1: Data preparation: collect high-quality face data and perform preprocessing and enhancement operations;
[0071] Step 2: Model initialization: Initialize the parameters of the generator and discriminator, and set the training hyperparameters;
[0072] Step 3: adversarial training to reach the preset training rounds or convergence conditions;
[0073] Step 4: Interaction and feedback: During or after training, allow users to interact with the generated virtual digital human, and further adjust the output of the generator based on user feedback;
[0074] Step 5, Ethics and Privacy Protection: During the generation process, strictly abide by ethical standards to protect user privacy and data security.
[0075] In this embodiment, the adversarial training process also includes:
[0076] Step 1: The generator generates a virtual digital human image or video.
[0077] In step 2, the discriminator determines whether the generated image or video is a real human or a virtual digital human.
[0078] Step three: According to the feedback from the discriminator, adjust the parameters of the generator to generate a more realistic virtual digital human.
[0079] Step 4: Repeat steps 1 to 3 until the preset training rounds or convergence conditions are reached.
[0080] The above shows and describes the basic principles, main features and advantages of the present invention. Technical personnel in this industry should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A system for automatically generating morphable digital humans based on generative adversarial networks, characterized in that: It includes a data preprocessing module for collecting and processing high-quality face data, the data preprocessing module is coupled to a generator module for receiving random noise or latent space vectors as input to generate realistic virtual digital human images or videos, the generator module is coupled to a discriminator module for judging whether the input image or video is a real human or a virtual digital human generated by the generator, the discriminator module is coupled to an interaction and feedback module for interaction between a user and the generated virtual digital human to adjust the output of the generator according to user feedback, and the interaction and feedback module is coupled to an ethics and privacy module for ensuring that the generation process complies with ethical standards and protects user privacy.
2. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The collecting and processing of high-quality face data includes cleaning, labeling, and enhancement.
3. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The generator module includes an input processing module, the input processing module is coupled to a feature generation module, the feature generation module is coupled to an image generation module, the image generation module is coupled to a detail enhancement module, and the detail enhancement module is coupled to a loss calculation and feedback module.
4. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The feature generation module generates features through a deep neural network algorithm.
5. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The deep neural network algorithm also includes: S1. Build a deep neural network model according to the data set requirements, randomly initialize the parameters in the deep neural network model, and divide the data set into a training set and a test set; S2, calculate the gradient of the deep neural network model corresponding to the current training data; S3. Calculate the update amount of AdamW and SGDM parts in the deep neural network optimization algorithm according to the gradient of the current training data; S4, discard part of the update amount of AdamW in the deep neural network optimization algorithm; S5. If the deep neural network stops optimizing, adjust the learning rate of the SGDM part; S6, multiply the processed AdamW and SGDM parts in S3 and S4 by their respective learning rates and add them together to obtain the update amount of each parameter; S7, complete the weight decay part and update the parameters in the deep neural network according to the update amount in S6, and the number of iterations increases by 1; if the training of all data in the training set has been completed, the epoch increases by 1; S8. When the epoch reaches the set value, the current model parameters are output and the process ends; otherwise, return to S2.
6. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The S1 includes the following sub-steps: S11: reading preset hyperparameters and random seeds; S12: According to the random seed, the parameters of the deep neural network model are randomly initialized according to a normal distribution with a mean of 0 and a variance of 1.
7. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The S3 comprises the following sub-steps: S31: calculating the accumulation of gradients and the accumulation of squared gradients; S32: Calculate the adaptive learning rate part of AdamW; S33: The cumulative gradient, the inverse of the cumulative square of the gradient and the adaptive learning rate are multiplied together to form the update amount of AdamW. At the same time, the cumulative gradient also constitutes the update amount of SGDM The S4 includes the following sub-steps: S41: generating a number sequence that is consistent with the total amount of parameters included in the deep learning model and conforms to the Bernoulli distribution according to the hyperparameters; S42: Multiply the update amount of AdamW by the numbers in the sequence in sequence to obtain the processed update amount of AdamW.
8. The system for automatically generating morphable digital humans based on generative adversarial networks according to claim 1 is characterized in that: The S5 includes the following sub-steps: S51: after the last training of each epoch is completed, the training set accuracy of the current training is recorded. If the number of recorded accuracies reaches 5 at this time, enter S52, otherwise enter S6; S52: Delete the earliest recorded training set accuracy. If the difference between the maximum and minimum values in the record is less than 1%, the current epoch number is recorded as epst, and the last epoch after 5 epochs is recorded as eped. S53: If epst≤epoch, increase the learning rate corresponding to the SGDM part; otherwise, restore the learning rate to the initial value of the hyperparameter.
9. A method for automatically generating morphable digital humans based on generative adversarial networks, characterized in that: The steps include: Step 1: Data preparation: collect high-quality face data and perform preprocessing and enhancement operations; Step 2: Model initialization: Initialize the parameters of the generator and discriminator, and set the training hyperparameters; Step 3: adversarial training to reach the preset training rounds or convergence conditions; Step 4: Interaction and feedback: During or after training, allow users to interact with the generated virtual digital human, and further adjust the output of the generator based on user feedback; Step 5, Ethics and Privacy Protection: During the generation process, strictly abide by ethical standards to protect user privacy and data security.
10. The method for automatically generating a morphable digital human based on a generative adversarial network according to claim 1, characterized in that: The adversarial training process also includes: Step 1: The generator generates a virtual digital human image or video. In step 2, the discriminator determines whether the generated image or video is a real human or a virtual digital human. Step three: According to the feedback from the discriminator, adjust the parameters of the generator to generate a more realistic virtual digital human. Step 4: Repeat steps 1 to 3 until the preset training rounds or convergence conditions are reached.