Picture digital human generation method based on deep learning

Through a deep learning-based method, combining speech and facial feature point models, high-quality dynamic digital human videos are generated, which solves the problems of insufficient generation efficiency and quality in the existing technology, and improves fidelity and reality.

CN120070683APending Publication Date: 2025-05-30天津金汇科技股份有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510089689.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing digital life generation technology of picture is difficult to generate high-quality dynamic videos, and the generated digital people are insufficient in realism and reality. Traditional methods require a lot of manual participation, and efficiency and quality are limited.

Method used

Using a deep learning-based method, the character facial movements and lip feature point sequences are automatically generated through the speech-face feature point model, the speech-lip feature point model, the facial reenactment model and the character repair model, to realize dynamic picture generation, and improve facial details and realism through the repair model.

Benefits of technology

It improves the efficiency and quality of digital life generation, realizes high-quality dynamic video generation, improves facial details and sense of reality, simplifies manual participation, and is suitable for various devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070683A_ABST
    Figure CN120070683A_ABST
Patent Text Reader

Abstract

A picture digital human generation method based on deep learning is characterized by comprising the following steps: S1, through a voice-facial feature point model, generating a feature point sequence of human facial actions based on voice; s2, through the voice-lip shape feature point model, generating a lip shape feature point sequence corresponding to the voice; s3, through a face replay model, making the figure picture perform corresponding actions according to the face actions and the lip shape feature points; s4, through a character restoration model, improving character face details in the video, and restoring artifacts; according to the invention, a deep learning technology is utilized to automatically generate a character face action and lip shape feature point sequence, so that the generation efficiency is improved; and meanwhile, the generated dynamic character picture is optimized through the character repairing model, so that the quality of the picture is improved. In addition, the method is simple and easy to implement, can be implemented on various devices, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and speech recognition, and in particular to a method for generating digital humans from images based on deep learning. Background Art

[0002] With the advancement of science and technology and the rapid development of artificial intelligence, digital human generation technology has become one of the hot topics of current research. Digital human generation technology aims to generate realistic human pictures or videos through computers to achieve more intuitive and vivid expression and communication. This technology can be applied to social media, virtual reality, games, movies and other fields, and has broad application prospects and market value.

[0003] Existing digital human generation technologies based on images can only generate static images, but not dynamic videos. Moreover, the fidelity and realism of the generated digital human images need to be improved. In addition, traditional digital human generation technologies based on images mainly use scanning or modeling to map images or models of the real world into the digital world. However, this method has certain limitations. First, they require a lot of manual participation, such as manually setting feature points or performing post-repair, which not only increases the workload, but also limits the generation efficiency and quality. Second, the quality of the generated images or videos is not high and cannot achieve realistic effects, especially in terms of facial details and movements. In addition, some methods are not accurate enough in converting speech and cannot accurately express the user's intentions.

[0004] To solve the above problems, a method for generating digital humans from images based on deep learning is proposed. Summary of the invention

[0005] In view of the above technical problems, the present invention provides a method for generating digital human from images based on deep learning, which is characterized by comprising the following steps:

[0006] S1: Generate a sequence of facial action feature points based on speech through a speech-facial feature point model;

[0007] S2: Generate a lip shape feature point sequence corresponding to the speech through the speech-lip shape feature point model;

[0008] S3: Use the facial reenactment model to make the character image perform corresponding actions based on facial movements and lip shape feature points;

[0009] S4: Use the character restoration model to enhance facial details and repair artifacts in the video;

[0010] The speech-facial landmark model in S1 includes the following steps:

[0011] A1: The model receives the input speech signal and preprocesses the speech signal through a preprocessing module;

[0012] A2: The model inputs the preprocessed speech signal into an acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks, and converts the speech signal into a feature vector;

[0013] A3: The feature vector is input into a facial action encoder, which is based on a multi-head self-attention module, captures the relationship between phonetic units and feature points, and models the temporal dependence relationship, converting the feature vector into a sequence of facial action feature points;

[0014] A4: In the facial action encoder, a convolutional neural network is used to reduce the dimension and extract features of the feature vector, and a multi-head self-attention module is used to model the sequence of the reduced-dimensional features, and finally outputs a sequence of facial action feature points;

[0015] The speech-lip feature point model in S2 includes the following steps:

[0016] B1: The model receives the input speech signal and preprocesses the speech signal through a preprocessing module;

[0017] B2: The model inputs the preprocessed speech signal into an acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks, and converts the speech signal into a feature vector;

[0018] B3: The feature vector is input into a lip encoder, which is composed of a convolutional neural network and a recurrent neural network, and converts the feature vector into a sequence of lip feature points;

[0019] B4: The encoder is based on a generative adversarial network, including a lip generator and a lip synchronization discriminator. The lip generator generates lip feature points based on the speech, and the lip synchronization discriminator judges whether the generated lips are correct, and finally outputs a correct sequence of lip feature points;

[0020] The facial reenactment model in S3 includes the following steps:

[0021] C1: The model receives the input person picture, sequence of facial action feature points and sequence of lip feature points;

[0022] C2: These input information are input into an image processing module, which includes a motion field alignment module and a facial image conversion module, and can dynamically adjust the person picture according to the input facial actions and lip feature points to achieve real-time reenactment of the person picture;

[0023] The person repair model in S4 includes the following steps:

[0024] D1: The model receives input character images or video frames, as well as information about the target repair area;

[0025] D2: A 5-layer convolutional neural network is used to identify and locate the areas that need to be repaired;

[0026] D3: Generate new pixel values ​​using a generative adversarial network to repair artifacts and enhance facial details;

[0027] D4: The final output is obtained by mixing the restored image with the original image to enhance the clarity and realism of the facial details;

[0028] A system for generating digital human from images based on deep learning, characterized in that it comprises an integration module, which can integrate the above four models into one system to achieve high-quality digital human generation from images;

[0029] It also includes a data collection module that can collect a large amount of character speech and facial action data, character speech and lip shape data, character facial action and lip shape data, and character facial detail data for training a speech-facial feature point model, a speech-lip shape feature point generation model, a facial reenactment model, and a character restoration model;

[0030] Also includes an output module, which can output high-quality picture digital people;

[0031] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements a system for generating digital humans from images based on deep learning.

[0032] Beneficial effects of the present invention:

[0033] The present invention uses deep learning technology to automatically generate a sequence of facial movements and lip shape feature points, thereby improving the generation efficiency; at the same time, the generated dynamic character pictures are optimized through a character restoration model, thereby improving the quality of the pictures. In addition, the method of the present invention is simple and easy to implement, can be implemented on various devices, and has a wide range of application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is an overall flow chart of a method for generating digital humans from images based on deep learning according to the present invention;

[0035] Figure 2 It is an overall flow chart of a voice-facial feature point model of a method for generating digital human from a picture based on deep learning in the present invention;

[0036] Figure 3It is an overall flow chart of the speech-lip shape feature point model of a method for generating digital human from a picture based on deep learning in the present invention;

[0037] Figure 4 It is an overall flow chart of a facial reenactment model of a method for generating digital human from an image based on deep learning according to the present invention;

[0038] Figure 5 This is an overall flow chart of a character restoration model of a method for generating digital humans from images based on deep learning according to the present invention. DETAILED DESCRIPTION

[0039] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0041] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0042] Example 1

[0043] The present invention provides a method for generating a digital human from a picture based on deep learning, and the method is characterized in that it comprises the following steps:

[0044] S1: Generate a sequence of facial action feature points based on speech through a speech-facial feature point model;

[0045] S2: Generate a lip shape feature point sequence corresponding to the speech through the speech-lip shape feature point model;

[0046] S3: Use the facial reenactment model to make the character image perform corresponding actions based on facial movements and lip shape feature points;

[0047] S4: Use the character restoration model to enhance facial details and repair artifacts in the video;

[0048] The speech-facial landmark model in S1 includes the following steps:

[0049] A1: The model receives the input speech signal and preprocesses the speech signal through the preprocessing module;

[0050] A2: The model inputs the preprocessed speech signal into the acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks to convert the speech signal into a feature vector;

[0051] A3: The feature vector is input into the facial action encoder, which is based on a multi-head self-attention module to capture the relationship between phoneme units and feature points, and model the temporal dependency to convert the feature vector into a sequence of facial action feature points;

[0052] A4: In the facial action encoder, a convolutional neural network is used to reduce the dimension of the feature vector and extract features. A multi-head self-attention module is used to perform sequence modeling on the features after dimension reduction, and finally a sequence of facial action feature points is output.

[0053] The speech-lip shape feature point model in S2 includes the following steps:

[0054] B1: The model receives the input speech signal and preprocesses the speech signal through the preprocessing module;

[0055] B2: The model inputs the preprocessed speech signal into the acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks to convert the speech signal into a feature vector;

[0056] B3: The feature vector is input into the lip shape encoder, which is composed of a convolutional neural network and a recurrent neural network to convert the feature vector into a sequence of lip shape feature points;

[0057] B4: The encoder is based on a generative adversarial network and includes a lip shape generator and a lip shape synchronization discriminator. The lip shape generator generates lip shape feature points based on speech, and the lip shape synchronization discriminator determines whether the generated lip shape is correct, and finally outputs the correct lip shape feature point sequence;

[0058] The facial reenactment model in S3 includes the following steps:

[0059] C1: The model receives the input person picture, the sequence of facial action feature points, and the sequence of lip feature points.

[0060] C2: Input these input information into the image processing module. The image processing module includes a motion field alignment module and a facial image conversion module, which can dynamically adjust the person picture according to the input facial actions and lip feature points to achieve real-time reenactment of the person picture.

[0061] The person restoration model in S4 includes the following steps:

[0062] D1: The model receives the input person picture or video frame, and the information of the target restoration area.

[0063] D2: Identify and locate the area to be restored through a 5-layer convolutional neural network.

[0064] D3: Use a generative adversarial network to generate new pixel values to repair artifacts and enhance the details of the person's face.

[0065] D4: Obtain the final output by mixing the restored image with the original image to enhance the clarity and realism of the details of the person's face.

[0066] A picture digital human generation system based on deep learning, characterized in that it includes an integration module, which can integrate the above four models into a system to achieve high-quality picture digital human generation.

[0067] It also includes a data collection module, which can collect a large amount of person voice and facial action data, person voice and lip data, person facial action and lip data, and person facial detail data for training the voice-facial feature point model, the voice-lip feature point generation model, the facial reenactment model, and the person restoration model.

[0068] It also includes an output module, which can output high-quality picture digital humans.

[0069] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements a picture digital human generation system based on deep learning.

[0070] Embodiment 2

[0071] S1: Based on the speech-facial feature point model, generate a sequence of feature points (landmarks) for the facial movements of a person from speech. First, the model receives the input speech signal and preprocesses the speech signal through a preprocessing module, including operations such as noise removal and normalization. Then, the model inputs the preprocessed speech signal into an acoustic model, which consists of a 7-layer convolutional neural network (CNN) and 12 Transformer blocks and can convert the speech signal into a feature vector. Next, the feature vector is input into a facial movement encoder, which is based on a multi-head self-attention module, can better capture the relationship between phoneme units and landmarks, and model the temporal dependence relationship, and converts the feature vector into a sequence of facial movement landmarks. In the facial movement encoder, first use CNN to reduce the dimension and extract features of the feature vector, and then use the multi-head self-attention module to perform sequence modeling on the reduced-dimensional features, and finally output a sequence of facial movement landmarks. To improve the accuracy and generalization ability of the model, the model also includes a loss function and an optimizer for calculating the loss of the model and updating the model parameters. The loss function includes mean squared error (MSE) and cross-entropy, etc., for measuring the difference between the model prediction result and the true result. During the training process, the model is trained using a large amount of data corresponding to speech and facial movements. These data can be manually labeled or obtained through other data collection methods. By continuously adjusting the model parameters, the prediction result of the model gets closer and closer to the true result. Compared with traditional technologies, this method can generate the feature point sequence more automatically, improving the generation efficiency and quality;

[0072] S2: Generate lip feature points corresponding to the speech through the speech-lip feature point model. The acoustic model part is the same as in step 1 and is used to convert the speech into a feature vector. Next, the feature vector is input into a lip encoder, which is composed of a series of convolutional neural networks (CNN) and recurrent neural networks (RNN) and can convert the feature vector into a sequence of lip landmarks. This encoder is based on a generative adversarial network (GAN) and includes a lip generator and a lip synchronization discriminator. The lip generator generates lip landmarks based on the speech, and the lip synchronization discriminator judges whether the generated lips are correct, and finally outputs a correct sequence of lip landmarks. During the training process, the model is trained using a large amount of data corresponding to speech and lips. These data can be manually labeled or obtained through other data collection methods. By continuously adjusting the model parameters, the prediction result of the model gets closer and closer to the true result;

[0073] S3: Use the facial reenactment model to make the person picture perform corresponding actions according to the facial movements and lip feature points. First, the model receives the input person picture, the facial movement landmarks sequence, and the lip landmarks sequence. Then, the model inputs this information into the image processing module, which includes a motion field alignment module and a facial image conversion module, and can dynamically adjust the person picture according to the input facial movements and lip landmarks to achieve real-time reenactment of the person picture. Specifically, the image processing module will perform deformation and displacement operations on the person picture according to the facial movement landmarks sequence, so that the facial expression and movements of the person picture correspond to the input facial movement landmarks sequence. At the same time, the image processing module will also dynamically adjust the mouth shape of the person picture according to the lip landmarks sequence, so that the mouth shape of the person picture corresponds to the input lip landmarks sequence. During the training process, the model is trained using a large number of person pictures, facial movement landmarks sequences, and lip landmarks sequences. These data can be manually labeled or obtained through other data collection methods. By continuously adjusting the model parameters, the prediction results of the model are getting closer and closer to the real results;

[0074] S4: Use the person restoration model to enhance the facial details of the person in the video and repair artifacts. First, the model receives the input person picture or video frame, and the information of the target repair area; then uses a 5-layer CNN to identify and locate the area to be repaired, and then uses GAN to generate new pixel values to repair the artifacts and enhance the facial details of the person; finally, the final output is obtained by mixing the repaired image with the original image to enhance the clarity and realism of the facial details of the person. This model is trained with a large amount of dynamic person picture data after repair, and can perform refined optimization and repair on the generated pictures while ensuring the overall generation efficiency. Compared with traditional technologies, this method can more effectively repair the artifacts and detail problems in the generated pictures and improve the quality of the generated pictures or videos.

[0075] The above shows and describes the basic principles, main features and advantages of the present invention. Each component mentioned in the present invention is a common technology in the existing field. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating digital human from images based on deep learning, characterized in that: The following steps are involved: S1: Generate a sequence of facial action feature points based on speech through a speech-facial feature point model; S2: Generate a lip shape feature point sequence corresponding to the speech through the speech-lip shape feature point model; S3: Use the facial reenactment model to make the character image perform corresponding actions based on facial movements and lip shape feature points; S4: Enhance facial details and repair artifacts in the video through the character restoration model.

2. The method for generating digital human from images based on deep learning according to claim 1, characterized in that: The speech-facial landmark model in S1 includes the following steps: A1: The model receives the input speech signal and preprocesses the speech signal through the preprocessing module; A2: The model inputs the preprocessed speech signal into the acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks to convert the speech signal into a feature vector; A3: The feature vector is input into the facial action encoder, which is based on a multi-head self-attention module to capture the relationship between phoneme units and feature points, and model the temporal dependency to convert the feature vector into a sequence of facial action feature points; A4: In the facial action encoder, a convolutional neural network is used to reduce the dimension of the feature vector and extract features, and a multi-head self-attention module is used to perform sequence modeling on the reduced features, and finally a sequence of facial action feature points is output.

3. The method for generating digital human from images based on deep learning according to claim 2, characterized in that: The speech-lip shape feature point model in S2 includes the following steps: B1: The model receives the input speech signal and preprocesses the speech signal through the preprocessing module; B2: The model inputs the preprocessed speech signal into the acoustic model, which consists of a 7-layer convolutional neural network and 12 Transformer blocks to convert the speech signal into a feature vector; B3: The feature vector is input into the lip shape encoder, which is composed of a convolutional neural network and a recurrent neural network to convert the feature vector into a sequence of lip shape feature points; B4: The encoder is based on a generative adversarial network and includes a lip shape generator and a lip shape synchronization discriminator. The lip shape generator generates lip shape feature points based on speech, and the lip shape synchronization discriminator determines whether the generated lip shape is correct, and finally outputs the correct lip shape feature point sequence.

4. The method for generating digital human from images based on deep learning according to claim 3, characterized in that: The facial reenactment model in S3 consists of the following steps: C1: The model receives input character pictures, facial action feature point sequences, and lip shape feature point sequences; C2: These input information are input into the image processing module, which includes a motion field alignment module and a facial image conversion module. The module can dynamically adjust the character image according to the input facial movements and lip shape feature points to achieve real-time replay of the character image.

5. The method for generating digital human from images based on deep learning according to claim 4, characterized in that: The character repair model in S4 includes the following steps: D1: The model receives input character images or video frames, as well as information about the target repair area; D2: A 5-layer convolutional neural network is used to identify and locate the areas that need to be repaired; D3: Generate new pixel values ​​using a generative adversarial network to repair artifacts and enhance facial details; D4: The final output is obtained by blending the restored image with the original image to enhance the clarity and realism of the facial details.

6. A system for generating digital human from images based on deep learning according to any one of claims 1 to 5, characterized in that: It also includes an integration module that can integrate the above four models into one system to achieve high-quality image digital human generation.

7. The system for generating digital human from images based on deep learning according to claim 6, characterized in that: It also includes a data collection module that can collect a large amount of character speech and facial movement data, character speech and lip shape data, character facial movement and lip shape data, and character facial detail data for training speech-facial feature point models, speech-lip shape feature point generation models, facial reenactment models, and character restoration models.

8. The system for generating digital human from images based on deep learning according to claim 7, characterized in that: It also includes an output module that can output high-quality picture digital people.

9. The computer-readable storage medium according to claim 8, wherein: A computer program is stored thereon for running a system for generating digital humans from images based on deep learning.

Citation Information

Patent Citations

  • Video generation method and device, server and storage medium

    CN113901894A

  • Portrait generation method and device

    CN115631269A

  • Digital face video generation method based on audio driving

    CN118864665A

  • Efficient method for driving quadratic element picture expressions and actions based on audio

    CN119048655A

  • Digital human construction method based on voice driving

    CN119049468A