Voice-driven digital human image generation method, apparatus, and storage medium

CN116758189BActive Publication Date: 2026-08-14CHINA MERCHANTS BANK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]本发明的主要目的在于提供一种基于语音驱动的数字人图像生成方法、装置及存储介质,旨在解决现有数字人唇形与语音不匹配的技术问题

Benefits of technology

[0040]This invention predicts multiple first 53-dimensional facial expression lip coefficients by inputting the speech to be predicted into a target lip coefficient inference model. Then, it uses the parameters other than the 53-dimensional facial expression in the preset FLAME general head model parameters, along with the first 53-dimensional facial expression lip coefficients, to perform 3D facial reconstruction, obtaining multiple face rendering images. Each face rendering image is then input into a target generative adversarial network model to generate a target face image corresponding to each face rendering image. By learning the mapping relationship between speech and digital human lip shape changes through the lip coefficient inference model, the temporal dependency relationship of speech features within a certain time window can be effectively obtained, enabling precise matching and synchronization between digital human lip shape and speech. By decoupling the lip shape change coefficients from the face shape coefficients through the preset FLAME general head model parameters, only the model that infers lip shape change coefficients from speech needs to be trained to generalize to different character images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758189B_ABST
    Figure CN116758189B_ABST
Patent Text Reader

Abstract

This invention discloses a voice-driven digital human image generation method, apparatus, and storage medium. The method includes: inputting the voice to be predicted into a target lip shape coefficient inference model to predict multiple first 53-dimensional facial expression lip shape coefficients; performing 3D facial reconstruction using parameters other than the 53-dimensional facial expressions from a preset FLAME general head model and the first 53-dimensional facial expression lip shape coefficients to obtain multiple face rendering images; and inputting each of the face rendering images into a target generative adversarial network model to generate a target face image corresponding to each face rendering image. This invention learns the mapping relationship between voice and digital human lip shape changes through a lip shape coefficient inference model, enabling precise matching and synchronization between digital human lip shape and voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital human technology, and in particular to a voice-driven digital human image generation method, apparatus and storage medium. Background Technology

[0002] A digital human is a composite entity existing in the non-physical world, created and used by computers, possessing multiple personality traits (physical features, interactive abilities, performance abilities, etc.). Digital humans integrate numerous cutting-edge artificial intelligence technologies such as human image simulation, voice cloning, and natural language processing, and have been widely applied in scenarios such as news broadcasting, sign language generation, virtual actors, and online education. In the financial sector, digital human technology can be used to generate intelligent financial advisors, intelligent customer service representatives, and other roles, providing customer-centric, intelligent, efficient, and personalized services.

[0003] Virtual digital humans can be categorized by their appearance type into 2D, 3D cartoon, and 3D hyper-realistic. Among them, 2D digital humans refer to virtual avatars that are trained to resemble real people by collecting video data of real people speaking in a professional recording studio and can speak according to given voice.

[0004] Currently, methods for generating 2D digital humans include: directly generating a digital human face with corresponding lip shapes based on speech and reference facial images; however, this method directly learns the mapping from speech to pixel space, making it difficult to generate high-quality corresponding lip shapes. Another method involves generating intermediate expressions of the corresponding lip shapes from speech, and then using these intermediate expressions to generate facial images. This method effectively utilizes the geometric constraints of lip shapes, but changes in lip shapes are coupled with facial shape and head movements, making it difficult to generate synchronized and detailed lip shapes based on speech. Existing methods for generating digital humans result in a lack of perfect matching between lip shapes and speech, insufficient lip shape detail, and a significant difference from the lip shapes of real people speaking.

[0005] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this invention is to provide a voice-driven digital human image generation method, apparatus, and storage medium, aiming to solve the technical problem of mismatch between existing digital human lip shapes and speech.

[0007] To achieve the above objectives, the present invention provides a voice-driven digital human image generation method, which includes the following steps:

[0008] The target lip shape coefficient inference model is used to predict the speech input to obtain multiple first 53-dimensional facial expression lip shape coefficients.

[0009] The parameters other than the 53-dimensional expression in the preset FLAME general head model parameters, as well as the lip shape coefficient of the first 53-dimensional expression, are used to perform 3D facial reconstruction to obtain multiple face rendering images.

[0010] Each of the face rendering images is input into the target generative adversarial network model to generate the target face image corresponding to each of the face rendering images.

[0011] Furthermore, prior to the step of predicting the target lip shape coefficient of the speech input to be predicted using the inference model to obtain multiple first 53-dimensional facial expression lip shape coefficients, the speech-driven digital human image generation method further includes:

[0012] Obtain the audio data corresponding to the audio and video data to be trained and each video frame in the video data, and obtain the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image.

[0013] The 53-dimensional facial expressions and audio data from the preset FLAME general head model parameters are input into the initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient.

[0014] Based on the second 53-dimensional facial lip shape coefficient, the preset FLAME general head model parameters, and the face image, a composite face image corresponding to each video frame is generated.

[0015] The composite face image and face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the target generative adversarial network model.

[0016] Further, the step of inputting the 53-dimensional facial expressions and audio data from the preset FLAME general head model parameters into the initial lip shape coefficient inference model for model training to obtain the target lip shape coefficient inference model and the second 53-dimensional facial expression lip shape coefficient includes:

[0017] The 53-dimensional facial expressions and audio data in each of the preset FLAME general head model parameters are aligned with time frames to obtain training data.

[0018] The training data is input into the initial lip coefficient inference model for model training to obtain the trained lip coefficient inference model.

[0019] If the first loss function of the trained lip shape coefficient inference model is less than the first preset value, then the trained lip shape coefficient inference model is used as the target lip shape coefficient inference model, and the output data of the current model training is used as the second 53-dimensional facial expression lip shape coefficient.

[0020] The initial lip coefficient inference model includes a position encoder, a Transformer encoder, and a linear layer. The Transformer encoder includes six coding layers.

[0021] Furthermore, the step of generating a composite face image corresponding to each video frame based on the second 53-dimensional lip shape coefficient, the preset FLAME general head model parameters, and the face image includes:

[0022] The second 53-dimensional facial lip coefficient and other parameters in the preset FLAME general head model parameters other than the 53-dimensional facial expression are rendered to obtain the face rendering map corresponding to each video frame.

[0023] Based on the facial key points of each of the aforementioned facial images, a facial mask corresponding to each video frame is generated;

[0024] Based on the face rendering image, face mask, and face image, a composite face image corresponding to each video frame is generated.

[0025] Further, the step of inputting the composite face image and face image corresponding to each video frame into the initial generative adversarial network model for model training to obtain the target generative adversarial network model includes:

[0026] The composite face image and face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the trained generative adversarial network model.

[0027] If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

[0028] Further, the generative adversarial network model includes a face discriminator and a facial component discriminator for the teeth region; the step of using the trained generative adversarial network model as the target generative adversarial network model if the second loss function of the trained generative adversarial network model is less than a second preset value includes:

[0029] Based on the trained generative adversarial network model, the face reconstruction error, the face discrimination error corresponding to the face discriminator, and the face component error corresponding to the facial component discriminator in the tooth region are obtained.

[0030] The second loss function is determined based on the face reconstruction error, face discrimination error, and facial component error.

[0031] If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

[0032] Furthermore, the steps of obtaining the audio data corresponding to the audio and video data to be trained and each video frame in the video data, and obtaining the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image include:

[0033] Obtain the audio and video data corresponding to the audio and video data to be processed;

[0034] Obtain the face images corresponding to each video frame in the video data;

[0035] Each of the aforementioned face images is input into the FLAME general head model to obtain the preset FLAME general head model parameters corresponding to each video frame. The preset FLAME general head model parameters include: 100-dimensional shape, 53-dimensional expression, 50-dimensional texture, 6-dimensional lighting, and 3-dimensional projection parameters.

[0036] Furthermore, after the step of inputting each of the face rendering images into the target generative adversarial network model to generate the target face image corresponding to each of the face rendering images, the voice-driven digital human image generation method further includes:

[0037] The speech to be predicted is aligned with each of the target face images in time frames to generate a digital human audio-visual video corresponding to the speech to be predicted.

[0038] Furthermore, to achieve the above objectives, the present invention also provides a voice-driven digital human image generation apparatus, which includes: a memory, a processor, and a voice-driven digital human image generation program stored in the memory and executable on the processor. When the voice-driven digital human image generation program is executed by the processor, it implements the steps of the voice-driven digital human image generation method described in any of the preceding claims.

[0039] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a voice-driven digital human image generation program, wherein the voice-driven digital human image generation program, when executed by a processor, implements the steps of the voice-driven digital human image generation method described in any of the preceding claims.

[0040] This invention predicts multiple first 53-dimensional facial expression lip coefficients by inputting the speech to be predicted into a target lip coefficient inference model. Then, it uses the parameters other than the 53-dimensional facial expression in the preset FLAME general head model parameters, along with the first 53-dimensional facial expression lip coefficients, to perform 3D facial reconstruction, obtaining multiple face rendering images. Each face rendering image is then input into a target generative adversarial network model to generate a target face image corresponding to each face rendering image. By learning the mapping relationship between speech and digital human lip shape changes through the lip coefficient inference model, the temporal dependency relationship of speech features within a certain time window can be effectively obtained, enabling precise matching and synchronization between digital human lip shape and speech. By decoupling the lip shape change coefficients from the face shape coefficients through the preset FLAME general head model parameters, only the model that infers lip shape change coefficients from speech needs to be trained to generalize to different character images. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the structure of a voice-driven digital human image generation device in the hardware operating environment involved in the embodiments of the present invention;

[0042] Figure 2 This is a flowchart illustrating the first embodiment of the voice-driven digital human image generation method of the present invention.

[0043] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0044] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0045] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of a voice-driven digital human image generation device in the hardware operating environment involved in the embodiments of the present invention.

[0046] The voice-driven digital human image generation device of this invention can be a PC, or a smartphone, tablet computer, e-book reader, MP3 (Moving Picture Experts Group Audio Layer III) player, MP4 (Moving Picture Experts Group Audio Layer IV) player, portable computer, or other portable terminal devices with display functions.

[0047] like Figure 1As shown, the voice-driven digital human image generation device may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to establish communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM or a stable, non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0048] Optionally, the voice-driven digital human image generation device may also include a camera, RF (Radio Frequency) circuitry, and sensors. Of course, the voice-driven digital human image generation device may also be equipped with other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, which will not be elaborated here.

[0049] Those skilled in the art will understand that Figure 1 The terminal structure shown does not constitute a limitation on a voice-driven digital human image generation device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0050] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a voice-driven digital human image generation program.

[0051] exist Figure 1 In the voice-driven digital human image generation device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with the backend server; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; and the processor 1001 can be used to call the voice-driven digital human image generation program stored in the memory 1005.

[0052] In this embodiment, the voice-driven digital human image generation device includes: a memory 1005, a processor 1001, and a voice-driven digital human image generation program stored in the memory 1005 and executable on the processor 1001. When the processor 1001 calls the voice-driven digital human image generation program stored in the memory 1005, it executes the steps of the voice-driven digital human image generation method in the following embodiments.

[0053] This invention also provides a voice-driven digital human image generation method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the voice-driven digital human image generation method of the present invention.

[0054] In this embodiment, the voice-driven digital human image generation method includes:

[0055] Step S101: The lip shape coefficient inference model of the speech input to be predicted is used to predict multiple first 53-dimensional facial expression lip shape coefficients.

[0056] In this embodiment, when it is necessary to predict a digital human image through speech, the speech to be predicted is obtained. When the speech to be predicted is obtained, a target lip shape coefficient inference model is obtained, and the speech to be predicted is input into the target lip shape coefficient inference model for prediction. The output of the target lip shape coefficient inference model is multiple first 53-dimensional facial expression lip shape coefficients. The number of first 53-dimensional facial expression lip shape coefficients corresponds to the duration of the speech to be predicted.

[0057] It should be noted that the target lip shape coefficient inference model is a pre-trained lip shape coefficient inference model. This model is trained using the 53-dimensional facial expressions in the preset FLAME general head model parameters and the audio data in the corresponding audio-video data. Each of the first 53-dimensional facial expression lip shape coefficients corresponds to a video frame in the audio-video data.

[0058] Step S103: Perform three-dimensional facial reconstruction using the parameters other than the 53-dimensional facial expression in the preset FLAME general head model parameters and the lip shape coefficient of the first 53-dimensional facial expression to obtain multiple facial rendering images.

[0059] The preset FLAME general head model parameters are the FLAME general head model parameters corresponding to the 53-dimensional expressions of the training target lip coefficient inference model.

[0060] In this embodiment, when the first 53-dimensional expression lip shape coefficient is obtained, the other parameters in the preset FLAME general head model parameters other than the 53-dimensional expression, along with the first 53-dimensional expression lip shape coefficient, are used to perform 3D facial reconstruction. Specifically, the first 53-dimensional expression lip shape coefficient and the other parameters in the preset FLAME general head model parameters other than the 53-dimensional expression can be rendered to obtain a rendering image corresponding to each first 53-dimensional expression lip shape coefficient; a facial mask is generated based on the facial key points of the face image to be predicted; and a facial rendering image corresponding to the first 53-dimensional expression lip shape coefficient is generated based on each rendering image, the facial mask, and the face image to be predicted.

[0061] Specifically, in this embodiment, if the duration of the speech to be predicted differs from the duration of the audio-visual data corresponding to the preset FLAME general head model parameters, then when the duration of the speech to be predicted is greater than the duration of the audio-visual data, the preset parameters are rendered cyclically with the first 53-dimensional facial expression lip-shape coefficients. For example, if the number of video frames in the audio-visual data corresponding to the preset FLAME general head model parameters is n, and the number of video frames corresponding to the speech to be predicted is 2n, then the two sets of preset parameters are rendered with the first 53-dimensional facial expression lip-shape coefficients; if the number of video frames corresponding to the speech to be predicted is 1.5n, then... The system renders a set of preset parameters, the first half of a set of preset parameters, and a first 53-dimensional facial expression lip coefficient. When the duration of the speech to be predicted is less than the duration of the audio / video data, the preset parameters are repeatedly rendered with the first 53-dimensional facial expression lip coefficient. For example, if the number of video frames in the audio / video data corresponding to the preset FLAME general head model parameters is n, and the number of video frames corresponding to the speech to be predicted is 0.5n, then the first half of the preset parameters is rendered with the first 53-dimensional facial expression lip coefficient. These preset parameters are all parameters in the preset FLAME general head model parameters except for the 53-dimensional facial expression lip coefficient.

[0062] Step S104: Input each of the face rendering images into the target generative adversarial network model to generate the target face image corresponding to each of the face rendering images.

[0063] In this embodiment, after obtaining each face rendering image, each face rendering image is input into a target generative adversarial network (GAN) model to generate a target face image corresponding to each face rendering image. That is, the target GAN model predicts the corresponding target face image by inputting each face rendering image. The target face GAN model is a pre-trained GAN model.

[0064] Furthermore, in one possible implementation, after step S103, the voice-driven digital human image generation method further includes:

[0065] Step S104: Align the speech to be predicted with each of the target face images in time frames to generate a digital human audio-visual video corresponding to the speech to be predicted.

[0066] In this embodiment, after acquiring each target face image, the speech to be predicted is aligned with each target face image in time frames, and a digital human audio-visual video is generated based on the aligned speech to be predicted and the aligned target face images.

[0067] The speech-driven digital human image generation method proposed in this embodiment predicts multiple first 53-dimensional facial expression lip coefficients by inputting the speech to be predicted into a target lip coefficient inference model. Then, it performs 3D facial reconstruction using the parameters other than the 53-dimensional facial expression parameters in the preset FLAME general head model parameters and the first 53-dimensional facial expression lip coefficients to obtain multiple face rendering images. Then, it inputs each face rendering image into a target generative adversarial network model to generate the target face image corresponding to each face rendering image. By learning the mapping relationship between speech and digital human lip shape changes through the lip coefficient inference model, it can effectively obtain the temporal dependency relationship of speech features within a certain time window, enabling the matching and accurate synchronization of digital human lip shape and speech. By decoupling the lip shape change coefficient from the face shape coefficient through the preset FLAME general head model parameters, it is only necessary to train the model that infers the lip shape change coefficient from speech to generalize to different human figures.

[0068] Based on the first embodiment, a second embodiment of the voice-driven digital human image generation method of the present invention is proposed. In this embodiment, before step S101, the voice-driven digital human image generation method further includes:

[0069] Step S201: Obtain the audio data corresponding to the audio and video data to be trained and each video frame in the video data, and obtain the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image.

[0070] In this embodiment, before generating digital human images based on voice-driven technology, it is necessary to capture audio and video data that meets the requirements. During video recording, it is necessary to ensure that the green screen background light is uniform, the camera position is in the center of the left and right sides of the frame, the person's gaze is straight ahead, the recommended lens focal length is around 50mm, the aperture should not be too large (5.6 is recommended), the frame rate is 25, and the resolution is 4K. Recording should be done while the user is facing the camera, and a segment of audio and video data with a duration of, for example, 5 minutes can be recorded. It is necessary to ensure that the audio and video data is clear and the audio and video are synchronized (i.e., the video sound and the person's lip movements must match). The captured audio and video data is the audio and video data to be trained.

[0071] Furthermore, in one possible implementation, step S201 includes:

[0072] Step S2011: Obtain the audio data and video data corresponding to the audio and video data to be trained;

[0073] Step S2012: Obtain the face images corresponding to each video frame in the video data;

[0074] Step S2013: Input each of the face images into the FLAME general head model to obtain the preset FLAME general head model parameters corresponding to each video frame. The preset FLAME general head model parameters include: 100-dimensional shape, 53-dimensional expression, 50-dimensional texture, 6-dimensional lighting, and 3-dimensional projection parameters.

[0075] After obtaining the audio and video data to be trained, data processing is required. This involves acquiring the corresponding audio data and video frames, as well as the face images and preset FLAME general head model parameters for each video frame. Specifically, the audio and video data are extracted, and a foreground segmentation algorithm is used to cut out each video frame to obtain the cutout mask. The face alignment algorithm is then used to extract the 2D facial key points from each video frame, calculate the center coordinates of the key points, and crop the face image by 1.4 times the length of the longest side of the minimum bounding box of the key points. The cropping position and original size are recorded in the original image. Finally, the cropped face images are uniformly resized to 256x256 to obtain the face images corresponding to each video frame. The face images corresponding to each video frame are input into the FLAME general head model to obtain the preset FLAME general head model parameters corresponding to each video frame. This FLAME general head model can be used to reconstruct the DECA face 3D model. The preset FLAME general head model parameters include 100-dimensional shape, 53-dimensional expression, 50-dimensional texture, 6-dimensional lighting, and 3-dimensional projection parameters. Based on the shape, expression, and texture parameters, the face 3D model is reconstructed. Then, according to the lighting and projection parameters, the pytorch3d rendering tool is used to render it on the corresponding original face image. During the rendering process, only the area below the eyes of the face model is retained.

[0076] Step S202: Input the 53-dimensional facial expressions and audio data from the preset FLAME general head model parameters into the initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient.

[0077] In this embodiment, after obtaining the preset FLAME general head model parameters, the 53-dimensional facial expressions and audio data in each of the preset FLAME general head model parameters are input into the initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient. Specifically, the 53-dimensional facial expressions and audio data are first time-frame aligned to obtain training data, and the training data is input into the lip coefficient inference model for model training. When the trained lip coefficient inference model converges, the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient are obtained.

[0078] Step S203: Based on the second 53-dimensional facial expression lip shape coefficient, the preset FLAME general head model parameters, and the face image, generate a composite face image corresponding to each video frame;

[0079] In this embodiment, after obtaining the second 53-dimensional expression lip shape coefficient, the face composite image corresponding to each video frame is generated through the initial digital human model based on the second 53-dimensional expression lip shape coefficient, the preset FLAME general head model parameters, and the face image. Specifically, the second 53-dimensional expression lip shape coefficient and other parameters in the preset FLAME general head model parameters other than the 53-dimensional expression are first input into the head model for rendering to obtain the face rendering image corresponding to each video frame. The face mask is obtained based on the facial key points of the face image, and the face composite image corresponding to each video frame is generated based on the face rendering image, the face mask, and the face image.

[0080] Step S204: Input the composite face image and face image corresponding to each video frame into the initial generative adversarial network model for model training to obtain the target generative adversarial network model.

[0081] In this embodiment, after obtaining the composite face image, the composite face image and face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the trained generative adversarial network model. When the trained generative adversarial network model converges, the trained generative adversarial network model is used as the target generative adversarial network model.

[0082] The voice-driven digital human image generation method proposed in this embodiment obtains the audio data corresponding to the audio and video data to be trained, as well as each video frame in the video data, and obtains the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image; then, it inputs the 53-dimensional expression data from each preset FLAME general head model parameter and the audio data into an initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and a second 53-dimensional expression lip coefficient; then, based on the second 53-dimensional expression lip coefficient, the preset FLAME general head model parameters, and the... The process involves generating composite face images for each video frame, then inputting the composite face images and the face images into an initial generative adversarial network (GAN) model for training to obtain the target GAN model. This model learns the mapping relationship between speech and lip shape changes in a digital human through lip shape coefficient inference, effectively acquiring the temporal dependencies of speech features within a time window. This ensures accurate matching and synchronization between the digital human's lip shape and speech. By pre-setting FLAME general head model parameters, the lip shape change coefficients are decoupled from the face shape coefficients. Therefore, training only the model that infers lip shape change coefficients from speech is sufficient to generalize to different human figures.

[0083] Based on the second embodiment, a third embodiment of the speech-driven digital human image generation method of the present invention is proposed. In this embodiment, step S202 includes:

[0084] Step S301: Align the 53-dimensional facial expressions and audio data in the preset FLAME general head model parameters with time frames to obtain training data;

[0085] Step S302: Input the training data into the lip shape coefficient inference model for model training to obtain the trained lip shape coefficient inference model;

[0086] Step S303: If the first loss function of the trained lip shape coefficient inference model is less than the first preset value, then the trained lip shape coefficient inference model is used as the target lip shape coefficient inference model, and the output data of the current model training is used as the second 53-dimensional facial expression lip shape coefficient.

[0087] The initial lip coefficient inference model includes a position encoder, a Transformer encoder, and a linear layer. The Transformer encoder includes six coding layers.

[0088] The initial lip shape coefficient inference model employs a Transformer network architecture suitable for sequence modeling tasks, comprising: one position encoder, one Transformer encoder, and two linear layers. The Transformer encoder contains six encoding layers, each with eight attention heads, a hidden layer feature dimension of 512, uses ReLU activation, and has a dropout layer deactivation probability of 0.5.

[0089] In this embodiment, after obtaining the preset FLAME general head model parameters, the 53-dimensional facial expressions and audio data in each preset FLAME general head model parameter are first aligned with time frames to obtain training data.

[0090] Next, the training data is input into the initial lip-shape coefficient inference model for training to obtain the trained lip-shape coefficient inference model. Specifically, the training data is used to extract a 29-dimensional feature vector through the deepspeech speech recognition pre-trained model. The 29-dimensional feature vector is then segmented using a sequence length of 25 frames as input to the model. The segmented data is then passed through a linear layer to increase the dimensionality to 512 dimensions, and then encoded sequentially by a position encoder to incorporate temporal information, resulting in position-encoded feature data. The position-encoded features are input into a Transformer encoder to obtain 512-dimensional features. These 512-dimensional features are then reduced to 53 dimensions through a linear layer, corresponding to the output of the 53-dimensional expression coefficient, i.e., the second 53-dimensional expression lip-shape coefficient. The deepspeech speech recognition pre-trained model is used to extract the feature vector of the input speech. Since the deepspeech pre-trained model has been trained on a large-scale speech recognition dataset, it possesses generalization ability for different timbres.

[0091] Next, the first loss function of the trained lip coefficient inference model is obtained. The formula for the first loss function is:

[0092]

[0093] Where t represents a time frame. b represents the predicted expression coefficient. t This represents the true value of the facial expression coefficients obtained from 3D facial reconstruction.

[0094] If the first loss function of the trained lip shape coefficient inference model is less than the first preset value, then the trained lip shape coefficient inference model is used as the target lip shape coefficient inference model, and the output data of the current model training is used as the second 53-dimensional facial expression lip shape coefficient; otherwise, the trained lip shape coefficient inference model is used as the initial lip shape coefficient inference model, and the process returns to step S2021. During the training of the model parameters, the Adam optimizer is used to optimize the parameters. The learner rate is set to 0.0002, and the beta value is set to (0.5, 0.99). The training batch size is set to 64, allowing for 20 training rounds. After the 5th round, the learner rate linearly decreases to 0 based on the remaining rounds.

[0095] The speech-driven digital human image generation method proposed in this embodiment obtains training data by aligning the 53-dimensional facial expressions in the preset FLAME general head model parameters with audio data over time frames. The training data is then input into the initial lip shape coefficient inference model for training, resulting in a trained lip shape coefficient inference model. If the first loss function of the trained lip shape coefficient inference model is less than a first preset value, the trained lip shape coefficient inference model is used as the target lip shape coefficient inference model, and the output data of the current model training is used as the second 53-dimensional facial expression lip shape coefficient. By learning the mapping relationship between speech and digital human lip shape changes through the lip shape coefficient inference model, the temporal dependency relationship of speech features within a time window can be effectively obtained, enabling precise matching and synchronization between digital human lip shape and speech. By decoupling the lip shape change coefficient from the face shape coefficient through the preset FLAME general head model parameters, only the model inferring the lip shape change coefficient from speech needs to be trained to generalize to different human figures.

[0096] Based on the second embodiment, a fourth embodiment of the speech-driven digital human image generation method of the present invention is proposed. In this embodiment, step S203 includes:

[0097] Step S401: Render the second 53-dimensional facial expression lip coefficient and other parameters in the preset FLAME general head model parameters other than the 53-dimensional facial expression to obtain the face rendering map corresponding to each video frame.

[0098] Step S402: Based on the facial key points of each of the aforementioned face images, generate a face mask corresponding to each video frame;

[0099] Step S403: Based on the face rendering image, face mask, and face image, generate a composite face image corresponding to each video frame.

[0100] In this embodiment, after obtaining the second 53-dimensional facial expression lip shape coefficient, other parameters besides the 53-dimensional facial expression in the preset FLAME general head model parameters are obtained. The second 53-dimensional facial expression lip shape coefficient and other parameters besides the 53-dimensional facial expression in the preset FLAME general head model parameters are input into the head model for rendering to obtain the face rendering image corresponding to each video frame. The head model can be reconstructed first using the shape, expression, and texture parameters in the preset FLAME general head model parameters, and then the second 53-dimensional facial expression lip shape coefficient and other parameters besides the 53-dimensional facial expression in the preset FLAME general head model parameters are input into the head model.

[0101] Next, facial key points based on each face image are obtained, i.e., facial key points obtained during the data processing. Based on the facial key points of each face image, a face mask corresponding to each video frame is generated.

[0102] Finally, based on the face rendering image, face mask, and face image, a composite face image corresponding to each video frame is generated. That is, the composite face image is obtained by applying a texture to the face rendering image and face image using a face mask.

[0103] The voice-driven digital human image generation method proposed in this embodiment renders the second 53-dimensional lip shape coefficient and other parameters in the preset FLAME general head model parameters other than the 53-dimensional expression to obtain the face rendering map corresponding to each video frame; then, based on the facial key points of each face image, a face mask corresponding to each video frame is generated; then, based on the face rendering map, the face mask, and the face image, a composite face image corresponding to each video frame is generated. By decoupling the lip shape change coefficient from the face shape coefficient through the preset FLAME general head model parameters, only the model that infers the lip shape change coefficient from the voice needs to be trained to generalize to different human figures.

[0104] Based on the second embodiment, a fifth embodiment of the speech-driven digital human image generation method of the present invention is proposed. In this embodiment, step S204 includes:

[0105] Step S501: Input the composite face image and face image corresponding to each video frame into the initial generative adversarial network model for model training to obtain the trained generative adversarial network model.

[0106] Step S502: If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

[0107] In this embodiment, after obtaining the composite face image, the composite face image and the face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the trained generative adversarial network model.

[0108] Obtain the second loss function of the trained generative adversarial network model and determine whether the second loss function is less than a second preset value. If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model. Otherwise, the trained generative adversarial network model is used as the initial generative adversarial network model, and the process returns to step S501.

[0109] Furthermore, in one possible implementation, the generative adversarial network model includes a face discriminator and a facial component discriminator for the teeth region; step S402 includes:

[0110] Step S4021: Based on the trained generative adversarial network model, obtain the face reconstruction error, the face discrimination error corresponding to the face discriminator, and the face component error corresponding to the facial component discriminator in the tooth region.

[0111] Step S4022: Determine the second loss function based on the face reconstruction error, face discrimination error, and facial component error;

[0112] Step S4023: If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

[0113] In this embodiment, the generative adversarial network (GAN) model is used, which includes one generator and two discriminators. The generator adopts a ResNet architecture, containing two downsampling modules, nine ResBlock residual modules, and two upsampling modules. The feature dimensions are 3, 64, 128, and 256 during downsampling, 256 in the ResBlock residual modules, and 256, 128, 64, and 3 during upsampling. The regularization method is instance regularization, and the activation function is ReLU. One of the discriminators is a face discriminator, whose input is the generated complete face image (the composite face image and face image corresponding to each video frame). The network structure is PatchGAN, with three downsampling layers. The feature dimensions from input to output are 3, 64, 128, 256, 512, and 1. The regularization method is instance regularization, and the activation function is LeakyReLU. Another discriminator is the tooth region facial component discriminator. The input image is the tooth region extracted by ROI alignment (the tooth region of the face composite image corresponding to each video frame and the tooth region of the face image). The network structure is similar to the face discriminator.

[0114] In this embodiment, after obtaining the trained generative adversarial network (GAN) model, the face reconstruction error, the face discrimination error corresponding to the face discriminator, and the face component error corresponding to the facial component discriminator in the tooth region are obtained based on the trained GAN model; and the second loss function is determined based on the face reconstruction error, the face discrimination error, and the face component error; its formula is:

[0115] L total =L rec +L adv +L comp ;

[0116]

[0117]

[0118]

[0119] Among them, L total For the second loss function, L rec For face reconstruction error, L adv For face recognition error, L comp For facial component errors, y is the composite face image, D is the face image, and D is the face discriminator. For the teeth region in a composite face image, y ROI Let Ψ be the teeth region of a face image, Ψ be the intermediate feature map obtained by the facial component discriminator for the teeth region, Gram be the gram matrix, and λ be the gram matrix. l1 , λ adv , λ fs These are the weight parameters.

[0120] Both the face recognition error and the facial component error can be optimized as follows:

[0121]

[0122] Finally, if the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model. If the second loss function is greater than or equal to the second preset value, then the trained generative adversarial network model is used as the generative adversarial network model, and the process returns to step S501.

[0123] During model parameter training, the Adam optimizer was used to optimize the parameters. The learner rate was set to 0.0002, and the beta value was set to (0.5, 0.99). The training batch size was set to 16, and the training lasted for 70 epochs. After the 30th epoch, the learner rate was linearly reduced to 0 according to the remaining epochs.

[0124] The voice-driven digital human image generation method proposed in this embodiment trains an initial generative adversarial network (GAN) model by inputting the composite face image and face image corresponding to each video frame into the model. If the second loss function of the trained GAN model is less than a second preset value, the trained GAN model is used as the target GAN model. Through a generation model that refines the lip coefficient inference model rendering image into a high-fidelity photorealistic face image, combined with a carefully designed generative adversarial training method and parameter settings, as well as additional discrimination constraints and loss function design for the tooth region, the trained generation model can generate high-fidelity face images and can generate clear and stable tooth regions.

[0125] Furthermore, this embodiment of the invention also proposes a computer-readable storage medium storing a voice-driven digital human image generation program, which, when executed by the processor, implements the steps of the voice-driven digital human image generation method as described above.

[0126] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0127] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0129] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A voice-driven digital human image generation method, characterized in that, The voice-driven digital human image generation method includes the following steps: The target lip shape coefficient inference model is used to predict the speech input to obtain multiple first 53-dimensional facial expression lip shape coefficients. Using the parameters other than the 53-dimensional expression in the preset FLAME general head model parameters and the lip shape coefficient of the first 53-dimensional expression, a three-dimensional face reconstruction is performed to obtain multiple face rendering images; Each of the face rendering images is input into the target generative adversarial network model to generate the target face image corresponding to each of the face rendering images; Prior to the step of predicting the target lip shape coefficient inference model of the speech input to obtain multiple first 53-dimensional facial expression lip shape coefficients, the speech-driven digital human image generation method further includes: Obtain the audio data corresponding to the audio and video data to be trained and each video frame in the video data, and obtain the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image. The 53-dimensional facial expressions and audio data from the preset FLAME general head model parameters are input into the initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient. Based on the second 53-dimensional facial lip shape coefficient, the preset FLAME general head model parameters, and the face image, a composite face image corresponding to each video frame is generated. The composite face image and face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the target generative adversarial network model. The step of inputting the 53-dimensional facial expressions and audio data from the preset FLAME general head model parameters into the initial lip coefficient inference model for model training to obtain the target lip coefficient inference model and the second 53-dimensional facial expression lip coefficient includes: The 53-dimensional facial expressions and audio data in each of the preset FLAME general head model parameters are aligned with time frames to obtain training data. The training data is input into the initial lip coefficient inference model for model training to obtain the trained lip coefficient inference model. If the first loss function of the trained lip shape coefficient inference model is less than the first preset value, then the trained lip shape coefficient inference model is used as the target lip shape coefficient inference model, and the output data of the current model training is used as the second 53-dimensional facial expression lip shape coefficient. The initial lip coefficient inference model includes a position encoder, a Transformer encoder, and a linear layer. The Transformer encoder includes six coding layers. The steps of obtaining the audio data corresponding to the audio and video data to be trained and each video frame in the video data, and obtaining the face image corresponding to each video frame and the preset FLAME general head model parameters corresponding to the face image include: Obtain the audio and video data corresponding to the audio and video data to be processed; Obtain the face images corresponding to each video frame in the video data; Each of the aforementioned face images is input into the FLAME general head model to obtain the preset FLAME general head model parameters corresponding to each video frame. The preset FLAME general head model parameters include: 100-dimensional shape, 53-dimensional expression, 50-dimensional texture, 6-dimensional lighting, and 3-dimensional projection parameters.

2. The voice-driven digital human image generation method as described in claim 1, characterized in that, The step of generating a composite face image corresponding to each video frame based on the second 53-dimensional lip shape coefficient, the preset FLAME general head model parameters, and the face image includes: The second 53-dimensional facial lip coefficient and other parameters in the preset FLAME general head model parameters other than the 53-dimensional facial expression are rendered to obtain the face rendering map corresponding to each video frame. Based on the facial key points of each of the aforementioned facial images, a facial mask corresponding to each video frame is generated; Based on the face rendering image, face mask, and face image, a composite face image corresponding to each video frame is generated.

3. The voice-driven digital human image generation method as described in claim 1, characterized in that, The step of inputting the composite face image and face image corresponding to each video frame into the initial generative adversarial network model for model training to obtain the target generative adversarial network model includes: The composite face image and face image corresponding to each video frame are input into the initial generative adversarial network model for model training to obtain the trained generative adversarial network model. If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

4. The voice-driven digital human image generation method as described in claim 3, characterized in that, The generative adversarial network (GAN) model includes a face discriminator and a facial component discriminator for the teeth region; the step of using the trained GAN model as the target GAN model if the second loss function of the trained GAN model is less than a second preset value includes: Based on the trained generative adversarial network model, the face reconstruction error, the face discrimination error corresponding to the face discriminator, and the face component error corresponding to the facial component discriminator in the tooth region are obtained. The second loss function is determined based on the face reconstruction error, face discrimination error, and facial component error. If the second loss function of the trained generative adversarial network model is less than the second preset value, then the trained generative adversarial network model is used as the target generative adversarial network model.

5. The voice-driven digital human image generation method according to any one of claims 1 to 4, characterized in that, After the step of inputting each of the face rendering images into the target generative adversarial network model to generate the target face image corresponding to each of the face rendering images, the voice-driven digital human image generation method further includes: The speech to be predicted is aligned with each of the target face images in time frames to generate a digital human audio-visual video corresponding to the speech to be predicted.

6. A voice-driven digital human image generation device, characterized in that, The voice-driven digital human image generation device includes: a memory, a processor, and a voice-driven digital human image generation program stored in the memory and executable on the processor. When the voice-driven digital human image generation program is executed by the processor, it implements the steps of the voice-driven digital human image generation method as described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a voice-driven digital human image generation program, which, when executed by a processor, implements the steps of the voice-driven digital human image generation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Three-dimensional virtual image lip shape generation method and device and electronic equipment

    CN113256821A

  • Virtual face generation method, apparatus and device, and readable storage medium

    CN115423908A