Digital human generation method, device, equipment and medium

In the process of digital life generation, the mouth area information is determined using target audio and picture reconstruction parameters to generate natural digital human postures, which solves the problem of unnatural postures in the existing technology and improves the user experience.

CN113886641BActive Publication Date: 2025-08-26SHENZHEN ZHUIYI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111165980.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-08-26
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

In the existing solution to generate digital people based on sound reasoning, the facial posture is unnatural, resulting in the generated digital people's posture unnatural, affecting the user experience.

Method used

By inputting the target audio into the pre-trained first generator, the target expression parameters of the character are obtained, the target 3D face reconstruction parameters are extracted from the target picture, and the image is processed to remove the mouth area. The target expression parameters and 3D face reconstruction parameters are used to determine the target mouth area information, and input it into the pre-trained second generator to generate a digital human picture.

Benefits of technology

The generated digital human posture is more natural and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886641B_ABST
    Figure CN113886641B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, device, and medium for generating a digital human, and relates to the field of computer technology. The method comprises: inputting target audio into a pre-trained first generator to obtain target expression parameters of the human; extracting target 3D face reconstruction parameters of the human from a target image, and processing the target image to obtain a first intermediate image that does not include the human's mouth area; determining target mouth area information based on the target expression parameters and the target 3D face reconstruction parameters; and inputting the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image. This method can make the posture of the digital human generated based on sound inference more natural, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for generating a digital human. Background Art

[0002] Digital humans utilize information science methods to simulate the human body at various levels of form and function. With the rapid development of computer technology, digital human generation technology is becoming increasingly mature. In practice, if digital human generation technology is to be applied commercially and achieve real-time interaction, the digital human generation solution must meet at least two requirements: good generation results and high inference efficiency. Good generation results are a prerequisite, while high inference efficiency is a commercial necessity.

[0003] At present, in order to improve the reasoning efficiency of digital humans, a solution for generating digital humans based on sound reasoning has emerged. It mainly uses a sound reasoning model to generate facial key points, then draws the facial key points into contour lines, inputs them into the generative adversarial network model, and finally generates a digital human.

[0004] However, the facial key points inferred based on sound contain facial posture information, and this facial posture information has an angle problem. Therefore, the posture of the digital human finally generated by applying the above solution is unnatural. Summary of the Invention

[0005] In view of this, the present application provides a method, apparatus, device and medium for generating a digital human, so that the posture of the digital human generated based on sound reasoning can be more natural, thereby improving the user experience.

[0006] In a first aspect, an embodiment of the present application provides a method for generating a digital human, comprising:

[0007] Input the target audio into the pre-trained first generator to obtain the target expression parameters of the character;

[0008] Extracting target 3D face reconstruction parameters of the person from the target image, and processing the target image to obtain a first intermediate image that does not include the person's mouth area;

[0009] Determining target mouth region information based on the target expression parameter and the target 3D face reconstruction parameter;

[0010] The target mouth area information and the first intermediate image are input into a pre-trained second generator to obtain a digital human image.

[0011] Optionally, determining target mouth region information based on the target expression parameter and the target 3D face reconstruction parameter includes:

[0012] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model to obtain a target three-dimensional face mesh;

[0013] Determining a target three-dimensional mouth region mesh from the target three-dimensional face mesh;

[0014] The target three-dimensional mouth region mesh is determined as target mouth region information.

[0015] Optionally, determining target mouth region information based on the target expression parameter and the target 3D face reconstruction parameter includes:

[0016] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D face deformation statistical model to obtain a plurality of target mouth area key points;

[0017] A plurality of target mouth region key points are determined as target mouth region information.

[0018] Optionally, inputting the target mouth region information and the first intermediate image into a pre-trained second generator to obtain a digital human image includes:

[0019] Merging the target mouth area information and the first intermediate image into a second intermediate image in a channel merging manner;

[0020] The second intermediate image is input into a pre-trained second generator to obtain a digital human image.

[0021] Optionally, processing the target image to obtain an intermediate image that does not include the person's mouth area includes:

[0022] Use the preset image detection algorithm to determine the mouth area of ​​the person from the target image;

[0023] The pixel values ​​of the pixels in the character's mouth area are set to preset values ​​to obtain an intermediate image that does not include the character's mouth area.

[0024] Optionally, the method further includes:

[0025] According to the time sequence of the target audios, the digital human images corresponding to the target audios are combined to generate a digital human video.

[0026] Optionally, the first generator is trained in the following manner:

[0027] Extracting a plurality of picture frames and audio frames corresponding to the picture frames from a video stream;

[0028] Perform the following operations on each of the image frames to obtain a plurality of first sample data:

[0029] Extracting sample expression parameters of a person from the picture frame; extracting sample audio features from the audio frame corresponding to the picture frame; determining the sample audio features and the sample expression parameters as first sample data;

[0030] Model training is performed based on a number of the first sample data to obtain the first generator.

[0031] Optionally, extracting sample expression parameters of a person from the picture frame includes:

[0032] Extracting a number of facial key points from the image frame using a preset key point detection algorithm;

[0033] A number of the facial key points are input into a preset facial 3D deformation statistical model to obtain sample expression parameters of the person.

[0034] Optionally, extracting sample audio features from the audio frame corresponding to the picture frame includes:

[0035] Extracting Mel-frequency cepstral coefficients using Fourier transform as sample audio features of the audio frame corresponding to the picture frame; or,

[0036] A preset speech recognition model is used to extract sample audio features from the audio frame corresponding to the image frame.

[0037] Optionally, performing model training based on a plurality of first sample data to obtain the first generator includes:

[0038] Inputting the sample audio features in each of the first sample data into an initial first generator to obtain corresponding predicted expression parameters;

[0039] Determining a model loss value according to the predicted expression parameters and the sample expression parameters corresponding to each of the first sample data;

[0040] If the model loss value does not meet the preset model convergence conditions, the model parameters of the first generator are updated based on the model loss value, and the first generator after the updated model parameters is iteratively trained until the model loss value meets the model convergence conditions, thereby obtaining the first generator.

[0041] In a second aspect, an embodiment of the present application provides a digital human generation device, comprising:

[0042] An expression parameter extraction module is used to input the target audio into a pre-trained first generator to obtain the target expression parameters of the character;

[0043] 3D face reconstruction parameter extraction module, used to extract the target 3D face reconstruction parameters of the person from the target image;

[0044] An image processing module is used to process the target image to obtain a first intermediate image that does not contain the character's mouth area;

[0045] A mouth region information determination module, configured to determine target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters;

[0046] The digital human generation module is used to input the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image.

[0047] Optionally, the mouth region information determination module is specifically configured to:

[0048] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model to obtain a target three-dimensional face mesh;

[0049] Determining a target three-dimensional mouth region mesh from the target three-dimensional face mesh;

[0050] The target three-dimensional mouth region mesh is determined as target mouth region information.

[0051] Optionally, the mouth region information determination module is specifically configured to:

[0052] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D face deformation statistical model to obtain a plurality of target mouth area key points;

[0053] A plurality of target mouth region key points are determined as target mouth region information.

[0054] Optionally, the digital human generation module is specifically used to:

[0055] Merging the target mouth area information and the first intermediate image into a second intermediate image in a channel merging manner;

[0056] The second intermediate image is input into a pre-trained second generator to obtain a digital human image.

[0057] Optionally, the image processing module is specifically configured to:

[0058] Use the preset image detection algorithm to determine the mouth area of ​​the person from the target image;

[0059] The pixel values ​​of the pixels in the character's mouth area are set to preset values ​​to obtain an intermediate image that does not include the character's mouth area.

[0060] Optionally, the device further includes:

[0061] The video generation module is used to combine the digital human images corresponding to the target audios according to the time sequence of the target audios to generate a digital human video.

[0062] Optionally, the device further includes:

[0063] The model training module is used to extract a number of picture frames and audio frames corresponding to the picture frames from a video stream; perform the following operations on each of the picture frames to obtain a number of first sample data: extract sample expression parameters of the character from the picture frame; extract sample audio features from the audio frames corresponding to the picture frame; determine the sample audio features and the sample expression parameters as first sample data; perform model training based on the number of first sample data to obtain the first generator.

[0064] Optionally, the model training module extracts sample expression parameters of the person from the picture frame, including:

[0065] A preset key point detection algorithm is used to extract a number of facial key points from the image frame; the number of facial key points is input into a preset facial 3D deformation statistical model to obtain sample expression parameters of the person.

[0066] Optionally, the model training module extracts sample audio features from the audio frame corresponding to the image frame, including:

[0067] Mel-frequency cepstral coefficients are extracted using Fourier transform as sample audio features of the audio frame corresponding to the picture frame; or, sample audio features are extracted from the audio frame corresponding to the picture frame using a preset speech recognition model.

[0068] Optionally, the model training module performs model training based on a plurality of the first sample data to obtain the first generator, including:

[0069] The sample audio features in each of the first sample data are input into the initial first generator to obtain the corresponding predicted expression parameters; the model loss value is determined based on the predicted expression parameters and the sample expression parameters corresponding to each of the first sample data; if the model loss value does not meet the preset model convergence conditions, the model parameters of the first generator are updated based on the model loss value, and the first generator after the updated model parameters is iteratively trained until the model loss value meets the model convergence conditions, thereby obtaining the first generator.

[0070] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the steps of the digital human generation method as described in any one of the first aspects when executing the program stored in the memory.

[0071] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the digital human generation method as described in any one of the first aspects.

[0072] The technical solution provided by the embodiment of the present application obtains the target expression parameters of the character by inputting the target audio into a pre-trained first generator, extracts the target 3D face reconstruction parameters of the character from the target image, and processes the target image to obtain a first intermediate image that does not include the character's mouth area. The target mouth area information is determined based on the target expression parameters and the target 3D face reconstruction parameters, and the target mouth area information and the first intermediate image are input into a pre-trained second generator to obtain a digital human image. Since the target mouth area information includes the opening and closing state information of the character's mouth but does not include facial posture information, the target mouth area information does not have an angle problem. Therefore, the posture of the digital human generated based on the target mouth area information and the image that does not include the character's mouth area is more natural, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0075] Figure 1 This is a flowchart of the steps of a method for generating a digital human provided in an embodiment of the present application;

[0076] Figure 2 This is a flowchart of another method for generating a digital human provided by an embodiment of the present application;

[0077] Figure 3 A flowchart of a method for generating a digital human provided in an optional embodiment of the present application;

[0078] Figure 4 This is a flowchart for the specific implementation of step 320 in a method for generating a digital human provided in an optional embodiment of the present application;

[0079] Figure 5 A flowchart for the specific implementation of step 330 in a method for generating a digital human provided in an optional embodiment of the present application;

[0080] Figure 6 This is a structural block diagram of a digital human generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0081] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0082] See also Figure 1 , shows a flowchart of the steps of a digital human generation method provided by an embodiment of the present application. Specifically, the digital human generation method provided by the present application can be applied to video generation scenarios, such as the case of generating a virtual image of a video based on a real image; wherein the virtual image can be an image of a digital human in a video, specifically used to represent a digital human in a digital human video. Specifically, as an example, the video generation scenario is a live video generation scenario. In this scenario, the digital human generation method provided by the present application is applied to the user audio and the image of the live broadcast host (such as a virtual host), so that the live broadcast host's image can be driven by the user audio, thereby generating a live video of the virtual host. As another example, the video generation scenario is an online education video generation scenario. In this scenario, the digital human generation method provided by the present application is applied to the lecturer's audio and the lecturer's image (such as a virtual lecturer), so that the lecturer's image can be driven by the lecturer's audio, thereby generating a video of the virtual lecturer giving an online lecture. Of course, the above application scenarios are merely exemplary application scenarios of the digital human generation method provided by the present application, and the embodiments of the present application do not limit the specific application scenarios.

[0083] like Figure 1 As shown, the digital human generation method in the embodiment of the present application may specifically include the following steps:

[0084] Step 110: Input the target audio into a pre-trained first generator to obtain the target expression parameters of the character.

[0085] The target audio can refer to the real audio to be processed, such as audio recorded by the user or audio from a video recorded by the user. Furthermore, in the embodiment of the present application, the target audio can be the audio contained in the final generated digital human video, that is, the audio output by the final generated digital human.

[0086] The first generator is a pre-trained machine learning model used to infer the character's expression parameters based on audio. The expression parameters are used to represent the opening and closing state of the character's mouth. As for how the first generator is trained, we will explain it in the following. Figure 3 、 Figure 4 as well as Figure 5 The process shown is explained below and will not be described in detail here.

[0087] Based on the above description, in step 110, the target audio is input into the pre-trained first generator to obtain the character's expression parameters (for the convenience of description, referred to as target expression parameters).

[0088] Step 120: extract target 3D face reconstruction parameters of the person from the target image.

[0089] The target image may refer to a real image to be processed. In a specific implementation, the target image may be obtained by capturing an image or video of a person through an image acquisition device.

[0090] 3D face reconstruction parameters include but are not limited to: face shape information, reflection information, texture information, lighting information, etc.

[0091] In a specific implementation, a preset 3D face reconstruction parameter extraction algorithm or 3D face reconstruction parameter extraction model can be used to extract 3D face reconstruction parameters from the target image. As for the specific 3D face reconstruction parameter extraction algorithm or 3D face reconstruction parameter extraction model, this embodiment of the application does not elaborate in detail.

[0092] Step 130: Process the target image to obtain a first intermediate image that does not include the person's mouth area.

[0093] In one embodiment, the specific implementation of step 130 includes: using a preset image detection algorithm to determine the mouth area of ​​the person from the target image, and then setting the pixel values ​​of the pixels in the mouth area to a preset value (for example, 0) to obtain an image that does not include the mouth area of ​​the person (for convenience of description, hereinafter referred to as the first intermediate image).

[0094] Step 140 : Determine target mouth region information based on target expression parameters and target 3D face reconstruction parameters.

[0095] First, it should be noted that the target mouth area information in the embodiment of the present application includes the opening and closing state information of the person's mouth, and does not include facial posture information.

[0096] Specifically, in one embodiment, step 140 includes inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model (e.g., 3DMM) to obtain a 3D face mesh (hereinafter referred to as the target 3D face mesh for ease of description). It will be appreciated that since the target 3D face mesh is obtained based on both the target expression parameters and the target 3D face reconstruction parameters, the target 3D face mesh includes information about the opening and closing state of the person's mouth.

[0097] Afterwards, a target three-dimensional mouth region mesh is determined from the target three-dimensional face mesh, and the target three-dimensional mouth region mesh is determined as target mouth region information.

[0098] In another embodiment, the specific implementation of step 140 includes: inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model (such as 3DMM), obtaining a number of target mouth area key points, and determining the several target mouth area key points as the target mouth area information.

[0099] It should be noted that the target mouth area key points in the embodiment of the present application are determined based on the target expression parameters and the target 3D face reconstruction parameters, and the target 3D face reconstruction parameters are extracted from the target image. Therefore, this is different from the facial key points determined simply based on audio in the prior art.

[0100] Step 150: Input the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image.

[0101] The above-mentioned second generator is pre-trained and is a machine learning model used to infer and generate digital human images based on the person's mouth area information and the picture that does not contain the person's mouth area (or the picture obtained by merging the two in a channel merging manner).

[0102] Based on this, in one embodiment, the target mouth region information and the first intermediate image can be merged using a channel merging method (for ease of description, the image obtained by the merging process is referred to as the second intermediate image). The second intermediate image is then input into a pre-trained second generator to obtain a digital human image.

[0103] In another embodiment, the target mouth region information and the first intermediate image may be directly input into a pre-trained second generator to obtain a digital human image.

[0104] It can be seen that the technical solution provided by the embodiment of the present application obtains the target expression parameters of the character by inputting the target audio into a pre-trained first generator, extracts the target 3D face reconstruction parameters of the character from the target image, and processes the target image to obtain a first intermediate image that does not contain the character's mouth area. Based on the target expression parameters and the target 3D face reconstruction parameters, the target mouth area information is determined, and the target mouth area information and the first intermediate image are input into a pre-trained second generator to obtain a digital human image. Since the target mouth area information includes the opening and closing state information of the character's mouth but does not include facial posture information, the target mouth area information does not have an angle problem. Therefore, the posture of the digital human generated based on the target mouth area information and the image that does not contain the character's mouth area is more natural, thereby improving the user experience.

[0105] See also Figure 2 , shows a flowchart of another method for generating a digital human provided by an embodiment of the present application. Figure 2 As shown, the digital human generation method in the embodiment of the present application may specifically include the following steps:

[0106] Step 210: Input the target audio into a pre-trained first generator to obtain the target expression parameters of the character.

[0107] Step 220: extract target 3D face reconstruction parameters of the person from the target image.

[0108] Step 230: Process the target image to obtain a first intermediate image that does not include the person's mouth area.

[0109] Step 240 : Determine target mouth region information based on target expression parameters and target 3D face reconstruction parameters.

[0110] Step 250: Input the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image.

[0111] For detailed description of steps 210 to 250, please refer to the above Figure 1 The description of the process shown will not be repeated here.

[0112] Step 260: Combine the digital human images corresponding to the target audios according to the time sequence of the target audios to generate a digital human video.

[0113] In the embodiment of the present application, the digital human generation method provided in the embodiment of the present application can be applied to a segment of audio (such as a segment of audio in wav format with a frame rate of 100) to ultimately generate a digital human video.

[0114] Specifically, each audio frame in the audio segment can be identified as a target audio segment. Then, steps 240 to 250 are performed for each target audio segment to obtain a digital human image corresponding to each target audio segment. Finally, the digital human images corresponding to the target audio segments are combined in the time sequence of the target audio segments to generate a digital human video.

[0115] It can be seen that the technical solution provided by the embodiment of the present application obtains the target expression parameters of the character by inputting the target audio into a pre-trained first generator, extracts the target 3D face reconstruction parameters of the character from the target image, and processes the target image to obtain a first intermediate image that does not contain the character's mouth area. Based on the target expression parameters and the target 3D face reconstruction parameters, the target mouth area information is determined, and the target mouth area information and the first intermediate image are input into a pre-trained second generator to obtain a digital human image. Since the target mouth area information includes the opening and closing state information of the character's mouth but does not include facial posture information, the target mouth area information does not have an angle problem. Therefore, the posture of the digital human generated based on the target mouth area information and the image that does not contain the character's mouth area is more natural, thereby improving the user experience.

[0116] Furthermore, the technical solution provided in the embodiment of the present application can generate a digital human video by combining digital human images corresponding to several target audios in the time sequence of the several target audios. This enables the technical solution provided in the embodiment of the present application to be applied to a variety of video generation scenarios, such as the live video generation scenario and online education video generation scenario exemplified above.

[0117] See also Figure 3 , shows a flowchart of a digital human generation method provided by an optional embodiment of the present application. Specifically, the digital human generation method provided by the embodiment of the present application may include the following steps in the model training phase:

[0118] Step 310: extract a number of picture frames and audio frames corresponding to the picture frames from the video stream.

[0119] The video stream can refer to a real video stream to be processed, such as a video stream recorded by a user. In a video stream, each video frame contains an audio frame and an image frame. For example, if a one-second video stream contains five video frames, then the video stream contains five audio frames and five image frames, that is, there is a one-to-one correspondence between audio frames and image frames.

[0120] Step 320: Obtain a first sample data for each picture frame.

[0121] See also Figure 4 The specific implementation of step 320 may include the following steps:

[0122] Step 3201: extract the sample expression parameters of the character from the picture frame.

[0123] In the specific implementation, a preset key point detection algorithm can be used to extract several facial key points (such as 68 facial key points) from the picture frame. Then, several facial key points are input into a preset facial 3D deformation statistical model to obtain the character's expression parameters (for convenience of description, called facial expression parameters).

[0124] Step 3202: extract sample audio features from the audio frame corresponding to the picture frame.

[0125] In one embodiment, Fourier transform is used to extract Mel-frequency cepstral coefficients as audio features of an audio frame corresponding to a picture frame (referred to as sample audio features for ease of description).

[0126] In another embodiment, a preset speech recognition model is used to extract sample audio features from the audio frame corresponding to the picture frame.

[0127] Step 3203: Determine the sample audio features and sample expression parameters as first sample data.

[0128] It should be noted that, in the first sample data, the sample audio feature is the input value, and the sample expression parameter is the label value.

[0129] Step 330: Perform model training based on a plurality of first sample data to obtain a first generator.

[0130] See also Figure 5 The specific implementation of step 330 may include the following steps:

[0131] Step 3401: Input the sample audio features in each first sample data into the initial first generator to obtain the corresponding predicted expression parameters.

[0132] In an embodiment of the present application, the first generator may be a CNN model, an LSTM (Long Short-Term Memory) model, etc. The embodiment of the present invention does not limit the model structure adopted by the first generator.

[0133] Step 3402: Determine the model loss value based on the predicted expression parameters and sample expression parameters corresponding to each first sample data.

[0134] In one embodiment, a L2 Loss loss function or an L1 Loss loss function may be used to determine the model loss value according to the predicted expression parameters and the sample expression parameters corresponding to each first sample data.

[0135] In another embodiment, the Wing Loss loss function and the L1 Loss loss function can be used to determine the model loss value based on the predicted expression parameters and sample expression parameters corresponding to each first sample data. Specifically, the loss value information obtained using different loss functions can be directly summed or weighted summed, and the summed result can be determined as the model loss value.

[0136] Of course, in a specific implementation, other loss functions can also be used to determine the model loss value. The above is only an exemplary description and the embodiments of the present application do not limit this.

[0137] Step 3403: If the model loss value does not meet the preset model convergence conditions, the model parameters of the first generator are updated based on the model loss value, and the first generator after the updated model parameters is iteratively trained until the model loss value meets the model convergence conditions, thereby obtaining the first generator.

[0138] It can be seen that in the embodiment of the present application, the first generator can be obtained based on real video stream training, so that in actual applications, the first generator can be used to infer the facial expression parameters of the character based on the audio.

[0139] The present application also provides a digital human generation device. Figure 6 As shown, the digital human generation device 600 provided in this embodiment of the application may include the following modules:

[0140] The expression parameter extraction module 610 is used to input the target audio into the pre-trained first generator to obtain the target expression parameters of the character;

[0141] 3D face reconstruction parameter extraction module 620, for extracting target 3D face reconstruction parameters of a person from a target image;

[0142] An image processing module 630 is configured to process the target image to obtain a first intermediate image that does not include the mouth area of ​​the person;

[0143] A mouth region information determination module 640 is configured to determine target mouth region information based on the target expression parameters and the target 3D face reconstruction parameters;

[0144] The digital human generation module 650 is configured to input the target mouth region information and the first intermediate image into a pre-trained second generator to obtain a digital human image.

[0145] Optionally, the mouth region information determination module 640 is specifically configured to:

[0146] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model to obtain a target three-dimensional face mesh;

[0147] Determining a target three-dimensional mouth region mesh from the target three-dimensional face mesh;

[0148] The target three-dimensional mouth region mesh is determined as target mouth region information.

[0149] Optionally, the mouth region information determination module 640 is specifically configured to:

[0150] Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D face deformation statistical model to obtain a plurality of target mouth area key points;

[0151] A plurality of target mouth region key points are determined as target mouth region information.

[0152] Optionally, the digital human generation module 650 is specifically configured to:

[0153] Merging the target mouth area information and the first intermediate image into a second intermediate image in a channel merging manner;

[0154] The second intermediate image is input into a pre-trained second generator to obtain a digital human image.

[0155] Optionally, the image processing module 630 is specifically configured to:

[0156] Use the preset image detection algorithm to determine the mouth area of ​​the person from the target image;

[0157] The pixel values ​​of the pixels in the character's mouth area are set to preset values ​​to obtain an intermediate image that does not include the character's mouth area.

[0158] Optionally, the device further includes (not shown in the figure):

[0159] The video generation module is used to combine the digital human images corresponding to the target audios according to the time sequence of the target audios to generate a digital human video.

[0160] Optionally, the device further includes (not shown in the figure):

[0161] The model training module is used to extract a number of picture frames and audio frames corresponding to the picture frames from a video stream; perform the following operations on each of the picture frames to obtain a number of first sample data: extract sample expression parameters of the character from the picture frame; extract sample audio features from the audio frames corresponding to the picture frame; determine the sample audio features and the sample expression parameters as first sample data; perform model training based on the number of first sample data to obtain the first generator.

[0162] Optionally, the model training module extracts sample expression parameters of the person from the picture frame, including:

[0163] A preset key point detection algorithm is used to extract a number of facial key points from the image frame; the number of facial key points is input into a preset facial 3D deformation statistical model to obtain sample expression parameters of the person.

[0164] Optionally, the model training module extracts sample audio features from the audio frame corresponding to the image frame, including:

[0165] Mel-frequency cepstral coefficients are extracted using Fourier transform as sample audio features of the audio frame corresponding to the picture frame; or, sample audio features are extracted from the audio frame corresponding to the picture frame using a preset speech recognition model.

[0166] Optionally, the model training module performs model training based on a plurality of the first sample data to obtain the first generator, including:

[0167] The sample audio features in each of the first sample data are input into the initial first generator to obtain the corresponding predicted expression parameters; the model loss value is determined based on the predicted expression parameters and the sample expression parameters corresponding to each of the first sample data; if the model loss value does not meet the preset model convergence conditions, the model parameters of the first generator are updated based on the model loss value, and the first generator after the updated model parameters is iteratively trained until the model loss value meets the model convergence conditions, thereby obtaining the first generator.

[0168] It should be noted that the digital human generation device provided above can execute the image processing method provided in any embodiment of the present application, and has the corresponding functions and beneficial effects of the execution method.

[0169] In a specific implementation, the above-mentioned digital human generation device can be applied to electronic devices such as personal computers and servers, so that the electronic devices can act as image processing devices to generate digital humans based on target audio, and make the generated digital humans' postures more natural, thereby improving the user experience.

[0170] Furthermore, an embodiment of the present application also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the steps of the digital human generation method described in any of the above method embodiments when executing the program stored in the memory.

[0171] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the digital human generation method described in any one of the above method embodiments are implemented.

[0172] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between the various embodiments can be referred to in conjunction with each other. For the device, equipment, and storage medium embodiments, since they are generally similar to the method embodiments, their descriptions are relatively simple. For relevant parts, refer to the descriptions of the method embodiments.

[0173] In this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0174] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A method for generating a digital human, characterized in that: include: Input the target audio into the pre-trained first generator to obtain the target expression parameters of the character; Extracting target 3D face reconstruction parameters of the person from the target image, and processing the target image to obtain a first intermediate image that does not include the person's mouth area; Determining target mouth region information based on the target expression parameter and the target 3D face reconstruction parameter, wherein the target mouth region information includes information about the opening and closing state of the person's mouth, but does not include facial posture information; Inputting the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image; Wherein, determining target mouth area information based on the target expression parameter and the target 3D face reconstruction parameter includes: Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D face deformation statistical model to obtain a plurality of target mouth area key points; determining a plurality of target mouth region key points as target mouth region information; The step of inputting the target mouth area information and the first intermediate image into a pre-trained second generator to obtain a digital human image includes: Merging the target mouth area information and the first intermediate image into a second intermediate image in a channel merging manner; The second intermediate image is input into a pre-trained second generator to obtain a digital human image.

2. The method according to claim 1, characterized in that The determining target mouth region information according to the target expression parameter and the target 3D face reconstruction parameter includes: Inputting the target expression parameters and the target 3D face reconstruction parameters into a preset 3D facial deformation statistical model to obtain a target three-dimensional face mesh; Determining a target three-dimensional mouth region mesh from the target three-dimensional face mesh; The target three-dimensional mouth region mesh is determined as target mouth region information.

3. The method according to claim 1, characterized in that The target image is processed to obtain an intermediate image excluding the mouth area of ​​the person, including: Use the preset image detection algorithm to determine the mouth area of ​​the person from the target image; The pixel values ​​of the pixels in the character's mouth area are set to preset values ​​to obtain an intermediate image that does not include the character's mouth area.

4. The method according to claim 1, wherein The method further comprises: According to the time sequence of the target audios, the digital human images corresponding to the target audios are combined to generate a digital human video.

5. The method according to claim 1, characterized in that The first generator is trained in the following way: Extracting a plurality of picture frames and audio frames corresponding to the picture frames from a video stream; Perform the following operations on each of the image frames to obtain a plurality of first sample data: Extracting sample expression parameters of the character from the picture frame; extracting sample audio features from the audio frame corresponding to the picture frame; Determining the sample audio feature and the sample expression parameter as first sample data; Model training is performed based on a number of the first sample data to obtain the first generator.

6. The method according to claim 5, characterized in that The step of extracting sample expression parameters of a person from the picture frame includes: Extracting a number of facial key points from the image frame using a preset key point detection algorithm; A number of the facial key points are input into a preset facial 3D deformation statistical model to obtain sample expression parameters of the person.

7. The method according to claim 5, characterized in that The extracting the sample audio features from the audio frame corresponding to the picture frame includes: Extracting Mel-frequency cepstral coefficients using Fourier transform as sample audio features of the audio frame corresponding to the picture frame; or, A preset speech recognition model is used to extract sample audio features from the audio frame corresponding to the image frame.

8. The method according to claim 5, characterized in that The performing model training based on the plurality of first sample data to obtain the first generator includes: Inputting the sample audio features in each of the first sample data into an initial first generator to obtain corresponding predicted expression parameters; Determining a model loss value according to the predicted expression parameters and the sample expression parameters corresponding to each of the first sample data; If the model loss value does not meet the preset model convergence conditions, the model parameters of the first generator are updated based on the model loss value, and the first generator after the updated model parameters is iteratively trained until the model loss value meets the model convergence conditions, thereby obtaining the first generator.

9. A digital human generation device, characterized in that: include: An expression parameter extraction module is used to input the target audio into a pre-trained first generator to obtain the target expression parameters of the character; 3D face reconstruction parameter extraction module, used to extract the target 3D face reconstruction parameters of the person from the target image; An image processing module is used to process the target image to obtain a first intermediate image that does not contain the character's mouth area; a mouth region information determination module, configured to determine target mouth region information based on the target expression parameter and the target 3D face reconstruction parameter, wherein the target mouth region information includes information on the opening and closing state of the person's mouth, but does not include facial posture information; a digital human generation module, configured to input the target mouth region information and the first intermediate image into a pre-trained second generator to obtain a digital human image; The mouth region information determination module is specifically configured to input the target expression parameter and the target 3D face reconstruction parameter into a preset 3D face deformation statistical model to obtain a plurality of target mouth region key points; and determine the plurality of target mouth region key points as target mouth region information; The digital human generation module is specifically used to: Merging the target mouth area information and the first intermediate image into a second intermediate image in a channel merging manner; The second intermediate image is input into a pre-trained second generator to obtain a digital human image.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the steps of the digital human generation method according to any one of claims 1 to 8 when executing the program stored in the memory.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the digital human generation method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and computer storage medium

    CN110677598A