Digital human generation method and apparatus, device, medium, and product
Patent Information
- Application Number
- PCT/CN2026/073086
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-01-16
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026073086_01102026_PF_FP_ABST
Abstract
Description
Methods, apparatus, equipment, media and products for digital human generation
[0001] Related applications
[0002] This application claims priority to Chinese patent application filed on March 25, 2025, with application number 202510371811.0 and entitled “Digital Human Generation Method, Apparatus, Device, Medium and Product”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of video processing technology, and more specifically, to a digital human generation method, a digital human generation device, an electronic device, a computer-readable storage medium, and a computer product. Background Technology
[0004] Due to the development of artificial intelligence technology, digital humans, as an emerging technological product, are digital human figures created using digital technology that closely resemble human appearances. However, in the current process of generating digital humans, due to the drawbacks of the technology and the defects of artificial intelligence models, there are still shortcomings in the facial features of digital humans, such as the lack of facial details leading to facial blurring. Summary of the Invention
[0005] Embodiments of this application provide a digital human generation method, a digital human generation apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the embodiments of this application, a digital human generation method is provided, comprising: acquiring input audio; inputting the input audio into a pre-trained generator network, wherein the generator network is obtained by adversarial training of the generator network based on a sample digital human image output by a network to be trained and a quality discrimination result output by a quality discriminator for the sample digital human image; the sample digital human image is generated by the network to be trained based on a sample facial video and sample audio of the sample facial video; and acquiring an output digital human image generated by the generator network corresponding to the input audio, wherein the output digital human image is used for a video content creation task based on the input audio.
[0007] According to one aspect of the embodiments of this application, a digital human generation apparatus is provided, comprising: an acquisition module for acquiring input audio; an input module for inputting the input audio into a pre-trained generation network, wherein the generation network is obtained by adversarial training of the network under training based on a sample digital human image output by a network under training and a quality discrimination result output by a quality discriminator for the sample digital human image; the sample digital human image is generated by the network under training based on a sample facial video and sample audio of the sample facial video; the acquisition module is further configured to acquire an output digital human image generated by the generation network corresponding to the input audio, wherein the output digital human image is used to process a video content creation task corresponding to the input audio.
[0008] According to one aspect of the embodiments of this application, an electronic device is provided, including one or more processors; and a storage device for storing one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the digital human generation method as described above.
[0009] According to one aspect of the embodiments of this application, an embodiment of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform the digital human generation method as described above.
[0010] According to one aspect of the embodiments of this application, an embodiment of this application provides a computer program product, including a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the digital human generation method described above.
[0011] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0013] Figure 1 is a schematic diagram of an implementation environment involved in this application;
[0014] Figure 2 is a flowchart illustrating a digital human generation method in an exemplary embodiment of this application;
[0015] Figure 3 is a schematic diagram illustrating a network training and generation process in an exemplary embodiment of this application;
[0016] Figure 4 is a schematic diagram illustrating the training process of a generative network in an exemplary embodiment of this application;
[0017] Figure 5 is a comparative schematic diagram illustrating a facial super-resolution reconstruction in an exemplary embodiment of this application;
[0018] Figure 6 is a flowchart illustrating another digital human generation method in an exemplary embodiment of this application;
[0019] Figure 7 is a flowchart illustrating a digital human generation method in an exemplary embodiment of this application;
[0020] Figure 8 is a schematic diagram illustrating a live streaming creation interface in an exemplary embodiment of this application;
[0021] Figure 9 is a schematic diagram illustrating a live streaming settings interface in an exemplary embodiment of this application;
[0022] Figure 10 is a schematic diagram of a template interface shown in an exemplary embodiment of this application;
[0023] Figure 11 is a structural block diagram of a digital human generation device illustrating an exemplary embodiment of this application;
[0024] Figure 12 shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device according to the embodiments of this application. Detailed Implementation
[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0026] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations, nor do they necessarily have to be executed in the described order. For example, some operations can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0028] It should also be noted that "multiple" as mentioned in this application refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0029] The technical solutions of the embodiments of this application will be described in detail below.
[0030] Please refer to Figure 1, which is a schematic diagram of an implementation environment related to this application. The implementation environment includes a terminal 10 and a server 20.
[0031] Terminal 10 is used to input audio and send the input audio to server 20.
[0032] Server 20 is used to input the input audio into a pre-trained generator network. The generator network is obtained by adversarially training the network under test based on the sample digital human image output by the network under test and the quality discrimination result of the quality discriminator outputting the sample digital human image. The sample digital human image is generated by the network under test based on the sample facial video and the sample audio of the sample facial video. The server 20 also obtains the output digital human image generated by the generator network that corresponds to the input audio.
[0033] The server can also send the output digital human image to the terminal, so that the terminal can perform the task of creating video content corresponding to the input audio based on the output digital human image.
[0034] In some embodiments, the server 20 can pre-train a generator network and send the generator network to the terminal for storage, so that the terminal can also independently implement the digital human generation process.
[0035] In some embodiments, the server may also acquire the input audio itself to generate the output digital human image through a pre-trained generative network.
[0036] The aforementioned terminal 10 can be any electronic device capable of acquiring modal data of the target object, such as a smartphone, tablet, laptop, computer, smart voice interaction device, smart home appliance, vehicle terminal, or aircraft. The server 20 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This document does not impose any restrictions on this.
[0037] Terminal 10 and server 20 establish a communication connection via a network beforehand, enabling them to communicate with each other. The network can be a wired network or a wireless network, and this is not a limitation.
[0038] It should be noted that in the specific embodiments of this application, facial videos, audio, digital human images, and other data related to the subject are involved. When the embodiments of this application are applied to specific products or technologies, permission or consent from the subject is required, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0039] The following describes in detail the various implementation details of the technical solutions in the embodiments of this application.
[0040] As shown in Figure 2, which is a flowchart of a digital human generation method according to an embodiment of this application, the method can be applied to the implementation environment shown in Figure 1. The method can be executed by a terminal or a server, or by both a terminal and a server. In this embodiment of the application, the method is described as being executed by a server. The digital human generation method can include S210 to S230, which are described in detail below.
[0041] S210, Obtain input audio.
[0042] In this embodiment of the application, the input audio refers to audio that includes voice information, which includes language content in Chinese and English. The number of input audio files can be one or more.
[0043] In one example, the input audio could be audio uploaded by an object or audio fetched from the network.
[0044] S220. Input the input audio into the pre-trained generator network. The generator network is obtained by adversarial training of the network to be trained based on the sample digital human image output by the network to be trained and the quality discrimination result output by the quality discriminator on the sample digital human image. The sample digital human image is generated by the network to be trained based on the sample facial video and the sample audio of the sample facial video.
[0045] S230. Obtain the output digital human image generated by the generative network corresponding to the input audio. The output digital human image is used for video content creation tasks based on the input audio.
[0046] In this embodiment of the application, a generative network is pre-trained to generate a corresponding digital human image based on audio, so as to simulate the appearance, expression and movement of a person. The digital human image includes detailed facial expressions and dynamic movements. Therefore, as shown in Figure 3, the corresponding output digital human image can be obtained by inputting the input audio into the generative network. The output digital human image includes facial expressions or lip-synchronized dynamic effects that are synchronized with the audio.
[0047] As shown in Figure 3, the network to be trained is subjected to adversarial training to obtain a generative network. During the training process, the network to be trained generates sample digital human images based on sample facial videos and sample audio of the sample facial videos. The sample facial videos refer to sample videos that include any type of facial images. The sample videos contain facial images from different frames, and the facial images include facial structure, texture, features, etc. The facial types of the facial data include human faces and non-human faces, such as animal faces and human-animal hybrid faces. The quality discriminator outputs a quality discrimination result for the sample digital human images. This quality discrimination result is used to indicate whether the face of the sample digital human image is clear and realistic enough. Then, adversarial training is performed on the network to be trained based on the quality discrimination result and the sample digital human images. That is, the network to be trained acts as a generator, and the quality discriminator acts as a discriminator. Through adversarial training, the digital human images generated by the network to be trained can deceive the discriminator, so that the generated faces are as close as possible to real high definition and realism.
[0048] Optionally, the network to be trained includes Neural Radiance Fields (NeRF) networks, GAN-based facial animation networks (GANimation), generative models (Stable Diffusion, SD), etc.
[0049] In one example, the output digital human avatar generated by the generative network is used for video content creation tasks based on the input audio. Video content creation tasks refer to creating video content from the output digital human avatar, such as using the output digital human avatar as a character in virtual reality, augmented reality (VR / AR), or video games; another example is using the output digital human avatar for virtual live streaming and real-time interaction in scenarios such as digital humans, virtual anchors, and virtual customer service; yet another example is using the output digital human avatar for video creation in distance education and film and television production.
[0050] It should be noted that although a quality discriminator is added during training, no additional computation is required when generating the network application. Therefore, it does not increase the inference time of the generator network, meeting the requirements of scenarios with high real-time requirements such as live streaming. Optionally, the input audio is audio used to introduce the item to be live streamed; the video content creation task includes a live streaming content creation task. After step S210, it further includes: obtaining the corresponding live streaming scene based on the item to be live streamed and the live streaming content; after step S230, it further includes: projecting the output digital human image onto the live streaming scene to generate the video live streaming content corresponding to the live streaming content creation task.
[0051] Among them, the item to be live-streamed refers to the specific item that needs to be displayed in the current live stream, which may be a product, model, prop, etc.; the live-streaming scene refers to the virtual environment or background setting related to these items; for example, if the live stream is an electronic product, then the item is the product itself, and the scene is a virtual store or demonstration space that displays the product; multiple virtual environment templates can be pre-stored, and the matching live-streaming scene can be selected from multiple virtual environment templates according to the item type of the item to be live-streamed and the live stream content.
[0052] The input audio used to introduce the items to be live-streamed is fed into the generator network, which can quickly generate an output digital human avatar. The output digital human avatar is then placed in a virtual live-streaming scene and rendered in real time to obtain the live video content. At this point, the output digital human avatar will blend with the scene to form a complete live video content. The actions, expressions, and language of the digital human avatar will be adjusted synchronously with the audio, presenting a smooth and highly interactive live-streaming effect while ensuring real-time performance.
[0053] 3D modeling software can be used to create 3D models for the live streaming scene and corresponding 3D character models for the output digital avatar. The 3D model of the digital avatar is then placed into the 3D model of the live streaming scene according to the set position and pose information. A real-time rendering engine (such as Unreal Engine or Unity) is used to render the scene and the digital avatar in real time based on the lighting conditions, material properties, and the digital avatar's movements and expressions. During the rendering process, an audio synchronization algorithm is used to synchronize the digital avatar's movements, expressions, and speech with the input audio, ultimately generating smooth and highly interactive live video content.
[0054] Understandably, one input audio corresponds to one output digital human image. In addition to generating digital human images corresponding to audio in real time, multiple different types of input audio can be input into the generation network to obtain multiple different types of output digital human images. These multiple different types of output digital human images are stored. When it is necessary to live stream an item, the live streaming digital human image that matches the item to be live streamed can be selected from multiple output digital human images based on the item to be live streamed and the live stream content. For example, if the item to be live streamed is women's clothing and the live stream content is to showcase women's clothing, then a female digital human image with a suitable figure can be selected from multiple types of output digital human images as the live streaming digital human image. The live streaming digital human image is then projected onto the corresponding live stream scene for real-time rendering.
[0055] Specifically, attribute tags can be pre-defined for each output digital avatar, including information such as gender, body shape, and style. Simultaneously, matching rules are defined for the items to be live-streamed and the live-stream content. For example, if the item to be live-streamed is women's clothing and the live-stream content is showcasing women's clothing, the matching rule is set to select digital avatars that are female, have a body shape that matches the clothing size, and a style that matches the clothing. Based on these rules, the attribute tags of multiple output digital avatars are traversed to filter out the live-stream digital avatars that meet the criteria.
[0056] In this embodiment, the input audio is fed into a pre-trained generator network. This generator network is obtained through adversarial training based on the sample digital human image output by the network to be trained and the quality discrimination result output by the quality discriminator for the sample digital human image. The sample digital human image is generated by the network to be trained based on sample facial videos and sample audio of the sample facial videos. That is, the network to be trained is supervised by the quality discriminator to determine the key facial areas that the network to be trained should focus on, reducing blurring and loss of detail. Through adversarial training, the network to be trained generates higher definition and more natural digital human faces during the continuous optimization process. Furthermore, by combining audio-driven training, the generated digital human can accurately match the input audio, achieving high-precision lip-sync and effectively avoiding problems such as lip distortion, delay, or misalignment. This improves the clarity and realism of the digital human face generated by the generator network, thereby obtaining the output digital human image generated by the generator network corresponding to the input audio. This image can be directly used for video content creation tasks without additional high-definition processing after generation, reducing computational overhead and improving real-time performance and applicability.
[0057] In one embodiment of this application, another digital human generation method is provided. This digital human generation method can be applied to the implementation environment shown in Figure 1. Taking the method executed by the server as an example, as shown in Figure 4, this digital human generation method adds a training process for the generation network based on S210 to S230 shown in Figure 2, including steps S410 to S440, which are described in detail below.
[0058] S410: Obtain sample facial videos, sample audio of the sample facial videos, and the network to be trained.
[0059] In this embodiment of the application, a video set containing a speaking face can be obtained, and sample facial videos without speaking sound can be extracted from the video set. The sample facial videos contain faces but do not contain facial expressions or speaking actions. The sample audio can be audio extracted from the sample speech information of the speaking in the video set.
[0060] Optionally, video processing tools (such as FFmpeg) can be used to process the video set, separating the audio tracks from the video using audio separation functionality. For extracting sample facial videos, a video frame extraction algorithm is used to extract image frames from the video at a fixed frame rate, removing frames containing facial expressions and speaking actions, and combining the remaining image frames into a sample facial video without speaking sound. For extracting sample audio, speech recognition and segmentation are performed on the separated audio tracks to extract the sample speech information, which is then saved as an audio file.
[0061] Optionally, taking facial video as an example, during the training of the generative network, the input facial video and audio are key data for generating digital human images. It is understood that training samples under single angle and lighting conditions may lead to poor performance of the generated faces under different environments and conditions, failing to fully capture the diversity and complexity of faces. Therefore, in this embodiment, obtaining sample facial video includes: obtaining the initial sample facial video corresponding to the sample audio; and rendering the initial sample facial video according to multiple viewpoints and multiple different lighting conditions to obtain the sample facial video.
[0062] If the initial sample facial video is extracted from a video set containing faces but excluding facial expressions and speech, sample facial image frames are extracted from the initial sample facial video. These sample facial image frames are then rendered in 3D space from different viewpoints (e.g., 0°, 30°, etc.) to generate image frames of the same face from different angles. Specifically, 3D reconstruction techniques can be used to reconstruct a 3D facial model from the sample facial image frames. In 3D space, with the 3D facial model as the center, the camera position and orientation are determined according to the set viewpoint angles (e.g., 0°, 30°, etc.). Using a rendering engine (such as Blender's Cycles renderer), the 3D facial model is rendered based on the camera position and orientation, generating image frames of the same face at different angles. During the rendering process, appropriate lighting conditions, material properties, and other parameters can be set to improve the realism of the image.
[0063] Data augmentation techniques, such as rotation, affine, and perspective transformations, can be used to rotate, shift, and stretch sample face image frames at different angles, thus expanding the diversity of training samples. For rotation transformations, rotation functions in image processing libraries (such as OpenCV) can be used, specifying the rotation center, rotation angle, and scaling factor to rotate the sample face image frames. For affine transformations, the correspondence between three non-collinear points in the original and transformed images can be defined, the affine transformation matrix can be calculated, and then the affine transformation function can be used to transform the image, achieving translation, rotation, and scaling. For perspective transformations, similarly, the correspondence between four points in the original and transformed images can be defined, the perspective transformation matrix can be calculated, and the perspective transformation function can be used to transform the image, achieving a perspective effect. These transformation methods allow sample face image frames to be rotated, shifted, and stretched at different angles, expanding the diversity of training samples.
[0064] Different lighting conditions can be simulated by applying high dynamic range (HDR) rendering technology to sample face image frames, or by adding different local light sources (e.g., simulating facial light reflection and shadows) to the sample face image frames, thereby obtaining face image frames with different brightness, shadows, and highlight contrast. This can be achieved by adding different lighting conditions to face image frames at different angles, or by adding different lighting conditions to sample face image frames. Local light sources are light sources used in image processing to simulate lighting effects in specific areas. In this application, by adding different local light sources to sample face image frames, facial light reflection and shadows can be simulated, resulting in images with different brightness, shadows, and highlight contrast, thereby expanding the diversity of training samples.
[0065] The sample facial video is obtained by adding facial image frames from different angles and under different lighting conditions to the initial sample facial video.
[0066] Through the above approach, the amplification of multiple angles and lighting allows the network to better capture facial details (such as skin texture, light reflection, and facial shadows). By using various lighting and angles, the network learns how to preserve subtle facial details, such as the shape of the lips, the expression of the eyes, and the tiny shadows on the face, thereby improving the quality of the network's generation and making the generated digital human image more realistic.
[0067] S420. Perform facial super-resolution reconstruction on the facial regions in the sample facial video to obtain the super-resolution sample facial video.
[0068] It is understandable that, due to the difficulty of capturing or acquiring high-resolution data, it is sometimes difficult to obtain high-definition face training data. Networks trained with low-resolution data can only generate low-resolution faces. Therefore, in this embodiment, the facial region, i.e. the face, in the sample facial video is first subjected to facial super-resolution reconstruction to directly improve the clarity of the original training data and obtain a super-resolution sample facial video, as shown in Figure 5, which is a schematic diagram of the effect of facial super-resolution reconstruction.
[0069] Among them, facial super-resolution reconstruction refers to the image processing technology of face super-resolution (FSR), which aims to recover high-resolution detail information from low-resolution facial images.
[0070] Optionally, facial super-resolution reconstruction of the facial region includes: extracting global facial features from the facial region through a first convolutional layer; extracting local facial features from the global facial features through a second convolutional layer, wherein the convolutional kernel of the second convolutional layer is smaller than the convolutional kernel of the first convolutional layer; and reconstructing the facial region based on the local facial features to obtain a super-resolution sample facial video.
[0071] The process involves reconstructing details by progressively extracting features through multiple convolutional layers. The first convolutional layer uses a large kernel (e.g., 9×9) to capture global contextual information of the facial region and obtain global facial features. If the dimension of the facial region is 3×H×W, representing the RGB channels, then the feature dimension of the global facial features is 64×H×W, representing low-level features in the facial region, such as global background information and main contours and edge features (e.g., the outer contour line of the face). The second convolutional layer uses a small kernel (e.g., 3×3) to further extract finer-grained local facial features based on the global facial features. For example, the output dimension of the local facial features is 32×H×W.
[0072] Optionally, there can be multiple second convolutional layers. Each second convolutional layer can focus on capturing details at different levels, such as the edge shape of the eyes, the curvature of the lips, and the shadow of the bridge of the nose. The local features extracted by each second convolutional layer are fused to enhance regional contrast and obtain the final facial local features.
[0073] Optionally, residual modules can be added on top of each convolutional layer. The residual blocks are used to learn the differences between the input and output, rather than learning the entire mapping directly, in order to better learn high-frequency details.
[0074] In one example, the third convolutional layer maps local facial features to an RGB image to generate a high-resolution image. That is, based on the extraction of local and global features, the third convolutional layer gradually restores the details in the image (such as facial contours, facial features, and skin texture).
[0075] In another example, the global facial features output from the second convolutional layer are weighted by an attention layer. Then, a third convolutional layer reconstructs a higher-resolution image from these weighted global facial features. The attention layer generates an attention map and weights it pixel-by-pixel with the global facial features, allowing the third convolutional layer to focus more on key regions in the image. The attention map, generated by the attention layer, assigns a weight value to each pixel in the image, representing its importance in the image generation process. By weighting the attention map pixel-by-pixel with the global facial features, subsequent convolutional layers can focus more on key regions in the image, thereby improving the quality of the generated image.
[0076] The above method involves performing facial super-resolution reconstruction on the facial region, which is then fed into the training network to ensure that the network can capture richer details and more accurate scene structure, thereby directly improving the clarity of the facial region in the final generated digital human image.
[0077] S430: Input the super-resolution sample facial video and sample audio into the network to be trained to obtain sample digital human images.
[0078] In this embodiment of the application, a super-resolution sample facial video containing a clear facial region and sample audio are input into the network to be trained to obtain a sample digital human image.
[0079] Optionally, the network to be trained is used to extract features from the sample audio and the super-resolution sample facial video to obtain sample audio features and sample facial video features respectively, and to perform position encoding on the spatial coordinates and viewpoint direction corresponding to the super-resolution sample facial video to obtain sample geometric encoding features. The sample audio features and sample facial video features are aligned according to the time scale to obtain target sample audio features. Based on the target sample audio features, sample geometric encoding features, and sample facial video features, a sample digital human image is generated.
[0080] The process involves extracting features from sample audio data using the feature layer of the network under training. These features include intonation, speech rate, pronunciation variations, and rhythm information, reflecting the dynamic changes in the vocalization process. This sample audio feature is a multi-dimensional feature vector (e.g., in time-series format) used to guide the dynamic changes in facial movement generation. Similarly, the process also involves extracting features from super-resolution sample facial videos using the feature layer. This extraction includes both static and dynamic facial information, such as expressions and posture. Static information refers to relatively fixed facial features, such as the basic facial contours and the approximate positions of facial features. Dynamic information refers to the changing features of the face at different times, such as changes in expression and adjustments in posture.
[0081] The facial video features of the sample are also a multi-dimensional feature vector (such as in time series form); spatial coordinates (x, y, z), viewpoint direction Used to construct 3D scenes and capture viewpoint information, the hidden layers of the network under training store spatial coordinates (x, y, z) and viewpoint direction. Positional encoding is performed to obtain geometrically encoded features of the samples through high-dimensional expansion. Positional encoding enhances the network's ability to learn complex details, especially facial texture and lighting variations, by introducing high-frequency information. Positional encoding can employ a combination of sine and cosine functions; for each dimension i of the spatial coordinates, the encoding formula is as follows: and Where pos is the position and d is the encoding dimension. In this way, low-dimensional spatial coordinates and viewpoint orientation information are extended to high dimensions to obtain geometric encoding features of the samples, thereby enhancing the network's ability to learn complex details.
[0082] Audio is based on time-series features, while the network to be trained processes point sampling and facial features in three-dimensional space during the generation process. In order to effectively map the dynamic information in the sample audio to the generation process, the sample audio features and sample facial features are aligned in time scale. The time dimension of the audio features is aligned with the frame generation cycle of the network to be trained, ensuring that the audio and lip movements at the same time point are matched. In one example, the sample facial features and sample geometric coding features have the same time dimension. The alignment process makes the audio features at each time step match the three-dimensional coordinate features of the current frame. The alignment process can be achieved by compressing or expanding the audio sequence through linear transformation or recurrent neural network (RNN).
[0083] Optionally, a sample digital human image is generated based on the target sample audio features, sample geometric coding features, and sample facial video features, including: extracting sample spatial features from the sample geometric coding features using a multilayer perceptron of the network to be trained; mapping the sample spatial features to obtain the volume density of each spatial point using a fully connected layer of the network to be trained; generating the color value of each spatial point based on the sample viewpoint coding features from the target sample audio features, sample facial video features, and sample geometric coding features; and generating the sample digital human image based on the volume density and color value of each spatial point.
[0084] The multilayer perceptron includes multiple hidden layers, each containing a certain number of neurons (e.g., 256). Features of spatial points are extracted using a non-linear activation function (e.g., ReLU). The high-dimensional vector obtained from the position encoding, i.e., the sample geometric encoding features, is input into the multilayer perceptron for spatial feature extraction to understand the spatial distribution and dynamic changes of the face. Then, a fully connected layer maps the sample spatial features to output the volume density of each spatial point. The volume density represents whether a point is on the object's surface and its probability of occupancy, used for subsequent ray integration and volume rendering. Ray integration, in the process of generating the sample digital human image, is the process of mathematically calculating and accumulating the propagation of light in space and its interaction with objects based on the volume density and RGB color values of each spatial point, thus obtaining the final pixel color. It is one of the key steps in achieving volume rendering. Volume rendering is a rendering technique that generates two-dimensional images based on three-dimensional data fields. In this application, based on the volume density and color values of each spatial point, the three-dimensional sample digital human image data is converted into a two-dimensional visual image using methods such as ray integration, thereby generating the final sample digital human image.
[0085] As described earlier, the viewpoint direction is extended to a high-dimensional representation through positional encoding. Therefore, the viewpoint encoding features of the samples can be extracted from the geometric encoding features of the samples. These viewpoint encoding features reflect viewpoint-related light reflection and texture changes. The sample facial video features, target sample audio features, and sample viewpoint encoding features are fused to obtain the sample fusion features. These fusion features are high-dimensional feature representations with orientation sensitivity. The sample fusion features are feature-mapped through a fully connected layer to output the RGB colors of spatial points. These RGB colors reflect the lighting and color information of the sample facial features under a specific viewpoint. Combining the lighting changes and facial dynamics in the video, a sample digital face is generated based on the volume density and RGB color value of each spatial point. For example, for each spatial point, a volume-weighted integral is performed based on its volume density and RGB value to obtain the final pixel color, which is then rendered to obtain the sample digital human image.
[0086] Through the above scheme, by using position encoding, audio feature extraction, and a multilayer perceptron (MLP) structure, the network to be trained can simultaneously process spatial and perspective information, and combine audio information to achieve precise synchronization of facial expressions and speech, ultimately generating a dynamic and realistic digital human image.
[0087] Optionally, a sample digital human image is generated based on the target sample audio features, sample geometric coding features, and sample facial video features, including: calculating the attention weight between the target sample audio features and sample geometric coding features; weighting the input audio features according to the attention weight to obtain weighted sample audio features; and generating a sample digital human image based on the weighted sample audio features, sample geometric coding features, and sample facial video features.
[0088] The attention weights are calculated by determining the correlation between the target sample audio features and the sample geometric coding features. For example, a query vector is generated based on the sample geometric coding features, and key and value vectors are generated based on the target sample audio features. Then, a dot product attention mechanism is used to calculate the attention weights between the target sample audio features and the sample geometric coding features. The input audio features are then weighted according to these attention weights; features with higher weights have a more significant impact on image generation. The key and value vectors are two types of vectors generated in the attention mechanism calculation based on the target sample audio features through different linear transformations. The key vector is used to perform a dot product operation with the query vector to calculate the attention weights, while the value vectors are weighted and summed according to the attention weights, ultimately participating in the subsequent image generation process. The linear transformation W is performed based on the sample geometric coding features G. q Generate query vector Q = W q G, based on the target sample audio features A, is transformed by linear transformation W. k and W vGenerate the key vectors K = W respectively. k A and value vector V = W v A. Then, the attention weights are calculated based on the dot product attention mechanism, using the following formula: Where d k It is the dimension of the key vector.
[0089] Optionally, weighted audio features, sample geometric coding features, and sample facial video features can be fused to generate a sample digital human image based on the fused features; alternatively, the sample geometric coding features can be processed through a multilayer perceptron and a fully connected layer to obtain the volume density of each spatial point, and the color value of each spatial point can be generated based on the sample viewpoint coding features in the weighted sample audio features, sample facial video features, and sample geometric coding features; the sample digital human image can be generated based on the volume density and color value of each spatial point.
[0090] S440. Input the sample digital human image into the quality discriminator to obtain the quality discrimination result, and perform adversarial training on the network to be trained based on the sample digital human image and the quality discrimination result to obtain the generator network.
[0091] In this embodiment, the sample digital human image is input into the quality discriminator to obtain a discrimination result based on the clarity of the face and the accuracy of the mouth to determine whether the sample digital human image is real. The network to be trained is then subjected to adversarial training based on the sample digital human image and the quality discrimination result to obtain the generator network.
[0092] The above approach first performs super-resolution on the facial regions of the sample facial videos, and then feeds them into the training network to increase the clarity of the trained faces. On this basis, an adversarial loss is added to the training network using a facial / lip quality discriminator to focus on high-definition facial details, thereby further increasing the clarity of the generated faces.
[0093] This application provides another method for generating a digital human. This method can be applied to the implementation environment shown in Figure 1. The method can be executed by a terminal or a server, or by both a terminal and a server. In this application embodiment, the method is described using the server as an example. As shown in Figure 6, this digital human generation method expands S440 shown in Figure 4 to S610-S630 based on the method shown in Figure 4. S610-S630 are described in detail below:
[0094] S610. Obtain the real-life image corresponding to the sample facial video and sample audio.
[0095] Specifically, video frames from the same video segment as the sample facial video and sample audio can be selected from the original video set. The human figure presented in the video frame is the real human figure corresponding to the sample facial video and sample audio.
[0096] S620. Calculate and generate a loss function based on the difference between the sample digital human image and the real human image.
[0097] Among them, sample facial videos and sample audio can be extracted from the video set corresponding to real human images. The sample digital human image is the "fake" image generated by the network to be trained. Therefore, the generation loss function can be calculated based on the difference between the sample digital human image and the real human image. For example, the generation loss function can be obtained by calculating the cross-entropy function based on the difference.
[0098] Optionally, the generation loss function includes multiple task losses. Calculating the generation loss function includes: generating a first generation loss based on the pixel differences between the sample digital human image corresponding to the sample digital human image and the real digital human image corresponding to the real human image; calculating a first optical flow between adjacent digital human image frames of the sample digital human image and a second optical flow between adjacent digital human image frames corresponding to the real human image; generating a second generation loss based on the optical flow difference between the first and second optical flows; and calculating the generation loss function based on the first and second generation losses.
[0099] Optical flow is a concept used to describe the motion of objects in an image. It represents the speed and direction of motion of each pixel in an image across different image frames over time. In this application, by calculating the first optical flow between adjacent digital human image frames of the sample digital human image and the second optical flow between adjacent digital human image frames corresponding to the real human image, and generating a second generation loss based on the difference in optical flow between the two, the dynamic facial features of the sample digital human image have smooth motion changes between frames, avoiding jitter or unnatural mouth movements.
[0100] The first generation loss is calculated based on the mean squared difference of pixel errors between the sample digital human image corresponding to the sample digital human image and the real digital human image corresponding to the real human image, to ensure basic shape fitting. For example, the first generation loss can be the mean squared error loss.
[0101] Where n is the number of digital human figures in the digital sample, y i It is the real digital human image of the i-th sample. It is the sample digital human image of the i-th sample.
[0102] It is understandable that the image frames of the sample digital human image may have temporal discontinuities, causing jitter in the mouth or facial expressions between adjacent frames. Therefore, to ensure that the generated face is not only clear but also consistent with the input sample audio, this embodiment calculates optical flow loss, i.e., the second generation loss, to ensure that the dynamic facial features of the sample digital human image have smooth motion changes between frames, avoiding jitter or unnatural mouth movements. Specifically, two adjacent digital human image frames corresponding to the sample digital human image are selected, and the first optical flow is calculated using an optical flow estimation algorithm; two adjacent digital human image frames corresponding to the real human image are selected, and the second optical flow of the adjacent digital human image frames is calculated using an optical flow estimation algorithm. The second generation loss is generated based on the difference between the first and second optical flows. For example, the first optical flow is: The second optical flow is: The second generation loss is: FlowNet is an optical flow estimation algorithm, and ||·||1 represents the L1 norm, which is used to measure the difference in optical flow fields.
[0103] In one example, the generation loss function can be obtained by directly summing the first generation loss and the second generation loss, or by weighted summing the first generation loss and the second generation loss. The weights corresponding to the first generation loss and the second generation loss can be flexibly adjusted according to the actual situation.
[0104] S630. Input the sample digital human image into the quality discriminator to obtain the quality discrimination result. Generate an adversarial loss function based on the quality discrimination result. Adjust the network parameters of the network to be trained based on the generation loss function and the adversarial loss function to obtain the generation network.
[0105] In this embodiment of the application, the sample digital human image is input into the quality discriminator to obtain the quality discrimination result. The quality discrimination result is the probability that the sample digital human image is fake. The adversarial loss function is the negative log-likelihood of the probability that the sample digital human image is fake, or 1 minus the average probability that the sample digital human image is fake.
[0106] Optionally, the quality discriminator includes a face quality discriminator and a lip quality discriminator. The face quality discriminator targets the entire face and is used to determine whether the face of the sample digital human image is realistic enough. The lip quality discriminator targets the lips and is used to determine whether the lips of the sample digital human image are realistic enough. Inputting the sample digital human image into the quality discriminator includes: acquiring sample face images and sample lip images of the sample digital human image. For example, the sample face image can be extracted from the sample digital human image, and the sample lip image can be cropped from the sample face image. Then, the sample face image is input into the face quality discriminator to obtain the face quality discrimination result, and the sample lip image is input into the lip quality discriminator to obtain the lip quality discrimination result. The face quality discrimination result ensures that the network to be trained generates a clear face, and the lip quality discrimination result ensures that the network to be trained can generate clear lip details, making the lip shape more natural and matching the audio.
[0107] Through the above scheme, the facial quality discriminator makes a judgment on the face as a whole. Based on the judgment of facial quality, since the lips are the most critical part of speech-driven facial generation, a dedicated lip quality discriminator can ensure that the network to be trained learns the details of the mouth more accurately, making the lip shape more natural and matching the audio, thereby improving the final generated facial quality.
[0108] Optionally, the sample lip image is input to the lip quality discriminator to obtain the lip quality discrimination result, including: inputting the sample lip image to the lip quality discriminator, the lip quality discriminator is used to downsample the sample lip image to generate multiple sub-sample lip images of different resolutions, extracting features from the multiple sub-sample lip images to obtain multiple sample lip features, and generating the lip quality discrimination result based on the multiple sample lip features.
[0109] When judging the quality of sample lip images, the lip quality discriminator can introduce multi-scale features. The lip quality discriminator makes judgments at different resolutions and scales, allowing the network to be trained to focus on local details.
[0110] The lip quality discriminator progressively downsamples the sample lip images, generating multiple sub-sample lip images at different resolutions (e.g., original resolution, 1 / 2 resolution, 1 / 4 resolution, etc.). The downsampling is performed a maximum of n times, with a minimum resolution of at least m×m pixels to ensure effective feature extraction. The lip quality discriminator can design a separate branch for each resolution, with each branch consisting of convolutional layers. Each branch extracts features at its corresponding resolution to obtain sample lip features, and then makes a judgment based on these features. The judgment results from different resolution branches are then fused to obtain the final lip quality judgment result.
[0111] Optionally, an attention mechanism can be used to focus on important lip features of samples at different resolutions, and the discrimination results of different resolution branches can be fused based on attention weights to obtain the lip quality discrimination result.
[0112] In the case where the quality discriminator includes a face quality discriminator and a lip quality discriminator, an adversarial loss function is generated based on the quality discrimination results, including: generating a face quality loss function based on the face quality discrimination results and generating a lip quality loss function based on the lip quality discrimination results; and generating an adversarial loss function based on the face quality loss function and the lip quality loss function.
[0113] Specifically, the quality loss function of the facial quality discriminator is obtained by calculating the cross-entropy function based on the facial quality discrimination results, and the facial quality loss function is obtained based on the quality loss function; for example, the quality loss function of the facial quality discriminator is: The facial quality loss function for the network to be trained is: Wherein, D1(x fake D1(x) represents the probability that the face quality discriminator outputs a false result. real The probability of a face quality discriminator being true is given by the face quality discriminator; similarly, the quality loss function of the lip quality discriminator is: The lip quality loss function of the network to be trained is D2(x fake D2(x) represents the probability of a false positive output by the lip quality discriminator. real ) represents the probability of the lip quality discriminator outputting a true judgment.
[0114] The adversarial loss function is obtained by weighted summation of the facial quality loss function and the lip quality loss function.
[0115] In this embodiment, the generative loss function and the adversarial loss function are weighted and fused to obtain the total loss function. The model parameters of the network to be trained are adjusted according to the total loss function to obtain the generative network. The weights corresponding to the generative loss function and the adversarial loss function can be flexibly adjusted according to the actual situation.
[0116] Optionally, adjusting the model parameters of the network to be trained to obtain the generative network includes: S1, generating a discriminant loss function based on the quality discrimination result, fixing the network parameters of the network to be trained, and adjusting the model parameters of the quality discriminator based on the discriminant loss function to obtain a new quality discriminator; S2, fixing the model parameters of the new quality discriminator, and adjusting the network parameters of the network to be trained based on the total loss function to obtain a primary generative network; S3, inputting the new super-resolution sample facial video and the corresponding new sample audio into the primary generative network to obtain a new sample digital human image, and inputting the new sample digital human image into the new quality discriminator to obtain a new quality discrimination result; S4, repeating the alternating training process of the new quality discriminator and the primary generative network based on the new sample digital human image and the new quality discrimination result until the primary generative network converges to obtain the generative network.
[0117] In step S3, new super-resolution sample facial videos and corresponding new sample audios are obtained from the pre-prepared new training dataset using the same processing method as in the initial training phase. The new super-resolution sample facial videos and new sample audios are preprocessed according to the format and dimensions required by the input layer of the primary generator network, and then the processed data is accurately input into the corresponding input port of the primary generator network to obtain the new sample digital human image.
[0118] In step S4, based on the new sample digitized human image and the new quality discrimination result, the following alternating training process is repeated: First, a discriminative loss function is generated based on the new quality discrimination result, and the network parameters of the primary generator network are fixed. Backpropagation and gradient descent are used to adjust the model parameters of the new quality discriminator according to the discriminative loss function. Then, the model parameters of the new quality discriminator are fixed, and the generation loss function and adversarial loss function are recalculated based on the new sample digitized human image and the new quality discrimination result. The two are weighted and summed to obtain a new total loss function. Backpropagation and gradient descent are used to adjust the network parameters of the primary generator network according to the new total loss function. These alternating training steps are repeated until the primary generator network converges to obtain the generator network.
[0119] In this embodiment, during adversarial training, the model parameters of the network to be trained and the model parameters of the quality discriminator are updated. The training of the quality discriminator and the network to be trained is performed alternately, forming an iterative training process. First, the quality discriminator is trained to accurately distinguish between real and generated data. A discriminant loss function, which can be a binary cross-entropy loss, is generated based on the quality discrimination result. With the network parameters of the network to be trained fixed, the model parameters of the quality discriminator are updated using backpropagation and gradient descent to minimize the discriminant loss function, resulting in a new quality discriminator. Then, the network to be trained is trained, and the generated fake data can deceive the discriminator. With the parameters of the new quality discriminator fixed, the parameters of the network to be trained are updated using backpropagation and gradient descent to minimize the total loss function, resulting in a primary generator network. This completes one training cycle, i.e., the first iteration of training.
[0120] When the network parameters of the network to be trained are fixed, the gradient of the discriminant loss function with respect to the parameters (weights and biases) of each layer of the new quality discriminator is first calculated. Taking the weight W of a certain layer as an example, the gradient is calculated using the chain rule. Where L discriminator To determine the loss function, the same calculation is performed for bias b. Then, the parameters are updated using gradient descent, with the update formula as follows: and Where η is the learning rate. A new quality discriminator is obtained by iteratively updating the discriminant loss function multiple times.
[0121] During the subsequent training of the network to be trained, the parameters of the new quality discriminator are fixed, and the total loss function L is calculated. total The gradients of the parameters (weights and biases) of each layer of the network to be trained are calculated using the chain rule. and Where W′ and b′ are the weights and biases of the network to be trained. Then, gradient descent is used to update the parameters, with the update formula being W′=W′- and Where η′ is the learning rate. The initial generative network is obtained by iteratively updating the network multiple times to minimize the total loss function.
[0122] Suppose the quality discriminator outputs a series of probability values P = {p1, p2, ..., p...} that indicate a sample belongs to the true sample. n The actual labels are Y = {y1, y2, ..., y}. n}, where y i The value can be 0 (indicating the sample is fake) or 1 (indicating the sample is real). The discriminant loss function uses binary cross-entropy loss, and the formula is as follows: Substituting the probability values from the quality discrimination results and the corresponding true labels into this formula generates the discriminative loss function. Here, n is the number of samples, and p... i y is the probability that the i-th sample is identified as a real sample. i is the true label of the i-th sample, and the result obtained by calculating this formula is the value of the discriminant loss function.
[0123] In the second iteration of training, it is necessary to reacquire training samples. Therefore, new super-resolution sample facial videos and corresponding new sample audio are acquired. New sample digital human images are generated through the primary generative network. New quality discrimination results are obtained through the new quality discrimination results. S1 and S2 are repeated, that is, the primary generative network is fixed, the model parameters of the new quality discrimination are updated according to the new quality discrimination results, the model parameters of the new quality discriminator are fixed, and the network parameters of the primary generative network are updated according to the new total loss function generated by the new quality discrimination results and the new sample digital human images, so as to complete one training cycle.
[0124] Optionally, after the number of iterations reaches a preset number, the performance of the updated primary generator network can be checked to see if it has reached the convergence condition. The convergence condition is that the primary generator network can generate high-quality fake samples, and the discriminator has difficulty distinguishing between real and fake samples. If the convergence condition is not reached, iterative training continues. If the convergence condition is reached, the latest primary generator network obtained from the last iteration is used as the final generator network. Through continuous iterative training and parameter adjustment, the primary generator network will gradually learn to generate fake samples similar to the distribution of real data. Finally, when the generator network can successfully deceive the quality discriminator, it can generate digital human images with clear and realistic faces.
[0125] The performance of the updated primary generator network can be tested to determine whether it has reached convergence. If the discrimination accuracy is close to 50%, the primary generator network is considered to be able to generate high-quality fake samples, and the discriminator has difficulty distinguishing between real and fake samples, thus reaching convergence. If the discrimination accuracy deviates significantly from 50%, then convergence has not been reached.
[0126] To facilitate understanding, this application also provides a digital human generation method, including a training process and an application process for the generation network. The method uses a NeRF network as the network to be trained and a face video as an example, as shown in Figure 7. The model training process is as follows: A face super-resolution network is used to perform facial super-resolution reconstruction on the sample face video to obtain a target sample face video. The target sample face video, along with matching sample audio and sample face video, are input into the NeRF network for training, thereby increasing the clarity of the trained face. The facial difference between the sample digital human image generated by the NeRF network and the real digital human image corresponding to the sample face video is calculated to obtain the generation loss. A face quality discriminator and a lip quality discriminator are used to perform quality discrimination on the sample digital human image. An adversarial loss is generated based on the quality discrimination result. The generation loss and adversarial loss are then used to supervise the NeRF network to focus on high-definition details of the face, thereby further increasing the clarity of the generated face.
[0127] After the network is trained, during inference, relevant audio can be directly input into the NeRF network to generate a digital human image containing a high-definition face.
[0128] The core of the NeRF network is a multilayer perceptron, as shown in Table 1.
[0129] Table 1
[0130] Position encoding: After the input 3D coordinates and view direction are position encoded, they are usually extended to higher dimensions (e.g., 60 dimensions for coordinates and 24 dimensions for direction). The position encoding is then used to transform and improve the ability to express details.
[0131] Hidden layers: Multiple hidden layers make up the main body of MLP. Each layer typically has 256 neurons. The role of multiple hidden layers is to extract spatial features of the image and understand the spatial distribution and dynamic changes of the face.
[0132] Density branch: The features output by the MLP output volume density through the last fully connected layer.
[0133] Feature layer: Extracts facial and audio features from the hidden layer.
[0134] RGB color branch: The feature information extracted from the hidden layer is further combined with the viewpoint direction and extended to a high-dimensional representation through position encoding. The RGB color is then output through a fully connected network of the RGB color branch.
[0135] The face quality discriminator and the lip quality discriminator have the same model structure. The difference is that the input of the face quality discriminator is the image corresponding to the face, while the input of the lip quality discriminator is the image corresponding to the lips. The structure of the quality discriminator is shown in Table 2.
[0136] Table 2
[0137] Input layer: Accepts input images, usually RGB images, with dimensions of 3×H×W.
[0138] Convolutional layers: These layers extract features through convolution operations. The parameters of each layer, such as kernel size, stride, and padding, can be adjusted according to specific needs. Convolutional layer 1 extracts low-level edge features (such as texture and contrast); Convolutional layer 2 extracts mid-level structural features (such as facial contours and local details); Convolutional layer 3 extracts high-level semantic features (such as the overall shape of the facial region); and Convolutional layer 4 further extracts more abstract features for the final discrimination decision.
[0139] Batch normalization layer: used to accelerate the training process and stabilize the model.
[0140] Activation layer: The LeakyReLU activation function is typically used to provide a nonlinear transformation.
[0141] Output layer: The last layer uses 1×1 convolution to aggregate features and output a single value. After passing through the Sigmoid activation function, the value is mapped to the range [0,1], representing the probability that the input image is real or generated.
[0142] The loss function of the quality discriminator can be a binary cross-entropy loss:
[0143] Among them, y i It is a binary label 0 or 1, where 0 represents false and 1 represents true, p(y i ) indicates the output y i The probability of a label, where N is the total number of samples.
[0144] The network structure of the face super-resolution network is shown in Table 3.
[0145] Table 3
[0146] Input layer: Accepts low-resolution face images as input, typically RGB images with dimensions of 3×H×W.
[0147] Convolutional layers: These layers extract features and reconstruct images through convolution operations. Convolutional layer 1 uses a large kernel to capture global contextual information of the input image, such as low-level features in the image (e.g., edges, textures, and basic shapes). Convolutional layer 2 further extracts local features, and convolutional layer 3 maps the deep features extracted from the first two layers back into the final high-resolution image.
[0148] Activation layer: Use the ReLU activation function to increase non-linearity.
[0149] Output layer: The Tanh activation function is used for non-linear transformation. The Tanh function limits the output value to the range of [-1,1]. The Tanh function can better standardize the pixel values of the image, especially in super-resolution tasks, which helps to keep the color and brightness of the image consistent and obtain the super-resolution face image.
[0150] In this embodiment, the NeRF network includes two types of loss functions: one is a loss function based on the difference between the NeRF network-generated and the real network, and the other is a loss function based on the discrimination result output by the quality discriminator. For details, please refer to the foregoing embodiments, which will not be repeated here.
[0151] The generative network training method provided in this application is based on the idea of generative adversarial networks. It adds a face quality discriminator and a lip quality discriminator during the NeRF network training process. The clarity of the generated face is significantly improved through adversarial training without adding extra time during inference. By performing face super-resolution on the original training data, the clarity of the original training data is directly improved, thereby improving the clarity of the final generated face.
[0152] In the application of NeRF networks, just like the original NeRF network, audio is directly input into the NeRF network to generate high-definition faces. No additional steps are added during inference, thus ensuring the real-time requirements of live digital humans.
[0153] As shown in Figure 8, the terminal displays the live stream creation interface, which includes a creation control and a list of already created live streams. Users can start a live stream from the list. In response to the creation control's trigger, the terminal displays the live stream settings interface, as shown in Figure 9. The live stream settings interface includes live stream items and live stream modes. In response to the "Start Creation" control in the live stream settings interface, a template interface is displayed, as shown in Figure 10. This template interface includes different types of digital avatars of the target broadcaster, pre-generated based on different audio. For a specific target broadcaster avatar, a preview of the target avatar and live stream items can be displayed in the template interface's display area. It should be noted that the timbre of the target broadcaster avatar's voice can also be adjusted in the template interface without affecting the facial details of the avatar. After creation, the live stream list in the live stream creation interface includes the newly created live stream, allowing users to start a live stream from the list and push it to the video account's live stream room.
[0154] The template interface allows users to set a timbre adjustment control. Clicking this control will bring up a timbre selection window. The window offers a variety of preset timbre options, such as calm, lively, and gentle. After the user selects a target timbre, the system applies the corresponding timbre parameters (such as pitch and frequency distribution) to the voice generation model of the target anchor avatar, regenerating a voice with the new timbre. This process does not affect the facial details of the anchor avatar.
[0155] This application describes an apparatus embodiment that can be used to execute the digital human generation method described above. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the digital human generation method described above.
[0156] This application provides a digital human generation device, as shown in FIG11, the device comprising:
[0157] Module 1110 is used to acquire input audio;
[0158] The input module 1120 is used to input the input audio into a pre-trained generator network. The generator network is obtained by adversarial training of the network to be trained based on the sample digital human image output by the network to be trained and the quality discrimination result output by the quality discriminator for the sample digital human image. The sample digital human image is generated by the network to be trained based on the sample facial video and the sample audio of the sample facial video.
[0159] The acquisition module 1110 is also used to acquire the output digital human image generated by the generation network corresponding to the input audio, and the output digital human image is used to process the video content creation task corresponding to the input audio.
[0160] In one embodiment of this application, based on the foregoing scheme, the device further includes a training module for acquiring sample facial videos, sample audio of the sample facial videos, and a network to be trained; performing facial super-resolution reconstruction on the facial regions in the sample facial videos to obtain super-resolution sample facial videos; inputting the super-resolution sample facial videos and the sample audio into the network to be trained to obtain sample digital human images; inputting the sample digital human images into the quality discriminator to obtain the quality discrimination result, and performing adversarial training on the network to be trained based on the sample digital human images and the quality discrimination result to obtain the generator network.
[0161] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to input the super-resolution sample facial video and the sample audio into the network to be trained. The network to be trained is configured to extract features from the sample audio and the super-resolution sample facial video to obtain sample audio features and sample facial video features, respectively, and to perform position encoding on the spatial coordinates and viewpoint direction corresponding to the super-resolution sample facial video to obtain sample geometric encoding features. The sample audio features and the sample facial video features are aligned according to the time scale to obtain target sample audio features. The sample digital human image is generated based on the target sample audio features, the sample geometric encoding features, and the sample facial video features.
[0162] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to extract sample spatial features from the sample geometric coding features through the multilayer perceptron of the network to be trained, and to perform feature mapping on the sample spatial features through the fully connected layer of the network to be trained to obtain the volume density of each spatial point; generate the color value of each spatial point according to the target sample audio features, sample facial video features, and sample viewpoint coding features in the sample geometric coding features; and generate the sample digital human image according to the volume density of each spatial point and the color value.
[0163] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to perform adversarial training on the network to be trained according to the sample digital human image and the quality discrimination result to obtain the generator network, including: obtaining a real human image corresponding to the sample facial video and the sample audio; calculating a generation loss function based on the difference between the sample digital human image and the real human image; generating an adversarial loss function based on the quality discrimination result; and adjusting the network parameters of the network to be trained according to the generation loss function and the adversarial loss function to obtain the generator network.
[0164] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to generate a first generation loss based on the pixel difference between the sample digital human image corresponding to the sample digital human image and the real digital human image corresponding to the real human image; calculate a first optical flow between adjacent digital human image frames of the sample digital human image and a second optical flow between adjacent digital human image frames corresponding to the real human image; generate a second generation loss based on the optical flow difference between the first optical flow and the second optical flow; and calculate the generation loss function based on the first generation loss and the second generation loss.
[0165] In one embodiment of this application, based on the aforementioned scheme, the quality discriminator includes a face quality discriminator and a lip quality discriminator. The training module is further configured to acquire sample face images and sample lip images corresponding to the sample digital human image; input the sample face image into the face quality discriminator to obtain a face quality discrimination result, and input the sample lip image into the lip quality discriminator to obtain a lip quality discrimination result; generate a face quality loss function based on the face quality discrimination result, and generate a lip quality loss function based on the lip quality discrimination result; and generate the adversarial loss function based on the face quality loss function and the lip quality loss function.
[0166] In one embodiment of this application, based on the aforementioned scheme, the training module is further configured to input the sample lip image into the lip quality discriminator, the lip quality discriminator is configured to perform downsampling processing on the sample lip image to generate multiple sub-sample lip images of different resolutions, extract features from the multiple sub-sample lip images to obtain multiple sample lip features, and generate the lip quality discrimination result based on the multiple sample lip features.
[0167] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to generate a discriminant loss function according to the quality discrimination result, fix the network parameters of the network to be trained, adjust the model parameters of the quality discriminator according to the discriminant loss function to obtain a new quality discriminator; perform weighted summation of the generation loss function and the adversarial loss function to obtain a total loss function, fix the network parameters of the quality discriminator, adjust the model parameters of the network to be trained according to the total loss function to obtain a primary generation network; input the new super-resolution sample facial video and the corresponding new sample audio into the primary generation network to obtain a new sample digital human image, and input the new sample digital human image into the new quality discriminator to obtain a new quality discrimination result; based on the new sample digital human image and the new quality discrimination result, repeat the alternating training process of the new quality discriminator and the primary generation network until the primary generation network converges to obtain the generation network.
[0168] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to obtain an initial sample facial video corresponding to the sample audio; and to render the initial sample facial video according to multiple viewpoints and multiple different lighting conditions to obtain a sample facial video.
[0169] In one embodiment of this application, based on the foregoing scheme, the training module is further configured to extract global facial features from the facial region through a first convolutional layer; extract local facial features from the global facial features through a second convolutional layer, wherein the convolutional kernel of the second convolutional layer is smaller than the convolutional kernel of the first convolutional layer; and reconstruct the facial region based on the local facial features to obtain the super-resolution sample facial video.
[0170] In one embodiment of this application, based on the foregoing scheme, the input audio is audio used to introduce the item to be live-streamed; the video content creation task includes a live-streaming task; the acquisition module is further used to acquire the corresponding live-streaming scene based on the item to be live-streamed and the live-streaming content; the device further includes a generation module, used to project the output digital human image onto the live-streaming scene to generate the live-streaming video content corresponding to the live-streaming task.
[0171] It should be noted that the apparatus provided in the above embodiments and the method provided in the above embodiments belong to the same concept, and the specific way in which each module and unit performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0172] The device provided in the above embodiments can be located in a terminal or in a server.
[0173] Embodiments of this application also provide an electronic device, including one or more processors and a storage device, wherein the storage device is used to store one or more computer programs, which, when executed by one or more processors, cause the electronic device to implement the digital human generation method described above.
[0174] Figure 12 shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device according to the embodiments of this application.
[0175] It should be noted that the computer system 1200 of the electronic device shown in Figure 12 is only an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0176] As shown in Figure 12, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in read-only memory (ROM) 1202 or programs loaded from storage portion 1208 into random access memory (RAM) 1203. The RAM 1203 also stores various programs and data required for system operation. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0177] In some embodiments, the following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.
[0178] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by processor (CPU) 1201, it performs various functions defined in the system of this application.
[0179] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory, flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and a computer program.
[0181] The units or modules described in the embodiments of this application can be implemented in software or hardware, and can also be located in a processor. The names of these units or modules do not necessarily limit the specific unit or module itself.
[0182] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the digital human generation method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0183] Another aspect of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium and executes the computer program, causing the electronic device to perform the digital human generation method as described above in the various embodiments.
[0184] In summary, this application provides a method, apparatus, device, computer-readable storage medium, and computer program product for generating digital humans. It acquires input audio via an electronic device and then feeds it into a pre-trained generative network. This generative network is obtained through adversarial training, continuously optimizing the network based on sample digital human images generated by the network under test from sample facial videos and sample audio, and the quality discrimination results given by a quality discriminator. This adversarial training mechanism encourages the generative network to focus more on key facial features when generating digital human images, improving the fitting accuracy of the facial model and directly enhancing the clarity and realism of the generated digital human faces. Simultaneously, it avoids additional high-definition processing steps after generation, reducing computational resource consumption, improving computational resource utilization and processing efficiency, and ensuring the generation of compliant digital human images in a short time.
[0185] Furthermore, before inputting the audio into the generator network, the electronic device acquires sample facial videos, sample audio, and the network to be trained. Facial super-resolution reconstruction is then performed on the facial regions of the sample facial videos to obtain super-resolution sample facial videos. This reconstruction process utilizes multi-layer convolutions to progressively extract and reconstruct features. The first convolutional layer uses a larger kernel to extract global features, while the second convolutional layer uses a smaller kernel to extract local features. This approach accurately captures feature information at different levels of the face, effectively improving the clarity and quality of the training data. High-quality training data helps the network to be trained learn richer and more accurate facial features, resulting in digital human images with richer and clearer facial details generated in subsequent training, thus improving the learning ability and accuracy of the generator network's generation.
[0186] Furthermore, when inputting super-resolution sample facial videos and sample audio into the network to be trained to obtain sample digital human images, the network extracts features from both the sample audio and the super-resolution sample facial videos, obtaining sample audio features and sample facial video features respectively. Simultaneously, positional encoding is performed on the spatial coordinates and viewpoint direction corresponding to the super-resolution sample facial videos to obtain sample geometric encoding features. Positional encoding, by introducing high-frequency information, expands low-dimensional spatial and viewpoint information to high dimensions, enhancing the network's ability to learn complex details. Then, temporal alignment processing is performed on the sample audio features and sample facial video features to obtain target sample audio features, ensuring precise temporal matching between audio and facial movements. Finally, sample digital human images are generated based on these features, achieving high-precision lip-sync, improving the accuracy and realism of digital human image generation, and making the generated digital human images more consistent with human visual and auditory perception.
[0187] Furthermore, when generating a sample digital human image based on the target sample's audio features, sample geometric coding features, and sample facial video features, the multilayer perceptron of the network to be trained extracts sample spatial features from the sample geometric coding features. Multiple hidden layers of the multilayer perceptron can deeply explore the spatial distribution and dynamic changes of the face. Then, a fully connected layer performs feature mapping on the sample spatial features to obtain the volume density of each spatial point. Color values for each spatial point are generated based on the sample viewpoint coding features from the target sample's audio features, sample facial video features, and sample geometric coding features. Finally, the sample digital human image is generated based on the volume density and color values. This generation method based on spatial point features and color values can accurately construct a 3D model of the digital human, enabling the generated digital human image to have rich details and a realistic appearance under different viewpoints and lighting conditions, improving the quality and realism of the generated digital human image and enhancing its visual effect.
[0188] Furthermore, when generating a generative network by adversarial training of the network to be trained based on the sample digital human images and quality discrimination results, the electronic device first acquires real human images corresponding to the sample facial videos and sample audio. Then, it calculates a generation loss function based on the difference between the sample digital human images and the real human images, and simultaneously generates an adversarial loss function based on the quality discrimination results. The network parameters of the network to be trained are adjusted through these two loss functions. This parameter adjustment method, which combines generation loss and adversarial loss, allows the network to be trained to continuously optimize and adjust based on real human images, gradually narrowing the gap with the distribution of real human images, thereby generating more realistic and clearer digital human images and improving the generalization ability of the generative network and the quality of the generated digital human images.
[0189] Furthermore, when calculating the generation loss function, a first generation loss is generated based on the pixel differences between the sample digital human image corresponding to the sample digital human image and the real digital human image corresponding to the real human image, ensuring the good fit between the basic shape of the digital human and the real human image. Simultaneously, a first optical flow between adjacent digital human image frames of the sample digital human image and a second optical flow between adjacent digital human image frames corresponding to the real human image are calculated. A second generation loss is generated based on the optical flow differences, ensuring the smoothness and naturalness of the digital human image during dynamic changes. By combining these two losses to obtain the generation loss function, the differences between the sample digital human image and the real human image in both static and dynamic aspects can be comprehensively measured, more accurately guiding the optimization direction of the network to be trained, avoiding jitter or unnatural movements in the generated digital human image, and improving the stability and realism of the generated digital human image.
[0190] Furthermore, when the quality discriminator includes a facial quality discriminator and a lip quality discriminator, during the process of inputting the sample digital human image into the quality discriminator to obtain the quality discrimination result, the sample facial image and sample lip image corresponding to the sample digital human image are respectively input into the facial quality discriminator and the lip quality discriminator. The facial quality discriminator controls the overall realism of the face, while the lip quality discriminator focuses on lip details, ensuring that the network to be trained learns mouth features more accurately. Based on these two discrimination results, facial quality loss functions and lip quality loss functions are generated respectively, and then an adversarial loss function is generated. This clearly defined discrimination method and loss function generation method can more effectively optimize the network to be trained, improve the quality of the final generated face, especially improve the matching degree of speech and lip movements of the digital human image, and enhance the interactivity and realism of the digital human image.
[0191] Furthermore, when inputting the sample lip image into the lip quality discriminator to obtain the lip quality discrimination result, the lip quality discriminator downsamples the sample lip image to generate multiple sub-sample lip images at different resolutions. Features are extracted from these sub-sample lip images to obtain multiple sample lip features, and finally, the lip quality discrimination result is generated based on these features. By introducing multi-scale features, the lip quality discriminator can judge lip quality at different resolutions and scales, allowing the network to be trained to pay more attention to the local details of the lips, improving the quality of lip detail generation, making the lip shape of the digital human more realistic and natural, and enhancing the realism and vividness of the digital human image.
[0192] Furthermore, when adjusting the network parameters of the training network to obtain the generator network based on the generative loss function and the adversarial loss function, a discriminative loss function is first generated based on the quality discrimination result. The network parameters of the training network are then fixed, and the model parameters of the quality discriminator are adjusted to obtain a new quality discriminator. Then, the generative loss function and the adversarial loss function are weighted and summed to obtain the total loss function. The model parameters of the new quality discriminator are fixed, and the network parameters of the training network are adjusted based on the total loss function to obtain the primary generator network. Next, new super-resolution sample facial videos and corresponding new sample audio are input into the primary generator network to obtain new sample digital human images. These are then input into the new quality discriminator to obtain new quality discrimination results. This alternating training process is repeated until the primary generator network converges to obtain the generator network. This alternating training method allows the generator network and the quality discriminator to mutually promote and continuously optimize each other. The generator network can gradually learn the distribution characteristics of real data, enhancing its robustness and stability, thereby generating clearer and more realistic digital human images and improving the overall performance of the generator network.
[0193] Furthermore, when acquiring sample facial videos, sample audio, and the network to be trained, the initial sample facial video corresponding to the sample audio is first acquired. Then, the initial sample facial video is rendered based on multiple viewpoints and different lighting conditions to obtain the sample facial video. Multi-angle and multi-lighting rendering expands the diversity of the training data, enabling the network to learn the feature changes of the face under different viewpoints and lighting conditions, better capturing facial details such as skin texture, light reflection, and facial shadows. Through training with various lighting and angles, the network can learn to preserve subtle facial details, improving its generalization ability and resulting in a more realistic digital human image in different scenarios, thus enhancing the adaptability of the digital human image.
[0194] Furthermore, when the input audio is used to introduce the item to be live-streamed and the video content creation task includes live-stream content creation, after acquiring the input audio, the electronic device obtains the corresponding live-stream scene based on the item to be live-streamed and the live-stream content. After acquiring the output digital human image generated by the generative network corresponding to the input audio, it is projected onto the live-stream scene to generate live-stream video content. The generative network does not incur additional time consumption during inference, meeting the high real-time requirements of live streaming. This method of rapidly integrating the digital human image and the live-stream scene ensures the real-time nature and smoothness of the live-stream content, improves the system's real-time processing capabilities, and enhances the immersiveness and appeal of the live-stream content by accurately synchronizing the digital human image's actions, expressions, and language with the audio.
[0195] Furthermore, during the training of the generative network, when acquiring sample facial videos, initial sample facial videos corresponding to sample audio can be obtained and rendered based on multiple viewpoints and different lighting conditions to obtain sample facial videos. Multi-angle rendering renders sample face image frames by setting different viewpoints in 3D space. Data augmentation techniques such as rotation transformation, affine transformation, and perspective transformation expand the diversity of training samples in terms of angle. Multi-lighting rendering applies high dynamic range rendering technology and adds local light sources to simulate different lighting conditions, increasing the diversity of training samples in terms of lighting. This multi-angle and multi-lighting amplification method enables the network to learn more comprehensive facial features, especially its performance under different environments and conditions. This improves the network's adaptability to complex scenes, resulting in highly realistic and detailed digital human images under different viewpoints and lighting conditions. It avoids the problem of poor performance of generated images under different conditions due to a single training sample, thus improving the generalization ability of the generative network and the quality of generated digital human images.
[0196] Meanwhile, different reconstruction methods were mentioned when performing facial super-resolution reconstruction. One method involves progressively extracting features through multiple convolutional layers. The first convolutional layer uses a large kernel to capture global contextual information, while the second convolutional layer uses a small kernel to extract finer-grained local features. A residual module can also be added to learn the differences between input and output, better learning high-frequency details. Alternatively, a third convolutional layer can map local facial features to an RGB image, or an attention layer can be used to weight the global facial features output from the second convolutional layer before a third convolutional layer recovers a higher-resolution image. These reconstruction methods can extract and enhance facial features at multiple levels, optimizing image reconstruction results, improving the clarity and detail richness of facial regions, providing a better training data foundation for generating high-quality digital human images, and enhancing the learning ability of the generative network and the facial quality of the final generated digital human image.
[0197] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0198] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0199] The above content is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.
Claims
1. A method for generating a digital human, executed by an electronic device, comprising: Get the input audio; The input audio is fed into a pre-trained generator network, which is obtained by adversarial training of the network under training based on the sample digital human image output by the network under training and the quality discrimination result output by the quality discriminator for the sample digital human image; the sample digital human image is generated by the network under training based on the sample facial video and the sample audio of the sample facial video. and Obtain the output digital human image generated by the generative network corresponding to the input audio, and use the output digital human image for video content creation tasks based on the input audio.
2. The method according to claim 1, further comprising, before inputting the input audio into the pre-trained generative network: Acquire sample facial videos, sample audio of the sample facial videos, and the network to be trained; Facial super-resolution reconstruction is performed on the facial regions in the sample facial video to obtain a super-resolution sample facial video. The super-resolution sample facial video and the sample audio are input into the network to be trained to obtain sample digital human images; The sample digital human image is input into the quality discriminator to obtain the quality discrimination result, and the network to be trained is subjected to adversarial training based on the sample digital human image and the quality discrimination result to obtain the generator network.
3. The method according to claim 2, wherein inputting the super-resolution sample facial video and the sample audio into the network to be trained to obtain sample digital human images includes: The super-resolution sample facial video and the sample audio are input into the network to be trained. The network to be trained is used to extract features from the sample audio and the super-resolution sample facial video to obtain sample audio features and sample facial video features, respectively. The network also performs position encoding on the spatial coordinates and viewpoint direction corresponding to the super-resolution sample facial video to obtain sample geometric encoding features. The network performs time-scale alignment processing on the sample audio features and the sample facial video features to obtain target sample audio features. The network then generates the sample digital human image based on the target sample audio features, the sample geometric encoding features, and the sample facial video features.
4. The method according to claim 3, generating the sample digital human image based on the target sample audio features, the sample geometric coding features, and the sample facial video features, comprising: The sample spatial features are obtained by extracting features from the geometric coding features of the sample through the multilayer perceptron of the network to be trained, and the volume density of each spatial point is obtained by performing feature mapping on the sample spatial features through the fully connected layer of the network to be trained. The color value of each spatial point is generated based on the target sample audio features, the sample facial video features, and the sample viewpoint coding features in the sample geometric coding features; The sample digital human image is generated based on the volume density of each spatial point and the color value.
5. The method according to any one of claims 2 to 4, wherein the step of performing adversarial training on the network to be trained based on the sample digital human image and the quality discrimination result to obtain the generator network comprises: Obtain the real-life image corresponding to the sample facial video and the sample audio; A loss function is generated based on the difference between the sample digital human image and the real human image; An adversarial loss function is generated based on the quality discrimination result, and the network parameters of the network to be trained are adjusted according to the generated loss function and the adversarial loss function to obtain the generated network.
6. The method according to claim 5, wherein calculating and generating a loss function based on the difference between the sample digital human image and the real human image includes: A first generation loss is generated based on the pixel differences between the sample digital human image corresponding to the sample digital human image and the real digital human image corresponding to the real human image. Calculate the first optical flow between adjacent digital human image frames of the sample digital human image, and the second optical flow between adjacent digital human image frames corresponding to the real human image; A second generation loss is generated based on the optical flow difference between the first optical flow and the second optical flow; The generation loss function is calculated based on the first generation loss and the second generation loss.
7. The method according to claim 5 or 6, wherein the quality discriminator includes a face quality discriminator and a lip quality discriminator, and the step of inputting the sample digital human image into the quality discriminator to obtain the quality discrimination result includes: Obtain the sample facial image and sample lip image corresponding to the sample digital human image; The sample facial image is input into the facial quality discriminator to obtain the facial quality discrimination result, and the sample lip image is input into the lip quality discriminator to obtain the lip quality discrimination result; Generate an adversarial loss function based on the quality discrimination result, including: A facial quality loss function is generated based on the facial quality discrimination result, and a lip quality loss function is generated based on the lip quality discrimination result; The adversarial loss function is generated based on the facial quality loss function and the lip quality loss function.
8. The method according to claim 7, wherein inputting the sample lip image into the lip quality discriminator to obtain the lip quality discrimination result includes: The sample lip image is input into the lip quality discriminator, which performs downsampling processing on the sample lip image to generate multiple sub-sample lip images of different resolutions. Feature extraction is performed on the multiple sub-sample lip images to obtain multiple sample lip features. The lip quality discrimination result is generated based on the multiple sample lip features.
9. The method according to any one of claims 5 to 8, wherein the generated network is obtained by adjusting the network parameters of the network to be trained according to the generation loss function and the adversarial loss function, comprising: A discriminant loss function is generated based on the quality discrimination result, and the network parameters of the network to be trained are fixed. The model parameters of the quality discriminator are adjusted according to the discriminant loss function to obtain a new quality discriminator. The total loss function is obtained by weighted summation of the generation loss function and the adversarial loss function, and the model parameters of the new quality discriminator are fixed. The network parameters of the network to be trained are adjusted according to the total loss function to obtain the primary generation network. The new super-resolution sample facial video and the corresponding new sample audio are input into the primary generative network to obtain a new sample digital human image, and the new sample digital human image is input into the new quality discriminator to obtain a new quality discrimination result; Based on the new sample digital human image and the new quality discrimination result, the alternating training process of the new quality discriminator and the primary generator network is repeated until the primary generator network converges to obtain the generator network.
10. The method according to any one of claims 2 to 9, wherein performing facial super-resolution reconstruction on the facial regions in the sample facial video to obtain a super-resolution sample facial video comprises: The global facial features are obtained by performing global feature extraction on the facial region through the first convolutional layer. Local facial features are obtained by extracting local features from the global facial features through a second convolutional layer, wherein the convolutional kernel of the second convolutional layer is smaller than the convolutional kernel of the first convolutional layer; The super-resolution sample facial video is obtained by reconstructing the facial region based on the local facial features.
11. The method according to any one of claims 2 to 10, wherein acquiring the sample facial video, the sample audio of the sample facial video, and the network to be trained comprises: Obtain the initial sample facial video corresponding to the sample audio; The initial sample facial video is rendered using multiple viewpoints and different lighting conditions to obtain a sample facial video.
12. The method according to any one of claims 1 to 11, wherein the input audio is audio used to introduce the item to be live-streamed; The video content creation task includes live streaming content creation task; After acquiring the input audio, the method further includes: The corresponding live streaming scene is obtained based on the items to be live streamed and the live streaming content. After obtaining the output digital human image generated by the generative network corresponding to the input audio, the method further includes: The output digital human image is projected onto the live streaming scene to generate the video live streaming content corresponding to the live streaming content creation task.
13. A digital human generation device, comprising: The acquisition module is used to acquire input audio. The input module is used to input the input audio into a pre-trained generator network. The generator network is obtained by adversarial training of the network to be trained based on the sample digital human image output by the network to be trained and the quality discrimination result output by the quality discriminator for the sample digital human image. The sample digital human image is generated by the network to be trained based on the sample facial video and the sample audio of the sample facial video. and The acquisition module is also used to acquire the output digital human image generated by the generation network corresponding to the input audio, and the output digital human image is used to process the video content creation task corresponding to the input audio.
14. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method of any one of claims 1 to 12.
15. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform the method of any one of claims 1 to 12.
16. A computer program product comprising a computer program stored in a computer-readable storage medium, wherein a processor of an electronic device reads from and executes the computer program to cause the electronic device to perform the method of any one of claims 1 to 12.