Digital human video generation method and system

This digital human video generation method, which employs multi-dimensional loss constraints and lightweight processing, addresses the issues of insufficient facial expression capabilities and high training costs in existing technologies. It achieves high-precision, natural digital human video generation and supports applications in multiple scenarios.

CN122340327APending Publication Date: 2026-07-03JINAN INSTITUTE OF SUPERCOMPUTING TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610265445.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-05
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing digital human generation technologies suffer from limited facial geometry and texture expression capabilities, stiff facial expressions, high training costs, and difficulty in large-scale deployment, making it difficult to generate high-fidelity, natural digital human videos.

Method used

A digital human video generation method with multi-dimensional loss constraints is proposed. Through volumetric rendering network and lightweight processing, combined with the first to fourth loss modules, appearance details, geometric accuracy, temporal continuity and model stability are quantified to generate personalized weights. It supports independent driving control of audio, pose and style. It utilizes LoRA low-rank adaptive fine-tuning technology and neural rendering processing to quickly generate high-precision digital human videos.

Benefits of technology

It achieves high-precision and natural digital human video generation, has good robustness, can be trained quickly, and the generated digital human videos have realistic facial details, accurate matching of lip movements and speech, natural and smooth expressions, and support multiple application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340327A_ABST
    Figure CN122340327A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for generating digital human videos, belonging to the field of digital human technology. Addressing the problems of poor facial detail reproduction, stiff expression timing, and low modeling efficiency in existing digital human videos, this invention employs a volumetric rendering network, which undergoes lightweight processing and is optimized using a multi-dimensional loss module with parallel constraints to generate digital human videos. The digital human videos generated by this invention feature realistic facial details and natural, smooth expressions. Furthermore, personalized modeling can be completed in minutes using only a single image or a short video clip, making it suitable for scenarios such as virtual anchors, online education, and digital customer service, achieving highly realistic and efficient digital human video generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital human technology, and in particular to a method and system for generating digital human videos. Background Technology

[0002] The statements in this section are merely to provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of artificial intelligence, digital humans have been widely applied in various fields such as intelligent customer service, virtual anchors, digital education, film and television production, and cultural dissemination. Among them, audio-driven digital human facial synthesis technology is a crucial foundational capability for digital human production. It enables virtual characters to display lip movements and facial expressions synchronized with speech by inputting sound signals, achieving natural expression and emotion transmission. Existing technologies mainly include various approaches such as two-dimensional image-driven methods, three-dimensional parametric model methods, and implicit field-based generation methods.

[0004] However, in practical applications, these methods all have obvious shortcomings, mainly in the following aspects: 1) Limited facial geometry and texture representation capabilities: 2D methods are prone to distortion due to inconsistent lighting and excessive angle changes. Traditional 3D parametric models are insufficient in handling facial expression details, lip micro-movements, skin texture, and hair edge processing, making it difficult to achieve the detail accuracy and realism required for high-fidelity scenes. 2) Most audio-driven methods only predict mouth opening and closing parameters, resulting in overly simplistic structures that lead to stiff and unnatural facial expressions during long speech inputs. 3) High training costs: requiring extensive parameter optimization that takes hours to days, making large-scale deployment difficult and unable to meet the needs of ordinary users for quickly generating digital humans. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a digital human video generation method and system that makes the generated digital human videos more realistic and natural, and allows for rapid fine-tuning, combining high-precision fidelity with efficient video generation.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the present invention provides a method for generating digital human videos, comprising: Acquire the original images / videos of the people and the original audio; The images / videos are preprocessed to generate an original parameter sequence; The original parameter sequence is passed through a volumetric rendering network to generate an initial rendering result; and the initial rendering result is then lightweighted to generate a reconstructed rendering result. The reconstructed rendering results are processed in parallel through a first loss module, a second loss module, a third loss module, and a fourth loss module. The first loss module generates a first loss value to quantify appearance details; The second loss module generates a second loss value to quantify geometric precision. The third loss module generates a third loss value to quantify temporal continuity. The fourth loss module generates a fourth loss value to quantify model stability. The reconstructed rendering result is updated based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; The original audio is processed to generate a target parameter sequence; The personalized weights and the target parameter sequence are loaded, and the digital human video is generated.

[0007] In a second aspect, the present invention provides a digital human video generation system, comprising: Data acquisition module: used to acquire raw human images / videos and raw audio; First processing module: used to preprocess the image / video to generate an original parameter sequence; Data reconstruction module: used to generate an initial rendering result from the original parameter sequence through a volumetric rendering network; and to perform lightweight processing on the initial rendering result to generate a reconstructed rendering result; First Loss Module: Used to generate a first loss value from the reconstructed rendering results to quantify appearance details; Second loss module: used to generate a second loss value from the reconstructed rendering results to quantify geometric accuracy; The third loss module is used to generate a third loss value from the reconstructed rendering results to quantify temporal continuity. The fourth loss module is used to generate a fourth loss value from the reconstructed rendering results to quantify the model's stability. Training module: used to update the reconstructed rendering result based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; The second processing module is used to process the original audio to generate a target parameter sequence. Video generation module: used to load the personalized weights and the target parameter sequence, and process them to generate digital human videos.

[0008] Compared with the prior art, the beneficial effects of the present invention are: 1) This invention accurately restores facial geometry, texture, and lip micro-movements through multi-dimensional loss constraints, solving the problems of lighting distortion and insufficient detail in traditional methods; it ensures the temporal continuity of lip movements, blinking, and facial muscle movements, avoiding stiffness and roboticity, making them more realistic and natural; through lightweight processing, it improves the efficiency of personalized modeling, requiring only a single image or short video to complete training in minutes, and significantly reducing the number of trainable parameters; through neural rendering processing, it achieves high-precision facial rendering, high-definition output, temporal smoothing, and audio-visual alignment, ultimately generating realistic and natural high-resolution digital human videos with synchronized lip movements and expressions.

[0009] 2) This invention supports independent drive control of audio, posture, and style, and has good robustness to changes in lighting and angle. It can generate high-resolution digital human videos with precise audio-visual synchronization from end to end, meeting the needs of multiple scenarios such as virtual anchors, digital customer service, online education, and film and television production. It balances production efficiency, generation quality, and practical application flexibility, realizing the implementation of realistic, credible, and scalable digital human production. Attached Figure Description

[0010] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0011] Figure 1 This is a flowchart of the method according to Embodiment 1 of the present invention; Figure 2 This is a diagram of the overall architecture of the present invention; Figure 3 This is a flowchart of the data preprocessing module of the present invention; Figure 4 This is a structural diagram of the fine-tuning module of the present invention; Figure 5 This is a flowchart of the reasoning stage of the present invention; Figure 6 This is a structural diagram of the ICS-A2M action generation module of the present invention; Figure 7 This is a structural diagram of the dynamic neural rendering module of the present invention. Detailed Implementation

[0012] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0013] Example 1 It should be understood that the database sources for this application were obtained legally and compliantly, and with the consent of the applicant.

[0014] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0015] 3DMM (3D Morphable Model): A parametric 3D face modeling method that uses linear combinations of identity and expression basis vectors to represent the shape and texture of a face.

[0016] Basel Face Model (BFM): A high-precision parametric 3D face model built from 3D scan data, including identity and expression parameters.

[0017] Tri-plane Representation: A highly efficient implicit representation method for three-dimensional scenes, consisting of three mutually orthogonal two-dimensional feature planes, used to encode the geometric and appearance information of a three-dimensional scene.

[0018] Neural Radiance Fields (NeRF): A deep learning-based volumetric rendering technique that maps 3D coordinates and viewpoint orientation to color and density using a multilayer perceptron.

[0019] SECC (Semantic Expression Conditioned Color): A semantic expression conditional color, an intermediate representation based on 3D face parameter rendering, used to bridge parametric models and neural renderers.

[0020] LoRA (Low-Rank Adaptation): A fine-tuning technique for large models, which adds a trainable side matrix to the original weight matrix through low-rank decomposition, significantly reducing the number of trainable parameters.

[0021] HuBERT (Hidden-Unit BERT): Hidden Unit BERT is a self-supervised speech representation learning model that extracts semantic features from audio through mask prediction tasks.

[0022] MFCC (Mel-Frequency Cepstral Coefficients): A speech feature extraction method based on the characteristics of human hearing.

[0023] LPIPS (Learned Perceptual Image Patch Similarity): A measure of image similarity based on deep feature space.

[0024] AdamW: An improved Adam optimizer that enhances model generalization by decoupling weight decay.

[0025] EG3D (Efficient Geometry-aware 3D Generative Adversarial Network): A framework for 3D generative models based on three-plane representation.

[0026] Transformer: A deep learning model architecture based on self-attention mechanism, widely used in sequence-to-sequence tasks.

[0027] In-Context Learning: A learning method that allows a model to learn new tasks by referring to examples without updating model parameters.

[0028] Classifier-Free Guidance: A control technique for conditional generative models that improves generation quality by mixing conditional and unconditional predictions.

[0029] MediaPipe: A multimedia machine learning model framework that provides real-time face detection, key point localization, and semantic segmentation.

[0030] Volume Rendering: A rendering technique that uses ray tracing and integration to calculate the appearance of transparent or translucent media.

[0031] Optical Flow: Optical flow represents the velocity field of pixels in an image as they move over time.

[0032] Dense Motion Estimation: A technique for predicting motion vectors between images pixel by pixel.

[0033] Semantic Segmentation is a task that classifies each pixel in an image and outputs pixel-level category labels.

[0034] F0 (Fundamental Frequency): The fundamental frequency, the lowest frequency component in sound, corresponding to the frequency of vocal cord vibration.

[0035] Denoising: The process of removing noise components from a signal through iterative optimization.

[0036] Temporal Smoothing: A technique for filtering time series data to reduce jitter.

[0037] Moving Average Filter: A time-series smoothing method that calculates local average values ​​using a sliding window.

[0038] Entropy Regularization is a technique that constrains the distribution of a model by maximizing or minimizing information entropy.

[0039] Compared to the shortcomings of existing digital human generation technologies, such as creating virtual anchors for self-media and small and medium-sized businesses, existing technologies often require multi-view high-definition acquisition, massive data training and professional computing power support, with modeling cycles lasting several days and resulting in stiff expressions and audio-visual asynchrony. In contrast, this invention only requires users to provide a single portrait or a 20-second short video, which can complete personalized digital human modeling within minutes. The generated virtual anchor has realistic facial details, accurate lip-sync with voice, and natural, smooth expressions.

[0040] For example, in the scenario of designing digital teachers for online education institutions, this invention can quickly generate stable and usable digital human images in batches on ordinary computing devices. It has good robustness to changes in lighting and posture, solves the problems of low modeling efficiency and difficulty in large-scale implementation of existing technologies, and ensures the naturalness of the digital human's face and the smoothness of interaction when teaching, truly realizing highly realistic and efficient digital human applications.

[0041] In a typical embodiment of the present invention, such as Figure 1 As shown, this embodiment discloses a method for generating digital human videos, which specifically includes the following steps: S1: Acquire the original human image / video and original audio; S2: The image / video is preprocessed to generate an original parameter sequence; S3: The original parameter sequence is processed by a volumetric rendering network to generate an initial rendering result; and the initial rendering result is then lightweighted to generate a reconstructed rendering result. The reconstructed rendering results are processed in parallel through a first loss module, a second loss module, a third loss module, and a fourth loss module. The first loss module generates a first loss value to quantify appearance details; The second loss module generates a second loss value to quantify geometric precision. The third loss module generates a third loss value to quantify temporal continuity. The fourth loss module generates a fourth loss value to quantify model stability. The reconstructed rendering result is updated based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; S4: The original audio is processed to generate a target parameter sequence; S5: Load the personalized weights and the target parameter sequence, and process to generate a digital human video.

[0042] The above-mentioned digital human video generation method will be described in detail below with reference to specific implementation methods.

[0043] like Figure 2 As shown, the process first inputs a target image / video and driving audio, and can also input a style reference video to drive a 3D model, extracting identity / appearance features, facial expression / speaking habit style features, and semantic / phoneme audio features respectively; then, through a Few-shot rapid fine-tuning mechanism, personalized adaptation can be completed in just minutes; combined with the ICS-A2M model of Flow Matching, motion generation is realized, outputting 3DMM facial expression / posture and other motion parameters and specific person weights; finally, based on the 3D rendering technology of Tri-plane and NeRF neural radiation field, a personalized speaking video is generated.

[0044] S1: Obtain the original human image / video and the original audio.

[0045] The original audio can be obtained in the following ways, but is not limited to: for example, users can submit a single clear frontal face image in JPG / PNG format or a short video in MP4 format of about 20 seconds through the system upload interface, or record in real time through shooting devices such as mobile phones and cameras; the face and neck should be unobstructed, and no hair should cover the face. The person in the recorded picture should be centered, the lighting should be uniform and there should be no shadows, and the movements should not cover the head and neck.

[0046] The original audio can be obtained in the following ways, but is not limited to: 1) extracting the audio track from user-uploaded videos containing audio using audio-video separation technology; 2) having users directly upload independent WAV / MP3 format audio files of the target person; 3) recording the target person's voice synchronously with the video or separately in real time; the audio must be noise-free and free of microphone pops, and will be uniformly resampled to 16kHz in the future.

[0047] S2: The image / video is preprocessed to generate an original parameter sequence.

[0048] Specifically, preprocessing is performed on single images / short videos, including frame sequence extraction, face detection, semantic segmentation, and illumination and color calibration to standardize the raw data. Based on a BFM 3D face model, initial 3D face reconstruction is completed, and fitting parameters are optimized. Camera parameters are calculated to establish a basic parameter framework for the 3D morphology of the person. An original parameter sequence containing 80-dimensional identity parameters, 64-dimensional expression parameters, and 25-dimensional camera parameters is generated. The data source for this sequence directly depends on the aforementioned obtained and approved original person images / videos, ensuring that the parameters match the geometric features and expression habits of the target person. Specifically, the camera parameter vector is obtained by calculating the camera intrinsic and extrinsic parameter matrices through least-squares optimization fitting of frame-by-frame parameters followed by temporal smoothing filtering.

[0049] like Figure 3 As shown in the figure, this diagram illustrates the preprocessing flow from raw data to 3D parameter extraction. In the original human video, there are two parallel paths: an audio stream and a video stream. The audio stream uses a pre-trained audio encoder to transform the raw target audio into an audio feature vector containing semantics and phonemes. The video stream first performs facial landmark detection on the frame sequence of the original target video, and then completes FLAME parameter fitting through inverse rendering. Finally, it decomposes the shape parameter (Shapeβ) that determines the person's weight / face shape, the pose parameter (Poseθ) that controls head rotation, and the expression parameter (Expψ) that is responsible for mouth shape / blinking.

[0050] The input short video of about 20 seconds is decoded at a preset frame rate (preferably 25 frames / second) to extract a complete image frame sequence. Resolution standardization is performed on each frame to uniformly adjust it to a fixed size (preferably 512×512 pixels). A pre-trained face detection model is used to detect face regions in each frame, accurately extracting the two-dimensional coordinate information of 68 or more facial key points. Based on the two-dimensional coordinate information of the facial key points, face cropping, rotation correction, and center alignment are performed to ensure uniform face posture. A semantic segmentation network is used to perform pixel-level semantic classification on each frame, generating a segmentation mask containing target regions such as head, torso, and background. The head region image, torso restoration image, and background image are separated and extracted based on the segmentation mask. Illumination normalization is performed on the extracted image sequence to effectively eliminate color deviations caused by illumination changes. Color space standardization is performed to normalize pixel values ​​to the [-1,1] range, providing standardized input for subsequent modeling.

[0051] A parametric 3D face model based on the Basel Face Model (BFM) is employed. This model includes 80-dimensional identity parameters, 64-dimensional expression parameters, and head pose parameters composed of 3-dimensional Euler angles and 3-dimensional translation vectors. For each preprocessed image frame, an optimization algorithm is used to solve for the optimal parameter combination, ensuring accurate matching between the reconstructed 3D face projection and the detected 2D keypoints. The fitting process employs a least-squares optimization method to minimize the 2D keypoint reprojection error. Temporal smoothing filtering is performed on the fitted parameter sequence, preferably using a moving average filter with a kernel size of 7 to effectively eliminate parameter jitter and ensure temporal continuity. Then, based on the pose parameters, a 4×4 camera extrinsic matrix and a 3×3 camera intrinsic matrix conforming to the EG3D rendering specification are calculated. These two types of matrices are flattened into a 25-dimensional camera parameter vector containing 16-dimensional extrinsic parameters and 9-dimensional intrinsic parameters, providing standardized input for subsequent rendering.

[0052] S3: After generating the initial rendering result from the original parameter sequence and obtaining the reconstructed rendering result through lightweight processing, the reconstructed rendering result is input into the first to fourth loss modules for parallel processing. Each module generates loss values ​​for quantizing appearance details, geometric accuracy, temporal continuity, and model stability, respectively. The reconstructed rendering result is updated based on the sum of these loss values ​​to generate personalized weights.

[0053] S3.1: Generate an initial rendering result from the original parameter sequence. The initial rendering result is generated by a volumetric rendering network consisting of a combination mechanism of SECC conditional encoding, three-plane feature fusion, and NeRF neural volumetric rendering. First, the original parameter sequence is used to generate three types of SECC semantic expression condition maps through SECC conditional encoding, transforming abstract three-dimensional parameters into visual features recognizable by neural rendering. Then, the facial expression feature plane extracted by SECC is added to the three-plane features through three-plane feature fusion to generate a fused feature that combines personalized attributes and dynamic facial expression features. Finally, the fused feature is subjected to ray tracing sampling and volume integration using NeRF neural volumetric rendering technology to generate a low-resolution result. The initial result is then processed by a super-resolution network to generate the initial rendering result.

[0054] Specifically, the Tri-plane is a three-dimensional implicit representation structure composed of three mutually orthogonal two-dimensional feature planes along XY, XZ, and YZ. Each feature plane is set to a size of 256×256, with a hidden feature dimension of C×D, where C is the number of hidden layer channels and D is the number of depth layers. The complete dimensions of the tri-plane tensor are [1,3,C×D,256,256]. Using the first frame of the training video as input, an initial standard tri-plane representation is generated through a pre-trained image encoding network. This standard tri-plane encodes the three-dimensional geometric shape and texture features of a person in a neutral expression state, and the initial tri-plane is set as learnable parameters, allowing for dynamic updates in subsequent processes. A facial expression control mechanism based on SECC (Semantic Expression Conditioned Color) is constructed. For each training sample, three types of SECC conditional maps are generated: a standard SECC map (cano_secc), which is a neutral facial expression image rendered by a 3D model after setting the expression parameters to zero; a source reference SECC map (src_secc), which is a facial expression image rendered using the expression parameters of the source reference frame; and a target-driven SECC map (drv_secc), which is a target facial expression image rendered using the expression parameters of the current driving frame. Combined with the SECC-based facial expression control mechanism, a volumetric rendering network based on three-plane representation is constructed as the model to be trained. This model includes four parts: a SECC conditional encoder, a three-plane feature fusion module, a volumetric rendering module, and a super-resolution upsampling module. The SECC conditional encoder inputs three types of SECC condition maps into a convolutional neural network to extract expression-related SECC feature planes (secc_plane); the three-plane feature fusion module adds and fuses the learnable standard three planes with the SECC feature planes to output the three-plane features required for final rendering; the volume rendering module uses Neural Radiation Field (NeRF) technology to perform ray tracing sampling and volume integration on the fused three-plane features to generate a low-resolution rendering image of 128×128 pixels and a density weight map; the super-resolution upsampling module upsamples the low-resolution rendering result to a high-resolution output of 512×512 pixels through a convolutional neural network.

[0055] To address the need for digital human torso synthesis, a keypoint-driven motion deformation network is employed to achieve torso synthesis: 68 facial keypoint coordinates (kp_source) from the source reference frame and keypoint coordinates (kp_driving) from the current driving frame are extracted; a dense motion estimation network predicts pixel-level motion fields and occlusion masks; optical flow deformation is performed on the reference torso image to generate a torso image matching the current expression; and the generated head image and the deformed torso image are weighted and fused based on the segmentation mask to ensure head-body coordination.

[0056] S3.2: To accelerate training efficiency, the initial rendering result is lightweighted by combining the LoRA low-rank adaptive fine-tuning mechanism of the Tri-plane lightweight design to generate the reconstructed rendering result. Specifically, the LoRA low-rank decomposition module is embedded in the key layers of the SECC encoder and the super-resolution network to perform low-rank decomposition on the weight matrix. Only the LoRA parameters are set to a trainable state and the other parameters of the pre-trained model are frozen. At the same time, a lightweight Tri-plane tensor design of [1,3,C×D,256,256] and a batch size configuration adapted to ordinary hardware are adopted. The AdamW optimizer is used to optimize the Tri-plane learnable parameters and the LoRA low-rank matrix with differentiated learning rates. While reducing the computation and storage costs, the personalized geometric and texture features of the target character are preserved, and the reconstructed rendering result is finally generated.

[0057] like Figure 4 As shown in the attached diagram, this visually illustrates the parameter update logic and training process for personalized fine-tuning of the target person. Using a short video of the target person as input, audio / expression is extracted as input conditions, while ground truth (Ground Truth) is extracted for loss calculation. In the general base model, the gray dashed lines representing the general audio / expression encoder and general Tri-plane are frozen parameters (not updated), while only the Tri-plane and Neural Renderer (MLP) are training parameters (updated) as shown by the red solid lines. After the model outputs the predicted image, the training parameters are updated via gradient backpropagation by calculating L1 loss and perceptual loss (L_perceptual).

[0058] The LoRA low-rank adaptive fine-tuning mechanism is as follows: A LoRA low-rank decomposition module is embedded in the key layers of the SECC encoder and the super-resolution network. Here, the key layers refer to the core convolutional layer in the SECC encoder responsible for facial feature extraction and fusion, and the core network layer in the super-resolution upsampling module that implements image resolution enhancement and detail restoration. Then, a low-rank decomposition is performed on the weight matrix W, ΔW=BA, where B and A are both reduced-rank matrices, and the rank r is preferably set to 2. During training, only the LoRA parameters are set to a trainable state, while the remaining parameters of the pre-trained model are frozen, reducing the number of trainable parameters to less than 5% of the original model, thus reducing computational overhead. The LoRA mode supports flexible configuration, including full mode, only SECC2Plane mode (secc2plane_sr), or no LoRA mode (none), to meet the fine-tuning needs of different scenarios.

[0059] The parameter optimization strategy employs the AdamW optimizer, with differentiated learning rates set for different parameter types: a three-plane learning rate of 0.005 for video input, a three-plane learning rate of 0.001 for single image input, and a uniform LoRA parameter learning rate of 0.001. The optimizer hyperparameters are configured with β1=0.9, β2=0.98, and a weight decay coefficient of 0.01. Simultaneously, the Tri-plane feature plane size is set to 256×256, using a lightweight tensor design of [1,3,C×D,256,256]. Batch sizes support either 1 or 2, corresponding to 8GB and 15GB of video memory requirements, respectively. The number of training iterations is dynamically adjusted according to the type of input data. Video input undergoes 2000-10000 iterations, taking about 10 minutes, while a single image input undergoes 3-10 iterations. Verification is performed every 2000 iterations, and model checkpoints are saved to ensure the traceability of the training process. While significantly reducing computation and storage costs, it accurately preserves the personalized geometric and texture features of the target person, ultimately generating a reconstruction rendering result that is both lightweight and personalized.

[0060] S3.3: The reconstructed rendering results are processed in parallel by the first loss module, the second loss module, the third loss module and the fourth loss module.

[0061] 1) The first loss module is an image reconstruction loss module. It generates a first loss value by calculating the weighted sum of the L1 pixel loss between the generated image and the real image, the LPIPS perceptual loss based on a pre-trained network (such as VGG / AlexNet), the low-resolution head L1 loss, and the local L1 and LPIPS losses of the lip region. The first loss value quantifies overall pixel consistency by calculating a first sub-loss, quantifies deep feature similarity by calculating a second sub-loss, quantifies low-resolution reconstruction accuracy by calculating a third sub-loss, and quantifies local detail restoration by calculating a fourth sub-loss. The first, second, third, and fourth sub-losses are pixel-level loss, perceptual loss, low-resolution head loss, and local region enhancement loss, respectively.

[0062] The fourth sub-loss quantifies the pixel-level reconstruction accuracy of the lip shape by calculating the local first sub-loss and quantifies the deep feature similarity of the lip shape by calculating the local second sub-loss; the local first sub-loss and the local second sub-loss are the L1 loss and the LPIPS loss of the lip region, respectively.

[0063] Image reconstruction loss: directly constrains the pixel-level and feature-level similarity between the generated image and the real image. In particular, the specific loss for the lip region is directly related to the core requirement of subsequent audio-lip matching and is the foundation for ensuring the realism of digital human faces.

[0064] L1 pixel loss: Calculates the average absolute error between the generated image and the real image, with a weighting coefficient set to 1.0;

[0065] in, Pixel loss is used to optimize the model, making the generated image closer to the real image; N is the total number of pixels in the image; for an image with width W, height H, and number of channels C (e.g., C=3 for an RGB image), N = W×H×C, which represents the range of the summation. p is the pixel index, an integer from 1 to N, used to traverse each pixel in the image, representing the position of the single pixel currently being processed. This represents the pixel value at the p-th pixel position in the generated image; Let be the pixel value at the P-th pixel position in the original image. The ground truth is the ideal reference image, typically the true label from the dataset. These are the weighting coefficients for pixel loss.

[0066] Perceived loss Deep features are extracted using a pre-trained VGG or AlexNet network, and the distance in the feature space is calculated with a weight coefficient set to 0.5. In addition to VGG or AlexNet, other networks can also be used for pre-training (mainly using feature extraction networks to calculate the difference between the generated image and the target image; any network with similar functions is acceptable). This application does not impose specific limitations on the implementation.

[0067]

[0068] in ( ) ( () represents the feature extractor of a pre-trained network (such as VGG / AlexNet). ( , () is a feature space distance metric, which is the perceptual distance between two images calculated by the model; The weighting coefficients for perceived loss; Low-resolution head loss : Calculate the L1 loss on the low-resolution image output by neural rendering, with the weighting coefficient set to 0.2;

[0069] in and For low-resolution images, Its pixel count; These are the weighting coefficients for low-resolution head loss.

[0070] Lip region enhancement loss: The rectangular bounding box of the lip region is located based on facial key points. The L1 loss and LPIPS loss are calculated separately in this local region, with weight coefficients set to 1.0 and 1.0 respectively, to enhance the accuracy of lip shape. Local L1 loss :

[0071] in, These are the weighting coefficients for the local L1 loss.

[0072] Local LPIPS loss :

[0073] in The bounding box of the lip region defined by facial key points. This represents the number of pixels within the region. Represents an image patch corresponding to the region; These are the weighting coefficients for the local LPIPS loss.

[0074] 2) The second loss module is a geometric consistency loss module. It generates a second loss value to quantify geometric accuracy by calculating the weighted sum of the L2 density supervision loss of the neural rendering output density weight map and the supervised target, and the binary entropy regularization loss of the density weights. The second loss value quantifies the foreground-background separation accuracy by calculating at least the fifth sub-loss, and the density distribution rationality by calculating the sixth sub-loss; the fifth sub-loss and the sixth sub-loss are the density supervision loss and the density entropy regularization loss, respectively.

[0075] Geometric consistency loss: constrains the geometric accuracy of 3D reconstruction, ensuring density separation between facial and non-facial areas and that the 3D structure conforms to the logic of the real human body. If the geometric structure is distorted, subsequent rendering and audio driving will lose their 3D foundation, affecting the overall realism.

[0076] Density supervised loss: The density weight map of the neural rendering output is supervised, and the density weight of the facial region is constrained to approach 1, while that of the non-facial region is constrained to approach 0. L2 loss is used, and the weight coefficient is set to 0.01.

[0077] in, For density monitoring loss, The density weight map output by neural rendering. The corresponding supervision targets are (facial areas approach 1, non-facial areas approach 0). This represents the number of elements in the density map; represents the weighting coefficients for density-supervised loss.

[0078] Density entropy regularization loss: Calculate the binary entropy of the density weights to promote a clear foreground-background separation in the weight distribution, with the weight coefficient set to 0.001;

[0079]

[0080] in, For density entropy regularization loss, These are the weighting coefficients for density entropy regularization loss. The logarithmic term, in the form of binary cross-entropy, represents the density weight map of the neural rendering output. The logarithmic density weight map is obtained by performing a natural logarithmic operation on the density value corresponding to each pixel position m, where m represents the pixel coordinates on the density weight map.

[0081] 3) The third loss module is the temporal smoothness loss module. It generates a third loss value to quantify temporal continuity by calculating the weighted sum of the L1 blink smoothness loss of the linear interpolation result of the interpolated expression coding features and endpoint features, and the L1 stability loss of the generated results before and after the SECC representation perturbation.

[0082] The third loss value is at least calculated by the blink smoothing loss to quantify the naturalness of eye movements, and by the SECC stability loss to quantify the temporal coherence of facial expressions; Temporal smoothness constraints: Solving the problems of inter-frame jitter and facial expression jumps, and ensuring the temporal continuity of video, is the key to the natural and smooth audio-driven digital human.

[0083] Blink smoothness loss: SECC sequences with different blink rates (blink ratios of 0-0.5, 0.5-1.0 and their intermediate values) are constructed by interpolation. The encoding features of the interpolated expressions are constrained to be linear interpolation of the endpoint features. L1 loss is used and the weight coefficient is set to 0.003.

[0084] in, For loss of blink smoothness, Let the blink rate α ∈ [0, 0.5, 1.0, median] be the value of the blink rate. α SECC sequence coding features obtained by under-interpolation in the range [0, 0.5, 1.0, median value]. and As endpoint features, For feature dimensions; The weighting coefficients for the loss of blink smoothness.

[0085] SECC stability loss Add a small random perturbation (standard deviation configurable) to the SECC representation, constrain the consistency of the generated results before and after the perturbation, and adjust the SECC stability loss weight coefficient. Set to 0.01;

[0086] in To add to SECC Small random perturbations, The standard deviation is configurable.

[0087] 4) The fourth loss module is the regularization loss module. It generates a fourth loss value to quantify the model stability by calculating the weighted sum of the L1 distance loss between the current Tri-plane and the initial Tri-plane, the motion field occlusion L1 regularization loss of the torso synthesis network, and the occlusion weight entropy constraint loss.

[0088] The fourth loss value is quantified at least by calculating the three-plane regularization loss to quantify the constraint degree of the three-dimensional representation parameters, and by calculating the torso module regularization loss to quantify the generalization ability of torso synthesis.

[0089] Regularization constraints: Their function is to prevent the model from overfitting and to ensure generalization ability.

[0090] Triplane regularization loss Calculate the L1 distance between the current three planes and the initial three planes to prevent overfitting. The weight coefficients are adaptively adjusted according to the number of training samples: set to 1.0 for a single image, 1.0 for up to 5 frames, 0.1 for up to 250 frames, and 0 for more than 250 frames.

[0091] in For the current c-th feature plane, Let W, H, and C be the initial values ​​for the plane, respectively, representing the width, height, and number of channels. The weighting coefficients of the three-plane regularization loss are... Based on the number of training frames Adaptive adjustment: where, The number of frames.

[0092]

[0093] The torso module regularization loss includes motion field occlusion L1 regularization (weight 0.001) and occlusion weight entropy constraint (weight configurable). L1 regularization of sports field shading :

[0094] Occlusion weight entropy constraint :

[0095] in and These are the occlusion weights for sports fields and general occlusion weights, respectively. S For the corresponding number of elements, For configurable weighting coefficients, Weighting coefficients for L1 regularization of occlusion in sports fields; The weight coefficients are used to occlude the weight entropy constraint.

[0096] 5) The total loss is the sum of all weighted losses:

[0097] in These are the weighting coefficients for each loss term. This represents the corresponding loss value.

[0098] The weighted sum of the first, second, third, and fourth loss values ​​calculated using preset weight coefficients is used to update the trainable parameters (including the LoRA low-rank matrix and Tri-plane learnable parameters) corresponding to the reconstructed rendering result through backpropagation of the AdamW optimizer. The optimization is iterated until the total loss value converges, generating personalized weights adapted to the target character.

[0099] After generating personalized weights, the original audio is processed through stream matching to generate a target parameter sequence. Finally, the personalized weights and target parameter sequence are loaded and processed by neural rendering to generate a digital human video.

[0100] S4: The original audio is processed by stream matching to generate a target parameter sequence. The stream matching is implemented by an ICS-A2M network model that incorporates an in-context learning mechanism. The ICS-A2M network model is a sequence-to-sequence mapping model built on Transformer or recurrent neural network. Transformer or recurrent neural network can also be replaced by other networks. This application embodiment does not make specific limitations. Specifically, the original audio is first resampled to a standard sampling rate of 16kHz, and HuBERT / MFCC semantic features and F0 fundamental frequency prosodic features are extracted and temporal alignment is performed. Then, the audio features are combined with the 80-dimensional identity parameters of the target person and input into the three-dimensional motion mapping network. After temperature sampling, noise reduction optimization and other processing, a target parameter sequence that accurately matches the audio semantics and prosody is finally generated.

[0101] Specifically, the input audio is resampled to a standard sampling rate of 16kHz; the HuBERT (Hidden-Unit BERT) model or the MFCC (Mel-Frequency Cepstral Coefficients) algorithm is used to extract audio semantic features, generating a feature sequence with dimensions [T, 1024], where T is the number of audio frames; the fundamental frequency F0 feature is extracted simultaneously to generate a prosodic feature sequence with dimensions [T, 1]; and temporal alignment processing is performed to ensure a precise synchronous mapping relationship between the 50Hz audio frame rate and the 25fps video frame rate, ensuring audio-visual consistency.

[0102] Construct a sequence-to-sequence mapping model based on Transformer or recurrent neural networks; the model input includes HuBERT / MFCC semantic feature sequences, F0 prosodic feature sequences, and an 80-dimensional identity encoding vector; the model output is a 64-dimensional expression parameter sequence synchronized with the audio in real time, as well as an optional blink control signal; introduce a context learning mechanism to enable the model to adapt to different speaking styles based on reference samples; use a temperature parameter to control generation diversity, with a preferred value of 0.3; support classifier-free guidance technology, and balance generation quality and diversity through a conditional strength parameter (cfg_scale, with a preferred value of 1.5) to improve the naturalness of expressions.

[0103] After training, save the following core model components and parameters to the specified directory: learnable triplane representation parameters specific to the character; SECC encoding network and LoRA fine-tuning parameters; neural rendering network, including volumetric rendering module and super-resolution module; torso synthesis network parameters; AdamW optimizer state dictionary; training configuration file (YAML format); character-specific core data, including 80-dimensional identity parameters, reference image, source keypoint coordinates, and video identifiers, to ensure quick loading and reuse later.

[0104] S5: During the inference phase, the personalized weights and the target parameter sequence are loaded, and neural rendering is used to generate a digital human video. Specifically, the neural rendering process involves first converting the target parameter sequence into a SECC semantic expression conditional graph, fusing it with the Tri-plane 3D implicit representation parameters in the personalized weights, and then performing ray tracing sampling and volume integration on the fused features using NeRF neural volume rendering technology to generate a low-resolution facial image. The low-resolution image is then input into a super-resolution network to output a high-resolution facial frame. Simultaneously, a torso image adapted to the current expression is generated through dense motion estimation and optical flow deformation. The high-resolution facial frame and torso image are then fused using a segmentation mask. Subsequently, the fused high-resolution frame sequence undergoes temporal smoothing to eliminate inter-frame jitter. Finally, the processed high-resolution frame sequence is temporally aligned and encoded with the original audio to generate the digital human video.

[0105] Specifically, such as Figure 5 As shown in the attached diagram, this process demonstrates the end-to-end personalized speaking video generation workflow, integrating three core stages: driving signal, style transfer, and rendering output. The driving audio is processed by an audio encoder (Wav2Vec / HuBERT) to extract audio features, while the style reference video is processed by a style encoder to obtain style features. Both types of features are input into the ICS-A2M model (based on a Flow Matching generator). Combined with fine-tuned weights for the target person (personalized checkpoints), the model generates a 3D motion parameter sequence including facial expressions, jaw movements, and posture. Subsequently, the person's identity texture is loaded, and a low-resolution feature map is generated using a personalized NeRF / Tri-plane 3D renderer. This map is then enhanced with a 2D augmentation network (super-resolution) to improve image quality, ultimately outputting a high-definition personalized speaking video.

[0106] S5.1: Model and personalized data initialization. First, load the trained personalized model file from the specified checkpoint directory and parse the configuration to restore the hyperparameters. Instantiate core model components such as Audio2SECC and SECC2Video and migrate them to the GPU to set them to evaluation mode. Then, extract the target person's exclusive data and the optimized Tri-plane representation and store them in the cache. Finally, perform semantic segmentation on the reference image to extract the head / torso image. If there is no custom background, extract the original video background through the KNN algorithm to complete the resource preparation before inference.

[0107] Model component loading: Load the personalized model files saved during the training phase from the specified checkpoint directory, parse the configuration file to restore the model hyperparameters; instantiate and load the Audio2SECC model (audio-driven facial expression decoding), SECC2Video model (three-plane neural rendering), MediaPipeSegmenter (semantic segmentation), SECC renderer and Face3DHelper (3D face assistance), migrate all components to the GPU device and set them to evaluation (eval) mode, and disable gradient calculation to improve inference efficiency.

[0108] Person data cache loading: Extract the target person's exclusive data (person_ds) from the checkpoint, including a 512x512x3 reference image, 80-dimensional identity parameters, a 68x3 source keypoint matrix, and video identifiers; load the optimized personalized Tri-plane representation and assign it to the rendering model's _internal cache (_last_cano_planes) to provide basic feature support for subsequent rendering.

[0109] Auxiliary resource preparation: Semantic segmentation is performed on the reference image to generate a segmentation mask, accurately separating and extracting the head region image and the repaired torso image; when no custom background is provided, the original video background is extracted based on the segmentation results using the KNN algorithm to ensure the integrity of the scene.

[0110] S5.2: During drive signal processing, the target's speaking audio is first resampled to 16kHz, and HuBERT / MFCC semantic features and F0 fundamental frequency prosodic features are extracted. After aligning the frame numbers, the Audio2SECC model is used for temperature sampling prediction and denoising optimization to obtain the expression parameter sequence. At the same time, the pose input in the form of static pose, video file or .npy parameter file is processed, Euler angles and translation vector are extracted and adapted to length, and then smoothed to generate the camera parameter sequence. It can also receive style reference video to extract relevant features to achieve style transfer, and generate blink signal sequence according to blink_mode parameter.

[0111] Specifically, such as Figure 6 As shown in the attached figure, this illustrates the expression parameter generation logic of the ICS-A2M core network, achieving accurate action prediction based on the Flow Matching framework. Style reference segments are processed by a style encoder to extract style features, which are then input together with audio features into the ICS-A2M flow matching generation network. The model's noise state is initialized, and the velocity field is predicted through a Transformer / Conformer backbone network to guide the denoising path. The noise state is then progressively optimized by integrating along the vector field using an ODE solver (Euler / RK4). Based on the action generation mechanism of feature fusion, velocity field guidance, and denoising optimization, the final output is a sequence of expression parameters that is synchronized with the audio and meets style requirements.

[0112] Audio-driven signal processing: The target speech audio is resampled to a standard sampling rate of 16kHz, and the HuBERT / MFCC semantic features of dimension [T_audio,1024] and the F0 fundamental frequency prosodic features of dimension [T_audio,1] are extracted. After aligning the number of frames of the two types of features, the number of video frames T_video=T_audio / 2 is determined, and audio frame masks and video frame masks are generated. The audio features, identity parameters and frame masks are input into the Audio2SECC model, and temperature sampling temperature=0.3 is used to predict the expression parameter sequence of 64-dimensional xT_video frames. After 20 steps of denoising optimization denoising_steps=20, the final expression parameter sequence is obtained.

[0113] Attitude-driven signal processing: Supports receiving attitude control input in the form of static attitude, video file, or .npy parameter file, extracts Euler angle sequence (euler) and translation vector sequence (trans), adapts them to the target video length through mirror index mapping, and fixes the Z-axis translation component; optionally enables the map_to_init_pose parameter to align the deviation between the reference attitude and the driving attitude, calculates the EG3D format camera matrix based on the processed Euler angle sequence (euler) and translation vector sequence (trans), and generates the camera parameter sequence [T_video,25] through temporal smoothing filtering with a kernel size of 7.

[0114] Style-driven signal processing (optional): Receives a speech style reference video, extracts style-related facial expression parameters or latent variables, integrates them into the conditional input of the Audio2SECC model, and achieves speech style transfer adaptation.

[0115] Blink control signal generation: Select the none / period strategy according to the blink_mode parameter to generate a blink signal sequence of [T_video,1]; optionally enable hold_eye_opened mode to force the eyes to remain open to meet the needs of special scenarios.

[0116] S5.3: In the SECC conditional graph generation stage, a standard neutral expression image representing the standard geometric appearance is generated based on the identity parameters and the zero expression parameters, and a source expression image representing the expression baseline is generated based on the expression parameters of the reference frame. For each video frame i, a target expression image is generated based on the identity parameters, the expression parameters of the i-th frame, and the zero pose parameters, generating a SECC sequence. The key point sequence is calculated and smoothed, including key point reconstruction, source key point preparation, and key point sequence smoothing, to eliminate motion jitter.

[0117] Specifically, the generation of the standard SECC map: Based on the identity parameters and zero-expression parameters, a standard neutral-expression image (cano_secc) of 512x512x3 is generated through the SECC renderer, representing the standard geometric appearance of the person.

[0118] Generation of the source reference SECC map: Based on the identity parameters and the reference-frame expression parameters, a source-expression image (src_secc) is generated as the expression benchmark reference for the rendering process.

[0119] Generation of the driving target SECC map sequence: For each video frame i (0 ≤ i < T_video), based on the identity parameters, the expression parameters of the i-th frame, and the zero-pose parameters, a target-expression image drv_secc[i] is generated through the SECC renderer; when blink control is enabled, the eye-expression parameters are synchronously adjusted, and finally a SECC sequence of [T_video, 3, 512, 512] is generated to provide intermediate representation support for expression driving.

[0120] Based on the identity, expression, and pose parameters, 68 three-dimensional key points are reconstructed through a three-dimensional face model, projected onto a two-dimensional plane and normalized to the range of [-1, 1] to generate a key-point tensor of [T_video, 68, 3]; then the source reference key points (src_kp) are extracted from person_ds and copied to the T_video frames to form a source key-point sequence to ensure the consistency of the inter-frame benchmark; finally, the 68x2 coordinates of the driving key-point sequence (kp_drv) are flattened into a 136-dimensional vector, and after being reshaped into a [T_video, 68, 2] sequence through moving average filtering with a kernel size of 7, the motion jitter is eliminated to ensure the visual continuity of the digital human's facial motion.

[0121] S5.4: In the temporal rendering stage, initialize the image output list and perform single-frame rendering. After the rendering loop is completed, splice the output results of all frames. Through the rendering loop, perform frame-by-frame neural rendering, convert the SECC conditions, camera parameters, etc. into high-resolution image frames, which directly determine the picture quality and detail authenticity of the output video. During single-frame rendering, construct a conditional input dictionary containing various SECC maps, reference torso / background images, segmentation masks, and source / driving key points, call the SECC2Video model for forward inference, complete the SECC condition encoding, Tri-plane fusion, ray tracing sampling and volume rendering, torso synthesis, and super-resolution upsampling, output a low-resolution image of 128x128, a final image of 512x512, and a depth map, and collect the results of all frames into the corresponding list.

[0122] As Figure 7As shown in the attached figure, the neural rendering process of Tri-plane combined with NeRF is illustrated. After inputting the target motion parameter sequence, personalized data is loaded. Using a spatial deformation and feature mapping mechanism, the sampling point x in the standard space (canonical space) is first transformed into the observation point x' in the observed space (observed space) through a deformation network. Then, features are queried by projecting them into a static Tri-plane feature library that stores the texture of the person's appearance. Combined with the camera's view direction, color and density are calculated through a lightweight MLP decoder. A low-resolution feature map is generated through volume rendering. Finally, the final high-resolution frame is output through a super-resolution network, completing the conversion process from 3D parameters to a high-resolution image.

[0123] The rendering loop is initialized, including initializing the output lists of img_lst (final image), img_raw_lst (low-resolution image), and depth_img_lst (depth map). The torch.no_grad() context is enabled to disable gradient calculation to reduce inference overhead.

[0124] Then, single-frame rendering is performed. The process is as follows: Construct a conditional input dictionary, which includes the standard SECC map (cond_cano), the source reference SECC map (cond_src), the current frame SECC map (cond_tgt), the reference torso image (ref_torso_img), the background image (bg_img), the segmentation mask (segmap), the source keypoints (kp_s), and the driving keypoints (kp_d); Then, the SECC2Video model is called for forward inference, with the input of an empty image encoder (img=None), the current frame camera parameters (camera[i:i+1]), and the conditional dictionary. The cache_backbone=False and use_cached_backbone=True are set to reuse the preloaded Tri-plane cache; The model internally completes SECC conditional encoding, Tri-plane fusion (learnable_triplane+secc_plane), ray tracing sampling and volume rendering, torso synthesis, and super-resolution upsampling, and outputs a 128x128 low-resolution image, a 512x512 final image, and a depth map. Add the single-frame image, low-resolution image, and depth map to the corresponding list, and collect all frame output results.

[0125] After the rendering loop is complete, all frame outputs are concatenated to generate tensors imgs([T_video,3,512,512]), imgs_raw([T_video,3,128,128]), and depth_imgs([T_video,1,128,128]), providing a data foundation for subsequent video synthesis.

[0126] S5.5: Through mode selection, video encoding, and audio-video synchronization, video synthesis and output are completed, converting image frames into a final digital human video that can be used directly.

[0127] Specifically, the output mode is selected based on the `out_mode` parameter, including `final` (final video), `concat_debug` (debug video stitching), and `debug` (debug video). In `concat_debug` mode, the reference image, SECC map, low-resolution rendered image, depth map, and final image are horizontally stitched together for debugging during inference. When encoding the video, tensor data is converted to NumPy arrays, and pixel values ​​are mapped from [-1,1] to the [0,255] range. The video writer is initialized using the `imageio` library (fps=25, format='FFMPEG', codec='h264'), writing image data frame by frame to generate a temporary video without an audio track. Finally, audio track synthesis is performed. When the driving signal is an audio file, the temporary video is synthesized with 16kHz audio using FFmpeg (sampling rate 16kHz, encoding libmp3lame, bitrate 2000k, shortest option for duration alignment). When the driving signal is a video file, its audio track is extracted for synthesis; if synthesis fails, a video without audio is output.

[0128] This invention utilizes deep fusion of LoRA, Tri-plane, and SECC to enable the generated digital human videos to support 512×512 high-resolution output. Through 3D implicit representation and neural rendering technology, it accurately restores details such as facial geometry, skin texture, and lip micro-movements. Combined with parallel constraints of multiple loss modules, it ensures realistic appearance details, reliable geometric accuracy, smooth temporal continuity, and stable model parameters. Furthermore, it requires only a single image or a 20-second short video and original audio to complete personalized digital human modeling within minutes, significantly reducing training costs and data requirements. Simultaneously, through precise flow matching of audio features and 3D facial expression parameters, it achieves synchronized linkage of lip movements, facial expressions, and speech, avoiding robotic stiffness. It also supports independent control of audio, posture, and style, ultimately generating realistic, credible, and synchronized digital human videos suitable for multiple scenarios.

[0129] Example 2 In a typical embodiment of the present invention, a digital human video generation system is provided, comprising: Data acquisition module: used to acquire raw human images / videos and raw audio; First processing module: used to preprocess the image / video to generate an original parameter sequence; Data reconstruction module: used to generate an initial rendering result from the original parameter sequence through a volumetric rendering network; and to perform lightweight processing on the initial rendering result to generate a reconstructed rendering result; First loss module: used to generate a first loss value from the reconstructed rendering result to quantify appearance details; Second loss module: used to generate a second loss value from the reconstructed rendering result to quantify geometric accuracy; Third loss module: used to generate a third loss value from the reconstructed rendering result to quantify temporal continuity; Fourth loss module: used to generate a fourth loss value from the reconstructed rendering result to quantify model stability. Training module: used to update the reconstructed rendering result based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; The second processing module is used to process the original audio to generate a target parameter sequence. Video generation module: used to load the personalized weights and the target parameter sequence, and process them to generate digital human videos.

[0130] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A digital human video generation method, characterized by, include: Acquire the original images / videos of the people and the original audio; The images / videos are preprocessed to generate an original parameter sequence; The original parameter sequence is processed by a volume rendering network to generate the initial rendering result; The initial rendering results are then lightweighted to generate reconstructed rendering results. The reconstructed rendering results are processed in parallel through a first loss module, a second loss module, a third loss module, and a fourth loss module. The first loss module generates a first loss value to quantify appearance details; A second loss value is generated through the second loss module to quantify geometric precision; The third loss module generates a third loss value to quantify temporal continuity. The fourth loss module generates a fourth loss value to quantify model stability. The reconstructed rendering result is updated based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; The original audio is processed to generate a target parameter sequence; The personalized weights and the target parameter sequence are loaded, and the process generates a digital human video.

2. The method of claim 1, wherein, The initial rendering result is to transform the original parameter sequence into visual features recognizable by neural rendering through SECC conditional coding in the volume rendering network, and then generate fused features by fusing the visual features through three-plane feature fusion, and finally process the fused features to generate the final result.

3. The digital human video generation method as described in claim 1, characterized in that, The reconstructed rendering result is generated from the initial rendering result through a lightweight process using the LoRA low-rank adaptive fine-tuning mechanism.

4. The digital human video generation method as described in claim 1, characterized in that, The first loss value is obtained by at least calculating a first sub-loss to quantify overall pixel consistency, calculating a second sub-loss to quantify deep feature similarity, and calculating a third sub-loss to quantify low-resolution reconstruction accuracy.

5. The digital human video generation method as described in claim 4, characterized in that, The first loss value is further quantified by calculating a fourth sub-loss to determine the degree of local detail restoration.

6. The digital human video generation method as described in claim 5, characterized in that, The fourth sub-loss is quantified at least by calculating the local first sub-loss to quantify the pixel-level lip shape reconstruction accuracy, and by calculating the local second sub-loss to quantify the deep feature similarity of the lip shape.

7. The digital human video generation method as described in claim 1, characterized in that, The second loss value is quantified at least by calculating the fifth sub-loss to quantify the foreground-background separation accuracy, and by calculating the sixth sub-loss to quantify the reasonableness of the density distribution.

8. The digital human video generation method as described in claim 1, characterized in that, The target parameter sequence is generated by stream matching processing of the original audio.

9. A digital human video generation method as described in claim 1, characterized in that, The digital human video was generated through neural rendering.

10. A digital human video generation system, characterized by, include: Data acquisition module: used to acquire raw human images / videos and raw audio; First processing module: used to preprocess the image / video to generate an original parameter sequence; Data reconstruction module: used to generate initial rendering results from the original parameter sequence via a volume rendering network; The initial rendering results are then lightweighted to generate reconstructed rendering results. First loss module: used to generate a first loss value from the reconstructed rendering result to quantify appearance details; Second loss module: used to generate a second loss value from the reconstructed rendering result to quantify geometric accuracy; Third loss module: used to generate a third loss value from the reconstructed rendering result to quantify temporal continuity; The fourth loss module is used to generate a fourth loss value from the reconstructed rendering results to quantify model stability. Training module: used to update the reconstructed rendering result based on the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to generate personalized weights; The second processing module is used to process the original audio to generate a target parameter sequence. Video generation module: used to load the personalized weights and the target parameter sequence, and process them to generate digital human videos.