Digital human video generation method based on multi-modal large model

Through multimodal large model training and adaptation, combined with 3D scanning and GAN generation technology, the problem of insufficient semantic correlation in digital human video generation is solved, efficient and realistic digital human video generation is achieved, diverse scenario needs are supported, automation level and generation quality are improved, vertical field specifications are met, and production costs and thresholds are reduced.

CN120472059APending Publication Date: 2025-08-12ZHE JIANG YAN HUANG KE JI YOU XIAN GONG SI
View PDF 0 Cites 23 Cited by

Patent Information

Application Number
CN202510546751.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the existing digital human video generation methods, text, images, audio, and action data lack semantic correlation. The generated content has problems such as lip-synchronization and speech, and mismatch between actions and semantics. It is difficult to generate high-precision 3D digital human models. Facial expressions and body movements are mechanical, lack of micro-expressions and physical realism. The general model is difficult to meet the professional needs of vertical fields. Action specifications and semantic expressions do not meet industry standards. There are many manual interventions, low degree of automation, long content production cycles and high cost, and lack of visual editing tools. It is difficult for users to dynamically adjust digital human expressions, actions and scene parameters. The feedback iteration efficiency is low. The traditional rendering pipeline is not optimized for lip-sync. The lip-sync delay often exceeds 100ms, which affects the viewing experience. User feedback is difficult to effectively integrate into model optimization. The generation effect has been stagnant for a long time, and it is impossible to continuously improve the naturalness and business adaptability.

Method used

Multimodal large model training and adaptation are adopted. Through methods such as multimodal data system construction, digital human three-dimensional model construction, semantic analysis and modal mapping, timing action and lip generation, virtual scene construction and rendering, audio and video synchronous rendering and synthesis, quality optimization and defect repair, user interaction and iterative optimization, multi-dimensional data training of text, images, audio, and actions is realized, parameterized adjustment and real-time preview are supported, and high-precision digital human models are created using 3D scanning, GAN generation and other technologies. Skeleton binding and BlendShape technology achieve smooth movements, providing a visual editing interface, combining reinforcement learning and fine-tuning models to continuously improve the generation quality.

Benefits of technology

It realizes efficient and realistic digital video generation, supports diverse scenario needs, and increases the degree of automation by more than 80%, and reaches 90% of the naturalness of digital human expressions and movements, reducing manual intervention, reducing production costs, and meeting vertical field specifications. Users can quickly adjust parameters, and continuously improve the naturalness of generated content and business adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472059A_ABST
    Figure CN120472059A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of virtual person generation, and particularly relates to a digital person video generation method based on a multi-modal large model, and the method comprises the following steps: 1, constructing a multi-modal data system; 2, multi-modal large model training and adaptation are carried out; 3, constructing a digital human three-dimensional model; step 4, performing semantic analysis and modal mapping; 5, generating a time sequence action and a mouth shape; step 6, building and rendering a virtual scene; step 7, audio and video synchronous rendering and synthesis; step 8, quality optimization and defect repair; and step 9, performing user interaction and iterative optimization. Through technical innovation and engineering, the core pain point in digital human video generation is solved, efficient, vivid and customizable content production capacity is provided for virtual anchors, intelligent customer service, enterprise training and other scenes, and the AI digital human technology is promoted to be applied to large-scale business from experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of virtual character generation, and in particular relates to a digital human video generation method based on a multimodal large model. Background Art

[0002] Digital human video generation is the process of using artificial intelligence to create virtual human avatars, imitating their appearance, behavior, and language abilities in videos with the help of artificial intelligence. Its core features include a rich library of digital human models (covering images of different genders, ages, ethnicities, and styles), the ability to simulate realistic movements and expressions (such as natural body movements and delicate facial expressions), and intelligent speech synthesis. The process is as follows: After registering and logging into the platform, users select a digital human from the model library, adjust its appearance details (such as hairstyle and clothing), set the action sequence and voice content, and add background scenes, props, and other elements. Finally, they click the "Generate" button, and the platform outputs a complete digital human video based on the settings. This technology is widely used in various fields: in education, teaching videos featuring digital human teachers explaining knowledge points can be created to enhance student learning interest; in corporate promotions, digital humans can serve as spokespersons for product introductions or corporate culture videos, enhancing their appeal; and in the entertainment industry, they can be used to create virtual idol performance videos and short virtual dramas, injecting new formats into the industry.

[0003] In the existing digital human video generation methods, text, images, audio, and action data lack semantic correlation. The generated content has problems such as lip-sync and voice asynchrony, and action and semantic mismatch. It is difficult to generate high-precision 3D digital human models. Facial expressions and body movements are mechanical, lacking micro-expressions and physical realism. General models are difficult to meet the professional needs of vertical fields. Action specifications and semantic expressions do not meet industry standards. There is a lot of manual intervention and a low degree of automation. The entire process from modeling to rendering relies on designers to adjust frame by frame. The content production cycle is long and the cost is high. There is a lack of visual editing tools. It is difficult for users to dynamically adjust digital human expressions, actions and scene parameters. The feedback iteration efficiency is low. The traditional rendering pipeline is not optimized for lip-sync. The lip-sync and voice delay often exceeds 100ms, affecting the viewing experience. User feedback is difficult to effectively integrate into model optimization. The generation effect stagnates for a long time, and it is impossible to continuously improve naturalness and business adaptability. Summary of the Invention

[0004] The purpose of this invention is to provide a digital human video generation method based on a multimodal large model. Through technological innovation and engineering implementation, it can solve the core pain points in digital human video generation, provide efficient, realistic, and customizable content production capabilities for scenarios such as virtual anchors, intelligent customer service, and corporate training, and promote AI digital human technology from experiments to large-scale commercial applications.

[0005] The technical solutions adopted by the present invention are as follows:

[0006] A method for generating a digital human video based on a multimodal large model, the method comprising the following steps:

[0007] Step 1. Construction of multimodal data system;

[0008] Step 2. Multimodal large model training and adaptation;

[0009] Step 3. Construction of digital human 3D model;

[0010] Step 4. Semantic parsing and modality mapping;

[0011] Step 5. Generate timing actions and lip shapes;

[0012] Step 6. Virtual scene construction and rendering;

[0013] Step 7. Audio and video synchronization rendering and synthesis;

[0014] Step 8. Quality optimization and defect repair;

[0015] Step 9. User interaction and iterative optimization.

[0016] In a preferred embodiment, the construction of the multimodal data system includes data acquisition, collecting text, image, audio, motion capture data, unifying the resolution, lighting correction, and feature point annotation of images and videos, denoising audio, framing, and transcribing speech into text, mapping motion data to a skeletal binding model, removing abnormal frames, performing intent classification and emotion annotation on text, collecting digital human images, using a 3D scanner to obtain facial and body geometry data, taking photos under multiple lighting conditions, training lighting robustness, taking cloth textures and prop images for later material mapping, recording the same text in different tones, recording background sounds of offices, conference rooms, and outdoor scenes for noise reduction model training, optical motion capture, and collecting full-body movements by actors wearing marker point suits in a green screen studio. , use Faceware equipment to record micro-expressions, add semantic labels and emotional labels to action sequences, unify the resolution and format to 1920×1080 or 4K resolution, store as PNG / TIFF, perform lighting correction and feature point annotation, use OpenCV for histogram equalization, eliminate shadow differences, mark facial key points and body joints, use Webrtc noise reduction algorithm or SpectralSubtraction to remove environmental noise, analyze text emotions through the BERTNLP model, generate emotion vectors, skeleton mapping and anomaly removal, map the motion capture data to the digital human skeleton binding model, eliminate abnormal frames with sudden changes in motion trajectory, supplement missing frames through linear interpolation, and scale the action data to the body size range of the digital human model.

[0017] In a preferred embodiment, the multimodal large model training and adaptation adopts CLIP, DALL·E, and Whisper basic models, focusing on training the semantic alignment ability of text-image-audio-action, injecting domain-specific data, strengthening the association between modalities through contrastive learning, and adopting the encoder-decoder architecture, BERT / XLNet, extracting semantic features, ResNet / ViT, extracting visual features, Mel-Spectrogram+CNN, extracting acoustic features, skeleton joint coordinate sequence+LSTM, extracting motion features, realizing inter-modal information fusion through cross-attention mechanism, and loading from the pre-trained model. The weights are used as the starting point for training. The action encoder is randomly initialized or zero-initialized. Modal alignment preprocessing is performed. Text, image, audio, and action data are grouped according to semantics. Token embedding vectors are generated after word segmentation. Image patch feature vectors are generated through ViT. Mel-spectrograms are input into CNN, and acoustic feature vectors are output. The bone coordinate sequence is encoded into motion feature vectors through LSTM. The domain test set is used to calculate the retrieval accuracy of text-image / audio / action. The domain text is input. The generated voice intonation and digital human facial expressions are checked to see if they meet the scene specifications. The features of each modality are projected into two-dimensional space to observe the degree of aggregation of similar semantic samples.

[0018] In a preferred embodiment, the digital human three-dimensional model is constructed based on text description or image input, using a GAN network to generate a high-precision 3D human face model, or obtaining real-life geometric data through scanning and reconstruction, customizing clothing, hairstyle, and makeup details, supporting parametric adjustment, adding a skeleton system to the model, binding facial BlendShape and body joints, establishing a motion control interface, using text description to generate a 3D model, parsing keywords through natural language processing, inputting them into a Text-to-3D model generator, generating an initial mesh model, using MakeHuman or MetahumanCreator to generate a basic human face model based on a single photo, reconstructing 3D point clouds from multi-angle photos through SFM technology, and then using PoissonSurfaceReconstruction to generate a mesh, using ArtecEva or Shining3DEinScan to obtain real-life full-body point cloud data, and adjusting PCA principal component parameters to match input features for face modeling, and using Mixamo for body modeling. Or use Daz3D to generate a standard human body model, adjust the body shape parametrically, sculpt personalized features such as muscle lines and fat distribution, use Blender or ZBrush to sculpt facial wrinkles, pores, and hair texture details, simulate cloth folds and jewelry details for clothing, use Topogun or Blender's Remesh function to convert high-poly models into low-poly models while retaining key details, follow the direction of the muscles for facial topology, split UV maps to avoid texture stretching, use RizomUV to automatically unfold complex structures, use SSS subsurface scattering material for the skin, set diffuse reflection, roughness, and scattering radius, add Fresnel reflection to the iris, set anisotropic pupils to simulate light refraction, import MixamoFBX skeletons or Humanoid skeletons as skeleton templates, which include 24 main joints, add BlendShape controllers to the facial skeleton extensions, corresponding to basic expressions such as opening eyes, smiling, and frowning, automatic skinning uses the Blender plug-in to automatically bind bones and meshes, generate initial weights, and correct penetrating areas in weight painting mode.

[0019] In a preferred solution, the semantic parsing and modal mapping is to perform semantic analysis on the text / audio input by the user, extract emotional tendencies, action instructions, and scene requirements, and convert semantics into multimodal instructions through a trained large model. The input text T is encoded using a BERT-like model to generate a context-aware word embedding vector. The emotional label is mapped to a facial expression parameter vector through a linear layer. According to the text T and emotion E, an emotion-controllable TTS model is used to generate speech:

[0020] A synth =TTS(T,E;θ)

[0021] Among them, θ is a model parameter that controls intonation by adjusting the emotion encoding;

[0022] The retrieved actions are weighted and fused, with the weight being the similarity score sim(M i ):

[0023]

[0024] Through the time interpolation algorithm f interp Smooth action transitions;

[0025] Combine the emotion parameter p with real-time speech features to dynamically adjust the expression:

[0026] P dynamic =P×(1+α·Voicing)

[0027] Where α is the adjustment coefficient.

[0028] In a preferred embodiment, the timed action and lip shape generation is based on the speech features and text transcription results, the lip shape animation is generated by the TTS model, the semantic instructions and the action database are combined to generate coherent body movements, the kinematic algorithm is used to optimize the smoothness of the joint transition, the facial expression parameters are mixed according to the emotion label, the eyebrows, eyes, and mouth parts are driven to change in real time, the speech signal is processed, the input speech is cut into short frames of 50 milliseconds, the Mel spectrum and energy value are extracted, and the speech is converted into text by the speech transcription tool, the pronunciation time point of each word is marked, and each word is decomposed into phonemes according to the transcribed text, matched with the predefined lip shape template, and the phonemes are converted into The corresponding lip shape parameters are input into the BlendShape controller of the digital human facial model. The delay between voice playback and image rendering is calculated in advance, and the delay time is reserved when generating the lip shape. The action keywords and emotional requirements are extracted from the input text or audio. According to the semantics and emotions, the matching action clips are retrieved from the action library, and the action clips are combined in semantic order. The kinematic algorithm is used to calculate the joint angle transition of adjacent action clips to avoid freezes or penetration when switching actions. The emotional label of the current content is determined through text sentiment analysis or voice intonation recognition. According to the emotional label, the facial expression parameters are mixed and combined with the emotional fluctuations of the voice to dynamically fine-tune the expression intensity.

[0029] In a preferred embodiment, the virtual scene is built and rendered based on text description or image reference to generate a virtual background, support material, lighting, and camera angle adjustment, import the digital human model, action sequence, and background scene into the rendering engine, set the camera movement path, complete the spatial composition, use Blender / 3dsMax to create a basic model, and generate the walls, ground, and ceiling by stretching the Box geometry, add chamfer details, and optimize the real-time rendering to turn on screen space ambient occlusion, turn off anti-aliasing, enable occlusion culling, hide objects outside the camera's field of view, separate the rendering layer, and import the rendered video stream and voice track into the editing software synchronously, add background music, subtitles, and transition effects.

[0030] In a preferred solution, the audio and video synchronous rendering and synthesis is to accelerate the rendering through GPU, generate frame-by-frame video images, synchronously output the audio stream, mix digital human voice, environmental sound effects, background music, superimpose subtitles or special effects, output the complete video stream, import the video stream and audio stream in the engine or post-production software, align the timestamps, calculate the audio and video delay through the phase correlation algorithm, and perform batch correction.

[0031] In a preferred solution, the quality optimization and defect repair are to improve image quality by applying super-resolution and denoising algorithms, optimize shadow and reflection effects through ray tracing, simulate human body dynamics using a physics engine, correct unnatural movements, check the matching degree between speech content and lip shape, movement and semantics, automatically mark abnormal frames and trigger regeneration, apply the Real-ESRGAN model to low-resolution video frames, reconstruct 4K resolution images through deep learning, focus on enhancing the facial details of digital humans, enhance the text and icons in the virtual scene separately with a sharpening filter, use collision volume detection to mark the penetration area between digital human limbs and scene props, trigger action redirection, and regenerate lip shape parameters for abnormal frames through the Wav2Lip model based on the original text and speech features to overwrite erroneous data.

[0032] In a preferred solution, the user interaction and iterative optimization provides a visual editing interface, allowing users to modify the digital human's expression, movement speed, and scene parameters, preview the effects in real time, collect user ratings and error reports, fine-tune the model through reinforcement learning, continuously improve the naturalness and business adaptability of the generated content, collect feedback from multiple channels, and the rating system provides a 1-5 star rating with open comments, recording the frequency and magnitude of user parameter adjustments.

[0033] The technical effects achieved by the present invention are:

[0034] Through multi-dimensional data training including text, images, audio, and actions, the model can capture the deep connections between semantics, vision, and hearing, avoiding the modal fragmentation problem of traditional methods. Multimodal data is cleaned, labeled, and pre-processed to ensure controllable data quality. After injecting vertical domain data, the model can generate content that meets industry standards.

[0035] Users only need to provide text or simple instructions, and the system automatically completes the entire process, including semantic analysis, action generation, and scene rendering. There is no need for manual frame-by-frame adjustment, which improves efficiency by over 80% compared to traditional animation production. Leveraging the real-time rendering capabilities of the Unity / Unreal engine, the system supports dynamic adjustment of parameters such as digital human expressions, action speed, and scene lighting, and allows for real-time preview of the effects, reducing waiting time for repeated rendering. The parametric model design supports rapid reuse, and the same digital human can be adapted to multiple scenes.

[0036] Through 3D scanning, GAN generation and other technologies, digital human models with millimeter-level precision are created, supporting the rendering of micro-details such as pores and hair. Skeletal rigging and BlendShape technology enable over 60 facial expressions and smooth body movements, with expressions more than 90% as natural as real people. Physical engine simulation is added to hair and clothing, and movements conform to human dynamics, avoiding a "mechanical" feel.

[0037] It supports generating customized virtual backgrounds from text / images and seamlessly integrates them with digital human movements to meet the needs of diverse scenarios. The visual editing interface allows non-technical users to quickly modify parameters such as digital human expressions and scene lighting, and preview the effects in real time, lowering the usage threshold. It collects real feedback through user ratings and error reports, and combines reinforcement learning to fine-tune the model to continuously improve the generation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of a digital human video generation method based on a multimodal large model of the present invention. DETAILED DESCRIPTION

[0039] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0040] See also Figure 1 As shown, the present invention provides a method for generating digital human videos based on a multimodal large model, the video generation method comprising the following steps:

[0041] Step 1. Construction of multimodal data system;

[0042] Step 2. Multimodal large model training and adaptation;

[0043] Step 3. Construction of digital human 3D model;

[0044] Step 4. Semantic parsing and modality mapping;

[0045] Step 5. Generate timing actions and lip shapes;

[0046] Step 6. Virtual scene construction and rendering;

[0047] Step 7. Audio and video synchronization rendering and synthesis;

[0048] Step 8. Quality optimization and defect repair;

[0049] Step 9. User interaction and iterative optimization.

[0050] The construction of a multimodal data system includes data collection, collecting text, image, audio, motion capture data, unifying the resolution of images and videos, lighting correction, feature point annotation, audio noise reduction, framing, and speech transcription into text, mapping motion data to a skeletal binding model, removing abnormal frames, intent classification and emotion annotation of text, digital human image collection, using a 3D scanner to obtain facial and body geometry data, taking photos under multiple lighting conditions, training lighting robustness, taking cloth textures and prop images for later material mapping, recording the same text in different tones, recording background sounds of offices, conference rooms, and outdoor scenes for noise reduction model training, optical motion capture, in a green screen studio, actors wearing marker point suits, collecting full-body movements, using Face ID to capture the full-body movements, and using Face ID to capture the full-body movements. eware equipment, records micro-expressions, adds semantic and emotional labels to action sequences, unifies resolution and format to 1920×1080 or 4K resolution, stores in PNG / TIFF, performs lighting correction and feature point annotation, uses OpenCV for histogram equalization to eliminate shadow differences, annotates facial key points and body joints, uses Webrtc noise reduction algorithm or SpectralSubtraction to remove environmental noise, analyzes text emotions through the BERTNLP model, generates emotion vectors, skeleton mapping and anomaly removal, maps motion capture data to the digital human skeleton binding model, removes abnormal frames with sudden changes in motion trajectory, supplements missing frames through linear interpolation, and scales the action data to the body size range of the digital human model.

[0051] The multimodal large model training and adaptation uses CLIP, DALL·E, and Whisper as basic models, focusing on training the semantic alignment ability of text, image, audio, and action, injecting domain-specific data, and strengthening the association between modalities through contrastive learning. The encoder-decoder architecture is used, BERT / XLNet is used to extract semantic features, ResNet / ViT is used to extract visual features, Mel-Spectrogram+CNN is used to extract acoustic features, and skeleton joint coordinate sequence + LSTM is used to extract motion features. The cross-attention mechanism is used to achieve inter-modal information fusion, and weights are loaded from the pre-trained model as At the training start point, the action encoder is randomly initialized or zero-initialized, and modality alignment preprocessing is performed. Text, image, audio, and action data are grouped according to semantics. Token embedding vectors are generated after word segmentation. Image patch feature vectors are generated through ViT. The Mel-spectrogram is input into CNN, and the acoustic feature vector is output. The bone coordinate sequence is encoded into a motion feature vector through LSTM. Using the domain test set, the retrieval accuracy of text-image / audio / action is calculated. The domain text is input, and the generated speech intonation and digital human facial expressions and movements are checked to see if they meet the scene specifications. The features of each modality are projected into two-dimensional space to observe the degree of aggregation of similar semantic samples.

[0052] The digital human 3D model is constructed based on text description or image input, using the GAN network to generate a high-precision 3D face model, or obtaining real-life geometric data through scanning and reconstruction, customizing clothing, hairstyle, and makeup details, supporting parametric adjustment, adding a skeletal system to the model, binding facial BlendShape and body joints, establishing a motion control interface, using text description to generate a 3D model, parsing keywords through natural language processing, inputting them into the Text-to-3D model generator, and generating an initial mesh model. A single photo uses MakeHuman or MetahumanCreator to generate a basic face model based on a single image, multi-angle photos use SFM technology to reconstruct 3D point clouds, and then use PoissonSurfaceReconstruction to generate meshes. ArtecEva or Shining3DEinScan is used to obtain real-life full-body point cloud data. Face modeling matches input features by adjusting PCA principal component parameters, and body modeling uses Mixamo or Daz3D Generate a standard human body model, adjust the body shape parametrically, sculpt personalized features such as muscle lines and fat distribution, use Blender or ZBrush to sculpt facial wrinkles, pores, and hair texture details, simulate cloth folds and jewelry details for clothing, use Topogun or Blender's Remesh function to convert high-poly models into low-poly models while retaining key details, follow the direction of the muscles in the facial topology, split the UV map to avoid texture stretching, use RizomUV to automatically unfold complex structures, use SSS subsurface scattering material for the skin, set diffuse reflection, roughness, and scattering radius, add Fresnel reflection to the iris, set anisotropy in the pupil to simulate light refraction, import MixamoFBX skeletons or Humanoid skeletons as skeleton templates, including 24 main joints, add BlendShape controllers to the facial skeleton extension, corresponding to basic expressions such as opening eyes, smiling, and frowning, automatic skinning uses the Blender plug-in to automatically bind bones and meshes, generate initial weights, and correct penetrating areas in weight painting mode.

[0053] Semantic parsing and modal mapping involves performing semantic analysis on user input text / audio to extract emotional tendencies, action instructions, and scenario requirements. Using a trained large model, semantics are converted into multimodal instructions. A BERT-like model is used to encode the input text T and generate context-aware word embedding vectors. Emotion labels are mapped to facial expression parameter vectors via a linear layer. Based on the text T and emotion E, an emotion-controlled TTS model is used to generate speech:

[0054] A synth =TTS(T,E;θ)

[0055] Among them, θ is a model parameter that controls intonation by adjusting the emotion encoding;

[0056] The retrieved actions are weighted and fused, with the weight being the similarity score sim(M i ):

[0057]

[0058] Through the time interpolation algorithm f interp Smooth action transitions;

[0059] Combine the emotion parameter p with real-time speech features to dynamically adjust the expression:

[0060] P dynamic =P×(1+α·Voicing)

[0061] Where α is the adjustment coefficient.

[0062] The timed action and lip shape generation is based on speech features and text transcription results. The TTS model is used to generate lip shape animation. Combining semantic instructions with the action database, coherent body movements are generated. The kinematic algorithm is used to optimize the smoothness of joint transitions. According to the emotional label, the facial expression parameters are mixed to drive the real-time changes of eyebrows, eyes, and mouth parts. The speech signal is processed and the input speech is cut into short frames of 50 milliseconds. The Mel spectrum and energy value are extracted. At the same time, the speech is converted into text through the speech transcription tool, and the pronunciation time point of each word is marked. According to the transcribed text, each word is decomposed into phonemes, matched with the predefined lip shape template, and the lip shape corresponding to the phoneme is obtained. The parameters of the digital human facial model are input into the BlendShape controller, which calculates the delay between voice playback and image rendering in advance, reserves delay time when generating lip shapes, extracts action keywords and emotional requirements from the input text or audio, retrieves matching action clips from the action library based on semantics and emotions, combines action clips in semantic order, and uses kinematic algorithms to calculate the joint angle transitions of adjacent action clips to avoid freezes or model penetration when switching actions. Through text sentiment analysis or voice intonation recognition, the emotional label of the current content is determined. Based on the emotional label, the facial expression parameters are mixed, combined with the emotional fluctuations of the voice, and the expression intensity is dynamically fine-tuned.

[0063] Virtual scene construction and rendering are based on text descriptions or image references to generate virtual backgrounds, support material, lighting, and camera angle adjustments, import digital human models, action sequences, and background scenes into the rendering engine, set the camera movement path, complete the spatial composition, use Blender / 3dsMax to create basic models, and generate walls, floors, and ceilings by stretching Box geometry, adding chamfer details, and real-time rendering optimization. Turn on screen space ambient occlusion, turn off anti-aliasing, enable occlusion culling, hide objects outside the camera's field of view, separate rendering layers, and synchronously import the rendered video stream and voice track into the editing software to add background music, subtitles, and transition effects.

[0064] Audio and video synchronous rendering and synthesis uses GPU accelerated rendering to generate frame-by-frame video images, synchronously output audio streams, mix digital human voice, environmental sound effects, background music, overlay subtitles or special effects, output complete video streams, import video streams and audio streams into the engine or post-production software, align based on timestamps, calculate audio and video delays through phase correlation algorithms, and perform batch corrections.

[0065] Quality optimization and defect repair include applying super-resolution and denoising algorithms to improve image quality, optimizing shadow and reflection effects through ray tracing, using a physics engine to simulate human dynamics, correcting unnatural movements, checking the matching between speech content and lip shape, movement and semantics, automatically marking abnormal frames and triggering regeneration, applying the Real-ESRGAN model to low-resolution video frames, reconstructing 4K resolution images through deep learning, focusing on enhancing digital human facial details, and using sharpening filters to enhance text and icons in virtual scenes separately. Collision volume detection is used to mark the penetration area between digital human limbs and scene props, triggering action redirection, and regenerating lip shape parameters for abnormal frames through the Wav2Lip model based on the original text and speech features to overwrite erroneous data.

[0066] User interaction and iterative optimization provide a visual editing interface that allows users to modify digital human expressions, movement speed, and scene parameters, preview the effects in real time, collect user ratings and error reports, fine-tune the model through reinforcement learning, and continuously improve the naturalness and business adaptability of the generated content. Multi-channel feedback is collected, and the rating system provides a 1-5 star rating with open comments, recording the frequency and magnitude of user parameter adjustments.

[0067] In this invention, through multi-dimensional data training of text, images, audio, and actions, the model can capture the deep connection between "semantics, vision, and hearing", avoiding the problem of modal separation in traditional methods. The multimodal data is cleaned, labeled, and pre-processed to ensure controllable data quality. After injecting vertical field data, the model can generate content that meets industry standards.

[0068] Users only need to provide text or simple instructions, and the system automatically completes the entire process, including semantic analysis, action generation, and scene rendering. There is no need for manual frame-by-frame adjustment, which improves efficiency by over 80% compared to traditional animation production. Leveraging the real-time rendering capabilities of the Unity / Unreal engine, the system supports dynamic adjustment of parameters such as digital human expressions, action speed, and scene lighting, and allows for real-time preview of the effects, reducing waiting time for repeated rendering. The parametric model design supports rapid reuse, and the same digital human can be adapted to multiple scenes.

[0069] Through 3D scanning, GAN generation and other technologies, digital human models with millimeter-level precision are created, supporting the rendering of micro-details such as pores and hair. Skeletal rigging and BlendShape technology enable over 60 facial expressions and smooth body movements, with expressions more than 90% as natural as real people. Physical engine simulation is added to hair and clothing, and movements conform to human dynamics, avoiding a "mechanical" feel.

[0070] It supports generating customized virtual backgrounds from text / images and seamlessly integrates them with digital human movements to meet the needs of diverse scenarios. The visual editing interface allows non-technical users to quickly modify parameters such as digital human expressions and scene lighting, and preview the effects in real time, lowering the usage threshold. It collects real feedback through user ratings and error reports, and combines reinforcement learning to fine-tune the model to continuously improve the generation quality.

[0071] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.

Claims

1. A method for generating digital human videos based on a multimodal large model, characterized by: The video generation method comprises the following steps: Step 1. Construction of multimodal data system; Step 2. Multimodal large model training and adaptation; Step 3. Construction of digital human 3D model; Step 4. Semantic parsing and modality mapping; Step 5. Generate timing actions and lip shapes; Step 6. Virtual scene construction and rendering; Step 7. Audio and video synchronization rendering and synthesis; Step 8. Quality optimization and defect repair; Step 9. User interaction and iterative optimization.

2. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The construction of the multimodal data system includes data acquisition, collecting text, image, audio, motion capture data, unifying the resolution of images and videos, lighting correction, feature point annotation, audio noise reduction, framing, and speech transcription into text, mapping motion data to a skeletal binding model, eliminating abnormal frames, intent classification and emotion annotation of text, digital human image acquisition, using a 3D scanner to obtain facial and body geometry data, taking photos under multiple lighting conditions, training lighting robustness, taking cloth textures and prop images for later material mapping, recording the same text in different tones, recording background sounds of offices, conference rooms, and outdoor scenes for noise reduction model training, optical motion capture, in a green screen studio, actors wearing marker point suits, collecting full-body movements, using Fa ceware equipment records micro-expressions, adds semantic and emotional labels to action sequences, unifies resolution and format to 1920×1080 or 4K resolution, stores in PNG / TIFF, performs lighting correction and feature point annotation, uses OpenCV for histogram equalization to eliminate shadow differences, annotates facial key points and body joints, uses Webrtc noise reduction algorithm or SpectralSubtraction to remove environmental noise, analyzes text emotions through the BERTNLP model, generates emotion vectors, performs skeleton mapping and anomaly removal, maps motion capture data to the digital human skeleton binding model, removes abnormal frames with sudden changes in motion trajectory, supplements missing frames through linear interpolation, and scales the action data to the body size range of the digital human model.

3. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The multimodal large model training and adaptation adopts CLIP, DALL·E, and Whisper basic models, focusing on training the semantic alignment ability of text, image, audio, and action, injecting domain-specific data, strengthening the association between modalities through contrastive learning, and adopting an encoder-decoder architecture. BERT / XLNet is used to extract semantic features, ResNet / ViT is used to extract visual features, Mel-Spectrogram+CNN is used to extract acoustic features, and skeleton joint coordinate sequence+LSTM is used to extract motion features. The cross-attention mechanism is used to achieve inter-modal information fusion, and weights are loaded from the pre-trained model. As the starting point for training, the action encoder is randomly initialized or zero-initialized, and modality alignment preprocessing is performed. The text, image, audio, and action data are grouped according to semantics. Token embedding vectors are generated after word segmentation. Image patch feature vectors are generated through ViT. The Mel-spectrogram is input into CNN, and the acoustic feature vector is output. The bone coordinate sequence is encoded into a motion feature vector through LSTM. Using the domain test set, the retrieval accuracy of text-image / audio / action is calculated. The domain text is input, and the generated speech intonation and digital human facial expressions and movements are checked to see if they meet the scene specifications. The features of each modality are projected into two-dimensional space to observe the degree of aggregation of similar semantic samples.

4. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The digital human 3D model is constructed based on text description or image input, using a GAN network to generate a high-precision 3D face model, or obtaining real-life geometric data through scanning and reconstruction, customizing clothing, hairstyle, and makeup details, supporting parametric adjustment, adding a skeleton system to the model, binding facial BlendShape and body joints, establishing a motion control interface, using text description to generate a 3D model, parsing keywords through natural language processing, inputting them into a Text-to-3D model generator, generating an initial mesh model, using MakeHuman or MetahumanCreator to generate a basic face model based on a single photo, reconstructing 3D point clouds from multi-angle photos through SFM technology, and then using PoissonSurfaceReconstruction to generate meshes, using ArtecEva or Shining3DEinScan to obtain real-life full-body point cloud data, and adjusting PCA principal component parameters to match input features for face modeling, and using Mixamo or Daz3 for body modeling. D generates a standard human body model, adjusts the body shape parametrically, sculpts personalized features such as muscle lines and fat distribution, uses Blender or ZBrush to sculpt facial wrinkles, pores, and hair texture details, simulates cloth folds and jewelry details for clothing, uses Topogun or Blender's Remesh function to convert high-poly models into low-poly models while retaining key details, and the facial topology follows the direction of the muscles. Split UV maps to avoid texture stretching, use RizomUV to automatically unfold complex structures, use SSS subsurface scattering material for the skin, set diffuse reflection, roughness, and scattering radius, add Fresnel reflection to the iris, and set the pupil to anisotropic to simulate light refraction. The skeleton template chooses to import MixamoFBX skeletons or Humanoid skeletons, which contain 24 main joints. The facial skeleton extension adds a BlendShape controller to correspond to basic expressions such as opening eyes, smiling, and frowning. Automatic skinning uses the Blender plug-in to automatically bind bones and meshes, generate initial weights, and correct the penetrating areas in weight painting mode.

5. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The semantic parsing and modal mapping is to perform semantic analysis on the text / audio input by the user, extract emotional tendencies, action instructions, and scene requirements, and convert the semantics into multimodal instructions through a trained large model. The input text T is encoded using a BERT-like model to generate a context-aware word embedding vector. The emotional label is mapped to a facial expression parameter vector through a linear layer. Based on the text T and emotion E, an emotion-controllable TTS model is used to generate speech: TO synth =TTS(T,E;θ) Among them, θ is a model parameter that controls intonation by adjusting the emotion encoding; The retrieved actions are weighted and fused, with the weight being the similarity score sim(M i ): Through the time interpolation algorithm f interp Smooth action transitions; Combine the emotion parameter p with real-time speech features to dynamically adjust the expression: P dynamic =P×(1+α·Voicing) Where α is the adjustment coefficient.

6. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The timed action and lip shape generation is based on speech features and text transcription results, and lip shape animation is generated through the TTS model. Semantic instructions and action database are combined to generate coherent body movements. The kinematic algorithm is used to optimize the smoothness of joint transitions. According to the emotion label, facial expression parameters are mixed to drive the real-time changes of eyebrows, eyes, and mouth parts. Speech signal processing cuts the input speech into short frames of 50 milliseconds, extracts the Mel spectrum and energy value, and converts the speech into text through the speech transcription tool, marking the pronunciation time point of each word. According to the transcribed text, each word is decomposed into phonemes, matched with the predefined lip shape template, and the corresponding mouth of the phoneme is converted into a single word. The BlendShape controller of the digital human facial model inputs shape parameters, calculates the delay between voice playback and image rendering in advance, reserves delay time when generating lip shape, extracts action keywords and emotional requirements from the input text or audio, retrieves matching action clips from the action library based on semantics and emotions, combines action clips in semantic order, uses kinematic algorithms to calculate the joint angle transition of adjacent action clips, avoids freezes or model penetration when switching actions, determines the emotional label of the current content through text emotion analysis or voice intonation recognition, and dynamically fine-tunes the expression intensity by mixing facial expression parameters based on the emotional label and combining the emotional fluctuations of the voice.

7. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The virtual scene construction and rendering is based on text description or image reference to generate a virtual background, support material, lighting, and camera angle adjustment, import the digital human model, action sequence, and background scene into the rendering engine, set the camera movement path, complete the spatial composition, use Blender / 3dsMax to create the basic model, and generate the walls, ground, and ceiling by stretching the Box geometry, add chamfer details, and optimize the real-time rendering by turning on screen space ambient occlusion, turning off anti-aliasing, enabling occlusion culling, hiding objects outside the camera's field of view, separating the rendering layers, and synchronously importing the rendered video stream and voice track into the editing software, adding background music, subtitles, and transition effects.

8. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The audio and video synchronous rendering and synthesis is to generate frame-by-frame video images through GPU accelerated rendering, synchronously output audio streams, mix digital human voices, environmental sound effects, background music, overlay subtitles or special effects, output complete video streams, import video streams and audio streams into the engine or post-production software, align based on timestamps, calculate audio and video delays through phase correlation algorithms, and perform batch corrections.

9. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The quality optimization and defect repair are to improve the image quality by applying super-resolution and denoising algorithms, optimize the shadow and reflection effects through ray tracing, use the physics engine to simulate human dynamics, correct unnatural movements, check the matching degree between speech content and lip shape, movement and semantics, automatically mark abnormal frames and trigger regeneration, apply the Real-ESRGAN model to low-resolution video frames, reconstruct 4K resolution images through deep learning, focus on enhancing the facial details of digital humans, and use sharpening filters to enhance the text and icons in the virtual scene separately. Collision volume detection is used to mark the penetration area between digital human limbs and scene props, trigger action redirection, and regenerate lip shape parameters for abnormal frames through the Wav2Lip model based on the original text and speech features to overwrite the erroneous data.

10. The method for generating digital human videos based on a multimodal large model according to claim 1, characterized in that: The user interaction and iterative optimization provide a visual editing interface that allows users to modify digital human expressions, movement speed, and scene parameters, preview the effects in real time, collect user ratings and error reports, fine-tune the model through reinforcement learning, and continuously improve the naturalness and business adaptability of the generated content. Multi-channel feedback is collected, and the rating system provides a 1-5 star rating with open comments, recording the frequency and magnitude of user parameter adjustments.

Citation Information

Cited By

  • Humanoid robot interaction system

    CN120715912A

  • Character-driven digital character speaking video generation method and system

    CN120812364A

  • A method and system for generating digital human speaking videos via text-driven technology

    CN120812364B

  • Three-dimensional Gaussian digital human generation system and method and electronic equipment

    CN120823342A

  • Three-dimensional digital human generation method and system capable of voice interaction

    CN120931773A