Digital Human Video Generation With Speech-Action Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital human videos in scenarios such as e-commerce live streaming and film production suffer from insufficient expressiveness in actions, leading to poor video quality and a negative user experience.
Innovation Solution
A method involving a first large model to process requirement information to generate a target script with matching speech segment text, and a second large model to process the script and action video segment to create a target video where the digital human performs speech delivery aligned with the specified action, enhancing expressiveness and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional digital human video generation methods are used, then the production process is simple, but the expressiveness of actions and video quality are insufficient
Solution Approach 1:
The patent divides the video generation system into multiple specialized large models: a first large model for generating speech segment text from requirement information, a second large model for generating action video segments, and a third large model for integrating them into the final video. This segmentation allows each model to specialize in specific tasks, improving overall video quality and expressiveness while managing system complexity through modular architecture
Solution Approach 2:
The patent introduces a temporal dimension by generating and synchronizing multiple video segments (speech segments and action segments) along the time axis. The system creates a timeline where speech video segments and action video segments are precisely aligned and integrated, adding temporal coordination as a new dimension to the video generation process, thereby enhancing both quality and expressiveness
2Reliability
If digital human videos are generated with basic speech delivery, then the generation process is fast, but the naturalness and expressiveness are poor
Solution Approach 1:
The patent performs preliminary generation of speech segment text using the first large model before video generation. This pre-generated text serves as input for the second large model to create action video segments, and subsequently for the third large model to integrate everything. This preliminary text generation ensures that speech and action are coordinated from the outset, improving naturalness and consistency while maintaining efficiency through the structured workflow
Solution Approach 2:
The patent introduces speech segment text as an intermediary element that mediates between requirement information and video generation. The first large model generates this intermediary text, which then guides the second large model to create corresponding action video segments. This intermediary ensures precise alignment between speech content and action expressions, enhancing naturalness without significantly impacting generation efficiency
Data Source
AI summary
A method for generating a digital human video based on a large model, an electronic device, and a storage medium are provided, which relate to a field of artificial intelligence technologies, and may be applied to scenarios such as video livestreaming, advertisement production, and e-commerce sales. The method includes: acquiring a requirement information including an action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes a target speech segment text matching the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying a target digital human performing a speech delivery based on the target speech segment text while performing the specified action.


