A multi-role digital human podcast generation method and system based on multi-modal AI
By automating the entire process of generating multi-role digital human podcasts through multimodal AI technology, the problem of comprehensive understanding and intelligent arrangement of multimodal content in the generation of multi-role digital human podcasts is solved, improving the efficiency and quality of content creation, and is applicable to a variety of application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for generating multi-role digital human podcasts suffer from several problems, including insufficient comprehensive understanding of multimodal content, lack of intelligent content arrangement, simplistic character interaction mechanisms, monotonous performance, low system integration, and insufficient personalization capabilities. These issues result in generated content that lacks naturalness and logic, and is inefficient.
Employing multimodal AI technology, through multimodal information conversion, semantic understanding, structured outline design, digital human role setting and allocation, dialogue unit script optimization, scene synthesis and audiovisual synchronization, it achieves fully automated generation of multi-source heterogeneous data into multi-role interactive videos.
It significantly improves content creation efficiency, automates the entire process of multi-role interactive videos, enhances the flexibility and adaptability of content processing, and is suitable for various scenarios such as education and training, corporate publicity, and news broadcasting, possessing high versatility and scalability.
Smart Images

Figure CN121099162B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of artificial intelligence, multimodal content processing and digital human synthesis, and specifically relates to a method and system for generating multi-role digital human podcasts based on multimodal AI. Background Technology
[0002] Currently, various digital human and podcast generation technologies have emerged in the field of digital content creation. Existing digital human generation technologies can already achieve text-driven generation of single digital human videos, enabling digital humans to present content according to a preset script through speech synthesis and lip-syncing technologies. These technologies are mainly used in scenarios such as virtual anchors and digital assistants. Traditional podcast production typically requires professionals to handle content planning, scriptwriting, recording, and post-production editing to create a complete audio-visual program. Some systems have introduced automated tools to assist in content generation and editing.
[0003] Despite some progress in the aforementioned technologies, significant shortcomings remain in the generation of multi-role digital human podcasts: existing technologies have limited capabilities in comprehensively understanding multimodal content, struggling to effectively process and integrate heterogeneous data from multiple sources such as documents, audio, and images. Furthermore, existing systems lack mechanisms for intelligently orchestrating content into multi-role dialogues, resulting in digital human interactions that lack naturalness, expressiveness, and the ability to capture the emotional nuances and interactive characteristics of real conversations. Moreover, existing technologies consist largely of fragmented functional modules, lacking a complete integrated workflow from content input to video synthesis. Users must navigate multiple systems to complete production, severely impacting content creation efficiency and quality. Specifically:
[0004] (1) Limited multimodal input processing capability: Existing technologies are mostly designed for processing single-modal content (such as plain text or plain audio), lacking the ability to comprehensively understand and integrate multimodal content such as documents, audio, and images, resulting in limited content sources and the inability to fully utilize diverse information.
[0005] (2) Insufficient intelligent content arrangement: When converting the original content into a multi-role dialogue format, the existing system lacks an intelligent content allocation and arrangement mechanism, and cannot make reasonable allocation according to the content characteristics and role settings, resulting in a lack of naturalness and logic in the generated dialogue.
[0006] (3) Simple role interaction mechanism: Most existing multi-role digital human systems are simple turn-by-turn speaking modes, lacking interactive features in real dialogue, such as interruption, supplementation, and rebuttal, which makes the generated content lack realism and appeal.
[0007] (4) Monotonous character performance: Existing digital human performances are mostly limited to lip-syncing and simple expressions, lacking facial expressions, body movements and interactive performances that match the emotions of the content, which affects the viewing experience.
[0008] (5) Low system integration: Existing technologies are mostly scattered functional modules, lacking a complete process integration from multimodal input to final video synthesis. Users need to use multiple systems and tools to complete the entire production process, which is inefficient.
[0009] (6) Insufficient personalization capabilities: The existing system is difficult to adjust flexibly according to different themes, styles and audience needs, and lacks optimization mechanisms for specific content characteristics.
[0010] These problems mainly stem from the fact that existing technologies are mostly breakthroughs in a single aspect, lacking systematic integration and optimization; at the same time, the integration and application of multimodal AI technology and digital human generation technology is still in its early stages, lacking specialized design and optimization for multi-role interaction scenarios. Summary of the Invention
[0011] Purpose of the invention: The purpose of this invention is to provide a method for generating multi-role digital human podcasts based on multimodal AI, which automates the entire process from input of multi-source heterogeneous data to output of interactive videos by multiple characters, and significantly improves the efficiency and quality of digital content creation.
[0012] Technical solution: The present invention provides a method for generating multi-role digital human podcasts based on multimodal AI, comprising the following steps:
[0013] The input multimodal information is converted into text information, semantic understanding is performed on the text information and key information is extracted, and the text information and key information are combined with the program style and content length requirements set by the user to design a structured outline, and node content is generated based on the structured outline.
[0014] Based on key information and user selections, digital human roles are set and assigned, and the structured outline and node content are combined with the digital human role settings to generate dialogue unit scripts and content tags;
[0015] Interaction points are designed and scripts are optimized for dialogue unit scripts. Based on the dialogue unit scripts and character settings, matching voice content is generated for each character, and accurate lip-sync animation sequences are generated based on the voice content, i.e., digital human character lip-sync videos.
[0016] Based on content tags, select or generate background scenes, and place the lip-syncing video of the digital human character into the background scene for scene compositing;
[0017] Intelligently add auxiliary effects and synchronize audio and video;
[0018] The composited scene and added auxiliary effects are then rendered to generate a complete multi-character digital human podcast video, and files in the specified format and resolution are output according to user requirements.
[0019] Furthermore, the input multimodal information is converted into text information, and semantic understanding is performed on the text information to extract key information, including:
[0020] Multimodal information includes documents, audio, and images; preprocessing is performed according to different formats, including text extraction, audio transcription, and image analysis, to convert various types of content into text information;
[0021] Based on large models, semantic understanding of textual information is performed to identify content features such as core themes, key information points, and sentiment tendencies.
[0022] Furthermore, a structured outline is designed, and node content is generated based on the structured outline, including:
[0023] The text information and key information, combined with the user-inputted program style requirements and content length requirements, are organized into prompt words as input to the large model. The large model outputs a structured outline. The key information includes the core theme, key information points, and emotional tendency. The structured outline includes the overall framework of the program and the core theme, key information points, and emotional tendency of each node. The text information, core theme, key information points, program style requirements, content length requirements, and emotional tendency are input into the large model as prompt words, and the node plan is output. Each node plan and text information are input into the large model, and the content of that node is output, which includes a complete description of the key information points.
[0024] Furthermore, digital human roles are set and assigned based on key information and user choices, including:
[0025] Based on the core theme, emotional tone, and user selection, determine the types and number of characters needed for the podcast, and assign a persona, language style, and performance characteristics to each character to match the core theme and target audience.
[0026] Furthermore, the structured outline and node content are combined with digital human role settings to generate dialogue unit scripts and content tags. This includes: firstly, based on the structured outline and node content, each node of the structured outline and its corresponding node content are combined with the digital human role settings as prompt words to input into the large language model to design multi-role dialogue formats, including the opening remarks of the first unit and the closing remarks of the last unit, and the dialogue content is assigned according to the role settings; finally, multiple dialogue unit scripts and content tags are generated.
[0027] Furthermore, the dialogue unit scripts undergo interaction point design and script optimization, including:
[0028] Based on multiple dialogue unit scripts, interactive design was added to generate collaborative performances between characters. Interruption mechanisms, dialogue overlap, response delays, interjections and thought words were added to enhance the tone of rhetorical questions and improve the naturalness and realism of interactions between characters.
[0029] The method for generating lip-sync videos of digital human characters is as follows:
[0030] For each line of dialogue assigned to a digital human character in the dialogue unit script, audio for each line of dialogue is generated using the character's voice.
[0031] The audio file synthesized from each dialogue is then used to generate a digital lip-sync video that matches the lip movements of the corresponding digital human character.
[0032] The specific process for generating each frame is as follows: First, the audio features are extracted using the Hubert model. The audio features are then input into the audio encoder to obtain the audio encoded features. Next, the cropped face image is input into the face image encoder to obtain the face feature map. The audio encoded features and the face feature map are then concatenated. Finally, the concatenated features are input into the lip image generator to obtain an image that matches the audio to the lips. The generated lip image is then pasted back into the original image to obtain the final modified lip-synchronized face image.
[0033] Each generated frame of image is then combined with audio to create a digital lip-sync video.
[0034] Furthermore, based on content tags, a background scene is selected or generated, and the lip-sync video of the digital human character is placed within the background scene for scene compositing; including:
[0035] Based on the lip-sync video of the digital human character and the content tags of each unit in the dialogue unit script, the background image is automatically selected and transition effects are automatically added; and the camera is automatically zoomed in when the digital human is speaking alone and automatically zoomed out when multiple people are discussing, so as to realize intelligent camera switching.
[0036] Another embodiment of the present invention provides a multi-role digital human podcast generation system based on multimodal AI, comprising:
[0037] The data processing unit is used to convert the input multimodal information into text information, perform semantic understanding on the text information and extract key information, combine the text information and key information with the program style and content length requirements set by the user, design a structured outline, and generate node content based on the structured outline.
[0038] The dialogue script production unit is used to set and assign digital human roles based on key information and user selections, and to generate dialogue unit scripts and content tags by combining structured outlines and node content with digital human role settings.
[0039] The digital human character lip-sync video production unit is used to design interaction points and optimize the script of the dialogue unit. Based on the dialogue unit script and character settings, it generates matching voice content for each character and generates an accurate lip-sync animation sequence based on the voice content, i.e., digital human character lip-sync video.
[0040] The scene design and compositing unit is used to select or generate background scenes based on content tags, and place the lip-syncing video of the digital human character in the background scene for scene compositing;
[0041] The audiovisual synchronization unit is used to intelligently add auxiliary effects and synchronize audiovisual content.
[0042] The video rendering and output unit is used to perform final rendering of the composited scene and added auxiliary effects, generate a complete multi-role digital human podcast video, and output files in the specified format and resolution according to user needs.
[0043] An electronic device according to another embodiment of the present invention includes a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor. When the computer program / instructions are executed by the processor, they implement the steps of the multi-role digital human podcast generation method based on multimodal AI.
[0044] In another embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, which, when invoked, are used to execute the steps of the multi-role digital human podcast generation method based on multimodal AI.
[0045] Beneficial effects: Compared with the prior art, the significant technical effects of the present invention are: (1) It greatly improves the efficiency of content creation, realizes the full-process automated generation from multi-source heterogeneous data to multi-role interactive videos, shortens the podcast production process that traditionally requires multiple teams to complete for several days to complete within a few hours, improves the efficiency of creation, and greatly reduces the cost of content production; (2) It supports multiple formats such as documents, audio, and images. Through the collaboration of multimodal encoders and large language models, it realizes the deep understanding and integration of cross-modal content, greatly improving the flexibility and adaptability of content processing; (3) It is not only suitable for traditional podcast production, but can also be widely used in various scenarios such as education and training, corporate publicity, news broadcasting, and product demonstration. It has extremely high versatility and scalability, and provides a powerful tool for digital content innovation in various industries. Attached Figure Description
[0046] Figure 1 This is a flowchart of the method of the present invention;
[0047] Figure 2 Flowchart for multimodal content input and preprocessing;
[0048] Figure 3 Intelligently generate flowcharts for dialogue scripts. Detailed Implementation
[0049] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0050] Example 1:
[0051] like Figure 1 As shown, the present invention provides a multi-role digital human podcast generation method based on multimodal AI. Through key technologies such as multimodal content understanding, intelligent dialogue arrangement, digital human synthesis, and interactive performance generation, it achieves fully automated generation of multi-role interactive videos from multi-source heterogeneous data. Specifically, it includes the following steps:
[0052] Step 1: Multimodal information input, content extraction and structuring. Receive multimodal information (documents, audio, images) uploaded by users, and preprocess it according to different formats, including text extraction, audio transcription, image analysis, etc., to extract and structure various types of content into text information.
[0053] Specific operations: The PaddleOCR text recognition model is used to extract text information from the submitted PDF document content; the FunASR speech recognition model is used to transcribe audio content into text; and the Qwen2.5-VL-72B multimodal large-scale model is used to understand image content and obtain text information. This operation is used to process three types of input content into plain text content (i.e., text information). Figure 2 As shown.
[0054] Step 2: Semantic understanding and key information extraction. The Deepseek-R1 model is called to perform semantic understanding on the text content processed in Step 1, and to identify the content features of core themes, key information points and sentiment.
[0055] Specific operation: Input the text content as a prompt into the large model. The large model performs semantic understanding of the text content and identifies the core theme, key information points and sentiment.
[0056] Step 3: Design a structured outline and generate node content. Based on the text content obtained in Step 1 and the core theme, key information points and emotional tendencies of the text content obtained in Step 2, and according to the program style requirements and content length requirements set by the user, design a structured outline and generate node content based on the structured outline.
[0057] Specific operation: Take the core theme, key information points and emotional tendency of the text content obtained in step 1 and step 2, and combine them with the program style requirements and content length requirements input by the user to organize them into prompt words as input to the large model. The large model outputs a structured outline and node content.
[0058] The structured outline design and node content generation process is as follows:
[0059] 1. Generate a structured outline: Based on the input, generate a structured outline, including the overall framework of the program and the core theme, key information points, and emotional tone of each node.
[0060] Node segmentation: Input the large model with text content, core theme, key information points, program style requirements, content length requirements, and emotional tendency as prompts, and output a node scheme. The output includes: ① number of nodes; ② subtitle of each node; ③ list of key information points covered by the node; ④ suggested word count; ⑤ main emotional tags (e.g., lighthearted / serious / emotional / humorous); ⑥ main style tags (e.g., documentary / popular science / interview / commentary); ⑦ optional modifier tags (e.g., turning point / suspense / summary, 0-1); ⑧ placeholder descriptions (e.g., "humorous scene / comparative case").
[0061] Constraints: Each key information point must appear at least once; the same information point should not be repeatedly assigned to too many nodes (no more than twice); and adjacent nodes should avoid high content overlap.
[0062] 2. Generate node content: After the structured outline is determined, input the obtained ②-⑧ of each node together with the text content obtained in step 1 into the large model, and directly output the content of the node. The content includes a complete description of the key information points (including necessary examples / definitions / comparisons, and short transition sentences between the preceding and following nodes, with natural language forming a coherent paragraph).
[0063] Step 4: Role setting and assignment. Based on the characteristics of the content (core theme, emotional tone) and user selection (i.e., digital human configuration), determine the types and number of roles required for the podcast, and assign suitable personas, language styles and performance characteristics to each role to match the core theme and target audience.
[0064] In this embodiment of the invention, the professional field, background setting, emotional tone, language style and performance characteristics of the digital human character are set by combining the digital human character settings input by the user with the core theme and emotional characteristics of step 2, so as to match the core theme and target audience.
[0065] Step 5: Intelligent generation of dialogue scripts. The structured outline and node content from Step 3 are combined with the digital human role settings from Step 4 to transform into dialogue unit scripts and content tags for multi-role dialogues, including opening remarks and closing remarks.
[0066] In this embodiment of the invention, firstly, based on the structured outline and node content from step 3, each node and its corresponding content in the structured outline, along with the digital human role settings from step 4, are input as prompts into a large language model (such as DeepSeek-R1) to design a multi-role dialogue format. This includes the opening remarks for the first unit, the closing remarks for the last unit, and other parts. Dialogue content is then assigned according to the role settings. Finally, multiple dialogue unit scripts and content tags are generated. The process flow is as follows: Figure 3 As shown, specifically:
[0067] Input the structured outline, node content, and digital human role settings into the large language model;
[0068] The large language model generates dialogue unit scripts and content tags. The dialogue unit scripts include dialogue unit script 1, dialogue unit script 2, ..., dialogue unit script N.
[0069] Dialogue Unit Script 1 generates an opening and multi-role dialogue using a large language model, serving as the first unit.
[0070] Dialogue unit scripts 2 to N-1 generate multi-role dialogues through a large language model, serving as intermediate units.
[0071] The dialogue unit script N generates closing remarks and multi-role dialogues through a large language model, serving as the ending unit.
[0072] Step 6: Interaction point design and script optimization. Based on the multiple dialogue unit scripts generated in Step 5, add interaction design, generate collaborative performances between characters, add interruption mechanisms, dialogue overlap, response delays, add interjections and thinking words, enhance the tone of rhetorical questions, and enhance the naturalness and realism of the interaction between characters.
[0073] In this embodiment of the invention, interruption mechanisms, dialogue overlap, response delays, addition of interjections and thought words, enhancement of rhetorical question tone, and calling of a large model to generate dialogue content for different roles are added to the multiple dialogue unit scripts generated in step 5.
[0074] For each dialogue unit script, the DeepSeek-R1 large model is called to input the dialogue unit script as a prompt word, and the large model outputs a dialogue unit script with interruption mechanism, dialogue overlap, response delay, added tone words and thinking words, enhanced rhetorical question tone, and role-based features.
[0075] Each dialogue unit script can be input as a prompt into the large DeepSeek-R1 model, which will output a dialogue script containing the following:
[0076] Interruption mechanism: The model will automatically identify appropriate interruption opportunities and generate natural dialogue interruptions.
[0077] Overlapping dialogue: Simulates multiple characters speaking or reacting quickly at the same time, enhancing realism.
[0078] Response delay: Add appropriate pauses or thinking time to simulate the character's state of thinking or hesitation.
[0079] Interjections and thought words: Appropriately insert words such as "um," "ah," and "uh" into the dialogue to express the character's emotional fluctuations or thought process.
[0080] Rhetorical question intonation: Enhance the intonation of rhetorical questions through phonetic features and grammatical structure to emphasize emotions and stances.
[0081] Role-based dialogue content: Dialogue content is designed to suit each character's personality, emotions, background, and other factors. Each character's speaking style, tone, and reactions influence the atmosphere and emotional expression of the dialogue. By calling the large model DeepSeek-R1, the settings for different characters are used as prompt inputs, and the output dialogue scripts that match the character's traits are rendered, making the dialogue more layered and emotionally resonant.
[0082] Step 7: Speech synthesis and digital lip-sync video generation. Based on the dialogue unit script and role settings, generate matching speech content for each role, including features such as intonation, speech rate, and stress, and generate accurate lip-sync animation sequences based on the speech content.
[0083] In this embodiment of the invention, the script is subjected to role-specific audio synthesis, and then the synthesized audio file is used to generate the corresponding digital lip-sync video.
[0084] For each dialogue unit script optimized in step 6, role-specific audio files are generated, and then the generated audio files are used to generate digital lip-sync video. The specific steps are as follows:
[0085] 1. For each dialogue in the dialogue unit script, a digital human character is assigned to each dialogue. Using the voice of this character, F5-TTS is used to generate the audio for each dialogue.
[0086] 2. Combine the audio files of each dialogue sentence and generate a digital lip-sync video that matches the lip movements of the corresponding digital human character.
[0087] 3. The specific process of generating each frame: First, the Hubert model is used to extract audio features. The audio features are then input into the audio encoder to obtain audio encoded features. Next, the cropped face image is input into the face image encoder to obtain a face feature map. The audio encoded features and the face feature map are then concatenated. Finally, the concatenated features are input into the lip image generator to obtain an image that matches the audio to the lips. The generated lip image is then pasted back into the original image to obtain the final modified lip-synchronized face image.
[0088] 4. Combine each generated frame with audio to create a digital lip-sync video.
[0089] Step 8: Scene Design and Compositing. Based on the content tags generated for each unit in Step 5, select or generate suitable background scenes, and place each digital lip-sync video within the background scenes. Design appropriate character positions and camera layouts. Select environments that match the content tags from the scene library, or create custom scenes using a generative model. Design camera transitions and scene changes based on the rhythm and emphasis of the dialogue content.
[0090] In this embodiment of the invention, the digital human mouth-printing video generated in step 7 is used, and then a suitable background image is automatically selected based on the content tags of each unit generated in step 5, and transition effects are automatically added; and intelligent camera switching is achieved based on mechanisms such as automatically zooming in when the digital human is speaking alone and automatically zooming out when multiple people are discussing.
[0091] Step 9: Audiovisual synchronization and intelligent effects addition. Intelligently add necessary transition effects, subtitles, background music and other auxiliary effects to enhance visual expressiveness.
[0092] Transition effects: Automatically select transition effects based on the rhythm of camera cuts or changes in audio to avoid abrupt transitions. For example, fast-paced action scenes may use quick transitions, while slow-paced scenes may use gentle fade-in and fade-out effects.
[0093] Subtitles: Automatically add subtitles to each character's dialogue, ensuring precise synchronization between subtitles and audio.
[0094] Background music: Based on the core theme, appropriate background music is generated by calling the Suno API. For example, slow and melancholic music is used for sad scenes, while fast-paced music is used for intense scenes.
[0095] Step 10: Final video rendering and output. The system performs final rendering of all elements (scene compositing in step 8 and special effects in step 9) to generate a complete multi-role digital human podcast video, and outputs files in the specified format and resolution according to user requirements.
[0096] In a specific embodiment of the present invention, a subtitle SRT file is generated based on the text and the generated audio file, and then rendered into the final video using the ffmpeg tool.
[0097] Example 2:
[0098] A multi-role digital human podcast generation system based on multimodal AI includes:
[0099] The data processing unit is used to convert the input multimodal information into text information, perform semantic understanding on the text information and extract key information, combine the text information and key information with the program style and content length requirements set by the user, design a structured outline, and generate node content based on the structured outline.
[0100] The dialogue script production unit is used to set and assign digital human roles based on key information and user selections, and to generate dialogue unit scripts and content tags by combining structured outlines and node content with digital human role settings.
[0101] The digital human character lip-sync video production unit is used to design interaction points and optimize the script of the dialogue unit. Based on the dialogue unit script and character settings, it generates matching voice content for each character and generates an accurate lip-sync animation sequence based on the voice content, i.e., digital human character lip-sync video.
[0102] The scene design and compositing unit is used to select or generate background scenes based on content tags, and place the lip-syncing video of the digital human character in the background scene for scene compositing;
[0103] The audiovisual synchronization unit is used to intelligently add auxiliary effects and synchronize audiovisual content.
[0104] The video rendering and output unit is used to perform final rendering of the composited scene and added auxiliary effects, generate a complete multi-role digital human podcast video, and output files in the specified format and resolution according to user needs.
[0105] Example 3: An electronic device includes a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor, wherein the computer program / instructions, when executed by the processor, implement the steps of the multi-role digital human podcast generation method based on multimodal AI.
[0106] Example 4: A computer-readable storage medium storing computer instructions, which, when invoked, are used to execute the steps of the multi-role digital human podcast generation method based on multimodal AI.
Claims
1. A method for generating multi-role digital human podcasts based on multimodal AI, characterized in that, Includes the following steps: The process involves converting input multimodal information into text, performing semantic understanding on the text, extracting key information, and combining the text and key information with user-defined program style and content length requirements to design a structured outline. Based on this structured outline, node content is generated. This includes: organizing text and key information, along with user-input program style and content length requirements, into prompts as input to the main model; the main model outputting a structured outline; key information including the core theme, key information points, and emotional tendency; the structured outline including the overall program framework and the core theme, key information points, and emotional tendency of each node; inputting text, core theme, key information points, program style requirements, content length requirements, and emotional tendency as prompts into the main model, outputting node schemes; and inputting each node scheme and text information into the main model, outputting the content of that node, including a complete description of the key information points. Digital human roles are set and assigned based on key information and user selection, including: determining the types and number of roles required for the podcast based on the core theme, emotional inclination, and user selection, and assigning a persona, language style, and performance characteristics to each role to match the core theme and target audience; and generating dialogue unit scripts and content tags by combining the structured outline and node content with the digital human role settings. The dialogue unit scripts undergo interaction point design and script optimization, including: adding interactive design based on multiple dialogue unit scripts, generating collaborative performances between characters, adding interruption mechanisms, dialogue overlap, response delays, adding interjections and thought words, enhancing the tone of rhetorical questions, enhancing the naturalness and realism of interactions between characters, and calling a large model to generate dialogue content for each character; based on the dialogue unit scripts and character settings, generating matching voice content for each character, and generating precise lip-sync animation sequences based on the voice content, i.e., digital human character lip-sync videos; Based on content tags, select or generate background scenes, and place the lip-sync video of the digital human character in the background scene for scene compositing; including: automatically selecting background images and automatically adding transition effects based on the lip-sync video of the digital human character and the content tags of each unit in the dialogue unit script; and achieving intelligent camera switching based on the automatic zoom-in mechanism when the digital human is speaking alone and the automatic zoom-out mechanism when multiple people are discussing. Intelligently add auxiliary effects and synchronize audio and video; The composited scene and added auxiliary effects are then rendered to generate a complete multi-character digital human podcast video, and files in the specified format and resolution are output according to user requirements.
2. The method for generating multi-role digital human podcasts based on multimodal AI according to claim 1, characterized in that, The input multimodal information is converted into text information, and semantic understanding and key information are extracted from the text information, including: Multimodal information includes documents, audio, and images; preprocessing is performed according to different formats, including text extraction, audio transcription, and image analysis, to convert various types of content into text information; Based on large models, semantic understanding of textual information is performed to identify content features such as core themes, key information points, and sentiment tendencies.
3. The method for generating multi-role digital human podcasts based on multimodal AI according to claim 1, characterized in that, The process involves combining structured outlines and node content with digital human role settings to generate dialogue unit scripts and content tags. This includes: firstly, based on the structured outlines and node content, adding digital human role settings as prompt words to each node of the structured outline and its corresponding node content, and then using the large language model to design multi-role dialogue formats, including the opening remarks of the first unit and the closing remarks of the last unit, and assigning dialogue content according to role settings; finally, generating multiple dialogue unit scripts and content tags.
4. The method for generating multi-role digital human podcasts based on multimodal AI according to claim 1, characterized in that, The method for generating lip-sync videos of digital human characters is as follows: For each line of dialogue assigned to a digital human character in the dialogue unit script, audio for each line of dialogue is generated using the character's voice. The audio file synthesized from each dialogue is then used to generate a digital lip-sync video that matches the lip movements of the corresponding digital human character. The specific process for generating each frame is as follows: First, the audio features are extracted using the Hubert model. The audio features are then input into the audio encoder to obtain the audio encoded features. Next, the cropped face image is input into the face image encoder to obtain the face feature map. The audio encoded features and the face feature map are then concatenated. Finally, the concatenated features are input into the lip image generator to obtain an image that matches the audio to the lips. The generated lip image is then pasted back into the original image to obtain the final modified lip-synchronized face image. Each generated frame of image is then combined with audio to create a digital lip-sync video.
5. A multi-role digital human podcast generation system based on multimodal AI, characterized in that, include: The data processing unit is used to convert the input multimodal information into text information, perform semantic understanding on the text information and extract key information, combine the text information and key information with the user-defined program style and content length requirements, design a structured outline, and generate node content based on the structured outline. This includes: organizing text information and key information, combined with the user-input program style and content length requirements, into prompt words as input to the large model; the large model outputs a structured outline, where key information includes the core theme, key information points, and emotional tendency; the structured outline includes the overall program framework and the core theme, key information points, and emotional tendency of each node; inputting text information, core theme, key information points, program style requirements, content length requirements, and emotional tendency as prompt words into the large model, and outputting node schemes; inputting each node scheme and text information into the large model, and outputting the content of that node, which includes a complete description of the key information points; The dialogue script production unit is used to set and assign digital human roles based on key information and user selections. This includes: determining the types and number of roles required for the podcast based on the core theme, emotional tone, and user selections, and assigning a persona, language style, and performance characteristics to each role to match the core theme and target audience; and combining the structured outline and node content with the digital human role settings to generate dialogue unit scripts and content tags. The digital human character lip-sync video production unit is used to design interaction points and optimize dialogue unit scripts, including: adding interactive designs based on multiple dialogue unit scripts, generating collaborative performances between characters, adding interruption mechanisms, dialogue overlap, response delays, adding interjections and thought words, enhancing the tone of rhetorical questions, and enhancing the naturalness and realism of interactions between characters; generating matching voice content for each character based on the dialogue unit script and character settings, and generating precise lip-sync animation sequences based on the voice content, i.e., digital human character lip-sync videos; The scene design and compositing unit is used to select or generate background scenes based on content tags, and place the lip-sync video of the digital human character in the background scene for scene compositing; including: automatically selecting background images and automatically adding transition effects based on the lip-sync video of the digital human character and the content tags of each unit in the dialogue unit script; and achieving intelligent camera switching based on the automatic zoom-in mechanism when the digital human is speaking alone and the automatic zoom-out mechanism when multiple people are discussing. The audiovisual synchronization unit is used to intelligently add auxiliary effects and synchronize audiovisual content. The video rendering and output unit is used to perform final rendering of the composited scene and added auxiliary effects, generate a complete multi-role digital human podcast video, and output files in the specified format and resolution according to user needs.
6. An electronic device, characterized in that, It includes a memory, a processor, and a computer program / instructions stored in the memory and executable on the processor, wherein the computer program / instructions, when executed by the processor, implement the steps of the multi-role digital human podcast generation method based on multimodal AI according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when invoked, are used to perform the steps of the multi-role digital human podcast generation method based on multimodal AI as described in any one of claims 1-4.
Citation Information
Patent Citations
Voice dialogue script generation method and device and electronic equipment
CN116312456A
Multi-person dialogue video generation method and device, electronic equipment and storage medium
CN118158453A
Generative question answering method and device based on AI
CN119322831A
Automated video creation-oriented script creation method based on large language model
CN119854597A