Text-to-video full-link generation method and system based on multimodal large model

Through the full-link text-to-video generation method of a multimodal large model, multiple intelligent agents are used to collaboratively build a cross-modal memory library, which solves the problems of low efficiency and poor consistency in video production in existing technologies, realizes efficient and automated video generation, and ensures the unity and immersion of video and audio.

CN120512591BActive Publication Date: 2025-09-23INSPUR QILU SOFTWARE IND
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510991328.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-23
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing video content production relies on manual editing and special effects synthesis, which has problems of low efficiency, high cost, and long time consumption. In addition, existing automated video generation technology has problems such as poor feature consistency, poor emotional logic, difficulty in multimodal synchronization, and high rate of manual intervention.

Method used

A full-link text-to-video generation method based on a multimodal large model is adopted. Through the collaborative work of multiple intelligent agents, a cross-modal memory library is built to realize the automatic generation of the entire process from text to video, ensuring the unity and consistency of the generated video and audio, including steps such as text analysis, storyboard generation, and audio and video synthesis.

Benefits of technology

It achieves narrative coherence in long video generation, improves the feature consistency of storyboards and the consistency of cross-modal emotions, reduces manual intervention, and improves the efficiency of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512591B_ABST
    Figure CN120512591B_ABST
Patent Text Reader

Abstract

The present invention discloses a full-link text-to-video generation method and system based on a multimodal large model, which belongs to the field of artificial intelligence content generation technology. Through the collaborative work of multiple intelligent agents, user input text is analyzed, a cross-modal memory library is constructed, and the unified video and audio of the generated storyboard is ensured based on the memory library content, thereby realizing the full-process automatic generation from text to video. The implementation of this method includes the following steps: obtaining user text input; text analysis, through collaborative agents, dynamically extracting, analyzing, generating, associating, and storing multimodal information of images, texts, and audios from the input text to construct a multimodal memory library; generating storyboards, generating storyboard videos and audios according to the memory library; audio and video synthesis, and forming the final video after synchronous alignment of audio and video. The present invention can achieve narrative coherence in long video generation, improve the feature consistency of storyboards, enhance the consistency of cross-modal emotions, reduce manual intervention, and improve the efficiency of video production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence content generation technology, and specifically to a full-link text-to-video generation method and system based on a multimodal large model. Background Art

[0002] Current video content production relies primarily on manual editing and special effects synthesis, which is inefficient, costly, and time-consuming. Existing automated video generation technology has gone through the following stages:

[0003] Template splicing: matches text keywords based on preset templates, but has poor flexibility and is difficult to adapt to complex narrative needs;

[0004] Single-modal generation: Relying on a single model (such as GAN) to generate video clips, this can lead to problems such as audio and video separation, awkward emotional expression, short video generation time, character breakdown, and choppy movement.

[0005] Multi-model serialization: different AI models are called in stages (script generation → keyframe generation → video generation → smoothing mechanism). This method enhances the narrative coherence of multiple shots, but due to the lack of a coordination mechanism, style discontinuities are prone to occur, such as the break in space-time continuity, cross-shot character drift, and difficulty maintaining the consistency of character appearance when switching between multiple shots.

[0006] With the development of large multimodal models, there is an urgent need to build an end-to-end generation framework to solve the problems of poor feature consistency, poor emotional logic, difficulty in multimodal synchronization, and high rate of manual intervention in the above methods. Summary of the Invention

[0007] The technical task of the present invention is to provide a full-link text-to-video generation method and system based on a multimodal large model, which can realize the full-process automatic generation of text to video, achieve narrative coherence in long video generation, improve the feature consistency of storyboards, enhance cross-modal emotional consistency, reduce manual intervention, and improve the efficiency of video production.

[0008] The technical solution adopted by the present invention to solve its technical problem is:

[0009] A full-link text-to-video generation method based on a multimodal large model uses multiple agents to collaborate, analyze user input text, build a cross-modal memory library, and ensure the consistency of video and audio generated storyboards based on the memory library content, thus achieving full-process automatic generation from text to video. Implementation of this method includes the following steps:

[0010] Step 1: Get user text input;

[0011] Step 2: Text analysis: Through collaborative agents, we dynamically extract, analyze, generate, associate, and store multimodal information from the input text, building a structured multimodal memory library that guides subsequent video generation. The content in the memory library includes images, sounds, and text.

[0012] Step 3: Generate storyboards. Generate storyboard video and audio based on the memory library. When generating storyboard audio and video, use emotional cues to simultaneously and synchronously guide video generation and speech synthesis. Ensure that visual expression (such as character expression, action rhythm, and image color) is highly consistent with speech emotion (intonation and speed) to enhance immersion.

[0013] Step 4: Audio and video synthesis, after the audio and video are synchronized and aligned, the final video is formed.

[0014] Furthermore, the images in the memory library include: image style reference images, reference images of scenes, characters, props, costumes, and other entities; the sounds include: sample audio of narration and other characters’ timbre; the text includes: text type, style, theme, descriptions and features of scenes, characters, props, costumes, and other entities, and timbre features of narration and other characters;

[0015] The text analysis process involves inputting user text into the agent to analyze the text type, style, and theme, and then determining the video style and narration timbre. The process also involves analyzing entities that appear in the text, analyzing the characteristics of each entity, and then creating or selecting its reference image and timbre and storing them in a memory library. Specifically, this process includes:

[0016] Text metadata analysis: Agent 1 analyzes the user text's type (novel / drama / essay), style (plain / heroic / humorous), and theme (emotional / historical / science fiction);

[0017] Stylized resource generation: Agent 2 generates image style prompts and, through Agent 6, generates reference images (automatically stored in the image style library); Agent 3 matches the voiceover tone (professional broadcast voice / energetic youth voice, etc.);

[0018] Entity feature extraction and enhancement: Agent 4 extracts entities such as scenes, characters, props, or clothing. Agent 5 generates text descriptions, features, and classification prompt word templates for each entity (customizing image generation instructions based on entity type). Agent 6 generates entity reference images and stores them in the design library. Agent 7 implements intelligent character-clothing synthesis (input character image + clothing image → output dressed character image).

[0019] Global style unification preprocessing: Agent8 performs style transfer on entity images (input style reference image + entity image → output unified style entity image).

[0020] Furthermore, the text analysis is specifically implemented in the following steps:

[0021] (2.1) Through agent 1, input the user input text and obtain the text type (novel, drama, prose...), style (plain, heroic, humorous...) and theme (emotional, historical, science fiction, European, historical...) information;

[0022] (2.2) Automatically select appropriate image style reference images from the image style library based on the text type, style, and theme. If there is no corresponding style, Agent 2 inputs the text type, style, theme, and user input text to generate image style reference image prompt words (including style, theme, image quality description, artistic style, composition, light / tone, environment description, and texture). Agent 6 then generates image style reference images and stores them in the image style library.

[0023] (2.3) Through Agent 3, input the text type, style, and theme to obtain appropriate narration voice characteristics (professional broadcast voice, energetic youth voice, etc.), and select the appropriate voice from the voice library;

[0024] (2.4) Through agent 4, input the user text input, obtain all scenes, characters, props, costumes, and other entities that appear in the text input, and output them in JSON format;

[0025] (2.5) For each entity output by agent 4, agent 5 takes the user text input, text style, and entity name, obtains the corresponding description of the entity in the text, entity features, and image to generate prompt words. For each type of entity, different prompt words are used to improve the generation effect;

[0026] (2.6) For each entity output by agent4, the corresponding entity reference image is obtained from the design library based on the entity features. If there is no corresponding entity, agent6 is used to input the image generated by agenet5 to generate a prompt word, obtain the corresponding reference image and add it to the design library;

[0027] (2.7) For each "person" entity output by agent 4, input the person reference image obtained in step (2.6) and all "clothing" reference images through agent 7 to obtain the person reference image wearing the specified clothing;

[0028] (2.8) For each entity reference image generated in step (2.6) and the person reference image generated in step (2.7), input the image style reference image and the entity reference image through agent8 to obtain entity images with a unified style after image style transfer; ensure that the visual style of all entity reference images is consistent;

[0029] (2.9) For the entities with “lines” output by agent4, agent3 inputs the text style and entity features, obtains the appropriate dubbing timbre features, and selects the appropriate timbre from the timbre library.

[0030] Furthermore, the storyboard is generated by inputting the user text into the storyboard parsing agent to obtain a storyboard list and a list of entities appearing in the storyboard, and then performing the following operations on each storyboard:

[0031] Generate the storyboard prompt words, video prompt words, and emotional prompt words for the storyboard through the storyboard description and the feature description of the entity in the memory library. Then, generate the storyboard through the prompt words and the image style reference pictures in the memory library and the reference pictures of the entity, and then generate the storyboard video. Then, generate the audio used for the storyboard through the timbre and emotional prompt words corresponding to the entity in the memory library. This includes:

[0032] Storyboard structured analysis: Agent9 divides shots into scenes / duration / scenes, and outputs: shot descriptions and entity lists;

[0033] Storyboard resource binding: automatically associate style reference images, entity images, and timbre characteristics in the memory library;

[0034] Storyboard content generation: Agent 10 generates three key prompts: storyboard prompts, video prompts, and emotional prompts. Agent 11 generates storyboard reference images (by inputting a physical image and prompts). Agent 8 performs secondary style transfer on the storyboard reference images to ensure a consistent style across multiple shots.

[0035] Emotion-driven audio and video synchronization generation: Agent 12 generates storyboard videos (input storyboard images + video prompts + emotion prompts); Agent 13 generates character dubbing (input timbre + lines + the same emotion prompts).

[0036] Furthermore, the specific implementation steps of step three are as follows:

[0037] (3.1) Through Agent 9, the user input text and the text type, style, and theme in the memory library are input, the narrative structure is analyzed, the scenes or shots are divided, and a storyboard list containing shot descriptions (shot number, duration, setting, image content, shot size, perspective, camera movement, dialogue, background sound effects, etc.) is generated. The background, characters, costumes, props, and other entities that appear in the storyboard are counted.

[0038] (3.2) For each storyboard, obtain the image style reference image and the background, characters, costumes, props and other entity list information appearing in the storyboard from the memory library; then execute steps (3.3) to (3.7);

[0039] (3.3) Through agent 10, input the storyboard description and the entity information description appearing in the storyboard to obtain the storyboard image prompt word, storyboard video prompt word, and the emotional prompt word of each dialogue;

[0040] (3.4) For each storyboard, input the storyboard prompt words, background image, character image, and other entity images through agent 11 to obtain the storyboard reference image;

[0041] (3.5) For each storyboard, Agent8 inputs the image style reference image and the storyboard reference image. Using image style transfer technology, a storyboard with a unified style is obtained. This ensures that all storyboards have a consistent visual style and solves the problem of inconsistent styles across multiple shots.

[0042] (3.6) For each storyboard, input the storyboard image, storyboard video prompt words, and emotional prompt words through agent 12 to obtain the storyboard video;

[0043] (3.7) For each line of dialogue, the character’s corresponding timbre, lines, and emotional cues are input through agent 13 to obtain the dialogue audio. The same emotional cues are used for audio and video generation to ensure consistency between the visual representation and the speech emotion.

[0044] Furthermore, the audio and video synthesis includes:

[0045] Timeline framework construction: Create a video draft framework through Agent14;

[0046] Audio and video alignment: Agent15 analyzes video content and outputs precise timestamps (in milliseconds) of dialogue or sound effects.

[0047] Multi-track synthesis: Use Agent14 to insert storyboard videos into tracks according to timestamps, and inject dubbing audio and subtitles synchronously (based on Agent15 time data);

[0048] Adaptive background music: Agent16 matches the mood of the storyboard and recommends background music and time intervals; Agent14 injects the background music into the soundtrack.

[0049] Furthermore, the specific implementation steps of step 4 are as follows:

[0050] (4.1) Create a draft: Create a video draft through agent 14, and then perform steps (4.2) to (4.3) for each storyboard;

[0051] (4.2) Action audio matching: For each storyboard, agent 15 is used to input the storyboard shot description and storyboard video to determine the start and end time of the dialogue and background sound effects;

[0052] (4.3) Insert video, audio, and text into the timeline. Through agent 14, input video, audio, text, and corresponding start and end times, insert the video into the timeline, and then insert the audio and subtitles into the timeline according to the dialogue start time offset;

[0053] (4.4) Select background music. Through agent16, input the storyboard list, select the appropriate background music list from the background music library and provide the corresponding start and end time of the background music;

[0054] (4.5) Insert the background music into the timeline. Through agent 14, input the background music and the corresponding start and end time of the background music, and insert the background music into the timeline of the video draft;

[0055] (4.6) Rendering output.

[0056] The present invention also claims protection for a full-link text-to-video generation system based on a multimodal large model, comprising:

[0057] Text acquisition module, used to obtain user text input;

[0058] The text analysis module is used to dynamically extract, analyze, generate, associate, and store multimodal information (text, audio, and images) from input text through collaborative agents, creating a cross-modal memory library.

[0059] A storyboard generation module is used to generate storyboard videos and audios based on a memory library;

[0060] The audio and video synthesis module is used to synchronize the audio and video to form the final video;

[0061] The system realizes the full-link generation of text to video through the above method.

[0062] The present invention also claims protection for a text-to-video full-link generation device based on a multimodal large model, comprising: at least one memory and at least one processor;

[0063] The at least one memory is configured to store a machine-readable program;

[0064] The at least one processor is configured to call the machine-readable program to implement the above method.

[0065] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, can implement the above method.

[0066] Compared with the existing technology, the text-to-video full-link generation method and system based on a multimodal large model of the present invention has the following beneficial effects:

[0067] This invention constructs a closed-loop cross-modal memory library architecture for the first time. Through the specialized division of labor and collaboration of 16 intelligent agents, it realizes end-to-end automatic generation from text to high-quality video. By establishing a persistent cross-modal memory library of scenery, props, costumes, characters, etc., it achieves narrative coherence in long video generation, improves the feature consistency of storyboards, enhances cross-modal emotional consistency, reduces manual intervention, and improves the efficiency of video production. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a flowchart of a method for generating a full-link text-to-video link based on a multimodal large model provided by an embodiment of the present invention;

[0069] Figure 2 is a flowchart of building a memory library provided by an embodiment of the present invention;

[0070] Figure 3 is a flowchart of generating a video from text provided by an embodiment of the present invention;

[0071] Figure 4 It is a diagram of the intelligent agent collaborative workflow provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The present invention will be further described below with reference to specific embodiments.

[0073] An embodiment of the present invention provides a full-link text-to-video generation method based on a multimodal large model. Through the collaborative work of multiple intelligent agents, user input text is analyzed, a cross-modal memory library is constructed, and the unity of the generated storyboard video and audio is ensured based on the content of the memory library, thereby realizing the automatic generation of the entire process from text to video.

[0074] This method achieves element reuse and style unification by dynamically building a cross-modal memory library (image / audio / text). It consists of three core stages: cross-modal memory library construction, storyboard generation and emotional coordination, and audio and video synthesis. The entire process is carried out by 16 specialized agents, forming a closed-loop workflow. The details are as follows:

[0075] Step 1: Get user text input.

[0076] Step 2: Text Analysis, Dynamically Creating a Cross-Modal Memory Bank. User text is input into the agent to analyze the text type, style, and theme, thereby determining the video style and voiceover tone. The agent also analyzes entities appearing in the text, analyzes the characteristics of each entity, and then creates or selects reference images and tones for each entity and stores them in the memory bank.

[0077] The contents of the memory library include pictures (reference pictures of image style, reference pictures of scenes, characters, props, costumes, and other entities), sounds (audio samples of narration and other characters’ tones), and text (text type, style, theme, descriptions and features of scenes, characters, props, costumes, and other entities, and tonal features of narration and other characters). Specifically, it includes:

[0078] (1) Text metadata analysis: Agent 1 analyzes the type (novel / drama / prose), style (plain / heroic / humorous), and theme (emotional / historical / science fiction) of the user's text.

[0079] (2) Stylized resource generation: Agent 2 generates image style prompts and generates reference images through Agent 6 (automatically stored in the image style library); Agent 3 matches the narration tone (professional broadcast voice / energetic youth voice, etc.).

[0080] (3) Entity feature extraction and enhancement: Agent 4 extracts entities including scenes, characters, props, or clothing; Agent 5 generates text descriptions, features, and classification prompt word templates for each entity (customizing image generation instructions according to entity type); Agent 6 generates entity reference images and stores them in the design library; Agent 7 implements intelligent character-clothing synthesis (input character image + clothing image → output dressed character image).

[0081] (4) Global style unification preprocessing: Agent8 performs style transfer on the entity image (input style reference image + entity image → output unified style entity image).

[0082] Step 3: Generate storyboards. Generate visually / audio-consistent storyboards based on the memory library. Input the user text into the storyboard parsing agent to obtain a list of storyboards and a list of entities that appear in the storyboards. Then, perform the following operations on each storyboard:

[0083] The storyboard description and the entity feature description in the memory library are used to generate the storyboard prompt words, video prompt words, and emotional prompt words for the storyboard. Then, the storyboard is generated using the prompt words and the image style reference pictures in the memory library and the reference pictures of the entities that appear. Then, the storyboard video is generated. The audio used for the storyboard is then generated using the timbre and emotional prompt words corresponding to the entities in the memory library. Specifically, this includes:

[0084] (1) Storyboard structured analysis: Agent9 is used to divide shots (scenes / duration / scenes), and outputs: shot description and entity list.

[0085] (2) Storyboard resource binding: automatically associate style reference images, entity images, and timbre features in the memory library.

[0086] (3) Storyboard content generation: Agent 10 generates three key elements: storyboard image prompt, video prompt, and emotional prompt. Agent 11 generates a storyboard reference image (input entity image + prompt). Agent 8 performs secondary style transfer on the storyboard reference image to ensure a consistent style across multiple shots.

[0087] (4) Emotion-driven audio and video synchronization generation: Agent 12 generates storyboard videos (input storyboard images + video prompt words + emotion prompt words); Agent 13 generates character dubbing (input timbre + lines + the same emotion prompt words).

[0088] Step 4: Audio and video synthesis, aligning the audio and video to form the final video. This includes:

[0089] (1) Timeline architecture construction: Create a video draft framework through Agent14.

[0090] (2) Audio and video alignment: Agent15 analyzes the video content and outputs the precise timestamp (millisecond level) of the dialogue or sound effects.

[0091] (3) Multi-track synthesis: Use Agent14 to insert the storyboard video into the track according to the timestamp, and inject the dubbing audio and subtitles synchronously (based on the time data of Agent15).

[0092] (4) Adaptive background music: Agent 16 matches the mood of the storyboard and recommends background music and time intervals; Agent 14 injects the background music into the soundtrack.

[0093] Through the specialized division of labor and collaboration of 16 intelligent agents, end-to-end automated generation from text to high-quality videos is achieved. By establishing a persistent cross-modal memory library of sets, props, costumes, characters, etc., narrative coherence in long video generation is achieved, feature consistency of storyboards is improved, cross-modal emotional consistency is enhanced, manual intervention is reduced, and the efficiency of video production is improved.

[0094] The implementation of this method will be further described in detail below with reference to the accompanying drawings.

[0095] like Figure 1 The figure shows the overall process of this method.

[0096] Step 1: Get user text input.

[0097] Step 2: Text analysis to create a cross-modal memory library. This memory library includes images (reference images for image style, scenes, characters, props, costumes, and other entities), sounds (audio examples of narration and other characters' tones), and text (text type, style, and theme; descriptions and features of scenes, characters, props, costumes, and other entities; and tonal features of narration and other characters).

[0098] The specific steps are as follows:

[0099] 2.1. Through agent 1, input the user input text and obtain the text type (novel, drama, prose...), style (plain, heroic, humorous...) and theme (emotional, historical, science fiction, European, historical...) information.

[0100] 2.2. Automatically select appropriate image style reference images from the image style library based on the text type, style, and theme. If there is no corresponding style, agent 2 inputs the text type, style, theme, and user input text to generate image style reference image prompt words (including style, theme, image quality description, artistic style, composition, light / tone, environment description, and texture). Agent 6 then generates image style reference images and stores them in the image style library.

[0101] 2.3. Through Agent 3, input the text type, style, and theme to obtain the appropriate narration tone characteristics (professional broadcast voice, energetic youth voice, etc.), and select the appropriate tone from the tone library.

[0102] 2.4. Through agent4, input user text input, obtain all scenes, characters, props, costumes, and other entities that appear in the text input, and output them in JSON format.

[0103] 2.5. For each entity output by agent 4, agent 5 inputs the user text input, text style, and entity name to obtain the corresponding description of the entity in the text, entity features, and image to generate prompt words. For each type of entity, different prompt words are used to improve the generation effect.

[0104] 2.6. For each entity output by agent4, the corresponding entity reference image is obtained from the design library based on the entity features. If there is no corresponding entity, agent6 is used to input the image generated by agenet5 to generate a prompt word, obtain the corresponding reference image and add it to the design library.

[0105] 2.7. For each "person" entity output by agent 4, input the person reference image obtained in step (2.6) and all "clothing" reference images through agent 7 to obtain the person reference image wearing the specified clothing.

[0106] 2.8. For each entity reference image generated in step (2.6) and the person reference image generated in step (2.7), input the image style reference image and the entity reference image through agent8 to obtain entity images with a unified style after image style transfer; ensure that the visual style of all entity reference images is consistent.

[0107] 2.9. For the entities with "lines" output by agent4, agent3 inputs the text style and entity features, obtains the appropriate dubbing timbre features, and selects the appropriate timbre from the timbre library.

[0108] Step 3: Generate storyboards and generate storyboard videos and audio based on the memory library. The specific steps are as follows:

[0109] 3.1. Through Agent9, user input text and text type, style, and theme in the memory library are input, narrative structure is analyzed, scenes or shots are divided, and a storyboard list containing shot descriptions (shot number, duration, scene, image content, shot size, perspective, camera movement, dialogue, background sound effects, etc.) is generated. The background, characters, costumes, props, and other entities that appear in the storyboard are counted.

[0110] 3.2. For each storyboard, obtain the image style reference image and the background, characters, costumes, props and other entity list information that appear in the storyboard from the memory library; then execute steps 3.3 to 3.7.

[0111] 3.3. Through agent 10, input the storyboard description and the entity information description appearing in the storyboard to obtain the storyboard image prompt word, storyboard video prompt word, and the emotional prompt word of each dialogue.

[0112] 3.4. For each storyboard, input the storyboard prompt words, background image, character image, and other entity images through agent 11 to obtain the storyboard reference image.

[0113] 3.5. For each storyboard, Agent8 inputs the image style reference image and the storyboard reference image. Using image style transfer technology, a storyboard with a unified style is obtained to ensure the visual style of all storyboards is consistent and solve the problem of inconsistent styles among multiple shots.

[0114] 3.6. For each storyboard, input the storyboard image, storyboard video prompt words, and emotional prompt words through agent 12 to obtain the storyboard video.

[0115] 3.7. For each line of dialogue, agent 13 is used to input the character's corresponding timbre, lines, and emotional cues to obtain the dialogue audio. The same emotional cues are used for audio and video generation to ensure consistency between the visual expression and the speech emotion.

[0116] Step 4: Synthesize audio and video to form the final video.

[0117] 4.1. Create a draft: Create a video draft through agent14, and then perform steps 4.2 to 4.3 for each storyboard.

[0118] 4.2. Action audio matching. For each storyboard, the storyboard shot description and storyboard video are input through agent15 to determine the start and end time of the dialogue and background sound effects.

[0119] 4.3. Insert video, audio, and text into the timeline. Through agent 14, input video, audio, text, and corresponding start and end times, insert the video into the timeline, and then insert audio and subtitles into the timeline according to the dialogue start time offset.

[0120] 4.4. Select background music. Through agent16, input the storyboard list, select the appropriate background music list from the background music library and provide the corresponding start and end time of the background music.

[0121] 4.5. Insert the background music into the timeline. Through agent14, input the background music and the corresponding start and end time of the background music, and insert the background music into the timeline of the video draft.

[0122] 4.6. Rendering output.

[0123] Among them, the list of agents and their functions used are as follows:

[0124] ‌agent1: text analysis agent‌.

[0125] Function: Analyze text type, style and theme;

[0126] Input: User enters text;

[0127] Output: text type (novel / drama / essay), style (plain / heroic / humorous), theme (emotional / historical / science fiction).

[0128] ‌agent2: Image style cue word generation agent‌.

[0129] Function: Generate image style reference prompt words;

[0130] Input: text type, style, theme, and user input text;

[0131] Output: Image style reference picture prompt words.

[0132] ‌agent3: timbre selection agent‌.

[0133] Function: Responsible for the identification and selection of timbre characteristics of narration and character dubbing;

[0134] Input: text type, style, theme, or text style + entity features;

[0135] Output: Tone characteristics (professional broadcast voice / energetic youth voice, etc.).

[0136] ‌agent4: Entity extraction agent‌.

[0137] Function: Extract entity information from text;

[0138] Input: user text input;

[0139] Output: A list of entities such as scenes, characters, props, and costumes in JSON format.

[0140] ‌agent5: Entity feature extraction agent‌.

[0141] Function: Generate entity description and image prompt words;

[0142] Input: user text input, text style, entity name;

[0143] Output: entity description, features, and image generation prompts;

[0144] ‌agent6: Image generation agent‌.

[0145] Function: Generate reference pictures;

[0146] Input: picture to generate prompt words;

[0147] Output: Reference images (stored in the image style library and design library).

[0148] ‌agent7: Character clothing synthesis agent‌.

[0149] Function: Synthesize a picture of a person wearing specified clothing;

[0150] Input: character reference picture + clothing reference picture;

[0151] Output: Reference image of the dressed character.

[0152] ‌agent8: Image style transfer agent‌.

[0153] Function: Apply image style transfer algorithm to unify image style;

[0154] Input: image style reference picture + input picture;

[0155] Output: The image after style transfer.

[0156] ‌agent9: Storyboard analysis agent‌.

[0157] Function: Analyze storyboard elements;

[0158] Input: user input text and text features in the memory bank;

[0159] Output: Storyboard list and entity list.

[0160] ‌agent10: An agent that generates storyboard prompts‌.

[0161] Function: Generate storyboard production prompts;

[0162] Input: Storyboard description + entity features in the memory bank;

[0163] Output: storyboard prompt words, video prompt words, and emotional prompt words.

[0164] ‌agent11: Storyboard generation agent‌.

[0165] Function: Generate storyboard reference pictures;

[0166] Input: storyboard prompt word + entity picture;

[0167] Output: Storyboard reference.

[0168] ‌agent12: Storyboard video generation agent‌.

[0169] Function: Generate storyboard video;

[0170] Input: storyboard + storyboard video prompt words;

[0171] Output: Storyboard video.

[0172] ‌agent13: Conversational audio synthesis agent.

[0173] Function: Generate dubbing audio;

[0174] Input: timbre features + lines + emotional prompts;

[0175] Output: Conversation audio file.

[0176] ‌Agent14, the core agent for video editing.

[0177] Function: Responsible for video editing, such as creating drafts, adding videos, adding audio, adding subtitles, adding special effects, etc.

[0178] ‌Input: video editing specific operations;

[0179] Output: Edit the video draft accordingly.

[0180] ‌Agent15, audio and video synchronization agent.

[0181] ‌Function‌: Analyze the timing alignment of video content and audio;

[0182] Input: storyboard description text + video clip;

[0183] Output: Structured temporal data (with precise start and end timestamps of dialogue / sound effects).

[0184] ‌Agent16, music recommendation agent.

[0185] ‌Function‌: Match background music based on the scene;

[0186] Input: Shot list (including emotional labels / pacing requirements);

[0187] Output: recommended music list + suitable time interval suggestions.

[0188] This method uses a series of dedicated, collaborative agents (Agent1-Agent9) to dynamically extract, analyze, generate, associate, and store multimodal information including text, image, and audio from the input text, thereby constructing a structured multimodal memory library for guiding subsequent video generation.

[0189] When generating storyboard audio and video, emotional cues are used simultaneously and synchronously to guide video generation (Agent 12) and speech synthesis (Agent 13), ensuring a high degree of consistency between visual expression (such as character expressions, action rhythm, and screen color) and speech emotion (intonation and speaking speed), thereby enhancing immersion.

[0190] This method uses multiple dedicated agents to achieve fine-grained automated decision-making (such as style selection, entity description extraction, prompt word generation, image generation / transfer, and timbre selection), and proposes a modular, flexibly scalable automation architecture based on multi-agent collaboration.

[0191] This method uses image style transfer technology to improve the feature consistency of storyboards, obtains text features through text analysis, and then selects or generates image style reference images. Through the image style reference images, image style transfer technology is used to ensure that the image style of each storyboard is consistent.

[0192] An embodiment of the present invention also provides a text-to-video full-link generation system based on a multimodal large model, which realizes text-to-video full-link generation through the text-to-video full-link generation method based on a multimodal large model described in the above embodiment.

[0193] The system includes:

[0194] 1. Text acquisition module, used to obtain user text input.

[0195] 2. The text analysis module is used to dynamically extract, analyze, generate, associate, and store multimodal information including text, image, and audio from the input text through collaborative agents, and to create a cross-modal memory library.

[0196] The user text is input into the intelligent agent to analyze the text type, style, and theme, and then determine the video style and narration tone; analyze the entities appearing in the text, analyze the characteristics of each entity, and then create or select its reference image and tone and store them in the memory library.

[0197] The contents of the memory library include pictures (reference pictures of image style, reference pictures of scenes, characters, props, costumes, and other entities), sounds (audio samples of narration and other characters’ tones), and text (text type, style, theme, descriptions and features of scenes, characters, props, costumes, and other entities, and tonal features of narration and other characters). Specifically, it includes:

[0198] (2.1) Text metadata analysis: Agent 1 analyzes the user text’s type (novel / drama / essay), style (plain / heroic / humorous), and theme (emotional / historical / science fiction).

[0199] (2.2) Stylized Resource Generation: Agent 2 generates image style prompts and, through Agent 6, generates reference images (automatically stored in the image style library); Agent 3 matches the voiceover timbre (professional broadcast voice / energized youth voice, etc.).

[0200] (2.3) Entity Feature Extraction and Enhancement: Agent 4 extracts entities such as scenes, characters, props, or clothing. Agent 5 generates text descriptions, features, and classification prompt word templates for each entity (customizing image generation instructions by entity type). Agent 6 generates entity reference images and stores them in the design library. Agent 7 implements intelligent character-clothing synthesis (input character image + clothing image → output dressed character image).

[0201] (2.4) Global style unification preprocessing: Agent8 performs style transfer on the entity image (input style reference image + entity image → output unified style entity image).

[0202] 3. Generate storyboard module, which is used to generate storyboard video and audio based on the memory library.

[0203] The storyboard generation module generates visually / audio-consistent storyboards based on the memory library. The user text is input into the storyboard parsing agent to obtain a list of storyboards and a list of entities that appear in the storyboards. The following operations are then performed on each storyboard:

[0204] The storyboard description and the entity feature description in the memory library are used to generate the storyboard prompt words, video prompt words, and emotional prompt words for the storyboard. Then, the storyboard is generated using the prompt words and the image style reference pictures in the memory library and the reference pictures of the entities that appear. Then, the storyboard video is generated. The audio used for the storyboard is then generated using the timbre and emotional prompt words corresponding to the entities in the memory library. Specifically, this includes:

[0205] (3.1) Storyboard structured analysis: Agent9 is used to divide shots (scenes / duration / scenes), and outputs: shot description and entity list.

[0206] (3.2) Storyboard resource binding: Automatically associate style reference images, entity images, and timbre characteristics in the memory library.

[0207] (3.3) Storyboard Content Generation: Agent 10 generates three key prompts: storyboard image prompt, video prompt, and emotional prompt. Agent 11 generates a storyboard reference image (input: physical image + prompt). Agent 8 performs secondary style transfer on the storyboard reference image to ensure a consistent style across multiple shots.

[0208] (3.4) Emotion-driven audio and video synchronization generation: Agent 12 generates a storyboard video (input storyboard image + video prompt word + emotion prompt word); Agent 13 generates the character dubbing (input timbre + lines + the same emotion prompt word).

[0209] 4. Audio and video synthesis module, used to synchronize audio and video to form the final video. Specifically includes:

[0210] (4.1) Timeline architecture construction: Create a video draft framework through Agent14.

[0211] (4.2) Audio and video alignment: Agent15 analyzes the video content and outputs the precise timestamp (in milliseconds) of the dialogue or sound effects.

[0212] (4.3) Multi-track synthesis: Use Agent14 to insert the storyboard video into the track according to the timestamp, and inject the dubbing audio and subtitles synchronously (based on the time data of Agent15).

[0213] (4.4) Adaptive Background Music: Agent 16 matches the mood of the storyboard and recommends background music and time intervals; Agent 14 injects the background music into the soundtrack.

[0214] An embodiment of the present invention further provides a text-to-video full-link generation device based on a multimodal large model, comprising: at least one memory and at least one processor;

[0215] The at least one memory is configured to store a machine-readable program;

[0216] The at least one processor is used to call the machine-readable program to implement the full-link text-to-video generation method based on a multimodal large model described in the above embodiment.

[0217] An embodiment of the present invention further provides a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions cause the processor to execute the method for generating a full-link text-to-video image based on a multimodal large model as described in the above-described embodiment. Specifically, a system or device equipped with a storage medium can be provided. The storage medium stores software program code that implements the functions of any of the above-described embodiments, and causes a computer (or CPU or MPU) of the system or device to read and execute the program code stored in the storage medium.

[0218] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.

[0219] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, and DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.

[0220] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.

[0221] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.

[0222] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.

Claims

1. A full-link text-to-video generation method based on a multimodal large model, characterized by: Through the collaborative work of multiple intelligent agents, user input text is analyzed, a cross-modal memory library is constructed, and the video and audio of the generated storyboards are unified based on the memory library content, thus realizing the automatic generation of the entire process from text to video. The implementation of this method includes the following steps: Step 1: Get user text input; Step 2: Text analysis: Through collaborative agents, we dynamically extract, analyze, generate, associate, and store multimodal information from the input text, building a structured multimodal memory library that guides subsequent video generation. The content in the memory library includes images, sounds, and text. Step 3: Generate storyboards, generating storyboard video and audio based on the memory library; when generating storyboard audio and video, use the emotional prompt words to simultaneously and synchronously guide video generation and speech synthesis; Step 4: Audio and video synthesis, after the audio and video are synchronized and aligned, the final video is formed; The text analysis is specifically implemented in the following steps: (2.1) Through agent 1, input the user input text and obtain the type, style and theme information of the text; (2.2) Automatically select image style reference images from the image style library based on the text type, style, and theme. If there is no corresponding style, agent 2 inputs the text type, style, theme, and user input text to generate image style reference image prompt words, and agent 6 generates image style reference images and stores them in the image style library; (2.3) Through agent 3, input text type, style, and theme, obtain narration timbre characteristics, and select timbre from the timbre library; (2.4) Through agent 4, input the user text input, obtain all scenes, characters, props, costumes, and other entities that appear in the text input, and output them in JSON format; (2.5) For each entity output by agent 4, agent 5 takes the user text input, text style, and entity name, obtains the corresponding description in the text, entity features, and image, and generates prompt words. Different prompt words are used for each type of entity. (2.6) For each entity output by agent4, the corresponding entity reference image is obtained from the design library based on the entity features. If there is no corresponding entity, agent6 is used to input the image generated by agenet5 to generate a prompt word, obtain the corresponding reference image and add it to the design library; (2.7) For each "person" entity output by agent 4, input the person reference image obtained in step (2.6) and all "clothing" reference images through agent 7 to obtain a reference image of the person wearing the specified clothing; (2.8) For each entity reference image generated in step (2.6) and the person reference image generated in step (2.7), input the image style reference image and the entity reference image through agent8 to obtain the entity image with a unified style after image style transfer; (2.9) For the entity with "lines" output by agent4, agent3 inputs the text style and entity features, obtains the voiceover timbre features, and selects a timbre from the timbre library; The specific implementation steps of step three are as follows: (3.1) Through Agent 9, the user input text and the text type, style, and theme in the memory library are input, the narrative structure is analyzed, the scenes or shots are divided, a storyboard list containing shot descriptions is generated, and a list of backgrounds, characters, costumes, props, and other entities that appear in the storyboard is counted; (3.2) For each storyboard, obtain the image style reference image and the background, characters, costumes, props and other entity list information appearing in the storyboard from the memory library; then execute steps (3.3) to (3.7); (3.3) Through agent 10, input the storyboard description and the entity information description appearing in the storyboard to obtain the storyboard image prompt word, storyboard video prompt word, and the emotional prompt word of each dialogue; (3.4) For each storyboard, input the storyboard prompt words, background image, character image, and other entity images through agent 11 to obtain the storyboard reference image; (3.5) For each storyboard, Agent8 inputs the image style reference image and the storyboard reference image. Using image style transfer technology, a storyboard with a unified style is obtained. This ensures that all storyboards have a consistent visual style and solves the problem of inconsistent styles across multiple shots. (3.6) For each storyboard, input the storyboard image, storyboard video prompt words, and emotional prompt words through agent 12 to obtain the storyboard video; (3.7) For each line of dialogue, input the character's corresponding timbre, lines, and emotional cues through agent 13 to obtain the dialogue audio. The same emotional cues are used for audio and video generation. The specific implementation steps of step 4 are as follows: (4.1) Create a draft: Create a video draft through agent 14, and then perform steps (4.2) to (4.3) for each storyboard; (4.2) Action audio matching: For each storyboard, agent 15 is used to input the storyboard shot description and storyboard video to determine the start and end time of the dialogue and background sound effects; (4.3) Insert video, audio, and text into the timeline. Through agent 14, input video, audio, text, and corresponding start and end times, insert the video into the timeline, and then insert the audio and subtitles into the timeline according to the dialogue start time offset; (4.4) Select background music. Through agent16, input the storyboard list, select the appropriate background music list from the background music library and provide the corresponding start and end time of the background music; (4.5) Insert the background music into the timeline. Through agent 14, input the background music and the corresponding start and end time of the background music, and insert the background music into the timeline of the video draft; (4.6) Rendering output.

2. The method for generating a full-link text-to-video link based on a multimodal large model according to claim 1 is characterized in that: The images in the memory library include: image style reference images, reference images of scenes, characters, props, costumes, and other entities; the sounds include: sample audio of narration and other characters’ timbre; the text includes: text type, style, theme, descriptions and features of scenes, characters, props, costumes, and other entities, and timbre features of narration and other characters; The text analysis inputs the user text into the intelligent agent to analyze the text type, style, and theme, and then determine the video style and narration tone; analyzes the entities appearing in the text, analyzes the characteristics of each entity, and then creates or selects its reference image and tone and stores them in the memory library; specifically includes: Text metadata analysis: Agent 1 analyzes the type, style, and theme of user texts; Stylized resource generation: Agent 2 generates image style cues and Agent 6 generates reference images; Agent 3 matches the narration tone; Entity feature extraction and enhancement: Agent 4 extracts entities including scenes, characters, props, or clothing; Agent 5 generates text descriptions, features, and classification prompt word templates for each entity; Agent 6 generates entity reference images and stores them in the design library; Agent 7 implements intelligent character-clothing synthesis; Global style unified preprocessing: Agent8 performs style transfer on entity graphs.

3. According to the multimodal large model-based full-link text-to-video generation method of claim 1, the storyboard generation comprises inputting the user text into the storyboard parsing agent to obtain a storyboard list and a list of entities appearing in the storyboard, and then performing the following operations on each storyboard: Generate the storyboard prompt words, video prompt words, and emotional prompt words for the storyboard through the storyboard description and the feature description of the entity in the memory library. Then, generate the storyboard through the prompt words and the image style reference pictures in the memory library and the reference pictures of the entity, and then generate the storyboard video. Then, generate the audio used for the storyboard through the timbre and emotional prompt words corresponding to the entity in the memory library. This includes: Storyboard structured analysis: Agent9 is used to divide shots and output: shot description and entity list; Storyboard resource binding: automatically associate style reference images, entity images, and timbre characteristics in the memory library; Storyboard content generation: Agent 10 generates three key prompts: storyboard prompts, video prompts, and emotional prompts. Agent 11 generates storyboard reference images. Agent 8 performs secondary style transfer on the storyboard reference images to ensure a consistent style across multiple shots. Emotion-driven audio and video synchronization generation: Agent12 generates storyboard videos; Agent13 generates character dubbing.

4. The method for generating a full-link text-to-video link based on a multimodal large model according to claim 3 is characterized in that: The audio and video synthesis includes: Timeline framework construction: Create a video draft framework through Agent14; Audio and video alignment: Agent15 analyzes video content and outputs precise timestamps for dialogue or sound effects. Multi-track synthesis: Use Agent14 to insert storyboard videos into tracks according to timestamps, and inject dubbing audio and subtitles synchronously; Adaptive background music: Agent16 matches the mood of the storyboard and recommends background music and time intervals; Agent14 injects the background music into the soundtrack.

5. A full-link text-to-video generation system based on a multimodal large model, characterized by: include: Text acquisition module, used to obtain user text input; The text analysis module is used to dynamically extract, analyze, generate, associate, and store multimodal information (text, audio, and images) from input text through collaborative agents, creating a cross-modal memory library. A storyboard generation module is used to generate storyboard videos and audios based on a memory library; The audio and video synthesis module is used to synchronize the audio and video to form the final video; The system realizes full-link text-to-video generation through the method described in any one of claims 1 to 4.

6. A full-link text-to-video generation device based on a multimodal large model, characterized by: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 4.

7. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video generation control method and computer readable storage medium

    CN120017931A