An AIGC long video stabilization generation method and system based on agent cooperation
Through the agent collaborative process, the script is converted into a storyboard, character and scene setting diagrams, dialogue audio and video clips, which solves the problems of low efficiency, high cost and poor consistency in long video generation in existing technologies, and realizes efficient and low-cost multimodal long video generation.
Patent Information
- Application Number
- CN202510100849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing technologies have problems such as low efficiency, high cost, and inconsistent roles and scenes in the process of converting video scripts into finished videos. In particular, it is difficult to achieve the unification of multimodal content when generating long videos.
The AIGC long video stable generation method based on agent collaboration is adopted. Through preset prompts and large language models, the agent collaboration process is used to convert the script text into storyboard scripts, character and scene setting diagrams, dialogue audio and video clips, and finally edit them into a film to achieve automatic generation of multimodal content.
It improves the efficiency and quality of long video generation, reduces costs, solves the problem of inconsistency between characters and scenes in traditional methods, and realizes fast and efficient multimodal content generation.
Smart Images

Figure CN119893206B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an agent-collaboration-based AIGC long video stable generation method and system. Background Art
[0002] There are currently two methods for converting video scripts into finished films. The first is the traditional workflow of pre-production planning, mid-production shooting, and post-production. This requires collaboration among the director, producer, actors, editors, and other crew members. However, the shooting and production process is very lengthy and costly. The second method involves AIGC to convert text into video, which requires generating short videos of individual scenes and then stitching them together. This requires the use of multiple AI tools, and the characters and scenes are inconsistent, making it difficult to maintain consistency. This can lead to bloopers in the shots, making them unusable. The third method is that the videos currently generated by AIGC are silent, making it difficult to generate multimodal content (audio and video) in a one-stop manner.
[0003] Therefore, how to improve the efficiency and quality of long video generation and reduce the cost of long video generation has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The purpose of the present invention is to address the above-mentioned technical problems existing in the prior art and provide an AIGC long video stable generation method based on agent collaboration, so as to improve the efficiency and quality of long video generation and reduce the cost of long video generation.
[0005] Another object of the present invention is to provide an AIGC long video stable generation system based on agent collaboration that can achieve fast, efficient and automatic generation of long videos.
[0006] In a first aspect, the present invention provides an AIGC long video stable generation method based on agent collaboration, comprising the following steps:
[0007] 1) Get the script text content entered by the user, verify and organize its format, and report an error if it is empty;
[0008] 2) Using the pre-installed prompt "Crew Member" Agent, based on a large language model, and fine-tuned using thought chaining and low-frequency prompting techniques, the agent invokes application-layer tools to process the script text in different roles.
[0009] 3) The casting agent disassembles the script and extracts descriptions of characters and scenes;
[0010] 4) The script agent uses a large language model to convert the script text into a storyboard, including information such as scene number, shot number, and shot size.
[0011] 5) The prompt agent converts the role and scenario description into a standard prompt of the diffusion model and adds prompt words;
[0012] 6) Input the character and scene prompts into the diffusion model in the application layer to generate character and scene setting images. The multi-view diffusion model is then used to generate character multi-view reference images and scene panoramas.
[0013] 7) Select appropriate parts from the character's multi-view reference image according to the storyboard script and merge them with the scene image to generate a storyboard;
[0014] 8) The voice agent selects a timbre that matches the character's settings from the application-layer timbre library based on the character description. The pre-trained timbre cloning model generates dialogue audio based on the timbre and dialogue text.
[0015] 9) Actor Agent uses lip-activated technology to combine dialogue audio and character images to generate dialogue video clips;
[0016] 10) Editing Agent Edits the storyboard and dialogue video clips into a film.
[0017] In step 1), the script text content input by the user is received through the user interface. This method accepts both story text and Hollywood standard format scripts. The script text content includes elements such as the plot, character dialogue, and scene descriptions. The input script text content is initially formatted and verified to ensure smooth subsequent processing. If the received script text content is empty, the process will stop at this step and report an error.
[0018] In step 2), the prompt engineering method offers a flexible architecture. Large language models can utilize self-deployed Generative Pre-training Transformers (GPTs), such as Qwen and LLaMA, or utilize Multi-LLM through lightweight and flexible API calls. Pre-configured prompts fine-tune the large language model using chain-of-thought and few-shot prompt engineering techniques. Each agent uses a pre-configured prompt, plays a different role, possesses diverse expertise, and can call upon various application-layer tools to generate multimodal content. Through pre-configured prompts, the agent learns how to understand and process information in the script. This prompt engineering technique improves the agent's understanding of the script content and its efficiency in executing subsequent tasks.
[0019] In step 3), the casting agent extracts character and scene descriptions from the script input by the user. The agent first breaks down the script, extracts character and scene elements, and stores them in JSON format. Furthermore, the agent generates detailed information for each character and scene, using the script, character, and scene elements. This information serves as a character biography and location scouting reference for subsequent image processing.
[0020] In step 4), the script agent converts the script text or story content entered by the user into a storyboard. Specifically, the script agent uses the text understanding capabilities of the Large Language Model to convert the story or script into a text storyboard and stores it in a CSV format. This storyboard includes scene number, shot number, shot size, camera movement, composition instructions, scene arrangement, dialogue information, and sound effects.
[0021] In step 5, the prompt agent converts the character and scene descriptions into standard prompts for the diffusion model. Specifically, the prompt agent converts the character and scene descriptions extracted and refined by the casting agent into a format that the diffusion model can understand and process. To facilitate the generation of multi-angle images, the prompt includes prompts such as "white background, front view, facing the camera."
[0022] In step 6, the character and scene prompts are fed into the diffusion model to generate character and scene reference images. Specifically, the character and scene prompts generated by the prompt agent are fed into the diffusion model, which then generates character and scene reference images (512px*512px) that match the script content. The character reference images use multi-view diffusion to generate images with consistent appearance from different angles. Relative to the input image, a set of multi-view images with azimuth angles {+0, +60, +120, +180, +240, +300} are generated. The scene reference image generates a panoramic image (1752px*584px) of the scene to facilitate subsequent composition.
[0023] In step 7), the storyboard script selects appropriate portions from the character multi-angle reference image and, based on pre-set storyboard knowledge, fuses them with the scene image to generate a storyboard. For example, a character dialogue storyboard should first introduce the characters in the dialogue, followed by a reverse shot to create close-up shots of the characters A and B. The first introduction shot should have a 1:1 ratio of characters A to B within the frame, with character A selected at a 60° angle and character B selected at a 240° angle. In the close-up shot, character A should account for 70% of the frame, with character A selected at a 0° angle and character B selected at a 300° angle. Character A should be positioned at 1 / 3 of the frame width to ensure optimal character placement and positioning.
[0024] In step 8), the voice agent selects a timbre that matches the character's settings from a timbre library based on the character's description. Specifically, the voice agent selects a timbre that matches the character's settings from the timbre library based on the character's description. The pre-trained timbre cloning model generates dialogue audio from the timbre and dialogue text. Specifically, the pre-trained timbre cloning model uses natural language processing and audio synthesis technologies to convert the dialogue text into audio with the character's characteristics based on the timbre selected by the voice agent and the dialogue text extracted by the script agent.
[0025] In step 9), the actor agent generates a video of the dialogue based on the dialogue audio and the character map. Specifically, the actor agent generates a video of the dialogue using lip shape driving technology in combination with the dialogue audio and the character map.
[0026] In step 10), the editing agent edits the storyboard and the dialogue video into a film. Specifically, the editing agent uses a video processing tool to stitch the video clips together according to the corresponding numbers of the video clips and the order of the storyboard based on the storyboard and the dialogue video clips generated by the actor agent, to form a complete long video.
[0027] Furthermore, the steps 2) to 10) are specifically as follows:
[0028] Agent collaboration can be viewed as a parallelization of Chain of Thought (CoT) technology within a large language model. It consists of script, casting, prompt, audio, actor, and editing agents. Each agent utilizes prompt engineering technology, possesses film and television expertise, and has tool-calling capabilities. It possesses multimodal information processing capabilities (a multimodal-large language model), enabling efficient training-free processing of multimodal information.
[0029] Furthermore, the steps 1) to 4) are specifically as follows:
[0030] Steps 1)-4) are designed to generate a text storyboard from the script. The script (either a narrative or a standard Hollywood format script) is input, and the large language model is used to break down the core elements (characters, locations) in the script and store them as sequences. Leveraging the text understanding capabilities of the large language model, the generated text storyboards are mapped between visual elements (characters, locations) and audio elements (dialogue), providing a reference for subsequent synthesis and editing.
[0031] Furthermore, the steps 5) to 7) are specifically as follows:
[0032] Steps 5-7 aim to generate a picture storyboard from the text storyboard generated in the previous stage. First, using the script and the extracted elements (characters and scenes), and taking into account the GPT token limit, character and scene descriptions are refined one by one. The prompt agent adds generation requirements such as a "white background" to the character and scene descriptions, standardizing them into prompts. The prompts are then fed into the application-layer diffusion model to generate character and scene reference images, further generating multi-view character images and panoramic images of the scene. Based on storyboard expertise, appropriate sections are selected and synthesized into the picture storyboard.
[0033] Furthermore, the steps 8)-9) are specifically as follows:
[0034] Step 8) Based on the character description from step 3), an appropriate timbre is selected from the character's sound library. The character's dialogue audio is then generated based on the timbre and the character's dialogue from the Chinese storyboard from step 4). In step 9), the actor agent invokes the application layer's Wav2Lip lip syncing method to generate video from the character storyboard and audio.
[0035] Furthermore, the step 10) is specifically as follows:
[0036] The editing agent calls the application layer video editing method, and synthesizes the video into a complete long video according to the corresponding labels of the video segments and the text storyboard.
[0037] In a second aspect, the present invention provides an AIGC long video stable generation system based on agent collaboration, comprising the following modules:
[0038] The script-to-text storyboard module is used to obtain the script input by the user and convert it into text storyboards;
[0039] The image storyboard generation module is used to convert text storyboards into image storyboards. It first breaks down the characters and scene elements in the script, refines the descriptions to form character and scene reference images, and then further generates multi-angle character reference images and scene panoramas. The storyboard agent selects appropriate parts based on its professional knowledge and synthesizes them into image storyboards.
[0040] The video generation module is used to generate video clips from the storyboards. It also includes a dialogue audio generation module, which generates dialogue audio from the video clips based on character descriptions and script dialogue.
[0041] The video stitching module is used to stitch video clips into a complete video.
[0042] Furthermore, the script to text storyboard module is specifically as follows:
[0043] Obtain script data and break it down into elements such as characters and scenes. Convert it into a shot-by-shot script. The shot-by-shot script is stored in CSV format. It includes scene number, shot number, shot size, camera movement, composition guidance, scene arrangement, dialogue information, and sound effects. It also includes character and scene information.
[0044] Furthermore, the picture storyboard generation module is specifically:
[0045] Extract characters and scenes from the script and further generate descriptive information. Create a text-to-image model based on a diffusion model, and generate reference images and multi-angle reference images from the descriptive information. Crop the multi-angle reference images into sections, and select appropriate sections to synthesize the storyboard.
[0046] Furthermore, the video generation module is specifically:
[0047] According to the character description, a suitable timbre is selected for the character from the timbre library, and dialogue audio is generated based on the dialogue text and timbre in the text storyboard; further, Wav2Lip is used to generate a video clip of the character dialogue with a shot number.
[0048] Furthermore, the video splicing module is specifically:
[0049] Based on the video clips with shot numbers and text storyboards, the video clips are assembled into a complete video.
[0050] The advantages of the present invention are:
[0051] Long videos are created through a collaborative process using pre-trained "crew member" agents using pre-set prompts. The script agent converts the user-entered script text into a storyboard and extracts dialogue from the user-entered script. The casting agent extracts character and scene descriptions from the user-entered script. The prompt agent converts the character and scene descriptions into standard prompts for the diffusion model. The character and scene prompts are input into the diffusion model to generate storyboard keyframes. The sound agent generates timbre that matches the character description. The pre-trained voice cloning model generates dialogue audio from timbre and dialogue text. The actor agent generates dialogue video based on the dialogue audio and character images. The editing agent automatically edits the storyboard and dialogue video into a film. Applying agent collaboration to long video generation overcomes the tedious multi-tool switching process and the high cost of actual short film shooting in traditional processes. It also effectively solves the problem of inconsistent characters and scenes in traditional short film production. Ultimately, this greatly improves the efficiency and quality of long video generation and significantly reduces the cost of long video production. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a flow chart of an agent-collaboration-based AIGC long video stable generation method and system of the present invention.
[0053] Figure 2 This is a schematic diagram of an AIGC long video stable generation system based on Agent collaboration of the present invention. DETAILED DESCRIPTION
[0054] The technical solution in the embodiments of this application has the following overall concept: applying agent collaboration to the production of long videos overcomes the tedious multi-tool switching process of traditional processes and the high cost of actual short film shooting. It also effectively solves the problem of inconsistent characters and scenes in traditional short film production. Ultimately, this greatly improves the efficiency and quality of long video production and significantly reduces the cost of long video production.
[0055] Please refer to Figure 1 As shown, a preferred embodiment of the present invention is a method and system for stabilizing long video generation using AIGC based on agent collaboration, comprising the following steps:
[0056] Step S10: Obtain the script text content input by the user;
[0057] The method receives the script text content input by the user through the user interface, which is compatible with story text and Hollywood standard format scripts. The script text content covers the core elements of the plot, character dialogue, scene description, etc. The input script text content is strictly formatted and verified to check if the text meets the common script format specifications, such as correct format of dialogue, scene transition, character name annotation, etc. At the same time, the text content is sorted and redundant spaces, special characters (random codes unrelated to scripts, etc.) are removed to ensure smooth subsequent processing. If the received script text content is empty, the system will immediately feedback an error prompt and require the user to re-enter.
[0058] Step S20, use the preset prompt "crew staff" Agent, the LLM-based Agent has text understanding ability, and can also call application tools in the application layer;
[0059] The method architecture has high flexibility, can use self-deployed Generative Pre-training Tranformer (GPT), such as Qwen, LLaMA, etc., fully utilizes local computing resources, guarantees data security and processing efficiency. Also can use the light and flexible way of calling API to use Multi-LLM, with the help of powerful computing power of cloud, quickly responds to processing needs. The preset prompt uses Chain-of-Thought and few-shot prompting engineering technology to fine-tune the large language model. Each Agent is assigned a unique preset prompt, simulating different crew roles such as director, casting director, and screenwriter, thus possessing different professional knowledge backgrounds. For example, the prompt of the script Agent will guide it to understand the overall structure of the script, the logic of the plot development, and how to convert the text into visualized storyboard scripts; the prompt of the casting Agent focuses on professional knowledge such as character feature analysis and image matching. These Agents can call various application tools in the application layer, such as text processing tools, image generation tools, and audio processing tools, to generate multi-modal content. Through the preset prompt, the Agent can deeply learn how to accurately understand and efficiently process the information in the script, greatly improving the understanding ability of the script content and the execution efficiency of subsequent tasks.
[0060] Step S30, the casting Agent extracts the description of the characters and scenes from the user input script;
[0061] The casting agent first deeply disassembles the script, accurately identifies the character elements and scene elements in the script through natural language processing technology, and stores them in json format for subsequent data calling and processing. Further, using the rich information provided by the script, combined with the character and scene elements, a detailed character profile is generated for each character, including the character's personality traits, physical characteristics, background, and growth experience. A comprehensive scene reference is generated for each scene, such as the scene's geographical location, environmental atmosphere, and spatial layout. In the generation process, semantic analysis and sentiment analysis techniques are used to ensure the accuracy and completeness of the character and scene information, providing a solid foundation for subsequent reference picture production.
[0062] Step S40, the script agent converts the user input script text content into a split script;
[0063] With the powerful text understanding ability of large language models, the script agent deeply analyzes the plot development, character relationships, and scene transitions of the script. The story or script is carefully converted into a professional text split, stored in csv format for easy data organization and analysis. The split script records the scene, shot number, scene type (such as full shot, medium shot, close-up, close-up, etc.), camera movement (push, pull, pan, shift, follow, etc.), composition guidance (position of characters in the frame, symmetry and balance of the frame, etc.), scene scheduling (character movement, action, etc.), dialogue information, sound effects, and other key information. In the generation process, the script agent will refer to classic film and television split cases and professional split production specifications to ensure the professionalism and practicality of the split script.
[0064] Step S50, the prompt agent converts the character and scene descriptions into a standard prompt for Stable Diffusion;
[0065] The prompt agent uses semantic conversion and format adaptation techniques to convert the descriptions into a standard format that the diffusion model can accurately understand and efficiently process based on the character and scene descriptions extracted and refined by the casting agent. To generate high-quality, multi-perspective pictures, the prompt is cleverly added with key words such as "white background, front, facing the camera." At the same time, according to the character's personality traits and the scene's atmosphere requirements, descriptions such as "firm expression" and "mysterious light effects" are added to guide the diffusion model to generate images that meet the expectations.
[0066] Step S60, input the prompt of the character and the scene into Stable Diffusion in the application layer to generate a setting picture of the character and the scene; and further generate a multi-view reference picture of the character and a panoramic picture of the scene using a multi-view diffusion model.
[0067] The prompt Agent carefully generated character and scene prompts are input into the Stable Diffusion model in the application layer. According to these prompts, the model uses deep learning algorithms and image generation techniques to generate character reference pictures and scene reference pictures (512px * 512px) that are highly consistent with the content of the script. For character reference pictures, advanced multi-view diffusion technology is used to generate pictures with consistent appearance characteristics at different angles. Compared with input pictures, a set of multi-view pictures with azimuth angles of {+0, +60, +120, +180, +240, +300} are generated to fully display the appearance of the character. Scene reference pictures are generated by an optimized diffusion model to provide a broad and rich scene basis for further composition. During the generation process, the parameters of the model are finely adjusted according to different script styles and requirements, such as science fiction, historical costume, and modern, to optimize the generation strategy of the model and ensure that the quality and style of the generated images meet the overall settings of the script.
[0068] Step S70, select appropriate parts from the character multi-view reference picture according to the split screen script and fuse them with the scene picture to generate a split screen picture.
[0069] According to the pre-set split screen knowledge and professional composition principles, appropriate parts are accurately selected from the character multi-view reference picture and skillfully fused with the scene picture to generate high-quality split screen pictures. For example, in a character dialogue split screen, first, follow the conventional methods of film and television shooting to introduce the dialogue characters, and then use the AB character close-up split screen method to form a dialogue split screen. In the first split screen, the proportion of the characters A and B in the picture should be 1:1, and the azimuth angle of character A should be 60° and the azimuth angle of character B should be 240° to show the relative position and communication state of the two characters. In the character close-up split screen, the proportion of character A should be 70%, and the azimuth angle of character A should be 0° and the azimuth angle of character B should be 300°, and character A should be located at 1 / 3 of the picture width. Through this precise selection and layout, the characters are reasonably arranged and accurately positioned in the picture, enhancing the expressiveness and narrative nature of the split screen picture. During the fusion process, image fusion algorithms are used to ensure that the light and shadow, color tone, and other elements between the characters and the scene are consistent, creating a realistic visual effect.
[0070] Step S80: The voice agent selects a timbre that matches the character's setting from the application layer timbre library according to the character description; the timbre cloning pre-trained model generates dialogue audio from the timbre and dialogue text;
[0071] Based on the character's detailed description, including personality traits, age, and identity, the Voice Agent intelligently selects a voice that closely matches the character from a rich library of application-level voices. The Voice Cloning pre-trained model leverages advanced natural language processing and audio synthesis technologies to synthesize the voice selected by the Voice Agent and the dialogue text extracted by the Script Agent. During the synthesis process, the model fully considers factors such as intonation, speech rate, and emotional expression. Dynamically adjusting voice parameters based on the plot and emotional changes of the character, the generated dialogue audio retains distinct character characteristics. This Voice Cloning pre-trained model is trained on a large amount of voice data, covering a wide range of language styles, emotional types, and voice characteristics to ensure the generation of high-quality, natural and fluent dialogue audio.
[0072] Step S90: The actor Agent generates a video clip of the dialogue based on the dialogue audio and the character map;
[0073] The Actor Agent uses advanced lip-activated technology, combined with dialogue audio and character images, to generate realistic dialogue video clips. Through speech analysis of the dialogue audio, it accurately extracts information such as phonemes and prosody. This information is then used to drive the lip movements in the character image, achieving precise synchronization between lip shape and speech. During the generation process, lip movements are personalized to account for the speaking habits and facial expressions of different characters, making the characters in the video clips more vivid and natural. Furthermore, image rendering and animation technologies are used to add appropriate facial expressions and body movements to the characters, enhancing the expressiveness and appeal of the videos.
[0074] Step S100: The editing agent edits the storyboard and the dialogue video clips into a film.
[0075] Based on the detailed storyboard and the dialogue clips generated by the actor agents, the editing agent uses professional video processing tools to precisely stitch the clips together according to the storyboard sequence. During the stitching process, the clips are labeled to ensure their correct placement on the timeline. Transition effects between shots, such as fades, flash cuts, and rotations, are also considered to ensure a smooth and natural transition. The overall rhythm of the video is controlled, adjusting the playback speed and editing rhythm based on the plot's intensity and emotional ups and downs to enhance the video's visual appeal. Finally, the stitched video undergoes comprehensive optimization and adjustments, including color correction, audio mixing, and other post-processing, to create a complete, high-quality, long-form video.
[0076] like Figure 2 , a preferred embodiment of the AIGC long video stable generation system based on agent collaboration of the present invention includes the following modules in sequence: script to text storyboard module, picture storyboard generation module, dialogue audio generation module, video clip generation module, video splicing module;
[0077] The script to text storyboard module: as the starting module of the entire system, it is responsible for processing the script text input by the user and breaking the script into elements such as characters and scenes. Convert it into a storyboard script. The storyboard script is stored in the form of csv. It includes scene number, shot number, shot size, camera movement, composition guidance, scene scheduling, line information, and sound effects. It also includes character information and scene information; its core is to carefully break down the various elements in the script and present them in a clear and organized manner, providing a basic framework for the creation of long videos. In addition, the module performs a preliminary logical check on the generated storyboard script to ensure the coherence between the storyboards and the rationality of the plot development. For possible logical loopholes, the script structure is optimized through intelligent prompts or automatic adjustments to ensure the smoothness of subsequent video creation.
[0078] The picture storyboard generation module creates corresponding visual images for each storyboard based on the storyboard script output by the script-to-text storyboard module. It also creates a text-to-image model based on the diffusion model, generating reference images and multi-angle reference images from the descriptive information. The multi-angle reference images are then cropped into sections, and appropriate sections are selected to synthesize the picture storyboards. This provides a foundation for the video's visual presentation and ensures that the images match the script content.
[0079] The dialogue audio generation module selects a suitable timbre for the character from a timbre library based on the character description, and generates dialogue audio based on the dialogue text and timbre in the text storyboard; it gives the character in the video a voice with personality and emotion, making the character's language expression more vivid and more in line with the character image and plot atmosphere.
[0080] The video clip generation module: according to the dialogue audio and the character image, the character dialogue video clip with mirror number is generated by Wav2Lip, the language expression and action expression of the character are organically combined, and the character is lifelike in the video.
[0081] The video splicing module: according to the video clip with mirror number and the text breakdown, each video clip is spliced into a complete long video, and the whole video is coherent and smooth after post-processing.
[0082] Each module transmits and shares information through a unified data interface. The output data of each module is stored in a standardized format, facilitating the calling and processing of other modules. The breakdown script output by the script to text breakdown module is stored in CSV format, and other modules can easily read the information therein; the image data generated by the picture breakdown generation module adopts a common image format and contains corresponding metadata, so that subsequent modules can operate it.
[0083] During the operation of the system, an error monitoring and processing mechanism can be set. Real-time monitoring is performed for possible input errors, model running errors, data processing errors, etc. When an error occurs, the system provides explicit error information to the user according to the error type, and attempts to automatically repair or provide corresponding solutions. If the image generated by the table Diffusion type has quality problems, the system will automatically adjust the model parameters or use a backup image generation algorithm; if the audio synthesis has problems, different audio synthesis parameters or voice selection are tried.
[0084] The following takes making an animated long video about "Little Red Riding Hood" as an example to illustrate the present application. The script is divided into several clips: Little Red Riding Hood and her mother say goodbye; the Big Bad Wolf eats the grandmother; Little Red Riding Hood goes to her grandmother's house and is eventually eaten by the Big Bad Wolf; the hunter arrests the Big Bad Wolf and rescues Little Red Riding Hood and her grandmother.
[0085] Step 1: Obtain the user input script text content:
[0086] The user inputs the script text content about "Little Red Riding Hood" through the user interface. The script contains detailed plot, character dialogue, scene description and other elements. The system performs preliminary format verification and arrangement on the input script to ensure smooth subsequent processing.
[0087] Step 2: Pre-trained "Crew" Agent with preset prompt
[0088] The system pre-trains the script agent, casting agent, prompt agent, sound agent, actor agent, and editing agent using pre-set prompts. These agents learn how to understand and process script information through pre-set prompts, improving their understanding of the script content and the efficiency of executing subsequent tasks. Application layer tools include stable diffusion and multi-view diffusion, a character timbre library, a timbre cloning model, and a Wav2Lip lip sync driver.
[0089] Step 3: The casting agent extracts core elements from the script text entered by the user: characters and scenes, and then describes each one in detail. For example, from "Little Red Riding Hood," the characters are extracted and numbered: Little Red Riding Hood 0, Mom 1, Grandma 2, Big Bad Wolf 3, Hunter 4. The scenes are extracted: Little Red Riding Hood's House 0, Forest 1, Grandma's House 2. Next, the descriptions are refined:
[0090]
[0091] Step 4: The script agent uses the large language model to convert the script into text storyboards.
[0092] For example, (Scene 1) Little Red Riding Hood says goodbye to her mother.
[0093]
[0094] And disassembled into storyboard script
[0095]
[0096] Step 5: The prompt agent converts the description of the role and scene into a standard prompt for Stable Diffusion.
[0097]
[0098] Step 6: Input the character and scene prompts into Stable Diffusion in the application layer to generate a set image of the character and scene (512px*512px). Furthermore, a multi-view diffusion model is used to generate a multi-view reference image of the character (with azimuth angles of {+0, +60, +120, +180, +240, +300}) and a panoramic image of the scene (1080px*584px).
[0099] Step 7: Select the appropriate part from the character's multi-view reference image according to the storyboard script, and merge it with the scene image to generate a storyboard.
[0100]
[0101] Step 8, the sound Agent selects the tone that meets the role setting from the application layer tone library according to the role description; for example, a clear and pleasant child female voice is selected for Little Red Riding Hood; a young female voice is selected for the mother. The tone cloning pre-training model generates dialogue audio from the tone and dialogue text; for example: audio 0 is generated for shot 2: "Mom, goodbye!" Audio 1 is generated for shot 2: "Be careful on the road."
[0102] Step 9, the actor Agent generates the video clip of the dialogue according to the dialogue audio and the role graph; audio 0 is used on shot picture 2 for lip matching. Video 1 is generated. Audio 1 is used on shot picture 3 for lip matching. Video 2 is generated.
[0103] Step 10, the editing Agent edits the film according to the shot script and the dialogue video clip. That is, the scenes and shots are rented in order.
[0104] The present application adopts multi-Agent division of labor and cooperation, modularizes the long video generation process, and avoids tedious manual operation. The large language model combines prompt engineering technology, can quickly process script information, accelerate the generation of shot script, and shorten the production cycle. Reduce the dependence on manpower and reduce labor costs; without actual shooting, save expenses such as venue, props, equipment, and actor compensation. The information extracted from the script throughout the present application ensures the consistency of the role and the scene; using advanced technology to ensure the high quality and stability of the image, audio, and video, optimizing the transition and post-processing when editing. The present application has flexible architecture and can be deployed locally or in the cloud; it can adapt to different script styles and languages, and can extend functional modules or update algorithms. The present application facilitates users to quickly convert ideas into long videos. Effectively solve the problem of inconsistent roles and scenes in the traditional short film production process, improve the efficiency and quality of long video generation, and reduce the cost of long video generation.
[0105] Although the specific embodiments of the present application are described above, those skilled in the art should understand that the specific examples described are only illustrative, and are not intended to limit the scope of the present application. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present application should be covered within the scope of the claims of the present application.
Claims
1. A stable long video generation method based on AIGC agent collaboration, characterized by The steps include: 1) Obtain the script text content input by the user, verify its format and organize it, and report an error if it is empty; the script text content includes the storyline, character dialogue, and scene description; 2) Using the pre-installed prompt "Crew Member" Agent, based on a large language model, and fine-tuned using thought chaining and low-frequency prompting techniques, different roles invoke application-layer tools to process the script text. 3) The casting agent disassembles the script and extracts descriptions of characters and scenes; 4) The script agent uses a large language model to convert the script text into a storyboard, including scene number, shot number, shot size, camera movement, composition guidance, scene arrangement, dialogue information, and sound effects. 5) The prompt agent converts the role and scenario description into a standard prompt of the diffusion model and adds prompt words; 6) Input the character and scene prompts into the diffusion model in the application layer to generate character and scene setting images. The multi-view diffusion model is then used to generate character multi-view reference images and scene panoramas. 7) Select appropriate parts from the character's multi-view reference image according to the storyboard script and merge them with the scene image to generate a storyboard; 8) The voice agent selects a timbre that matches the character's settings from the application-layer timbre library based on the character description. The pre-trained timbre cloning model generates dialogue audio based on the timbre and dialogue text. 9) Actor Agent uses lip-activated technology to combine dialogue audio and character images to generate dialogue video clips; 10) Editing Agent Edits the storyboard and dialogue video clips into a film.
2. The method for generating stable long videos using AIGC based on agent collaboration as claimed in claim 1, characterized in that: In step 2), the large language model is a self-deployed GPT, using Qwen or LLaMA; Or by calling the Multi-LLM API. Pre-set prompts use thought chaining and low-prompt engineering technology to fine-tune the large language model. Each agent uses a preset prompt, plays a different role, has different professional knowledge backgrounds, and calls different tools in the application layer to generate multimodal content. The agent learns how to understand and process the information in the script through preset prompts. Through prompt engineering technology, the agent's ability to understand the script content and the efficiency of executing subsequent tasks are improved.
3. The method for generating stable long videos using AIGC based on agent collaboration as claimed in claim 1, characterized in that: In step 3), the description of the characters and scenes includes character biographies and location scouting references, covering the characters' personalities, appearances, background information, and the location, atmosphere, and spatial layout of the scenes.
4. The method for stabilizing long video generation based on AIGC and agent collaboration as claimed in claim 1, characterized in that: In step 5), the prompt agent converts the character and scene descriptions into a standard prompt for the diffusion model and adds prompt words. Specifically, the prompt agent converts the character and scene descriptions extracted and refined by the casting agent into a format that the diffusion model can understand and process. To facilitate the generation of multi-perspective images, the prompt wording "white background, front view, facing the camera" is added to the prompt.
5. The method for stabilizing long video generation based on AIGC and agent collaboration according to claim 1, characterized in that: In step 6), the character multi-view reference images and scene panoramic images are generated. Specifically, the character and scene prompts generated by the prompt agent are input into the diffusion model. The diffusion model generates character reference images and scene reference images that match the script content based on these prompts. The character reference images use multi-view diffusion to generate images with consistent appearance features at different angles. Relative to the input image, a set of multi-view images with azimuth angles of {+0, +60, +120, +180, +240, +300} are generated. The scene reference images generate a panoramic image of the scene to facilitate subsequent further composition.
6. The method for stabilizing long video generation based on AIGC and agent collaboration according to claim 1, characterized in that: In step 7), when generating the storyboard, for the character dialogue storyboard, first explain the dialogue characters, and use the method of positive and negative shots to form the A and B character close-up storyboards. In the first storyboard, the character A and B screen ratio is 1:1, and character A with an azimuth angle of 60° and character B with an azimuth angle of 240° are selected respectively; in the character close-up storyboard, character A accounts for 70%, and character A with an azimuth angle of 0° and character B with an azimuth angle of 300° are selected, and character A is located at 1 / 3 of the width of the screen to achieve reasonable arrangement and positioning of the characters.
7. The method for stabilizing long video generation based on AIGC and agent collaboration as claimed in claim 1, characterized in that: In step 9), the actor agent adds facial expressions and body movements that match the lines and emotions to the character when generating the dialogue video clip.
8. The method for generating stable long videos using AIGC based on agent collaboration as claimed in claim 1, characterized in that: In step 10), the editing agent edits the storyboard and the dialogue video clips into a film. Based on the storyboard and the dialogue video clips generated by the actor agent, the editing agent uses video processing tools to splice the video clips according to their corresponding labels and in the order of the storyboard to synthesize a complete long video. When splicing the videos, the editing agent uses various transition effects, including fade in and fade out, flash cut or rotation, and performs color correction and audio mixing on the videos.
9. A generation system based on the agent-cooperation-based AIGC long video stable generation method according to any one of claims 1 to 8, characterized in that: Includes the following modules: The script-to-text storyboard module is used to obtain the script input by the user and convert it into a text storyboard containing scene numbers, shot numbers, shot sizes, camera movements, composition instructions, scene scheduling, dialogue information, and sound effect information; The picture storyboard generation module is used to convert the characters and scene elements in the text storyboard into multi-angle reference images of the characters and panoramic images of the scene, and synthesize the picture storyboard based on the storyboard knowledge; A video generation module is used to generate video clips from storyboards. The video generation module includes a dialogue audio generation module, which is used to generate dialogue audio in the video clips based on character descriptions and script dialogues. The video splicing module is used to splice video clips with shot numbers into a complete video according to text segmentation.
10. The generation system of the agent-cooperative AIGC long video stable generation method according to claim 9, characterized in that: Each module exchanges data through a data interface, and data transmission between modules follows a unified standard format.
Citation Information
Patent Citations
Advertisement landing page generation system and method based on AIGC, medium and equipment
CN118115208A
Digital human video generation method and device based on AIGC technology
CN118945440A