Synthesizing method and system of split video and related equipment
By generating storyboard scripts and replacing text descriptions with reference identifiers, the problem of character consistency in storyboard videos was solved, achieving stable generation of unified character images and multi-character interactions, and improving the automation and controllability of video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MAIFENG TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to ensure cross-camera consistency of characters and precise control over multi-character interaction scenes when generating storyboard videos, lacking automated and systematic solutions.
By acquiring script data to generate structured storyboards, associating each character with a unique visual reference image, and replacing text descriptions with reference identifiers when generating storyboard images, an image generation model is used to ensure consistency in character appearance, and finally, the storyboard videos are synthesized in sequence.
It achieves a high degree of consistency in the image of the same character under different shots, improves the stability and automation of multi-character video content generation, and reduces repetitive work for creators.
Smart Images

Figure CN122002102A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video editing technology, and in particular to a method, system, and related equipment for compositing storyboard videos. Background Technology
[0002] With the development of generative artificial intelligence, the technology of generating videos based on generative models has gradually matured. There are two main types of existing technical solutions for generating videos based on generative models: end-to-end video generation: directly inputting text or reference images and generating videos by the model; frame-by-frame image generation + stitching: obtaining each frame or key frame through text generation or image generation, and then performing temporal stitching.
[0003] In practical applications such as film, animation, and advertising production, users often want to control video content through storyboards. In application scenarios involving characters appearing consecutively in multiple storyboards, whether it is end-to-end generation or frame-by-frame generation, when the model processes the cue words of different storyboards, even if the text descriptions are exactly the same, the model will reimagine and render the character's appearance each time it is generated. Existing video generation methods still have obvious shortcomings in ensuring the consistency of video characters across shots and in accurately controlling multi-character interaction scenes. Existing technologies urgently need a storyboard video generation solution that can automate and systematically solve the problem of character image consistency.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] This invention provides a method, system, and related equipment for synthesizing storyboard videos. The main objective of this invention is to solve the technical problems mentioned in the background section of the prior art.
[0006] The first aspect of this invention provides a method for synthesizing storyboard videos, comprising: The script data is obtained from input, and the script data is processed using a large language model to generate a structured storyboard. Read the global character list from the storyboard script, and associate a reference image as the unique visual reference for each character in the global character list; The storyboard is traversed, and for each storyboard, a target text-based image prompt is generated based on the initial text-based image prompt in the storyboard. In the target text-based image prompt, the text description of the character appearing in the initial text-based image prompt is replaced with a reference identifier of the reference image associated with the character. For each of the storyboards, the target text-based image prompts and the reference images of the characters involved in the storyboard are input into the image generation model to generate a storyboard image corresponding to the content of the storyboard; According to the order of the storyboard script, multiple storyboard images are combined into a storyboard video.
[0007] In an optional embodiment of the first aspect of the present invention, the step of acquiring input script data and processing the script data using a large language model to generate a structured storyboard includes: Using a large language model, the script data is converted into a structured data format containing a global character list, storyboard index, scene description, storyboard character information, and initial text-to-image prompts according to a preset instruction template, thus obtaining the storyboard script.
[0008] In an optional embodiment of the first aspect of the invention, the reference image for each character in the character list is determined by the following method: Based on the character descriptions extracted from the character list, a text image generation model is invoked to generate the image. Alternatively, a match can be made from a pre-defined character image library based on the character's information, or the character can be directly specified.
[0009] In an optional embodiment of the first aspect of the present invention, the reference identifier is a specific placeholder or code that can be recognized by the image generation model, and the reference identifier is used to instruct the model to render the associated character in the storyboard image in a specified area of the storyboard image according to the reference character in the corresponding reference image.
[0010] In an optional embodiment of the first aspect of the invention, the step of replacing the text description of the character appearing in the initial text-based image prompt with a reference identifier of the reference image associated with the character in the target text-based image prompt includes: Determine whether the currently processed storyboard contains at least two characters from the character list; The replacement of the text description of a character with a reference identifier of the reference image associated with that character is performed only if the currently processed storyboard contains at least two characters from the character list.
[0011] In an optional embodiment of the first aspect of the present invention, the target text image prompt includes stylized prompt phrases for defining the overall artistic style of the scene, so as to ensure that all the storyboard images have a consistent style.
[0012] In an optional embodiment of the first aspect of the present invention, the step of combining multiple storyboard images into a storyboard video according to the order of the storyboard script includes: The generated multiple storyboard images are combined with image-generated video cues used to describe dynamic effects, and then input into the video generation model to generate a sequence of video clips containing dynamic transition effects. The generated video segment sequence is spliced together to obtain a storyboard video.
[0013] A second aspect of the present invention provides a system for synthesizing storyboard videos, the system comprising: The script data processing module is used to acquire input script data, process the script data using a large language model, and generate a structured storyboard. A reference image determination module is used to read a global character list from the storyboard script and associate a reference image as a unique visual reference for each character in the global character list; The prompt word generation module is used to traverse the storyboard script and generate target text-based image prompt words for each storyboard script based on the initial text-based image prompt words in the storyboard script. In the target text-based image prompt words, the text description of the character appearing in the initial text-based image prompt words is replaced with the reference identifier of the reference image associated with the character. The storyboard image generation module is used to input the target text image prompt and the reference image of the character involved in the storyboard into the image generation model for each storyboard, and generate a storyboard image corresponding to the content of the storyboard. The storyboard video generation module is used to synthesize multiple storyboard images into a storyboard video according to the order of the storyboard script.
[0014] A third aspect of the present invention provides a video generation device, the video generation device comprising: a memory and at least one processor, the memory storing instructions, and the memory and the at least one processor being interconnected via a circuit; The at least one processor invokes the instructions in the memory to cause the video generation device to perform a method for compositing storyboard videos as described in any one of the first aspects of the invention.
[0015] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for synthesizing storyboard videos as described in any one of the first aspects of the present invention.
[0016] Beneficial Effects: This invention provides a method, system, and related equipment for compositing storyboard videos. The method includes first parsing the input script data to generate a structured storyboard, and associating each character in the global character list of the storyboard with a reference image as a unique visual benchmark; then, before generating storyboard images for each scene, performing key transformation on the initial text-based image prompts in the storyboard, replacing the text descriptions of characters in the initial text-based image prompts with reference identifiers pointing to the reference images; subsequently, inputting the transformed target text-based image prompts and the corresponding reference images into an image generation model to generate storyboard images; finally, compositing the generated storyboard images into a video in sequence. This invention's storyboard video compositing method, by strongly binding characters to fixed reference images, fundamentally ensures a high degree of consistency in the image of the same character in different shots, significantly improving the stability of multi-character video content generation. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an embodiment of a method for synthesizing storyboard videos according to the present invention; Figure 2 This is a schematic diagram of an embodiment of a video storyboard synthesis system according to the present invention; Figure 3 This is a schematic diagram of an embodiment of a video generation device according to the present invention. Detailed Implementation
[0018] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] See Figure 1 The first aspect of the present invention provides a method for synthesizing storyboard videos, comprising: S100. Obtain the input script data, process the script data using a large language model, and generate a structured storyboard. In this invention, this step may specifically include: using a large language model, according to a preset instruction template, converting the script data into a structured data format containing a global character list, storyboard index, scene description, storyboard character information, and initial text-to-image prompts, to obtain the storyboard.
[0020] For example, a user inputs the text "Xiaoming and Xiaohua met on a park bench. At first, they were a little nervous, but Xiaohua took the initiative to greet them, and they chatted happily." The system then combines this text with a preset instruction template and sends it to a large language model (such as GPT-4 or an equivalent model). The instruction template instructs the model to identify the main characters in the story and create unique characterIds and a global character list for them, as well as to break the story into multiple logically coherent scenes. For each scene, the model generates a scene description, initial text-to-image prompts, and initial image-to-video prompts. Furthermore, each scene clearly labels the character IDs. The large language model returns structured data in JSON format (i.e., the scene script). An example could be as follows: { "roles": [ { "name": "Xiaoming", "prompt": "A boy with black hair wearing a blue T-shirt", "characterId": "char_id_001", "categoryName": "human"} { "name": "Xiaohua", "prompt": "A blonde girl wearing a red dress", "characterId": "char_id_002", "categoryName": "woman"} ] "scences": [ { index: "1" "useCharacterId": ["char_id_001", "char_id_002"] "subtitleText": "Xiaoming and Xiaohua met on a park bench, and Xiaoming looked nervous." "text2imagePrompt": "A boy and a girl are sitting on a park bench. The boy looks nervous, and the girl is looking at him." "image2videoPrompt": "The boy's eyes flickered uneasily, and the girl's head turned slightly." } { "index": "2", "useCharacterId": ["char_id_001", "char_id_002"] "subtitleText": "Xiaohua greeted them with a smile, and the two chatted happily." "text2imagePrompt": "The girl on the bench smiled and spoke to the boy, and the boy smiled back." "image2videoPrompt": "The two people's lips are moving, as if they are talking, and their expressions are pleasant." } ] }
[0021] S200. Read the global character list from the storyboard script, and associate a reference image as the unique visual reference for each character in the global character list. In this invention, the reference image for each character in the character list is determined by: generating it using a text image generation model based on the character description extracted from the character list; or matching or directly specifying it from a preset character image library based on the character information.
[0022] Specifically, there are two main methods for acquiring the reference image in this invention, taking the scenario exemplified in step S100 as an example: Method 1 (Dynamic Generation): Extract the character description prompt (a boy with black hair wearing a blue T-shirt), call a text-to-image generation model (such as the Stable Diffusion model or the Midjourney model), generate a standard character image, and save it as ref_img_001.jpg.
[0023] Method 2 (Preset Library Specification): The system provides a character library interface, allowing users to upload an existing image of "Xiaoming" or select one from the library and bind it to char_id_001. Similarly, a reference image ref_img_002.jpg is also associated with "Xiaohua" for char_id_002. In this way, each character has a unique and fixed visual reference.
[0024] S300. Traverse the storyboard script, and for each storyboard, generate a target text-to-image prompt based on the initial text-to-image prompt in the storyboard script. In the target text-to-image prompt, replace the text description of the character appearing in the initial text-to-image prompt with a reference identifier of the reference image associated with the character.
[0025] Specifically, step S300 is the core of this invention. This invention ensures character consistency by replacing the text description of the character in the initial prompt. For example, if the system detects that the useCharacterId list of scene 1 contains char_id_001 and char_id_002, which is a multi-character scene, the system reads the initial text2imagePrompt: "A boy and a girl are sitting on a park bench. The boy looks nervous, and the girl is looking at him." The system replaces "a boy" in the text with a reference identifier [REF_CHAR_1] pointing to ref_img_001.jpg. The reference identifier is a specific placeholder or code that can be recognized by the image generation model. The reference identifier is used to instruct the model to render the associated character in the scene image in a specified area of the scene image according to the reference character in the corresponding reference image.
[0026] Similarly, the system will replace "a girl" with [REF_CHAR_2], and during the replacement process, the system will remove any appearance-related words from the original character description (such as "black-haired boy", where "black-haired" will be removed) to prevent conflicts with the reference image.
[0027] It should be noted that the reference identifier replacement of the present invention is mainly aimed at scenarios with multiple characters. That is, in an optional embodiment of the first aspect of the present invention, before replacing the text description of the character appearing in the initial text image prompt in the target text image prompt with the reference identifier of the reference image associated with the character, the method includes: determining whether the currently processed storyboard contains at least two characters in the character list; and only when the currently processed storyboard contains at least two characters in the character list, performing the replacement of the text description of the character with the reference identifier of the reference image associated with the character.
[0028] Furthermore, in order to make the generated image style more consistent, the system loads a preset stylized prompt phrase, such as "cinematic style, 8k, best quality, masterpiece", and appends it to the converted prompt phrase. That is, in an optional embodiment of the first aspect of the present invention, the target text image prompt phrase includes a stylized prompt phrase for defining the overall artistic style of the picture, so as to ensure that the style of all the storyboard images is consistent.
[0029] After the above steps, the present invention can obtain a set of text-based image prompts, such as: [REF_CHAR_1] and [REF_CHAR_2] are sitting on a park bench, [REF_CHAR_1] looks nervous, [REF_CHAR_2] is looking at him, cinematic style, 8k, best quality, masterpiece.
[0030] S400. For each of the storyboards, the target text-based image prompt and the reference image of the character involved in the storyboard are input into the image generation model to generate a storyboard image corresponding to the content of the storyboard.
[0031] In this step, taking the processing of storyboard 1 as an example, the system provides the text-based image prompts generated in the previous step, as well as the two reference images ref_img_001.jpg (as input to [REF_CHAR_1]) and ref_img_002.jpg (as input to [REF_CHAR_2]), to an image generation model that supports image references (such as an image diffusion model integrated with IP-Adapter or ReVision). When generating the image, the model understands [REF_CHAR_1] and [REF_CHAR_2] as reference placeholders, then reads ref_img_001.jpg to obtain the appearance of the first character, and draws the image of the first character based on the action description "tense expression" in the prompt; at the same time, it reads ref_img_002.jpg to obtain the appearance of the second character, and draws the image of the second character based on the action description "looking at him" in the prompt, finally obtaining a static storyboard image scene_01.png that matches the content of storyboard 1 and accurately depicts the character.
[0032] For scene 2, the system will process it in the same way as scene 1, generating the image scene_02.png. In this invention, since both scenes use the same reference image, the appearance of "Xiaoming" and "Xiaohua" in the two images will remain highly consistent.
[0033] S500. According to the order of the storyboard script, multiple storyboard images are combined into a storyboard video. After the processing in step S400, the system processes the generated static storyboard images sequentially according to the storyboard index order. In this invention, this step may include: combining the generated multiple storyboard images with image-generated video prompts used to describe dynamic effects, and inputting them together into a video generation model to generate a video segment sequence containing dynamic transition effects; splicing the generated video segment sequence to obtain the storyboard video.
[0034] Specifically, in step S500, taking the static storyboard image scene_01.png as an example, for the static storyboard image scene_01.png, the system will extract the corresponding image-generated video prompt image2videoPrompt from the storyboard script: "The boy's eyes flickered uneasily, and the girl's head turned slightly." The system will input scene_01.png and this image-generated video prompt together into the video generation model (such as the SVD model and the Runway Gen-2 model). Based on the static storyboard image scene_01.png, the model will generate a short video clip clip_01.mp4 according to the dynamic description in the image-generated video prompt. The system will repeat the above operation on the storyboard image scene_02.png to obtain the short video clip clip_02.mp4. Finally, the system will use a video processing tool (such as the FFmpeg library) to stitch the two short video clips together in sequence to obtain the final storyboard video file.
[0035] In summary, the main key steps of the storyboard video synthesis method of this invention are as follows: Step 1: Storyboard Generation and Analysis 1.1 Enter the text description provided by the user (such as a script outline or scene description).
[0036] 1.2 A pre-trained large language model is used to generate multiple storyboard scripts, which include scene settings, character actions and dialogues.
[0037] 1.3 Automatically parse the list of characters involved from the storyboard text, record the character information contained in each shot, and establish a correspondence table between characters and storyboards.
[0038] Step 2: Character Image Generation 2.1 For each parsed character, call the Wensheng graph model to generate a standard character image.
[0039] 2.2 Character images serve as reference information for subsequent storyboard generation, ensuring that the same character maintains a consistent appearance in different shots.
[0040] 2.3 Character images can be saved as searchable files as needed for subsequent video compositing and secondary editing.
[0041] Step 3: Storyboard Image Generation 3.1 Single-character / no-character scene: When the number of characters involved in a storyboard is ≤1, directly input the storyboard description and the character reference image (if it exists) into the image-generated image model to generate the corresponding storyboard image.
[0042] 3.2 Multi-character scenes: When a storyboard involves ≥2 characters, perform the following operations: Design prompt word generation rules to bind character reference images with storyboard descriptions; Replace the original character names and appearance descriptions in the storyboard descriptions with "the character in picture n"; The prompts clearly define the correspondence between different characters to ensure that the identities and appearances of different characters are consistent in the generated storyboard images.
[0043] In multi-role scenarios, the system uses the following rules to generate prompts: Input: Current storyboard content.
[0044] Replacement rules: If a character in the storyboard corresponds to an already generated character image, replace the character description with "the character in image n"; remove the original character's appearance descriptions such as hair color, skin color, and clothing from the storyboard to avoid conflicts with the reference image.
[0045] Output: Chinese prompt, including visuals and character rigging information.
[0046] 3.3 Based on the generated prompts, the system calls the image-to-image model to generate storyboard images.
[0047] Step 4: Video Compositing 4.1 The images from each storyboard are spliced together according to the time sequence of the storyboard script to generate a complete video sequence.
[0048] The synthesis method of the storyboard video of the present invention can achieve at least the following technical effects: (1) the image of the characters is consistent across shots; (2) the natural generation of multi-character interactive scenes is supported; (3) the repetitive work of creators in designing prompts is reduced; and (4) the controllability and automation of video synthesis are improved.
[0049] See Figure 2 A second aspect of the present invention provides a system for synthesizing storyboard videos, the system comprising: The script data processing module 10 is used to acquire the input script data, process the script data using a large language model, and generate a structured storyboard. The reference image determination module 20 is used to read the global character list from the storyboard and associate a reference image as the unique visual reference for each character in the global character list; The prompt word generation module 30 is used to traverse the storyboard script and generate target text-based image prompt words for each storyboard script based on the initial text-based image prompt words in the storyboard script. In the target text-based image prompt words, the text description of the character appearing in the initial text-based image prompt words is replaced with the reference identifier of the reference image associated with the character. Storyboard image generation module 40 is used to input the target text image prompt and the reference image of the character involved in the storyboard into the image generation model for each storyboard, and generate a storyboard image corresponding to the content of the storyboard. Storyboard video generation module 50 is used to synthesize multiple storyboard images into a storyboard video according to the order of the storyboard script.
[0050] In an optional embodiment of the second aspect of the present invention, the script data processing module includes: The data structuring and conversion unit is used to convert the script data into a structured data format containing a global character list, storyboard index, scene description, storyboard character information, and initial text-to-image prompts, based on a large language model and a preset instruction template, to obtain the storyboard script.
[0051] In an optional embodiment of the second aspect of the invention, the reference image for each character in the character list is determined by the following method: Based on the character descriptions extracted from the character list, a text image generation model is invoked to generate the image. Alternatively, a match can be made from a pre-defined character image library based on the character's information, or the character can be directly specified.
[0052] In an optional embodiment of the second aspect of the present invention, the reference identifier is a specific placeholder or code that can be recognized by the image generation model, and the reference identifier is used to instruct the model to render the associated character in the storyboard image in a specified area of the storyboard image according to the reference character in the corresponding reference image.
[0053] In an optional embodiment of the second aspect of the present invention, the prompt word generation module includes: The storyboard character quantity determination unit is used to determine whether the currently processed storyboard contains at least two characters in the character list; The text description replacement unit is configured to replace the text description of a character with a reference identifier of the reference image associated with the character only when the currently processed storyboard contains at least two characters in the character list.
[0054] In an optional embodiment of the second aspect of the invention, the target text image prompt includes stylized prompt phrases for defining the overall artistic style of the scene, so as to ensure that all the storyboard images have a consistent style.
[0055] In an optional embodiment of the second aspect of the present invention, the storyboard video generation module includes: The dynamic effects fusion unit is used to combine multiple generated storyboard images with image-generated video prompts used to describe dynamic effects, and input them together into the video generation model to generate a sequence of video clips containing dynamic transition effects. The video splicing unit is used to splice the generated video segment sequence to obtain a storyboard video.
[0056] Figure 3 This is a schematic diagram of a video generation device according to an embodiment of the present invention. The video generation device can vary significantly due to differences in configuration or performance, and may include one or more processors 60 (central processing units, CPUs) (e.g., one or more processors) and memory 70, and one or more storage media 80 (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the video generation device. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the video generation device.
[0057] The video generation device of this invention may further include one or more power supplies 90, one or more wired or wireless network interfaces 100, one or more input / output interfaces 110, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated video generation device structure does not constitute a limitation on the video generation device and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0058] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the method for synthesizing the storyboard video.
[0059] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system or system / unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0060] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0061] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for synthesizing storyboard videos, characterized in that, include: The script data is obtained from input, and the script data is processed using a large language model to generate a structured storyboard. Read the global character list from the storyboard script, and associate a reference image as the unique visual reference for each character in the global character list; The storyboard is traversed, and for each storyboard, a target text-based image prompt is generated based on the initial text-based image prompt in the storyboard. In the target text-based image prompt, the text description of the character appearing in the initial text-based image prompt is replaced with a reference identifier of the reference image associated with the character. For each of the storyboards, the target text-based image prompts and the reference images of the characters involved in the storyboard are input into the image generation model to generate a storyboard image corresponding to the content of the storyboard; According to the order of the storyboard script, multiple storyboard images are combined into a storyboard video.
2. The method for synthesizing storyboard videos according to claim 1, characterized in that, The process of acquiring input script data and processing the script data using a large language model to generate a structured storyboard includes: Using a large language model, the script data is converted into a structured data format containing a global character list, storyboard index, scene description, storyboard character information, and initial text-to-image prompts according to a preset instruction template, thus obtaining the storyboard script.
3. The method for synthesizing storyboard videos according to claim 1, characterized in that, The reference image for each character in the character list is determined in the following manner: Based on the character descriptions extracted from the character list, a text image generation model is invoked to generate the image. Alternatively, a match can be made from a pre-defined character image library based on the character's information, or the character can be directly specified.
4. The method for synthesizing storyboard videos according to claim 1, characterized in that, The reference identifier is a specific placeholder or code that can be recognized by the image generation model. The reference identifier is used to instruct the model to render the associated character in the storyboard image in a specified area of the storyboard image based on the reference character in the corresponding reference image.
5. The method for synthesizing storyboard videos according to claim 1, characterized in that, The step of replacing the text description of the character appearing in the initial text-based image prompt with a reference identifier of the reference image associated with the character in the target text-based image prompt includes: Determine whether the currently processed storyboard contains at least two characters from the character list; The replacement of the text description of a character with a reference identifier of the reference image associated with that character is performed only if the currently processed storyboard contains at least two characters from the character list.
6. The method for synthesizing storyboard videos according to claim 1, characterized in that, The target text image prompts include stylized prompt phrases used to define the overall artistic style of the scene, ensuring stylistic consistency across all storyboard images.
7. The method for synthesizing storyboard videos according to claim 1, characterized in that, The step of combining multiple storyboard images into a storyboard video according to the order of the storyboard script includes: The generated multiple storyboard images are combined with the image-generated video prompts in the storyboard script used to describe dynamic effects, and then input into the video generation model to generate a sequence of video clips containing dynamic transition effects. The generated video segment sequence is spliced together to obtain a storyboard video.
8. A system for synthesizing storyboard videos, characterized in that, The system for synthesizing the storyboard video includes: The script data processing module is used to acquire input script data, process the script data using a large language model, and generate a structured storyboard. A reference image determination module is used to read a global character list from the storyboard script and associate a reference image as a unique visual reference for each character in the global character list; The prompt word generation module is used to traverse the storyboard script and generate target text-based image prompt words for each storyboard script based on the initial text-based image prompt words in the storyboard script. In the target text-based image prompt words, the text description of the character appearing in the initial text-based image prompt words is replaced with the reference identifier of the reference image associated with the character. The storyboard image generation module is used to input the target text image prompt and the reference image of the character involved in the storyboard into the image generation model for each storyboard, and generate a storyboard image corresponding to the content of the storyboard. The storyboard video generation module is used to synthesize multiple storyboard images into a storyboard video according to the order of the storyboard script.
9. A video generation device, characterized in that, The video generation device includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor invokes the instructions in the memory to cause the video generation device to perform the method for compositing storyboard videos as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the method for synthesizing the storyboard video as described in any one of claims 1-7.