Short video intelligent generation method and device, storage medium and program product

By using a large language model to process creative text to generate story outlines and extract scene features, the entire short video production process is automated, solving the problems of complex operation and low efficiency of existing tools, and improving the efficiency and content coherence of short video production.

CN122053934APending Publication Date: 2026-05-15BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing short video production tools are complex and time-consuming to operate. AI-assisted tools cannot achieve an end-to-end closed loop from idea to finished product, resulting in poor content coherence, difficulty in improving production efficiency, and a high barrier to entry for short video production.

Method used

By processing user-generated creative text using a large language model, a story outline text is generated and key features of the storyboard are extracted. Based on the storyboard script, storyboard images and videos are generated, achieving fully automated short video generation.

Benefits of technology

It lowers the barrier to entry for short video production, improves creative efficiency, ensures the logical coherence and quality of generated content, simplifies multi-stage collaborative operations, and meets the needs of efficient creation and rapid delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053934A_ABST
    Figure CN122053934A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a short video intelligent generation method and device, a storage medium and a program product. The method comprises the steps that a creative text containing at least one sentence of demand content of a user is preprocessed and then input into a large language model, and a story infarction text output by the large language model is obtained; extracting a split key feature from the story infarction text, and generating a split script based on the split key feature; on the basis of each split content contained in the split script, generating an adaptive text picture prompt word and a picture video prompt word; generating a split picture based on the split content and the corresponding text picture cue word; generating a split video based on the split picture and the corresponding picture generation video cue word; and combining the generated split videos to obtain a target short video and feeding back the target short video to the user terminal. According to the method, through full-process automatic conversion from the creative text to the short video, the production threshold is greatly reduced, the creation efficiency is improved, and meanwhile, the logic continuity and the quality stability of the generated content are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of natural language processing technology, and in particular to a method, apparatus, storage medium, and program product for intelligent short video generation. Background Technology

[0002] With the development of mobile internet and the popularization of short video platforms, short videos have become a mainstream medium for information dissemination and creation, with a continuously growing user base. Individuals and SMEs have a strong demand for short videos, and their application scenarios cover multiple fields such as daily life, science popularization, and promotion. Existing production technology tools are mainly traditional single-function module tools (such as scripting, editing, and drawing tools) and some AI-assisted tools. The former requires users to manually switch to complete the entire process, while the latter can only automate a single step (such as text-to-image conversion and intelligent editing).

[0003] Traditional tools lead to a fragmented production process, which is complex and time-consuming. They not only require high levels of professional skills from users, but AI-assisted tools also cannot achieve an end-to-end closed loop from idea to finished product. This results in poor content coherence and difficulty in improving production efficiency, making the threshold for short video production high and hindering the development of the industry. Summary of the Invention

[0004] In view of this, the present disclosure provides a method, apparatus, storage medium, and program product for intelligent short video generation, which can significantly reduce the production threshold and improve the creation efficiency through fully automated conversion of creative text into short videos, while ensuring the logical coherence and quality stability of the generated content.

[0005] In a first aspect, this disclosure provides a method for intelligently generating short videos, employing the following technical solution: Obtain the creative text input by the user, wherein the creative text contains at least one sentence of requirement content; The creative text is preprocessed and then input into a large language model to obtain the story synopsis text output by the large language model; Extract key features of storyboard scenes from the story synopsis text, and generate storyboard scripts based on these key features; Based on the content of each scene in the storyboard, generate appropriate text-to-image prompts and image-to-video prompts; Based on the storyboard content and the corresponding text-to-image prompts, generate storyboard images; Based on the storyboard images and corresponding image-generated video prompts, a storyboard video is generated; The generated storyboard videos are merged to obtain the target short video, which is then sent back to the user's terminal.

[0006] Optionally, the large language model includes an input layer, an inference layer, an iterative validation layer, and an output layer. The step of preprocessing the creative text and inputting it into the large language model to obtain the story synopsis text output by the large language model includes: The creative text is preprocessed to generate a creative feature vector; The story outline generation rules are transformed into constraint prompts that can be recognized by the large language model; The input layer performs semantic embedding and tensor mapping processing on the creative feature vector and the prompt words to generate a semantic feature tensor. The reasoning layer, based on the semantic feature tensor, performs plot expansion and structured organization through context modeling and text generation algorithms to obtain the initial story outline content; The iterative verification layer performs multiple rounds of verification on the story outline content and generates targeted adjustment prompts. The reasoning layer optimizes the current story outline based on the targeted adjustment prompts to obtain the iterated story outline. After multiple rounds of verification are completed, the output layer converts the story synopsis content into story synopsis text and outputs it.

[0007] Optionally, the iterative verification layer performs multiple rounds of verification on the story outline content and generates targeted adjustment prompts, including: The iterative verification layer calls the basic constraint rules at the story level to perform compliance checks on the story outline content and generates targeted adjustment prompts for the first round of verification. Receive personalized adjustment instructions input by the user, and generate targeted adjustment prompts for the second round of verification based on the personalized adjustment instructions; The detailed optimization rules at the production level are invoked to perform short video adaptation detection on the story outline content, generating targeted adjustment prompts for the third round of verification.

[0008] Optionally, the step of extracting key storyboard features from the story synopsis text and generating a storyboard script based on the key storyboard features includes: The story synopsis text is structured and parsed to extract key features of the storyboard, including scene transition points, global plot nodes, and global key actions. Using the scene switching point as the dividing boundary, the story outline text is split into several storyboard segments, each segment corresponding to a continuous scene; Extract local plot nodes and local key actions from global plot nodes and global key actions to adapt to the continuous scenes within the scene segment unit; Based on the local plot nodes and the local key actions, multiple storyboard contents are generated; Based on the story complexity of the storyboard content and the target duration of the short video, the storyboard content is merged or split; All the revised storyboard content was sorted and integrated into a standardized storyboard script according to the narrative logic.

[0009] Optionally, the step of generating suitable text-based image prompts and image-based video prompts based on each scene content included in the storyboard includes: Extract static and dynamic elements from each storyboard; Add compositional and dynamic details to each storyboard; Based on the static elements and compositional details of each storyboard, generate text-based image prompts that are adapted to the storyboard content; Based on the dynamic elements and dynamic scene details of each storyboard content, generate image-generated video prompts that are adapted to the storyboard content.

[0010] Optionally, generating storyboard images based on the storyboard content and corresponding text-based image prompts includes: The storyboard content and the text-based image prompts that match the storyboard content are input into the text-based image model, and the text-based image model outputs at least one storyboard image that matches the storyboard content.

[0011] Optionally, generating a storyboard video based on the storyboard images and corresponding image-generated video prompts includes: The storyboard content, the image-generated video prompts that match the storyboard content, and the storyboard images that match the storyboard content are input into the image-to-video model, and the image-to-video model outputs a storyboard video that matches the storyboard content.

[0012] Secondly, this disclosure also provides a short video intelligent generation system, which adopts the following technical solution: The creative text acquisition module is used to acquire the creative text input by the user, wherein the creative text contains at least one sentence of requirement content; The story synopsis acquisition module is used to preprocess the creative text and input it into the large language model to obtain the story synopsis text output by the large language model. The storyboard generation module is used to extract key storyboard features from the story synopsis text and generate storyboards based on the key storyboard features. The prompt word generation module is used to generate appropriate text-based image prompt words and image-based video prompt words based on the content of each scene contained in the storyboard; The storyboard image generation module is used to generate storyboard images based on the storyboard content and corresponding text-based image prompts; The storyboard video generation module is used to generate storyboard videos based on the storyboard images and corresponding image-generated video prompts; The storyboard video merging module is used to merge the generated storyboard videos to obtain the target short video and send it back to the user terminal.

[0013] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform any of the above-described short video intelligent generation methods.

[0014] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the short video intelligent generation methods described above.

[0015] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0016] The short video intelligent generation method provided in this embodiment preprocesses and expands the simple creative text input by the user into a story outline using a large language model, and then extracts storyboard features to generate a storyboard script. This allows users without professional creative experience to quickly start short video production. At the same time, the storyboard script provides clear content guidance for subsequent steps, ensuring the coherence and rationality of the short video from story logic to visual presentation, and effectively improving the quality of the generated content.

[0017] By leveraging a standardized mechanism for generating text-to-image and image-to-video prompts, along with a step-by-step conversion process from storyboard images to storyboard videos, the production efficiency of short videos is significantly improved. The precise matching of storyboard content with corresponding prompts ensures that the generated images and videos align with the design requirements of the storyboard script. Finally, the final product is output by merging the storyboard videos, simplifying the complexity of multi-stage collaboration and quickly transforming user ideas into visual video content, thus meeting the demands for efficient creation and rapid delivery.

[0018] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating the short video intelligent generation method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating the method for obtaining story synopsis text provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating the targeted adjustment prompt generation method provided in this embodiment of the disclosure; Figure 4 A flowchart illustrating the storyboard generation method provided in this embodiment of the disclosure; Figure 5 This is a flowchart illustrating the method for generating text-based image prompts and image-based video prompts according to embodiments of this disclosure. Figure 6 A schematic diagram of the short video intelligent generation system provided in this embodiment of the disclosure; Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0021] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0022] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0023] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0024] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0025] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0026] Reference Figure 1 This disclosure provides a method for intelligent generation of short videos, including the following steps: S1: Obtain the creative text input by the user. The creative text must contain at least one sentence of the requirement. S2: Input the preprocessed creative text into the large language model to obtain the story outline text output by the large language model; S3: Extract key storyboard features from the story synopsis text, and generate storyboard scripts based on these key features; S4: Based on each scene content contained in the storyboard, generate appropriate text-to-image prompts and image-to-video prompts; S5: Generate storyboard images based on the storyboard content and corresponding text-to-image prompts; S6: Generate storyboard video based on storyboard images and corresponding image-generated video prompts; S7: Merge the generated storyboard videos to obtain the target short video and send it back to the user terminal.

[0027] The intelligent short video generation method disclosed herein automates the entire short video creation process from idea to finished product, significantly lowering the barrier to entry for short video production. By preprocessing simple creative text input by users and expanding it into a story outline using a large language model, and then extracting storyboard features to generate a storyboard script, users lacking professional creative experience can quickly start short video production. At the same time, the storyboard script provides clear content guidance for subsequent stages, ensuring the coherence and rationality of the short video from story logic to visual presentation, effectively improving the quality of the generated content.

[0028] By leveraging a standardized mechanism for generating text-to-image and image-to-video prompts, along with a step-by-step conversion process from storyboard images to storyboard videos, the production efficiency of short videos is significantly improved. The precise matching of storyboard content with corresponding prompts ensures that the generated images and videos align with the design requirements of the storyboard script. Finally, the final product is output by merging the storyboard videos, simplifying the complexity of multi-stage collaboration and quickly transforming user ideas into visual video content, thus meeting the demands for efficient creation and rapid delivery.

[0029] In S1, an agent is pre-configured, which works with various models to complete the entire process from a user's sentence to the generation of a short video. The agent's input receiving module obtains creative text, which is composed of a sentence input by the user, including but not limited to the core theme, key elements (characters, scenes, events) and style preferences.

[0030] In S2, refer to Figure 2 The flowchart illustrating the method for obtaining the story synopsis text shows the steps involved in "preprocessing the creative text and inputting it into a large language model to obtain the story synopsis text output by the large language model": S21: Preprocess the creative text to generate creative feature vectors; S22: Transform the story outline generation rules into constraint prompts that can be recognized by a large language model; S23: The input layer performs semantic embedding and tensor mapping on the creative feature vector and prompt words to generate a semantic feature tensor; S24: The reasoning layer, based on semantic feature tensors, performs plot expansion and structured organization through context modeling and text generation algorithms to obtain the initial story outline content; S25: The iterative verification layer performs multiple rounds of verification on the story outline content and generates targeted adjustment prompts; S26: The reasoning layer optimizes the current story outline based on the targeted adjustment prompts to obtain the iterative story outline. S27: After multiple rounds of validation are completed, the output layer will convert the story synopsis content into story synopsis text and output it.

[0031] In S21, the preprocessing of creative text includes multi-dimensional feature analysis and text feature combination. Multi-dimensional feature analysis is performed on the creative text to extract multi-dimensional text features. Since creative text supports any combination of core themes (e.g., "campus environmental protection public welfare"), key elements (characters: students, scenes: playground, events: waste recycling competition), and style preferences (e.g., "relaxing and healing"), feature analysis is performed using Natural Language Processing (NLP) technology, including entity parsing, intent recognition, and keyword weighting. Entity parsing refers to identifying and extracting core entities such as characters, scenes, and events from the text; intent recognition refers to determining the user's creative style preferences (e.g., humorous, science popularization, heartwarming); and keyword weighting refers to assigning high weights to core theme words (e.g., "environmental protection") and secondary weights to modifiers, forming a structured creative feature vector. The text features are then sequentially integrated according to dimensional priority to generate a creative feature vector. For example, based on the dimensional priority of theme, style, and core elements, the format of the creative feature vector is: {Theme: XXX, Core Elements: [Character A, Scene B, Event C], Style: XXX}.

[0032] In S22, the implementation scheme of combining prompt engineering with rule templates is adopted to transform unstructured story outline generation rules into structured constraint prompts that can be parsed by a large language model. First, the hierarchy and core requirements of the story synopsis generation rules are clarified. The basic constraints at the story level (plot completeness, logical coherence, and short video adaptability) and the detailed optimization rules at the production level are broken down into quantifiable and describable rule factors. Then, based on the instruction comprehension habits of the large language model, a standardized prompt word template is constructed. The template includes three parts: rule constraint prefix, core rule description, and output format requirements. The rule constraint prefix is ​​used to clarify the model's generation boundaries (e.g., "Strictly follow the following rules to generate a short video story synopsis:"). The core rule description transforms abstract rules into specific instructions (e.g., "The plot must include four parts: beginning, development, climax, and ending, with a total duration of 60-300 seconds"). The output format requirements clarify the structured presentation of the synopsis (e.g., "It must include the core plot outline, character settings, scene information, and emotional tone module"). Finally, all rule factors are integrated to generate constraint-type prompt words that adapt to the model's input specifications. The input of this step is the story synopsis generation rule set, and the output is the structured constraint-type prompt words. The core function of this step is to transform abstract rules into executable generation instructions for the model, ensuring that the subsequently generated content meets the preset requirements.

[0033] Taking an environmental-themed creative idea adapted for a 60-second short video as an example, the converted constraint prompts are: "Strictly follow the following rules to generate a 60-second short video story outline: 1. Plot completeness: It must include four parts: beginning (discovery of environmental problems), development (attempts to solve the problem), climax (key point in solving the problem), and ending (presentation of environmental achievements); 2. Logical coherence: The characters' actions must be strongly related to the campus scene and the environmental theme, with no logical gaps; 3. Short video adaptability: The total length must strictly adapt to 60 seconds, with the climax accounting for no less than 30% (18 seconds); 4. Production adaptability: It must include visually descriptive shots, concise dialogue (no more than 15 characters per sentence), and reasonable scene transitions (no more than 3 scenes); The output format must include four modules: core plot, character settings, scene information, and emotional tone, with a light and inspirational language style." In S23, the input layer of the large language model uses the Transformer architecture's Embedding Layer as its core implementation module, specifically responsible for the unified processing of creative feature vectors and constraint prompts. The input consists of the creative feature vectors generated in step S21 and the pre-transformed constraint prompts. First, the Embedding Layer transforms both types of input into fixed-dimensional word embedding vectors, where the creative feature vectors correspond to semantic feature embeddings, and the constraint prompts correspond to rule constraint embeddings. Next, a sinusoidal positional encoding algorithm adds sequence position information to the two types of embedding vectors, ensuring the model captures the sequential correlation of semantics and avoiding feature confusion. Subsequently, a dimension-wise alignment fusion algorithm is used to fuse the creative semantic embeddings and rule constraint embeddings element-wise, eliminating their format heterogeneity and strengthening the semantic binding between creativity and rules. Finally, the tensor mapping module transforms the fused embedding vectors into semantic feature tensors conforming to the inference layer's input specifications (dimension adapted to the Transformer decoder's input dimensions). The core function of this layer is to standardize and fuse the input data, providing the inference layer with unified, parsable structured data.

[0034] In S24, the initial story outline generation in the reasoning layer is built on the Transformer decoder architecture, relying on context modeling and autoregressive text generation capabilities to achieve initial expansion. The input is the semantic feature tensor output by the input layer. First, the tensor is used to perform contextual semantic modeling through a multi-head attention mechanism (the number of heads is set to 12, and the dimension of each head is 64 after dimensional splitting). This accurately captures the deep relationship between creative features and constraint rules, ensuring that the generated content fits the core idea and follows the preset rules. Subsequently, an autoregressive text generation process is performed using a bundle search decoding algorithm (with a bundle width of 5 to balance generation efficiency and content quality). During the generation process, the fit between the content and the constraint rules is checked in real time. If problems such as missing plot points or exceeding the time limit occur, a local regeneration mechanism is triggered. Finally, the plot is expanded and structured according to the preset format, and the initial story outline, which includes the core plot, main character settings, key scene information, and emotional tone, is output. This layer ensures semantic relevance through a multi-head attention mechanism, the bundle search algorithm improves the fluency and rationality of the generated content, and the local regeneration mechanism further strengthens the rule adaptability. Its core function is to transform the semantic feature tensor into a structured initial story outline that meets the requirements of the basic rules.

[0035] In S25, refer to Figure 3 The flowchart illustrating the method for generating targeted adjustment prompts shows that "the iterative verification layer performs multiple rounds of verification on the story outline content to generate targeted adjustment prompts" includes the following steps: S251: The iterative verification layer calls the basic constraint rules at the story level to perform compliance checks on the story outline content and generates targeted adjustment prompts for the first round of verification; S252: Receive personalized adjustment instructions input by the user, and generate targeted adjustment prompts for the second round of verification based on the personalized adjustment instructions; S253: Invoke the detailed optimization rules at the production level to perform short video adaptation detection on the story outline content, and generate targeted adjustment prompts for the third round of verification.

[0036] In S251, the first round of verification is triggered by the rule engine module calling a preset story-level basic constraint rule library. The rule library includes three core rules: plot completeness, logical coherence, and duration suitability. The rule engine module first structurally breaks down the current version's story synopsis content, extracting key elements such as the core plot, character settings, and duration allocation, and then performs compliance checks on each element. For plot completeness, it verifies whether it contains the four elements of "beginning-development-climax-ending". If any element is missing, the defect location and type are marked. For logical coherence, it verifies the correlation between character behavior and scenes / events through semantic similarity calculation and causal reasoning algorithms. If logical breaks occur, contradiction nodes are recorded. For duration suitability, it verifies whether the story synopsis content can support the generation of a short video with a total duration between 60-300 seconds, and whether the climax accounts for no less than 30%. If information overload leads to excessive length or insufficient information leads to excessive length, it clarifies the direction of content to be added or deleted to avoid content redundancy or insufficiency.

[0037] Based on the above verification results, the feedback prompt generation module transforms the defect information into standardized targeted adjustment prompts. These prompts must include three core components: verification dimension, defect description, and modification requirements. This ensures the inference layer can accurately pinpoint the problem and perform targeted optimizations. The input for this round of verification is the initial story outline, and the output is the first round of targeted adjustment prompts based on compliance defects. Its core function is to ensure the basic narrative compliance of the story outline.

[0038] In S252, the second round of verification is initiated after receiving the user's personalized adjustment instructions. It serves as an intermediary link between basic compliance and short video production adaptability, with the core function of personalized content customization. The rules engine module first performs intent parsing through a BERT pre-trained model, semantically decomposing the user's instructions to accurately extract core adjustment needs (such as "add plot twists," "strengthen character personality," "adjust style to warm and healing"), modification scope (plot module, character module, style module, etc.), and boundary constraints (such as "do not change the core theme," "do not adjust the determined duration allocation"). Subsequently, the decomposed needs are matched module by module with the current version's story outline to locate the target areas that need optimization. At the same time, it verifies whether the personalized needs conflict with the basic constraint rules that have been met in the first round (especially duration adaptability). If a conflict exists, the priority is clearly stated in the prompt (prioritizing compliance with basic rules).

[0039] The feedback prompt generation module generates structured, targeted adjustment prompts based on the requirements analysis results and conflict verification conclusions. These prompts must use an imperative style, clearly defining the priority of the requirements, modification details, and output standards. This ensures that personalized user needs are implemented while avoiding disruption of the already optimized framework. The input for this round is the current story outline and the user's personalized adjustment instructions. The output is a second round of targeted adjustment prompts adapted to the user's needs. Its core function is to align the story outline with user expectations and enhance the content's uniqueness.

[0040] In S253, the third round of verification is executed by the rule engine module, which calls the production-level refined and optimized rule library. The core revolves around the adaptability of short video production, specifically targeting three rules: "camera feel, concise dialogue, and reasonable scene transitions." This forms a clear distinction in level and dimension from the first round's "duration adaptability," focusing on the feasibility of turning a story outline into a short video production. The rules engine module breaks down and analyzes the current story outline from production dimensions. Regarding the visual appeal, it checks whether it includes visually identifiable descriptions (such as shot type, shooting angle, and image details). If the description is abstract and cannot be directly converted into a storyboard, it marks the content that needs to be visualized and provides suggestions for shot representation. Regarding the conciseness of dialogue, it uses text length statistics and semantic density analysis to verify whether character dialogue and narration conform to the short, concise characteristics of short videos. If the dialogue is lengthy or semantically repetitive, it clarifies the simplification ratio and the requirements for retaining core semantics. Regarding the rationality of scene transitions, it combines the results of time-appropriateness analysis to verify whether the number of scenes matches the short video length (e.g., no more than 3 scenes in a 60-second short video), and whether scene transitions fit the plot rhythm and whether there are any abrupt transitions. If not, it provides suggestions for scene deletion or reordering.

[0041] The feedback prompt generation module transforms the detection results into third-round targeted adjustment prompts, highlighting production implementation requirements, clarifying shot expression standards, dialogue simplification standards, and scene optimization directions. This ensures that the optimized content from the reasoning layer can be directly integrated into subsequent scriptwriting and editing stages. This round's input is the current story synopsis, and the output is targeted adjustment prompts focusing on production suitability. Its core function is to bridge the gap between creative generation and production implementation, lowering the subsequent production threshold.

[0042] This iterative verification layer adopts a dual-module architecture of "rule engine module + feedback prompt word generation module" to build a closed-loop verification mechanism. The rule engine module has a built-in basic constraint rule library at the story level and a refined optimization rule library at the short video production level. It is responsible for executing multi-dimensional verification logic. The feedback prompt word generation module generates structured and targeted adjustment prompt words based on the verification results through prompt word engineering algorithms. Overall, it realizes three rounds of progressive verification and prompt word generation.

[0043] In S26, the inference layer uses the Transformer decoder architecture to perform precise incremental optimization based on the targeted adjustment prompts, rather than regenerating the entire text. The input consists of the previous story synopsis and the targeted adjustment prompts generated in S25. First, the prompt parsing module extracts the core adjustment requirements, scope of modification, and constraints, locating the content modules to be optimized (such as plot modules and shot description modules). Then, a prompt-guided incremental modification algorithm is used to rewrite and optimize only the target modules. Simultaneously, a semantic consistency check algorithm compares the modified content with the overall logic and style of the original text to avoid semantic gaps or stylistic disjointness. During the optimization process, the requirements of the targeted adjustment prompts are strictly followed, balancing story-level compliance with production-level adaptability, ensuring that the iterated content addresses existing problems while maintaining overall integrity. This step takes the previous story synopsis and targeted adjustment prompts as input and outputs the iteratively optimized story synopsis, achieving progressive content improvement and gradually enriching the story synopsis to meet the user's personalized requirements.

[0044] In S27, the output layer employs a combined architecture of a format standardization module and an information completion module. Its core objective is to bridge the gap between creative generation and short video production. The input is the final story synopsis content, validated by the iterative verification layer. First, the information completion module adds auxiliary information for short video production, such as shot annotations (e.g., "close-up character expressions," "overhead shot of the entire scene"), suggestions for concise dialogue, and scene transition durations, enhancing the content's feasibility. Then, a text formatting algorithm organizes the content into a preset standardized format, arranged in the order of "core plot - character settings - scene information - emotional tone - duration allocation - production tips," eliminating format redundancy and ensuring a clear structure. Finally, a text encoding conversion module transforms the formatted content into standardized story synopsis text (supporting common formats such as TXT and DOCX) that can be directly integrated into subsequent short video script production and editing stages. This layer takes the final compliant story synopsis as input and outputs standardized story synopsis text; its core function is to make the generated content practical and standardized, connecting it to subsequent production processes.

[0045] For a large language model comprising an input layer, inference layer, iterative validation layer, and output layer, the training method employs a strategy of pre-training fine-tuning combined with hierarchical adaptation training to achieve accurate support for generating story outlines from creative text expansion. First, the focus is on basic pre-training and adaptation training of the input and inference layers. A training dataset is constructed based on massive amounts of short video story outline texts and creative text-story outline pairing data. Basic large language models (such as the GPT series and LLaMA series) are further pre-trained using this dataset. Through masked language modeling (MLM) tasks and autoregressive generation tasks, the model learns the narrative logic, plot structure, and element organization of short video stories. Regarding the semantic embedding and tensor mapping capabilities of the input layer, paired training samples containing creative feature vectors, constraint prompts, and semantic feature tensors are constructed. The embedding matrix of the input layer is optimized through contrastive learning, enabling the input layer to accurately integrate creative features and constraint rules into a unified semantic feature tensor. Meanwhile, to address the plot expansion and structured organization capabilities of the reasoning layer, a supervised fine-tuning task was designed. Using labeled "semantic feature tensor-structured story outline" data pairs as training samples, an autoregressive generation loss function was adopted to enable the reasoning layer to learn the generation path from fused features to a story outline that meets the requirements of plot completeness, logical coherence, and duration adaptability.

[0046] The subsequent focus is on the joint collaborative training of the iterative verification layer, inference layer, and output layer to enhance the interoperability and adaptability among these three layers. A five-element linked training dataset is constructed, comprising "initial story outline - rule verification defects - targeted adjustment prompts - optimized story outline - standardized text." The iterative verification layer, inference layer, and output layer are treated as a single collaborative training unit, and joint training is conducted end-to-end. For the iterative verification layer, the compliance and quality of the optimized story outline generated by the inference layer based on the targeted adjustment prompts are used as reward signals. Reinforcement learning enables the iterative verification layer to learn to accurately identify defects in the story outline and generate targeted adjustment prompts that guide the inference layer to optimize efficiently. For the inference layer, the targeted adjustment prompts output by the iterative verification layer are used as conditional constraints. Constraint loss terms are added to the supervised fine-tuning task to ensure that the inference layer can make targeted modifications to the current story outline based on the prompts, rather than regenerating it without any connection. For the output layer, the optimized story outline finally generated by the inference layer is used as input, and standardized format text is used as the supervision target. The format conversion capability is trained using a sequence labeling loss function. During training, a closed-loop feedback is formed among the three layers. The quality of the prompts in the iterative verification layer directly affects the optimization effect of the inference layer. The optimization results of the inference layer iteratively update the rule judgment model of the verification layer. The format standardization results of the output layer further verify the usability of the overall generated content. Finally, the three layers are collaboratively adapted to ensure that the model stably outputs standardized story summary text that meets the requirements.

[0047] In S3, refer to Figure 4The flowchart illustrating the storyboard generation method, which involves "extracting key storyboard features from the story synopsis text and generating a storyboard based on these features," includes the following steps: S31: Perform structured analysis on the story outline text and extract key features of the storyboard. Key features of the storyboard include scene transition points, global plot nodes, and global key actions. S32: Using scene transition points as dividing boundaries, the story outline text is split into several storyboard segments, each segment corresponding to a continuous scene; S33: Extract local plot nodes and local key actions from global plot nodes and global key actions to adapt to the local plot nodes and key actions of continuous scenes within the scene segment unit. S34: Generate multiple storyboard content based on local plot nodes and key local actions; S35: Based on the story complexity of the storyboard content and the target duration of the short video, merge or split the storyboard content; S36: Sort all the revised storyboard content according to the narrative logic and integrate it into a standardized storyboard script.

[0048] In S31, the intelligent agent also includes a storyboard generation module. The core of this module, a built-in NLP structured analysis submodule, relies on semantic segmentation algorithms and key information extraction models to extract key features for the storyboard. First, this submodule transforms the story outline text into a structured semantic tree. Using scene keyword recognition algorithms (such as a BERT-based named entity recognition model), it locates words or phrases indicating spatial transitions (e.g., "the camera pans to the playground," "back to the classroom," "the community square in the evening"), thus marking scene transition points. Second, based on narrative logic analysis algorithms, it breaks down the story outline into four parts: "beginning, development, climax, and ending," extracting the core events of each part (e.g., "proposing an environmental protection plan," "failed to transform waste," "winning an award for showcasing the finished product") as global plot nodes. Finally, through action verb extraction and character behavior association algorithms, it selects core behaviors that drive the plot forward (e.g., "distributing tools," "pasting labels," "lifting up the transformed flowerpot") as global key actions. After extraction, the submodule stores the three types of features in association, forming a three-dimensional feature mapping table of "scene-plot-action," providing data support for subsequent storyboard division.

[0049] In S32, based on the generated 3D feature mapping table, the scene segmentation submodule performs the division of storyboard segments. First, this submodule sorts all scene transition points according to narrative order, using adjacent scene transition points as start and end boundaries to segment the story synopsis text. If there are no explicit scene transition points at the beginning or end of the text, the default boundary is the beginning and end of the text. Second, each segmented text is assigned a scene attribute label (e.g., "classroom scene," "playground scene"), clearly defining each segment unit as corresponding to a continuous and independent spatial scene. Finally, each storyboard segment unit is associated with its corresponding scene in the 3D feature mapping table, ensuring that each segment unit matches its associated global plot node and global key action, forming a one-to-one correspondence between "scene segment" and "feature," avoiding deviations in subsequent local feature extraction.

[0050] In S33, the feature filtering and mapping submodule achieves precise adaptation from global features to local features. First, this submodule reads the scene attribute tags of the storyboard segment unit and, combined with the corresponding text content, filters out plot nodes and key actions belonging to the current scene from the 3D feature mapping table. Second, through a plot-scene matching degree calculation algorithm, it removes global plot nodes and key actions irrelevant to the current scene (e.g., removing the key action "distributing tools on the playground" from the "classroom scene" segment). Finally, the filtered features are refined and broken down, decomposing macro-level global plot nodes into specific local plot nodes adapted to the current scene (e.g., breaking down the global plot node "remodeling waste" into two local plot nodes under the "classroom scene": "discussing the renovation plan" and "drawing design sketches"). Simultaneously, global key actions are refined into local key actions (e.g., remodeling waste is refined into "cutting plastic bottles" and "pasting decorative stickers"), and the association between local plot nodes and local key actions is established to ensure that local features can support the generation of storyboard content for a single scene.

[0051] In S34, the storyboard content generation submodule strictly follows the language specifications of short video shots, transforming local features into standardized storyboard content. First, this submodule assigns 1-2 storyboards to each local plot node. Each storyboard corresponds to a core local key action. Then, according to the short video shot expression logic, basic elements are configured for each storyboard: the shot number is assigned sequentially according to the generation order of the storyboards (e.g., 1, 2, 3...), the shot type is matched according to the action attributes and plot importance (e.g., a long shot to show the whole scene, a close-up to show the action details, and a close-up to depict the character's expression), the shot duration is set to N seconds by default (e.g., 5 seconds), the storyboard duration of core plot nodes can be extended to M (e.g., 8 seconds), and the storyboard duration of transition plots is shortened to R seconds (e.g., 3 seconds). The screen description is generated based on the local key action and scene attributes to generate a concrete expression (e.g., "Close-up: Xiaolin is holding scissors and cutting a plastic bottle, with a focused expression"). The sound effects / dubbing prompts match the screen content (e.g., "Tool collision sound + narration: 'Every cut embodies creativity'"). The subtitle content extracts the core information of the screen (e.g., "The first step of creative transformation: cutting the bottle body"). The six key elements of each storyboard are integrated into a standardized storyboard entry, forming an initial set of storyboard content. Each entry has a unique scene number, which facilitates subsequent adjustments.

[0052] In S35, a submodule for optimizing the number of storyboards dynamically adjusts the number of storyboards to match the target duration and content complexity, updating the storyboard number synchronously during the adjustment process. First, this submodule calculates the total duration of the initial storyboard content (cumulative sum of single-scene durations) and compares it with the target duration of the short video. If the total duration exceeds the target duration and the story complexity is low (no excessive details), a storyboard merging mechanism is activated. This merges adjacent storyboards within the same scene with strong plot connections (e.g., merging the "cutting plastic bottles" and "pasting decorative stickers" storyboards into one, integrating visual descriptions, sound effects, and subtitles). After merging, the storyboard number of the merged storyboard is deleted, and subsequent storyboard numbers are sequentially updated. If the total duration is lower than the target duration and the story complexity is low, the submodule optimizes the number of storyboards. If the content is high (containing many details that can be refined), then the storyboard splitting mechanism is activated, splitting the storyboard corresponding to the core plot node into multiple detailed storyboards (e.g., splitting "showing the finished product and winning the award" into three storyboards: "holding the finished product and walking onto the stage", "raising the finished product and showing it", and "applause from the audience"). After splitting, the new storyboards are inserted with consecutive shot numbers (e.g., the original shot number 5 is split into 5, 6, and 7, and the subsequent storyboard shot numbers are moved to the next position). During the adjustment process, it is necessary to ensure that the core plot information is not lost when merging storyboards, and that the storyboard splitting does not disrupt the narrative logic, so as to achieve a precise match between the total storyboard duration and the target duration.

[0053] In S36, the storyboard script integration and output submodule is used to standardize the layout of storyboard content. This submodule reorganizes all adjusted storyboard content according to the narrative order of the story, ensuring that the storyboard order is consistent with the plot development logic and avoiding order confusion caused by merging or splitting. All storyboards are uniformly calibrated with shot numbers, and assigned consecutive shot numbers (such as 1, 2, 3...N) in the sorted order to eliminate possible gaps or repetitions in shot numbers during the adjustment process. According to the preset standardized storyboard script format, the shot number, shot size, shot duration, visual description, sound effects / dubbing prompts, and subtitle content of each storyboard item are arranged in sequence to generate a structured storyboard script file, which can be directly exported to a format that can be used for short video shooting or editing (such as table format, TXT format).

[0054] In S4, refer to Figure 5 The flowchart illustrating the methods for generating text-based image prompts and image-based video prompts shows that "generating appropriate text-based image prompts and image-based video prompts based on the content of each scene in the storyboard" includes the following steps: S41: Extract static and dynamic elements from each storyboard content; S42: Add compositional and dynamic details to each storyboard; S43: Based on the static elements and compositional details of each storyboard, generate text-based image prompts that are adapted to the storyboard content; S44: Generate image-based video prompts that are adapted to the content of each storyboard, based on the dynamic elements and dynamic scene details of each storyboard.

[0055] In S41, based on BERT's semantic parsing algorithm and regular expression matching rules, static and dynamic elements are accurately separated from individual storyboard entries in the storyboard script. First, for static elements, four core information categories are extracted from the storyboard entries: shot type (long shot / medium shot / close-up / extreme close-up), subject (characters, objects, scenes, etc.), scene attributes (e.g., classroom, playground, cozy wooden house), and preset style tendency (e.g., cartoon, realistic, Chinese style). These elements are crucial to the basic visual form of the scene. For dynamic elements, three types of information are extracted from the storyboard entries: shot duration (e.g., 5 seconds, 8 seconds), sound effects / dubbing association information (e.g., tool collision sounds, narration), and plot rhythm annotations (e.g., climax, transition). The sound effects / dubbing association information can indirectly infer the speed of the dynamic rhythm of the scene, while the plot rhythm annotations provide a basis for subsequent addition of dynamic details. After extraction, the two types of elements are bound to the corresponding storyboard shot number and stored in a database linking shot number and elements, ensuring that the elements of each storyboard are traceable and retrievable.

[0056] In S42, based on the basic elements extracted in step 1, and combined with the built-in knowledge base and intelligent matching algorithm, the composition details and dynamic scene details are supplemented to realize the visualization of the prompt words. First, regarding the supplementation of composition details, the corresponding composition method is matched according to the shot size and scene attributes (such as matching a "rule of thirds composition to highlight the entire playground" for a distant playground scene, and matching a "central composition to focus on the character's hand movements" for a close-up of a person's actions). At the same time, details such as lighting conditions (such as warm yellow natural light, cool white indoor lighting), color tone (such as fresh low-saturation tones, retro high-saturation tones), and image texture (such as delicate and smooth, grainy film style) are supplemented. The supplementation rules follow the principle that the shot size determines the composition framework, and the scene determines the light and shadow colors. To supplement details of dynamic scenes, the corresponding camera movement and amplitude are marked and matched according to the plot rhythm of the storyboard (e.g., "the camera slowly advances with moderate amplitude" for climax scenes, and "the camera slightly pans with gentle amplitude" for transition scenes). At the same time, details of transition effects (e.g., fade-in and fade-out, hard cut without transition) and dynamic effects of the scene (e.g., the smoothness of character movements, slight shaking of objects) are added. If there are key action descriptions in the storyboard (e.g., "cutting plastic bottles" or "lifting up the finished product"), the dynamic details of the actions are further added (e.g., "the jarring feeling of scissors cutting plastic bottles" or "the slow trajectory of the arm being raised").

[0057] When supplementing compositional and dynamic details, the quality parameters must be set to accommodate both text-based images and image-based videos, achieving both uniformity and differentiation between static and dynamic parameters. For compositional details serving the text-based image stage, the corresponding static quality parameters must strictly match the short video playback specifications, uniformly setting the screen resolution to 1080×1920 pixels. Then, the clarity standards are differentiated according to the plot priority of the storyboard scenes. For example, core plot storyboard scenes are configured with 4K high-definition parameters, while transitional plot storyboard scenes are configured with 1080P high-definition parameters. Simultaneously, image quality requirements are supplemented; for example, realistic style storyboard scenes are given parameters for delicate, noise-free detail, while cartoon-style storyboard scenes are given parameters for smooth, jagged lines, ensuring that the visual presentation of static storyboard images accurately matches the storyboard's expectations. For dynamic image details serving the image-to-video stage, the corresponding dynamic quality parameters should adhere to the 1080×1920 pixel resolution baseline, supplemented with specific parameters such as frame rate (30fps) and camera motion smoothness (no stuttering, no screen tearing). Furthermore, the dynamic quality parameters must be deeply correlated with the storyboard length; for example, short storyboards under 5 seconds should match parameters for fast camera motion without motion blur, while long storyboards over 8 seconds should match parameters for slow camera motion without blurring. Ultimately, this ensures that the quality parameters support both high-definition presentation of still images and smooth playback of dynamic videos. After supplementation, compositional details are integrated with static elements, and dynamic image details are integrated with dynamic elements to form a complete set of elements.

[0058] In S43, the integrated static elements and compositional details are transformed into structured, parsable prompts according to the input specifications of the text-based image model. First, the set of static elements and compositional details bound to the current shot number are retrieved and sorted according to the standardized format of "style + shot size + subject + action + scene + compositional details + quality requirements". Then, a natural language refining algorithm is used to splice the sorted elements into a smooth and coherent statement, avoiding mechanical piling up and deleting redundant expressions to ensure that the prompts are concise and accurate. The prompts are then bound to the shot number again to generate the final text-based image prompts. For example, for the storyboard of "close-up, young Xiaolin, cutting plastic bottles, classroom scene", the generated prompts are: "Realistic style, close-up, young Xiaolin holding scissors cutting plastic bottles, warm classroom scene, warm yellow natural light, central composition focusing on hand movements, plastic bottle edges curled, debris scattered on the table, 1080×1920 pixels, 4K high definition, detailed picture, 9:16 aspect ratio."

[0059] In S44, the integrated dynamic elements and dynamic scene details are transformed into standardized prompts that control the movement and effects of video shots. First, the set of dynamic elements and dynamic scene details bound to the current shot number are retrieved, and core information such as shot duration and plot rhythm is extracted to determine the core framework of the dynamic prompts. Then, following the format of "shot movement method + movement amplitude + transition effect + duration matching constraint + quality requirements", the dynamic scene details are filled into the framework. Among them, the shot movement method and amplitude must match the plot rhythm (e.g., a climax corresponds to "slow camera movement with moderate amplitude"), the transition effect must match the scene attributes (e.g., indoor scene switching corresponds to "fade in and fade out"), and the duration matching constraint must clearly state "the shot movement duration is consistent with the shot duration". At the same time, quality requirements (e.g., resolution 1080×1920 pixels, frame rate 30fps, no stuttering) are added and bound to the shot number to generate the final image-generated video prompts. For example, for the above storyboard, the generated prompt is: "The camera zooms in and out slightly, with no hard cuts in the transitions. The duration of the camera movement matches the storyboard duration of 5 seconds. The resolution is 1080×1920 pixels, 30fps, and the video is smooth without any stuttering."

[0060] This process enables the precise conversion of storyboard content into dual-modal cue words for both text-based images and image-based videos. It generates text-based image cue words that meet visual style and image quality requirements based on static elements and compositional details, ensuring a high degree of alignment between static storyboard images and the intended storyboard. Simultaneously, it generates image-based video cue words that control camera movement and transition rhythm based on dynamic elements and image details, ensuring that the subsequently generated storyboard video clips are of matching length and dynamically smooth. Ultimately, this achieves end-to-end adaptation from "storyboard script → static image → dynamic video," improving the accuracy of content generation and implementation efficiency.

[0061] In S5, the storyboard content and the text-to-image prompts that match the storyboard content are input into the text-to-image model, and the text-to-image model outputs at least one storyboard image that matches the storyboard content.

[0062] The intelligent agent also includes a storyboard image generation module. The generation of storyboard images is accomplished by calling a mature text-based image model interface. Supported text-based image models include Stable Diffusion, Midjourney, and DALL·E3. The storyboard image generation module performs format validation on the text-based image prompts that are bound to each storyboard content, ensuring that the prompts contain core elements such as style, shot size, subject, action, scene, details, and quality requirements. Then, according to the storyboard shot number, the prompts are input into the selected text-based image model in batches or sequentially. The text-based image model generates at least one corresponding storyboard image for each storyboard content based on the style parameters (user presets or the intelligent agent's default styles such as cartoon, realistic, and Chinese style) and quality parameters (1080×1920 pixel short video adaptation specifications) in the prompts. During the generation process, the image fidelity can be controlled by adjusting the model's generation parameters (such as sampling steps and CFG scaling values) to ensure that the visual content of the output image highly matches the storyboard script description.

[0063] Meanwhile, the storyboard generation module has a built-in storyboard quality verification submodule. It uses an image semantic matching algorithm to compare the core elements of the generated image with the storyboard script. If there are problems such as missing main subject, style deviation, or insufficient resolution, the verification submodule will automatically optimize the text-based image prompts (such as strengthening detailed descriptions and adjusting style weights) and re-call the text-based image model to generate new storyboard images. The verification and regeneration process is repeated until the output storyboard images fully meet the adaptation requirements of the storyboard content.

[0064] In S6, the storyboard content, image-generated video prompts that match the storyboard content, and storyboard images that match the storyboard content are input into the image-to-video model, and the image-to-video model outputs a storyboard video that matches the storyboard content.

[0065] The intelligent agent also includes a storyboard video generation module, which is responsible for generating the storyboard video. This module calls mainstream image-to-video models to convert static storyboard images into dynamic video clips. Supported image-to-video models include Pika Labs, Runway Gen-2, and Stable Video Diffusion. The module first performs format validation on the storyboard images and image-to-video prompts that are bound to the storyboard content to ensure that the prompts contain core dynamic elements such as camera movement, movement amplitude, transition effects, and duration constraints. Then, the storyboard images are used as the visual base and the image-to-video prompts are used as dynamic control instructions, which are simultaneously input into the selected image-to-video model. Based on the dynamic parameters in the prompts (such as fade-in / fade-out transitions, slight scaling and panning, and camera movement duration), the image-to-video model gives the static storyboard images dynamic effects that match the requirements of the storyboard script, and overlays the pre-set subtitle content in the storyboard script to generate storyboard video clips that are adapted to the storyboard content.

[0066] At the same time, the storyboard video generation module automatically calls speech synthesis models (such as TTS and Azure Speech) to generate corresponding dubbing audio based on the sound effects / dubbing prompts of the storyboard script, or matches the preset background sound effects library to select suitable ambient sounds and background music, and synchronizes the audio with the generated storyboard video clips in terms of duration to ensure that the audio playback rhythm and the video dynamics are perfectly matched. The final output storyboard video clips strictly follow the duration set in the storyboard script, and the content and dynamic effects of the screen accurately match the expectations of the storyboard.

[0067] In S7, the agent's video merging module reads all independent storyboard video segments in the order of the storyboard script's shot number, automatically adding transition effects such as flash white, blur gradient, and push-pull transitions to adjacent segments. The transition duration is set to 0.1-0.3 seconds by default and can be customized by the user. Then, the merged video undergoes end-to-end optimization processing, including audio volume equalization adjustment, uniform calibration of screen resolution to short video adaptation specifications, and format standardization conversion to MP4 format. Finally, a complete target short video is generated and fed back to the user terminal.

[0068] In summary, this solution significantly lowers the barrier to entry for short video production and dramatically improves creative efficiency. Users do not need professional skills such as scriptwriting, storyboard design, or video editing; they only need to input a single creative idea, and the intelligent agent will automate the entire process from idea expansion to storyboard generation, material production, and video compositing. Furthermore, leveraging the advantages of batch generation of storyboard prompts and diagrams, seamless integration of various functional modules, and rapid AI model computation, the traditional production cycle of several hours or even days is shortened to minutes, fully meeting the lightweight creation needs of ordinary users, self-media creators, and small and medium-sized enterprises.

[0069] This technical solution outputs short video content with strong coherence and consistency, as well as high flexibility and adaptability. The intelligent agent operates based on the same creative concept throughout the entire process, ensuring that the core theme, style, and logic of the story outline, storyboard, storyboard diagrams, and final video remain consistent, effectively solving the content discontinuity problem caused by switching between existing tools. It also supports style customization, multi-round interactive modifications, and multilingual input. Users can add style commands such as suspense or healing themes, and can also provide modification suggestions for the story outline, storyboard, or storyboard diagrams. The intelligent agent will iteratively optimize based on feedback, adapting to the creative needs of various scenarios such as life recording, knowledge popularization, and product promotion.

[0070] Reference Figure 6 This disclosure provides a short video intelligent generation system, including: The creative text acquisition module 101 is used to acquire the creative text input by the user. The creative text must contain at least one sentence of the requirement content. The story outline acquisition module 102 is used to input the preprocessed creative text into the large language model and obtain the story outline text output by the large language model. Storyboard generation module 103 is used to extract key storyboard features from the story outline text and generate storyboards based on the key storyboard features. The prompt generation module 104 is used to generate appropriate text-based image prompts and image-based video prompts based on the content of each storyboard included in the storyboard script. Storyboard image generation module 105 is used to generate storyboard images based on storyboard content and corresponding text-to-image prompts; Storyboard video generation module 106 is used to generate storyboard videos based on storyboard images and corresponding image-generated video prompts; The storyboard video merging module 107 is used to merge the generated storyboard videos to obtain the target short video and send it back to the user terminal.

[0071] The various variations and specific examples of the short video intelligent generation method provided above are also applicable to the short video intelligent generation system provided in this disclosure. Through the foregoing detailed description of the short video intelligent generation method, those skilled in the art can clearly understand the implementation method of the short video intelligent generation system. For the sake of brevity, they will not be described in detail here.

[0072] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0073] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the short video intelligent generation method described in the foregoing embodiments of this disclosure.

[0074] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0075] like Figure 7 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 7 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0076] like Figure 7 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0077] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 7A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0078] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the short video intelligent generation method of embodiments of this disclosure are performed.

[0079] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0080] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the short video intelligent generation methods described in the foregoing embodiments of the present disclosure are performed.

[0081] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0082] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0083] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0084] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0085] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0086] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0087] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0088] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0089] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for intelligently generating short videos, characterized in that, include: Obtain the creative text input by the user, wherein the creative text contains at least one sentence of requirement content; The creative text is preprocessed and then input into a large language model to obtain the story synopsis text output by the large language model; Extract key features of storyboard scenes from the story synopsis text, and generate storyboard scripts based on these key features; Based on the content of each scene in the storyboard, generate appropriate text-to-image prompts and image-to-video prompts; Based on the storyboard content and the corresponding text-to-image prompts, generate storyboard images; Based on the storyboard images and corresponding image-generated video prompts, a storyboard video is generated; The generated storyboard videos are merged to obtain the target short video, which is then sent back to the user's terminal.

2. The short video intelligent generation method according to claim 1, characterized in that, The large language model includes an input layer, an inference layer, an iterative validation layer, and an output layer. The process of preprocessing the creative text and inputting it into the large language model to obtain the story synopsis text output by the large language model includes: The creative text is preprocessed to generate a creative feature vector; The story outline generation rules are transformed into constraint prompts that can be recognized by the large language model; The input layer performs semantic embedding and tensor mapping processing on the creative feature vector and the prompt words to generate a semantic feature tensor. The reasoning layer, based on the semantic feature tensor, performs plot expansion and structured organization through context modeling and text generation algorithms to obtain the initial story outline content; The iterative verification layer performs multiple rounds of verification on the story outline content and generates targeted adjustment prompts. The reasoning layer optimizes the current story outline based on the targeted adjustment prompts to obtain the iterated story outline. After multiple rounds of verification are completed, the output layer converts the story synopsis content into story synopsis text and outputs it.

3. The short video intelligent generation method according to claim 1, characterized in that, The iterative verification layer performs multiple rounds of verification on the story outline content, generating targeted adjustment prompts, including: The iterative verification layer calls the basic constraint rules at the story level to perform compliance checks on the story outline content and generates targeted adjustment prompts for the first round of verification. Receive personalized adjustment instructions input by the user, and generate targeted adjustment prompts for the second round of verification based on the personalized adjustment instructions; The detailed optimization rules at the production level are invoked to perform short video adaptation detection on the story outline content, generating targeted adjustment prompts for the third round of verification.

4. The short video intelligent generation method according to claim 1, characterized in that, The step of extracting key storyboard features from the story synopsis text and generating a storyboard script based on the key storyboard features includes: The story synopsis text is structured and parsed to extract key features of the storyboard, including scene transition points, global plot nodes, and global key actions. Using the scene switching point as the dividing boundary, the story outline text is split into several storyboard segments, each segment corresponding to a continuous scene; Extract local plot nodes and local key actions from global plot nodes and global key actions to adapt to the continuous scenes within the scene segment unit; Based on the local plot nodes and the local key actions, multiple storyboard contents are generated; Based on the story complexity of the storyboard content and the target duration of the short video, the storyboard content is merged or split; All the revised storyboard content was sorted and integrated into a standardized storyboard script according to the narrative logic.

5. The short video intelligent generation method according to claim 1, characterized in that, The process of generating appropriate text-based image prompts and image-based video prompts based on each scene content contained in the storyboard includes: Extract static and dynamic elements from each storyboard; Add compositional and dynamic details to each storyboard; Based on the static elements and compositional details of each storyboard, generate text-based image prompts that are adapted to the storyboard content; Based on the dynamic elements and dynamic scene details of each storyboard content, generate image-generated video prompts that are adapted to the storyboard content.

6. The short video intelligent generation method according to claim 5, characterized in that, The process of generating storyboard images based on the storyboard content and corresponding text-based image prompts includes: The storyboard content and the text-based image prompts that match the storyboard content are input into the text-based image model, and the text-based image model outputs at least one storyboard image that matches the storyboard content.

7. The short video intelligent generation method according to claim 5, characterized in that, The step of generating a storyboard video based on the storyboard images and corresponding image-generated video prompts includes: The storyboard content, the image-generated video prompts that match the storyboard content, and the storyboard images that match the storyboard content are input into the image-to-video model, and the image-to-video model outputs a storyboard video that matches the storyboard content.

8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the short video intelligent generation method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the short video intelligent generation method according to any one of claims 1-7.

10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the short video intelligent generation method according to any one of claims 1-7.