Short drama creation method, device, equipment, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]但目前的短剧创作多由人工完成,依赖个人经验和团队协作,往往存在效率低、结构不稳定、质量难以保证等问题
[0010]The short drama creation solution provided in this embodiment first obtains the original plot text and short drama requirement parameters; based on the original plot text and short drama requirement parameters, a script outline is generated to create episode plans, and each episode script is generated based on the creative evidence provided by a preset reference script template library and each episode plan; then, based on each episode script and a preset director's script design template library and director's script content template library, a complete director's script corresponding to each episode script is generated; next, based on each complete director's script and a preset multi-level cascading video generation mechanism, each episode short video is generated; finally, based on each complete director's script, the corresponding episode short videos are edited, and the target short drama is generated based on the edited episode short videos. This solution significantly improves the efficiency and quality of short drama creation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the professionalism and filmability of short drama creation, making it suitable for large-scale short drama production.
Smart Images

Figure CN121509776B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence technologies such as large language models, intelligent agents, video generation, and script generation, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for creating short dramas. Background Technology
[0002] With the explosive growth of short video content, vertical short dramas, as an emerging content format, are becoming a key focus for the film, television, and internet content industries. Short drama creation is characterized by short production cycles, fast-paced content, and strong emotional conflicts, placing higher demands on the structural rationality and pacing control of scriptwriting.
[0003] However, most short drama creation is currently done manually, relying on personal experience and teamwork, which often results in problems such as low efficiency, unstable structure, and difficulty in guaranteeing quality. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for creating short dramas.
[0005] In a first aspect, this disclosure proposes a method for creating short dramas, including: obtaining original plot text and short drama requirement parameters; generating episode plans based on script outlines generated from the original plot text and short drama requirement parameters, and generating episode scripts based on creative evidence provided by a preset reference script template library and episode plans; generating complete director scripts corresponding to each episode script based on each episode script and a preset director script design template library and director script content template library; wherein, the reference script templates, director script design templates, and director script content templates are all obtained by templated processing of the corresponding content of historical film and television works with actual scores exceeding preset scores; generating short videos for each episode based on each complete director script and a preset multi-level cascading video generation mechanism; editing the corresponding short videos for each episode based on each complete director script, and generating a target short drama based on the edited short videos for each episode.
[0006] Secondly, this disclosure proposes a short drama creation device, comprising: an input information acquisition unit configured to acquire original plot text and short drama requirement parameters; a script creation unit configured to generate episode plans based on a script outline generated from the original plot text and short drama requirement parameters, and to generate episode scripts based on creative evidence provided by a preset reference script template library and episode plans; a director script generation unit configured to generate complete director scripts corresponding to each episode script based on each episode script and a preset director script design template library and director script content template library; wherein the reference script templates, director script design templates, and director script content templates are all obtained by templated processing of corresponding content from historical film and television works with actual scores exceeding preset scores; a video generation unit configured to generate episode short videos based on each complete director script and a preset multi-layer cascaded video generation mechanism; and an editing unit configured to edit the corresponding episode short videos based on each complete director script, and to generate a target short drama based on the edited episode short videos.
[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the short drama creation method as described in the first aspect when executed by the at least one processor.
[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the short drama creation method as described in the first aspect when executed.
[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the short drama creation method as described in the first aspect.
[0010] The short drama creation solution provided in this embodiment first obtains the original plot text and short drama requirement parameters; based on the original plot text and short drama requirement parameters, a script outline is generated to create episode plans, and each episode script is generated based on the creative evidence provided by a preset reference script template library and each episode plan; then, based on each episode script and a preset director's script design template library and director's script content template library, a complete director's script corresponding to each episode script is generated; next, based on each complete director's script and a preset multi-level cascading video generation mechanism, each episode short video is generated; finally, based on each complete director's script, the corresponding episode short videos are edited, and the target short drama is generated based on the edited episode short videos. This solution significantly improves the efficiency and quality of short drama creation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the professionalism and filmability of short drama creation, making it suitable for large-scale short drama production.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart of a short drama creation method provided in this embodiment of the disclosure; Figure 3 A flowchart of a short drama script generation method provided in this disclosure embodiment; Figure 4 A flowchart of a method for generating episode planning based on a script outline is provided as an embodiment of this disclosure; Figure 5 A flowchart illustrating a short drama video generation method provided in this embodiment of the disclosure; Figure 6 A flowchart illustrating a method for generating and filtering multiple alternative complete director's scripts for a specific scene, provided in this embodiment of the disclosure; Figure 7 A flowchart of a multi-layer cascaded video generation method provided in this disclosure embodiment; Figure 8 A flowchart of a short video editing method provided in this embodiment of the disclosure; Figure 9 A flowchart illustrating an image frame editing method provided in this disclosure embodiment; Figure 10 A flowchart illustrating an image frame editing method based on plot rhythm balance provided in this disclosure embodiment; Figure 11 A flowchart illustrating a semantic-based audio addition method provided in this disclosure embodiment; Figure 12 A flowchart for adding matching effects based on visual language tags and scene tags is provided in this embodiment of the disclosure; Figure 13 A flowchart illustrating a short drama creation method in an application scenario provided by an embodiment of this disclosure; Figure 14 A structural block diagram of a short drama creation device provided in this embodiment of the disclosure; Figure 15 This is a schematic diagram of the structure of an electronic device suitable for performing a short drama creation method, provided as an embodiment of the present disclosure. Detailed Implementation
[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0014] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0015] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the short drama creation methods, apparatuses, electronic devices, and computer-readable storage media disclosed herein can be applied.
[0016] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0017] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include short drama creation applications, script generation applications, and instant messaging applications.
[0018] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0019] Server 105 can provide various services through its built-in applications. Taking a short drama creation application that provides script creation services as an example, when running this application, Server 105 can achieve the following effects: First, it obtains the original plot text and short drama requirement parameters; based on the original plot text and short drama requirement parameters, it generates a script outline to create episode plans, and generates episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans; then, based on the episode scripts and the preset director script design template library and director script content template library, it generates complete director scripts corresponding to each episode script; next, based on the complete director scripts and the preset multi-level cascading video generation mechanism, it generates short videos for each episode; based on the complete director scripts, it edits the corresponding short videos for each episode, and generates the target short drama based on the edited short videos for each episode.
[0020] It should be noted that the original plot text and short drama requirement parameters can be obtained from terminal devices 101, 102, and 103 via network 104, or they can be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when it begins processing previously stored short drama creation tasks), it can choose to retrieve this data directly from the local storage. In this case, the exemplary system architecture 100 may not include terminal devices 101, 102, and 103 and network 104.
[0021] Because short drama creation requires significant computing resources and power, the short drama creation methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the short drama creation device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also complete the aforementioned calculations performed by the server 105 through their installed short drama creation applications, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the short drama creation application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the short drama creation device can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0022] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0023] Please refer to Figure 2 , Figure 2 A flowchart of a short drama creation method provided in this disclosure embodiment, wherein process 200 includes the following steps: Step 201: Obtain the original plot text and short drama requirement parameters.
[0024] This step is intended for the execution body of the short drama script generation method (e.g., Figure 1 The server 105 shown obtains the original plot text and short drama requirement parameters. The original plot text, as the basic material for content generation, typically refers to the unedited source text (such as a novel or story outline). The executing entity uses natural language processing technology to analyze its semantic structure and extract key narrative elements (such as characters, scenes, and plot). Taking a novel as an example, the executing entity can obtain the original novel text and short drama requirement parameters, including the number of episodes, episode length, video style, and target audience.
[0025] In this embodiment, the executing entity can obtain the original plot text either by having the user upload a complete novel document or by having the user submit episode fragments in batches. The original plot text can also be preprocessed; for example, a pre-trained text embedding model can be used to convert the original plot text into a high-dimensional vector while preserving semantic relevance. Simultaneously, a chapter segmentation algorithm is used to automatically divide the text into logical paragraphs, providing a structured basis for the pre-generation of subsequent episodes. For instance, when inputting a 500,000-word novel, the executing entity will first segment it by chapter, then extract the core plot, character relationships, and other elements from each chapter, ultimately organizing them into indexable structured data.
[0026] The short drama requirement parameters describe the requirements for the short drama generated from the input original plot text. These parameters may include, for example, the number of episodes to determine narrative density (e.g., 80 episodes require high-frequency conflict), the episode length to constrain the pace of a single scene (1-2 minutes requires strong conflict design), the video style to influence visual language selection (e.g., "ancient style martial arts" requires ink-wash tones), and the target audience type to drive content preference adaptation (e.g., "Generation Z" requires fast cuts and internet memes). These short drama requirement parameters can be stored in memory as key-value pairs, forming a callable creative constraint matrix. For example, when a user selects the "suspense" style, it can automatically associate parameters such as dark tones and fast cuts. Furthermore, the aforementioned short drama requirement parameters can be flexibly configured through an interactive interface. For instance, users can select the drama type using checkboxes, set the number of episodes using sliders, and select a visual style template from a drop-down menu. In addition, when users do not specify the duration of a single episode, the industry average can be automatically used as the duration of a single episode; when the audience type is not clearly defined, age preferences can be inferred based on platform data; the executing entity can also perform coupling verification of short drama requirement parameters, automatically reject contradictory combinations, and avoid the risk of generation failure from the source.
[0027] Specifically, for serialized creation scenarios, cross-session parameter inheritance is also supported. This means that when a user pauses and resumes creation, the previous short drama requirement parameter template can be automatically loaded. The execution entity can optimize short drama requirement parameters through comparative testing, such as generating two versions of a short drama and then optimizing the short drama requirement parameter library based on click data. Furthermore, short drama requirement recommendation parameters can be derived based on text features. For example, the number of novel chapters can be counted to automatically recommend the number of matching episodes, or the distribution of text sentiment can be analyzed to recommend suitable video styles.
[0028] Furthermore, to enhance usability, the system can also include built-in industry template presets. For example, for common genres such as martial arts dramas and suspense dramas, there are preset template packages containing complete short drama requirement parameters such as the number of episodes, duration, and style. When a user selects the "Martial Arts Drama Template," the system automatically loads short drama requirement parameters such as an 80-episode setting and a cool color scheme, significantly reducing the configuration cost of short drama requirement parameters. These templates can be continuously updated by operations personnel to ensure they are in sync with market trends.
[0029] Step 202: Generate episode plans based on the script outline generated from the original plot text and short drama requirement parameters, and generate episode scripts based on the creative evidence provided by the preset reference script template library and episode plans.
[0030] Building upon step 201, this step aims to generate episode plans from the script outlines created by the aforementioned executing entity based on the original plot text and short drama requirement parameters, and to generate episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans. The script outline typically includes the overall storyline, main character arcs, core plot sequence, and preliminary episode point suggestions. Episode planning, based on the script outline, is a specific plan that precisely divides the entire story into each episode, typically including episode objectives, core conflicts / events, plot development, main scenes, and character appearances and status changes. The reference script template library is constructed based on templated historical scripts with actual scores exceeding the preset rating. These templates provide evidence of creative intent. Creative evidence refers to reference script templates extracted from the reference script template library that highly match the specific plots, scenes, character relationships, and emotional needs in the current episode plan. Episode scripts are script texts in a standard format (or near-standard format) automatically generated under the guidance of the episode plans and by integrating creative evidence.
[0031] Step 203: Based on each episode script and the preset director's script design template library and director's script content template library, generate a complete director's script corresponding to each episode script.
[0032] In this embodiment, the executing entity generates a complete director's script corresponding to each episode script based on the scripts for each episode and a preset director's script design template library and director's script content template library. The reference script template, director's script design template, and director's script content template are all obtained by templated processing of the corresponding content of historical film and television works with actual ratings exceeding the preset score. The director's script design template refers to a standardized visual design pattern and structural framework extracted and abstracted from a large number of actual director's scripts or storyboards of historical film and television works with actual ratings exceeding the preset score. The director's script content template refers to specific content descriptions, performance instructions, and artistic details extracted and abstracted from a large number of actual director's scripts or storyboards of historical film and television works with actual ratings exceeding the preset score, used to fill the director's script design template. The complete director's script specifies the specific shooting requirements for each scene, and even each storyboard, including framing, character actions and performances, lighting atmosphere, environmental details, dialogue, sound design, etc.
[0033] Step 204: Generate short videos for each episode based on each complete director's script and the preset multi-layered cascading video generation mechanism.
[0034] In this embodiment, the executing entity generates episode short videos based on each complete director's script and a preset multi-layered cascaded video generation mechanism. The multi-layered cascaded video generation mechanism achieves high-quality conversion from director's scripts to episode short videos through a layered and progressive synthesis strategy. Each episode short video is a video file output by the multi-layered cascaded video generation mechanism after processing the complete director's script corresponding to an episode; it represents the initial visual composition of that episode.
[0035] In this embodiment, the execution entity generates episode-specific short videos based on a complete director's script and a pre-defined multi-layered cascading video generation mechanism. Specifically, static visual anchor points are generated based on the storyboard descriptions in the director's script. A text-to-image intelligent agent parses the text instructions into spatial composition parameters and outputs basic keyframes. A graph-to-image intelligent agent then performs cross-storyboard consistency optimization, that is, it identifies core elements through a feature extraction model and compares them with a global parameter library. If a conflict is detected, feature mapping correction is automatically triggered, replacing abnormal elements while preserving the composition.
[0036] Specifically, an action interpolation engine can be used to calculate pixel displacement trajectories (such as arm waving paths) between the first and last keyframes based on optical flow. Redundant actions are compressed using a time warp algorithm, ensuring precise matching of single-scene duration to the script timecode. A strong correlation model between audio tracks and lip movements is established for the dialogue audio binding process: dialogue is decomposed into phoneme sequences to drive a 3D facial model, ensuring strict synchronization between lip and tooth movements at 30 frames per second and the audio waveform. An environmental sound field simulation module can also be introduced, calling acoustic databases to mix environmental sound effects (such as rain and footsteps) based on scene tags, and matching spatial acoustic characteristics (such as echo attenuation coefficients in narrow alleys) using a convolutional reverberation algorithm. Furthermore, inter-frame transition functions can be calculated based on the scene sequence relationship through episode-level splicing. Dynamic frame sampling technology is used for fast-cut scenes, inserting fade-in and fade-out at switching points to avoid visual jumps, and tracking key parameters through a global state recorder. When the executing entity detects scene continuity between adjacent episodes, it automatically inherits terminal lighting parameters to eliminate color banding.
[0037] Step 205: Based on each complete director's script, edit the corresponding episode short videos, and generate the target short drama based on the edited episode short videos.
[0038] In this embodiment, the executing entity edits the corresponding episode short videos based on each complete director's script, and generates a target short drama based on the edited episode short videos. Editing refers to a series of intelligent and refined post-editing operations performed on each episode short video according to the complete director's script, with the aim of realizing the director's intentions and improving the quality of the final product. The target short drama refers to a complete short drama product formed by integrating and packaging all the edited episode short videos in episode order.
[0039] The short drama creation method provided in this disclosure first obtains the original plot text and short drama requirement parameters; based on the original plot text and short drama requirement parameters, a script outline is generated to create episode plans, and each episode script is generated based on the creative evidence provided by a preset reference script template library and each episode plan; then, based on each episode script and a preset director's script design template library and director's script content template library, a complete director's script corresponding to each episode script is generated; next, based on each complete director's script and a preset multi-layer cascading video generation mechanism, each episode short video is generated; finally, based on each complete director's script, the corresponding episode short videos are edited, and the target short drama is generated based on the edited episode short videos. This solution significantly improves the efficiency and quality of short drama creation through automated processes, thereby reducing manual intervention, lowering production costs, and also improving the professionalism and filmability of short drama creation, making it suitable for large-scale short drama production.
[0040] Building upon the short drama creation methodology disclosed in Process 200, please refer to the following for improving the efficiency and quality of short drama script generation: Figure 3 , Figure 3 A flowchart of a short drama script generation method provided in this disclosure embodiment, wherein process 300 includes the following steps: Step 301: Generate a complete plot summary based on the key plot elements extracted from the original plot text.
[0041] In this embodiment, the executing entity generates a complete plot summary based on key plot elements extracted from the original plot text. Specifically, the executing entity can perform deep semantic analysis on the input original plot text using a pre-trained model, locate named entities such as characters, scenes, and props through entity recognition, construct an interaction network between characters through relation extraction, and mark key plot turning points through event detection. This yields key plot elements including chapter summaries, named entities (i.e., character entities, scene entities, prop entities, etc.), the interaction network between characters, and key plot turning points. Furthermore, the extraction process can incorporate a graph attention mechanism to construct a traversable knowledge graph from these key plot elements using a graph neural network. Nodes represent entities, and edges represent relation weights, providing quantifiable association criteria for subsequent creation.
[0042] For the key plot elements obtained, the executing entity can use dynamic template filling and adaptive recombination strategies to generate a structured complete plot summary. For example, firstly, the core event types of the key plot elements are extracted and matched with a preset narrative template library; then, an attention mechanism is used to calculate the relevance score between each key plot element and the template slot, and high-weight key plot elements are filled into the corresponding positions. To improve fluency, the executing entity can also call a fine-tuned language model to perform grammatical calibration on the filling results, ensuring that the output conforms to the concise narrative style required by the script. The structured complete plot summary can provide multi-dimensional constraints for subsequent episode planning. For example, when generating the "poisoning in the Imperial Garden" summary, parameters such as "high conflict intensity, time limit, and spatial closure" are output simultaneously. These data will directly affect the pacing control when dividing subsequent episodes.
[0043] Furthermore, frequently identified named entities (such as the protagonist's name) can be cached in memory to reduce redundant computations, while complex relational reasoning is processed in parallel using a distributed computing framework. For novels with millions of words, a sliding window approach is used for segmentation, and contextual attention is employed to maintain cross-segment coherence.
[0044] Step 302: Generate a script outline based on the complete plot summary, and generate an episode plan based on the script outline that matches the number of episodes set in the short drama requirement parameters.
[0045] In this embodiment, the executing entity generates a script outline based on a complete plot summary, and then generates an episode plan consistent with the number of episodes set in the short drama's requirements parameters based on the script outline. Specifically, the executing entity can first identify the core narrative elements in the complete plot summary (such as the main storyline, key conflict points, and character development arcs), and then match the optimal structural framework (such as a four-act structure of "introduction, development, climax, and conclusion" or a three-act structure of "goal-obstacle-reversal") through a preset short drama narrative template library, thereby obtaining the script outline. During this process, the inter-act pacing ratio can also be dynamically adjusted to meet the single-episode length constraint in the short drama's requirements parameters. For example, if a single episode is set to 1.5 minutes, then the event density of each act must be adapted to a 90-second narrative capacity to avoid lengthy dialogues or slow setups.
[0046] After obtaining the script outline, the implementing entity can perform the following steps to obtain the episode plans: First, the script outline is initially divided according to the set number of episodes. For example, for an 80-episode script, the conflict density of each episode is calculated as "total number of conflicts / number of episodes," and episode boundaries are set at key plot points. Then, all episode units are scanned to identify the first appearance of characters / props / scenes, and their characteristic parameters (such as character appearance descriptions, weapon attributes, and architectural styles) are stored in a global object library to implement a global parameter inheritance mechanism. Finally, cross-episode consistency is ensured. That is, when the same object appears again in subsequent episodes, the parameters are automatically retrieved from the object library and copied. If a parameter needs to be modified in an episode (such as a character changing clothes), the new parameter version is only bound to that episode, and other episodes still inherit the original settings. For example, if the female lead is set to have "long, straight black hair" in episode 1, and changes to curly hair in episode 30, the change will be recorded as only affecting episode 30 and subsequent episodes, while the original hairstyle remains unchanged in the first 29 episodes.
[0047] For serialized content creation, when a user adds new episode text, the system can automatically compare it with the historical object library and perform feature alignment for duplicate characters in the new content. Simultaneously, a conflict density monitoring algorithm can optimize the episode structure in real time. Specifically, when an episode is detected to have insufficient conflict intensity, it automatically compresses transitional scenes or merges adjacent episodes to ensure that each episode meets the fast-paced requirements of a short drama.
[0048] Step 303: Use the reference script template library to provide matching creative evidence for each episode plan, and obtain the scripts for each episode after the content is completed with the matching creative evidence, which are consistent with the set number of episodes.
[0049] In this embodiment, the executing entity utilizes a pre-defined reference script template library to provide matching creative evidence for each episode plan, resulting in episode scripts that are complete with the matched creative evidence and match the set number of episodes. The reference script template library is constructed based on templated historical scripts with actual ratings exceeding a preset score, and these templates provide evidence of creative merit.
[0050] In this embodiment, for each episode planning in the episode planning, a target reference script template that matches both the semantic level and the requirement level of the short sentence requirement parameters is retrieved from the reference script template library, and the corresponding episode planning is supplemented with content based on the creative evidence provided by the target reference script template to obtain an episode script that matches the set number of episodes.
[0051] The reference script template library is built from in-depth research on high-quality historical scripts. Specifically, it can be constructed in the following way: short drama scripts with actual ratings exceeding a preset threshold are selected. Then, their structured features are extracted through a script deconstruction model. The script deconstruction model is based on the Transformer architecture and can identify high-value narrative elements in the script, such as scene transition patterns, conflict design templates, and character dialogue paradigms.
[0052] To obtain the scripts for each episode, the executing entity can first perform semantic matching, transforming the core plot of the current episode plan into semantic vectors, and then retrieving the template with the closest narrative structure from the template library using cosine similarity calculation. Next, demand-level matching is performed, filtering out target templates that meet market preferences based on factors such as audience type and video style in the short drama's demand parameters. Then, the target templates are broken down into fillable narrative units, and elements are replaced and logically calibrated according to the specific scenes in the episode plan, thus obtaining the scripts for each episode. Furthermore, an adversarial generation mechanism can be introduced, with a discriminator model detecting logical conflicts to ensure that new content is coherent and connected to the original plot.
[0053] Furthermore, a feedback channel can be established. When a script generated after content completion is marked as high-quality output by the evaluation object, the template optimization process can be automatically triggered. This involves extracting the script's innovative narrative structure, de-identifying it, and adding it to the library as a new template. It also supports users manually importing scripts from popular TV series and using a difference comparison algorithm to identify improvements compared to older templates, enabling continuous iteration of the template library. This "use-feedback-evolution" mechanism allows the template library to keep up with market trends and avoids rigid creative models.
[0054] When generating episode scripts based on a series plan, matching reference script templates are needed to provide creative evidence for content completion. To improve the final quality of the episode scripts, a target reference script template that matches both semantically and in terms of short drama requirements can be retrieved from the reference script template library. Then, based on the creative evidence provided by the target reference script template, the corresponding series plan is completed, resulting in episode scripts that match the set number of episodes. Semantically, the executing entity can convert key plot descriptions of the current series plan into high-dimensional semantic vectors and match them with scene description vectors in the template library using a cosine similarity algorithm. In terms of requirements, templates can be filtered based on short drama requirements to eliminate options with incompatible styles.
[0055] The reference script template library specifically includes scene narrative templates, character profile templates, and emotional atmosphere templates, which provide evidence of creation for scene narrative plots, character profile settings, and emotional atmosphere rendering plots, respectively. Based on this, the target reference script templates can provide the following three types of structured creative evidence: Scene narrative templates can provide creative evidence for the standardized construction of plot structure. Their technical principle is to deconstruct classic narrative patterns into fillable skeleton models, including parameterized fields such as scene spatiotemporal coordinates, event sequences, and conflict intensity indices; Character profile templates can provide character behavior logic chains, ensuring the depth and consistency of character development. Their core lies in establishing a multi-dimensional vector model of character characteristics, including physiological characteristics, social attributes, psychological motivations, and behavioral patterns; Emotional atmosphere templates can associate emotion types with audiovisual elements to provide emotional rendering formulas, achieving industrialized control of emotional expression.
[0056] Furthermore, a slot filling mechanism can be used in the content completion process. For example, abstract nodes in the series planning can be mapped to the event slots of the template to automatically fill in the pre-set detailed descriptions of the template. At the same time, a style calibrator is introduced to ensure that the filled content is consistent with the style of the original text. For example, if the original series uses colloquial dialogue, the execution entity will filter out overly formal lines in the template.
[0057] During content completion, the implementing entity can supplement the content of each episode constituting the corresponding drama series based on the creative evidence provided by the target reference script template. That is, each drama series contains at least one episode, and the scenes and characters in each episode remain unchanged. Different episodes correspond to different conflict events that occur in different scenes involving the characters. An episode refers to a series of events completed by the same group of characters in a fixed scene. It usually serves as the basic narrative unit of the script to ensure that content completion begins at the episode level.
[0058] Specifically, the implementing entity can complete specific episodes through steps such as template matching, element replacement, and dynamic calibration. Specifically, it can retrieve matching templates from a reference template library based on conflict event types and scene tags; then, it extracts the general narrative skeleton from the templates and replaces abstract slots with the current instances: character ID mapping, prop binding, and dialogue conversion; finally, it uses a pre-trained language model to detect logical conflicts and automatically filter incompatible elements. Simultaneously, it can compress dialogue turns based on single-episode duration constraints to ensure event density meets industry standards.
[0059] The short drama script generation method provided in this disclosure first ensures that the script creation closely matches the user-defined requirements such as the number of episodes and episode length by acquiring the original plot text and short drama requirement parameters, thus achieving personalized customization. Secondly, it generates a complete plot summary based on key plot elements, effectively preserving the core story and refining the narrative structure, laying the foundation for subsequent creation. Thirdly, it ensures that each episode is compact and conforms to the overall drama framework by generating a script outline and episode planning. Finally, it utilizes a reference template library built based on high-scoring historical scripts to provide professional-level creative evidence for episode planning, thereby enhancing the script's logic and consistency. This solution significantly improves the efficiency and quality of short drama script generation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the script's professionalism and filmability, making it suitable for large-scale short drama production.
[0060] Based on the short drama script generation method disclosed in Process 300, this disclosure provides a specific implementation scheme for generating episode plans based on the script outline. Please refer to [link / reference]. Figure 4 The flowchart, wherein process 400 includes the following steps: Step 401: Divide the script outline according to the set number of episodes to obtain the original episode outlines.
[0061] This step involves the aforementioned execution entity dividing the script outline into original episode outlines based on narrative nodes, according to the number of episodes set by the user. During this process, a sliding window algorithm can be used to dynamically detect key plot turning points as episode boundaries, ensuring that each episode contains independent conflict events.
[0062] Step 402: Extract the setting parameters of each object that appears for the first time from each original episode outline in the order of episode growth, and summarize them to obtain the object setting parameter library.
[0063] Building upon step 401, this step aims to have the executing entity extract the setting parameters of each object appearing for the first time from each original episode outline in the order of episode progression, and compile them into an object setting parameter library. That is, when a new character, prop, or scene is detected to appear for the first time, its setting parameters (material, pattern, size, etc.) will be automatically extracted, and these parameters can be encapsulated into a structured object model and stored in the global object library.
[0064] The object setting parameter library can adopt a spatiotemporal index design. Each entry contains basic parameters, effective range, and version branch. The basic parameters record the original characteristics when they first appear. The effective range is bound to the "starting episode - ending episode" interval by default. The version branch refers to the generation of a new version when the characteristics of subsequent episodes change.
[0065] Step 403: Based on the object setting parameter library, inherit the setting parameters of objects that do not appear for the first time in each original episode outline in the order of episode growth, so as to obtain each episode plan that makes the same object have consistent setting parameters.
[0066] Based on step 402, this step aims to have the executing entity inherit the setting parameters of objects that are not appearing for the first time in each original episode outline according to the order of episode growth, so as to obtain each episode plan that makes the same objects have consistent setting parameters.
[0067] Specifically, the episode planning can be determined by performing the following two phases: In the inheritance detection phase, the object identifiers in the current episode outline are scanned. If an object with the same name exists in the global library, parameter inheritance is triggered. In the version matching phase, the effective range of the object is retrieved based on the current episode number. If it is within the effective range of the basic parameters, the basic parameters are directly copied. If it enters a new version range, a new parameter set is loaded to handle boundary cases. When the same object has multiple version parameters in the same episode, a temporary copy is automatically created and a time-space marker is added to ensure that the main storyline is not disturbed.
[0068] Further, the implementing entity can extract new setting parameters for non-first-time appearances of objects from each original episode outline in chronological order of episode growth, and determine the target episode number affected by the new setting parameters. This target episode number is used to limit the portion of the original episode outline that inherits the setting parameters of the corresponding object using the updated setting parameters. Then, new setting parameters corresponding to the target episode number are added to the object's setting parameter library, and other episode numbers (not the target episode number) are added to the original setting parameters of the corresponding object. These other episode numbers are used to limit the portion of the original episode outline that inherits the setting parameters of the corresponding object using the original setting parameters. In other words, by scanning the original episode outlines in episode order, new setting parameters for non-first-time appearances of objects are identified. Then, a semantic comparison algorithm is used to detect the difference between the new parameters and the original parameters, and they are automatically marked as valid changes. Subsequently, the effective scope of the new parameters is determined, relevant plot context is extracted, and its continuous impact is analyzed to generate a set of target episode numbers. If the new parameters only affect a local plot, the target episode number is precisely limited.
[0069] The method for generating episode plans based on a script outline provided in this embodiment first divides the script outline according to a set number of episodes to obtain original episode outlines. Then, in the order of episode growth, the setting parameters of each first-appearing object are extracted from each original episode outline, and a set of object setting parameters is compiled to obtain an object setting parameter library. Finally, based on the object setting parameter library, the setting parameters of non-first-appearing objects in each original episode outline in the order of episode growth are inherited to obtain episode plans that ensure that the same objects have consistent setting parameters. This embodiment enforces parameter inheritance in the order of episodes, ensuring a high degree of consistency in the setting parameters of non-first-appearing objects in all subsequent episodes. By dynamically capturing the first-appearing objects to build a global object parameter library, the maintenance of episode script setting consistency is automatically completed, significantly improving the creation efficiency and quality of the multi-episode script basic framework.
[0070] Based on any of the above embodiments, the executing entity also discloses a specific implementation method for generating episode scripts: First, a preset content understanding agent extracts key plot elements from the original plot text to generate a complete plot summary; then, a preset script outline generation agent generates a script outline based on the complete plot summary, and a preset episode planning agent generates episode plans consistent with the number of episodes set in the short drama requirement parameters based on the script outline; finally, a preset episode script creation agent uses a reference script template library to provide matching creative evidence for each episode plan, resulting in episode scripts with content completion based on the matching creative evidence, consistent with the set number of episodes. This implementation method utilizes multiple agents to generate high-quality episode scripts with consistent logic and coherent content, reducing manual intervention and improving creation efficiency.
[0071] Based on the short drama creation method disclosed in Process 200, this disclosure provides a specific implementation scheme for a short drama video generation method in order to achieve efficient production of short drama videos. Please refer to [the relevant documentation]. Figure 5 The flowchart, wherein process 500 includes the following steps: Step 501: Determine the story content of each scene in the episode script and the corresponding shooting type.
[0072] In this embodiment, the executing entity determines the story content of each act in the episode script and the corresponding plot shooting type. For example, the story content of each act in the episode script can be semantically parsed, and the plot shooting type corresponding to the episode script can be determined using type mapping technology. Here, semantic parsing refers to converting natural language text into a machine-understandable structured representation, and type mapping technology maps information in words or sentences to specific types or categories.
[0073] Specifically, the story content of each act can be analyzed in the following way: First, the scene boundaries are identified by the spatiotemporal labeling algorithm, and the entities such as characters and props are located by named entity recognition to construct an entity relationship graph. Then, the action chain is analyzed by the event extraction model to form a "subject-action-object" triple structure. Finally, the conflict intensity is quantified by the sentiment analysis model, and the story content of each act containing spatiotemporal coordinates, character behavior chains and sentiment tags is output in the script.
[0074] Specifically, the shooting type of the episode script can be determined based on a multi-dimensional classification model. Standardized scene descriptions can be input into the multi-dimensional classification model, which will output matching items in the preset type system. In the multi-dimensional classification model, the classification criteria include conflict feature dimension (distinguishing between physical confrontation and verbal confrontation through action verb analysis), spatial enclosure dimension (judging between narrow spaces and open scenes based on scene descriptions, affecting the setting of camera movement range), and emotional intensity dimension (the shooting style is divided into soothing or intense types based on the level of emotional value).
[0075] Another specific implementation method is as follows: First, based on the script content of each act contained in the episode script, determine the actual plot of the corresponding character in the corresponding scene of the corresponding act; then determine the shooting type to be used to shoot the video content of the actual plot of the corresponding task in the corresponding scene for each act, and obtain the plot shooting type corresponding to each act.
[0076] Furthermore, when a matching scene shooting type cannot be detected or determined, temporary type tags can be automatically created and their visual characteristics recorded. Cross-screen correlation analysis is also supported; if three consecutive scenes are of the same type, optimization suggestions are generated (such as "switch to an overhead shot in the third scene to avoid visual fatigue").
[0077] Step 502: Select the director script design template and director script content template that match the shooting type of the plot from the director script design template library and the director script content template library.
[0078] Building upon step 201, this step aims to have the aforementioned executing entity determine, from the director's script design template library and director's script content template library, the director's script content template that matches the plot shooting type. The correspondence between different plot shooting types and different director's script design templates and different director's script content templates is pre-recorded.
[0079] Specifically, the director's script design template may include at least one of the following template configuration items to be determined: scene structure, shot breakdown corresponding to each scene, temporal relationship between shots, and shooting techniques. The scene structure defines the overall narrative framework of the script unit, including the scene type (e.g., "beginning-development-climax-ending"), emotional intensity curve (e.g., the climax scene must contain a conflict value of 0.8 or higher), and time constraints (e.g., the development scene must account for 40% of the total duration). The shot breakdown corresponding to each scene allows for a refined decomposition of scene visual elements. Each scene can be broken down into executable shot units, containing three dimensions of parametric constraints: basic parameters (e.g., number of shots ≥ 5 / scene), type parameters (e.g., close-up shots ≥ 30%), and spatial parameters (e.g., it must contain at least one cross-axis shot). The temporal relationship between shots... Relationships are used to ensure the logic and rhythm of shot transitions. By modeling shot dependencies using a directed acyclic graph, three types of mandatory constraints can be defined: continuity constraints (e.g., close-up A must be followed by a wide shot B), time constraints (e.g., flashback shot duration ≤ 3 seconds), and causal constraints (e.g., the "drawing a gun" action must be preceded by a "loading a bullet" shot). The shooting techniques used can transform abstract director instructions into quantifiable parameters. Basic techniques can be defined through parameters (e.g., push / pull / pan / track), mechanical indicators can be set through parameters (e.g., movement speed 1.2m / s ± 0.3), and environmental variables can be bound through parameters (rainy weather shooting requires the lens waterproof marking to be turned on).
[0080] The director's script content template can preset multiple key parameters, such as shooting techniques, character dialogue, action logic, and video length. In practice, the storyboard requirements in the actual director's script design can be analyzed first, and the corresponding parameter groups can be extracted from the director's script content template. Then, placeholders can be replaced with specific values through entity binding, and the speech rate can be adjusted according to the emotional intensity model. For action logic, the temporal rationality can be verified through a behavior rule library. For example, if it is detected that "a punch misses" but there is no subsequent action, the shot instruction "Character B counterattacks" can be automatically completed.
[0081] During template matching, the executing entity can use the plot shooting type as an index and calculate the template relevance based on multiple dimensions, including conflict intensity, spatial closure, and emotional tone, to determine the director's script design template and director's script content template that match the plot shooting type. For example, when a scene is classified as "ancient costume martial arts," the primary search will activate the action design template with the highest relevance in the basic template library; if it is detected that the current scene's emotional intensity deviates from the template's preset value by more than a preset threshold, then semantic enhancement search is initiated, dynamically replacing unsuitable modules through cosine similarity calculation.
[0082] Furthermore, when the rating of a newly released film or television work exceeds a threshold, an incremental training process can be automatically triggered to extract its innovative visual language and add it to the template library; when the audience is detected to be positioned in a specific regional market, a cultural adapter is activated to strengthen the regional aesthetic characteristics.
[0083] Step 503: Generate an actual director's script design using the story content of each act and the matching director's script design template, and generate a complete director's script containing the actual director's script content corresponding to each act using the actual director's script design and the matching director's script content template.
[0084] Building upon step 502, this step aims to enable the executing entity to generate an actual director's script design using the story content of each act and a matching director's script design template, thus transforming the script into executable director's instructions. Specifically, key event nodes for each act are first extracted from the story content. Then, an event importance analysis algorithm maps each key event node to the corresponding structural position in the director's script design template, binding entity parameters and inheriting the duration constraints of the template to generate the actual director's script design. The event importance analysis algorithm analyzes historical events or data using methods such as frequency, weighting, impact, and clustering to calculate the influence of events.
[0085] Specifically, the actual director's script design for each scene can be generated by using a director's script content template that matches the corresponding scene, resulting in a complete director's script for each scene. The actual script content includes at least one of the following: shooting techniques to be used in different shots, character dialogue, action logic, and video length. In other words, template matching is performed on a scene-by-scene basis.
[0086] In practice, the abstract instructions in the actual director's script design can be transformed into quantifiable parameters, character behavior chains can be constructed, and dialogue trigger nodes can be bound. Then, the content density can be optimized in reverse according to the length of the scene. If the dialogue is too long, redundant descriptions can be compressed and pause marks can be inserted to generate a complete director's script containing the content of the actual director's script for each scene.
[0087] Furthermore, when the actual director's script involves the collaboration of multiple elements, a timestamp mapping relationship can be established, and the video duration can be dynamically adjusted through a reverse constraint mechanism. For example, when the generated dialogue text exceeds the preset duration of the template, the pause interval will be automatically compressed or the action description will be simplified to ensure that the storyboard duration strictly matches the script design.
[0088] For non-linear narratives (such as flashbacks), multiple sets of spatiotemporal parameters (cool colors in reality / warm colors in memories) can be created and automatically activated during splicing; and a physical feasibility verification module can be introduced to automatically add a "green screen compositing" mark and corresponding special effects instructions when a "fall from a height" without safety measures is detected; a cross-screen parameter inheritance mechanism can also be established to track key visual elements (such as the color of the protagonist's coat), and if the color parameters are modified in subsequent scenes, correction instructions will be automatically added to the close-up of the front-facing camera.
[0089] To address discrepancies between the template and the design, an adaptive compensation process can be initiated: historical parameters from similar scenes can be retrieved to generate supplementary instructions; if no reference data is available, reasonable parameters can be calculated based on a physics simulator. Manual intervention is also supported, allowing directors to directly modify the generated script content using annotation tools and record adjustment points, feeding them back to the template library to create a closed loop of continuous optimization.
[0090] The short drama video generation method provided in this disclosure first matches a determined plot shooting type with director's script design templates and director's script content templates derived from highly-rated historical films and television works. This transforms professional directing experience into reusable parameters, ensuring that the camera language conforms to industry standards. Finally, the templates are used to generate actual director's script designs and fill in the script content, thus solving the problems of fragmented camera language and uncontrolled pacing in the generated videos. This solution achieves efficient short drama video production through the use of structured templates and a layered generation mechanism.
[0091] Building upon the aforementioned implementation, the executing entity can further extract structured reference script templates, director's script design templates, and director's script content templates from the films and television works to be processed using a pre-defined deconstruction model. This model learns script extraction knowledge from training samples comprised of historical films and television works with actual scores exceeding a preset rating and sample templates extracted from these historical works. Its core function is to analyze high-rated historical films and television works and extract structured narrative templates. The sample templates include: sample reference script templates, sample director's script design templates, and sample director's script content templates. The executing entity can also extract structured reference script templates from the films and television works to be processed using a pre-defined script deconstruction model, achieving more efficient acquisition of reference script templates.
[0092] Based on the short drama video generation method disclosed in Process 500, and considering that videos generated using large language models are not stable, even with a relatively accurate complete director's script, there is no guarantee that the generated video will meet expectations. Therefore, to ensure that the generated video is as close to expectations as possible and to better evaluate the generated complete director's script, please refer to... Figure 6 , Figure 6A flowchart of a method for generating and filtering multiple alternative complete director's scripts for a specific scene, provided in this disclosure embodiment, wherein process 600 includes the following steps: Step 601: For the target scenes corresponding to key plot points in multiple scenes, design the actual director's script for each target scene and multiple matching director's script content templates, and use a tree search algorithm to generate multiple alternative complete director's scripts corresponding to each target scene in parallel.
[0093] This step aims to have the aforementioned executing entity, for target scenes corresponding to key plot points across multiple scenes, utilize the actual director's script design for each target scene and multiple matching director's script content templates, and then use a tree search algorithm in parallel to generate multiple alternative complete director's scripts corresponding to each target scene. The tree search algorithm is an algorithm used for searching within graphical or tree structures.
[0094] In this embodiment, the executing entity first identifies target scenes in the script with high conflict or emotional intensity exceeding a threshold. The actual director's script design for each target scene is combined with multiple matching content templates to create an independent search tree for each target scene. The root node is the original script design, and each branch represents a content template filling scheme.
[0095] Specifically, the implementing entity can identify key scenes through a plot importance analysis algorithm. The judgment criteria include three dimensions: first, the emotional intensity value (e.g., the emotional value of a conflict scene is ≥0.8); second, the narrative node weight (e.g., "twist" and "climax" are marked as key events); and third, the character participation density (scenes where the main character's screen time accounts for more than 70%).
[0096] Step 602: For the target scenes corresponding to key plot points in multiple scenes, generate pending storyboard videos corresponding to each of the multiple alternative complete director scripts for each target scene according to the multi-level cascaded video generation mechanism. Determine the target storyboard video based on the actual effect evaluation of each pending storyboard video, and cut off all alternative complete director scripts that do not correspond to the target storyboard video.
[0097] For each target scene within a multi-scene framework, the executing entity generates a pending storyboard video corresponding to each of the multiple candidate complete director's scripts for that scene, using a multi-layered cascaded video generation mechanism. Specifically, for each target scene, multiple candidate complete director's scripts are created using a variant algorithm (e.g., A / B testing-based text generation). Each selected complete director's script is then input into the multi-layered cascaded mechanism, where the script text is first parsed into structured visual data, then a dynamic model is applied to add motion trajectories and synchronize temporal data, and finally, the rendering engine is invoked to integrate audio and video elements.
[0098] The evaluation and cropping process can be carried out through a two-stage screening: in the first round, versions with scores below the threshold (e.g., <0.7) are quickly eliminated through a pre-trained quality prediction model; the remaining candidate videos enter the human simulation evaluation stage, and the final score is calculated through an audience reaction prediction algorithm.
[0099] This embodiment discloses a method for generating and filtering multiple alternative complete director's scripts for specific scenes. For target scenes corresponding to key plot points across multiple scenes, it utilizes the actual director's script design for each target scene and multiple matching director's script content templates. A tree search algorithm is used to generate multiple alternative complete director's scripts corresponding to each target scene in parallel, thereby automating the generation of alternative complete director's scripts, reducing repetitive manual labor, shortening the pre-production cycle, and enhancing the feasibility of creative ideas. For target scenes corresponding to key plot points across multiple scenes, the multiple alternative complete director's scripts for each target scene are processed according to a multi-layered cascaded video generation mechanism to generate pending storyboard videos corresponding to each alternative complete director's script. The target storyboard video is determined based on the actual effect evaluation of each pending storyboard video, and all alternative complete director's scripts that do not correspond to the target storyboard video are trimmed. This ensures creative diversity and improves the quality of the final film through precise filtering, achieving intelligent production and optimization of target storyboard videos.
[0100] Furthermore, the executing entity can also provide feedback to the evaluation object on the complete director's script generated for the target scene corresponding to the key plot in multiple scenes, and adjust the actual director's script design and / or actual director's script content used to generate the complete director's script based on the feedback received from the evaluation object.
[0101] The evaluation subjects can submit feedback through a standardized form, which includes three types of actionable items: first, shot-level parameter adjustments (e.g., "extend close-up duration from 3 seconds to 5 seconds"), second, narrative logic corrections (e.g., "remove redundant actions of the protagonist in scene A"), and third, style preference annotations (e.g., "increase the intensity of the cool-toned filter"). The implementing entity can use a natural language processing model to transform the textual feedback into structured instructions: for vague descriptions, it automatically associates quantifiable parameters such as shot duration and editing frequency to generate specific adjustment suggestions.
[0102] Based on any of the above embodiments, the executing entity may specifically utilize a preset scene analysis agent to determine the story content and corresponding plot shooting type of each scene in the episode script; utilize a preset director script design agent to determine a matching director script design template based on the story content of each scene, and generate an actual director script design based on the matching director script design template; utilize a preset director script content generation agent to determine a matching director script design template based on the story content of each scene, and generate a complete director script containing the actual director script content corresponding to each scene based on the actual director script design of each scene and the matching director script content template.
[0103] In this embodiment, different intelligent agents are used collaboratively to complete the complete director's script, which improves the reliability of the results of a single stage and strengthens the connection between different intelligent agents, thus helping to better control the overall quality of the generated complete director's script.
[0104] Based on the specific implementation scheme of the short drama video generation method disclosed in Process 500, to deepen the understanding of how to generate episode videos from a complete director's script, please also refer to [link to relevant documentation]. Figure 7 , Figure 7 A flowchart of a multi-layer cascaded video generation method provided in this disclosure embodiment is included in process 700, comprising the following steps: Step 701: Generate the keyframes for each shot based on the script content for each shot in the complete director's script for each scene.
[0105] The aforementioned execution entity can generate key image frames for each scene based on the script content for each scene in the complete director's script for each scene.
[0106] Specifically, the executing entity can use text-based and graph-based graph agents to generate key image frames for each scene from the script content of each shot in the complete director's script for each scene. For example, the input text instructions are first parsed into composition parameters, and the basic image is output. Then, the graph-based graph agent performs cross-shot consistency checks on the basic image.
[0107] Step 702: Based on the key image frames of each storyboard, generate the corresponding storyboard video carrying action behavior information and interactive dialogue audio.
[0108] Based on step 701, the aforementioned execution entity can generate a storyboard video carrying action behavior information and interactive dialogue audio for each storyboard based on the key image frames of each storyboard. Specifically, the execution entity can use an action interpolation engine to calculate pixel displacement paths based on optical flow, compress redundant frames through a time warp algorithm to ensure that the duration of a single storyboard precisely matches the duration set in the script, and drive three-dimensional lip-sync through a phoneme decomposition model for interactive dialogue audio to ensure that the lip shape changes every second are strictly aligned with the audio waveform.
[0109] Step 703: According to the time information determined for each shot in the complete director's script for each scene, stitch together the shot videos of each shot in a timeline manner to obtain the scene video corresponding to each scene.
[0110] Based on step 702, the aforementioned execution entity can stitch together the storyboard videos of each scene in a timeline manner according to the time information determined for each scene in the complete director's script for each scene, thereby obtaining the scene video corresponding to each scene. When stitching together the storyboard videos of each scene, a preset transition strategy can be adopted: when the editing rhythm of adjacent scenes exceeds a preset threshold, a fade-in / fade-out effect is automatically inserted to avoid visual jumps; for non-linear narratives, an independent timeline is created and color parameters are marked, and the corresponding filter is activated during stitching.
[0111] Step 704: Based on the time information of each scene, stitch together the scene videos of each scene in a timeline manner to obtain the episode short videos corresponding to the episode script.
[0112] Based on step 703, the aforementioned execution entity can splice the segment videos of each segment in a timeline manner according to the time information of each segment to obtain the episode short video corresponding to the episode script.
[0113] The multi-layered cascaded video generation method disclosed in this embodiment first generates key image frames for each scene based on the script content of each shot in the complete director's script for each scene. Then, based on the key image frames of each shot, a shot video carrying action behavior information and interactive dialogue audio for the corresponding shot is generated. Next, according to the time information determined for each shot in the complete director's script for each scene, the shot videos of each shot are spliced together in a timeline manner to obtain the scene video corresponding to each scene. Finally, according to the time information of each scene, the scene videos of each scene are spliced together in a timeline manner to obtain the episode short video corresponding to the episode script. This embodiment, by processing script content and synthesizing audiovisual elements, finally outputs video results that meet the requirements of the episode script, realizing the fully automated generation process from a complete director's script to episode short videos, reducing manual labor and improving the efficiency of video production.
[0114] Furthermore, to maximize the quality of the generated video, the implementing entity can also conduct a video quality evaluation for each storyboard video and / or each segment video, and generate a second modification suggestion if the video quality evaluation fails. The video quality evaluated includes at least one quality indicator among: smoothness of image, consistency of shot transitions, naturalness of movement, and audio-visual synchronization. The currently generated storyboard video and / or segment video is adjusted according to the second modification suggestion until the regenerated storyboard video and / or segment video content passes the video quality evaluation.
[0115] Similar to the example above, the aforementioned execution entity can also specifically evaluate the video quality of each storyboard video and / or each segment video through a preset human-like video quality evaluation intelligent agent, and generate a second modification suggestion when the video quality evaluation fails. The currently generated storyboard video and / or segment video is then adjusted according to the second modification suggestion until the content of the regenerated storyboard video and / or segment video passes the video quality evaluation of the human-like video quality intelligent agent.
[0116] In this embodiment, the video quality evaluation and dynamic correction mechanism achieves quality control of generated content through multi-dimensional quantitative analysis and closed-loop feedback optimization. This technical solution first establishes a multi-level evaluation system for video quality: image smoothness is detected through inter-frame optical flow rate change; shot continuity is verified using a feature point matching algorithm; motion naturalness is physically validated based on a biomechanical model; and audio-visual synchronization is precisely measured by comparing the millisecond-level timestamps of audio waveforms and lip movements. When the video quality evaluation fails, the executing entity can initiate the following correction strategies: for image smoothness issues, a frame interpolation algorithm is used to generate transition frames; for shot continuity issues, a scene re-rendering module is triggered to eliminate visual jumps by adjusting lighting parameters and spatial coordinates; for motion naturalness defects, a physics simulator is called to recalculate the motion trajectory; and for audio-visual synchronization issues, the audio track timeline is dynamically adjusted or lip-sync animation data is regenerated. After each correction, an incremental evaluation is automatically performed, only partially regenerating substandard segments to avoid wasting resources on full reconstruction.
[0117] Based on the short drama creation method disclosed in Process 200, this disclosure provides a specific implementation scheme for a short video editing method to improve the production quality of short dramas. Please refer to the following embodiments. Figure 8 The flowchart, wherein process 800 includes the following steps: Step 801: Based on the storyboard design information contained in the complete director's script of the episode short video to be edited, edit the image frames of the episode short video to be edited.
[0118] In this implementation, the executing entity edits the image frames of the episode short video to be edited based on the storyboard design information contained in the complete director's script. The storyboard design information in the complete director's script includes shot planning (such as splicing order and transition methods) and character introduction planning (such as the timing of introducing new characters).
[0119] Specifically, the process begins by extracting shot planning and character entrance planning using a structured analysis engine. A temporal relationship analysis algorithm is then employed to identify shot sequences and transition markers. An entity recognition model is used to locate the keyframe positions where a new character first appears. Finally, the original video frames corresponding to each shot are arranged sequentially. Optical flow analysis can also be used to detect the continuity of motion trajectories between adjacent shots. For example, if a break in motion is detected (such as switching to shot 2 before the waving motion in shot 1 is completed), automatic interpolation is used to generate transition frames (inserting 3 frames to smooth the arm trajectory). Simultaneously, a storyboard association database is established. When a subsequent storyboard references a previous scene (such as a flashback), the lighting parameters and composition ratios of the previous storyboard are forcibly inherited to eliminate any sense of disjointed editing.
[0120] Step 802: Based on the semantic understanding results of the plot keywords contained in the complete director's script of the episode short video to be edited, add matching audio to the episode short video to be edited.
[0121] Building upon step 801, this step aims to have the aforementioned executing entity add matching audio to the episode short video to be edited, based on the semantic understanding results of the plot keywords contained in the complete director's script of the episode short video to be edited.
[0122] Specifically, the execution entity can first extract the emotional tone, identify environmental features, and dynamic events from the complete director's script of the episode-by-episode short video to be edited; then, it performs a multimodal search in a preset sound effects library, matching existing audio resources through acoustic feature vectors (such as low-frequency energy distribution and pulse density). If there is no direct match in the library, basic sound effects are generated based on a physical simulator; then, emotional features are injected through an emotion enhancement module; finally, after decomposing the dialogue into phoneme sequences, the speech rate is adjusted according to the emotion value (angry speech rate increased by 40%) and pitch fluctuations (sad scenes decreased by 15%). Furthermore, environmental changes can be simulated through an acoustic propagation model. For example, when a scene change between adjacent shots is detected (such as "indoor → outdoor"), the difference in reverberation parameters is automatically calculated to generate a gradual audio transition.
[0123] In this embodiment, the executing entity can extract voiceprint features when a character's lines first appear, establish a character ID-voiceprint mapping library, and ensure the consistency of the voice features of the same character in subsequent shots.
[0124] Step 803: Based on the visual language tags and scene tags contained in the complete director's script of the episode short video to be edited, add matching special effects to the episode short video to be edited.
[0125] Building upon step 801, this step aims to have the aforementioned executing entity add matching special effects to the episode-specific short video to be edited, based on the visual language tags and scene tags contained in the complete director's script. The addition of matching special effects based on the visual language tags and scene tags of the complete director's script for the episode-specific short video to be edited is achieved through a semantically driven special effects generation mechanism. Visual language tags can define basic stylistic attributes of the visuals, such as "cool color tone" or "motion blur," while scene tags are used to trigger environmental-level special effects combinations, such as "rain particles + ground reflection / holographic projection + laser beam."
[0126] In this embodiment, the executing entity can identify the type of special effects through a label classification model, then locate the effective range of the special effects based on timestamps, and finally perform layered compositing through the rendering engine. Specifically, the base layer adjusts the global color curve, the overlay layer injects dynamic elements, and physical simulation is applied to enhance realism. Furthermore, when label correlation is detected in consecutive scenes, gradient filters can be automatically generated to avoid visual jumps. For example, if the plot rhythm label shows "climax conflict," the special effects' expressiveness is enhanced; if it's "calm dialogue," the particle density is reduced to the base value.
[0127] Step 804: Align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited episode short videos, and generate the target short drama based on the edited episode short videos.
[0128] Building upon steps 801-804, this step aims to align the edited image frames, the added audio stream, and special effects along the timeline, resulting in edited episode-by-episode short videos. Based on these edited episode-by-episode short videos, a target short drama is generated. Specifically, the executing entity can first establish a main timeline based on the complete director's script timecode, binding each material element with an independent timestamp: image frame sequences inherit the start time of the storyboard, audio streams are anchored according to dialogue nodes, and special effects layers are associated with scene transition times. Then, through a dynamic delay compensation algorithm, the system detects the processing delay of different modal data in real time and automatically adjusts the audio track pre-compensation to ensure strict alignment of lip movements and pronunciation.
[0129] The rendering process employs a layered rendering strategy to align the edited image frames, added audio streams, and effects along the timeline. Specifically, the bottom layer arranges the video tracks in frame sequence order, the middle layer overlays the effects particle system, and the top layer mounts the audio waveform. When multi-track overlap is detected, the rendering process activates a conflict resolution mechanism, prioritizing core narrative elements (dialogue audio has absolute priority) and dynamically compressing the duration of non-critical effects. During mobile rendering, particle density is automatically simplified, and keyframe thinning technology maintains smoothness on low-power devices.
[0130] Furthermore, when the audio stream is unexpectedly interrupted, supplementary speech can be generated based on contextual semantics; when special effects rendering times out, a backup template can be enabled (e.g., replacing complex particles with simplified light effects).
[0131] The short video editing method provided in this disclosure firstly provides precise shot instructions for image frame editing through storyboard design information, thereby eliminating logical deviations in shot transitions in traditional editing. Secondly, it drives intelligent audio matching through semantic understanding of plot keywords, achieving audio-visual context synchronization through sentiment analysis and environmental analysis. Simultaneously, it guides the dynamic generation of special effects through visual language tags and scene tags to strongly bind them to scene semantics, avoiding stylistic disjointedness. Finally, it synchronizes the editing frames, audio stream, and special effects layers along the timeline using a timeline alignment engine, significantly reducing rhythm errors in the final product. This solution achieves multimodal collaborative optimization of video editing through structured analysis of the complete director's script, significantly improving video editing efficiency and the industry standard of the edited results, as well as significantly enhancing the production quality of short dramas.
[0132] Based on the specific implementation scheme of the short video editing method disclosed in Process 800, in order to accurately achieve the mapping from storyboard description to physical editing actions, please also refer to... Figure 9 , Figure 9 A flowchart of an image frame editing method provided for an embodiment of this disclosure is provided, and the process 900 includes the following steps: Step 901: Determine the splicing order and transition method of the image frames of each shot that constitute the corresponding episode short video based on the shot planning contained in the complete director's script of the episode short video to be edited; This step aims to determine, by the aforementioned executing entity, the splicing order and transition methods of the image frames constituting the corresponding episode short video, based on the shot planning contained in the complete director's script of the episode short video to be edited. Specifically, the shot planning contained in the complete director's script can be analyzed, a timeline mapping relationship can be established using the shot transition point as an anchor point, the splicing order of the image frames of each shot can be determined through a temporal relationship analysis algorithm, and the scene attributes of adjacent shots can be detected. If the difference in scene type exceeds a threshold, the transition instruction generation process is triggered to generate the transition method.
[0133] Step 902: Stitch the image frames of each shot in the stitching order, and add a transition frame between the last frame and the first frame of two consecutive shots with scene changes after stitching.
[0134] In this embodiment, the execution entity stitches the image frames of each shot in the stitching order, and adds a transition frame between the last frame and the first frame of two consecutive shots that have a scene change after stitching, according to the transition method. Specifically, when a shot pair that needs a transition is detected (such as the last frame of shot A and the first frame of shot B), the execution entity can generate a transition frame according to the transition method specified in the script.
[0135] Furthermore, the executing entity can also generate subtitles describing the newly introduced character based on the character entrance plan included in the director's script, for the first image frame in the target short video that presents the new character. The generated subtitles have a style that matches the style of the first image frame presenting the new character and / or the style of the new character, and are placed in a position within the first image frame that does not overlap with any other subject.
[0136] When a new character's appearance frame is detected, the main subject area is defined through face detection. A safe zone is calculated using compositional aesthetic rules. Style matching is then used to determine a style that matches the style of the first image frame presenting the new character and / or the style of the new character. Corresponding subtitles are then generated within the safe zone according to this style. The style of the subtitles can be determined by extracting the main color tone and texture features of the image. Subsequent appearances of the same character automatically inherit the initially determined subtitle style, ensuring visual consistency.
[0137] Furthermore, users can manually adjust transition durations or subtitle positions during editing. In addition, when new characters appear consecutively in the same scene, the subtitle background transparency and animation are automatically standardized to avoid visual interference.
[0138] The image frame editing method provided in this embodiment first determines the splicing order and transition methods of the image frames constituting the corresponding episode short video based on the shot planning contained in the complete director's script of the episode short video to be edited; then, the image frames of each shot are spliced in the splicing order, and transition frames are added between the last frame and the first frame of two consecutive shots with scene transitions after splicing, according to the transition method. This embodiment determines the shot timing and transition methods by deeply analyzing the structured data of the director's script, ensuring the smoothness of the subsequently generated video, and transforming the error-prone editing process into a high-precision, reusable automated process, significantly shortening the video production cycle.
[0139] Based on the specific implementation scheme of the short video editing method disclosed in Process 800, in order to achieve precise optimization of video narrative rhythm control based on plot rhythm, please refer to... Figure 10 , Figure 10 A flowchart of an image frame editing method based on plot rhythm balance provided in this disclosure embodiment, wherein process 1000 includes the following steps: Step 1001: Determine the plot pacing information based on the emotional changes of the characters in the complete director's script of the episode short video to be edited.
[0140] In this embodiment, the executing entity determines the plot rhythm information based on the emotional change information of the characters appearing in the complete director's script of the episode short video to be edited. Specifically, the emotional change information of the characters in the complete director's script of the episode short video to be edited can be analyzed first. The text description can be converted into a numerical sequence to generate an emotion curve through an emotion intensity quantification model. Then, the emotion curve can be scanned in time units of preset duration to generate plot rhythm information.
[0141] Step 1002: Generate plot rhythm tags for the corresponding episode short videos based on the plot rhythm information.
[0142] Building upon step 1001, this step aims to have the aforementioned executing entity generate plot rhythm tags for the corresponding episode short videos based on the plot rhythm information. Specifically, when an emotional abrupt change is detected, the image frame corresponding to the emotional abrupt change in the corresponding episode short video can be marked as a rhythm keyframe, and plot rhythm tags can be automatically generated. For example, a "conflict peak" can be marked for a climax segment, and a "buffer zone" can be marked for a transition segment.
[0143] Step 1003: Calculate the rhythm index based on the video duration in the middle of different plot rhythm tags, and adjust the rhythm of the corresponding episode short videos to the desired balanced rhythm based on the rhythm index.
[0144] Building upon step 1002, this step aims to have the aforementioned executing entity calculate a rhythm index based on the video duration within different plot rhythm tags, and adjust the rhythm of the corresponding episode short videos to the desired balanced rhythm based on the rhythm index. Specifically, the rhythm index can be calculated by statistically analyzing the cumulative duration percentage corresponding to each plot rhythm tag. This rhythm index can include the amplitude of emotional fluctuations and the frequency of segment transitions. When the rhythm index exceeds a preset threshold, the plot rhythm is deemed unbalanced.
[0145] Adjusting the pacing of the corresponding episode short videos to the desired balanced pacing based on the pacing index can include the following methods: For segments with too fast a pacing, dynamic blur frames can be inserted to prolong visual dwell time; for segments with a sluggish pacing, an intelligent editing engine can be enabled to automatically remove redundant actions; when the emotional lines of multiple characters intertwine, a composite pacing curve can be generated through an emotion overlay algorithm to force the peak of the main character's emotion to align with the shot transition point; and models can also be trained based on historical data to automatically insert environmental empty shots in emotionally depressing segments to avoid viewer fatigue.
[0146] This embodiment discloses an image frame editing method based on plot rhythm balance. First, it determines plot rhythm information based on the emotional changes of the characters in the complete director's script of the episode short video to be edited. Then, it generates plot rhythm tags for the corresponding episode short videos based on the plot rhythm information. Finally, it calculates a rhythm index based on the video duration between different plot rhythm tags and adjusts the rhythm of the corresponding episode short videos to the desired balanced rhythm. This embodiment, through emotion-driven rhythm generation, transforms the psychological changes of characters into a visual narrative rhythm, ensuring that the video rhythm always approaches the optimal experience threshold. This solves the problem of the disconnect between emotional expression and rhythm control, achieving intelligent optimization of plot rhythm.
[0147] Building upon the aforementioned implementation, to optimize for redundant and abnormal frames, the execution entity can first identify redundant image frames and / or abnormal frames in the episode short video to be edited; then, it can remove the identified redundant image frames; finally, it can repair the identified abnormal frames according to their abnormality type using appropriate repair methods. The abnormality types include: still frames, blurred frames, and dialogue misalignment frames. For redundant frame detection, the execution entity uses a time sliding window algorithm to calculate the visual similarity of consecutive frames; when more than 5 consecutive frames meet the condition, it is determined to be a redundant sequence. For abnormal frame detection, the execution entity establishes a categorized recognition model: still frames are captured by using a pixel change rate threshold to detect image stillness; blurred frames are calculated using a Laplacian operator to calculate image sharpness; and dialogue misalignment frames are detected by an audio-visual synchronization analyzer to detect lip movements and audio timestamp deviations.
[0148] Based on the above implementation, the execution entity can perform intelligent deletion of redundant frames in the following ways: retain key action node frames, generate intermediate transition frames through motion trajectory interpolation algorithms to ensure action continuity; for still frames, an optical flow prediction model can be called to generate a compensation frame sequence based on the motion vectors of the preceding and following frames; for blurred frames, a Generative Adversarial Network (GAN) super-resolution reconstruction network can be used to improve detail clarity, while constraining texture generation to conform to scene consistency; for dialogue misalignment frames, a two-way adjustment mechanism can be adopted, if the video is lagging, the frame sequence is moved forward, if the audio is lagging, the silent segment is dynamically cropped, and the lip-sync animation is fine-tuned through a phoneme alignment model.
[0149] Building upon the aforementioned implementation, the executing entity can also automatically record frequently occurring anomaly types and preload optimization parameters in subsequent similar scenarios. Simultaneously, the executing entity can establish a cross-segment repair knowledge base, triggering global parameter optimization when similar defect patterns are detected.
[0150] Based on the specific implementation scheme of the short video editing method disclosed in Process 800, in order to achieve a high degree of integration between audio and video, please refer to... Figure 11 , Figure 11 A flowchart of a semantic-based audio addition method provided for embodiments of this disclosure, wherein process 1100 includes the following steps: Step 1101: Determine the semantics of the plot keywords extracted from the complete director's script of the episode short videos to be edited; In this embodiment, the executing entity determines the semantics of the plot keywords extracted from the complete director's script of the episode-based short video to be edited. Specifically, the plot keywords are first extracted from the complete director's script of the episode-based short video to be edited, then semantic parsing is performed using a pre-trained natural language processing model, and word vector encoding technology is used to convert the text into semantic vectors, which serve as the semantics of the plot keywords.
[0151] In this embodiment, the executing entity can construct a multi-level semantic index. The first-level index is associated with the basic scene type (e.g., nature, city, science fiction, etc.), the second-level index is bound to dynamic parameters (e.g., rain sound intensity 0-1 scale), and the third-level index is linked to spatiotemporal attributes (e.g., cave echo attenuation coefficient).
[0152] Step 1102: Determine existing background sounds and / or existing ambient sounds that match the semantics from the preset sound effects library; In this embodiment, the executing entity can determine existing background sounds and / or existing ambient sounds that match the semantics from a preset sound effects library. Specifically, the executing entity compares the semantic vector with the audio feature vectors of sound effects in the preset sound effects library using cosine similarity. When the similarity is greater than a preset threshold, the current sound effect is determined to be an existing background sound and / or existing ambient sound that matches the semantics.
[0153] Step 1103: In response to the absence of a semantically matching background sound and / or ambient sound in the sound effects library, generate a new semantically matching background sound and / or ambient sound using a preset audio generation model; In this embodiment, when there is no background sound and / or ambient sound in the sound effect library that matches the semantics, the execution subject inputs the semantic vector and physical parameter constraints into the preset audio generation model to generate a new background sound and / or a new ambient sound that matches the semantics.
[0154] Step 1104: Add the semantically matched audio to the video frame corresponding to the relevant plot keyword.
[0155] In this embodiment, the executing entity adds semantically matching audio to the video frame corresponding to the relevant plot keyword. Specifically, the executing entity inserts semantically matching audio into the audio track of the video frame corresponding to the relevant plot keyword, based on the timecode and keyframes in the director's script.
[0156] In this embodiment, when multiple audio frequencies are detected to overlap, a priority strategy can be automatically activated to adjust the multiple audio frequencies. For example, when dialogue and explosion sounds overlap, the sound pressure level of the dialogue is increased, and the amplitude of the explosion sounds is reduced. The executing entity can also adjust the volume curve of the rain sound in real time according to the visual elements on the screen. For example, when the visual element on the screen is raindrops, the volume curve of the rain sound is adjusted in real time according to the density of the raindrops.
[0157] In this embodiment, to ensure optimal performance of each audio element in different plot scenarios, the executing entity can determine a mixing balance strategy for the dialogue audio, background sound, and ambient sound in the target short video, and then perform mixing processing on the dialogue audio, background sound, and ambient sound according to the mixing balance strategy. The specific mixing processing method is as follows: First, the fundamental frequency range of the dialogue audio is extracted, the energy distribution of the main melody of the background music is identified, and the intensity of the ambient sound's continuous noise is determined. Then, an adaptive masking algorithm is used to automatically reduce the gain of the corresponding frequency band of the background music by calculating the spectral overlap between the dialogue and the background sound, and adjust the ambient sound according to the noise intensity to avoid acoustic masking effects that reduce the clarity of the dialogue. For example, when the plot tag is "intense conflict," the maximum loudness of the ambient sound is limited to -6dBFS, the dialogue sound pressure level is increased by 15%, and the dynamic range is expanded so that special effects such as explosions do not drown out key lines; in "lyrical passages," the high-frequency overtones of the background music can be increased, while the ambient sound is faded in and out to enhance the immersive atmosphere. In addition, when the scene of an adjacent shot changes, the system automatically inherits the noise floor of the previous shot (such as the sound of an air conditioner at 30dB) and superimposes the ambient sound of the new scene (the sound of traffic at 40dB), eliminating the auditory jump through linear gradation.
[0158] The semantic-based audio addition method disclosed in this embodiment first determines the semantics of plot keywords extracted from the complete director's script of the episode short video to be edited. Then, it determines existing background sounds and / or existing ambient sounds that match the semantics from a preset sound effects library. When no matching background sounds and / or ambient sounds are found in the sound effects library, a new matching background sounds and / or ambient sounds are generated using a preset audio generation model. Finally, the semantically matching audio is added to the video frames. This embodiment, by deeply binding plot semantics with auditory experience, transforms sound effects design from being dominated by human experience to data-driven automated production, significantly improving the emotional immersion and production efficiency of video content. Furthermore, it generates customized audio when no matching sound effects are found in the sound effects library, ensuring audio-visual consistency.
[0159] Furthermore, in mobile playback scenarios, the frequency response defects of mobile phone speakers can be compensated by enhancing the dialogue frequency band; while for headphone users, the binaural rendering algorithm can be activated to enhance the spatial orientation of ambient sounds.
[0160] Based on the specific implementation scheme of the short video editing method disclosed in Process 800, in order to achieve a deep adaptation between visual effects and plot scenes, please refer to... Figure 12 , Figure 12 A flowchart of a method for adding matching effects based on visual language tags and scene tags provided in this disclosure embodiment, wherein process 1200 includes the following steps: Step 1201: Generate text animations, character effects, background effects, and filter effects based on visual language tags; In this embodiment, the executing entity generates text animations, character effects, background effects, and filter effects based on visual language tags. Specifically, the executing entity parses the visual language tags to generate text animations, character effects, background effects, and filter effects. Text animations refer to generating text animation trajectories using a path algorithm; for example, the keyword "unveil" triggers a text fragmentation and reconstruction animation. Character effects refer to adding dynamic elements based on a skeletal binding model; for example, the tag "magician" triggers fingertip particle streams. Background effects refer to simulating environmental interactions through a physics engine; for example, the tag "rainstorm" generates ripples spreading as raindrops collide with the ground. Filter effects refer to applying a color transfer model; for example, the tag "nostalgia" overlays sepia tones and grainy noise.
[0161] Step 1202: Determine the scene change time based on the scene tags, and generate highlight visual effects for the scene change time; In this embodiment, the executing entity determines the scene change time based on the scene label and generates a highlight visual effect for the scene change time. Specifically, the executing entity determines the scene change time based on the scene label, generates a highlight effect for the scene change time, and generates a transition effect sequence before and after the switch.
[0162] Step 1203: Determine the plot rhythm based on the scene tags, and control the style of each special effect within each plot rhythm to match the corresponding plot rhythm.
[0163] In this embodiment, the executing entity determines the plot rhythm based on scene tags and controls the style of various special effects within each plot rhythm to match the corresponding plot rhythm. For example, when a "intense conflict" rhythm is identified based on the scene tag, the particle effect emissivity is increased and the filter contrast is enhanced; when a "lyrical passage" rhythm is identified based on the scene tag, a low-interference mode is activated, reducing the background particle density. In addition, the same plot passage can inherit the main color tone and motion blur parameters, and deviations are automatically corrected through optical flow consistency detection.
[0164] In this embodiment, the execution entity can achieve device-adaptive rendering, such as automatically simplifying the particle system on mobile devices and maintaining performance through keyframe thinning. Furthermore, when a conflict between special effects and scene semantics is detected, special effects correction operations can be performed; for example, when "raindrops" appear in a "desert" scene, the system automatically switches to a sandstorm particle model.
[0165] This embodiment discloses a method for adding matching special effects using visual language tags and scene tags. First, it generates text animations, character effects, background effects, and filter effects based on the visual language tags. Then, it determines the scene change time based on the scene tags and generates highlight visual effects for that time. Finally, it determines the plot rhythm based on the scene tags and controls the style of each effect within each plot rhythm to match the corresponding rhythm. This embodiment significantly lowers the barrier to special effects production through large-scale automated production of effects. By using tags as a medium to concretize abstract plots into effect parameters, it achieves a precise visual transformation of narrative intent, providing a fully automated and scalable special effects solution for content creation.
[0166] Based on the above implementation, the executing entity can also utilize multiple intelligent agents to add special effects to the target short video. For example, the executing entity can first use a preset video editing intelligent agent to edit the image frames of the target short video based on the storyboard design information contained in the director's script; then use a preset audio adding intelligent agent to add matching audio to the target short video based on the semantic understanding results of the plot keywords contained in the director's script; finally, use a preset special effects adding intelligent agent to add matching special effects to the target short video based on the visual language and scene tags contained in the director's script.
[0167] Based on the above implementation method, the implementing entity can also evaluate the generated products to be evaluated and generate modification suggestions if the evaluation fails. Among them, the products to be evaluated include at least one of the following: script outline, episode script, actual director script design and / or actual director script content used to generate a complete director's script, each storyboard video and / or each scene video used to generate episode short videos to be edited, and edited episode short videos. The corresponding products to be evaluated are adjusted according to the modification suggestions until the modified products to be evaluated pass the evaluation.
[0168] Based on the above implementation method, the implementing entity can evaluate the script outline and / or each episode script according to the short drama requirement parameters and complete plot summary, and generate modification suggestions if the evaluation fails; adjust the corresponding script outline and / or episode script according to the modification suggestions until the revised script outline and / or episode script pass the evaluation.
[0169] Specifically, the implementing entity can utilize a pre-defined humanoid script evaluation agent to understand the complete plot summary output by the agent based on the content of the short drama's requirement parameters, and evaluate the script outline and / or each episode's script. This humanoid script evaluation agent quantitatively assesses the script outline and episodes' scripts based on a pre-defined multi-dimensional evaluation model. Evaluation indicators may include conflict density, character consistency, pacing suitability, and market preference matching. The process of generating modification suggestions may include: first, locating the root cause of defects through attention weight analysis; then, matching the repair strategy template in the film and television knowledge base, i.e., calling the modification plan for the problem and generating modification suggestions. The priority of modification suggestions can be dynamically divided according to the scope of impact; for example, problems within a single episode are marked as low priority, while logical breaks across episodes are marked as urgent corrections.
[0170] The implementing entity can utilize the script outline generation agent to adjust the script outline and / or the episode script creation agent to adjust the episode scripts based on the modification feedback, until the revised script outline and / or episode scripts pass the evaluation of the humanoid script evaluation agent. The adjustments made by the implementing entity to the corresponding script outline and / or episode scripts can take several forms: For structural defects at the script outline level, the script design agent can reallocate narrative weights and rearrange the event sequence according to conflict intensity; for localized issues in the episode scripts, the creation agent can perform precise replacements based on the modification feedback. Furthermore, incremental evaluation can be initiated after each modification, re-evaluating only the affected episodes to avoid wasting resources on full-scale testing. For problematic segments that still fail after three rounds of optimization, a manual intervention channel is activated, and these are recorded as negative samples and fed back to the training model.
[0171] Based on the above implementation, the executing entity can evaluate the correctness of the actual director's script design and / or actual director's script content generated for each scene, and generate a first modification suggestion if the correctness evaluation fails; adjust the currently generated actual director's script design and / or actual director's script content according to the first modification suggestion until the modified actual director's script design and / or actual director's script content passes the correctness evaluation.
[0172] Specifically, the executing entity evaluates the correctness of the actual director's script design and / or actual director's script content generated for each scene, and generates a first modification suggestion if the correctness evaluation fails. The correctness evaluation assesses whether the actual similarity between the generated actual director's script design and the matching director's script design template used is lower than a preset similarity threshold, and / or whether the actual similarity between the actual director's script content and the matching director's script content design template used is lower than a preset similarity threshold. The executing entity then adjusts the currently generated actual director's script design and / or actual director's script content based on the first modification suggestion until the modified actual director's script design and / or actual director's script content pass the correctness evaluation.
[0173] The adjustments include: adjusting the actual director's script design and / or content based on the first set of feedback, without changing the corresponding template used; and regenerating a new actual director's script design and / or new actual director's script content using a new corresponding template and the first set of feedback. Quality control of the director's script can be achieved through the closed-loop evaluation and dynamic correction mechanism described above, thereby ensuring that the generated content meets professional film and television production standards. This mechanism triggers quality alarms by calculating the similarity between the generated script and the matching template, and performs a structured evaluation of the actual director's script design and content. When the correctness evaluation fails, adjustments are made to the currently generated actual director's script design and / or actual director's script content.
[0174] Specifically, for local deviations, the original template is retained and modification suggestions are injected, and the parameter fine-tuning engine automatically optimizes it; for global inconsistencies (such as conflicting emotional tones), a template replacement process is initiated, which involves retrieving candidate templates with higher suitability from the template library based on the scene tags and the current defect type, and regenerating the script framework. A version comparison tool is introduced during the correction process to visually display the differences before and after the modification, assisting in manual confirmation.
[0175] Based on the above implementation, the executing entity can evaluate the video quality of each storyboard video and / or each segment video, and generate a second modification suggestion if the video quality evaluation fails. The video quality evaluated includes at least one of the following quality indicators: smoothness of image, consistency of shot transitions, naturalness of movement, and audio-visual synchronization. The currently generated storyboard video and / or segment video are adjusted according to the second modification suggestion until the regenerated storyboard video and / or segment video content passes the video quality evaluation. Similarly, the executing entity can also specifically use a pre-defined human-like video quality evaluation agent to evaluate the video quality of each storyboard video and / or each segment video, and generate a second modification suggestion if the video quality evaluation fails. The currently generated storyboard video and / or segment video are then adjusted according to the second modification suggestion until the regenerated storyboard video and / or segment video content passes the video quality evaluation of the human-like video quality agent.
[0176] Building upon the aforementioned implementation, the executing entity can further conduct a multi-dimensional quality assessment of the edited episode short videos and generate modification suggestions if the multi-dimensional quality assessment fails. Based on these suggestions, adjustments can be made to at least one of the video images, audio, and special effects until the adjusted target short video passes the multi-dimensional quality assessment. The quality dimensions assessed in the multi-dimensional quality assessment include at least two of the following: frame continuity and image consistency when connecting two consecutive shots, audio-visual synchronization, sound field equalization, composition score, color distribution score, and rhythmic integrity.
[0177] Specifically, the implementing entity can establish the following multi-dimensional quality assessment system: 1) Use optical flow analysis to detect the continuity of image frames, and identify smoothness issues in shot transitions by calculating the mutation rate of pixel motion vectors between adjacent image frames; 2) Evaluate the consistency of the scene based on feature point matching algorithms, and trigger a scene transition alarm if the matching degree of key feature points in consecutive shots is lower than a preset threshold; 3) Detect audio-visual synchronization by aligning lip movements and phoneme timestamps at the millisecond level; 4) Analyze the spectral energy distribution of dialogue, background sounds, and ambient sounds using sound field equalization, for example, if the 200-4000Hz frequency band of dialogue is obscured by more than 30% by background music, it is considered unbalanced; 5) Calculate the center-of-gravity shift of the scene based on the rule of thirds and the subject focus algorithm to score the composition; 6) Score the color distribution by comparing the main color difference values of adjacent shots using HSV (hue, saturation, brightness) space histograms; 7) Combine the emotional curve of the plot with the shot duration distribution to judge the completeness of the rhythm and detect whether the climax segment reaches the preset duration percentage.
[0178] Building upon the aforementioned implementation, when a multi-dimensional quality assessment fails, the executing entity activates an intelligent diagnostic engine to pinpoint the root cause and generate suggested modifications. For example, upon detecting a frame consistency defect, it automatically marks the range of problematic frames and suggests adding a 0.3-second transition effect; when sound field imbalance is detected, it outputs an audio correction parameter of "4dB attenuation in the mid-frequency band of background music." Furthermore, the executing entity can employ a layered repair strategy: for frame continuity issues, it calls a frame interpolation algorithm to generate a transition sequence; when compositional defects exist, it triggers an automatic cropping and dynamic reconstruction module; and when rhythm is missing, it inserts empty shots or compresses redundant segments based on an emotional intensity model.
[0179] Building upon the above implementation, when audio-visual asynchrony and abnormal color distribution are detected simultaneously, audio repair may affect the timeline; therefore, the audio-visual synchronization issue is repaired first, followed by color adjustment. The execution entity can also record frequently occurring problem scenarios, allowing for pre-loading of optimization parameters for subsequent similar scenarios.
[0180] Based on the above implementation, the executing entity can also specifically use a preset editing evaluation agent to conduct a multi-dimensional quality assessment of the edited target short video, and generate modification suggestions when the multi-dimensional quality assessment fails. Based on the modification suggestions, at least one of the video images, audio and special effects is adjusted until the adjusted target short video passes the multi-dimensional quality assessment.
[0181] Based on the above implementation, the executing entity can also specifically utilize multiple intelligent agents to generate the target short drama: A pre-defined script creation intelligent agent generates episode plans based on the original plot text and short drama requirement parameters, and generates episode scripts based on the creative evidence provided by a pre-defined reference script template library and the episode plans; a pre-defined director's script generation intelligent agent generates complete director's scripts corresponding to each episode script based on the episode scripts and a pre-defined director's script design template library and director's script content template library; a pre-defined video generation intelligent agent generates short videos for each episode based on the complete director's scripts and a pre-defined multi-layered cascading video generation mechanism; and a pre-defined editing intelligent agent edits the corresponding short videos for each episode based on the complete director's scripts, and generates the target short drama based on the edited short videos for each episode.
[0182] Based on the above implementation, the executing entity can also specifically use a human-like evaluation agent to evaluate the generated product to be evaluated, and generate modification opinions when the evaluation fails. The product to be evaluated is then adjusted according to the modification opinions until the modified product to be evaluated passes the evaluation.
[0183] To enhance understanding, this disclosure also includes... Figure 13This paper presents an end-to-end short drama creation system based on multi-agent collaboration. The system provides semantic and visual reference support for the creation process by introducing a heterogeneous intelligent agent hybrid collaboration mechanism, a knowledge enhancement mechanism for human-intelligent agent collaboration, a film and television professional knowledge base, and a multimodal material library. Through a unified intelligent agent evaluation and feedback mechanism, it constructs a quality closed loop across stages and modalities, realizes self-optimization of creative content and style consistency control, and achieves full-process automation from script creation, director script design, storyboard video generation to video editing and compositing.
[0184] In this implementation, the system includes four core creative stages and two supporting mechanisms: script creation stage, director's script design stage, storyboard video generation stage, video editing and compositing stage, as well as a knowledge enhancement mechanism and a unified evaluation and feedback mechanism to support these four stages.
[0185] First, during the scriptwriting stage, the system introduces a collaborative mechanism between a "screenwriting design agent" and multiple "screenwriting creation agents," allowing users to import popular scripts, film and television novels, and other texts as a knowledge base. The agents can enhance knowledge generation based on semantic retrieval and content transfer, thereby improving the narrative quality and market relevance of the script.
[0186] 1. The script design AI is responsible for generating an overall script design scheme based on the theme, type, style, duration requirements, or imported text (such as novels, stories, etc.) provided by the user, including story outline, character settings, plot structure, and emotional development.
[0187] 2. The screenwriting AI agents are responsible for creating different chapters, dialogues, or plot developments based on this design scheme, and are eventually automatically integrated into a complete script by the system.
[0188] To ensure the quality of creation, the system can be equipped with a built-in script evaluation agent to automatically evaluate the output script from dimensions such as story completeness, character development, and narrative tension, and provide real-time feedback to the creation agent to form an adaptive optimization loop.
[0189] Second, in the director's script design stage, the system introduces professional director's scripts and classic film examples as references through a knowledge enhancement mechanism, supports the transfer of camera language and the comparison of rhythm and style, and consists of a "director design intelligent agent" and multiple "director creation intelligent agents".
[0190] 1. The director design agent is responsible for transforming the script into a visual plan from the director's perspective, including camera style, scene design, character performance, pacing, and visual language.
[0191] 2. The director's creative AI further refines the director's script based on this plan, generates the structural content of scenes, sequences, and shots, and defines the shot size, camera position, motion trajectory, lighting relationship and dialogue coordination for each shot.
[0192] Meanwhile, the system has a built-in director script evaluation agent that evaluates visual appeal, pacing, shot fluency, and narrative consistency from multiple dimensions, and transmits the feedback results to the corresponding agent in real time for optimization.
[0193] III. Storyboard Video Generation Stage: This stage is the core of visual content production and includes two sub-stages: image generation and video generation.
[0194] 1. Image Generation Stage
[0195] The system is comprised of a multimodal material matching agent, a text-to-image agent, and an image-to-image agent working collaboratively. Based on the director's script, the system automatically retrieves the most suitable image materials from the character and scene libraries. If materials are missing, the text-to-image agent generates character and scene images, and the image-to-image agent extends and refines the image. The system also includes an "image evaluation agent" that comprehensively evaluates the generated results based on dimensions such as clarity, stylistic consistency, and compositional rationality, providing feedback for optimization.
[0196] 2. Video generation stage
[0197] The "Storyboard Video Generation Agent" is responsible for generating short video clips based on the director's script and available footage. The system automatically selects the starting frame based on the shot's position: the first shot uses a text-to-speech agent to generate the image, while subsequent shots use the ending frame of the previous shot to ensure smooth transitions. If a shot contains dialogue, the "Text-to-Speech Agent" generates the speech, and the "Lip-Sync Agent" synchronizes the audio and video. The generated video clips are then evaluated by the "Video Evaluation Agent" and further optimized based on feedback to ensure overall smoothness and audiovisual consistency.
[0198] IV. Video editing and compositing stage: This stage is completed collaboratively by the video editing agent, the audio mixing agent, and the special effects editing agent.
[0199] 1. The video editing AI is responsible for shot splicing, transition rhythm control, and rhythm balance; 2. The audio mixing AI automatically retrieves and matches background music and ambient sound effects, and performs mixing and balancing according to the rhythm of the plot; 3. The special effects editing AI automatically adds visual effects, subtitles, and filters based on the director's script.
[0200] The system has a built-in video editing evaluation agent that comprehensively evaluates the output video in terms of shot transitions, sound effects coordination, and visual aesthetics, generating quality feedback and guiding each agent to continuously optimize.
[0201] V. Supporting Mechanism 1, namely, the knowledge enhancement mechanism of human-intelligent agent collaboration
[0202] This mechanism allows creators to import multimodal materials, including script texts, director's scripts, storyboard examples, and video clips, to build an indexable multimodal knowledge base. The system employs cross-modal retrieval technology, automatically extracting relevant knowledge fragments as reference input based on the current task context during the generation process, enabling each agent to possess industry-specific and professional semantic background during creation.
[0203] In addition, the mechanism supports dynamic learning and knowledge updates. Users can feed newly generated high-quality content back into the knowledge base in the later stages of creation, forming a continuous enhancement loop of "knowledge-creation-feedback-relearning", which greatly improves the system's creative intelligence level and style adaptability.
[0204] Supporting Mechanism 2, namely a unified agent evaluation and feedback mechanism
[0205] This invention innovatively proposes a unified evaluation and feedback system based on a Mixture-of-Agents (MoA) architecture (the core of which involves multiple independent agents collaborating to process the same task). This system unifies the management of different types of evaluation agents (such as script evaluation, director's script evaluation, image evaluation, video evaluation, editing evaluation, etc.) and shares evaluation criteria, prompt templates, and historical records.
[0206] The MoA architecture can intelligently select the appropriate evaluation agent based on the characteristics of the current task, and conduct a comprehensive evaluation of the output content from multiple dimensions and perspectives. The evaluation results are transmitted to the corresponding creative agent in real time through a hierarchical feedback channel, enabling automatic optimization and style calibration.
[0207] The system also supports a historical context sharing mechanism between stages, which means that the creative style and goal setting of the previous stage can be referenced when evaluating the subsequent stage, so as to ensure the overall consistency of the short drama in terms of narrative logic, visual expression and rhythm style.
[0208] This disclosed end-to-end short drama creation system based on multi-agent collaboration first ensures that the script creation closely matches the user's requirements, such as the number of episodes and episode length, by acquiring the original plot text and short drama requirement parameters, thus achieving personalized customization. Secondly, it generates a complete plot summary based on key plot elements, effectively preserving the core story and refining the narrative structure, laying the foundation for subsequent creation. Thirdly, it ensures that each episode is compact and conforms to the overall drama framework by generating a script outline and episode plan. Finally, it utilizes a reference template library built based on high-scoring historical scripts to provide professional-level creative evidence for episode planning, thereby enhancing the script's logic and consistency. This solution significantly improves the efficiency and quality of short drama script generation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the script's professionalism and filmability, making it suitable for large-scale short drama production.
[0209] Further reference Figure 14 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a short drama creation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0210] like Figure 14 As shown, the short drama creation device 1400 of this embodiment may include: an input information acquisition unit 1401, a script creation unit 1402, a director script generation unit 1403, a video generation unit 1404, and an editing unit 1405. The system includes the following components: Input Information Acquisition Unit 1401, configured to acquire the original plot text and short drama requirement parameters; Script Creation Unit 1402, configured to generate episode plans based on the script outlines generated from the original plot text and short drama requirement parameters, and generate episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans; Director's Script Generation Unit 1403, configured to generate complete director's scripts corresponding to each episode script based on each episode script and the preset director's script design template library and director's script content template library; wherein the reference script templates, director's script design templates, and director's script content templates are all obtained by templated processing of the corresponding content of historical film and television works with actual scores exceeding the preset scores; Video Generation Unit 1404, configured to generate short videos for each episode based on each complete director's script and the preset multi-layer cascading video generation mechanism; and Editing Unit 1405, configured to edit the corresponding short videos for each episode based on each complete director's script, and generate the target short drama based on the edited short videos for each episode.
[0211] In this embodiment, the specific processing and technical effects of the input information acquisition unit 1401, script creation unit 1402, director's script generation unit 1403, video generation unit 1404, and editing unit 1405 in the short drama creation device 1400 can be found in reference to [reference needed]. Figure 2 The relevant descriptions of steps 201-205 in the corresponding embodiments will not be repeated here.
[0212] In some other optional implementations of this embodiment, the script creation unit 1402 includes: a complete plot summary generation subunit, configured to generate a complete plot summary based on key plot elements extracted from the original plot text; a script outline and episode planning generation subunit, configured to generate a script outline based on the complete plot summary, and generate an episode plan consistent with the number of episodes set in the short drama requirement parameters based on the script outline; and an episode script generation subunit, configured to provide matching creative evidence for each episode plan using a reference script template library, and obtain each episode script consistent with the set number of episodes after content completion with the matching creative evidence.
[0213] In some other optional implementations of this embodiment, the script outline and episode planning generation subunit includes an episode planning generation module configured to generate episode plans consistent with the number of episodes set in the short drama requirement parameters based on the script outline. The episode planning generation module includes: an episode division submodule configured to divide the script outline according to the set number of episodes to obtain each original episode outline; a setting parameter extraction submodule configured to extract the setting parameters of each first-appearing object from each original episode outline in the order of episode growth and summarize them to obtain an object setting parameter library; and a setting parameter inheritance submodule configured to inherit the setting parameters of non-first-appearing objects in each original episode outline in the order of episode growth according to the object setting parameter library, so as to obtain each episode plan that makes the same object have consistent setting parameters.
[0214] In some other optional implementations of this embodiment, the episode planning generation module may further include: a new setting parameter extraction submodule, configured to extract new setting parameters of non-first-time-appearing objects sequentially from each original episode outline according to the episode growth order, and determine the target episode number affected by the new setting parameters; wherein, the target episode number is used to limit the portion of the original episode outline that uses the setting parameters updated by the new setting parameters to inherit the setting parameters of the corresponding object; and an impact-attached episode number submodule, configured to add new setting parameters corresponding to the target episode number to the corresponding object in the object setting parameter library, and attach other episode numbers that are not the target episode number to the original setting parameters of the corresponding object; wherein, other episode numbers are used to limit the portion of the original episode outline that uses the original setting parameters to inherit the setting parameters of the corresponding object.
[0215] In some other optional implementations of this embodiment, the episode script generation subunit is further configured to: for each episode plan in the episode planning, retrieve a target reference script template from the reference script template library that matches both the semantic level and the requirement level of the short sentence requirement parameters, and complete the content of the corresponding episode plan based on the creative evidence provided by the target reference script template to obtain an episode script consistent with the set number of episodes.
[0216] In some other optional implementations of this embodiment, the director's script generation unit 1403 includes: an episode script analysis subunit, configured to determine the story content of each scene in the episode script and the corresponding plot shooting type; a template matching subunit, configured to determine the director's script design template and director's script content template that match the plot shooting type in the director's script design template library and the director's script content template library; wherein, the correspondence between different plot shooting types and different director's script design templates and different director's script content templates is pre-recorded; and a complete director's script generation subunit, configured to generate an actual director's script design using the story content of each scene and the matched director's script design template, and generate a complete director's script containing the actual director's script content corresponding to each scene using the actual director's script design and the matched director's script content template.
[0217] In some other optional implementations of this embodiment, the complete director's script generation subunit includes: a parallel generation module for multiple alternative scripts, configured to generate multiple alternative complete director's scripts corresponding to each target scene using a tree search algorithm for each target scene's actual director's script design and multiple matching director's script content templates; correspondingly, the video generation unit 1404: a storyboard video actual effect trimming subunit, configured to generate, for each target scene's target scene corresponding to the key plot, generate a pending storyboard video corresponding to each alternative complete director's script using a multi-layer cascaded video generation mechanism, and determine the target storyboard video based on the actual effect evaluation of each pending storyboard video, and trim all alternative complete director's scripts that do not correspond to the target storyboard video.
[0218] In some other optional implementations of this embodiment, the complete director's script generation subunit further includes: an object feedback adjustment module, which is configured to feed back the complete director's script generated for the target scene corresponding to the key plot in multiple scenes to the evaluation object, and adjust the actual director's script design and / or actual director's script content used to generate the complete director's script based on the feedback received from the evaluation object.
[0219] In some other optional implementations of this embodiment, the video generation unit 1404 is further configured to: generate key image frames for each scene based on the script content for each shot in the complete director's script for each scene; generate a scene video carrying action behavior information and interactive dialogue audio for each scene based on the key image frames for each scene; splice the scene videos of each scene in a timeline manner according to the time information determined for each scene in the complete director's script for each scene to obtain a scene video corresponding to each scene; and splice the scene videos of each scene in a timeline manner according to the time information of each scene to obtain a short episode video corresponding to the episode script.
[0220] In some other optional implementations of this embodiment, the short drama creation device 1400 may further include: a template deconstruction unit 1406, configured to extract structured reference script templates, director script design templates, and director script content templates from the film and television works to be processed through a preset deconstruction model; wherein, the deconstruction model learns script extraction knowledge from training samples consisting of historical film and television works with actual scores exceeding preset scores and sample templates parsed from historical film and television works, and the sample templates include: sample reference script templates, sample director script design templates, and sample director script content templates.
[0221] In some other optional implementations of this embodiment, the editing unit 1405 includes: a video editing subunit, configured to edit image frames of the episode short video to be edited based on the storyboard design information contained in the complete director's script of the episode short video to be edited; an audio adding subunit, configured to add matching audio to the episode short video to be edited based on the semantic understanding results of the plot keywords contained in the complete director's script of the episode short video to be edited; a special effects adding subunit, configured to add matching special effects to the episode short video to be edited based on the visual language tags and scene tags contained in the complete director's script of the episode short video to be edited; and a timeline alignment subunit, configured to align the edited image frames, the added audio stream, and the special effects along the timeline to obtain the edited episode short video, and generate a target short drama based on the edited episode short videos.
[0222] In some other optional implementations of this embodiment, the video editing subunit includes: a splicing order and transition method determination module, configured to determine the splicing order and transition method of the image frames of each shot constituting the corresponding episode short video according to the shot planning contained in the complete director's script of the episode short video to be edited; and a splicing and transition frame addition module, configured to splice the image frames of each shot in the splicing order, and add transition frames between the last frame and the first frame of two consecutive shots with scene transitions after splicing, according to the transition method.
[0223] In some other optional implementations of this embodiment, the video editing subunit further includes: a plot rhythm information determination module, configured to determine plot rhythm information based on the emotional change information of the characters appearing in the complete director's script of the episode short video to be edited; a plot rhythm tag generation module, configured to generate plot rhythm tags for the corresponding episode short video according to the plot rhythm information; and a plot rhythm balancing module, configured to calculate a rhythm index based on the video duration between different plot rhythm tags, and adjust the rhythm of the corresponding episode short video to the desired balanced rhythm according to the rhythm index.
[0224] In some other optional implementations of this embodiment, the audio adding subunit includes: a semantic determination module, configured to determine the semantics of plot keywords extracted from the complete director's script of the episode short video to be edited; an existing audio matching module, configured to determine existing background sounds and / or existing ambient sounds that match the semantics from a preset sound effects library; a new audio generation module, configured to generate new background sounds and / or new ambient sounds that match the semantics using a preset audio generation model in response to the absence of background sounds and / or ambient sounds that match the semantics in the sound effects library; and an audio adding module, configured to add the semantically matched audio to the video frame to which the corresponding plot keyword belongs.
[0225] In some other optional implementations of this embodiment, the audio adding subunit may further include: a mixing balance strategy determination module, configured to determine the mixing balance strategy for dialogue audio, background sound and ambient sound in the episode short video to be edited; and a mixing balance processing module, configured to perform mixing processing on the dialogue audio, background sound and ambient sound according to the mixing balance strategy.
[0226] In some other optional implementations of this embodiment, the special effects addition subunit is further configured to: generate text animation, character effects, background effects and filter effects based on visual language tags; determine the scene change time based on scene tags and generate highlight visual effects for the scene change time; determine the plot rhythm based on scene tags and control the style of each special effect within each plot rhythm to match the plot rhythm.
[0227] In some other optional implementations of this embodiment, the short drama creation device 1400 may further include: a rating unit 1407, configured to evaluate the generated product to be evaluated and generate modification suggestions when the evaluation fails; wherein the product to be evaluated includes at least one of the following: a script outline, episode scripts, actual director script design and / or actual director script content for generating a complete director's script, each storyboard video and / or each scene video for generating episode short videos to be edited, and edited episode short videos; and a modification suggestion adjustment control unit 1408, configured to adjust the corresponding product to be evaluated according to the modification suggestions until the modified product to be evaluated passes the evaluation.
[0228] In some other optional implementations of this embodiment, the script creation unit 1402 is further configured to: generate episode plans using a preset script creation intelligent agent based on the original plot text and short drama requirement parameters, and generate episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans; the director script generation unit 1403 is further configured to: generate complete director scripts corresponding to each episode script using a preset director script generation intelligent agent based on each episode script and the preset director script design template library and director script content template library; the video generation unit 1404 is further configured to: generate episode short videos using a preset video generation intelligent agent based on each complete director script and the preset multi-layer cascaded video generation mechanism; and the editing unit 1405 is further configured to: edit the corresponding episode short videos using a preset editing intelligent agent based on each complete director script, and generate the target short drama based on the edited episode short videos.
[0229] In some other optional implementations of this embodiment, the evaluation unit and the control unit for adjusting according to modification opinions are further configured to: evaluate the generated product to be evaluated using a preset human-like evaluation agent, generate modification opinions when the evaluation fails, and adjust the corresponding product to be evaluated according to the modification opinions until the modified product to be evaluated passes the evaluation.
[0230] This embodiment exists as a device embodiment corresponding to the method embodiment described above. The short drama creation device provided in this embodiment first obtains the original plot text and short drama requirement parameters to ensure that the script creation closely matches the user's requirements for the number of episodes, episode length, etc., achieving personalized customization. Secondly, it generates a complete plot summary based on key plot elements, effectively preserving the core of the story and refining the narrative structure, laying the foundation for subsequent creation. Furthermore, it generates a script outline and episode planning to ensure that each episode is compact and conforms to the overall drama framework. Finally, it utilizes a reference template library built based on high-scoring historical scripts to provide professional-level creative evidence for episode planning, thereby enhancing the script's logic and consistency. This solution significantly improves the efficiency and quality of short drama script generation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the script's professionalism and filmability, making it suitable for large-scale short drama production.
[0231] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the short drama creation method described in any of the above embodiments.
[0232] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the short drama creation method described in any of the above embodiments when executed.
[0233] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the short drama creation method described in any of the above embodiments.
[0234] Figure 15 A schematic block diagram of an example electronic device 1500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0235] like Figure 15As shown, device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1502 or a computer program loaded from storage unit 1508 into random access memory (RAM) 1503. The RAM 1503 may also store various programs and data required for the operation of device 1500. The computing unit 1501, ROM 1502, and RAM 1503 are interconnected via bus 1504. Input / output (I / O) interface 1505 is also connected to bus 1504.
[0236] Multiple components in device 1500 are connected to I / O interface 1505, including: input unit 1506, such as keyboard, mouse, etc.; output unit 1507, such as various types of monitors, speakers, etc.; storage unit 1508, such as disk, optical disk, etc.; and communication unit 1509, such as network card, modem, wireless transceiver, etc. Communication unit 1509 allows device 1500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0237] The computing unit 1501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, such as the short drama creation method. For example, in some embodiments, the short drama creation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1500 via ROM 1502 and / or communication unit 1509. When the computer program is loaded into RAM 1503 and executed by the computing unit 1501, one or more steps of the short drama creation method described above may be performed. Alternatively, in other embodiments, computing unit 1501 may be configured to perform a short drama creation method by any other suitable means (e.g., by means of firmware).
[0238] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0239] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0240] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0241] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0242] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0243] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0244] According to the technical solution of this disclosure, firstly, by acquiring the original plot text and short drama requirement parameters, the script creation is ensured to closely match the user's requirements for the number of episodes, episode length, etc., achieving personalized customization. Secondly, a complete plot summary is generated based on key plot elements, effectively preserving the core of the story and refining the narrative structure, laying the foundation for subsequent creation. Furthermore, by generating a script outline and episode planning, each episode's content is ensured to be compact and conform to the overall drama framework. Finally, a reference template library built based on high-scoring historical scripts provides professional-level creative evidence for episode planning, thereby enhancing the script's logic and consistency. This solution significantly improves the efficiency and quality of short drama script generation through automated processes, thereby reducing manual intervention, lowering production costs, and enhancing the script's professionalism and filmability, making it suitable for large-scale short drama production.
[0245] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0246] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for creating short dramas, comprising: Obtain the original plot text and short drama requirement parameters; Based on the original plot text and the short drama requirement parameters, the script outline is generated to generate each episode plan, and based on the creative evidence provided by the preset reference script template library and each episode plan, each episode script is generated. Based on the episode scripts and the preset director's script design template library and director's script content template library, a complete director's script corresponding to each episode script is generated, including: determining the story content and corresponding plot shooting type of each scene in each episode script; determining the director's script design template and director's script content template matching the plot shooting type in the director's script design template library and director's script content template library, wherein the script template, director's script design template and director's script content template are all templated based on the corresponding content of historical film and television works with actual scores exceeding the preset score, and the correspondence between different plot shooting types and different director's script design templates and different director's script content templates is pre-recorded; generating an actual director's script design using the story content of each scene and the matching director's script design template, and generating a complete director's script containing the actual director's script content corresponding to each scene using the actual director's script design and the matching director's script content template; Based on the complete director's scripts and the preset multi-layered cascaded video generation mechanism, short videos for each episode are generated. The multi-layered cascaded video generation mechanism generates static visual anchor points based on the storyboard descriptions in the director's script, uses a text-to-image intelligent agent to parse text instructions into spatial composition parameters and output basic keyframes, uses an image-to-image intelligent agent to perform cross-storyboard consistency optimization, identifies core elements through a feature extraction model and compares them with a global parameter library, and triggers a feature mapping correction mechanism that replaces abnormal elements while retaining the composition when a conflict is detected. Based on the complete director's script, the corresponding episode short videos are edited, and the target short drama is generated based on the edited episode short videos.
2. The method according to claim 1, wherein, The process of generating episode plans based on the script outlines generated from the original plot text and the short drama requirement parameters, and generating episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans, includes: A complete plot summary is generated based on the key plot elements extracted from the original plot text; A script outline is generated based on the complete plot summary, and an episode plan is generated based on the script outline, which is consistent with the number of episodes set in the short drama requirement parameters; The reference script template library is used to provide matching creative evidence for each episode plan, resulting in episode scripts that are completed with the matching creative evidence and are consistent with the set number of episodes.
3. The method according to claim 2, wherein, The generation of episode planning based on the script outline, which matches the number of episodes set in the short drama requirement parameters, includes: The script outline is divided according to the set number of episodes to obtain the original episode outlines; The setting parameters of each object appearing for the first time are extracted from each original episode outline in the order of episode growth, and the object setting parameter library is obtained by summarizing them. Based on the object setting parameter library, the setting parameters of objects that are not the first appearance in each original episode outline according to the growth order of the series are inherited in sequence to obtain each episode plan that makes the same object have consistent setting parameters.
4. The method according to claim 3, further comprising: According to the order of episode growth, new setting parameters of objects that are not appearing for the first time are extracted from each of the original episode outlines, and the target episode number affected by the new setting parameters is determined; wherein, the target episode number is used to limit the part of the original episode outline that uses the setting parameters updated by the new setting parameters to inherit the setting parameters of the corresponding object. In the object setting parameter library, new setting parameters corresponding to the target series number are added to the corresponding object, and other series numbers other than the target series number are added to the original setting parameters of the corresponding object; wherein, the other series numbers are used to limit the original episode outlines that inherit the setting parameters of the corresponding object using the original setting parameters.
5. The method according to claim 2, wherein, The step of using the reference script template library to provide matching creative evidence for each episode plan, and obtaining episode scripts that are completed with the matched creative evidence and consistent with the set number of episodes, includes: For each episode in the episode planning, a target reference script template that matches both the semantic level and the requirement level of the short drama requirement parameters is retrieved from the reference script template library. The corresponding episode planning is then supplemented with content based on the creative evidence provided by the target reference script template to obtain an episode script that matches the set number of episodes.
6. The method according to claim 1, wherein, The process of generating a complete director's script corresponding to each episode script, based on the episode scripts and a preset director's script design template library and director's script content template library, includes: Determine the story content of each act in the episode script and the corresponding shooting type; In the director's script design template library and the director's script content template library, a director's script design template and a director's script content template that match the plot shooting type are determined; wherein, the correspondence between different plot shooting types and different director's script design templates and different director's script content templates is recorded in advance; The actual director's script design is generated by using the story content of each act and the matching director's script design template. Then, the complete director's script containing the actual director's script content is generated for each act by using the actual director's script design and the matching director's script content template.
7. The method according to claim 6, wherein, The process involves using the actual director's script design and matching director's script content template for each scene to generate a complete director's script corresponding to each scene, including: For target scenes with key plot points in multiple scenes, multiple matching director script content templates are designed using the actual director script of each target scene, and a tree search algorithm is used to generate multiple alternative complete director scripts corresponding to each target scene in parallel. Correspondingly, the generation of each episode's short video, based on the complete director's script and the preset multi-layered cascading video generation mechanism, includes: For a target scene corresponding to a key plot point in multiple scenes, multiple candidate complete director scripts for each target scene are generated into a pending storyboard video corresponding to each candidate complete director script according to the multi-layer cascaded video generation mechanism. The target storyboard video is determined based on the actual effect evaluation of each pending storyboard video, and all candidate complete director scripts that do not correspond to the target storyboard video are cut off.
8. The method according to claim 6, further comprising: For a target scene corresponding to a key plot point in multiple scenes, the complete director's script generated for the target scene will be fed back to the evaluation object, and the actual director's script design and / or actual director's script content used to generate the complete director's script will be adjusted according to the feedback received from the evaluation object.
9. The method according to claim 6, wherein, The generation of short videos for each episode, based on the complete director's scripts and the preset multi-layered cascading video generation mechanism, includes: Based on the script content for each shot in the complete director's script for each scene, generate the keyframes for the corresponding shot. Based on the key image frames of each storyboard, generate a storyboard video that carries action and behavior information and interactive dialogue audio for the corresponding storyboard. Based on the time information determined for each shot in the complete director's script for each scene, the shot videos of each shot are spliced together in a timeline manner to obtain the scene video corresponding to each scene. Based on the time information of each scene, the scene videos of each scene are spliced together in a timeline manner to obtain the episode short videos corresponding to the episode script.
10. The method according to claim 1, further comprising: The pre-defined deconstruction model extracts structured reference script templates, director's script design templates, and director's script content templates from the films and television works to be processed. The deconstruction model learns script extraction knowledge from training samples consisting of historical films and television works with actual scores exceeding the preset score and sample templates parsed from the historical films and television works. The sample templates include: sample reference script templates, sample director's script design templates, and sample director's script content templates.
11. The method according to claim 1, wherein, The process of editing the corresponding episode short videos based on the complete director's scripts, and generating the target short drama based on the edited episode short videos, includes: Based on the storyboard design information contained in the complete director's script of the episode short video to be edited, the image frames of the episode short video to be edited are edited. Based on the semantic understanding results of the plot keywords contained in the complete director's script of the episode short video to be edited, add matching audio to the episode short video to be edited; Based on the visual language tags and scene tags contained in the complete director's script of the episode short video to be edited, add matching special effects to the episode short video to be edited; The edited image frames, the added audio stream, and the special effects are aligned along the timeline to obtain the edited episode short videos, and the target short drama is generated based on the edited episode short videos.
12. The method according to claim 11, wherein, The editing of image frames from the episode-specific short video to be edited, based on the storyboard design information contained in the complete director's script, includes: The stitching order and transition method of the image frames that make up the corresponding episode short video are determined based on the shot planning contained in the complete director's script of the episode short video to be edited. The image frames of each of the aforementioned shots are stitched together in the stitching order, and a transition frame is added between the last frame and the first frame of two consecutive shots with scene changes after stitching, according to the transition method.
13. The method according to claim 12, wherein, The process of editing the image frames of the episode-specific short video, based on the storyboard design information contained in the complete director's script, also includes: Based on the emotional changes of the characters in the complete director's script of the episode short videos to be edited, determine the plot rhythm information; Based on the plot rhythm information, generate plot rhythm tags for the corresponding episode short videos; The rhythm index is calculated based on the video duration in the middle of different plot rhythm tags, and the rhythm of the corresponding episode short videos is adjusted to the desired balanced rhythm according to the rhythm index.
14. The method according to claim 12, wherein, Based on the semantic understanding results of the plot keywords contained in the complete director's script of the episode-based short video to be edited, matching audio is added to the episode-based short video to be edited, including: Determine the semantics of the plot keywords extracted from the complete director's script of the episode short videos to be edited; Determine existing background sounds and / or existing ambient sounds from a preset sound effects library that match the semantics; In response to the absence of background sound and / or ambient sound matching the semantics in the sound effects library, a new background sound and / or new ambient sound matching the semantics is generated using a preset audio generation model; Add the audio that matches the semantics to the video frame that corresponds to the plot keyword.
15. The method according to claim 14, wherein, The semantic understanding results of the plot keywords contained in the complete director's script of the episode-based short video to be edited, adding matching audio to the episode-based short video to be edited, also include: Determine the mixing balance strategy for dialogue audio, background sound, and ambient sound in the episode short videos to be edited; The dialogue audio, background sound, and ambient sound are mixed according to the mixing balance strategy.
16. The method according to claim 12, wherein, The visual language tags and scene tags contained in the complete director's script of the episode-based short video to be edited are used to add matching special effects to the episode-based short video to be edited, including: Generate text animations, character effects, background effects, and filter effects based on the visual language tags; The scene change time is determined based on the scene label, and a highlight visual effect is generated for the scene change time; The plot rhythm is determined based on the scene tags, and the style of each special effect within each plot rhythm is controlled to match the corresponding plot rhythm.
17. The method according to any one of claims 1-16, further comprising: The generated products to be evaluated are evaluated, and modification suggestions are generated if the evaluation fails; wherein, the products to be evaluated include at least one of the following: the script outline, the episode script, the actual director script design and / or the actual director script content used to generate the complete director script, each storyboard video and / or each scene video used to generate the episode short video to be edited, and the edited episode short video. The corresponding products to be evaluated are adjusted according to the proposed modifications until the modified products pass the evaluation.
18. The method according to claim 17, wherein, The process of generating episode plans based on the script outlines generated from the original plot text and the short drama requirement parameters, and generating episode scripts based on the creative evidence provided by the preset reference script template library and the episode plans, includes: Using a pre-defined script creation agent, script outlines are generated based on the original plot text and the short drama requirement parameters to generate episode plans, and episode scripts are generated based on the creation evidence provided by the pre-defined reference script template library and each episode plan. The process of generating a complete director's script corresponding to each episode script, based on the episode scripts and a preset director's script design template library and director's script content template library, includes: Using a preset director script, an intelligent agent generates a complete director script corresponding to each episode script based on each episode script and a preset director script design template library and director script content template library. The generation of short videos for each episode, based on the complete director's scripts and the preset multi-layered cascading video generation mechanism, includes: Using a pre-defined video generation agent based on the complete director's script and a pre-defined multi-layered cascaded video generation mechanism, short videos for each episode are generated. The process of editing the corresponding episode short videos based on the complete director's scripts, and generating the target short drama based on the edited episode short videos, includes: Using a pre-defined editing agent, the corresponding episode short videos are edited based on the complete director's script, and a target short drama is generated based on the edited episode short videos.
19. The method according to claim 18, wherein, The process of evaluating the generated product to be evaluated, generating modification suggestions when the evaluation fails, and adjusting the corresponding product to be evaluated according to the modification suggestions until the modified product to be evaluated passes the evaluation includes: The generated product to be evaluated is evaluated using a pre-defined human-like evaluation agent. If the evaluation fails, modification suggestions are generated. The product to be evaluated is then adjusted according to the modification suggestions until the modified product passes the evaluation.
20. A short drama creation device, comprising: The input information acquisition unit is configured to acquire the original plot text and short drama requirement parameters; The scriptwriting unit is configured to generate episode plans based on the script outline generated from the original plot text and the short drama requirement parameters, and to generate episode scripts based on the creative evidence provided by the preset reference script template library and each episode plan. The director's script generation unit is configured to generate a complete director's script corresponding to each episode script based on the episode scripts and a preset director's script design template library and director's script content template library. This includes: determining the story content and corresponding plot shooting type for each scene in each episode script; determining director's script design templates and director's script content templates matching the plot shooting type in the director's script design template library and director's script content template library, where the script templates, director's script design templates, and director's script content templates are all templated based on the corresponding content of historical film and television works with actual scores exceeding preset scores, and pre-recording the correspondence between different plot shooting types and different director's script design templates and different director's script content templates; generating an actual director's script design using the story content of each scene and the matching director's script design template, and generating a complete director's script containing the actual director's script content corresponding to each scene using the actual director's script design and the matching director's script content template; The video generation unit is configured to generate short videos for each episode based on the complete director's script and a preset multi-layer cascaded video generation mechanism. The multi-layer cascaded video generation mechanism generates static visual anchor points based on the storyboard description in the director's script, uses a text-to-image intelligent agent to parse text instructions into spatial composition parameters and output basic keyframes, uses an image-to-image intelligent agent to perform cross-storyboard consistency optimization, identifies core elements through a feature extraction model and compares them with a global parameter library, and triggers a feature mapping correction mechanism that replaces abnormal elements while retaining the composition when a conflict is detected. The editing unit is configured to edit the corresponding episode short videos based on the complete director's script, and generate the target short drama based on the edited episode short videos.
21. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the short drama creation method according to any one of claims 1-19.
22. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the short drama creation method according to any one of claims 1-19.
23. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the short drama creation method according to any one of claims 1-19.
Citation Information
Patent Citations
Artificial intelligence-based script creation method, apparatus and device, and chip
CN117521628A
Intelligent novel-script conversion method and system
CN119494320A
Automated video creation-oriented script creation method based on large language model
CN119854597A