Whole-process AI (Artificial Intelligence) short drama creation system

The end-to-end AI short drama creation system uses the AI ​​assistant core module to schedule script generation, visual generation, and post-production optimization modules, solving the problems of fragmented processes and unstable quality in traditional short drama creation. It achieves efficient and low-cost creation, adapting to diverse market demands.

CN121836619APending Publication Date: 2026-04-10JIANGSU CUDATEC TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU CUDATEC TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional short drama creation models suffer from fragmented processes, high communication costs, high expenses, limited creativity, and unstable quality. Furthermore, existing AI tools cannot achieve full-process collaboration, resulting in low creation efficiency and inconsistent work quality.

Method used

The system adopts a full-process AI short drama creation system, which uses the AI ​​assistant core module to schedule script generation, visual generation and post-production optimization modules to form a closed-loop creation chain. It utilizes a multi-task intelligent agent to achieve seamless data flow and AI collaborative decision-making in each stage.

Benefits of technology

It enables highly efficient collaboration throughout the entire creative process, reduces communication costs, improves creative efficiency, ensures creative consistency and work quality, supports rapid creation by small and medium-sized teams, and adapts to diverse market demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121836619A_ABST
    Figure CN121836619A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of artificial intelligence, discloses a full-process AI short episode creation system comprising an AI assistant core module, a script generation module, a vision generation module, a post optimization module and a data support module, the script generation module, the vision generation module, the post optimization module and the data support module are scheduled by the AI assistant core module, and the AI assistant core module is used to create a short episode. Real-time linkage is realized, and a closed-loop creation link is formed. According to the full-process AI short drama creation system, an intelligent assistant constructed by an AI intelligent agent serves as a center, a linear creation process is reconstructed, and seamless circulation of data in all links and AI collaborative decision making are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to artificial intelligence, and relates to a full-process AI short drama creation system. BACKGROUND

[0002] The current short drama industry is showing explosive growth, but the traditional creation mode has significant pain points, which has become the core bottleneck restricting the development of the industry: first, the process is fragmented, short drama creation needs to go through scriptwriting, character design, scene building, shooting and recording, voice and music dubbing, editing and packaging, etc. Each link relies on different professional post personnel (scriptwriter, artist, actor, editor, etc.) to collaborate, with high communication costs and low connection efficiency, and the full-process cycle generally takes several months. Second, the cost is high, the real actor's remuneration, the rental of physical sites, the use of professional equipment and the cost of multi-post personnel make the creation cost of a single short drama high, which is difficult for small and medium-sized creation teams to bear. Third, the creativity is limited, artificial creation is easily affected by personal experience and style fixation, and the creative iteration speed is slow, which is difficult to quickly adapt to market diversification needs. Fourth, the standardization degree is low, the difference in professional level of different post personnel leads to unstable work quality, and there is a lack of unified creation standards and quality control system.

[0003] In the prior art, although some single-function AI tools (such as AI script generation, AI painting, and AI video generation) have appeared, these tools are independent modules and do not form full-process collaboration capabilities. When these tools are used in series, new technical problems such as difficulty in maintaining consistent character image and scene style, loss of plot logic in modal conversion, and inability of the output results of each link to be directly reused by downstream links will be encountered. SUMMARY

[0004] To solve the above technical problems, the present application aims to provide a full-process AI short drama creation system. The full-process AI short drama creation system adopts an architecture design of "AI intelligent agent hub + core function module", including an AI assistant core module, a script generation module, a visual generation module, a post-optimization module, and a data support module. Each function module realizes real-time linkage through intelligent scheduling of the AI assistant core module, forming a closed-loop creation link.

[0005] To achieve the above purpose, the technical solution adopted by the present application is as follows: a full-process AI short drama creation system, comprising an AI assistant core module, a script generation module, a visual generation module, a post-optimization module, and a data support module. The script generation module, the visual generation module, the post-optimization module, and the data support module accept the scheduling of the AI assistant core module, realize real-time linkage, and form a closed-loop creation link.

[0006] As a preferred mode of the present application, the AI assistant core module is constructed by a multi-task intelligent agent, including a demand analysis intelligent agent, a task allocation intelligent agent, a supervision intelligent agent and a collaborative intelligent agent; The demand analysis intelligent agent is used for intelligent demand analysis, and the specific workflow is demand receiving-demand analysis-standardized instruction generation. The task allocation intelligent agent is used for task allocation, and the specific workflow is task decomposition-module matching-parameter issuing-priority sorting. The supervision intelligent agent is used for monitoring the process and results, and the specific workflow is real-time tracking of progress-dynamic detection of quality-compliance audit-data log recording. The collaborative intelligent agent is used for seamless linkage between intelligent agents and functional modules, and the specific workflow is data transfer-cross-module collaboration-iterative response.

[0007] As a preferred mode of the present application, the script generation module integrates a large language model application program interface, including a demand analysis unit, a version control unit and a quality verification unit. The demand analysis unit: through a standardized interface, the LLM is called to receive the standardized instructions of the task allocation intelligent agent, to obtain the unstructured story text input by the user, and to automatically convert the unstructured demand into a standardized storyboard script containing role setting-scene setting-shot instruction by using the narrative logic and shot language analysis capability of the LLM. The version control unit: receiving the user operation instructions transmitted by the collaborative intelligent agent, automatically creating independent version identifiers for the initial script generated by the LLM API and the script edited by the user, and each version being associated with a complete project data package, covering the script text, role setting, scene data, storyboard parameters and modification log in the current state; the version data is synchronously stored in the version management database of the data support module. The quality verification unit: a double mechanism of input / output verification is constructed, and the verification rules and processing logic are linked with the supervision intelligent agent. The input verification mechanism is aimed at the story text and demand instructions input by the user, and performs text format checking, minimum length checking and sensitive word checking; if the input verification fails, the demand analysis intelligent agent is immediately fed back to the user through the collaborative intelligent agent, and the error information containing the failure type is returned to the user, and the LLM API calling is not performed. The output verification mechanism is aimed at the standardized storyboard script generated by the LLM API, and performs role consistency and scene continuity checking; if the output verification fails, the system automatically starts a retry strategy to control the LLM API to generate a script based on the failure reason, and the upper limit of the retry times is 3 times; if the three times of retry all fail, the manual filling mode is triggered by the supervision intelligent agent, and the filling prompt and problem script annotation are pushed to the user; the verified script is synchronized to the supervision intelligent agent for acceptance through the collaborative intelligent agent, and is output to the visual generation module after passing.

[0008] As a preferred mode of the present application, the visual generation module is the core execution module of visual content creation, including a reference picture generation unit, a storyboard generation unit, a dynamic video generation unit and a batch creation unit; The reference picture generation unit receives the creation instructions issued by the task allocation agent, synchronously acquires the standardized storyboard script output by the script generation module, extracts the character setting description and scene description therefrom through a text analysis component, and automatically converts them into structured prompt words; calls a text-to-picture large model through a standardized interface to generate character reference pictures and scene reference pictures in accordance with the set style - the character reference pictures include close-up views of the face and front, side and back views; the scene reference pictures include panoramic, medium and close-up level views; after the style consistency check is completed by the supervision agent calling the effect evaluation library, the reference pictures are synchronously transmitted to the storyboard generation unit and the resource material library of the data support module; The storyboard generation unit receives the script storyboard description information transmitted by the collaborative agent, based on the multi-source input data structure of the character reference picture + scene reference picture + storyboard description, inputs the multi-source data to the picture-to-picture large model through the picture-to-picture API, and the model generates a storyboard picture matched with the plot based on the visual features of the character reference picture, the spatial logic of the scene reference picture and the composition requirements of the storyboard description; the storyboard picture automatically labels the technical parameters of the shot sequence number, single shot duration and camera movement mode, supports the content matching degree check by the supervision agent of the AI assistant core module, and pushes the storyboard picture to the dynamic video generation unit after the check is passed; The dynamic video generation unit is a dynamic content core output component, which receives the video generation instructions issued by the task allocation agent, integrates the picture-to-video API to realize the conversion from static to dynamic; for a single static storyboard picture, the API calls the video generation model to convert it into a dynamic video segment of 5-10 seconds; after generating multiple sets of dynamic segments corresponding to the storyboards, the system starts the picture splicing component, automatically extracts the tail frame of the previous storyboard segment and the head frame of the next storyboard segment, extracts the feature points through the image feature matching algorithm, adjusts the picture angle, color parameters and character action posture of the head frame of the next storyboard, so that the feature matching degree of the head and tail frames is ≥90%; The batch creation unit receives the batch generation requirements transmitted by the collaborative agent, supports the batch creation of storyboards based on the same set of character / scene prompt words; the user sets the generation quantity through the system interface, the unit starts the random seed allocation mechanism to allocate a unique random seed for each generation task; under the premise of keeping the core style consistent, the picture-to-picture large model generates differential storyboards in picture details through seed differentiation; the batch-generated storyboards are automatically associated with version identifiers and synchronously transmitted to the version control unit and the data support module for subsequent calling and backtracking.

[0009] As a preferred mode of the present application, the post-production optimization module includes an intelligent voice matching unit, a sound effect matching unit, and an automatic editing unit. The intelligent voice matching unit receives the standardized shot script delivered by the collaborative agent, and constructs a voice line matching-emotion adaptation-parameter adjustable voice generation link through AI voice synthesis technology. First, the basic voice line in the preset voice line library is matched based on the role setting. Second, the emotion label corresponding to the lines is extracted, and the tone, stress, and pause rhythm of the voice are adjusted through an emotion adjustment algorithm. At the same time, the user can adjust the speed and tone through the AI assistant, and the generated voice audio is automatically associated with the shot timestamp, and the lip synchronization data is output to the automatic editing unit; The sound effect matching unit receives the dynamic video clip and shot emotion annotation data delivered by the collaborative agent, and constructs an automatic sound effect generation mechanism of scene analysis-emotion matching-material calling. The scene type of the video clip is identified through the scene analysis component, and the sound effect demand label is generated in combination with the emotion node. Based on the label, three types of sound effect resources-background music, environmental sound, and special effect sound are automatically matched from the resource material library of the data support module. The generated sound effect material is automatically labeled with a timestamp, and the sound effect track data is synchronously pushed to the automatic editing unit. The automatic editing unit integrates a visual timeline editor, which includes five core tracks and a unified millisecond time scale: the video track is used to load the dynamic video clip output by the shot generation module and the video generation module; the voice track is used to mount the voice audio generated by the intelligent voice matching unit; the subtitle track is used to automatically load the synchronized subtitles generated based on the voice lines; the sound effect track is used to embed various sound effect materials of the sound effect matching unit; and the BGM track is used to carry background music resources. The unified millisecond time scale supports user manual intervention adjustment. The user can directly drag the track material to modify the time position through the visual interface, or adjust the track attributes through the parameter setting panel. All modification operations are real-time feedback to the preview screen. Finally, through multi-track collaborative rendering, a film that meets the quality requirements is output, which is pushed to the export after being checked by the supervisory agent.

[0010] As a preferred mode of the present application, the data support module provides data support for the AI assistant core module, including training data set, resource material library, and effect evaluation library. The training data set includes short play scripts, character images, scene pictures, shot cases, dynamic video clips, and audience feedback data. Through the interface of the collaborative agent, training data support is provided for the text-to-image, image-to-image, and image-to-video API, and the agent optimization of the AI assistant core module. The resource material library includes classified sound effects, background music, font styles and special effect templates, and is linked with the collaborative intelligent agent through a standardized interface, and according to the instructions issued by the task allocation intelligent agent, provides accurate material support for the intelligent dubbing unit and the sound effect matching unit of the post-production optimization module. The effect evaluation library includes industry creation standards, audience preference models, quality evaluation indexes and video continuity determination rules, and is deeply linked with the supervision intelligent agent to provide the basis for quality detection - when monitoring the operation of each module, the supervision intelligent agent calls the evaluation library rules in real time, and performs real-time quality detection on the generated storyboard, dynamic video segments and finished film content, and if the detection is not up to standard, a warning is triggered and passed to the corresponding module by the collaborative intelligent agent for optimization.

[0011] Compared with the prior art, the present application has the following technical effects: The full-process AI short drama creation system of the present application takes the intelligent assistant constructed by the AI intelligent agent as the center, restructures the linear creation process, and realizes seamless data flow and AI collaborative decision-making in each link. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 It is a structure schematic diagram of a full-process AI short drama creation system in the embodiment.

[0013] Figure 2 It is a creation process schematic diagram taking the creative word as the initial input in the embodiment.

[0014] Figure 3 It is a creation process schematic diagram taking the story text as the initial input in the embodiment.

[0015] Figure 4 It is a creation process schematic diagram taking the reference video as the initial input in the embodiment. DETAILED DESCRIPTION

[0016] The technical solutions of the present application will be further introduced below in combination with the drawings.

[0017] As shown in the drawings, Figure 1 The present application discloses a full-process AI short drama creation system, which adopts the architecture design of "AI intelligent agent center + core function module", including AI assistant core module, script generation module, visual generation module, post-production optimization module and data support module, each function module realizes real-time linkage through intelligent scheduling of the AI assistant core module, forming a closed-loop creation link. The specific content is as follows: The AI assistant core module is constructed by a multi-task intelligent agent, specifically including a demand analysis intelligent agent, a task allocation intelligent agent, a supervision intelligent agent and a collaborative intelligent agent.

[0018] (1) Demand analysis agent realizes intelligent demand analysis. The specific workflow is demand receiving → demand analysis → standardized instruction generation.

[0019] Among them, demand receiving refers to that the user inputs the story text through the text input box or file upload. Demand analysis refers to extracting key elements such as style, duration, and characters. Standardized instruction generation is to convert the analyzed demand into machine-identifiable structured instructions.

[0020] (2) Task allocation agent realizes the optimal allocation of tasks, and its workflow is task decomposition → module matching → parameter delivery → priority sorting.

[0021] Among them, task decomposition is to decompose the total instruction output by the demand analysis agent into subtasks such as script generation, character and scene generation, storyboard and video generation, and to clarify the output standards and time nodes of each task. Module matching is to match the optimal execution module for the subtasks according to the real-time load of each functional sub-module. Parameter delivery is to deliver task instructions and associated data to the corresponding module. Priority sorting is to support users to mark urgent demands, and the agent automatically adjusts the task priority to allocate more computing resources for urgent tasks and shorten the execution cycle.

[0022] (3) Supervision agent ensures controllable process and qualified results, and its workflow is progress real-time tracking → quality dynamic detection → compliance audit → data log recording.

[0023] Among them, progress real-time tracking synchronously collects task execution progress through the state interface of each module, and feeds back to the user through visual charts, and automatically warns when the task is delayed. Quality dynamic detection is the effect evaluation library connected with the data support module, which performs real-time checking on the output results of each link (such as whether the script meets the short drama narrative rhythm, whether the character model matches the setting, and whether the video picture exists), and generates rectification instructions if it does not meet the standard. Compliance audit refers to the built-in short drama content compliance database, which automatically detects whether the script dialogue and picture elements contain illegal content (such as vulgar expression and sensitive scene), to ensure that the finished film meets the platform audit requirements. Data log recording refers to the retention of task data, modification records and quality detection results in each link throughout the process, providing basis for subsequent process optimization and problem tracing.

[0024] (4) Coordination agent realizes seamless linkage between each agent and functional module. Its workflow is data transfer → cross-module collaboration → iterative response.

[0025] The instruction channel of the demand analysis agent and the task allocation agent is connected, and the task allocation result is synchronized to the supervision agent to ensure real-time synchronization of data of each agent. Cross-module collaboration is automatically issuing a collaboration instruction to the associated module when the output of a module needs to be associated and adjusted (such as updating the role lines after the script is modified). Iterative response is to quickly locate the corresponding responsible agent after receiving the user's modification feedback, and coordinate related modules to complete optimization.

[0026] 2. The script generation module integrates a large language model application program interface (LLM API), including a demand analysis unit, a version control unit, and a quality verification unit. Specifically as follows: (1) Demand analysis unit: Through a standardized interface, call LLM (such as Doubao and Qwen large model), receive the standardized instructions of the task allocation agent, obtain the unstructured story text input by the user, and use the narrative logic and shot language analysis capabilities of LLM to automatically convert the unstructured demand into a standardized shot script containing "role setting-scene setting-shot instruction".

[0027] (2) Version control unit: Receive user operation instructions transmitted by the collaboration agent, automatically create independent version identifiers (including generation time and operation type) for the initial script generated by the LLM API and the script edited by the user, each version is associated with a complete project data package, covering the current state of the script text, role setting, scene data, shot parameters and modification log. Through the collaboration agent, respond to the user's version switching demand, support fast backtracking and difference comparison of historical versions, load the corresponding complete data, and avoid the risk of loss of creative data. The version data is stored in the version management database of the data support module.

[0028] (3) Quality verification unit: Build an input / output verification dual mechanism, and the verification rules and processing logic are linked with the supervision agent. The input verification mechanism checks the story text and demand instructions input by the user, and performs three core checks, namely text format check, minimum length check and sensitive word check. If the input verification fails, immediately feedback to the demand analysis agent through the collaboration agent, and return the error information containing the failure type to the user, without executing the LLM API call, to avoid invalid calculation. The output verification mechanism checks the standardized shot script generated by the LLM API for role consistency and scene continuity. If the output verification fails, the system automatically starts a retry strategy to control the LLM API to generate a script based on the failure reason, with a maximum of 3 retries; if 3 retries fail, trigger the manual filling mode through the supervision agent, and push the filling prompt and problem script label to the user; the verified script is synchronized to the supervision agent for acceptance through the collaboration agent, and is output to the visual generation module after passing.

[0029] 3. The visual generation module is the core execution module for visual content creation, including a reference picture generation unit, a storyboard generation unit, a dynamic video generation unit, and a batch creation unit. The details are as follows: (1) Reference picture generation unit: receive the creation instruction issued by the task allocation agent, synchronize the standardized storyboard script output by the script generation module, extract the character setting description and scene description through the text analysis component, and automatically convert them into structured prompt words. Call the picture generation model through the standardized interface to generate character reference pictures and scene reference pictures that meet the set style - the character reference picture includes close-up, front, side, and back views to ensure the integrity and recognizability of the character image; the scene reference picture includes panoramic, medium, and close-up level views to clearly define the scene space layout and core elements. After the style consistency check by the supervision agent calling the effect evaluation library, it is synchronized to the storyboard generation unit and the resource material library of the data support module.

[0030] (2) Storyboard generation unit: receive the script storyboard description information transmitted by the collaborative agent, based on the multi-source input data structure of character reference picture + scene reference picture + storyboard description, input the multi-source data to the picture generation model through the picture generation API, the model generates a storyboard picture that accurately matches the plot based on the visual features of the character reference picture, the spatial logic of the scene reference picture, and the composition requirements of the storyboard description. The storyboard picture automatically labels the shot number, single shot duration, camera movement mode (push / pull / pan / shift), and other technical parameters, supports the content matching degree check by the supervision agent of the AI assistant core module, and pushes it to the dynamic video generation unit after the check is passed.

[0031] (3) Dynamic video generation unit: dynamic content core output component, receive the video generation instruction issued by the task allocation agent, integrate the picture generation video API to realize the conversion from static to dynamic. For a single static storyboard picture, call the video generation model through the API to convert it into a 5-10 second dynamic video segment. After generating multiple dynamic segments corresponding to the storyboard, the system starts the picture connection component, automatically extracts the tail frame of the previous storyboard segment and the first frame of the next storyboard segment, extracts the key feature points through image feature matching algorithm, adjusts the picture angle, color parameters, and character action posture of the first frame of the next storyboard, so that the feature matching degree of the first and last frames is ≥90%, realizes the overall visual continuity of the spliced different storyboard segments, and avoids picture jumping.

[0032] (4) Batch creation unit: receives the batch generation requirements passed by the collaborative agent, supports the batch creation of storyboard based on the same set of role / scene prompt words. Users can set the generation quantity through the system interface, and the unit starts the random seed allocation mechanism to allocate a unique random seed for each generation task. Under the premise of maintaining the consistency of the core style, the seed difference drives the picture generation model to generate differentiated storyboard pictures in detail, meeting the user's multi-solution selection needs. The batch-generated storyboard pictures are automatically associated with version identifiers and synchronized to the version management unit and data support module for subsequent calling and backtracking.

[0033] 4. The post-production optimization module includes an intelligent dubbing unit, an audio effect matching unit, and an automatic editing unit.

[0034] (1) Intelligent dubbing unit: receives the standardized storyboard script passed by the collaborative agent, and constructs a "voice line matching-emotion adaptation-parameter adjustable" dubbing generation link through AI speech synthesis technology. First, match the basic voice line in the preset voice line library based on the role setting; second, extract the emotion label corresponding to the lines, adjust the intonation, stress and pause rhythm of the voice through the emotion adjustment algorithm, and realize accurate emotion matching; at the same time, support users to put forward speed and tone adjustment requirements through AI assistant, and automatically associate the generated dubbing audio with the storyboard timestamp, and output the lip synchronization data to the automatic editing unit to ensure audio-visual collaboration.

[0035] (2) Audio effect matching unit: receives dynamic video clips and storyboard emotion annotation data passed by the collaborative agent, and constructs an automatic audio effect generation mechanism of "scene analysis-emotion matching-material calling". Through the scene analysis component, identify the scene type of the video clip, and generate audio effect demand labels combined with emotion nodes; based on the labels, automatically match three types of audio effect resources - background music, environment sound, and special effect sound from the resource material library of the data support module; the generated audio effect materials are automatically labeled with timestamps to ensure precise synchronization with video picture actions, and the audio track data is synchronized to the automatic editing unit.

[0036] (3) Automatic editing unit: integrates visual timeline editor, builds the editing ability of "multi-track cooperation-precise scale synchronization-human-computer interaction optimization". The timeline editor contains five core tracks and unified millisecond time scale: video track is used to load the storyboard and dynamic video segments generated by the video generation module; voice track is used to load the voice audio generated by the intelligent voice unit; subtitle track is used to automatically load the synchronized subtitles generated based on the voice script; sound effect track is used to embed various sound effect materials of the sound effect matching unit; BGM track is used to carry background music resources. The unified millisecond time scale ensures the precise alignment of the timestamp of each track material, avoiding problems such as audio-visual misalignment and subtitle delay. At the same time, manual intervention is supported, and the user can directly drag each track material to modify the time position through the visual interface, or adjust the track attributes through the parameter setting panel. All modified operations are fed back to the preview screen in real time, realizing the editing effect of what you see is what you get. Finally, through multi-track cooperative rendering, the final film is output, which meets the quality requirements and is pushed to the export after being checked by the supervision agent.

[0037] 5. The data support module provides data support for the AI assistant core module, including training data set, resource material library and effect evaluation library.

[0038] (1) The training data set contains a large number of short drama scripts, character images, scene pictures, storyboard cases, dynamic video segments and audience feedback data, which provides training data support for the intelligent agent optimization of the text-to-image, image-to-image and image-to-video API and the AI assistant core module through the interface of the collaborative intelligent agent.

[0039] (2) The resource material library contains classified sound effects, background music, font styles and special effect templates, which provides precise material support for the intelligent voice unit and sound effect matching unit of the post-production optimization module through the standardized interface and the linkage with the collaborative intelligent agent, and supports quick retrieval and calling of materials according to scene and emotion tags.

[0040] (3) The effect evaluation library contains industry creation standards, audience preference models, quality evaluation indicators and video continuity determination rules (such as the threshold of the matching degree of the first and last frame characteristics), which provides the core basis for quality detection through deep linkage with the supervision intelligent agent. When monitoring the operation of each module, the supervision intelligent agent can real-time call the evaluation library rules to conduct real-time quality detection on the generated storyboard, dynamic video segments and final film content. If the detection is not up to standard, the pre-warning will be triggered and delivered to the corresponding module for optimization by the collaborative intelligent agent.

[0041] The closed-loop process based on the whole-process AI short drama creation system is as follows: For example, Figure 2As shown, the creation process using creative words as initial input is suitable for core creative concept scenarios. First, creative analysis: the user submits creative words (e.g., "sweet romance + workplace + coffee accident"). The requirement analysis agent calls training data to expand and generate structured requirements containing core conflicts, styles, scenes, and characters. After confirmation, these are converted into standard instructions. Second, script generation: the visual generation module calls the LLM API to generate the initial script. The version unit creates an identifier associated with all data; the quality verification unit performs double verification, retrying three times if it fails. Qualified scripts are then pushed out after being reviewed by the supervisory intelligent experience system.

[0042] Then, visual generation: the visual module extracts script descriptions to generate character / scene baseline images, and generates storyboard images via the image-generated image API (random seeds ensure batch differentiation); the image-generated video API converts 5-10 second clips, with the first and last frames matching ≥90% to ensure continuity, and pushes them to the post-production optimization module after verification.

[0043] In the final post-production output, the intelligent dubbing unit generates emotion-adapted audio, the sound effects matching unit calls up resources from the material library; the automatic editing unit synchronizes materials through a multi-track editor, with audio-visual errors ≤50ms; the finished product is exported in optimized platform format, and logs are archived.

[0044] like Figure 3 As shown, the creation process, which uses story text as the initial input, is suitable for scenarios with a complete story outline, with the core being the accurate conversion of unstructured text. First, the text is structured: the user uploads the story text, and the requirement analysis agent extracts characters, scenes, and conflict nodes, generating a standard requirement list and synchronizing it. Next, script management: the script module calls the LLM API to generate a script containing dialogue and camera instructions; version units create multiple version identifiers; quality verification ensures the absence of sensitive words, and only after passing the verification is the script pushed out. Finally, the creative output: the visual module generates a baseline image, storyboard, and 5-8 second clips (with a 92% matching rate between the first and last frames); the post-production optimization module completes the audio-visual synthesis; and after passing regulatory verification, the final product is exported.

[0045] like Figure 4 As shown, the creation process using reference videos as input is suitable for secondary creation scenarios. First, feature analysis: the user uploads a reference video, and the requirement analysis agent activates the video feature extraction component to accurately extract character, scene, and style features. Through a cross-modal transformation model, these features are converted into structured text instructions for character templates, scene templates, and style tags, which are then synchronized to the task allocation agent and data support module. Second, secondary creation is triggered and the script is generated. After receiving the instructions, the collaborative agent retains the core character features and basic scene features, incorporates the need for innovative plot conflicts, and inputs new text instructions. The task allocation agent sends the instructions that integrate the requirements to the script generation module. The module calls LLMAPI and, combined with the structured text instructions converted from the reference video, generates a new storyboard script.

[0046] Then, the visual generation module generates a reference picture highly matching the reference character appearance, a scene reference picture consistent with the reference scene layout, and integrates the camera operation features of the reference video into the shot list generation, etc. The post-production optimization module matches the sound effect style of the reference video (light and slow piano BGM) to complete the sound and picture synthesis.

[0047] Finally, the regulatory intelligent agent calls the "character consistency index" and "scene matching degree standard" in the effect evaluation library to perform double verification on the newly generated script: first, the character feature matching degree (the female protagonist's dress and the male protagonist's temperament are consistent with the reference video setting with a deviation of less than 90%); second, the plot innovation rationality. If the verification finds that the "male protagonist's lines are too gentle, deviating from the "cool temperament" of the reference video", the collaborative intelligent agent feeds back to the script generation module, triggering the LLM API to iteratively optimize based on the "cool tone + guide tone" parameters; the optimized script is verified by the regulatory intelligent agent to have a style matching degree of 88%, and the qualified script is associated with the reference video feature data and the secondary creation log, and is synchronized to the visual generation module, and finally output as a film.

[0048] The above is only a preferred embodiment of the present application, and it should be understood that the present application is not limited to the form disclosed herein, and should not be considered as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concepts described herein by the above teachings or related technical or knowledge.

Claims

1. A full-process AI short drama creation system, characterized in that, It includes an AI assistant core module, a script generation module, a visual generation module, a post-production optimization module, and a data support module. The script generation module, visual generation module, post-production optimization module, and data support module are scheduled by the AI ​​assistant core module to achieve real-time linkage and form a closed-loop creation chain.

2. The full-process AI short drama creation system according to claim 1, characterized in that: The core module of the AI ​​assistant is built from multi-task intelligent agents, including a demand analysis intelligent agent, a task allocation intelligent agent, a monitoring intelligent agent, and a collaborative intelligent agent; The demand parsing intelligent agent is used for intelligent demand parsing, and the specific workflow is demand reception → demand parsing → standardized instruction generation; The task allocation agent is used for task allocation, and the specific workflow is task decomposition → module matching → parameter distribution → priority sorting. The regulatory intelligence agent is used to monitor processes and results. The specific workflow is: real-time progress tracking → dynamic quality detection → compliance review → data log recording. The collaborative intelligent agent is used for seamless linkage between various intelligent agents and functional modules. The specific workflow is data transfer → cross-module collaboration → iterative response.

3. The full-process AI short drama creation system according to claim 2, characterized in that: Request reception refers to users inputting story text through text input boxes or by uploading files; request parsing refers to extracting key elements, including style, duration, and characters; standardized instruction generation is the process of converting the parsed requests into structured instructions that can be recognized by machines. Task decomposition breaks down the overall instructions output by the requirement analysis agent into sub-tasks such as script generation, character and scene generation, and storyboard and video generation, clarifying the output standards and time nodes for each task; module matching matches execution modules to sub-tasks based on the real-time load of each functional sub-module; parameter delivery transmits task instructions and related data to the corresponding modules; priority sorting supports users marking urgent requirements, and the agent automatically adjusts task priorities, allocating more computing resources to urgent tasks; Real-time progress tracking is achieved by synchronously collecting task execution progress through the status interfaces of each module, providing users with visual charts, and automatically issuing warnings when tasks are delayed. Dynamic quality monitoring is an effect evaluation library that connects to the data support module. It verifies the output of each stage in real time, including whether the script conforms to the narrative rhythm of the short drama, whether the character models match the settings, and whether there are any continuity errors in the video. If the standards are not met, rectification instructions are generated. Compliance review refers to the built-in short drama content compliance database, which automatically detects whether there is any illegal content in the script, dialogue, and visual elements. Data log recording refers to the full retention of task data, modification records, and quality inspection results at each stage. Data transfer is the instruction channel that connects the demand analysis agent and the task allocation agent, and synchronizes the task allocation results to the supervision agent. Cross-module collaboration automatically sends collaboration instructions to related modules when the output of a module needs to be adjusted; iterative response involves receiving user feedback on modifications, locating the responsible agent, and coordinating relevant modules to complete the optimization.

4. The full-process AI short drama creation system according to claim 1, characterized in that: The script generation module integrates a large language model application programming interface, including a requirements parsing unit, a version control unit, and a quality verification unit; Requirements Analysis Unit: Through standardized interfaces, it calls LLM, receives standardized instructions from the task allocation agent, obtains unstructured story text input by the user, and uses LLM's narrative logic and camera language analysis capabilities to automatically transform unstructured requirements into standardized storyboards containing character settings, scene settings, and camera instructions. Version control unit: Receives user operation instructions from the collaborative agent, automatically creates independent version identifiers for the initial script generated by the LLM API and the script edited by the user, and associates each version with a complete project data package, covering the script text, character settings, scene data, storyboard parameters and modification logs in the current state; version data is synchronously stored in the version management database of the data support module; Quality verification unit: Constructs a dual input / output verification mechanism, with verification rules and processing logic linked with the supervisory intelligent agent; The input validation mechanism performs text format checks, minimum length checks, and sensitive word checks on the story text and requirement instructions entered by the user. If the input validation fails, it immediately feeds back to the requirement parsing agent through the collaborative agent, which then returns an error message containing the failure type to the user and does not execute the LLM API call. The output verification mechanism performs role consistency and scene continuity checks on the standardized storyboard scripts generated by the LLM API. If the output verification fails, the system automatically starts a retry strategy, controlling the LLM API to regenerate the script based on the reason for the failure, with a maximum of 3 retries. If all three retries fail, the manual completion mode is triggered through the supervisory intelligent agent, which pushes completion prompts and problem script annotations to the user. The script that passes the verification is synchronized to the supervisory intelligent agent for acceptance by the collaborative intelligent agent, and output to the visual generation module after passing the test.

5. The full-process AI short drama creation system according to claim 4, characterized in that: The visual generation module is the core execution module for visual content creation, including the baseline image generation unit, storyboard image generation unit, dynamic video generation unit, and batch creation unit. Baseline Image Generation Unit: Receives creative instructions from the task allocation agent, synchronously acquires the standardized storyboard output by the script generation module, extracts character setting descriptions and scene descriptions through the text parsing component, and automatically converts them into structured prompts; it calls the text-generated image large model through a standardized interface to generate character baseline images and scene baseline images that conform to the set style—the character baseline image includes facial close-ups and front, side, and back views; the scene baseline image includes panoramic, medium, and close-up layered views; After the regulatory intelligent agent calls the effect evaluation library to complete the style consistency verification, it is synchronized to the resource material library of the storyboard generation unit and the data support module; Storyboard Generation Unit: Receives storyboard description information from the collaborative intelligent agent. Based on a multi-source input data structure of character baseline image + scene baseline image + storyboard description, it inputs the multi-source data into the graph-generated image large model via the graph-generated image API. The model generates storyboards that match the plot based on the visual features of the character baseline image, the spatial logic of the scene baseline image, and the composition requirements of the storyboard description. The storyboards automatically label the technical parameters of shot number, single shot duration, and camera movement method. It supports the monitoring intelligent agent of the AI ​​assistant core module to verify the content matching degree. After successful verification, it is pushed to the dynamic video generation unit. Dynamic Video Generation Unit: The core output component for dynamic content, receiving video generation instructions from the task allocation agent and integrating the image-generated video API to achieve static-to-dynamic conversion; for a single static storyboard image, it calls the video generation model through the API to convert it into a 5-10 second dynamic video clip; after generating dynamic clips corresponding to multiple storyboard images, the system activates the screen transition component, automatically extracting the last frame of the previous storyboard clip and the first frame of the next storyboard clip, extracting feature points through an image feature matching algorithm, and adjusting the screen angle, color parameters, and character action posture of the first frame of the next storyboard clip to ensure that the feature matching degree of the first and last frames is ≥90%; Batch Creation Unit: Receives batch generation requests from the collaborative intelligent agent, supporting batch creation of storyboards based on the same set of character / scene prompts; users set the generation quantity through the system interface, and the unit activates a random seed allocation mechanism to assign a unique random seed to each generation task; while maintaining core style consistency, seed differentiation drives the generation of differentiated storyboards with varying visual details from the large-scale image model; batch-generated storyboards are automatically associated with version identifiers and synchronized to the version control unit and data support module for subsequent retrieval and retrospective analysis.

6. The full-process AI short drama creation system according to claim 1, characterized in that: The post-production optimization module includes an intelligent dubbing unit, a sound effect matching unit, and an automatic editing unit; Intelligent Dubbing Unit: Receives standardized storyboard scripts from the collaborative intelligent agent and constructs a dubbing generation link through AI speech synthesis technology, encompassing voice matching, emotion adaptation, and adjustable parameters. First, it matches basic voices from a preset voice library based on the character setting. Second, it extracts the emotion tags corresponding to the lines and adjusts the tone, emphasis, and pause rhythm of the speech through an emotion adjustment algorithm. Simultaneously, it supports users submitting speech speed and tone adjustment requests through an AI assistant. The generated dubbing audio is automatically associated with the storyboard timestamp and simultaneously outputs lip-sync data to the automatic editing unit. Sound Effect Matching Unit: Receives dynamic video clips and storyboard emotion annotation data from the collaborative intelligent agent, and constructs an automated sound effect generation mechanism of scene analysis, emotion matching, and material retrieval; identifies the scene type of the video clip through the scene analysis component, and generates sound effect requirement tags by combining emotion nodes; automatically matches three types of sound effect resources—background music, ambient sound, and special effects sound—from the resource material library of the data support module based on the tags; automatically timestamps the generated sound effect materials, and pushes the sound effect track data synchronously to the automatic editing unit; Automatic Editing Unit: Integrates a visual timeline editor, which includes five interconnected core tracks and a unified millisecond-level time scale: the video track is used to load storyboards and dynamic video clips output by the video generation module; the dubbing track is used to mount dubbing audio generated by the intelligent dubbing unit; and the subtitle track is used to automatically load synchronized subtitles generated based on the dubbing dialogue. The sound effects track is used to embed various sound effects materials of the sound effects matching unit; the BGM track is used to carry background music resources. The unified millisecond-level time scale allows users to manually adjust it. Users can drag and drop the materials on each track to modify the time position through the visual interface, or adjust the track attributes through the parameter setting panel. All modifications are reflected in the preview screen in real time. Finally, through multi-track collaborative rendering, the final product that meets the quality requirements is output and pushed to the export stage after being verified by the monitoring intelligent agent.

7. The full-process AI short drama creation system according to claim 1, characterized in that: The data support module provides data assurance for the core module of the AI ​​assistant, including training datasets, resource material libraries, and effect evaluation libraries; The training dataset includes short drama scripts, character images, scene images, storyboard examples, dynamic video clips, and audience feedback data. Through the interface of the collaborative agent, it provides training data support for the agent optimization of text-to-image, image-to-image large model, image-to-video API, and AI assistant core modules. The resource library contains categorized sound effects, background music, font styles, and special effects templates. It works in conjunction with a collaborative intelligent agent through a standardized interface. Based on the instructions issued by the task allocation intelligent agent, it provides accurate material support for the intelligent dubbing unit and sound effect matching unit of the later optimization module. The effect evaluation library includes industry creation standards, audience preference models, quality evaluation indicators, and video continuity judgment rules. It is deeply integrated with the regulatory intelligent agent to provide a basis for quality detection. When monitoring the operation of each module, the regulatory intelligent agent retrieves the evaluation library rules in real time to perform real-time quality detection on the generated storyboards, dynamic video clips, and finished content. If the detection fails to meet the standards, an early warning is triggered, which is then passed to the corresponding module by the collaborative intelligent agent for optimization.

8. A full-process AI short drama creation system according to any one of claims 1-7, characterized in that: The initial input data for the system includes creative words, story text, and video; When the initial input is a creative word, the process begins with creative parsing. The user submits the creative word, and the requirement parsing agent expands it by calling the training data, generating structured requirements that include core conflicts, styles, scenes, and roles. After confirmation, these requirements are converted into standard instructions. Next, the script is generated. The script generation module calls the LLM API to generate the initial script, and the version control unit creates an identifier that is associated with all the data. The quality verification unit performs double verification, retrying three times if it fails. Qualified scripts are pushed to the supervisory intelligent experience system after acceptance. Then, content generation occurs: the visual generation module extracts script descriptions to generate character / scene baseline images, which are then used to generate storyboards via the image-generated image API; the image-generated video API converts these into 5-10 second clips with a first and last frame matching degree of ≥90%, and after verification, they are pushed to the post-production optimization module; finally, post-production output occurs: the intelligent dubbing unit generates emotion-adapted audio, the sound effects matching unit calls resources from the material library; the automatic editing unit synchronizes materials through a multi-track editor, with an audio-visual error of ≤50ms; the final product is optimized and exported according to the platform format, and logs are archived.

9. The full-process AI short drama creation system according to claim 8, characterized in that: When the initial input is a story text, the text is first structured. The user uploads the story text, and the requirement parsing agent extracts the characters, scenes and conflict nodes, generates a standard requirement list and synchronizes it. Then, the script is managed. The script generation module calls the LLMAPI to generate a script containing dialogue and camera instructions, and the version control unit creates multiple version identifiers. The quality verification unit ensures the absence of sensitive words and pushes the video after it passes the verification. In the final creative output, the visual generation module generates a baseline image, storyboard image, and 5-8 second clips, with a first and last frame matching degree of ≥92%. The post-production optimization module completes the audio-visual synthesis, and the final film is exported after passing the regulatory verification.

10. The full-process AI short drama creation system according to claim 8, characterized in that: When using a reference video as input, this approach is suitable for secondary creation scenarios. First, feature analysis: the user uploads a reference video, and the requirement analysis agent activates the video feature extraction component to extract character features, scene features, and style features. Through a cross-modal transformation model, these features are converted into structured text instructions for character templates, scene templates, and style tags, which are then synchronized to the task allocation agent and data support module. Second, secondary creation is triggered and the script is generated: after receiving the instructions, the collaborative agent retains the core character features and baseline scene features, innovates the need for plot conflict, and inputs new text instructions. The task allocation agent sends instructions that integrate the requirements to the script generation module, calls the LLM API, and combines the structured text instructions converted from the reference video to generate a new storyboard script. Then, the visual generation module generates a baseline image that matches the appearance of the reference character and a scene baseline image that matches the layout of the reference scene based on the script and the features of the reference video; the camera movement features of the reference video are incorporated when the storyboard is generated; the post-production optimization module matches the sound effect style of the reference video to complete the audio-visual synthesis. Finally, the regulatory agent retrieves the role consistency index and scene matching standard from the effect evaluation library to perform dual verification on the newly generated script: first, the matching degree of role characteristics; second, the rationality of plot innovation. If the verification fails, the collaborative agent will provide feedback to the script generation module, triggering iterative optimization of the LLM API parameters; qualified scripts will be associated with reference video feature data and secondary creation logs, synchronized to the visual generation module, and finally output as finished products.

Citation Information

Cited By

  • An ai video production system and method based on multi-model abstraction layer

    CN122205197A

  • An ai automatic script structuring and video parsing integrated generation method and system

    CN122205200A