Automatic short play creation and local editing method based on multi-agent collaborative mechanism

CN122513637APending Publication Date: 2026-08-04BEIJING XINGMAI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XINGMAI INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-05-11
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

然而,该类系统存在明显缺陷:一是缺乏视频的局部微调与编辑能力,系统为高度封装的黑盒,一旦生成的视频中存在不符合预期的微小细节,用户无法通过简单的自然语言指令进行局部修改,通常只能修改原始提示词并让系统完全重新生成,这不仅耗费高昂的算力资源,且重新生成的视频画面与上一版往往截然不同,缺乏画面与逻辑的连贯一致性;二是无法编排长视频并确保角色形象一致性,此类系统大多只能生成单一维度的极短素材,无法自动化地串联配音、背景音乐和多段分镜以形成具有完整剧情的长视频,且在连续生成多个包含同一角色的分镜视频时,由于缺乏对角色特征的全局约束与记忆机制,主角的容貌和服饰在不同镜头中常常发生严重的跳变现象

Benefits of technology

1、本发明通过引入剧情张力曲线向量与动态身份特征融合权重,实现了在无需模型训练微调的前提下,对长视频中角色面部特征的高精度一致性约束,有效解决了传统生成式模型中连续分镜角色形象跳变的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513637A_ABST
    Figure CN122513637A_ABST
Patent Text Reader

Abstract

This invention discloses an automated short drama creation and partial editing method based on a multi-agent collaborative mechanism, specifically relating to the field of short drama creation and editing technology. The method involves: acquiring natural language commands input by the user; generating a structured script containing plot tension parameters and node verification information through an agent; dynamically determining the character feature fusion weights based on the plot tension parameters and generating storyboards by combining preset character reference images to maintain character consistency; generating audio based on the structured script and extracting timestamp information from the vocal units; constraining video generation with audio duration and applying lip-sync enhancement processing to the video frame intervals corresponding to the timestamp information; and, upon receiving an editing command, triggering partial regeneration or reusing existing materials based on the comparison results between the target modification node and the corresponding node verification information in the structured script to synthesize the final video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of short drama creation and editing technology, and more specifically, to an automated short drama creation and partial editing method based on a multi-agent collaborative mechanism. Background Technology

[0002] Currently, AI-based video generation technologies and products are mainly divided into two categories.

[0003] The first type is generative large-scale model systems that generate video clips based on plain text input. These systems are end-to-end generation engines where users input descriptive text prompts, and the system outputs short video clips in one go after inference and computation. However, this type of system has significant drawbacks: First, it lacks the ability to fine-tune and edit the video locally. The system is a highly encapsulated black box, and if there are minor details in the generated video that do not meet expectations, users cannot make local modifications through simple natural language commands. Usually, they can only modify the original prompts and have the system completely regenerate the video. This not only consumes high computing resources, but the regenerated video is often completely different from the previous version, lacking consistency in visuals and logic. Second, it cannot arrange long videos and ensure the consistency of character appearance. Most of these systems can only generate very short clips in a single dimension and cannot automatically connect dubbing, background music, and multiple storyboards to form a long video with a complete storyline. Moreover, when generating multiple storyboard videos containing the same character, due to the lack of global constraints and memory mechanisms for character characteristics, the protagonist's appearance and clothing often show serious jumps in different shots.

[0004] The second type is text-driven intelligent video editors. These tools use speech recognition technology to convert existing videos into text, allowing users to trim, splice, and replace existing video footage by deleting, searching, or modifying text. However, their main drawbacks are: they are limited to splicing existing footage and lack the ability to create from scratch, making them ineffective for short drama scenes that require entirely new AI characters or specific storylines; moreover, the human operation threshold and cost remain high. Existing editing tools still heavily rely on traditional timeline editing modes when making complex modifications, and users inevitably need to master professional video editing concepts such as tracks, keyframes, and layers, making it impossible to achieve fully automatic generation and modification of a single sentence driven by pure natural language. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an automated short drama creation and partial editing method based on a multi-agent collaborative mechanism to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: An automated short drama creation and partial editing method based on a multi-agent collaborative mechanism includes the following steps: obtaining natural language creation instructions and editing instructions input by the user; The system receives the natural language creation instructions through the intelligent agent center, injects preset plot tension scoring constraint rules into the system prompt words, calls the large language model to generate the storyboard script for the natural language creation instructions, forces the large language model to output a plot tension scalar when outputting each storyboard text, combines the plot tension scalars corresponding to each storyboard into a plot tension curve vector, and serializes the storyboard script into a structured intermediate script with node semantic text vectors. The intelligent agent central system calls the graph-generating tool to generate character reference images for each character in the storyboard script and persist them as global anchor files. When subsequent storyboard scenes are generated, the identity feature fusion weights are dynamically calculated through a preset mapping function based on the plot tension scalar of the corresponding storyboard in the plot tension curve vector. The identity feature fusion weights and the global anchor files are then passed into the graph-generating interface to achieve consistency constraints on the facial features of characters in the storyboard scenes without the need for fine-tuning the model. The intelligent agent central system calls the speech synthesis tool to convert the dialogue text corresponding to each scene into audio files according to the storyboard script, and parses the alignment metadata returned by the speech synthesis tool to extract the timestamp of each pronunciation unit and construct a phoneme timestamp sequence as a strong constraint input parameter for subsequent video generation. The video generation tool is invoked through the intelligent agent's central hub. The total physical duration of the audio file is used as the duration constraint parameter for video generation. Based on the phoneme timestamp sequence, the timestamp of each vocal unit is mapped to a video frame index using a preset frame mapping formula. When the video generation interface is invoked, lip-sync enhancement prompts or local motion amplitude parameters are injected into the keyframe interval corresponding to the video frame index to achieve high-precision synchronization between lip-sync and audio. When the editing instruction is received, the editing agent parses the editing instruction to determine the target modification node. The semantic features of the core parameter dictionary of the modified target node are compared with the node semantic text vector of the corresponding node in the structured intermediate script. Only when the distance change exceeds the preset semantic distance tolerance threshold is the node determined to be in a dirty state and the local regeneration of the corresponding material is triggered. Unchanged nodes directly reuse existing cached materials. The video track, audio track, and subtitle track corresponding to each scene are multi-track mixed and compressed to output the final film.

[0007] In a preferred embodiment, the plot tension parameter is the plot tension scalar corresponding to each scene. The agent injects preset scoring constraint rules into the system prompts of the large language model, forcing the large language model to output the plot tension scalar when outputting the text of each scene, and combines the plot tension scalars of each scene into a plot tension curve vector.

[0008] In a preferred embodiment, the character feature fusion weight is dynamically calculated using the following formula: α = 1.0 - k × V_tension, where α is the character feature fusion weight, V_tension is the plot tension scalar, and k is the experience adjustment coefficient.

[0009] In a preferred embodiment, the node verification information is a semantic text vector calculated by serializing the core parameter dictionary of the storyboard node; the structured script uses YAML format to carry the storyboard structure tree, and the core parameter dictionary of each storyboard node includes the scene prompts, timbre and action intensity parameters of that storyboard.

[0010] In a preferred embodiment, the comparison result is obtained by: extracting the text feature vector of the core parameter dictionary of the modified target node, calculating the cosine distance between the text feature vector and the node semantic text vector of the corresponding node in the structured script, and determining whether the cosine distance exceeds a preset semantic distance tolerance threshold.

[0011] In a preferred embodiment, the timestamp information of the pronunciation unit is a phoneme timestamp sequence, which is extracted by parsing the alignment metadata returned by the speech synthesis tool.

[0012] In a preferred embodiment, the lip-sync enhancement process includes: mapping the timestamps of each articulation unit to video frame indices according to the phoneme timestamp sequence using a preset frame mapping formula, and injecting lip-sync enhancement prompts into the keyframe intervals corresponding to the video frame indices when the video generation tool is invoked.

[0013] In a preferred embodiment, the frame mapping formula is: Frameidx=(Tms / 1000)×FPS, where Tms is the timestamp of the vocal unit in the audio in milliseconds, FPS is the frame rate of the target generated video, and Frameidx is the specific frame sequence number corresponding to the vocal unit in the video.

[0014] In a preferred embodiment, the agent is an agent hub based on the ReAct paradigm, which can perform logical reasoning after receiving user instructions, autonomously decide the specific external tools to be invoked in the current step, and decide the next action based on the execution results returned by the tools, until the closed loop of the entire multimodal video task is completed.

[0015] In a preferred embodiment, the structured script may use JSON, XML, or a custom binary format instead of YAML to achieve decoupling between structured expression and local parameter modification.

[0016] The technical effects and advantages of this invention are as follows: 1. This invention introduces the plot tension curve vector and dynamic identity feature fusion weights to achieve high-precision consistency constraints on the facial features of characters in long videos without the need for model training and fine-tuning, effectively solving the problem of character image jumps in continuous shot splits in traditional generative models.

[0017] 2. This invention constructs a phoneme timestamp sequence by parsing the alignment metadata of the speech synthesis tool, and realizes dynamic keyframe prompt word injection at the phoneme level by combining the frame mapping formula. It is the first to create a reverse timing mechanism that constrains the video duration before audio, which greatly improves the synchronization accuracy between lip movements and audio, and eliminates the tedious steps of manually aligning audio and video.

[0018] 3. This invention transforms traditional black-box redrawing into white-box editing by using structured intermediate scripts with node semantic text vectors. When a modification instruction is received, it can accurately compare semantic changes and trigger local regeneration only for nodes that have undergone substantial changes, which greatly reduces computing power consumption and modification waiting time, while ensuring the consistency of unmodified parts.

[0019] 4. This invention uses a ReAct-based intelligent agent hub to abstract and encapsulate all underlying atomic capabilities such as music acquisition, dubbing, image generation, and video synthesis into callable external tools. Users are completely freed from complex timeline operation interfaces and can achieve fully automatic short drama creation and precise local editing simply by communicating in natural language. Attached Figure Description

[0020] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is a flowchart illustrating the workflow of the automated short drama creation and partial editing method based on a multi-agent collaborative mechanism of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Example: Please refer to Figure 1 An automated short drama creation and partial editing method based on a multi-agent collaborative mechanism includes the following steps: Obtain user-input natural language creation and editing commands; The system receives the natural language creation instructions through the intelligent agent center, injects preset plot tension scoring constraint rules into the system prompt words, calls the large language model to generate the storyboard script for the natural language creation instructions, forces the large language model to output a plot tension scalar when outputting each storyboard text, combines the plot tension scalars corresponding to each storyboard into a plot tension curve vector, and serializes the storyboard script into a structured intermediate script with node semantic text vectors. The intelligent agent central system calls the graph-generating tool to generate character reference images for each character in the storyboard script and persist them as global anchor files. When subsequent storyboard scenes are generated, the identity feature fusion weights are dynamically calculated through a preset mapping function based on the plot tension scalar of the corresponding storyboard in the plot tension curve vector. The identity feature fusion weights and the global anchor files are then passed into the graph-generating interface to achieve consistency constraints on the facial features of characters in the storyboard scenes without the need for fine-tuning the model. The intelligent agent central system calls the speech synthesis tool to convert the dialogue text corresponding to each scene into audio files according to the storyboard script, and parses the alignment metadata returned by the speech synthesis tool to extract the timestamp of each pronunciation unit and construct a phoneme timestamp sequence as a strong constraint input parameter for subsequent video generation. The video generation tool is invoked through the intelligent agent's central hub. The total physical duration of the audio file is used as the duration constraint parameter for video generation. Based on the phoneme timestamp sequence, the timestamp of each vocal unit is mapped to a video frame index using a preset frame mapping formula. When the video generation interface is invoked, lip-sync enhancement prompts or local motion amplitude parameters are injected into the keyframe interval corresponding to the video frame index to achieve high-precision synchronization between lip-sync and audio. When the editing instruction is received, the editing agent parses the editing instruction to determine the target modification node. The semantic features of the core parameter dictionary of the modified target node are compared with the node semantic text vector of the corresponding node in the structured intermediate script. Only when the distance change exceeds the preset semantic distance tolerance threshold is the node determined to be in a dirty state and the local regeneration of the corresponding material is triggered. Unchanged nodes directly reuse existing cached materials. The video track, audio track, and subtitle track corresponding to each scene are multi-track mixed and compressed to output the final film.

[0023] In this embodiment, the intelligent agent hub refers to an intelligent agent based on the ReAct paradigm, which can perform logical reasoning after receiving user instructions, autonomously decide which specific external tool to call in the current step, and decide the next action based on the execution result returned by the tool, continuously iterating until the entire multimodal video task loop is completed.

[0024] The aforementioned plot tension curve vector refers to a numerical sequence composed of plot tension scalars corresponding to each scene, arranged in the order of the scenes. This numerical sequence is used to quantify the emotional intensity of each scene and serves as a key control parameter for subsequent image and video generation interface calls. The plot tension scalar is a scalar value ranging from [0.0, 1.0]. By injecting preset scoring constraint rules into the system prompts of the large language model, the large language model is forced to calculate and output this value when outputting the text of each scene.

[0025] The dynamic identity feature fusion weight refers to the control weight parameter obtained by converting the plot tension scalar through a preset mapping function, instead of passing a hard-coded default weight when calling the graph-generated graph interface. The calculation formula is as follows: ,in, For identity feature fusion weights, As a scalar measure of plot tension, This is an empirical adjustment coefficient, ranging from 0.1 to 0.3. This coefficient is the optimal statistical threshold range determined through extensive experiments in generating raw images, manually comparing facial distortion rates and motion stiffness. It is used to control the rate attenuation of tension on the weights. When the plot tension is intense, it automatically decreases. The value is adjusted to release the amplitude of the generated model's actions; when the plot is calm, it is automatically increased. Values ​​are used to enhance the consistency of facial features. This ensures a high degree of facial consistency across long videos without requiring any model tweaking.

[0026] The phoneme timestamp sequence refers to an array of timestamps accurate to the millisecond level, precisely extracted by deeply analyzing the aligned metadata returned by a third-party speech synthesis interface to accurately extract the start and end timestamps of the pronunciation of each word or phoneme. This array is explicitly returned to the agent as a strong constraint parameter for subsequent video generation tools to inject frame-level keyframe cue words.

[0027] The frame mapping formula is: ,in This refers to the timestamp of the specific speech unit in the audio returned by the speech synthesis interface, in milliseconds. Set the frame rate for generating the target video, such as 24 frames per second. This refers to the calculated specific frame sequence number corresponding to the pronunciation unit in the video. When calling the video generation interface that supports keyframe control, enhanced prompts or increased local motion amplitude parameters are precisely injected into the keyframe interval corresponding to the video frame index, thereby achieving extremely high-precision lip-sync through engineering methods of pure interface scheduling.

[0028] The node semantic text vector refers to the process of serializing the core parameter dictionary of each storyboard node and calculating its overall text vector, which is then attached to the node, when the storyboard script is serialized into a structured intermediate script. The structured intermediate script preferably uses YAML format to carry the storyboard structure tree. The core parameter dictionary of each storyboard node includes parameters such as visual cues, timbre, and action intensity for that storyboard. This semantic text vector is used for rigorous state comparison in subsequent editing stages, transforming the originally static script into a robust state machine with verification capabilities.

[0029] When a user submits natural language modification suggestions, the editing agent parses the modification instructions to determine the target modification node. It then compares the semantic features of the core parameter dictionary of the modified target node with the semantic text vector of the corresponding node in the previous round of structured intermediate script. The semantic feature extraction can be achieved by extracting text feature vectors from the core parameter dictionary, and the distance comparison uses cosine distance or Euclidean distance in vector space. Only when the distance change exceeds a preset semantic distance tolerance threshold is the node determined to be in a "dirty" state, triggering a local regeneration of the corresponding material; unchanged nodes directly reuse existing cached material. This modification completely breaks away from the traditional black-box redrawing state where every modification is wrong, achieving high-precision local modification with extremely low computational power consumption.

[0030] The tools invoked by the intelligent agent's central processing unit include, but are not limited to: context-aware text generation tools, graph-to-graph tools, speech synthesis tools, video generation tools, intermediate state protocol tools, final video rendering tools, semantic similarity matching tools, and environment configuration and persistence tools. Among these, the final video rendering tool, through automated orchestration of the FFMpeg process, performs multi-threaded hybrid compression rendering of aligned video, audio, and subtitle tracks to output the final video. While the external models or third-party interfaces invoked by the aforementioned tools are themselves known technologies, this invention independently improves the data transmission scheme when scheduling these interfaces, introducing novel parameter variables such as plot tension curve vectors, dynamic identity feature fusion weights, phoneme timestamp sequences, and node semantic text vectors. This reorganization of input-output dependencies and the workflow orchestration combination method based on self-created variables constitute the core innovation of this invention.

[0031] As an alternative, the structured intermediate script can also use JSON, XML, or a custom binary format to carry the storyboard structure tree, as long as it can decouple the structured expression from local parameter modifications. Another alternative is that character consistency anchors can be implemented by calling the fine-tuning module during system initialization to temporarily train a very lightweight character-specific model file using multiple generated reference images; however, this alternative will significantly increase the pipeline's computational power consumption and waiting time.

[0032] Existing AI video generation systems have significant limitations in short drama creation scenarios: generative large-scale model systems lack local editing capabilities and character consistency guarantees, while intelligent video editors are limited to existing materials and rely on professional operations. This invention constructs an intelligent agent hub based on the ReAct (Reasoning and Acting) paradigm, combined with innovative mechanisms such as plot tension curve vectors, dynamic identity feature fusion weights, phoneme timestamp sequences, and node semantic text vectors. This abstracts the entire video creation process into a toolchain that can be scheduled by the intelligent agent, achieving a fully automated closed loop from script generation, character anchoring, audio-visual synchronization to local editing, completely breaking down the professional barriers and computing power bottlenecks of traditional video creation.

[0033] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0034] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0035] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An automated short drama creation and partial editing method based on a multi-agent collaborative mechanism, characterized by: This includes the following steps: obtaining natural language commands input by the user; A structured script containing plot tension parameters and node verification information is generated through an intelligent agent; The character feature fusion weights are dynamically determined based on the aforementioned plot tension parameters, and storyboards are generated by combining them with preset character reference images. Audio is generated based on the structured script, and the timestamp information of the pronunciation units is extracted; The video generation is constrained by the duration of the audio, and lip-sync enhancement processing is applied to the video frame interval corresponding to the timestamp information; When an editing instruction is received, based on the comparison result between the target modification node and the corresponding node verification information in the structured script, local regeneration and reuse of existing materials are triggered. Combine the video track, audio track, and subtitle track of each scene into the final video.

2. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The plot tension parameter is the plot tension scalar corresponding to each scene. The agent injects preset scoring constraint rules into the system prompts of the large language model, forcing the large language model to output the plot tension scalar when outputting the text of each scene, and combines the plot tension scalars of each scene into a plot tension curve vector.

3. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The character feature fusion weight is dynamically calculated using the following formula: α = 1.0 - k × V_tension, where α is the character feature fusion weight, V_tension is the plot tension scalar, and k is the experience adjustment coefficient.

4. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The node verification information is a semantic text vector calculated by serializing the core parameter dictionary of the storyboard node; the structured script uses YAML format to carry the storyboard structure tree, and the core parameter dictionary of each storyboard node includes the on-screen prompts, timbre and action intensity parameters of that storyboard.

5. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The comparison results are obtained by: extracting the text feature vector of the core parameter dictionary of the modified target node, calculating the cosine distance between the text feature vector and the node semantic text vector of the corresponding node in the structured script, and determining whether the cosine distance exceeds a preset semantic distance tolerance threshold.

6. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The timestamp information of the pronunciation unit is a phoneme timestamp sequence, which is extracted by parsing the alignment metadata returned by the speech synthesis tool.

7. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 6, characterized in that: The lip-sync enhancement process includes: mapping the timestamps of each articulation unit to video frame indices using a preset frame mapping formula based on the phoneme timestamp sequence, and injecting lip-sync enhancement prompts into the keyframe intervals corresponding to the video frame indices when the video generation tool is invoked.

8. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 7, characterized in that: The frame mapping formula is: Frameidx=(Tms / 1000)×FPS, where Tms is the timestamp of the vocal unit in the audio in milliseconds, FPS is the frame rate of the target generated video, and Frameidx is the specific frame sequence number corresponding to the vocal unit in the video.

9. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The intelligent agent is a central intelligent agent based on the ReAct paradigm. After receiving user instructions, it can perform logical reasoning, autonomously decide the specific external tools to be invoked in the current step, and decide the next action based on the execution results returned by the tools, until the entire multimodal video task is completed.

10. The automated short drama creation and partial editing method based on a multi-agent collaborative mechanism according to claim 1, characterized in that: The structured script can use JSON, XML, or a custom binary format instead of YAML to achieve decoupling between structured expression and local parameter modification.