Structured editing instruction generation method and system based on director language model

By using a structured editing instruction generation method based on a director's language model, the problem of existing automated video editing technologies being unable to understand the director's intentions is solved, enabling efficient and personalized video production while reducing manual intervention and costs.

CN121509742APending Publication Date: 2026-02-10SHANGHAI FRAME VIEW TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511653120.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing automated video editing technologies cannot deeply understand the director's intentions and artistic style, resulting in videos that appear stiff and template-like, failing to meet the demands of high-quality commercial use, and relying heavily on manual post-production intervention, which is costly and inefficient.

Method used

A structured editing instruction generation method based on the director's language model is adopted. The script text is parsed by a large language model, and shot planning parameters are generated by combining director rules. Structured editing instructions are automatically generated, and intelligent output from script to finished film is achieved through material scheduling and editing execution.

Benefits of technology

It achieves deep semantic understanding, generates highly personalized and stylized editing instructions, automates the editing process, reduces human intervention, and significantly improves video production efficiency and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509742A_ABST
    Figure CN121509742A_ABST
Patent Text Reader

Abstract

The invention discloses a structured editing instruction generation method and system based on a director language model. The method comprises the following steps: analyzing a script text by utilizing a large language model, and outputting a structured semantic unit; based on a semantic unit, combining a director rule and model reasoning to generate lens planning parameters; fusing the parameters and the script into a machine-readable structured editing instruction; scheduling or generating a visual material according to the instruction; and finally, converting the instruction and the material into an editing project file and outputting a video. The method also supports style migration and multi-modal feedback learning, can adjust the output according to the reference style, and continuously optimizes according to the user feedback. According to the system, automatic and high-quality editing from scripts to slices is realized, and the system is particularly suitable for commercial video production scenes with high requirements on efficiency and creativity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and digital video processing technology, specifically to a method and system for generating structured editing instructions based on a director's language model. Background Technology

[0002] With the explosive growth of digital media content, the market has placed higher demands on the production efficiency of high-quality video content. Currently, automated video editing technology mainly relies on predefined templates and simple rule matching. While these systems can achieve basic editing functions, they have significant limitations: they lack a deep understanding of script semantics, cannot capture the director's creative intentions and artistic style preferences, and therefore struggle to generate videos with professional narrative logic and emotional expressiveness.

[0003] Specifically, existing technologies cannot accurately translate unstructured, emotionally rich script text into structured editing instructions that include a range of professional elements such as camera language, pacing, transitions, and special effects. This results in automated editing products often appearing stiff and template-based, failing to meet the demands of commercial scenarios such as advertising and short dramas, which require high-quality content and stylistic precision. These still necessitate significant manual post-production intervention, leading to high costs and low efficiency.

[0004] Therefore, there is an urgent need in this field for a technical solution that can understand the director's intentions and automatically generate precise and controllable structured editing instructions. Summary of the Invention

[0005] This invention provides a method and system for generating structured editing instructions based on a director's language model. It can deeply integrate script text, original materials and director's intentions, and automatically generate structured instructions that can directly drive professional editing software, thereby achieving intelligent and high-quality output from script to finished film.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for generating structured editing instructions based on a director's language model, comprising the following steps:

[0007] Script semantic parsing steps: Use a pre-trained large language model to parse the input script text and output structured data containing multiple semantic units. Each semantic unit contains at least paragraph type, text content and sentiment attribute.

[0008] Shot planning steps: Based on the semantic unit information in the structured data, by combining preset director rules and model reasoning, corresponding shot planning parameters are generated for each semantic unit;

[0009] Editing instruction structure generation steps: The shot planning parameters are fused with the corresponding script text to generate machine-readable structured editing instruction data;

[0010] Material scheduling steps: Based on the requirements of the structured editing instruction data, schedule local materials or create new materials through the generation model;

[0011] Editing execution steps: Convert the structured editing instruction data and materials into a video editing project file, and render and output the video.

[0012] Preferably, the shot planning parameters include at least one of shot type, shot duration, shooting angle, transition effect type, and visual effects.

[0013] Preferably, the director rules are a predefined set of rules that map emotional attributes or segment types to recommended shot parameters.

[0014] Preferably, in the material scheduling step, creating new materials through the generation model specifically includes: constructing descriptive text prompts according to the needs of the editing instructions, and calling an image generation or video generation model to synthesize visual materials.

[0015] Preferably, the method further includes a style transfer step: receiving style reference input, extracting its visual and rhythmic features, and using the features to adjust the parameter generation strategy in the shot planning step.

[0016] Preferably, it also includes a multimodal feedback learning step: receiving user feedback information on the output video and using the feedback information to fine-tune the model parameters used for shot planning.

[0017] Preferably, a system for implementing a structured editing instruction generation method based on a director's language model includes:

[0018] The script semantic parsing module is used to execute the script semantic parsing steps;

[0019] A lens planning module is used to execute the lens planning steps;

[0020] The editing instruction structure generation module is used to execute the editing instruction structure generation step;

[0021] The material scheduling interface module is used to execute the material scheduling steps.

[0022] The editing executor interface module is used to execute the editing execution steps.

[0023] Preferably, the shot planning module has a built-in director rule library, which is used to provide shot planning suggestions based on the emotional attributes of semantic units or the type of paragraphs.

[0024] The beneficial effects of this invention are as follows:

[0025] Deep semantic understanding: Through deep analysis of scripts using a large language model, abstract text is transformed into concrete, executable film and television language elements.

[0026] Structured and executable instructions: The generated editing instructions are highly structured machine-readable data that can seamlessly connect to and drive downstream editing tools, achieving true automated production.

[0027] Highly personalized and stylized: Breaking through the limitations of traditional templates, it can generate diverse editing schemes based on script content and style preferences, resulting in high creativity.

[0028] Significantly improved efficiency: Freeing workers from tedious shot design and timeline arrangement, greatly shortening the video production cycle and reducing labor costs.

[0029] It has the ability to continuously evolve: by introducing a feedback learning mechanism, the system can continuously optimize its output and become more adaptable. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a schematic diagram of the overall system structure and workflow of the present invention;

[0032] Figure 2 This is a detailed flowchart illustrating the workflow of the lens planning module of the present invention.

[0033] Figure 3 This is a schematic diagram of the system closed-loop optimization of the feedback learning mechanism of the present invention. Detailed Implementation

[0034] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Example 1: Basic System Implementation

[0036] like Figure 1 As shown, this embodiment provides a basic structured editing instruction generation system based on a director's language model. The system's workflow is as follows:

[0037] First, the user inputs a script text, such as an advertising copy, into the script semantic parsing module. This module then uses its built-in large language model to deconstruct the text. The model identifies sections such as "opening hooks," "emotional build-up," "product highlight presentations," and "calls to action," and parses out sentiment tags such as "tension" and "curiosity," while also marking specific statements mentioning the product. The parsing results are organized into a structured data list, where each list item represents a semantic unit and includes information such as its type, text content, sentiment, and product mention status.

[0038] Subsequently, as Figure 2 As shown, the structured semantic data is fed into the shot planning module. This module embeds a set of director rules based on film and television theory. For example, when the rule base recognizes the emotional tag of "tension," it triggers a rule suggesting the use of a "close-up" shot, a shorter shot duration, and a "slow-motion" transition effect to enhance the tense atmosphere. When "product mention" is recognized, it suggests using a "cut-in" shot with a superimposed "highlight frame" effect to highlight the product. The module integrates these rules with the model's contextual reasoning capabilities to output a set of optimal shot planning parameters for each semantic unit.

[0039] Next, the editing instruction structure generation module integrates the received shot planning parameters with the original script text. Following a predefined data structure, this module generates a complete, machine-readable, structured editing instruction. This instruction details the video's timeline structure: for example, the first scene begins at second 0, lasts 3.5 seconds, uses a close-up shot, and is accompanied by the caption "He walks into the office, looking nervous," with the caption style "bold white at the bottom." The scene begins with a fade-in transition. The entire instruction file is organized in a structured data format, such as JSON, clearly defining the visual and auditory elements of each frame.

[0040] Then, the material scheduling interface module analyzes this structured instruction. If a scene in the instruction is marked as "requires specific material", and the system does not find a match in the local material library, the module will automatically construct a detailed text description, such as: "a 3.5-second close-up shot, the person in the picture looks anxious, office background, dim lighting, hand holding a medicine bottle", and call an external image generation or video generation model to create the required visual material based on the description.

[0041] Finally, the editing executor interface module packages the final structured editing instructions and all generated materials, and transmits them to the backend editing executor via API. This executor can parse the instructions, accurately map them to project files in professional editing software such as MLT and Kdenlive, automatically arrange shots on the timeline, add subtitles and effects, and finally render and generate a finished video file that can be distributed.

[0042] Example 2: System Implementation with Creativity Enhancement Function

[0043] like Figure 3 As shown, this embodiment, based on embodiment one, introduces a style transfer and multimodal feedback learning mechanism, further enhancing the system's creativity and adaptability.

[0044] Regarding the style transfer function: Users can provide a style reference to the system, such as inputting text like "imitating the visual style of Wong Kar-wai's films," or directly uploading a reference video with a fast-paced editing style. The system's style encoder will extract key features from this reference, such as color preferences, motion rhythm, and transition characteristics. These features will be fed into the shot planning module, enabling it to consider not only script semantics but also style features when generating shot parameters. For example, for the same "tense" scene, under the "Wong Kar-wai style," it might tend to use "slow shutter speed" and "high saturation" effects, rather than conventional close-ups.

[0045] Regarding the multimodal feedback learning function: After generating a video and delivering it to the user, the system can receive textual feedback from the user, such as "the second transition is too abrupt" or "the product demonstration time is too short." The system converts this natural language feedback into machine-understandable feature vectors through a feedback encoder. This vector is then used to fine-tune the director's language model. For example, in response to feedback about "abrupt transitions," the model will learn to use smoother transitions in similar situations. Through this closed-loop learning mechanism, the system can continuously learn from user feedback, constantly optimizing its editing decisions, making the generated videos increasingly meet the user's personalized expectations, possessing an evolutionary capability that traditional systems lack.

[0046] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating structured editing instructions based on a director's language model, characterized in that, Includes the following steps: Script semantic parsing steps: Use a pre-trained large language model to parse the input script text and output structured data containing multiple semantic units. Each semantic unit contains at least paragraph type, text content and sentiment attribute. Shot planning steps: Based on the semantic unit information in the structured data, by combining preset director rules and model reasoning, corresponding shot planning parameters are generated for each semantic unit; Editing instruction structure generation steps: The shot planning parameters are fused with the corresponding script text to generate machine-readable structured editing instruction data; Material scheduling steps: Based on the requirements of the structured editing instruction data, schedule local materials or create new materials through the generation model; Editing execution steps: Convert the structured editing instruction data and materials into a video editing project file, and render and output the video.

2. The method according to claim 1, characterized in that, The shot planning parameters include at least one of the following: shot type, shot duration, shooting angle, transition effect type, and visual effects.

3. The method according to claim 1, characterized in that, The director rules are a predefined set of rules that map emotional attributes or segment types to recommended shot parameters.

4. The method according to claim 1, characterized in that, In the material scheduling step, creating new materials through the generation model specifically includes: constructing descriptive text prompts according to the needs of the editing instructions, and calling image generation or video generation models to synthesize visual materials.

5. The method according to claim 1, characterized in that, It also includes a style transfer step: receiving style reference input, extracting its visual and rhythmic features, and using the features to adjust the parameter generation strategy in the shot planning step.

6. The method according to claim 1, characterized in that, It also includes a multimodal feedback learning step: receiving user feedback on the output video and using this feedback to fine-tune the model parameters used for shot planning.

7. A structured editing instruction generation system based on a director's language model for implementing the method of any one of claims 1 to 6, characterized in that, include: The script semantic parsing module is used to execute the script semantic parsing steps; A lens planning module is used to execute the lens planning steps; The editing instruction structure generation module is used to execute the editing instruction structure generation step; The material scheduling interface module is used to execute the material scheduling steps. The editing executor interface module is used to execute the editing execution steps.

8. The system according to claim 7, characterized in that, The shot planning module has a built-in director rule library, which is used to provide shot planning suggestions based on the emotional attributes of semantic units or the type of paragraphs.