Method for automatic generation of structured shot parameters based on natural language keywords
Patent Information
- Application Number
- CN202610953953.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]有鉴于此,本发明提供了一种基于自然语言的结构化视频分镜生成与增量编辑方法,用以解决现有技术存在文本、图像、动画、音频多模态资产仅做简单时序拼接,缺少跨模态参数实时联动调度的问题
[0024]第一,通过以分镜数据为单位的结构化中间数据层的第一提示约束条件、第二提示约束条件、第三提示约束条件,使分镜视频片段、配音音频数据和字幕时间轴数据在生成阶段即基于同一结构化中间数据层实现跨模态参数协同,步骤S4的拼接操作是对已协同对齐的素材进行顺序组合,而非在拼接阶段进行强行时序对齐。
Smart Images

Figure CN122802753A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital media technology, and more specifically, to a method for automatically generating structured storyboard parameters based on natural language keywords. The patent title of this invention may also be "Method for Generating and Incrementally Editing Structured Video Storyboards Based on Natural Language". Background Technology
[0002] Existing technology (application date: 2025114122977) discloses an AI-based method for automatic generation of animation storyboards and visual pre-visualization. This method receives natural language script text, uses natural language processing to perform semantic structured parsing, and extracts key narrative elements such as scenes, characters, and shots, further enabling static storyboard generation, dynamic pre-visualization synthesis, and interactive editing iteration. However, this existing technology only performs simple temporal splicing of multimodal assets (text, images, animation, and audio), lacking real-time cross-modal parameter linkage and scheduling. Therefore, this is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0003] In view of this, the present invention provides a structured video storyboard generation and incremental editing method based on natural language, which solves the problem that existing technologies only perform simple temporal splicing of multimodal assets such as text, images, animation and audio, and lack real-time linkage scheduling of cross-modal parameters.
[0004] This application provides a method for generating and incrementally editing structured video storyboards based on natural language, characterized by the following steps:
[0005] The natural language keyword description data and the first prompt constraint are input into the video script agent to generate and output structured video script data; the structured video script data includes: a storyboard sequence, each storyboard data corresponding to a shot number, storyboard duration data, storyboard semantic description data, character action description data, scene description data, initial camera movement description parameters, character dialogue data and character identity identifier;
[0006] A structured intermediate data layer is constructed using the storyboard data as the basic unit. The structured video script data is stored in the structured intermediate data layer in units of the storyboard data, and a unique index is established for each storyboard data. The structured intermediate data layer is a tree-like or graph-like data structure, which supports local reading, modification, and write-back operations on any data in any storyboard data.
[0007] Based on the semantic description data of each storyboard, the character action description data, the scene description data, and the initial camera movement description parameters in the structured intermediate data layer, a second prompt constraint is constructed. This second prompt constraint is then input into the video generation agent to generate a storyboard video clip. Furthermore, based on the character dialogue data and character identity identifier of each storyboard, a third prompt constraint is constructed. This third prompt constraint is then input into the dubbing generation agent and the subtitle generation agent, respectively, to generate corresponding dubbing audio data and subtitle timeline data. Finally, the storyboard video clip, the dubbing audio data, and the subtitle timeline data are stored in association with each storyboard data in the corresponding storyboard data record within the structured intermediate data layer.
[0008] The storyboard video clips, the dubbing audio data, and the subtitle timeline data are spliced together in the order of the storyboard sequence to generate a complete preview video. A visual editing interface is provided, which allows users to visually modify any data of any storyboard data in the structured intermediate data layer, and allows users to add, delete, adjust the order, and adjust the splicing point position of the storyboard sequence.
[0009] Optionally, the first prompt constraint includes:
[0010] The sum of all the segment duration data is constrained to be equal to the total duration of the target video specified by the user, or by default equal to the total semantic duration implied by the natural language keyword description data; the segment duration data of each segment is between the preset shortest single-segment duration and the preset longest single-segment duration.
[0011] The semantic description data of the storyboard between two adjacent storyboards must satisfy temporal and spatial continuity in terms of action connection, scene transition and character position, and the scene description data of any adjacent storyboard must be accompanied by a scene transition type label when switching.
[0012] The language style of the character action description data, the character dialogue data, and the character spatial position in the scene description data of the same character identity identifier are kept semantically consistent in all the storyboard sequences to avoid character attribute conflicts.
[0013] The values of the initial camera movement description parameters are constrained to be within a preset physically executable range. The initial camera movement description parameters include lens movement type, lens movement speed, lens pitch angle, and lens rotation angle.
[0014] When any of the storyboard data contains the character dialogue data, the storyboard duration data of the storyboard is not less than the duration required for the character dialogue data to be read at a preset speech rate, and the start timestamp of the character dialogue data is after the start timestamp of the storyboard data and the end timestamp is before the end timestamp of the storyboard data.
[0015] Optionally, the second prompt constraint includes:
[0016] The video generation agent is constrained to use the role identity identifier and scene description data recorded in the structured intermediate data layer as the main control conditions, and to keep the main object, main clothing, and main color features in the generated storyboard video clips consistent with the corresponding records in all other storyboards;
[0017] For multiple consecutive scenes with the same scene description data in the scene sequence, the background environment of the scene video clip generated by the video generation agent is constrained to maintain a smooth inter-frame transition in terms of lighting direction, texture layout and depth parameters, and the structural similarity index of any two adjacent frames is not lower than a preset threshold.
[0018] The video generation agent is constrained to perform camera movements according to the initial camera movement description parameters recorded in the structured intermediate data layer, and the root mean square error between the camera movement trajectory of the generated video and the initial camera movement description parameters is less than a preset error threshold.
[0019] Optionally, the second prompt constraint further includes:
[0020] Input a list of negative prompt words into the video generation agent to prevent the generation of frames containing blurry, distorted, overlapping multiple subjects, abnormal subject organs, or torn backgrounds;
[0021] The video generation agent is constrained to output a sequence of video frames with a fixed resolution and a fixed frame rate.
[0022] Optionally, the third prompt constraint includes: character and timbre mapping constraint, speech rate and duration matching constraint, subtitle timeline synchronization constraint, and emotional rhythm consistency constraint.
[0023] Compared with existing technologies, the natural language-based structured video storyboard generation and incremental editing method provided by this invention achieves at least the following beneficial effects:
[0024] First, by using the first, second, and third cue constraints of the structured intermediate data layer based on storyboard data, the storyboard video clips, dubbing audio data, and subtitle timeline data achieve cross-modal parameter collaboration based on the same structured intermediate data layer during the generation stage. The splicing operation in step S4 is to sequentially combine the already coordinated and aligned materials, rather than forcibly aligning the timing during the splicing stage.
[0025] Second, the natural language keyword description data undergoes a four-layer progressive transformation, and the final output is a complete preview video with three modal alignment of visuals, voice-over, and subtitles. The entire process requires no manual intervention, achieving a complete automated transformation from natural language keyword description data to audiovisual content.
[0026] Third, the same structured intermediate data layer is read in parallel by two task lines (generating video clips corresponding to each segment data and generating audio data and subtitle timeline data corresponding to each segment data). The video clips, audio data and subtitle timeline data of the segment data are read and then spliced together. Data is obtained from the same structured intermediate data layer to ensure data version consistency and no copy conflicts.
[0027] Of course, any product implementing the present invention does not need to achieve all of the above additional technical effects while solving the background technical problem. It is sufficient for the product to solve the background technical problem first. The additional technical effects are effects that are beyond the understanding of those skilled in the art when the specific structure of the present invention is combined with a specific environment.
[0028] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0030] Figure 1 This is a flowchart illustrating the structured video storyboard generation and incremental editing method based on natural language provided by the present invention. Detailed Implementation
[0031] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0034] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0035] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0036] Figure 1 This is a flowchart illustrating the structured video storyboard generation and incremental editing method based on natural language provided by the present invention. This embodiment provides a structured video storyboard generation and incremental editing method based on natural language, including the following steps:
[0037] Step S1: Input the natural language keyword description data and the first prompt constraint into the video script agent to generate and output structured video script data; the structured video script data includes: a storyboard sequence, each storyboard data corresponding to a shot number, storyboard duration data, storyboard semantic description data, character action description data, scene description data, initial camera movement description parameters, character dialogue data and character identity identifier;
[0038] Specifically, before executing step S1, natural language keyword description data (such as text) is obtained, and the natural language keyword description data and the first prompt constraint are input into the video script agent to generate and output structured video script data.
[0039] Step S1 is the entry point for the entire solution from natural language keyword description data to structured data. It is responsible for the core function of transforming a sentence entered by the user (such as "I want to generate a video of a kitten cooking in the kitchen") into structured video script data that can be accurately parsed and processed programmatically by the processor.
[0040] Natural language refers to descriptive text input by users using everyday human language (rather than programming languages or specialized storyboard terminology). For example, a picture of an orange kitten cooking fish soup in a sunny kitchen is natural language, while shot 1 (medium shot, 3 seconds, cat character's action = chopping vegetables) is not. This lowers the barrier to entry for users.
[0041] The keyword description text emphasizes that the input can be short, incomplete narrative text fragments. Users are not required to provide a complete script format (like a traditional storyboard artist writing scene 1, interior, kitchen, day, etc.). Users only need to provide the core concepts and key elements, and the system will handle the completion and development. For example, entering the phrase "cat stealing fish" will trigger the complete generation process.
[0042] Video script intelligent agents refer to general-purpose language models pre-trained on large-scale text corpora, such as Generative Pre-trained Large Model 4, Claude, and Wenxin Yiyan. These large language models have the ability to understand text, reason, generate and format outputs, and can infer complete narrative structures, character relationships, emotional trends and cinematic language from short keywords.
[0043] In practice, natural language keyword descriptions and initial constraint conditions are transmitted together to the video script agent. This limits the output space of the video script agent, preventing it from generating unusable scripts due to unchecked behavior.
[0044] Optionally, the first hint constraint includes:
[0045] Duration consistency constraints, such as constraining the sum of all segment duration data to equal the total duration of the target video specified by the user, or by default equaling the total semantic duration implied by the natural language keyword description data; the segment duration data of each segment data is between the preset shortest single-shot duration and the preset longest single-shot duration;
[0046] Specifically, the total duration of all shots must be equal to the total duration specified by the user or the reasonable total duration inferred by the system (e.g., by default, 10 words of description are approximately equal to 3 seconds of video); the duration of each single shot should be between 0.5 seconds and 15 seconds (the lower limit should avoid switching as soon as the image appears, and the upper limit should avoid viewer visual fatigue); the duration of dialogue shots should not be shorter than the time required for the lines to be read at a normal speaking speed.
[0047] Consistency constraints on duration are used to prevent the total duration of video script agents from getting out of control (e.g., generating a 2-hour long film from the input of a cat cooking), or the duration of a single shot from being extremely unreasonable (e.g., a 0.1-second flash or a 60-second still shot).
[0048] The logic of storyboard continuity is constrained, such as the constraint that the semantic description data of the storyboard between two adjacent storyboard data must satisfy temporal continuity and spatial continuity in terms of action connection, scene transition and character position, and the scene transition type label must be attached when the scene description data of any adjacent storyboard data is switched.
[0049] By constraining the continuity of the scene sequence logic, narrative breaks between shots generated by the video script agent are prevented, ensuring that the entire video is a coherent story flow rather than a disordered splicing of fragments.
[0050] Role consistency constraints, such as constraining the same role identity to maintain semantic consistency in the language style of the character action description data, the character dialogue data, and the character spatial location in the scene description data across all parting sequences, so as not to generate character attribute conflicts.
[0051] By constraining the consistency of roles, we prevent the video script agent from changing its behavior in different storyboard data, which could lead to a sudden change in the appearance / clothing of the characters generated by the video generation agent in different storyboard data, thus destroying the character's identity.
[0052] Legality constraints on camera movement parameters, such as limiting the values of initial camera movement description parameters to a preset physically executable range. Initial camera movement description parameters include camera movement type, camera movement speed, camera pitch angle, and camera rotation angle.
[0053] By constraining the legality of camera movement parameters, we prevent the video script agent from generating physically infeasible camera movement parameters, which would otherwise prevent the video generation agent from executing or generate distorted images.
[0054] The dialogue and action timing matching constraint stipulates that when any storyboard data contains character dialogue data, the storyboard duration of the storyboard data is not less than the duration required for the character dialogue data to be read at a preset speech rate, and the start timestamp of the character dialogue data is after the start timestamp of the storyboard data and the end timestamp is before the end timestamp of the storyboard data.
[0055] By using dialogue and action timing matching constraints, we prevent subtitles generated by the video script agent from going out of bounds on the timeline (e.g., assigning a 10-word line of dialogue to a 2-second shot to ensure audio-visual synchronization).
[0056] Step S2: Construct a structured intermediate data layer using the storyboard data as the basic unit. Store the structured video script data into the structured intermediate data layer in units of storyboard data, and establish a unique index identifier for each storyboard data. The structured intermediate data layer is a tree-like or graph-like data structure, which supports local reading, modification, and write-back operations on any data in any storyboard data.
[0057] Specifically, step S2 is the data hub construction stage of the entire solution. It takes the structured video script data generated in S1 and transforms it from the structured video script data output by the video script agent into a persistent, addressable, and evolvable structured intermediate data layer. It solves the core pain points of the traditional artificial intelligence video generation process, such as the inaccessibility, inaccessibility, and lack of backtracking of intermediate states, and provides a unified data foundation for subsequent steps S3 and S4.
[0058] The construction of a structured intermediate data layer based on storyboard data is not simply about saving the structured video script data output from step S1 as a file. Instead, it establishes an operable data container with query, indexing, and transaction capabilities, providing a unified data access interface for all subsequent modules. Using storyboard data as the basic unit means that each storyboard data is an independent data record that can be individually addressed and manipulated. This granularity is the technical prerequisite for all subsequent local operations (such as modifying a storyboard data in step S4). Structured data is the opposite of unstructured data (such as plain text or semi-structured data). Structured data means that the system can perform type checking at compile time or runtime to ensure data integrity. The structured intermediate data layer is not a temporary variable or cache, but the sole source of data read and written by all modules in the entire generation process. The video generation agent reads data from here, and the dubbing generation agent reads data from here, rather than directly calling each other between the video generation agent and the dubbing generation agent.
[0059] The complete storyboard sequence output in step S1 is broken down into individual records for storage. For example, if step S1 outputs 10 storyboard data, then 10 storyboard data records are created in the structured intermediate data layer. Each storyboard data record is stored, indexed, and managed independently, rather than packaging the 10 storyboard data records into a single data set.
[0060] The structured intermediate data layer not only stores the structured script data generated in step S1, but also needs to pre-create empty slots for data to be filled in the subsequent stage (step S3).
[0061] A unique index identifier means that each segment of data has a globally unique, never-repeating primary key in the structured intermediate data layer.
[0062] A structured intermediate data layer, characterized by parsable and editable tree or graph-like data structures, means that the data stored in this layer has a defined, self-describing format. Any module can read and interpret the data content using standard parsers (such as JSON parsers or database query engines) without relying on external context or heuristics. For example, the storyboard semantic description data is explicitly a string representing the narrative content of that shot, and the dubbing module will not mistakenly interpret it as dubbing text.
[0063] The data in the structured intermediate data layer is not read-only; it supports four basic operations: add, delete, modify, and query. Any module (such as the video generation agent in the visual editing interface) can send modification requests to the structured intermediate data layer, which is responsible for ensuring the atomicity of modifications and data consistency.
[0064] Tree-based or graph-based data structures are two available data organization methods, each corresponding to different project complexities.
[0065] Building a structured intermediate data layer with storyboard data as the basic unit means creating a new data storage structure in system memory or persistent storage (such as databases or file systems); the smallest granularity of data organization is a storyboard data, rather than a video or a scene.
[0066] Partial reading: Achieved through projection queries using Structured Query Language or document databases; for tree structures, precise location is achieved using XPath or tree structures of key-value pairs and arrays (such as JSON Path).
[0067] Local modifications: achieved through atomic data update operations; for document-oriented databases, data-level update operators are used; for tree structures, node attributes are modified after locating the node via a path.
[0068] Write-back operation: Transactional write, ensuring that no other operations cause dirty reads or dirty writes during the reading, modification and writing process; after the write-back is completed, a broadcast notification is sent to all modules that have subscribed to the data.
[0069] Instead of simply saving the structured video script data output from step S1 as a file, we establish an operable data container with query, indexing, and transaction capabilities to provide a unified data access interface for all subsequent modules.
[0070] Step S3: Construct a second prompt constraint based on the storyboard semantic description data, character action description data, scene description data, and initial camera movement description parameters of each storyboard data in the structured intermediate data layer; input the second prompt constraint to the video generation agent to generate a storyboard video clip; and construct a third prompt constraint based on the character dialogue data and character identity identifier of each storyboard data; input the third prompt constraint to the dubbing generation agent and the subtitle generation agent respectively to generate dubbing audio data and subtitle timeline data respectively; and associate and store the storyboard video clip, the dubbing audio data, and the subtitle timeline data in the corresponding storyboard data record in the structured intermediate data layer, with each storyboard data as a unit.
[0071] Specifically, step S3 is the multimodal data generation stage of the entire scheme. It follows the structured intermediate data layer built in S2, reads the structured description data of each storyboard data from it, and schedules the video generation agent, dubbing generation agent and subtitle generation agent respectively to convert the text-based structured video script data into playable audiovisual materials (storyboard video clips, dubbing audio data and subtitle timeline data). The generated results are then backfilled into the structured intermediate data layer in units of storyboard data, providing complete material preparation for the splicing and synthesis in S4.
[0072] Step S3 contains two parallel task lines that are independent of each other and can be executed concurrently. Specifically, the first task line generates storyboard semantic description data, character action description data, scene description data, and initial camera movement description parameters. It calls the video generation agent to output the storyboard video clip corresponding to each storyboard data. The second task line generates character dialogue data, character identity identifiers, and storyboard duration data (implicitly used for duration verification). It calls the dubbing generation agent and the subtitle generation agent to generate dubbing audio data and subtitle timeline data corresponding to each storyboard data.
[0073] The target of the read operation is the structured intermediate data layer, not the raw input text or other external data sources. It iterates through all the scene records in the structured intermediate data layer, independently constructing second cue constraints for each scene data.
[0074] The storyboard semantic description data, character action description data, scene description data, and initial camera movement description parameters are respectively derived from the storyboard semantic description data, character action description data, scene description data, and initial camera movement description parameters recorded in the structured intermediate data layer, describing the core narrative content to be presented in this shot;
[0075] The aforementioned storyboard semantic description data, character action description data, scene description data, and initial camera movement description parameter values are assembled according to a preset template and preset rules into a set of input prompts with a second cue constraint specifically for the video generation agent. This second cue constraint is not simply a matter of concatenating text; rather, it includes mandatory screen control instructions to constrain the output space of the video generation agent.
[0076] Optionally, the second hint constraint includes:
[0077] Consistency constraints on the main subject of the image, such as constraining the video generation agent to use the character identity identifiers and scene description data recorded in the structured intermediate data layer as the main control conditions, to keep the main object, main clothing, and main color features in the generated storyboard video clips consistent with the corresponding records in all other storyboards;
[0078] The video generation agent is constrained to maintain consistency in the main subject, clothing, and color features of the generated storyboard video clips with the corresponding records in all other storyboard data.
[0079] Background scene continuity constraints, such as for multiple consecutive scenes with the same scene description data in a storyboard sequence, constrain the background environment of the storyboard video clips generated by the video generation agent to maintain a smooth transition between frames in terms of lighting direction, texture layout and depth parameters, and the structural similarity index of any two adjacent frames is not lower than a preset threshold.
[0080] By constraining the continuity of background scenes, the background environment of the video clips generated by the video generation agent is constrained to maintain a smooth transition between frames in terms of lighting direction, texture layout, and depth parameters for multiple consecutive segment data with the same scene description data in the segment sequence.
[0081] The camera movement parameters are precisely executed under constraints, such as constraining the video generation agent to perform camera movements according to the initial camera movement description parameters recorded in the structured intermediate data layer, and constraining the root mean square error between the generated video's camera movement trajectory and the initial camera movement description parameters to be less than a preset error threshold.
[0082] The constraint video generation agent precisely executes camera movement parameters to perform camera movements according to the initial camera movement description parameters recorded in the structured intermediate data layer. The root mean square error between the camera movement trajectory of the generated video and the initial camera movement description parameters is less than a preset error threshold. The preset error threshold can be designed according to actual conditions, and this embodiment does not impose specific limitations on it.
[0083] The second suggestion constraint also includes:
[0084] Content constraints are prohibited, such as inputting a list of negative prompt words into the video generation agent, and prohibiting the generation of frames containing blurry, distorted, overlapping multiple subjects, abnormal subject organs, or torn backgrounds.
[0085] By prohibiting the input of a list of negative prompt words to the video generation agent through content constraints, the generation of frames containing blurry, distorted, overlapping multiple subjects, abnormal subject organs, or torn backgrounds is prohibited.
[0086] Resolution and frame rate constraints, such as constraining the video generation agent to output a fixed resolution. A sequence of frames with a fixed frame rate (fps).
[0087] The video generation agent is constrained to output a fixed resolution by limiting resolution and frame rate. A sequence of frames with a fixed frame rate (fps), where W is the width in pixels, H is the height in pixels, fps is the number of frames per second, and the segment duration data is also included. satisfy: ,in This represents the total number of frames actually generated for this storyboard. are positive integers and The absolute deviation from the storyboard duration data set in step S1 shall not exceed the preset time tolerance. .
[0088] Simultaneously, based on the character dialogue data and character identity identifiers for each storyboard, a third cue constraint is constructed. This third cue constraint serves to generate the dubbing audio data and subtitle timeline data corresponding to each storyboard, but different constraints within the third cue constraint apply to the dubbing generation agent and the subtitle generation agent respectively:
[0089] The third set of constraints mentioned above includes:
[0090] The character and timbre mapping constraint, such as the constraint that the dubbing generation agent selects the target timbre parameter uniquely bound to the character identity from the preset timbre library based on the character identity recorded in the structured intermediate data layer for dubbing synthesis, and the cosine similarity of the timbre feature vector of the dubbing audio data output by the same character identity in all storyboards is not lower than the preset timbre similarity threshold.
[0091] The voice-over generation agent, constrained by the mapping between roles and timbres, selects target timbre parameters uniquely bound to the role's identity from a preset timbre library for voice-over synthesis based on the role's identity identifier recorded in the structured intermediate data layer.
[0092] Speech rate and duration matching constraints, such as constraining the actual duration of the dubbing audio data output by the dubbing generation agent. satisfy ,in This refers to the storyboard duration data corresponding to the scene, where Δ represents the preset silence buffer duration before and after; if the character dialogue data is calculated based on the preset baseline speech rate, the baseline duration is... Exceed This triggers the speech rate adaptive adjustment factor α, making And α is used as a constraint condition input into the subtitle generation agent, controlling the speech rate adjustment range to be... ;
[0093] The actual duration of the dubbing audio data output by the dubbing agent is constrained by matching speech rate and duration. satisfy If the dialogue text is calculated at the baseline speaking speed... Exceed This triggers the speech rate adaptive adjustment factor α.
[0094] Subtitle timeline synchronization constraints, such as the constraint that the start and end timestamps of each subtitle in the subtitle timeline data generated by the subtitle generation agent must fall within the range of the start and end timestamps of the scene, and the presentation duration of each subtitle is not less than the preset minimum subtitle dwell time. The subtitle timeline data is segmented according to the semantic sentence break position of the character dialogue data, and the number of characters in each subtitle does not exceed the preset maximum number of subtitle characters.
[0095] The subtitle timeline synchronization constraint ensures that the start and end timestamps of each subtitle in the subtitle timeline data generated by the subtitle generation agent strictly fall within the range of the start and end timestamps of the scene, and is segmented according to the semantic sentence break position.
[0096] Emotional prosody consistency constraints, such as constraining the voice-over generation agent to select corresponding emotional prosody parameters based on the emotional tags in the storyboard semantic description data, include the fundamental frequency mean, fundamental frequency variation range, and speech rate variation curve, so that the emotional expression of the voice-over audio data matches the emotional tendency of the storyboard semantic description data.
[0097] The voice-over agent is constrained by the emotional prosody consistency constraint and selects the corresponding emotional prosody parameters (mean fundamental frequency, range of fundamental frequency variation, and speech rate variation curve) based on the emotional tags in the storyboard semantic description data.
[0098] Step S4: Combine the storyboard video clips, dubbing audio data, and subtitle timeline data according to the sequence of the storyboard scenes to generate a complete preview video. A visual editing interface is provided, which allows users to visually modify any data in any storyboard data in the structured intermediate data layer. It also allows users to add, delete, and adjust the order of the storyboard sequence and the position of the splicing points.
[0099] Specifically, step S4 is the entry point for the synthesis preview and editing of the entire solution. It follows all the multimodal materials generated and backfilled into the structured intermediate data layer in step S3, and assembles the discrete storyboard video clips, dubbing audio data and subtitle timeline data into a complete and playable preview video for users to watch according to the storyboard sequence. At the same time, it provides a visual editing interface, allowing users to modify any data in the structured intermediate data layer, thereby entering the iterative optimization process in the human loop.
[0100] In step S4, the scope of the splicing operation for all storyboard data is the full set of storyboard records stored in the structured intermediate data layer, that is, all storyboards from the first storyboard to the Nth storyboard participate in the splicing.
[0101] The order of the storyboard sequence is maintained by a global index (such as...) maintained by the structured intermediate data layer. The decision is made. During splicing, starting from the first position of the index, the second, third, and so on, are taken sequentially until the Nth position, and the corresponding video clips for each position are connected end to end in order.
[0102] The splicing and compositing process comprises three sub-operations: video track splicing, which connects the video clips of each scene sequentially end-to-end to merge them into a continuous video stream; audio track splicing, which overlays the dubbing audio data of each scene onto the audio track of the video stream at the same timeline position, with the audio of each scene aligned chronologically; and subtitle track overlay, which overlays each subtitle in the subtitle timeline data of each scene onto the corresponding time position of the video frame based on its start / end timestamp (converted from the relative time offset of each scene to a global absolute time offset). The final output is a playable, complete video file containing a continuous video stream, a synchronized audio track, and an overlaid subtitle track. This video serves as a complete preview video that the user can play and view in the editing interface.
[0103] An additional visual editing interface is provided, and the two are provided simultaneously—the video preview playback window and the editing interface are integrated in the same user interaction view, allowing users to view and modify the corresponding structured data while watching the video;
[0104] The system outputs visual interactive components to the front-end user interface layer, including but not limited to: video player window, data editing panel (text box, drop-down selector, numerical slider, etc.), storyboard sequence timeline (displaying all storyboard data in the form of thumbnails / progress bars), and operation buttons (save changes, trigger regeneration, export video, etc.).
[0105] A graphical user interface allows users to modify content through intuitive operations such as clicking, dragging, and filling out forms, without writing code or editing raw structured data files. The backend of this interface is synchronized bidirectionally with the structured intermediate data layer in real time—all changes on the interface are instantly reflected in the structured intermediate data layer, and any updates to the structured intermediate data layer are also instantly refreshed on the interface.
[0106] User actions in the timeline or storyboard list view: Add: Insert a new blank storyboard data at any position in the sequence (at this time, all data in the new storyboard data is empty or default value, which needs to be filled in by the user or generated by the video script intelligent agent in step S1); Delete: Remove one or more storyboard data from the sequence, and at the same time, the entire record of the corresponding storyboard data in the structured intermediate data layer is marked as deleted or physically removed, and the deleted storyboard data will no longer be included in subsequent splicing.
[0107] Users can change the order of storyboard data on the timeline by dragging storyboard thumbnails or clicking the move up / down buttons. The essence of this order adjustment operation is updating the global sequence index maintained in the structured intermediate data layer (e.g., moving...). Adjusted to Each shot records its own unique index identifier and shot number data, which can be updated synchronously to maintain consistency of the user's perspective, or left unchanged while only the sequence order index is adjusted.
[0108] Users can make fine adjustments at the splicing point between two adjacent storyboards (i.e., between the end frame of the previous storyboard and the start frame of the next storyboard), including: splicing point offset: moving the splicing point forward or backward to advance or delay the start position of the next storyboard (creating a dissolve effect, i.e., the image of the previous storyboard gradually fades out, and the image of the next storyboard gradually fades in); adding transition effects: inserting or changing the transition type (such as cut, fade in / out, wipe, erase, zoom, etc.) between two storyboards to improve the smoothness of the visual transition. The essence of adjusting the splicing point position is to modify the transition parameters in the dependency data of adjacent storyboards, or to add transition commands in the global splicing configuration of the structured intermediate data layer.
[0109] Steps S1 to S4, through a four-layer progressive collaborative workflow of data format transformation, realize a complete transformation chain from user-input natural language keyword description data to structured storyboard records, then to multimodal video materials, and finally to playable preview videos: Step S1 applies a first prompt constraint, transforming the natural language keyword description data into parsable structured script data, with storyboard duration data, storyboard semantic description data, character action description data, scene description data, initial camera movement description parameters, character dialogue data, and character identity identifiers, in units of storyboard data; Step S2 follows the output of Step S1, using the structured script data of each storyboard data generated in Step S1 as a basis, establishing a unique index identifier for each storyboard data, and storing all data in a tree or graph data structure, so that each storyboard data is transformed from the structured video script data output by the video script agent into persistent data records that can be continuously read and written by subsequent modules; Step S3 reads each storyboard data from the structured intermediate data layer constructed in Step S2. The structured description data of the storyboard data (storyboard semantic description data, character action description data, scene description data, initial camera movement description parameters, character dialogue data, and character identity identifier) are used to construct second and third prompt constraints and schedule video generation agents, dubbing generation agents, and subtitle generation agents to generate corresponding storyboard video clips, dubbing audio data, and subtitle timeline data. Then, using the unique index identifier established in step S2 as the key, the data is written back to the corresponding storyboard data record in the structured intermediate data layer, so that the data record of each storyboard data evolves from structured video script data into a complete record that is associated with storyboard video clips, dubbing audio data, and subtitle timeline data. In step S4, all records are read from the structured intermediate data layer constructed in step S2 and, guided by the storyboard sequence order maintained in step S2, the storyboard video clips, dubbing audio data, and subtitle timeline data of each storyboard data are spliced and synthesized to generate a complete preview video that can be played, so that the data form completes the final leap from structured record to playable complete preview video.
[0110] Through the collaboration of steps S1 to S4, the data format achieves a four-layer progressive leap from the input natural language keyword description data to the output complete preview video. Throughout the process, the data always takes the structured intermediate data layer constructed in step S2 as the sole carrier. Each step adds a new data dimension on this carrier, and the complete transformation from text to audiovisual content can be completed without any human intervention.
[0111] The structured intermediate data layer constructed in step S2 decouples the data requirements between each step, allowing the two task lines in step S3 (image generation, dubbing and subtitle generation) to call their respective required data in parallel. Step S4 can independently read the complete multimodal data of all the parting data from the structured intermediate data layer for splicing. All downstream steps do not need to care about where the upstream data comes from or how it is generated.
[0112] Compared with existing technologies, the natural language-based structured video storyboard generation and incremental editing method provided in this embodiment achieves at least the following beneficial effects:
[0113] First, steps S1 to S4 use the first, second, and third prompt constraints of the structured intermediate data layer based on storyboard data to enable cross-modal parameter collaboration of storyboard video clips, dubbing audio data, and subtitle timeline data during the generation stage, based on the same structured intermediate data layer. The splicing operation in step S4 is to sequentially combine the already coordinated and aligned materials, rather than forcibly aligning the timing during the splicing stage.
[0114] Second, the natural language keyword description data undergoes a four-layer progressive transformation, and the final output is a complete preview video with three modal alignment of visuals, voice-over, and subtitles. The entire process requires no manual intervention, achieving a complete automated transformation from natural language keyword description data to audiovisual content.
[0115] Third, the same structured intermediate data layer is read in parallel by the two task lines of step S3 (generating the video clips of each segment data, generating the dubbing audio data and subtitle timeline data corresponding to each segment data), and the segment video clips, dubbing audio data and subtitle timeline data of all segment data are read and spliced together by step S4. Step S4 obtains data from the same structured intermediate data layer to ensure that the data version is consistent and there are no duplicate conflicts.
[0116] Fourth, the structured intermediate data layer evolves from an initial state where plain text data is filled and multimodal data is empty, to a complete state where all multimodal materials are filled. Step S4 ultimately consumes the complete state as a complete preview video; each stage can be interrupted, resumed, and verified.
[0117] Fifth, the first constraint ensures the rationality of the text layer of the structured video script data (scene length data, character dialogue data, character action description data); the second constraint ensures the splicability of the visual layer of the images; and the third constraint ensures the auditory synergy of the dubbing / subtitles. The constraints among the first, second, and third constraints progress from coarse to fine, layer by layer, jointly ensuring that the complete preview video synthesized in step S4 meets usable standards in terms of text structure, visual presentation, and audio expression.
[0118] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalent elements of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0119] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating and incrementally editing structured video storyboards based on natural language, characterized in that, Includes the following steps: Input the natural language keyword description data and the first prompt constraint into the video script agent to generate and output structured video script data; The structured video script data includes: a storyboard sequence, each storyboard data corresponding to a shot number, storyboard duration data, storyboard semantic description data, character action description data, scene description data, initial camera movement description parameters, character dialogue data, and character identity identifier; A structured intermediate data layer is constructed using the storyboard data as the basic unit. The structured video script data is stored in the structured intermediate data layer in units of the storyboard data, and a unique index is established for each storyboard data. The structured intermediate data layer is a tree-like or graph-like data structure, which supports local reading, modification, and write-back operations on any data in any storyboard data. Based on the semantic description data of each storyboard data in the structured intermediate data layer, the character action description data, the scene description data, and the initial camera movement description parameters, a second prompting constraint is constructed. The second prompting constraint is then input into the video generation agent to generate storyboard video clips. A third prompt constraint is constructed based on the character dialogue data and character identity identifier of each storyboard data. The third prompt constraint is then input into the dubbing generation agent and the subtitle generation agent to generate corresponding dubbing audio data and subtitle timeline data. The storyboard video clip, the dubbing audio data, and the subtitle timeline data are then associated and stored in the corresponding storyboard data record in the structured intermediate data layer, with each storyboard data as a unit. The storyboard video clips, the dubbing audio data, and the subtitle timeline data are spliced together in the order of the storyboard sequence to generate a complete preview video. A visual editing interface is provided, which allows users to visually modify any data of any storyboard data in the structured intermediate data layer, and allows users to add, delete, adjust the order, and adjust the splicing point position of the storyboard sequence.
2. The method for generating and incrementally editing structured video storyboards based on natural language according to claim 1, characterized in that, The first prompt constraint includes: The sum of all the segment duration data is constrained to be equal to the total duration of the target video specified by the user, or by default equal to the total semantic duration implied by the natural language keyword description data; the segment duration data of each segment is between the preset shortest single-segment duration and the preset longest single-segment duration. The semantic description data of the storyboard between two adjacent storyboards must satisfy temporal and spatial continuity in terms of action connection, scene transition and character position, and the scene description data of any adjacent storyboard must be accompanied by a scene transition type label when switching. The language style of the character action description data, the character dialogue data, and the character spatial position in the scene description data of the same character identity identifier are kept semantically consistent in all the storyboard sequences to avoid character attribute conflicts. The values of the initial camera movement description parameters are constrained to be within a preset physically executable range. The initial camera movement description parameters include lens movement type, lens movement speed, lens pitch angle, and lens rotation angle. When any of the storyboard data contains the character dialogue data, the storyboard duration data of the storyboard is not less than the duration required for the character dialogue data to be read at a preset speech rate, and the start timestamp of the character dialogue data is after the start timestamp of the storyboard data and the end timestamp is before the end timestamp of the storyboard data.
3. The method for generating and incrementally editing structured video storyboards based on natural language according to claim 1, characterized in that, The second set of constraints includes: The video generation agent is constrained to use the role identity identifier and scene description data recorded in the structured intermediate data layer as the main control conditions, and to keep the main object, main clothing, and main color features in the generated storyboard video clips consistent with the corresponding records in all other storyboards; For multiple consecutive scenes with the same scene description data in the scene sequence, the background environment of the scene video clip generated by the video generation agent is constrained to maintain a smooth inter-frame transition in terms of lighting direction, texture layout and depth parameters, and the structural similarity index of any two adjacent frames is not lower than a preset threshold. The video generation agent is constrained to perform camera movements according to the initial camera movement description parameters recorded in the structured intermediate data layer, and the root mean square error between the camera movement trajectory of the generated video and the initial camera movement description parameters is less than a preset error threshold.
4. The method for generating and incrementally editing structured video storyboards based on natural language according to claim 3, characterized in that, The second hint constraint also includes: Input a list of negative prompt words into the video generation agent to prevent the generation of frames containing blurry, distorted, overlapping multiple subjects, abnormal subject organs, or torn backgrounds; The video generation agent is constrained to output a sequence of video frames with a fixed resolution and a fixed frame rate.
5. The method for generating and incrementally editing structured video storyboards based on natural language according to claim 1, characterized in that, The third set of constraints includes: character and timbre mapping constraints, speech rate and duration matching constraints, subtitle timeline synchronization constraints, and emotional rhythm consistency constraints.