Dynamic split script automatic generation method and system based on diffusion model
By using a diffusion model-based method for automatically generating dynamic storyboards, the problem of limited diversity and creativity in the generated results in existing technologies has been solved. This method enables efficient and intelligent conversion of scripts into dynamic storyboards, thereby improving the efficiency and quality of film and television production.
Patent Information
- Application Number
- CN202511104894.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-12-16
AI Technical Summary
Existing technologies for the automated generation of dynamic storyboard scripts are unable to flexibly cope with complex and ever-changing narrative needs. The diversity and creativity of the generated results are limited, and they are highly dependent on keyframe settings. The degree of automation is limited, making it difficult to meet the high efficiency and accuracy requirements of professional-grade dynamic storyboard production.
A dynamic storyboard generation method based on a diffusion model is adopted. Key information of the script text is extracted through a pre-trained natural language processing model, a spatiotemporal relationship framework is constructed, and a dynamic storyboard sequence that conforms to the visual narrative logic is generated iteratively. The storyboard rhythm and screen composition are optimized through an adaptive feedback mechanism, and the final output storyboard meets the standards of the film and television industry.
It achieves efficient and intelligent conversion of scripts into dynamic storyboards, significantly improving film and television production efficiency, reducing the need for manual intervention, and generating storyboards with high sequential fluency, artistic expressiveness, and industrial compatibility, supporting rapid iteration and version management across multiple scenarios.
Smart Images

Figure CN121145804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of dynamic storyboard technology, specifically to a method and system for automatically generating dynamic storyboards based on a diffusion model. Background Technology
[0002] Dynamic storyboards are a crucial step in transforming static storyboard images into dynamic clips. By adding animation effects, voice-over, and music, the storyboards come alive, providing a visual reference for subsequent production. Dynamic storyboards are animated storyboards that connect static storyboards and add dynamic elements (such as camera zoom, rotation, and movement). They help directors control the pacing, assess workload, and avoid rework in post-production. They include shot numbers, scene content, camera movement techniques (push, pull, pan, tilt, etc.), animation effects (such as zoom and rotation), voice-over, music, and duration.
[0003] Existing methods for automatically generating dynamic storyboards typically employ the following technical approaches: First, there's the automated generation method based on preset templates and rule libraries. This method pre-builds standardized libraries of camera movements, transitions, and sound effects, then matches corresponding dynamic templates with input static storyboard elements (such as shot numbers and content descriptions). While relatively simple to implement, the diversity and creativity of the generated results are limited by the size and richness of the preset library, making it difficult to flexibly address complex and ever-changing narrative needs. Second, there's the semi-automated method based on keyframe animation technology. This method requires manual setting of key shot states and parameters (such as start and end positions, angles, and time points), and the system automatically generates smooth animation effects by interpolating intermediate frames based on the keyframes. This method allows for finer control, but it's highly dependent on keyframe settings, has limited automation, and requires users to have some animation production knowledge. Third, there's the method combining traditional computer vision and simple machine learning algorithms (such as motion trajectory prediction), attempting to infer possible dynamic elements from static images or text descriptions. However, such methods still have significant technical bottlenecks in handling complex camera movements (such as camera movements with emotional expression) or generating high-quality synchronized audio-visual effects, making it difficult to meet the high efficiency and accuracy requirements of professional-grade dynamic storyboard production. To address this, we propose an automated dynamic storyboard script generation method and system based on a diffusion model. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for automatically generating dynamic storyboard scripts based on a diffusion model, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention specifically adopts the following technical solution:
[0006] A method for automatically generating dynamic storyboard scripts based on a diffusion model includes the following steps:
[0007] Step 1: Input the script text data to be processed, and use a pre-trained natural language processing model to perform semantic analysis and key information extraction on the script text to identify scene descriptions, character actions, dialogue content and plot turning points.
[0008] Step 2: Based on the key information extracted in Step 1, construct the initial spatiotemporal relationship framework of the storyboard script, and determine the scene spatial layout, character positions and basic camera movement direction corresponding to each storyboard.
[0009] Step 3: Using a pre-trained diffusion model, based on the spatiotemporal relationship framework and the needs of the script plot development, iteratively generate a dynamic sequence of shots that conforms to the visual narrative logic. The iterative generation process includes noise addition and noise reduction optimization mechanisms.
[0010] Step 4: Perform temporal coherence verification and visual aesthetic scoring on the dynamic shot sequence generated in Step 3, and adjust the diffusion model parameters through an adaptive feedback mechanism to optimize the shot rhythm and screen composition.
[0011] Step 5: Output a complete dynamic storyboard script that includes camera angles, durations, camera movement paths, and visual effects annotations.
[0012] Furthermore, the key information extraction in step 1 includes: labeling scene elements using named entity recognition technology, constructing a character interaction relationship graph using dependency parsing, and quantifying the emotional intensity of dialogue using a sentiment analysis model.
[0013] Furthermore, the diffusion model in step 3 adopts a hierarchical condition control mechanism, and its input conditions include: scene topology constraint matrix, character motion trajectory probability field, and camera language rule template library.
[0014] Furthermore, the adaptive feedback mechanism in step 4 includes three-dimensional modules: a motion coherence detection module based on optical flow, a visual balance evaluation module based on compositional design principles, and a focus trajectory analysis module based on audience attention models.
[0015] Furthermore, the output format of step 5 is compatible with film and television industry standards, including but not limited to: generating FinalDraft parsable XML metadata, automatically binding Unreal Engine camera track data, and exporting DaVinciResolve compatible EDL editing decision tables.
[0016] Furthermore, it also includes step 6: establishing a storyboard version iteration knowledge base, generating a difference heatmap by comparing historical versions, and optimizing subsequent generation strategies based on reinforcement learning mechanisms.
[0017] An automated dynamic storyboard generation system based on a diffusion model includes:
[0018] The script input interface is used to receive script text input by the user and convert it into a structured data format;
[0019] The semantic analysis unit is used to parse script scene elements, character behaviors, and time markers to generate a spatiotemporal semantic graph.
[0020] The diffusion model-driven unit is used to iteratively generate a sequence of scenes based on the semantic map, optimize shot transitions through the motion coherence detection module, adjust composition parameters through the visual balance evaluation module, and dynamically allocate visual weights through the focus trajectory analysis module.
[0021] The spatiotemporal framework building unit is used to integrate the storyboard sequence and build a framework structure that includes a timeline, spatial coordinates, and camera motion trajectory.
[0022] The diffusion model generation unit is used to execute the iterative generation algorithm of the diffusion model, including noise addition and denoising optimization mechanisms, to generate a dynamic sequence of shots that conforms to the visual narrative logic.
[0023] The adaptive feedback unit is used to integrate the 3D module. It detects motion coherence through optical flow, evaluates visual balance by defining the composition, and allocates attention weights by analyzing focus trajectory, thereby dynamically adjusting the rhythm and composition of the storyboard.
[0024] The script output unit is used to implement the output function, converting the optimized storyboard into a format compatible with film and television industry standards.
[0025] The verification unit is used to establish a knowledge base for storyboard version iterations. It generates a difference heatmap by comparing historical versions and optimizes the subsequent generation strategy based on the reinforcement learning mechanism. This includes storing historical storyboard script data, calculating the dynamic parameter differences between versions, generating a visual heatmap to identify optimization areas, and applying reinforcement learning algorithms to adjust the diffusion model parameters to improve the quality of subsequent storyboard generation.
[0026] The user interaction interface is used to receive user feedback data and integrate it into the adaptive feedback unit, including scene adjustment suggestions, aesthetic preference input, and rhythm control instructions, to achieve human-computer collaborative optimization.
[0027] The script output interface is used to implement the function of the script output unit and transmit the storyboard script data to the external film and television production system or user terminal through a standardized interface to ensure seamless integration into the actual production process.
[0028] The storyboard version iteration knowledge base unit is used to establish and manage the storyboard version iteration knowledge base. It generates a difference heatmap by comparing historical versions and optimizes subsequent generation strategies based on reinforcement learning mechanisms.
[0029] Furthermore, the user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback.
[0030] Furthermore, the user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback.
[0031] Furthermore, the storyboard version iteration knowledge base unit integrates a federated learning framework, which optimizes model parameters through distributed node collaboration and generates a dynamic decision tree visualization report.
[0032] The beneficial effects of this invention are as follows:
[0033] 1. This invention can automatically capture the deep semantic information of the script, efficiently identify scene transitions, changes in character emotions, and implied intentions in dialogue, thereby providing structured input data for subsequent storyboard generation based on the diffusion model; it can efficiently construct the underlying structure of the storyboard script, optimize the smoothness of scene transitions and character interaction logic, thereby supporting the dynamic generation of the subsequent diffusion model.
[0034] 2. This invention can efficiently generate dynamic storyboard sequences that conform to visual narrative logic, significantly improving the visual fluency and plot expressiveness of the storyboard; it can efficiently identify and correct temporal misalignments and visual defects in the storyboard sequence, significantly improving the overall fluency and artistic expressiveness of the storyboard; it can efficiently output complete dynamic storyboards that can be directly used in film and television production, thereby greatly reducing the need for manual intervention and improving the integration efficiency of the production process.
[0035] 3. This invention achieves accurate analysis and visual expression of script text by deeply integrating natural language processing, diffusion models, and adaptive feedback mechanisms, thereby ensuring that the generated storyboard has high temporal fluency, artistic expressiveness, and industrial compatibility; through reinforcement learning mechanisms, the generation strategy is continuously optimized to ensure that the storyboard enhances visual impact and audience immersion while maintaining narrative logic consistency, ultimately achieving seamless collaboration with the post-production process and comprehensively improving the development efficiency and cost-effectiveness of film and television projects.
[0036] 4. This invention enables efficient and intelligent conversion from script to dynamic storyboard, significantly improving the efficiency of pre-production for film and animation. By automatically parsing script semantics, constructing a spatiotemporal graph, iteratively generating and optimizing storyboard sequences, and integrating three-dimensional spatial analysis and adaptive feedback mechanisms, it ensures that the generated storyboards reach professional standards in narrative logic, visual fluency, and artistic expression. The version iteration knowledge base and reinforcement learning mechanism continuously optimize the generation strategy, and the user interaction interface supports human-machine collaborative adjustment. Finally, it outputs a standardized industrial-format storyboard, seamlessly connecting to downstream production processes and significantly reducing the cost and time investment of manual storyboard design. Attached Figure Description
[0037] Figure 1 This is a flowchart of the method of the present invention;
[0038] Figure 2 This is a system architecture diagram of the present invention;
[0039] Figure 3 This is a flowchart of the system's workflow in this invention. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0041] Please see Figure 1 This invention provides an automated method for generating dynamic storyboard scripts based on a diffusion model, comprising the following steps:
[0042] Step 1: Input the script text data to be processed. A pre-trained natural language processing (NLP) model is used to perform semantic analysis and extract key information from the script text, identifying scene descriptions, character actions, dialogue content, and plot turning points. This step automatically captures the deep semantic information of the script, efficiently identifying scene transitions, changes in character emotions, and implicit intentions in dialogue, thus providing structured input data for subsequent storyboard generation based on a diffusion model. Specifically, the pre-trained NLP model comprehensively analyzes the script text through word segmentation, entity recognition, and dependency parsing techniques, ensuring that key elements such as plot climaxes and action sequences are accurately extracted, improving the processing efficiency and accuracy of subsequent steps.
[0043] Step 2: Based on the key information extracted in Step 1, construct the initial spatiotemporal framework of the storyboard, determining the scene spatial layout, character positions, and basic camera movement direction for each scene. This step efficiently constructs the underlying structure of the storyboard, optimizes the smoothness of scene transitions and character interaction logic, thereby supporting the dynamic generation of subsequent diffusion models. Specifically, based on the extracted key information, a graph neural network is used to model the spatiotemporal relationships, mapping the scene spatial layout to a coordinate grid. Character positions are accurately located using a keypoint detection algorithm, and the basic camera movement direction is predicted by a probability distribution model to determine the camera movement trajectory. This process ensures that the storyboard framework is adaptive and scalable, significantly improving the visual coherence and narrative efficiency of the generated script.
[0044] Step 3: Using a pre-trained diffusion model, a dynamic storyboard sequence conforming to visual narrative logic is iteratively generated based on the spatiotemporal framework and the needs of the script's plot development. The iterative generation process includes noise addition and denoising optimization mechanisms. This step can efficiently generate dynamic storyboard sequences that conform to visual narrative logic, significantly improving the visual fluency and plot expressiveness of the storyboard. Specifically, the diffusion model simulates potential shot changes through multiple rounds of noise addition and performs conditional denoising optimization based on the spatiotemporal framework, precisely controlling shot angles, motion trajectories, and tempo. At the same time, it adaptively adjusts storyboard elements in conjunction with the needs of the script's plot development, ensuring that the generated sequence has high coherence and artistic expressiveness, thereby greatly reducing the need for manual intervention and enhancing the practicality and adaptability of the generated script.
[0045] Step 4: Perform temporal coherence verification and visual aesthetic scoring on the dynamic storyboard sequence generated in Step 3. Adjust the diffusion model parameters through an adaptive feedback mechanism to optimize the storyboard rhythm and composition. This step can efficiently identify and correct temporal misalignments and visual defects in the storyboard sequence, significantly improving the overall fluency and artistic expression of the storyboard script. Specifically, a temporal analysis algorithm is used to detect the coherence of shot transitions, and a pre-trained visual aesthetic evaluation model is used to quantify the composition, color balance, and lighting effects. The adaptive feedback mechanism dynamically adjusts the noise parameters and denoising strategies of the diffusion model based on the scoring results, iteratively optimizing the rhythmic distribution and visual focus of the storyboard elements. This ensures that the generated script has a high degree of professionalism and adaptability, significantly reduces the cost of manual correction, and enhances the robustness of the system.
[0046] Step 5: Output a complete dynamic storyboard script containing shot angles, durations, camera movement paths, and visual effects annotations. This step efficiently outputs a complete dynamic storyboard script that can be directly used in film and television production, significantly reducing the need for manual intervention and improving the integration efficiency of the production process. Specifically, the system uses standardized data formats (such as JSON or XML) to encapsulate shot angles, durations, camera movement paths, and visual effects annotations. Seamless integration with post-production software is achieved through API interfaces or file export functions, ensuring the immediate usability and cross-platform compatibility of the generated script, further optimizing production efficiency and reducing resource consumption.
[0047] In this embodiment, preferably, the key information extraction in step 1 includes: labeling scene elements through named entity recognition technology, constructing a character interaction relationship graph through dependency parsing, and quantifying the emotional intensity of lines through a sentiment analysis model, thereby comprehensively covering the spatiotemporal and emotional dimensions of the script and providing multimodal structured data support for subsequent storyboard generation.
[0048] In this embodiment, preferably, the diffusion model in step 3 adopts a hierarchical conditional control mechanism. Its input conditions include: scene topology constraint matrix, character motion trajectory probability field, and shot language rule template library. This enables efficient integration of multi-source constraints and optimizes the accuracy and diversity of storyboard generation. Specifically, the hierarchical conditional control mechanism ensures spatial consistency by fusing the scene topology constraint matrix, dynamically predicts the character's movement path by the character motion trajectory probability field, and embeds professional shooting specifications into the shot language rule template library. This adaptively adjusts the noise distribution and denoising strategy, significantly enhancing the narrative logic and visual impact of the storyboard sequence.
[0049] In this embodiment, preferably, the adaptive feedback mechanism in step 4 includes three-dimensional modules: a motion coherence detection module based on optical flow, a visual balance evaluation module based on compositional positioning principles, and a focus trajectory analysis module based on audience attention models. Specifically, the motion coherence detection module based on optical flow is used to quantify the smooth transition of object motion between adjacent shots; the visual balance evaluation module based on compositional positioning principles is used to detect compliance with the golden ratio and visual center of gravity distribution; and the focus trajectory analysis module based on audience attention models is used to track whether the visual focus transfer path of key elements conforms to narrative logic. The motion coherence detection module calculates the directional consistency error of optical flow vectors between adjacent frames, the visual balance evaluation module uses spatial frequency analysis and regional weight mapping techniques to quantify compositional deviation, and the focus trajectory analysis module combines an eye-tracking hotspot prediction algorithm to generate an attention heatmap. The adaptive feedback mechanism dynamically integrates the quantitative scoring results of the above three modules to generate multi-dimensional optimization signals and reversely adjust the latent variable distribution of the diffusion model. Specifically, it includes: adjusting the shot switching time interval parameters according to motion coherence error, correcting the spatial arrangement weight of scene elements according to composition deviation, and optimizing the lens focal length and depth of field configuration by referring to the outlier of the focus trajectory. This systematically improves the temporal smoothness of the storyboard sequence, the aesthetic quality of the picture, and the efficiency of guiding the narrative focus, so that the generated script has industrial-grade production standards.
[0050] In this embodiment, preferably, the output format of step 5 is compatible with film and television industry standards, including but not limited to: generating FinalDraft parsable XML metadata, automatically binding Unreal Engine camera track data, and exporting DaVinciResolve compatible EDL editing decision tables, so that it can be seamlessly integrated into the entire film and television production process, supporting automated collaboration of script writing, scene preview, camera motion simulation and post-editing, ensuring that the generated storyboard data can directly drive actual shooting and post-processing, reducing manual intervention, and improving overall production efficiency and industrial compatibility.
[0051] In this embodiment, preferably, it also includes step 6: establishing a storyboard version iteration knowledge base, generating a difference heatmap by comparing historical versions, and optimizing subsequent generation strategies based on a reinforcement learning mechanism. This knowledge base systematically stores historically generated storyboard script versions and their corresponding parameter configurations. It analyzes the differences between versions through a multi-dimensional feature comparison algorithm, generates a visual difference heatmap to highlight key change areas (such as shot switching frequency, element layout offset, focus trajectory deviation, etc.), and constructs a strategy optimization model based on a reinforcement learning mechanism. It trains a deep neural network through reward functions (such as improved temporal fluency, aesthetic score gain, and improved focus guidance efficiency), dynamically adjusts the latent variable sampling strategy and feedback parameter weights of the diffusion model, thereby iteratively optimizing the subsequent storyboard generation process and achieving adaptive version evolution and continuous improvement in generation quality.
[0052] The aforementioned method enables highly efficient and fully automated generation of dynamic storyboards, significantly improving the efficiency and quality of film and television production. Specifically, this method deeply integrates natural language processing, diffusion models, and adaptive feedback mechanisms to achieve accurate analysis and visual expression of the script text, ensuring that the generated storyboards possess high temporal fluency, artistic expressiveness, and industrial compatibility. In practical applications, the system can adapt to different script types (such as action films, dramas, or animations), automatically optimizing shot transition rhythm and frame composition, greatly reducing the need for manual intervention, while supporting rapid iteration and version management across multiple scenarios. Furthermore, this method continuously optimizes the generation strategy through reinforcement learning mechanisms, ensuring that the storyboards maintain narrative logic consistency while enhancing visual impact and audience immersion, ultimately achieving seamless collaboration with the post-production process and comprehensively improving the development efficiency and cost-effectiveness of film and television projects.
[0053] Please see Figures 2-3 The present invention also provides an automated generation system for dynamic storyboard scripts based on a diffusion model, comprising:
[0054] The script input interface receives user-input script text and converts it into a structured data format. It supports multiple input formats, including plain text files, rich text documents, and files exported from common screenwriting software. By integrating a natural language processing engine, it automatically parses elements such as scene divisions, character dialogues, action descriptions, and camera directions from the script, generating standardized structured data. The conversion process employs a deep learning-based semantic segmentation algorithm to ensure high-precision extraction of key narrative elements and outputs JSON or XML formats compatible with subsequent diffusion model processing. Simultaneously, the interface provides real-time data verification and error feedback mechanisms, allowing users to interactively correct the parsing results, improving the reliability and completeness of the input data.
[0055] The semantic analysis unit is used to parse script scene elements, character behaviors, and time markers to generate a spatiotemporal semantic graph. The semantic analysis unit employs multimodal feature fusion technology, combining convolutional neural networks and temporal modeling algorithms to perform in-depth analysis of the scene spatial layout, character behavior patterns, and time series in the script. Specifically, it includes: using an entity relationship graph convolutional network to identify the topological connections between scene elements, modeling the temporal evolution path of character behaviors through a long short-term memory network, and integrating an attention mechanism to capture key time markers. The generated spatiotemporal semantic graph is stored in the form of a directed weighted graph, where nodes represent scene entities (such as locations and props), edges represent spatiotemporal relationships (such as distance and temporal dependencies), and weights quantify the strength of these relationships. The graph output is in JSON-LD format, compatible with W3C standards, ensuring data structure consistency and scalability.
[0056] The diffusion model-driven unit iteratively generates a storyboard sequence based on a semantic graph. It optimizes shot transitions through a motion coherence detection module, adjusts composition parameters through a visual balance evaluation module, and dynamically allocates visual weights through a focus trajectory analysis module. The diffusion model-driven unit employs a hierarchical iterative generation mechanism. First, it initializes the storyboard layout based on the semantic graph, then progressively optimizes the storyboard sequence through multiple rounds of denoising. During iteration, the motion coherence detection module integrates optical flow algorithms and long short-term memory networks to quantify the displacement continuity between adjacent shots and automatically adjust the switching sequence to reduce visual discontinuities. The visual balance evaluation module applies the golden ratio and color distribution model to calculate the symmetry and saturation weights of scene elements in real time, dynamically correcting the composition proportions. The focus trajectory analysis module combines character movement paths and gaze vectors, utilizing an attention weight allocation mechanism to optimize the visual prominence of key narrative elements. The final generated storyboard sequence is output as a structured sequence file, supporting JSON or XML formats to ensure seamless compatibility with downstream rendering engines.
[0057] The spatiotemporal framework building unit is used to integrate storyboard sequences and construct a framework structure that includes a timeline, spatial coordinates, and camera motion trajectories. The timeline module employs a timestamp alignment algorithm to precisely arrange the start points and durations of the storyboard sequences, ensuring that shot transitions are synchronized with the narrative rhythm. The spatial coordinate module applies 3D scene reconstruction technology, generating Cartesian coordinate system parameters based on the positional data of storyboard elements to define the relative positions of characters, props, and the background. The camera motion trajectory module integrates a Bézier curve interpolation algorithm to calculate smooth transition paths from the camera's perspective and avoids spatial conflicts through a collision detection mechanism. During framework construction, a real-time optimization module dynamically adjusts trajectory parameters using a kinematic model to enhance visual smoothness. The final output framework structure is encoded in a standardized data format, supporting JSON or XML serialization for direct use by downstream rendering engines.
[0058] The diffusion model generation unit is responsible for executing the iterative generation algorithm of the diffusion model, including noise addition and denoising optimization mechanisms, to generate dynamic storyboard sequences that conform to visual narrative logic. The noise addition stage employs Gaussian noise injection technology to gradually perturb the initial storyboard sketch data, simulating a random diffusion process. The denoising optimization mechanism applies a conditional diffusion model, gradually restoring key visual elements through back-diffusion iteration, optimizing shot composition, character actions, and scene details. This unit integrates a pre-trained visual language model as a constraint module to ensure that the generated storyboard sequence meets script requirements in terms of narrative rhythm, emotional expression, and plot coherence. During iteration, an adaptive learning rate scheduler dynamically adjusts the optimization step size to balance generation speed and quality. The final output dynamic storyboard sequence is encoded in a standardized keyframe data format, supporting seamless integration with the spatiotemporal framework construction unit, enabling real-time transmission and integration of storyboard elements.
[0059] The Adaptive Feedback Unit integrates the 3D module and uses optical flow to detect motion coherence, assess visual balance in composition determination, and allocate attention weights through focus trajectory analysis, dynamically adjusting the storyboard rhythm and composition. The Adaptive Feedback Unit analyzes the optical flow characteristics of the storyboard sequence in real time, calculating motion vector fields to quantify the smoothness of transitions between shots, ensuring natural action connections. In composition determination, it applies the golden ratio and visual center of gravity algorithms to evaluate the spatial distribution of elements in the frame and automatically correct unbalanced compositions. Focus trajectory analysis utilizes an attention heatmap model to track the viewer's eye path and optimize the focus position allocation of keyframes. The Adaptive Feedback Unit also integrates 3D spatial data to reconstruct scene depth information and dynamically adjusts lens focal length and angle of view to enhance the sense of depth. Feedback results are input into the diffusion model generation unit in real time, forming a closed-loop optimization mechanism to improve the narrative fluency and visual appeal of the storyboard. Finally, the adjusted storyboard parameters are encoded as metadata, compatible with standardized keyframe data formats, achieving seamless system integration.
[0060] The script output unit is used to implement the output function, converting the optimized storyboard into a format compatible with film and television industry standards. Specifically, it parses the metadata output by the adaptive feedback unit, applies a template-based conversion engine, and supports multiple industry standard protocols such as EDL (Editorial Decision List), AAF (Advanced Authoring Format), and XML. This unit integrates a format verification module to automatically detect the syntax integrity and compatibility of the output file, ensuring seamless integration with non-linear editing systems such as Adobe Premiere or Final Cut Pro. At the same time, the output results can be exported in real time as executable script files or transmitted to a cloud rendering platform via API interface, facilitating subsequent integration and collaboration in film and television production workflows.
[0061] The verification unit is used to establish a knowledge base for storyboard version iterations. It generates a difference heatmap by comparing historical versions and optimizes subsequent generation strategies based on reinforcement learning mechanisms. This includes storing historical storyboard script data, calculating dynamic parameter differences between versions, generating a visual heatmap to identify optimization areas, and applying reinforcement learning algorithms to adjust diffusion model parameters to improve the quality of subsequent storyboard generation. The verification unit uses a distributed database to store historical storyboard script data, supporting efficient indexing and querying. When calculating dynamic parameter differences between versions, it applies the Euclidean distance algorithm to quantify the magnitude of change for core parameters such as spatial depth, focal length, and perspective of keyframes, and generates a multi-dimensional difference matrix. When generating a visual heatmap, it uses color coding technology (e.g., red to indicate high-difference areas and blue to indicate low-difference areas) to intuitively display optimization priorities, and supports manual annotation by users through an interactive interface. When applying reinforcement learning algorithms, it analyzes heatmap feedback based on the Q-learning mechanism and dynamically adjusts the weight parameters of the diffusion model, such as noise scheduling rules and attention mechanisms, to iteratively optimize the storyboard generation strategy. At the same time, the verification unit and the adaptive feedback unit collaborate in real time, feeding the optimization results back to the knowledge base to form a closed-loop learning system, significantly improving the efficiency and accuracy of storyboard iteration and reducing manual intervention.
[0062] The user interaction interface is used to receive user feedback data and integrate it into the adaptive feedback unit, including scene adjustment suggestions, aesthetic preference input, and rhythm control instructions, to achieve human-computer collaborative optimization. The user interface adopts a multimodal input design, including a graphical user interface (GUI) and an application programming interface (API). It supports real-time reception of user feedback data, such as drag-and-drop suggestions for scene adjustments, drop-down menus for aesthetic preferences (including lighting effects, color saturation, and composition style), and slider controls for adjusting rhythm control commands (such as timeline scaling and frame rate adjustment). The interface converts user input into structured JSON format through a data parsing module and automatically maps it to the knowledge graph nodes of the adaptive feedback unit, ensuring seamless alignment between feedback data and system parameters. Simultaneously, the interface integrates a natural language processing (NLP) component, supporting semantic parsing of voice or text commands to handle ambiguous feedback and enhance human-computer interaction efficiency. During integration, the user interface and the adaptive feedback unit are synchronized in real time, triggering optimization loops through an event-driven mechanism. For example, user-annotated aesthetic preferences are fed back to the attention layer weights of the diffusion model, thereby strengthening the human-computer collaborative optimization effect and significantly improving the adaptability and personalization of scene generation.
[0063] The script output interface is used to implement the functionality of the script output unit and transmit storyboard script data to external film and television production systems or user terminals through a standardized interface, ensuring seamless integration into the actual production workflow. The script output interface supports multiple standardized data formats (such as JSON, XML, or Protobuf) to meet the compatibility requirements of different external systems, and achieves efficient, low-latency data transmission through RESTful API or WebSocket protocols. The interface also has built-in error detection and retry mechanisms to automatically handle network interruptions or format anomalies, ensuring the integrity and reliability of the storyboard script during transmission. Simultaneously, the interface provides configurable output options, allowing users to specify the script's resolution, encoding standard, and metadata embedding, facilitating direct import into editing software (such as Adobe Premiere or DaVinci Resolve) by film and television production systems, and providing real-time feedback on the generation status to the user terminal to enhance the visualization and controllability of the production workflow.
[0064] The storyboard version iteration knowledge base unit is used to establish and manage the storyboard version iteration knowledge base. It generates difference heatmaps by comparing historical versions and optimizes subsequent generation strategies based on reinforcement learning mechanisms. The unit uses a distributed database architecture to store all historical storyboard script versions, supporting version snapshot records and timestamp indexes for easy backtracking and comparison. When generating difference heatmaps, this unit uses image processing algorithms to automatically identify areas of change in key elements (such as scene layout, character actions, or camera angles) and highlights iteration differences in a visual heatmap format, helping users intuitively evaluate the effects of modifications. Simultaneously, the reinforcement learning mechanism collects user preference data and generation quality indicators (such as script coherence or creative matching degree) through feedback loops, dynamically adjusting generation strategy parameters (such as diffusion model weights or iteration steps) to improve the accuracy and efficiency of subsequent storyboard generation. Furthermore, this unit integrates version conflict detection functionality, automatically merging user input and system suggestions to ensure data consistency and production process optimization during knowledge base updates.
[0065] The system described above enables efficient and intelligent conversion from scripts to dynamic storyboards, significantly improving the efficiency of pre-production in film and animation. By automatically parsing script semantics, constructing a spatiotemporal graph, and iteratively generating optimized storyboard sequences, the system integrates 3D spatial analysis and adaptive feedback mechanisms to ensure that the generated storyboards meet professional standards in narrative logic, visual fluency, and artistic expression. Simultaneously, a version iteration knowledge base and reinforcement learning mechanism continuously optimize the generation strategy, and the user interface supports human-machine collaborative adjustments. Ultimately, it outputs standardized, industrial-format storyboards that seamlessly integrate with downstream production processes, significantly reducing the cost and time investment of manual storyboard design.
[0066] In this embodiment, preferably, the user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback. The gesture recognition sensor captures user gestures, such as translation, scaling, or rotation, to intuitively adjust camera angles, character actions, or scene elements in the storyboard script. The eye-tracking module monitors the user's eye movements in real time, analyzes the distribution of gaze hotspots, automatically highlights areas of user focus in the storyboard, and triggers local script optimization. These biofeedback signals are transmitted to the reinforcement learning mechanism through a low-latency interface to dynamically adjust model weights, ensuring that the fine-tuning process is seamlessly integrated into the overall generation process, while simultaneously improving user engagement and generation efficiency.
[0067] In this embodiment, preferably, the user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback. The gesture recognition sensor captures user gestures in three-dimensional space, such as pinching, sliding, or rotating, using a depth camera, directly mapping them to the adjustment of shot parameters in the storyboard script. The eye-tracking module uses infrared sensing technology to track eye movements in real time, identify focal points, and automatically magnify relevant areas, triggering local script optimization. This biofeedback data is transmitted to the adaptive algorithm unit via a high-speed data bus to optimize model parameters in real time, ensuring low latency and seamless integration with the overall generation process during fine-tuning, thereby significantly improving user creation efficiency and system response accuracy.
[0068] In this embodiment, preferably, the storyboard version iteration knowledge base unit integrates a federated learning framework. This framework enables distributed nodes to collaboratively optimize model parameters and generate a dynamic decision tree visualization report. The federated learning framework supports multiple distributed nodes sharing model updates while protecting data privacy. It optimizes global parameters through an asynchronous gradient aggregation mechanism, effectively reducing communication overhead. Simultaneously, the collaborative training process between nodes generates a dynamic decision tree visualization report. This report maps decision branches and optimization paths in real-time using a tree structure and supports interactive exploration, helping users intuitively evaluate the version iteration effect and model convergence status, thereby improving the adaptability and decision transparency of the storyboard generation system.
[0069] Based on the system architecture, the working principle and usage process of this invention are as follows:
[0070] First, the user inputs initial script elements (such as scene description and character action instructions) through the user interface. After the system parses the semantics, it initiates the spatiotemporal semantic graph construction unit. This unit extracts key entities (characters, props) and their spatiotemporal relationships, enhances time marker recognition through an attention mechanism, and generates a structured JSON-LD graph. Subsequently, the diffusion model-driven unit receives the graph data and performs hierarchical iterative generation: after initializing the storyboard layout, it gradually refines the shot sequence through multiple rounds of noise addition and denoising optimization (using Gaussian noise perturbation and a conditional diffusion model); during this process, the motion coherence detection module evaluates displacement continuity through an optical flow algorithm and an LSTM network, automatically correcting the switching timing; the visual balance evaluation module applies the golden ratio rule to adjust the composition parameters; and the focus trajectory analysis module assigns visual weights based on the character's motion path, finally outputting the storyboard sequence in JSON / XML format.
[0071] Secondly, the spatiotemporal framework construction unit integrates the storyboard sequence: the timeline module arranges the timing of shots using a timestamp alignment algorithm; the spatial coordinate module generates Cartesian coordinate system parameters based on 3D reconstruction technology to locate scene elements; the camera motion trajectory module uses Bézier curve interpolation to calculate a smooth path and avoids spatial conflicts through collision detection; the real-time optimization module dynamically adjusts trajectory parameters and finally outputs standardized framework data.
[0072] Subsequently, the adaptive feedback unit initiates closed-loop optimization: motion coherence is quantified using optical flow, the image composition module applies a visual center-of-gravity algorithm to correct image balance, focus trajectory analysis generates an attention heatmap to adjust visual weights, and 3D depth data is integrated to dynamically adjust focal length and perspective. The feedback results are transmitted back to the diffusion model generation unit in real time, triggering parameter adjustments (such as noise scheduling rules or attention layer weights) to improve narrative fluency.
[0073] Finally, the script output unit converts the optimized storyboard sequence into industry-standard formats such as EDL and AAF. After compatibility checks by the format verification module, it is transmitted to a non-linear editing system or cloud platform via the script output interface. Simultaneously, the verification unit continues to operate: a distributed database stores historical versions, calculates keyframe parameter differences and generates color-coded heatmaps; a reinforcement learning mechanism (based on Q-learning) analyzes the difference data, dynamically optimizes the diffusion model strategy, and forms a knowledge-based iterative upgrade cycle. Users can fine-tune the process in real time via gesture recognition or eye tracking, and the system updates the federated learning framework parameters accordingly, ultimately generating a decision tree report, achieving end-to-end adaptive optimization.
[0074] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for automatically generating dynamic storyboard scripts based on a diffusion model, characterized in that, Includes the following steps: Step 1: Input the script text data to be processed, and use a pre-trained natural language processing model to perform semantic analysis and key information extraction on the script text to identify scene descriptions, character actions, dialogue content and plot turning points. Step 2: Based on the key information extracted in Step 1, construct the initial spatiotemporal relationship framework of the storyboard script, and determine the scene spatial layout, character positions and basic camera movement direction corresponding to each storyboard. Step 3: Using a pre-trained diffusion model, based on the spatiotemporal relationship framework and the needs of the script plot development, iteratively generate a dynamic sequence of shots that conforms to the visual narrative logic. The iterative generation process includes noise addition and noise reduction optimization mechanisms. Step 4: Perform temporal coherence verification and visual aesthetic scoring on the dynamic shot sequence generated in Step 3, and adjust the diffusion model parameters through an adaptive feedback mechanism to optimize the shot rhythm and screen composition. Step 5: Output a complete dynamic storyboard script that includes camera angles, durations, camera movement paths, and visual effects annotations.
2. The method for automatically generating dynamic storyboard scripts based on a diffusion model according to claim 1, characterized in that, The key information extraction in step 1 includes: labeling scene elements using named entity recognition technology, constructing a character interaction relationship graph using dependency parsing, and quantifying the emotional intensity of dialogue using a sentiment analysis model.
3. The method for automatically generating dynamic storyboard scripts based on a diffusion model according to claim 1, characterized in that, The diffusion model in step 3 adopts a hierarchical conditional control mechanism, and its input conditions include: scene topology constraint matrix, character motion trajectory probability field, and camera language rule template library.
4. The method for automatically generating dynamic storyboard scripts based on a diffusion model according to claim 1, characterized in that, The adaptive feedback mechanism in step 4 includes three-dimensional modules: a motion coherence detection module based on optical flow, a visual balance evaluation module based on compositional design principles, and a focus trajectory analysis module based on audience attention models.
5. The method for automatically generating dynamic storyboard scripts based on a diffusion model according to claim 1, characterized in that, The output format of step 5 is compatible with film and television industry standards, including but not limited to: generating FinalDraft parsable XML metadata, automatically binding Unreal Engine camera track data, and exporting DaVinciResolve compatible EDL editing decision tables.
6. The method for automatically generating dynamic storyboard scripts based on a diffusion model according to claim 1, characterized in that, It also includes step 6, establishing a storyboard version iteration knowledge base, generating a difference heatmap by comparing historical versions, and optimizing subsequent generation strategies based on reinforcement learning mechanisms.
7. A dynamic storyboard script automated generation system based on a diffusion model, characterized in that, include: The script input interface is used to receive script text input by the user and convert it into a structured data format; The semantic analysis unit is used to parse script scene elements, character behaviors, and time markers to generate a spatiotemporal semantic graph. The diffusion model-driven unit is used to iteratively generate a sequence of scenes based on the semantic map, optimize shot transitions through the motion coherence detection module, adjust composition parameters through the visual balance evaluation module, and dynamically allocate visual weights through the focus trajectory analysis module. The spatiotemporal framework building unit is used to integrate the storyboard sequence and build a framework structure that includes a timeline, spatial coordinates, and camera motion trajectory. The diffusion model generation unit is used to execute the iterative generation algorithm of the diffusion model, including noise addition and denoising optimization mechanisms, to generate dynamic shot sequences that conform to visual narrative logic. The adaptive feedback unit is used to integrate the 3D module. It detects motion coherence through optical flow, evaluates visual balance by defining the composition, and allocates attention weights by analyzing focus trajectory, thereby dynamically adjusting the rhythm and composition of the storyboard. The script output unit is used to implement the output function, converting the optimized storyboard into a format compatible with film and television industry standards. The verification unit is used to establish a knowledge base for storyboard version iterations. It generates a difference heatmap by comparing historical versions and optimizes the subsequent generation strategy based on the reinforcement learning mechanism. This includes storing historical storyboard script data, calculating the dynamic parameter differences between versions, generating a visual heatmap to identify optimization areas, and applying reinforcement learning algorithms to adjust the diffusion model parameters to improve the quality of subsequent storyboard generation. The user interaction interface is used to receive user feedback data and integrate it into the adaptive feedback unit, including scene adjustment suggestions, aesthetic preference input, and rhythm control instructions, to achieve human-computer collaborative optimization. The script output interface is used to implement the function of the script output unit and transmit the storyboard script data to the external film and television production system or user terminal through a standardized interface to ensure seamless integration into the actual production process. The storyboard version iteration knowledge base unit is used to establish and manage the storyboard version iteration knowledge base. It generates a difference heatmap by comparing historical versions and optimizes subsequent generation strategies based on reinforcement learning mechanisms.
8. The automated dynamic storyboard generation system based on a diffusion model according to claim 7, characterized in that, The user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback.
9. The automated dynamic storyboard generation system based on a diffusion model according to claim 7, characterized in that, The user interaction interface integrates a multimodal input channel, including a gesture recognition sensor and an eye-tracking module, to achieve real-time fine-tuning of the storyboard based on biofeedback.
10. The automated generation system for dynamic storyboard scripts based on a diffusion model according to claim 7, characterized in that, The storyboard version iteration knowledge base unit integrates a federated learning framework, which optimizes model parameters through distributed node collaboration and generates a dynamic decision tree visualization report.
Citation Information
Cited By
Virtual-real fusion scene automatic construction method based on AIGC script generation
CN121861245A
AI video filming system and method based on multi-model abstraction layer
CN122205197A