An interactive storyboard, a pre-visualization processing method, and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本发明实施例的主要目的在于提出一种交互式故事板和预可视化处理方法、装置、电子设备、存储介质及程序产品,旨在解决现有技术中的至少一种问题
[0016]本发明实施例至少包括以下有益效果:本发明提供一种交互式故事板和预可视化处理方法、装置、电子设备、存储介质及程序产品,该方案通过获取目标对象输入的用于描述目标场景的结构化属性数据和包含角色对话序列的非结构化剧本数据;其中,结构化属性数据包括第一控制指令集和第二控制指令集;将非结构化剧本数据解析为场景元数据,进而结合结构化属性数据在预设的电影镜头数据库中执行语义匹配检索,得到初始影片帧序列;其中,初始影片帧序列中的各帧与非结构化剧本数据中的对话行一一对应;将初始影片帧序列中的每一帧作为基础画布,利用第一控制指令集驱动预设的模型工具对基础画布执行重打光及风格化处理,生成中间故事板帧序列;将中间故事板帧序列中的每一帧作为待处理图像,利用第二控制指令集驱动扩散模型对待处理图像中的人物区域执行局部重绘,生成目标故事板帧序列;在实时预览界面中显示目标故事板帧序列,以获得目标对象的反馈指令,根据反馈指令输出故事板序列结果。本发明实施例通过分层获取结构化属性数据与非结构化剧本数据,可以将复杂、模糊的抽象视觉描述分解为离散、标准化的控制指令,能够显著降低文本构思转化为视觉方案的技术门槛;其次,本发明实施例利用语义匹配检索得到的初始影片帧序列与剧本对话一一对应,并结合模型驱动的重打光、风格化及角色局部重绘处理,能够确保从剧本到多镜头故事板的视觉连续性;不同于传统手绘易产生前后不一致的问题,本发明实施例可以在保持人物身份和场景构图一致的基础上,进一步实现场景属性的动态统一;并且,本发明实施例最终通过引入实时预览与反馈机制,可即时修改任何参数并观察到相应的视觉效果更新,能够实现高效的闭环交互流程,极大提升了创作探索的效率。
Smart Images

Figure CN122574170A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an interactive storyboard, a pre-visualization processing method, and related equipment. Background Technology
[0002] In the pre-production stage of filmmaking, effective visual communication between the director and the director of photography is crucial to ensuring the artistic quality of the film. However, transforming the abstract text of the screenwriter's script into a concrete and executable visual solution always presents significant technical challenges.
[0003] Traditionally, the film industry has relied on hand-drawn storyboards for this transformation. The process of creating storyboards requires directors or professional artists to manually sketch the composition, character positions, and general atmosphere of each key shot, based on the script description. This traditional method has significant drawbacks. First, it is highly demanding and inefficient, heavily reliant on the artist's personal drawing skills and artistic expression; incomplete or poorly executed sketches can easily lead to misunderstandings between departments. Second, this method lacks dynamic interactivity; once a storyboard is completed, modifications to lighting effects, color schemes, or character designs often require redrawing, resulting in a time-consuming, labor-intensive, and costly creative iteration process. Furthermore, hand-drawn storyboards inherently struggle to maintain visual continuity across multiple shots and scenes, especially in complex multi-character dialogues, where inconsistencies in character appearance, costume details, and ambient lighting can disrupt overall visual coherence. Summary of the Invention
[0004] The main objective of this invention is to provide an interactive storyboard and pre-visualization processing method, apparatus, electronic device, storage medium, and program product, aiming to solve at least one problem in the prior art.
[0005] To achieve the above objectives, one aspect of this invention provides an interactive storyboard and a pre-visualization processing method, the method comprising: The system acquires structured attribute data describing the target scene and unstructured script data containing character dialogue sequences from the input of the target object; wherein the structured attribute data includes a first control instruction set and a second control instruction set. Unstructured script data is parsed into scene metadata, and then combined with structured attribute data to perform semantic matching retrieval in a pre-set movie shot database to obtain an initial film frame sequence; wherein, each frame in the initial film frame sequence corresponds one-to-one with the dialogue lines in the unstructured script data; Each frame in the initial film frame sequence is used as the base canvas. The first control instruction set drives the preset model tools to perform relighting and stylization processing on the base canvas to generate the intermediate storyboard frame sequence. Each frame in the intermediate storyboard frame sequence is taken as the image to be processed. The second control instruction set is used to drive the diffusion model to perform local redrawing of the character area in the image to be processed, and generate the target storyboard frame sequence. Display the target storyboard frame sequence in the real-time preview interface to obtain feedback instructions from the target object, and output the storyboard sequence result based on the feedback instructions.
[0006] In some embodiments, obtaining structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object includes the following steps: The structured attribute data is obtained by receiving the first and second control instruction sets from the target object through a hierarchical menu. Receive unstructured script data input by the target object through the text editing area.
[0007] In some embodiments, the hierarchical menu includes a first-level menu interface and a second-level menu interface. Receiving a first set of control instructions and a second set of control instructions input by a target object through the hierarchical menu includes the following steps: The system receives the target object's selections for scene environment, light source direction, time attribute, director style, and lighting effect type from the first-level menu interface, and generates the first set of control instructions. The second-level menu interface receives the target object's selections of facial feature details, hairstyle parameters, clothing style, clothing material, and clothing color, and generates a second set of control instructions.
[0008] In some embodiments, each film frame sequence in the film shot database is associated with annotation information, including scene, time, and character dialogue relationships. Unstructured script data is parsed into scene metadata, and then combined with structured attribute data to perform semantic matching retrieval in a preset film shot database to obtain an initial film frame sequence, including the following steps: Template parsing is performed on unstructured script data to extract scene metadata, including the number of characters, character gender, dialogue rounds, and scene location information. Combine scene metadata with scene environment and time attributes in structured attribute data as query conditions; The query conditions are used to perform similarity matching with the annotation information of each film frame sequence in the film lens database, and the film frame sequence with the highest similarity is output as the initial film frame sequence.
[0009] In some embodiments, the first control instruction set includes scene environment, light source direction, time attribute, director style, and lighting effect type. Using the first control instruction set to drive preset model tools to perform relighting and stylization processing on the base canvas, generating an intermediate storyboard frame sequence, includes the following steps: Use the scene environment as the background scene, import the light source direction, time attribute, light effect type and background scene into the preset image relighting tool, and use the image relighting tool to relight the base canvas to obtain the first canvas. Identify the target director based on their style, index key shots from the target director's representative works from the material library as style-supporting materials, and incorporate the target director into the cue word template to construct stylized cue words. Stylistic aids and stylization tips are input as instructions into a large visual model to guide it in stylizing the first canvas and obtaining the second canvas. The second canvas corresponding to each frame in the initial movie frame sequence is aggregated to form the intermediate storyboard frame sequence.
[0010] In some embodiments, the diffusion model is driven by a second set of control instructions to perform local redrawing of the character region in the image to be processed, generating a target storyboard frame sequence, including the following steps: Obtain the character selection command for the intermediate storyboard frame sequence input by the target object, and determine the target character; The second set of control instructions is used as the local generation conditions for the target character; wherein, the second set of control instructions includes facial feature details, hairstyle parameters, clothing style, clothing material and clothing color, and the local generation conditions include the character design local conditions organized according to facial feature details and hairstyle parameters, and the clothing design local conditions organized according to clothing style, clothing material and clothing color. The local generation condition-driven diffusion model is used to perform local redrawing operations on the corresponding pixel regions of the target character in the image to be processed, generating the target design frame; where the local conditions of the character design correspond to the pixel region of the target character's head, and the local conditions of the clothing design correspond to the pixel region of the target character's body. The target design frames corresponding to each frame in the intermediate storyboard frame sequence are summarized as the target storyboard frame sequence.
[0011] In some embodiments, the feedback instructions include adjustment instructions and confirmation instructions. Outputting the storyboard sequence result based on the feedback instructions includes the following steps: If the feedback instruction is a confirmation instruction, the target storyboard frame sequence will be used as the storyboard sequence result; If the feedback instruction is an adjustment instruction, in response to the adjustment instruction, the corresponding parameters in the first control instruction set or the second control instruction set are adjusted to update the structured attribute data. Then, the process returns to execute the step of using each frame in the initial movie frame sequence as the base canvas until the feedback instruction is a confirmation instruction. Finally, the target storyboard frame sequence of the last iteration is used as the storyboard sequence result.
[0012] To achieve the above objectives, another aspect of the present invention provides an interactive storyboard and a pre-visualization processing device, the device comprising: The first module is used to acquire structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object; wherein, the structured attribute data includes a first control instruction set and a second control instruction set; The second module is used to parse unstructured script data into scene metadata, and then combine it with structured attribute data to perform semantic matching retrieval in a preset movie shot database to obtain an initial film frame sequence; wherein, each frame in the initial film frame sequence corresponds one-to-one with the dialogue lines in the unstructured script data; The third module is used to take each frame in the initial film frame sequence as the basic canvas, and use the first control instruction set to drive the preset model tools to perform relighting and stylization processing on the basic canvas to generate the intermediate storyboard frame sequence. The fourth module is used to take each frame in the intermediate storyboard frame sequence as an image to be processed, and use the second control instruction set to drive the diffusion model to perform local redrawing of the character area in the image to be processed, thereby generating the target storyboard frame sequence. The fifth module is used to display the target storyboard frame sequence in the real-time preview interface to obtain feedback instructions from the target object and output the storyboard sequence result based on the feedback instructions.
[0013] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.
[0014] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0015] To achieve the above objectives, another aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0016] The embodiments of the present invention include at least the following beneficial effects: The present invention provides an interactive storyboard and a pre-visualization processing method, apparatus, electronic device, storage medium, and program product. This solution obtains structured attribute data describing a target scene and unstructured script data containing character dialogue sequences input by a target object. The structured attribute data includes a first control instruction set and a second control instruction set. The unstructured script data is parsed into scene metadata, and then semantic matching retrieval is performed in a preset movie shot database in conjunction with the structured attribute data to obtain an initial film frame sequence. Each frame in the initial film frame sequence corresponds one-to-one with a dialogue line in the unstructured script data. Each frame in the initial film frame sequence is used as a basic canvas, and a preset model tool driven by the first control instruction set is used to perform relighting and stylization processing on the basic canvas to generate an intermediate storyboard frame sequence. Each frame in the intermediate storyboard frame sequence is used as an image to be processed, and a diffusion model driven by the second control instruction set is used to perform local redrawing of the character areas in the image to be processed to generate a target storyboard frame sequence. The target storyboard frame sequence is displayed in a real-time preview interface to obtain feedback instructions from the target object, and the storyboard sequence result is output based on the feedback instructions. This invention, through hierarchical acquisition of structured attribute data and unstructured script data, can decompose complex and ambiguous abstract visual descriptions into discrete and standardized control instructions, significantly reducing the technical threshold for transforming textual ideas into visual solutions. Secondly, this invention utilizes semantic matching retrieval to obtain an initial film frame sequence that corresponds one-to-one with script dialogues, and combines model-driven relighting, stylization, and local character redrawing to ensure visual continuity from script to multi-shot storyboards. Unlike traditional hand-drawn animation, which is prone to inconsistencies, this invention can maintain character identities and scene structures. Figure 1 Building upon this foundation, the dynamic unification of scene attributes is further achieved. Furthermore, by introducing a real-time preview and feedback mechanism, this embodiment of the invention allows for the immediate modification of any parameter and the observation of corresponding visual effect updates, enabling an efficient closed-loop interactive process and greatly improving the efficiency of creative exploration. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an implementation environment for the interactive storyboard and pre-visualization processing method provided in this embodiment of the invention; Figure 2 This is a flowchart illustrating the interactive storyboard and pre-visualization processing method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the workflow architecture principle of the interactive storyboard and pre-visualization processing provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of the interactive storyboard and pre-visualization processing device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0019] It is understood that the terms "first," "second," etc., used in this invention may be used to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of embodiments of this invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to determination," or "in the event of a determination."
[0020] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0021] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this invention is for descriptive purposes only and is not intended to limit the invention.
[0022] While some automated storyboard generation tools based on generative adversarial networks or variational autoencoders have emerged in related technologies, they still rely on matching text with preset static image templates. They cannot achieve real-time and precise interactive control of lighting, environmental atmosphere, character details, etc., and are also difficult to meet the director's need for flexible and rapid experimentation with different visual styles during the creative exploration period.
[0023] In view of this, this embodiment of the invention provides an interactive storyboard and a pre-visualization processing method and related equipment. This solution obtains structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object. The structured attribute data includes a first control instruction set and a second control instruction set. The unstructured script data is parsed into scene metadata, and then semantic matching retrieval is performed in a preset movie shot database in conjunction with the structured attribute data to obtain an initial film frame sequence. Each frame in the initial film frame sequence corresponds one-to-one with a dialogue line in the unstructured script data. Each frame in the initial film frame sequence is used as a basic canvas, and the first control instruction set drives a preset model tool to perform relighting and stylization processing on the basic canvas, generating an intermediate storyboard frame sequence. Each frame in the intermediate storyboard frame sequence is used as an image to be processed, and the second control instruction set drives a diffusion model to perform local redrawing of the character areas in the image to be processed, generating a target storyboard frame sequence. The target storyboard frame sequence is displayed in a real-time preview interface to obtain feedback instructions from the target object, and the storyboard sequence result is output based on the feedback instructions. This invention, through hierarchical acquisition of structured attribute data and unstructured script data, can decompose complex and ambiguous abstract visual descriptions into discrete and standardized control instructions, significantly reducing the technical threshold for transforming textual ideas into visual solutions. Secondly, this invention utilizes semantic matching retrieval to obtain an initial film frame sequence that corresponds one-to-one with script dialogues, and combines model-driven relighting, stylization, and local character redrawing to ensure visual continuity from script to multi-shot storyboards. Unlike traditional hand-drawn animation, which is prone to inconsistencies, this invention can maintain character identities and scene structures. Figure 1 Building upon this foundation, the dynamic unification of scene attributes is further achieved. Furthermore, by introducing a real-time preview and feedback mechanism, this embodiment of the invention allows for the immediate modification of any parameter and the observation of corresponding visual effect updates, enabling an efficient closed-loop interactive process and greatly improving the efficiency of creative exploration.
[0024] It is understood that the interactive storyboard and pre-visualization processing method provided by this invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet, laptop, or desktop computer, but it is not limited to these.
[0025] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0026] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0027] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0028] Terminal 102 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment of the invention does not impose any limitations.
[0029] For example, based on Figure 1The implementation environment shown in this embodiment of the invention provides an interactive storyboard and a pre-visualization processing method. The following description uses the application of this interactive storyboard and pre-visualization processing method in server 101 as an example. It can be understood that this interactive storyboard and pre-visualization processing method can also be applied in terminal 102.
[0030] Reference Figure 2 , Figure 2 This is an optional flowchart of an interactive storyboard and pre-visualization processing method provided in an embodiment of the present invention. The executing entity of the interactive storyboard and pre-visualization processing method can be any of the aforementioned computer devices (including servers or terminals). Figure 2 The method may include, but is not limited to, steps S100 to S500.
[0031] Step S100: Obtain structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object; wherein, the structured attribute data includes a first control instruction set and a second control instruction set; It should be noted that in some embodiments, step S100 may include the following steps: receiving a first control instruction set and a second control instruction set input by the target object through a hierarchical menu to obtain structured attribute data; and receiving unstructured script data input by the target object through a text editing area.
[0032] For example, in some implementations, the user first inputs the script and an SQL statement defining fixed or variable attributes (i.e., structured attribute data). Specifically, this invention proposes a hierarchical input strategy that decomposes complex visual descriptions into multiple independent layers: a base layer (covering basic visual elements such as scene type, time, and lighting direction), a character and action layer (defining the number, position, expression, and actions of characters), and a detail and style layer (handling clothing, specific lighting effects, color tones, and special effects).
[0033] It should be noted that the layered menu includes a first-level menu interface and a second-level menu interface. In some embodiments, receiving the first and second control instruction sets input by the target object through the layered menu may include the following steps: receiving the target object's selections of scene environment, light source direction, time attribute, director style, and lighting effect type from the first-level menu interface, and generating the first control instruction set; receiving the target object's selections of facial feature details, hairstyle parameters, clothing style, clothing material, and clothing color from the second-level menu interface, and generating the second control instruction set.
[0034] For example, in some implementations, in the first-level menu, the director can select basic scene attributes, including environment settings (e.g., bedroom or airport), light source direction (e.g., left or right), time (e.g., noon, night, or sunset), director style (10 well-known director styles are available for selection), and lighting effects (e.g., soft light, hard light, key light). Additionally, basic actor attributes (e.g., expressions such as anger, joy, or sadness, and facial features such as eyes, nose, and mouth) and hairstyles can be set. To further simplify input, the system provides the director with a list of commonly used and effective descriptive terms to choose from, eliminating the need for detailed manual descriptions.
[0035] In the second-level menu, directors can further refine the character's facial features, hairstyle length and color, and clothing style (including material and color). This structured and intuitive input method not only simplifies the workflow but also effectively avoids the ambiguity and complexity of traditional methods. Directors can focus more on the creative aspects rather than technical details. By selecting the most relevant basic descriptions for the script scenes, the system automatically generates accurate storyboards, improving communication efficiency and ensuring that the final visual outcome aligns with the director's vision.
[0036] Step S200: Parse the unstructured script data into scene metadata, and then combine it with structured attribute data to perform semantic matching retrieval in the preset movie shot database to obtain the initial film frame sequence; Each frame in the initial film frame sequence corresponds one-to-one with a line of dialogue in the unstructured script data; It should be noted that each film frame sequence in the film shot database is associated with annotation information, which includes scene, time, and character dialogue relationships. In some embodiments, the following steps are included: performing template parsing on unstructured script data to extract information such as the number of characters, character gender, dialogue rounds, and scene location to form scene metadata; merging the scene metadata with the scene environment and time attributes in the structured attribute data as query conditions; and using the query conditions to perform similarity matching with the annotation information of each film frame sequence in the film shot database, outputting the film frame sequence with the highest similarity as the initial film frame sequence.
[0037] Exemplary, in some embodiments, the method of the present invention utilizes a comprehensive film database to match shots based on time sequence and character dialogue. This method follows a traditional storyboard workflow: from script to visuals (framing and shot composition), and then from visuals to character interaction (dialogue and interaction), ensuring natural continuity between shots. Specifically, script information is first filtered: based on the director's vision, the director provides specific details for dialogue scenes, such as the number of characters, their genders, time, and location of the dialogue; then, text-to-film database matching is performed: the system searches the film database for storyboards that meet given conditions, where multiple outputs can be generated (e.g., matching the Top K most similar film frame sequences), each output consisting of an image corresponding to a line of dialogue in the script.
[0038] Step S300: Take each frame in the initial film frame sequence as the basic canvas, and use the first control instruction set to drive the preset model tool to perform relighting and stylization processing on the basic canvas to generate an intermediate storyboard frame sequence. It should be noted that the first control instruction set includes scene environment, light source direction, time attribute, director style, and lighting effect type. In some embodiments, step S300 may include the following steps: using the scene environment as the background scene, importing the light source direction, time attribute, lighting effect type, and background scene into a preset image relighting tool, and using the image relighting tool to relight the base canvas to obtain the first canvas; determining the target director based on the director style, indexing key shots of the target director's representative works from the material library as style auxiliary materials, and substituting the target director into the prompt word template to construct stylized prompts; inputting the style auxiliary materials and stylized prompts as instruction conditions into a large visual model to guide the large visual model to perform stylized visual processing on the first canvas to obtain the second canvas; and summarizing the second canvas corresponding to each frame in the initial film frame sequence as an intermediate storyboard frame sequence.
[0039] For example, in some specific implementations, relighting and stylization can be achieved as follows: Lighting is a key element in conveying the emotional tone and atmosphere of a scene in the script. To enable directors and cinematographers to better communicate and confirm the expected emotions, this embodiment of the invention integrates advanced relighting functionality based on IC-Light (an open-source AI image relighting tool, short for "Imposing Consistent Light," designed to solve the problem of lighting consistency when foreground and background blend in an image. It can intelligently adjust the lighting direction, intensity, color temperature, etc. of the subject based on text cues or background images, making the composite result more natural and realistic). This module allows users to adjust the direction, type, texture, and color of light sources in existing film visual content. Recognizing that lighting in a shot often involves multiple interacting sources, the system provides a wide range of control options. In addition, a "Director's Master Style" function is provided. Directors can simulate the style of a specific filmmaker by applying pre-configured style cues—extracted from key works and refined through multiple rounds of testing using the GPT-4O large visual model—thereby unifying the lighting, tone, and overall visual style.
[0040] Step S400: Take each frame in the intermediate storyboard frame sequence as the image to be processed, and use the second control instruction set to drive the diffusion model to perform local redrawing of the character area in the image to be processed, thereby generating the target storyboard frame sequence. It should be noted that in some embodiments, step S400 may include the following steps: obtaining the character selection instruction for the intermediate storyboard frame sequence input by the target object, and determining the target character; using the second control instruction set as the local generation condition for the target character; wherein, the second control instruction set includes facial feature details, hairstyle parameters, clothing style, clothing material, and clothing color, and the local generation condition includes the character design local condition organized based on the facial feature details and hairstyle parameters, and the clothing design local condition organized based on the clothing style, clothing material, and clothing color; using the local generation condition to drive the diffusion model to perform a local redrawing operation on the corresponding pixel area of the target character in the image to be processed, and generating the target design frame; wherein, the character design local condition corresponds to the head pixel area of the target character, and the clothing design local condition corresponds to the body pixel area of the target character; and summing the target design frames corresponding to each frame in the intermediate storyboard frame sequence as the target storyboard frame sequence.
[0041] For example, in some specific embodiments, character design and costume creation can be achieved as follows: In this embodiment of the invention, character design and costume creation is a crucial stage in translating the director's vision into visually appealing and realistic character portrayals. This process involves selecting appropriate attributes, utilizing filmmaking expertise, and employing diffusion models to create character models and costumes that conform to the overall narrative and aesthetic style of the film.
[0042] Character design begins with understanding the personality, role, and visual characteristics of each character in the story. The system allows directors to modify facial features such as expressions, eyes, nose, and mouth based on visual references of actors, ensuring that the character's appearance matches the emotions and personality they wish to convey. Furthermore, the system offers options to adjust hair features, such as hair length, hairstyle, and hair color, making it easy to create unique characters that align with the director's vision and meet audience expectations. Directors can fine-tune individual attributes within the user interface, selecting specific eye shapes, nose shapes, and mouth expressions. This level of customization allows directors to adjust each character's appearance to reflect their personality and narrative role, ensuring that the character's visual design seamlessly integrates with their function in the story.
[0043] Costume Design: Equally important as character design, costume creation further reinforces a character's identity, role, and the world they inhabit. In this embodiment of the invention, costume design is a flexible and intuitive process. The system provides a variety of options to adjust clothing elements, such as tops, trousers / skirts, dresses, and other accessories, allowing directors to design character appearances based on the desired historical context, social status, and thematic tone of the film. Customizable features include fabric texture, clothing color, and style, all designed to capture the essence of the world and the character's traits throughout their journey. The system also allows directors to customize the overall look by selecting clothing style (e.g., modern, retro, formal, casual), color (e.g., bold, soft, muted), and material texture (e.g., leather, silk, denim), ensuring the character's appearance reflects their personality and fits the scene's atmosphere. Once customized, the system applies the selected clothing to the character in a visual reference, ensuring the character's appearance is consistent with their personality and the overall tone of the scene.
[0044] Step S500: Display the target storyboard frame sequence in the real-time preview interface to obtain feedback instructions from the target object, and output the storyboard sequence result according to the feedback instructions; It should be noted that the feedback instructions include adjustment instructions and confirmation instructions. In some embodiments, outputting the storyboard sequence result according to the feedback instructions may include the following steps: if the feedback instruction is a confirmation instruction, the target storyboard frame sequence is used as the storyboard sequence result; if the feedback instruction is an adjustment instruction, in response to the adjustment instruction, the corresponding parameters in the first control instruction set or the second control instruction set are adjusted to update the structured attribute data, and the step of using each frame in the initial movie frame sequence as the base canvas is returned to be executed until the feedback instruction is a confirmation instruction, and the target storyboard frame sequence of the last iteration is used as the storyboard sequence result.
[0045] For example, in some specific implementations, the effect of the target storyboard frame sequence can be viewed in a real-time display interface, enabling rapid and flexible generation and customization of the visualization from script to image. If the effect does not meet the user's needs, the user can also modify the image by adjusting the options of various parameters in the structured attribute data (such as selecting a background from the background library, adjusting lighting effects (e.g., soft light, hard light, key light), and selecting the direction of light). For character design and costume creation, the user can adjust facial features, hairstyle, hair color, and clothing style (e.g., top, pants / skirt, dress). These adjustments allow the creation of characters that meet the director's vision and script requirements. This ensures that the effect meets the requirements.
[0046] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0047] First and foremost, it's important to note that effective communication between the director and cinematographer is fundamental to filmmaking. However, traditional methods, relying on visual references and hand-drawn storyboards, often lack the efficiency and precision required in pre-production. Filmmaking is a highly visual medium, with the ultimate goal of transforming the script (the written form of the story) into a rich, immersive cinematic experience. However, scripts inherently have limitations in conveying visual elements crucial to realizing the director's creative vision. Elements such as lighting, camera angles, character positioning, and atmosphere are vital to the film's overall impact, yet are often not adequately or explicitly represented in the script. Like screenwriters, directors rely on visual imagination when conceiving scenes, envisioning the tones, color schemes, and emotional tensions that will bring the narrative to life. These mental images guide their decisions about how the story should be presented and felt on screen. However, translating these mental ideas into a shared language that allows for effective communication with key production departments such as cinematography, art direction, costume design, and lighting remains a significant challenge. Poor or ambiguous communication of these visual concepts can lead to inefficiencies, costly delays, and a final product that deviates from the director's original vision.
[0048] Traditionally, storyboards have been used to address this problem, allowing directors to sketch out key moments and visually represent scenes and shots. While storyboards are effective at conveying shot composition, angles, and framing, they are often limited by the director's drawing skills and the level of detail achievable with the medium. Incomplete or poorly drawn storyboards can cause confusion, leading to repeated clarifications between the director and key departments. Furthermore, creating traditional storyboards is time-consuming and disconnected from the dynamic nature of visual elements such as lighting and character design, which often require frequent revisions early in production. This disconnect between the static nature of storyboards and the fluidity of the visual design process further exacerbates inefficiency and communication breakdowns.
[0049] In light of this, and to address these challenges, this invention introduces CineVision (an interactive AI-powered storyboard and pre-visualization system for collaboration between directors and cinematographers). This platform seamlessly integrates scriptwriting with visual pre-visualization, providing directors with a dynamic and efficient storyboard creation system that enables them to visualize and optimize scenes in real-time as their scripts develop. CineVision leverages a vast film database to match script descriptions with accurate, context-relevant visual content, capturing the spatial and emotional dynamics of each scene. This bridges the gap between written narrative and the director's visual concept, allowing directors to see their scripts come to life. It also provides a wealth of reference material to facilitate collaboration with key production teams such as the director of photography. CineVision's core innovation lies in its ability to directly integrate image relighting, style control, and customizable character design into the pre-visualization workflow. Utilizing an AI-driven diffusion model, directors can dynamically adjust lighting—changing intensity, direction, and color—to evoke specific emotions and enhance the emotional tone of a scene. The system also applies the visual art styles of master directors, providing them with instant visual feedback on how their style choices affect the overall look and feel. Furthermore, CineVision supports real-time customization of character designs, allowing directors to modify facial expressions, hairstyles, clothing, and accessories to adapt to an evolving script. This integration ensures that all visual elements—from lighting to character design—evolve in sync with the narrative, maintaining alignment with the director's vision throughout pre-production. By providing this seamless, interactive environment, CineVision reduces the need for repetitive revisions, enhances the director's creative control, streamlines interdepartmental communication, and improves overall production efficiency.
[0050] In some embodiments, the preparatory work involved in the technical solution of the present invention includes: 1. Story visualization and storyboard generation: Storyboards are crucial in film and media production, providing a visual representation of scenes and shots before filming. They enable directors to communicate their vision to key departments such as cinematography, art direction, costume design, and acting. Traditionally, creating storyboards involved manually drawing each frame, a time-consuming process limited by the artist's skill and production timeline. However, recent advances in automation aim to streamline this process. Generative techniques such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) are now used to analyze scripts and generate corresponding visual scenes. These techniques improve the realism and consistency of generated images. Large-scale datasets, such as the VWP dataset, have become invaluable resources for training character-based visual narrative models, further driving the development of automated storyboard generation.
[0051] The development of intelligent writing assistants has accelerated this process. These tools enhance the creative process by generating text proposals from keywords, engaging in interactive dialogues, or processing multimodal input. Research on these systems has demonstrated their effectiveness in supporting story writing, leading to the emergence of many new intelligent writing tools. For example, ScriptViz, developed by Rao et al., focuses on script analysis and image retrieval, matching text descriptions with pre-existing images. However, these methods rely on predefined templates, limiting the flexibility in generating novel visual content. With advancements in machine learning and computer vision, methods such as Stable Diffusion have shown great potential in automatically generating high-quality, context-accurate visual content based on text descriptions. For instance, a director might describe a scene as "nighttime, a car chase through wet city streets," and the model would generate images that accurately reflect the scene's lighting, emotion, and composition, providing greater creative flexibility.
[0052] These models can render fine details such as lighting effects and complex camera angles, which were previously challenging to achieve manually. The integration of style transfer techniques also allows directors to apply specific visual styles to generated scenes, enhancing their creative control. However, challenges remain in maintaining stylistic consistency and adapting to diverse directorial preferences. Furthermore, integrating complex narrative structures into consecutive storyboard shots presents challenges, such as ensuring continuity of actor movement, background consistency, and smooth lighting transitions. In light of these issues, the work of this invention aims to address these challenges by enabling directors to define visual preferences and control stylistic elements within consecutive storyboard shots, providing a personalized, director-centric pre-visualization approach.
[0053] 2. Image relighting and cinematic lighting effects: Lighting is a fundamental aspect of filmmaking, shaping the mood and visual aesthetics of a scene. It has long been an essential tool for cinematographers, enabling them to evoke specific emotions and emphasize key elements within a scene. Meanwhile, image relighting—the process of altering lighting conditions in an image—has become a crucial technique in computer vision and computational photography. This method allows for manipulation of lighting in digital images, providing new opportunities for visual experimentation and enhancing realism in image processing.
[0054] Traditional image relighting methods involve manual adjustments using photo editing software, processes that are often slow and require significant expertise. Therefore, there is growing interest in automating image relighting using machine learning techniques. Recent advancements in deep learning have significantly advanced the automation of image relighting, allowing for realistic and dynamic lighting effects. Techniques such as Transformers and diffusion models have proven effective in simulating complex lighting conditions, including variations in light intensity, direction, and color. These models learn from large datasets of real-world lighting scenes, enabling them to generate plausible lighting effects that conform to photophysics. Furthermore, combining physically based rendering models with machine learning algorithms further enhances the physical accuracy of relighting techniques, ensuring visual appeal and realism. In filmmaking, image relighting plays a crucial role in pre-visualization, allowing directors to experiment with different lighting setups before shooting. For example, systems like IC-Light utilize real-time relighting based on scene metadata, allowing creators to visualize how various lighting configurations will affect a shot. The system of this invention integrates image relighting technology into the storyboard generation process, enabling directors to specify lighting parameters in the pre-production stage and ensure that the generated visual effects match their artistic vision.
[0055] 3. AI-driven storyboard character design and costume creation: Character design, including costumes, makeup, and facial expressions, plays a crucial role in pre-production as it helps bring characters in the narrative to life. Traditionally, creating detailed character designs for storyboards is a manual and labor-intensive process, requiring highly skilled artists to visualize and draw characters in specific environments and costumes. This approach is not only time-consuming but also often limited by the artist's ability to translate textual descriptions into visual representations. However, recent advances in AI and machine learning are transforming this process by automatically generating highly detailed character designs and costume variations based on textual descriptions and even sketches.
[0056] AI-driven models such as GANs and diffusion models have demonstrated remarkable capabilities in generating realistic human figures, costume designs, and makeup styles. These models have been trained on large datasets of fashion, historical costumes, and character design, enabling them to create diverse and context-appropriate character visuals, ranging from modern fashion to fantasy attire. Furthermore, the integration of AI techniques for generating makeup and facial features is gaining attention. Diffusion models can adjust facial attributes such as age, gender, expression, and even specific makeup styles (e.g., makeup for period dramas or fantasy settings). This allows for dynamic customization of characters, enabling directors to quickly explore different visual interpretations of their characters. A key area of research in this field is generating multi-dimensional character features that go beyond costumes and makeup. For example, diffusion models now allow for detailed customization of facial features, hairstyles, accessories, and even the texture and material of clothing. AI can interactively modify these attributes, providing filmmakers with a flexible and creative space to refine their character designs during pre-production.
[0057] The system of this invention leverages these advancements in diffusion models to provide a multi-dimensional approach to character and costume creation. Directors can specify parameters such as facial features, hairstyles, clothing textures, accessories, and even emotions, which are subsequently incorporated into the generated visuals. This allows for a higher level of personalization and creative control over characters, ensuring their design aligns with the director's vision and the narrative needs of the storyboard. By combining diffusion models for character and costume creation with other aspects of pre-visualization, such as scene composition and lighting, the system of this invention provides a comprehensive solution for film directors and cinematographers.
[0058] In some specific implementations, the design process and goal achievement of CINEVISION provided by the embodiments of the present invention are as follows: In filmmaking, visualizing textual information and effectively communicating shot instructions and visual requirements to key departments such as cinematography, costume design, makeup, and art direction is crucial. Traditional communication often relies on the director's ability to convey these ideas through sketches, which can lead to misunderstandings and repeated clarifications. For example, aspects such as shot angles, lighting design, overall style, and character design typically require extensive back-and-forth communication between the director and cinematographer. This communication process is both time-consuming and costly.
[0059] To address the aforementioned challenges, this invention involved user-centric iterative design with directors and cinematographers, ultimately resulting in CineVision, a system that supports directors and cinematographers in creating storyboards during the pre-visualization phase. With CineVision, the goal of this invention is to establish a more efficient and clearer way of visual communication between directors and cinematographers. The design process comprised three phases: 1) Identifying user pain points and needs: revealing the strategies used by directors and cinematographers when communicating visual concepts through storyboards. 2) Prototyping and walkthroughs: designing and developing CineVision based on design goals and user needs, then gathering user feedback and iterating accordingly. 3) Deployment and evaluation: conducting user research to evaluate how directors and cinematographers interact with CineVision and their perceived usability and ease of use.
[0060] The following section elaborates on the relevant strategies and objectives of the design process of this invention, and reports on the strategies and design objectives that guided the design and development of CineVision (see Table 1 below).
[0061] Table 1
[0062] This invention identifies a key problem: verbal or written descriptions, due to their abstract nature, often lead to ambiguity and poor communication (C1). Specifically, accurately conveying visual details such as shot composition and lighting effects through text is difficult, resulting in increased communication costs and frequent misunderstandings between departments. To address this issue, this invention proposes a layered input strategy that decomposes complex visual descriptions into three independent layers: a base layer (covering basic visual elements such as scene type, time, and lighting direction), a character and action layer (clarifying the number, position, expression, and actions of characters), and a detail and style layer (handling clothing, specific lighting effects, color tones, and special effects) (S1). The director of photography recommends using standardized terminology and preset templates within the system to minimize ambiguity. Real-time visual preview during the input process is crucial for quickly adjusting and refining elements, thereby reducing the need for extensive modifications and clarifications.
[0063] Film database matching: Maintaining consistency of key visual elements (such as character appearance, lighting, and background) across multiple shots or scenes is particularly challenging and often disrupts overall visual coherence (C2). Inconsistencies in visual style and detail, especially in complex narratives or multi-character dialogues, require significant additional communication and adjustments. To mitigate this problem, this invention proposes an integrated film database matching strategy involving the automatic retrieval of film references that closely match the script description and the use of AI-driven regeneration techniques to ensure visual consistency across scenes (S2).
[0064] Real-time Visual Adjustment: Traditional hand-drawn sketches and verbal descriptions often lead to indirect communication, prolonged confirmation processes, and misunderstandings between production departments, severely slowing down pre-production (C3). Iterative communication not only incurs high time costs but also increases the risk of misunderstandings. To address this, this invention introduces a real-time visual adjustment strategy, enabling directors to directly adjust background, lighting, and character designs within the system and preview the results instantly. This method allows for immediate feedback and rapid adjustments, thereby streamlining the pre-production workflow (S3).
[0065] Flexible Customization of Characters and Costumes: The storyboarding process requires frequent experimentation with various visual styles and cinematic techniques to achieve optimal creative expression. However, traditional tools typically offer limited flexibility, hindering rapid creative iteration and restricting visual diversity (C4). To address this issue, embodiments of this invention incorporate more flexible customization options to freely adjust character facial features, clothing materials, and colors. Therefore, this invention develops a flexible customization strategy for characters and costumes, utilizing a diffusion model for creative generation and customizing diverse materials to achieve dynamic replacement and targeted modifications (S4).
[0066] Specifically, this invention provides four design guidelines (DG1-DG4) to support strategies (S1-S4) proposed to address the challenges of reporting in the pre-visualization phase (C1-C4).
[0067] Lowering the creative threshold (DG1): Facing the C1 challenge of translating abstract textual descriptions into precise shot composition, relying solely on natural language can lead to vague and incomplete visual representations. To address this issue, the strategy (S1) of this invention employs a layered input system, breaking down the overall vision into distinct elements—such as scene type, basic lighting, and initial character placement—using standardized menus and preset options. By leveraging a comprehensive film database (e.g., the "Text-to-Film Database Matching" section), the system automatically matches script elements with appropriate visual references. Furthermore, the integration of automated character design and costume generation (from the character design and costume creation modules) through a diffusion model significantly reduces the burden of manually refining details. In summary, this approach fulfills DG1: lowering the creative threshold by simplifying the process, ensuring that directors can easily externalize their ideas into coherent storyboard visuals without getting bogged down in technical details.
[0068] Ensuring Visual Continuity (DG2): A key challenge identified in this research is C2: maintaining visual consistency across multiple shots and scenes. Directors have noted that traditional storyboards often lack continuity, leading to inconsistencies in character appearance, lighting, and scene composition. To overcome this, the strategy (S2) of this invention incorporates automated semantic matching using a comprehensive film database. As seen in the “Text-to-Film Database Matching” section, this process matches shots based on time sequences and character dialogue, simulating the natural progression of a scene—from framing and shot composition to capturing the dynamics of dialogue and interaction. Furthermore, the system of this invention employs advanced AI-driven relighting technology to ensure that lighting—its direction, type, texture, and color—remains consistent throughout the storyboard. These methods collectively achieve DG2—ensuring visual continuity—faithfully reflecting the director's intentions across all shots by providing a seamless and coherent visual narrative.
[0069] Improving Communication Efficiency (DG3): Furthermore, a recurring pain point is C3, namely the inefficiency caused by the need for repeated confirmations and clarifications between the director, cinematographer, and other production departments. To address this issue, the strategy (S3) of this invention emphasizes real-time visual adjustments and immediate feedback. By combining real-time preview functionality with instant editing (e.g., instant relighting adjustments using IC-light technology in the relighting and stylization modules), the director can immediately modify lighting and style parameters and see their impact on the storyboard. Additionally, the system's intuitive interface integrates input from the film database and character customization modules, ensuring all team members have clear and shared references. This directly contributes to DG3—improved communication efficiency—by reducing the need for back-and-forth interactions, enabling all stakeholders to quickly agree on visual details and advance the production process.
[0070] Supporting Creative Exploration (DG4): Finally, the research of this invention emphasizes C4: the need for flexible, iterative creative exploration. Directors and cinematographers have expressed a desire to quickly test different visual styles, character designs, and lighting effects without being constrained by rigid workflows. To this end, the strategy (S4) of this invention provides a wide range of customization options. The system allows users to experiment with various directorial styles through the "Director Master Style" option and supports dynamic character and costume adjustments through a diffusion model, as detailed in the Character Design and Costume Creation module. Furthermore, the system provides industry expert insights guiding features such as actor number control, precise spatial positioning, and specific emotional adjustments. These capabilities collectively fulfill DG4—Supporting Creative Exploration—by empowering filmmakers to iterate rapidly and experiment with a wide range of aesthetic possibilities, thereby achieving an optimal balance between creative freedom and production feasibility.
[0071] In some specific application scenarios, the CINEVISION system can achieve the following: The CineVision system aims to provide an efficient and coherent pre-visualization workflow, enabling a seamless transition from script to storyboard. The system integrates three core modules: text-to-film database matching, relighting and stylization, and character design and costume creation. These modules are integrated through a simplified two-level menu system, ensuring that the final storyboard not only aligns with the director's creative vision but also maintains a high degree of visual consistency and artistic quality.
[0072] 1. Core system components: 1.1 Text-to-Film Database Matching. In traditional storyboard creation, a scene typically contains multiple shots—including dialogue, various angles, and detailed compositions. Existing systems, emotional storyboards, and traditional storyboards often lack continuity within a scene. Their visual content is usually just a rough reference between the director and cinematographer, failing to accurately reflect the detailed shot composition of the complete scene. The method of this invention utilizes a comprehensive film database to match shots based on time sequence and character dialogue. This method follows the traditional storyboarding process: from script to visuals (framing and shot composition), and then from visuals to character interaction (dialogue and interaction), ensuring natural continuity between shots. In some optional implementations, embodiments of this invention employ "structured semantic parsing → SQL". A four-level progressive matching mechanism—"coarse filtering → multi-dimensional weighted scoring → scene complexity verification"—is used to achieve text-to-movie database matching. First, a finely tuned GPT-4O large language model automatically parses unstructured script text into standardized attribute tags, covering fixed attributes with strong matching constraints such as scene type, time, and number of characters, as well as variable attributes with soft matching constraints such as emotional tone and character interaction methods. Then, the extracted fixed attributes are converted into standard SQL query statements, performing millisecond-level coarse filtering on the preprocessed MovieNet database to select a set of candidate shots that meet basic conditions. Finally, a two-dimensional weighted scoring mechanism is used to evaluate the candidate shots... The shots are meticulously sorted, with CLIP text-visual cross-modal semantic similarity accounting for the first weight (e.g., 60%), focusing on matching the semantic consistency of scene environment, character interaction relationships and overall atmosphere. Character recognizability and narrative fit account for the second weight (e.g., 40%). Face detection algorithms are used to exclude shots where characters' faces are obscured and to verify the consistency of character identities in consecutive shots within the same scene. Finally, a scene complexity filtering algorithm automatically excludes complex group shots with 6 or more characters (the preset number can be adjusted according to actual needs), and the verified shots are sorted according to the classic narrative logic of movie dialogue scenes to ensure natural visual continuity between multiple shots.
[0073] 1.2 Relighting and Stylization. Lighting is a key element in conveying the emotional tone and atmosphere of a scripted scene. To enable directors and cinematographers to better communicate and confirm the intended mood, the CineVision system integrates advanced relighting capabilities built on IC-Light. This module allows users to adjust the direction, type, texture, and color of light sources in existing cinematic visual content. Recognizing that lighting in a shot often involves multiple interacting sources, the system offers extensive control options. Additionally, a "Director's Style" feature is provided. Directors can unify lighting, tone, and overall visual style by applying pre-configured style cues—extracted from key works and refined through multiple rounds of testing using the GPT-4O large visual model—to simulate the style of a specific filmmaker.
[0074] 1.3 Character Design and Costume Creation. Character design and costume creation are integral parts of storytelling. Traditional storyboarding methods often separate character and costume design from shot planning, leading to inconsistencies or delays in the visual presentation of characters on screen. CineVision integrates character design and costume creation directly into the pre-visualization process. Utilizing a diffusion model, the system generates character features and costume designs that align with the overall visual direction and tone of the scene. Directors can adjust facial features (such as expressions, eyes, nose, and mouth) and hairstyles, and fine-tune costume details, including type, texture, and color. This ensures that each character's design is coordinated with shot composition and lighting, achieving a seamless transition from concept to final production.
[0075] In some optional implementations, the diffusion model can employ a two-stage fine-tuning application architecture based on a Stable Diffusion 1.5 pre-trained model. The first stage uses a visual style training set (e.g., 1000 MovieNet movie scene images) to complete the adaptation to a general visual style in the movie domain. The second stage uses character and clothing training data for specific optimization of the character and clothing generation task. Specifically, the training data can consist of 2000 close-up and mid-range images of characters covering various styles such as modern, retro, and fantasy, selected from the MovieNet dataset. Priority is given to samples with clear facial expressions, well-defined clothing features, and natural character interactions. The GPT-4O large language model generates refined text descriptions of the corresponding character's facial features, hairstyle, clothing material, and color, constructing a one-to-one text-image pairing training set. This is achieved on a single NVIDIA A100 80GB... The model is trained on a GPU for 5 epochs to improve the accuracy of character detail generation. At the same time, IC-Light relighting technology is integrated to give the model the ability to adjust the local lighting of the character. The real-time adjustment function is implemented by mapping the facial features, hairstyle and clothing parameters in the system's hierarchical menu to the model's conditional control vectors. It supports targeted local modifications to the generated results, and the global relighting module will be automatically linked during the modification process to ensure that the visual features of the adjusted character are consistent with the overall lighting and composition of the scene.
[0076] 2. System workflow: CineVision provides an efficient and simplified method for matching script descriptions with actual storyboard frames and regenerating creations, building on the technology used by ScriptViz and IC-Light. Figure 3 As a workflow example of CineVision, in some specific implementations, the system can be divided into the following five parts: (1) text information input; (2) script information filtering; (3) text to film database matching; (4) film storyboard regeneration and stylization; (5) cinematography execution stage.
[0077] Specifically, the CineVision system consists of five basic components: (1) Text information input: The director inputs script information into the system. (2) Script information filtering: Based on the director's vision, the director provides specific details for dialogue scenes, such as the number of characters, their genders, the time, and the location where the dialogue takes place. (3) Text to film database matching: The system searches the film database for storyboards that meet the given conditions. (4) Film storyboard regeneration and stylization: The director defines visual parameters, including scene composition, light direction, lighting effects, actors' facial expressions, and actors' costumes and makeup. In addition, the director can also choose the visual style of well-known directors to generate customized storyboards, thereby achieving stylization customization. (5) Cinematographer execution stage: The approved storyboard is sent to the cinematographer to convey the director's vision and guide the lighting and camera settings.
[0078] 2.1 Script to Storyboard Conversion. Traditional storyboard creation typically requires directors and cinematographers to provide detailed descriptions for each scene, including actions, positions, camera angles, and other details, which can be complex and prone to misunderstanding. Furthermore, while AI-driven storyboard generation systems have emerged in recent years, they often face several challenges, such as overly complex cues, inconsistent results, difficulty maintaining character continuity between consecutive frames, poor background consistency, and seamless lighting transitions. These challenges can ultimately affect the final visual presentation. To address these issues, CineVision employs a simplified two-level menu system that significantly reduces input complexity while ensuring the consistency and accuracy of the generated results.
[0079] In the first-level menu, directors can select basic scene attributes, including environment settings (e.g., bedroom or airport), light source direction (e.g., left or right), time (e.g., noon, night, or sunset), director style (10 well-known director styles are available), and lighting effects (e.g., soft light, hard light, key light). Additionally, basic actor attributes can be set (e.g., expressions such as anger, joy, or sadness, and facial features such as eyes, nose, and mouth) and hairstyles. To further simplify input, the system provides directors with a list of commonly used and effective descriptive terms to choose from, eliminating the need for detailed manual descriptions.
[0080] In the second-level menu, directors can further refine the character's facial features, hairstyle length and color, and clothing style (including material and color). This structured and intuitive input method not only simplifies the workflow but also effectively avoids the ambiguity and complexity of traditional methods. Directors can focus more on the creative aspects rather than technical details. By selecting the most relevant basic descriptions for the script scenes, the system automatically generates accurate storyboards, improving communication efficiency and ensuring that the final visual outcome aligns with the director's vision.
[0081] 2.2 Relighting and Stylization. In the CineVision system, relighting and stylization are core tools for enhancing the visual effects of a film and accurately conveying the director's intentions. By fine-tuning the lighting and mimicking specific directorial styles, directors can precisely set the lighting atmosphere and visual style for each scene during pre-production, ensuring more efficient communication with the cinematographer and contributing to the consistency and artistic quality of the final work.
[0082] Relighting: In CineVision, the relighting function is based on IC-Light, providing reference options that align with cinematic styles. It offers directors multiple adjustable options to customize lighting effects according to script requirements. The system offers three common time options: "Noon" provides even, strong sunlight, emphasizing clarity and brightness; "Night" presents low-light sources and cool-toned lighting; and "Sunrise or Sunset" uses warm gold and orange light to create a romantic or emotionally rich atmosphere. Regarding light type selection, considering that directors may not be familiar with professional lighting equipment, CineVision provides intuitive and easy-to-understand options such as "Soft Light," "Hard Light," and "Key Light," which align with the communication preferences of directors and cinematographers, simplifying the complexity of lighting design. In addition, the system provides simple light source direction selection, such as "Left Light" or "Right Light," to help directors establish different lighting effects. To further assist directors in precisely setting up the shooting environment, CineVision provides over 100 common background scenes, covering everyday life, indoor and outdoor environments, ensuring that lighting effects coordinate with the background for perfect visual results.
[0083] Stylization: Every director has a unique visual style, with lighting, color, and shot style often becoming hallmarks of their work. To help directors quickly achieve these styles during the creative process, CineVision offers stylization adjustment capabilities. With 10 customizable style prompts, directors can choose from various classic directorial styles and apply them to their scenes, precisely shaping the visual atmosphere and enhancing the film's artistic expression. In this process, one of the authors, with five years of filmmaking experience, analyzed one of each director's major works and extracted five key shots as a raw material library. These materials were then input into the GPT-4O Large Visual Model (LVM), with instructions such as, "This is a shot from a Wes Anderson film; please help me generate a description of this shot in their style, focusing on lighting, color, and visual style." After three rounds of testing and optimization, the invention ultimately determined the style description for each director and provided specific stylization prompts. These finely tuned and verified prompts are presented as buttons with the director's name and style, allowing users to apply them directly within the system.
[0084] 2.3 Character Design and Costume Creation. In the CineVision system, character design and costume creation are crucial stages in translating the director's vision into visually appealing and realistic character portrayals. This process involves selecting appropriate attributes, leveraging filmmaking expertise, and using diffusion models to create character models and costumes that align with the film's overall narrative and aesthetic style.
[0085] Character Design: Creating characters in CineVision is a meticulous process that begins with understanding the personality, role, and visual characteristics of each character in the story. The system allows directors to modify facial features such as expressions, eyes, nose, and mouth based on visual references of actors, ensuring that the character's appearance matches the emotions and personality they wish to convey. Furthermore, the system offers options to adjust hair features such as hair length, hairstyle, and hair color, making it easy to create unique characters that align with the director's vision and meet audience expectations. Directors can fine-tune individual attributes within the user interface, selecting specific eye shapes, nose shapes, and mouth expressions. This level of customization allows directors to adjust each character's appearance to reflect their personality and narrative role, ensuring that the character's visual design seamlessly integrates with their function in the story.
[0086] Costume Design: Equally important as character design, costume creation further reinforces a character's identity, their role, and the world they inhabit. In CineVision, costume design is a flexible and intuitive process. The system offers a variety of options to adjust clothing elements such as tops, pants / skirts, dresses, and other accessories, allowing directors to design character looks based on the desired historical context, social status, and the film's thematic tone. Customizable features include fabric textures, clothing colors, and styles, all designed to capture the essence of the world and the character's traits throughout their journey. The system also allows directors to customize the overall look by selecting clothing styles (e.g., modern, retro, formal, casual), colors (e.g., bold, soft, muted), and material textures (e.g., leather, silk, denim), ensuring the character's appearance reflects their personality and fits the scene's atmosphere. Once customized, the system applies the selected clothing to the character in visual references, ensuring the character's appearance aligns with their personality and the overall tone of the scene.
[0087] 3. System Operation: CineVision provides a user-friendly interface that allows directors and cinematographers to modify scripts, review visual effects, and refine individual frames in real time, including adjustments to backgrounds, lighting, and character design. This ensures consistency with the dialogue and visual coherence throughout the project, with all images taken from the same scene to maintain consistency.
[0088] Input and Output: The user inputs the script and SQL statements defining fixed or variable attributes. Upon submission, the system generates multiple outputs, each consisting of an image corresponding to a line of dialogue in the script. Users can input their own prompts and select backgrounds, lighting effects, directing style, character designs, adjust system parameters, and view a live display. The SQL statements specify parameters such as settings, character details, and other relevant attributes, while the script determines the number of images generated.
[0089] Editing: Users can modify images by selecting backgrounds from a library of over 100 options, adjusting lighting effects (e.g., soft light, hard light, key light), and choosing the direction of the light. The system ensures that the lighting matches the background for visual consistency. For character design and costume creation, users can adjust facial features, hairstyles, hair color, and clothing styles (e.g., tops, pants / skirts, dresses). These adjustments allow for the creation of characters that meet the director's vision and script requirements.
[0090] In summary, this invention, by acquiring structured attribute data and unstructured script data in a layered manner, decomposes complex and ambiguous abstract visual descriptions into discrete and standardized control instructions. This significantly lowers the technical threshold for transforming textual ideas into precise visual solutions and effectively avoids communication barriers caused by ambiguity in natural language descriptions in traditional methods. Secondly, this invention utilizes semantic matching retrieval to obtain an initial film frame sequence that corresponds one-to-one with script dialogues. Combined with AI-driven relighting, stylization, and partial character redrawing, it ensures visual continuity from the script to the multi-shot storyboard. Unlike traditional hand-drawn animation, which is prone to inconsistencies, this solution maintains character identities and scene structures... Figure 1 Based on this foundation, dynamic unity was achieved in lighting atmosphere, directorial style, and character design. Furthermore, by introducing a real-time preview and feedback mechanism, the present invention allows directors to instantly modify any parameter and observe corresponding visual effect updates. This compresses the traditional linear iterative process of "description-waiting-review-re-modification" into a highly efficient "what you see is what you get" closed-loop interactive process, greatly improving the efficiency of creative exploration and the accuracy of cross-departmental communication.
[0091] like Figure 4 As shown, this embodiment of the invention also provides an interactive storyboard and a pre-visualization processing device 900, which can implement the above-described method. This device may include: The first module 901 is used to acquire structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object; wherein, the structured attribute data includes a first control instruction set and a second control instruction set; The second module 902 is used to parse the unstructured script data into scene metadata, and then combine it with structured attribute data to perform semantic matching retrieval in a preset movie shot database to obtain an initial film frame sequence; wherein, each frame in the initial film frame sequence corresponds one-to-one with the dialogue lines in the unstructured script data. The third module 903 is used to take each frame in the initial film frame sequence as the basic canvas, and use the first control instruction set to drive the preset model tool to perform relighting and stylization processing on the basic canvas to generate an intermediate storyboard frame sequence. The fourth module 904 is used to take each frame in the intermediate storyboard frame sequence as an image to be processed, and use the second control instruction set to drive the diffusion model to perform local redrawing of the character area in the image to be processed, thereby generating the target storyboard frame sequence. The fifth module 905 is used to display the target storyboard frame sequence in the real-time preview interface to obtain feedback instructions from the target object and output the storyboard sequence result based on the feedback instructions.
[0092] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0093] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0094] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0095] like Figure 5 As shown, Figure 5 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (aSIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention. The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RaM). The memory 1002 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0096] The electronic device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0098] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0099] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0100] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0101] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0102] This invention provides an interactive storyboard and pre-visualization processing method, apparatus, electronic device, storage medium, and program product. It acquires structured attribute data describing a target scene and unstructured script data containing character dialogue sequences input by a target object. The structured attribute data includes a first control instruction set and a second control instruction set. The unstructured script data is parsed into scene metadata, and then semantic matching retrieval is performed in a preset movie shot database in conjunction with the structured attribute data to obtain an initial film frame sequence. Each frame in the initial film frame sequence corresponds one-to-one with a dialogue line in the unstructured script data. Each frame in the initial film frame sequence is used as a base canvas, and a preset model tool driven by the first control instruction set performs relighting and stylization processing on the base canvas to generate an intermediate storyboard frame sequence. Each frame in the intermediate storyboard frame sequence is used as an image to be processed, and a diffusion model driven by the second control instruction set performs local redrawing of the character areas in the image to be processed to generate a target storyboard frame sequence. The target storyboard frame sequence is displayed in a real-time preview interface to obtain feedback instructions from the target object, and the storyboard sequence result is output based on the feedback instructions. This invention, through hierarchical acquisition of structured attribute data and unstructured script data, can decompose complex and ambiguous abstract visual descriptions into discrete and standardized control instructions, significantly reducing the technical threshold for transforming textual ideas into visual solutions. Secondly, this invention utilizes semantic matching retrieval to obtain an initial film frame sequence that corresponds one-to-one with script dialogues, and combines model-driven relighting, stylization, and local character redrawing to ensure visual continuity from script to multi-shot storyboards. Unlike traditional hand-drawn animation, which is prone to inconsistencies, this invention can maintain character identities and scene structures. Figure 1 Building upon this foundation, the dynamic unification of scene attributes is further achieved. Furthermore, by introducing a real-time preview and feedback mechanism, this embodiment of the invention allows for the immediate modification of any parameter and the observation of corresponding visual effect updates, enabling an efficient closed-loop interactive process and greatly improving the efficiency of creative exploration.
[0103] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0104] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0107] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. An interactive storyboard and pre-visualization processing method, characterized in that, The method includes the following steps: The system acquires structured attribute data describing the target scene and unstructured script data containing character dialogue sequences from the input of the target object; wherein the structured attribute data includes a first control instruction set and a second control instruction set. The unstructured script data is parsed into scene metadata, and then semantic matching retrieval is performed in a preset movie shot database in combination with the structured attribute data to obtain an initial film frame sequence; wherein, each frame in the initial film frame sequence corresponds one-to-one with a dialogue line in the unstructured script data; Each frame in the initial video frame sequence is used as the base canvas. The first control instruction set drives the preset model tool to perform relighting and stylization processing on the base canvas to generate an intermediate storyboard frame sequence. Each frame in the intermediate storyboard frame sequence is taken as an image to be processed. The second control instruction set is used to drive the diffusion model to perform local redrawing on the character area in the image to be processed, thereby generating the target storyboard frame sequence. The target storyboard frame sequence is displayed in the real-time preview interface to obtain feedback instructions from the target object, and the storyboard sequence result is output according to the feedback instructions.
2. The method according to claim 1, characterized in that, The process of obtaining structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object includes the following steps: The structured attribute data is obtained by receiving the first control instruction set and the second control instruction set input by the target object through a hierarchical menu. The unstructured script data input by the target object is received through the text editing area.
3. The method according to claim 2, characterized in that, The hierarchical menu includes a first-level menu interface and a second-level menu interface. Receiving the first control instruction set and the second control instruction set input by the target object through the hierarchical menu includes the following steps: The first control instruction set is generated by receiving the target object's selection of scene environment, light source direction, time attribute, director style and lighting effect type from the first-level menu interface. The second control instruction set is generated by receiving the target object's selection of facial feature details, hairstyle parameters, clothing style, clothing material, and clothing color from the second-level menu interface.
4. The method according to claim 1, characterized in that, Each film frame sequence in the film shot database is associated with annotation information, including scene, time, and character dialogue relationships. The process of parsing the unstructured script data into scene metadata, and then combining this metadata with the structured attribute data to perform semantic matching retrieval in the preset film shot database to obtain the initial film frame sequence includes the following steps: The unstructured script data is parsed using templates to extract information such as the number of characters, character gender, dialogue rounds, and scene location to form the scene metadata. The scene metadata is combined with the scene environment and time attributes in the structured attribute data as query conditions; The query conditions are used to perform similarity matching with the annotation information of each film frame sequence in the film shot database, and the film frame sequence with the highest similarity is output as the initial film frame sequence.
5. The method according to claim 1, characterized in that, The first control instruction set includes scene environment, light source direction, time attribute, director style, and lighting effect type. The step of using the first control instruction set to drive a preset model tool to perform relighting and stylization processing on the basic canvas to generate an intermediate storyboard frame sequence includes the following steps: Using the scene environment as the background scene, the light source direction, the time attribute, the light effect type and the background scene are imported into a preset image relighting tool. The image relighting tool is then used to relight the base canvas to obtain the first canvas. Based on the director's style, a target director is identified, and key shots from the target director's representative works are indexed from the material library as style-supporting materials. The target director is then substituted into the prompt word template to construct a stylized prompt. The style auxiliary materials and the stylization prompts are input as instruction conditions into the large visual model to guide the large visual model to perform stylization visual processing on the first canvas to obtain the second canvas. The second canvas corresponding to each frame in the initial movie frame sequence is aggregated to form the intermediate storyboard frame sequence.
6. The method according to claim 1, characterized in that, The step of using the second control instruction set to drive the diffusion model to perform local redrawing of the character region in the image to be processed, generating a target storyboard frame sequence, includes the following steps: Obtain the character selection instruction for the intermediate storyboard frame sequence input by the target object, and determine the target character; The second control instruction set is used as the local generation condition for the target character; wherein, the second control instruction set includes facial feature details, hairstyle parameters, clothing style, clothing material and clothing color, and the local generation condition includes character design local conditions organized according to the facial feature details and hairstyle parameters, and clothing design local conditions organized according to the clothing style, clothing material and clothing color; The local generation condition-driven diffusion model is used to perform local redrawing operations on the corresponding pixel regions of the target character in the image to be processed, thereby generating a target design frame; wherein, the local conditions of the character design correspond to the head pixel region of the target character, and the local conditions of the clothing design correspond to the body pixel region of the target character. The target design frames corresponding to each frame in the intermediate storyboard frame sequence are aggregated to form the target storyboard frame sequence.
7. The method according to claim 1, characterized in that, The feedback instructions include adjustment instructions and confirmation instructions. The step of outputting the storyboard sequence result based on the feedback instructions includes the following steps: If the feedback instruction is the confirmation instruction, the target storyboard frame sequence is taken as the storyboard sequence result; If the feedback instruction is the adjustment instruction, in response to the adjustment instruction, the corresponding parameters in the first control instruction set or the second control instruction set are adjusted to update the structured attribute data, and the process returns to the step of using each frame in the initial movie frame sequence as the base canvas, until the feedback instruction is the confirmation instruction, and the target storyboard frame sequence of the last iteration is used as the storyboard sequence result.
8. An interactive storyboard and pre-visualization processing device, characterized in that, The device includes: The first module is used to acquire structured attribute data describing the target scene and unstructured script data containing character dialogue sequences input by the target object; wherein, the structured attribute data includes a first control instruction set and a second control instruction set; The second module is used to parse the unstructured script data into scene metadata, and then combine the structured attribute data to perform semantic matching retrieval in a preset movie shot database to obtain an initial film frame sequence; wherein, each frame in the initial film frame sequence corresponds one-to-one with the dialogue lines in the unstructured script data; The third module is used to take each frame in the initial film frame sequence as a basic canvas, and use the first control instruction set to drive the preset model tool to perform relighting and stylization processing on the basic canvas to generate an intermediate storyboard frame sequence. The fourth module is used to take each frame in the intermediate storyboard frame sequence as an image to be processed, and use the second control instruction set to drive the diffusion model to perform local redrawing on the character area in the image to be processed, thereby generating the target storyboard frame sequence. The fifth module is used to display the target storyboard frame sequence in the real-time preview interface to obtain feedback instructions from the target object and output the storyboard sequence result according to the feedback instructions.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.