AI-based movie production system and method

By combining AI-assisted storyboarding, 3D digitization, virtual camera control, and LoRA models, the issues of controllability and scalability in AI filmmaking have been resolved, enabling the generation of high-quality long videos and complex shot sequences, and improving the efficiency and quality of creative collaboration networks.

CN121644926APending Publication Date: 2026-03-10SHENZHEN TCL HIGH TECH DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511249596.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-27
Filing Date
2025-09-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing AI tools lack controllability and scalability in filmmaking, making it difficult to achieve precise control over film quality and narrative coherence, especially in the generation of long videos and complex shot sequences.

Method used

The film production system employs an AI-based approach, including an AI-assisted storyboard module, a 3D digitization module, a virtual camera controller, an AI animation module, an AI-assisted compositing module, and a post-production module. It reconstructs 3D environments and characters using digital technology, combines LoRA models for video compositing, and achieves precise control over camera movement and character actions.

Benefits of technology

It improves the controllability and consistency of filmmaking, supports complex visual storytelling techniques, enhances the efficiency and quality of creative collaboration networks, and enables the generation of high-quality long videos and complex shot sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644926A_ABST
    Figure CN121644926A_ABST
Patent Text Reader

Abstract

The invention provides an AI-based movie production workflow. The AI-based movie production workflow comprises an AI auxiliary story board, an AI animation and a post-production process. The workflow provides techniques to reconstruct 3D digital environments and roles as well as virtual camera control. The process also includes techniques of capturing 2D live performance and extracting visual cues. The AI animation process uses virtual camera control-based cues to generate composite images and videos (for 3D digitization) and / or visual cues-based generation (for 2D photography). Furthermore, the workflow provides AI-assisted synthesis techniques that generate a synthesized video based on the AI animation video and input from 3D digitization and / or 2D video processing. The process also provides advanced post-processing techniques to generate a complete movie based on the composite video. The framework aims to promote a creative collaborative network and improve consistency, controllability and expandability in AI movie production by using a hybrid digitization method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Application No. 63 / 690,207, filed September 3, 2024, and U.S. Application No. 18 / 961,455, filed November 27, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to an AI-based film production system, and more specifically, to an AI-powered film production workflow platform that enables a creative collaboration network. Background Technology

[0003] In recent years, several artificial intelligence (AI) projects have been explored for automating the filmmaking process. By 2023, AI tools such as Runway and Pika enabled users to create short, realistic, professional-quality videos using text or image prompts. In February 2024, OpenAI launched Sora, capable of generating high-quality videos up to 60 seconds long. Subsequently, tools such as Luma, Kling, and Vidu were developed, further solidifying the foundation for practical AI-powered filmmaking workflows.

[0004] However, at least two major unresolved and often overlooked challenges remain in the field of AI-assisted filmmaking: controllability and scalability. Current AI tools fall short in addressing controllability, limiting the control and precision filmmakers need for film quality. These tools primarily focus on generating short, realistic segments, often neglecting the precise control and consistency required to maintain narrative coherence. Scalability, on the other hand, involves creating a filmmaking workflow that facilitates seamless remote collaboration between filmmakers and artists, crucial for realizing the full potential of AI-assisted filmmaking. The lack of mature workflows in these areas may be attributed to the limitations of current AI technology, which still struggles to support artists in producing high-quality films relatively easily. This may explain the scarcity of AI-generated video content in traditional film studios and distributors. Summary of the Invention

[0005] To address the aforementioned needs and overcome the shortcomings of existing AI tools, this disclosure provides a system and method for integrating digital technologies into an AI-assisted filmmaking workflow, which can be used to enable creative collaboration networks.

[0006] According to one aspect of this disclosure, an artificial intelligence (AI)-based film production system is provided. The system includes: an AI-assisted storyboard module for generating one or more images of a storyboard, one or more cues associated with the storyboard, and script-based guidelines; a 3D digitization module for generating a 3D model of a scene based on the script and the guidelines; a virtual camera controller for controlling camera settings to generate cues associated with the 3D model of the scene; an AI animation module for generating a first video based on the one or more images of the storyboard, the one or more cues associated with the storyboard, and the cues associated with the 3D model of the scene; and an AI-assisted compositing module for generating a composite video based on the first video generated by the AI ​​animation module and a low-rank adaptation (LoRA) model.

[0007] In some examples, the 3D digitization module is also used to generate LoRA models based on the script and the guidelines, wherein the 3D model of the scene represents a digitized 3D environment and the LoRA model represents digitized characters.

[0008] In some examples, the system further includes: a 2D camera capture module for generating a second video based on the script and the guide; and a visual cue extraction module for generating cues related to the second video.

[0009] In some examples, the AI ​​animation module is also used to generate the first video based on the one or more images of the storyboard and the one or more cues associated with the storyboard, as well as the cues associated with the second video.

[0010] In some examples, the visual cue extraction module is also used to generate one or more visual cues based on the second video; the system further includes a video processing module for generating a third video based on the second video and the one or more visual cues.

[0011] In some examples, the AI-assisted synthesis module is also used to generate a synthesized video based on the first video generated by the AI ​​animation module and the third video generated by the video processing module.

[0012] In some examples, the AI-assisted synthesis module is also used to add visual effects, sound effects, or a combination of both to the synthesized video.

[0013] In some examples, the system further includes a post-production module for generating and outputting a complete movie based on the composite video generated by the AI-assisted compositing module.

[0014] In some examples, the post-production module is used to perform one or more post-processing operations on the composite video generated by the AI-assisted compositing module to generate a complete movie.

[0015] According to another aspect of this disclosure, an AI-based filmmaking method is provided. The method includes: generating one or more images of a storyboard, one or more cues associated with the storyboard, and script-based guidelines via an AI-assisted storyboard module; and generating 3D models of scenes based on the script and the guidelines via a 3D digitization module. The virtual camera settings are controlled by a virtual camera controller to generate cues associated with the 3D model of the scene; a first video is generated by an AI animation module based on one or more images of the storyboard, one or more cues associated with the storyboard, and the cues associated with the 3D model of the scene generated by controlling the virtual camera settings; and a synthesized video based on the first video and a low-rank adaptation (LoRA) model is generated by an AI-assisted synthesis module.

[0016] In some examples, the method further includes: generating a LoRA model based on the script and the guide using the 3D digitization module; wherein the 3D model of the scene represents a digitized 3D environment, and the LoRA model represents a digitized character.

[0017] In some examples, the method further includes: generating a second video based on the script and the guide using a 2D camera capture module; and generating cues related to the second video using a visual cue extraction module.

[0018] In some examples, the method further includes generating the first video using an AI animation module, based on one or more images of the storyboard and one or more cues associated with the storyboard, as well as the cues associated with the second video.

[0019] In some examples, the method further includes: generating one or more visual cues based on the second video using the visual cue extraction module; and generating a third video based on the second video and the one or more visual cues using the video processing module.

[0020] In some examples, the AI-assisted synthesis module generates a synthesized video based on the first video and the third video.

[0021] In some examples, the method further includes adding visual effects, sound effects, or a combination of both to the synthesized video via the AI-assisted synthesis module.

[0022] In some examples, the method further includes: generating a full film from the synthesized video using a post-production module; and outputting the full film for review.

[0023] In some examples, the method further includes performing one or more post-processing operations via the post-production module to generate the complete film.

[0024] Beneficial Effects: This disclosure provides an AI-based filmmaking workflow, including AI-assisted storyboarding, AI animation, and post-production processes for creating the film. The workflow provides techniques for reconstructing 3D digital environments and characters, as well as virtual camera control. It also provides techniques for capturing 2D live-action performances and extracting visual cues. The AI ​​animation process uses cues based on virtual camera control (in the case of 3D digitization) and cues based on visual cues (in the case of 2D camera capture) to generate composite images and videos. Furthermore, the workflow provides AI-assisted compositing techniques to generate composite videos based on AI-animated videos and inputs from 3D digitization and / or 2D video processing. The workflow also provides advanced post-processing techniques to generate a complete film based on the composite video. This framework aims to improve consistency, maneuverability, and scalability in AI filmmaking by fostering creative collaboration networks through the use of hybrid digital methods. Attached Figure Description

[0025] Figure 1 The algorithm diagram shows the current state-of-the-art AI-based film production workflow;

[0026] Figure 2 An algorithmic diagram of an AI-based film production workflow, as illustrated in a first exemplary embodiment of this disclosure, is shown.

[0027] Figure 3 An algorithmic diagram of an AI-based film production workflow, as illustrated in a second exemplary embodiment of this disclosure, is shown.

[0028] Figure 4 An algorithmic diagram of an AI-based film production workflow, as illustrated in a third exemplary embodiment of this disclosure, is shown.

[0029] Figure 5A and Figure 5B It is an image of a three-dimensional scene, each image having a different background created by Gaussian scattering, and the same human body elements are added to the three-dimensional space according to one aspect of this disclosure;

[0030] Figure 6A It is the 2D image before processing;

[0031] Figure 6B These are 2D images reconstructed using different camera settings;

[0032] Figure 6C It is a 3D depth model reconstructed under camera control using monocular depth estimation;

[0033] Figure 6D This is one aspect of the disclosure, providing a reconstructed three-dimensional space under camera control (virtual camera);

[0034] Figure 7A This disclosure provides, as one aspect, a set of rows and columns of images of digitally customized human appearances using various different roles, wherein each row corresponds to one person and each column corresponds to a different appearance;

[0035] Figure 7B This is one aspect of the disclosure, which provides a set of three images of a human face replaced with a digital 3D model;

[0036] Figure 8A This is an image of an actor's performance captured by a 2D camera, provided for reference as one aspect of this disclosure;

[0037] Figure 8B This disclosure provides an aspect of using motion cues as action prompts in different contexts. Figure 8A Images of AI-generated characters whose actors perform the same movement styles;

[0038] Figure 9A This is a set of images provided as one aspect of the disclosure to illustrate the compositing process of integrating the performance of a real human actor, a stable AI-generated background, and a moving window scene into a single frame;

[0039] Figure 9B and Figure 9C This is one aspect of the disclosure providing two different perspectives on multiple object layers in the composition process;

[0040] Figure 10 This is a conceptual diagram of an AI film compositing process using an AI filmmaking workflow in a collaborative network, provided as one aspect of this disclosure;

[0041] Figure 11 This is a structural diagram of the creative collaboration network of the AI ​​film production workflow platform provided as one aspect of this disclosure;

[0042] Figure 12 This is an exemplary flowchart of the steps of the first method for conducting AI-based filmmaking provided in this disclosure;

[0043] Figure 13 This is an exemplary flowchart of the steps of the second method for conducting AI-based filmmaking provided in this disclosure;

[0044] Figure 14 This is an exemplary flowchart of the steps of a third method for AI-based filmmaking provided in this disclosure. Detailed Implementation 1. Overview

[0045] This application introduces a novel AI filmmaking framework designed to build a future-oriented creative collaboration network. The AI ​​filmmaking workflow platform employs a hybrid digital approach, involving core steps such as reconstructing 3D digital environments, capturing 2D live-action performances, generating composite images and videos using AI tools, and performing AI-assisted compositing. This AI filmmaking workflow platform and its corresponding AI filmmaking technologies offer a comprehensive solution to effectively address the major challenges currently facing AI video generation, including but not limited to issues of consistency, controllability, scalability, and human motion and interaction, and have the potential to become pioneers in the AI ​​filmmaking industry.

[0046] Digitalization offers a powerful solution to the maneuverability and scalability challenges in AI filmmaking. By digitizing 3D elements such as human characters and backgrounds, AI filmmaking workflows can reconstruct 3D environments, enabling virtual cameras to generate dynamic footage from any location. This approach allows filmmakers to bypass the camera control limitations of current AI tools, providing greater flexibility, precision, and creativity, supporting more complex and innovative visual storytelling techniques. Furthermore, digitalization transforms collaborative networks, enabling more dynamic, inclusive, and efficient teamwork. Digital tools eliminate the constraints of time zones and work hours, allowing teams to work asynchronously and seamlessly hand over tasks across different time zones, ensuring project continuity.

[0047] As previously mentioned, digitization is a key component in enhancing maneuverability and providing greater control in the filmmaking process. The exemplary embodiments of this disclosure effectively address the challenges of coherence and camera control in keyframe animation production by integrating various cues from the digitization process. These cues guide camera parameter settings, narrative logic construction, and the combination of character design and background elements. By integrating these diverse cues, filmmakers can achieve more precise control over the AI-driven creative process.

[0048] To the best of the inventors' knowledge, the AI ​​film production workflow described in this disclosure is the first comprehensive solution to address the three core challenges of consistency, controllability, and scalability in AI film production. The main contributions of the AI ​​film production workflow platform described in this application are as follows: • New Framework: The AI ​​film production workflow platform introduces a new framework that effectively addresses the issues of coherence and controllability in current AI film production solutions. • Higher visual quality: By combining live-action performances with AI-generated backgrounds and objects, the AI ​​filmmaking workflow platform achieves higher visual quality in synthetic films than the most advanced AI tools available. • Enhanced camera control: The AI ​​film production workflow platform significantly improves camera control by integrating virtual camera settings from the digitization process into the AI ​​animation production workflow.

[0049] Before further describing the details of the AI ​​filmmaking framework of various exemplary embodiments of this disclosure, the following will further explain the background information of work related to this field and the various challenges faced by existing AI-based content generation tools to provide additional background and revelation information. 2. Related work

[0050] 2.1 Reconstructing the 3D World from 2D Images

[0051] In AI filmmaking, precise camera movement control is crucial, as directors' creative visions for each scene often require specific camera positions and movements. However, current high-quality image and video generation models face challenges in maintaining scene consistency. While AI-generated scenes may be visually stunning, compositing the same scene from multiple perspectives can lead to issues such as changes in the 3D environment, object removal, and differences in background color between scenes. An effective solution is to separate 3D environment generation from foreground elements. By reconstructing the 3D environment, more consistent and flexible camera movement can be achieved. Therefore, the framework described in this paper relies on reconstructing 3D scenes from images to maintain scene coherence.

[0052] Within the technical framework disclosed herein, advanced Gaussian point cloud technology can be effectively integrated for background scene digitization. Gaussian point clouds, as an innovative method for 3D scene representation and rendering, have gained widespread attention in the fields of computer vision and graphics. This technology provides a highly attractive solution for digitizing film scenes, enabling precise camera control while ensuring the consistency and conditional characteristics of generated content within the shooting scene. Specifically, Gaussian point clouds construct novel visual effects by decomposing a 3D scene into a series of Gaussian point cloud elements and projecting them onto a 2D plane. The rendering process then projects these 3D Gaussian point cloud elements onto the image plane, thereby achieving a realistic 3D scene presentation.

[0053] By employing these new perspectives, we can simulate camera movement within the reconstructed 3D environment, ensuring a consistent background throughout the film. Furthermore, we can use the Stable Diffusion video style transfer model to adjust background textures, adding detail while streamlining the image and maintaining overall harmony.

[0054] 2.2 Stable Diffusion

[0055] Diffusion models, as a powerful class of generative models, have emerged in the field of computer vision in recent years, providing novel solutions for high-fidelity image generation. These models are based on the theoretical framework of Markov chains and stochastic processes. Their core principle lies in the fact that diffusion models generate coherent images from noise by inversely deriving the noise generation process. The basic principle of this model is to first progressively apply Gaussian noise to the data points (forward diffusion process), and then inversely derive this diffusion process by learning an iterative denoising process (backward diffusion process).

[0056] Diffusion models have developed rapidly in recent years and have been widely applied in various fields such as text-to-image (T2I), image-to-image (I2I), text-to-video (T2V), and 3D compositing. With the advent of tools such as DALL-E 2, Stable Diffusion, Midjourney, and Google Imagen, machine learning technology has truly entered every household, allowing users to easily generate various creative images simply by entering text commands.

[0057] Stable Diffusion models typically operate in a latent space to efficiently process high-dimensional data. Unlike standard diffusion models that operate directly in pixel space, Stable Diffusion utilizes the latent image encoding space, reducing memory requirements and computational costs while maintaining high fidelity in the generated images.

[0058] Through training, the diffusion model learns to predict the noise added during the forward diffusion process and reverses this diffusion process to remove noise from the image, generating a realistic image from a random seed. This method, known as parameterization, has been shown to improve training stability and sample quality. During inference, Stable Diffusion employs a sampling process that starts with random noise and iteratively applies the learned inverse process. This basic sampling process can be accelerated using techniques such as Denoising Diffusion Implicit Model (DDIM) or Pseudo-Linear Multistep (PLMS) methods, thereby reducing the number of sampling steps without significantly degrading sample quality. After the iterative denoising operation is complete, the actual generation process occurs in the latent space of a pre-trained variational autoencoder (VAE). Once the final latent representation is obtained, it is passed through the VAE decoder to generate the final high-resolution image.

[0059] In other words, a trained diffusion model can start with a random noisy image and some conditional information (e.g., text input from a user describing the desired image, pose vectors from motion capture, hand-drawn sketches, or other reference videos or images), and then gradually “denoise” the input signal through a learned iterative denoising procedure to finally obtain a realistic output image.

[0060] Diffusion models traditionally rely on the U-Net architecture, which encodes the input image layer by layer, transforming it into a low-dimensional representation before decoding it back into the original pixel space. Most diffusion models alternate between ResNet modules and visual transformer modules across layers. Furthermore, diffusion models based on pure visual transformers have emerged as an alternative to the U-Net architecture, demonstrating unique advantages in generating videos of varying lengths. Recent advancements have extended diffusion models to the field of video generation, bringing revolutionary possibilities to content creation. However, this also brings new challenges: ensuring spatiotemporal consistency, controlling computational costs, and generating long video sequences remain to be solved.

[0061] To achieve temporal consistency, the model needs to share information between frames. This is typically achieved using a 3D architecture or decomposition methods to reduce computational costs, while preprocessing features such as depth estimation can guide the denoising process for better results. Improvements are often made to the self-attention layers in the U-Net architecture, including the use of temporal attention, all-temporal attention, causal attention, and sparse causal attention. Each form differs in terms of computational requirements and motion capture capabilities.

[0062] In film production, video length has always been a challenging issue. While short clips are sufficient for trailers or commercials, they fall short of a full-length film. OpenAI's recently released Sora model is a benchmark in this field, capable of generating videos up to one minute long while maintaining visual fidelity and accurately reproducing user input. Another key challenge lies in achieving fine-grained control over content and motion synthesis. Human-based animation generation plays a central role in maintaining character consistency across scenes, thereby enhancing immersion and narrative coherence. By employing techniques based on human reference images and motion guidance, live-action animated videos can be directly generated, ensuring seamless character presentation throughout the film.

[0063] Low-rank adaptive learning (LoRA) is a technique designed to improve the efficiency of fine-tuning large-scale models, particularly suitable for transfer learning scenarios. Traditional fine-tuning methods require adjusting all parameters of the pre-trained model, which is not only computationally expensive but also prone to overfitting or catastrophic forgetting—that is, the model loses the knowledge accumulated during the pre-training phase. LoRA addresses these problems by introducing additional trainable parameters in the form of low-rank matrices while keeping the original model weights unchanged. This approach not only reduces computational costs and memory usage but also effectively preserves pre-trained knowledge, thereby minimizing the risk of catastrophic forgetting. Specifically, LoRA adds a low-rank factor to the original weight matrix of the denoising network, generating an updated weight matrix as a low-rank approximation superimposed on the original weights. This technique reduces the number of trainable parameters, making fine-tuning on small datasets easy and convenient. In practice, LoRA is often selectively applied to specific parts of the model (such as the attention layer in the Transformer architecture) to further reduce the computational and memory resources required for fine-tuning.

[0064] In AI-driven filmmaking, LoRA models play a crucial role because they can fine-tune stable diffusion models to capture specific actors and environments. In filmmaking, digital actors are key to generating content that describes scenes. However, collecting large datasets for each actor or environment is often impractical (or even impossible). LoRA technology allows for the fine-tuning of customized stable diffusion models to suit individual actors and environments, ensuring consistency in content generation.

[0065] Following breakthroughs in image synthesis, the application of stable diffusion models has expanded to conditional video generation. One approach involves transforming pre-trained large stable diffusion image models by converting the network into a 3D model (using dilated convolutional layers) and fine-tuning it based on video datasets. This method generates acceptable results for short video clips without requiring expensive video models to be trained from scratch. Another approach focuses on enhancing the ability of stable diffusion models to synthesize complete video sequences based on text prompts. These models generate consecutive frame sequences based on initial noisy input and text prompts. While this approach can produce high-quality short video clips, generating long, coherent videos remains challenging due to limitations in GPU memory. Some commercial products (such as Sora, Runway Gen3, Kling, and Luma) have attempted to address the temporal coherence issue, but their generated content remains primarily limited to short video clips rather than long takes or full-length films.

[0066] 2.3 Currently Available AI Content Generation Tools

[0067] Existing AI content generation tools can be divided into two categories: tools based on open-source models and proprietary models that are typically not publicly available. The main advantage of open-source models is that they provide developers with the opportunity to build custom tools and workflows based on them, enabling customized development according to the specific needs of film production.

[0068] Popular AI-generated content creation tools include, but are not limited to, the following: Midjourney is an advanced AI program that generates images using natural language descriptions. Users access the program through a Discord bot, inputting prompts to obtain image sets, enabling rapid prototyping. • OpenAI’s DALL-E 2 and its successor DALL-E 3, DALL-E 3 is built on ChatGPT, which allows users to use ChatGPT for brainstorming and hint refinement. Stable Diffusion is a similar open-source diffusion model released by Stability AI, comprising Stable Video Diffusion for video generation and Stable Video 3D for creating 3D videos. The source code and model weights are publicly available, allowing users to use and fine-tune the model, and develop various extensions and improvements. Furthermore, this open-source approach allows for customization for specific needs, such as those required in film production workflows. Numerous additional tools have been developed based on the Stable Diffusion model, including graphical interface tools such as Stable Diffusion WebUI and ComfyUI, providing user-friendly interfaces for designing and executing Stable Diffusion processes. ComfyUI offers great flexibility through a graphical / node / flowchart approach. These tools facilitate the integration of various additional features, such as retouching, super-resolution, and various generation guide techniques, giving users greater control over the generation of images or videos. Runway is a tool that supports AI video generation. Its underlying video diffusion model generates new videos based on provided structural and content information. Structural consistency is maintained through depth estimation, while content is controlled by image or natural language cues. Runway allows users to enrich cinematic experiences by incorporating horizontal and vertical motion, camera scrolling and zoom effects into animations. Runway also provides a motion brush to add dynamic effects to specific areas of animated scenes. Leonardo.ai offers video and image generation based on text or image input, real-time canvas editing, 3D texture generation, one-click video asset creation, custom model training, and the ability to use negative prompts to guide the generation process. Pika is a free AI tool that can generate videos based on text or image prompts. OpenAI's Sora is a new, state-of-the-art diffusion model that uses a transformer architecture to generate videos up to one minute long while maintaining visual quality and adherence to user cues. Sora is not yet publicly available, but OpenAI has granted access to a small group of professional artists and filmmakers to see what they are capable of creating. Viggle is an AI engine that automatically creates 3D character videos, offering options for text-to-character and text-to-motion animation, and generating character animations based on input images and guide motion videos.

[0069] However, existing tools remain inadequate, ineffective, or limited in addressing key challenges such as consistency, operability, and scalability, while also presenting problems with human actions and interactions. 3. Challenges in controllability

[0070] The concept of "controllability" in filmmaking refers to a director's control and precision over elements such as pacing, visual style, tone, camera angles, and actor performances. While advanced AI tools like Sora and Kling can quickly generate video content, existing tools often lack the fine-tuning required for high-quality filmmaking—a creative endeavor demanding creative flexibility, originality, and real-time decision-making. Regarding creative control, existing AI tools typically offer pre-made templates and automated processes, which limits the diversity of artistic expression, the realization of specific artistic visions, and the conveyance of subtle narrative nuances. Furthermore, filmmaking often requires real-time adjustments based on plot development or actor performances. However, the probabilistic nature of pre-set AI algorithms often hinders real-time adjustments aimed at capturing emotional impact or driving the story forward.

[0071] Camera control is a crucial element in narrative creation, but it also presents significant challenges when using AI tools in filmmaking. If directors cannot precisely control camera movement, angles, and composition, they will struggle to fully realize their creative vision, ultimately impacting the film's quality and expressiveness. Existing AI tools struggle to reproduce complex camera techniques such as tracking shots, zooms, or handheld shooting, which are essential for conveying narrative elements like emotional tension. Current AI tools are often not advanced enough to accurately reproduce these techniques, resulting in films that lose their intended effect or subtle emotional nuances. Furthermore, AI directors are often limited by general or preset camera settings, making it difficult to customize visual narratives according to specific narrative needs. Even more challenging is the complex sequence of shots, such as action scenes, which often require rapid editing, multi-angle switching, and precise timing—requirements that current AI tools struggle to effectively fulfill, leading to insufficient scene impact or poor visual coherence.

[0072] The AI-powered film production workflow platform described in this article addresses these limitations by introducing a modular approach, breaking down the production process into independent components, each managed by a dedicated AI model. This method allows for fine-grained control over narrative pacing, visual style, and camera work, more closely resembling traditional filmmaking practices.

[0073] This disclosed AI filmmaking framework enhances creative control by providing customizable AI models applicable to various production stages, including character animation, background generation, and scene composition. LoRA models can be used for actor digitization, while 3D scene reconstruction technology enables the digital representation of the background environment. Compared to rigid, template-based creative methods, this flexibility allows for more nuanced narrative expression and greater artistic tension. The AI ​​filmmaking platform features real-time adjustment capabilities, allowing directors to dynamically modify scenes based on plot development or actor performances. This feature provides the necessary responsiveness to capture the evolution of the film's emotional world.

[0074] The core advantage of the AI ​​filmmaking workflow lies in its integration of a virtual camera system within a digital 3D environment. Through 3D scene reconstruction technology, directors can precisely control the movement and angle of the virtual camera, enabling complex shot execution while ensuring narrative continuity in long takes. To guarantee consistency in extended content, the AI ​​filmmaking workflow breaks down scenes into manageable units and leverages specialized AI models to maintain visual and narrative coherence. This innovative approach overcomes the memory limitations of traditional diffused video models, making it possible to generate longer, more coherent shot sequences.

[0075] By addressing these core challenges, the AI ​​filmmaking workflow platform of this application enhances the controllability of AI-assisted filmmaking and paves the way for complex, high-quality works that can rival traditional filmmaking techniques. 4. Digitalization of Collaborative Networks

[0076] By digitizing all 3D elements, such as human characters and backgrounds, filmmakers can overcome many of the limitations of camera control associated with AI tools. This digital approach offers greater flexibility, precision, and creativity in the cinematography process, enabling complex and innovative visual storytelling techniques. Simultaneously, it enhances collaboration between AI and human filmmakers, elevating the overall quality and impact of the film.

[0077] Current AI tools often fall short in maintaining consistency across different scenes, especially when multiple artists are involved in the creation. Each creator typically focuses on producing an independent scene, but ensuring that the same actor's appearance or environmental elements remain consistent across different scenes is extremely difficult. This situation often leads to a linear workflow that relies on previous scenes to maintain coherence, thus significantly slowing down the creative process.

[0078] In a fully digital 3D environment, virtual cameras can be freely placed anywhere in the scene, providing filmmakers with complete creative freedom to experiment with different angles, movement trajectories, and compositions. Unlike physical cameras, virtual cameras are not limited by space, enabling more creative and complex shots. The digital environment supports highly precise and fluid camera movement, allowing for easy fine-tuning and animation. This allows directors or AI tools to create dynamic and fluid shots that are difficult to achieve with traditional shooting techniques. In digital scenes, adjustments to camera angles, lighting, or object positions can be previewed and modified in real time, allowing filmmakers to continuously optimize shot effects without spending time and money on on-location reshoots.

[0079] Digital 3D environments also enable the creation of complex and innovative camera techniques that are often difficult or impossible to achieve in the real world. For example, cameras can seamlessly penetrate walls, change shooting angles, or track motion in ways that defy physical limitations. Action scenes that typically require complex shot arrangements can be precisely controlled in digital environments, ensuring that every movement is captured perfectly. This precision is also reflected in special effects production, where cameras can interact with CGI elements in a highly controllable manner.

[0080] In a 3D digital environment, camera settings such as focal length, depth of field, and lighting can be kept consistent across different scenes, ensuring continuity of visual style and quality, which is difficult to achieve with traditional cameras, especially when shooting in multiple locations or at different times.

[0081] Once all elements are digitized, AI tools can analyze the 3D environment and enhance the narrative through intelligent suggestions or automatically generated camera movements. AI optimizes shooting angles based on scene composition, character actions, and lighting effects, ultimately delivering a more coherent and visually impactful film. Directors can also set specific camera behavior instructions for the AI, such as focusing on the protagonist during emotional climaxes or maintaining a wide-angle shot in action scenes. These preset behaviors ensure that the AI's decisions always align with the director's creative intentions.

[0082] Digital 3D elements also eliminate physical constraints such as location accessibility, lighting conditions, or equipment limitations, allowing the creation of scenes that would be logistically challenging or too costly to film in the real world. 5. AI-assisted film production framework

[0083] Figure 1 An algorithmic diagram of an AI-based film production workflow 100 is shown, representing current state-of-the-art technology and detailing the three main steps from script input to final film output. Once the script 110 is finalized, artists create storyboards using one or more AI-assisted storyboarding tools 120 (e.g., Midjourney, DALL-E, etc.). These AI tools help create storyboards that define key scenes, camera angles, character movements, and other important elements of the production process. AI analyzes the script 110, identifying important scenes, actions, dialogue, and emotions to generate a coherent visual flow for the story. AI also optimizes the layout and composition of each shot, adhering to various film principles such as the rule of thirds, depth of field, and focus. Furthermore, AI helps position characters within the frame, suggests their movements, and simulates motion based on the script or predefined behavioral models. AI also suggests camera angles and movements, determining the optimal positions, movements, and perspectives for capturing actions. The director then reviews and approves these storyboards before production begins.

[0084] A key step in storyboarding is breaking down the script into individual visual segments or "shots" to be filmed. Each shot typically lasts a few seconds, usually less than 10 seconds. This process ensures that every moment of the script is visually presented in a way that aligns with the film's narrative and artistic vision. AI-assisted storyboarding 120 can output one or more images for the storyboard, along with one or more associated cues, which can be understood as one or more annotations to the storyboard, hereinafter referred to as images / cues 125. Figure 1In the example, the steps following the storyboard are performed within the context of each specific shot. Each shot begins with a 2D keyframe, typically generated by an AI tool (e.g., Midjourney, DALL-E, etc.). With this keyframe, one or more AI animation tools 130 (e.g., Luma, Runway, etc.) can animate the image into video 135 (e.g., a short video clip), which then enters the post-production module 140 to complete a typical AI video generation process. After the post-production process is complete, the post-production module 140 outputs a film 150.

[0085] However, as mentioned above, current advanced AI filmmaking tools face several challenges, such as maintaining consistency between generated characters and backgrounds, controlling camera movement, and creating animations of complex character interactions or dynamic actions. To address these issues, this paper presents a novel AI-assisted filmmaking framework that integrates four key technological approaches into a unified AI filmmaking workflow:

[0086] (1) Digitizing all content in a 3D scene: This process involves converting a single or set of 2D images or videos (whether captured from the physical world or generated by AI) into a 3D digital space, a concept commonly found in animated film production. However, the purpose and techniques used here differ. In traditional animated film production, every object and detail must be meticulously crafted, making 3D modeling expensive. However, in the new AI-assisted filmmaking framework described here, the focus is on capturing the 3D structure of scenes, objects, and layouts. The primary goal is to provide AI with workflow guidance through depth maps or edge maps, enabling precise camera control settings when generating images and videos. Therefore, digitization can be achieved at a relatively low cost by utilizing methods such as Gaussian point clouds or monocular depth estimation, or by reusing or revising existing 3D models. On the other hand, by using LoRA models, faces can maintain consistency across multiple shots.

[0087] (2) Capturing digitized humans using 2D cameras: This technology is similar to how actors perform in front of a green screen during live filming. By recording actors' facial expressions and body movements, the captured data can be used to replace AI-generated faces with more realistic human features, or to map actors' movements onto AI-generated characters through style transfer. This addresses the challenges AI faces in mimicking complex physical character interactions and expressing human emotions.

[0088] (3) Better AI Animation Control: To address consistency and camera control issues, the enhanced AI animation process utilizes multiple cues from the digitization process to more effectively animate keyframes. These cues provide guidance for camera setup, narrative, and defining the style or attributes of specific characters or background elements. By integrating these diverse cues, filmmakers gain greater control over the AI ​​pipeline.

[0089] (4) AI-assisted compositing of complex scenes: This process involves managing multiple virtual cameras simultaneously, capturing elements from different spaces, such as simulating virtual environments and recording physical locations, and seamlessly combining them into a single frame. Visual effects (VFX) can also be used to create realistic or fantasy environments, characters, and effects that are difficult, costly, or impossible to achieve with traditional technologies.

[0090] Now for reference Figure 2 , Figure 3 and Figure 4 Various exemplary embodiments of this disclosure can be enhanced by adding various new functionalities (represented by corresponding computer-implemented "modules" in the figures and the following description). Figure 1 The existing AI filmmaking process 100 is shown. As further described below, the various operations of the methods described in this specification can be implemented using hardware, software, or a combination of both.

[0091] Figure 2 This is an algorithm diagram 200 of the first exemplary embodiment of an AI-based film production workflow provided in this disclosure. In addition to the AI-assisted storyboard module 220, the AI ​​animation module 230, and the post-production module 240, Figure 2 The AI ​​film production workflow 200 also includes a 3D digitization module 221, a virtual camera controller 224, and an AI-assisted compositing module 236.

[0092] like Figure 2 As shown, script 210 (e.g., one or more of its images) is input into AI-assisted storyboard module 220. Based on script 210 (e.g., its images), AI-assisted storyboard module 220 outputs guides 221 to 3D digitization module 222 and images / cues 225 to AI animation module 230.

[0093] Based on script 210 (e.g., its images) and guidelines 221, 3D digitization module 222 generates 3D environment 223 (i.e., a digitized 3D scene representation, including background, foreground, objects, people, etc.) and outputs 3D environment 223 to virtual camera controller 224. Virtual camera controller 224 receives 3D environment 223 (digitized 3D scene representation) from 3D digitization module 222 and outputs cues 229A related to 3D environment 223 to AI animation module 230.

[0094] In this exemplary embodiment, the AI ​​animation module 230 generates a first video 235 (Video 1) based on images / cues 225 from the AI-assisted storyboard module 220 and cues 229A related to the 3D environment 223 from the virtual camera controller 224. The AI ​​animation module 230 outputs the first video 235 (Video 1) to the AI-assisted synthesis module 236 for further processing.

[0095] Meanwhile, the 3D digitization module 222 can also generate or acquire a LoRA model 234 (i.e., a digitized 3D character representation) based on the script 210 and guide 221, and output the LoRA model 234 to the AI-assisted compositing module 236 for further processing with the first video 235 (Video 1).

[0096] In this exemplary embodiment, the AI-assisted compositing module 236 is configured to perform various compositing operations on a first video 235 (Video 1) from the AI ​​animation module 230 and a LoRA model 234 (digitized 3D character representation) from the 3D digitization module 222 to generate a composite video 237. In some embodiments, the AI-assisted compositing module 236 may also enhance the composite video 237 with visual effects (VFX), sound effects (SFX), and / or various combinations thereof.

[0097] The AI-assisted compositing module 236 outputs the composite video 237 to the post-production module 240 for processing. The post-production module 240 then performs one or more post-processing operations on the composite video 237 from the AI-assisted compositing module 236 to generate a movie 250.

[0098] according to Figure 2 In an exemplary implementation, the film 250 (i.e., the final version of the video or extended film clip) generated by the AI ​​filmmaking workflow 200 has multiple enhancements, which are achieved through the addition of a 3D digitization module 222, a virtual camera controller 224, and an AI-assisted compositing module 236, respectively. The 3D digitization module 222 uses a cost-effective method to convert various elements into a 3D environment 223 (a digitized 3D scene representation), enabling the virtual camera controller 224 to dynamically adjust camera settings in 3D space. The AI ​​animation module 230 receives multiple cues to effectively animate keyframes. These cues may include narrative input (e.g., images / cues 225) from the AI-assisted storyboard module 220, and camera settings (i.e., cues 229A) from the virtual camera controller 224.

[0099] The integration of these diverse inputs significantly enhances filmmakers' control over the AI ​​pipeline, marking a key innovation of the new framework described in this paper. Furthermore, Figure 2 It also demonstrates how the AI-assisted compositing module 236 can blend visual elements from different sources into different depth layers within a single frame or sequence. This includes videos generated by the AI ​​animation module 230 and elements created using the LoRA model 234 (Digital 3D Character Representation) from the 3D digitization process.

[0100] therefore, Figure 2 The AI-based film production workflow 200 shown provides the following functions for directors and artists: (1) allowing directors to determine the camera angle and movement for each shot using depth maps from a virtual camera in digitized 3D space, thanks to the 3D digitization module 222 and the virtual camera controller 224; and (2) ensuring consistency of characters and backgrounds across multiple shots, achieved using a LoRA model 234 trained on 3D digitization process data. This enables consistent reconstruction of objects and scene integration via the AI-assisted compositing module 236.

[0101] However, when Figure 2 The 3D digitization module 222, virtual camera controller 224, and AI animation tool 230 described herein cannot meet the director's specific needs, such as complex human behavior or interaction. In such cases, a 2D camera capture module, such as... Figure 3 As shown. In some example embodiments, this 2D camera capture module can be used in conjunction with the 3D digitization module 222. However, it should be understood that in other example embodiments, the 2D camera capture module can be used as an alternative to the 3D digitization module 222, depending on the director's needs and the appropriate 3D digitization mode or the 2D camera capture mode described below selected to obtain the desired results.

[0102] Figure 3 An algorithm diagram of an AI-based film production workflow 300, as shown in the second exemplary embodiment provided in this disclosure, is illustrated. In addition to the AI-assisted storyboard module 320, the AI ​​animation module 330, and the post-production module 340, Figure 3 The AI ​​film production workflow 300 also includes a 2D camera capture module 326, a visual cue extraction module 328, a video processing module 332, and an AI-assisted compositing module 336.

[0103] like Figure 3As shown, script 310 (e.g., one or more images) is input into AI-assisted storyboard module 320. Based on script 310 (e.g., its images), AI-assisted storyboard module 320 outputs guide 321 to 2D camera capture module 326 and outputs images / cues 325 to AI animation module 330 respectively.

[0104] Based on script 310 (e.g., its images) and guidelines 321, 2D camera capture module 326 generates a second video 327 (Video 2) and outputs the second video 327 to visual cue extraction module 328. Visual cue extraction module 328 receives the second video 327 (Video 2) from 2D camera capture module 326 and outputs cues 329B related to the second video 327 to AI animation module 330.

[0105] In this example embodiment, the AI ​​animation module 330 generates a first video 335 (Video 1) based not only on the image / clue 325 received from the AI-assisted storyboard module 320, but also on a cue 329B related to the second video 327 (Video 2) received from the visual cue extraction module 328. The AI ​​animation module 330 outputs the first video 335 (Video 1) to the AI-assisted synthesis module 336 for further processing.

[0106] Simultaneously, the visual cue extraction module 328 can also generate cue 331 based on the second video 327 (Video 2) from the 2D camera capture module 326, and output cue 331 to the video processing module 332. The video processing module 332 receives the second video 327 (Video 2) from the 2D camera capture module 326 and the cue 331 from the visual cue extraction module 328 as input, and generates a third video 333 (Video 3) based on the second video 327 (Video 2) and the cue 331. The video processing module 332 outputs the third video (Video 3) to the AI-assisted synthesis module 336 for further processing with the first video 335 (Video 1) from the AI ​​animation module 330.

[0107] In this example embodiment, the AI-assisted compositing module 336 is configured to perform various compositing operations on a first video 335 (Video 1) from the AI ​​animation module 330 and a third video 333 (Video 3) from the video processing module 332 to generate a composite video 337. In some embodiments, the AI-assisted compositing module 336 may also enhance the composite video 337 through visual effects (VFX), sound effects (SFX), and / or various combinations thereof.

[0108] The AI-assisted compositing module 336 outputs the composite video 337 to the post-production module 340 for processing. Subsequently, the post-production module 340 performs one or more post-processing operations on the composite video 337 from the AI-assisted compositing module 336 to generate the film 350.

[0109] according to Figure 3 In an exemplary embodiment, the film 350 (i.e., the final version or extended segment of the video) produced by the AI ​​filmmaking workflow 300 has various enhancements enabled by the addition of a 2D camera capture module 326, a visual cue extraction module 328, a video processing module 332, and an AI-assisted compositing module 336. The visual cue extraction module 328 extracts visual cues, such as human poses, skeletal movements, and facial expressions, from captured images or videos to assist the AI ​​animation module 330. The AI ​​animation module 330 receives multiple cues to effectively animate keyframes. These cues may include narrative input (e.g., image / cue 325) from the AI-assisted storyboard module 320, and style or attribute definitions (i.e., cue 329B) for specific characters or background elements from the visual cue extraction module 328.

[0110] The integration of these diverse inputs significantly enhances filmmakers' control over the AI ​​pipeline, marking a key innovation of the new framework described in this paper. Furthermore, Figure 3 This demonstrates how the AI-assisted synthesis module 336 blends visual elements from various sources into different depth layers within a single frame or sequence. This includes video generated by the AI ​​animation module 330 and components from the 2D camera capture module 326 that have been processed by the video processing module 332.

[0111] therefore, Figure 3 The AI ​​film production workflow 300 shown provides a new framework that introduces the following capabilities to directors and artists: (1) enabling directors to modify characters’ faces, clothing, hairstyles, poses and movements through cues, thanks to the visual cue extraction module 328 that supports the AI ​​animation module 330; (2) allowing directors to use the 2D camera capture module 326 to capture human performances or interactions that cannot be synthesized by AI, and seamlessly integrate them into the digital space through the visual cue extraction module 328, the video processing module 332 and the AI-assisted compositing module 336.

[0112] Furthermore, the exemplary embodiments disclosed herein are not limited to those described above. Figure 2 and Figure 3 Either of the two different solutions described herein. To better meet the director's needs, Figure 2 AI Film Production Workflow 200 and Figure 3The various aspects of the AI ​​filmmaking workflow 300 can be integrated, and the aforementioned functions and technologies can be combined in a variety of ways, as follows: Figure 4 This provides a more comprehensive solution than using any of the above embodiments alone.

[0113] Figure 4 This is an algorithm diagram of an AI-based film production workflow 400, which is a third exemplary embodiment provided in this disclosure. In addition to the AI-assisted storyboard module 420, the AI ​​animation module 430, and the post-production module 440, Figure 4 The AI ​​film production workflow 400 also includes a 3D digitization module 422, a virtual camera controller 424, a 2D camera capture module 426, a visual cue extraction module 428, a video processing module 432, and an AI-assisted compositing module 436.

[0114] like Figure 4 As shown, script 410 (e.g., one or more of its images) is input into AI-assisted storyboard module 420. Based on script 410 (e.g., its images), AI-assisted storyboard module 420 outputs guide 421 to 3D digitization module 422 and outputs images / cues 425 to AI animation module 430 respectively.

[0115] Based on script 410 (e.g., its images) and guide 421, 3D digitization module 422 generates 3D environment 423 (i.e., a digitized 3D scene representation, including background, foreground, objects, characters, etc.) and outputs 3D environment 423 to virtual camera controller 424. Virtual camera controller 424 receives 3D environment 423 (digitized 3D scene representation) from 3D digitization module 422 as input and outputs cues 429A related to 3D environment 423 to AI animation module 430.

[0116] The AI ​​animation module 430 generates a first video 435 (Video 1) based not only on images / cues 425 received from the AI-assisted storyboard module 420, but also on cues 429A received from the virtual camera controller 424 associated with the 3D environment 423. The AI ​​animation module 430 outputs the first video 435 (Video 1) to the AI-assisted synthesis module 436 for further processing.

[0117] Meanwhile, the 3D digitization module 422 can also generate or acquire a LoRA model 434 (i.e., a digitized 3D character representation) based on the script 410 and guide 421, and output the LoRA model 434 to the AI-assisted synthesis module 436 for further processing with the first video 435 (Video 1).

[0118] In this example, the AI-assisted compositing module 436 is configured to perform various compositing operations on a first video 435 (Video 1) from the AI ​​animation module 430 and a LoRA model 434 (digitized 3D character representation) from the 3D digitization module 422 to generate a composite video 437. The AI-assisted compositing module 436 can also enhance the composite video 437 through visual effects (VFX), sound effects (SFX), and various combinations thereof.

[0119] However, in the 3D digitization module 422, the virtual camera control module 424, and the AI ​​animation tool 430 (as described above) Figure 2 When the 2D camera capture module (as shown above) cannot meet the director's specific needs, such as complex human behavior or interaction, it can be used. Figure 3 (As shown). In some example embodiments, the 2D camera capture module can be used in conjunction with the 3D digitization module 422. Furthermore, it should be understood that the 2D camera capture module can be used as an alternative to the 3D digitization module 422, depending on the director's needs and the applicability of the 3D digitization mode described above or the 2D camera capture mode described below, to obtain the desired results.

[0120] Therefore, in some additional or alternative example embodiments, the AI-assisted storyboard module 420 outputs guidance 421 to the 2D camera capture module 426 and outputs images / cues 425A to the AI ​​animation module 430. Based on the script 410 (e.g., its images) and guidance 421, the 2D camera capture module 426 generates a second video 427 (Video 2) and outputs it to the visual cue extraction module 428. The visual cue extraction module 428 receives the second video 427 (Video 2) from the 2D camera capture module 426 as input and outputs cues 429B related to the second video 427 (Video 2) to the AI ​​animation module 430.

[0121] In this example embodiment, the AI ​​animation module 430 generates a first video 435 (Video 1) based not only on the image / clue 425 received from the AI-assisted storyboard module 420, but also on the clue 429B related to the second video 427 (Video 2) received from the visual cue extraction module 428. The AI ​​animation module 430 outputs the first video 435 (Video 1) to the AI-assisted synthesis module 436 for further processing.

[0122] Simultaneously, the visual cue extraction module 428 can also generate cue 431 based on the second video 427 (Video 2) received from the 2D camera capture module 426, and output cue 431 to the video processing module 432. The video processing module 432 receives the second video 427 (Video 2) from the 2D camera capture module 426 and the cue 431 from the visual cue extraction module 428 as input, and generates a third video 433 (Video 3) based on the second video 427 (Video 2) and the cue 431. The video processing module 432 outputs the third video (Video 3) to the AI-assisted synthesis module 436 for further processing with the first video 435 (Video 1) from the AI ​​animation module 430.

[0123] In this additional or alternative example, the AI-assisted compositing module 436 is configured to perform various compositing operations on a first video 435 (Video 1) received from the AI ​​animation module 430 and a third video 433 (Video 3) received from the video processing module 432 to generate a composite video 437. In some embodiments, the AI-assisted compositing module 436 may also enhance the composite video 437 with visual effects (VFX), sound effects (SFX), and / or various combinations thereof.

[0124] In any of the above example embodiments, the AI-assisted compositing module 436 outputs the composite video 437 to the post-production module 440 for processing. The post-production module 440 then performs one or more post-processing operations on the composite video 437 from the AI-assisted compositing module 436 to generate a movie 450 (i.e., the final version of the video or an extended movie clip).

[0125] according to Figure 4 In an example embodiment, the film 450 generated by the AI-based film production workflow 400 has a variety of enhancements, which are achieved by the addition of a 3D digitization module 422, a virtual camera controller 424 and an AI-assisted compositing module 436, and / or by the addition of a 2D camera capture module 426, a visual cue extraction module 428, a video processing module 432 and an AI-assisted compositing module 436.

[0126] Specifically, the 3D digitization module 422 uses a cost-effective method to convert various elements into a 3D environment 423 (digitized 3D scene representation), allowing the virtual camera controller 424 to dynamically adjust camera settings in 3D space. Furthermore, the visual cue extraction module 428 extracts visual cues, such as human posture, skeletal movement, and facial expressions, from captured images or videos to assist the AI ​​animation module 430. The AI ​​animation module 430 receives multiple cues to effectively animate keyframes. These cues may include narrative input from the AI-assisted storyboard module 420 (e.g., images / cues 425) and camera settings from the virtual camera controller module 424 (i.e., cues 429A), as well as style or attribute definitions for specific characters or background elements from the visual cue extraction module 428 (i.e., cues 429B).

[0127] As mentioned earlier, the integration of these diverse inputs significantly enhances filmmakers' control over the AI ​​process, marking a key innovation of the new framework described in this paper. Furthermore, Figure 4 This demonstrates how the AI-assisted compositing module 436 blends visual elements from different sources into different depth layers of a single frame or sequence. This can include videos generated by the AI ​​animation module 430, elements created using the LoRA model 434 (Digital 3D Character Representation), and / or components from the 2D camera capture module 426 processed by the video processing module 432.

[0128] therefore, Figure 4 The AI ​​filmmaking workflow 400, as shown, provides a new framework that introduces the following capabilities for directors and artists: • Allows the director to determine the camera angle and movement for each shot by using a depth map from a virtual camera in digital 3D space, thanks to the 3D digitization module 422 and the virtual camera controller module 424; • Enables the director to modify the character's facial expressions, clothing, hairstyle, posture and movements by providing prompts, supported by the visual cue extraction module 428, which is input into the AI ​​animation module 430. • Ensure consistency between human characters and backgrounds across multiple shots by training the LoRA model 434 using 3D digitized process data, thereby consistently reconstructing and integrating objects in the scene through the AI-assisted compositing module 436; • Allows directors to use the 2D camera capture module 426 to capture human performances or interactions that cannot be synthesized by AI, and then seamlessly integrate them into the digital space through the visual cue extraction module 428, video processing module 432 and AI-assisted synthesis module 436.

[0129] therefore, Figure 4 The AI-based film production workflow 400 shown integrates data from... Figure 2 AI film production workflow 200 and Figure 3 The AI ​​film production workflow 300 better meets the needs of directors by combining the aforementioned functions and technologies in various ways to provide a more comprehensive solution to address the problems associated with existing AI film production workflows and tools.

[0130] Next, we will refer to Figures 5A-5B , Figures 6A-6D , Figures 7A-7B , Figures 8A-8B and Figures 9A-9C This section will explain certain aspects of the AI ​​film production workflow described above. Then, it will combine... Figure 10 and Figure 11 Describe the application of these AI-assisted digital technologies in implementing collaborative networks.

[0131] 5.1 Digitalization

[0132] Digitizing 3D space, especially when dealing with complex backgrounds, is a challenging task. However, it remains the most efficient method if a cost-effective solution is available to achieve the necessary flexibility for camera control. In most cases, when a person can be physically present, a 3D background model can be reconstructed using Gaussian sputtering techniques from a series of images captured on-site, such as... Figures 5A-5B As shown.

[0133] Figure 5A and Figure 5B The images show background images created using Gaussian point clouds, and then the same human elements are added to the 3D space, enabling flexible camera control in the 3D scene.

[0134] However, when it is impossible to have someone physically present, monocular depth estimation can be used to predict the depth of individual points in a two-dimensional image, such as... Figures 6A-6D As shown, these images can be 2D images captured by a camera or generated by AI. By training on a large dataset containing corresponding depth maps, this method allows for reasonable "2.5D" reconstructions, especially in enclosed spaces such as indoors.

[0135] Figures 6A-6D An example implementation of how to create a 2.5D image from a 2D image using monocular depth estimation is shown. Figure 6A It is a two-dimensional image before processing. Figure 6B These are regenerated two-dimensional images from different camera settings. Figure 6C It showcases a 3D depth model reconstructed under camera control. Figure 6D It showcases a 3D space reconstructed under camera control.

[0136] During the digitization process, faces are 3D scanned from different angles and with different expressions. These scans are translated into specialized models to ensure consistent appearance across shots. Specifically, low-rank adaptation (LoRA) models (e.g., references 234, 434) are fine-tuned to capture different character features and expressions. By adjusting the weights of the pre-trained model using low-rank decomposition, LoRA can adapt to different scene conditions while maintaining visual continuity throughout the film. This approach consistently preserves facial features and expressions without requiring a complete retraining of the network. Figure 7A and 7B Each example showcases a digital human appearance to demonstrate this capability.

[0137] Figure 7A This demonstrates the LoRA model (e.g.) Figure 2 LoRA model 234 or Figure 4 The image generated by the LoRA model 434, arranged in rows and columns. Figure 7A In the example, a single LoRA model generated four rows of faces—a smiling face at age 10, a sad face at age 15, a smiling face at age 40, and a smiling male face at age 50, showing different genders. Figure 7A This demonstrates the LoRA model's ability to maintain character consistency. Using the LoRA model, users can adjust age, gender, hairstyle, clothing, facial expressions, and more to customize the generated character. Figure 7A As shown.

[0138] Figure 7B A set of images is shown to demonstrate how to use the LoRA model to replace the faces of human characters with digital 3D models while preserving the original movements and expressions. Figure 7B In the example, the left panel 701 displays the original image (the portrait of the first person), the right panel 703 displays the face to be applied by the LoRA model (the face of the second person), and the middle panel 705 displays a modified version of the original image that integrates the LoRA model's face replacement. Therefore, the LoRA model can replace the first person's face with a second person's face that is different from the first person's, while maintaining consistency in overall appearance.

[0139] 5.2 Style Transfer

[0140] When human actions in AI-generated scenes fail to meet the director's vision, the AI ​​film production platform offers a hybrid approach, using human performance to guide and refine the AI-generated content. The platform implements a strategy that combines human performance capture with AI-assisted optimization.

[0141] If a director is not satisfied with the human movements or actions generated by AI tools, an effective alternative is to have one or more actors perform the required movements as a reference. These performances are recorded, capturing precise postures and movements, serving as a blueprint for AI content generation. The recorded postures and movements can be used to guide the AI ​​in replicating these actions, such as... Figures 8A-8B As shown. This technology allows for more detailed action sequences that meet the director's requirements.

[0142] Figures 8A-8B It demonstrates motion guidance cues extracted from performances captured by cameras. Figure 8A It showed the performance of the actors on site. Figure 8B The video shows AI-generated characters that use the same style of movement as the live actors.

[0143] By using recorded human performances as a foundation, AI tools can reconstruct and optimize scenes. This process allows for the integration of a director's specific vision with the capabilities of AI generation. Furthermore, this framework enhances various details to increase the director's control. For example, details such as facial expressions, clothing, and hairstyles can be refined and provided as cues to enhance the director's maneuverability. Captured facial expressions from actors can be used to guide AI-generated performances that are more realistic and emotionally resonant. Specific clothing and hairstyle designs can serve as cues to ensure that the AI-generated character perfectly matches the director's vision. Additionally, real-world references can be used to fine-tune the AI-generated backgrounds and props.

[0144] 5.3 Camera Control

[0145] As mentioned above, the digitization process involves reconstructing a 3D (or 2.5D) model of the environment, which can be derived from a variety of sources, including 2D AI-generated images, 2D camera captures, or 3D camera scans. Figure 2 and Figure 4 It provides visualizations of how the 3D digitization module creates a comprehensive 3D scene representation, which is then used in the camera control phase. In this 3D space, a "virtual camera" (such as...) Figure 6D (As shown) can simulate viewpoints from any location. These simulated viewpoints and their corresponding depth or edge maps are then input into the aforementioned AI tool, enabling the AI-assisted video content to be precisely aligned with the intended camera settings.

[0146] While some short video generation tools based on stable diffusion offer the ability to incorporate camera motion as a textual condition or provide pre-trained camera motion LoRA models, these methods often have significant limitations. The repeatability of continuous motion poses a major challenge; attempting to create new shots of the same scene from different angles often results in inconsistent additions or deletions of details. Resampling the random number seed until a satisfactory result is generated is a time-consuming process requiring multiple iterations.

[0147] In contrast, the model disclosed herein, which reconstructs 3D backgrounds using Gaussian point clouds, offers significant advantages. This method allows for consistent background representation while providing unrestricted camera movement capabilities. By leveraging this 3D reconstruction technique, the limitations of existing AI tools in maintaining scene consistency across different camera angles and movements can be overcome. This not only enhances the flexibility of shot composition but also significantly reduces the time and computational resources required to achieve the desired results. Therefore, the AI ​​filmmaking framework described herein can bridge the gap between the creative freedom required by filmmakers and the technological limitations of current AI-based video generation tools, providing a more powerful and efficient solution for the creation of dynamic multi-angle scenes in AI-assisted filmmaking.

[0148] 5.4 AI-assisted synthesis

[0149] If a director is dissatisfied with the AI-generated video of certain elements during filmmaking, the AI-assisted compositing module offers an alternative solution. This module can replace foreground or background objects and apply style transfer to enhance both. For example, if AI-generated character movements lack realism, the AI-assisted compositing module can integrate the performance of real actors into the scene, such as... Figures 9A-9C As shown.

[0150] Figures 9A-9C The scene depicts a director deciding to composite the performance of real actors (green screen 902), a stable background (AI image 904), and a moving window scene (background 906) into a single image (composite 908). Figure 9A Multiple layers of objects (902, 904, 906) are composited into a single image (908). Figure 9B The first viewpoint 909A shows the compositing process, where multiple layers of objects are composited into a single frame (908). Figure 9C This shows the second perspective 909B of the compositing process, where multiple layers of objects are also composited into a single frame (908).

[0151] Therefore, a typical compositing process, according to exemplary embodiments of this disclosure, merges multiple visual elements (such as video, images, and graphics) into a unified frame or sequence. (Refer to the above...) Figure 2, Figure 3 and Figure 4 The various input sources for the AI-assisted synthesis module may include any of the following: (1) one or more videos (at different depth levels) generated through the AI ​​animation process, supported by the 3D digitization module and the virtual camera control module (see [link to documentation]). Figure 2 and Figure 4 (2) Lenses captured by a 2D camera are processed by a visual cue extraction module and a video processing module to create an alpha channel and a mask for integrating visible objects into the scene (see [link]). Figure 3 and Figure 4 (3) 3D characters created by the 3D digitization module (see Figure 2 and Figure 4 (4) Graphical elements generated through VFX processes (see Figure 2 , Figure 3 , Figure 4 (any one of them).

[0152] According to the example embodiments described above, the AI-assisted compositing module synchronizes multiple or all of these elements with the main camera (such as a camera used by a 2D camera capture module) to ensure that the virtual background, newly created 3D characters, and visual effects are seamlessly integrated with the live-action footage when creating the composite video.

[0153] 5.5 Post-production

[0154] Converting AI-generated footage (such as 237, 337, and 437) into cinematic-quality films (such as 250, 350, and 450) requires extensive post-processing. Current limitations of AI models often result in inconsistencies in style, lighting, and resolution among the generated short video clips. While AI filmmaking frameworks have successfully addressed issues of overall coherence and narrative continuity, achieving a unified style, composition, and lighting throughout the entire production process remains a significant challenge.

[0155] To address stylistic inconsistencies between scenes, recent advancements in video style transfer have demonstrated promising results. These techniques allow for the application of a consistent style throughout the video, based on a small number of stylized keyframes. This approach enables filmmakers to maintain a consistent visual aesthetic throughout production, thereby enhancing the overall quality and artistic vision of the film.

[0156] Lighting plays a crucial role in cinematic storytelling, but traditionally it has required labor-intensive manual adjustments using existing tools such as Adobe After Effects. For example, lighting conditions must be manually recalibrated when an object moves from one background to another. AI-based relighting technologies are changing this process, providing a fast, automated solution for still images that retains the emotional depth and realism of professional lighting while significantly reducing production time and costs. Leveraging a digital framework, current AI filmmaking workflows treat lighting as an actionable element in video. AI can extract lighting from the background and seamlessly apply it to foreground objects. For the first time, artists can “paint” lighting effects as needed, achieving effects previously unattainable through color gamut and bit depth. With these advancements, artists have unprecedented control over how they use lighting to tell stories.

[0157] Generating high-resolution video content presents a particular challenge to AI systems, as models like Stable Diffusion require significant amounts of GPU memory. This limitation complicates the production of industry-standard long-format 4K video. A promising solution is to first generate low-resolution content and then apply AI-driven super-resolution techniques. These models intelligently enhance video resolution and texture details, making it possible to create high-quality content within the constraints of current GPU technology.

[0158] Integrating these advanced post-processing technologies into a comprehensive framework, such as the AI-based filmmaking workflow platform disclosed herein, marks a significant step toward narrowing the quality gap between AI-generated film content and traditionally produced film content, and has the potential to transform the landscape of modern filmmaking.

[0159] An AI-based film production workflow platform introduces a new perspective on AI-driven filmmaking by breaking down film elements and processing them using specialized models, rather than relying on a single model to generate the entire film. This approach effectively addresses many challenges faced by existing models, including issues of maneuverability, video consistency, and scalability. 6. Creative Collaboration Network

[0160] 6.1 Digitalization of Collaborative Networks

[0161] As mentioned earlier, current AI tools often struggle with maintaining consistency across different scenes, especially when multiple artists are involved. Each artist typically focuses on the production of a single scene, but ensuring consistency in actor appearances or environments across scenes can become problematic. This often results in a linear, wait-dependent workflow where artists must review previous scenes to maintain coherence, significantly slowing down the production process.

[0162] This disclosed AI-based filmmaking framework addresses these challenges through a comprehensive digital approach. The process begins with the digitization of key elements such as environment, actors, scene style, and lighting models. Figure 10 The overall framework of the collaborative network is presented, and will be explained in more detail later. Crucially, these digitization processes are independent of each other and can be executed in parallel, enabling artists to work separately. This structure also allows for the direct integration of video footage of real actors into AI-based workflows. This initial setup lays the foundation for consistent, collaborative AI filmmaking.

[0163] Once digitized, this disclosed AI-based film production workflow platform and collaborative network allows artists to independently create individual scenes or shots according to the director's requirements, without strict sequential dependencies. Using digitized LoRA models of actors and 3D reconstructions of background environments ensures consistency between different scenes, even if they are produced by different artists or teams. This approach significantly reduces the need for artists to wait for each other to maintain coherence, as the models themselves provide the necessary consistency.

[0164] Figure 10 This is a conceptual diagram of the overall structure of an AI-based film creation workflow within a collaborative network, one aspect of which is disclosed. Figure 10 This demonstrates how to create different scenes (e.g., Scene 1, Scene 2, etc.) that share similar actor elements (e.g., digital actors 1-N) and background elements (digital backgrounds 1-M), while incorporating unique objects (e.g., digital objects 1, 2, etc.) or other style variations (e.g., facial styles, body styles, style models 1, 2, etc.), lighting (e.g., lighting settings 1-N), and / or camera angles (e.g., camera settings 1-N). For example, background motion can be represented through Gaussian stippling, stable diffusion, green screen, etc.; body style can be represented through skeletal animation, stable diffusion, motion capture, etc.; facial style can be represented through emotion control, synchronized lip-syncing, age control, etc. This design fosters a true collaborative network, theoretically allowing each scene to be produced largely in parallel, significantly improving efficiency and reducing production time.

[0165] like Figure 10As shown, the exemplary implementation of this disclosure provides an innovative AI-based filmmaking method that utilizes a unique digital approach. The framework employs LoRA fine-tuning technology to digitize actors, ensuring consistent character performance. Background elements are reconstructed into 3D form using advanced AI models, creating fully controllable digital scene versions. The framework integrates style transfer, relighting, and multiple post-processing models applicable to all scenes to achieve visual aesthetic unity. Once these fundamental components are trained and fine-tuned, multiple artists can simultaneously utilize these digital assets to create the desired scenes. This parallel workflow fosters a collaborative environment, significantly reducing overall filmmaking time and increasing creative flexibility.

[0166] Therefore, by digitizing all 3D elements, such as human characters and backgrounds, filmmakers can overcome many of the limitations of cinematic control associated with AI tools. This digital approach offers greater flexibility, precision, and creativity, enabling complex and innovative visual storytelling techniques in the cinematography process. It also enhances collaboration between AI and human filmmakers, improving the overall quality and impact of films.

[0167] 6.2 Expanding AI Film Production Through Collaborative Networks

[0168] Due to current limitations of AI technology, traditional 2D cinematography combined with live actors, as well as remote 3D digital support for specific landscapes or outdoor scenes, remain valuable. Meanwhile, AI filmmaking allows for remote collaboration among artists, emphasizing the importance of exploring creative collaboration networks.

[0169] Figure 11 This is a structural diagram of the publicly available creative collaboration network 1100. (See diagram below.) Figure 11 As shown, the Creative Collaboration Network 1100 is designed with several key components to improve the efficiency and effectiveness of remote collaboration, including but not limited to: one or more digital collaboration tools 1110; one or more cloud-based asset management systems 1120; an AI-based film production workflow platform 1130; and a security and intellectual property management system 1140.

[0170] Digital Collaboration Tools 1110: These tools provide a virtual workspace where team members can communicate, brainstorm, and share ideas in real time. Platforms such as video conferencing, chat applications, and digital whiteboards are essential for maintaining a continuous flow of communication, allowing dynamic discussions and rapid decision-making, which is crucial in the creative process. For example, tools like Slack, Microsoft Teams, and Zoom integrate multiple communication methods (such as chat, video calls, and file sharing) into a single platform, making collaboration simpler and more efficient. These platforms also integrate with other digital tools to create seamless workflows, directly linking communication with project management, file storage, and more. Digital tools support real-time collaboration of documents, designs, and code, enabling multiple people to work simultaneously, reducing lengthy back-and-forth communication, and accelerating the collaboration process.

[0171] Cloud-based asset management systems (1120): Cloud systems are essential for organizing, storing, and sharing large amounts of digital assets such as scripts, storyboards, 3D models, and original footage. These systems enable teams to access and update assets from anywhere, ensuring everyone is using the latest materials. This not only streamlines workflows but also reduces the risk of version control issues and data loss. For example, cloud computing allows resources such as files, data, and tools to be centrally stored and accessed by collaborators anytime, anywhere. This facilitates the sharing of large datasets, software, and collaborative environments, which is crucial for complex projects such as software development, research, and the creative industries.

[0172] AI-based Film Production Workflow Platform 1130: The various AI tools (“modules”) of the aforementioned AI-based film production workflows (200, 300, 400) can be provided on the AI-based film production workflow platform 1130 and integrated into the network, enabling artists and filmmakers to access the functionality needed to complete the creative process. For example, artists can find tools to capture realistic elements (such as actors’ performances) and can access virtual production environments, allowing remote teams to direct and shoot scenes as if they were on set.

[0173] Security and Intellectual Property Management System 1140: To protect creative works and ensure respect for intellectual property rights, the network is equipped with advanced security protocols and digital rights management tools. This ensures the security of all shared assets and communications, maintaining the integrity and confidentiality of projects.

[0174] like Figure 11As shown, the creative collaboration network 1100 also includes a communication network 150 (e.g., wired, wireless, mobile, internet, etc.), one or more user devices 1160 (e.g., user devices 1-N), and one or more servers 1170 (e.g., servers 1-N), which are communicatively connected to the communication network 1150 to implement the techniques described herein. The communication network 150 enables directors and other participants to create, communicate, store, and share files, data, and information, as well as access the AI-based film production workflow platform 1130. In some example embodiments, the AI-based film production workflow platform 1130 (or at least a portion thereof) can be accessed (e.g., downloaded) and executed locally on user devices 1160. In other example embodiments, the AI-based film production workflow platform 1130 (or at least a portion thereof) can be remotely accessed through one or more servers 1170 and executed remotely on behalf of users on one or more servers 1170.

[0175] It should be understood that Figure 11 The one or more components and devices shown, as well as one or more aspects of the aforementioned graphics-related technologies, can be implemented in hardware (e.g., computers, mobile devices, tablets, etc.), including one or more processors (e.g., CPUs, GPUs, processors, microprocessors, etc.) and one or more memories (e.g., storage devices), or in software (e.g., applications, programs, instructions, algorithms, models, etc.), or a combination of hardware and software. Although in Figure 11 The items are displayed as separate boxes only for ease of explanation. It should be understood that in some examples, various features, tools, modules, components, functions, etc., may be provided or accessed together on the same computing device, or in other examples they may be separate and distributed across multiple devices.

[0176] These platforms, devices, components, and users together form a comprehensive network that supports the diverse and dynamic needs of AI-assisted filmmaking, enabling a more collaborative, flexible, and efficient filmmaking process.

[0177] Therefore, the aforementioned digital technologies can fundamentally transform collaborative networks, breaking down traditional barriers, improving communication, and integrating advanced technologies like artificial intelligence. This creates new opportunities for innovation, efficiency, and creativity, enabling teams to collaborate in a more dynamic, inclusive, and efficient manner. Digital platforms enable individuals and organizations from around the world to collaborate in real time, regardless of location. This global interconnectedness fosters diverse collaborations, bringing together different cultures, expertise, and perspectives, thereby driving innovation and creativity. With digital tools, collaboration is no longer limited by time zones or office hours. Teams can work asynchronously, passing tasks between time zones to keep projects on track.

[0178] Furthermore, AI tools can automate and optimize task allocation within collaborative networks, ensuring the right people are assigned the right tasks based on their skills, availability, and past performance. This makes collaboration more efficient and reduces bottlenecks. For example, AI-driven tools like ChatGPT can assist in brainstorming, content creation, data analysis, and more. These tools can act as collaborators, offering suggestions, automating routine tasks, and even generating new ideas, thereby extending the capabilities of human teams. Digitalization also enables teams to collaboratively analyze large datasets in real time using Google Analytics, Tableau, or custom machine learning models. This shared access to data insights drives data-driven decision-making and more effective collaboration. By analyzing team members' skills, work habits, and preferences, AI tools can help create personalized collaborative networks, ensuring team members are matched with tasks and collaborators that align with their strengths, leading to more efficient and satisfying collaboration.

[0179] Figure 12 This is a flowchart of step 1200 of the first method for AI-based filmmaking, which is publicly available.

[0180] In step 1220, the method includes performing AI-assisted storyboard creation, generating one or more images of the storyboard, one or more storyboard-related prompts, and script-based guidelines.

[0181] In step 1222, the method includes 3D digitization to generate 3D models of scenes based on guidelines generated from the script and AI-assisted storyboard production process.

[0182] In step 1224, the method includes performing virtual camera control to generate cues based on guidelines and the 3D model generated during the 3D digitization process.

[0183] In step 1230, the method includes performing an AI animation process to generate a first video using one or more images based on a storyboard and one or more cues generated by AI-assisted storyboard production, as well as cues generated by a 3D digitization process and virtual camera control.

[0184] In step 1236, the method includes performing AI-assisted synthesis to generate a synthesized video based on a first video generated by an AI animation process and a LoRA model generated by a 3D digitization process.

[0185] In step 1240, the method includes performing one or more post-production processes on the synthesized video generated by the AI-assisted compositing process to generate a complete film (or extended video clips).

[0186] Therefore, the first method 1200 can be considered as a "3D digitization model" related to the first example embodiment described above, which relates to the AI ​​film production workflow 200. The first method 1200 can use the aforementioned... Figure 11 The related creative collaboration network 1100 utilizes computing devices and components, including a combination of hardware and software, to realize various "modules" of the AI ​​filmmaking workflow platform.

[0187] Figure 13 This is a flowchart of the steps of the second method 1300 for AI-based filmmaking, which is publicly available.

[0188] In step 1320, the method includes performing AI-assisted storyboard creation to generate one or more images of the storyboard, one or more cues associated with the storyboard, and script-based guidelines.

[0189] In step 1326, the method includes performing 2D camera capture to generate a second video based on guidelines generated from the script and the AI-assisted storyboard production process.

[0190] In step 1328, the method includes performing visual cue extraction to generate cues based on a guide and a second video captured by a 2D camera.

[0191] In step 1330, the method includes performing AI animation to generate a first video using one or more images based on a storyboard and one or more cues generated by AI-assisted storyboard production, as well as cues generated based on a 2D camera capture and visual cue extraction process.

[0192] In step 1332, the method includes performing video processing to generate a third video based on a second video captured from a 2D camera and one or more visual cues from a visual cue extraction process.

[0193] In step 1336, the method includes performing AI-assisted synthesis to generate a synthesized video based on a first video from an AI animation process and a third video from video processing.

[0194] In step 1340, the method includes performing one or more post-production processes on the synthesized video from the AI-assisted compositing process to generate a complete film (or an extended video clip).

[0195] Therefore, the second method 1300 can be regarded as the same as the above reference. Figure 3 The first example embodiment relates to the "2D camera capture mode". The second method 1300 can utilize the aforementioned... Figure 11The creative collaboration network 1100 described herein is implemented using computing devices and components, which include a combination of hardware and software, as corresponding structures for the various “modules” of the AI ​​film production workflow platform.

[0196] Figure 14 This is a flowchart illustrating the steps of a third method 1400 for AI-based movie production according to a third exemplary embodiment of this disclosure.

[0197] In step 1420, the method includes performing AI-assisted storyboard creation to generate one or more images of the storyboard, one or more cues associated with the storyboard, and script-based guidelines.

[0198] In step 1422, the method includes performing 3D digitization to generate a 3D model of the scene based on the script and guidelines generated by the AI-assisted storyboarding process. In step 1424, the method includes performing virtual camera control to generate cues based on the guidelines and the 3D model generated by the 3D digitization process.

[0199] In step 1426, the method includes performing 2D camera capture to generate a second video based on the script and guidelines generated by the AI-assisted storyboarding process. In step 1428, the method includes performing visual cue extraction to generate a cue based on the guidelines and the second video generated by the 2D camera capture. In step 1432, the method includes performing video processing to generate a third video based on the second video generated by the 2D camera capture and one or more visual cues generated by the visual cue extraction process.

[0200] In step 1430, the method includes performing AI animation to generate a first video based on one or more images based on a storyboard and one or more cues generated by AI-assisted storyboard production, as well as cues generated by a 3D digitization process, virtual camera control, and 2D video capture and visual cue extraction process.

[0201] In step 1436, the method includes performing AI-assisted synthesis to generate a synthesized video based on a first video generated by an AI animation process, a LoRA model generated by a 3D digitization process, and a third video generated by video processing.

[0202] In step 1440, the method includes performing one or more post-production processes on the synthetic video generated by the AI-assisted compositing process to generate a complete film (or extended video clip).

[0203] Therefore, the third method 1400 and according to the above Figure 4 The third example embodiment relates to the AI ​​film production workflow 400. The third method 1400 can utilize the above... Figure 11 The creative collaboration network 1100 described herein is implemented using computing devices and components, which include a combination of hardware and software, as corresponding structures for various "modules" that enable an AI-based film production workflow platform.

[0204] exist Figure 14 In the examples, it should be understood that some operations in the 3D digitization mode (1422, 1424) can be performed, and additionally, or alternatively, some operations in the 2D camera capture mode (1426, 1428, 1432) can also be performed. These processes can occur simultaneously or at different times, and can be performed by one person (e.g., the director) or multiple different individuals. These processes can be performed for the same scene or only for a specific scene using either the 3D digitization mode or the 2D camera capture mode. Whether to use the 3D digitization mode, the 2D camera capture mode, or both at a specific time or in a specific scene or sequence will depend on the director's narrative needs and creative perspective, and which technology is more suitable for the specific situation.

[0205] In summary, this disclosure introduces a novel AI-based filmmaking framework designed to foster future creative collaboration networks, as illustrated in the accompanying diagrams. This AI-based filmmaking workflow platform has already been used to create groundbreaking AI-driven films such as the love story *Next Stop Paris* and the science fiction short film *Robot Message*. The framework's innovative approach to the digitization, decomposition, and combination of AI elements enables unprecedented levels of control and flexibility in the creative process, setting a new standard for AI-based filmmaking. As filmmakers increasingly adopt AI-based filmmaking frameworks and related AI technologies, creative collaboration networks will play a vital role in developing communities and the future of filmmaking, characterized by an unprecedented tenfold increase in efficiency in a short time.

[0206] It should be noted that the above-described exemplary embodiments are intended to be illustrative only and should not be construed as limiting the scope of this disclosure, the inventive concept, or the appended claims in any way.

Claims

1. A video production system, characterized by, The system comprises: an auxiliary storyboard module for generating one or more images of a storyboard, one or more cues associated with the storyboard, and a guide based on a script; a digitization module for generating a scene model based on the script and the guide; a virtual camera controller for controlling camera settings to generate cues associated with the scene model; an animation module for generating a first video based on the one or more images of the storyboard, the one or more cues associated with the storyboard, and the cues associated with the scene model; an auxiliary synthesis module for generating a synthesized video based on the first video generated by the animation module and a low-rank adaptation (LoRA) model.

2. The system of claim 1, wherein, The digitization module is further configured to generate the LoRA model based on the script and the guide, wherein the scene model is configured to represent a digitized environment and the LoRA model is configured to represent a digitized character.

3. The system of claim 1, wherein, The system further comprises: a camera capture module for generating a second video based on the script and the guide; and a visual cue extraction module for generating cues related to the second video.

4. The system of claim 3, wherein, The animation module is further configured to generate the first video based on the one or more images of the storyboard, the one or more cues associated with the storyboard, and the cues related to the second video.

5. The system of claim 4, wherein, The visual cue extraction module is further configured to generate one or more visual cues based on the second video; the system further comprises: a video processing module for generating a third video based on the second video and the one or more visual cues.

6. The system of claim 5, wherein, The auxiliary synthesis module is further configured to generate a synthesized video based on the first video generated by the animation module and the third video generated by the video processing module.

7. The system of claim 1, wherein, The auxiliary synthesis module is further configured to add visual effects, audio effects, or a combination thereof in the synthesized video.

8. The system of claim 1, wherein, The system further comprises: a post-production module for generating and outputting a complete movie based on the synthesized video generated by the auxiliary synthesis module.

9. The system of claim 8, wherein, The post-production module is configured to perform one or more post-processing operations on the synthesized video generated by the auxiliary synthesis module to generate a complete movie.

10. The system of claim 1, wherein, The auxiliary storyboard module is an artificial intelligence (AI)-based auxiliary storyboard module, the digitization module is a 3D digitization module, the animation module is an AI-based animation module, and the auxiliary synthesis module is an AI-based auxiliary synthesis module.

11. The system of claim 3, wherein, The camera capture module is a 2D camera capture module.

12. A video production method characterized by, The system comprises: generating, by an auxiliary storyboard module, one or more images of a storyboard, one or more cues associated with the storyboard, and a guide based on a script; generating, by a digitization module, a scene model based on the script and the guide; controlling, by a virtual camera controller, virtual camera settings to generate cues associated with the scene model; generating, by an animation module, a first video based on the one or more images of the storyboard, the one or more cues associated with the storyboard, and the cues associated with the scene model generated by controlling the virtual camera settings; and generating, by an auxiliary synthesis module, a synthesized video based on the first video generated by the animation module and a low-rank adaptation (LoRA) model. A synthetic video is generated based on the first video and a low-rank adaptive (LoRA) model by an assisted synthesis module.

13. The method of claim 12, wherein, The method further includes: The LoRA model is generated based on the script and the guide by the digitization module. The scene model is used to represent a digitized environment and the LoRA model is used to represent a digitized character.

14. The method of claim 12, wherein, The method further includes: A second video is generated based on the script and the guide by a camera capture module; and Visual cues are extracted related to the second video by a visual cue extraction module.

15. The method of claim 14, wherein, The method further includes: The first video is generated based on the one or more images of the storyboard and the one or more cues associated with the storyboard, and the cues related to the second video by an animation module.

16. The method of claim 15, wherein, The method further includes: One or more visual cues are generated based on the second video by the visual cue extraction module; and A third video is generated based on the second video and the one or more visual cues by a video processing module.

17. The method of claim 16, wherein, The method further includes: A synthetic video is generated based on the first video and the third video by the assisted synthesis module.

18. The method of claim 12, wherein, The method further includes: Visual effects, sound effects, or a combination of both are added in the synthetic video by the assisted synthesis module.

19. The method of claim 12, wherein, The method further includes: A complete movie is generated based on the synthetic video by a post-production module; and The complete movie is output for review.

20. The method of claim 19, wherein, The method further includes: One or more post-processing operations are performed to generate the complete movie by the post-production module.

21. The method of claim 12, wherein, The assisted storyboard module is an artificial intelligence (AI)-based assisted storyboard module, the digitization module is a 3D digitization module, the animation module is an AI-based animation module, and the assisted synthesis module is an AI-based assisted synthesis module.

22. The method of claim 14, wherein, The camera capture module is a 2D camera capture module.