Video generation method and device, terminal, electronic equipment and storage medium
By using intelligent agents for end-to-end intelligent scheduling and leveraging large-scale generative language models and Unreal Engine to automate the generation of animated videos, the problem of low production efficiency and unstable quality in animated video production is solved, achieving fast and efficient animation video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU BULLET FINGER UNIVERSE TECH CO LTD
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-14
AI Technical Summary
Current technologies rely on manual creation for animation video generation, resulting in insufficient production capabilities, low efficiency, and unstable quality, making it difficult to meet the industry's rapid iteration needs.
Intelligent agents are used for end-to-end intelligent scheduling, including plot text processing, automatic scene composition, digital resource management and track arrangement, and large-scale generative language models and Unreal Engine are used to achieve automated video generation.
It significantly improves the production efficiency and stability of animated videos, reducing the time from days or months to just a few minutes, lowering the technical threshold and labor costs, and adapting to the needs of rapid iteration.
Smart Images

Figure CN121865053A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a video generation method, apparatus, terminal, electronic device, and storage medium. Background Technology
[0002] In fields requiring animated video generation, currently, after obtaining the script for the animated video, it is usually necessary to rely on manual creation to generate the animated video. The manual creation process requires the collaborative work of many developers such as storyboard designers, animators, and audio designers, and it takes several days or even months to complete the animated video. This results in insufficient animation video production capabilities, low generation efficiency, unstable quality, and difficulty in meeting industry needs. Summary of the Invention
[0003] This disclosure provides a video generation method, apparatus, terminal, electronic device, and storage medium to solve problems in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a video generation method is provided, applied to an intelligent agent, the method comprising: Receive plot text; Based on the aforementioned plot text, a storyboard sequence is generated; Based on the storyboard information corresponding to each storyboard in the storyboard sequence, the corresponding digital resources are obtained through interaction with the digital resource management component; By interacting with the track arrangement component, the track arrangement result obtained by performing corresponding track arrangement based on the digital resources is obtained; The output is a video obtained by rendering the track arrangement results.
[0004] In one exemplary implementation, generating a storyboard sequence based on the plot text includes: By interacting with a large generative language model, the large generative language model is triggered to perform literary expansion on the plot text to obtain the target text; The storyboard sequence is obtained based on the plot breakdown of the target text.
[0005] In one exemplary embodiment, obtaining the corresponding digital resources by interacting with the digital resource management component based on the storyboard information corresponding to each storyboard in the storyboard sequence includes performing the following operations for the storyboard information corresponding to each storyboard in the storyboard sequence: Generate at least one digital resource acquisition instruction based on the storyboard information; The at least one digital resource acquisition instruction is sent to the first communication component, which triggers the first communication component to generate or search for the digital resource pointed to by the at least one digital resource acquisition instruction by calling the digital resource management component.
[0006] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action generation interface is invoked to trigger the action generation model corresponding to the action generation interface to generate the action resources. When the digital resource acquisition instruction is used to acquire facial expression resources, the facial expression generation interface is invoked to trigger the facial expression generation model corresponding to the facial expression generation interface to generate the facial expression resources. When the digital resource acquisition instruction is used to acquire audio resources, the audio generation interface is invoked to trigger the audio generation model corresponding to the audio generation interface to generate the audio resources; The motion generation model, the facial expression generation model, and the audio generation model all belong to the digital resource management component.
[0007] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action retrieval interface is invoked to retrieve the action resources in the action resource library; When the digital resource acquisition instruction is used to acquire emoticon resources, the emoticon retrieval interface is invoked to retrieve the emoticon resources in the emoticon resource library; When the digital resource acquisition instruction is used to acquire audio resources, the audio retrieval interface is invoked to retrieve the audio resources in the audio resource library; The motion resource library, the facial expression resource library, and the audio resource library all belong to the digital resource management component.
[0008] In one exemplary implementation, obtaining the track arrangement result based on the digital resources through interaction with the track arrangement component includes: Based on the storyboard information and the digital resources, create track arrangement instructions; The track arrangement command is sent to the second communication component, which triggers the second communication component to obtain the track arrangement result by calling the track arrangement service of Unreal Engine. The track arrangement service of Unreal Engine belongs to the track arrangement component.
[0009] In one exemplary embodiment, the second communication component performs the following operations: The orbit arrangement instructions are converted into standardized instructions adapted to the Unreal Engine, and the standardized instructions are sent to the Unreal Engine's engine interface service so that the engine interface service can call the orbit arrangement service based on the standardized instructions.
[0010] In one exemplary implementation, the track scheduling service performs the following operations: Create a track; Mount the digital resources; A timeline is created based on the storyboard information. Based on the timeline, the digital resources are arranged on the track, and the track arrangement result is output.
[0011] In one exemplary embodiment, the output is a video obtained by rendering each of the track arrangement results, including: If the engine interface service detects that the track arrangement result is output by the engine interface service, the rendering component is invoked to render the track arrangement result to obtain the video.
[0012] According to a second aspect of the present disclosure, a video generation apparatus is provided for use with an intelligent agent, the apparatus comprising: The plot text receiving module is configured to receive plot text. The storyboard module is configured to generate a storyboard sequence based on the plot text; The digital resource acquisition module is configured to obtain the corresponding digital resources by interacting with the digital resource management component based on the storyboard information corresponding to each storyboard in the storyboard sequence. The track arrangement module is configured to perform track arrangement results based on the digital resources by interacting with the track arrangement component. The video output module is configured to output the video obtained by rendering the track arrangement results.
[0013] In one exemplary implementation, the storyboard module is configured to perform: By interacting with a large generative language model, the large generative language model is triggered to perform literary expansion on the plot text to obtain the target text; The storyboard sequence is obtained based on the plot breakdown of the target text.
[0014] In one exemplary embodiment, the digital resource acquisition module is configured to perform the following operations for the storyboard information corresponding to each storyboard in the storyboard sequence: Generate at least one digital resource acquisition instruction based on the storyboard information; The at least one digital resource acquisition instruction is sent to the first communication component, which triggers the first communication component to generate or search for the digital resource pointed to by the at least one digital resource acquisition instruction by calling the digital resource management component.
[0015] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action generation interface is invoked to trigger the action generation model corresponding to the action generation interface to generate the action resources. When the digital resource acquisition instruction is used to acquire facial expression resources, the facial expression generation interface is invoked to trigger the facial expression generation model corresponding to the facial expression generation interface to generate the facial expression resources. When the digital resource acquisition instruction is used to acquire audio resources, the audio generation interface is invoked to trigger the audio generation model corresponding to the audio generation interface to generate the audio resources; The motion generation model, the facial expression generation model, and the audio generation model all belong to the digital resource management component.
[0016] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action retrieval interface is invoked to retrieve the action resources in the action resource library; When the digital resource acquisition instruction is used to acquire emoticon resources, the emoticon retrieval interface is invoked to retrieve the emoticon resources in the emoticon resource library; When the digital resource acquisition instruction is used to acquire audio resources, the audio retrieval interface is invoked to retrieve the audio resources in the audio resource library; The motion resource library, the facial expression resource library, and the audio resource library all belong to the digital resource management component.
[0017] In one exemplary implementation, the track orchestration module is configured to perform: Based on the storyboard information and the digital resources, create track arrangement instructions; The track arrangement command is sent to the second communication component, which triggers the second communication component to obtain the track arrangement result by calling the track arrangement service of Unreal Engine. The track arrangement service of Unreal Engine belongs to the track arrangement component.
[0018] In one exemplary embodiment, the second communication component performs the following operations: The orbit arrangement instructions are converted into standardized instructions adapted to the Unreal Engine, and the standardized instructions are sent to the Unreal Engine's engine interface service so that the engine interface service can call the orbit arrangement service based on the standardized instructions.
[0019] In one exemplary implementation, the track scheduling service performs the following operations: Create a track; Mount the digital resources; A timeline is created based on the storyboard information. Based on the timeline, the digital resources are arranged on the track, and the track arrangement result is output.
[0020] In one exemplary implementation, the video output module is configured to perform: If the engine interface service detects that the track arrangement result is output by the engine interface service, the rendering component is invoked to render the track arrangement result to obtain the video.
[0021] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the video generation method as described in any of the above embodiments.
[0022] According to a fourth aspect of the present disclosure, a computer storage medium is provided, wherein when instructions in the computer storage medium are executed by a processor of an electronic device, the electronic device performs the video generation method described in any of the above embodiments.
[0023] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including a computer program that, when executed by a processor, implements the video generation method described in any of the above embodiments.
[0024] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: The video generation method provided in this disclosure uses an intelligent agent that can achieve intelligent scheduling across the entire chain, from the automatic understanding of the initial plot text to the final generation of animated videos. Specifically, the intelligent agent achieves full-chain understanding and automated driving of the plot's "text processing - automatic storyboarding - digital resources - track arrangement," breaking through the limitations of traditional large-scale generative language models that only focus on natural language processing. It realizes intelligent scheduling and direct driving of various tools along the link from plot text to video output. With the support of the intelligent agent, an automated animated video production process led by the intelligent agent is realized, eliminating reliance on manual labor and significantly improving the efficiency and stability of video production. As a result, animated videos that originally took days or even months to complete can be generated in just a few minutes, significantly improving the production capabilities of animated videos.
[0025] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0026] The accompanying drawings, which are incorporated in and form part of this disclosure, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0027] Figure 1 This is a flowchart illustrating a video generation method according to an exemplary embodiment; Figure 2 This is a schematic diagram illustrating the storyboard sequence generation process according to an exemplary embodiment; Figure 3 This is a schematic diagram illustrating a digital resource acquisition process according to an exemplary embodiment; Figure 4 This is a schematic diagram illustrating a track arrangement process according to an exemplary embodiment; Figure 5 This is a schematic diagram illustrating the operation process of a track orchestration service according to an exemplary embodiment; Figure 6 This is a schematic diagram illustrating an architecture for completing the entire video creation process based on an intelligent agent, according to an exemplary embodiment. Figure 7 This is a block diagram of a video generation apparatus according to an exemplary embodiment; Figure 8 This is a structural block diagram of a computer device according to an exemplary embodiment. Figure 1 ; Figure 9 This is a structural block diagram of a computer device according to an exemplary embodiment. Figure 2 . Detailed Implementation
[0028] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0029] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0030] To facilitate understanding of this disclosure, a brief introduction to the background technology of this disclosure is provided first: With the rapid development of computer graphics and artificial intelligence, many innovative technologies have emerged in the research field of digital content production. These technologies have not only improved the efficiency of content production but also enriched its forms of expression and expanded its application scenarios, thus promoting the development of related industries, such as the game industry and the production of 2D or 3D animated films and television series.
[0031] With the explosive growth of 2D and 3D narrative animation production in the gaming and film industries, the production efficiency of narrative animation, as a core carrier of storytelling and immersion, has become a prominent pain point in the industry. Traditional narrative animation production relies on manual frame-by-frame editing and multi-position collaboration. For example, in the production of cutscenes for games characterized by high cost, large scale, and high quality, teams need to invest a large number of animators and storyboard designers, spending months to complete key narrative animations. In the realm of lightweight content such as virtual short dramas and game promotional videos, the demand for "rapid iteration and low-cost trial and error" further highlights the inefficiency of traditional processes. Therefore, this publication proposes a highly automated video production solution to meet the industry's demand for full-process automation from script to narrative animation. This type of production solution has broad application value in scenarios such as game cutscenes, virtual short dramas, and animated trailers.
[0032] Typically, taking 3D narrative animation production as an example, those skilled in the art can only develop 3D animated videos through a manually-led 3D narrative animation production workflow. This development process can utilize Unreal Engine (UE), a real-time 3D creation tool developed by Epic Games, possessing capabilities such as level editing, animation choreography, and real-time rendering, and is widely used in the game and film industries. Taking the UE engine ecosystem as an example, the manually-led 3D narrative animation production workflow includes the following: First, after the screenwriter completes the script, the storyboard designer breaks the script down into several storyboards, clarifying the camera angle (such as close-up, wide shot), character actions, camera movement, duration, and other information for each storyboard. Next, the animators manually create character tracks and camera tracks in the UE engine's Sequencer, bind pre-made body animations (such as selecting the "fighting sword swing" animation from the motion library) and facial animations (such as the "determined" facial sequence) to the characters, and adjust the camera keyframes frame by frame to achieve motion effects. Here, Sequencer refers to the level sequence editor built into the UE engine, which is used to arrange the timelines of character, camera, audio and other tracks, and is the core tool for 3D animation track arrangement.
[0033] Meanwhile, audio designers need to create or select corresponding sound effects and voices, create audio tracks in Sequencer, and manually synchronize them with visual elements; Finally, the rendering capabilities of the UE engine are used to output a 3D story animation video.
[0034] This disclosure proposes that the aforementioned human-driven 3D narrative animation production process has at least the following drawbacks: The process is cumbersome and time-consuming: from script to animation, it requires multiple steps such as "manual storyboarding → manual choreography track by track → asset debugging". Taking a 1-minute 3D story animation as an example, the traditional process requires 3-5 professionals to work together for 1-2 weeks to complete, which cannot adapt to the industry's rapid iteration needs.
[0035] High technical barriers and high labor costs: It requires practitioners to be proficient in multiple fields such as UE engine operation, storyboard design, animation production, and audio design, and cross-position collaboration is required. Labor costs account for more than 60% of the total project cost.
[0036] Poor iteration flexibility: If the script needs to be modified, the storyboard and track parameters must be manually readjusted, which is time-consuming and prone to errors. For example, a detailed adjustment to a character's "sword-wielding action" in the script may lead to a chain of modifications to the animation track and camera track, taking several days.
[0037] To overcome the aforementioned drawbacks of the manually-driven 3D narrative animation production process, this disclosure provides a video generation scheme that can automatically generate both 3D and 2D narrative animations. This video generation scheme aims to achieve the following technical objectives: Solve the problems of cumbersome, time-consuming, and labor-intensive traditional 3D story animation production processes; To address the issues of high technical requirements for practitioners and high labor costs; To address the issues of poor flexibility and high modification costs in adjusting the animation production chain during script iteration; This invention addresses the issue of insufficient deep integration between intelligent agents and 3D production tools such as UX engines, which prevents them from directly driving automated production. The intelligent agent, as described in this disclosure, can be constructed based on a large-scale generative language model and possesses task planning, tool invocation, and autonomous decision-making capabilities, enabling it to drive automated 3D animation production processes.
[0038] Figure 1 This is a flowchart illustrating a video generation method according to an exemplary embodiment. The video generation method can be applied to an electronic device, which can be implemented independently by a server or a terminal, or jointly by a terminal and a server. The terminal can be, but is not limited to, physical devices such as smartphones, tablets, laptops, desktop computers, smart speakers, smart wearable devices, digital assistants, augmented reality devices, and virtual reality devices, and can also include software such as applications running on the physical device. The server can be, but is not limited to, a standalone server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, and big data and artificial intelligence platforms, among others.
[0039] The electronic device is equipped with an intelligent agent that can be used to perform... Figure 1 The video generation method in [the document / reference]. Figure 1 As shown, the method includes the following steps.
[0040] In S110, receive the story text.
[0041] This disclosure does not limit the narrative text. For example, the narrative text can be any story content input by the user, or narrative text generated by an electronic device. For instance, the narrative text could describe a science fiction adventure story with multiple characters and a complex plot development. Alternatively, it could be a simple slice of daily life, such as an afternoon spent in a café. This disclosure can automatically create videos that express the content of the narrative text based on it.
[0042] In S120, a storyboard sequence is generated based on the plot text; In this disclosure, the intelligent agent can be configured as an intelligent proxy with automatic shot creation capabilities. For example, the agent can analyze key plot points and character actions in the narrative text to determine the perspective, duration, and transitions of each shot. For instance, when the narrative text describes a tense chase scene, the agent might generate a sequence of shots containing rapidly switching close-ups and extreme close-ups to enhance the audience's tension. Furthermore, the agent can adjust the language style of the shots according to the emotional tone of the narrative, such as using soft lighting effects and slow camera movements in warm scenes. In this way, the agent not only improves the efficiency of shot creation but also ensures the artistry and narrative coherence of the video production.
[0043] In one exemplary implementation, please refer to Figure 2 The diagram illustrates the storyboard sequence generation process of this disclosure. The generation of the storyboard sequence based on the plot text includes: S210. By interacting with a large generative language model, the large generative language model is triggered to perform literary expansion on the plot text to obtain the target text.
[0044] This disclosure utilizes Large Language Models (LLMs) for literary expansion, thereby enriching and polishing narrative text to generate target text with more details. Large Language Models, or LLMs for short, possess powerful natural language understanding and generation capabilities. They can automatically generate context-appropriate descriptive language from the input narrative text, further enriching plot content and character portrayal. For example, when describing a scene of a warrior confronting a dragon, a large LLM can supplement environmental details, such as dark clouds in the sky, cracks in the castle walls, and the cold light reflected from the dragon's scales in dim light. These expansions not only enhance the visual appeal of the text but also provide richer material support for subsequent storyboard generation, digital asset creation, and trajectory arrangement. Furthermore, large LLMs can adjust the language style according to the emotional needs of the plot, making it more aligned with the narrative rhythm, thus achieving a seamless transition from text to visual presentation.
[0045] This disclosure does not limit the specific model structure or model type of large-scale generative language models. For example, LLM Backend, Llama series, and QWEN series models can be used. LLM Backend is a backend language model architecture that supports various natural language processing tasks. Its flexibility and efficiency make it suitable for text generation needs in different scenarios. The Llama series models are developed by leading research institutions and are known for their open-source nature and strong multilingual support capabilities, especially excelling in complex contexts. The QWEN series models are language models specifically optimized for the Chinese environment, possessing deep industry knowledge and high-precision generation capabilities, meeting diverse needs from creative writing to professional document processing. The choice of these models depends on the specific application scenario and the quality requirements of the generated content.
[0046] S220. Based on the plot decomposition of the target text, the storyboard sequence is obtained.
[0047] This intelligent agent can decompose the plot based on the target text and create storyboards based on the decomposition results, or it can decompose the plot and create storyboards through interaction with a large generative language model. This disclosure does not limit the specific decomposition method, the decomposition results, or the process of obtaining a storyboard sequence based on the decomposition results.
[0048] In one exemplary implementation, the agent can extract core elements suitable for visual expression by analyzing key plot points and character interactions in the target text. These core elements are then organized into a logically coherent sequence of storyboards to provide a clear guiding framework for subsequent animation production. During this process, the agent may combine a pre-defined rule base or utilize machine learning algorithms to optimize the generation of storyboards, ensuring that the final output content conforms to both the semantics of the original text and the needs of visual narrative. This disclosure does not limit the core elements; for example, these core elements may include character actions, scene transitions, emotional changes, and the use of key props. When generating the storyboard sequence, the agent dynamically adjusts the details of each storyboard based on the semantic features of the target text, such as the choice of camera angle, control of the scene rhythm, and setting of visual focus. This design not only improves the efficiency of storyboard generation but also enhances the expressiveness and narrative depth of the final animation. Furthermore, by introducing user-defined parameters or interactive adjustment functions, the agent can further meet the personalized needs of different creators, thereby achieving efficient conversion from text to visual content.
[0049] In another implementation, the agent can also dynamically adjust the generation strategy of the storyboard sequence according to the user's needs. For example, when the user wants to highlight certain specific plot points, the agent can automatically reallocate visual resources to focus on these plot points while maintaining the coherence of the overall narrative.
[0050] For example, first, the agent receives simple plot text input by the user (such as "The warrior confronts the dragon before the castle"), then sends a request to LLM Backend, completing two core tasks based on the interaction with LLM Backendjioahu: Plot polishing: Expand the input text with literary flair, adding details about characters, scene atmosphere, etc. (e.g., expand "the warrior drew his sword" to "the warrior in silver armor suddenly drew his long sword that gleamed coldly, the blade cutting through the air with a hum"). Storyboard structure generation: The polished plot is broken down into a structured storyboard list of "storyboard number, camera angle, character action, camera movement, duration, and audio requirements" (e.g., storyboard 1: "wide shot, warrior draws sword and roars, camera slowly advances, duration 3 seconds, audio is roar + sword drawing sound effect").
[0051] Clearly, in this approach, the intelligent agent replaces the human storyboard artist, automating the storyboard design process. By fully utilizing the natural language understanding capabilities of large-scale generated content, it not only expands and polishes the script but also automates and professionally breaks down the storyboards, resulting in higher-quality storyboards and significantly improved efficiency—at least 10 times faster than manual storyboarding in actual use. The core guarantee of this open-source approach's ease of use lies in its ability to transform unprofessional scripts into structured storyboards directly usable in video production through intelligent agents, thus lowering the user barrier.
[0052] In S130, based on the storyboard information corresponding to each storyboard in the storyboard sequence, the corresponding digital resources are obtained through interaction with the digital resource management component; This disclosure allows for the identification of storyboard information for each storyboard shot, typically encompassing key elements such as the shot's visual content, scene description, and character actions. Then, by interacting with a digital asset management component, this storyboard information is used as input to query relevant digital assets or trigger related models to generate them. These digital assets are the digital assets or media materials used to generate the video.
[0053] In one exemplary implementation, please refer to Figure 3 This diagram illustrates the digital resource acquisition process of this disclosure. The step of obtaining the corresponding digital resources based on the storyboard information corresponding to each storyboard in the storyboard sequence through interaction with the digital resource management component includes performing the following operations for the storyboard information corresponding to each storyboard in the storyboard sequence: S310. Generate at least one digital resource acquisition instruction based on the storyboard information; S320. Send the at least one digital resource acquisition instruction to the first communication component, triggering the first communication component to generate or search for the digital resource pointed to by the at least one digital resource acquisition instruction by calling the digital resource management component.
[0054] This disclosure does not limit the method of generating at least one digital resource acquisition instruction by parsing the storyboard information. For example, the object involved in the storyboard information, the action performed by the object, the environment in which the object is located, the sound effects configured for the object, etc., can be parsed. Based on this information, the digital resources required to generate the video segment corresponding to the storyboard information can be determined. Digital resource acquisition instructions are constructed for these digital resources, and these digital resource acquisition instructions are transmitted to the first communication component, so that the first communication component sends these digital resource acquisition instructions to the corresponding digital resource management component to trigger the digital resource management component to generate or search for relevant digital resources.
[0055] This design encapsulates the generation and query capabilities of various types of digital resources within a single logical layer. This logical layer is visible to the first communication component, which acts as a relay hub, interacting with it to achieve efficient scheduling and management of digital resources. This encapsulation not only simplifies system complexity but also improves the speed and accuracy of resource acquisition. The first communication component translates the specific instructions issued by the agent into a form acceptable to the standardized interface exposed by the logical layer, ensuring that different types of digital resources can be requested and processed in a unified manner. The agent only needs to focus on the logic without concern for the specific implementation details. This design gives the system greater flexibility and scalability, enabling it to adapt to the dynamic needs of different types of digital resources.
[0056] This disclosure does not limit the first communication component, which can be developed based on MCP, which stands for ModelContext Protocol. MCP is a modular communication platform that can be used as an intermediate layer to encapsulate and schedule the application programming interface of Unreal Engine and various artificial intelligence models, thereby achieving standardized network communication.
[0057] For example, based on the storyboard information, the intelligent agent can invoke the capabilities of the artificial intelligence model through the first communication component to generate corresponding actions, expressions, and audio assets. Specific examples may include: Action asset generation: The agent sends a digital resource acquisition instruction to the first communication component, which is "Character A performs the 'draw sword and roar' action". The first communication component calls the action generation interface of the action generation model through the HTTP protocol to trigger the action generation model to generate skeletal animation data and return the data to the first communication component.
[0058] Facial expression asset generation: The agent sends a digital resource acquisition instruction for "character A's angry expression" to the first communication component. The first communication component calls the facial expression generation model via the HTTP protocol to trigger the facial expression generation model to generate the corresponding facial expression sequence data and return it to the first communication component.
[0059] Audio asset generation: The agent sends a digital resource acquisition command of "roaring sound + sword drawing sound effect" to the first communication component to trigger the audio generation model to generate emotional speech and sound effect files and return them to the first communication component.
[0060] In this way, these digital resources are returned to the first communication component, which can then return this data or its storage address to the agent. The agent can then use this data to drive downstream components, ultimately completing the video generation task. This design utilizes the first communication component to standardize the scheduling of various digital resource generation or query capabilities, thereby replacing the manual creation / selection of digital resources and reducing the time required for this step from "hours" to "minutes".
[0061] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action generation interface is invoked to trigger the action generation model corresponding to the action generation interface to generate the action resources. When the digital resource acquisition instruction is used to acquire facial expression resources, the facial expression generation interface is invoked to trigger the facial expression generation model corresponding to the facial expression generation interface to generate the facial expression resources. When the digital resource acquisition instruction is used to acquire audio resources, the audio generation interface is invoked to trigger the audio generation model corresponding to the audio generation interface to generate the audio resources; The motion generation model, the facial expression generation model, and the audio generation model all belong to the digital resource management component.
[0062] For example, when creating an animated scene and needing to obtain character motion resources, the first communication component calls the motion generation interface. Suppose in a scene where a warrior confronts a dragon in front of a castle, to generate the action of the warrior drawing his sword, the motion generation interface is called, triggering the corresponding motion generation model to begin working. This motion generation model can generate a physically accurate and natural sword-drawing action based on a preset algorithm and dataset.
[0063] Similarly, for acquiring facial expression resources, such as to express the warrior's resolute expression when facing the dragon, the first communication component calls the facial expression generation interface to trigger the facial expression generation model to generate delicate and emotional expressions based on parameters such as the character's facial structure and emotional settings, making the character more vivid and lifelike.
[0064] When audio resources are involved, such as the sound effects of a warrior drawing his sword or background music, the first communication component calls the audio generation interface to trigger the audio generation model to synthesize realistic sword-drawing sounds or tense and exciting background music, enhancing the atmosphere of the entire animation scene.
[0065] These generative models all belong to the digital asset management component, and each model can be maintained, updated, and optimized independently. Because each model focuses on generating a specific type of asset, it can more accurately meet the needs of different types of assets in the animation production process, improving the quality and efficiency of asset generation.
[0066] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action retrieval interface is invoked to retrieve the action resources in the action resource library; When the digital resource acquisition instruction is used to acquire emoticon resources, the emoticon retrieval interface is invoked to retrieve the emoticon resources in the emoticon resource library; When the digital resource acquisition instruction is used to acquire audio resources, the audio retrieval interface is invoked to retrieve the audio resources in the audio resource library; The motion resource library, the facial expression resource library, and the audio resource library all belong to the digital resource management component.
[0067] For example, when creating an animated scene, if a sword-drawing action needs to be designed for the warrior, the first communication component will call the action retrieval interface based on the digital resource acquisition command. This interface will quickly locate the required action resources in the action resource library, such as finding action data related to sword drawing from the preset "weapon operation" category. Similarly, when it is necessary to depict the complex facial expressions of a warrior facing a dragon, the first communication component will call the expression retrieval interface to filter out expression resources that can express emotions such as determination and tension from the expression resource library. For background music or environmental sound effects, the audio retrieval interface is used to find suitable materials in the audio resource library, such as a deep cello melody or the sound of whistling wind. By clearly defining different types of resource libraries and providing corresponding retrieval interfaces, the efficiency of resource searching can be greatly improved, relevant resources can be quickly obtained, time costs can be reduced, and existing digital resources can be fully utilized to quickly acquire the relevant materials needed for video creation.
[0068] The retrieval interface disclosed herein can work in conjunction with the generation interface described above to rapidly improve the speed of digital resource generation or retrieval. Specifically, when users need to create complex animation scenes, the retrieval interface can quickly find matching materials from various resource libraries, thereby improving the speed and quality of the generation interface in generating specific materials. For example, when designing the action of a warrior drawing his sword, the generation interface can adjust the action details based on the retrieval results, making the entire process more efficient. Similarly, for facial expressions and audio resources, the retrieval interface can accurately locate the required content and work in conjunction with the generation interface to optimize the output effect. The various specific materials generated can then be used to populate relevant resource libraries, further enriching their content. This collaborative mechanism not only improves resource utilization efficiency but also provides more possibilities for subsequent creation. For example, when creating a scene of a warrior confronting a dragon, the retrieval interface can quickly find background music and environmental sound effects that match the atmosphere, while the generation interface can further optimize audio details based on these materials, making the overall effect more immersive. Furthermore, by continuously accumulating generated specific materials, the diversity and quality of the resource library are continuously improved, thereby supporting more complex and refined animation production needs. This virtuous cycle provides strong technical support for animation creation.
[0069] In S140, through interaction with the track arrangement component, the track arrangement result obtained by performing corresponding track arrangement based on the digital resources is obtained.
[0070] In one exemplary implementation, please refer to Figure 4 This illustration shows a schematic diagram of a track arrangement process according to an exemplary embodiment of the present disclosure. The step of obtaining the track arrangement result based on the digital resources through interaction with the track arrangement component includes: S410. Create track arrangement instructions based on the storyboard information and the digital resources; For example, a track orchestration instruction can be created based on storyboard information and the storage location of these digital resources. This track orchestration instruction is used to drive downstream components to use the storyboard information and perform track orchestration by mounting these digital resources.
[0071] S420. The track arrangement instruction is sent to the second communication component, triggering the second communication component to obtain the track arrangement result by calling the track arrangement service of Unreal Engine, wherein the track arrangement service of Unreal Engine belongs to the track arrangement component.
[0072] For example, the second communication component performs the following operations, including: converting the track orchestration instructions into standardized instructions adapted to the Unreal Engine, and sending the standardized instructions to the Unreal Engine's engine interface service, so that the engine interface service invokes the track orchestration service based on the standardized instructions.
[0073] This disclosure does not limit the second communication component, which can be developed based on MCP, which stands for ModelContext Protocol. MCP is a modular communication platform that can be used as an intermediate layer to encapsulate and schedule the application programming interface of Unreal Engine and various artificial intelligence models, thereby achieving standardized network communication.
[0074] For example, when creating an animated short film, a user might need to choreograph a complex battle scene. At this point, the second communication component receives choreographing instructions from upstream (the agent), which may include information such as character movements, camera paths, and special effects triggers. To ensure that Unreal Engine can accurately understand and execute these instructions, the second communication component first converts them into standardized instructions that conform to Unreal Engine specifications. This standardization process not only improves compatibility but also significantly reduces the error rate caused by mismatched instruction formats.
[0075] Next, these standardized instructions are sent to Unreal Engine's engine interface service. Upon receiving the instructions, the engine interface service parses their content and invokes the trajectory orchestration service to perform the specific operations. For example, it might generate a precise camera movement trajectory based on the instructions, or synchronize the movements of multiple characters to achieve smooth combat effects. In this way, the entire trajectory orchestration process becomes more efficient and easier to manage.
[0076] This design standardizes instructions, making collaboration between different components smoother and reducing the need for human intervention. By calling the orbital orchestration service through the engine interface service, the powerful capabilities of Unreal Engine can be fully utilized to achieve high-quality animation effects, thereby improving production efficiency.
[0077] In one exemplary implementation, please refer to Figure 5 This illustration shows a schematic diagram of the operation process performed by the track orchestration service in an exemplary embodiment of this disclosure. The track orchestration service performs the following operations: S510. Create a track.
[0078] Creating a track can be understood as building a basic framework that carries all animation and interactive elements. For example, in a scene of "the warrior confronting the dragon," track creation might involve defining the trajectory of changes in the camera's perspective, keyframes for character movements, and the timing of environmental effects.
[0079] S520. Mount the digital resource.
[0080] Mounting digital assets is the process of binding pre-made digital assets to their corresponding tracks. In practice, digital assets might include 3D models of warriors, animated dragons, textures of castle backgrounds, and fire effects. Through this process, each asset is precisely assigned to its corresponding track, ensuring it functions correctly at the right time and place.
[0081] S530. Create a timeline based on the storyboard information, arrange the digital resources on the track based on the timeline, and output the track arrangement result.
[0082] Creating a timeline and arranging digital assets based on storyboard information is the core of the entire process. Storyboard information provides detailed guidance on the timing and pacing of plot developments. For example, according to the storyboard, the warrior's sword-drawing action needs to be completed at the 5th second, while the dragon's fire-breathing effect needs to be triggered at the 7th second. Through the arrangement of the timeline, these events are precisely mapped onto the timeline, thus achieving smooth animation effects.
[0083] For example, an agent interacts with a second communication component, issuing track orchestration instructions. These instructions could be: "Create a track for character A, attaching the 'Sword Drawing Roar' animation, timeline 0-3 seconds; create a camera track, setting a 'Slow Progression' keyframe, timeline 0-3 seconds; create an audio track, attaching 'Roar + Sword Drawing Sound Effect,' timeline 0-3 seconds." The second communication component converts these into standardized instructions and sends them to the engine interface service. After parsing the instructions, the engine interface service automatically completes track creation, asset mounting, and timeline synchronization using the Unreal Engine's Sequencer, achieving track orchestration without human intervention. The Sequencer is a powerful tool provided by Unreal Engine for handling complex animation sequences and resource management. It automatically adjusts elements on each track according to preset rules and instructions, ensuring they trigger at the correct times and work together. This automated process significantly reduces the need for manual operation and improves overall efficiency and accuracy. Furthermore, Sequencer supports multi-track parallel processing, allowing for detailed simultaneous choreography of character movements, camera motion, special effects, and sound effects, thereby creating more vivid and realistic scene effects. In this way, creators can focus more on creative expression without being bogged down in tedious technical details. This disclosure uses an intelligent agent to ultimately drive the Sequencer, performing automatic track-by-track choreography, reducing track choreography time from "days" to "minutes".
[0084] In one exemplary embodiment, the first communication component of this disclosure can support HTTP communication, where HTTP refers to Hypertext Transfer Protocol. The first communication component can communicate with the Digital Resource Management component based on HTTP. In another exemplary embodiment, the second communication component of this disclosure can support WebSocket communication, where WebSocket refers to a bidirectional communication protocol. The second communication component can interact with Unreal Engine based on WebSocket. In some embodiments, the communication protocol between the second communication component and the Unreal Engine's engine interface service can be replaced with a high-performance protocol such as gRPC, improving instruction transmission efficiency in high-concurrency scenarios without changing the core logic of automated orchestration. gRPC refers to a high-efficiency remote procedure call framework that can automatically generate client-server communication code by defining service interfaces and message types. Both the first and second communication components of this disclosure solve the communication problem between different components, breaking down communication barriers between them, enabling intelligent agents to call various components like calling natural language interfaces, and serving as a key link in technological integration.
[0085] In S150, the video obtained by rendering the track arrangement results is output.
[0086] In one exemplary implementation, the output is a video obtained by rendering each of the track arrangement results, including: upon detecting that the track arrangement result is output by the engine interface service, calling the rendering component to render the track arrangement result to obtain the video. In this implementation, once the Unreal Engine's Sequencer completes track arrangement, the rendering process can be automatically triggered to generate a two-dimensional or three-dimensional animated video. This design can improve the efficiency of video production and reduce the need for manual intervention.
[0087] The video generation method provided in this disclosure uses an intelligent agent that can achieve intelligent scheduling across the entire chain, from the automatic understanding of the initial plot text to the final generation of animated videos. Specifically, the intelligent agent achieves full-chain understanding and automated driving of the plot's "text processing - automatic storyboarding - digital resources - track arrangement," breaking through the limitations of traditional large-scale generative language models that only focus on natural language processing. It realizes intelligent scheduling and direct driving of various tools along the link from plot text to video output. With the support of the intelligent agent, an automated animated video production process led by the intelligent agent is realized, eliminating reliance on manual labor and significantly improving the efficiency and stability of video production. As a result, animated videos that originally took days or even months to complete can be generated in just a few minutes, significantly improving the production capabilities of animated videos.
[0088] In one exemplary implementation, please refer to Figure 6 It illustrates an exemplary embodiment of the present disclosure of an architecture diagram for completing the entire video creation process based on an intelligent agent.
[0089] This architecture can be expressed as a collaborative architecture of "Agent + LLMBackend + MCP layer + AI capability service + UE engine". This collaborative architecture realizes full-process automation from script text to animation. Among them, LLM Backend is a large-scale generative language model used for literary expansion of script and auxiliary storyboarding. The MCP layer deploys the first and second communication components. The AI capability service is encapsulated in the digital resource management component. The UE engine refers to Unreal Engine, or it can refer to the track arrangement service provided by Unreal Engine.
[0090] The MCP layer comprises mcpserver1 and mcpserver2, which are the first and second communication components, respectively. The first communication component interacts with the AI capability service via HTTP, while the second communication component interacts with the Unreal Engine via WebSocket. The agent is the core of the automated script generation process, responsible for script understanding, storyboard generation, and intelligent invocation of related components. The LLM Backend provides the agent with reasoning capabilities based on a large generative language model. The MCP layer acts as an intermediate communication bridge, encapsulating communication capabilities between the AI capability service and the UE engine's related interfaces. The AI capability service provides the ability to generate or query assets such as actions, expressions, and audio. The engine / art capabilities (UE engine) are responsible for the final animation choreography and rendering.
[0091] Below is a specific example of a video generation process implemented using this framework: 1. Script Input: The user inputs the text "The warrior confronts the dragon in front of the castle. The warrior draws his sword and roars. The dragon breathes fire. The warrior dodges and then swings his sword to cut the dragon's wings."
[0092] 2. Plot polishing and storyboard generation: The agent made a request to LLM Backend to polish the target script: "A warrior clad in silver armor stood before the ruins of an ancient castle, his eyes burning with rage. He suddenly drew his longsword, which gleamed coldly at his waist, and roared at the sky, 'Die!' A hundred-meter-tall dragon swooped down from the clouds, opened its huge mouth and spewed flames, and the heat wave swept across the ground; the warrior nimbly rolled away from the flames, and as he stood up, his longsword cut through the air and slashed fiercely at the dragon's left wing membrane..."
[0093] 3. Generate a structured storyboard list: Scene 1: Panoramic view, the warrior stands in front of the castle, drawing his sword and roaring. The camera slowly zooms in, lasting 3 seconds, with the audio consisting of the roar and the sound of the sword being drawn.
[0094] Scene 2: Medium shot, the dragon dives and breathes fire, the camera follows the dragon's movement, 4 seconds long, audio is dragon roar + fire sound effect.
[0095] Scene 3: Close-up, the warrior rolls to dodge, then swings his sword at the dragon's wing. The camera focuses on the warrior's action, lasting 5 seconds, with audio consisting of rolling sound effects and sword-striking sound effects.
[0096] 4. Asset Generation and Track Arrangement: The intelligent agent calls the AI capability service through mcpserver1 to generate body animations such as "drawing a sword and roaring", "diving and spitting fire", and "rolling and swinging a sword", as well as corresponding facial animations, sound effects and voice for each scene.
[0097] The intelligent agent sends track arrangement instructions to the UE engine through mcpserver2. UESequencer automatically creates character tracks (warrior, dragon), camera tracks, and audio tracks, and attaches the corresponding assets to complete timeline synchronization.
[0098] 5. Rendering output: The UE engine renders and generates a 3D story animation video that includes character movements, camera motion, and sound effects. The entire process takes about 30 minutes (the traditional process takes more than a week).
[0099] Clearly, the generation process of this animated video achieves the following technical effects: First, production efficiency is significantly improved: the traditional process requires 1-2 weeks to produce animation, while this invention only takes tens of minutes, improving efficiency by more than 100 times.
[0100] Secondly, the technical threshold and cost are reduced: no professional animators or storyboard designers are needed. Ordinary users only need to input the script to generate animated videos, reducing labor costs by more than 80%.
[0101] Third, the flexibility of iteration is greatly improved: after the script is modified, the large model agent can complete the synchronous adjustment of the storyboard and track within minutes, and the modification cost is close to zero.
[0102] Fourth, expanding creative boundaries: The ability of large models to refine the plot can enrich the details of the story and provide creators with more inspiration. At the same time, the automated process allows creators to focus on the plot idea itself rather than the technical implementation.
[0103] Figure 7 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. The apparatus is applied to an intelligent agent and includes: The plot text receiving module 710 is configured to receive plot text; Storyboard module 720 is configured to generate a storyboard sequence based on the plot text; The digital resource acquisition module 730 is configured to obtain the corresponding digital resources by interacting with the digital resource management component based on the storyboard information corresponding to each storyboard in the storyboard sequence. The track arrangement module 740 is configured to perform track arrangement results based on the digital resources by interacting with the track arrangement component. The video output module 750 is configured to output the video obtained by rendering the track arrangement results.
[0104] In one exemplary embodiment, the storyboard module 720 is configured to perform: By interacting with a large generative language model, the large generative language model is triggered to perform literary expansion on the plot text to obtain the target text; The storyboard sequence is obtained based on the plot breakdown of the target text.
[0105] In one exemplary embodiment, the digital resource acquisition module 730 is configured to perform the following operations for the storyboard information corresponding to each storyboard in the storyboard sequence: Generate at least one digital resource acquisition instruction based on the storyboard information; The at least one digital resource acquisition instruction is sent to the first communication component, which triggers the first communication component to generate or search for the digital resource pointed to by the at least one digital resource acquisition instruction by calling the digital resource management component.
[0106] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action generation interface is invoked to trigger the action generation model corresponding to the action generation interface to generate the action resources. When the digital resource acquisition instruction is used to acquire facial expression resources, the facial expression generation interface is invoked to trigger the facial expression generation model corresponding to the facial expression generation interface to generate the facial expression resources. When the digital resource acquisition instruction is used to acquire audio resources, the audio generation interface is invoked to trigger the audio generation model corresponding to the audio generation interface to generate the audio resources; The motion generation model, the facial expression generation model, and the audio generation model all belong to the digital resource management component.
[0107] In one exemplary implementation, the first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action retrieval interface is invoked to retrieve the action resources in the action resource library; When the digital resource acquisition instruction is used to acquire emoticon resources, the emoticon retrieval interface is invoked to retrieve the emoticon resources in the emoticon resource library; When the digital resource acquisition instruction is used to acquire audio resources, the audio retrieval interface is invoked to retrieve the audio resources in the audio resource library; The motion resource library, the facial expression resource library, and the audio resource library all belong to the digital resource management component.
[0108] In one exemplary embodiment, the track arrangement module 740 is configured to perform: Based on the storyboard information and the digital resources, create track arrangement instructions; The track arrangement command is sent to the second communication component, which triggers the second communication component to obtain the track arrangement result by calling the track arrangement service of Unreal Engine. The track arrangement service of Unreal Engine belongs to the track arrangement component.
[0109] In one exemplary embodiment, the second communication component performs the following operations: The orbit arrangement instructions are converted into standardized instructions adapted to the Unreal Engine, and the standardized instructions are sent to the Unreal Engine's engine interface service so that the engine interface service can call the orbit arrangement service based on the standardized instructions.
[0110] In one exemplary implementation, the track scheduling service performs the following operations: Create a track; Mount the digital resources; A timeline is created based on the storyboard information. Based on the timeline, the digital resources are arranged on the track, and the track arrangement result is output.
[0111] In one exemplary embodiment, the video output module 750 is configured to perform: If the engine interface service detects that the track arrangement result is output by the engine interface service, the rendering component is invoked to render the track arrangement result to obtain the video.
[0112] Regarding the apparatus in the above embodiments, the specific manner of each step has been described in detail in the embodiments of the foregoing method, and will not be elaborated here.
[0113] Please refer to Figure 8 It illustrates the structural block of a computer device provided in an exemplary embodiment of this disclosure. Figure 1 The computer device can be a terminal. This computer device is used to implement the video generation method provided in the above embodiments. Specifically: Typically, computer device 800 includes a processor 801 and a memory 802.
[0114] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In an exemplary embodiment, processor 801 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In an exemplary embodiment, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0115] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In an exemplary embodiment, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one instruction, at least one program, code set, or instruction set, configured to be executed by one or more processors to implement the video generation method described above.
[0116] In one exemplary embodiment, the computer device 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 804, a touch display screen 805, a camera assembly 806, an audio circuit 807, a positioning assembly 808, and a power supply 809.
[0117] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0118] Please refer to Figure 9 It illustrates the structural block of a computer device provided in another exemplary embodiment of this disclosure. Figure 2 The computer device can be a server for executing the video generation method described above. Specifically: Computer device 900 includes a Central Processing Unit (CPU) 901, a system memory 904 including Random Access Memory (RAM) 902 and Read Only Memory (ROM) 903, and a system bus 905 connecting the system memory 904 and the CPU 901. Computer device 900 also includes a basic input / output system (I / O system) 906 that facilitates information transfer between various devices within the computer, and a mass storage device 907 for storing the operating system 913, application programs 914, and other program modules 911.
[0119] The basic input / output system 906 includes a display 908 for displaying information and an input device 909 for user input, such as a mouse or keyboard. Both the display 908 and the input device 909 are connected to the central processing unit 901 via an input / output controller 190 connected to the system bus 905. The basic input / output system 906 may also include the input / output controller 190 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 190 also provides output to a display screen, printer, or other types of output devices.
[0120] Mass storage device 907 is connected to central processing unit 901 via a mass storage controller (not shown) connected to system bus 905. Mass storage device 907 and its associated computer-readable media provide non-volatile storage for computer device 900. That is, mass storage device 907 may include computer-readable media (not shown) such as hard disk or CD-ROM (CompactDisc Read-Only Memory) drive.
[0121] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 904 and mass storage device 907 described above can be collectively referred to as memory.
[0122] According to various embodiments of this disclosure, the computer device 900 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 900 can be connected to a network 912 via a network interface unit 911 connected to a system bus 905, or the network interface unit 911 can be used to connect to other types of networks or remote computer systems (not shown).
[0123] The aforementioned memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the aforementioned video generation method.
[0124] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is executed by a processor to implement the video generation method.
[0125] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0126] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory including program code, which can be executed by a processor to complete the video generation method described above. Optionally, the computer-readable storage medium may be read-only memory (ROM), random access memory (RAM), compact-disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0127] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the video generation method described above.
[0128] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0129] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video generation method, characterized in that, Applied to intelligent agents, the method includes: Receive plot text; Based on the aforementioned plot text, a storyboard sequence is generated; Based on the storyboard information corresponding to each storyboard in the storyboard sequence, the corresponding digital resources are obtained through interaction with the digital resource management component; By interacting with the track arrangement component, the track arrangement result obtained by performing corresponding track arrangement based on the digital resources is obtained; The output is a video obtained by rendering the track arrangement results.
2. The method according to claim 1, characterized in that, The step of generating a storyboard sequence based on the plot text includes: By interacting with a large generative language model, the large generative language model is triggered to perform literary expansion on the plot text to obtain the target text; The storyboard sequence is obtained based on the plot breakdown of the target text.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining corresponding digital resources based on the storyboard information corresponding to each storyboard in the storyboard sequence through interaction with the digital resource management component includes performing the following operations for the storyboard information corresponding to each storyboard in the storyboard sequence: Generate at least one digital resource acquisition instruction based on the storyboard information; The at least one digital resource acquisition instruction is sent to the first communication component, which triggers the first communication component to generate or search for the digital resource pointed to by the at least one digital resource acquisition instruction by calling the digital resource management component.
4. The method according to claim 3, characterized in that, The first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action generation interface is invoked to trigger the action generation model corresponding to the action generation interface to generate the action resources. When the digital resource acquisition instruction is used to acquire facial expression resources, the facial expression generation interface is invoked to trigger the facial expression generation model corresponding to the facial expression generation interface to generate the facial expression resources. When the digital resource acquisition instruction is used to acquire audio resources, the audio generation interface is invoked to trigger the audio generation model corresponding to the audio generation interface to generate the audio resources; The motion generation model, the facial expression generation model, and the audio generation model all belong to the digital resource management component.
5. The method according to claim 3, characterized in that, The first communication component performs at least one of the following operations: When the digital resource acquisition instruction is used to acquire action resources, the action retrieval interface is invoked to retrieve the action resources in the action resource library; When the digital resource acquisition instruction is used to acquire emoticon resources, the emoticon retrieval interface is invoked to retrieve the emoticon resources in the emoticon resource library; When the digital resource acquisition instruction is used to acquire audio resources, the audio retrieval interface is invoked to retrieve the audio resources in the audio resource library; The motion resource library, the facial expression resource library, and the audio resource library all belong to the digital resource management component.
6. The method according to claim 1, characterized in that, The process of obtaining the track arrangement result based on the digital resources through interaction with the track arrangement component includes: Based on the storyboard information and the digital resources, create track arrangement instructions; The track arrangement command is sent to the second communication component, which triggers the second communication component to obtain the track arrangement result by calling the track arrangement service of Unreal Engine. The track arrangement service of Unreal Engine belongs to the track arrangement component.
7. The method according to claim 6, characterized in that, The second communication component performs the following operations, including: The orbit arrangement instructions are converted into standardized instructions adapted to the Unreal Engine, and the standardized instructions are sent to the Unreal Engine's engine interface service so that the engine interface service can call the orbit arrangement service based on the standardized instructions.
8. The method according to claim 7, characterized in that, The track orchestration service performs the following operations: Create a track; Mount the digital resources; A timeline is created based on the storyboard information. Based on the timeline, the digital resources are arranged on the track, and the track arrangement result is output.
9. The method according to claim 7, characterized in that, The output is a video obtained by rendering the track arrangement results, including: If the engine interface service detects that the track arrangement result is output by the engine interface service, the rendering component is invoked to render the track arrangement result to obtain the video.
10. A video generation apparatus, characterized in that, Applied to intelligent agents, the device includes: The plot text receiving module is configured to receive plot text. The storyboard module is configured to generate a storyboard sequence based on the plot text; The digital resource acquisition module is configured to obtain the corresponding digital resources by interacting with the digital resource management component based on the storyboard information corresponding to each storyboard in the storyboard sequence. The track arrangement module is configured to perform track arrangement results based on the digital resources by interacting with the track arrangement component. The video output module is configured to output the video obtained by rendering the track arrangement results.
11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the video generation method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device performs the video generation method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program stored in a readable storage medium, wherein at least one processor of a computer device reads from the readable storage medium and executes the computer program, causing the device to perform the video generation method as described in any one of claims 1 to 9.