Video generation method and electronic device

By using multiple AI agents to collaboratively generate scripts in virtual production and combining them with a 3D development engine to simulate video shooting, the problems of low video quality and poor character consistency in virtual production have been solved, achieving high-quality and coherent video generation.

WO2026092118A1PCT designated stage Publication Date: 2026-05-07HANGZHOU ALIBABA INT INTERNET IND CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HANGZHOU ALIBABA INT INTERNET IND CO LTD
Filing Date
2025-10-13
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In existing technologies, virtual production methods produce videos of low quality, poor character consistency, and limited video length, making it difficult to maintain the continuity and high quality of the storyline.

Method used

By creating a virtual film set environment based on a 3D development engine, multiple AI agents are used to simulate the collaboration of functional roles such as director, screenwriter, and cinematographer, generating a script and performing structured processing. Combined with the 3D virtual film set environment to simulate the video shooting process, the target video is generated.

Benefits of technology

It improves the quality of video generation, ensures consistency of characters, avoids video segmentation and splicing, and enhances the overall quality and coherence of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127330_07052026_PF_FP_ABST
    Figure CN2025127330_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a video generation method and an electronic device. The method comprises: creating a 3D virtual set environment on the basis of a 3D development engine; after receiving a video generation requirement text, generating a script by means of mutual coordination between multiple intelligent agents on the basis of an artificial intelligence (AI) large model, the multiple intelligent agents being used for simulating multiple functional jobs involved in an artificial video production process, the script comprising a character introduction in the script, actor lines, and a character position, action and shot types corresponding to the actor lines; performing structural processing on the script, and simulating a video shooting process in the 3D virtual set environment on the basis of the structured script, so as to generate a target video. By means of the embodiments of the present disclosure, video generation quality and character consistency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method and electronic device for generating video

[0001] The present disclosure claims priority to Chinese Patent Application No. 202411547518.7, filed on October 31, 2024, with the Chinese Patent Office, entitled “Method and electronic device for generating video”, the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to the technical field of video generation, and in particular, to a method and an electronic device for generating video. BACKGROUND

[0003] Virtual production is a production method for creating film and television content by using virtual environment and digital technology, which can greatly reduce costs and improve efficiency. At present, virtual production has been widely used in the shooting of films and television series. With the rapid development of artificial intelligence technology in recent years, people have begun to think about replacing part of the manual operation in the virtual film production process with an automated process in order to further reduce the workload of film production.

[0004] There are generally two methods for applying artificial intelligence to virtual production in the prior art. One is to directly use a “text-to-video” generative model to generate a video. For example, by inputting a piece of text, the model can generate a piece of video based on the text. However, such a large text-to-video model can usually only generate a video of less than 10 seconds, and the video quality is low. There may also be problems such as violation of common sense, poor consistency of characters (the same character may have inconsistent appearance features in different video frames), and no plot in the generated video.

[0005] The other way is to form a video by planning a plot, linking picture generation and video generation tools based on the above-mentioned text-to-video large model for video generation. For example, first, a “text-to-text” large model is used to generate a script. Then, a “text-to-image” large model is used to generate some images for different scenes in the script, i.e., a director's breakdown similar to a breakdown. Then, the images are given to a “image-to-image” large model to enrich some details and generate higher-quality images, which are then converted into moving images. Then, these short videos are spliced together to obtain the final generated video. However, this method only solves the problem of plot, and the video generation quality is still limited by the base model. Moreover, since the video content is generated in segments, it is difficult to maintain consistency between different video segments. SUMMARY

[0006] The present disclosure provides a method and an electronic device for generating video, which can improve the video generation quality and the consistency of characters.

[0007] The present disclosure provides the following solutions:

[0008] A method for generating a video, comprising:

[0009] creating a 3D virtual studio environment based on a 3D development engine, the 3D virtual studio environment including at least one virtual space scene, at least one virtual character image, and a plurality of positions, actions, and shot types in the scene;

[0010] After receiving a video generation requirement text, generating a script through mutual cooperation between a plurality of intelligent agents based on an artificial intelligence (AI) large model, the AI large model being a text generation text AI large model, the plurality of intelligent agents being used to simulate a plurality of functional positions involved in an artificial video production process, the script including character profiles, dialogue texts, and dialogue annotation text information in the script, the dialogue annotation text information including character positions, actions, and shot types corresponding to the dialogue;

[0011] Simulating a video shooting process in the 3D virtual studio environment based on the structured script to generate a target video.

[0012] The script is generated through mutual cooperation between a plurality of intelligent agents based on an artificial intelligence (AI) large model, comprising:

[0013] By simulating an artificial video production process, the script generation process is divided into a plurality of continuous and interdependent stages, and the script is generated through mutual cooperation between the plurality of intelligent agents at each stage.

[0014] The plurality of functional positions involved in the artificial video production process simulated by the plurality of intelligent agents at least include a director, a screenwriter, and a photographer; and the plurality of continuous and interdependent stages include planning, script creation, and photography.

[0015] The script is generated through mutual cooperation between the plurality of intelligent agents at each stage, comprising:

[0016] In the planning stage, the video generation requirement text input by the user is provided to a director intelligent agent corresponding to a director position, so that the director intelligent agent plans the characters involved in the script and their profiles, as well as the script outline;

[0017] In the script creation stage, the planning results of the director intelligent agent and the information of the positions and action types in the 3D virtual studio environment are provided to a screenwriter intelligent agent corresponding to a screenwriter position, so that the screenwriter intelligent agent generates dialogue in the script and annotates the positions and action types of the characters for the dialogue in the script;

[0018] In the photography stage, the lines in the script and the corresponding role positions, action type information and shot type information in the 3D virtual film set environment are provided to the photographer agent corresponding to the photographer post, so as to mark the corresponding shot type for the lines in the script by the photographer agent.

[0019] Further comprising:

[0020] In the script creation stage, the generated script is subjected to a criticism, correction and verification cycle by the scriptwriter agent, so as to optimize the generated script.

[0021] The criticism of the script output by the director agent;

[0022] The scriptwriter agent performs a correction task to correct the script according to the criticism output by the director agent;

[0023] The director agent performs a verification task to verify the corrected script and determines whether further adjustment is needed.

[0024] The criticism of the script output by the director agent includes:

[0025] After the scriptwriter agent initially generates the script, the director agent performs a criticism task to comprehensively review the script and provide criticism opinions on the coherence of the plot and / or the appropriateness of the role positions and actions.

[0026] The plurality of functional posts further includes an actor;

[0027] The criticism of the script output by the director agent includes:

[0028] The role profile information generated by the director agent and the script generated by the scriptwriter agent are provided to the actor agent, so that the actor agent provides feedback on the script according to the understanding of the role, so that the script is consistent with the role profile;

[0029] The director agent generates criticism of the script by aggregating the feedback information generated by the actor agent.

[0030] The photographer agent is at least two;

[0031] In the photography stage, the lines in the script and the corresponding role positions, action type information and shot type information in the 3D virtual film set environment are provided to the at least two photographer agents, and the at least two photographer agents independently select the corresponding shot type for the lines in the script;

[0032] differences existing in the shot type selection results by executing a debate task;

[0033] a decision task is executed by the director agent to summarize the debate results of the at least two cinematographers and determine the final shot type selection result.

[0034] wherein further comprising:

[0035] using an AI large model of the text-to-audio class to generate corresponding speech for the lines in the script, and synchronizing the durations of the shots and actions in the generated video with the corresponding speech segments to synthesize the speech and the video.

[0036] A method for providing video content, comprising:

[0037] receiving a video generation requirement text input by a user;

[0038] submitting the video generation requirement text to a server, the server being configured to create a 3D virtual studio environment based on a 3D development engine, the 3D virtual studio environment including at least one virtual space scene, at least one virtual character image, and a plurality of positions, actions, and shot types in the scene, after receiving the video generation requirement text, generating a script through mutual cooperation between a plurality of agents based on an AI large model, the AI large model being a text-to-text AI large model, the plurality of agents being configured to simulate a plurality of functional positions involved in the process of artificially producing a video, the script including character profiles, line texts, and line corresponding character position, action, and shot type annotation text information in the script; structuring the script and simulating a video shooting process in the 3D virtual studio environment according to the structured script to generate a target video;

[0039] receiving the video generation result of the server and providing it to the user.

[0040] A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the method of any of the preceding items.

[0041] An electronic device, comprising:

[0042] one or more processors; and

[0043] a memory associated with the one or more processors, the memory being configured to store program instructions, which, when read and executed by the one or more processors, perform the steps of the method of any of the preceding items.

[0044] A computer program product comprising computer program / computer executable instructions to implement the steps of any of the preceding method when executed by a processor in an electronic device.

[0045] According to specific embodiments provided by the present disclosure, the present disclosure discloses the following technical effects:

[0046] According to the embodiments of the present disclosure, a video generation service can be provided, and through the service, a 3D (3 Dimensional) virtual studio environment can be created based on a 3D development engine, which can include at least one virtual space scene, at least one virtual character image, and a variety of positions, actions, and shot types in the scene. After receiving a video generation requirement text input by a user, a script can be generated through mutual cooperation between a plurality of intelligent agents based on an AI (Artificial Intelligence) large model, wherein the specific AI large model is a text generation text type AI large model, the plurality of intelligent agents are used to simulate a plurality of functional positions involved in the artificial video production process, and the script includes character profiles, dialogues, and dialogue annotation text information in the script. The dialogue annotation text information includes the character positions, actions, and shot types corresponding to the dialogue. Then, the script can be structured, and a video shooting process can be simulated in the 3D virtual studio environment according to the structured script to generate a target video. In this way, a "text-to-video" type large model is not directly used for video generation, but a "text-to-text" type AI large model is used to complete the generation of script dialogues and the annotation of character positions, actions, shot types, and other information about specific dialogues. Finally, the video is generated according to the annotation information in the structured script in the 3D development engine. This makes the video quality depend on the performance of the 3D development engine, so it is easier to ensure the video quality, and the video generation process is not limited by the time length, and does not involve the segmented generation and splicing of multiple videos, so the consistency of the characters in the video is guaranteed.

[0047] In the process of specifically generating the script, a cooperation strategy between the plurality of intelligent agents is also provided, which can specifically include a criticism-correction-verification and debate-judge strategy to optimize the script.

[0048] Of course, implementing any product of the present disclosure does not necessarily require all the advantages described above to be achieved at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only constitute some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0050] FIG. 1 is a schematic diagram of a system architecture provided by an embodiment of the present disclosure;

[0051] FIG. 2 is a flowchart of a server method provided by an embodiment of the present disclosure;

[0052] FIG. 3 is a schematic diagram of a 3D virtual studio environment provided by an embodiment of the present disclosure;

[0053] FIG. 4 is a script generation schematic diagram provided by an embodiment of the present disclosure;

[0054] FIG. 5 is a flowchart of a client method provided by an embodiment of the present disclosure;

[0055] FIG. 6 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments only constitute some embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art belong to the scope of protection of the present disclosure.

[0057] In order to facilitate understanding of the schemes provided by the embodiments of the present disclosure, the concept of an agent will be briefly introduced first. An agent is a concept that has existed for a very long time. Generally, it refers to an agent that can perceive the environment and take actions to achieve a specific goal. In other words, any existence that can interact with the real world and produce an impact can be called an agent. For example, some robots that can plan their own travel paths can be an agent. Some software that can automatically operate a mouse on a webpage to complete shopping operations for humans can also be called an agent, and so on. However, before the emergence of AI large models, the intelligence of an agent was usually achieved through reinforcement learning based on the knowledge of a specific field. This way made the agent only play a role in a specific field and could not realize cross-field capabilities. With the emergence of AI large models, their outstanding emergent capabilities led to the birth of “language agents”. The significant breakthrough of such agents based on AI large models lies in their ability to use language as the core medium for thinking and communication. This feature is close to the unique ability of humans, making language agents capable of cross-field capabilities.

[0058] This language-based intelligent agent framework has also provided a completely new approach to virtual production. However, virtual production is a highly complex task. If implemented using a single intelligent agent, it may still only be able to generate short videos, or it may be difficult to achieve good performance in terms of video quality and character consistency.

[0059] Based on the above, this embodiment of the disclosure constructs a virtual production framework based on multi-agent cooperation. Specifically, a 3D development engine (e.g., some 3D game engines) can first be used to create a 3D virtual film set environment. For example, multiple spatial scenes can be constructed, and various positions (where characters can appear), action types, and shot types can be defined, etc. After the user inputs specific video generation requirements (which can be expressed through text), multiple agents based on a large AI model can respectively play the roles of director, screenwriter, actor, cinematographer, etc., and the script is generated through the collaboration between these agents. The generated script can include required character introductions and specific lines, determining the corresponding character position, action type, shot type, etc., for each line. This information can then be structured and configured into the 3D virtual film set environment. The video shooting process is simulated using functions provided by the 3D development engine, and the target video is generated. In a preferred approach, the script generation process can be divided into multiple continuous and interdependent stages (including planning, scriptwriting, and cinematography) by simulating the efficient human on-set workflow. Then, the scriptwriting process can be completed through the collaboration of multiple intelligent agents at each stage.

[0060] This implementation scheme can be applied in various scenarios, especially suitable for creating IP (Intellectual Property) works (a general term encompassing all established cultural and creative works, including literature, film, animation, and games). For example, if an IP's animation has ended, but players or viewers may still want to watch derivative videos, a 3D virtual studio can be built, arranging key scenes and characters from the animation within the game engine. Players or viewers can input their requests, such as "I want to see a certain scene," and then the specific scene can be reenacted in the 3D virtual studio using the scheme provided in this disclosure.

[0061] From a system architecture perspective, referring to Figure 1, this embodiment of the disclosure can provide a video generation service, which includes a 3D virtual studio environment setup service and multiple intelligent agents. The users of this video generation service can be consumers, professionals in the video production industry, and so on. To facilitate user interaction, relevant service pages can be provided, or mobile terminal apps, mini-programs, or lightweight applications can be offered. In short, a specific page provides an entry point for users to input their specific video generation needs. If multiple different 3D virtual studio environments are pre-built, a list can be displayed on the page for users to choose from. Users can select the 3D virtual studio environment they are interested in and input their requirements, including the desired plot. Correspondingly, after receiving the user's video generation requirements, the video generation service can coordinate collaboration among various intelligent agents (including directors, screenwriters, actors, cinematographers, etc.) to complete specific script generation and annotation tasks, simulate the specific video shooting process in the 3D virtual studio environment, and finally return the generated target video to the user.

[0062] The specific implementation schemes provided by the embodiments of this disclosure will be described in detail below.

[0063] Example 1

[0064] First, this embodiment of the present disclosure provides a method for generating video from the perspective of the server side of the aforementioned video generation service. Referring to Figure 2, the method may specifically include:

[0065] S201: Create a 3D virtual film set environment based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions, and shot types in the scene.

[0066] In this embodiment, a 3D virtual film set environment can first be built based on a 3D development engine (e.g., a game engine like Unity). This 3D virtual film set environment can be described as a collection of one or more virtual space scenes. For example, a 3D virtual film set environment can contain 15 scenes, covering various everyday scenarios such as the living room, kitchen, office, and roadside. Of course, other types of 3D virtual film set environments can also be built, including more technologically advanced virtual space scenes, and so on.

[0067] Besides building a 3D virtual film set environment, various actor positions, action types, and shot types can be defined within the environment. Actor positions can be categorized into standing points and sitting positions, etc. For example, in the aforementioned lifelike 3D virtual film set environment, there could be 32 standing points and 33 sitting positions, etc. Each actor position can be numbered and described using text, specifically outlining its relative positional relationship to an object or other position in the scene. For example, position A: located next to the sofa; position B: between position C, etc.

[0068] Regarding the types of movements, this specifically refers to the various different movements an actor can perform, which can be defined separately as needed. For example, in the example above, each actor can perform 21 different movements, including basic movement movements such as sitting down and walking, as well as more expressive movements such as jumping, shaking their head angrily, and so on.

[0069] Regarding lens types, that is, the specific types of lenses that can be used when taking photos. For example, in the scenario mentioned above, nine lens types can be defined, including three static lenses shot from different distances (such as close-up, medium shot, and wide shot) and six dynamic lenses that track or surround a character (such as a lens that follows a character, a panning shot, a zoom lens, a curved lens, etc.).

[0070] Additionally, at least one virtual character (which can be a human or an animal, etc.) can be configured in the 3D virtual film studio environment. For example, five male characters and five female characters can be pre-created, and during video generation, these characters can be selected to perform specific actions. The specific character can be arbitrarily set, or, if it corresponds to a particular anime or similar work, the specific character can correspond to a character prototype from that work.

[0071] After defining the various character appearances, positions, action types, and camera types mentioned above, multiple actor positions and camera settings can be configured for each specific scene. For example, Figure 3 shows a vertical view of a living room scene in a 3D virtual film studio environment, which is configured with three positions (A, B, and C) and multiple camera types, etc.

[0072] S202: After receiving the video generation requirement text, a script is generated through the collaboration of multiple intelligent agents based on an artificial intelligence (AI) big model. The AI ​​big model is a text generation AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual production of videos. The script includes character introductions, dialogue text, and dialogue annotation text information. The dialogue annotation text information includes the character's position, actions, and shot type corresponding to the dialogue.

[0073] In addition to creating a 3D virtual film set environment, multiple intelligent agents can be pre-created to simulate various functional roles involved in the manual video production process. For example, at least intelligent agents can be included for the roles of director, screenwriter, and cinematographer. In a preferred implementation, intelligent agents for actors can also be included, and so on. Different intelligent agents can use the same base model; for example, they can all be the same large AI language model. A large language model refers to a deep learning model trained with a large amount of text data that can generate natural language text or understand the meaning of language text. That is, in this embodiment, a large model of "text-to-video" generation is not directly used; instead, a large model of "text-to-text (text-to-text)" generation is used to generate the script. The generated script annotates each line of dialogue with the actor's position, actions, and shot type (even during the filming stage, the cinematographer intelligent agent only needs to annotate the shot type for specific lines, without actually performing video generation processing). This information can then be structured and associated with functions in the 3D development engine. By calling specific functions in the 3D development engine, the information configured in the script can be applied to the specific 3D virtual film set environment, generating the corresponding target video.

[0074] Within this framework, the AI ​​agents corresponding to different functional roles can be used to perform the functions of those roles. For example, the director AI agent is responsible for initiating and supervising the entire video production project. The basic functions of this role include setting character introductions and planning the video outline. In addition, optionally, it can also provide feedback on the script, discuss with other staff members, and make final decisions in case of conflicts, and so on.

[0075] The screenwriter AI is responsible for working closely under the director's guidance. Specifically, the screenwriter is not only responsible for writing dialogue, but also for assigning character positions and action types to each line of dialogue. In the best-case scenario, the screenwriter can also continuously update the script based on the director's feedback to ensure it is coherent, engaging, and structurally sound.

[0076] The actor agent is optional; that is, the agent may not exist. If it does exist, the actor agent can be responsible for fine-tuning the lines based on its own character description to ensure that the dialogue is consistent with the character, and informing the director of any necessary modifications.

[0077] The cinematographer agent's function is to select camera settings for each line of dialogue based on the shot usage guidelines. In a preferred manner, these camera settings can also be compared and discussed with other cinematographer agents on set to ensure the appropriateness of the shot settings.

[0078] It should be noted that, as mentioned above, since the multiple agents in this embodiment are all used to perform "text-to-text" processing, the agents corresponding to different functional positions can use the same base model, that is, the same large language model of the same text-to-text category. However, under different functional requirements, different prompts can be input to different agents, so that each agent can complete its own different tasks.

[0079] Specifically, based on the generation of multiple intelligent agents corresponding to different functional positions, in a preferred implementation, the process of generating a video can be simulated by dividing the script generation process into multiple continuous and interdependent stages. For example, in this embodiment, it can be divided into three stages: planning, scriptwriting, and photography. Specific collaboration strategies are applied in these stages to enable collaboration between multiple different intelligent agents to jointly complete the scriptwriting process.

[0080] During the planning phase, the user-inputted video generation requirement text can be provided to the director's agent. The director agent then plans the characters and their brief introductions, as well as the script outline. Since the user-inputted video generation requirement text typically describes the required plot and can be considered a brief story idea, the director agent can start from this brief story idea and generate various character introductions related to the story. Specifically, character introductions can include key attributes such as gender, occupation, and personality traits. Based on these character introductions and multiple predefined scenes in a pre-created 3D virtual film set environment, the director agent can expand the initial story idea into a detailed scene outline, specifying the location, plot, and characters of each segment.

[0081] During the scriptwriting stage, the planning results of the director's AI agent, along with information such as positioning and action types in the 3D virtual film set environment (all of which can exist in text form), can be provided to the screenwriter's AI agent. This allows the screenwriter's AI agent to generate dialogue in the script and assign character positioning and action types to the dialogue. Since the initial script generated by the screenwriter's AI agent may have some flaws, in a preferred embodiment of this disclosure, multiple AI agents can collaborate to complete the scriptwriting process.

[0082] To achieve collaboration among different agents, embodiments of this disclosure also provide two collaboration strategies: a critique-correct-verify strategy and a debate-judge strategy. In the critique-correct-verify strategy, the action agent P generates a response R based on a given context C and instruction I. Then, the critique agent Q reviews the response R and writes critiques F, pointing out potential areas for improvement. Next, the action agent P integrates these critiques and corrects the response. Finally, the critique agent Q evaluates the updated response R to determine whether the critiques have been adequately addressed or whether further iterations are needed.

[0083] The debate-referee collaboration refers to a process in which two or more peer agents M and N independently generate their respective responses in each iteration and critique each other's work. Based on the criticisms received, each agent may modify its response or maintain the original response. After several rounds of debate, the referee agent J summarizes the discussion and makes the final decision. It should be noted that, in this embodiment, although different agents may use the same base model, the randomness inherent in the content generated by large AI models means that even the same model under the same input conditions may generate different content. Therefore, the generated results of two or more peer agents may not be entirely the same. By comparing the differences and debating them, it is beneficial to output higher-quality results.

[0084] Specifically, during the scriptwriting stage, the aforementioned critique-revision-verification strategy can be employed. First, a screenwriting AI agent can write a preliminary script, including character dialogue, positions, and action types. Then, the script generated by the screenwriting AI agent can be critiqued, revised, and verified in a loop to optimize it. Specifically, a director AI agent can generate critiques of the script, the screenwriting AI agent can perform revision tasks to correct the script based on the director's feedback, and then the director AI agent can perform verification tasks to validate the revised script and determine if further adjustments are needed.

[0085] In this context, the director's AI agent can generate critiques of the script in various ways. For example, one approach is for the director's AI agent to perform the critique task after the screenwriter's AI agent has initially generated the script. This allows for a comprehensive review of the script, providing feedback on plot coherence and / or the appropriateness of character positioning and actions. For instance, suppose the initially generated script includes the line "(Standing and thinking) Wait, what?", where the screenwriter's AI agent labels the emotion as "thinking." However, the director's AI agent, through comprehensive review, might find that this line should reflect confusion or surprise. The screenwriter's AI agent could then revise the line to: "(Standing in surprise) Wait, what?".

[0086] Alternatively, if actor agents also exist, the character descriptions generated by the director agent and the script generated by the screenwriter agent can be provided to the actor agents. This allows the actor agents to provide feedback on the script based on their understanding of the characters, ensuring the script aligns with the character descriptions. Since there may be multiple characters, multiple actor agents can be assigned to each role, each providing feedback on the script from their own perspective. The director agent can then aggregate this feedback to generate critiques of the script.

[0087] Of course, in practice, the two methods mentioned above can be combined. That is, the director's agent and the screenwriter's agent can first conduct one or more criticism-correction-verification cycles, and then the actor's agent can provide feedback based on their understanding of the role. The director's agent can then filter and summarize this feedback, and then work with the screenwriter's agent to optimize the script through the same criticism-correction-verification cycle.

[0088] During the filming phase, the script's dialogue, corresponding character positions and action types, as well as shot type information from the 3D virtual film set environment, can be provided to the cinematographer's AI agent. This allows the AI ​​agent to label the dialogue with the appropriate shot type. Alternatively, specific shot usage guidelines can be provided to the AI ​​agent, enabling it to complete the task of labeling the dialogue with shot types under this more specialized guidance.

[0089] As mentioned earlier, during the photography phase, it is not a true conversion of text into video, but rather the labeling of dialogue with shot types. This labeling information still exists in the form of text. Therefore, for the photography intelligent agent, the large AI model it uses can still be a large language model that generates text from text.

[0090] Similar to the potential flaws in the initial script generation by the screenwriter AI, the cinematographer AI may also make inappropriate shot selections when labeling shot type information. Therefore, a preferred approach is to employ a collaborative strategy among multiple AI agents to optimize shot type selection. Specifically, in this shooting stage, a debate-judge collaboration strategy can be used. At least two cinematographer AI agents can be provided with the script's dialogue, corresponding character positions and action types, and shot type information from the 3D virtual film set environment. Each of the at least two cinematographer AI agents independently selects the corresponding shot type for the script's dialogue. Then, the at least two cinematographer AI agents can resolve discrepancies in the shot type selection results by performing a debate task. For example, for a particular line of dialogue, if the first cinematographer AI agent selects a medium shot and the second selects a zoom shot, they can debate, each explaining their reasons for choosing the corresponding shot. Finally, the director AI agent can perform a judgment task to summarize the debate results of the at least two cinematographers and determine the final shot type selection result.

[0091] S203: By structuring the script and simulating the video shooting process in the 3D virtual film set environment based on the structured script, a target video is generated.

[0092] After completing the character setup and generating the script, the specific lines, character positions, actions, and shot types corresponding to the lines have been determined. For example, as shown in Figure 4, it illustrates the various functional agents and their corresponding processing results at each stage in a specific example. It can also be clearly seen from the figure that these lines and annotations are described through text (the images corresponding to each line or dialogue in Figure 4 are only used to illustrate the images after the final video is generated; the generation of these images is not included in the script generation process using agents). Therefore, this script text can be structured so that the positions, actions, and shots in the structured script information can correspond to functions in the 3D development engine. This allows the video shooting process to be simulated in a 3D virtual film set environment by calling functions in the 3D development engine, and the target video can be generated. The generated target video can be a 2D video. In other words, the video in this embodiment is not directly generated by the AI ​​model. The AI ​​model is only responsible for generating the script lines and annotating the text with information such as character positions, actions, and shot types for specific lines. Finally, the 3D development engine generates the 2D video based on the annotated information in the structured script. The video quality depends on the performance of the 3D development engine, thus making it easier to ensure video quality. In addition, the video generation process is not limited by duration, so it does not involve the segmented generation and splicing of multiple videos, thereby ensuring the consistency of the characters in the video.

[0093] It's worth noting that in practical applications, since the script includes dialogue, speech generation may also be involved. Specifically, to create more natural and expressive audio, a large AI model for "text-to-audio" generation can be used to generate the speech for each line of dialogue. The duration of each shot and action in the video can be synchronized with the corresponding audio segment; that is, the duration of each line of dialogue in the video can be determined by the length of the corresponding audio segment. Finally, the generated video and audio are combined to produce the final video content, which is then returned to the user.

[0094] As can be seen, this disclosure systematically studies automated virtual film production from multiple aspects, including basic scene construction, the proposal of a full-process method, and evaluation methods, providing a solid foundation for the intelligent and efficient transformation of the future film production field. From a technical perspective, this disclosure uses a virtual film set based on a game physics engine to shoot videos, effectively solving problems such as missing plots, low video quality, violations of common sense, and inconsistencies in characters that exist in the current field of textual content creation videos. Videos generated under the multi-agent framework surpass the previous single-agent baseline in terms of script fluency, character movements, dialogue, and camera settings, demonstrating the effectiveness of the framework. Furthermore, through manual comparison of the script and camera settings before and after agent collaboration, it was found that the results after collaboration were more popular, demonstrating the important role of multi-agent collaboration in refining scripts and improving decision-making. In the virtual film production scenario, the collaboration of multiple agents based on a large textual content creation AI model can surpass the performance of a more capable single model, demonstrating that weaker models can achieve or even surpass the performance of stronger models by reasonably arranging the agent workflow.

[0095] In summary, this disclosure provides a video generation service that enables the creation of a 3D virtual film set environment based on a 3D development engine. This environment may include at least one virtual space scene, at least one virtual character, and various positions, actions, and shot types within the scene. Upon receiving a user's input text requesting video generation, a script can be generated through collaboration among multiple intelligent agents based on an AI (Artificial Intelligence) model. Specifically, this AI model is a text-to-text AI model. The multiple intelligent agents simulate various functional roles involved in the manual video production process. The script includes character introductions, dialogue, and dialogue annotation text information. The dialogue annotation text information includes the character's position, action, and shot type corresponding to the dialogue. Subsequently, the script can be structured, and the video shooting process can be simulated within the 3D virtual film set environment based on the structured script to generate the target video. This approach avoids directly using large "text-to-video" models for video generation. Instead, it employs large "text-to-text" AI models to generate script dialogue and annotate information such as character positioning, actions, and shot types for specific lines. Finally, the video is generated within the 3D development engine based on the annotated information in the structured script. This makes video quality dependent on the performance of the 3D development engine, thus ensuring higher quality. Furthermore, the video generation process is not limited by duration and does not involve segmenting and stitching multiple videos, ensuring consistency among characters within the video.

[0096] In the process of generating the script, a collaborative strategy among multiple agents is provided, which may include the Critique-Correct-Verify and Debate-Judge strategies to optimize the script.

[0097] Example 2

[0098] This second embodiment corresponds to the first embodiment described above. From the client's perspective, it provides a method for providing video content. Referring to Figure 5, the method may include:

[0099] S501: Receive user input for video generation request text;

[0100] S502: The video generation requirement text is submitted to the server. The server is used to create a 3D virtual film set environment based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions, and shot types in the scene. After receiving the video generation requirement text, the server generates a script through the cooperation of multiple intelligent agents based on an AI big model. The AI ​​big model is a text-to-text AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual video production process. The script includes character introductions, dialogue text, and text information annotating the character positions, actions, and shot types corresponding to the dialogue. The script is structured, and the video shooting process is simulated in the 3D virtual film set environment according to the structured script to generate the target video.

[0101] S502: Receives the video generation results from the server and provides them to the user.

[0102] For the parts of this embodiment that are not described in detail, please refer to the description in embodiment one and other parts of this specification, which will not be repeated here.

[0103] It should be noted that the embodiments disclosed herein may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country in which it is located (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0104] Corresponding to Embodiment 1, this disclosure also provides an apparatus for generating video, which may include:

[0105] The 3D virtual film set environment creation unit is used to create a 3D virtual film set environment based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions, and shot types in the scene.

[0106] The script generation unit is used to generate a script through the collaboration of multiple intelligent agents based on an artificial intelligence (AI) big model after receiving the video generation requirement text. The AI ​​big model is a text generation AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual production of videos. The script includes character introductions, dialogue text, and dialogue annotation text information. The dialogue annotation text information includes the character's position, actions, and shot type corresponding to the dialogue.

[0107] The video generation unit is used to generate a target video by structuring the script and simulating the video shooting process in the 3D virtual film set environment based on the structured script.

[0108] In specific implementation, the script generation unit can be used for:

[0109] By simulating the process of producing videos manually, the script generation process is divided into multiple continuous and interdependent stages, and the script is generated through the cooperation between the multiple intelligent agents at each stage.

[0110] The multiple functional positions involved in the artificial video production process simulated by the multiple intelligent agents include at least: director, screenwriter, and cinematographer; the multiple continuous and interdependent stages include: planning, scriptwriting, and cinematography.

[0111] At this point, the script generation unit can be specifically used for:

[0112] During the planning phase, the video generation requirement text input by the user is provided to the director's AI agent corresponding to the director position, so that the director AI agent can plan the characters involved in the script and their introductions, as well as the script outline.

[0113] During the scriptwriting stage, the planning results of the director's intelligent agent and the information on the standing position and action type in the 3D virtual film set environment are provided to the screenwriter's intelligent agent, so that the screenwriter's intelligent agent can generate the lines in the script and mark the characters' standing position and action type for the lines in the script.

[0114] During the filming stage, the script's lines, corresponding character positions, action types, and shot types in the 3D virtual film set environment are provided to the cinematographer's AI agent, so that the cinematographer AI agent can label the script's lines with the corresponding shot types.

[0115] In addition, to further enhance the quality of the script, the device may also include:

[0116] The optimization processing unit is used to optimize the script generated by the screenwriter agent in the script creation stage by criticizing, revising and verifying the script in a loop.

[0117] Among these, the director's AI agent outputs critical opinions about the script;

[0118] The scriptwriting agent performs the revision task to revise the script based on the criticisms output by the director agent.

[0119] The director agent performs a verification task to validate the revised script and determine whether further adjustments are needed.

[0120] Specifically, after the scriptwriting agent initially generates the script, the director agent can perform a critique task to conduct a comprehensive review of the script and provide criticisms on plot coherence and / or character positioning and appropriateness of actions.

[0121] Alternatively, the multiple functional positions may also include: actor; in this case, the character introduction information generated by the director's agent and the script generated by the screenwriter's agent can be provided to the actor's agent so that the actor's agent can provide feedback on the script based on its understanding of the character, so that the script is consistent with the character introduction; then, the director's agent can generate criticisms of the script by summarizing the feedback information generated by the actor's agent.

[0122] Additionally, there can be at least two cinematographer agents. In this case, during the filming phase, the script's dialogue, corresponding character positions and action types, as well as shot type information in the 3D virtual film set environment, can be provided to the at least two cinematographer agents. Each of the at least two cinematographer agents can independently select the corresponding shot type for the script's dialogue. Then, the at least two cinematographer agents can resolve the differences in the shot type selection results by performing a debate task. Finally, the director agent can perform a ruling task to summarize the debate results of the at least two cinematographers and determine the final shot type selection result.

[0123] Furthermore, the device may also include:

[0124] The speech generation unit is used to generate corresponding speech for the lines in the script using a large AI model that generates audio from text, and to synchronize the duration of shots and actions in the generated video with the corresponding speech segments so as to synthesize the speech and video.

[0125] Corresponding to Embodiment 2, this disclosure also provides an apparatus for providing video content, which may include:

[0126] The requirement text receiving unit is used to receive video input from the user and generate requirement text.

[0127] The submission unit is used to submit the video generation requirement text to the server. The server is used to create a 3D virtual film set environment based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions, and shot types in the scene. After receiving the video generation requirement text, the server generates a script through the cooperation of multiple intelligent agents based on an AI big model. The AI ​​big model is a text-to-text AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual video production process. The script includes character introductions, dialogue text, and text information annotating the character positions, actions, and shot types corresponding to the dialogue. By structuring the script, the server simulates the video shooting process in the 3D virtual film set environment based on the structured script to generate the target video.

[0128] The video return unit is used to receive the video generation results from the server and provide them to the user.

[0129] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0130] And an electronic device, comprising:

[0131] One or more processors; and

[0132] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0133] A computer program product includes a computer program / computer executable instructions that, when executed by a processor in an electronic device, implement the steps of the method described in the foregoing method embodiments.

[0134] Figure 6 illustrates the architecture of an electronic device, which may include a processor 610, a video display adapter 611, a disk drive 612, an input / output interface 613, a network interface 614, and a memory 620. The processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, and memory 620 can communicate with each other via a communication bus 630.

[0135] The processor 610 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this disclosure.

[0136] The memory 620 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 620 can store the operating system 621 for controlling the operation of the electronic device 600, and the basic input / output system (BIOS) for controlling the low-level operations of the electronic device 600. Additionally, it can store a web browser 623, a data storage management system 624, and a video generation and processing system 625, etc. The aforementioned video generation and processing system 625 can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when the technical solution provided by this disclosure is implemented through software or firmware, the relevant program code is stored in the memory 620 and is called and executed by the processor 610.

[0137] Input / output interface 613 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0138] Network interface 614 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0139] Bus 630 includes a pathway for transmitting information between various components of the device, such as processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, and memory 620.

[0140] It should be noted that although the above-described device only shows the processor 610, video display adapter 611, disk drive 612, input / output interface 613, network interface 614, memory 620, bus 630, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the present disclosure, and need not include all the components shown in the figures.

[0141] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this disclosure can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this disclosure.

[0142] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0143] The method and electronic device for generating video provided in this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this disclosure, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this disclosure.

Claims

1. A method for generating video, wherein, include: A 3D virtual film set environment is created based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions and shot types in the scene. After receiving the video generation requirement text, a script is generated through the collaboration of multiple intelligent agents based on an artificial intelligence (AI) big model. The AI ​​big model is a text generation AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual production of videos. The script includes character introductions, dialogue text, and dialogue annotation text information. The dialogue annotation text information includes the character's position, actions, and shot type corresponding to the dialogue. The target video is generated by structuring the script and simulating the video shooting process in the 3D virtual film set environment based on the structured script.

2. The method according to claim 1, wherein, The script generation process, which involves collaboration among multiple intelligent agents based on a large AI model, includes: By simulating the process of producing videos manually, the script generation process is divided into multiple continuous and interdependent stages, and the script is generated through the cooperation between the multiple intelligent agents at each stage.

3. The method according to claim 2, wherein, The multiple functional roles involved in the artificial video production process simulated by the multiple intelligent agents include at least: director, screenwriter, and cinematographer; the multiple continuous and interdependent stages include: planning, scriptwriting, and cinematography; The generation of scripts through the cooperation among the multiple intelligent agents at each stage includes: During the planning phase, the video generation requirement text input by the user is provided to the director's AI agent corresponding to the director position, so that the director AI agent can plan the characters involved in the script and their introductions, as well as the script outline. During the scriptwriting stage, the planning results of the director's intelligent agent and the information on the standing position and action type in the 3D virtual film set environment are provided to the screenwriter's intelligent agent, so that the screenwriter's intelligent agent can generate the lines in the script and mark the characters' standing position and action type for the lines in the script. During the filming stage, the script's lines, corresponding character positions, action types, and shot types in the 3D virtual film set environment are provided to the cinematographer's AI agent, so that the cinematographer AI agent can label the script's lines with the corresponding shot types.

4. The method according to claim 3, wherein, Also includes: During the scriptwriting stage, the script generated by the screenwriter AI is repeatedly criticized, revised, and verified in order to optimize the generated script. Among these, the director's AI agent outputs critical opinions about the script; The scriptwriting agent performs the revision task to revise the script based on the criticisms output by the director agent. The director agent performs a verification task to validate the revised script and determine whether further adjustments are needed.

5. The method according to claim 4, wherein, The criticisms of the script output by the director's AI agent include: After the scriptwriting agent initially generates the script, the director agent performs a critique task to conduct a comprehensive review of the script and provide criticism on plot coherence and / or character positioning and appropriateness of actions.

6. The method according to claim 4, wherein, The aforementioned multiple job positions also include: actor; The criticisms of the script output by the director's AI agent include: The character profile information generated by the director's intelligent agent and the script generated by the screenwriter's intelligent agent are provided to the actor's intelligent agent so that the actor's intelligent agent can provide feedback on the script based on its understanding of the character, so that the script is consistent with the character profile. The director's AI agent generates critiques of the script by aggregating feedback from the actor's AI agent.

7. The method according to any one of claims 3 to 6, wherein, The photographer's agent must consist of at least two agents; During the filming stage, the script's lines, corresponding character positions, action types, and shot types in the 3D virtual film set environment are provided to at least two cinematographer agents, who then independently select the corresponding shot types for the script's lines. The differences in lens type selection results are handled by the at least two photographer agents through a debate task. The director agent performs the adjudication task to summarize the results of the debate between the at least two cinematographers and determine the final shot type selection.

8. The method according to claim 7, characterized in that, The process of handling differences in lens type selection results by the at least two photographer agents through a debate task includes: Each cinematographer agent generates selection criteria for different shots based on a lens language specification library. These criteria include plot suitability analysis, visual emphasis logic, and explanation of the rationality of shot transitions. Multi-round debates between photographer agents are realized through a natural language interaction interface, and the points of disagreement and the basis for the arguments are automatically recorded during the debate. When the number of debate rounds reaches a preset threshold or the point of disagreement is eliminated, the debate ends and a text of the debate process is generated.

9. The method according to any one of claims 3 to 8, characterized in that, The aforementioned functional positions also include an art direction intelligent agent; Between the planning stage and the scriptwriting stage, there is also an art design stage: Provide the script outline, scene requirements, and basic information of the 3D virtual film set environment output by the director's AI agent to the art director's AI agent; Based on the scene style positioning, the art director AI agent generates scene decoration schemes, suggestions for adjusting character appearance parameters, and color tone annotations, and feeds the schemes and suggestions back to the director AI agent; After being verified by the director's AI agent, it is synchronized to the configuration module of the 3D virtual film set environment and the screenwriter's AI agent to guide the creation of script details.

10. The method according to any one of claims 1 to 9, wherein, Also includes: A large AI model for text-to-audio generation is used to generate corresponding audio for the lines in the script. The duration of shots and actions in the generated video is synchronized with the corresponding audio segments to synthesize the audio and video.

11. The method according to claim 10, characterized in that, The AI ​​model that uses text to generate audio to generate corresponding audio for the lines in the script also includes: Configure a set of voice feature parameters for each character, which includes timbre, speech rate, tone, and emotional fluctuation threshold; After receiving the user's customization request for the character's voice, adjust the corresponding character's voice feature parameters; The emotional adaptation interface of the text-to-audio AI model is used to input the dialogue text and adjusted parameters into the model to generate audio segments with emotional tags, which match the character's actions and emotions marked in the script.

12. The method according to any one of claims 1 to 11, characterized in that, The creation of a 3D virtual film studio environment based on a 3D development engine also includes: Establish a virtual prop resource library and an environmental dynamic parameter system. The virtual prop resource library includes at least one interactive virtual object and object attribute parameters. The environmental dynamic parameter system covers light intensity, weather effects, and spatial sound effect parameters. After receiving the user's instructions to select virtual props and adjust environmental parameters, the system loads the selected virtual props into the target virtual space scene by calling the attribute configuration interface of the 3D development engine, and adjusts the dynamic environmental parameters in real time to match the video creation atmosphere.

13. The method according to any one of claims 1 to 12, characterized in that, The process of structuring the script includes: The script is parsed into structured data frames containing scene identifiers, character IDs, dialogue content, and annotation information, where the annotation information is associated with function call parameters of the 3D development engine; A mapping relationship between data frame sequences and timelines is established, and a timestamp is configured for each data frame. The timestamp is generated based on the estimated duration of dialogue and the execution cycle of actions. The integrity of structured data frames is checked. If there is missing annotation information or parameter mismatch, the screenwriter agent or cinematographer agent is triggered to supplement and correct it.

14. The method according to any one of claims 1 to 13, characterized in that, After generating the target video, the following is also included: A video editing interface is generated, which includes scene jump markers, character action timeline, camera switching nodes, and voice adjustment controls. Receive modification instructions input by the user through the interface, and parse the target modification type and parameters corresponding to the instructions; If the modification type is a shot or motion adjustment, the cinematographer's agent or screenwriter's agent is invoked to regenerate the annotation information; if the modification type is a voice or scene parameter adjustment, the corresponding AI model or 3D engine interface is directly triggered to update the parameters, and the target video clip is regenerated based on the modified content to replace the original clip.

15. The method according to any one of claims 1 to 14, characterized in that, It also includes: establishing an evaluation model for the collaborative effect of intelligent agents, with script approval rate, shot selection accuracy rate, and number of modification iterations as the core evaluation indicators; After each video generation process is completed, the output results, collaborative interaction data and user feedback information of each intelligent agent are automatically collected. The collected data is input into the evaluation model to generate an agent performance score and suggestions for optimizing the collaboration strategy. Based on the suggestions, the agent's prompt word template and collaboration trigger conditions are updated.

16. A method for providing video content, wherein, include: Receive video input from the user and generate a request text; The video generation requirement text is submitted to the server. The server is used to create a 3D virtual film set environment based on a 3D development engine. The 3D virtual film set environment includes at least one virtual space scene, at least one virtual character image, and various positions, actions, and shot types in the scene. After receiving the video generation requirement text, the server generates a script through the cooperation of multiple intelligent agents based on an AI big model. The AI ​​big model is a text-to-text AI big model. The multiple intelligent agents are used to simulate multiple functional positions involved in the manual video production process. The script includes character introductions, dialogue text, and text information annotating the character's position, actions, and shot types corresponding to the dialogue. The script is structured, and the video shooting process is simulated in the 3D virtual film set environment based on the structured script to generate the target video. Receive the video generation results from the server and provide them to the user.

17. The method according to claim 16, characterized in that, The method of receiving the video generation request text input by the user also includes: A structured guide interface for request text is provided, which includes scene type options, number of characters input box, plot style tag set and key plot description area; The natural language understanding model performs semantic parsing of user input. If there is ambiguity in the requirements or logical conflict, it generates guiding follow-up questions. The parsed structured requirement text is matched with the user's historical creation preference data to recommend suitable 3D virtual film set environment and character image combination.

18. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the program performs the steps of the method described in any one of claims 1 to 17.

19. An electronic device, wherein, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 17.

20. A computer program product comprising a computer program / computer-executable instructions, wherein, When the computer program / computer-executable instructions are executed by a processor in an electronic device, they implement the steps of the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Video processing method and device, storage medium and computer equipment

    CN114697730A

  • Virtual reality content making system based on artificial intelligence image recognition

    CN116489203A

  • Language model role playing-based long script automatic generation method

    CN118394926A

  • Method and system for generating video by means of three-dimensional rendering based on text information

    CN118803387A

  • Video generation method and electronic equipment

    CN119600156A

Cited By

  • AI Agent-based workflow decision-making and execution method, device, and storage medium for text graphs.

    CN122156338A