Intelligent agent-based video generation system and method
By utilizing a video generation system based on intelligent agents, and leveraging virtual dialogue between user agent agents and assistant agents, as well as external function libraries, the system solves the problem of large language models being unable to efficiently generate complex videos. It achieves high-quality video generation and efficient content generation, and is applicable to fields such as video editing, advertising generation, and content creation.
Patent Information
- Application Number
- CN202411476530.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-10-22
AI Technical Summary
Existing Large Language Models (LLMs) cannot independently complete complex and customized video generation tasks, resulting in low video generation efficiency and inaccurate results, making it difficult to meet professional-level video generation needs.
A video generation system based on intelligent agents is adopted. Through virtual dialogue between the user agent and the assistant agent, the system uses external function libraries to break down and refine the text and video requirements. The virtual dialogue between the user agent and the assistant agent converts the user's text and video requirements into structured instructions, making full use of external function libraries to improve the quality of video generation.
It improves the quality and efficiency of video generation, reduces the number of human-computer interactions, and enhances the human-computer interaction experience. It is applicable to fields such as video editing, advertising generation, and content creation, providing efficient and accurate content generation services.
Smart Images

Figure CN119364105B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video generation technology, and in particular to a video generation system and method based on intelligent agents. Background Technology
[0002] With the rise of Large Language Models (LLMs) such as ChatGenerative Pre-trained Transformer (ChatGPT), users' demand for rapidly generating personalized videos is increasing. However, LLMs have limited built-in functionality and cannot independently complete highly complex and customized tasks, making it difficult to meet professional-level video generation needs. This results in low efficiency and inaccurate video generation results when handling complex tasks. Summary of the Invention
[0003] The purpose of this application is to provide a video generation system and method based on intelligent agents, which can improve the quality of video generation.
[0004] To achieve the above objectives, this application provides the following solution:
[0005] In a first aspect, this application provides a video generation system based on intelligent agents, the video generation system based on intelligent agents comprising: a user agent intelligent agent, an assistant intelligent agent, and an external function library; the user agent intelligent agent and the assistant intelligent agent are both intelligent agents based on a large language model; the user agent intelligent agent and the assistant intelligent agent both import functions from the external function library into their own function library;
[0006] The user agent is used to send the user's text and video requests to the assistant agent.
[0007] The assistant agent is used to select a function from the function library according to the text and video requirements, the execution result returned by the user agent agent, or the next work command, and send the selected function and its corresponding parameters to the user agent agent.
[0008] The user agent is used to execute the function selected by the assistant agent and return the execution result of the function to the assistant agent;
[0009] The assistant agent is also used to determine whether to continue executing the function based on the execution result returned by the user agent agent. If yes, it continues to select a function from the function library based on the current execution result returned by the user agent agent. Otherwise, the user agent agent determines whether to proceed to the next step until the final video is generated. If the user agent agent determines to proceed to the next step, it sends the next step command to the assistant agent.
[0010] Optionally, the video generation system based on intelligent agents further includes an intelligent agent framework, which is used to record the virtual dialogue content between the user agent and the assistant agent, the virtual dialogue content including function call chains.
[0011] Optionally, the video generation system based on intelligent agents further includes a result output module; the result output module is used to output the final answer, function directory and function call chain of the assistant agent, wherein the final answer is the generated final video and the function directory is a function directory consisting of functions executed by the user agent agent in the process of generating the final video.
[0012] Optionally, the intelligent agent-based video generation system further includes a question acquisition module, which is used to acquire the text video requirements input by the user and the selected large language model.
[0013] Optionally, the large language model is ChatGPT, LLaMA, Doubao Large Model, or Tongyi Qianwen.
[0014] Optionally, the external function library includes an information crawling module, a text-to-speech module, a template editing module, a template filling module, a lip alignment module, and a video enhancement module.
[0015] Secondly, this application provides a video generation method based on intelligent agents, the video generation method based on intelligent agents including:
[0016] The user agent sends the user's text and video requests to the assistant agent; both the user agent and the assistant agent are agents based on a large language model; both the user agent and the assistant agent import functions from the external function library into their own function library;
[0017] The assistant agent selects a function from the function library based on the text / video requirements, or the execution result or next step command returned by the user agent agent, and sends the selected function and its corresponding parameters to the user agent agent.
[0018] The user agent executes the function selected by the assistant agent and returns the execution result of the function to the assistant agent;
[0019] The assistant agent determines whether to continue executing the function based on the execution result returned by the user agent. If yes, it selects a function from the function library based on the current execution result returned by the user agent, sends the selected function and its corresponding parameters to the user agent, and returns the step of the user agent executing the function selected by the assistant agent and returning the execution result of the function to the assistant agent. Otherwise, the user agent determines whether to proceed to the next step. If the user agent determines to proceed to the next step, it sends the next step command to the assistant agent.
[0020] The assistant agent selects a function from the function library according to the next work command, sends the selected function and its corresponding parameters to the user agent, and returns the execution result of the function to the assistant agent until the final video is generated.
[0021] Optionally, the video generation method based on intelligent agents further includes: recording virtual dialogue content between the user agent and the assistant agent through an intelligent agent framework, wherein the virtual dialogue content includes function call chains.
[0022] Optionally, after generating the final video, the video generation method based on intelligent agents further includes:
[0023] Output the final answer, function directory, and function call chain of the assistant agent. The final answer is the generated final video, and the function directory is a list of functions executed by the user agent agent during the generation of the final video.
[0024] Optionally, before the user agent sends the user's text video request to the assistant agent, the video generation method based on the intelligent agent further includes: obtaining the user's input text video request and the selected large language model.
[0025] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0026] This application provides a video generation system and method based on intelligent agents. Both the user agent and the assistant agent are intelligent agents based on large language models. By utilizing virtual dialogue between the user agent and the assistant agent, the user's text video requirements are converted into structured instructions, thereby realizing the breakdown and refinement of text video requirements, making full use of external function libraries, and improving the quality of video generation. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This application provides a schematic diagram of the workflow of a video generation system based on intelligent agents, as an embodiment of the present application.
[0029] Figure 2 This is a schematic diagram of the collaboration process of various modules in an external function library provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] This application provides a video generation system based on intelligent agents. The system includes a user agent, an assistant agent, and an external function library. Both the user agent and the assistant agent are agents based on a large language model. The computing power of both the user agent and the assistant agent comes from the large language model. Both the user agent and the assistant agent import functions from the external function library into their own function library.
[0033] The user agent is used to send the user's text and video requests to the assistant agent.
[0034] The assistant agent is used to select a function from the function library according to the text / video requirements, the execution result returned by the user agent agent, or the next work command, and send the selected function and its corresponding parameters to the user agent agent.
[0035] The user agent is used to execute the function selected by the assistant agent and return the execution result of the function to the assistant agent.
[0036] The assistant agent is also used to understand the large language model and determine whether to continue executing the function based on the execution result returned by the user agent. If yes, it continues to select functions from the function library based on the current execution result returned by the user agent. Otherwise, the user agent determines whether to proceed to the next step until the final video is generated. If the user agent determines to proceed to the next step, it sends the next step command to the assistant agent.
[0037] User agent agents are used to control the start and end of virtual dialogue interactions between assistant agents and user agent agents.
[0038] In this application, the user agent and assistant agent generate the final video based on the user's text and video requests through virtual dialogue interaction.
[0039] Both the user agent and the assistant agent in this application are based on a large language model. By utilizing the virtual dialogue between the user agent and the assistant agent, the user's text and video requests are converted into structured instructions, thereby breaking down and refining the text and video requests, making full use of external function libraries, and improving the quality of video generation.
[0040] like Figure 1 As shown, the workflow of the video generation system based on intelligent agents in this application is as follows:
[0041] Step 101: Obtain the video requirements and selected LLM input by the user, initialize the system and create two agents: the user agent agent and the assistant agent, and provide the external function library to the two agents for access.
[0042] The code for creating an assistant agent by the intelligent agent is shown in Table 1.
[0043] Table 1 Creating Assistant Agents
[0044]
[0045] The code for creating the user agent (user_proxy) by the smart agent is shown in Table 2.
[0046] Table 2 Creating Assistant Agents
[0047]
[0048]
[0049] After creating the user agent and assistant agent, the function registration process begins. First, external function libraries are imported, and then the functions from those libraries are registered with both the user agent and assistant agent. This allows the assistant agent to call these functions when needed, while the user agent is responsible for executing them and returning the results.
[0050] Step 102: The user agent initiates a dialogue, sending the user's text and video requests to the assistant agent for analysis.
[0051] Step 103: The assistant agent analyzes and understands the information provided by the user agent, and automatically selects appropriate functions and parameters from the external function library as needed, and sends them to the user agent for execution.
[0052] Step 104: The user agent executes the function and parameters provided by the assistant agent and provides the function return value to the assistant agent. The assistant agent determines whether to continue executing the function. If yes, the function is executed according to the methods in steps 103 and 104; otherwise, the user agent determines the next step.
[0053] In this application, the user agent initiates a dialogue based on the video requirements input by the user. The assistant agent selects appropriate functions and parameters based on the dialogue content, executes these functions, and returns the results. The entire dialogue process and the function call chain are recorded to facilitate debugging and optimization by technical personnel. The assistant agent automatically selects and executes functions from a defined external function library based on the text and video requirements to process the question content.
[0054] Step 105: Use the intelligent agent framework to record the virtual dialogue content between the assistant agent and the user agent agent. The virtual dialogue content includes available functions and function call chains.
[0055] Step 106: The user agent detects that the assistant agent has completed video generation and ends the virtual dialogue; the intelligent agent framework outputs the generated final video, function directory, and function call chain.
[0056] In one exemplary embodiment, the intelligent agent-based video generation system further includes an intelligent agent framework for recording virtual dialogue content between the user agent and the assistant agent, the virtual dialogue content including function call chains. During system initialization, the intelligent agent framework transforms external function libraries into function libraries executable by the agent and provides them to the agent during agent initialization.
[0057] In an exemplary embodiment, the video generation system based on intelligent agents further includes a result output module; the result output module is used to output the final answer, function directory and function call chain of the assistant agent, wherein the final answer is the generated final video and the function directory is a function directory consisting of functions executed by the user agent agent in the process of generating the final video.
[0058] In an exemplary embodiment, the video generation system based on intelligent agents further includes a question acquisition module. The question acquisition module is used to acquire the text video requirements input by the user and the selected large language model after system initialization, and provide the text video requirements and large language model to the user agent intelligent agent to start analysis.
[0059] In one exemplary embodiment, the large language model is ChatGPT, LLaMA, Doubao large model, or Tongyi Qianwen, etc.
[0060] In one exemplary embodiment, the external function library includes an information crawling module, a text-to-speech (TTS) module, a template editing module, a template filling module, a lip alignment module, and a video enhancement module.
[0061] In one exemplary embodiment, the video generation process includes the following steps: generating a director's script based on the user's text video requirements; generating audio based on the director's script; generating a video template based on the audio; generating a first draft of the video based on the video template; lip-syncing the first draft of the video; and enhancing the lip-synced video to generate a final draft of the video, i.e., the final video.
[0062] The following is a specific example illustrating the workflow of a video generation system based on intelligent agents according to this application. Using this intelligent agent-based video generation system, the following tasks are accomplished: A user needs to generate a customized digital human advertising video for a product. The submitted title, i.e., the text video requirement, is: "I want to generate a female digital human voice-over video." The overall product content is based on this link: PRETTYGARDEN Women's Summer Tie One Shoulder BohoFloral Dress Elastic Waist Tiered Ruffle ALine Flowy Mini Dresses at AmazonWomen's Clothing store. The template and background are required to be light and lively, the language to be English, and the target audience to be young women in California, USA.
[0063] Therefore, the URL TO VIDEO function library needs to be called. This library is an external function library that includes functions that help users generate digital human advertising videos based on the Uniform Resource Locator (URL) of the shopping website. The workflow of the various modules in this function library working together to complete this task is as follows:
[0064] Step 201: The user inputs the video requirements (text video requirements), selects LLM, the intelligent agent starts, and two intelligent agents are created: the user agent intelligent agent and the assistant intelligent agent; during this process, the intelligent agent automatically reads the imported function library and converts it into a function library that can be executed by the intelligent agent, and provides it to the intelligent agent for access while initializing the intelligent agent.
[0065] Step 202: The user agent receives the user's video request and sends it to the assistant agent for analysis.
[0066] Step 203: The assistant agent automatically determines that the information crawling module needs to be called, with the URL link in the video requirement as the parameter, and submits the information crawling module and the parameter to the user agent agent.
[0067] Step 204: The user agent executes the information crawling module according to the function (information crawling module) and parameters provided by the assistant agent and obtains the return value (the crawled product description and images), and submits it to the assistant agent.
[0068] Step 205: The assistant agent analyzes the product description and, in conjunction with the language and audience requirements (English, young women in California, USA) in the user's video requirements, generates a script and submits it to the user agent (this process is completed automatically by the large language model in the assistant agent; if a separate script generation module exists in the function library, the assistant agent will submit this module and its corresponding parameters to the user agent).
[0069] Step 206: The user agent determines that audio needs to be generated based on the director's instructions and issues this command to the assistant agent.
[0070] Step 207: The assistant agent determines that the audio generation requires the use of a TTS module with parameters including director's remarks, female digital human, and a light and lively tone, and submits the above module and parameters to the user agent.
[0071] Step 208: The user agent executes the TTS module to generate the final audio transcript and submits it to the assistant agent.
[0072] Step 209: The assistant agent recognizes the above return value as the final audio draft, reports to the user agent that the final audio draft has been generated, and requests the next step.
[0073] Step 210: The user agent agent does not detect that the video has been completed, determines that video needs to be generated based on the audio, and issues this command to the assistant agent.
[0074] Step 211: The assistant agent determines that the template editing module needs to be called, with parameters including product image, director's words (subtitles), audio, and selected digital human (if the user does not select one, the default digital human of the corresponding gender will be used), and submits it to the user agent agent.
[0075] Step 212: The user agent executes the template editing module, generates a video template, and submits it to the assistant agent.
[0076] Step 213: The assistant agent determines that the template filling module needs to be called, with parameters including the edited template, product image, director's words (subtitles), audio, and selected digital human, and submits it to the user agent agent.
[0077] Step 214: The user agent executes the template filling module to generate a first draft of the video (without lip alignment) and submits it to the assistant agent.
[0078] Step 215: The assistant agent determines that the lip alignment module needs to be called, with the parameters being the initial video draft and audio, and submits it to the user agent.
[0079] Step 216: The user agent executes the lip alignment module to generate a lip-aligned video and submits it to the assistant agent.
[0080] Step 217: The assistant agent determines that the video enhancement module needs to be called, with the parameter being the video after lip alignment, and submits it to the user agent.
[0081] Step 218: The user agent executes the video enhancement module, generates the final video draft, and submits it to the assistant agent.
[0082] Step 219: The assistant agent recognizes the above return value as the final video draft, reports to the user agent that the final video draft has been generated, and requests the next step.
[0083] Step 220: The user agent detects that the video is complete and sends a TERMINATE command to end the session.
[0084] Step 221: The intelligent agent framework service now includes all session information. Extract the last session of the assistant agent (final video draft), the first session of the user agent agent (user video request), all sessions containing function calls and function return values (these sessions do not include other content besides function name, parameters, and return value), and available function libraries. Summarize and output the above information, and the process ends.
[0085] In completing this task, the agent automatically determines the workflow and decides whether to call functions. Therefore, the agent can automatically identify other added modules or switch to other function libraries, ensuring continuity and multifunctionality. For example, to further improve video quality, the intelligent agent-based video generation system can automatically identify and utilize other externally added modules or function libraries, such as a director's speech quality evaluation module, an audio quality evaluation module, and a video quality evaluation module.
[0086] Repeating the above process, due to the existence of the evaluation module, the user agent AI determines after step 206 that the generated director's speech needs to be evaluated, and the assistant AI determines after step 209 that the audio evaluation module needs to be executed, and after step 215 that the video evaluation module needs to be executed. Therefore, it is evident that a rich set of external function libraries can be automatically identified and utilized by this application, greatly improving its performance.
[0087] This application aims to improve the efficiency, accuracy, and user-friendliness of LLM in processing text-based video tasks. Through an intelligent agent framework, LLM can automatically call external functions, thereby enhancing its functionality and effectiveness in extracting text requirements and generating high-quality videos. The interaction between two intelligent agents reduces the number of human-computer interactions, allowing users to obtain the final video draft simply by inputting video requirements, thus improving the user experience and making it suitable for various scenarios requiring efficient content generation. This system can also simultaneously report identified function libraries and function call chains to technical personnel, helping them identify and resolve problems, monitor the operational process, improve debugging efficiency, further optimize issues, and enhance the quality of LLM's responses.
[0088] This application can be widely used in video editing, ad generation, content creation and other fields to provide users with efficient and accurate content generation services.
[0089] Based on the same inventive concept, this application also provides a method for implementing the intelligent agent-based video generation system described above. The solution provided by this method is similar to the implementation described in the above system; therefore, the specific limitations in one or more embodiments of the intelligent agent-based video generation method provided below can be found in the limitations of the intelligent agent-based video generation system described above, and will not be repeated here.
[0090] In one exemplary embodiment, a video generation method based on intelligent agents is provided, the video generation method based on intelligent agents comprising:
[0091] The user agent sends the user's text and video requests to the assistant agent; both the user agent and the assistant agent are agents based on a large language model; both the user agent and the assistant agent import functions from the external function library into their own function library.
[0092] The assistant agent selects a function from the function library based on the text / video request, the execution result returned by the user agent agent, or the next work command, and sends the selected function and its corresponding parameters to the user agent agent.
[0093] The user agent executes the function selected by the assistant agent and returns the execution result of the function to the assistant agent.
[0094] The assistant agent determines whether to continue executing the function based on the execution result returned by the user agent. If yes, it selects a function from the function library based on the current execution result returned by the user agent, sends the selected function and its corresponding parameters to the user agent, and returns the step of the user agent executing the function selected by the assistant agent and returning the execution result of the function to the assistant agent. Otherwise, the user agent determines whether to proceed to the next step. If the user agent determines to proceed to the next step, it sends the next step command to the assistant agent.
[0095] The assistant agent selects a function from the function library according to the next work command, sends the selected function and its corresponding parameters to the user agent, and returns the execution result of the function to the assistant agent until the final video is generated.
[0096] In an exemplary embodiment, the video generation method based on intelligent agents further includes: recording virtual dialogue content between the user agent and the assistant agent through an intelligent agent framework, wherein the virtual dialogue content includes function call chains.
[0097] In an exemplary embodiment, after generating the final video, the video generation method based on the intelligent agent further includes: outputting the final answer, function directory, and function call chain of the assistant agent, wherein the final answer is the generated final video, and the function directory is a directory of functions executed by the user agent agent in the process of generating the final video.
[0098] In an exemplary embodiment, before the user agent sends the user's text video request to the assistant agent, the video generation method based on the intelligent agent further includes: obtaining the user's input text video request and the selected large language model.
[0099] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0100] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An intelligent agent based video generation system, characterized by, The intelligent agent-based video generation system comprises a user agent intelligent agent, an assistant intelligent agent and an external function library; the user agent intelligent agent and the assistant intelligent agent are both intelligent agents based on a large language model, a virtual dialogue between the user agent intelligent agent and the assistant intelligent agent is used to convert a user text video requirement into a structured instruction, the text video requirement is disassembled and refined, and the external function library is used to improve the video generation quality; the user agent intelligent agent and the assistant intelligent agent both import functions in the external function library into the function library of the user agent intelligent agent and the assistant intelligent agent; the external function library comprises an information crawling module, a text-to-speech module, a template editing module, a template filling module, a mouth shape alignment module and a video enhancement module; The user agent intelligent agent is configured to send the text video requirement of a user to the assistant intelligent agent; The assistant intelligent agent is configured to select a function from the function library according to the text video requirement, an execution result returned by the user agent intelligent agent or a next step work command, and send the selected function and corresponding parameters to the user agent intelligent agent; The user agent intelligent agent is configured to execute the function selected by the assistant intelligent agent, and return an execution result of the function to the assistant intelligent agent; The assistant intelligent agent is further configured to determine whether to continue executing the function according to the execution result returned by the user agent intelligent agent, if yes, continue to select a function from the function library according to the execution result returned by the user agent intelligent agent, and if no, determine whether to execute a next step work until a final video is generated by the user agent intelligent agent, if the user agent intelligent agent determines to execute the next step work, send a next step work command to the assistant intelligent agent; The intelligent agent-based video generation system further comprises an intelligent agent framework, which is configured to record virtual dialogue content between the user agent intelligent agent and the assistant intelligent agent, and the virtual dialogue content comprises a function call link; The intelligent agent-based video generation system further comprises a result output module; the result output module is configured to output a final answer of the assistant intelligent agent, a function directory and a function call link, the final answer being a final video generated, and the function directory being a function directory constituted by functions executed by the user agent intelligent agent in the process of generating the final video; The intelligent agent automatically determines a work flow and judges whether to call a function, and the intelligent agent can automatically identify other modules added or other function libraries used; the external function library further comprises a director speech quality evaluation module, an audio quality evaluation module and a video quality evaluation module.
2. The intelligent agent-based video generation system of claim 1, wherein, The intelligent agent-based video generation system further comprises a topic acquisition module, which is configured to acquire a text video requirement input by a user and a selected large language model.
3. The intelligent agent-based video generation system of claim 1, wherein, The large language model is ChatGPT, LLaMA, a bean bag large model or a general-purpose thousand questions model.
4. A method for generating a video based on an intelligent agent, the method comprising: The intelligent agent-based video generation method comprises: The user agent intelligent agent sends the user's text video requirement to the assistant intelligent agent; the user agent intelligent agent and the assistant intelligent agent are both large language model-based intelligent agents; the user agent intelligent agent and the assistant intelligent agent both import functions in an external function library into their own function libraries; The assistant intelligent agent selects a function from the function library according to the text video requirement, the execution result returned by the user agent intelligent agent, or the next step work command, and sends the selected function and corresponding parameters to the user agent intelligent agent; The user agent intelligent agent executes the function selected by the assistant intelligent agent and returns the execution result of the function to the assistant intelligent agent; The assistant intelligent agent judges whether to continue executing the function according to the execution result returned by the user agent intelligent agent, if yes, continues to select a function from the function library according to the execution result returned by the current user agent intelligent agent, and sends the selected function and corresponding parameters to the user agent intelligent agent, returns the user agent intelligent agent to execute the function selected by the assistant intelligent agent, and returns the execution result of the function to the assistant intelligent agent, if not, judges whether to execute the next step work by the user agent intelligent agent, if the user agent intelligent agent judges to execute the next step work, sends the next step work command to the assistant intelligent agent; The assistant intelligent agent selects a function from the function library according to the next step work command, and sends the selected function and corresponding parameters to the user agent intelligent agent, and returns the user agent intelligent agent to execute the function selected by the assistant intelligent agent, and returns the execution result of the function to the assistant intelligent agent, until the final video is generated; The intelligent agent-based video generation method further comprises: recording the virtual dialogue content between the user agent intelligent agent and the assistant intelligent agent through an intelligent agent framework, the virtual dialogue content comprising a function call link; The intelligent agent-based video generation method further comprises: Output the final answer of the assistant intelligent agent, the function directory and the function call link, the final answer being the generated final video, and the function directory being a function directory composed of functions executed by the user agent intelligent agent in the process of generating the final video.
5. The smart agent based video generation method of claim 4, wherein, Before the user agent intelligent agent sends the user's text video requirement to the assistant intelligent agent, the intelligent agent-based video generation method further comprises: obtaining the user input text video requirement and the selected large language model.
Citation Information
Patent Citations
Data generation method, electronic equipment and storage medium
CN117556026A
Video generation method and system based on large language model
CN117676195A