Video generation method and device based on multi-agent cooperation and agent

This video generation method, which utilizes a multi-agent collaborative approach, leverages the target agent to acquire and analyze user needs, controls the execution process of subtasks, and generates high-quality videos. This solves the problems of long processing times and low quality found in existing video production tools, achieving efficient and accurate video generation.

CN120151560BActive Publication Date: 2025-12-26BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510354039.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-12-26
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing video production tools are time-consuming and struggle to accurately generate video data that meets user needs, resulting in low video quality that fails to satisfy users' actual requirements.

Method used

The video generation method using multi-agent collaboration utilizes the target agent to obtain sub-task requirement information, the first sub-agent to determine the sub-task execution elements, and the second sub-agent to execute the target sub-task, generating a video that meets the user's needs.

Benefits of technology

It improved the matching degree and quality of videos with user needs, and increased the efficiency and accuracy of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151560B_ABST
    Figure CN120151560B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-agent cooperation-based video generation method and device and an agent, relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large models, agents, AIGC, metaverse, etc., and is applied to the application scenarios of intelligent education, video production, animation production, etc. The method comprises the following steps: performing a target subtask in a target task by using at least one target agent to obtain a target subexecution result; and generating a target video based on the target subexecution result, wherein the target agent performs the following operation: obtaining subtask demand information for the target subtask; determining a subtask execution element for controlling the execution process of the target subtask based on the subtask demand information by using a first subagent of the target agent; and executing the target subtask based on the subtask execution element by using a second subagent of the target agent to obtain the target subexecution result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large models, agents, AIGC (Artificial Intelligence Generated Content), metaverse, and the like, and is applied to application scenarios such as smart education, video production, and animation production. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, a large model with generative capability such as a multi-modal large model can be used to generate video data to meet the needs of users to quickly produce video data. SUMMARY

[0003] The present disclosure provides a video generation method and device based on multi-agent cooperation, an agent, an electronic device, and a storage medium.

[0004] According to an aspect of the present disclosure, a video generation method based on multi-agent cooperation is provided, including: using at least one target agent to perform a target subtask in a target task to obtain a target subexecution result; and generating a target video based on the target subexecution result, wherein the target agent performs the following operations: obtaining subtask demand information for the target subtask; using a first subagent of the target agent to determine a subtask execution element for controlling an execution process of the target subtask based on the subtask demand information; and using a second subagent of the target agent to perform the target subtask based on the subtask execution element to obtain the target subexecution result.

[0005] According to another aspect of the present disclosure, a video generation device based on multi-agent cooperation is provided, including: an execution module configured to use at least one target agent to perform a target subtask in a target task to obtain a target subexecution result; and a target video generation module configured to generate a target video based on the target subexecution result, wherein the target agent performs the following operations: obtaining subtask demand information for the target subtask; using a first subagent of the target agent to determine a subtask execution element for controlling an execution process of the target subtask based on the subtask demand information; and using a second subagent of the target agent to perform the target subtask based on the subtask execution element to obtain the target subexecution result.

[0006] According to another aspect of the present disclosure, an artificial intelligence agent configured to perform a video generation method based on multi-agent cooperation is also provided.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method for video generation based on multi-agent cooperation.

[0008] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method for video generation based on multi-agent cooperation.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method for video generation based on multi-agent cooperation.

[0010] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0012] Figure 1 An exemplary system architecture to which the method and apparatus for video generation based on multi-agent cooperation according to embodiments of the present disclosure can be applied is schematically shown;

[0013] Figure 2A A flowchart of the method for video generation based on multi-agent cooperation according to embodiments of the present disclosure is schematically shown;

[0014] Figure 2B A flowchart of performing a target subtask in a target task by at least one target agent according to embodiments of the present disclosure is schematically shown.

[0015] Figure 3 A schematic diagram of performing a target subtask by a target agent according to embodiments of the present disclosure is schematically shown;

[0016] Figure 4 A schematic diagram of generating a task orchestration result by a designated agent according to embodiments of the present disclosure is schematically shown;

[0017] Figure 5 A schematic diagram of performing a target subtask by a target agent according to embodiments of the present disclosure is schematically shown;

[0018] Figure 6A schematic diagram of a video generation method based on multi-agent cooperation according to an embodiment of the present disclosure is shown schematically.

[0019] Figure 7A An effect schematic diagram of a target video according to an embodiment of the present disclosure is shown schematically.

[0020] Figure 7B An effect schematic diagram of a health popular science video according to an embodiment of the present disclosure is shown schematically.

[0021] Figure 7C An effect schematic diagram of a legal popularization video according to an embodiment of the present disclosure is shown schematically.

[0022] Figure 8 A block diagram of a video generation apparatus based on multi-agent cooperation according to an embodiment of the present disclosure is shown schematically.

[0023] Figure 9 A structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure is shown schematically.

[0024] Figure 10 A schematic block diagram of an example electronic device that can be used to implement a video generation method based on multi-agent cooperation according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.

[0026] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information involved comply with relevant laws and regulations, necessary security measures are taken, and do not violate public order and good customs.

[0027] The inventors have found that, with the rapid development of video platforms, more and more users are making video products based on personal needs. The video production process usually takes a long time, and related video production tools are difficult to accurately generate video data that meets user needs, or the quality of the generated video is low, which is difficult to meet the actual needs of users.

[0028] The disclosure provides a multi-agent cooperation-based video generation method and device, an agent, an electronic device, and a storage medium. The multi-agent cooperation-based video generation method includes: performing a target subtask in a target task by using at least one target agent to obtain a target subexecution result; and generating a target video based on the target subexecution result, wherein the target agent performs the following operations:

[0029] obtaining subtask requirement information for the target subtask; determining a subtask execution element for controlling an execution process of the target subtask based on the subtask requirement information by using a first subagent of the target agent; and performing the target subtask based on the subtask execution element by using a second subagent of the target agent to obtain the target subexecution result.

[0030] According to an embodiment of the disclosure, by using the first subagent in the target agent to determine the subtask execution element based on the subtask requirement information for the target subtask, the first agent can accurately analyze the subtask requirement information, and the subtask execution element can accurately represent the execution process of the target subtask to enable the target subtask to be executed according to the demand attribute represented by the subtask requirement information. Furthermore, by using the second subagent to perform the target subtask based on the subtask execution element, the second subagent can be accurately controlled to perform the target subtask according to the demand attribute represented by the subtask requirement information, so that the target subexecution result obtained conforms to the demand attribute represented by the subtask requirement information. In this way, the target video that can accurately meet the user demand can be generated based on the target subexecution result obtained by performing the target subtask by using the at least one target agent, the matching degree between the target video and the user demand is improved, and the quality of the target video is improved.

[0031] Figure 1 An exemplary system architecture to which the multi-agent cooperation-based video generation method and device according to an embodiment of the disclosure can be applied is schematically shown.

[0032] It should be noted that, Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the disclosure can be applied, to help those skilled in the art understand the technical content of the disclosure, but does not mean that the embodiments of the disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the multi-agent cooperation-based video generation method and device can be applied can include a terminal device, but the terminal device can not need to interact with the server to implement the multi-agent cooperation-based video generation method and device provided by the embodiments of the disclosure.

[0033] As Figure 1As shown, the system architecture 100 according to this embodiment can include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, and the like.

[0034] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, and the like. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, and the like (only as examples).

[0035] The terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and the like.

[0036] The server 105 can be a server providing various services, such as a background management server providing support for content browsed by a user using a terminal device 101, 102, 103 (only as an example). The background management server can analyze and process received user requests and the like, and feed back the processing results (such as web pages, information, or data, and the like obtained or generated according to user requests) to the terminal device.

[0037] The server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server 105 can also be a server of a distributed system, or a server combined with a blockchain.

[0038] The server 105 can obtain the target agent by deploying the generative model. It should be noted that the first sub-agent and the second sub-agent in the target agent can be deployed in the same server, or the first sub-agent and the second sub-agent in the target agent can also be deployed in different servers.

[0039] It should be noted that the method for generating a video based on multi-agent cooperation provided in the embodiments of the present disclosure can also be executed by the server 105. Accordingly, the apparatus for generating a video based on multi-agent cooperation provided in the embodiments of the present disclosure can be arranged in the server 105. The method for generating a video based on multi-agent cooperation provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the apparatus for generating a video based on multi-agent cooperation provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0040] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system is merely illustrative. Any number of terminal devices, networks and servers can be provided according to implementation needs.

[0041] Figure 2A A flowchart of the method for generating a video based on multi-agent cooperation according to an embodiment of the present disclosure is schematically shown.

[0042] As Figure 2A The method for generating a video based on multi-agent cooperation includes operations S210-S220.

[0043] In operation S210, a target sub-task in a target task is executed by using at least one target agent, and a target sub-execution result is obtained.

[0044] In operation S220, a target video is generated based on the target sub-execution result.

[0045] According to an embodiment of the present disclosure, the target agent can be an execution subject for executing the target sub-task constructed based on a large model having a generation function such as a generative model. The target agent can be deployed in a server, or can also be respectively deployed in a plurality of servers connected in communication.

[0046] According to an embodiment of the present disclosure, the target task can be used for generating a target video, the target task can include one or more target sub-tasks, and the target sub-task can be used for generating an intermediate result determining the target video. The target sub-task can be a processing process including a plurality of execution steps, which can involve intelligent reasoning links such as language understanding, context analysis, logical reasoning, etc. The target sub-task can include any task type, such as a copywriting generation type, a speech synthesis type, etc.

[0047] According to an embodiment of the present disclosure, the target sub-execution result can include an intermediate result for generating a video, for example, a video script generation result, an image rendering result, synthesized speech data, and the like. Generating a target video based on the target sub-execution result as video material can include video rendering based on one or more target sub-execution results to obtain the target video.

[0048] In one example, the video script, the multiple frames of images in the target sub-execution result can be processed based on a visual large model to obtain the target video.

[0049] In one example, the target sub-execution result can be synthesized speech data and a narration script. The synthesized speech data and the narration script can be filled into a video compilation script template, so that the updated video compilation script can be rendered to obtain the target video.

[0050] In one example, the one or more target sub-execution results can also be processed based on an agent with video generation capability to generate the target video.

[0051] It should be noted that the embodiments of the present disclosure do not limit the specific way of generating the target video, as long as at least one target sub-execution result generated by the target agent is used.

[0052] According to an embodiment of the present disclosure, the target agent can include a plurality of sub-agents. The sub-agent can be an agent with independent task execution function, and the target agent can be a system including a plurality of sub-agents. The target sub-task is executed by utilizing the cooperation of the plurality of sub-agents to improve the execution efficiency of the target sub-task and the accuracy and generation effect of the target sub-execution result.

[0053] Figure 2B A flowchart for executing a target sub-task in a target task by using at least one target agent according to an embodiment of the present disclosure is schematically shown.

[0054] As shown in Figure 2B The target agent performs the following operations S211-S213 to execute the target sub-task.

[0055] In operation S211, sub-task requirement information for the target sub-task is obtained.

[0056] In operation S212, a first sub-agent of the target agent determines a sub-task execution element for controlling the execution process of the target sub-task based on the sub-task requirement information.

[0057] In operation S213, a second sub-agent of the target agent executes the target sub-task based on the sub-task execution element to obtain the target sub-execution result.

[0058] According to an embodiment of the present disclosure, the subtask demand information for the target subtask can be information for representing a demand intention of the target object for the target video, which can represent a demand intention for a length, a style, a theme, etc. of the target video. The subtask demand information can represent a demand intention that needs to be focused on by the target agent performing the target subtask, for example, a demand intention for a video theme, a video length, etc. that needs to be focused on by a script generation agent performing a script generation subtask.

[0059] The execution subject for obtaining the subtask demand information can be any subagent in the target agent, for example, the first subagent can be used as the execution subject to obtain the subtask demand information. However, it is not limited thereto, and the subtask demand information can also be obtained based on other subagents or functional modules included in the target agent, and transmitted to the first subagent. Embodiments of the present disclosure do not limit the specific execution subject of step S211, as long as it is a functional module included in the target agent.

[0060] According to an embodiment of the present disclosure, the subtask demand information can include part or all of the demand intention. In the case where the subtask demand information includes all the demand intention, the demand intention of the target object can be fully analyzed based on the first subagent, so as to improve the matching degree between the subtask execution element and the demand intention, and avoid the intention missing or conflict between the subtask execution element and the demand intention.

[0061] According to an embodiment of the present disclosure, the first subagent and the second subagent can be two different agents included in the target agent. The first subagent can perform semantic analysis based on the subtask demand information, and generate a subtask execution element for controlling the execution process of the target subtask in combination with the task type of the target subtask. The subtask execution element can control the execution steps, execution parameters, execution logic, information acquisition, etc. of the target subtask, so that the second subagent having the target subtask execution function can accurately execute the target subtask according to the demand intention under the control of the subtask execution element, so as to make the target subtask execution result match the demand intention of the target object, and avoid the target video display effect not meeting the actual demand of the user due to the conflict between the subtask execution result and the demand intention.

[0062] In one example, the subtask execution element can include an execution code script, a task configuration parameter, etc. for controlling the second subagent to execute the target subtask. So that the second subagent can accurately execute the script based on the task configuration parameter to generate a target subtask execution result meeting the demand intention.

[0063] According to an embodiment of the present disclosure, by decomposing a target task of generating a target video into target sub-tasks, and by a first sub-agent in the target agent analyzing sub-task demand information to obtain a sub-task execution element for controlling a sub-task execution process, and by a second sub-agent precisely executing a target sub-task based on the control of the sub-task execution element, the cooperation between multiple sub-agents can be realized to finely control the execution process of the target sub-task, improve the execution efficiency of the target sub-task, and improve the matching degree between the target sub-execution result and the demand intention of the target object, thereby improving the video quality of the target video by fine-grained control of the quality of the target sub-execution result, and meeting the video generation demand of the target object.

[0064] According to an embodiment of the present disclosure, the sub-task execution element includes at least one of a style consistency condition and a logical coherence condition.

[0065] According to an embodiment of the present disclosure, the style consistency condition represents an expression style of the target sub-execution result, which matches a demand style represented by the object demand information. The target sub-execution result can include any type of text, voice, video picture, etc. The expression style includes at least one of a text expression style, a voice expression style, and a video picture expression style.

[0066] The text expression style can be any type of expression style such as lively, serious, colloquial, dialect style, etc. The voice expression style can include the same style types as the text expression style, such as lively, serious, etc., in addition to the voice attribute related styles such as gender, tone, etc. The video picture style can represent the image style of the picture displayed in the video background picture, video illustration, video frame image, etc. The video picture style can be the same as the text expression style, and the voice expression style can be the same as the video picture style.

[0067] It should be noted that the expression style can be represented based on a label. By training multiple sub-agents in the target agent based on sample data labeled with style labels, the sub-agents can learn the relationship between the sub-task execution result of the corresponding task type and the style label. Thus, the sub-task execution element output by the first sub-agent can represent the expression style corresponding to the sub-task demand information, and the second sub-agent can execute the generation task such as copywriting generation and image generation based on the relatively accurate style consistency condition.

[0068] According to an embodiment of the present disclosure, the logical coherence condition represents that the plurality of target sub-execution results of the target task have logical coherence. The target task can include a plurality of target sub-tasks, and the plurality of target sub-tasks can be cooperatively executed by different plurality of target agents to obtain the plurality of target sub-execution results. The logical coherence between the plurality of target sub-execution results can represent that different target sub-execution results such as text, image, voice, etc. express the same theme, demand intention. Or the logical coherence can represent that the plurality of target sub-execution results have the same or similar logical structure, avoiding logical confusion of the plurality of target sub-execution results.

[0069] For example, the plurality of target sub-execution results are video script and video illustration respectively, the video script represents the text content of the introduction of the prevention method for diabetes, and the video illustration appears the picture of a pet cat, and it can be considered that the video script and the video illustration do not have logical coherence.

[0070] For another example, the plurality of target sub-execution results are video script and video subtitle respectively, the plurality of operation steps represented by the video script of the power equipment are different from the operation execution order of the plurality of operation steps represented by the video subtitle, and it can be considered that the video script and the video illustration do not have logical coherence.

[0071] It should be noted that the sub-task execution element can include any type of information such as sub-task demand condition, demand condition explanation text, demand condition example, etc. The sub-task execution element can be used as prompt information for the second sub-agent to execute the target sub-task.

[0072] According to an embodiment of the present disclosure, the target sub-task includes at least one of the following: a script generation sub-task, a split screen data generation sub-task, a graphic element generation sub-task, and a subtitle generation sub-task.

[0073] The second sub-agent can execute the script generation sub-task based on the script generation sub-task execution element for the script generation sub-task to obtain a target script matched with the demand intention of the target object. The target script can be used for the composition elements of the target video such as explanation voice, inserted text, etc. The script generation sub-task execution element can include text theme condition, text word number demand condition, text expression style demand condition, etc.

[0074] In one example, the target script can be used to generate voice data to make the target video introduce the script content such as product information, device implementation steps, etc. by representing the target script based on the voice data.

[0075] Figure 3 The schematic diagram shows the principle diagram of executing the target sub-task by the target agent according to an embodiment of the present disclosure.

[0076] AsFigure 3 As shown, the target agent 310 for performing the script generation subtask can include a first sub-agent 311 and a second sub-agent 312. The first sub-agent 311 obtains the requirement intention 301 as subtask requirement information, and performs semantic analysis on the requirement intention 301 "A disease, prevention, long text, step-by-step", to obtain a subtask execution element for the script generation subtask. The second sub-agent 312 performs the script generation subtask based on the subtask execution element, to obtain the target script 302 as a target sub-execution result. The second sub-agent 312 controls the execution process of the script generation subtask based on the subtask execution element, can generate a plurality of step texts containing step identifiers, and the target script 302 is also a long text indicating prevention of the A disease.

[0077] The shot data generation subtask is used to generate target shot data, which can indicate switching time of different video frame images in a target video, display time, display duration, etc. of video elements such as images, texts inserted in the target video, and the like. The shot data generation subtask execution element for the shot data generation subtask can include constraints on display time, display duration, etc. of display properties of video elements, or the shot data generation subtask execution element can also indicate that the display properties of the video elements are adapted to the voice content in the target video. For example, the shot data generation subtask execution element can indicate that the voice content "eat less high-sugar food" is the same as the display time of the image element indicating that chocolate is prohibited.

[0078] According to an embodiment of the present disclosure, the second sub-agent can generate a graph element by performing a graph element generation subtask, and the graph element can include a background image of a video, an inserted image, an inserted animation, an inserted artistic character, and the like. The graph element generation subtask execution element for the graph element generation subtask can represent requirements related to the graph element, such as image content classification, image element style, graph element size, graph element arrangement, etc. of the graph element to be generated.

[0079] It should be noted that the second sub-agent can generate a graph element based on the image generation capability of the visual large model, or can also enhance retrieval based on semantic understanding capability, retrieve a graph element, and embodiments of the present disclosure do not limit the specific execution manner of the second sub-agent performing the graph element generation subtask.

[0080] According to an embodiment of the present disclosure, the subtitle generation subtask can be understood as generating a target subtitle in a target video by a second sub-execution body according to a subtitle generation subtask execution element. The target subtitle can include subtitle text content, a timestamp of the subtitle text, and a text word in the subtitle text that needs to be displayed by an emphasis display mode such as bold or underlined. The subtitle generation subtask execution element can be output by the first sub-intelligent agent based on understanding and analysis of the subtask demand information, so as to control the second sub-intelligent agent to accurately execute the subtitle generation subtask based on the subtitle generation subtask execution element. In addition, the subtitle generation subtask execution element can also include at least one of the logical coherence condition, the expression style consistency condition, and other subtask demand conditions, to control the logic and expression manner consistency between the target subtitle and other target sub-execution results.

[0081] According to an embodiment of the present disclosure, obtaining the subtask demand information for the target subtask includes: performing demand information detection on object demand information of the target object by the target intelligent agent based on a subtask type of the target subtask, to obtain the subtask demand information.

[0082] According to an embodiment of the present disclosure, the subtask type can include any subtask type such as a script generation type and a graph element generation type. The demand information detection can be performed on the object demand information of the target object by the first sub-intelligent agent based on the subtask type of the target subtask, to obtain subtask demand information related to the subtask type. Or the demand information detection can also be performed on the object demand information of the target object by other sub-intelligent agents or functional modules in the target intelligent agent, to obtain the subtask demand information.

[0083] According to an embodiment of the present disclosure, by performing the demand information detection on the object demand information based on the subtask type by the target intelligent agent, to obtain the subtask demand information for the target subtask, the subtask demand information can be matched with the execution process of the target subtask having the subtask type, so as to realize the preliminary demand intention understanding of the object demand information by the target intelligent agent, to enable the target intelligent agent to more comprehensively learn the demand intention of the target object. The subtask execution element generated by the first sub-intelligent agent based on the subtask demand information can more accurately represent the demand intention of the target object, the subtask execution element can closely relate the execution process of the target subtask to the demand intention, and the quality of the target sub-execution result is improved.

[0084] According to an embodiment of the present disclosure, the target task includes a plurality of target subtasks. The plurality of target subtasks can be assigned to a plurality of target intelligent agents for execution, or the plurality of target subtasks can be assigned to the target intelligent agent and a non-intelligent agent functional module for executing the target subtasks for execution. Embodiments of the present disclosure do not limit the specific assignment manner of the plurality of target subtasks.

[0085] According to an embodiment of the present disclosure, the multi-agent cooperation based video generation method further includes utilizing the specified agent to perform the following operations: performing intent recognition on the object demand information of the target object to obtain a demand intent; determining a plurality of target sub-tasks and a dependency relationship between the plurality of target sub-tasks based on the demand intent; and issuing the plurality of target sub-tasks and the dependency relationship between the plurality of target sub-tasks to the plurality of target agents.

[0086] According to an embodiment of the present disclosure, the specified agent can be any one of the plurality of target agents, or the specified agent can also be an agent different from the target agent used to perform the target sub-task.

[0087] According to an embodiment of the present disclosure, the specified agent can perform intent recognition by performing semantic understanding on the object demand information, so as to determine a plurality of sub-task types capable of meeting the demand intent based on the demand intent represented by the intent recognition result, and then determine the plurality of target sub-tasks, thereby achieving mapping to a plurality of specific sub-tasks based on the demand intent. The specified agent can also determine the execution logic between the plurality of target sub-tasks based on the specific execution elements of the target task used to generate the target video, and then obtain the dependency relationship between the plurality of target sub-tasks. By issuing the target sub-task corresponding to the target agent and the dependency relationship between the plurality of target sub-tasks to the plurality of target agents, the plurality of target agents can each obtain an associated sub-execution result from an associated target agent having a dependency relationship based on the dependency relationship, and then automatically execute the target sub-task based on the execution logic of the plurality of target sub-tasks, thereby improving the execution efficiency of the target task.

[0088] According to an embodiment of the present disclosure, the dependency relationship between the plurality of target sub-tasks can be represented based on a topology structure representing the execution logic of the target sub-task. Alternatively, the dependency relationship between the plurality of target sub-tasks can also be represented based on a flowchart or other forms, and the embodiments of the present disclosure do not limit this.

[0089] In one example, the specified agent is a task orchestration agent different from the target agent for executing the target subtask. The task orchestration agent can determine the target task for generating the target video according to the demand intention of the target object by performing intention recognition on the object demand information, and split the target task into multiple target subtasks according to the task elements of the target task. The task orchestration agent can deliver the multiple target subtasks and the dependency relationship to the target agents corresponding to the subtask types through the respective subtask types of the target agents, so as to implement task orchestration of the multiple target agents, and make the target agents transmit the target subexecution results generated by the target agents to the next target agent according to the execution logic of the target task based on the dependency relationship, so as to realize automatic intelligent flow of the target subexecution results of the target subtasks between the multiple target agents, and improve the execution efficiency of the target task.

[0090] According to an embodiment of the present disclosure, the subtask demand information for the target subtask is determined by the specified agent based on the demand intention.

[0091] According to an embodiment of the present disclosure, the demand intention can be the subtask demand information, and the specified agent delivers the subtask demand information to the target agent, or the target agent obtains the subtask demand information from the specified agent. By delivering the demand intention to each target agent as the subtask demand information, the target agent can reduce the process of intention recognition on the object demand information, and reduce the computational overhead generated by the target agent. Meanwhile, different target agents may Figure 1 have semantic conflicts or semantic loss in the demand intention represented by the subtask demand information identified by each target agent from the same object demand information, resulting in that the target subexecution result does not match the demand of the target object. Therefore, by sending the demand intention to multiple target agents as the subtask demand information, the multiple target agents can execute their respective target subtasks based on the same or similar demand intention, and the consistency and logical coherence between the multiple target subexecution results can be realized.

[0092] It should be noted that the subtask demand information delivered to the target agent can be a user personalized subtask demand for representing the demand intention, and the target agent can also obtain basic subtask demand information required by the target subtask of the subtask type, such as text coherence demand condition, sensitive word avoidance demand condition, etc.

[0093] According to an embodiment of the present disclosure, the determining the plurality of target sub-tasks based on the demand intention and the dependency relationship between the plurality of target sub-tasks comprises: performing, by a first specified sub-agent of the specified agent, a task scheduling task based on the demand intention to obtain a plurality of initial sub-tasks and an initial dependency relationship; performing, by a second specified sub-agent of the specified agent, a task scheduling detection sub-task on the plurality of initial sub-tasks and the initial dependency relationship based on a video demand condition represented by the demand intention to obtain a task scheduling detection result; in a case where the task scheduling detection result does not satisfy the video demand condition, performing, by the first specified sub-agent, a task scheduling sub-task based on the demand intention and the task scheduling detection result to obtain the plurality of target sub-tasks satisfying the video demand condition and the dependency relationship between the plurality of target sub-tasks.

[0094] According to an embodiment of the present disclosure, the first specified sub-agent and the second specified sub-agent can be agents for performing different generation functions. The first specified sub-agent generates the plurality of initial sub-tasks and the initial dependency relationship between the plurality of initial sub-tasks based on the understanding of the demand intention. The demand intention may, for example, represent a video length demand, a video theme demand, a video style demand, and the like of the video demand condition for the target video. By utilizing the second specified sub-agent to perform the task scheduling detection sub-task on the plurality of initial sub-tasks and the initial dependency relationship based on the video demand condition as prompt information, it is possible to detect, by the second specified sub-agent, whether the initial sub-tasks and the initial dependency relationship satisfy the video demand condition, and indicate, based on a defect type represented by the task scheduling detection result, which video demand conditions the initial sub-tasks and the initial dependency relationship specifically do not satisfy. Thus, in a case where the task scheduling detection result does not satisfy the video demand condition, it is possible to correct, by the first specified sub-agent, the understanding defects of the plurality of initial sub-tasks and the initial scheduling logic for the demand intention or the basic condition by processing the task scheduling detection result, so as to process the demand intention to perform the task scheduling task under the condition of correcting the understanding defects, to obtain the plurality of target sub-tasks satisfying the video demand condition and the dependency relationship between the plurality of target sub-tasks.

[0095] According to an embodiment of the present disclosure, in a case where the current task scheduling detection result indicates that the plurality of intermediate sub-tasks and the intermediate dependency relationship obtained by the first specified sub-agent performing the task scheduling sub-task based on the demand intention and the task scheduling detection result still do not satisfy the video demand condition, the first specified sub-agent can be utilized to perform the task scheduling detection sub-task based on the current task scheduling detection result and the demand intention in a loop until the current task scheduling detection result indicates that the plurality of intermediate sub-tasks and the intermediate dependency relationship generated satisfy the video demand condition, and the plurality of intermediate sub-tasks and the intermediate dependency relationship currently generated are taken as the plurality of target sub-tasks and the dependency relationship generated.

[0096] According to an embodiment of the present disclosure, by generating a plurality of initial sub-tasks and initial dependency relationships through a first designated sub-agent, and by detecting task arrangement through a second designated sub-agent, the task arrangement result can be checked and generated in a manner of cooperation of a plurality of agents, avoiding the same agent generating a plurality of target sub-task arrangement logics and self-reflection of the arrangement logic, so as to avoid the difficulty in detecting the defect type of the arrangement result, thereby the first designated sub-agent can be instructed to correct errors based on the task arrangement detection result generated by the second designated sub-agent, and the initial sub-tasks and initial dependency relationships can be accurately corrected.

[0097] According to an embodiment of the present disclosure, the task arrangement detection result represents at least one of the following defect types: the initial dependency relationship has an execution logic conflict; and a plurality of initial sub-task types of a plurality of initial sub-tasks do not match the demand type represented by the demand intention.

[0098] The initial dependency relationship has an execution logic conflict, which can be understood as an initial task executed in front of a plurality of initial sub-tasks requiring a sub-execution result of an initial sub-task executed later. For example, the copywriting generation task is the last executed initial sub-task, and the voice synthesis sub-task is the first executed initial sub-task, but the voice synthesis sub-task needs the target copywriting generated by the copywriting generation task to execute the voice generation sub-task. Alternatively, the initial dependency relationship has an execution logic conflict, which can also be understood as other types of sub-task execution logic defects, such as a dead loop in the initial sub-task execution and other execution logic defect types.

[0099] The plurality of initial sub-task types of the plurality of initial sub-tasks do not match the demand type represented by the demand intention, for example, a plurality of initial sub-tasks can lack a graphic element generation sub-task matching the video illustration demand represented by the demand intention. However, it is not limited to this, as long as the task arrangement detection result can represent which demand types do not match the demand intention.

[0100] Figure 4 The schematic diagram illustrates the principle of generating a task arrangement result by a designated agent according to an embodiment of the present disclosure.

[0101] As Figure 4As shown, the designated intelligent agent 410 can include an intention recognition sub-intelligent agent 411, a first designation sub-intelligent agent 412, and a second designation sub-intelligent agent 413. The object requirement information 401 of the target object can be "need to generate a short video for preventing disease A for popular science". The intention recognition sub-intelligent agent 411 is used to process the object requirement information 401 to obtain the requirement intention "short video", "preventing disease A", and "popular science". The first designation sub-intelligent agent 412 is used to process the requirement intention to perform a task arrangement task to obtain a plurality of initial sub-tasks and initial dependency relationships. The second designation sub-intelligent agent 413 is used to perform a task arrangement detection sub-task on the plurality of initial sub-tasks and the initial dependency relationships based on the short video condition represented by the requirement intention, the video theme requirement condition "preventing disease A", and the scene requirement condition "popular science" to obtain a task arrangement detection result. In the case that the task arrangement detection result does not satisfy any of the video requirement conditions, the second designation sub-intelligent agent 413 sends the task arrangement detection result to the first designation sub-intelligent agent 412 to inform the first designation sub-intelligent agent 412 that it needs to perform the task arrangement task again according to the defect type indicated by the task arrangement detection result, so as to obtain the current plurality of intermediate sub-tasks and intermediate arrangement logic. The second designation sub-intelligent agent 413 performs a task arrangement detection check on the current plurality of intermediate sub-tasks and the intermediate arrangement logic based on the video requirement condition represented by the requirement intention. In the case that the task arrangement detection result indicates that the video requirement condition is satisfied, the designated intelligent agent 410 sends the task arrangement result 402 satisfying the video requirement condition to the plurality of target intelligent agents. The task arrangement result 402 can include a plurality of target sub-tasks represented by nodes and dependency relationships between the plurality of target sub-tasks represented by directed edges.

[0102] According to an embodiment of the present disclosure, the first sub-intelligent agent of the target intelligent agent determines the sub-task execution element for controlling the execution process of the target sub-task based on the sub-task requirement information, which includes: the first sub-intelligent agent processes the sub-task requirement information and the associated sub-execution result of the associated sub-task in the target task to obtain the sub-task execution element.

[0103] According to an embodiment of the present disclosure, the associated sub-task and the target sub-task have a dependency relationship. For example, the associated sub-task can be a target sub-task that is executed and completed before the target sub-task in the task arrangement result. The associated sub-execution result can be obtained from an associated target intelligent agent having a dependency relationship with the current target intelligent agent.

[0104] According to an embodiment of the present disclosure, processing the sub-task demand information and the associated sub-execution result of the associated sub-task in the target task by the first sub-agent can include: taking the sub-task demand information and the associated sub-execution result as prompt information, and receiving the input prompt information by the first sub-agent to prompt the first sub-agent to understand the demand points of the current target sub-task according to the associated sub-execution result, improve the description accuracy of the sub-task execution elements to the task attributes of the sub-task, and then make the second sub-agent execute the target sub-task based on the relatively accurate sub-task execution elements, and improve the accuracy of the subsequent target sub-execution result.

[0105] According to an embodiment of the present disclosure, processing the sub-task demand information and the associated sub-execution result of the associated sub-task in the target task by the first sub-agent can include: taking the sub-task demand information and the associated sub-execution result as prompt information, and receiving the input prompt information by the first sub-agent to prompt the first sub-agent to understand the demand points of the current target sub-task according to the associated sub-execution result, improve the description accuracy of the sub-task execution elements to the task attributes of the sub-task, and then make the second sub-agent execute the target sub-task based on the relatively accurate sub-task execution elements, and improve the accuracy of the subsequent target sub-execution result.

[0106] According to an embodiment of the present disclosure, the object demand information represents the demand attributes of the target object for the target video. The target object can be a user object obtaining the target video.

[0107] In one example, the sub-task demand information can include the full demand intention of the target object, or can also include the object demand information, so that the first sub-agent can process based on the full demand information of the target object for the target task demand as prompt information, improve the representation accuracy of the sub-task execution elements to the demand intention, and ensure that the sub-task execution elements can maintain logical coherence and demand consistency with the sub-task demand conditions represented by the associated sub-tasks having a dependency relationship, avoid conflicts between the sub-task execution elements, and improve the accuracy and adaptability of the target sub-execution result.

[0108] According to an embodiment of the present disclosure, the target sub-task includes a shot data sub-task. Wherein, processing the sub-task demand information and the associated sub-execution result of the associated sub-task in the target task by the first sub-agent can include: processing the target script and the shot sub-task demand information for the shot data sub-task by the first sub-agent to obtain the shot sub-task execution element.

[0109] According to an embodiment of the present disclosure, the target script is determined by executing a script generation sub-task.

[0110] In one example, the target script can be generated by a script generation agent processing script generation sub-task demand information, and the target script satisfies the script demand condition in the script generation sub-task demand information.

[0111] According to an embodiment of the present disclosure, the sub-shooting task execution element is used to indicate that the display attribute of the target element in the target video, including at least one of the image element and the subtitle text, is adapted to the target text content of the target script. The display attribute includes at least one of the display duration and the display opportunity. For example, the sub-shooting task execution element can indicate that the sub-shooting data generated by the second intelligent agent will have the same display opportunity for the subtitle text, image element and script voice that represent the same content or semantics. Or, for another example, the sub-shooting task execution element can indicate that the sub-shooting data generated by the second intelligent agent will have the same display duration and display opportunity for the subtitle text, image element and script voice that represent the same demand intention or the same semantic content, so as to maintain the display consistency and coordination of the target video.

[0112] In one example, the first sub-intelligent agent can be used to process the sub-shooting task execution element and the associated sub-execution result of the target script, and at least one of the object demand information and the demand intention, to generate target sub-shooting data, so that the target sub-shooting data can display the key script text subtitle, the voice data corresponding to the key script text, and the key text explanation image element in the target video through the same display attribute for the key script text in the script text that matches the demand intention, so as to improve the display effect of the target video.

[0113] According to an embodiment of the present disclosure, the second sub-intelligent agent of the target intelligent agent executes the target sub-task based on the sub-task execution element to obtain the target sub-execution result can include: the second sub-intelligent agent of the target intelligent agent executes the target sub-task based on the sub-task execution element to obtain the sub-execution result; the third sub-intelligent agent of the target intelligent agent detects the sub-execution result to obtain the sub-task detection result; in the case that the sub-task detection result does not satisfy the sub-task demand condition in the sub-task execution element, the second sub-intelligent agent executes the target sub-task based on the sub-task detection result to obtain the target sub-execution result that satisfies the sub-task demand condition.

[0114] According to an embodiment of the present disclosure, the third sub-intelligent agent can be an intelligent agent different from the first sub-intelligent agent and the second sub-intelligent agent, and the first sub-intelligent agent, the second sub-intelligent agent and the third sub-intelligent agent in the target intelligent agent can be intelligent agents for executing different reasoning functions. The sub-execution result generated by the second sub-intelligent agent can be the execution result of the target sub-task without detection. By using the third intelligent agent to detect the sub-execution result, in the case that the sub-task detection result does not satisfy the sub-task demand condition, the second sub-intelligent agent is controlled to execute the target sub-task under the condition of paying more attention to the defect type indicated by the sub-task detection result, so that the generated target sub-task satisfies the sub-task demand condition.

[0115] According to an embodiment of the present disclosure, the sub-task requirement condition can be understood as a condition for indicating that the sub-execution result of the second sub-agent output needs to meet the sub-task execution element. The sub-task requirement condition can be input into the second sub-agent as prompt information to control the second sub-agent to execute the target sub-task. By detecting the sub-execution result based on the sub-task requirement condition through the third sub-agent, the sub-task detection result can more accurately represent the semantic understanding defects of the second sub-agent for the sub-task requirement condition, and the second sub-agent can be enabled to focus on the understanding defect type and the execution defect type for the sub-task requirement condition by executing the target sub-task based on the defect type indicated by the sub-task detection result through the second sub-agent, so as to realize that the second sub-agent is controlled to accurately execute the target sub-task based on the sub-task execution element based on the sub-task detection result as prompt information, and obtain the target sub-execution result meeting the sub-task requirement condition.

[0116] In one example, executing the target sub-task based on the sub-task detection result through the second sub-agent can include controlling the second sub-agent to execute the target sub-task based on the sub-task detection result and the sub-task execution element as prompt information, and obtaining a current intermediate sub-execution result. By detecting the current intermediate sub-execution result based on the sub-task requirement condition through the third sub-agent, in a case where the sub-task detection result indicates that the current intermediate sub-execution result meets the sub-task requirement condition, the current intermediate sub-execution result is determined as the target sub-execution result. In a case where the sub-task detection result indicates that the current intermediate sub-execution result does not meet the sub-task requirement condition, the second sub-agent can be utilized to iteratively execute the target sub-task based on the current sub-task detection result until the intermediate sub-execution result meets the sub-task requirement condition.

[0117] According to an embodiment of the present disclosure, by utilizing the semantic understanding ability and execution function of the first sub-agent, the second sub-agent and the third sub-agent respectively, the sub-task execution element reasoning, the target sub-task execution and the sub-execution result detection are respectively performed, which can realize that different agents are respectively used for task execution point analysis, task accurate execution and execution result detection and feedback for the target sub-task, and the target sub-task is executed in a fine-grained manner through multi-agent cooperation to avoid the execution illusion, semantic understanding error and the like of the generative model for constructing the same agent, and the execution accuracy and execution efficiency of the target sub-task are improved.

[0118] According to an embodiment of the present disclosure, detecting the sub-execution result through the third sub-agent of the target agent to obtain the sub-task detection result includes detecting the sub-execution result based on the sub-task requirement condition and the associated sub-execution result of the associated sub-task in the target task through the third sub-agent to obtain the sub-task detection result.

[0119] According to an embodiment of the present disclosure, the associated sub-task has a dependency relationship with the target sub-task, and the associated sub-execution result can represent a target sub-execution result meeting a sub-task requirement condition of the associated sub-task. The performing the target sub-task by the second sub-agent based on the sub-task detection result can include performing the target sub-task by the second sub-agent based on the sub-task detection result, the associated sub-execution result meeting the sub-task requirement condition of the associated sub-task, the object requirement information, and at least one of the requirement intention as prompt information. In this way, the second sub-agent can sufficiently learn the global information in the target task and the generated target sub-execution result, thereby improving the global information understanding ability of the second sub-agent, maintaining logical coherence, semantic consistency, and style coordination between the generated target sub-execution result and other target sub-execution results of the target sub-tasks, and further improving the display effect of the target video.

[0120] In one example, the first sub-agent, the second sub-agent, and the third sub-agent in the target agent can perform an inference function based on the full amount of information generated in the current target task execution process. For example, the second sub-agent can perform the target sub-task based on the full amount of information generated in the current target task execution process, and the third sub-agent can detect the sub-task execution result based on the full amount of information generated in the current target task execution process to obtain the sub-task detection result. The full amount of information generated in the current target task execution process includes object requirement information, requirement intention, sub-task execution elements, sub-execution results, sub-task execution elements of associated sub-tasks, and sub-execution results meeting the requirement conditions of the associated sub-tasks. In this way, the multiple sub-agents can pay attention to the sub-task execution elements and the target sub-execution results of other target sub-tasks based on the full amount of information for the target task, so as to maintain logical coherence and style consistency among the multiple target sub-tasks, avoid missing or misunderstanding the requirement attributes represented by the requirement information, and improve the display effect of the target video and the requirement matching degree with the target object.

[0121] According to an embodiment of the present disclosure, the target sub-task includes a graph element generation sub-task, and the sub-execution result includes a graph element for display in the target video. The graph element is displayed in synchronization with video caption text and voice data representing the same or similar meaning in the target video to improve the display effect of the target video for key information.

[0122] According to an embodiment of the present disclosure, the detecting the sub-execution result by the third sub-agent of the target agent includes detecting the graph element based on at least one of the target script and the target shot data, and a graph element requirement condition for the graph element generation sub-task by the third agent to obtain a graph element detection result.

[0123] According to an embodiment of the present disclosure, at least one of the target script and the target screenplay data is a correlation sub-execution result determined by executing a correlation subtask in the target task, and the correlation subtask includes a script generation subtask and / or a screenplay data generation subtask having a dependency relationship with the graph element generation subtask.

[0124] According to an embodiment of the present disclosure, by controlling the third intelligent agent to detect the graph element based on at least one of the target script and the target screenplay data and the graph element requirement condition for the graph element generation subtask as prompt information, the graph element detection result can more accurately represent the relevance between the currently generated graph element and the target script text and the target screenplay data that have been generated, so that the second intelligent agent can be controlled to execute the graph element generation subtask to generate the target graph element meeting the graph element requirement condition according to the defect type indicated by the graph element detection result. In this way, the target graph element can meet the same requirement intention as the target script and the target screenplay data, avoiding the occurrence of semantically irrelevant images and texts in the target video, maintaining the logical coherence and consistency of the expression of the content of the target video, and improving the display effect of the target video.

[0125] It should be noted that the target script can be determined by a script generation intelligent agent for executing the script generation subtask, or the target script can also be obtained based on other manners. The target screenplay data can be determined by a screenplay generation intelligent agent for executing the screenplay data generation subtask, or the target screenplay data can also be obtained based on other manners.

[0126] Figure 5 A schematic diagram illustrating the principle of executing a target subtask by a target intelligent agent according to an embodiment of the present disclosure is shown.

[0127] As shown in Figure 5 The target intelligent agent 510 can include a first sub-intelligent agent 511, a second sub-intelligent agent 512, and a third sub-intelligent agent 513. The target intelligent agent 510 is a graph element generation sub-intelligent agent for generating a target graph element. The target script 501 is taken as a correlation sub-execution result, and the first sub-intelligent agent 511 performs task element analysis based on the target script 501 and the object requirement information to obtain a plurality of subtask execution elements. The subtask execution elements can include “motion graph”, “sleep graph”, video theme requirement condition “preventing A disease”, style requirement condition “popular science”, and the like. The subtask execution elements can also include any task parameter, task execution step, and the like for controlling the execution process of the script generation subtask, such as literature type for reference, authoritative website name, and the like.

[0128] The second sub-agent 512 executes a graph element generation subtask based on the plurality of subtask execution elements to obtain a plurality of graph elements as an execution result. The third sub-agent 513 detects the plurality of graph elements based on the target script 501 and the subtask requirement condition to obtain a subtask detection result. In a case where the subtask detection result indicates that the graph element requirement condition is not met, the third sub-agent 513 feeds back the subtask detection result to the second sub-agent 512. The second sub-agent 512 executes a graph element generation subtask based on the subtask detection result and the subtask execution element to obtain a new graph element. In this way, the third sub-agent 513 can detect the new graph element based on the target script 501 and the subtask requirement condition until the graph element currently generated by the second sub-agent 512 meets the subtask requirement condition of the graph element, and the target graph element 502 is obtained.

[0129] According to an embodiment of the present disclosure, the subtask detection result includes a graph element detection result, the graph element detection result being obtained by detecting the graph element determined by executing the graph element generation subtask,

[0130] According to an embodiment of the present disclosure, the graph element detection result represents at least one of the following defect types: the image style of the graph element does not match the script style of the target script; the image content of the graph element does not meet the semantic similarity condition with the text semantic difference of the target script; and the arrangement manner of the graph element does not meet the preset arrangement manner condition.

[0131] According to an embodiment of the present disclosure, the image style and the script style can be any preset style attribute such as lively, serious, and two-dimensional. By fine-tuning the sub-agent using sample images or sample texts labeled with style attribute tags, the sub-agent can understand the image style of the graph element or understand the script style of the target script. By indicating that the image style of the current graph element of the second sub-agent does not match the script style, the second sub-agent is controlled to pay more attention to the consistency of the graph element style during the execution of the target subtask, and the efficiency of generating the target graph element is improved.

[0132] According to an embodiment of the present disclosure, the image content can be an element object such as a character, an animal, or the like represented by the graph element. The third sub-agent can obtain the image content by using a text recognition or object detection algorithm, and compare the image content with the corresponding script text content to determine whether the script text and the image content represent the same or similar element objects, thereby avoiding a situation where the displayed text or voice content in the target video is completely irrelevant to the image content.

[0133] The semantic similarity condition is not met when the image content of the graph element and the text semantic difference of the target script do not meet the semantic similarity condition. For example, the image content is a half-body photo of "Zhang San", and the text content corresponding to the target script is "Li Si has ever won the championship". It can be known that the text difference between the image content and the target script is large. The semantic similarity condition is not met when the image content of the graph element and the text semantic difference of the target script do not meet the semantic similarity condition.

[0134] According to an embodiment of the present disclosure, the arrangement manner of the graph element does not meet the preset arrangement manner condition. It can be understood that the arrangement shape, arrangement interval, and arrangement position of the plurality of sub-elements in the graph element do not meet the preset arrangement manner condition. By detecting the arrangement manner of the graph element, the target graph element can be displayed according to the demand intention of the target object in the target video, so as to improve the display effect of the target video, and avoid ambiguity caused by the arrangement manner of the graph element not meeting the preset arrangement manner condition, and improve the quality of the target video.

[0135] According to an embodiment of the present disclosure, the sub-task detection result includes a shot data detection result. The shot data detection result is obtained by detecting the shot data determined by executing the shot data generation sub-task.

[0136] According to an embodiment of the present disclosure, the shot data detection result represents at least one of the following defect types: the display time of the graph element indicated by the shot data does not match the playing time of the target voice segment; and the display duration of the graph element indicated by the shot data does not match the playing duration of the target voice segment.

[0137] According to an embodiment of the present disclosure, the shot data detection result indicates that the display time of the graph element focused by the second sub-agent matches the playing time of the target voice segment, and the display duration of the graph element matches the playing duration of the target voice segment. Therefore, the target voice segment and the target graph element in the target video can be synchronously displayed by using the more accurate target shot data, so as to improve the consistency and fluency of the target video.

[0138] Figure 6 A schematic diagram of a video generation method based on multi-agent cooperation according to an embodiment of the present disclosure is schematically shown.

[0139] As Figure 6As shown, the target object can send a video generation request based on a smart terminal device such as a notebook computer. The video generation request carries object demand information 601, which can record natural language information "need to generate a short video for preventing A disease for popular science". The task arrangement intelligent agent 611 as a specified intelligent agent performs intent recognition on the object demand information 601 to obtain a demand intent. And the first specified sub-intelligent agent of the task arrangement intelligent agent 611 performs a task arrangement task to obtain a plurality of initial sub-tasks and an initial dependency relationship. Through the second specified sub-intelligent agent in the task arrangement intelligent agent 611, the plurality of initial sub-tasks and the initial dependency relationship are detected to obtain a task arrangement detection result. In the case that the task arrangement detection result does not meet the video demand condition, the first specified sub-intelligent agent of the task arrangement intelligent agent 611 performs the task arrangement task based on the task arrangement detection result and the demand intent or the object demand information 601 in a loop until a plurality of target sub-tasks and a dependency relationship between the plurality of target sub-tasks meeting the video demand condition are obtained. According to the task types of the plurality of target sub-tasks and the ability description information of the plurality of candidate intelligent agents in the intelligent agent library, the task arrangement intelligent agent 611 selects a plurality of target sub-tasks adapted to the demand intent from the candidate intelligent agents in the intelligent agent library. And the plurality of target sub-tasks and the dependency relationship based on the topological structure are distributed to the plurality of target intelligent agents, so that the target intelligent agents can know the execution order of the plurality of target sub-tasks and the interaction relationship between the target intelligent agents and other target intelligent agents.

[0140] The plurality of target sub-tasks in the target task for generating the target video can include a script generation sub-task, a subtitle generation sub-task, a shot data generation sub-task, a graphic element generation sub-task, and a dynamic image generation sub-task. The plurality of target agents can be a script generation agent 621, a subtitle generation agent 622, a shot generation agent 623, a graphic element generation agent 624, and a dynamic image generation agent 625, respectively. The script generation agent 621, the subtitle generation agent 622, the shot generation agent 623, the graphic element generation agent 624, and the dynamic image generation agent 625 can each obtain an associated sub-execution result satisfying an associated sub-task requirement condition from an adjacent target agent having a dependency relationship based on a target sub-task execution order indicated by the dependency relationship, to execute the target sub-task and obtain a target sub-execution result. Meanwhile, the plurality of target sub-agents can share their respective sub-task execution elements, sub-task requirement conditions, target sub-execution results, search database contents, and object requirement information 601 input by the target object. In this way, each target agent performing a target sub-task later can obtain all the information generated in the process of executing the target task, so that the target agent can select the required content from the right information based on the functions and requirements of the plurality of sub-agents, and generate a target sub-execution result based on the obtained content. In this way, the execution order of the target sub-tasks represented based on the dependency relationship can be realized, the right information shared by each target agent can be constantly updated using the target sub-execution result, and the target sub-execution result can be automatically transmitted to the target agent performing the next target sub-task according to the dependency relationship. Finally, the target sub-execution results of the plurality of target sub-tasks are summarized according to the dependency relationship in the task arrangement result, and the plurality of target sub-execution results are transmitted to the video template compiling component 631 as video generation materials to generate a video. The video materials after the plurality of target sub-execution results are summarized can include a target script, a target subtitle text, target shot data, target graphic elements, and target dynamic images.

[0141] The video template compiling component 631 automatically selects the best video generation template according to the requirement intention of the target object and the summarized video materials, and renders a target video 602. A file recording the target video 602 is sent to the smart terminal device of the target object, so that the user can retrieve and enjoy or share the target video at any time.

[0142] As Figure 6As shown, each target intelligent agent internally includes a first sub-intelligent agent for performing sub-task execution element reasoning, a second sub-intelligent agent for performing a target sub-task execution function, and a third sub-intelligent agent for performing a sub-execution result detection function. The first sub-intelligent agent in each target intelligent agent internally acts as a task analysis sub-intelligent agent to analyze the full amount of information such as associated sub-execution results or object demand information transmitted by the upper-level intelligent agent, and obtain sub-task demand elements, so that the sub-task demand elements can plan and control the execution process of the target sub-task, so that the second sub-intelligent agent can accurately execute the target sub-task. The third sub-intelligent agent is used to detect the sub-execution result output by the second sub-intelligent agent, and inform the second sub-intelligent agent to improve the execution process of the execution target sub-task through the defect type indicated by the sub-task detection result, so as to form an analysis-execution-feedback closed-loop task execution mechanism through the cyclic execution of the target sub-task and the detection feedback, thereby quickly generating a target sub-execution result that meets the sub-task demand conditions. In this way, the target video can be quickly generated by quickly generating the target sub-execution result, the video industrial production can be realized based on the cooperation of multiple target intelligent agents and the cooperation of multiple sub-intelligent agents, and the video production can be realized on a large scale and customized.

[0143] The sub-intelligent agents involved in the present embodiment, including but not limited to the first designated sub-intelligent agent, the first sub-intelligent agent, etc., can be used as the basic unit of the intelligent agent. The sub-intelligent agent is constructed based on a large model such as a large language model or a visual large model, and can interact with external tools such as search engines, material libraries, and code compilers to execute the reasoning function of the intelligent agent, and can efficiently complete tasks such as video script writing, creative copy generation, and multi-modal video material production.

[0144] The target intelligent agent and the designated intelligent agent involved in the present embodiment can be based on the execution functions of the multiple sub-intelligent agents to realize the cooperation of the multiple sub-intelligent agents to complete relatively complex target sub-tasks or task arrangement tasks, etc. The multiple sub-intelligent agents in the target intelligent agent can assume the functional roles of task analysis, task execution, and result feedback, realize the analysis-execution-feedback closed-loop function for the target sub-task or task arrangement, realize the continuous optimization and adjustment of the target sub-task execution result based on the defect type of the feedback, and realize the fast and efficient execution of the target sub-task.

[0145] According to the embodiments of the present disclosure, a general intelligent agent can be used to detect the sub-execution result of each target sub-task, which is suitable for executing a general detection task. However, it is not limited thereto. A plurality of sub-specialist intelligent agents can also be set for the main intelligent agent, and the sub-specialist intelligent agents are suitable for executing an individualized detection task.

[0146] The model structure configured in the general agent or the sub-specialist agent is not limited, for example, the deep learning model can include one or more of attention mechanism, convolution network, recurrent network, long short-term memory network, but is not limited thereto, and can be a large language model with more powerful functions. As long as it is a model that can be used to perform detection tasks. The difference is that by using multiple sub-specialist agents, a mixture of experts mechanism can be formed, and multiple sub-specialist agents with pertinence and individualization are combined, and different sub-specialist agents are selected to perform detection tasks for different sub-task requirements of detection tasks.

[0147] Figure 7A An effect schematic diagram of a target video is schematically shown according to an embodiment of the present disclosure.

[0148] As Figure 7A shown, the target video 700 can be a target voice representing a target script by driving the lip action of a virtual object 701. The target voice will display the subtitle text "don't eat sugar, don't drink sugar-containing beverages" representing the target voice in the video screen of the target video 700 during the playing process, wherein the subtitle text can be highlighted based on underlining and bolding. The target graph element 702 can be an image content representing the text content of the subtitle text "don't eat sugar, don't drink sugar-containing beverages". The subtitle text "don't eat sugar, don't drink sugar-containing beverages" representing the same meaning, the voice data segment and the target graph element 702 will be displayed synchronously in the video screen of the target video based on the same display time and display duration indicated by the target shot data.

[0149] Figure 7B An effect schematic diagram of a health popular science video is schematically shown according to an embodiment of the present disclosure.

[0150] As Figure 7BAs shown, the target video can be a health popularization video T710, and the virtual image in the health popularization video T710 can explain knowledge related to taking medicine based on target voice data corresponding to the target script. The first image element T711 and the second image element T712 in the health popularization video T710 can be obtained by the image element generation agent performing the image element generation subtask. The first image element T711 and the second image element T712 satisfy the theme requirement condition of the user requirement “knowledge explanation of medicine taking”. The subtitle text “recall whether special food has been eaten within 48 hours…” in the health popularization video T710 is a target sub-execution result obtained by the subtitle generation agent performing the subtitle generation subtask. The first image element T711 and the subtitle text “recall whether special food has been eaten within 48 hours…” in the health popularization video T710 can be displayed based on a suitable display time and a display duration, and the font size and color of the key information “48 hours” in the subtitle text can be modified to make the health popularization video T710 more intuitive for the audience to pay attention to whether the dragon fruit and the medicine displayed in the first image element T711 are taken at the same time in the process of taking medicine or treatment, improve the display effect of the health popularization video, and improve the matching degree between the target video and the user requirement.

[0151] It should be noted that, Figure 7B The black box in the health popularization video T710 shown is to protect relevant private information or important information, and is not a display defect of the target video.

[0152] Figure 7C An effect schematic diagram of a legal popularization video according to an embodiment of the present disclosure is schematically shown.

[0153] As Figure 7CAs shown, the target video can be a legal popularization video T720, in which a virtual image can explain relevant legal knowledge about prohibition of carrying restricted knives based on target voice data corresponding to the target script. The third image element T721 and the fourth image element T722 in the legal popularization video T720 can be obtained by the image element generation agent performing the image element generation subtask. The third image element T721 and the fourth image element T722 satisfy the theme requirement condition of "restricted knife explanation" of the user requirement. The subtitle text "They are dangerous because of their potential" in the legal popularization video T720 is a target sub-execution result obtained by the subtitle generation agent performing the subtitle generation subtask. The third image element T721 and the subtitle text "They are dangerous because of their potential" in the legal popularization video T720 can be displayed based on the adapted display time and display duration, so as to make the legal popularization video T720 more intuitive and timely for the audience to focus on various types of restricted knives in the third image element T721 in the explanation process of the virtual image, to strengthen the understanding of the relevant legal knowledge, to improve the display effect of the legal popularization video, and to improve the matching degree between the target video and the user requirement.

[0154] It should be noted that, Figure 7C The black square in the legal popularization video T720 shown is to protect relevant private information or important information, and is not a display defect of the target video.

[0155] It should be noted that the acquisition, processing and display of data involved in any embodiment of the present disclosure are implemented under the condition that relevant user or agency authorization is obtained, and necessary encryption or desensitization measures are adopted in the implementation of the embodiments of the present disclosure to avoid information leakage.

[0156] Figure 8 A block diagram of a video generation apparatus based on multi-agent cooperation according to an embodiment of the present disclosure is schematically shown.

[0157] As Figure 8 shown, the video generation apparatus based on multi-agent cooperation 800 includes an execution module 810 and a target video generation module 820.

[0158] The execution module 810 is configured to execute at least one target agent to execute a target subtask in a target task, and obtain a target sub-execution result.

[0159] The target video generation module 820 is configured to generate a target video based on a target sub-execution result, wherein the target agent performs the following operations: obtaining sub-task demand information of a target sub-task; determining, by a first sub-agent of the target agent, a sub-task execution element for controlling an execution process of the target sub-task based on the sub-task demand information; and executing, by a second sub-agent of the target agent, the target sub-task based on the sub-task execution element to obtain the target sub-execution result.

[0160] According to an embodiment of the present disclosure, the target agent is configured to perform the following operations to execute the target sub-task based on the sub-task execution element by the second sub-agent of the target agent to obtain the target sub-execution result: executing, by the second sub-agent of the target agent, the target sub-task based on the sub-task execution element to obtain a sub-execution result; detecting, by a third sub-agent of the target agent, the sub-execution result to obtain a sub-task detection result; and in a case where the sub-task detection result does not satisfy a sub-task demand condition in the sub-task execution element, executing, by the second sub-agent, the target sub-task based on the sub-task detection result to obtain the target sub-execution result satisfying the sub-task demand condition.

[0161] According to an embodiment of the present disclosure, the target agent is configured to perform the following operations to detect, by the third sub-agent of the target agent, the sub-execution result to obtain the sub-task detection result: detecting, by the third sub-agent, the sub-execution result based on the sub-task demand condition and an associated sub-execution result of an associated sub-task in the target task to obtain the sub-task detection result, wherein the associated sub-task has a dependency relationship with the target sub-task.

[0162] According to an embodiment of the present disclosure, the target sub-task includes a graph element generation sub-task, and the sub-execution result includes a graph element for display in the target video.

[0163] According to an embodiment of the present disclosure, the target agent is configured to perform the following operations to detect, by the third sub-agent of the target agent, the sub-execution result: detecting, by the third agent, the graph element based on at least one of a target script and target shot data and a graph element demand condition for the graph element generation sub-task to obtain a graph element detection result, wherein the at least one of the target script and the target shot data is an associated sub-execution result, the associated sub-execution result is determined by executing an associated sub-task in the target task, and the associated sub-task includes a script generation sub-task and / or a shot data generation sub-task having a dependency relationship with the graph element generation sub-task.

[0164] According to an embodiment of the present disclosure, the sub-task detection result includes the graph element detection result, and the graph element detection result is obtained by detecting the graph element determined by executing the graph element generation sub-task.

[0165] According to an embodiment of the present disclosure, the image element detection result represents at least one of the following defect types: an image style of the image element does not match a script style of the target script; an image content of the image element does not satisfy a semantic similarity condition in terms of a semantic difference degree with respect to a text semantic of the target script; and an arrangement manner of the image element does not satisfy a preset arrangement manner condition.

[0166] According to an embodiment of the present disclosure, the subtask detection result includes a shot data detection result, which is obtained by detecting shot data determined by performing a shot data generation subtask; and the shot data detection result represents at least one of the following defect types: a display time of the image element indicated by the shot data does not match a playing time of the target voice segment; and a display duration of the image element indicated by the shot data does not match a playing duration of the target voice segment.

[0167] According to an embodiment of the present disclosure, the target intelligent agent is configured to perform the following operation to determine, by using a first sub-intelligent agent of the target intelligent agent, a subtask execution element for controlling an execution process of a target subtask based on subtask demand information: processing, by using the first sub-intelligent agent, the subtask demand information and an associated sub-execution result of an associated subtask in the target task to obtain the subtask execution element, wherein the associated subtask and the target subtask have a dependency relationship.

[0168] According to an embodiment of the present disclosure, the target intelligent agent is configured to perform the following operation to determine, by using a first sub-intelligent agent of the target intelligent agent, a subtask execution element for controlling an execution process of a target subtask based on subtask demand information: processing, by using the first sub-intelligent agent, the subtask demand information, an associated sub-execution result of an associated subtask in the target task, and object demand information of a target object to obtain the subtask execution element, wherein the object demand information represents a demand attribute of the target object for the target video.

[0169] According to an embodiment of the present disclosure, the target subtask includes a shot data subtask.

[0170] According to an embodiment of the present disclosure, the target intelligent agent is configured to perform the following operation to determine, by using a first sub-intelligent agent of the target intelligent agent, a subtask execution element for controlling an execution process of a target subtask based on subtask demand information:

[0171] processing, by using the first sub-intelligent agent, the target script and shot subtask demand information for the shot data subtask to obtain a shot subtask execution element, wherein the target script is determined by performing a script generation subtask; and wherein the shot subtask execution element is used to indicate that a display attribute of a target element in the target video is adapted to target text content of the target script, the target element including at least one of an image element and a subtitle text, and the display attribute including at least one of a display duration and a display time.

[0172] According to an embodiment of the present disclosure, the subtask execution element includes at least one of the following subtask requirement conditions: a style consistency condition, representing an expression style of a target sub-execution result, which matches a requirement style represented by the object requirement information, wherein the expression style includes at least one of a text expression style, a voice expression style, and a video picture expression style; and a logical coherence condition, representing that the target task has logical coherence between a plurality of target sub-execution results.

[0173] According to an embodiment of the present disclosure, the target task includes a plurality of target subtasks.

[0174] The video generation apparatus 800 based on multi-agent cooperation is further configured to perform the following operations by using the specified agent: performing intent recognition on the object requirement information of the target object to obtain a requirement intent; determining a plurality of target subtasks and a dependency relationship between the plurality of target subtasks based on the requirement intent; and issuing the plurality of target subtasks and the dependency relationship between the plurality of target subtasks to the plurality of target agents.

[0175] According to an embodiment of the present disclosure, determining the plurality of target subtasks and the dependency relationship between the plurality of target subtasks based on the requirement intent includes: performing a task arrangement task by using a first specified sub-agent of the specified agent based on the requirement intent to obtain a plurality of initial subtasks and an initial dependency relationship; performing a task arrangement detection subtask by using a second specified sub-agent of the specified agent based on a video requirement condition represented by the requirement intent to perform task arrangement detection on the plurality of initial subtasks and the initial dependency relationship to obtain a task arrangement detection result; and in a case where the task arrangement detection result does not satisfy the video requirement condition, performing a task arrangement subtask by using the first specified sub-agent based on the requirement intent and the task arrangement detection result to obtain the plurality of target subtasks and the dependency relationship between the plurality of target subtasks that satisfy the video requirement condition.

[0176] According to an embodiment of the present disclosure, the task arrangement detection result represents at least one of the following defect types: an execution logic conflict exists in the initial dependency relationship; and a plurality of initial subtask types of the plurality of initial subtasks do not match a requirement type represented by the requirement intent.

[0177] According to an embodiment of the present disclosure, the subtask requirement information for the target subtask is determined by using the specified agent based on the requirement intent.

[0178] According to an embodiment of the present disclosure, obtaining the subtask requirement information for the target subtask includes: performing requirement information detection on the object requirement information of the target object by using the target agent based on a subtask type of the target subtask to obtain the subtask requirement information.

[0179] According to an embodiment of the present disclosure, the target subtask includes at least one of the following: a script generation subtask, a shot data generation subtask, a graph element generation subtask, and a subtitle generation subtask.

[0180] Figure 9 A structural block diagram of an AI agent of artificial intelligence is schematically shown according to an embodiment of the present disclosure.

[0181] In an embodiment of the present disclosure, as shown in Figure 9 The AI agent 900 can include an input module 910, a processing module 920, and an output module 930.

[0182] The input module 910 is configured to receive input information.

[0183] The processing module 920 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the video generation method based on multi-agent collaboration provided according to an embodiment of the present disclosure.

[0184] The output module 930 is configured to output the output information obtained by the processing module.

[0185] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (for example, a user or an external environment), and converting them into a format that the AI agent 900 can understand and process. The input module 910 is the first link for the AI agent 900 to interact with the outside world, which enables the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to it.

[0186] In an example, the input module 910 can input the object demand information, subtask demand information, and the like described in the foregoing.

[0187] In an example, the processing module 920 is the core support for the AI agent 900 to process complex tasks. The processing module 920 can execute the video generation method based on multi-agent collaboration described in the foregoing.

[0188] In an example, the performance of the processing module 920 can be closely related to the large model on which the AI agent 900 is based. In order to fully exert the capabilities of the large model, the internal structure of the processing module 920 can be designed to be highly configurable and extensible in order to cope with various different types of tasks and demands in real scenarios.

[0189] In an example, after obtaining the demand voice, the processing module 920 can process the object demand information by using the target agent based on the large model to obtain a target video, and deliver the target video to the output module 930.

[0190] It can be understood that, although the large language model has excellent language understanding and generation capabilities, it, like a human, can solve only a limited number of tasks without the aid of any tools. When the AI agent 900 is endowed with the ability to call tools, it can implement tasks such as completing mathematical operations with the aid of a calculator, completing data analysis with the aid of Python, and completing weather forecasts with the aid of a search engine.

[0191] In an example, the output module 930 can output the target video described in the foregoing.

[0192] The AI agent 900 according to the embodiments of the present disclosure can simply and effectively improve the intelligent degree and improve the flexibility and versatility.

[0193] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0194] According to the embodiments of the present disclosure, an electronic device comprises at least one processor, and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0195] According to the embodiments of the present disclosure, a non-transitory computer readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to perform the method described above.

[0196] According to the embodiments of the present disclosure, a computer program product comprises a computer program, and the computer program implements the method described above when executed by a processor.

[0197] Figure 10 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0198] As Figure 10As shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0199] A plurality of components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006 such as a keyboard, a mouse, etc., an output unit 1007 such as various types of displays, speakers, etc., a storage unit 1008 such as a magnetic disk, an optical disk, etc., and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0200] The computing unit 1001 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the multi-agent collaboration based video generation method. For example, in some embodiments, the multi-agent collaboration based video generation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the multi-agent collaboration based video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the multi-agent collaboration based video generation method by any other appropriate means, e.g., by means of firmware.

[0201] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0202] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0203] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0204] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0205] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0206] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0207] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation, as long as the desired results of the technology disclosed in the present disclosure are achieved.

[0208] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above.

Claims

1. A method for generating a video based on multi-agent cooperation, comprising: executing a target sub-task in a target task by using at least one target agent to obtain a target sub-execution result; and generating a target video based on the target sub-execution result, wherein the target agent performs the following operations: obtaining sub-task requirement information for the target sub-task; processing the sub-task requirement information, an associated sub-execution result of an associated sub-task in the target task, and object requirement information of a target object by using a first sub-agent to obtain a sub-task execution element, the associated sub-task having a dependency relationship with the target sub-task, and the object requirement information representing a demand attribute of the target object for the target video; executing the target sub-task based on the sub-task execution element by using a second sub-agent of the target agent to obtain the target sub-execution result. The execution of the target sub-task based on the sub-task execution element by using the second sub-agent of the target agent to obtain the target sub-execution result comprises:

2. The method of claim 1, wherein, executing the target sub-task based on the sub-task execution element by using the second sub-agent of the target agent to obtain a sub-execution result; detecting the sub-execution result by using a third sub-agent of the target agent to obtain a sub-task detection result; in a case where the sub-task detection result does not satisfy a sub-task requirement condition in the sub-task execution element, executing the target sub-task based on the sub-task detection result by using the second sub-agent to obtain a target sub-execution result satisfying the sub-task requirement condition. The detection of the sub-execution result by using the third sub-agent of the target agent to obtain a sub-task detection result comprises:

3. The method of claim 2, wherein, detecting the sub-execution result based on the sub-task requirement condition and an associated sub-execution result of an associated sub-task in the target task by using the third sub-agent to obtain the sub-task detection result, wherein the associated sub-task has a dependency relationship with the target sub-task. The target sub-task comprises a graph element generation sub-task, and the sub-execution result comprises a graph element for display in the target video.

4. The method of claim 2, wherein, The detection of the sub-execution result by using the third sub-agent of the target agent comprises: detecting the graph element based on at least one of a target script and target shot data and a graph element requirement condition for the graph element generation sub-task by using the third sub-agent to obtain a graph element detection result, wherein the at least one of the target script and the target shot data is an associated sub-execution result, the associated sub-execution result being determined by executing an associated sub-task in the target task, and the associated sub-task comprises a script generation sub-task and / or a shot data generation sub-task having a dependency relationship with the graph element generation sub-task. The sub-task detection result comprises a graph element detection result, which is obtained by detecting a graph element determined by executing the graph element generation sub-task.

5. The method of claim 2, wherein, The graph element detection result represents at least one of the following defect types: ​ An image style of the image element does not match a script style of the target script; An image content of the image element does not meet a semantic similarity condition in terms of a text semantic difference from the target script; An arrangement manner of the image element does not meet a preset arrangement manner condition.

6. The method of claim 2, wherein, The subtask detection result includes shot data detection result, which is obtained by detecting the shot data determined by executing the shot data generation subtask; The shot data detection result represents at least one of the following defect types: A display timing of the image element indicated by the shot data does not match a playing timing of the target voice segment; A display duration of the image element indicated by the shot data does not match a playing duration of the target voice segment.

7. The method of claim 1, wherein, The target subtask includes a shot data subtask; The processing of the subtask requirement information and the associated subexecution result of the associated subtask in the target task by the first subagent includes: processing the target script and the shot subtask requirement information for the shot data subtask by the first subagent to obtain a shot subtask execution element, wherein the target script is determined by executing the script generation subtask; The shot subtask execution element is used to indicate that a display attribute of a target element in the target video is adapted to target text content of the target script, the target element includes at least one of an image element and a subtitle text, and the display attribute includes at least one of a display duration and a display timing.

8. The method of claim 1 or 2, wherein, The subtask execution element includes at least one of the following subtask requirement conditions: A style consistency condition, representing that an expression style of the target subexecution result matches a requirement style represented by the object requirement information, wherein the expression style includes at least one of a text expression style, a voice expression style, and a video picture expression style; A logical coherence condition, representing that a plurality of target subexecution results of the target task have logical coherence.

9. The method of claim 1, wherein, The target task includes a plurality of target subtasks; the method further includes executing the following operations by using a specified agent: performing intent recognition on object requirement information of a target object to obtain a requirement intent; determining a plurality of target subtasks and a dependency relationship between the plurality of target subtasks based on the requirement intent; and issuing the plurality of target subtasks and the dependency relationship between the plurality of target subtasks to a plurality of target agents.

10. The method of claim 9, wherein, The determination of the plurality of target subtasks and the dependency relationship between the plurality of target subtasks based on the requirement intent includes: executing a task arrangement task by a first specified subagent of the specified agent based on the requirement intent to obtain a plurality of initial subtasks and an initial dependency relationship; executing a task arrangement detection subtask by a second specified subagent of the specified agent based on a video requirement condition represented by the requirement intent to obtain a task arrangement detection result. In a case where the task arrangement detection result does not satisfy the video demand condition, the first specified sub-agent is used to execute the task arrangement sub-task based on the demand intention and the task arrangement detection result, to obtain a plurality of target sub-tasks satisfying the video demand condition and a dependency relationship between the plurality of target sub-tasks.

11. The method of claim 10, wherein, The task arrangement detection result represents at least one of the following defect types: The initial dependency relationship has an execution logic conflict; A plurality of initial sub-task types of the plurality of initial sub-tasks do not match the demand type represented by the demand intention.

12. The method of claim 9, wherein, The sub-task demand information for the target sub-task is determined by the specified agent based on the demand intention.

13. The method of claim 1, wherein, The sub-task demand information for the target sub-task is determined by the specified agent based on the demand intention. The sub-task demand information for the target sub-task is determined by the specified agent based on the demand intention.

14. The method of claim 1, wherein, The target sub-task includes at least one of the following: a script generation sub-task, a shot data generation sub-task, a graph element generation sub-task, and a subtitle generation sub-task.

15. A video generation device based on multi-agent cooperation, comprising: an execution module configured to execute a target sub-task in a target task by using at least one target agent to obtain a target sub-execution result; and a target video generation module configured to generate a target video based on the target sub-execution result, wherein the target agent performs the following operations: obtain sub-task demand information for the target sub-task; use a first sub-agent to process the sub-task demand information, an associated sub-execution result of an associated sub-task in the target task, and object demand information of a target object, to obtain a sub-task execution element, the associated sub-task having a dependency relationship with the target sub-task, and the object demand information representing a demand attribute of the target object for a target video; use a second sub-agent of the target agent to execute the target sub-task based on the sub-task execution element to obtain the target sub-execution result.

16. An artificial intelligence agent, comprising: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute a method according to any one of claims 1 to 14 by calling the large model to obtain output information; an output module configured to output the output information obtained by the processing module.

17. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 14.

18. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to execute the method of any one of claims 1 to 14.

19. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Mathematical physics explanation video generation method and device based on multi-agent cooperation

    CN119277169A