Video generation method and device based on multi-agent cooperation, agents and equipment
Through the video generation method based on multi-agent collaboration, the execution results of the target agent are detected and adjusted, and the problem of insufficient accuracy and visual expression of animation teaching video content is solved, the needs of personalized learning and instant feedback are realized, and the video generation quality is improved.
Patent Information
- Application Number
- CN202510355095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is difficult to meet the needs of personalized learning and instant feedback, and the content accuracy and visual expression of animation teaching videos are insufficient.
The video generation method based on multi-agent collaboration is adopted to obtain the sub-execution results of the target agent performing the target sub-task through the main agent, detect whether the sub-task results meet the sub-task requirements conditions, and control the target agent to re-execute the task when it is not met, until the target video is generated after the conditions are met.
It improves the content accuracy and visual expression ability of the target video, meets the needs of personalized learning and instant feedback, and improves the quality of video generation.
Smart Images

Figure CN120151561A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to technologies such as deep learning, large models, human-computer interaction, teaching, and video generation. A video generation method, device, intelligent agent, electronic device, storage medium, and program product based on multi-agent collaboration. Background Art
[0002] With the rapid development of Internet technology, online education has been widely popularized. Knowledge can be imparted not only through live broadcasts and recordings by real people, but also by directly generating animated teaching videos using artificial intelligence technology and imparting knowledge using the animated teaching videos. The content accuracy and visual expression of the animated teaching videos generated based on artificial intelligence technology have become limiting factors for their development. Summary of the Invention
[0003] The present disclosure provides a video generation method, device, intelligent agent, electronic device, storage medium, and program product based on multi-agent collaboration.
[0004] According to one aspect of the present disclosure, there is provided a video generation method based on multi-agent collaboration, which is applied to a main intelligent agent and includes: obtaining sub-execution results determined by at least one target intelligent agent executing a target sub-task, where the target task for generating a target video includes the above-mentioned target sub-task; detecting the above-mentioned sub-execution results to obtain sub-task detection results; in the case where the sub-task detection results indicate that the above-mentioned sub-execution results do not meet the sub-task requirement conditions for the above-mentioned target sub-task, controlling the above-mentioned target intelligent agent to execute the above-mentioned target sub-task based on the above-mentioned sub-task detection results to obtain target sub-execution results that meet the above-mentioned sub-task requirement conditions; and generating the above-mentioned target video based on the above-mentioned target sub-execution results.
[0005] According to another aspect of the present disclosure, there is provided a video generation device based on multi-agent collaboration, which is applied to a main intelligent agent and includes: an obtaining module, configured to obtain sub-execution results determined by at least one target intelligent agent executing a target sub-task, where the target task for generating a target video includes the above-mentioned target sub-task; a detection module, configured to detect the above-mentioned sub-execution results to obtain sub-task detection results; a task execution module, configured to, in the case where the sub-task detection results indicate that the above-mentioned sub-execution results do not meet the sub-task requirement conditions for the above-mentioned target sub-task, control the above-mentioned target intelligent agent to execute the above-mentioned target sub-task based on the above-mentioned sub-task detection results to obtain target sub-execution results that meet the above-mentioned sub-task requirement conditions; and a video generation module, configured to generate the above-mentioned target video based on the above-mentioned target sub-execution results.
[0006] According to another aspect of the present disclosure, there is provided an intelligent agent for artificial intelligence, including: an input module for receiving input information; a processing module for executing the above method based on the input information received by the above input module to obtain output information; and an output module for outputting the output information obtained by the above processing module.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the above method when executed by a processor.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which a video generation method and apparatus based on multi-agent collaboration can be applied according to an embodiment of the present disclosure;
[0013] Figure 2 Schematically shows a flowchart of a video generation method based on multi-agent collaboration according to an embodiment of the present disclosure;
[0014] Figure 3A Schematically shows an application diagram of the execution result of a detection sub according to an embodiment of the present disclosure;
[0015] Figures 3B to 3F Schematically shows a schematic diagram of a target video according to an embodiment of the present disclosure;
[0016] Figure 4 Schematically shows a scene diagram of a lecture video according to an embodiment of the present disclosure;
[0017] Figure 5Schematically shows an application schematic diagram of generating target graphic data code according to an embodiment of the present disclosure;
[0018] Figure 6A Schematically shows a schematic diagram of target graphic data according to an embodiment of the present disclosure;
[0019] Figure 6B Schematically shows a scene schematic diagram of generating a target video according to an embodiment of the present disclosure;
[0020] Figure 7 Schematically shows an interaction scene diagram according to an embodiment of the present disclosure;
[0021] Figure 8 Schematically shows a block diagram of a video generation device based on multi-agent collaboration according to an embodiment of the present disclosure;
[0022] Figure 9 Schematically shows a structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure; and
[0023] Figure 10 Schematically shows a block diagram of an electronic device suitable for implementing a video generation method based on multi-agent collaboration according to an embodiment of the present disclosure. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0026] In the field of problem-solving teaching, although some online education platforms provide recorded or live courses, it is still difficult to meet the needs of personalized learning and instant feedback. In addition, artificial intelligence technology has also been used to generate animated teaching videos, and the animated teaching videos are published through the Internet. However, compared with the way of real-person teaching, the content accuracy and visual expression of the animated teaching videos cannot be satisfied.
[0027] In view of this, embodiments of the present disclosure provide a method for generating videos based on multi-agent collaboration, which can generate a problem-solving video for explaining a problem raised by a target object, such as a student, based on the collaboration of multiple agents. While meeting the needs of personalized learning and instant feedback, the main agent is used to detect the sub-execution results determined by the target sub-tasks executed by the target agents, so as to ensure that the target video is generated using the target sub-execution results that meet the requirements of the sub-tasks, thereby improving the content accuracy and visualization expression ability of the target video.
[0028] Figure 1 Schematically shows an exemplary system architecture to which the method and apparatus for generating videos based on multi-agent collaboration according to embodiments of the present disclosure can be applied.
[0029] It should be noted that Figure 1 The illustration is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the method and apparatus for generating videos based on multi-agent collaboration can be applied may include a terminal device, but the terminal device can implement the method and apparatus for generating videos based on multi-agent collaboration provided by embodiments of the present disclosure without interacting with the server.
[0030] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0031] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0032] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0033] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports the content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0034] It should be noted that the video generation method based on multi-agent collaboration provided by the embodiments of the present disclosure can generally be executed by terminal devices 101, 102, or 103 configured with a master agent. Correspondingly, the video generation device based on multi-agent collaboration provided by the embodiments of the present disclosure can also be set in terminal devices 101, 102, or 103.
[0035] Alternatively, the video generation method based on multi-agent collaboration provided by the embodiments of the present disclosure can generally also be executed by server 105 configured with a master agent. Correspondingly, the video generation device based on multi-agent collaboration provided by the embodiments of the present disclosure can generally be set in server 105. The video generation method based on multi-agent collaboration provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the video generation device based on multi-agent collaboration provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0036] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0037] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0038] As Figure 2 shown, the method includes operations S210 to S240.
[0039] In operation S210, obtain the sub-execution results determined by at least one target agent for executing the target sub-task.
[0040] In operation S220, detect the sub-execution results to obtain the sub-task detection results.
[0041] In operation S230, when the sub-task detection result indicates that the sub-execution result does not meet the sub-task requirement conditions for the target sub-task, the target agent is controlled based on the sub-task detection result to execute the target sub-task, and a target sub-execution result that meets the sub-task requirement conditions is obtained.
[0042] In operation S240, a target video is generated based on the target sub-execution result.
[0043] This method is applied to the main agent. The main agent is configured on a server or a terminal device. The target agent can be configured on the same or different electronic devices as the main agent, such as a server or a terminal device, as long as they can communicate and the main agent can control the target agent.
[0044] Both the main agent and the target agent can be agents that can perceive the environment and take actions to achieve specific goals. Both the main agent and the target agent can be designed to be highly scalable. For example, they can be designed as agents configured with large models. The type of the large model is not limited. For example, the large model can include one or more combinations of a large language model (LLM), a large vision model (LVM), or a multimodal large model (MLM). The difference between the main agent and at least one target agent lies in that the configured large models have different large model knowledge and can perform different types of tasks.
[0045] For example, the target video can be generated through the cooperation of the main agent and at least one target agent. The target task of producing the target video can be split into at least one target sub-task, and the at least one target agent executes the target sub-task to obtain a sub-execution result. The target agent detects the sub-execution result to determine whether the sub-execution result meets the sub-task requirement conditions. When the sub-task requirement conditions are met, the sub-execution result is used as the target sub-execution result. When the sub-execution result does not meet the sub-task requirement conditions, the main agent can control the target agent to re-execute the target sub-task until the obtained sub-execution result meets the sub-task requirement conditions. The target agent generates the target video based on at least one target sub-execution result.
[0046] Taking a lecture video as an example of the target video, the target task can be split into a problem-solving sub-task, a copywriting generation sub-task, a code generation sub-task, a video generation sub-task, etc. The target agents can include a problem-solving agent for executing the problem-solving sub-task, a copywriting generation agent for executing the copywriting generation sub-task, a code generation agent for executing the code generation sub-task, and a main agent for executing the video generation sub-task.
[0047] The problem-solving agent can execute the problem-solving subtask to obtain problem-solving data. The target agent can detect the problem-solving data to obtain a subtask detection result related to the problem-solving data. When the subtask detection result indicates that the problem-solving data meets the subtask requirement conditions corresponding to the problem-solving subtask, the problem-solving data can be used as the target problem-solving data; otherwise, the problem-solving agent is used to re-execute the problem-solving subtask to obtain updated problem-solving data until the updated problem-solving data meets the subtask requirement conditions corresponding to the problem-solving subtask, and the problem-solving data is used as the target problem-solving data. Similar to the cooperation mode of the problem-solving agent and the main agent, the target explanation text determined by executing the copywriting generation subtask based on the target problem-solving data, the target graphic data code determined by executing the code generation subtask based on the target explanation text, and the target video determined by executing the video generation subtask based on the target graphic data code can be obtained in sequence.
[0048] In the embodiments of the present disclosure, multiple agents are used to cooperate to generate a target video, which can use the main agent to obtain the sub-execution result determined by the target agent executing the target subtask, and determine whether the sub-execution result meets the subtask requirement conditions by detecting the sub-execution result, so that the target agent can re-execute the target subtask when the sub-execution result does not meet the subtask requirement conditions. Furthermore, through the cooperation of multiple agents, the target video generated according to the target sub-execution result can meet the requirements of content accuracy and visual expression, improving the quality of video generation.
[0049] Optionally, the main agent can control at least one target agent to execute at least one target subtask to obtain at least one sub-execution result. Without detecting the sub-execution result, the at least one sub-execution result is directly used as the at least one target sub-execution result. A target video is generated based on the at least one target sub-execution result.
[0050] Compared with the method of not detecting the sub-execution result, using the method provided by the embodiments of the present disclosure to detect the sub-execution result by the main agent to determine whether the sub-execution result meets the subtask requirement conditions based on the subtask detection result can improve the quality of the target sub-execution result, and further improve the quality of the target video generated from the target sub-execution result.
[0051] The above provides a general description of the video generation method based on multi-agent cooperation. The following will further explain the operation S220 as Figure 2 shown.
[0052] According to an embodiment of the present disclosure, detecting the sub-execution result may include: detecting the sub-execution result using the sub-task requirement conditions. However, it is not limited thereto. The sub-execution result may also be detected based on the sub-task requirement conditions and the associated target sub-execution result of the associated sub-task in the target task.
[0053] There is an execution dependency relationship between the associated sub-task and the target sub-task. The associated target sub-execution result satisfies the associated sub-task requirement conditions for the associated sub-task.
[0054] The execution dependency relationship can be understood as an influence relationship existing between the target sub-task and the associated sub-task. For example, the execution of the target sub-task depends on the execution of the associated sub-task. Specifically, the associated sub-task is the target sub-task before at least one step of the target sub-task, and the associated target sub-execution result output by the associated sub-task affects the execution of the target sub-task. For example, for the copywriting generation sub-task, the target execution sub-result of the solution sub-task, such as the target problem-solving data, is required as the input data, so there is an execution dependency relationship between the solution sub-task and the copywriting generation sub-task. However, it is not limited thereto. It can also be understood that the target sub-task and the associated sub-task jointly affect the execution of other sub-tasks. For example, the target sub-execution results of the image generation sub-task and the audio generation sub-task jointly affect the video synthesis sub-task.
[0055] The main agent is used to evaluate each sub-execution result jointly based on the associated target sub-execution result and the sub-task requirement conditions. In addition to using the sub-task requirement conditions to impose conditional restrictions on the sub-execution result, the sub-execution result is also evaluated by combining global information through the associated target sub-execution result, so that the target sub-execution result maintains the logical relevance with the associated target sub-execution result, thereby achieving global coherence among multiple target sub-tasks in the target video production process.
[0056] According to an embodiment of the present disclosure, a general agent may be used to detect the sub-execution result of each target sub-task, which is applicable to performing generalized detection tasks. However, it is not limited thereto. Multiple sub-expert agents may also be set for the main agent, and the sub-expert agents are applicable to performing personalized detection tasks.
[0057] The model structure configured in the general agent or sub-expert agent is not limited. For example, it may include a deep learning model with one or more of the attention mechanism, convolutional network, recurrent network, and long short-term memory network. However, it is not limited thereto, and it may also be a large language model with more powerful functions. As long as it is a model that can be used to perform the detection task. The difference is that by using multiple sub-expert agents, a mixture of experts mechanism can be formed, combining multiple targeted and personalized sub-expert agents, and for detection tasks with different sub-task requirement conditions, different sub-expert agents are selected to perform the detection task.
[0058] Optionally, a sub-expert agent corresponding to the sub-task requirement condition can be used to perform the detection task based on the sub-task requirement condition and the sub-execution result. However, it is not limited thereto. A sub-expert agent corresponding to the sub-task requirement condition can also be used to perform the detection task based on the sub-task requirement condition, the associated target sub-execution result, and the sub-execution result.
[0059] According to an embodiment of the present disclosure, a gating selector can be set for multiple sub-expert agents to determine, for each detection task, the sub-expert agent corresponding to the sub-task requirement condition from multiple sub-expert agents.
[0060] Figure 3A Schematically shows an application diagram of detecting the sub-execution result according to an embodiment of the present disclosure.
[0061] As Figure 3A shown, the target sub-task may include a copywriting generation sub-task, and the associated sub-task may include a problem-solving sub-task. Based on the associated target sub-execution result 310 generated for the problem-solving sub-task, such as target problem-solving data, a sub-execution result 320 is generated, such as an explanatory text. The main agent A300 may include a gating selector 340 and multiple sub-expert agents 350, and based on the sub-task requirement condition 330 of the copywriting generation sub-task, the gating selector 340 can be used to determine the first sub-expert agent 351 from multiple sub-expert agents 350. Control the first sub-expert agent 351 to detect the generated sub-execution result 320, such as the explanatory text, based on the associated target sub-execution result 310 and the sub-task requirement condition 330, to obtain a detection result 360.
[0062] Using the detection means provided by the embodiments of the present disclosure, different sub-expert agents can be used to perform the detection task for different sub-task requirement conditions, meeting the personalized needs of detection and improving the detection accuracy.
[0063] It should be noted that the number of target agents or main agents involved in the embodiments of the present disclosure can be one or more. For example, the target agent can also be a combination of multiple sub-expert agents with a Mixture of Experts (MOE) structure. Different sub-expert agents can perform the methods provided in the embodiments of the present disclosure through information interaction. By using multiple sub-expert agents, a Mixture of Experts mechanism can be formed, combining multiple targeted and personalized sub-expert agents, and selecting different sub-expert agents to perform the detection task for detection tasks with different sub-task requirement conditions.
[0064] The above has elaborated in detail on the associated target sub-execution results as associated data for performing the detection task and the sub-expert agents as detection tools. The following will describe the sub-task requirement conditions for performing the detection task.
[0065] According to the embodiments of the present disclosure, for different target sub-tasks, the sub-task requirement conditions are different. The object requirement information of the target object can be intention-recognized to obtain an intention recognition result. Based on the intention recognition result, the sub-task requirement conditions for at least one target sub-task are determined.
[0066] The target object can refer to a user for human-computer interaction. The object requirement information can include object input information, such as a query. However, it is not limited thereto. It can also include object attribute information while including object input information. For example, information such as the age, education level, and learning interests of the target object.
[0067] The object requirement information can be used for intention recognition to obtain an intention recognition result. Based on the intention recognition result, rule-based matching screening is performed among multiple sub-task requirement conditions to obtain the sub-task requirement conditions. However, it is not limited thereto. The sub-task requirement conditions can also be generated based on the intention recognition result in an Artificial Intelligence Generated Content (AIGC) manner.
[0068] The content obtained by the above methods can be used as the sub-task requirement conditions. It can also be used as personalized sub-task requirement conditions, and based on the combination of the personalized sub-task requirement conditions and the basic sub-task requirement conditions, the sub-task requirement conditions are generated.
[0069] By using the method for determining the sub-task requirement conditions provided in the embodiments of the present disclosure, it is possible to improve the generated target sub-execution results to meet personalized and basic requirements, and when applied to the teaching field, improve the accuracy of the content of the target video, the clarity of logic, the coherence of expression, and the pertinence of knowledge.
[0070] Figures 3B to 3F A schematic diagram showing a target video according to an embodiment of the present disclosure is schematically illustrated.
[0071] Combined with FIGS. 3B to Figure 3F As shown, the target video can be a problem-solving explanation video. In the problem-solving explanation video, the virtual character can be driven by target speech data corresponding to the target copywriting to explain knowledge related to taking medicine. The problem-solving explanation video can include Figure 3B the 1st video frame image T310 to the 5th video frame image T350 shown in FIG. The problem-solving explanation video can be based on the virtual character to drive the target speech data corresponding to the target copywriting to explain the problem, the analysis process, and the problem-solving answer. The analysis process can include the examination points and the key steps of problem-solving. The 1st video frame image T310 to the 5th video frame image T350 can be adapted to the progress information represented by the video progress bar at the bottom of the video frame.
[0072] As Figure 3B shown, in the 1st video frame image T310, the problem and the analysis process can be displayed in the video screen. The key information of the problem in the problem, such as "200", "4 red", "2 white", and "3 blue", can be obtained by annotating the key copywriting information in the target copywriting generated by the copywriting generation intelligent agent.
[0073] As Figure 3C shown, in the 2nd video frame image T320, the graphic element T301 can be displayed, and the progress information T302 represented by the video progress bar can indicate that the progress of the current problem-solving explanation video is in the "step writing" stage. Among them, the graphic element T301 can be a plurality of circular elements generated based on the target graphic code generated by the execution code generation intelligent agent. The plurality of circular elements meet the element requirement conditions for the key information of the problem, such as "4 red", "2 white", and "3 blue", so that the plurality of circular elements in the graphic element T301 can represent the semantic information of the key information of the problem and are arranged in a relatively neat array.
[0074] Combined with Figure 3D and Figure 3E shown, in the 3rd video frame image T330 and the 4th video frame image T340, the virtual character can be driven to explain the problem based on the explanation semantic process of the explanation copywriting, and the problem-solving steps for the problem can be explained in detail in the 4th video frame image T340. The problem-solving steps can be the target problem-solving data obtained by the problem-solving intelligent agent executing the problem-solving subtasks for the problem.
[0075] As Figure 3FAs shown, on the video progress bar of the problem-solving video of the 5th video frame image T350, it can indicate entering the summary stage of this problem. In the 5th video frame image T350, the problem content, analysis content, solution content, and summary content of this problem can be displayed based on the problem-solving copywriting, so that users can summarize and generalize knowledge points in a timely manner. At the same time, a problem-solving interaction instruction window T303 can also be displayed in the 5th video frame image T350, so that users can perform interaction operations such as scanning codes and clicking on the problem-solving interaction instruction window T303 to interact with the interactive intelligent agent, realizing the use of the interactive intelligent agent to explain knowledge or answer questions for users, and improving the firmness of users' knowledge learning.
[0076] It should be noted that Figures 3B to 3F The black box shown is to protect relevant privacy information or important information, and it is not a display defect of the target video.
[0077] For the acquisition and processing process of the information involved in the embodiments of the present disclosure, relevant users or institutions have been authorized, and necessary encryption measures have been provided to avoid information leakage, which complies with relevant laws and regulations and does not violate public order and good customs.
[0078] The detection tasks performed by the main intelligent agent were introduced above. Below, a brief introduction will be given on how to determine the target intelligent agent for performing the target subtask.
[0079] According to the embodiments of the present disclosure, before performing the operation S210 as Figure 2 shown, the video generation method based on multi-agent collaboration may further include: determining a target intelligent agent from multiple candidate intelligent agents. Using the target intelligent agent to perform the target subtask to obtain a sub-execution result.
[0080] Optionally, a target intelligent agent that matches the task requirement elements in the intention recognition can be determined from multiple candidate intelligent agents.
[0081] The object requirement information can be subjected to intention recognition to obtain an intention recognition result. Based on the intention recognition result, the target task is determined. The multiple task requirement elements that make up the target task are determined.
[0082] The task requirement elements may include the basic units or components that make up the target task. For example, the task requirement element can be a target subtask. However, it is not limited to this. It can also be that multiple task requirement elements form a target subtask. As long as it is a task requirement element for performing the target task.
[0083] A mapping relationship between the task requirement elements and the intelligent agent can be pre-constructed. Based on the task requirement elements and the mapping relationship, a target intelligent agent is determined from multiple candidate intelligent agents.
[0084] The target task is split into fine-grained tasks according to the task requirement elements, facilitating the combination of the task requirement elements according to the intention of the target object. Then, multiple target agents for executing the target subtasks are combined to jointly generate the target video, improving the flexibility and diversity of the target video generation.
[0085] Next, the target video will be taken as an example of the topic-explaining video for detailed description.
[0086] The embodiment of the present disclosure provides a method for generating a topic-explaining video based on multi-agent collaboration. Multiple target agents and a main agent can be used to generate a complete problem-solving process, personalized explanation text, and accurate graphic animation based on the problem input by the target object, and then automatically generate the content of the topic-explaining video. It realizes the automated, batch, and interactive generation of topic-explaining videos, improving the intelligent level of online education. This method can be widely applied to application scenarios such as intelligent tutoring, competition training, and autonomous learning.
[0087] Figure 4 Schematically shows a scene diagram of the topic-explaining video according to the embodiment of the present disclosure.
[0088] As Figure 4 shown, taking the topic-explaining video as an example of the target video, the main agent 410 obtains the intention recognition result according to the object requirement information input by the target object, such as a problem. Based on the intention recognition result, it determines the task of generating the topic-explaining video as the target task. The target task can be split into a problem-solving subtask, a text generation subtask, a code generation subtask, a video generation subtask, etc. The target agents can include a problem-solving agent 420 for executing the problem-solving subtask, a text generation agent 430 for executing the text generation subtask, a code generation agent 440 for executing the code generation subtask, and the main agent 410 for executing the video generation subtask.
[0089] The problem-solving agent 420 executes the problem-solving subtask to obtain the target problem-solving data 421. The text generation agent 430 determines the target explanation text 431 based on the target problem-solving data by executing the text generation subtask. The code generation agent 440 determines the target graphic data code 441 based on the target explanation text by executing the code generation subtask. The main agent 410 determines the target video 450 by executing the video generation subtask based on at least one of the target graphic data code, the target problem-solving data, and the target explanation text.
[0090] It should be noted that after each target agent executes the target subtask, the master agent needs to be used to detect the sub - execution result. When the sub - task detection result indicates that the sub - execution result meets the sub - task requirement conditions, the sub - execution result is determined as the target sub - execution result. When the sub - task detection result indicates that at least one sub - execution result does not meet the sub - task requirement conditions, the master agent can control multiple target agents to re - execute their respective target subtasks based on the defect type characterized by the sub - task detection result, so as to uniformly correct their respective sub - execution results and make the generated multiple target sub - execution results maintain consistency and logical coherence.
[0091] As Figure 4 shown, the master agent 410 detects the graphical data code output by the code - generation agent 440 to obtain a sub - task detection result for the code - generation subtask. When the sub - task detection result for the code - generation subtask indicates that the sub - task requirement conditions are not met, the master agent 410 can re - control the problem - solving agent 420 and the copywriting - generation agent 430 to re - execute the problem - solving subtask and the copywriting - generation subtask based on the defect type indicated by the sub - task detection result for the code - generation subtask, so as to obtain new target problem - solving data and target copywriting. The master agent 410 can detect the new target problem - solving data and target copywriting based on the sub - task requirement conditions corresponding to each target agent, so that the new target problem - solving data, target copywriting, and target graphical data code meet all the sub - task requirement conditions of the target task, maintain the logical coherence and consistency of multiple target execution results, improve the ability of the target video 450 (as a problem - solving video) to explain the problem logically and clearly, and improve the user's understanding of the problem - explanation information.
[0092] Next, taking the problem - solving video as an example of the target video, the execution of each target subtask will be described in detail.
[0093] According to an embodiment of the present disclosure, the target agent may include a problem - solving agent for executing a problem - solving subtask for a specified problem. The sub - execution result may include problem - solving data.
[0094] The master agent can extract information from the object - demand information input by the target object to obtain the problem specified by the target object. The master agent transmits the problem to the problem - solving agent. The problem - solving agent is controlled to execute the problem - solving subtask to obtain problem - solving data. The master agent receives the problem - solving data sent by the problem - solving agent and detects the problem - solving data based on the sub - task requirement conditions matching the problem - solving subtask to obtain a sub - task detection result related to the problem - solving data.
[0095] The sub - task requirement conditions may include at least one of the following: accuracy condition, required - knowledge - point condition, and required - quantity condition.
[0096] Optionally, the accuracy condition may characterize the accuracy condition related to the result, such as the answer, the problem-solving formula, etc. The required knowledge point condition may characterize the knowledge matching condition. For problem A, there are multiple problem-solving methods, and different problem-solving methods use different knowledge points. The required knowledge point condition can be determined according to the object demand information of the target object, such as the education level of the sixth grade of primary school, including the knowledge point condition of the linear equation with one unknown. The required quantity condition may include the quantity condition of the problem-solving method.
[0097] Correspondingly, the sub-task detection result related to the problem-solving data characterizes at least one of the following problem-solving data defects: The answer in the problem-solving data does not meet the accuracy condition. The knowledge points involved in the problem-solving data do not meet the required knowledge point condition. The number of solution methods of the problem-solving data does not meet the required quantity condition.
[0098] Detecting the problem-solving data using different types of multi-sub-task requirement conditions can improve the generation quality of the target problem-solving data that meets the sub-task requirement conditions, meet the needs of the target object from different dimensions, and improve the accuracy, richness of the target problem-solving data and its matching with the needs of the target object.
[0099] Optionally, taking the solution sub-task as the target sub-task as an example, for the operation S230 as shown in Figure 2 it may include: in the case where there is at least one of the above-mentioned problem-solving data defects in the problem-solving data, controlling the problem-solving agent to execute the solution sub-task based on the problem and the sub-task detection result related to the problem-solving data to obtain updated problem-solving data. Until the problem-solving data meets the sub-task requirement conditions. Taking the updated problem-solving data as the target problem-solving data.
[0100] It is possible to control the problem-solving agent to execute the solution sub-task based on the problem to obtain updated problem-solving data. Compared with this method, using the problem and the sub-task detection result to execute the solution sub-task can generate prompt information based on the sub-task detection result, improve the guiding ability of executing the solution sub-task towards the expected result, and thus reduce the number of repetitions of re-executing the solution sub-task and improve the generation effect of the problem-solving data.
[0101] The solution sub-task is described in detail above, and the copywriting generation sub-task that has a dependency relationship with the solution sub-task will be described below.
[0102] According to an embodiment of the present disclosure, the target sub-task may include a copywriting generation sub-task, the sub-execution result may include an explanation copywriting determined by the copywriting generation agent based on the target problem-solving data, and the associated target sub-execution result includes the target problem-solving data.
[0103] Based on the sub-task requirement conditions and the associated target sub-execution results of the associated sub-tasks, the detection of the sub-execution results may include: based on the target problem-solving data and the copywriting requirement conditions for the copywriting generation sub-task, detecting the explanatory copywriting to obtain a copywriting detection result.
[0104] Taking a math problem as an example, the target problem-solving data may include a problem-solving formula and an answer. The explanatory copywriting may refer to personalized problem-solving content that is clear and easy to understand. For example, adding illustrative examples, transitional terms between multiple steps, and extended knowledge, etc.
[0105] The explanatory copywriting is generated based on the target problem-solving data and is used to assist in explaining the target problem-solving data. Using the target problem-solving data and the copywriting requirement conditions to detect the explanatory copywriting can improve the detection scope and requirements of the copywriting detection result, and thus improve the generation quality of the target explanatory copywriting that meets the copywriting requirement conditions.
[0106] According to an embodiment of the present disclosure, the copywriting requirement conditions may include the matching conditions between the target problem-solving data and the explanatory copywriting. For example, the copywriting requirement conditions may include at least one of the following: the semantic similarity condition between the text segments of the explanatory copywriting and the problem-solving steps in the target problem-solving data; the solving logic relationship between multiple text segments in the explanatory copywriting and whether it is the same as the target solving logic relationship represented by the target problem-solving data; the adaptation condition between the problem-solving result in the explanatory copywriting and the target problem-solving result in the target problem-solving data.
[0107] Correspondingly, the copywriting detection result may characterize the copywriting defects that the explanatory copywriting does not match the target problem-solving data. Specifically, the copywriting detection result may characterize at least one of the following copywriting defects: the semantic similarity between the text segments of the explanatory copywriting and the problem-solving steps in the target problem-solving data does not meet the semantic similarity condition; the solving logic relationship between multiple text segments in the explanatory copywriting is different from the target solving logic relationship represented by the target problem-solving data; the problem-solving result in the explanatory copywriting is different from the target problem-solving result in the target problem-solving data.
[0108] By detecting the copywriting defects, the target explanatory copywriting can be adapted from multiple aspects such as semantic matching, logical matching, and accuracy with the target problem-solving data, improving the interestingness of the target explanatory copywriting while improving the logicality and accuracy of the problem-solving video in the teaching field.
[0109] Optionally, in the case where the sub-task detection result characterizes that the explanatory copywriting has copywriting defects, the copywriting generation agent may be controlled based on the copywriting detection result to execute the copywriting generation sub-task again to obtain the target explanatory copywriting that meets the copywriting requirement conditions.
[0110] Specifically, determine the copywriting defect prompt information based on the copywriting defects that can be characterized by the copywriting detection results. Control the copywriting generation agent to generate updated explanatory copy based on the copywriting defect prompt information.
[0111] The large model configured by the copywriting generation agent can be used to generate explanatory copy. The copywriting defects can be filled into the copywriting generation prompt template to obtain the copywriting defect prompt information. However, it is not limited to this. The target problem-solving data, copywriting defects, and copywriting requirement conditions can also be filled into the copywriting generation prompt template to obtain the copywriting defect prompt information.
[0112] Input the copywriting defect prompt information into the large model to generate updated explanatory copy.
[0113] Using the copywriting defect prompt information with added copywriting defects to generate updated explanatory copy can use the copywriting defects as additional knowledge to improve the ability of the copywriting generation agent to generate explanatory copy, avoid the problem of generating explanatory copy with copywriting defects again during the copywriting generation task execution, and improve the generation efficiency and accuracy of the target explanatory copy.
[0114] The generation of the target explanatory copy was described above. The next related target subtask - the code generation subtask will be described below.
[0115] According to an embodiment of the present disclosure, the target subtask may include a code generation subtask, and the sub-execution result may include the graphic data code determined by the code generation agent based on the target explanatory copy for executing the code generation subtask, and the target explanatory copy meets the copywriting requirement conditions.
[0116] Taking the code generation subtask as an example, the detection of the sub-execution result may include: detecting the graphic data code based on at least one of the target explanatory copy, the target problem-solving data, and the graphic code conditions for the code generation subtask to obtain the code detection result.
[0117] At least one of the target explanatory copy and the target problem-solving data is an associated target sub-execution result, and the associated target sub-execution result is determined by controlling the target intelligent agent to execute an associated subtask having an execution dependency relationship with the target subtask.
[0118] Optionally, the graphic data code can be detected using the associated target sub-execution result and the graphic code conditions to obtain the code detection result.
[0119] The target explanatory copy and the target problem-solving data meet the copywriting requirement conditions. Therefore, using the target explanatory copy or the target problem-solving data as the reference data for detection can achieve the corresponding detection effect.
[0120] Exemplarily, the graphic data code can be detected using the target explanation text and graphic code conditions to obtain a code detection result. Compared with the detection method using the target problem-solving data as the reference data, the detection method using the target explanation text as the reference data can utilize more abundant information in the target explanation text to improve the adaptability between the target graphic data generated using the graphic code conditions and the target explanation solution.
[0121] According to an embodiment of the present disclosure, the code detection result characterizes at least one of the following code defects: code format defect; occlusion defect of the image elements generated by executing the graphic data code; the color of the image elements generated by executing the graphic data code does not meet the color condition; the semantic relationship condition between the image elements generated by executing the graphic data code and the problem-solving steps in the target problem-solving data is not satisfied; the arrangement of multiple image elements generated by executing the graphic data code does not meet the arrangement condition.
[0122] Optionally, the graphic code conditions may include code format conditions and graphic data conditions generated based on the code. Specifically, the graphic data conditions generated based on the code may include at least one of the following: occlusion condition of the image elements generated by executing the graphic data code; color condition of the image elements generated by executing the graphic data code; semantic relationship condition between the image elements generated by executing the graphic data code and the problem-solving steps in the target problem-solving data; arrangement condition of the arrangement of multiple image elements generated by executing the graphic data code.
[0123] By detecting the above code defects, the image clarity of the target graphic data, the semantic relevance with the target problem-solving data, and the logical coherence between multiple target image data can be improved, thereby improving the logicality and accuracy of the problem-solving video in the teaching field.
[0124] According to an embodiment of the present disclosure, taking the update of the graphic data code as an example, controlling the target intelligent agent to execute the target sub-task based on the sub-task detection result, and obtaining the target sub-execution result that meets the sub-task requirement conditions may include: determining code defect prompt information based on the code defects characterized by the code detection result. Controlling the code generation intelligent agent to generate an updated graphic data code based on the code defect prompt information.
[0125] The graphic data code can be generated using the large model configured by the code generation intelligent agent. The code defects can be filled into the code generation prompt template to obtain the code defect prompt information. However, it is not limited to this. The target problem-solving data and / or the target explanation text, code defects, and code requirement conditions can also be filled into the code generation prompt template to obtain the code defect prompt information.
[0126] Input the code defect prompt information into the large model to generate an updated problem-solving data code.
[0127] Using the code defect hint information with added code defects to generate updated diagram data code can use code defects as additional knowledge to improve the ability of the code generation agent to generate diagram data code, avoid the problem of generating code diagram data with code defects again during the execution of the code generation subtask, and improve the generation efficiency and accuracy of the target diagram data code.
[0128] Figure 5 Schematically shows an application schematic diagram of generating target diagram data code according to an embodiment of the present disclosure.
[0129] As Figure 5 shown, the code generation agent 510 generates the diagram data code 530 based on at least one of the target explanation text 521 and the target problem-solving data 522 as the associated target sub-execution result.
[0130] As Figure 5 shown, the main agent 540 detects the diagram data code 530 based on at least one of the target explanation text 521, the target problem-solving data 522, and the diagram code condition 523 for the code generation subtask, and obtains the code detection result 550.
[0131] As Figure 5 shown, the code generation agent generates updated diagram data code based on at least one of the target explanation text 521 and the target problem-solving data 522 as the associated target sub-execution result and the code detection result 550.
[0132] Taking the updated diagram data code as the diagram data code, repeat the above operations until the updated diagram data code meets the diagram code condition, and take the updated diagram data code as the target diagram data code.
[0133] It should be noted that Figure 5 only the generation of the target diagram data code is taken as an example for illustration. However, it is not limited to the execution of the code generation subtask, and any target subtask can be executed according to the process as Figure 5 shown until the target sub-execution result is obtained.
[0134] The above has described the execution of the target subtask by each target agent. The following will describe the execution of the video generation subtask by the main agent.
[0135] According to an embodiment of the present disclosure, generating a target video based on at least one target sub-execution result may include: generating target graphic data based on target graphic data codes that meet the graphic code conditions. Generating target voice data based on target explanation texts that meet the copywriting requirement conditions. Driving a preset virtual character to perform target actions based on the target voice data, and displaying the target graphic data according to the display timing data to obtain the target video.
[0136] The target explanation text includes display timing data indicating the display timing of the target graphic data in the target video.
[0137] The target graphic data may include multiple graphic illustrations. The content displayed by each graphic illustration respectively matches the corresponding target explanation text, and the matching can be performed according to the display timing data.
[0138] Figure 6A A schematic diagram of a graphic illustration according to an embodiment of the present disclosure is schematically shown.
[0139] As Figure 6A shown, the graphic illustration 611 may include at least one of a question 6111, an analysis 6112, an explanation, an answer, etc. and display timing data 6113.
[0140] Optionally, the multiple graphic illustrations are arranged in sequence based on the display timing data as the target graphic data.
[0141] Optionally, the display timing data can be represented by a progress bar, but is not limited thereto. It can also be represented by time or by text description. As long as it can be adapted to the target voice data generated based on the target explanation text.
[0142] Figure 6B A schematic diagram of the scenario for generating the target video according to an embodiment of the present disclosure is schematically shown.
[0143] As Figure 6B shown, the target voice data 630 can be obtained by means of text-to-speech using the target explanation text 620. The preset virtual character 640 can also be driven to perform target actions using the target voice data. The preset virtual character carrying action information and the target graphic data 610 are combined according to the display timing data to obtain multiple video frames. The target voice data and the multiple video frames including the dynamic virtual character are combined to generate the target video 650.
[0144] Optionally, it can be generated using a large model in the main intelligent agent, or can be generated using voice generation code. The semantic generation code and the generation method of the target graphic data code can be similar.
[0145] Optionally, the preset virtual image may include a digital human, but is not limited thereto, and may also include other animated images, as long as the image can reflect the dynamic effect.
[0146] Generating a target video using the video generation method provided by the embodiments of the present disclosure can make the target video coherent and logically smooth, and at the same time, combine dynamic virtual images and voice data to improve the interest and authenticity of the target video.
[0147] According to an embodiment of the present disclosure, after performing the operation S240 as Figure 2 shown, the video generation method based on multi-agent collaboration may further include the following interaction operations.
[0148] For example, controlling the target video to be displayed on the interaction interface. Controlling the interaction agent to determine a reply message based on the input information of the target object for the target video, and controlling the reply message to be displayed on the interaction interface.
[0149] The interaction agent may be an agent for performing human-computer interaction with the target object. The main agent may be communicatively connected to the interaction agent, and control the interaction agent to determine a reply message based on the input information of the target object through a control instruction, and display the reply message on the interaction interface.
[0150] The information type of the reply message may include voice, text, or image.
[0151] Optionally, the interaction agent may be used to still interact with the target object during the process of the target object watching the target video, such as obtaining a detailed explanation of a certain problem-solving step, modifying the problem-solving idea, etc. as the reply message.
[0152] After the target video is generated, the reply message output in real-time interaction can be combined with the target video, improving the ability of dynamic interaction and personalized adjustment, thereby making the teaching interaction vivid and flexible.
[0153] Figure 7 Schematically shows an interaction scenario diagram according to an embodiment of the present disclosure.
[0154] As Figure 7 shown, the target object 710 may ask a question. The main agent 750 controls multiple target agents to collaborate to generate a target video 720. After the target object 710 watches the target video 720, the target object 710 may input input information through the input box of the interaction interface 730. The main agent responds to the input information of the target object 710 for the target video 720, controls the interaction agent 740 to generate a reply message based on the input information, and controls the reply message to be displayed on the interaction interface 730.
[0155] According to an embodiment of the present disclosure, the reply information may be obtained by directly performing intelligent generation processing based on the input information input by the target object. However, it is not limited thereto. It may also be to process the input information by controlling the interaction agent based on the target sub-execution result to obtain the reply information.
[0156] The target sub-execution result is not limited, as long as it is one of the sub-execution results used to generate the target video. It may also be possible to obtain the object intention recognition result of the target object based on the input information. Based on the object intention recognition result, the target sub-execution result is determined from multiple sub-execution results used to generate the target video. The target sub-execution result is used as reference knowledge to output the reply information for the input information.
[0157] Taking the input information as "Please explain the meaning of the 3rd formula". Based on the object intention recognition result of "explaining the problem-solving data", it can be determined that the target sub-execution result may include the target problem-solving data. Based on the target problem-solving data and the input information, the reply information of "specific explanation of the meaning of the 3rd formula" can be obtained.
[0158] Outputting the reply information in combination with the target sub-execution result can improve the accuracy and effectiveness of the reply information.
[0159] Figure 8 The block diagram of a video generation device based on multi-agent collaboration according to an embodiment of the present disclosure is schematically shown.
[0160] As Figure 8 shown, the video generation device 800 based on multi-agent collaboration includes an acquisition module 810, a detection module 820, a task execution module 830, and a video generation module 840.
[0161] The acquisition module 810 is configured to acquire the sub-execution result determined by at least one target agent executing a target subtask. The target task for generating the target video includes the target subtask.
[0162] The detection module 820 is configured to detect the sub-execution result to obtain the subtask detection result.
[0163] The task execution module 830 is configured to, when the subtask detection result indicates that the sub-execution result does not meet the subtask requirement condition for the target subtask, control the target agent to execute the target subtask based on the subtask detection result to obtain the target sub-execution result that meets the subtask requirement condition.
[0164] The video generation module 840 is configured to generate the target video based on the target sub-execution result.
[0165] According to an embodiment of the present disclosure, the detection module includes: a target detection sub-module.
[0166] A target detection sub-module, configured to detect the sub-execution result based on the sub-task requirement conditions and the associated target sub-execution result of the associated sub-task in the target task. There is an execution dependency relationship between the associated sub-task and the target sub-task, and the associated target sub-execution result meets the associated sub-task requirement conditions for the associated sub-task.
[0167] According to an embodiment of the present disclosure, the target agent includes a problem-solving agent for performing a problem-solving sub-task on a specified problem, and the sub-execution result includes problem-solving data.
[0168] According to an embodiment of the present disclosure, the sub-task detection result related to the problem-solving data characterizes at least one of the following problem-solving data defects: the answer in the problem-solving data does not meet the accuracy condition; the knowledge points involved in the problem-solving data do not meet the required knowledge point conditions; the number of solution methods of the problem-solving data does not meet the required quantity conditions.
[0169] According to an embodiment of the present disclosure, the target sub-task includes a copywriting generation sub-task, the sub-execution result includes an explanatory copy determined by the copywriting generation agent based on the target problem-solving data for performing the copywriting generation sub-task, and the associated target sub-execution result includes the target problem-solving data.
[0170] According to an embodiment of the present disclosure, the target detection sub-module includes: a copywriting detection unit.
[0171] The copywriting detection unit is configured to detect the explanatory copy based on the target problem-solving data and the copywriting requirement conditions for the copywriting generation sub-task, and obtain a copywriting detection result.
[0172] According to an embodiment of the present disclosure, the copywriting detection result characterizes at least one of the following copywriting defects: the semantic similarity between the text segment of the explanatory copy and the problem-solving steps in the target problem-solving data does not meet the semantic similarity condition; the solution logic relationship between multiple text segments in the explanatory copy is different from the target solution logic relationship characterized by the target problem-solving data; the problem-solving result in the explanatory copy is different from the target problem-solving result in the target problem-solving data.
[0173] According to an embodiment of the present disclosure, the task execution module includes: a copywriting prompt sub-module and a copywriting update sub-module.
[0174] The copywriting prompt sub-module is configured to determine copywriting defect prompt information based on the copywriting defects characterized by the copywriting detection result.
[0175] The copywriting update sub-module is configured to control the copywriting generation agent to generate an updated explanatory copy based on the copywriting defect prompt information.
[0176] According to an embodiment of the present disclosure, the target subtask includes a code generation subtask, the sub-execution result includes the graphical data code determined by the code generation agent based on the target explanation text for executing the code generation subtask, and the target explanation text meets the text requirement conditions.
[0177] According to an embodiment of the present disclosure, the detection module includes: a code detection sub-module.
[0178] The code detection sub-module is used to detect the graphical data code based on at least one of the target explanation text, the target problem-solving data, and the graphical code conditions for the code generation subtask, and obtain a code detection result.
[0179] According to an embodiment of the present disclosure, the code detection result characterizes at least one of the following code defects: code format defect; occlusion defect in the image elements generated by executing the graphical data code; the color of the image elements generated by executing the graphical data code does not meet the color conditions; the semantic relationship condition is not met between the image elements generated by executing the graphical data code and the problem-solving steps in the target problem-solving data; the arrangement of multiple image elements generated by executing the graphical data code does not meet the arrangement conditions.
[0180] According to an embodiment of the present disclosure, the task execution module includes: a code hint sub-module and a code update sub-module.
[0181] The code hint sub-module is used to determine code defect hint information based on the code defects characterized by the code detection result.
[0182] The code update sub-module is used to control the code generation agent to generate updated graphical data code based on the code defect hint information.
[0183] According to an embodiment of the present disclosure, the video generation module includes: a graph generation sub-module, a voice generation sub-module, and an action generation sub-module.
[0184] The graph generation sub-module is used to generate target graphical data based on the target graphical data code that meets the graphical code conditions. The target explanation text includes display timing data indicating the display timing of the target graphical data in the target video.
[0185] The voice generation sub-module is used to generate target voice data based on the target explanation text that meets the text requirement conditions.
[0186] The action generation sub-module is used to drive a preset virtual image to execute a target action based on the target voice data, and display the target graphical data according to the display timing data to obtain a target video.
[0187] According to an embodiment of the present disclosure, the main agent includes multiple sub-expert agents, and the sub-expert agents are suitable for executing detection tasks for sub-task requirement conditions.
[0188] According to an embodiment of the present disclosure, the detection module includes: a detection agent determination sub-module.
[0189] The agent determination sub-module is configured to use a sub-expert agent corresponding to the sub-task requirement conditions and perform a detection task based on the sub-task requirement conditions and sub-execution results.
[0190] According to an embodiment of the present disclosure, the video generation device based on multi-agent collaboration further includes: an intention recognition module and a requirement condition determination module.
[0191] The intention recognition module is configured to perform intention recognition on the object requirement information of the target object to obtain an intention recognition result.
[0192] The requirement condition determination module is configured to determine sub-task requirement conditions for at least one target sub-task based on the intention recognition result.
[0193] According to an embodiment of the present disclosure, the video generation device based on multi-agent collaboration further includes: a target agent determination module.
[0194] The target agent determination module is configured to determine a target agent that matches the task requirement elements in the intention recognition from multiple candidate agents.
[0195] According to an embodiment of the present disclosure, the video generation device based on multi-agent collaboration further includes: a display module and a reply module.
[0196] The display module is configured to control the target video to be displayed on the interaction interface.
[0197] The reply module is configured to control the interaction agent to determine a reply message based on the input information of the target object for the target video, and control the reply message to be displayed on the interaction interface.
[0198] According to an embodiment of the present disclosure, the reply module includes: a reply sub-module.
[0199] The reply sub-module is configured to control the interaction agent to process the input information based on the target sub-execution result to obtain a reply message.
[0200] Figure 9 Schematically shows a structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure.
[0201] In an embodiment of the present disclosure, as Figure 9 shown, the AI agent 900 may include an input module 910, a processing module 920, and an output module 930.
[0202] The input module 910 is configured to receive input information;
[0203] A processing module 920, configured to execute, based on input information received by an input module, output information obtained by a video generation method based on multi-agent collaboration provided in an embodiment of the present disclosure;
[0204] An output module 930, configured to output the output information obtained by the processing module.
[0205] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (such as a user or an external environment), and converting it into a format that the AI agent 900 can understand and process. The input module 910 is the primary link for the AI agent 900 to interact with the outside world, enabling the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0206] In an example, the input module 910 may input the object demand information or input information described above.
[0207] In an example, the processing module 920 is the core support for the AI agent 900 to handle complex tasks. The processing module 920 may execute the video generation method based on multi-agent collaboration described above.
[0208] In an example, the performance of the processing module 920 may be closely related to the large model on which the AI agent 900 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 920 may be designed to be highly configurable and extensible to cope with various different types of tasks and requirements in real scenarios.
[0209] In an example, after the AI agent 900 obtains the object demand information, the processing module 920 may obtain sub-execution results determined by at least one target agent to execute a target sub-task. The target tasks for generating a target video include the target sub-tasks; detect the sub-execution results to obtain sub-task detection results; in the case where the sub-task detection results indicate that the sub-execution results do not meet the sub-task requirement conditions for the target sub-task, control the target agent to execute the target sub-task based on the sub-task detection results to obtain target sub-execution results that meet the sub-task requirement conditions; and generate a target video based on the target sub-execution results. And transmit the target video to the output module 930.
[0210] It can be understood that although large language models have excellent language understanding and generation capabilities, like humans, the tasks they can solve without the help of any tools are very limited. When the AI agent 900 is given the ability to call tools, it can achieve tasks such as performing mathematical operations with the help of a calculator, performing data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine.
[0211] In the example, the output module 930 may output the target video described above.
[0212] The AI agent 900 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and enhance flexibility and versatility.
[0213] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0214] According to the embodiments of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0215] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0216] According to the embodiments of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0217] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0218] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0219] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as a keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as a disk, optical disc, etc.; and communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0220] Computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 executes the various methods and processes described above, such as the video generation method based on multi-agent collaboration. For example, in some embodiments, the video generation method based on multi-agent collaboration can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the video generation method based on multi-agent collaboration described above can be executed. Alternatively, in other embodiments, computing unit 1001 can be configured to execute the video generation method based on multi-agent collaboration by any other suitable means (e.g., by means of firmware).
[0221] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0222] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0223] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0224] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0225] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0226] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0227] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0228] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video generation method based on multi-agent collaboration, applied to a main agent, comprising: Obtaining a sub-execution result determined by at least one target agent executing a target subtask, wherein the target task for generating a target video includes the target subtask; Detecting the sub-execution result to obtain a sub-task detection result; In the case where the subtask detection result indicates that the sub-execution result does not satisfy the subtask requirement condition for the target subtask, controlling the target agent to execute the target subtask based on the subtask detection result to obtain a target sub-execution result that satisfies the subtask requirement condition; and The target video is generated based on the target sub-execution result.
2. The method according to claim 1, wherein: The detecting of the sub-execution result comprises: Based on the subtask requirement condition and the associated target sub-execution result of the associated subtask in the target task, the sub-execution result is detected, wherein the associated subtask has an execution dependency relationship with the target subtask, and the associated target sub-execution result satisfies the associated subtask requirement condition for the associated subtask.
3. The method according to claim 1 or 2, wherein: The target agent includes a problem-solving agent for executing a subtask of solving a specified problem, and the sub-execution result includes problem-solving data; The subtask detection result related to the problem-solving data indicates at least one of the following problem-solving data defects: The answer in the problem-solving data does not meet the accuracy condition; The knowledge points involved in the problem-solving data do not meet the required knowledge point conditions; The number of solution methods for the problem-solving data does not meet the required number conditions.
4. The method according to claim 2, wherein: The target subtask includes a text generation subtask, the sub-execution result includes the explanation text determined by the text generation agent executing the text generation subtask based on the target problem-solving data, and the associated target sub-execution result includes the target problem-solving data; The detecting of the sub-execution result based on the sub-task requirement condition and the associated target sub-execution result of the associated sub-task includes: Based on the target problem-solving data and the text requirement conditions for the text generation subtask, the explanation text is detected to obtain a text detection result.
5. The method according to claim 4, wherein: The text detection result indicates at least one of the following text defects: The semantic similarity between the text segment of the explanation copy and the problem-solving steps in the target problem-solving data does not meet the semantic similarity condition; The logical relationship between the multiple text segments in the explanation text is different from the target logical relationship represented by the target problem-solving data; The problem-solving result in the explanation text is different from the target problem-solving result in the target problem-solving data.
6. The method according to claim 4 or 5, wherein: The controlling the target agent to execute the target subtask based on the subtask detection result to obtain the target subtask execution result that meets the subtask requirement condition includes: Determining text defect prompt information based on the text defect represented by the text detection result; and The text generation agent is controlled to generate an updated explanation text based on the text defect prompt information.
7. The method according to claim 1, wherein: The target subtask includes a code generation subtask, and the sub-execution result includes the code generation agent executing the graphic data code determined by the code generation subtask based on the target explanation text, and the target explanation text meets the text requirement condition; The detecting of the sub-execution result includes: Based on at least one of the target explanation text, the target problem-solving data and the graphic code conditions for the code generation subtask, the graphic data code is detected to obtain a code detection result.
8. The method according to claim 7, wherein: The code detection result indicates at least one of the following code defects: Code formatting defects; The image elements generated by executing the graphic data code have occlusion defects; The color of the image element generated by executing the graphic data code does not meet the color condition; The image element generated by executing the graphic data code does not satisfy the semantic relationship condition with the problem-solving steps in the target problem-solving data; The arrangement of the plurality of image elements for executing the graphic data code generation does not satisfy an arrangement condition.
9. The method according to claim 7 or 8, wherein: The controlling the target agent to execute the target subtask based on the subtask detection result to obtain the target subtask execution result that meets the subtask requirement condition includes: Determining code defect prompt information based on the code defect represented by the code detection result; and The code generation agent is controlled to generate updated graphical data code based on the code defect prompt information.
10. The method according to claim 7 or 8, wherein: The generating the target video based on at least one of the target sub-execution results comprises: generating target graphic data based on the target graphic data code satisfying the graphic code condition, wherein the target explanation copy includes display timing data indicating a display timing of the target graphic data in the target video; Generate target voice data based on the target explanation text that meets the text requirement conditions; The target video is obtained by driving a preset virtual image to perform a target action based on the target voice data and displaying the target graphic data according to the display timing data.
11. The method according to claim 1, wherein: The main agent includes a plurality of sub-expert agents, and the sub-expert agents are suitable for performing detection tasks for the sub-task requirement conditions; The detecting of the sub-execution result includes: The detection task is performed based on the subtask requirement and the sub-execution result by using a sub-expert agent corresponding to the subtask requirement.
12. The method according to claim 1, further comprising: Performing intention recognition on the object demand information of the target object to obtain an intention recognition result; as well as A subtask requirement condition for at least one of the target subtasks is determined based on the intention recognition result.
13. The method according to claim 12, further comprising: From multiple candidate agents, determine the target agent that matches the task requirement elements in the intention recognition.
14. The method according to claim 1, further comprising: Controlling the target video to be displayed on an interactive interface; The interactive agent is controlled to determine reply information based on the input information of the target object for the target video, and control the reply information to be displayed on the interactive interface.
15. The method according to claim 14, wherein: The control interactive agent determines the reply information based on the input information of the target object for the target video, including: The interactive agent is controlled to process the input information based on the target sub-execution result to obtain the reply information.
16. A video generation device based on multi-agent collaboration, applied to a main agent, comprising: An acquisition module, used to acquire a sub-execution result determined by at least one target agent executing a target subtask, wherein the target task for generating a target video includes the target subtask; A detection module, used to detect the sub-execution result and obtain a sub-task detection result; a task execution module, configured to control the target agent to execute the target subtask based on the subtask detection result to obtain a target subtask execution result that satisfies the subtask requirement condition, if the subtask detection result indicates that the subtask execution result does not satisfy the subtask requirement condition for the target subtask; and A video generation module is used to generate the target video based on the target sub-execution result.
17. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to execute the method according to any one of claims 1 to 15 based on the input information received by the input module to obtain output information; An output module is used to output the output information obtained by the processing module.
18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.
20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Method, apparatus and computer program product for processing questions
CN117911730A
Mathematical physics explanation video generation method and device based on multi-agent cooperation
CN119277169A
Multi-agent cooperation-based explicit result acquisition class task processing method and related device
CN119558340A
Query answering method based on large model, electronic device, storage medium, and intelligent agent
US20250094460A1
Cited By
Intelligent agent capability description method, task execution method and device and intelligent agent
CN120688542A
Operator code generation method and system based on large model driving and multi-agent cooperation mechanism
CN121255173A
Operator code generation method and system based on large model driving and multi-agent cooperation mechanism
CN121255173B