Content generation method and apparatus based on artificial intelligence, intelligent agent system, and electronic device

The agent architecture for video production efficiently divides tasks among intelligent agents to enhance video generation quality and reduce human resource dependency, addressing limitations in existing AIGC video production methods.

JP2026015142AActive Publication Date: 2026-01-29BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024202153
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2024-11-20
Publication Date
2026-01-29
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

Current video production methods require significant human resources due to issues with controllability, time length, quality, and cost, limiting production efficiency, especially in the context of artificial intelligence-generated content (AIGC) where video generation quality needs improvement.

Method used

An agent architecture is utilized to efficiently produce video content by leveraging multi-type content resources on the Internet, employing intelligent agents to divide content generation tasks into multiple steps and invoking AIGC models, enabling high-quality content creation through a general and intelligent video production system.

Benefits of technology

This approach allows for the efficient generation of high-quality video content by dividing tasks, utilizing multiple intelligent agents to fulfill complex requests and improve content generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015142000001_ABST
    Figure 2026015142000001_ABST
Patent Text Reader

Abstract

The present disclosure provides a content generation method and apparatus based on artificial intelligence, an intelligent agent system, an electronic device, and a storage medium, and relates to the technical field of artificial intelligence.SOLUTION: A specific solution includes that a first intelligent agent sends a task execution request to a second intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a target task that the second intelligent agent needs to execute to generate the content, and the first intelligent agent receives a task execution result from the second intelligent agent, the task execution result including an execution result generated by the second intelligent agent executing the target task.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as computer vision, deep learning, large-scale models, and intelligent agents, and can be applied to scenarios such as content generation based on artificial intelligence (AIGC). [Background technology]

[0002] Artificial intelligence is a field that allows computers to simulate some human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware and software technologies. AI hardware technology generally includes sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, etc. AI software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and several directions such as machine learning / deep learning, big data processing technology, and knowledge spectrum technology.

[0003] With the development of computer and network technology, the application of deep learning models has become more widespread, and deep learning models have made breakthroughs in various fields. Here, AI generated content (AIGC) is an important direction in deep learning. Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure provides an artificial intelligence-based content generation method, apparatus, intelligent agent system, electronic device, storage medium, and program. [Means for solving the problem]

[0005] In one aspect of the present disclosure, there is provided an artificial intelligence based content generation method, the method comprising: The first intelligent agent sends a task execution request to a second intelligent agent based on task guidance information, where the task guidance information includes guidance information for generating content, and the task execution request includes a goal task that the second intelligent agent needs to perform to generate the content; The first intelligent agent receives a task execution result from the second intelligent agent, the task execution result including an execution result generated by the second intelligent agent executing the target task.

[0006] In another aspect of the present disclosure, there is provided an artificial intelligence based content generation method, the method comprising: A second intelligent agent receives a task execution request sent by the first intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate the content; the second intelligent agent executes the target task according to the task execution request to generate a task execution result; The second intelligent agent sends the task execution result to the first intelligent agent.

[0007] In another aspect of the present disclosure, there is provided an artificial intelligence based content generation apparatus for use with a first intelligent agent, the apparatus comprising: a first sending module for sending a task execution request to a second intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate the content; and a first receiving module for receiving a task execution result from the second intelligent agent, the task execution result including an execution result generated by the second intelligent agent executing the target task.

[0008] In another aspect of the present disclosure, there is provided an artificial intelligence based content generation apparatus for use with a second intelligent agent, the apparatus comprising: a third receiving module for receiving a task execution request sent by the first intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate content; a calling module for executing the target task based on the task execution request and generating a task execution result; and a second sending module for sending the task execution result to the first intelligent agent.

[0009] Another aspect of the present disclosure provides an intelligent agent system, the system comprising: a first intelligent agent; and a second intelligent agent; The first intelligent agent is used to send a task execution request to the second intelligent agent based on task guidance information and receive a task execution result from the second intelligent agent, wherein the task guidance information includes guidance information for generating content, the task execution request includes a target task that the second intelligent agent needs to execute to generate the content, and the task execution result includes an execution result generated by the second intelligent agent executing the target task; The second intelligent agent is used for receiving a task execution request sent by the first intelligent agent based on task guidance information, executing the target task based on the task execution request to generate a task execution result, and sending the task execution result to the first intelligent agent.

[0010] In another aspect of the present disclosure, there is provided an electronic device, the device comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the implementation of any one of the methods in the embodiments of the present disclosure.

[0011] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform any one of the methods in the embodiments of the present disclosure.

[0012] Another aspect of the present disclosure provides a program that, when executed by a processor, performs any of the methods in the embodiments of the present disclosure.

[0013] It should be understood that the contents described herein are not intended to describe key or important features of the embodiments of the present disclosure, nor are they intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will be better understood through the following specification.

[0014] The accompanying drawings are for better understanding of the solutions of the present disclosure and are not to be construed as limiting the present disclosure. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a flowchart of an artificial intelligence-based content generation method according to one embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram of an application scenario of an artificial intelligence-based content generation method according to an embodiment of the present disclosure; FIG. [Figure 3] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 4]1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 5] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 6] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 7] 1 is a flowchart of an artificial intelligence-based content generation method according to one embodiment of the present disclosure. [Figure 8] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 9] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 10] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 11] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 12] 1 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. [Figure 13] FIG. 1 is a system architecture diagram according to one embodiment of the present disclosure. [Figure 14] FIG. 1 is an agent architecture diagram according to one embodiment of the present disclosure. [Figure 15] FIG. 2 is a schematic diagram of a search matching enhancement module according to one embodiment of the present disclosure. [Figure 16] 1 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to an embodiment of the present disclosure. [Figure 17] FIG. 10 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to another embodiment of the present disclosure. [Figure 18] 1 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to an embodiment of the present disclosure. [Figure 19]FIG. 10 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to another embodiment of the present disclosure. [Figure 20] 1 is a schematic diagram illustrating the configuration of an intelligent agent system according to an embodiment of the present disclosure. [Figure 21] FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0016]

[0023] Exemplary embodiments of the present disclosure will now be described with reference to the accompanying drawings. These drawings include various details of the embodiments of the present disclosure to facilitate understanding, and should be considered as illustrative only. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known features and structures are omitted from the following description for clarity and conciseness.

[0017] Compared with text content, the visuals and audio in video content offer users a more intuitive and immersive information acquisition experience, resulting in better communication. In recent years, with the rapid rise of short video content, demand for video content has increased dramatically. There are various methods for creating content such as videos. For example, using video editing tools to edit original videos relies on the user's familiarity with the tools, resulting in low video productivity and high learning costs. AIGC video generation tools can automatically generate videos, but the quality of the generated videos needs to be improved. AIGC can involve the automatic creation of various forms of content, including text, images, audio, and video, using artificial intelligence technology. AIGC has already begun to have a profound impact on various industries, with many applications centered on text and conversational content, and image generation applications are also relatively mature. However, due to issues with controllability, time length, quality, and cost, current video production still requires a large number of human resources, severely limiting production efficiency.

[0018] To efficiently produce video content that meets user requirements, the embodiments of the present disclosure are based on an agent architecture, fully utilize the multi-type content resources on the Internet, and realize a general and intelligent video production agent system through deep mining and automatic invocation of the potential of the AIGC model. Agents are also called intelligent agents or intelligent agents.

[0019] 1 is a flowchart of an artificial intelligence-based content generation method according to an embodiment of the present disclosure.

[0020] In S101, a first intelligent agent sends a task execution request to a second intelligent agent based on task guidance information, where the task guidance information includes guidance information for generating content, and the task execution request includes a target task to be performed by the second intelligent agent to generate content.

[0021] In S102, the first intelligent agent receives a task execution result from the second intelligent agent, where the task execution result includes an execution result generated by the second intelligent agent when the second intelligent agent executes the target task.

[0022] In the embodiment of the present disclosure, the first intelligent agent can obtain task guidance information first; The task guidance information is information obtained based on the description text entered by the user, and the task guidance information includes guidance information for generating content. For example, it includes information on tasks that need to be performed and tools that need to be invoked to generate the content. In the embodiments of the present disclosure, the content can include one or more combinations of content, such as text, video, and audio. The first intelligent agent can obtain the task guidance information from another intelligent agent or generate the task guidance information itself. The first intelligent agent can generate a task execution request based on the task guidance information and send the task execution request to the second intelligent agent. The task execution request can include one or more goal tasks to be performed by the second intelligent agent. Depending on the characteristics of the goal task, the task execution request can also include one or more goal tools that need to be invoked to perform the one or more goal tasks. For example, if a complete content generation task needs to be decomposed into multiple goal tasks to be performed, the first intelligent agent can send one task execution request to the second intelligent agent each time, and the task execution request can include one goal task and one or more goal tools that need to be invoked to perform the goal task. Then, based on the task execution results returned by the second intelligent agent, it continues to send another task execution request to the second intelligent agent, and so on until the first intelligent agent no longer generates new task execution requests or the second intelligent agent has executed all the target tasks.

[0023] In addition, in the embodiment of the present disclosure, a maximum number of conversations can be set, and the period from sending a task execution request from the first intelligent agent to the second intelligent agent until the first intelligent agent receives the task execution result can be considered as one conversation. When the number of conversations exceeds the maximum number of conversations, the content generation process can be terminated even if the first intelligent agent has not generated all the target tasks required for the task guidance information.

[0024] In an application scenario, as shown in FIG. 2 , the intelligent agents may include a system agent, a user agent, an auxiliary agent, etc., but the present disclosure is not limited thereto. The system agent can generate task guidance information based on input text information and a system prompt template. The system agent can also send the generated task guidance information to the user agent and the auxiliary agent. Assume that the first intelligent agent is a user agent and the second intelligent agent is an auxiliary agent. When the user agent receives the task guidance information from the system agent, it generates a task execution request based on the task guidance information. When the user agent sends the task execution request to the auxiliary agent, the auxiliary agent can invoke a target tool to execute a target task based on the task guidance information and the task execution request to obtain a task execution result. When the auxiliary agent returns the task execution result to the user agent, the user agent can generate a new task execution request based on the task execution result, previous task guidance information, task execution requests, etc., and send it to the auxiliary agent again for processing. This process continues until the user agent no longer generates a new task execution request or until the auxiliary agent has executed all the target tasks.

[0025] According to the embodiments of the present disclosure, the first intelligent agent generates a task execution request based on the task guidance information and sends the task execution request to the second intelligent agent, thereby guiding the second intelligent agent to call the target tool to perform the target task, which helps to fully utilize the tools related to content generation to perform more tasks, fulfill more complex content generation requests, and generate high-quality content.

[0026] 3 is a flowchart of a content generation method based on artificial intelligence according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, in S101, the first intelligent agent sends a task execution request to the second intelligent agent based on the task guidance information.

[0027] In S301, the first intelligent agent adds a first system message generated based on the task guidance information to a first message pool, where the first message pool further includes a first local message.

[0028] In S302, the first intelligent agent generates a first task message based on a message in a first message pool.

[0029] In S303, the first intelligent agent sends a task execution request to the second intelligent agent based on the first task message, and the task execution request includes a target task that the second intelligent agent needs to execute to generate content.

[0030] In an embodiment of the present disclosure, the first intelligent agent can generate a first system message based on task guidance information and the first intelligent agent's prompt template. The first intelligent agent can add the first system message to its first message pool. The first message pool can provide the first intelligent agent with message storage and message retrieval functions.

[0031] In an embodiment of the present disclosure, a first message pool may be preset with an initial first local message of a first intelligent agent. The first intelligent agent may generate a first task message based on the first system message and the initial first local message. The first intelligent agent may add the first task message to the first message pool. The first intelligent agent may also parse the first task message to obtain a task execution request. The first intelligent agent may then send the task execution request to a second intelligent agent, and the second intelligent agent may invoke a target tool in the task execution request to execute a corresponding target task.

[0032] According to an embodiment of the present disclosure, the first intelligent agent generates first task information based on the first system message, the first local message, etc. in the first message pool, and then sends a task execution request to the second intelligent agent based on the first task information. Through the task execution request, the second intelligent agent can be controlled to fully utilize tools related to content generation to perform more tasks, fulfill more complex content generation requests, and further improve the quality of the generated content.

[0033] 4 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the first intelligent agent receiving the task execution result from the second intelligent agent in S102 includes:

[0034] In S401, the first intelligent agent receives a second task message from the second intelligent agent, where the second task message includes the task execution result and is generated by the second intelligent agent based on messages in the second message pool, the second message pool including a second system message and a second local message, the second system message is generated by the second intelligent agent based on task guidance information, and the second local message includes the task execution request from the first intelligent agent.

[0035] In an embodiment of the present disclosure, the second intelligent agent can generate a second system message based on the task guidance information and the second intelligent agent's prompt template. The second intelligent agent can add the second system message to its second message pool. The second message pool can provide the second intelligent agent with message storage and message retrieval functions. When the second intelligent agent receives a task execution request from the first intelligent agent, it can add the received task execution request to the second message pool as a second local message. Then, the second intelligent agent can invoke a target tool to execute the target task based on the second system message and the second local message, and generate a second task message including the task execution result. The second intelligent agent can add the second task to the second message pool and return the second task message including the task execution result to the first intelligent agent.

[0036] According to an embodiment of the present disclosure, by mutually communicating task execution requests and task execution results between a first intelligent agent and a second intelligent agent, content such as a video generation request can be divided into multiple task execution requests, more tasks can be executed, more complex content generation requests can be fulfilled, and higher quality content can be generated.

[0037] In one embodiment, the method further comprises:

[0038] In S402, the first intelligent agent adds the second task message to the first message pool as a newly added first local message, where the first message pool further includes a first task message corresponding to the processed first local message and the newly added first local message.

[0039] In an embodiment of the present disclosure, when the first intelligent agent receives a second task message from the second intelligent agent, the first intelligent agent can add the second task message to the first message pool as a new first local message. The first intelligent agent can return to S202, and the first intelligent agent can generate a new first task message based on messages in the first message pool, such as the first system message, the first local message, the first task message, and the new first local message. Then, the first intelligent agent can continue to execute S203, parsing the new first task message to obtain a new task execution request. Then, the first intelligent agent sends the new task execution request to the second intelligent agent, and the second intelligent agent invokes a target tool in the new task execution request to execute the corresponding target task.

[0040] In one example, referring to FIG. 2, the user agent may generate a first task message A11 based on a first system message S1 and an initial first local message U11 in a first message pool. The user agent adds A11 to the first message pool, parses A11, and sends the parsing result to the auxiliary agent. The auxiliary agent's second message pool may include a second system message S2 and a second local message U21 newly added based on the received parsing result. When the auxiliary agent invokes a target tool to perform a target task based on S2 and U21, the auxiliary agent may generate a second task message A21, add A21 to the second message pool, and send it back to the user agent. The user agent adds A21 to the first message pool as a new first local message U12, generates a new A12 based on S1, U11, A11, and U12, parses A12, and sends the parsing result to the auxiliary agent. The auxiliary agent adds a new second local message U22 to the second message pool based on the received parsing result. When the auxiliary agent invokes the target tool based on S2, U21, A21, and U22 to execute the target task, it can generate a second task message A22, add A22 to the second message pool, and send it back to the user agent. This process continues until the user agent no longer generates a new first task message task execution request or its analysis result, or until the auxiliary agent executes all the target tasks, the current content generation process is terminated, and the final task execution result is the content generation result.

[0041] According to the embodiments of the present disclosure, the first intelligent agent can add the task execution results of the second intelligent agent to its own message pool to determine subsequent task execution requests, which helps to divide the content generation request into multiple related task execution requests, and by repeatedly executing multiple target tasks, can fulfill more complex content generation requests and generate higher quality content.

[0042] 5 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the method further includes:

[0043] In S501, if the number of messages in the first message pool exceeds a set threshold, delete one or more sets of messages from the first message pool, where one set of messages includes one first local message and a corresponding first task message.

[0044] In an embodiment of the present disclosure, if the first intelligent agent detects that the number of messages in the first message pool is excessive, it can delete some of the messages. If a first local message and a subsequent first task message generated based on the first local message are considered as a pair, when messages need to be deleted, one or more pairs of messages at the top of the message pool can be deleted at a time. For example, referring to FIG. 2, in the first message pool, U11 and A11 can be deleted first, followed by U12 and A12. The second message pool can be processed in a similar manner. The thresholds for the number of messages in the first message pool and the second message pool can be the same or different. Controlling the number of messages in the message pool can effectively control the amount of computation and improve the content generation speed.

[0045] In an embodiment of the present disclosure, S302, S303, S401, S402, and S501 can be periodically executed when an execution condition is met. For example, if it is determined that the number of messages exceeds a threshold after S402, S501 can be executed to delete some of the messages. Further, S302 is executed, and new first task messages are continuously generated based on the deleted first message pool until the first intelligent agent no longer generates new first task messages or the second intelligent agent has completed all of the target tasks. Then, S303, S401, and S402 are continuously executed.

[0046] 6 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the method further includes:

[0047] In S601, a first intelligent agent receives task guidance information from a system intelligent agent, where the task guidance information is generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

[0048] In an embodiment of the present disclosure, the system intelligent agent, the first intelligent agent, and the second intelligent agent may be preset with respective prompt templates. The system intelligent agent generates task guidance information based on the input text information and its own prompt template. The input text information may include information such as text features and text vectors extracted based on the initial input text. The system intelligent agent can then send the task guidance information to the first intelligent agent and the second intelligent agent, respectively. The first intelligent agent can generate a first system message based on the task guidance information and its own prompt template (see S301). The second intelligent agent can generate a second system message based on the task guidance information and its own prompt template.

[0049] In one embodiment, the goal task includes at least one of a retrieval task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task.

[0050] In one embodiment, the target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool.

[0051] In one embodiment, the task execution results include at least one of a search result, a segmentation result, a shot division result, a matching result, a generation result, and a synthesis result.

[0052] In the embodiments of the present disclosure, the task execution results of the second intelligent agent can include intermediate task execution results and final task execution results. For example, search results, segmentation results, shot division results, matching results, and generation results are intermediate task execution results. The synthesis result is the final task execution result. The second intelligent agent can return the intermediate task execution results to the first intelligent agent for subsequent processing, and can also output or store the intermediate task execution results. The second intelligent agent can directly output the final task execution results or send the final task execution results to the first intelligent agent, and the first intelligent agent outputs the final task execution results.

[0053] In one embodiment, when the target task includes a shot division task, the target tool that needs to be called to execute the shot division task includes a shot division tool, and the task execution result includes a shot division result obtained by calling the shot division tool to execute the shot division task.

[0054] In an embodiment of the present disclosure, a task execution request sent from a first intelligent agent to a second intelligent agent may include information such as a task name, a task identifier, a tool name, and a tool identifier. The first task execution request sent from the first intelligent agent to the second intelligent agent may specify a shot division task and a shot division tool. The second intelligent agent calls the shot division tool to execute the shot division task, adds the shot division result to a second task pool as a second task message, and sends the shot division result to the first intelligent agent. The shot division result may include a text segment corresponding to each shot obtained by executing the shot division task on the text information to be processed. For example, shot 1 corresponds to shot text 1, shot 2 corresponds to shot text 2, and shot 3 corresponds to shot text 3. The text information to be processed may be text information extracted based on task guidance information.

[0055] For example, referring to FIG. 2, the analysis result of the first task message A11 sent by the user agent to the auxiliary agent may include a shot division task and a shot division tool. The auxiliary agent can call the shot division tool based on the received analysis result to obtain shot text corresponding to each shot in the text information waiting to be processed. The auxiliary agent generates a shot division result based on the second system message S1, a second local message U21 corresponding to the received analysis result, and the correspondence between the shot and the shot text, adds the shot division result to the second task pool as a second task message A21, and sends the shot division result to the user agent.

[0056] In one embodiment, when the target task includes a search task, the target tool that needs to be invoked to execute the search task includes a search tool, and the task execution results include search results obtained by invoking the search tool to execute the search task.

[0057] In an embodiment of the present disclosure, the first intelligent agent can send a second task execution request to the second intelligent agent based on the shot division result to instruct a search task and a search tool. The second intelligent agent can call the search tool to search and obtain search results based on the shot division result. The search results can include shot text and corresponding images, videos, video addresses, etc. For example, shot text 1 corresponds to video 1 and video 2, shot text 2 corresponds to video 3, and the search results can include shot text 1 corresponding to link addresses of video 1 and video 2, and shot text 2 corresponding to video 3. After generating search results based on messages in the second message pool, the search results are added to the second task pool as second task messages and sent to the first intelligent agent.

[0058] In another embodiment, the first intelligent agent can send a second task execution request to a second intelligent agent based on the pending text information, instructing the second intelligent agent to execute a search task and a search tool. The second intelligent agent can call the search tool to search and obtain search results based on the pending text information. The search results can include the pending text information and corresponding images, videos, video addresses, etc.

[0059] For example, referring to FIG. 2, the user agent may send a first task message A12 to the auxiliary agent as an analysis result, which may instruct a search task and a search tool. The auxiliary agent may call the search tool based on the received analysis result, execute a search task on the shot text in the shot segmentation result, and obtain content information corresponding to the shot text. The auxiliary agent may generate search results based on the second system message S1, the second local message U21, the second task message A21, the second local message U22 corresponding to the received analysis result, and the correspondence between the shot text and the content, add the search results to the second task pool as a new second task message A22, and send the search results to the user agent.

[0060] In one embodiment, if the target task includes a segmentation task, a target tool that needs to be called to perform the segmentation task includes a segmentation tool, and the task execution result includes invoking the computer vision tool to perform the segmentation task based on the search results to obtain content segments. For example, if the target task includes a segmentation task, a target tool that needs to be called to perform the segmentation task includes a computer vision tool, and the task execution result includes video-text pairs (alternatively referred to as text-video pairs) obtained by invoking the computer vision tool to perform the segmentation task based on the search results. The video-text pairs include segmented video segments and corresponding text segments. These video-text pairs can be stored in a designated location, such as a material library.

[0061] In an embodiment of the present disclosure, the first intelligent agent sends a third task execution request to the second intelligent agent based on the search results, instructing the segmentation task and a computer vision (CV)-based tool. The second intelligent agent can call the CV-based tool to perform a segmentation task on the search results (e.g., images, videos, etc.), generate segmentation results based on messages in the second message pool, add the segmentation results to the second task pool as second task messages, and send the segmentation results to the first intelligent agent. The segmentation results can include video-text pairs obtained by performing a segmentation task on the search results using the CV-based tool. The video-text pairs can include text segments and corresponding video segments. The video segments can be sub-materials suitable for video production. For example, the CV-based tool can be used to perform processes such as text extraction and semantic segmentation on the video to obtain correspondences between text segments and video segments, such as a [(T, V)] list. Here, (T, V) represents the video-text pair, T represents the text content extracted from the video, and V represents information such as the video sub-material cut out from the video and the address of the video sub-material. These video-text pairs can be temporarily stored in a pre-defined memory space.

[0062] For example, referring to FIG. 2, the user agent may send a first task message A13 to the auxiliary agent to indicate a segmentation task and a computer vision tool in the analysis result. The auxiliary agent may call the computer vision tool based on the received analysis result to obtain a description text corresponding to each video segment in the search result. The auxiliary agent may generate a segmentation result based on the second system message S1, the second local message U21, the second task message A21, the second local message U22, the second task message A22, the second local message U23 corresponding to the received analysis result, and the segmented text-video pairs. The auxiliary agent may add the segmentation result to the second task pool as the second task message A23 and send the segmentation result to the user agent.

[0063] In one embodiment, when the goal task includes a matching task, the goal tool that needs to be called to execute the matching task includes a matching tool, and the task execution result includes a matching result obtained by calling the matching tool and performing matching based on the shot division result.

[0064] In an embodiment of the present disclosure, the first intelligent agent sends a fourth task execution request to the second intelligent agent based on the segmentation result to instruct a matching task and a matching tool. For example, the second intelligent agent invokes the matching tool to perform a matching task on the segmentation result, semantically matching the shot text with the text segments in the segmented video-text pairs, and generating a matching result based on the messages in the second message pool. The second intelligent agent adds the matching result to the second task pool as a second task message and sends the matching result to the first intelligent agent. The matching result may include, for example, which shot texts have matched video-text pairs and which shot texts do not have matched video-text pairs. For shot texts that do not have matched video-text pairs, it may be necessary to generate corresponding video segments using another tool. Alternatively, the second intelligent agent invokes the matching tool to perform a matching task on the search results, semantically matching the shot text with the search results, and generating a matching result based on the messages in the second message pool. The matching result may include, for example, which shot texts have matched search results, such as content material and descriptive text, and which shot texts do not have matched search results. For shot text that does not have search results, it may be necessary to use another tool to generate the corresponding content segments.

[0065] For example, referring to FIG. 2, the user agent may send a first task message A14 to the auxiliary agent as an analysis result, which may indicate a matching task and a matching tool. The auxiliary agent may call the matching tool based on the received analysis result to obtain a video text pair matched with the shot text in the segmentation result. The auxiliary agent may generate a matching result based on the second system message S1, the second local message U21, the second task message A21, the second local message U22, the second task message A22, the second local message U23, the second task message A23, the second local message U24 corresponding to the received analysis result, and the matching relationship between the shot text and the video text pair, add the matching result to the second task pool as the second task message A24, and send the matching result to the user agent.

[0066] In one embodiment, if the goal task includes a generation task, the target tool that needs to be called to perform the generation task includes a generation tool, and the task execution result includes invoking the generation tool to generate content segments for the shot division results for which there are no matching results or for which content needs to be generated. For example, if the goal task includes a video segment generation task, the target tool that needs to be called to perform the video segment generation task includes a generative artificial intelligence tool. The task execution result includes invoking the generative artificial intelligence tool to generate video segments for the shot division results for which there are no matching results or for which video needs to be generated.

[0067] In an embodiment of the present disclosure, the first intelligent agent sends a fifth task execution request to the second intelligent agent based on the shot text that does not have a matching content segment in the matching result, instructing a generation task and a generation tool, such as an artificial intelligence generative computing (AIGC)-based tool. The second intelligent agent invokes the generation tool to execute the generation task on the shot text, and generates content segments based on messages in the second message pool. For example, the AIGC-based tool is used to generate video segment V4 based on shot text 2, video segment V5 based on shot text 3, etc. A second task message is generated based on the content segment or its address, identification, link, etc., and added to the second task pool, and the second task message is sent to the first intelligent agent.

[0068] Additionally, when an AIGC-based tool generates video, it can invoke other tools, such as audio-based tools, to assist in generating audio content within the video content.

[0069] For example, referring to FIG. 2, the user agent may send a first task message A15 to the auxiliary agent to indicate a generation task and an AIGC-based tool in the analysis result. The auxiliary agent may call the AIGC-based tool based on the received analysis result to obtain a video segment corresponding to the shot text. The auxiliary agent may call the generative AI-based tool based on the second system message S1, the second local message U21, the second task message A21, the second local message U22, the second task message A22, the second local message U23, the second task message A23, the second local message U24, the second task message A24, and the second local message U25 corresponding to the received analysis result to generate a video segment generation result, add the video segment generation result to the second task pool as the second task message A25, and send the video segment generation result to the user agent.

[0070] In one embodiment, when the goal task includes a content synthesis task, a target tool that needs to be called to perform the content synthesis task includes a content synthesis tool, and the task execution result includes a target content obtained by invoking the content synthesis tool to synthesize content segments corresponding to the multiple shot division results. For example, when the goal task includes a video synthesis task, a target tool that needs to be called to perform the video synthesis task includes a video synthesis tool, and the task execution result includes a target video obtained by invoking the video synthesis tool to synthesize at least one of video segments, text segments, and audio segments corresponding to the multiple shot division results.

[0071] In an embodiment of the present disclosure, the first intelligent agent sends a sixth execution request to the second intelligent agent based on information about the shot text and its corresponding content segments, such as video segments, text segments, and audio segments, to instruct the content synthesis task and the content synthesis tool. The second intelligent agent invokes the content synthesis tool to execute the content synthesis task on the matching result or generation result corresponding to the shot text, and then generates a content synthesis result corresponding to the shot text based on messages in the second message pool. Here, the matching result includes, for example, information about the video segments, audio segments, and text segments in the content text pair matching the shot text, and the generation result includes information about the video segments, audio segments, and text segments generated based on the shot text. The content synthesis results corresponding to multiple shot texts are further synthesized into target content. The target content can be directly output by the second intelligent agent. The target content can be added to the second task pool as a second task message and sent to the first intelligent agent.

[0072] For example, referring to FIG. 2, the user agent may send a first task message A16 to the auxiliary agent, which may include an analysis result indicating a video composition task and a video composition tool. Based on the received analysis result, the auxiliary agent may invoke the video composition tool to compose information such as video segments, audio segments, and text segments according to the shot text to obtain a target video. The auxiliary agent may invoke the video composition tool to generate a target video based on the second system message S1, the second local message U21, the second task message A21, the second local message U22, the second task message A22, the second local message U23, the second task message A23, the second local message U24, the second task message A24, the second local message U25, the second task message A25, and the second local message U26 corresponding to the received analysis result.

[0073] Furthermore, information on each content segment, such as a video segment, an audio segment, and a text segment, can be sorted and combined according to the timeline of the shot text to obtain a target content such as a target video.

[0074] The first to sixth task execution requests in the above embodiment are merely examples and are not limiting. In a practical application scenario, one or more task execution requests can be flexibly generated as needed. For example, the first task execution request indicates a search task and a search tool, the second task execution request indicates a shot segmentation task and a shot segmentation tool, the third task execution request indicates a matching task and a matching tool, and the fourth task execution request indicates a content synthesis task and a content synthesis tool. Furthermore, the first task execution request indicates a shot segmentation task and a shot segmentation tool, the second task execution request indicates a content segment generation task and a generation tool, and the third task execution request indicates a content synthesis task and a content synthesis tool.

[0075] According to the embodiments of the present disclosure, through the task guidance information, it is possible to instruct a shot division task, a search task, a division task, a matching task, a content segment generation task, a content synthesis task, etc., and to guide the second intelligent agent to call a target tool to perform the target task, which is helpful in fulfilling more complex content generation requirements and generating high-quality content.

[0076] 7 is a flowchart of an artificial intelligence-based content generation method according to one embodiment of the present disclosure.

[0077] In S701, the second intelligent agent receives a task execution request sent by the first intelligent agent based on task guidance information, where the task guidance information includes guidance information for generating content, and the task execution request includes a target task that the second intelligent agent needs to perform to generate content.

[0078] In S702, the second intelligent agent executes the target task according to the task execution request to generate a task execution result.

[0079] In S703, the second intelligent agent sends the task execution result to the first intelligent agent.

[0080] In the embodiment of the present disclosure, the second intelligent agent first receives the task execution request sent by the first intelligent agent, and obtains information about the tasks that need to be performed to generate the content and the tools that need to be called, etc., as indicated in the task execution request. The second intelligent agent can call the tools that perform the tasks to perform the corresponding target tasks, and obtain the task execution results. The second intelligent agent can send the task execution results to other intelligent agents.

[0081] According to the embodiments of the present disclosure, the second intelligent agent can call the target tool based on the task execution request to perform the target task, and can call tools related to content generation to perform more tasks, fulfill more complex content generation requests, and help generate high-quality content.

[0082] 8 is a flowchart of a content generation method based on artificial intelligence according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the step of the second intelligent agent executing the target task based on the task execution request to generate a task execution result in S702 includes:

[0083] In S801, the second intelligent agent adds a second system message generated based on the task guidance information to a second message pool.

[0084] In S802, the second intelligent agent adds the task execution request from the first intelligent agent to a second message pool as a second local message.

[0085] In S803, the second intelligent agent executes the target task based on the messages in the second message pool to generate the task execution result.

[0086] In an embodiment of the present disclosure, the second intelligent agent can generate a second system message based on the task guidance information and the second intelligent agent's prompt template. The second intelligent agent can add the second system message to its own second message pool. The second message pool can provide the second intelligent agent with message storage and message retrieval functions. The second intelligent agent receives a task execution request issued by the first intelligent agent and adds it to the second message pool as a second local message. The second intelligent agent can obtain the task execution result after completing the target task.

[0087] According to the embodiments of the present disclosure, the second intelligent agent can call a task execution tool to perform the corresponding target task based on the second system message and the second local message, and obtain a task execution result. Through the task execution request, the second intelligent agent can be controlled to fully utilize the tools related to content generation to perform more tasks, fulfill more complex content generation requests, and further improve the quality of the generated content.

[0088] 9 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the method further includes:

[0089] In S901, the second intelligent agent adds a second task message containing the task execution result to the second message pool.

[0090] In S703, the second intelligent agent sending the task execution result to the first intelligent agent includes the second intelligent agent sending the second task message to the first intelligent agent.

[0091] In one embodiment, the method further includes: in S903, the second intelligent agent receives a new task execution request from the first intelligent agent, and adds the task execution request to a second message pool as a new second local message, and then returns to S802 for execution.

[0092] In the embodiment of the present disclosure, S802, S803, S901, S902, and S903 can be periodically executed until the first intelligent agent no longer generates the first task message or its analysis result, or until the second intelligent agent has completed executing all target task positions.

[0093] In the embodiment of the present disclosure, the second intelligent agent can add the task execution result as a second task message to a second message pool, and the second intelligent agent can send the second task message in the second message pool to the first intelligent agent, and further send the task execution result in the second task message to the first intelligent agent as the basis for the next analysis process.

[0094] According to the embodiments of the present disclosure, the task execution request and the task execution result are transmitted between the first intelligent agent and the second intelligent agent, which helps to divide the content generation request into multiple task execution requests, thereby executing more tasks, fulfilling more complex content generation requests, and generating higher quality content.

[0095] 10 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the method further includes:

[0096] In S1001, if the number of messages in the second message pool exceeds a set threshold, delete one or more sets of messages from the second message pool, where one set of messages includes one second local message and a corresponding second task message.

[0097] In an embodiment of the present disclosure, S802, S803, S901, S902, S903, and S1001 can be periodically executed if the execution conditions are met. For example, if it is determined after S903 that the number of messages exceeds a threshold, S1001 can be executed to delete some of the messages. S802 can then be executed, and new first task messages can be generated based on the deleted first message pool, until the first intelligent agent no longer generates new first task messages or the second intelligent agent has completed all of the target tasks. S802 can then be executed, and new first task messages can then be generated based on the deleted first message pool, and S803, S901, S902, and S903 can then be executed. Furthermore, if it is determined after S901 that the number of messages exceeds a threshold, S1001 can be executed to delete some of the messages. S902 and S903 can then be executed, until the first intelligent agent no longer generates new first task messages or the second intelligent agent has completed all of the target tasks.

[0098] In an embodiment of the present disclosure, if the second intelligent agent detects that the number of messages in the second message pool is excessive, it can delete some of the messages. If a second local message and a second system message obtained by executing a task based on the second local message are considered as a set, when messages need to be deleted, one or more sets of messages at the top of the message pool can be deleted at a time. For example, referring to FIG. 2, in the second message pool, U21 and A21 can be deleted first, followed by U22 and A22. The thresholds for the number of messages in the first message pool and the second message pool can be the same or different. Controlling the number of messages in the message pool can effectively control the amount of computation and improve the content generation speed.

[0099] 11 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. This method may include one or more features of the content generation method described above. In one embodiment, the method further includes:

[0100] In S1101, the second intelligent agent receives task guidance information from the system intelligent agent, where the task guidance information is generated by the system intelligent agent according to input text information and a prompt template of the system intelligent agent.

[0101] In an embodiment of the present disclosure, a system intelligent agent, a first intelligent agent, and a second intelligent agent may be preset with respective prompt templates. The system intelligent agent generates task guidance information based on input text information and its own prompt template. The input text information may include information such as text features and text vectors extracted based on the initial input text. Then, the system intelligent agent can send the task guidance information to the first intelligent agent and the second intelligent agent, respectively. The first intelligent agent can generate a first system message based on the task guidance information and its own prompt template (see S301). The second intelligent agent can generate a second system message based on the task guidance information and its own prompt template (see S801).

[0102] 12 is a flowchart of an artificial intelligence-based content generation method according to another embodiment of the present disclosure. The method may include one or more features of the content generation methods described above. In one embodiment, the goal task includes at least one of a search task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task.

[0103] In one embodiment, the target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool.

[0104] In one embodiment, the task execution results include at least one of a search result, a segmentation result, a shot division result, a matching result, a generation result, and a synthesis result.

[0105] In one embodiment, if the target task includes a search task, the target tool that needs to be invoked to perform the search task includes a search tool, and S803 includes:

[0106] In S1201, the second intelligent agent invokes the search tool to execute the search task and obtain search results based on the message in the second message pool and the newly added second local message.

[0107] In one embodiment, if the goal task includes a segmentation task, the goal tool that needs to be invoked to perform the segmentation task includes a segmentation tool such as a computer vision tool, and S803 includes:

[0108] In S1202, the second intelligent agent, based on the message in the second message pool, calls the segmentation tool to perform the segmentation task based on the search result to obtain a content segment, for example, a video-text pair, where the video-text pair includes a segmented video segment and a corresponding text segment.

[0109] In one embodiment, if the target task includes a shot division task, and the target tool that needs to be called to perform the shot division task includes a shot division tool, S803 includes:

[0110] In S1203, the second intelligent agent, based on the message in the second message pool, calls the shot division tool to execute the shot division task and obtain a shot division result.

[0111] In one embodiment, if the target task includes a matching task, the target tool that needs to be invoked to perform the search task includes a matching tool, and S803 includes:

[0112] In S1204, the second intelligent agent calls the matching tool based on the message in the second message pool to perform matching based on the shot division result, and obtains a matching result.

[0113] In one embodiment, when the goal task includes a generation task, e.g., a video segment generation task, the goal tool that needs to be invoked to perform the generation task includes a generation tool, e.g., a generative artificial intelligence tool, and in S803 includes the following:

[0114] In S1205, the second intelligent agent calls the generation tool based on messages in the second message pool to generate content segments for the shot division results for which there are no matching results or for which content needs to be generated, for example, generating one or more video segments, text segments, audio segments, etc.

[0115] In one embodiment, if the goal task includes a content synthesis task, the goal tool that needs to be invoked to perform the content synthesis task includes a content synthesis tool, and S803 includes:

[0116] In S1206, the second intelligent agent, based on the messages in the second message pool, calls the content synthesis tool to synthesize at least one of content segments, such as video segments, text segments, and audio segments, corresponding to multiple shot division results to obtain a target content, such as a target video.

[0117] In the embodiment of the present disclosure, when the second intelligent agent executes S803, for specific examples of invoking various target tools to execute corresponding target tasks, please refer to the description of FIG. 2 in the above embodiment.

[0118] According to an embodiment of the present disclosure, shot division tasks, search tasks, division tasks, matching tasks, generation tasks, content synthesis tasks, etc. can be instructed through messages in the second message pool, and target tools can be called to perform target tasks, which helps to fulfill more complex content generation requirements and generate high-quality content.

[0119] A content generation method according to an embodiment of the present disclosure can be used in an efficient, versatile, and intelligent content creation agent system. The content can include one or more of video, audio, and text. A system framework is designed based on the Agent model, and a complete content creation tool component library is constructed by comprehensively utilizing multi-source materials from the Internet and proprietary materials. For example, large-scale model technology can be used to realize unified content creation planning, tool invocation, and content generation for creative content.

[0120] The overall framework of the system is as shown in Figure 13. The main intelligent agents of the system can include a system agent, a user agent, and an assistant agent.

[0121] The system agent performs task understanding (1) on the user's input (0), generates specific task-related information, and sends it to the user agent and auxiliary agent, respectively. The user agent and auxiliary agent generate their own system messages S to present and guarantee the core task goal that the user agent and auxiliary agent can always refer to when performing multi-steps. In this way, goal deviation / forgetting caused by the long execution process of a complex task can be avoided.

[0122] Based on the specific task information provided by the system agent, the user agent understands and plans the task, and breaks it down into specific processes that can be realized step by step. The user agent instructs the auxiliary agent to hand over the process that currently needs to be executed so that it can actually be processed, and determines the next process to be executed based on the processing results of the auxiliary agent. If the user agent's feedback result satisfies the task goal in the final system message, task completion (TASK_DONE) specification information is returned, the execution of the entire agent (Agent) task ends, and the user logs out.

[0123] The auxiliary agent analyzes the content that needs to be executed based on the prompt content of the user agent. If the execution content is related to a tool invocation, the auxiliary agent specifically executes the tool invocation, and the invocation result is sent back to the user agent for further planning.

[0124] The framework can understand and process multimodal data, mainly content such as video. Due to the complexity of the task, accurate semantic understanding and detailed processing of content data such as video data are required, and functions such as correct invocation of multiple tools under complex data structure interfaces, alignment and rendering of multimodal data are required.

[0125] The system mainly includes the following parts:

[0126] 1. Agent (Intelligent Agent) Framework

[0127] An example of an agent framework may include the following functional modules, as shown in FIG.

[0128] These are prompt templates. Sys-P, User-P, and Assit-P are prompt templates for system agents, user agents, and assistant agents, respectively, and are used to build content production systems for various needs, such as designing a video production system or understanding a specified production task.

[0129] For example, Task_specifier is a text description of the overall content production system, such as a video production system text description, Task_info is a text description of a specific content production task, such as a video production task, and Specified Task is a text description generated in the content production system after understanding the specific production task, which guides the development of the specific task.

[0130] System Message: Generates a system message (S) for the user agent based on the Specified Task and User-P, and generates a system message (S) for the auxiliary agent based on the Specified Task and Assistant-P.

[0131] This is the message pool [Messages]. The number of messages in each agent's message pool can be an odd number [S,U,A,U,A,…]. If the length of the message pool content exceeds the model input requirements, it can pop up [U,A] that was previously added to the message pool (messages). Here, S represents a system message, U represents a user type message, and A represents an assistant type message.

[0132] Parsing: Analyzes the contents of a message and includes tools (actions) and (inputs).

[0133] If the parsed tool name (Tool_name) exists in the deployed tool component library, the assistant agent (assit) will call the corresponding tool to complete the execution of the specified function.

[0134] It is also possible to set a maximum number of conversations. Before the maximum number of conversations is reached, the task is completed and the user returns a task completion message (TASK_DONE). For example, when a conversation between a user agent and a supplementary agent is completed, it can be recorded as one conversation. The maximum conversation test can be used to allow an agent system to jump out of a conversation loop more rationally.

[0135] 2. Tool Component Library

[0136] (1) Search tool components

[0137] The search tool component is used to search and retrieve relevant text, images, video, audio, and other materials based on the user's production requirements and production content.

[0138] (2) Computer Vision (CV) Tool Components

[0139] The computer vision tool components are mainly used to understand and analyze content from images, videos, etc., and sub-materials suitable for content production such as video production and editing are constructed. Among them, examples of CV tool components are as follows:

[0140] Human figure / audio broadcast detection is used to detect people in the material, which can avoid issues such as ID inconsistency when applying the material and can avoid audio / screen desynchronization issues.

[0141] Subtitle recognition is used to optimize material quality, and subtitle identification information can aid in material understanding and improve material matching similarity in the compilation process.

[0142] Caption description generation mainly relies on the implementation of multimodal models and is primarily used for content understanding such as images and videos.

[0143] Semantic shot segmentation is used to precisely process the retrieved material and divide long video material into semantically pure / accurate sub-materials, which are the basic content elements in video production tasks and can be organized and combined into new videos with different ingenuity.

[0144] (3) Artificial Intelligence Generated Content (AIGC) tool components

[0145] Digital people are used to generate digital people videos to make up for the lack of basic CV materials, while broadening the types of materials and improving the content richness of the generated videos.

[0146] Text-to-Image (T2I) generates image materials based on text, which can be used to generate more accurate and original materials, improving the originality of video generation and video quality.

[0147] Text-to-Video (T2V) fully utilizes current video generation capabilities and is used to generate short video material based on text, complementing the application of highly correlated video material.

[0148] (4) Audio-related tool components

[0149] Automatic Speech Recognition (ASR) is used to enhance the understanding of video material and assist video material matching applications.

[0150] Text to Speech (TTS) is used to generate audio based on a draft and to provide different levels of timestamp information, which is the basis for alignment in the compilation process of visual, textual and audio material.

[0151] (5) Video composition tool components

[0152] It is used to realize hybrid rendering of multi-source materials and involves various rendering effect tools, including but not limited to subtitle / keyword rendering, picture-in-picture / video, special effect animation, etc. It improves the rendering effect of the generated video and the feel of user-generated content (UGC).

[0153] 3. Enhanced material search matching in memory space

[0154] As shown in Figure 15, the present system can include a search and matching enhancement module based on the material in the storage space, which is used to achieve accurate understanding and matching of video material and improve the quality of the final video generation. The specific process is as follows:

[0155] (1) Extract text content from the related videos returned by the search using OCR and ASR, and perform semantic segmentation based on a large-scale language model to obtain [(T = text content, V = video sub-material), and ultimately obtain a [(T, V)] list. (T, V) is called the material in the memory space related to the task, and T in the material in the memory space is the search key, and V is the search value. Here, V may include the video itself or information such as the video address.

[0156] (2) For the draft with completed shot design, select the text content of the shot to be matched, search for the key in the material in the storage space, and obtain the corresponding value (video sub-material).

[0157] (3) The sub-materials are traced back to the original video using screen similarity, and the matching accurate timestamps are obtained. The optimal sub-materials are trimmed and used in the final material organization.

[0158] FIG. 16 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to an embodiment of the present disclosure, which can be used as a first intelligent agent, and which includes: a first sending module 1601 for sending a task execution request to a second intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate content; The system may include a first receiving module 1602 for receiving a task execution result from the second intelligent agent, where the task execution result includes an execution result generated by the second intelligent agent executing the target task.

[0159] 17 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to another embodiment of the present disclosure, which may include one or more features of the content generation device described above. In one embodiment, the first sending module 1601: a first generating sub-module 1701 for adding a first system message generated based on the task guidance information to a first message pool, the first message pool further including a first local message; a first generation sub-module 1702 for generating a first task message based on a message in the first message pool; a first sending sub-module 1703 for sending a task execution request to the second intelligent agent based on the first task message, the task execution request including a goal task that the second intelligent agent needs to execute to generate content;

[0160] In one embodiment, the first receiving module 1602 is further used for receiving a second task message from the second intelligent agent, the second task message including the task execution result, the second task message being generated by the second intelligent agent based on a message in the second message pool, the second message pool including a second system message and a second local message, the second system message being generated by the second intelligent agent based on task guidance information, and the second local message including the task execution request from the first intelligent agent.

[0161] In one embodiment, as shown in FIG. 17, the artificial intelligence based content generation device comprises: The system further includes a first adding module 1603 for adding the second task message to the first message pool as a newly added first local message, wherein the first message pool further includes a first task message corresponding to the processed first local message and the newly added first local message.

[0162] In one embodiment, as shown in FIG. 17, the artificial intelligence based content generation device comprises: The system further includes a first deletion module 1604 for deleting one or more sets of messages from the first message pool if the number of messages in the first message pool exceeds a set threshold, where the set of messages includes a first local message and a corresponding first task message.

[0163] In one embodiment, as shown in FIG. 17, the artificial intelligence based content generation device comprises: The system further includes a second receiving module 1605 for receiving the task guidance information from the system intelligent agent, where the task guidance information is generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

[0164] In one embodiment, the goal task includes at least one of a retrieval task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task.

[0165] In one embodiment, the target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool.

[0166] In one embodiment, the task execution results include at least one of a search result, a segmentation result, a shot division result, a matching result, a generation result, and a synthesis result.

[0167] In one embodiment, when the target task includes a search task, the target tool that needs to be invoked to execute the search task includes a search tool, and the task execution results include search results obtained by invoking the search tool to execute the search task.

[0168] In one embodiment, when the target task includes a segmentation task, the target tool that needs to be called to perform the segmentation task includes a segmentation tool, and the task execution result includes a content segment obtained by calling the segmentation tool and performing the segmentation task based on the search results.

[0169] In one embodiment, when the target task includes a shot division task, the target tool that needs to be called to execute the shot division task includes a shot division tool, and the task execution result includes a shot division result obtained by calling the shot division tool to execute the shot division task.

[0170] In one embodiment, when the goal task includes a matching task, the goal tool that needs to be called to execute the matching task includes a matching tool, and the task execution result includes a matching result obtained by calling the matching tool and performing matching based on the shot division result.

[0171] In one embodiment, when the goal task includes a content segment generation task, the goal tool that needs to be invoked to perform the content segment generation task includes a generation tool, such as a generative artificial intelligence tool, and the task execution result includes one or more generated content segments, such as a video segment, an audio segment, and a text segment, for which the generation tool is invoked and there is no matching result or content needs to be generated for the shot division result.

[0172] In one embodiment, when the target task includes a content synthesis task, a target tool that needs to be invoked to execute the content synthesis task includes a content synthesis tool, and the task execution result includes target content obtained by invoking the content synthesis tool and synthesizing content segments corresponding to multiple shot division results.

[0173] FIG. 18 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to an embodiment of the present disclosure, which can be used as a second intelligent agent, and includes: a third receiving module 1801 for receiving a task execution request sent by a first intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a target task that the second intelligent agent needs to perform to generate content; a calling module 1802 for executing the target task based on the task execution request and generating a task execution result; and a second sending module 1803 for sending the task execution result to the first intelligent agent.

[0174] 19 is a schematic diagram illustrating the configuration of an artificial intelligence-based content generation device according to another embodiment of the present disclosure, which may include one or more features of the content generation device described above. In one embodiment, the calling module 1802: a first adding sub-module 1901 for adding a second system message generated based on the task guidance information to a second message pool; a second adding sub-module 1902 for adding the task execution request from the first intelligent agent as a second local message to a second message pool; a third generating sub-module 1903 for executing the target task based on the messages in the second message pool and generating the task execution result.

[0175] In one embodiment, as shown in FIG. 19, the device comprises: a second adding module 1804 for adding a second task message including the task execution result to the second message pool; The second sending module 1803 is further used for sending the second task message to the first intelligent agent.

[0176] In one embodiment, as shown in FIG. 19, the device comprises: and a second deletion module 1805 for deleting one or more sets of messages from the second message pool when the number of messages in the second message pool exceeds a set threshold, where the set of messages includes one second local message and a corresponding second task message.

[0177] In one embodiment, as shown in FIG. 19, the device comprises: The system further includes a fourth receiving module 1806 for receiving the task guidance information from the system intelligent agent, the task guidance information being generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

[0178] In one embodiment, the goal task includes at least one of a retrieval task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task.

[0179] In one embodiment, the target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool.

[0180] In one embodiment, the task execution results include at least one of a search result, a segmentation result, a shot separation result, a matching result, a generation result, and a synthesis result.

[0181] In one embodiment, when the target task includes a search task, the target tool that needs to be called to perform the search task includes a search tool, and the third generation submodule is further used to call the search tool to perform the search task and obtain search results based on the message in the second message pool and the newly added second local message.

[0182] In one embodiment, when the target task includes a segmentation task, the target tool that needs to be invoked to perform the segmentation task includes a segmentation tool, for example, a computer vision tool, and based on the messages in the second message pool, the third generation sub-module is further used to invoke the segmentation tool to perform the segmentation task based on the search results to obtain content segments, for example, text video pairs, and the video text pairs include segmented video segments and corresponding text segments.

[0183] In one embodiment, when the target task includes a shot division task, the target tool that needs to be called to execute the shot division task includes a shot division tool, and the third generation sub-module is further used to call the shot division tool to execute the shot division task and obtain a shot division result based on a message in the second message pool.

[0184] In one embodiment, when the target task includes a matching task, the target tool that needs to be called to perform the matching task includes a matching tool, and the third generation submodule is further used to call the matching tool based on messages in the second message pool to perform matching based on the shot division result and obtain a matching result.

[0185] In one embodiment, when the target task includes a generation task, such as a video segment generation task, the target tool that needs to be invoked to perform the generation task includes a generative artificial intelligence tool, and the third generation sub-module is further used to invoke the generative artificial intelligence tool based on messages in the second message pool to generate one or more content segments, such as a video segment, an audio segment, and a text segment, for the shot division results for which there is no matching result or content needs to be generated.

[0186] In one embodiment, when the target task includes a content synthesis task, the target tool that needs to be called to perform the content synthesis task includes a content synthesis tool, and the third generation sub-module is further used to call the content synthesis tool based on messages in the second message pool to synthesize at least one of content segments, text segments, and audio segments corresponding to multiple shot division results to obtain target content.

[0187] For specific functions and exemplary descriptions of each module and sub-module of the apparatus according to the embodiments of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the above-mentioned method embodiments, and they will not be repeated here.

[0188] 20 is a schematic diagram showing the configuration of an intelligent agent system according to an embodiment of the present disclosure, the intelligent agent system includes a first intelligent agent 2001 and a second intelligent agent 2002, the first intelligent agent 2001 is used to send a task execution request to the second intelligent agent 2002 based on task guidance information, and receive a task execution result from the second intelligent agent 2002, the task guidance information includes guidance information for generating content, the task execution request includes a target task that the second intelligent agent needs to execute to generate the content, and the task execution result includes an execution result generated by the second intelligent agent executing the target task;

[0189] The second intelligent agent 2002 is used for receiving a task execution request sent by the first intelligent agent based on task guidance information, executing the target task based on the task execution request to generate a task execution result, and sending the task execution result to the first intelligent agent.

[0190] In one embodiment, the intelligent agent system comprises: The system further includes a system intelligent agent 2003 for sending the task guidance information to the first intelligent agent 2001 and the second intelligent agent 2002, respectively, wherein the task guidance information is generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

[0191] The specific functions and exemplary descriptions of each intelligent agent in the system according to the embodiments of the present disclosure may refer to the relevant descriptions of the corresponding steps in the above-mentioned method embodiments, and will not be repeated here.

[0192] In the technical solution of the present disclosure, the acquisition, storage, and application of users' personal information comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0193] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a non-transitory computer-readable storage medium, and a program product.

[0194] 21 is a block diagram of an electronic device 2100 for implementing an embodiment of the present disclosure. The electronic device refers to various types of digital computers, including, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device also refers to various types of mobile devices, including, for example, personal digital assistants, cellular phones, intelligent phones, wearable devices, and other similar computing devices. The components, their connections, and functions described in this disclosure are merely exemplary and do not limit the implementation of what is described and specified in this disclosure.

[0195] 21, device 2100 includes a computing unit 2101 that can perform various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 2102 or loaded from a storage unit 2108 into a random access memory (RAM) 2103. The RAM 2103 can further store various programs and data required for the operation of device 2100. The computing unit 2101, ROM 2102, and RAM 2103 are connected to each other via a bus 2104. An input / output (I / O) interface 2105 is also connected to the bus 2104.

[0196] Multiple components in device 2100 are connected to an I / O interface 2105, which includes input units 2106 such as a keyboard and a mouse, output units 2107 such as various displays and speakers, storage units 2108 such as a magnetic disk and an optical disk, and communication units 2109 such as a network card, a modem, a wireless communication transceiver, etc. The communication units 2109 allow device 2100 to exchange information / data with other devices via a computer network such as the Internet and / or various carrier networks.

[0197] The computing unit 2101 may be a variety of general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 2101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2101 performs each of the methods and processes described above, such as the content generation method. For example, in some embodiments, the content generation method may be implemented as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 2108. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 2100 via the ROM 2102 and / or the communication unit 2109. When the computer program is loaded into the RAM 2103 and executed by the computing unit 2101, it may perform one or more steps of the content generation method described above. Additionally, in other embodiments, the computing unit 2101 may be configured to perform the content generation method in any other suitable manner (eg, firmware).

[0198] Various embodiments of the systems or techniques described in this disclosure may be implemented using digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may involve execution by one or more computer programs executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, capable of receiving data and instructions from, and transferring data and instructions to, a storage system, at least one input device, and at least one output device.

[0199] Program code for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programming data processing apparatus, such that when the program code is executed by the processor or controller, it can perform the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on-site, partially on-site, partially on-site and partially on a remote site as a separate soft encapsulation, or entirely on a remote site or server.

[0200] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Further examples of machine-readable storage media include one or more hard-wired electrical connections, a portable computer disk cartridge, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination of the foregoing.

[0201] To provide for user interaction, the systems and techniques described herein can be implemented on a computer that includes a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, etc.) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball, etc.) for the user to provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback, etc.), and input from the user can be received in any form (e.g., acoustic input, voice input, tactile input, etc.).

[0202] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., as a data server), middleware components (e.g., an application server), front-end components (e.g., a user computer having a graphical user interface or network browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. Components of the system can be connected to each other via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0203] The computer system may include a client and a server. Typically, the client and server are remote from each other and generally interact via a communication network. The client-server relationship is created by a computer program running on a corresponding computer. The server may be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0204] It should be understood that steps can be newly ranked, added, or deleted using the various aspects of the flow shown above. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order. This disclosure is not limited thereto, as long as the technical solutions disclosed in this disclosure can achieve the desired results.

[0205] The above specific examples do not constitute limitations on the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions are possible depending on design considerations and other factors. Any modifications, equivalent replacements, improvements, etc. within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. 1. A method for artificial intelligence-based content generation, comprising: The first intelligent agent sends a task execution request to the second intelligent agent based on task guidance information, where the task guidance information includes guidance information for generating content, and the task execution request includes a goal task that the second intelligent agent needs to perform to generate the content; The first intelligent agent receives a task execution result from the second intelligent agent, the task execution result including an execution result generated by the second intelligent agent executing the target task; Artificial intelligence-based content generation method.

2. The first intelligent agent sends a task execution request to the second intelligent agent based on the task guidance information; the first intelligent agent adds a first system message generated based on the task guidance information to a first message pool, the first message pool further including a first local message; the first intelligent agent generating a first task message based on messages in a first message pool; the first intelligent agent sending a task execution request to the second intelligent agent based on the first task message, the task execution request including a goal task that the second intelligent agent needs to execute to generate content; The artificial intelligence based content generation method of claim 1.

3. The first intelligent agent receives a task execution result from the second intelligent agent, The first intelligent agent receives a second task message from the second intelligent agent, the second task message including the task execution result, the second task message being generated by the second intelligent agent based on a message in the second message pool, the second message pool including a second system message and a second local message, the second system message being generated by the second intelligent agent based on task guidance information, and the second local message including the task execution request from the first intelligent agent. The artificial intelligence based content generation method of claim 2.

4. The artificial intelligence-based content generation method includes: The method further includes: the first intelligent agent adding the second task message to the first message pool as a newly added first local message, wherein the first message pool further includes a first task message corresponding to the processed first local message and the newly added first local message. The artificial intelligence based content generation method of claim 3.

5. The artificial intelligence-based content generation method includes: and further comprising: removing one or more sets of messages from the first message pool when the number of messages in the first message pool exceeds a set threshold, where the set of messages includes a first local message and a corresponding first task message. The artificial intelligence based content generation method of claim 2.

6. The artificial intelligence-based content generation method includes: The first intelligent agent further includes receiving the task guidance information from a system intelligent agent, the task guidance information being generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent. The artificial intelligence based content generation method of claim 1.

7. the target task includes at least one of a search task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task; The target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool; The task execution results include at least one of a search result, a division result, a shot division result, a matching result, a generation result, and a synthesis result. The artificial intelligence based content generation method of claim 1.

8. When the target task includes a search task, a target tool that needs to be called to execute the search task includes a search tool, and the task execution result includes a search result obtained by executing the search task by calling the search tool.

8. The artificial intelligence based content generation method of claim 7.

9. If the target task includes a segmentation task, a target tool that needs to be called to execute the segmentation task includes a segmentation tool, and the task execution result includes a content segment obtained by calling the segmentation tool and executing the segmentation task based on the search result.

9. The artificial intelligence based content generation method of claim 8.

10. When the target task includes a shot division task, a target tool that needs to be called to execute the shot division task includes a shot division tool, and the task execution result includes a shot division result obtained by executing the shot division task by calling the shot division tool.

8. The artificial intelligence based content generation method of claim 7.

11. When the goal task includes a matching task, a goal tool that needs to be called to execute the matching task includes a matching tool, and the task execution result includes a matching result obtained by calling the matching tool and performing matching based on the shot division result; and / or If the goal task includes a generation task, a goal tool that needs to be called to execute the generation task includes a generation tool, and the task execution result includes a content segment generated by calling the generation tool for the shot division result for which there is no matching result or content needs to be generated.

11. The artificial intelligence based content generation method of claim 10.

12. When the goal task includes a content synthesis task, a goal tool that needs to be called to execute the content synthesis task includes a content synthesis tool, and the task execution result includes a goal content obtained by calling the content synthesis tool and synthesizing content segments corresponding to a plurality of shot division results; 12. The artificial intelligence based content generation method of claim 11.

13. 1. A method for artificial intelligence-based content generation, comprising: A second intelligent agent receives a task execution request sent by the first intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate the content; the second intelligent agent executes the target task based on the task execution request to generate a task execution result; the second intelligent agent transmitting the task execution result to the first intelligent agent; Artificial intelligence-based content generation method.

14. the second intelligent agent executes the target task based on the task execution request to generate a task execution result; the second intelligent agent adds a second system message generated based on the task guidance information to a second message pool; the second intelligent agent adds the task execution request from the first intelligent agent to a second message pool as a second local message; the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; 14. The artificial intelligence based content generation method of claim 13.

15. The artificial intelligence-based content generation method includes: The second intelligent agent further includes adding a second task message including the task execution result to the second message pool; The second intelligent agent sending the task execution result to the first intelligent agent includes the second intelligent agent sending the second task message to the first intelligent agent.

15. The artificial intelligence based content generation method of claim 14.

16. The artificial intelligence-based content generation method includes: and further including: deleting one or more sets of messages from the second message pool when the number of messages in the second message pool exceeds a set threshold, where the set of messages includes one second local message and a corresponding second task message.

15. The artificial intelligence based content generation method of claim 14.

17. The artificial intelligence-based content generation method includes: The second intelligent agent further includes receiving the task guidance information from a system intelligent agent, the task guidance information being generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

15. The artificial intelligence based content generation method of claim 14.

18. the target task includes at least one of a search task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task; The target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool; The task execution results include at least one of a search result, a division result, a shot division result, a matching result, a generation result, and a synthesis result.

15. The artificial intelligence based content generation method of claim 14.

19. the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; When the target task includes a search task, a target tool that needs to be called to execute the search task includes a search tool, and the second intelligent agent calls the search tool to execute the search task and obtain search results based on the message in the second message pool and the newly added second local message.

20. The artificial intelligence based content generation method of claim 18.

20. the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; If the goal task includes a segmentation task, a goal tool that needs to be invoked to perform the segmentation task includes a segmentation tool, and the second intelligent agent further includes, based on a message in the second message pool, invoking the segmentation tool to perform the segmentation task based on the search result to obtain content segments.

20. The artificial intelligence based content generation method of claim 19.

21. the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; When the target task includes a shot division task, a target tool that needs to be called to execute the shot division task includes a shot division tool, and the second intelligent agent further includes, based on a message in the second message pool, calling the shot division tool to execute the shot division task and obtain a shot division result.

20. The artificial intelligence based content generation method of claim 18.

22. the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; When the goal task includes a matching task, a goal tool that needs to be called to perform the matching task includes a matching tool, and the second intelligent agent calls the matching tool according to messages in the second message pool to perform matching based on the shot division result, thereby obtaining a matching result; If the goal task includes a generation task, a goal tool that needs to be invoked to perform the generation task includes a generation tool, and the second intelligent agent further includes at least one of invoking the generation tool based on messages in the second message pool to generate content segments for the shot division results for which there is no matching result or content needs to be generated.

22. The artificial intelligence based content generation method of claim 21.

23. the second intelligent agent executes the target task based on messages in the second message pool to generate the task execution result; When the goal task includes a content synthesis task, a goal tool that needs to be called to perform the content synthesis task includes a content synthesis tool, and the second intelligent agent further includes calling the content synthesis tool based on messages in the second message pool to synthesize content segments corresponding to a plurality of shot division results to obtain a goal content.

23. The artificial intelligence based content generation method of claim 22.

24. An artificial intelligence-based content generation device for use with a first intelligent agent, a first sending module for sending a task execution request to a second intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate the content; a first receiving module for receiving a task execution result from the second intelligent agent, the task execution result including an execution result generated by the second intelligent agent executing the target task; An artificial intelligence-based content generation device.

25. The first transmitting module: a first generating sub-module for adding a first system message generated based on the task guidance information to a first message pool, the first message pool further including a first local message; a first generation sub-module for generating a first task message based on messages in the first message pool; a first sending sub-module for sending a task execution request to the second intelligent agent based on the first task message, the task execution request including a goal task that the second intelligent agent needs to execute to generate content; 25. The artificial intelligence based content generation device of claim 24.

26. The first receiving module is receiving a second task message from the second intelligent agent, the second task message including the task execution result, the second task message being generated by the second intelligent agent based on a message in the second message pool, the second message pool including a second system message and a second local message, the second system message being generated by the second intelligent agent based on task guidance information, and the second local message including the task execution request from the first intelligent agent; 26. The artificial intelligence based content generation device of claim 25.

27. The artificial intelligence-based content generation device comprises: a first adding module for adding the second task message to the first message pool as a newly added first local message, the first message pool further including a first task message corresponding to the processed first local message and the newly added first local message; 27. The artificial intelligence based content generation device of claim 26.

28. The artificial intelligence-based content generation device comprises: a first deletion module for deleting one or more sets of messages from the first message pool when the number of messages in the first message pool exceeds a set threshold, where the set of messages includes a first local message and a corresponding first task message; 28. An artificial intelligence based content generation device according to any one of claims 25 to 27.

29. The artificial intelligence-based content generation device comprises: a second receiving module for receiving the task guidance information from a system intelligent agent, the task guidance information being generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent; 25. The artificial intelligence based content generation device of claim 24.

30. the target task includes at least one of a search task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task; The target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool; The task execution results include at least one of a search result, a division result, a shot division result, a matching result, a generation result, and a synthesis result.

25. The artificial intelligence based content generation device of claim 24.

31. When the target task includes a search task, a target tool that needs to be called to execute the search task includes a search tool, and the task execution result includes a search result obtained by executing the search task by calling the search tool.

31. The artificial intelligence based content generation device of claim 30.

32. If the target task includes a segmentation task, a target tool that needs to be called to execute the segmentation task includes a segmentation tool, and the task execution result includes a content segment obtained by calling the segmentation tool and executing the segmentation task based on the search result.

32. The artificial intelligence based content generation device of claim 31.

33. When the target task includes a shot division task, a target tool that needs to be called to execute the shot division task includes a shot division tool, and the task execution result includes a shot division result obtained by executing the shot division task by calling the shot division tool.

31. The artificial intelligence based content generation device of claim 30.

34. When the goal task includes a matching task, a goal tool that needs to be called to execute the matching task includes a matching tool, and the task execution result includes a matching result obtained by calling the matching tool and performing matching based on the shot division result; and / or If the goal task includes a generation task, a goal tool that needs to be called to execute the generation task includes a generation tool, and the task execution result includes a content segment generated by calling the generation tool for the shot division result for which there is no matching result or content needs to be generated.

34. The artificial intelligence based content generation device of claim 33.

35. When the goal task includes a content synthesis task, a goal tool that needs to be called to execute the content synthesis task includes a content synthesis tool, and the task execution result includes a goal content obtained by calling the content synthesis tool and synthesizing content segments corresponding to a plurality of shot division results; 35. The artificial intelligence based content generation device of claim 34.

36. an artificial intelligence-based content generation device for use with a second intelligent agent, a third receiving module for receiving a task execution request sent by a first intelligent agent based on task guidance information, the task guidance information including guidance information for generating content, and the task execution request including a goal task that the second intelligent agent needs to perform to generate the content; a calling module for executing the target task based on the task execution request and generating a task execution result; a second sending module for sending the task execution result to the first intelligent agent; An artificial intelligence-based content generation device.

37. The calling module: a first adding sub-module for adding a second system message generated based on the task guidance information to a second message pool; a second adding sub-module for adding the task execution request from the first intelligent agent as a second local message to a second message pool; a third generating sub-module for executing the target task based on the message in the second message pool and generating the task execution result; 37. The artificial intelligence based content generation device of claim 36.

38. The artificial intelligence-based content generation device comprises: a second adding module for adding a second task message including the task execution result to the second message pool; the second sending module is further used for sending the second task message to the first intelligent agent; 38. The artificial intelligence based content generation device of claim 37.

39. The artificial intelligence-based content generation device comprises: a second deletion module for deleting one or more sets of messages from the second message pool when the number of messages in the second message pool exceeds a set threshold, where the set of messages includes a second local message and a corresponding second task message; 38. The artificial intelligence based content generation device of claim 37.

40. The artificial intelligence-based content generation device comprises: a fourth receiving module for receiving the task guidance information from a system intelligent agent, the task guidance information being generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent; 38. The artificial intelligence based content generation device of claim 37.

41. the target task includes at least one of a search task, a segmentation task, a shot separation task, a matching task, a generation task, and a content synthesis task; The target tools that need to be invoked to perform the target task include at least one of a search tool, a segmentation tool, a shot separation tool, a matching tool, a generation tool, and a composition tool; The task execution results include at least one of a search result, a division result, a shot division result, a matching result, a generation result, and a synthesis result.

38. The artificial intelligence based content generation device of claim 37.

42. The third generation sub-module: When the target task includes a search task, a target tool that needs to be called to execute the search task includes a search tool, and is further used to call the search tool to execute the search task and obtain search results based on the messages in the second message pool and the newly added second local message.

42. The artificial intelligence based content generation device of claim 41.

43. The third generation sub-module: If the target task includes a segmentation task, a target tool that needs to be called to perform the segmentation task includes a segmentation tool, and is further used to call the segmentation tool based on messages in the second message pool to perform the segmentation task based on the search results to obtain content segments.

43. The artificial intelligence based content generation device of claim 42.

44. The third generation sub-module: When the target task includes a shot division task, a target tool that needs to be called to execute the shot division task includes a shot division tool, and is further used for calling the shot division tool to execute the shot division task and obtain a shot division result based on a message in the second message pool.

42. The artificial intelligence based content generation device of claim 41.

45. The third generation sub-module: When the target task includes a matching task, a target tool that needs to be called to perform the matching task includes a matching tool, and according to a message in the second message pool, call the matching tool to perform matching based on the shot division result to obtain a matching result; If the goal task includes a generation task, a goal tool that needs to be invoked to execute the generation task includes a generation tool, and the message pool is further used for at least one of invoking the generation tool based on messages in the second message pool to generate content segments for the shot division results for which there is no matching result or content needs to be generated.

45. The artificial intelligence based content generation device of claim 44.

46. The third generation sub-module: When the target task includes a content synthesis task, a target tool that needs to be called to perform the content synthesis task includes a content synthesis tool, and the target tool is further used for: calling the content synthesis tool to synthesize content segments corresponding to a plurality of shot division results according to messages in the second message pool to obtain a target content; 46. ​​The artificial intelligence based content generation device of claim 45.

47. An intelligent agent system comprising: a first intelligent agent; and a second intelligent agent; The first intelligent agent is used to send a task execution request to the second intelligent agent based on task guidance information, and receive a task execution result from the second intelligent agent, wherein the task guidance information includes guidance information for generating content, the task execution request includes a target task that the second intelligent agent needs to execute to generate the content, and the task execution result includes an execution result generated by the second intelligent agent executing the target task; The second intelligent agent is used for receiving a task execution request sent by the first intelligent agent based on task guidance information, executing the target task based on the task execution request to generate a task execution result, and sending the task execution result to the first intelligent agent. Intelligent agent systems.

48. The intelligent agent system comprises: The system further comprises a system intelligent agent for sending the task guidance information to the first intelligent agent and the second intelligent agent, respectively, wherein the task guidance information is generated by the system intelligent agent based on input text information and a prompt template of the system intelligent agent.

48. The intelligent agent system of claim 47.

49. at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform a method according to any one of claims 1 to 12 or any one of claims 13 to 23. Electronic devices.

50. A non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform the method of any one of claims 1 to 12 or any one of claims 13 to 23.

51. A program product comprising a program for implementing the method according to any one of claims 1 to 12 or any one of claims 13 to 23 when executed by a processor in a computer.

Citation Information

Patent Citations

  • Target content automatic generation method and device, electronic equipment and readable storage medium

    CN117332823A

  • Virtual human video generation system driven by large language model, control method and medium

    CN117827322A