Controllable content large model generation method and system based on multi-modal thinking chain reasoning
Through the multimodal thinking chain reasoning method, complex tasks are subdivided and resource transfer processes are planned, the problems of multi-agent systems in task dependence and execution order are solved, task planning and execution efficiency are improved, and multimodal data processing capabilities are enhanced.
Patent Information
- Application Number
- CN202411867569.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When dealing with complex tasks, multi-agent systems face the problem of task dependence and execution order that is difficult to accurately understand. Tool selection and coordination are complex, and the memory mechanism is insufficiently efficient.
The controllable content big model generation method based on multimodal thinking chain reasoning is adopted, complex tasks are subdivided into subtasks through the thinking chain method, resource transfer process is accurately planned, task planning complexity is reduced, and task planning and execution capabilities of multi-agent systems are improved.
The efficiency and flexibility of multi-agent systems in task planning and execution are improved, the processing capability of multi-modal data is enhanced, and a more intelligent response mechanism and efficient memory mechanism are realized.
Smart Images

Figure CN120012909A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of artificial intelligence technology, and in particular relates to a method and system for generating a large controllable content model based on multimodal thinking chain reasoning. Background Art
[0002] The research and development of multimodal systems is a challenging and innovative frontier topic in the field of AI, which can promote mutual understanding between different types of data and deeper human-computer interaction. Multimodal understanding and generation refers to the use of data from multiple different modalities by computing systems to process and understand information, while being able to generate new multimodal content or responses.
[0003] An agent is a system that can perceive its environment and make decisions based on the perceived information to achieve a specific goal. A multi-agent system consists of multiple interacting agents that can collaborate, compete, or work independently to complete complex tasks. A multi-agent-based multimodal system combines the concepts of multi-agent systems and multimodal understanding and generation, involving multiple agents working together to process and generate data of different modalities simultaneously. This type of system can effectively enhance the flexibility and efficiency of task processing.
[0004] When studying multi-agent systems, task planning is mainly used to decompose complex tasks into smaller, easier-to-perform subtasks, which are then executed or distributed to other sub-agents. However, the complexity and degree of subdivision of tasks increase the difficulty of task planning, especially when dealing with the execution order and dependency issues between tasks. In addition, as the number of tool types increases, multi-agent systems also face challenges when combining subtasks. Multi-agent systems must also handle more complex contextual information, which puts new requirements on the design of memory storage.
[0005] In view of the above analysis, the technical problems that need to be solved urgently in the prior art are:
[0006] 1. Task dependencies and execution order: Complex tasks usually require multiple subtasks to be executed in a specific order, and there are resource dependencies between these subtasks. Multi-agent systems must be able to accurately understand and handle the dependencies in these task flows to ensure that each step is carried out correctly and efficiently.
[0007] 2. Tool selection and coordination: As the number of available tools increases, the intelligent system encounters more difficulties in creating a suitable combination of subtasks. This requires not only the sub-agents to efficiently and reasonably assemble task execution tools, but also to predict the execution effects of different sub-agents during the initial task planning.
[0008] 3. Efficient memory mechanism: In multi-round dialogue scenarios, when processing instructions and resource information of multiple modal data including text, images, audio and video, the multi-agent system needs an efficient memory mechanism to effectively perceive, integrate and utilize this information to support the planning process and task execution. Summary of the invention
[0009] In view of the problems existing in the prior art, the present invention provides a method and system for generating a large controllable content model based on multimodal thinking chain reasoning.
[0010] The present invention is implemented as follows: a method for generating a large controllable content model based on multimodal thinking chain reasoning, comprising:
[0011] S1, improves the ability of multi-agents to plan task dependencies and execution order by adopting the mind chain method, which divides complex tasks into multiple subtasks and accurately plans the resource transfer process between subtasks, and then executes them in a chain manner;
[0012] S2, combining complex tools into sub-agents with different modal characteristics to reduce the number of optional sub-tasks during task planning and improve planning capabilities;
[0013] S3, with the help of an external database to save the contextual information of previous conversations, when encountering a new request, it improves the relevance and accuracy of the multi-agent response by referring to the historical interaction information and historical multimodal resources that have contextual similarities with the current request, thus achieving a more intelligent response mechanism.
[0014] Furthermore, in step S1, the subdivision of complex tasks includes:
[0015] Use task decomposition algorithms to break down complex tasks into multi-level subtasks;
[0016] The execution order of subtasks is determined by logical constraint rules and resource dependencies;
[0017] Each subtask in the task chain forms a directed graph structure of task execution by marking key inputs, outputs and dependency conditions.
[0018] Furthermore, in step S2, sub-agents of different modes are optimized in the following manner:
[0019] The language modality sub-agent uses a natural language processing model to extract semantic features;
[0020] The visual modality sub-agent extracts image features and generates image descriptions through convolutional neural networks;
[0021] The sound modality sub-agent uses audio signal processing algorithms to extract spectral features to analyze voiceprints or audio content;
[0022] Each modal sub-agent communicates and collaborates through a unified task interface protocol.
[0023] Furthermore, in step S3, the context information in the external database includes:
[0024] Historical conversation texts and their corresponding contextual semantic associations;
[0025] The input, output, execution time and resource consumption of each modal subtask;
[0026] The success rate and failure records of tasks in historical interactions;
[0027] Cosine similarity and BERT embedding technology are used to semantically match new requests with historical scenarios, thereby referencing the most relevant historical task data for response optimization.
[0028] Furthermore, the mind chain method makes multi-agent planning more controllable and flexible by breaking down complex tasks into multiple subtasks. The key to this method lies in the following aspects:
[0029] Task segmentation: Split complex overall tasks into simple and manageable subtasks, making the goals of each subtask clearer and reducing the complexity of execution.
[0030] Dependency analysis: After segmentation, clarify the dependencies between subtasks. Identify which subtasks need to be completed before other subtasks to help plan a reasonable execution order. Accurately plan the resource transfer process between subtasks to ensure that resources flow efficiently between different subtasks and avoid resource waste or conflicts.
[0031] Collaborative execution: Multiple agents can collaborate based on the dependencies and resource requirements of subtasks to improve the execution efficiency of the overall task.
[0032] Furthermore, based on the large language model and prompt engineering, the user request text is converted into a specific thought chain reasoning task to achieve task planning. According to the natural language instructions, the large model outputs four subtasks with dependencies in JSON format, namely, generating videos based on text, generating pictures based on videos, answering the main content of pictures, and writing stories based on pictures and videos. Each subtask can correspond to a specific tool or another intelligent agent. After obtaining the planning results, the multi-agent system traverses and executes all tasks, inferring in parallel for tasks that do not have dependencies, and inferring one by one in a thought chain manner for tasks that have dependencies according to the dependencies.
[0033] Furthermore, after the initial agent performs task planning, the tasks are assigned to different sub-agents to execute in the order of thought chain, and the sub-agents integrate more complex and segmented tool sets. The initial task planning includes the data transfer process between different sub-tasks, so the output modality type of the tool set of each sub-agent is consistent, ensuring that data can be correctly transferred between sub-agents.
[0034] Multimodal perception: In the process of multimodal perception task planning, the present invention realizes multimodal perception of large models by splicing resource addresses of different modes with prompt words. These resources include various types such as images, text, audio and video, so as to ensure that large models can make full use of these multimodal data during planning. When predicting the parameters of the task, we need to fill the addresses of multimodal resources into the corresponding parameter slots to ensure that the system can accurately locate and load the required data.
[0035] Task secondary planning: After the task is assigned to the sub-agent, secondary planning is performed to select the specific tool to implement the task, as well as the tool input parameters. In this example, the input content is similar to the task format of the first planning, but no dependency information is required. In the output content, "task" corresponds to the specific tool called, for example, "generate_speech_audio" corresponds to the Text-to-Speech tool. The "args" field contains the specific parameters passed into the tool. For this tool, a speech text and the character's voice type need to be passed in.
[0036] During the secondary planning process, a small number of prompt examples and instructions for planning precautions for each agent in Prompt can enable the large model to learn the ability to generate high-quality parameters. After the large model detects that the incoming content is "There is a fire in the building, use a woman's voice to persuade people to evacuate", it automatically generates the speech content in the "text" parameter in the planning result.
[0037] Multimodal question answering: When performing multimodal question answering tasks, it is necessary to combine the content of multiple modalities for summary output. Multimodal generation: When performing multimodal generation tasks, the sub-model requires text description as input. If the user instruction only has a text description, it can be directly used to generate multimodal resources. If the user uploads other modalities, it is a cross-modal generation task, such as generating a video based on a picture. When performing a cross-modal generation task, the multimodal question answering model obtains the text description of the existing resources, and then uses the large model to generate appropriate resource generation instructions.
[0038] Furthermore, in multi-round dialogue scenarios, the multi-agent system needs to process multimodal data (including text, images, audio, and video) uploaded by users and generated by the model to execute various user instructions. Existing methods usually save planning results as historical information, which makes it difficult to effectively utilize all multimodal resources. For example, the addresses of previously generated multimodal resources are not visible in historical information. Therefore, an efficient memory mechanism is proposed to improve the efficiency of the system in multimodal data processing and task execution by adding complete contextual information of past tasks to historical records.
[0039] Saving the history of context information means that in a multi-agent system, not only the planning information of each task is saved, but also all relevant details and situations in the context, including user input, task parameters, reasoning results, all uploaded and generated multimodal resources, etc. When planning tasks for new requests, historical context information is input into the big model at the same time as the new request, so as to support more intelligent historical dialogue functions.
[0040] The object of the present invention is to provide a controllable content large model generation system based on multimodal thinking chain reasoning to realize the controllable content large model generation method based on multimodal thinking chain reasoning, comprising:
[0041] Reasoning task module, used to plan thought chain reasoning tasks;
[0042] System building module, used to build multi-agent systems to achieve arbitrary modal input and output;
[0043] The memory mechanism module is used to implement the multi-round dialogue memory mechanism.
[0044] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method for generating a large model of controllable content based on multimodal thinking chain reasoning.
[0045] Another object of the present invention is to provide an information data processing terminal, which includes the controllable content large model generation system based on multimodal thinking chain reasoning.
[0046] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0047] First, the arbitrary modal combination input and output method based on multi-agent thinking chain reasoning of the present invention has the following beneficial effects compared with the prior art:
[0048] Improve task planning capabilities: Compared with single-agent methods, the combination of thought chain reasoning and multi-agent applications simplifies the complexity of single planning and improves the system's ability to handle dependencies and execution order when processing tasks.
[0049] Arbitrary combination of multimodal input and output: Integrate complex tools into sub-agents with different modal characteristics. Simplify the planning and decision-making process. And combined with prompt engineering, the arbitrary combination of input and output capabilities of text-image-audio-video complex tasks are more flexible than common multimodal systems.
[0050] Multi-round dialogue function that supports multimodal awareness: Compared with the method of directly saving planning results, saving all relevant details in the context improves the context association ability and multimodal support ability, and is more suitable for processing multi-round dialogue tasks involving multiple multimodal resources.
[0051] In summary, the present invention has significant improvements in the task planning, multi-modal arbitrary combination reasoning, and multi-modal multi-round dialogue capabilities of multi-agent systems, and has application value and advantages.
[0052] Multi-agent collaboration for primary and secondary planning: After the initial task planning, the tasks are assigned to different sub-agents to execute in the order of thought chain, and the sub-agents then subdivide the tasks into specific tools for execution. Compared with directly planning the execution process of all tools at once, the solution of the present invention simplifies the difficulty of task planning.
[0053] Multi-agents realize arbitrary modal input and output: Agents autonomously call different types of multimodal tools and assemble them reasonably to handle any combination of input and output.
[0054] Multi-round dialogue implementation solution: The method of the present invention for preserving contextual reasoning details and situations in historical dialogues improves the context association capability and multi-modal support capability, and effectively handles multi-round dialogue tasks involving multiple modal resources.
[0055] Second, (1) the expected benefits and commercial value of the technical solution of the present invention after transformation are:
[0056] The technical solution of the present invention is expected to improve the task planning and execution efficiency of multi-agent systems, especially when dealing with complex and highly dependent tasks. Through the multimodal thinking chain reasoning method, the intelligence and controllability of content generation can be improved, providing more accurate and flexible support for applications such as dialogue systems, intelligent customer service, and automated decision-making in the field of artificial intelligence. In terms of commercial value, this technology can promote the research and development of intelligent assistants, virtual customer service, and other products, and has broad market prospects and potential.
[0057] (2) The technical solution of the present invention fills the technical gap in the industry at home and abroad:
[0058] At present, in multimodal intelligent agent systems at home and abroad, most methods focus on the ability to perform a modal task decomposition in words, and it is difficult to efficiently handle complex, multi-modal combined input and output requirements. The present invention provides an innovative task planning and execution mechanism by combining multimodal thinking chain reasoning, filling the technical gap in the arbitrary combination of multi-modal data input and output.
[0059] (3) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never been able to solve successfully:
[0060] This invention optimizes the inefficiency and lack of flexibility of task planning in traditional intelligent agent systems, especially the challenges of combining multimodal data and task sequence planning. By introducing thought chain reasoning and multimodal sub-agent combination, it can effectively handle complex dependencies and task switching problems, and improve the system's responsiveness and intelligence level.
[0061] (4) The technical solution of the present invention overcomes technical prejudice:.
[0062] Traditional intelligent agent systems often rely on single-modal input, ignoring the importance of combining multimodal information and task decomposition. This invention breaks through this limitation and proposes a system architecture based on multimodal thinking chain reasoning, which enables intelligent agents to perform arbitrary modal combined input and output in more complex situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flow chart of a method for generating a controllable content large model based on multimodal thinking chain reasoning provided by an embodiment of the present invention;
[0064] Figure 2 It is a structural diagram of a controllable content large model generation system based on multi-modal thinking chain reasoning provided by an embodiment of the present invention;
[0065] Figure 3 is a schematic diagram of constructing a multi-agent system provided by an embodiment of the present invention;
[0066] Figure 4 It is a schematic diagram of comprehensive understanding and reasoning provided by an embodiment of the present invention;
[0067] Figure 5 is a schematic diagram of generating text instructions provided by an embodiment of the present invention;
[0068] Figure 6 It is a schematic diagram of a multi-step reasoning dependency scenario provided by an embodiment of the present invention;
[0069] Figure 7 A schematic diagram of generating an audio segment according to an instruction provided by an embodiment of the present invention;
[0070] Figure 8 It is a schematic diagram of video generation provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0072] like Figure 1 As shown, the method for generating a large controllable content model based on multimodal thinking chain reasoning includes:
[0073] S1, improves the ability of multi-agents to plan task dependencies and execution order by adopting the mind chain method, which divides complex tasks into multiple subtasks and accurately plans the resource transfer process between subtasks, and then executes them in a chain manner;
[0074] S2, combining complex tools into sub-agents with different modal characteristics to reduce the number of optional sub-tasks during task planning and improve planning capabilities;
[0075] S3, with the help of an external database to save the contextual information of previous conversations, when encountering a new request, it improves the relevance and accuracy of the multi-agent response by referring to the historical interaction information and historical multimodal resources that have contextual similarities with the current request, thus achieving a more intelligent response mechanism.
[0076] The present invention processes complex planning tasks through the thought chain method. Thought chain reasoning is a reasoning method based on progressive decomposition, which can decompose complex tasks into a series of subtasks with a logical order. By clarifying the goals, resource requirements and completion conditions of each subtask, the dependencies between tasks are established into a chain structure. The planning of tasks starts from the initial node and gradually advances along the chain structure, so that the multi-agent system can show strong planning capabilities in task dependencies and execution order.
[0077] By decomposing complex tasks, each subtask is given different modal features, including language modality, visual modality, sound modality or other modalities, to simulate the properties of different tasks. The subtask modal design is used to integrate tools into sub-agents with specific functions, which greatly reduces the number of subtasks that need to be considered during task planning and avoids overly complex task allocation problems. Sub-agents complete specific functions in their respective modalities, thereby improving the overall planning efficiency and execution accuracy of the system.
[0078] The present invention records and stores the context information of previous conversations and the multimodal resources used in tasks through an external database, and builds a dynamic historical interaction information library. When encountering a new request, the system can reference relevant historical data and multimodal resources by retrieving historical interaction information with contextual similarity to the current request in the database. This approach can significantly improve the relevance of the system's response to requests, while avoiding repeated processing of similar tasks and saving computing resources.
[0079] During the task execution phase, the system combines historical information with task modalities, enabling multi-agents to accurately call sub-agents of corresponding modalities when executing chained tasks, and make accurate decisions based on contextual information in external databases. For example, when a new task involves multi-step reasoning, the system will prioritize similar reasoning paths in history to optimize task planning and reduce potential reasoning errors. This approach enhances the response accuracy of the multi-agent system, making it more intelligent in complex task scenarios.
[0080] In order to improve the overall efficiency of the system, the present invention supports real-time task optimization. After each execution, the task chain will dynamically adjust the priority and dependency of subtasks according to the actual results. For example, when a bottleneck frequently occurs in the subtasks of a certain mode, the system will re-plan the execution order or mode allocation method of the task in the chain structure. Through this dynamic adjustment, the system can gradually optimize the task chain and improve the efficiency of task execution.
[0081] The present invention uses multimodal resource collaboration and intelligent reasoning methods to enable the system to autonomously complete task decomposition, resource call and task chain execution in complex scenarios. Combined with the contextual relevance of the external database, the system has a higher understanding of input requests and can make highly relevant responses in real time. For example, when a user initiates a complex cross-modal task request, the system can automatically coordinate sub-agents of different modalities to complete the task chain and ultimately generate high-quality controllable content output.
[0082] In summary, the present invention utilizes the combination of thought chain reasoning and multimodal features to significantly improve the planning ability, response efficiency and intelligence of large models in complex tasks.
[0083] like Figure 2 As shown, the controllable content large model generation system based on multimodal thinking chain reasoning includes:
[0084] Reasoning task module, used to plan thought chain reasoning tasks;
[0085] System building module, used to build multi-agent systems to achieve arbitrary modal input and output;
[0086] The memory mechanism module is used to implement the multi-round dialogue memory mechanism.
[0087] In response to the above problems and the solutions proposed in the present invention, the specific implementation principles are now explained from three aspects: planning thinking chain reasoning tasks, multi-agent construction, and multi-round dialogue memory mechanism.
[0088] 1. Planning the thought chain reasoning task
[0089] The thought chain method makes multi-agent planning more controllable and flexible by breaking down complex tasks into multiple subtasks. The key points of this method are as follows:
[0090] Task segmentation: Split complex overall tasks into simple and manageable subtasks, making the goals of each subtask clearer and reducing the complexity of execution.
[0091] Dependency analysis: After segmentation, clarify the dependencies between subtasks. Identify which subtasks need to be completed before other subtasks to help plan a reasonable execution order. Accurately plan the resource transfer process between subtasks to ensure that resources flow efficiently between different subtasks and avoid resource waste or conflicts.
[0092] Collaborative execution: Multiple agents can collaborate based on the dependencies and resource requirements of subtasks to improve the execution efficiency of the overall task.
[0093] Through these mechanisms, the mind chain method improves the efficiency, flexibility and success rate of multi-agents in planning tasks, allowing complex tasks to be completed more efficiently.
[0094] Based on the large language model and prompt engineering, the user request text is converted into a specific thought chain reasoning task to achieve task planning. The task planning result meets the following format:
[0095]
[0096]
[0097] In this standardized format, each task object contains the following key-value pairs:
[0098] "task": task name
[0099] "id": unique identifier of the task
[0100] "dep": the identifier of the dependent task
[0101] "args": An object containing task parameters, including:
[0102] "text": text content or dependent resource ID
[0103] "image": image address or dependent resource ID
[0104] "audio": audio address or dependent resource ID
[0105] "video": video address or dependent resource ID
[0106] In order to enable the large model to output the task planning results strictly in the above format, the prompt engineering design for task planning is as follows:
[0107] Parse_task_prompt = 'You are an AI assistant who is good at using tools to gradually complete user needs. The optional subtasks and their output types are as follows:<task_List> You can parse user input into one or more tasks and output them in the following json format: <task planning format description>. Please note the special tag " <generated>"-dep_id" means the text / image / audio / video resources previously generated by this round of dialogue are dependent (please consider whether the dependent task generates resources of this type), and "dep_id" must be in the "dep" list. The "dep" field indicates the "id" of the prerequisite tasks that the current task depends on, and these prerequisite tasks have generated some resources first. If the current task does not depend on the resources generated by the previous task, "dep" is [-1]. The key of the "args" field must be selected from ["text", "image", "audio", "video"], and the resources should be filled in the correct position according to their type. Consider all the tasks required to solve the user request step by step, and do not generate tasks that are not related to the user request. Parse as few tasks as possible while ensuring that the user request can be parsed. Pay attention to the dependencies and order between tasks, pay attention to the information in the historical dialogue, which saves previous user requests, planning results and reasoning results, and ensure that the output information is accurate and concise. It is necessary to ensure that the output conforms to the JSON format. If the user input cannot be parsed, reply directly with \'[]\'. Here are some specific examples: <example>. Note that when a task needs to rely on a previous resource, the type of the dependent resource is determined based on the output type of the previous task. Next, the user will enter the request text, and you will directly output the planning results of the task. '
[0108] Three special tags are inserted according to the contents of the configuration file when the program is executed. They are:
[0109] <Task_List> : The name of the callable subtask and a detailed description of its function.
[0110] <Task planning format description>: The standard format and field description of the above task planning.
[0111] <example>: Few-shot prompt examples are given to improve planning ability.
[0112] <Task_List> The subtask information inserted at the position contains four fields: "name", "desc", "input", and "output", corresponding to the name, detailed description of the function, input modality type, and output modality type. By modifying the content of the configuration file, you can quickly iterate the multi-agent framework with different functions.
[0113] <example>The location prompt examples help the large model quickly learn the correct task planning mode. The input and output forms of the examples are consistent with the actual user input and planning results. The examples are as follows:
[0114]
[0115]
[0116]
[0117] This example includes a thought chain reasoning task. According to the natural language instructions, the large model outputs four dependent subtasks in JSON format, namely, generating videos based on text, generating pictures based on videos, answering the main content of pictures, and writing stories based on pictures and videos. Each subtask can correspond to a specific tool or another intelligent agent. For example, "image_generation" can refer to a specific image generation model. When executing this task, the planned parameters are used as model parameters for reasoning. It can also refer to an intelligent agent with more subdivided functions specifically for image generation.
[0118] After obtaining the planning results, the multi-agent system traverses and executes all tasks. For tasks that do not have dependencies, they are reasoned in parallel. For tasks that have dependencies, they are reasoned one by one in a chain of thought according to the dependencies.
[0119] 2. Build a multi-agent system to achieve arbitrary modal input and output
[0120] The present invention aims to realize arbitrary modal input and output of complex tasks of text-image-audio-video, and designs and implements a multi-agent system. The system integrates a series of understanding, generation and processing tools involving four modalities, and the overall model parameter volume reaches 113B. The schematic diagram is shown in Figure 3 shown.
[0121] After the initial agent performs task planning, the tasks are assigned to different sub-agents to execute in the order of thought chain, and the sub-agents integrate more complex and segmented tool sets. The initial task planning includes the data transfer process between different sub-tasks, so the output modality type of each sub-agent's tool set is consistent, ensuring that data can be correctly transferred between sub-agents.
[0122] Multimodal perception: In the process of multimodal perception task planning, the present invention realizes multimodal perception of large models by splicing resource addresses of different modes with prompt words. These resources include various types such as images, text, audio and video, so as to ensure that large models can make full use of these multimodal data during planning. When predicting the parameters of the task, we need to fill the addresses of multimodal resources into the corresponding parameter slots to ensure that the system can accurately locate and load the required data.
[0123] Task secondary planning: After the task is assigned to the sub-agent, secondary planning is performed to select the specific tool to implement the task, as well as the tool input parameters. The secondary planning prompts the following:
[0124] Second_parse_task_prompt = 'You are a task planning assistant that can break down the user's tasks into more specific tasks. The user's input and your output follow the same json format: [{"task":"task_name","args":{...}}]. The task input and output of the available tools are as follows:<Tool_List> . Pay attention to the following points when planning: <Notes on planning for each agent>. Here are some specific examples: <example>. Now please plan the tasks entered by the user. If the user input contains irrelevant resource information, just ignore it. Note that the output must conform to the json format. '
[0125] Three special tags are inserted according to the contents of the configuration file when the program is executed. They are:
[0126] <Tool_List> : The name of the sub-tool that can be called and the detailed description of its function.
[0127] <Planning considerations for each agent>: Planning considerations for each agent.
[0128] <example>: Few-shot hint examples for sub-agent planning.
[0129] A specific planning example is as follows:
[0130]
[0131]
[0132] In this example, the input is similar to the first planned task format, but no dependency information is required. In the output, "task" corresponds to the specific tool called, for example, "generate_speech_audio" corresponds to the Text-to-Speech tool. The "args" field contains the specific parameters passed to the tool, for which a speech text and the character's voice type need to be passed.
[0133] During the secondary planning process, a small number of prompt examples and instructions for planning precautions for each agent in Prompt can enable the large model to learn the ability to generate high-quality parameters. After the large model detects that the incoming content is "There is a fire in the building, use a woman's voice to persuade people to evacuate", it automatically generates the speech content in the "text" parameter in the planning result.
[0134] Multimodal question answering: When performing multimodal question answering tasks, it is necessary to combine the content of multiple modes for summary output. The multimodal information summary prompt project is as follows:
[0135] Prompt = 'You are an intelligent question-answering system that can give accurate and precise replies in Chinese based on the information provided by the user. The user will provide a requirement, as well as additional information such as pictures, videos, and audio descriptions. If the user's requirement is clear, please refer to the uploaded additional information. Output around the user's needs, and note that not every additional information is necessarily useful. If you cannot determine the user's specific needs, directly output all the additional information. Note that your reply is accurate and can have line breaks. Now, the additional information provided by the user is: <resource description>. The user input requirement is: <user instruction>. Your answer is:'
[0136] <Resource description> is the text description information of all modalities summarized by calling the question-answering model of the corresponding modality. <User instruction> is the current task instruction. Finally, the large model completes the question-answering task of multimodal resources through the above prompt template.
[0137] Multimodal generation: When performing multimodal generation tasks, the sub-model requires a text description as input. If the user instruction only has a text description, it can be used directly to generate multimodal resources. If the user uploads other modalities, it is a cross-modal generation task, such as generating a video based on an image. When performing a cross-modal generation task, the multimodal question-answering model obtains the text description of the existing resources, and then uses the large model to generate appropriate resource generation instructions. The prompt engineering for cross-modal resource generation is as follows, taking image generation as an example:
[0138] Prompt = 'You are a scene description system that can generate a description of a picture scene in less than 30 words in English based on the user's input information. Please use your imagination to guess the user's ideas, directly output the description text of the picture scene in English, and try to retain the main content of the additional scene description. For example: the user's requirement is to generate a new video by combining all resources. The additional scene includes a video of a building on fire, the sound of a puppy barking, and a picture of two firefighters putting out a fire. Your output: a building on fire, two firefighters putting out a fire, and a cute puppy watching next to it. The user's requirement is to let the animals in the picture and video stand together. The additional scene includes a picture of a white kitten and a video of a black puppy. Your output: a white kitten and a black puppy standing together. The user's requirement is to generate pictures by combining uploaded resources. The additional scene includes a picture of a building on fire, the sound of a puppy barking, and a video of a building on fire. Your output: a building on fire, and a dog barking next to the building. Note that the output only contains the picture scene description you generated, and do not include any additional explanations. Now the user has uploaded some additional scene descriptions for your reference: <resource description>, and the user's requirement is: <user instruction>. Please output the description of the picture scene:'
[0139] In this prompt template, <resource description> is the text description of the upload mode, which is generated by the question-answering model. <user instruction> is the instruction generated across modalities. Finally, the large model outputs the text description of the image scene, and this scene description is passed as a parameter to the image generation model to generate an image that meets the user's needs.
[0140] The above content can realize arbitrary modal input and output of text, image, audio and video.
[0141] 3. Multi-round dialogue memory mechanism
[0142] In multi-round dialogue scenarios, the multi-agent system needs to process multimodal data (including text, images, audio, and video) uploaded by users and generated by models to execute various user instructions. Existing methods usually save planning results as historical information, which makes it difficult to effectively utilize all multimodal resources. For example, the addresses of previously generated multimodal resources are not visible in historical information. Therefore, an efficient memory mechanism is proposed to improve the efficiency of the system in multimodal data processing and task execution by adding complete contextual information of past tasks to historical records.
[0143] The memory information is saved in the database in JSON format. The specific format is as follows:
[0144]
[0145]
[0146]
[0147]
[0148] Saving the history of context information means that in a multi-agent system, not only the planning information of each task is saved, but also all relevant details and situations in the context, including user input, task parameters, reasoning results, all uploaded and generated multimodal resources, etc. When planning tasks for new requests, historical context information is input into the big model at the same time as the new request, so as to support more intelligent historical dialogue functions.
[0149] Based on the above key technical principles and implementation methods, a multimodal content understanding and generation framework based on multi-agent thought chain reasoning with multi-round dialogue function can be realized.
[0150] The invention is combined with the Gewu platform of our company, and the specific product embodiments are as follows:
[0151] (1) Figure 4 As shown in the figure, it can comprehensively understand and reason about multimodal data such as images, audio and video input by users.
[0152] (2) Figure 5 As shown, the three modes of image, audio and video are input and the three modes of image, audio and video are generated according to the text instructions.
[0153] (3) Figure 6 As shown, we can accurately understand the multi-step reasoning dependency scenarios.
[0154] (4) Figure 7 As shown, an audio segment is generated according to the instructions. It can be at the semantic level, or it can be a human repeating the original text, or it can be a model creating and then repeating an audio segment.
[0155] (5) Figure 8 As shown, the video generation has sound effect function, which can generate background sound and synchronize the speech of the characters in the input video.
[0156] The controllable content large model generation method based on multimodal thinking chain reasoning of the present invention has achieved remarkable technical effects through practical application and implementation. The specific evidence and results are as follows:
[0157] During implementation, the system is able to comprehensively process and understand data in multiple modalities (such as images, audio, video, etc.), and enhance the reasoning ability of the model through multimodal fusion. For example, Figure 4 As shown in the figure, the multimodal data such as images, audio and video input by the user can accurately capture the correlation between the data and provide accurate responses and feedback after the system's reasoning and processing. Compared with traditional single-modal processing, the system can better understand multiple information in complex scenes and improve the intelligence of user interaction.
[0158] The system divides complex tasks into multiple subtasks through the thought chain method, and accurately plans the resource transfer process between subtasks, thereby ensuring the rationality of the task execution order and the accuracy of the dependencies. By enhancing the multi-agent task planning ability, the system can efficiently handle multi-step reasoning tasks with high complexity. The experimental results show that the implemented system can more accurately grasp each link of the task when facing complex multi-step tasks, avoiding task conflicts and resource waste.
[0159] By combining complex tools into sub-agents with different modal characteristics (such as Figure 5 As shown in the figure, the system can effectively reduce the number of optional subtasks in task planning, improving the efficiency and accuracy of task planning. This approach significantly reduces redundant calculations and invalid reasoning, ensuring that the system can complete tasks within a reasonable time. Especially in application scenarios that require rapid response, the system performs much better than traditional models.
[0160] The system effectively improves the relevance and accuracy of multi-agent responses by referencing historical conversations and multimodal resources stored in external databases. When faced with new requests, the system can quickly locate the historical information most relevant to the current request based on reasoning about historical interaction data and contextual similarities, thereby providing users with more intelligent and personalized feedback. For example, in response to multiple questions from users, the system can perform intelligent reasoning based on contextual information in the conversation history and provide highly continuous and relevant responses, significantly improving the interactive experience.
[0161] In multimodal generation tasks, after the implementation of the present invention, the system can not only generate high-quality multimodal outputs such as images, audio, and video according to user instructions, but also accurately control the quality and content of the output. Figure 6 As shown in Figure 1, the system can accurately understand multi-step reasoning dependency scenarios, ensuring that the generated content meets user needs and has high relevance and accuracy. In particular, in audio and video generation tasks, the system can generate content with high semantic consistency according to instructions, and achieve advanced functions such as background sound and character voice synchronization when outputting audio (such as Figure 7 and Figure 8 ), which improves the professionalism and authenticity of the generated content.
[0162] In summary, the present invention has significantly improved the system's capabilities in multi-agent task planning, resource transfer, historical information reference, and multi-modal data generation through an innovative multi-modal thinking chain reasoning method, and has achieved relatively ideal technical effects. In practical applications, it has been successfully applied in multiple fields (such as image, audio, video generation and reasoning, intelligent dialogue systems, etc.).
[0163] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0164] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.< / example> < / example> < / example> < / example> < / example> < / generated>
Claims
1. A method for generating a large controllable content model based on multimodal thinking chain reasoning, characterized in that: include: S1, improves the ability of multi-agents to plan task dependencies and execution order by adopting the mind chain method, which divides complex tasks into multiple subtasks and accurately plans the resource transfer process between subtasks, and then executes them in a chain manner; S2, combining complex tools into sub-agents with different modal characteristics to reduce the number of optional sub-tasks during task planning and improve planning capabilities; S3, with the help of an external database to save the contextual information of previous conversations, when encountering a new request, it improves the relevance and accuracy of the multi-agent response by referring to the historical interaction information and historical multimodal resources that have contextual similarities with the current request, thus achieving a more intelligent response mechanism.
2. The method for generating a large controllable content model based on multimodal thought chain reasoning as claimed in claim 1, characterized in that: The thought chain method makes multi-agent planning more controllable and flexible by decomposing complex tasks into multiple subtasks; the key points of this method are as follows: Task segmentation: split the complex overall task into simple and manageable subtasks, making the goal of each subtask clearer and reducing the complexity of execution; Dependency analysis: After segmentation, clarify the dependencies between subtasks; Helps plan a reasonable execution order by identifying which subtasks need to be completed before other subtasks; Accurately plan the resource transfer process between subtasks to ensure efficient flow of resources between different subtasks and avoid waste or conflict of resources; Collaborative execution: Multiple agents can collaborate based on the dependencies and resource requirements of subtasks to improve the execution efficiency of the overall task.
3. The method for generating a large controllable content model based on multimodal thought chain reasoning as claimed in claim 1, characterized in that: Based on the large language model and prompt engineering, the user request text is converted into a specific thought chain reasoning task to realize task planning; according to the natural language instructions, the large model outputs four subtasks with dependencies in JSON format, namely generating videos based on text, generating pictures based on videos, answering the main content of pictures, and writing stories based on pictures and videos; each subtask can correspond to a specific tool or another intelligent agent; after obtaining the planning results, the multi-agent system traverses and executes all tasks, and parallel reasoning is performed for tasks that do not have dependencies, and reasoning is performed one by one in a thought chain manner for tasks that have dependencies according to the dependencies.
4. The method for generating a large controllable content model based on multimodal thought chain reasoning as claimed in claim 1, characterized in that: After the initial agent performs task planning, the tasks are assigned to different sub-agents to execute in the order of thought chain, and the sub-agents integrate more complex and segmented tool sets; the initial task planning includes the data transfer process between different sub-tasks, so the output modality type of the tool set of each sub-agent is consistent, ensuring that the data between sub-agents can be correctly transferred; Multimodal perception: In the process of multimodal perception task planning, the present invention realizes multimodal perception of large models by splicing resource addresses of different modes with prompt words; these resources include multiple types of images, texts, audios and videos, so as to ensure that large models can make full use of these multimodal data during planning; when predicting the parameters of the task, we need to fill the addresses of multimodal resources into the corresponding parameter slots to ensure that the system can accurately locate and load the required data; Task secondary planning: After the task is assigned to the sub-agent, secondary planning is performed to select the specific tool to implement the task and the tool input parameters. In this example, the input content is similar to the task format of the first planning, but no dependency information is required. In the output content, "task" corresponds to the specific tool called, for example, "generate_speech_audio" corresponds to the Text-to-Speech tool; the "args" field contains the specific parameters passed into the tool. For this tool, a speech text and the character's voice type need to be passed in. In the secondary planning process, a small number of prompt examples and instructions for planning precautions for each agent in Prompt can enable the large model to learn the ability to generate high-quality parameters. After the large model detects that the incoming content is "a fire broke out in the building, and a woman's voice was used to persuade people to evacuate", it automatically generates the speech content in the "text" parameter in the planning result; Multimodal question answering: When performing multimodal question answering tasks, it is necessary to combine the content of multiple modalities for summary output. Multimodal generation: When performing multimodal generation tasks, the sub-model requires a text description as input. If the user instruction only has a text description, it can be used directly to generate multimodal resources; if the user uploads other modalities, it is a cross-modal generation task, such as generating a video based on a picture; when performing a cross-modal generation task, the multimodal question answering model obtains the text description of the existing resources, and then uses the large model to generate appropriate resource generation instructions.
5. The method for generating a large controllable content model based on multimodal thought chain reasoning as claimed in claim 1, characterized in that: Saving historical records of situational information means that in a multi-agent system, not only the planning information of each task is saved, but also all relevant details and situations in the context, including user input, task parameters, reasoning results, all uploaded and generated multimodal resources, etc.; when planning tasks for new requests, historical situational information is input into the large model at the same time as the new request, so as to support more intelligent historical dialogue functions.
6. The method for generating a large controllable content model based on multimodal thought chain reasoning according to claim 1 is characterized in that: In step S1, the subdivision of complex tasks includes: Use task decomposition algorithms to break down complex tasks into multi-level subtasks; The execution order of subtasks is determined by logical constraint rules and resource dependencies; Each subtask in the task chain forms a directed graph structure of task execution by marking key inputs, outputs and dependency conditions.
7. The method for generating a large controllable content model based on multimodal thought chain reasoning according to claim 1 is characterized in that: In step S2, sub-agents of different modes are optimized in the following ways: The language modality sub-agent uses a natural language processing model to extract semantic features; The visual modality sub-agent extracts image features and generates image descriptions through convolutional neural networks; The sound modality sub-agent uses audio signal processing algorithms to extract spectral features to analyze voiceprints or audio content; Each modal sub-agent communicates and collaborates through a unified task interface protocol.
8. The method for generating a large controllable content model based on multimodal thought chain reasoning according to claim 1 is characterized in that: In step S3, the context information in the external database includes: Historical conversation texts and their corresponding contextual semantic associations; The input, output, execution time and resource consumption of each modal subtask; The success rate and failure records of tasks in historical interactions; Cosine similarity and BERT embedding technology are used to semantically match new requests with historical scenarios, thereby referencing the most relevant historical task data for response optimization.
9. A controllable content large model generation system based on multimodal thinking chain reasoning that implements the controllable content large model generation method based on multimodal thinking chain reasoning as described in any one of claims 1 to 8, characterized in that: include: Reasoning task module, used to plan thought chain reasoning tasks; System building module, used to build multi-agent systems to achieve arbitrary modal input and output; The memory mechanism module is used to implement the multi-round dialogue memory mechanism.
10. An information data processing terminal, comprising the controllable content large model generation system based on multimodal thinking chain reasoning as described in claim 9.
Citation Information
Cited By
Information extraction method and storage medium
CN120430420A
Multi-dimensional knowledge management method and system based on multi-agent collaboration
CN120494071A
Mine ventilation intelligent decision-making method and system driven by large model intelligent agent
CN120746187A
Gravitational dam design spatial feature extraction and data set creation method based on thinking chain
CN121234313A
Task planning and subtask execution optimization method, equipment and medium
CN121501457A