A dialectical knowledge transfer learning method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202410452961.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-04-16
AI Technical Summary
[0003]然而,现有的智能体通常由于其自回归训练机制和训练数据同质等问题,在多模态问答任务的推理过程中往往会出现幻觉(如智能体输出内容与提问内容明显不匹配或逻辑错误),导致现有的智能体在处理多模态问答任务时的可靠性较差,阻碍了其实际落地应用
[0029]本发明中任务数据库中包括多个多模态问答任务,基于此得到的辩证知识库有利于提高目标智能体处理复杂的多模态问答任务的能力。在利用多个预设的第一智能体针对多模态问答任务进行多角度决策的过程中,当每个第一智能体输出的决策信息满足预设置信条件时,再将每个决策信息整合成辩证知识,一方面有利于确保辩证知识的可靠性,另一方面使得辩证知识能够涵盖不同的第一智能体在面对多模态问答任务时所持有的不同“观点”(即决策信息),确保辩证知识的全面性。其中,多角度决策包括独立决策和/或交互决策,当多个第一智能体进行独立决策后即可满足预设置信条件时,将每个决策信息整合成辩证知识,在确保决策信息可靠性的同时,还能提高获取辩证知识的效率。而本发明中多个智能体之间进行交互决策有利于彼此纠正推理过程中出现的错误。因此,当直接令多个第一智能体进行交互决策,或者当多个第一智能体进行独立决策后无法满足预设置信条件时,令多个第一智能体进行交互决策,有利于提高每个第一智能体对应输出决策信息的准确性。由此,针对每个多模态问答任务都能得到全方位、多角度且可靠性强的辩证知识。在此基础上,遍历任务数据库得到多个辩证知识,并基于构建的辩证知识库训练预设的第二智能体,得到目标智能体,相当于将基于多个第一智能体得到的多角度辩证知识赋予目标智能体,使得单一的目标智能体能够具备多个第一智能体之间共同协作时(即独立决策和/或交互决策)才能具备的“辩证决策”能力。有效降低了单个目标智能体在进行复杂的多模态问答任务推理时出现幻觉的概率,缩小了单个目标智能体独立推理能力与多个第一智能体联合推理能力的差距,提高了目标智能体处理多模态问答任务的可靠性。
Smart Images

Figure CN118246535B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a dialectical knowledge transfer learning method, apparatus, electronic device, and storage medium. Background Technology
[0002] In recent years, with the rapid development of deep learning technology, large models such as GPT have achieved remarkable results in fields such as natural language question answering. However, in practical applications, the limitations of a single large language model when facing multimodal question answering tasks such as visual question answering have become apparent. Against this backdrop, MLLM (Multi-modal Large Language Model), which integrates multimodal data into a unified model, has emerged, and MLLM-based agents have also shown great potential in reasoning for multimodal question answering tasks.
[0003] However, existing agents often suffer from illusions during the reasoning process of multimodal question answering tasks due to problems such as their autoregressive training mechanism and homogeneous training data. These illusions include a significant mismatch between the agent's output and the question content or logical errors. As a result, existing agents have poor reliability in handling multimodal question answering tasks, which hinders their practical application. Summary of the Invention
[0004] The problem addressed by this invention is how to improve the reliability of intelligent agents in handling multimodal question-answering tasks.
[0005] To address the above problems, this invention provides a dialectical knowledge transfer learning method, comprising the following steps:
[0006] Step 1: Construct a task database, which includes multiple multimodal question-answering tasks;
[0007] Step 2: Utilize multiple preset first intelligent agents to make multi-angle decisions for the unprocessed multimodal question-answering task. When the decision information output by each first intelligent agent meets the preset confidence conditions, integrate each decision information into dialectical knowledge. The multi-angle decision includes independent decision and / or interactive decision.
[0008] Step 3: Return to Step 2 until all the multimodal question-answering tasks in the task database are traversed to obtain multiple dialectical knowledge points and construct a dialectical knowledge base;
[0009] Step 4: Train the preset second intelligent agent based on the dialectical knowledge base to obtain the target intelligent agent.
[0010] Optionally, the decision information includes the reasoning result; step two includes:
[0011] Each of the first intelligent agents makes independent decisions for the unprocessed multimodal question-answering task, thereby obtaining each decision information;
[0012] Determine whether each of the decision information satisfies the preset confidence condition, wherein the preset confidence condition includes that the reasoning result output by each of the first agents in the most recent time is consistent;
[0013] If so, then each piece of decision information will be integrated into the dialectical knowledge.
[0014] If not, then reference decision information is obtained based on the decision information most recently output by each of the first agents, and each of the first agents is instructed to make the interactive decision based on the reference decision information and then output the decision information, and the step of determining whether each decision information satisfies the preset information condition is returned.
[0015] Optionally, the decision information further includes a reasoning process; the preset information conditions further include:
[0016] The reasoning process corresponding to each piece of decision information is matched with the reasoning result.
[0017] Optionally, step two further includes:
[0018] When the number of interactive decisions exceeds a preset threshold, each decision information is integrated into the dialectical knowledge.
[0019] Optionally, after constructing the dialectical knowledge base, the following steps are also included:
[0020] Extract summary information corresponding to each dialectical knowledge in the dialectical knowledge base, wherein the summary information includes the final reasoning result of the last output of each first agent;
[0021] The dialectical knowledge is filtered based on the summarized information and the preset filtering strategy.
[0022] Optionally, filtering the dialectical knowledge based on the summarized information and a preset filtering strategy includes:
[0023] When each of the final reasoning results in the summary information is inconsistent, the dialectical knowledge corresponding to the summary information is discarded; and / or,
[0024] When each of the final inference results in the summary information is consistent, and the number of the final inference results is different from the preset number of inference results, the dialectical knowledge corresponding to the summary information is discarded; and / or,
[0025] When each of the final inference results in the summary information is consistent, and the preset inference result set does not contain the final inference result, the dialectical knowledge corresponding to the summary information is removed.
[0026] Optionally, step four includes:
[0027] The decision information that matches the final reasoning result is extracted from each of the dialectical knowledge to obtain a dialectical chain, and the dialectical chain is associated with the multimodal question-answering task corresponding to the dialectical knowledge to obtain a training database;
[0028] The second agent is trained using the training database to obtain the target agent.
[0029] The task database in this invention includes multiple multimodal question-answering tasks. The resulting dialectical knowledge base enhances the ability of the target agent to handle complex multimodal question-answering tasks. During the process of multiple pre-defined first agents making multi-perspective decisions on multimodal question-answering tasks, when the decision information output by each first agent meets pre-set confidence conditions, each decision is integrated into dialectical knowledge. This ensures the reliability of the dialectical knowledge and allows it to encompass the different "viewpoints" (i.e., decision information) held by different first agents when facing multimodal question-answering tasks, ensuring the comprehensiveness of the dialectical knowledge. Multi-perspective decision-making includes independent decision-making and / or interactive decision-making. When multiple first agents make independent decisions that meet pre-set confidence conditions, each decision is integrated into dialectical knowledge, ensuring the reliability of the decision information while improving the efficiency of acquiring dialectical knowledge. Furthermore, interactive decision-making among multiple agents in this invention facilitates mutual correction of errors occurring during the reasoning process. Therefore, when multiple first agents are directly instructed to make interactive decisions, or when independent decisions by multiple first agents fail to meet pre-set confidence conditions, instructing multiple first agents to make interactive decisions helps improve the accuracy of the output decision information of each first agent. Thus, comprehensive, multi-faceted, and highly reliable dialectical knowledge can be obtained for each multimodal question-answering task. Based on this, multiple dialectical knowledge bases are obtained by traversing the task database, and a pre-defined second agent is trained based on the constructed dialectical knowledge base to obtain the target agent. This is equivalent to endowing the target agent with multi-faceted dialectical knowledge obtained from multiple first agents, enabling a single target agent to possess the "dialectical decision-making" ability that is only possible when multiple first agents collaborate (i.e., independent decision-making and / or interactive decision-making). This effectively reduces the probability of a single target agent experiencing illusions when performing complex multimodal question-answering task reasoning, narrows the gap between the independent reasoning ability of a single target agent and the joint reasoning ability of multiple first agents, and improves the reliability of the target agent in handling multimodal question-answering tasks.
[0030] The present invention also provides a dialectical knowledge transfer learning device, comprising:
[0031] A task building module is used to build a task database, wherein the task database includes multiple multimodal question-answering tasks;
[0032] The knowledge generation module is used to make multi-angle decisions for the unprocessed multimodal question-answering task using multiple preset first intelligent agents. When the decision information output by each first intelligent agent meets the preset information conditions, each decision information is integrated into dialectical knowledge. The multi-angle decision includes independent decision and / or interactive decision.
[0033] The loop traversal module is used to return the steps of using multiple first agents to make multi-angle decisions for the unprocessed multimodal question answering tasks until all the multimodal question answering tasks in the task database are traversed, multiple dialectical knowledge is obtained, and a dialectical knowledge base is constructed.
[0034] The transfer learning module is used to train a pre-defined second agent based on the dialectical knowledge base to obtain the target agent.
[0035] The dialectical knowledge transfer learning device and the dialectical knowledge transfer learning method provided by this invention have essentially the same advantages over the prior art, and will not be repeated here.
[0036] The present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to implement the dialectical knowledge transfer learning method as described above when the computer program is executed.
[0037] The electronic device provided by this invention has essentially the same advantages as the dialectical knowledge transfer learning method compared to existing technologies, and will not be elaborated further here.
[0038] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the dialectical knowledge transfer learning method described above.
[0039] The computer-readable storage medium provided by this invention has essentially the same advantages as the dialectical knowledge transfer learning method compared to the prior art, and will not be repeated here. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the dialectical knowledge transfer learning method according to an embodiment of the present invention.
[0041] Figure 2 This is a flowchart illustrating another embodiment of the dialectical knowledge transfer learning method of the present invention.
[0042] Figure 3 This is a schematic diagram of the structure of the dialectical knowledge transfer learning device according to an embodiment of the present invention;
[0043] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0044] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0045] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0046] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0047] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0048] like Figure 1 As shown, this invention provides a dialectical knowledge transfer learning method, comprising the following steps:
[0049] Step 1: Build a task database, which includes multiple multimodal question-answering tasks.
[0050] Specifically, the multimodal question-answering task referred to in this embodiment may include a visual question-answering task. Visual question answering is an artificial intelligence task that combines computer vision and natural language processing techniques, aiming to enable an intelligent agent to understand and answer image-related questions. In a visual question-answering task, given an image and a natural language question related to the image, the intelligent agent needs to understand the content of the image and make a decision, answering the question in text form. In this embodiment, the multimodal question-answering task may include preset materials and a task description for the preset materials. For example, the preset material in a multimodal question-answering task may be an image, and the corresponding task description may be: "Which of the following products do you think this image belongs to the same category as: product A, product B, product C, or product D?"
[0051] Step 2: Utilize multiple pre-set first agents to make multi-angle decisions for unprocessed multimodal question-answering tasks. When the decision information output by each first agent meets the pre-set confidence conditions, integrate each decision information into dialectical knowledge. Among them, multi-angle decision-making includes independent decision-making and / or interactive decision-making.
[0052] Specifically, the intelligent agents referred to in this embodiment can be constructed based on multimodal large models (such as GPT-4Vision or Gemini). The multiple pre-defined first intelligent agents referred to in this embodiment can be constructed based on the same multimodal large model or on different types of multimodal large models. Intelligent agents based on multimodal large models can typically comprehensively process multiple types of data (such as text, images, and sound), and can analyze image or video content and combine it with relevant text descriptions to achieve accurate content understanding and other functions. For example, in autonomous vehicles, intelligent agents based on multimodal large models can simultaneously process visual data and voice commands, and then respond to voice commands (such as outputting driving assistance decisions).
[0053] In one embodiment, taking the construction of a first agent based on GPT-4Vision as an example, the construction method of the agent is illustrated: A prompt template can be set on the existing GPT-4Vision multimodal large model to obtain the first agent, thereby constraining the output content and format of each first agent. For example, the prompt template could be: "You are a debater. You need to thoroughly analyze the information presented in the image, answer the questions step by step as accurately as possible, and output your reasoning process and results. You can also debate with other debaters, improving each other's reasoning processes based on your historical output and the viewpoints of other debaters. During the debate, you can challenge other debaters' positions, defend your own position, or modify your own position, and output your reasoning process and results."
[0054] Specifically, in this embodiment, the pre-set confidence condition indicates that the decision information output by each first agent has a high degree of credibility. For example, to address the problem of agents frequently experiencing hallucinations when processing complex tasks, the pre-set confidence condition can be set as follows: the reasoning process in each decision information matches the reasoning result (which can be judged using a pre-set third agent).
[0055] In one embodiment, each first agent can make decisions on unprocessed multimodal question-answering tasks in a multi-task database, and then output corresponding decision information. It should be understood that even when making decisions on the same multimodal question-answering task, the decision information output by each first agent may differ. The multi-perspective decision-making in this embodiment includes independent decision-making and / or interactive decision-making. Independent decision-making means that only the multimodal question-answering task is input into each first agent, and the decisions between multiple first agents do not affect each other. Interactive decision-making means that each first agent will refer to the decision information output by other first agents before making its own decision. Therefore, this embodiment utilizes multiple preset first agents to make multi-perspective decisions on unprocessed multimodal question-answering tasks, which can correspond to the following different situations:
[0056] (1) Multiple first agents can be sorted, and each first agent can make interactive decisions in sequence. For example, assuming that this embodiment contains three first agents, denoted as A1, A2 and A3 respectively, A1 can first output decision information R11 for the multimodal question answering task; then input the multimodal question answering task and R11 into A2, and A2 outputs decision information R21; on this basis, input the multimodal question answering task, R11 and R21 into A3, and A3 outputs decision information R31; determine whether each decision information meets the preset confidence conditions. If not, continue to input the multimodal question answering task, R11, R21 and R31 into A1, and A1 outputs R12. Repeat this process of interactive decision-making between A1, A2 and A3 until the preset confidence conditions are met. Integrate each decision information output by A1, A2 and A3 to obtain dialectical knowledge, which corresponds to the case where multiple first agents only made interactive decisions.
[0057] (2) Alternatively, multiple first agents can make independent decisions (for ease of understanding and description, this is referred to as the first round of decision-making in this embodiment). If the preset confidence conditions are met after the first round of decision-making, each decision information is integrated into dialectical knowledge, which corresponds to the case where multiple first agents have only made independent decisions.
[0058] (3) If the preset confidence conditions are not met after the first round of decision-making, multiple first agents can be used for interactive decision-making (for ease of understanding and expression, this is referred to as the second round of decision-making). If the preset confidence conditions are met after the second round of decision-making, the decision information output by each first agent in each round of decision-making is integrated into dialectical knowledge; if the preset confidence conditions are still not met after the second round of decision-making, multiple first agents continue to be used for interactive decision-making, and so on, until the preset confidence conditions are met. At this point, the decision information output by each first agent in each round of decision-making is integrated into dialectical knowledge, corresponding to the cases where multiple first agents have made independent and interactive decisions.
[0059] In this embodiment, multiple pre-defined first agents are used to make multi-faceted decisions for unprocessed multimodal question-answering tasks. When the decision information output by each first agent meets the pre-set confidence conditions, each decision information is integrated into dialectical knowledge. This ensures the reliability of the dialectical knowledge and allows it to encompass the different "viewpoints" (i.e., decision information) held by different first agents when facing multimodal question-answering tasks, thus ensuring the comprehensiveness of the dialectical knowledge. Multi-faceted decision-making includes independent decision-making and / or interactive decision-making. When the pre-set confidence conditions are met after multiple first agents make independent decisions, each decision information is integrated into dialectical knowledge, ensuring the reliability of the decision information while improving the efficiency of acquiring dialectical knowledge. Interactive decision-making among multiple agents helps to correct errors that occur during the reasoning process. Therefore, when multiple first agents are directly instructed to make interactive decisions, or when the pre-set confidence conditions cannot be met after multiple first agents make independent decisions, interactive decision-making is used to improve the accuracy of the decision information output by each first agent. Therefore, comprehensive, multi-faceted, and highly reliable dialectical knowledge can be obtained for each multimodal question-answering task.
[0060] Step 3: Return to Step 2 until all multimodal question-answering tasks in the task database are traversed, obtain multiple dialectical knowledge points, and construct a dialectical knowledge base.
[0061] Specifically, in step two, multi-faceted decision-making is performed on unprocessed multimodal question-answering tasks in the task database to obtain corresponding dialectical knowledge. Step two can then be repeated, selecting any unprocessed multimodal question-answering task from the task database again for multi-faceted decision-making and obtaining corresponding dialectical knowledge, until all multimodal question-answering tasks in the task database are traversed, thus obtaining the dialectical knowledge corresponding to each multimodal question-answering task in the task database. Preferably, after obtaining the corresponding dialectical knowledge for a specific multimodal question-answering task through multi-faceted decision-making, that multimodal question-answering task can be deleted from the task database, which helps improve the efficiency of constructing the dialectical knowledge base.
[0062] Step 4: Train the pre-set second intelligent agent based on the dialectical knowledge base to obtain the target intelligent agent.
[0063] Specifically, to improve the reasoning efficiency of the target agent and reduce its deployment difficulty, the second agent in this embodiment is preferably constructed using a multimodal small model. The second agent built on the multimodal small model can be trained using a dialectical knowledge base. This is equivalent to assigning the dialectical knowledge obtained from the first agent built on multiple multimodal large models to the second agent built on the multimodal small model, thus obtaining the target agent. This enables a lightweight target agent to possess the "dialectical decision-making" ability that is only possible when multiple first agents collaborate (i.e., independent decision-making and / or interactive decision-making).
[0064] Optionally, taking a multimodal question answering task as an example of a visual question answering task, the dialectical knowledge transfer learning method provided in this embodiment is introduced. The method includes:
[0065] Step 1: Construct a task database, which includes multiple visual question answering tasks;
[0066] Step 2: Utilize multiple preset first intelligent agents to make multi-angle decisions for the unprocessed visual question-answering task. When the decision information output by each first intelligent agent meets the preset information conditions, integrate each decision information into dialectical knowledge. The multi-angle decision includes independent decision and / or interactive decision.
[0067] Step 3: Return to Step 2 until all the visual question-answering tasks in the task database are traversed to obtain multiple dialectical knowledge points and construct a dialectical knowledge base;
[0068] Step 4: Train the preset second intelligent agent based on the dialectical knowledge base to obtain the target intelligent agent.
[0069] In this embodiment, the task database includes multiple multimodal question-answering tasks. The resulting dialectical knowledge base is beneficial for improving the target agent's ability to handle complex multimodal question-answering tasks. During the process of multiple preset first agents making multi-angle decisions on unprocessed multimodal question-answering tasks, when the decision information output by each first agent meets preset confidence conditions, each decision information is integrated into dialectical knowledge. This ensures the reliability of the dialectical knowledge and allows it to encompass the different "viewpoints" (i.e., decision information) held by different first agents when facing multimodal question-answering tasks, ensuring the comprehensiveness of the dialectical knowledge. Multi-angle decision-making includes independent decision-making and / or interactive decision-making. When multiple first agents make independent decisions and the preset confidence conditions are met, each decision information is integrated into dialectical knowledge. This ensures the reliability of the decision information while improving the efficiency of acquiring dialectical knowledge. Furthermore, interactive decision-making among multiple agents in this embodiment helps to correct errors that occur during the reasoning process. Therefore, when multiple first agents are directly instructed to make interactive decisions, or when independent decisions by multiple first agents fail to meet pre-set confidence conditions, instructing multiple first agents to make interactive decisions helps improve the accuracy of the output decision information of each first agent. Thus, comprehensive, multi-faceted, and highly reliable dialectical knowledge can be obtained for each multimodal question-answering task. Based on this, multiple dialectical knowledge bases are obtained by traversing the task database, and a pre-defined second agent is trained based on the constructed dialectical knowledge base to obtain the target agent. This is equivalent to endowing the target agent with the multi-faceted dialectical knowledge obtained from multiple first agents, enabling a single target agent to possess the "dialectical decision-making" ability that is only possible when multiple first agents collaborate (i.e., independent decision-making and / or interactive decision-making). This effectively reduces the probability of a single target agent experiencing illusions when performing complex multimodal question-answering task reasoning, narrows the gap between the independent reasoning ability of a single target agent and the joint reasoning ability of multiple first agents, and improves the reliability of the target agent in handling multimodal question-answering tasks.
[0070] Optionally, the decision information includes the reasoning result; step two includes:
[0071] Each first agent makes independent decisions for unprocessed multimodal question-answering tasks, obtaining each decision information;
[0072] Determine whether each decision information satisfies the preset confidence conditions, wherein the preset confidence conditions include that the reasoning result of the most recent output of each first agent is consistent;
[0073] If so, then each piece of decision information will be integrated into dialectical knowledge;
[0074] If not, then reference decision information is obtained based on the decision information most recently output by each first agent, and each first agent is instructed to make interactive decisions based on the reference decision information and output decision information, and then return to the step of judging whether each decision information meets the preset information conditions.
[0075] Specifically, in this embodiment, a pre-defined third agent can be used to determine whether the most recent inference results output by each first agent are consistent. Compared to simply using string matching to extract inference results to determine consistency, this method helps ensure the accuracy of the judgment results. The construction method of the third agent is basically the same as that of the first agent, with only the prompt template settings differing, and will not be elaborated further here.
[0076] In one embodiment, assuming there are three first agents, A1, A2, and A3 respectively, after making independent decisions, the decision information output by the three first agents corresponds to R11, R21, and R31 respectively. It is then determined whether the reasoning results corresponding to R11, R21, and R31 are consistent. If they are inconsistent, reference decision information can be obtained based on the most recent output decision information of each first agent, which is equivalent to using R11, R21, and R31 as reference decision information. After this, the reference decision information can be input into A1, A2, and A3 respectively, causing them to output their interactive decision information: R12, R22, and R32, and the process returns to the step of determining whether the decision information meets the preset confidence conditions.
[0077] Preferably, to reduce the computational burden on the first agent and improve the efficiency of interactive decision-making, in the above process, when the inference results corresponding to R11, R21, and R31 are inconsistent, R11, R21, and R31 can be used as reference decision information for A1, causing A1 to output the interactive decision information R12; then, R21, R31, and R12 are used as reference decision information for A2, causing A2 to output the interactive decision information R22; finally, R31, R12, and R22 are used as reference decision information for A3, causing A3 to output the interactive decision information R32, and returning to the step of judging whether the decision information meets the preset confidence conditions. In this way, the reference decision information can always be composed of the decision information most recently output by each first agent.
[0078] In this embodiment, each first agent first makes independent decisions for an unprocessed multimodal question-answering task, obtaining multiple decision pieces of information. This is equivalent to each first agent expressing its own "viewpoint" in its decision regarding the multimodal question-answering task. Based on this, if the reasoning results obtained by each first agent after making independent decisions are consistent, it indicates that the currently obtained decision information has high reliability, and dialectical knowledge can be derived from it. If the reasoning results obtained by each first agent after making independent decisions are inconsistent, it indicates that there is a high probability of erroneous decision information. In this case, reference decision information is obtained based on the most recent output decision information of each first agent, and each first agent interacts and makes decisions based on the reference decision information. This realizes information interaction among multiple first agents, allowing them to correct each other and promote in-depth analysis of various decision information by each first agent, thereby generating more comprehensive and objective decision information.
[0079] Optionally, the decision information also includes the reasoning process; the pre-set information conditions also include:
[0080] The reasoning process and reasoning result are matched for each piece of decision information.
[0081] Specifically, in this embodiment, the decision information may include the reasoning process and reasoning result corresponding to the first intelligent agent making a decision on a multimodal question answering task. For example, when the multimodal question answering task is to ask a question about an opinion or a solution for a certain image, the corresponding reasoning process may be image content analysis and task analysis, and the corresponding reasoning result may be an opinion description or a solution description.
[0082] In this embodiment, a pre-defined third agent can be used to determine whether the reasoning process and reasoning result corresponding to each decision information match. Compared with the approach that only cares about whether the reasoning result is consistent, this embodiment also considers whether the reasoning process and reasoning result corresponding to each decision information match. This helps to ensure the accuracy of the reasoning process, thereby ensuring the reliability of decision information and dialectical knowledge, and effectively reducing the probability of reasoning logic errors and other problems occurring in the actual application of the target agent.
[0083] Optionally, step two also includes:
[0084] When the number of interactive decisions exceeds a preset threshold, each decision piece of information is integrated into dialectical knowledge.
[0085] In this embodiment, even if multiple first agents conduct multiple rounds of interactive decision-making, they may still not be able to reach a consensus (i.e., the reasoning results are inconsistent), which can easily lead to a lot of wasted computing power. In this embodiment, the number of interactive decision-makings between multiple first agents is constrained by judging whether the number of interactive decision-makings between multiple first agents exceeds a preset threshold (e.g., 5 times). When the number of interactive decision-makings exceeds the preset threshold, each decision information is integrated into dialectical knowledge, which helps to improve the efficiency of acquiring dialectical knowledge.
[0086] Optionally, after constructing the dialectical knowledge base, it also includes:
[0087] Extract the summary information corresponding to each dialectical knowledge in the dialectical knowledge base. The summary information includes the final reasoning result of the last output of each first agent.
[0088] The dialectical knowledge is filtered based on the summarized information and the preset filtering strategy.
[0089] In this embodiment, a pre-defined fourth agent can be used to extract summary information corresponding to each piece of dialectical knowledge. The construction method of the fourth agent is basically the same as that of the first agent, and will not be described in detail here. This embodiment first summarizes the dialectical knowledge to obtain concise summary information, and then filters the dialectical knowledge based on the summary information and a pre-defined filtering strategy. Compared with the method of directly filtering based on complex dialectical knowledge and a pre-defined filtering strategy, this method helps to reduce data pressure and improve filtering efficiency. At the same time, the summary information includes the final reasoning result output by each first agent, which is equivalent to the "viewpoint" finally output by each first agent after mutual correction and in-depth reasoning. This is beneficial for providing a reliable reference for the filtering of dialectical knowledge.
[0090] Optionally, dialectical knowledge can be filtered based on summarized information and preset filtering strategies, including:
[0091] When the final reasoning results in the summary information are inconsistent, the dialectical knowledge corresponding to the summary information is removed; and / or,
[0092] When each final reasoning result in the summary information is consistent, and the number of final reasoning results is different from the preset number of reasoning results, the dialectical knowledge corresponding to the summary information is discarded; and / or,
[0093] When each final reasoning result in the summary information is consistent, and the preset reasoning result set does not contain the final reasoning result, the dialectical knowledge corresponding to the summary information is removed.
[0094] Specifically, in this embodiment, the final inference result refers to the inference result contained in the last output decision information of each first agent when making multi-angle decisions for a multimodal question-answering task. When the number of interactive decisions exceeds a preset threshold, each decision information is integrated to generate dialectical information. However, at this time, the final inference result output by each first agent may be inconsistent, and the reliability of such data is poor. Before training the second agent based on the dialectical knowledge base, the dialectical knowledge corresponding to the inconsistent summary information of each final inference result is removed, which helps to improve the reliability of the inference result of the target agent in practical applications.
[0095] In one embodiment, the number of constraints on the reasoning results corresponding to each multimodal question-answering task can be preset, i.e., the preset number of reasoning results. For example, assuming the question description of the multimodal question-answering task is: "Please select an option related to artificial intelligence from the following images", the preset number of reasoning results is 1. If multiple first agents have the same final reasoning results, but the number of reasoning results exceeds 1, it indicates that the dialectical knowledge corresponding to the summarized information is unreliable and needs to be removed.
[0096] In one embodiment, a preset inference result set is pre-defined for each multimodal question-answering task. For example, suppose the question description for the multimodal question-answering task is: "Which of the following should the calculation result corresponding to the image be: Option A1, Option B2, Option C3, Option D4?" If multiple first agents arrive at the same inference result, but their corresponding inference result is 5, which is not among the options, or if the final inference result is: "The calculation result is 5, but there is no corresponding option, so the closest option D is selected," it indicates that dialectical knowledge is unreliable and needs to be eliminated.
[0097] In this embodiment, after constructing the dialectical knowledge base, the dialectical knowledge is summarized to obtain concise summary information. Without needing to pre-set correct reference reasoning results for each multimodal question-answering task, obviously unreliable dialectical knowledge can be eliminated based on the summary information and corresponding preset filtering strategies, thereby further improving the accuracy of dialectical knowledge and ensuring the reliability of the target intelligent agent.
[0098] Optionally, step four includes:
[0099] The decision information that matches the final reasoning result is extracted from each dialectical knowledge to obtain the dialectical chain, and the dialectical chain is associated with the multimodal question answering task corresponding to the dialectical knowledge to obtain the training database.
[0100] The second agent is trained using the training database to obtain the target agent.
[0101] Specifically, even if the final reasoning results contained in each piece of dialectical knowledge in the filtered dialectical knowledge base are consistent, in the case of multiple rounds of interactive decision-making, the decision information of the first agents undergoes a process of mutual correction. In some rounds, the reasoning results in the decision information output by some first agents may be inconsistent with the final reasoning result. This type of decision information is highly likely due to reasoning errors by the first agents. In this embodiment, when constructing the dialectical chain, a preset fifth agent can be used to extract decision information that matches the final reasoning result, condensing highly reliable dialectical knowledge into a dialectical chain. This helps to prevent the second agent from learning incorrect decision information, which could lead to incorrect decisions in subsequent complex task reasoning. The construction method of the preset fifth agent is basically the same as that of the first agent, and will not be described in detail here.
[0102] In one embodiment, after extracting the dialectical chain, the dialectical chain is associated with the multimodal question-answering task corresponding to dialectical knowledge to obtain a training database. The training database is then used to train a second agent to obtain the target agent. The training database can be represented as:
[0103]
[0104] Where D represents the training dataset, This represents a single sample data point; N represents the number of sample data points; img i Let x represent the preset materials (such as images) in the i-th multimodal question answering task. i This represents the question description in the i-th multimodal question answering task; This represents the reasoning process corresponding to the decision information in the dialectical chain associated with the i-th multimodal question-answering task; This represents the reasoning result corresponding to the decision information in the dialectical chain associated with the i-th multimodal question-answering task.
[0105] During the training of the second agent based on the training database, img can be used i and x i As input, As a pseudo-label, the output of the second agent continuously approaches... Maximize the generation of the second agent The possibility that the second agent, given a preset material (img), i and problem description x i In the case of reasoning process and reasoning result The goodness of fit can be expressed as:
[0106]
[0107] in, This means that for each sample data The corresponding expected value; Indicates that in a given img i x i Generate under the conditions The probability; Φ represents the parameters for constructing the multimodal small model of the second agent.
[0108] Therefore, by maximizing the above degree of fit, the reasoning performance of the second agent can be optimized, enabling the second agent to make more accurate decisions in multimodal question answering tasks, thereby obtaining the target agent.
[0109] In this embodiment, the dialectical chain represents a condensed result of dialectical knowledge for multimodal question-answering tasks, encompassing the reasoning process and results, and possessing comprehensiveness, consistency, and reliability. A training database constructed based on the dialectical chain is used to train a second agent, enabling the second agent to continuously learn the "dialectical decision-making" capabilities that can only be achieved through collaboration among multiple first agents. This facilitates the transfer learning of dialectical knowledge, effectively alleviating the illusion problem faced by a single target agent in handling complex multimodal question-answering tasks and improving the target agent's reasoning ability.
[0110] For example, such as Figure 2 As shown, a specific embodiment will now be used to further introduce the transfer learning method of dialectical knowledge. The method includes:
[0111] Construct a task database, which includes multiple multimodal question-answering tasks;
[0112] Multiple first-level agents are used to make independent decisions for unprocessed multimodal question-answering tasks, and each decision is output.
[0113] Judgment steps: Use a pre-set third-party intelligent agent to determine whether each decision information meets the pre-set information conditions, or whether the number of interactive decisions exceeds a pre-set threshold.
[0114] If not, then obtain reference decision information based on the decision information most recently output by each first agent, and have each first agent make interactive decisions based on the reference decision information and output decision information, and return to the judgment step;
[0115] If so, then each piece of decision information will be integrated into dialectical knowledge;
[0116] Determine whether to iterate through all multimodal tasks in the task database;
[0117] If not, return to the steps of using multiple first agents to make independent decisions for the multimodal question-answering task and outputting the decision information for each decision;
[0118] If so, then a dialectical knowledge base will be constructed based on dialectical knowledge;
[0119] The system uses a pre-defined fourth agent to extract summary information corresponding to each dialectical knowledge in the dialectical knowledge base, and filters the dialectical knowledge based on the summary information and a pre-defined filtering strategy.
[0120] By using a pre-defined fifth intelligent agent, decision information that matches the final reasoning result is extracted from each dialectical knowledge, thus obtaining the dialectical chain;
[0121] By associating the dialectical chain with the multimodal question-answering task corresponding to dialectical knowledge, a training database is obtained. The training database is then used to train a pre-set second agent to obtain the target agent.
[0122] like Figure 3 As shown, another embodiment of the present invention also provides a dialectical knowledge transfer learning device, comprising:
[0123] The task building module is used to build a task database, which includes multiple multimodal question-answering tasks.
[0124] The knowledge generation module is used to make multi-angle decisions for unprocessed multimodal question-answering tasks by utilizing multiple preset first intelligent agents. When the decision information output by each first intelligent agent meets the preset confidence conditions, each decision information is integrated into dialectical knowledge. Among them, multi-angle decision-making includes independent decision-making and / or interactive decision-making.
[0125] The loop traversal module is used to return the steps of making multi-angle decisions using multiple first agents for unprocessed multimodal question-answering tasks until all multimodal question-answering tasks in the task database are traversed, multiple dialectical knowledge is obtained, and a dialectical knowledge base is constructed.
[0126] The transfer learning module is used to train a pre-defined second agent based on a dialectical knowledge base to obtain the target agent.
[0127] The dialectical knowledge transfer learning device and the dialectical knowledge transfer learning method provided in this embodiment can produce basically the same technical effects, and will not be described again here.
[0128] Another embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to implement the dialectical knowledge transfer learning method as described above when the computer program is executed.
[0129] like Figure 4 As shown, the electronic device provided in this embodiment includes: at least one processor 401, and a memory 402 connected to at least one processor 401. This embodiment does not limit the specific connection medium between the processor 401 and the memory 402. Figure 4 The example shown is the connection between processor 401 and memory 402 via bus 400. Bus 400 is... Figure 4 The connections between other components are indicated by thick solid lines and are for illustrative purposes only, not as limiting information. The bus 400 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 4 The bus is represented by a single thick solid line, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 401 can also be called a controller; there is no restriction on the name.
[0130] In this embodiment, the memory 402 stores instructions executable by at least one processor 401. By executing the instructions stored in the memory 402, the at least one processor 401 can perform the dialectical knowledge transfer learning method discussed above. The processor 401 can implement... Figure 3 The functions of each module in the device shown.
[0131] The processor 401 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 402 and calling data stored in memory 402, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0132] In one possible design, processor 401 may include one or more processing units. Processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 401. In some embodiments, processor 401 and memory 402 may be implemented on the same chip; in some embodiments, they may also be implemented separately on separate chips.
[0133] Processor 401 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the dialectical knowledge transfer learning method disclosed in this embodiment can be directly manifested as execution by a hardware processor, or as execution by a combination of hardware and software modules within the processor.
[0134] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 402 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 402 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In this embodiment, memory 402 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0135] By designing and programming the processor 401, the code corresponding to the dialectical knowledge transfer learning method described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during operation. Figure 1 The steps of the dialectical knowledge transfer learning method in the illustrated embodiment are as follows. How to design and program the processor 401 is a technique well-known to those skilled in the art and will not be described further here.
[0136] The electronic device provided in this embodiment and the dialectical knowledge transfer learning method can produce basically the same technical effects, and will not be described again here.
[0137] Another embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the dialectical knowledge transfer learning method described above.
[0138] In some possible implementations, various aspects of the dialectical knowledge transfer learning method provided by the present invention can also be implemented in the form of a program product, which includes program code. When the program product is run on a device, the program code is used to cause the control device to perform the steps in the dialectical knowledge transfer learning method according to various exemplary embodiments of the present invention described above.
[0139] Those skilled in the art will understand that this embodiment can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0143] The computer-readable storage medium provided in this embodiment and the dialectical knowledge transfer learning method can produce essentially the same technical effects, which will not be repeated here.
[0144] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A dialectical knowledge transfer learning method, characterized in that, Includes the following steps: Step 1: Construct a task database, which includes multiple multimodal question-answering tasks; Step 2: Utilize multiple preset first intelligent agents to make multi-angle decisions for the unprocessed multimodal question-answering task. When the decision information output by each first intelligent agent meets the preset confidence conditions, integrate each decision information into dialectical knowledge. The multi-angle decision includes independent decision and / or interactive decision. Step 3: Return to Step 2 until all the multimodal question-answering tasks in the task database are traversed to obtain multiple dialectical knowledge points and construct a dialectical knowledge base; Step 4: Train the pre-defined second intelligent agent based on the dialectical knowledge base to obtain the target intelligent agent; Step four includes: The decision information that matches the final reasoning result is extracted from each of the dialectical knowledge to obtain the dialectical chain, and the dialectical chain is associated with the multimodal question answering task corresponding to the dialectical knowledge to obtain the training database. The second agent is trained using the training database to obtain the target agent; Following the construction of the dialectical knowledge base, the following is also included: Extract summary information corresponding to each dialectical knowledge in the dialectical knowledge base, wherein the summary information includes the final reasoning result of the last output of each first agent; The dialectical knowledge is filtered based on the summarized information and the preset filtering strategy.
2. The dialectical knowledge transfer learning method according to claim 1, characterized in that, The decision information includes the reasoning results; step two includes: Each of the first intelligent agents makes independent decisions for the unprocessed multimodal question-answering task, thereby obtaining each decision information; Determine whether each of the decision information satisfies the preset confidence condition, wherein the preset confidence condition includes that the reasoning result output by each of the first agents in the most recent time is consistent; If so, then each piece of decision information will be integrated into the dialectical knowledge. If not, then reference decision information is obtained based on the decision information most recently output by each of the first agents, and each of the first agents is instructed to make the interactive decision based on the reference decision information and then output the decision information, and the step of determining whether each decision information satisfies the preset information condition is returned.
3. The dialectical knowledge transfer learning method according to claim 2, characterized in that, The decision information also includes a reasoning process; the pre-set information conditions also include: The reasoning process corresponding to each piece of decision information is matched with the reasoning result.
4. The dialectical knowledge transfer learning method according to claim 2, characterized in that, Step two also includes: When the number of interactive decisions exceeds a preset threshold, each decision information is integrated into the dialectical knowledge.
5. The dialectical knowledge transfer learning method according to claim 1, characterized in that, The filtering of dialectical knowledge based on the summarized information and a preset filtering strategy includes: When each of the final reasoning results in the summary information is inconsistent, the dialectical knowledge corresponding to the summary information is discarded; and / or, When each of the final inference results in the summary information is consistent, and the number of the final inference results is different from the preset number of inference results, the dialectical knowledge corresponding to the summary information is discarded; and / or, When each of the final inference results in the summary information is consistent, and the preset inference result set does not contain the final inference result, the dialectical knowledge corresponding to the summary information is removed.
6. A dialectical knowledge transfer learning device, characterized in that, include: A task building module is used to build a task database, wherein the task database includes multiple multimodal question-answering tasks; The knowledge generation module is used to make multi-angle decisions for the unprocessed multimodal question-answering task using multiple preset first intelligent agents. When the decision information output by each first intelligent agent meets the preset information conditions, each decision information is integrated into dialectical knowledge. The multi-angle decision includes independent decision and / or interactive decision. The loop traversal module is used to return the steps of using multiple first agents to make multi-angle decisions for the unprocessed multimodal question answering tasks until all the multimodal question answering tasks in the task database are traversed, multiple dialectical knowledge is obtained, and a dialectical knowledge base is constructed. The transfer learning module is used to train a preset second agent based on the dialectical knowledge base to obtain a target agent. The loop traversal module is also used to extract summary information corresponding to each dialectical knowledge in the dialectical knowledge base, wherein the summary information includes the final reasoning result output by each first agent at the last time. The dialectical knowledge is filtered based on the summarized information and the preset filtering strategy; The transfer learning module is also used to extract the decision information that matches the final reasoning result from each of the dialectical knowledge, to obtain a dialectical chain, and associate the dialectical chain with the multimodal question answering task corresponding to the dialectical knowledge to obtain a training database; The second agent is trained using the training database to obtain the target agent.
7. An electronic device, characterized in that, It includes a memory and a processor, the memory being used to store a computer program, and the processor being used to implement the dialectical knowledge transfer learning method as described in any one of claims 1 to 5 when the computer program is executed.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the dialectical knowledge transfer learning method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Solution plan generation method and device, equipment and storage medium
CN117407514A
Multiagent debate
US20240104125A1