Multi-agent cooperation method for resisting cross-modal cognitive interference based on cognitive chain

By separating the perception and reasoning processes through a multi-agent collaborative approach, this method addresses the cross-modal cognitive interference problem in multimodal large language models, improves question-answering accuracy, adapts to different hardware resources, and is applicable to complex multimodal scenarios.

CN121882292APending Publication Date: 2026-04-17CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
Filing Date
2025-11-17
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing multimodal large language models suffer from poor question-answering accuracy in the face of cross-modal cognitive interference. Traditional methods have failed to effectively address the fusion of perception and reasoning modes, leading to hallucination phenomena generated by the models.

Method used

A multi-agent collaborative approach based on cognitive chain to combat cross-modal cognitive interference is adopted, deploying four types of agents: two multimodal large language models, a lightweight evaluation agent, a dedicated reasoning agent, and a decision agent. Through parallel processing and cognitive load judgment, the perception and reasoning processes are separated, multi-perspective observation results are generated and verified, and finally, a conclusion is output.

Benefits of technology

It effectively mitigates cross-modal cognitive interference, improves question-answering accuracy, ensures the purity and accuracy of the reasoning process, adapts to different hardware resources, and is applicable to complex multimodal scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882292A_ABST
    Figure CN121882292A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent cooperation method for resisting cross-modal cognitive interference based on a cognitive chain, and the method comprises the following steps: firstly deploying four types of agents, carrying out the parallel processing of two multi-modal large language model agents, and generating two groups of different observation results; the observation result is evaluated, cognitive load judgment is conducted on the observation result to obtain an evaluation result, if the load is low, the decision-making stage of the fifth step is directly entered, and if the load is high, the reasoning stage of the fourth step is entered; the method comprises the following steps: firstly, taking a question, an option and an observation result sum as input of a reasoning agent, and generating a detailed rational basis after the reasoning agent is processed; and the decision-making agent firstly receives the information, then the decision-making agent performs verification again according to the information and outputs a final conclusion. According to the method, multi-agent cooperation is achieved, cross-modal interference is relieved, and the multi-modal question and answer precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an improvement of a multi-agent collaboration method, belonging to the field of multimodal large language model collaboration, and particularly to a multi-agent collaboration method based on cognitive chain to counteract cross-modal cognitive interference. Background Technology

[0002] In recent years, multimodal large language models (MLLMs) have successfully extended the reasoning capabilities of large language models (LLMs) to multimodal scenarios by integrating training data from different modalities and fine-tuning the instructions. Among these capabilities, the thought chain (CoT) is a key ability that guides the model to generate a series of reasoning steps, helping to solve complex mathematical calculations and logical deductions. After fine-tuning with multimodal question-answering datasets containing reasoning processes, these models have acquired a certain degree of cross-modal complex reasoning ability. However, human cross-modal information perception and information-based reasoning rely on fundamentally different cognitive modes. Current advanced multimodal large language models mostly adopt an end-to-end architecture, fusing these two fundamentally different cognitive modes together. This fusion inevitably leads to cross-modal cognitive interference, causing hallucinations when the model generates thought chains due to the mutual influence between visual perception and reasoning abilities. Specifically, this manifests as biases in non-textual modal information perception or hallucinations during textual reasoning.

[0003] Traditional research has focused on two main approaches to address this issue, but neither has addressed the core problem. First, it focuses on achieving multimodal semantic alignment, using technical means to convert information from different modalities into features and map them to the same vector space. This only solves the fundamental problem of "multimodal information understanding" but fails to address the essential difference between "perception and reasoning requiring different cognitive modes," thus failing to alleviate cross-modal cognitive interference. Second, it attempts to further fine-tune multimodal large language models based on high-quality, human-selected multimodal question-and-answer datasets. While this can improve the model's complex reasoning capabilities to some extent, it still doesn't break the end-to-end architectural limitations of "perception and reasoning fusion," and the problem of cross-modal cognitive interference persists.

[0004] Chinese patent application CN202510374596.X, filed on March 27, 2025, discloses a method for optimizing agent strategies based on multimodal fusion, an agent, and an electronic device. The method includes: selecting the most suitable fusion strategy for a target task from a pre-built fusion strategy library; wherein the fusion strategy library contains multiple predefined fusion strategies; fusing the multimodal data required for the target task using the most suitable fusion strategy; and, when executing the target task, learning based on the fused multimodal data corresponding to the target task to generate task decisions. This scheme dynamically adjusts task decisions based on interactions with other agents, thereby achieving fast and efficient learning in complex scenarios and significantly improving the accuracy and adaptability of agent decisions. However, it does not solve the problem of cross-modal cognitive interference in multimodal large language models, leading to poor question-answering accuracy.

[0005] The information disclosed in this background section is intended only to enhance the understanding of the overall background of this patent application and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this invention is to overcome the problem of poor question-answering accuracy caused by cross-modal cognitive interference in multimodal large language models in the prior art. It provides a multi-agent collaborative method based on cognitive chain to mitigate cross-modal cognitive interference in multimodal large language models and improve question-answering accuracy.

[0007] To achieve the above objectives, the technical solution of the present invention is: a multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference, wherein the multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference includes the following steps:

[0008] The first step is to deploy four types of intelligent agents, including:

[0009] Two multimodal large language model agents A lightweight evaluation agent finely tuned by injecting target capabilities. A dedicated reasoning agent and a decision-making agent ;

[0010] The second step involves two multimodal large language model agents processing the data in parallel, simultaneously extracting images. Problems with input and options Based on relevant key visual features and contextual information, two different sets of observations are generated. and ;

[0011] The third step is to first evaluate the lightweight intelligent agent. As a metacognitive controller, then evaluate the observations. and Cognitive load is assessed based on the observations to obtain evaluation results. ,like If the cognitive load is low, proceed directly to the fifth step, the decision-making stage. If the cognitive load is high, then proceed to the fourth step, the reasoning stage;

[0012] Step 4: Start with a question and options Observation results and As input to the reasoning agent, the agent processes the data to generate detailed rational arguments. ;

[0013] Step 5: The decision-making agent first receives the original image. ,question Options Detailed and rational basis and evaluation results Then, the decision-making agent re-verifies based on the above information and outputs the final conclusion. .

[0014] In the second step, the multimodal large language model intelligent agent Fine-tuned using different multimodal instruction datasets.

[0015] The two multimodal large language models can intelligently capture different emphases in the image.

[0016] In the second step, observe the results. and This can be expressed by the following formula:

[0017] ;

[0018] in, Represents image input, It is a prompt used to guide the agent to perform only observation and description, and the generated result and This forms a multi-perspective, textual basis for visual perception.

[0019] In the third step, the evaluation results This can be expressed by the following formula:

[0020] ;

[0021] in, These are prompt words used to guide the assessment agent in making cognitive judgments.

[0022] In the fourth step, detailed rational evidence is generated. Specifically:

[0023] ;

[0024] in, It is a prompt used to guide the reasoning agent in generating structured reasoning steps.

[0025] like The reasoning load is high, based on the rational basis for the generation of the reasoning agent. ;like If the cognitive load is low, it is based on the integrated results. .

[0026] The final conclusion Based on the evaluation results It is generated in two cases:

[0027] like Under high cognitive load:

[0028] ;

[0029] like Under low cognitive load:

[0030] ;

[0031] in, This represents the integrated result, that is, the two observations. and Spelling together simple prompts. This indicates a prompt that guides the decision-making agent to make a final judgment.

[0032] The dedicated reasoning agent performs in-depth, long-chain logical deduction, causal analysis, and critical reasoning.

[0033] The target capability injection strategy for evaluating the agent uses a training dataset based on automated fine-grained annotation.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] 1. In this invention, a multi-agent collaborative method based on cognitive chain to combat cross-modal cognitive interference involves two multimodal large language model agents processing in parallel, synchronously extracting key visual features and contextual information related to the image, input question, and options, generating two different sets of observation results. The decision agent first receives the original image, question, options, detailed rational basis, and evaluation results, and then re-verifies based on the above information and outputs the final conclusion. The advantages of this design are as follows:

[0036] First, by clearly separating perception and reasoning through the cognitive chain, we can avoid mutual contamination between the two heterogeneous cognitive modes, reduce model perception bias and reasoning errors, exclude visual input in the reasoning stage, and perform pure logical deduction based solely on textual observation results to ensure the purity and accuracy of the reasoning process.

[0037] Secondly, the evaluation agent dynamically judges the cognitive load of the task. Low-load tasks skip resource-intensive reasoning, while high-load tasks accurately activate deep reasoning to achieve efficient resource utilization. The evaluation agent is built based on a lightweight model with fine-tuning, without relying on a large model, further controlling the computational cost.

[0038] Thirdly, in the perception stage, multiple agents process in parallel, generating multi-perspective observation results due to differences in knowledge background and data, providing rich textual clues for subsequent stages and enhancing the comprehensiveness of information; in the decision-making stage, the observation results or reasoning conclusions are reintegrated to complete the final verification, ensuring the consistency between the results and the questions and images.

[0039] Therefore, this invention alleviates cross-modal cognitive interference in multimodal large language models and improves question-answering accuracy.

[0040] 2. In this invention, a multi-agent collaborative method based on cognitive chains to combat cross-modal cognitive interference eliminates the need for manually labeled datasets. Training data is constructed through an automated, fine-grained labeling process, reducing data dependency costs. This method not only performs excellently in static image question answering tasks but also accurately captures semantic coherence in video question answering tasks with temporal correlations, adapting to complex multimodal scenarios. Therefore, this invention exhibits strong generalization ability and is applicable to a wide range of scenarios.

[0041] 3. In this invention, a multi-agent collaborative method based on cognitive chains to counter cross-modal cognitive interference designs three differentiated agent combinations, covering lightweight, medium, and high-performance requirements at different levels, allowing for flexible selection based on actual hardware resources. Therefore, this invention adapts to different hardware conditions, balancing performance and deployment flexibility. Attached Figure Description

[0042] Figure 1 This is a flowchart of the present invention.

[0043] Figure 2 This is a schematic diagram of collaboration in this invention.

[0044] Figure 3 This is a performance comparison chart of different MADC combinations in this invention. Detailed Implementation

[0045] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] See Figures 1 to 3 A multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference includes the following steps:

[0047] The first step is to deploy four types of intelligent agents, including:

[0048] Two multimodal large language model agents A lightweight evaluation agent finely tuned by injecting target capabilities. A dedicated reasoning agent and a decision-making agent ;

[0049] The second step involves two multimodal large language model agents processing the data in parallel, simultaneously extracting images. Problems with input and options Based on relevant key visual features and contextual information, two different sets of observations are generated. and ;

[0050] The third step is to first evaluate the lightweight intelligent agent. As a metacognitive controller, then evaluate the observations. and Cognitive load is assessed based on the observations to obtain evaluation results. ,like If the cognitive load is low, proceed directly to the fifth step, the decision-making stage. If the cognitive load is high, then proceed to the fourth step, the reasoning stage;

[0051] Step 4: Start with a question and options Observation results and As input to the reasoning agent, the agent processes the data to generate detailed rational arguments. ;

[0052] Step 5: The decision-making agent first receives the original image. ,question Options Detailed and rational basis and evaluation results Then, the decision-making agent re-verifies based on the above information and outputs the final conclusion. .

[0053] In the second step, the multimodal large language model intelligent agent Fine-tuned using different multimodal instruction datasets.

[0054] The two multimodal large language models can intelligently capture different emphases in the image.

[0055] In the second step, observe the results. and This can be expressed by the following formula:

[0056] ;

[0057] in, Represents image input, It is a prompt used to guide the agent to perform only observation and description, and the generated result and This forms a multi-perspective, textual basis for visual perception.

[0058] In the third step, the evaluation results This can be expressed by the following formula:

[0059] ;

[0060] in, These are prompt words used to guide the assessment agent in making cognitive judgments.

[0061] In the fourth step, detailed rational evidence is generated. Specifically:

[0062] ;

[0063] in, It is a prompt used to guide the reasoning agent in generating structured reasoning steps.

[0064] like The reasoning load is high, based on the rational basis for the generation of the reasoning agent. ;like If the cognitive load is low, it is based on the integrated results. .

[0065] The final conclusion Based on the evaluation results It is generated in two cases:

[0066] like Under high cognitive load:

[0067] ;

[0068] like Under low cognitive load:

[0069] ;

[0070] in, This represents the integrated result, that is, the two observations. and Spelling together simple prompts. This indicates a prompt that guides the decision-making agent to make a final judgment.

[0071] The dedicated reasoning agent performs in-depth, long-chain logical deduction, causal analysis, and critical reasoning.

[0072] The target capability injection strategy for evaluating the agent uses a training dataset based on automated fine-grained annotation.

[0073] The supplementary technical features of this design are as follows:

[0074] By physically and logically separating the time-consuming reasoning process from the susceptible visual perception process, this invention effectively mitigates cross-modal cognitive interference, ensures the purity and accuracy of the reasoning process, and thus solves the problem of mutual contamination between reasoning and perception in traditional end-to-end models.

[0075] Example 1:

[0076] A multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference includes the following steps:

[0077] The first step is to deploy four types of intelligent agents, including:

[0078] Two multimodal large language model agents A lightweight evaluation agent finely tuned by injecting target capabilities. A dedicated reasoning agent and a decision-making agent ;

[0079] The second step involves two multimodal large language model agents processing the data in parallel, simultaneously extracting images. Problems with input and options Based on relevant key visual features and contextual information, two different sets of observations are generated. and ;

[0080] The third step is to first evaluate the lightweight intelligent agent. As a metacognitive controller, then evaluate the observations. and Cognitive load is assessed based on the observations to obtain evaluation results. ,like If the cognitive load is low, proceed directly to the fifth step, the decision-making stage. If the cognitive load is high, then proceed to the fourth step, the reasoning stage;

[0081] Step 4: Start with a question and options Observation results and As input to the reasoning agent, the agent processes the data to generate detailed rational arguments. ;

[0082] Step 5: The decision-making agent first receives the original image. ,question Options Detailed and rational basis and evaluation results Then, the decision-making agent re-verifies based on the above information and outputs the final conclusion. .

[0083] Example 2:

[0084] Example 2 is basically the same as Example 1, except that:

[0085] Multimodal large language model intelligent agent After fine-tuning with different multimodal instruction datasets, due to the differences in knowledge background, pre-training corpus and fine-tuning data among different agents, they tend to capture different emphases in the images, which greatly enhances the diversity and robustness of the observation results and provides rich textual clues for subsequent reasoning.

[0086] Observation results and This can be expressed by the following formula:

[0087] ;

[0088] in, Represents image input, It is a prompt used to guide the agent to perform only observation and description, and the generated result and This forms a multi-perspective, textual basis for visual perception.

[0089] Example 3:

[0090] Example 3 is basically the same as Example 1, except that:

[0091] A dedicated evaluation agent is introduced during the evaluation phase. As a metacognitive controller, its core function is to dynamically determine the cognitive load of the current task, that is, based on the input question. and the observations provided by the perception stage and The goal is to determine whether these observations alone are sufficient to arrive at a high-confidence, accurate conclusion. If the evaluation agent can answer based solely on intuitive observations, it can skip the resource-intensive deep reasoning stage and proceed directly to the decision-making stage. To achieve this efficient and accurate evaluation of cognitive resource allocation, the evaluation agent is deployed using a target capability injection strategy: a dedicated cognitive load judgment dataset is constructed through an automated, fine-grained annotation process, and this dataset is used to fine-tune a lightweight large language model.

[0092] The automated fine-grained annotation process is as follows:

[0093] Rule-based annotation strategy: M 3 The CoT dataset includes subject category classifications for the samples, and those that do not require reasoning are marked as "no reasoning required" because solving problems in fields such as linguistics and social sciences often relies on retrieving declarative knowledge from perceptual output rather than complex reasoning.

[0094] Automatic labeling pipeline: includes three steps: perceptual simulation, automatic judgment, and label assignment, and processes Infinity-MM subsets.

[0095] Perceptual simulation: Instantiate Qwen2VL-7B-Instruct to process each multimodal input to generate descriptive textual observations. This simulates the exact inputs the evaluation agent would receive during operation.

[0096] Automatic judgment: Using GLM-4-flash as an expert annotator, it receives the original question and simulated perceptual observations, and determines whether this plain text information is sufficient to generate a correct, high-confidence answer.

[0097] Label assignment: If a question is considered to be answerable with high confidence and the generated answer is correct, the sample is labeled "no reasoning required" and all other instances are labeled "reasoning required".

[0098] This fine-tuning enables the evaluation agent to accurately classify tasks as "low cognitive load" or "high cognitive load," thereby achieving dynamic allocation of cognitive resources. This effectively reduces the risk of hallucinations in large models and significantly optimizes inference costs. The evaluation results... This can be expressed by the following formula:

[0099] ;

[0100] in, These are prompt words used to guide the assessment agent in making cognitive judgments.

[0101] Example 4:

[0102] Example 4 is basically the same as Example 1, except that:

[0103] The reasoning phase only occurs when evaluating the results. The activity is activated when the task is deemed "high cognitive load." The core of this stage is to completely exclude visual input and use only questions. and options and textual observations and As input, a dedicated reasoning agent performs in-depth, long-chain logical deduction, causal analysis, and critical reasoning to generate detailed rational evidence. By explicitly separating the time-consuming reasoning process from the susceptible visual perception process both physically and logically, this invention effectively mitigates cross-modal cognitive interference, ensuring the purity and accuracy of the reasoning process. This solves the problem of mutual contamination between reasoning and perception in traditional end-to-end models. This process is represented by the following formula, generating detailed rationale. Specifically:

[0104] ;

[0105] in, It is a prompt used to guide the reasoning agent in generating structured reasoning steps.

[0106] Example 5:

[0107] Example 5 is basically the same as Example 1, except that:

[0108] During the decision-making phase, regardless of whether the reasoning phase is activated, all tasks will ultimately reach their final results in this phase, determined by a single decision-making agent. To integrate all the information and make a final judgment, the agent receives the original visual image. ,question Options , and the textual basis generated in the previous stage, if The reasoning load is high, based on the rational basis for the generation of the reasoning agent. ;like If the cognitive load is low, it is based on the integrated results. The decision-making agent, acting as the ultimate arbitrator, re-examines visual information and the reasoning / observation chain, considering options... The final conclusion is drawn from this. This mechanism ensures that even after textual reasoning, the final decision remains consistent with the original image and question, and is ultimately validated. The final conclusion of this stage... Based on the evaluation results It is generated in two cases:

[0109] like Under high cognitive load:

[0110] ;

[0111] like Under low cognitive load:

[0112] ;

[0113] in, This represents the integrated result, that is, the two observations. and Spelling together simple prompts. This indicates a prompt that guides the decision-making agent to make a final judgment.

[0114] Example 6:

[0115] Example 6 is basically the same as Example 1, except that:

[0116] Dataset: The training dataset based on automated fine-grained annotation was used to evaluate the target capability injection strategy of the agent, which includes M 3 The dataset used is the training set of the CoT dataset and a subset of the Infinity-MM dataset. A rule-based annotation strategy was applied to the M3CoT training set, and an automatic annotation pipeline consisting of three steps—perceptual simulation, automatic judgment, and label assignment—was applied to the Infinity-MM subset.

[0117] Model Setup: This invention employs three different multi-agent cooperative combinations based on cognitive chains to counteract cross-modal interference, abbreviated as MADC-small, MADC-medium, and MADC-proper. Specifically, MADC-small uses the models Qwen2VL-7B-Instruct and InternVL2-8B in the perception phase, Qwen2.5-14B-Instruct in the inference phase, and Qwen2VL-7B-Instruct in the decision-making phase. MADC-medium uses Qwen2VL-72B-Instruct-AWQ-Int4 and InternVL2-40B in the perception phase, Qwen2.5-32B-Instruct in the inference phase, and Qwen2VL-72B-Instruct-AWQ-Int4 in the decision-making phase. MADC-proper uses GPT4o and Claude3.5-Sonnet in the perception phase, GPT4o in the inference phase, and GPT4o in the decision-making phase. All multi-agent collaborative combinations used a finely tuned Qwen2.5-0.5B-Instruct as the cognitive evaluation agent.

[0118] The test results on the image-based intelligent question answering benchmarks ScienceQA, M3CoT, and MMMU-Pro are shown in Tables 1, 2, and 3, respectively.

[0119] Table 1 is as follows:

[0120]

[0121] Table 2 is as follows:

[0122]

[0123] Table 3 is as follows:

[0124]

[0125] As shown in Tables 1, 2, and 3, the multi-agent collaborative method proposed in this invention demonstrates a significant performance improvement over directly using multimodal large models or other structured reasoning methods in image-based intelligent question answering benchmark tests. This advantage is particularly pronounced in problems involving mathematical calculations, logical reasoning, and other issues requiring multi-step reasoning and abstract thinking. This result indicates that the multi-agent collaborative mechanism can effectively alleviate common information interference and semantic bias problems in cross-modal information fusion, thereby reducing perceptual errors and reasoning illusions, proving the advantages of this invention. Furthermore, the multi-stage collaborative framework based on cognitive chains proposed in this invention fully simulates human cognitive and decision-making processes. Through staged information extraction, cognitive load assessment, and logical verification, it achieves complex multimodal reasoning that more closely resembles human thinking.

[0126] The test results on the Video-MME benchmark for intelligent video question answering are shown in Table 4:

[0127]

[0128] As shown in Table 4, the multi-agent collaborative method proposed in this invention achieves superior performance in video intelligent question answering benchmark tests compared to directly using a large multimodal model. Compared to static images, video data not only contains more complex visual information but also significant temporal correlation features. The dynamic changes and semantic connections between different frames enable video content to jointly construct a higher-level semantic expression in both spatial and temporal dimensions. The advantages of the multi-agent collaborative method in this task fully demonstrate that it can more accurately capture the logical evolution and semantic coherence of events in videos when processing time-dependent multimodal information. Through the division of labor and collaborative reasoning among agents, this method effectively improves the model's understanding depth and reasoning ability of video semantic structure, thereby achieving efficient modeling and accurate question answering of complex temporal information. This result further verifies the wide applicability and technical advantages of this invention in dynamic multimodal cognitive scenarios.

[0129] Example 7:

[0130] Example 7 is basically the same as Example 1, except that:

[0131] A multi-agent collaborative device based on cognitive chain-based adversarial cross-modal interference includes a memory and a processor;

[0132] The memory is used to store computer program code and transmit the computer program code to the processor;

[0133] The processor is configured to execute, according to instructions in the computer program code, a multi-agent cooperative method based on cognitive chain to counteract cross-modal cognitive interference as described above.

[0134] Example 8:

[0135] Example 8 is basically the same as Example 1, except that:

[0136] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-agent cooperative method based on cognitive chain-based adversarial cross-modal cognitive interference as described above.

[0137] The above description is only a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. Any equivalent modifications or changes made by those skilled in the art based on the content disclosed in the present invention should be included within the scope of protection set forth in the claims.

Claims

1. A multi-agent cooperative method based on cognitive chain to counteract cross-modal cognitive interference, characterized in that: The multi-agent cooperation method based on cognitive chain to counter cross-modal cognitive interference includes the following steps: The first step is to deploy four types of intelligent agents, including: Two multimodal large language model agents A lightweight evaluation agent finely tuned by injecting target capabilities. A dedicated reasoning agent and a decision-making agent ; The second step involves parallel processing of images by two multimodal large language model agents, simultaneously pre-processing them. Problems with input and options Based on relevant key visual features and contextual information, two different sets of observations are generated. and ; The third step is to first evaluate the lightweight intelligent agent. As a metacognitive controller, then evaluate the observations. and Cognitive load is assessed based on the observations to obtain evaluation results. ,like If the cognitive load is low, proceed directly to the fifth step, the decision-making stage. If the cognitive load is high, then proceed to the fourth step, the reasoning stage; Step 4: Start with a question and options Observation results and As input to the reasoning agent, the reasoning agent processes the data to generate detailed rational arguments. ; Step 5: The decision-making agent first receives the original image. ,question Options Detailed and rational basis and evaluation results Then, the decision-making agent re-verifies based on the above information and outputs the final conclusion. .

2. The multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference as described in claim 1, characterized in that: In the second step, the multimodal large language model intelligent agent Fine-tuned using different multimodal instruction datasets.

3. The multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference as described in claim 2, characterized in that: The two multimodal large language models can intelligently capture different emphases in the image.

4. The multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference as described in claim 3, characterized in that: In the second step, observe the results. and This can be expressed by the following formula: ; in, Represents image input, It is a prompt used to guide the agent to perform only observation and description, and the generated result and This forms a multi-perspective, textual basis for visual perception.

5. A multi-agent cooperation method based on cognitive chain to counter cross-modal cognitive interference as described in claim 1, characterized in that: In the third step, the evaluation results This can be expressed by the following formula: ; in, These are prompt words used to guide the assessment agent in making cognitive judgments.

6. A multi-agent cooperation method based on cognitive chain to counter cross-modal cognitive interference as described in claim 1, characterized in that: In the fourth step, detailed rational evidence is generated. Specifically: ; in, It is a prompt used to guide the reasoning agent in generating structured reasoning steps.

7. A multi-agent cooperative method based on cognitive chain to counter cross-modal cognitive interference as described in claim 1, characterized in that: like The reasoning load is high, based on the rational basis for the generation of the reasoning agent. ;like If the cognitive load is low, it is based on the integrated results. .

8. A multi-agent cooperation method based on cognitive chain to counter cross-modal cognitive interference as described in claim 1, characterized in that: The final conclusion Based on the evaluation results It is generated in two cases: like Under high cognitive load: ; like Under low cognitive load: ; in, This represents the integrated observation, that is, the two observations. and Spelling together simple prompts. This indicates a prompt that guides the decision-making agent to make a final judgment.

9. A multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference as described in claim 1, characterized in that: The dedicated reasoning agent performs in-depth, long-chain logical deduction, causal analysis, and critical reasoning.

10. A multi-agent cooperation method based on cognitive chain to counteract cross-modal cognitive interference according to claim 1, characterized in that: The target capability injection strategy for evaluating the agent uses a training dataset based on automated fine-grained annotation.

Citation Information

Patent Citations

  • Intelligent agent strategy optimization method based on multi-modal fusion, intelligent agent and electronic equipment

    CN120234759A