Intelligent agent reasoning method and system based on reliability and storage medium
By dividing the reasoning process of a multimodal agent into a clue acquisition and integration stage, and using high-entropy lexical units to evaluate reliability and perform weighted voting, the problem of unreliable answers by multimodal agents in visual question answering tasks is solved, thereby improving reasoning accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-12
AI Technical Summary
In visual question answering tasks, multimodal agents suffer from unreliable answers due to erroneous cue acquisition and integration. Existing methods ignore the reliability of the cue acquisition and integration process, resulting in low reasoning accuracy.
The reasoning process of a multimodal agent is divided into two stages: clue acquisition and clue integration. Reliability is evaluated by calculating high-entropy lexical units, a filtering threshold is set to filter unreliable paths, and a weighted voting based on reliability is performed to output the answer.
It effectively avoids error accumulation and significantly improves the reasoning accuracy and robustness of multimodal agents.
Smart Images

Figure CN122021918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal agent reasoning, and more specifically to a reliability-based agent reasoning method, system, and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, multimodal agents have been widely applied in visual question answering scenarios. Among them, multimodal thought chains have become one of the important paradigms for multimodal agents to solve complex visual question answering tasks, effectively improving the reasoning ability of multimodal agents. Specifically, the multimodal agent first searches for task-related visual information, then generates text tools to call instructions to obtain visual cues, subsequently uses the collected visual cues as intermediate steps in the thought chain, and finally generates a text thought chain to integrate the cues and deduce the answer. This reasoning process, which interweaves text and images, is the multimodal thought chain.
[0003] However, in practical applications, due to the limited perception and understanding capabilities of multimodal agents, errors often occur in the clue acquisition and integration processes of multimodal thought chains. These errors accumulate gradually in subsequent thinking, ultimately leading to unreliable output answers. To address the problem of unreliable agent reasoning answers, researchers have recently attempted to select the final answer by generating multiple reasoning paths and conducting a simple vote on the reasoning results. In short, existing methods typically assume that all reasoning answers have the same reliability and therefore consider the answer that appears most frequently as the correct answer, thus using a "majority rule" voting method to determine the final reasoning answer. However, existing methods ignore the fact that erroneous clue acquisition and integration processes in multimodal thought chains can lead to unreliable reasoning answers, directly incorporating unreliable reasoning answers into the final voting process, which severely reduces the reasoning accuracy of multimodal agents.
[0004] In summary, multimodal agent reasoning faces the problem of error accumulation caused by the erroneous clue acquisition process and the erroneous clue integration process. Summary of the Invention
[0005] To address the aforementioned shortcomings in existing technologies, this invention proposes a reliability-based agent reasoning method, system, and storage medium, which alleviates the error accumulation problem in multimodal agent reasoning tasks and improves the reasoning accuracy of multimodal agents.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This solution provides a reliability-based agent reasoning method, system, and storage medium, including the following steps: S1: Utilize a multimodal intelligent agent to process text and image inputs, generating multiple reasoning paths. Each reasoning path includes a clue acquisition stage and a clue integration stage. S2: Conduct a reliability assessment of the clue acquisition and clue integration stages of each reasoning path; S3: Based on the reliability of the clue acquisition stage and the clue integration stage, the reasoning path is screened to filter out unreliable reasoning paths; S4: Perform a reliability-based weighted vote on the selected reasoning paths and output the reasoning answer.
[0007] The beneficial effects of this invention are as follows: the reasoning path of a multimodal agent is divided into two stages: clue acquisition and clue integration. Reliability assessment and screening are performed on the two stages to avoid the accumulation of errors in the reasoning process. At the same time, the accuracy of multimodal agent reasoning is significantly improved by using a reliability-based weighted voting mechanism.
[0008] Further, step S1 includes the following sub-steps: Furthermore, the input text problem is... The corresponding image is The multimodal agent first generates... There are several reasoning paths, each including a clue acquisition stage and a clue integration stage. Specifically, for a given reasoning path... The multimodal agent first generates textual clues during the clue acquisition phase. Then, following the tool invocation instructions in the cue acquisition phase, visual cues are obtained from the image. in, For reasoning path Visual cues, A tool for obtaining visual cues. This is the clue acquisition phase. It involves acquiring visual cues. Afterwards, the multimodal agent generates textual clues for integration. And extract the answer from it. in, For reasoning path The answer, For the answer extraction function, This is the clue integration stage.
[0009] The beneficial effect of the above-mentioned further scheme is that it divides the thinking process of multimodal thinking chain into a clue acquisition stage and a clue integration stage, which can be used to study reliability assessment methods.
[0010] Furthermore, step S2 includes the following sub-steps: S21: Calculate the entropy of each word in the reasoning path; S22: Calculate the reliability of the clue acquisition stage and the clue integration stage separately using high-entropy lexical units.
[0011] Furthermore, the entropy of each word in step S21 is calculated as follows: in, For the first The entropy of each word element The th prediction probability vector A probability value, It is the set of all words.
[0012] Furthermore, in step S22, the reliability calculation method for the clue acquisition stage and the clue integration stage is as follows: in, This is the clue-gathering stage. This is the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, It is the number of high-entropy lexical units. For the clue acquisition stage China has the highest Entropy-based lexical index, For the clue integration stage China has the highest Entropy-based lexical index, For the first The entropy of each word element.
[0013] The beneficial effect of the above-mentioned further scheme is that by using high-entropy lexical units to estimate the reliability of the clue acquisition stage and the clue integration stage, unreliable clue acquisition stage and clue integration stage can be identified.
[0014] Furthermore, step S3 includes the following sub-steps: S31: Determine the filtering threshold based on the reliability of the clue acquisition stage and the clue integration stage; S32: Adaptively filter unreliable inference paths based on thresholds.
[0015] Furthermore, the filtering threshold calculation method for the clue acquisition stage and the clue integration stage in step S31 is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, To make the first from childhood to adulthood The percentage reliability was determined as the screening threshold. This is the set of inference paths used to estimate the screening threshold.
[0016] Furthermore, in step S32, the filtering method for unreliable inference paths is as follows: in, This is the filtered set of reasoning paths. For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage The reliability.
[0017] The beneficial effect of the above-mentioned further scheme is that reliable reasoning paths are screened by the reliability of the clue acquisition stage and the clue integration stage, and unreliable answers are eliminated before the final answer vote.
[0018] Further, step S4 involves performing a reliability-based weighted vote on the selected reasoning paths, ultimately outputting the voted reasoning answer. Specifically, the reliability-based weighted voting method is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the filtered set of reasoning paths. It is the answer based on a weighted vote. As potential answer candidates, For reasoning path Confidence weights It is 1 if and only if the condition is met. For reasoning path The answer, For the clue acquisition stage Reliability, For the clue integration stage Reliability, Take the larger of the two values. It is an exponential function. It is an absolute value function.
[0019] The beneficial effect of the above further scheme is that it estimates the confidence of the reasoning path after screening, performs weighted voting based on the estimated confidence, and obtains a reliable answer.
[0020] A computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method described above.
[0021] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the above-described method.
[0022] The beneficial effects of this invention are that it filters out erroneous reasoning paths from multiple reasoning paths generated by a multimodal agent, performs weighted voting on the filtered reasoning paths, and finally outputs a reliable reasoning answer. This alleviates the problem of error accumulation caused by the erroneous clue acquisition process and the erroneous clue integration process, and improves the reasoning accuracy of the multimodal agent. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. It should be understood that the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating the steps of a reliability-based intelligent agent reasoning method, system, and storage medium provided in an embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0026] like Figure 1As shown, this invention provides a reliability-based agent reasoning method, system, and storage medium, the implementation of which is as follows: S1: Utilize a multimodal intelligent agent to process text and image inputs, generating multiple reasoning paths. Each reasoning path includes a clue acquisition stage and a clue integration stage. S2: Conduct a reliability assessment of the clue acquisition and clue integration stages of each reasoning path; S3: Based on the reliability of the clue acquisition stage and the clue integration stage, the reasoning path is screened to filter out unreliable reasoning paths; S4: Perform a reliability-based weighted vote on the selected reasoning paths and output the reasoning answer.
[0027] In this embodiment, the input text problem is: The corresponding image is The multimodal agent first generates... There are several reasoning paths, each including a clue acquisition stage and a clue integration stage. Specifically, for a given reasoning path... The multimodal agent first generates textual clues during the clue acquisition phase. Then, following the tool invocation instructions in the cue acquisition phase, visual cues are obtained from the image. in, For reasoning path Visual cues, A tool for obtaining visual cues. This is the clue acquisition phase. It involves acquiring visual cues. Afterwards, the multimodal agent generates textual clues for integration. And extract the answer from it. in, For reasoning path The answer, For the answer extraction function, This is the clue integration stage.
[0028] In this embodiment, step S2 includes the following sub-steps: S21: Calculate the entropy of each word in the reasoning path; S22: Calculate the reliability of the clue acquisition stage and the clue integration stage separately using high-entropy lexical units.
[0029] In this embodiment, the entropy of each word in step S21 is calculated as follows: in, For the first The entropy of each word element The th prediction probability vector A probability value, It is the set of all words.
[0030] In this embodiment, the reliability calculation method for the clue acquisition stage and the clue integration stage in step S22 is as follows: in, This is the clue-gathering stage. This is the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, It is the number of high-entropy lexical units. For the clue acquisition stage China has the highest Entropy-based lexical index, For the clue integration stage China has the highest Entropy-based lexical index, For the first The entropy of each word element.
[0031] In this embodiment, step S3 includes the following sub-steps: S31: Determine the filtering threshold based on the reliability of the clue acquisition stage and the clue integration stage; S32: Adaptively filter unreliable inference paths based on thresholds.
[0032] In this embodiment, the filtering threshold calculation method for the clue acquisition stage and the clue integration stage in step S31 is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, To make the first from childhood to adulthood The percentage reliability was determined as the screening threshold. This is the set of inference paths used to estimate the screening threshold.
[0033] In this embodiment, the filtering method for unreliable inference paths in step S32 is as follows: in, This is the filtered set of reasoning paths. For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage The reliability.
[0034] In this embodiment, step S4 involves performing a reliability-based weighted vote on the filtered reasoning paths, and finally outputting the voted reasoning answer. Specifically, the reliability-based weighted voting method is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the filtered set of reasoning paths. It is the answer based on a weighted vote. As potential answer candidates, For reasoning path Confidence weights It is 1 if and only if the condition is met. For reasoning path The answer, For the clue acquisition stage Reliability, For the clue integration stage Reliability, Take the larger of the two values. It is an exponential function. It is an absolute value function.
[0035] To further verify the effectiveness of the multimodal agent reasoning method provided in this invention, experimental evaluation was conducted using the Vstar-Bench multimodal reasoning dataset. Vstar-Bench is a visual question-answering benchmark for high-resolution, visually detailed scenes, used to test the ability of multimodal models to locate key details and perform reasoning in complex images. In this embodiment, Vstar-Bench consists of 191 multiple-choice questions, including two types of tasks: attribute recognition and spatial relationship reasoning. The attribute recognition task contains 115 samples, requiring the model to identify the color, material, and other attribute information of the target, with 4 possible answers; the spatial relationship reasoning task contains 76 samples, requiring the model to determine the relative spatial relationship between two targets, with 2 possible answers.
[0036] In this embodiment, to evaluate the effectiveness of the multimodal agent reasoning method in offline reasoning scenarios, this invention utilizes the most mainstream multimodal agent model, Qwen3-VL, for experiments. Under the same input, data partitioning, and evaluation rules, it is compared with several cutting-edge test-time scaling strategies. The comparison methods include: the multimodal agent base model (Qwen3-VL), the self-consistent sampling method (SC), the iterative consistency method (CISC), the confidence-based method (Deepconf), and the self-verification method (Self-Cer).
[0037] In terms of evaluation metrics, this embodiment uses inference accuracy (ACC) as a quantitative indicator. ACC measures the hit rate of the multimodal agent model in reasoning multiple-choice questions; a higher ACC value indicates more accurate reasoning results and better model performance. ACC values were statistically analyzed on attribute recognition and spatial relationship reasoning tasks to reflect the effectiveness of this scheme in multimodal agent reasoning tasks. The experimental results are shown in Table 1.
[0038] Table 1. Inference performance of different test-time scaling methods on high-resolution visual question answering datasets.
[0039] As shown in Table 1, the method provided by this invention has a significant improvement in inference accuracy compared to other methods, demonstrating better robustness and generalization ability.
[0040] This embodiment provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the steps of the above-described multimodal intelligent agent reasoning method based on reliability assessment.
[0041] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0042] The memory includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or D-interface display memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device. Of course, the memory may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is often used to store the operating system and various application software installed on the computer device, such as the code of the multimodal intelligent agent inference method based on reliability assessment. In addition, the memory can also be used to temporarily store inference results that have been output or will be output.
[0043] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data, for example, to run the program code of the multimodal agent inference method based on reliability assessment.
[0044] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to perform the steps of the multimodal intelligent agent reasoning method based on reliability assessment described above.
[0045] The computer-readable storage medium stores an interface display program that can be executed by at least one processor to cause the at least one processor to perform the steps of the multimodal agent reasoning method based on reliability assessment as described above.
[0046] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the multimodal intelligent agent reasoning method based on reliability assessment described in the embodiments of this application.
[0047] This invention provides a reliability-based intelligent agent reasoning method, system, and storage medium, which alleviates the problem of error accumulation caused by the error clue acquisition stage and error clue integration stage in complex visual question answering tasks.
Claims
1. A reliability-based agent reasoning method, system, and storage medium, characterized in that, The method includes: S1: Utilize a multimodal intelligent agent to process text and image inputs, generating multiple reasoning paths. Each reasoning path includes a clue acquisition stage and a clue integration stage. S2: Conduct a reliability assessment of the clue acquisition and clue integration stages of each reasoning path; S3: Based on the reliability of the clue acquisition stage and the clue integration stage, the reasoning path is screened to filter out unreliable reasoning paths; S4: Perform a reliability-based weighted vote on the filtered reasoning paths and output the reasoning answer; Step S2 specifically includes: S21: Calculate the entropy of each word in the reasoning path; S22: Calculate the reliability of the clue acquisition stage and the clue integration stage separately using high-entropy lexical units; Step S3 specifically includes: S31: Determine the filtering threshold based on the reliability of the clue acquisition stage and the clue integration stage; S32: Adaptively filter unreliable inference paths based on thresholds.
2. The reliability-based agent reasoning method, system, and storage medium according to claim 1, characterized in that, In step S22, the reliability calculation method for the clue acquisition stage and the clue integration stage is as follows: in, This is the clue-gathering stage. This is the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, It is the number of high-entropy lexical units. For the clue acquisition stage China has the highest Entropy-based lexical index, For the clue integration stage China has the highest Entropy-based lexical index, For the first The entropy of each word element.
3. The reliability-based agent reasoning method, system, and storage medium according to claim 1, characterized in that, The filtering threshold calculation method for the clue acquisition stage and the clue integration stage in step S31 is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage Reliability, To make the first from childhood to adulthood The percentage reliability was determined as the screening threshold. This is the set of inference paths used to estimate the screening threshold.
4. The reliability-based agent reasoning method, system, and storage medium according to claim 1, characterized in that, The method for filtering unreliable inference paths in step S32 is as follows: in, This is the filtered set of reasoning paths. For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the screening threshold for the clue acquisition stage. This is the screening threshold for the clue integration stage. For the clue acquisition stage Reliability, For the clue integration stage The reliability.
5. The reliability-based agent reasoning method, system, and storage medium according to claim 1, characterized in that, In step S4, the filtered reasoning paths undergo a reliability-based weighted vote, and the final reasoning answer is output. Specifically, the reliability-based weighted voting method is as follows: in, For the reasoning path, This is the clue-gathering stage. This is the clue integration stage. This is the filtered set of reasoning paths. It is the answer based on a weighted vote. As potential answer candidates, For reasoning path Confidence weights It is 1 if and only if the condition is met. For reasoning path The answer, For the clue acquisition stage Reliability, For the clue integration stage Reliability, Take the larger of the two values. It is an exponential function. It is an absolute value function.