Visual data reasoning method, visual reasoning model training method, electronic device, storage medium and program product

CN122819482APending Publication Date: 2026-09-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611032294.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

此时,基于频率的简单选择机制不仅无法纠正偏差,反而会将错误答案固化为最终输出,导致推理准确性下降

Benefits of technology

[0010]本说明书实施例至少具有以下有益效果:通过利用多模态大语言模型对视觉数据和查询问题进行多次独立推理生成包含推理过程和回答结果的初始推理假设集合,并提取该集合的统计分布特征作为条件信号与视觉数据、查询问题以及初始推理假设集合一同输入整合推理模型进行联合推理,使得整合推理模型在联合推理过程中能够获得关于初始推理假设集合整体分布特性的信息。由于统计分布特征作为注意力引导信号直接参与整合推理模型对各初始推理假设的融合侧重调整,整合推理模型能够感知初始推理假设集合中各回答结果的分布状态,而非仅依赖出现频率的高低进行简单选择。当多次独立推理中出现高频错误共识时,整合推理模型基于统计分布特征所反映的共识集中程度,能够综合考量各初始推理假设的整体分布状况来确定融合侧重,从而降低了盲从高频错误答案的风险,提升了视觉数据推理的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819482A_ABST
    Figure CN122819482A_ABST
Patent Text Reader

Abstract

The specification provides a visual data reasoning method, a visual reasoning model training method, an electronic device, a storage medium and a program product. The method comprises: acquiring visual data, a query question and a pre-trained multi-modal large language model; using the multi-modal large language model to perform multiple independent reasoning on the visual data and the query question to generate an initial reasoning hypothesis set, wherein each initial reasoning hypothesis contains a reasoning process and an answer result; extracting statistical distribution features of the initial reasoning hypothesis set; using an integrated reasoning model, taking the statistical distribution features as a conditional signal, and jointly reasoning with the visual data, the query question and the initial reasoning hypothesis set to obtain a target reasoning path and a target answer result. Wherein the statistical distribution features are used as attention guiding signals to adjust the fusion emphasis of the integrated reasoning model on each initial reasoning hypothesis. The embodiments of the specification can effectively utilize the statistical distribution characteristics of the reasoning hypothesis set to guide the reasoning decision, and improve the accuracy of visual data reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to one or more embodiments in the field of multimodal artificial intelligence technology, specifically to a visual data reasoning method, a visual reasoning model training method, an electronic device, a storage medium, and a program product. Background Technology

[0002] With the rapid development of multimodal large language models, visual question answering technology, which performs joint reasoning based on visual data and natural language query questions, has been widely used. In this type of technology, the model needs to understand the content of the input visual data and combine it with the query question posed by the user to generate an output that includes the reasoning process and the answer result.

[0003] To improve reasoning accuracy, researchers have proposed a self-consistent reasoning method based on multiple sampling. The core idea is to use a multimodal large language model to perform multiple independent samplings on the same visual data and query question, generating multiple candidate answers. Then, by statistically analyzing the frequency of each candidate answer, the candidate answer with the highest frequency is selected as the final output. This method is based on the fundamental assumption that a correct answer is more likely to be repeated in multiple independent reasoning processes; therefore, the candidate answer with the highest frequency is the most reliable answer.

[0004] However, the aforementioned self-consistent reasoning method has significant limitations in practical applications. This method implicitly assumes that multiple independent reasoning results are unbiased, meaning that the probability of sampling correct and incorrect answers is equal. However, in real-world scenarios, multimodal large language models often suffer from inherent cognitive biases, such as over-reliance on certain visual patterns or systematic misjudgments of specific scene types. When such systematic biases exist, incorrect answers may be repeatedly generated in multiple independent reasoning sessions, forming a high-frequency consensus of errors. In this case, a simple frequency-based selection mechanism not only fails to correct the bias but also solidifies incorrect answers into the final output, leading to a decrease in reasoning accuracy. Summary of the Invention

[0005] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of the embodiments of this specification, a visual data reasoning method is provided, comprising: Acquire visual data, query questions, and pre-trained multimodal large language models; The multimodal large language model is used to perform multiple independent inferences on the visual data and the query question to generate an initial inference hypothesis set, wherein each initial inference hypothesis contains an independent inference process and answer result; Extract the statistical distribution characteristics of the initial inference hypothesis set; Using an integrated reasoning model, the statistical distribution features are used as conditional signals to perform joint reasoning with the visual data, the query question, and the initial set of reasoning hypotheses, to obtain the target reasoning path and target answer output by the integrated reasoning model. The statistical distribution features are used as attention-guiding signals to adjust the fusion emphasis of the integrated inference model on each initial inference hypothesis in the initial inference hypothesis set.

[0006] According to a second aspect of the embodiments of this specification, a method for training a visual reasoning model is provided, comprising: Obtain training samples, which include visual data, query questions, labeled answers, and a set of initial inference hypotheses generated by a pre-trained multimodal large language model. Each initial inference hypothesis contains an inference process and answer result for an independent inference of the visual data and the query question. Extract the statistical distribution characteristics of the initial inference hypothesis set of the sample; Using the integrated reasoning model to be trained, with the statistical distribution features as conditional signals, joint reasoning is performed with the visual data, the query question, and the initial inference hypothesis set of the samples to obtain the target reasoning path and the target answer result; The parameters of the integrated reasoning model are updated based on the target answer result, the target reasoning path, the labeled answer, and the statistical distribution characteristics. The statistical distribution features are used to guide the integrated reasoning model to adjust the fusion emphasis of the initial reasoning assumptions for each sample in joint reasoning.

[0007] According to a third aspect of the embodiments of this specification, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in the first or second aspect above.

[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the method described in the first or second aspect above.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the method described in the first or second aspect above.

[0010] The embodiments of this specification have at least the following beneficial effects: By utilizing a multimodal large language model to perform multiple independent inferences on visual data and query questions, an initial inference hypothesis set containing the inference process and answer results is generated. The statistical distribution characteristics of this set are extracted as conditional signals and input into an integrated inference model along with the visual data, query questions, and the initial inference hypothesis set for joint inference. This allows the integrated inference model to obtain information about the overall distribution characteristics of the initial inference hypothesis set during the joint inference process. Since the statistical distribution characteristics, as attention-guiding signals, directly participate in the integrated inference model's adjustment of the fusion emphasis of each initial inference hypothesis, the integrated inference model can perceive the distribution state of each answer result in the initial inference hypothesis set, rather than simply relying on the frequency of occurrence for selection. When high-frequency erroneous consensus occurs in multiple independent inferences, the integrated inference model, based on the consensus concentration reflected by the statistical distribution characteristics, can comprehensively consider the overall distribution of each initial inference hypothesis to determine the fusion emphasis, thereby reducing the risk of blindly following high-frequency erroneous answers and improving the accuracy of visual data inference. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of the overall architecture of the visual data reasoning system provided in the embodiments of this specification; Figure 2 A paradigm comparison diagram between the conventional solutions provided in the embodiments of this specification and the solutions in this specification; Figure 3 The main flowchart of the visual data reasoning method provided in the embodiments of this specification; Figure 4 This is an overview diagram of the Video-BCI framework provided in the embodiments of this specification; Figure 5 The main flowchart of the method for training a visual reasoning model provided in the embodiments of this specification; Figure 6 A flowchart for calculating the reward signal provided in the embodiments of this specification; Figure 7 The training process comparison graphs provided in the embodiments of this specification; Figure 8 Qualitative case study diagrams provided for embodiments of this specification; Figure 9 This is a block diagram of the visual data reasoning device provided in the embodiments of this specification; Figure 10 This is a block diagram of the device for training visual reasoning models provided in the embodiments of this specification; Figure 11 This is a schematic diagram of the electronic device structure provided in the embodiments of this specification. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions of the embodiments of this specification, the technical solutions of the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification should fall within the protection scope of the embodiments of this specification.

[0013] This specification provides a visual data reasoning method, a visual reasoning model training method, an electronic device, a storage medium, and a program product, which are described in detail below. It should be noted that the order of description of the following embodiments does not constitute a limitation on the preferred order of the embodiments.

[0014] Before introducing the technical solutions of the embodiments of this specification, a detailed analysis is first conducted on the visual data reasoning methods in the prior art and their existing defects, in order to better understand the technical contributions of the embodiments of this specification.

[0015] In existing technologies, self-consistency reasoning based on multiple sampling is a widely adopted visual question answering technique. The basic process of this method is as follows: First, visual data, query questions, and a pre-trained multimodal large language model are acquired. Then, the multimodal large language model is used to perform multiple independent sampling inferences on the same visual data and query questions. Each sampling generates an output containing the reasoning process and the answer result, resulting in multiple outputs forming a set of reasoning hypotheses. Finally, the frequency of each answer result in this set is counted, and the answer result with the highest frequency is selected as the final answer. For example, assuming five independent inferences are performed, with three answers being "knife," one answer being "oven," and one answer being "gas stove," then according to the frequency selection strategy, the final answer is "knife."

[0016] The above method is based on a core assumption: correct answers are more likely to be repeated in multiple independent inferences, therefore the most frequent answer is most likely to be correct. However, this assumption does not always hold true in practical applications, leading to inaccuracies in the method at the following three levels.

[0017] First, frequency selection strategies are susceptible to systematic biases in the model. Multimodal large language models may develop inherent cognitive biases towards certain visual patterns during training. When the model systematically misjudges a particular type of visual content, incorrect answers may be repeatedly generated in multiple independent inferences. For example, if the model has a bias in recognizing the visual features of a "knife" in a kitchen scene, even if the most dangerous object in the visual data is actually a "gas stove," the model may repeatedly answer "knife" in multiple inferences, causing the incorrect answer to dominate the frequency statistics and ultimately be selected as the answer. In this case, the frequency selection mechanism not only fails to correct this bias but also solidifies the error.

[0018] Second, existing methods completely ignore the use of information from the reasoning process. In multiple independent inferences, each inference contains a complete reasoning process (i.e., the deductive chain from visual data to the answer), which contains important information about how the model understands the visual data. Different reasoning processes may go through different chains of visual evidence, and even if they ultimately arrive at the same answer, the reliability of their reasoning paths may vary significantly. For example, for the same answer of "knife," one reasoning process is based on the correct visual evidence that "the knife has a sharp edge," while another reasoning process is based on the irrelevant visual cue that "the knife is in the center of the image," and the reliability of the two is drastically different. However, existing methods only focus on the frequency consistency of the final answer, completely discarding the information from these intermediate reasoning processes.

[0019] Third, existing methods fail to detect cognitive conflict within the set of inference hypotheses. When multiple inferences produce highly consistent answers, it indicates that the model's understanding of the problem is relatively certain; when answers are scattered across different options, it indicates significant cognitive uncertainty in the model. This state of cognitive conflict contains important information about the problem's difficulty and the model's confidence, but existing methods have failed to effectively extract and utilize this information to guide inference decisions. Especially in high-cognitive-conflict scenarios (i.e., when answers are highly scattered), simple frequency selection often fails to identify the correct inferences hidden behind the minority responses.

[0020] In summary, the core flaw of existing technologies lies in their focus on frequency statistics at the level of response results, failure to effectively extract and utilize the statistical distribution characteristics of the set of inference hypotheses, and inability to guide the model to dynamically adjust the fusion strategy based on the distribution status of each inference hypothesis.

[0021] To address the aforementioned shortcomings, this specification provides a visual data reasoning method. By extracting the statistical distribution characteristics of an initial set of reasoning hypotheses and using them as conditional signals in joint reasoning, the model can perceive the cognitive conflict state within the hypothesis set and dynamically adjust the fusion emphasis. The following detailed description is in conjunction with the accompanying drawings.

[0022] First, combine Figure 1 The overall architecture of the visual data reasoning system provided in the embodiments of this specification is described. Figure 1 This is a schematic diagram of the overall architecture of the visual data reasoning system provided in the embodiments of this specification.

[0023] like Figure 1 As shown in the embodiments of this specification, the visual data reasoning system includes two core components: a pre-trained multimodal large language model and an integrated reasoning model. The multimodal large language model is used to perform multiple independent reasoning operations on the input visual data and query question to generate an initial set of reasoning hypotheses. The integrated reasoning model is used to receive the visual data, query question, initial set of reasoning hypotheses and their statistical distribution characteristics, and perform joint reasoning to output the target reasoning path and target answer result.

[0024] In the embodiments described in this specification, the multimodal large language model is a large neural network model capable of processing both visual and textual modal inputs and generating text outputs. This model is obtained through pre-training on large-scale image and text data and possesses powerful visual understanding and text generation capabilities. In visual data reasoning tasks, the multimodal large language model acts as a "sampler," its main function being to generate multiple candidate answers and their corresponding reasoning processes through multiple independent inferences, given visual data and a query question.

[0025] The integrated reasoning model is the core innovative component of the embodiments in this specification. It also adopts the architecture of a multimodal large language model, but is trained with specific optimization objectives. Unlike conventional multimodal large language models, the input to the integrated reasoning model includes not only the original visual data and query question, but also an initial set of reasoning hypotheses generated during the sampling phase, along with the statistical distribution characteristics of that set. The task of the integrated reasoning model is to generate more accurate and reliable target reasoning paths and target answer results based on this rich contextual information during joint reasoning.

[0026] From the perspective of data flow, such as Figure 1As shown, the system's workflow can be divided into two main stages: the inference stage and the training stage. In the inference stage, visual data and the query question are first input into a multimodal large language model. This model performs multiple independent inferences, generating an initial inference hypothesis containing the inference process and the answer result each time. Multiple such hypotheses form an initial inference hypothesis set. Subsequently, the system extracts the statistical distribution features of this set and uses these features as conditional signals, inputting them along with the original visual data, the query question, and the initial inference hypothesis set into the integrated inference model. During joint inference, the integrated inference model uses the statistical distribution features as attention-guided signals to dynamically adjust the fusion emphasis on each initial inference hypothesis, ultimately outputting the target inference path and the target answer result.

[0027] During the training phase, the system employs a similar architecture, but the ensemble reasoning model is in a pre-training state. The system acquires training samples containing visual data, query questions, labeled answers, and a set of initial inference hypotheses. It extracts the statistical distribution features of these initial inference hypotheses and then uses the pre-trained ensemble reasoning model for joint inference to obtain the target inference path and target answer. Subsequently, a reward signal is calculated based on the target answer, target inference path, labeled answers, and statistical distribution features, and the parameters of the ensemble reasoning model are updated using reinforcement learning. Through iterative training with a large number of training samples, the ensemble reasoning model gradually learns to adjust the fusion emphasis on the initial inference hypotheses of each sample using statistical distribution features, thereby enabling it to more accurately complete visual data reasoning tasks during the inference phase.

[0028] The system architecture of the embodiments in this specification has the following significant features. First, the system adopts a two-stage design, separating the sampling stage and the integration stage, enabling the model to make secondary decisions based on the generation of multiple candidate answers, avoiding the bias that may arise from single-step reasoning. Second, statistical distribution characteristics are used as conditional signals in joint reasoning, allowing the integrated reasoning model to perceive the overall distribution characteristics of the initial set of reasoning hypotheses, thereby making more reasonable fusion decisions. Third, the training stage employs reinforcement learning, guiding the model to learn the optimal fusion strategy through reward signals, enabling the integrated reasoning model to adaptively adjust the fusion emphasis of each hypothesis during the reasoning stage.

[0029] In practical applications, the visual data reasoning system provided in the embodiments of this specification can be deployed on various computing devices, including but not limited to servers, cloud computing platforms, and edge computing devices. The integrated reasoning model needs to be specifically trained based on a multimodal large language model, following the training methods provided in the embodiments of this specification. The system can be applied to various visual question-answering scenarios, such as video understanding, image analysis, and document recognition, providing users with accurate and reliable intelligent reasoning services.

[0030] The technical solutions in the embodiments of this specification are based on Bayesian cognitive theory. Bayesian cognition is a probabilistic reasoning framework that simulates the human cognitive process. Its core idea is to model the cognitive process as three stages: the generation of prior hypotheses, the evaluation of prior hypotheses, and the posterior update.

[0031] In the Bayesian cognitive framework, a "prior hypothesis" represents a preliminary judgment formed by the model based on existing knowledge before observing specific evidence. In the embodiments of this specification, prior hypotheses are generated through multiple independent inferences on the same visual data and query question using a multimodal large language model. Each inference result includes the inference process and the answer result, constituting an initial set of inference hypotheses. These prior hypotheses reflect the model's multiple possible understandings and inference paths of the question.

[0032] "Statistical distribution characteristics" are a quantitative description of the overall distribution characteristics of the prior hypothesis set, including the frequency distribution of the responses and the divergence metric. The frequency distribution reflects the proportion of different responses in the prior hypothesis set and can be used to identify mainstream consensus hypotheses and non-mainstream hypotheses. The divergence metric (such as information entropy) reflects the degree of cognitive uncertainty within the prior hypothesis set. A high degree of divergence indicates that the model has significant cognitive conflict regarding the question and requires more careful reasoning and decision-making.

[0033] "Posterior update" refers to the process by which, after observing the statistical distribution characteristics of prior hypotheses, the model dynamically adjusts the fusion weights of each hypothesis based on these characteristics, thereby generating a more accurate and reliable target reasoning path and target answer result. In the embodiments of this specification, the integrated reasoning model uses statistical distribution characteristics as attention-guiding signals to perceive the state of cognitive conflict during joint reasoning and dynamically adjusts the fusion strategy according to the degree of conflict: when cognitive conflict is low (i.e., prior hypotheses are highly consistent), the model tends to trust the mainstream consensus; when cognitive conflict is high (i.e., prior hypotheses are highly dispersed), the model reduces its reliance on frequency statistics and pays more attention to the quality of the reasoning process and the reliability of visual evidence.

[0034] Compared with traditional frequency selection strategies, this Bayesian cognition-based reasoning method has the following advantages: First, it can effectively identify and correct systematic biases in the model, avoiding incorrect selection due to the high frequency of incorrect answers; second, it makes full use of information from the reasoning process, guiding decision-making by evaluating the quality of the reasoning path; third, it adaptively adjusts the fusion strategy according to the state of cognitive conflict, exhibiting stronger robustness in high uncertainty scenarios.

[0035] To better understand the technical solutions of the embodiments in this specification, Figure 2 This paper presents a paradigm comparison between traditional video understanding methods and the Bayesian cognitive integration framework provided in the embodiments of this specification.

[0036] like Figure 2As shown, the traditional method (the upper part) adopts a naive positivist paradigm, directly mapping video frames to answers, relying on simple AI tools (such as Captioner, Summary, etc.) for processing, which has drawbacks such as being prone to errors and requiring human thought processes.

[0037] In contrast, the Bayesian cognitive integration framework (lower half) provided in the embodiments of this specification introduces an intermediate inference layer: first, prior hypotheses are generated through multiple independent inferences, then integrated through a planner debate mechanism, and finally outputting robust inference results. This mechanism of "utilizing and challenging its own priors" enables the model to achieve adaptive robust inference, significantly improving the accuracy of complex video understanding tasks.

[0038] The following is combined Figure 3 The visual data reasoning method provided in the embodiments of this specification will be described in detail. Figure 3 This is the main flowchart of the visual data reasoning method provided in the embodiments of this specification, which includes steps S301 to S304.

[0039] Step S301: Acquire visual data, query questions, and pre-trained multimodal large language models.

[0040] In this step, the system receives the visual data to be processed. And related query issues Visual data Visual input can be in any form, such as images, video frame sequences, or scanned documents. (Query question) These are questions presented in natural language that require reasoning based on visual data to answer, such as multiple-choice questions, numerical estimation questions, and open-ended questions.

[0041] Pre-trained multimodal large language models This is a large-scale neural network model capable of processing both visual and textual modal inputs and generating text outputs. In the embodiments of this specification, this model is obtained through pre-training on large-scale image and text data, possessing powerful visual understanding and text generation capabilities. In this method, the model acts as a "sampler," its main function being to process given visual data... and query issues Multiple candidate answers and their corresponding reasoning processes are generated through multiple independent inferences.

[0042] In practical applications, multimodal large language models Existing open-source models, such as Qwen2.5-VL and LLaMA-VL, can be used. Model parameters... It remains fixed during the inference phase and is not updated.

[0043] Step S302: Use the multimodal large language model to perform multiple independent inferences on the visual data and the query question to generate an initial inference hypothesis set, wherein each initial inference hypothesis contains an independent inference process and answer result.

[0044] In this step, the system utilizes a pre-trained multimodal large language model. Visual data and query issues conduct Subindependent reasoning ( (A positive integer greater than 1). Each inference uses a different random sampling strategy (such as temperature sampling, kernel sampling, etc.) within the model, so even if the input is exactly the same, each inference may generate different inference processes and answer results. The results of this reasoning constitute the initial set of reasoning hypotheses. Each initial inference hypothesis Includes reasoning process and the answer results Two parts. Reasoning process It is the complete chain from which the model starts with visual data and the query question, goes through intermediate reasoning steps, and derives the answer result; the answer result It is the answer that is finally obtained through the reasoning process. Each inference is independent of the others, and the sampling process for each inference is not affected by the results of other inferences.

[0045] To aid understanding, let's illustrate this with a specific visual question-answering scenario. Assume the visual data is an image of a kitchen scene containing multiple objects such as knives, an oven, and a gas stove; the query is "Which object in the image is the most dangerous?" Let K=5. A multimodal large language model performs 5 independent inferences, potentially generating the following 5 initial inference hypotheses: Assumption 1: The reasoning process is "The knife has a sharp edge and may cause cuts," and the answer is "knife." Assumption 2: The reasoning process is "knives are the most dangerous tools in the kitchen," and the answer is "knives." Assumption 3: The reasoning process is "Knives can easily cause injury if used improperly," and the answer is "Knives." Assumption 4: The reasoning process is "Ovens may cause burns when operating at high temperatures," and the answer is "Oven." Assumption 5: The reasoning process is "the gas stove may leak, leading to poisoning," and the answer is "gas stove." As the example above shows, five independent inferences produced three different answers: "knife" appeared three times, "oven" appeared once, and "gas stove" appeared once. Furthermore, even the three inferences regarding the same answer, "knife," differed in their reasoning processes, each emphasizing different aspects such as sharp edges, danger, and improper use. This diversity contains important information about the model's cognitive state.

[0046] It should be noted that the reasoning processes in the initial set of inference assumptions are chain-of-thoughts that naturally arise when the model generates responses, reflecting the model's reasoning logic for the current visual data and query question. Each reasoning process may include multiple steps such as analyzing visual evidence, inferring causal relationships, and generating intermediate conclusions. These reasoning processes will be utilized by the integrated reasoning model in the subsequent joint reasoning stage.

[0047] In practice, the value of K can be flexibly adjusted according to task complexity and computing resources. In a preferred embodiment of this specification, K is set to 8. During the inference phase, the optimal value of K can be selected based on the characteristics of the specific benchmark. For example, for relatively simple benchmarks such as TemporalCompass, K=4 can achieve stable performance; while for complex inference benchmarks such as VSI-Bench, K=8 can bring more significant performance improvements.

[0048] In one specific implementation, the sampling cue template used to guide the multimodal large language model in generating initial inference hypotheses employs a specific cue format, requiring the model to... <analysis> and< / analysis> The tag outputs a detailed reasoning process, in <answer> and< / answer> The final answer is output within the label. This structured output format allows the system to programmatically extract the reasoning process and the answer from the model output, providing standardized data input for subsequent statistical distribution feature extraction and joint inference.

[0049] Step S303: Extract the statistical distribution characteristics of the initial inference hypothesis set.

[0050] In this step, the system performs statistical analysis on the initial inference hypothesis set generated in step S302, extracting statistical distribution features that characterize the overall distribution of the set. Statistical distribution features are characteristics obtained by quantifying the distribution of answers within the initial inference hypothesis set.

[0051] Specifically, statistical distribution characteristics can include the frequency distribution of response results and a measure of divergence among different response results. The frequency distribution characterizes the degree of consensus concentration among the response results in the initial set of inference hypotheses, distinguishing between mainstream and non-mainstream hypotheses. The divergence measure characterizes the degree of internal cognitive uncertainty within the initial set of inference hypotheses.

[0052] Let's continue with the kitchen scenario example above. Among the five initial inference hypotheses, "knife" appeared three times (frequency 3 / 5 = 0.6), "oven" appeared once (frequency 1 / 5 = 0.2), and "gas stove" appeared once (frequency 1 / 5 = 0.2). From the frequency distribution, "knife" is the dominant consensus hypothesis (most frequent), while "oven" and "gas stove" are less common hypotheses. The frequency distribution clearly depicts the degree of consensus among the various responses.

[0053] One way to determine the divergence metric is based on the information entropy of the distribution of different response outcomes in the initial set of inference hypotheses. Specifically, the information entropy H(H_a) is first calculated: in, The set of different answers from the initial set of inference hypotheses. To answer the results The empirical probability (i.e. frequency of occurrence) in the initial set of inference hypotheses.

[0054] Then, the information entropy is normalized to obtain the normalized information entropy as the divergence metric D(H_a): in, The number of different answers is represented by D(H_a). Normalization limits the range of the divergence metric to [0, 1], facilitating comparisons across questions. When all answers are completely identical, D(H_a) = 0, indicating no divergence; when all answers are distinct and occur with equal probability, D(H_a) = 1, indicating maximum divergence.

[0055] Calculate using the above kitchen scenario example: = {“knives”, “oven”, “gas stove”}, = 3. , , Substituting into formula (1) yields the result. Bit, ,therefore This high value indicates a significant degree of internal cognitive uncertainty in the initial set of inference assumptions.

[0056] In joint inference, the integrative inference model identifies mainstream and non-mainstream assumptions based on the degree of consensus concentration and internal cognitive uncertainty, and perceives the overall cognitive conflict state of the initial inference hypothesis set to guide adjustments to the integration emphasis of each initial inference hypothesis. When the divergence metric is high, it indicates that the model has significant cognitive uncertainty about the current problem. In this case, the integrative inference model needs to more carefully analyze the quality of each inference hypothesis, rather than simply trusting the majority vote result.

[0057] Step S304: Using the integrated reasoning model, the statistical distribution features are used as conditional signals to perform joint reasoning with the visual data, the query question, and the initial reasoning hypothesis set to obtain the target reasoning path and target answer result output by the integrated reasoning model; wherein, the statistical distribution features are used as attention guidance signals to adjust the fusion emphasis of the integrated reasoning model on each initial reasoning hypothesis in the initial reasoning hypothesis set.

[0058] In this step, the system uses the statistical distribution features extracted in step S303 as conditional signals, and inputs them into the integrated reasoning model along with the original visual data, the query question, and the initial inference hypothesis set generated in step S302. During the joint reasoning process, the integrated reasoning model uses the statistical distribution features as attention guidance signals to dynamically adjust the fusion emphasis on each initial inference hypothesis, and finally outputs the target reasoning path and the target answer result.

[0059] Within the integrated inference model, statistical distribution characteristics act as conditional signals, influencing the model's attention mechanism. Specifically, when calculating attention weights, the model considers the degree of consensus concentration and cognitive conflict state of each hypothesis as indicated by the statistical distribution characteristics. For high-consensus scenarios (low divergence metric), the model may tend to trust the mainstream consensus hypothesis and assign it a higher fusion weight; for high-divergence scenarios (high divergence metric), the model will more carefully analyze the reasoning quality of each hypothesis and may assign higher attention weights to non-mainstream hypotheses to avoid simply blindly following potentially erroneous majority consensus.

[0060] Continuing with the kitchen scenario example above, the integrated inference model receives the following input: Visual data: Kitchen scene images Question: "Which object in the picture is the most dangerous?" Statistical distribution characteristics: knives 3 / 5 (0.6), ovens 1 / 5 (0.2), gas stoves 1 / 5 (0.2); Disparity measure 0.87 Initial set of inference assumptions: 5 assumptions (Assumptions 1-5) In the joint reasoning process, the integrated reasoning model first detected a divergence metric of 0.87, indicating a high level of cognitive uncertainty. By analyzing the reasoning process of each hypothesis, the model found that while hypotheses 1-3 all point to "knives," their reasoning angles differ (sharp edges, danger, improper use), and hypotheses 4 and 5 point to "oven" and "gas stove," respectively. After comprehensive evaluation, the model might generate the following target reasoning path: "Analyzing the reasoning quality of each hypothesis, hypotheses 1-3 demonstrate the danger of knives from different angles, with consistent and complementary reasoning logic; while hypotheses 4 and 5 also have some merit, the danger of ovens and gas stoves usually requires specific conditions (high-temperature operation, leakage) to be established. In summary, knives are the most dangerous object." The final output target answer is "knives."

[0061] It should be noted that the output format of the integrated reasoning model also adopts a structured format, outputting the target reasoning path and the target answer separately. For example, the target reasoning path is output in... <think> and< / think> Within the tags, the target answer result is output in <answer> and< / answer> Within the tags. This format is consistent with the output format of the sampling phase, facilitating processing in subsequent training phases.

[0062] It is worth noting that the integrated reasoning model There are consistency constraints between the training and inference phases. Specifically, the integrated inference model is based on a multimodal large language model with fixed parameters during the training phase. The generated sample initial inference hypothesis set is optimized for parameters, and a multimodal large language model with fixed parameters is used during the inference phase. Generate an initial set of inference hypotheses. This consistency constraint means that the multimodal large language model used during training to generate the initial set of inference hypotheses is the same model used during inference, with identical and fixed parameters. This constraint is necessary because the correction strategies learned during training are tightly coupled to the output distribution of the sampling models—the statistical distribution features during training are extracted based on the output distribution of a specific sampling model. If a different sampling model is used during inference, its output distribution will change, causing the correction strategies learned during training to fail during inference. By keeping the sampling model parameters fixed, the distribution characteristics of the statistical distribution features are ensured to be consistent between training and inference, allowing the ensemble inference model to naturally utilize the correction strategies learned during training for reliable inference.

[0063] In a preferred embodiment of this specification, the multimodal large language model employs a fine-tuned version of Qwen2.5-VL, and the model parameters remain fixed during the generation of the initial inference hypothesis set (i.e., no parameter updates are performed). The integrated inference model is obtained by specialized training on the same architecture using the training method provided in the embodiments of this specification.

[0064] The following is combined Figure 4 This specification provides a detailed description of the overall architecture of the Video-BCI (Bayesian Cognitive Integration) framework provided in the embodiments. Figure 4 This is a schematic diagram of the architecture of the Video-BCI framework provided in the embodiments of this specification. This framework incorporates the aforementioned... Figure 3 The reasoning method shown and the subsequent Figure 5 The training methods shown are unified within a complete system architecture.

[0065] like Figure 4 As shown, the data flow of the Video-BCI framework, from left to right, includes the following core modules: (1) Prior hypothesis generation module (sampling phase) This module corresponds to Figure 4 The area within the upper left dashed box. Video data. and query issues As input to the system, it is fed into a pre-trained multimodal large language model. (i.e., sampler). For example... Figure 4 As shown, the process comprises two key steps: sampling and prior generation. In the sampling phase, the multimodal large language model... The parameters remain fixed and are not updated, regardless of the input video data. and query issues conduct Subindependent reasoning ( In a preferred embodiment of the present invention, the integer is a positive integer greater than 1. Each inference iteration employs a different random sampling strategy. During the prior generation phase, the model output... Each prior hypothesis. Includes a complete reasoning process and the answer results This constitutes the initial inference hypothesis set in step S302 above. .

[0066] like Figure 4As shown in the upper left area, multiple prior hypotheses are displayed side-by-side as blue rectangles, visually illustrating the diverse cognitive states the model exhibits when faced with the same video comprehension problem. This diversity forms the basis for subsequent cognitive conflict identification and inference path optimization.

[0067] (2) Statistical distribution feature extraction module (reward estimation) This module corresponds to Figure 4 The upper-middle region, also known as the reward estimation module during the training phase, receives the set of prior hypotheses output by the prior hypothesis generation module. For each prior hypothesis, a text quality assessment is performed, and a corresponding textual reward value is generated.

[0068] Specifically, the Reward Estimation module uses a text quality scoring function. For each prior hypothesis Reasoning process and the answer results Scoring is performed. For classification tasks (such as multiple-choice questions), the scoring function is... Returns discrete accuracy (0 or 1, indicating whether the answer is correct); for open tasks (such as open-ended generation), the scoring function... Returns continuous text similarity (such as ROUGE-L score, which indicates how similar the reasoning and answer are to the standard answer).

[0069] like Figure 4 As shown in the diagram, the output of the Reward Estimation module is represented by green rectangles, with each green rectangle corresponding to a text reward value for a prior hypothesis. A higher text reward value indicates better reasoning quality for that prior hypothesis, making it more likely to be used by the subsequent Cognitive Utility Function module to guide reasoning optimization. These text reward values ​​are then passed as input to the Cognitive Utility Function (CUF) module to calculate the Cognitive Conflict Incentive Signal (DUS) and Reasoning Path Tracking Signal (PTS).

[0070] (3) Integrated reasoning model (i.e.) Figure 4 Policy Model (in Chinese) This module corresponds to Figure 4 The area within the lower left dashed box is the core decision-making component of the entire Video-BCI framework. (Integrated Inference Model) Marked by a purple box, its input consists of three parts: Raw video data and query issues (External evidence); The set of prior hypotheses output by the prior hypothesis generation module (Internal cognitive state); The statistical distribution features (conditional signals) output by the statistical distribution feature extraction module.

[0071] Integrated reasoning model In the joint reasoning process, statistical distribution characteristics are used as attention-guiding signals to dynamically adjust the prior hypotheses. The fusion focus. Specifically, when the statistical distribution characteristics show high divergence (e.g. The model will analyze each hypothesis more carefully. The quality of reasoning may give higher weight to non-mainstream hypotheses; when the statistical distribution characteristics show low divergence (e.g. The model tends to trust mainstream consensus assumptions.

[0072] like Figure 4 As shown in the lower left area, the integrated reasoning model The output is purple ReasoningPaths, representing the final target reasoning path. and target response results This process corresponds to the joint reasoning operation described in step S304 above.

[0073] (4) Cognitive Utility Function (CUF) This module corresponds to Figure 4 The lower middle area (marked with a yellow border) and the area within the dashed box on the right are labeled "CUF" in the diagram. This module contains two sub-modules: the Dialectical Uncertainty Signal (DUS) sub-module and the Process Tracing Signal (PTS) sub-module. Figure 4 In the code, these two submodules are labeled "CUF: DUS" and "CUF: PTS" respectively.

[0074] DUS submodules such as Figure 4 The upper right half is shown, used for integrating inference models. The cognitive conflict stimulus signal is calculated when a correct answer is given. The DUS calculation utilizes the textual reward value output by the Reward Estimation module and is performed in two ways depending on the task type: Discrete task scenarios (such as multiple-choice questions): DUS is based on a set of prior hypotheses. The calculation of the frequency distribution of responses and the degree of cognitive conflict includes the following two indicators: Cognitive Conflict Level (Cognitive Conflict): Based on a set of prior hypotheses The normalized information entropy of different answer distributions is calculated to quantify the degree of cognitive uncertainty. The calculation formula is consistent with the aforementioned formula (2). The higher the value, the stronger the cognitive conflict within the set of prior hypotheses.

[0075] Consensus Challenge Strength (Consensus Challenge Strength): Calculates the target response result In the set of prior hypotheses Complementary values ​​appearing in the frequency . The higher the value, the scarcer the target response, and the greater the difficulty for the model to challenge mainstream consensus. The calculation formula can be found in formula (4) below.

[0076] Specifically, the DUS signal in this scenario is and The product of these factors is shown in formula (3) below. Specifically, when the model is in a state of high cognitive conflict (…), the product of these factors is shown in formula (3) below. High, high consensus challenge ( The fact that the model still gave the correct answer even in a high-level state indicates that it has a strong corrective reasoning ability and should be given a higher reward.

[0077] For non-discrete task scenarios (such as open-ended generation tasks): DUS uses the relative superiority method for calculation, directly utilizing the text reward value output by the Reward Estimation module. Specifically, by comparing the text reward value of the current answer with the text reward values ​​of each hypothesis in the prior hypothesis set, it calculates the proportion of prior hypotheses that the current answer is superior to. The higher this ratio, the better the reasoning quality of the current answer is relative to the prior set, and the higher the reward should be given. The specific calculation method can be found in the following formula (5).

[0078] PTS submodules such as Figure 4 The lower right half is shown, used in the integrated reasoning model. When an incorrect answer is given, the inference path tracing signal is calculated. This submodule first filters out answers with higher accuracy than the target answer. Advantageous sample initial reasoning hypothesis subset Then, based on the target reasoning path With the dominant hypothesis subset Each reasoning process The similarity is calculated using weighted averages. Specifically, this submodule includes the following calculation elements: weight coefficients. Based on the first The weights are calculated based on the accuracy improvement of each advantage hypothesis; the maximum value function. Used to calculate the maximum accuracy improvement within the dominant subset; summation operation. The weighted similarity of all hypotheses in the dominant subset is summed. The role of the PTS signal is: when integrating the inference model... When an answer is incorrect, guide the student to learn from the reasoning path based on their dominant hypothesis, gradually improving the quality of their reasoning. The specific formula for calculating the PTS signal will be discussed later. Figure 5 The steps are explained in detail in step S504.

[0079] During the training phase, the cognitive utility function module responds to the target based on the results. Answers with tags Based on the matching results, the DUS submodule or PTS submodule is selectively activated to generate a reward signal. This reward signal is correlated with the standard accuracy reward. and format rewards Together they constitute the complete training objective function. The ensemble inference model is updated using reinforcement learning algorithms (such as GRPO, Group Relative Policy Optimization). parameters ,form Figure 4 The training feedback loop is shown below. The specific training objective function will be discussed later. Figure 5 The steps are explained in detail in step S504.

[0080] like Figure 4 As shown by the red dashed arrow in the diagram, the training feedback loop feeds the reward signal output by the cognitive utility function module back to the integrated reasoning model. Update its parameters This feedback mechanism enables the integration of inference models. Able to learn step by step: In a state of cognitive conflict (high divergence measure) Identify and challenge potentially flawed majority consensus, and provide correct minority judgments; When the quality of the reasoning path is poor, learn from the reasoning path of the dominant hypothesis and gradually improve the quality of the reasoning path; Adaptively utilize statistical distribution characteristics to adjust the fusion emphasis of each hypothesis, trusting the majority in consensus scenarios and conducting prudent analysis in divergent scenarios.

[0081] Through the above architecture, the Video-BCI framework organically unifies the forward inference process in the inference phase with the feedback optimization process in the training phase. In the inference phase, the inference model is integrated... By leveraging the fusion strategy learned during training, accurate visual data inference is adaptively performed based on statistical distribution features. During training, the cognitive utility function guides the integrated inference model through two complementary reward signals: DUS and PTS. Learn to challenge erroneous consensuses when faced with cognitive conflict, and learn better reasoning logic from superior assumptions when the quality of reasoning paths is poor.

[0082] It is worth noting that there is a clear correspondence between the modules in the Video-BCI framework and the steps in the aforementioned method embodiments: The prior hypothesis generation module corresponds to the multiple independent reasoning processes in steps S302 / S501; The statistical distribution feature extraction module corresponds to the statistical distribution feature extraction process in steps S303 / S502; Integrated reasoning model This corresponds to the joint reasoning process in steps S304 / S503; The cognitive utility function module corresponds to the reward signal calculation in the parameter update process of step S504.

[0083] therefore, Figure 4 The Video-BCI framework shown is not only Figure 3 Reasoning methods and Figure 5 A concrete implementation example of the training method will be provided later. Figure 7 Training curve analysis and Figure 8 The technical basis of qualitative case studies.

[0084] The following is combined Figure 5 The method for training a visual reasoning model provided in the embodiments of this specification will be described in detail. Figure 5 This is the main flowchart of the method for training a visual reasoning model provided in the embodiments of this specification, which includes steps S501 to S504.

[0085] Step S501: Obtain training samples, which include visual data, query questions, labeled answers, and a set of initial inference hypotheses generated by a pre-trained multimodal large language model. Each initial inference hypothesis contains an inference process and answer result for an independent inference of the visual data and the query question.

[0086] In this step, the system obtains training samples for training the integrated reasoning model. Each training sample contains four components: visual data, query question, labeled answer, and a set of initial inference hypotheses for the sample.

[0087] Visual data and query questions are defined in the same way as in the inference phase, representing the visual input to be inferred and the natural language question to be answered, respectively. The labeled answer is the correct answer to the query question and is used to calculate the reward signal during training.

[0088] The initial inference hypothesis set for the samples is generated by a pre-trained multimodal large language model (with fixed parameters and no updates) through multiple independent inferences on the visual data and query questions in the training samples. Its generation method is completely consistent with the description in step S302 of the inference stage, that is, K independent inferences are performed on the same visual data and query questions, and each inference generates an initial inference hypothesis for the samples containing the inference process and the answer result.

[0089] Let's continue with the kitchen scene example above. Assuming the labeled answer is "knife", the pre-trained multimodal large language model performs 5 independent inferences on the kitchen scene image and the query question "Which object in the picture is the most dangerous?", generating 5 sample initial inference hypotheses (hypotheses 1-5), the content of which is consistent with the example described in step S302 of Part 2.

[0090] In a preferred embodiment of this specification, the visual data and query questions of the training samples are derived from the publicly available video understanding dataset Video-R1-260k. It should be noted that this embodiment does not perform supervised fine-tuning (SFT) before training; it only uses the hints and answers from the dataset. The initial inference hypothesis set for the samples is generated entirely by a multimodal large language model with fixed parameters through multiple independent inferences.

[0091] Step S502: Extract the statistical distribution characteristics of the initial inference hypothesis set of the sample.

[0092] In this step, the system performs statistical analysis on the initial inference hypothesis set obtained in step S501 and extracts statistical distribution features. The extraction method is completely consistent with that described in step S303 of the inference stage, including extracting the frequency distribution of the answer results and the discrepancy measure between different answer results.

[0093] Let's continue with the kitchen scenario example above. In the initial inference hypothesis of the 5 samples, the answer "knife" appeared 3 times (frequency 0.6), "oven" appeared once (frequency 0.2), and "gas stove" appeared once (frequency 0.2). The divergence metric D(H_a) ≈ 0.87, indicating a high level of cognitive uncertainty.

[0094] Step S503: Using the integrated reasoning model to be trained, with the statistical distribution features as conditional signals, perform joint reasoning with the visual data, the query question, and the initial inference hypothesis set of the samples to obtain the target reasoning path and the target answer result.

[0095] In this step, the system uses the statistical distribution features extracted in step S502 as conditional signals, and inputs them, along with the visual data, query question, and initial inference hypothesis set of the samples, into the ensemble inference model to be trained. During the joint inference process, the ensemble inference model to be trained uses the statistical distribution features as attention guidance signals to dynamically adjust the fusion emphasis on the initial inference hypotheses of each sample, and outputs the target inference path and target answer result.

[0096] The parameters of the ensemble inference model to be trained may be randomly initialized in the early stages of training, or they may be fine-tuned based on the pre-trained parameters of a multimodal large language model. As training progresses, the model gradually learns to adjust the fusion strategy using statistical distribution features, enabling it to naturally utilize statistical distribution characteristics for accurate inference during the inference phase.

[0097] Step S504: Update the parameters of the integrated reasoning model based on the target answer result, the target reasoning path, the labeled answer, and the statistical distribution characteristics.

[0098] In this step, the system calculates a reward signal based on the target answer result and target reasoning path output in step S503, combined with the labeled answers in the training samples and the statistical distribution features extracted in step S502, and updates the parameters of the integrated reasoning model based on the reward signal through reinforcement learning.

[0099] The specific implementation of parameter updates can be as follows: calculate a reward signal based on the target answer result, target inference path, labeled answer, and statistical distribution characteristics, and update the parameters of the integrated inference model based on this reward signal through reinforcement learning. Unlike traditional supervised learning where the loss function directly calculates the deviation between the output and the label and minimizes it through backpropagation, this embodiment uses a reinforcement learning framework to guide the model to learn the optimal inference strategy through the reward signal. In the embodiments of this specification, reinforcement learning can employ policy gradient algorithms such as GRPO (Group Relative Policy Optimization).

[0100] The overall design of the reward signal adopts a mutually exclusive selection architecture, including two types: cognitive conflict incentive signals and inference path tracking signals. These two signals do not take effect simultaneously, but are used selectively based on the matching relationship between the target answer and the label answer: when the target answer matches the label answer, a cognitive conflict incentive signal is generated based on statistical distribution characteristics and the target answer to incentivize the integrated reasoning model to perform corrective reasoning under cognitive conflict conditions; when the target answer does not match the label answer, an inference path tracking signal is generated based on the target inference path, the initial set of inference hypotheses for the sample, and the label answer to guide the integrated reasoning model to align with a better inference path.

[0101] The technical logic of this mutually exclusive selection architecture is as follows: When the model answers correctly, it is necessary to strengthen its ability to give correct answers even under cognitive conflict. Therefore, a cognitive conflict incentive signal is used. If the model still gives a correct answer in a set of hypotheses with high divergence and high scarcity, it indicates that the model has a strong corrective reasoning ability and should be given a higher reward. When the model answers incorrectly, it is necessary to guide it to find a better reasoning path from the set of hypotheses as a reference. Therefore, a reasoning path tracking signal is used. Advantageous hypotheses with higher accuracy than the target answer are selected from the set of hypotheses, and the model is encouraged to move closer to the reasoning path of these hypotheses.

[0102] Figure 6 This is a schematic diagram illustrating the mutual exclusion selection and generation mechanism for reward signals. (Example) Figure 6 As shown, the system first determines whether the target answer matches the labeled answer. If they match, it proceeds to the generation process of the cognitive conflict stimulus signal; if they do not match, it proceeds to the generation process of the reasoning path tracking signal. The two signals are mutually exclusive, and only one type is generated in each training iteration.

[0103] The generation processes of cognitive conflict stimulus signals and reasoning path tracking signals are explained in detail below.

[0104] The generation process of cognitive conflict stimulus signals employs different calculation methods depending on the task type.

[0105] In closed-ended tasks (such as multiple-choice questions), the generation of cognitive conflict stimulus signals includes the following steps: determining the degree of disagreement of the answer results and the scarcity of the target answer results based on statistical distribution characteristics; and generating cognitive conflict stimulus signals based on the product of the degree of disagreement and the degree of scarcity.

[0106] The higher the degree of disagreement and the scarcer the target answer, the more likely the model can still provide a correct answer under cognitive conflict, and thus should be given a higher incentive reward. The formula for calculating the cognitive conflict incentive signal is: in, As an indicator function, ensure that only the target answer result is used. Answers with tags The cognitive conflict stimulus signal is only calculated when the accuracy is greater than 0, in order to prevent the model from deliberately outputting wrong answers in order to obtain rewards. To the degree of disagreement, This refers to the degree of scarcity.

[0107] Degree of disagreement One specific calculation method is to determine the normalized information entropy by calculating the distribution of different answer results in the initial inference hypothesis set of the sample, and its calculation formula is consistent with formula (2).

[0108] In closed-loop mission scenarios, scarcity Determined by calculating the complementary values ​​of the frequencies of the target answer in the initial inference hypothesis set of the sample: in, Result of answering the target Initial inference hypothesis set of the sample The number of times it appears in This represents the total number of responses in the initial set of inference hypotheses for the sample. When the target response is a minority opinion, its scarcity level is close to 1; when the target response is the mainstream consensus, its scarcity level is close to 0.

[0109] Let's take the kitchen scenario example above as an illustration. Assuming the target answer is "knife" and the labeled answer is also "knife" (match), then the process of generating cognitive conflict stimulus signals begins. , , Degree of disagreement Cognitive conflict motivating signals .

[0110] In open-ended generation tasks, since the answer results are not discrete options, their frequency cannot be directly calculated. Therefore, the cognitive conflict stimulus signal is calculated using the relative superiority method. Specifically, the accuracy score (e.g., ROUGE-L score) of the target answer result is obtained. The number of hypotheses with accuracy scores lower than the target answer result's accuracy score is counted in the initial inference hypothesis set. The cognitive conflict stimulus signal is generated based on the ratio of this number of hypotheses to the total number of hypotheses in the initial inference hypothesis set. The calculation formula is as follows: in, An accuracy scoring function for non-multiple-choice tasks (such as ROUGE-L). This is an indicator function. The formula calculates the current answer. The accuracy score exceeds the accuracy scores of a given prior hypothesis, meaning the current answer surpasses a certain percentage of the prior hypotheses. A higher ratio indicates that the reasoning quality of the current answer is superior to the prior set, and therefore warrants a higher reward.

[0111] The generation process of the inference path tracing signal can be broken down into the following steps: First, a subset of advantageous initial inference hypotheses with higher accuracy than the target answer is selected from the initial set of inference hypotheses. Advantageous subset Defined as: in, For the first The accuracy of the answers to the initial inference hypothesis of each sample. The accuracy of the target answer. For classification problems, the accuracy function is... Returns 0 or 1 (1 for correct, 0 for incorrect).

[0112] Secondly, based on the difference between the accuracy of the initial inference hypothesis of each sample in the dominant sample initial inference hypothesis subset and the accuracy of the target answer, the weight corresponding to the initial inference hypothesis of each sample is determined. The calculation formula is: in, The highest accuracy among the initial inference hypothesis sets for the sample. To prevent zero constant (such as) The weight design assigns higher weights to hypotheses that improve accuracy more significantly, meaning the model is more strongly guided to mimic the inference paths that bring the greatest performance gains.

[0113] Finally, the inference path tracking signal is calculated based on the similarity between the weights and the target inference path and the inference processes of the initial inference hypotheses of each sample in the dominant sample initial inference hypothesis subset. The calculation formula is as follows: in, To normalize the weights, Target reasoning path With the The reasoning process of the initial inference hypothesis for each sample Similarity between them (such as textual similarity or semantic similarity).

[0114] Continuing with the kitchen scenario example above, assuming the target answer is "oven" (which doesn't match the label answer "knife"), the system proceeds to generate the inference path tracking signal. The system selects a subset of initial inference hypotheses from five samples, choosing those with higher accuracy than "oven". Since hypotheses 1-3 result in "knife" (matching the label answer, high accuracy), and hypotheses 5 result in "gas stove" (accuracy may be higher than or equal to "oven"), the subset of inference hypotheses may include hypotheses 1-3 (and possibly hypotheses 5). The system calculates the similarity between the target inference path and the inference processes of these inference hypotheses, weighted and summed according to the accuracy improvement, to obtain the inference path tracking signal. This signal guides the integrated inference model to generate inference paths that are closer to the inference logic of hypotheses 1-3 during subsequent training.

[0115] Before determining the weights mentioned above, a boundary case needs to be addressed: whether the initial inference hypothesis subset for the dominant samples is empty. If there are no hypotheses in the hypothesis set with accuracy higher than the target answer, it means the model has provided the best answer in the hypothesis set, but it still doesn't match the labeled answer. In this case, the model is facing an extremely difficult training sample. In this situation, the preset maximum reward value is determined as the inference path tracking signal. The engineering consideration for this design is that in extremely difficult scenarios, the model cannot find a better reference from the hypothesis set. Giving zero or negative rewards would inhibit the model's willingness to explore such samples; while giving a preset maximum reward value can encourage the model to continue exploring in difficult scenarios and avoid getting stuck in local optima during training.

[0116] Combining the above cognitive conflict stimulus signals and reasoning path tracking signals, the complete training objective function of the integrated reasoning model is: in, A standard accuracy reward (0 or 1 for categorized tasks). Format-based rewards (e.g., whether the output format is correct). It is a cognitive utility function that includes cognitive conflict stimulus signals. and inference path tracing signals In each training iteration, based on the matching results between the target answer and the labeled answer, Only activated in the middle or one.

[0117] Through the above training process, the integrated reasoning model gradually learns the following capabilities: First, in a state of cognitive conflict (high divergence metric), it can identify and challenge potentially erroneous majority consensus and give correct minority judgments; Second, when the quality of the reasoning path is poor, it can learn from the reasoning path of the dominant hypothesis and gradually improve the quality of the reasoning path; Third, it can adaptively utilize statistical distribution characteristics to adjust the fusion emphasis of each hypothesis, trusting the majority in consensus scenarios and conducting prudent analysis in divergence scenarios.

[0118] The visual data reasoning and training methods provided in the embodiments of this specification are verified and explained below with specific experimental data.

[0119] In a preferred implementation of the embodiments of this specification, the following training configuration is adopted: the hardware environment uses 16 NVIDIA A100 GPUs (totaling 1280GB of video memory). The base model uses a fine-tuned version of Qwen2.5-VL as a pre-trained multimodal large language model with 7B (7 billion) parameters. During the generation of the initial inference hypothesis set, the model parameters remain fixed and are not updated. The training data uses the publicly available video understanding dataset Video-R1-260k as the source of visual data and query questions for training samples. It should be noted that the embodiments of this specification do not perform supervised fine-tuning (SFT) before training; only the hints and answers from this dataset are used. The initial inference hypothesis set for the samples is generated entirely by Qwen2.5-VL with fixed parameters through multiple independent inferences. During the training phase, the video resolution is set to 128×28×28, with 16 frames sampled uniformly. During the inference phase, the resolution is set to 256×28×28 to ensure fair comparison. For hyperparameter settings, the training learning rate was set to 1e-6, and the total batch size was 16 (1 batch per device). The number of prior hypotheses K and the number of rollouts were both set to 8. The reinforcement learning algorithm used was GRPO (Group Relative Policy Optimization).

[0120] Regarding benchmarking and evaluation metrics, the embodiments in this specification have conducted extensive evaluations on six mainstream video understanding benchmarks, which are divided into two categories: inference benchmarks. For example, these may include: VSI-Bench, used to evaluate video scene understanding capabilities; VideoMMMU, used to evaluate multidisciplinary professional video understanding capabilities; and MMVU, used to evaluate expert-level multidisciplinary video understanding capabilities. General benchmarks. For example, these may include: MVBench, used to evaluate multidimensional video understanding capabilities; TempCompass, used to evaluate temporal understanding capabilities; and VideoMME, used to evaluate multimodal video understanding capabilities. All benchmarks use accuracy as the evaluation metric.

[0121] The embodiments in this specification were compared with various existing methods on six benchmarks at a 16-frame setting. The Video-BCI method (containing only the cognitive conflict stimulus signal DUS) and the Video-BCI-R method (containing both DUS and the reasoning path tracking signal PTS) provided in the embodiments of this specification significantly outperform the baseline model Qwen2.5-VL on all benchmarks.

[0122] The performance of the embodiments in this specification under different frame rate settings is described in detail below with reference to Table 1. Table 1 shows a performance comparison between the embodiments in this specification and recent best methods on video inference benchmarks and general benchmarks.

[0123] Table 1 In this document, "Proprietary" represents a closed-source business model, "SFT" represents a supervised fine-tuning method, "LLM Agent" represents an agent-based method based on a large language model, and "RL" represents a reinforcement learning method. Video-BCI-7B is the basic cognitive model (containing only DUS) in the embodiments of this specification, and Video-BCI-R-7B is the complete reasoning enhancement model (containing both DUS and PTS) in the embodiments of this specification.

[0124] As shown in Table 1, compared with the closed-source commercial model, the 64-frame Video-BCI-R-7B of this specification's embodiment demonstrates high competitiveness on inference benchmarks. It achieves a score of 39.4 on VSI-Bench, exceeding GPT-4o's 34.0; and a score of 69.6 on MMVU, approaching GPT-4o's 75.4. On general benchmarks, Video-BCI-R-7B achieves 67.8 on MVBench, 73.8 on TempCompass, and 62.9 on VideoMME, all approaching or exceeding the performance of GPT-4o and Gemini 1.5 Pro.

[0125] Compared to the existing state-of-the-art reinforcement learning method, Video-R1-7B, the embodiments in this specification achieve comprehensive improvements. At a 64-frame setting, Video-BCI-R-7B outperforms Video-R1-7B by 2.3 percentage points on VSI-Bench (39.4 vs 37.1), by 2.0 percentage points on VideoMMMU (54.4 vs 52.4), by 5.8 percentage points on MMVU (69.6 vs 63.8), by 3.0 percentage points on MVBench (67.8 vs 64.8), and by 1.5 percentage points on VideoMME (62.9 vs 61.4).

[0126] The ablation experiment results of the embodiments of this specification are described below with reference to Table 2, in order to verify the effectiveness of the cognitive conflict stimulus signal (DUS) and reasoning path tracking signal (PTS).

[0127] Table 2 Among them, Qwen2.5-VL-CoT: baseline model, using thought chain cues; Qwen2.5-BCI: only uses the Bayesian cognitive integration framework, without reinforcement learning training; Qwen2.5-VL-MJ: baseline model, using a majority voting strategy; Qwen2.5-BCI-GRPO: trained using standard GRPO, without DUS or PTS; Qwen2.5-BCI-DUS: only includes cognitive conflict stimulus signals (DUS); Qwen2.5-BCI-PTS: only includes reasoning path tracking signals (PTS); Qwen2.5-BCI-R: includes both DUS and PTS (complete model).

[0128] Experimental results show that the Bayesian cognitive integration framework itself brings significant performance improvements. Qwen2.5-BCI achieves a score of 28.8 on VSI-Bench, surpassing Qwen2.5-VL-CoT's 27.7; and on MMVU, it reaches 66.3, exceeding Qwen2.5-VL-CoT's 59.2. This indicates that even without reinforcement learning training, simply extracting statistical distribution features and performing joint inference can significantly improve inference accuracy.

[0129] Using either DUS or PTS alone yields significant performance improvements. Qwen2.5-BCI-DUS achieves a score of 36.2 on VSI-Bench, 68.0 on MMVU, 64.3 on MVBench, and 73.4 on TempCompass. Qwen2.5-BCI-PTS achieves a score of 34.9 on VSI-Bench, 68.2 on MMVU, 64.4 on MVBench, and 74.5 on TempCompass. This indicates that both cognitive conflict stimulus signals and inference path tracking signals can effectively guide model learning.

[0130] The complete model Qwen2.5-BCI-R, using both DUS and PTS simultaneously, achieved optimal performance on multiple benchmarks. It reached 52.4 on VideoMMMU and 58.7 on VideoMME, both exceeding variants using either DUS or PTS alone. This indicates that the two signals are complementary, and using them together can further improve model performance.

[0131] The following is combined Figure 7 The training process comparison graphs shown illustrate the performance comparison between the embodiment (BCI) and the GRPO baseline during training. The X-axis represents the number of training steps (Step × 100), and all curves are smoothed using EMA = 50. Figure 7 It contains three subgraphs: Left chart (Thinking Index): This index is the PTS score, representing the correlation between the inference path generated by the integrated inference model and the high-quality prior assumptions. For example... Figure 7 As shown in the left figure, the BCI (red line) steadily converges to a high plateau of approximately 0.49-0.50 during training, significantly higher than the GRPO baseline (orange line), which exhibits sharp fluctuations and a downward trend. This indicates that the Inference Path Tracking Signal (PTS) in the embodiments of this specification can stably guide the model to learn high-quality inference paths.

[0132] The middle figure (Superiority index): This index represents the relative accuracy ranking (percentile) of the responses generated by the integrated inference model within the set of prior hypotheses. For example... Figure 7 As shown in the middle figure, after the initial stage, the BCI curve consistently outperforms the GRPO curve and stabilizes near 1.0. This indicates that BCI can learn cognitive reflection abilities more quickly and generate responses that are superior to the initial prior assumptions.

[0133] The chart on the right (Accuracy metric): This metric represents the accuracy score. For example... Figure 7 As shown in the right figure, BCI not only converges faster but also stably converges to a high accuracy plateau of approximately 0.72-0.74, which is higher than the GRPO baseline. This indicates that the training method described in the embodiments of this specification can deliver superior final performance.

[0134] In the ablation experiments with hyperparameter K (number of prior hypotheses), the performance of the embodiments in this specification changed significantly across six benchmarks as the K value varied from 0 to 8. The experiments included two groups: Video-BCI (DUS only) and Video-BCI-R (DUS+PTS).

[0135] For Video-BCI (DUS only): When K=0 or K=1, the model lacks sufficient statistical samples to form a meaningful answer distribution, causing the DUS signal to fail; as K increases to 8, the model obtains more stable statistical data, and DUS can more accurately quantify cognitive conflict and reward behaviors that challenge erroneous consensus; on the VideoMMMU and MMVU benchmarks, the performance steadily improves with increasing K; on benchmarks such as TempCompass, the performance saturates when K≈4, indicating that 4 prior hypotheses are sufficient to establish a stable answer distribution.

[0136] For Video-BCI-R (DUS+PTS): When K=0, the model lacks traceable inference paths, and its performance (especially on VSI-Bench and MMVU) is even lower than the Video-BCI baseline with K=0. As K increases, PTS obtains richer teaching materials (high-quality inference paths), and its performance curve is steeper. On complex inference benchmarks (such as VSI-Bench), its peak performance is significantly higher than that of the DUS-only model. This shows that PTS successfully elevates the model's use of prior assumptions from statistical answer verification to knowledge distillation at the inference path level.

[0137] Experimental conclusions: Although the models in the embodiments of this specification uniformly use K=8 during training, the optimal K value during inference depends on the specific task. For example, TempCompass's performance saturates when K=4, while VSI-Bench still achieves significant improvements when K=8.

[0138] The following is combined Figure 8 The qualitative case study diagram shown illustrates the practical effectiveness of the embodiments in this specification on complex reasoning tasks. The task in this case is: "What is the size of this room (in square meters)? If multiple rooms are shown, estimate the size of the combined space." The standard answer is 15 square meters.

[0139] like Figure 8 As shown, the baseline model Qwen2.5-VL (MJ, majority voting) performs poorly when faced with ambiguity. Its reasoning process states, "Let me think, considering all visible elements, I feel that the area of ​​a single room here is approximately 12 square meters. As for the combined area, we can only estimate one at present, because more relevant information is needed to make an accurate judgment," ultimately giving the incorrect answer 12. This exposes the vulnerability of the majority voting strategy to ambiguity and erroneous prior consensus. In contrast, the Video-BCI (DUS only) embodiment in this specification successfully challenges erroneous consensus, giving the correct answer 15. This demonstrates that cognitive conflict incentive signals (DUS) can effectively guide the model to select the correct answer.

[0140] Furthermore, the complete model Video-BCI-R (DUS+PTS) in the embodiments of this specification not only provides the correct answer but also generates a high-quality inference path. For example... Figure 8As shown, the reasoning process is as follows: "The analysis provided by the assumptions consistently indicates that, whether for a single room or a combined space, the room area falls perfectly within the range of 10 to 15 square meters. This range ensures that the room can comfortably accommodate all visible equipment and space, while allowing for minor variations in depth and corner areas. This consistent range perfectly matches the estimated area, confirming that the room area and the combined space are ideally 15 square meters." This demonstrates that the inference path tracking signal (PTS) successfully guided the model to distill from the prior assumptions and reproduce a high-quality reasoning process.

[0141] In summary, in the embodiments provided in this specification, an initial set of inference hypotheses, including the inference process and answer results, is generated by using a multimodal large language model to perform multiple independent inferences on visual data and query questions. The statistical distribution characteristics of this set are extracted as conditional signals and input into an integrated inference model along with the visual data, query questions, and the initial set of inference hypotheses for joint inference. This allows the integrated inference model to obtain information about the overall distribution characteristics of the initial set of inference hypotheses during the joint inference process. Since the statistical distribution characteristics, as attention-guiding signals, directly participate in the integrated inference model's adjustment of the fusion emphasis of each initial inference hypothesis, the integrated inference model can perceive the distribution state of each answer result in the initial set of inference hypotheses, rather than simply relying on the frequency of occurrence. When high-frequency erroneous consensus occurs in multiple independent inferences, the integrated inference model, based on the consensus concentration reflected by the statistical distribution characteristics, can comprehensively consider the overall distribution of each initial inference hypothesis to determine the fusion emphasis, thereby reducing the risk of blindly following high-frequency erroneous answers and improving the accuracy of visual data inference.

[0142] This specification also provides a visual data reasoning device, which can be deployed in the following electronic devices as the execution subject of each method step.

[0143] like Figure 9 As shown, the visual data inference apparatus provided in the embodiments of this specification includes: The hypothesis generation unit 901 is used to perform multiple independent inferences on visual data and query questions using a pre-trained multimodal large language model to generate an initial inference hypothesis set, wherein each initial inference hypothesis contains an independent inference process and answer result.

[0144] Assume that the specific implementation of generation unit 901 is consistent with the description of step S302 in the aforementioned method embodiment. Assume that generation unit 901, by calling a pre-trained multimodal large language model, performs K independent inferences on the same visual data and query question, with each inference employing a different random sampling strategy, thereby generating a set containing K initial inference hypotheses. Each initial inference hypothesis contains a reasoning process and an answer result.

[0145] Feature extraction unit 902 is used to extract the statistical distribution features of the initial inference hypothesis set.

[0146] The specific implementation of the feature extraction unit 902 is consistent with the description of step S303 in the aforementioned method embodiment. The feature extraction unit 902 performs statistical analysis on the initial inference hypothesis set output by the hypothesis generation unit 901 and extracts statistical distribution features that can characterize the overall distribution characteristics of the set.

[0147] Optionally, the feature extraction unit 902 is specifically used to: extract statistical distribution features including the frequency distribution of the answer results and the divergence measure between different answer results. The frequency distribution is used to characterize the degree of consensus concentration among the answer results in the initial inference hypothesis set, to distinguish between mainstream consensus hypotheses and non-mainstream hypotheses. The divergence measure is used to characterize the degree of internal cognitive uncertainty of the initial inference hypothesis set. The joint inference unit 903 is specifically used to: perceive the overall cognitive conflict state of the initial inference hypothesis set based on the degree of consensus concentration and the degree of internal cognitive uncertainty, so as to guide the adjustment of the fusion emphasis of each initial inference hypothesis.

[0148] Optionally, the feature extraction unit 902 is specifically used to determine the divergence metric based on the information entropy of the distribution of different answer results in the initial inference hypothesis set. Specifically, the feature extraction unit 902 first calculates the information entropy, and then normalizes the information entropy to obtain the normalized information entropy as the divergence metric.

[0149] The joint reasoning unit 903 utilizes the integrated reasoning model to perform joint reasoning with visual data, the query question, and the initial set of reasoning hypotheses, using statistical distribution features as conditional signals, to obtain the target reasoning path and target answer output by the integrated reasoning model. The statistical distribution features serve as attention-guiding signals to adjust the fusion emphasis of the integrated reasoning model on each initial reasoning hypothesis in the initial set of reasoning hypotheses.

[0150] The specific implementation of the joint reasoning unit 903 is consistent with the description of step S304 in the aforementioned method embodiment. The joint reasoning unit 903 transforms the statistical distribution features extracted by the feature extraction unit 902 into structured descriptive text, and concatenates this text with the text content of each hypothesis in the initial reasoning hypothesis set to form a joint input sequence. This joint input sequence, along with the visual data and the query question, is input into the integrated reasoning model. During the joint reasoning process, the integrated reasoning model uses the statistical distribution features as attention guidance signals to dynamically adjust the fusion emphasis on each initial reasoning hypothesis, ultimately outputting the target reasoning path and the target answer result.

[0151] Optionally, the device satisfies the following configuration: during the training phase, the integrated inference model optimizes its parameters based on the initial inference hypothesis set generated by a multimodal large language model with fixed parameters, and during the inference phase, it uses the same multimodal large language model with fixed parameters to generate the initial inference hypothesis set. This consistency constraint ensures that the fusion strategy learned during training matches the actual hypothesis distribution characteristics encountered during inference.

[0152] It should be noted that the hypothesis generation unit 901, feature extraction unit 902, and joint inference unit 903 described above can be implemented in software, hardware, or a combination of both. In one specific implementation, each of these units can be executed on a processor by calling corresponding program code. In another specific implementation, each of these units can be integrated into one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other hardware devices.

[0153] like Figure 10 As shown in the embodiments of this specification, an apparatus for training a visual reasoning model is also provided, which includes the following functional units: The sample acquisition unit 1001 is used to acquire training samples, which include visual data, query questions, labeled answers, and a set of initial inference hypotheses generated by a pre-trained multimodal large language model. Each initial inference hypothesis contains an inference process and answer result for an independent inference of the visual data and the query question.

[0154] The training feature extraction unit 1002 is used to extract the statistical distribution features of the initial inference hypothesis set of the sample.

[0155] The training joint reasoning unit 1003 is used to perform joint reasoning with the visual data, query question and sample initial reasoning hypothesis set using the integrated reasoning model to be trained, with statistical distribution characteristics as conditional signals, to obtain the target reasoning path and target answer result.

[0156] The parameter update unit 1004 is used to update the parameters of the integrated inference model based on the target answer result, the target inference path, the labeled answer, and the statistical distribution characteristics. The statistical distribution characteristics are used to guide the integrated inference model to adjust the fusion emphasis of the initial inference assumptions for each sample during joint inference.

[0157] Optionally, the parameter update unit 1004 is specifically used to: calculate a reward signal based on the target answer result, the target reasoning path, the labeled answer, and the statistical distribution characteristics, and update the parameters of the integrated reasoning model based on the reward signal through reinforcement learning.

[0158] Optionally, the reward signal includes a cognitive conflict incentive signal and a reasoning path tracking signal. Specifically, the parameter update unit 1004 is used to: generate a cognitive conflict incentive signal based on statistical distribution characteristics and the target answer when the target answer matches the labeled answer, to incentivize the integrated reasoning model to perform corrective reasoning under cognitive conflict conditions; and generate a reasoning path tracking signal based on the target reasoning path, the initial set of inference hypotheses for the sample, and the labeled answer when the target answer does not match the labeled answer, to guide the integrated reasoning model to align with a better reasoning path.

[0159] Optionally, the parameter update unit 1004 is specifically used to generate the cognitive conflict stimulus signal when: in the case that the query question corresponds to a closed task, determine the degree of divergence of the answer results and the degree of scarcity of the target answer results based on statistical distribution characteristics; and generate the cognitive conflict stimulus signal based on the product of the degree of divergence and the degree of scarcity.

[0160] Optionally, the parameter update unit 1004 is specifically used to generate the cognitive conflict stimulus signal by: obtaining the accuracy score of the target answer result when the query question corresponds to an open generation task; counting the number of hypotheses with accuracy scores lower than the accuracy score of the target answer result in the initial inference hypothesis set of the sample; and generating the cognitive conflict stimulus signal based on the ratio of the number of hypotheses to the total number of hypotheses in the initial inference hypothesis set of the sample.

[0161] Optionally, the degree of divergence is determined by calculating the normalized information entropy of the distribution of different answer results in the initial inference hypothesis set of the sample.

[0162] Optionally, the parameter update unit 1004, when generating the inference path tracking signal, specifically performs the following: selecting a subset of advantageous sample initial inference hypotheses with higher accuracy than the target answer from the sample initial inference hypothesis set; determining the weight corresponding to each sample initial inference hypothesis based on the difference between the accuracy of each sample initial inference hypothesis in the advantageous sample initial inference hypothesis subset and the accuracy of the target answer; and calculating the inference path tracking signal based on the weight and the similarity between the inference process of the target inference path and the inference process of each sample initial inference hypothesis in the advantageous sample initial inference hypothesis subset.

[0163] Optionally, the parameter update unit 1004 is further configured to: determine whether the subset of initial inference hypotheses of the dominant samples is an empty set before determining the weights corresponding to the initial inference hypotheses of each sample; and if the subset of initial inference hypotheses of the dominant samples is an empty set, determine the preset maximum reward value as the inference path tracking signal.

[0164] This specification also provides an electronic device through its embodiments. For example... Figure 11As shown, the electronic device includes at least one processor 1101 and a memory 1102 communicatively connected to the processor 1101. The memory 1102 stores instructions executable by the at least one processor 1101, which, when executed, enable the at least one processor 1101 to perform the visual data inference method or the training method for the visual inference model described in the foregoing method embodiments. The electronic device may also include a communication interface 1103 for data communication with other devices. The processor 1101, memory 1102, and communication interface 1103 are interconnected via a bus 1104.

[0165] This specification also provides a non-transitory computer-readable storage medium storing computer-executable instructions. When these instructions are executed by a processor, they can implement the visual data reasoning method or the training method for the visual reasoning model described in the foregoing method embodiments. The non-transitory computer-readable storage medium can be a tangible storage medium, such as a disk, optical disk, semiconductor memory, etc.

[0166] This specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, can implement the visual data reasoning method or the training method for a visual reasoning model described in the foregoing method embodiments. This computer program product can be stored in a storage medium as a software installation package, or it can be downloaded and installed via a network.

[0167] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0168] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this specification may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0169] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0170] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0171] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0172] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0173] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0174] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0175] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. A visual data reasoning method, comprising: Acquire visual data, query questions, and pre-trained multimodal large language models; The multimodal large language model is used to perform multiple independent inferences on the visual data and the query question to generate an initial inference hypothesis set, wherein each initial inference hypothesis contains an independent inference process and answer result; Extract the statistical distribution characteristics of the initial inference hypothesis set; Using an integrated reasoning model, the statistical distribution features are used as conditional signals to perform joint reasoning with the visual data, the query question, and the initial set of reasoning hypotheses, to obtain the target reasoning path and target answer output by the integrated reasoning model. The statistical distribution features are used as attention-guiding signals to adjust the fusion emphasis of the integrated inference model on each initial inference hypothesis in the initial inference hypothesis set.

2. The method according to claim 1, wherein, The statistical distribution characteristics include the frequency distribution of the response results and the discrepancy measure between different response results; The frequency distribution is used to characterize the degree of consensus concentration among the responses in the initial set of inference hypotheses, in order to distinguish between mainstream consensus hypotheses and non-mainstream hypotheses; The divergence metric is used to characterize the degree of internal cognitive uncertainty of the initial set of inference hypotheses. In the joint reasoning, the integrated reasoning model identifies mainstream consensus assumptions and non-mainstream assumptions based on the degree of consensus concentration and the degree of internal cognitive uncertainty, and perceives the overall cognitive conflict state of the initial reasoning assumption set, so as to guide the adjustment of the integration emphasis of each initial reasoning assumption.

3. The method according to claim 2, wherein, The divergence metric is determined based on the information entropy of the distribution of different answer results in the initial set of inference hypotheses.

4. The method according to claim 1, wherein, The integrated reasoning model optimizes its parameters during the training phase based on the initial inference hypothesis set generated by the multimodal large language model with fixed parameters, and generates the initial inference hypothesis set using the multimodal large language model with fixed parameters during the inference phase.

5. A method for training a visual reasoning model, comprising: Obtain training samples, which include visual data, query questions, labeled answers, and a set of initial inference hypotheses generated by a pre-trained multimodal large language model. Each initial inference hypothesis contains an inference process and answer result for an independent inference of the visual data and the query question. Extract the statistical distribution characteristics of the initial inference hypothesis set of the sample; Using the integrated reasoning model to be trained, with the statistical distribution features as conditional signals, joint reasoning is performed with the visual data, the query question, and the initial inference hypothesis set of the samples to obtain the target reasoning path and the target answer result; The parameters of the integrated reasoning model are updated based on the target answer result, the target reasoning path, the labeled answer, and the statistical distribution characteristics. The statistical distribution features are used to guide the integrated reasoning model to adjust the fusion emphasis of the initial reasoning assumptions for each sample in joint reasoning.

6. The method according to claim 5, wherein, The step of updating the parameters of the integrated inference model based on the target answer result, the target inference path, the labeled answer, and the statistical distribution features includes: The reward signal is calculated based on the target answer result, the target reasoning path, the labeled answer, and the statistical distribution characteristics, and the parameters of the integrated reasoning model are updated through reinforcement learning based on the reward signal.

7. The method according to claim 6, wherein, The reward signal includes a cognitive conflict incentive signal and a reasoning path tracking signal; When the target answer matches the labeled answer, a cognitive conflict incentive signal is generated based on the statistical distribution characteristics and the target answer to incentivize the integrated reasoning model to perform corrective reasoning in a state of cognitive conflict. If the target answer does not match the labeled answer, a reasoning path tracking signal is generated based on the target reasoning path, the initial set of inference hypotheses for the sample, and the labeled answer, to guide the integrated reasoning model to align with a better reasoning path.

8. The method according to claim 7, wherein, The cognitive conflict stimulus signal is generated based on the statistical distribution characteristics and the target response result, including: In the case where the query question corresponds to a closed-ended task, the degree of divergence of the answer results and the scarcity of the target answer results are determined based on the statistical distribution characteristics; the cognitive conflict incentive signal is generated based on the product of the degree of divergence and the scarcity. In the case where the query question corresponds to an open-ended generation task, the accuracy score of the target answer result is obtained, the number of hypotheses with accuracy scores lower than the accuracy score of the target answer result is counted in the initial inference hypothesis set of the sample, and the cognitive conflict incentive signal is generated based on the ratio of the number of hypotheses to the total number of hypotheses in the initial inference hypothesis set of the sample.

9. The method according to claim 8, wherein, The degree of divergence is determined by calculating the normalized information entropy of the distribution of different answer results in the initial inference hypothesis set of the sample.

10. The method according to claim 7, wherein, The inference path tracking signal is generated based on the target inference path, the initial inference hypothesis set of the samples, and the labeled answer, including: Select a subset of advantageous initial inference hypotheses that have higher accuracy than the target answer from the initial set of sample inference hypotheses; Based on the difference between the accuracy of the initial inference hypothesis of each sample in the subset of the dominant samples and the accuracy of the target answer, the weight corresponding to the initial inference hypothesis of each sample is determined. The inference path tracking signal is calculated based on the weights and the similarity between the target inference path and the inference process of each sample's initial inference hypothesis in the subset of dominant samples' initial inference hypotheses.

11. The method according to claim 10, wherein, Before determining the weights corresponding to the initial inference hypotheses of each sample based on the difference between the accuracy of the initial inference hypotheses of each sample in the subset of dominant samples and the accuracy of the target answer, the method further includes: Determine whether the initial inference hypothesis subset of the dominant sample is an empty set; When the initial inference hypothesis subset of the dominant sample is an empty set, the preset maximum reward value is determined as the inference path tracking signal.

12. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1 to 4, or the method of any one of claims 5 to 11.

13. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 4, or the method of any one of claims 5 to 11.

14. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the method of any one of claims 1 to 4, or the method of any one of claims 5 to 11.