Method and device for enhancing reasoning capability of multi-modal reasoning model, storage medium, system and computer program product

By constructing multi-source data sets and decision sets, and using multi-modal inference models to generate and evaluate state transition data, the problem of insufficient inference capabilities of existing multi-modal large models is solved, and more accurate and reliable complex inference task processing capabilities are achieved.

CN119990332AActive Publication Date: 2025-05-13INST OF AUTOMATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510460498.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

When existing multimodal large models deal with complex inference tasks, they lack inference capabilities and are difficult to generate coherent and logically consistent inference chains, which affects the accuracy and reliability of the answers.

Method used

By constructing multi-source data sets and decision sets, using multi-modal inference models to generate state transition data, and determining data quality through multi-modal judgment and evaluation models, model training is carried out to enhance inference capabilities.

Benefits of technology

It significantly improves the inference ability of multimodal large models in visual question-and-answer tasks, ensures the accuracy and reliability of the models when dealing with complex problems, and enhances the adaptability and expressiveness of the models in diverse task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990332A_ABST
    Figure CN119990332A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device, a storage medium, a system and a computer program product for enhancing the reasoning capability of a multi-modal reasoning model. The method comprises the following steps: constructing a multi-source data set; determining a decision set for solving the problem; for each question and answer pair, executing the following processing: inputting the decision set and the current question in the current question and answer pair into a multi-modal reasoning model to obtain first state conversion data of the current question, the first state conversion data being logic data of a pre-estimated answer of the current question based on at least one decision in multiple decisions; based on a real answer in the current question and answer pair, a first data state of the first state conversion data is determined, and the first data state indicates that the first state conversion data is correct data or wrong data; and in response to all the question and answer pairs, executing the processing, and training the multi-modal reasoning model by using all the correct data and the question and answer pairs corresponding to the correct data to obtain a first multi-modal reasoning model with enhanced reasoning ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of artificial intelligence technology, and more specifically, to a method, device, storage medium, system, and computer program product for enhancing the reasoning capability of a multimodal reasoning model. Background Art

[0002] With the rapid development of artificial intelligence technology, multimodal large models have been widely used in visual question answering (VQA) tasks. VQA tasks require large models to be able to comprehensively process image and text information to generate accurate answers. However, existing technologies still face many challenges and limitations when dealing with complex reasoning tasks. For example, existing multimodal large models often show the defect of insufficient reasoning ability when facing complex problems, because traditional large models can usually only process simple visual queries. For complex problems that require deep reasoning, it is difficult to generate a coherent and logically consistent reasoning chain, resulting in the accuracy and reliability of the answer being affected. In this case, the reasoning process of the large model lacks transparency and explainability, and it is difficult to meet the needs of practical applications.

[0003] At present, the Chain-of-Thought (CoT) reasoning method has made some progress in enhancing the reasoning ability of large models, but there are still many challenges in its application in multimodal environments, especially in the zero-shot CoT method, where large models need to reason through prompts without clear examples, which places higher requirements on the logical consistency and reasoning quality of large models. However, when dealing with diverse and complex tasks, existing CoT methods often find it difficult to maintain the logical consistency and high quality of the reasoning process, resulting in unreliable reasoning results. In addition, when dealing with visual question-answering tasks, existing multimodal large models usually lack adaptive reasoning strategies. For example, when facing different types of questions, large models need to be able to dynamically adjust the reasoning path to adapt to the needs of different tasks, but large models in existing technologies often use fixed reasoning paths, lack flexibility and adaptability, and are difficult to cope with diverse task scenarios. This fixed path reasoning method limits the performance of the model in complex tasks and cannot give full play to the potential of multimodal models. Summary of the invention

[0004] The embodiments of the present disclosure provide a method, device, storage medium, system and computer program product for enhancing the reasoning capability of a multimodal reasoning model, which can effectively solve the problem that the prior art model has insufficient reasoning capability and is difficult to cope with diverse task scenarios.

[0005] In a general aspect, a method for enhancing the reasoning ability of a multimodal reasoning model is provided, comprising: constructing a multi-source data set, wherein the multi-source data set includes question-answer pairs of multiple modalities, each question-answer pair including a question and a true answer to the question; determining a decision set for solving the problem, wherein the decision set includes multiple decisions, and an estimated answer to any question can be determined based on at least one of the multiple decisions; for each question-answer pair of the multiple modalities, performing the following processing: inputting the decision set and the current question in the current question-answer pair into the multimodal reasoning model to obtain first state transition data of the current question, wherein the first state transition data is logical data for inferring an estimated answer to the current question based on at least one of the multiple decisions; determining a first data state of the first state transition data based on the true answer in the current question-answer pair, wherein the first data state indicates whether the first state transition data is correct data or incorrect data; in response to all question-answer pairs having completed the above processing, training the multimodal reasoning model using all correct data and the question-answer pairs corresponding to the correct data to obtain a first multimodal reasoning model with enhanced reasoning ability.

[0006] Optionally, the processing performed on each question and answer pair in the question and answer pairs of multiple modalities also includes: in response to the first data state indicating that the first state conversion data is correct data, inputting the first state conversion data into the multimodal evaluation model to obtain the state score and the scoring reason for each state data in the first state conversion data; in response to the average state score of all state data in the first state conversion data being greater than or equal to a preset threshold, determining that the first state conversion data is high-quality data; wherein, after obtaining the first multimodal reasoning model with enhanced reasoning capability, it also includes: using all high-quality data and the question and answer pairs corresponding to the high-quality data to train the first multimodal reasoning model to obtain the second multimodal reasoning model with enhanced reasoning capability.

[0007] Optionally, the processing performed on each question-and-answer pair in the question-and-answer pairs of multiple modalities also includes: in response to the average state score of all state data in the first state transition data being less than a preset threshold, inputting the first state transition data, the state score and the score reason into the multimodal correction model to obtain second state transition data; in response to the second state transition data being high-quality data, determining the first state transition data and the second state transition data as a first training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, it also includes: using all the first training sample pairs to train the second multimodal reasoning model to obtain a third multimodal reasoning model with enhanced reasoning capability.

[0008] Optionally, the processing performed on each question-and-answer pair in the multiple modal question-and-answer pairs also includes: in response to the first data state indicating that the first state conversion data is erroneous data, inputting the first state conversion data into the multimodal correction model to obtain third state conversion data; in response to the third state conversion data being high-quality data, determining the first state conversion data and the third state conversion data as a second training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, it also includes: using all the second training sample pairs to train the second multimodal reasoning model to obtain a fourth multimodal reasoning model with enhanced reasoning capability.

[0009] Optionally, the processing performed on each question-and-answer pair in the multiple modal question-and-answer pairs also includes: in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, inputting the current question into the second multimodal reasoning model to obtain fourth state transition data; in response to the fourth state transition data being erroneous data, determining the fourth state transition data and the first state transition data as a third training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, it also includes: using all third training sample pairs to train the second multimodal reasoning model to obtain a fifth multimodal reasoning model with enhanced reasoning capability.

[0010] Optionally, based on the true answer in the current question and answer pair, determining the first data state of the first state transition data includes: inputting the first state transition data and the true answer in the current question and answer pair into a multimodal judgment model to obtain the first data state of the first state transition data. In another general aspect, a device for enhancing the reasoning ability of a multimodal reasoning model is provided, comprising: a construction unit, configured to construct a multi-source data set, wherein the multi-source data set includes question-answer pairs of multiple modalities, each question-answer pair includes a question and a true answer to the question; a determination unit, configured to determine a decision set for solving the problem, wherein the decision set includes multiple decisions, and an estimated answer to any question can be determined based on at least one of the multiple decisions; an execution unit, configured to perform the following processing for each question-answer pair in the multiple modalities: input the decision set and the current question in the current question-answer pair into the multimodal data set; A multimodal reasoning model is used to obtain first state transition data of the current question, wherein the first state transition data is logical data of an estimated answer to the current question based on at least one decision among multiple decisions; a first data state of the first state transition data is determined based on a true answer in the current question and answer pair, wherein the first data state indicates that the first state transition data is correct data or incorrect data; and a training unit is configured to train the multimodal reasoning model in response to all question and answer pairs having performed the above processing, using all correct data and the question and answer pairs corresponding to the correct data, to obtain a first multimodal reasoning model with enhanced reasoning capability.

[0011] Optionally, the execution unit is further configured to, in response to the first data state indicating that the first state transition data is correct data, input the first state transition data into a multimodal evaluation model to obtain a state score and a scoring reason for each state data in the first state transition data; in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, determine that the first state transition data is high-quality data; wherein the training unit is further configured to, after obtaining the first multimodal reasoning model with enhanced reasoning capability, train the first multimodal reasoning model using all high-quality data and question-answer pairs corresponding to the high-quality data to obtain a second multimodal reasoning model with enhanced reasoning capability.

[0012] Optionally, the execution unit is further configured to, in response to an average state score of all state data in the first state transition data being less than a preset threshold, input the first state transition data, the state score and the score reason into the multimodal correction model to obtain second state transition data; in response to the second state transition data being high-quality data, determine the first state transition data and the second state transition data as a first training sample pair; wherein the training unit is further configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, use all the first training sample pairs to train the second multimodal reasoning model to obtain a third multimodal reasoning model with enhanced reasoning capability.

[0013] Optionally, the execution unit is further configured to perform processing on each question-answer pair in multiple modal question-answer pairs, and also includes: in response to the first data state indicating that the first state conversion data is erroneous data, inputting the first state conversion data into a multimodal correction model to obtain third state conversion data; in response to the third state conversion data being high-quality data, determining the first state conversion data and the third state conversion data as a second training sample pair; wherein the training unit is further configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, train the second multimodal reasoning model using all second training sample pairs to obtain a fourth multimodal reasoning model with enhanced reasoning capability.

[0014] Optionally, the execution unit is also configured to perform processing on each question-answer pair in multiple modal question-answer pairs, and also includes: in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, inputting the current question into the second multimodal reasoning model to obtain fourth state transition data; in response to the fourth state transition data being erroneous data, determining the fourth state transition data and the first state transition data as a third training sample pair; wherein the training unit is also configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, use all third training sample pairs to train the second multimodal reasoning model to obtain a fifth multimodal reasoning model with enhanced reasoning capability.

[0015] Optionally, the execution unit is configured to input the first state transition data and the true answer in the current question and answer pair into the multimodal judgment model to obtain the first data state of the first state transition data. In another general aspect, a computer-readable storage medium storing instructions is provided, wherein when the instructions are executed by at least one computing device, the at least one computing device is prompted to perform any of the methods for enhancing model reasoning capabilities as described above.

[0016] In another general aspect, a system is provided that includes at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, cause the at least one computing device to perform any of the methods for enhancing model reasoning capabilities as described above.

[0017] In another general aspect, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any of the above methods for enhancing model reasoning capabilities.

[0018] According to the method, device, storage medium, system and computer program product for enhancing the reasoning ability of the multimodal reasoning model of the embodiment of the present disclosure, a decision set containing multiple decisions is introduced, so that the model can adaptively adjust the reasoning path, and then flexibly respond to different types of task requirements, overcome the limitations of the traditional model using a fixed reasoning path, so that the model can effectively generate a coherent and logically consistent reasoning chain, and ensure the accuracy and reliability of the model's answers when dealing with complex problems; in addition, the present disclosure uses the correct data output by the multimodal reasoning model and its corresponding question and answer pairs to train the multimodal model, which can significantly improve the reasoning ability of the multimodal large model in the visual question and answer task; the present disclosure enhances the adaptability and expressiveness of the model in diversified task scenarios, and provides more solid technical support for the practical application of multimodal large models. Therefore, through the present disclosure, it is possible to effectively solve the problem that the prior art model has insufficient reasoning ability and is difficult to cope with diversified task scenarios.

[0019] Additional aspects and / or advantages of the present general inventive concept will be set forth in part in the following description and in part will be apparent from the description or may be learned through practice of the present general inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other objects and features of the embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings showing the embodiments, in which: Figure 1 is a flow chart showing a method for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure; Figure 2is a schematic diagram showing the overall flow of a method for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure; Figure 3 is a schematic diagram showing the overall structure of a method for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure; Figure 4 is a diagram showing a comparison of the inference effects of the Llama-CoMMoD and Llama-CoMMoD models according to an embodiment of the present disclosure; Figure 5 is a comparison diagram showing the inference effects of two models of the embodiments of the present disclosure, in which the method of the present disclosure is not applied, and in which the method of the present disclosure is applied; Figure 6 Detailed description is a block diagram showing an apparatus for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following specific embodiments are provided to help the reader obtain a comprehensive understanding of the methods, devices and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear. For example, the order of operations described herein is only an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and simplicity, the description of features known in the art may be omitted.

[0022] The features described herein can be implemented in different forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein have been provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will be clear after understanding the disclosure of the present application.

[0023] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.

[0024] Although terms such as "first", "second", and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are only used to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Therefore, without departing from the teachings of the examples described herein, the first member, first component, first region, first layer, or first portion referred to in the examples may also be referred to as the second member, second component, second region, second layer, or second portion.

[0025] In the specification, when an element (such as a layer, a region, or a substrate) is described as being “on”, “connected to”, or “coupled to” another element, the element may be directly “on”, “connected to”, or “coupled to” another element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on”, “directly connected to”, or “directly coupled to” another element, there may be no other elements present therebetween.

[0026] The terms used herein are only used to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. The terms "comprise", "include" and "have" indicate the presence of the described features, quantities, operations, components, elements and / or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, components, elements and / or combinations thereof.

[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which the present disclosure belongs after understanding the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.

[0028] Furthermore, in the description of examples, when it is considered that a detailed description of a well-known related structure or function would cause vague interpretation of the present disclosure, such a detailed description will be omitted.

[0029] The method and device for enhancing the reasoning capability of a multimodal reasoning model disclosed in the present invention are described in detail below in conjunction with the accompanying drawings.

[0030] This disclosure proposes a method for enhancing the reasoning ability of a multimodal reasoning model. Figure 1 is a flow chart showing a method for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure. Figure 1 , the method for enhancing the reasoning ability of the multimodal reasoning model comprises the following steps: In step S101, a multi-source dataset is constructed, wherein the multi-source dataset contains question-answer pairs of multiple modalities, and each question-answer pair contains a question and a true answer to the question.

[0031] As an example, the question-answer pairs in the above multi-source dataset can be derived from a variety of open source datasets, covering multimodal question-answer pairs of various types of multimodal question-answering tasks, where the various types may include but are not limited to: data questions, general knowledge, table questions, and comprehensive questions, etc. The above multimodal question-answer pairs generally contain images and texts, which are not limited to this disclosure. The form of the above question-answer pairs can include multiple-choice questions, fill-in-the-blank questions, and subjective questions and answers, which are not limited to this disclosure. It should be noted that the question options of the above multiple-choice questions can be randomly arranged to create new question-answer pairs, so that the model can reason about the same question from different angles to obtain the correct answer.

[0032] As an example, the above open source datasets may include but are not limited to: ChartQA, DocVQA, TextVQA, InfoVQA, ScienceQA, Geometry3K, GeoQA, Super-CLEVR and MathV360K, among which ChartQA, DocVQA, TextVQA, InfoVQA and ScienceQA focus on simple visual tasks, with a total of about 18,200 question-answer pairs; Geometry3K, GeoQA and Super-CLEVR are aimed at complex reasoning tasks; MathV360K is a multi-source dataset that includes simple visual tasks and complex reasoning tasks. The multi-source dataset is constructed through the above diverse datasets to ensure that the present disclosure is robust and can handle a wide range of multimodal question-answering tasks.

[0033] In step S102, a decision set for solving the problem is determined, wherein the decision set includes multiple decisions, and the estimated answer to any question can be determined based on at least one decision among the multiple decisions.

[0034] As an example, the above decision set can include multiple decisions, and the estimated answer to any question can be determined based on at least one of the multiple decisions. For example, the above decision set can include twelve different decisions, and these twelve decisions can be systematically divided into four components, namely, problem understanding, problem solving, verification, and answering, wherein each component reflects a key stage of human cognitive processing, thereby enhancing the model's ability to process and respond to complex queries. This structured decision set ensures that each stage of the problem reasoning logic is thoroughly resolved, facilitating a more detailed and accurate response to complex problems, but the present disclosure is not limited to this.

[0035] As an example, the above-mentioned problem understanding stage may include decisions such as identification, analysis, research, and formalization, which focus on the initial understanding and recognition of text and images in the question-answer pair, laying the foundation for more complex reasoning tasks; the above-mentioned problem solving stage may include decisions such as decomposition, planning, solving sub-problems, and solving parent problems, which help the model gradually disassemble complex problems and find solutions; the above-mentioned verification stage includes decisions such as re-examination, verification, and review, which allow the model to self-check, verify, and correct its reasoning process; the above-mentioned answering stage includes a decision, that is, giving the final answer to the question, in which the model synthesizes the entire reasoning process and outputs the answer according to the given paradigm. It should be noted that the decisions included in the above-mentioned stages are not limited to this, and the required decisions can be set as needed.

[0036] In step S103, for each question-answer pair in the multiple modal question-answer pairs, the following processing is performed: the decision set and the current question in the current question-answer pair are input into the multimodal reasoning model to obtain the first state transition data of the current question, wherein the first state transition data is logical data for inferring the estimated answer to the current question based on at least one decision among the multiple decisions; based on the true answer in the current question-answer pair, the first data state of the first state transition data is determined, wherein the first data state indicates whether the first state transition data is correct data or incorrect data.

[0037] As an example, the above decision set can be Integrate into the prompt of the multimodal reasoning model to guide the multimodal reasoning model By autonomously selecting an appropriate decision from a decision set at each state, the multimodal reasoning model generates state transition data that is both coherent and contextually relevant.

[0038] Specifically, it can be expressed as follows: "Describe each decision function and tell the model At each state of the problem solving, any of the above decisions can be made. Integrate into prompt , making it possible to utilize the above multi-source datasets; for each image in the question-answer pair and text , you can use and Adaptively generate state transition data , the process can be expressed as:

[0039] in, represents the number of question-answer pairs, Indicates the current question and answer pair.

[0040] It should be noted that the above-mentioned multimodal reasoning model is a large model, and the basic structure adopts GPT-4o, which is not limited in this disclosure.

[0041] As an example, after obtaining the state transition data of the current question, it can be compared with the real answer in the current question-answer pair to determine whether the first state transition data is correct data or wrong data. The comparison between the state transition data and the real answer can be performed manually or by using a model, which is not limited in this disclosure.

[0042] According to an embodiment of the present disclosure, the first state transition data and the true answer in the current question-answer pair can be input into a multimodal evaluation model to obtain the first data state of the first state transition data. Through this embodiment, the multimodal evaluation model is introduced to accurately judge whether the first state data is correct, thereby obtaining an accurate first data state.

[0043] As an example, the multimodal evaluation model may also be a large model, such as Qwen2-VL-7B, and the evaluation model may be trained in advance, which is not limited in the present disclosure.

[0044] Specifically, the first state transition data and the true answer in the current question-answer pair can be input into the multimodal evaluation model, and the multimodal evaluation model will Compare the real answer with the previous question and answer Compare and output status data and ,in, Indicates that the first state transition data is correct data, Indicates that the first state transition data is erroneous data, which can be expressed as follows:

[0045] in, Represents the hint of the multimodal judgment model, represents the number of question-answer pairs, Indicates the current question and answer pair.

[0046] In step S104, in response to the above processing being performed on all question-answer pairs, the multimodal reasoning model is trained using all correct data and the question-answer pairs corresponding to the correct data to obtain a first multimodal reasoning model with enhanced reasoning capability.

[0047] According to an embodiment of the present disclosure, the processing performed on each question and answer pair in a question and answer pair of multiple modalities may also include: in response to the first data state indicating that the first state conversion data is correct data, inputting the first state conversion data into a multimodal evaluation model to obtain a state score and a scoring reason for each state data in the first state conversion data; in response to the average state score of all state data in the first state conversion data being greater than or equal to a preset threshold, determining that the first state conversion data is high-quality data; wherein, after obtaining the first multimodal reasoning model with enhanced reasoning capability, the first multimodal reasoning model may also be trained using all high-quality data and the question and answer pairs corresponding to the high-quality data to obtain a second multimodal reasoning model with enhanced reasoning capability.

[0048] Through this embodiment, a state evaluation mechanism is introduced to determine the high-quality data in the correct answer, and the multimodal model is further trained using the high-quality data, so that the trained model can maintain high-quality information transmission during the learning process, avoiding information loss and logical inconsistency problems.

[0049] As an example, the multimodal evaluation model can also be a large model, such as GPT-4o, and the evaluation model can be trained in advance, which is not limited in this disclosure. The predetermined threshold can be set as needed, such as 0.8 or 0.7, which is not limited in this disclosure.

[0050] Specifically, a multimodal evaluation model can be used Evaluating the right data Each state in the problem is scored from -1 to 1 based on its positive or negative impact on problem solving, and a reason for the score is provided; then, manual verification can be performed, such as calculating The average score of all states in , if the average score does not exceed 0.8, the correct data is marked as low-quality data If the average score is equal to or greater than 0.8, the correct data is marked as high-quality data. , which can be expressed as follows: ) + Human Verification in, Represents a hint for the multimodal evaluation model.

[0051] Then, using this high-quality data, we can obtain high-quality SFT data. , which can be expressed as follows:

[0052] It should be noted that the full name of SFT is Supervised Fine-Tuning, which means supervised fine-tuning.

[0053] According to an embodiment of the present disclosure, the processing performed on each question-answer pair in multiple modal question-answer pairs may also include: in response to the average state score of all state data in the first state transition data being less than a preset threshold, inputting the first state transition data, the state score and the score reason into a multimodal correction model to obtain second state transition data; in response to the second state transition data being high-quality data, determining the first state transition data and the second state transition data as a first training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, all the first training sample pairs may also be used to train the second multimodal reasoning model to obtain a third multimodal reasoning model with enhanced reasoning capability.

[0054] Through this embodiment, the low-quality data generated in the previous training process is used to obtain the corresponding second state transition data through the multimodal correction model, and the high-quality second state transition data and the low-quality first state transition data are selected to continue training the second multimodal reasoning model, so as to obtain a model with better reasoning ability, which can avoid generating the above-mentioned low-quality data.

[0055] As an example, the multimodal correction model can also be a large model, such as GPT-4o, and the evaluation model can be trained in advance, which is not limited in the present disclosure.

[0056] As an example, when determining that the first state data is low-quality data When low-quality data and status evaluation results (rating and rating rationale for each status) are added to the prompt and feed the prompt into the multimodal correction model In the process, the second state transition data is generated. If the second state transition data is high-quality data, the second state transition data is retained. If the second state transition data is low-quality data, the second state transition data and its corresponding question-answer pair are deleted. Specifically, it can be expressed as:

[0057] in, Hints representing multimodal correction models.

[0058] After obtaining high-quality second state transition data, the second state transition data and low-quality first state transition data can be used as a training pair to continue training the above-mentioned second multimodal reasoning model to obtain a third multimodal reasoning model with better reasoning ability.

[0059] It should be noted that the method for judging the quality of the second state transition data may adopt the method mentioned in the above embodiment, and the present disclosure does not limit this.

[0060] According to an embodiment of the present disclosure, the processing performed on each question-answer pair in multiple modal question-answer pairs may also include: in response to the first data state indicating that the first state conversion data is erroneous data, inputting the first state conversion data into a multimodal correction model to obtain third state conversion data; in response to the third state conversion data being high-quality data, determining the first state conversion data and the third state conversion data as a second training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, all second training sample pairs can also be used to train the second multimodal reasoning model to obtain a fourth multimodal reasoning model with enhanced reasoning capability.

[0061] Through this embodiment, the erroneous data generated in the previous training process is used to obtain the corresponding third state transition data through the multimodal correction model, and the high-quality third state transition data and erroneous data are selected to continue training the second multimodal reasoning model, so as to obtain a model with better reasoning ability, which can avoid generating the above-mentioned erroneous data.

[0062] As an example, in the first state, the conversion data is error data. In this case, the error data can be Input the multimodal correction model, the multimodal correction model identifies the wrong step and continues the previous correct step to generate the third state transition data. If the third state transition data is the correct answer The third state transition data is retained. If the third state transition data is an incorrect answer, the third state transition data and its corresponding question-answer pair are deleted. Specifically, it can be expressed as:

[0063] in, Hints representing multimodal correction models.

[0064] After obtaining the correct third state transition data, the third state transition data and the erroneous first state transition data can be used as a training pair to continue training the above-mentioned second multimodal reasoning model to obtain a fourth multimodal reasoning model with better reasoning ability.

[0065] It should be noted that the method for judging whether the third state transition data is correct or not may adopt the method mentioned in the above embodiment, and the present disclosure does not limit this.

[0066] According to an embodiment of the present disclosure, the processing performed on each question-answer pair in a question-answer pair of multiple modalities may also include: in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, inputting the current question into the second multimodal reasoning model to obtain fourth state transition data; in response to the fourth state transition data being erroneous data, determining the fourth state transition data and the first state transition data as a third training sample pair; wherein, after obtaining the second multimodal reasoning model with enhanced reasoning capability, all third training sample pairs may also be used to train the second multimodal reasoning model to obtain a fifth multimodal reasoning model with enhanced reasoning capability.

[0067] Through this embodiment, the high-quality data generated in the previous training process is input as test data into the first multimodal reasoning model to obtain the corresponding fourth state transition data, and the erroneous fourth state transition data and the high-quality first state transition data are selected to continue training the second multimodal reasoning model. The trained model can avoid generating the above-mentioned erroneous fourth state transition data, further optimizing the model and making the model's reasoning ability even better.

[0068] As an example, when determining that the first state data is high-quality data When the high-quality data is input into the second multimodal reasoning model, the question contained in the question-answer pair corresponding to the high-quality data is input into the second multimodal reasoning model. , get the fourth state transition data, if the fourth state transition data is error data , then keep the fourth state transition data, if the fourth state transition data is correct data, then delete the fourth state transition data. The above processing is performed, which can be specifically expressed as:

[0069] in, Hints representing multimodal correction models.

[0070] After obtaining the erroneous fourth state transition data, the fourth state transition data and the high-quality first state transition data are combined into a training pair, and the above-mentioned second multimodal reasoning model is further trained to obtain a fifth multimodal reasoning model with better reasoning ability.

[0071] It should be noted that the first training pair, the second training pair and the third training pair can be combined to obtain high-quality DPO data. . This process is formalized in the following equation:

[0072] The full English spelling of DPO is Direct Preference Optimization. Then, through three sets of training (i.e. ) The second multimodal model is trained simultaneously to obtain a multimodal reasoning model with further enhanced reasoning ability.

[0073] In order to better understand the present disclosure, Figure 2 and Figure 3 Provide a description of the system.

[0074] Figure 2 The overall process of the present disclosure is shown, as Figure 2 As shown, the overall process includes the following steps: S201: Integrate multiple open source datasets to ensure data diversity, and rearrange the options of multiple-choice questions to create new data so that the model can reason from different perspectives to derive the correct answer.

[0075] S202: Based on the integration results of the previous step, the decision set is defined as four stages: problem understanding, problem solving, verification and answering. Among them, the problem understanding stage focuses on the initial understanding and recognition of pictures and texts; the problem solving stage achieves step-by-step reasoning through decomposition and planning; the verification stage allows the model to self-check, check and correct its reasoning process; the answering stage integrates the reasoning process to output the final answer.

[0076] S203: The decision set of the previous step is embedded into the prompt template, driving the multimodal reasoning model to dynamically select decisions from the decision set and generate state transition data according to the image and text in the input question, that is, by adaptively generating a coherent chain of thoughts, a context-related reasoning path is realized, ensuring that the generated state transition data is deeply adapted to the multimodal task.

[0077] S204: Construct high-quality SFT data: Introduce the state transition data obtained in the previous step of the multimodal judgment model and compare it with the real answer in the question-answer pair to determine whether the state transition data is correct data or incorrect data. Then, use the state evaluation mechanism to quantify the score (-1 to 1) for each state of the correct data, and manually calculate the average of the states. Mark the state transition data with an average score ≥ 0.8 as high-quality SFT data, and form a high-quality supervised fine-tuning dataset based on the SFT data.

[0078] S205: Construct high-quality DPO data through three parts: 1) For the erroneous data in step S204, correct the erroneous steps and generate the corresponding correct answers following the erroneous steps. The specific steps have been described in detail above and will not be discussed in detail here; then, the erroneous data and the corresponding correct data are regarded as a set of training pairs. 2) For the low-quality data that does not reach 0.8 in step S204, regenerate the optimized answer in combination with the evaluation results, that is, generate the corresponding high-quality data. The specific steps have been described in detail above and will not be discussed in detail here; then, the low-quality data and the corresponding high-quality data form a set of training pairs. 3) For SFT data, use the SFT model to re-reason the problem corresponding to the SFT data to obtain new state transition data, and extract the erroneous state transition data and the corresponding SFT data to form a pair of training pairs, where the SFT model is the model trained using the SFT data, that is, the second multimodal reasoning model mentioned above. .

[0079] S206, model training: after obtaining the correct data in S204, the correct data can be used to train the multimodal reasoning model to obtain a first multimodal reasoning model with enhanced reasoning ability; then the first multimodal reasoning model can be further trained with high-quality data to obtain a second multimodal reasoning model with further enhanced reasoning ability; then, the second multimodal reasoning model can be trained separately using three groups of training pairs to obtain a third multimodal reasoning model, a fourth multimodal reasoning model and a fifth multimodal reasoning model with further enhanced reasoning ability; finally, the second multimodal reasoning model can be trained simultaneously using three groups of training pairs to obtain a multimodal reasoning model with further enhanced reasoning ability.

[0080] Figure 3 The overall structure of the present disclosure is shown as Figure 3 As shown, the overall structure of the present disclosure may include the following parts: a state transition data generation part, a high-quality SFT data construction part, a high-quality DPO data construction part and a model training part, wherein the state transition data generation part includes constructing a multi-task data set, determining a decision set, and generating and judging state transition data, and the high-quality SFT data construction part includes evaluating the state transition data to obtain corresponding scores and reasons for the scores; the high-quality DPO construction part includes the correction of erroneous data and low-quality data. It should be noted that this part also includes the re-inference of the SFT data, but it is not shown; the model training part includes training the multimodal reasoning model using SFT data and re-training the SFT model using DPO data. It should be noted that this part also includes the training of other data, but it is not shown.

[0081] In summary, the present disclosure proposes a new method for enhancing the reasoning ability of multimodal large models, which can be called CoMMoD (Controlled Multi-Modal Decision-Making for Chain-of-ThoughtReasoning in Multi-Modal Language Models). By introducing the decision set and state evaluation mechanism of adaptive problem solving, the reasoning ability of multimodal large models in visual question answering tasks is significantly improved. Compared with the prior art, the present disclosure can effectively generate a coherent and logically consistent reasoning chain to ensure the accuracy and reliability of the answers of the trained model when dealing with complex problems; moreover, by adaptively adjusting the reasoning path and ensuring the integrity of the intermediate reasoning steps, the trained model can flexibly respond to different types of task requirements, overcoming the limitations of the fixed reasoning path of the traditional model; in addition, the introduction of the state evaluation mechanism enables the trained model to maintain high-quality information transmission during the learning process, avoiding information loss and logical inconsistency; therefore, the present disclosure enhances the adaptability and expressiveness of the model in diverse task scenarios, and provides more solid technical support for the practical application of multimodal large models.

[0082] Figure 4 The state transition data output by the Llama-CoMMoD and Llama-CoMMoD models for the same question are shown. It can be seen that after applying the method disclosed in the present invention, coherent and logically consistent state transition data can be generated, and accurate and reliable answers are given.

[0083] Figure 5 The following is a comparison chart of the inference effects of two models without the application of the disclosed method and the model with the application of the disclosed method, as shown in FIG. Figure 5 As shown, the model to which the disclosed method is applied provides coherent and logically consistent state transition data, and gives accurate and reliable answers.

[0084] The experimental results show that compared with the existing technology, the present invention significantly improves the model's reasoning ability and answer accuracy in complex scenarios, verifying the effectiveness and superiority of the method of the present invention.

[0085] Figure 6 is a block diagram showing an apparatus for enhancing the reasoning capability of a multimodal reasoning model according to an embodiment of the present disclosure, such as Figure 6 As shown, the device includes a construction unit 60 , a determination unit 62 , an execution unit 64 and a training unit 66 .

[0086] A construction unit 60 is configured to construct a multi-source data set, wherein the multi-source data set includes question-answer pairs of multiple modalities, and each question-answer pair includes a question and a true answer to the question; a determination unit 62 is configured to determine a decision set for solving the problem, wherein the decision set includes multiple decisions, and the estimated answer to any question can be determined based on at least one of the multiple decisions; an execution unit 64 is configured to perform the following processing for each question-answer pair in the multiple modalities: input the decision set and the current question in the current question-answer pair into the multi-modal reasoning model to obtain first state transition data of the current question, wherein the first state transition data is logical data for inferring the estimated answer to the current question based on at least one of the multiple decisions; based on the true answer in the current question-answer pair, determine a first data state of the first state transition data, wherein the first data state indicates that the first state transition data is correct data or incorrect data; a training unit 66 is configured to, in response to all question-answer pairs having completed the above processing, train the multi-modal reasoning model using all correct data and the question-answer pairs corresponding to the correct data to obtain a first multi-modal reasoning model with enhanced reasoning capability.

[0087] According to an embodiment of the present disclosure, the execution unit 64 is further configured to, in response to the first data state indicating that the first state transition data is correct data, input the first state transition data into the multimodal evaluation model to obtain a state score and a scoring reason for each state data in the first state transition data; in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, determine that the first state transition data is high-quality data; wherein the training unit 66 is further configured to, after obtaining the first multimodal reasoning model with enhanced reasoning capability, train the first multimodal reasoning model using all high-quality data and question-answer pairs corresponding to the high-quality data to obtain a second multimodal reasoning model with enhanced reasoning capability.

[0088] According to an embodiment of the present disclosure, the execution unit 64 is further configured to, in response to an average state score of all state data in the first state transition data being less than a preset threshold, input the first state transition data, the state score and the score reason into the multimodal correction model to obtain second state transition data; in response to the second state transition data being high-quality data, determine the first state transition data and the second state transition data as a first training sample pair; wherein the training unit 66 is further configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, train the second multimodal reasoning model using all the first training sample pairs to obtain a third multimodal reasoning model with enhanced reasoning capability.

[0089] According to an embodiment of the present disclosure, the execution unit 64 is also configured to perform processing on each question-answer pair in multiple modal question-answer pairs, and also includes: in response to the first data state indicating that the first state conversion data is erroneous data, inputting the first state conversion data into a multimodal correction model to obtain third state conversion data; in response to the third state conversion data being high-quality data, determining the first state conversion data and the third state conversion data as a second training sample pair; wherein the training unit 66 is also configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, train the second multimodal reasoning model using all second training sample pairs to obtain a fourth multimodal reasoning model with enhanced reasoning capability.

[0090] According to an embodiment of the present disclosure, the execution unit 64 is also configured to perform processing on each question-answer pair in multiple modal question-answer pairs, and also includes: in response to the average state score of all state data in the first state transition data being greater than or equal to a preset threshold, inputting the current question into the second multimodal reasoning model to obtain fourth state transition data; in response to the fourth state transition data being erroneous data, determining the fourth state transition data and the first state transition data as a third training sample pair; wherein the training unit 66 is also configured to, after obtaining the second multimodal reasoning model with enhanced reasoning capability, train the second multimodal reasoning model using all third training sample pairs to obtain a fifth multimodal reasoning model with enhanced reasoning capability.

[0091] According to an embodiment of the present disclosure, the execution unit 64 is configured to input the first state transition data and the true answer in the current question and answer pair into the multimodal judgment model to obtain the first data state of the first state transition data. According to an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute a method for enhancing the reasoning capability of a multimodal reasoning model as described in any of the above embodiments.

[0092] According to an embodiment of the present disclosure, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, prompt the at least one computing device to execute a method for enhancing the reasoning capability of a multimodal reasoning model as in any of the above embodiments.

[0093] According to an embodiment of the present disclosure, a computer program product is provided, including computer instructions, which, when executed by a processor, implement any of the above methods for enhancing the reasoning capability of a multimodal reasoning model.

[0094] Although some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that modifications may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for enhancing the reasoning capability of a multimodal reasoning model, characterized in that: include: Constructing a multi-source dataset, wherein the multi-source dataset contains question-answer pairs of multiple modalities, and each question-answer pair contains a question and a true answer to the question; Determine a decision set for solving the problem, wherein the decision set includes multiple decisions, and an estimated answer to any question can be determined based on at least one decision among the multiple decisions; For each question-answer pair in the multiple modal question-answer pairs, the following processing is performed: Inputting the decision set and the current question in the current question-answer pair into a multimodal reasoning model to obtain first state transition data of the current question, wherein the first state transition data is logical data for inferring an estimated answer to the current question based on at least one decision among the multiple decisions; Based on the true answer in the current question-answer pair, determining a first data state of the first state transition data, wherein the first data state indicates whether the first state transition data is correct data or incorrect data; In response to all question-answer pairs having completed the above processing, the multimodal reasoning model is trained using all correct data and the question-answer pairs corresponding to the correct data to obtain a first multimodal reasoning model with enhanced reasoning capability.

2. The method according to claim 1, characterized in that The processing performed on each question-answer pair in the multiple modal question-answer pairs further includes: In response to the first data state indicating that the first state conversion data is correct data, the first state conversion data is input into a multimodal evaluation model to obtain a state score and a score reason for each state data in the first state conversion data; In response to an average state score of all state data in the first state transition data being greater than or equal to a preset threshold, determining that the first state transition data is high-quality data; After obtaining the first multimodal reasoning model with enhanced reasoning capability, the method further includes: The first multimodal reasoning model is trained using all high-quality data and question-answer pairs corresponding to the high-quality data to obtain a second multimodal reasoning model with enhanced reasoning capability.

3. The method according to claim 2, characterized in that The processing performed on each question-answer pair in the multiple modal question-answer pairs further includes: In response to an average state score of all state data in the first state transition data being less than the preset threshold, inputting the first state transition data, the state score and the score reason into a multimodal correction model to obtain second state transition data; In response to the second state transition data being high-quality data, determining the first state transition data and the second state transition data as a first training sample pair; After obtaining the second multimodal reasoning model with enhanced reasoning capability, the method further includes: The second multimodal reasoning model is trained using all first training sample pairs to obtain a third multimodal reasoning model with enhanced reasoning capability.

4. The method according to claim 2, characterized in that The processing performed on each question-answer pair in the multiple modal question-answer pairs further includes: In response to the first data state indicating that the first state conversion data is erroneous data, inputting the first state conversion data into a multimodal correction model to obtain third state conversion data; In response to the third state transition data being high-quality data, determining the first state transition data and the third state transition data as a second training sample pair; After obtaining the second multimodal reasoning model with enhanced reasoning capability, the method further includes: The second multimodal reasoning model is trained using all second training sample pairs to obtain a fourth multimodal reasoning model with enhanced reasoning capability.

5. The method according to claim 2, characterized in that The processing performed on each question-answer pair in the multiple modal question-answer pairs further includes: In response to an average state score of all state data in the first state transition data being greater than or equal to the preset threshold, inputting the current question into the second multimodal reasoning model to obtain fourth state transition data; In response to the fourth state transition data being erroneous data, determining the fourth state transition data and the first state transition data as a third training sample pair; After obtaining the second multimodal reasoning model with enhanced reasoning capability, the method further includes: The second multimodal reasoning model is trained using all third training sample pairs to obtain a fifth multimodal reasoning model with enhanced reasoning capability.

6. The method according to claim 1, characterized in that The determining, based on the true answer in the current question-answer pair, the first data state of the first state conversion data comprises: The first state transition data and the true answer in the current question-answer pair are input into a multimodal judgment model to obtain a first data state of the first state transition data.

7. A device for enhancing the reasoning capability of a multimodal reasoning model, characterized in that: include: A construction unit is configured to construct a multi-source dataset, wherein the multi-source dataset includes question-answer pairs of multiple modalities, and each question-answer pair includes a question and a true answer to the question; A determination unit, configured to determine a decision set for solving the problem, wherein the decision set includes multiple decisions, and an estimated answer to any question can be determined based on at least one decision among the multiple decisions; The execution unit is configured to perform the following processing for each question-answer pair in the multiple modalities: Inputting the decision set and the current question in the current question-answer pair into a multimodal reasoning model to obtain first state transition data of the current question, wherein the first state transition data is logical data for inferring an estimated answer to the current question based on at least one decision among the multiple decisions; Based on the true answer in the current question-answer pair, determining a first data state of the first state transition data, wherein the first data state indicates whether the first state transition data is correct data or incorrect data; The training unit is configured to train the multimodal reasoning model in response to all question-answer pairs having completed the above processing, using all correct data and the question-answer pairs corresponding to the correct data, to obtain a first multimodal reasoning model with enhanced reasoning capability.

8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the method for enhancing the reasoning capability of a multimodal reasoning model as claimed in any one of claims 1 to 6.

9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, the at least one computing device is prompted to perform the method for enhancing the reasoning capability of a multimodal reasoning model as claimed in any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method for enhancing the reasoning capability of a multimodal reasoning model as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Question and answer pair scoring model training method and device, equipment and storage medium

    CN114528391A

  • Dialectical knowledge transfer learning method and device, electronic equipment and storage medium

    CN118246535A

  • Large language model reasoning optimization method and system based on error strategy feedback

    CN119358663A

  • Question and answer task processing model training method and device, equipment and storage medium

    CN119493849A

  • Answer generation method, device and system, electronic equipment and nonvolatile storage medium

    CN119740648A

Cited By

  • Question and answer model training method and question and answer task processing method

    CN120744074A