Big language model multi-modal reasoning method and device based on decoding guidance

By deconstructing the reasoning problem into a set of sub-problems and using a decoding algorithm to filter the answers, the hallucination phenomenon and high cost problems of multimodal large models are solved, and efficient and accurate visual reasoning is achieved, which is suitable for fields such as intelligent assistants, teaching assistance and autonomous driving.

CN120633830APending Publication Date: 2025-09-12WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510620717.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing large multimodal models suffer from hallucinations in visual reasoning tasks, causing the generated image descriptions to contain misleading information. Existing methods also rely on manual labeling and computing power, which is costly.

Method used

A decoding-guided large language model multimodal reasoning method is adopted. By deconstructing the reasoning problem into a set of sub-problems, and using the beam search decoding algorithm to generate candidate word units, the confidence is calculated by combining joint conditional probability and content consistency to screen the final answer.

Benefits of technology

It improves the accuracy and robustness of visual information extraction, reduces the cost of reasoning learning, realizes zero-sample reasoning in complex visual scenes, and enhances the accuracy and generalization ability of reasoning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633830A_ABST
    Figure CN120633830A_ABST
Patent Text Reader

Abstract

The invention provides a large language model multi-modal reasoning method and device based on decoding guidance, and belongs to the field of natural language process.The method comprises the steps that picture description is generated based on a target problem and a corresponding target image, and the target problem is deconstructed into a sub-problem set according to the picture description; traversing the sub-question set, generating an answer by adopting a cluster search decoding algorithm to obtain a plurality of candidate sub-answers corresponding to each sub-question, calculating confidence, and determining the candidate sub-answer with the highest confidence as the sub-answer corresponding to the sub-question; and constructing a multi-modal reasoning prompt based on the sub-question-sub-answer pair, and inputting the multi-modal reasoning prompt to a large language model for reasoning to obtain a reasoning answer. Therefore, error accumulation of a multi-modal large model is relieved, the robustness of wrong visual information during large model reasoning is enhanced, the accuracy of the visual information is guaranteed, the final reasoning effect is effectively improved, a training data set does not need to be constructed manually, and the reasoning learning cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and in particular relates to a decoding-guided large language model multimodal reasoning method and device. Background Art

[0002] Multimodal reasoning is an important direction in the field of natural language processing. Its goal is to perceive information from multiple modalities, such as vision and language, and to reason and draw conclusions. Application scenarios involve many fields such as intelligent assistants, teaching assistance, and autonomous driving. Traditional pre-trained visual language models are limited by the scale of training data and model parameters. Their language expression capabilities are insufficient, and their understanding and generalization capabilities of complex visual scenes are limited. When deployed in new visual scenarios, retraining is usually required. With the development of large language models and large multimodal models, large multimodal models have demonstrated better generalization and language expression capabilities in visual reasoning tasks. However, due to insufficient modal alignment and imbalance in inter-modal capabilities, large multimodal models suffer from serious hallucination phenomena, that is, they often generate text that is inconsistent with the input image, making it difficult to generate high-quality long-term reasoning content.

[0003] The existing technology usually combines a large language model with a large multimodal model, using the large multimodal model for visual perception and generating image descriptions, and then the large language model performs reasoning. However, this method is limited by the hallucination problem of the large multimodal model. The obtained image descriptions often contain a lot of misleading information, resulting in inaccurate reasoning results.

[0004] Therefore, in order to improve the accuracy of inference results, the existing approach is to fine-tune large multimodal models to enhance their reasoning capabilities, such as using reinforcement learning based on human preferences or retraining regional-level image-text pairs. However, such methods rely on manual labeling and computing power, which is very costly. Summary of the Invention

[0005] In order to solve the problem of high cost caused by improving the accuracy of reasoning results in the prior art, the present invention provides a large language model multimodal reasoning method, device, computer equipment and storage medium based on decoding guidance.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] First, a decoding-guided multimodal reasoning method for large language models is provided, including:

[0008] Generate an image description based on the target question and the corresponding target image, and deconstruct the target question into a set of sub-questions based on the image description;

[0009] Traverse the set of sub-questions, input each sub-question into the multimodal large model in turn, use the beam search decoding algorithm to generate multiple related word-grams as candidate word-grams based on the sub-questions, calculate the probability distribution of the candidate word-grams, and generate sequences based on the candidate word-grams with the highest joint conditional probabilities. Decoding terminates when all sequences have generated end words or reached the maximum generation length, determine the joint conditional probability of each sequence, and output sub-answers based on the joint conditional probabilities.

[0010] Based on the sub-question-sub-answer pairs, multimodal reasoning prompts are constructed and input into the large language model for reasoning to obtain the reasoning answer.

[0011] Optionally, outputting a sub-answer according to the joint conditional probability includes:

[0012] Output multiple candidate sub-answers according to the joint conditional probability;

[0013] The confidence of each candidate sub-answer is obtained by performing a weighted summation based on the joint conditional probability, initial conditional probability, generated sequence length, and content consistency score of each candidate sub-answer; the content consistency score is obtained by calculating the Levenshtein distance between the two sequences;

[0014] The candidate sub-answer with the highest confidence is used as the sub-answer corresponding to the sub-question.

[0015] Optionally, the method generates a picture description based on the target question and the corresponding target image, and deconstructs the target question into a set of sub-questions according to the picture description, including: inputting the target image into a large multimodal model to generate a picture description describing the target image; inputting the target question and the picture description into the large language model to generate multiple sub-questions.

[0016] Optionally, the method also includes: when the confidence of the inference answer does not reach a preset confidence, re-deconstructing the target problem into a set of sub-problems, and using the multimodal large model and the large language model to perform iterative reasoning until the confidence of the obtained inference answer reaches a preset confidence or the number of iterations reaches a maximum number.

[0017] Optionally, for non-first round planning, deconstructing the target question into a set of sub-questions based on the image description includes: inputting the target question, the image description, the set of sub-questions in the previous round, and the set of corresponding answers into the large language model to generate multiple sub-questions.

[0018] Optionally, during the first round of planning, the formula for reasoning planning of the large language model is:

[0019]

[0020] For non-first-round planning, the formula for reasoning planning using the large language model is:

[0021]

[0022] Among them, P LLM Indicates the probability distribution generated by LLM, Prompt s Indicates the prompt template for reasoning planning, sq i represents the i-th subproblem, sq 1:(i-1) represents the first i-1 sub-questions generated, c represents the image description, q represents the reasoning question, A represents the answer option set, SubQ 1:(t-1) SubA represents the set of subproblems in the first t-1 rounds. 1:(t-1) Indicates that it corresponds to SubQ 1:(t-1) A collection of answers.

[0023] Secondly, a decoding-guided large language model multimodal reasoning device is provided, the device comprising:

[0024] The generation module generates an image description based on the target question and the corresponding target image, and deconstructs the target question into a set of sub-questions based on the image description;

[0025] The answer module is used to traverse the sub-question set, input each sub-question into the multimodal large model in sequence, use the beam search decoding algorithm to generate multiple related word-grams as candidate word-grams based on the sub-questions, calculate the probability distribution of the candidate word-grams, and generate sequences based on the candidate word-grams with the highest joint conditional probabilities. Decoding terminates when all sequences have generated end words or reached the maximum generation length, determines the joint conditional probability of each sequence, and outputs the sub-answer based on the joint conditional probabilities;

[0026] The reasoning module is used to construct multimodal reasoning prompts based on sub-question-sub-answer pairs, input them into the large language model for reasoning, and obtain the reasoning answer.

[0027] In addition, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned large language model multimodal reasoning method based on decoding guidance.

[0028] Finally, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned large language model multimodal reasoning method based on decoding guidance is implemented.

[0029] The decoding-guided large language model multimodal reasoning method provided by the present invention has the following beneficial effects:

[0030] First, the present invention structures the reasoning problem into a set of more fine-grained sub-problems through inference planning, making the overall extracted visual text information richer and more accurate, alleviating the error accumulation phenomenon of the multimodal large model when generating long reasoning content, and enhancing the robustness of the large model to erroneous visual information during reasoning. Secondly, the present invention uses a decoding algorithm to enhance the reasoning response of the multimodal large model, and screens the reasoning answers by performing confidence calculations from the perspectives of joint probability, sequence length, and content consistency, further ensuring the accuracy of the visual information and effectively improving the final reasoning effect. Finally, the present invention performs multimodal reasoning based on a large language model and a decoding algorithm, which can achieve zero-sample reasoning in different complex visual scenes. It does not require manpower to construct a training data set, greatly reducing the cost of reasoning learning and having good generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] To more clearly illustrate the embodiments of the present invention and its design, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0032] Figure 1 A schematic diagram of a multimodal reasoning process provided by the present invention according to an exemplary embodiment.

[0033] Figure 2 The present invention provides a flowchart of a decoding-guided large language model multimodal reasoning method according to an exemplary embodiment of the present invention.

[0034] Figure 3 This is a block diagram of a decoding-guided large language model multimodal reasoning device provided by the present invention according to an exemplary embodiment. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the technical solution of the present invention and to be able to implement it, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.

[0036] Since fine-tuning a large multimodal model requires expensive labeled data and computing power, the present invention combines a large language model with a large multimodal model to perform multimodal reasoning under zero-sample conditions. Figure 1As shown, to improve the accuracy of visual information extracted by a large multimodal model, this invention, based on the divide-and-conquer principle, utilizes language planning to deconstruct the reasoning problem into multiple sub-problems, thereby mitigating the error accumulation caused by hallucinations during long inference processes in the large multimodal model. To further obtain reliable visual information, this invention combines a decoding algorithm with a beam search to obtain a set of candidate answers during the decoding process of the large multimodal model. The confidence of the candidate answers is calculated from multiple perspectives, such as cumulative probability, answer length, and repetition frequency, to obtain the answer with the highest confidence. Finally, the multiple sub-question-answer pairs are input into the large language model to obtain the inference answer.

[0037] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0038] First, the present invention provides a large language model multimodal reasoning method based on decoding guidance, specifically as follows Figure 2 As shown, the following steps are included:

[0039] S201. Generate an image description based on the target question and the corresponding target image, and deconstruct the target question into a set of sub-questions according to the image description.

[0040] The target problem and the corresponding target image are input into the multimodal large model. The multimodal large model MLLM generates the corresponding image description c for the image I of the target reasoning problem. The formula is as follows:

[0041] c~P MLLM (c|Prompt c (I))

[0042] Among them, Prompt c Represents the instruction template used to generate image description, P MLLM Represents the probability distribution generated by the MLLM. Image descriptions are generated in an autoregressive manner. When the end token appears, the image description is generated. The multimodal MLLM model can be one of several different models such as GPT-4V, LLaVA, or Claude 3.5Sonnet.

[0043] The target question q and the corresponding image description c are input into the large language model LLM, which can be one of the GPT series (such as GPT-3, GPT-4, ChatGPT, etc.), PaLM, LLaMA and BLOOM models. The large language model is based on the target question q, the answer option set A = {a1,…,a n} and image description c for reasoning planning, deconstructing the target problem into a set of sub-problems.

[0044] Among them, reasoning planning refers to deconstructing the main problem q into several sub-problems sq based on the divide-and-conquer idea.i , to obtain the sub-problem set SubQ={sq1,…,sq m In this step, the target problem is the main problem, and a set of subproblems is obtained by performing reasoning planning on the target problem.

[0045] Reasoning planning can simplify the original problem from difficult to complex, obtaining fine-grained reasoning sub-problems. When subsequently input into a large multimodal model, the answers required by the reasoning sub-problems are simpler and shorter. This can effectively alleviate the error accumulation phenomenon when the large multimodal model outputs long texts, helping to improve robustness and accuracy.

[0046] Furthermore, since performing reasoning planning only once may result in the final answer not meeting the expected results, this step may need to be performed multiple times, dividing the target question from different perspectives and iterating to obtain the expected answer. During the first round of planning, the large language model generates several sub-questions based on the target question, the set of answer options, and the image description. For non-first-round planning, the large language model's input also includes the set of sub-questions and their corresponding answers from the previous round.

[0047] Specifically, in the first round of planning, the formula for reasoning planning of the large language model is as follows, where SubQ1 is the set of sub-problems in the first round of planning, sq i For the generated i-th sub-problem, Prompt s Represents the hint template for reasoning planning, P LLM represents the probability distribution generated by LLM, sq 1:(i-1) Represents the first i-1 subproblems generated:

[0048]

[0049] For the t-th (t>1) round of planning, that is, not the first round of planning, the input to the large language model also includes the sub-problem set SubQ of the previous t-1 rounds 1:(t-1) and the corresponding answer set SubA 1:(t-1) :

[0050]

[0051] S202: traverse the sub-problem set and use the beam search decoding algorithm to generate a sub-answer corresponding to each sub-problem.

[0052] In this step, it is necessary to obtain the answer corresponding to each sub-question, traverse the sub-question set, input each sub-question into the multimodal large model in turn, use the beam search decoding algorithm, generate multiple related words as candidate words according to the sub-questions, calculate the probability distribution of the candidate words, and generate sequences based on multiple candidate words with the highest joint conditional probabilities. When all sequences generate the end word or reach the maximum generation length, the decoding is terminated, the joint conditional probability of each sequence is determined, and the sub-answer is output based on the joint conditional probability.

[0053] Specifically, the target image is input into the multimodal large model to generate a picture description describing the target image; the target question and the picture description are input into the large language model to generate multiple sub-questions.

[0054] For example, the current subproblem sq i The input of the multimodal model is the target image I. The model uses the beam search decoding algorithm to generate the answer in the answer process. The input is represented as follows, where Prompt a This is the question instruction template to be input to the multimodal large model:

[0055] input=Prompt a (I,sq i ).

[0056] In one embodiment, a beam search decoding algorithm may be employed to maintain a fixed number of generated sequences during the decoding process; relevant word-grams are generated based on the sub-problems, and when the next word-gram is generated, the probability distribution of the remaining candidate generated sequences is calculated, and multiple sequences with the highest joint conditional probability are selected as new candidate sequences; word-grams corresponding to each sequence are generated in sequence until all sequences have generated an end word-gram or the maximum generated length is reached, at which point decoding is terminated; the joint conditional probability of each candidate sequence is determined, and multiple candidate sub-answers are output based on the joint conditional probability.

[0057] Beam search decoding is to maintain a fixed number of generated sequences during the decoding process Where k is the beam width, which can be set to any value between 5 and 10 in this step, and t is the time step, that is, the length of the current generated sequence. refers to the word at the tth time step in the i-th sequence, It refers to the entire i-th sequence currently maintained. When generating the next word, it is necessary to calculate the probability distribution of each candidate generated sequence, as shown below:

[0058]

[0059] in, is the vector output by the model, and V is the size of the vocabulary. Then, the first k new sequences are obtained from all probability distributions according to the conditional probability as new candidate sequences. The formula is as follows:

[0060]

[0061] When all sequences generate end tokens (" <eos>”) or when the maximum generated length is reached, the decoding is terminated. Then the sum of the likelihood conditional probabilities of all candidate sequences that have appeared so far is calculated, and the n sequences with the highest probability are selected as the candidate sub-answer set In the present invention, n can be set to any value between 5 and 7. The formula is as follows, where S(y 1:L )→R represents the probability calculation function, L represents the length of the candidate sequence, and α is used as the length penalty term, which is usually set to 0.75:

[0062]

[0063] In another embodiment, to improve the accuracy of sub-answers, a beam search decoding algorithm is used to output multiple candidate sub-answers based on joint conditional probabilities. For example, multiple candidate sequences with the highest joint conditional probabilities can be used as sub-answers, and the sub-answer corresponding to the sub-question is determined by calculating the confidence score of each candidate sub-answer. The confidence of each candidate sub-answer is obtained by performing a weighted sum based on the joint conditional probability, initial conditional probability, generated sequence length, and content consistency score of each candidate sub-answer. The candidate sub-answer with the highest confidence is selected as the sub-answer corresponding to the sub-question. The content consistency score is obtained by calculating the Levenshtein distance between two sequences.

[0064] In order to select the most credible answer from the set of candidate sub-answers and improve the robustness and accuracy of reasoning, this paper designs a confidence calculation formula from the perspectives of joint conditional probability, initial conditional probability, generated sequence length, and content consistency. The confidence is calculated for each candidate answer, and the answer with the highest confidence is selected as the final sub-answer. The formula is as follows, where the joint conditional probability calculation function S1 is detailed in formula (8), S2 represents the initial conditional probability function, S3 represents the length compensation calculation function, S4 represents the consistency score calculation function, and conf represents the confidence score of the sequence:

[0065] S2(y 1:L )=logP MLLM (y1|input)→R;

[0066] S3(y 1:L )=log(1+L)→R;

[0067] S4(y 1:L )=log(1+N i )→R;

[0068]

[0069]

[0070] Among them 1 {…} is an indicator function, which is 1 when the condition is met and 0 otherwise; Levenshtein function refers to the Levenshtein distance calculation formula, which is used to calculate the distance between two sequences. α is 0.15, β is 0.8, γ is 0.02, ε is 0.03, and θ represents the similarity threshold, which is 0.7. Then, the candidate sub-answer with the highest confidence score is selected as sq i The final sub-answer formula is as follows:

[0071]

[0072] Finally, determine whether the sub-problem set has been traversed. If so, perform large language model inference. Otherwise, continue traversing and generate the corresponding answer for the next sub-problem.

[0073] S203: Construct a multimodal reasoning prompt based on the sub-question-sub-answer pair, input it into the large language model for reasoning, and obtain the reasoning answer.

[0074] In this step, the large language model is used for reasoning. The existing sub-question-answer pairs are used to construct multimodal reasoning prompts and input into the large language model for reasoning. The formula is as follows, where SubQ 1:t and SubA 1:t Respectively represent the sub-questions and corresponding answers from the first round of planning to the tth round of planning, Prompt r The instruction template for performing inference, a is the answer generated by the model:

[0075] a~P LLM (a|Prompt r (c,q,A,SubQ 1:t ,SubA 1:t )).

[0076] In addition, in order to improve the accuracy of the generated reasoning answer, the reasoning planning process of the reasoning answer can also be iterated. When the confidence of the reasoning answer does not reach the preset confidence, the target problem is reconstructed into a set of sub-problems, and the multimodal large model and the large language model are used to perform iterative reasoning until the confidence of the obtained reasoning answer reaches the preset confidence or the number of iterations reaches the maximum number, then the iteration is stopped and the reasoning answer is output.

[0077] Specifically, whether to perform iteration is determined based on the confidence of the answer inferred by the large language model and the maximum number of planning times.

[0078] First, the confidence level of the answer a inferred by the large language model is determined. If it is determined that the confidence level meets the preset requirements, the answer is output and the process ends. If it does not meet the preset requirements, the target problem is re-planned and the iteration continues. If the maximum number of planning times is reached, a random answer is output and the process ends. For non-first-round planning, the target problem is deconstructed into a set of sub-problems based on the image description, including: inputting the target problem, the image description, and the set of sub-problems and corresponding answers from the previous round into the large language model to generate multiple sub-problems.

[0079] It should be noted that multimodal reasoning tasks in different fields may require different numbers of iterations for the maximum number of planning iterations and the beamwidth parameter settings during decoding. More complex multimodal reasoning tasks may require a larger number of planning iterations. Those skilled in the art can preset these parameters based on their own experience.

[0080] Using the above method, first, the present invention structures the reasoning problem into a set of more fine-grained sub-problems through inference planning, so that the overall extracted visual text information is richer and more accurate, alleviates the error accumulation phenomenon of the multimodal large model when generating long reasoning content, and enhances the robustness of the large model to erroneous visual information during reasoning. Secondly, the present invention uses a decoding algorithm to enhance the reasoning response of the multimodal large model, and screens the reasoning answers by performing confidence calculations from the perspectives of joint probability, sequence length, and content consistency, further ensuring the accuracy of the visual information and effectively improving the final reasoning effect. Finally, the present invention performs multimodal reasoning based on a large language model and a decoding algorithm, which can achieve zero-sample reasoning in different complex visual scenes. It does not require manpower to construct a training data set, greatly reduces the cost of reasoning learning, and has good generalization.

[0081] Secondly, the present invention also provides a large language model multimodal reasoning device based on decoding guidance, such as Figure 3 Shown, including:

[0082] The generation module 301 generates an image description based on the target question and the corresponding target image, and deconstructs the target question into a set of sub-questions according to the image description.

[0083] The answer module 302 is used to traverse the sub-question set, input each sub-question into the multimodal large model in turn, use the beam search decoding algorithm to generate multiple related word-grams as candidate word-grams based on the sub-questions, calculate the probability distribution of the candidate word-grams, and generate sequences based on multiple candidate word-grams with the highest joint conditional probabilities. When all sequences generate end words or reach the maximum generation length, the decoding is terminated, the joint conditional probability of each sequence is determined, and the sub-answer is output based on the joint conditional probability.

[0084] The reasoning module 303 is used to construct multimodal reasoning prompts based on sub-question-sub-answer pairs, input them into the large language model for reasoning, and obtain reasoning answers.

[0085] Using the above-mentioned device, firstly, the present invention structures the reasoning problem into a set of more fine-grained sub-problems through inference planning, so that the overall extracted visual text information is richer and more accurate, alleviates the error accumulation phenomenon of the multimodal large model when generating long reasoning content, and enhances the robustness of the large model to erroneous visual information during reasoning. Secondly, the present invention uses a decoding algorithm to enhance the reasoning response of the multimodal large model, and screens the reasoning answers by performing confidence calculations from the perspectives of joint probability, sequence length, and content consistency, further ensuring the accuracy of the visual information and effectively improving the final reasoning effect. Finally, the present invention performs multimodal reasoning based on a large language model and a decoding algorithm, which can achieve zero-sample reasoning in different complex visual scenes. It does not require manpower to construct a training data set, greatly reduces the cost of reasoning learning, and has good generalization.

[0086] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 The steps of the decoding-guided large language model multimodal inference method are provided.

[0087] The present invention also provides a computer device. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The steps of the decoding-guided large language model multimodal inference method are provided.

[0088] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0089] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0090] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0092] It should be noted that the above specific embodiments can enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail, those skilled in the art should understand that the present invention can still be modified or replaced with equivalents; and all technical solutions and improvements that do not depart from the spirit and scope of the present invention are included in the scope of protection of the patent for the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.< / eos>

Claims

1. A decoding-guided large language model multimodal reasoning method, characterized in that: include: Generate an image description based on the target question and the corresponding target image, and deconstruct the target question into a set of sub-questions based on the image description; Traverse the set of sub-questions, input each sub-question into the multimodal large model in turn, use the beam search decoding algorithm to generate multiple related word-grams as candidate word-grams based on the sub-questions, calculate the probability distribution of the candidate word-grams, and generate sequences based on the candidate word-grams with the highest joint conditional probabilities. Decoding terminates when all sequences have generated end words or reached the maximum generation length, determine the joint conditional probability of each sequence, and output sub-answers based on the joint conditional probabilities. Based on the sub-question-sub-answer pairs, multimodal reasoning prompts are constructed and input into the large language model for reasoning to obtain the reasoning answer.

2. The decoding-guided large language model multimodal reasoning method according to claim 1, characterized in that: Outputting the sub-answer according to the joint conditional probability includes: Output multiple candidate sub-answers according to the joint conditional probability; The confidence of each candidate sub-answer is obtained by performing a weighted summation based on the joint conditional probability, initial conditional probability, generated sequence length, and content consistency score of each candidate sub-answer; the content consistency score is obtained by calculating the Levenshtein distance between the two sequences; The candidate sub-answer with the highest confidence is used as the sub-answer corresponding to the sub-question.

3. The decoding-guided large language model multimodal reasoning method according to claim 1, characterized in that: The process of generating an image description based on the target question and the corresponding target image, and deconstructing the target question into a set of sub-questions based on the image description, includes: Inputting the target image into the multimodal large model to generate a picture description describing the target image; The target question and image description are input into the large language model to generate multiple sub-questions.

4. The decoding-guided large language model multimodal reasoning method according to claim 1, characterized in that: The method further comprises: When the confidence of the inference answer does not reach the preset confidence, the target problem is reconstructed into a set of sub-problems, and iterative reasoning is performed using the multimodal large model and the large language model until the confidence of the obtained inference answer reaches the preset confidence or the number of iterations reaches the maximum number.

5. The decoding-guided large language model multimodal reasoning method according to claim 4, characterized in that: For non-first-round planning, the target problem is decomposed into a set of sub-problems according to the image description, including: The target question, image description, and a set of sub-questions and corresponding answers from the previous round are input into the large language model to generate multiple sub-questions.

6. The decoding-guided large language model multimodal reasoning method according to claim 5, characterized in that: In the first round of planning, the formula for reasoning planning of the large language model is: ; For non-first-round planning, the formula for reasoning planning using the large language model is: ; in, express The generated probability distribution, Represents a hint template for reasoning planning, represents the i-th subproblem, Indicates the generated sub-questions, c represents the picture description, q represents the reasoning question, and A represents the set of answer options. Before The set of subproblems of the wheel, Indicates that it corresponds to A collection of answers.

7. A decoding-guided large language model multimodal reasoning device, characterized in that: The device comprises: The generation module generates an image description based on the target question and the corresponding target image, and deconstructs the target question into a set of sub-questions based on the image description; The answer module is used to traverse the sub-question set, input each sub-question into the multimodal large model in sequence, use the beam search decoding algorithm to generate multiple related word-grams as candidate word-grams based on the sub-questions, calculate the probability distribution of the candidate word-grams, and generate sequences based on the candidate word-grams with the highest joint conditional probabilities. Decoding terminates when all sequences have generated end words or reached the maximum generation length, determines the joint conditional probability of each sequence, and outputs the sub-answer based on the joint conditional probabilities; The reasoning module is used to construct multimodal reasoning prompts based on sub-question-sub-answer pairs, input them into the large language model for reasoning, and obtain the reasoning answer.

8. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 6 when executing the program.