Multi-modal large model optimization method and device and electronic equipment

By introducing process reward model and process supervision data into the multimodal large model, combined with SFT and DPO algorithm optimization models, the problem of hallucinatory content in image description is solved, and higher accuracy and stability are achieved.

CN119940557AActive Publication Date: 2025-05-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510436082.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Existing multimodal large models have hallucinatory content in image descriptions, resulting in instability in output and low accuracy.

Method used

Through process reward model training based on basic multimodal large models, corrected label data is generated sentence by sentence, and is used to optimize the model to suppress illusions. The specific methods include: using the process reward model to determine whether the candidate description of the current sentence is correct, generating paired correct and incorrect image descriptions as process supervision data, and optimizing the model with SFT and DPO algorithms.

Benefits of technology

It effectively improves the accuracy and stability of image description, reduces hallucination content, and improves the reliability of model output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940557A_ABST
    Figure CN119940557A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large model optimization method and device and electronic equipment, and the method comprises the steps: training a basic multi-modal large model, and obtaining a process reward model; performing sentence-by-sentence reasoning of image description on the training image by using the basic multi-modal large model; for each current sentence obtained by reasoning, determining whether each candidate description of the current sentence is correct or not by using the process reward model, and applying the correct candidate description of the current sentence to next sentence reasoning of image description; based on paired correct candidate descriptions and paired error candidate descriptions in the sentence descriptions obtained through sentence-by-sentence reasoning, paired correct image descriptions and paired error image descriptions are determined to serve as process supervision data; and optimizing the basic multi-modal large model based on a training image and the process supervision data to obtain a hallucination-suppressed multi-modal large model. According to the invention, the illusion suppression performance can be effectively improved during image description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to neural network technology, and in particular to an optimization method, device and electronic device for a multimodal large model. Background Art

[0002] Image description has always been an important research field in the field of vision, which can give a text description of an image. After the rise of multimodal large models, this field has made great progress: the level of detail and richness of image description has been greatly improved, and image description is no longer just a simple description of the main target of the image. But on the other hand, the application of multimodal large models to image description also introduces more hallucination content to the description, and there is a phenomenon that the more you "say", the more you get it right, and the more you get it wrong. In order to suppress the problem of hallucination in image description, the current mainstream technologies include contrastive decoding and outcome-supervised reward models, which optimize multimodal large models to achieve hallucination suppression of image description. At present, contrastive decoding has low inference efficiency and unstable output; the outcome-supervised reward supervision signal is unclear, and the hallucination suppression effect is not good. Summary of the invention

[0003] The present application provides a method, device and electronic device for optimizing a multimodal large model, which can effectively improve the performance of hallucination suppression when describing an image.

[0004] To achieve the above purpose, this application adopts the following technical solutions: A multi-modal large model optimization method, comprising: Based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, the basic multimodal large model is trained to obtain a process reward model; wherein the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; The basic multimodal large model is used to perform sentence-by-sentence reasoning of image description on the training image; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description; Based on the pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, pairs of correct image descriptions and incorrect image descriptions are determined as process supervision data; The basic multimodal large model is optimized based on the training images and the process supervision data to obtain a multimodal large model that suppresses hallucinations.

[0005] Preferably, the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and a correct current sentence corresponding to the correction of the wrong current sentence; When training the basic multimodal large model to obtain the process reward model, the input data includes: a positive sample consisting of the input image, the preceding correct description in the image description, and a correct current sentence, and a negative sample consisting of the input image, the preceding correct description in the image description, and an incorrect current sentence; the output data includes: information on whether the current sentence is correct.

[0006] Preferably, for any current sentence, the pair of correct candidate description and incorrect candidate description is a pair, including: a correct candidate description and an incorrect candidate description selected from all candidate descriptions of the current sentence; or, For any current sentence, the pairs of correct candidate descriptions and incorrect candidate descriptions are multiple pairs, including: all paired combinations of any correct candidate descriptions and any incorrect candidate descriptions in all candidate descriptions of the current sentence.

[0007] Preferably, the determining of a pair of correct image descriptions and incorrect image descriptions comprises: The pairs of correct candidate descriptions and incorrect candidate descriptions corresponding to the same sentence are used as the pairs of correct image descriptions and incorrect image descriptions, and organized into a pair of positive and negative samples in the process supervision data.

[0008] Preferably, the determining of a pair of correct image descriptions and incorrect image descriptions comprises: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the first wrong candidate description and other descriptions following it as the first wrong subsequent description corresponding to the current sentence; Using the first correct candidate description paired with the first incorrect candidate description and other descriptions following it in the complete image description of the training image obtained by the sentence-by-sentence reasoning as the first correct subsequent description corresponding to the current sentence; The first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0009] Preferably, generating pairs of correct image descriptions and incorrect image descriptions as process supervision data comprises: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the complete image description obtained by reasoning as the first wrong full image description corresponding to the current sentence; Using the complete image description of the training image obtained by the sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct full image description corresponding to the current sentence; The first correct full-image description and the first incorrect full-image description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0010] Preferably, the optimizing the basic multimodal large model based on the training image and the process supervision data comprises: Based on the training images and the process supervision data, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm.

[0011] Preferably, the combined SFT algorithm and DPO algorithm are used to optimize the basic multimodal large model, including: Based on the correct image description in the process supervision data, using the SFT algorithm to determine the value of the first loss function; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, the DPO algorithm is used to determine the value of the second loss function, the weighted sum of the value of the first loss function and the value of the second loss function is calculated, and the parameters of the basic multimodal large model are updated based on the result of the weighted sum.

[0012] An optimization device for a multimodal large model for image description, comprising: a process reward model generation unit, a process supervision data generation unit and a model optimization unit; The process reward model generation unit is used to train the basic multimodal large model based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, so as to obtain the process reward model; wherein the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; The process supervision data generating unit is used to perform sentence-by-sentence reasoning of image descriptions on training images using the basic multimodal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description; and is also used to determine pairs of correct image descriptions and incorrect image descriptions based on pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning as process supervision data; The model optimization unit is used to optimize the basic multimodal large model based on the training image and the process supervision data to obtain a multimodal large model for image description that suppresses hallucinations.

[0013] Preferably, the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and a correct current sentence obtained after correcting an erroneous current sentence.

[0014] In the process reward model generation unit, when the basic multimodal large model is trained to obtain the process reward model, the input data includes: a positive sample consisting of the input image, the preceding correct description in the image description, and a correct current sentence, and a negative sample consisting of the input image, the preceding correct description in the image description, and an incorrect current sentence; the output data includes: information on whether the current sentence is correct.

[0015] Preferably, the pair of correct candidate description and incorrect candidate description of the current sentence is a pair, including: a correct candidate description and an incorrect candidate description selected from all candidate descriptions of the current sentence; or, The correct candidate descriptions and incorrect candidate descriptions of the current sentence are multiple pairs, including: all paired combinations of any correct candidate description and any incorrect candidate description among all candidate descriptions of the current sentence.

[0016] Preferably, in the process supervision data generating unit, the determining of the paired correct image description and the incorrect image description comprises: The pairs of correct candidate descriptions and incorrect candidate descriptions corresponding to the same sentence are used as the pairs of correct image descriptions and incorrect image descriptions, and organized into a pair of positive and negative samples in the process supervision data.

[0017] Preferably, in the process supervision data generating unit, the determining of the paired correct image description and the incorrect image description comprises: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the first wrong candidate description and other descriptions following it as the first wrong subsequent description corresponding to the current sentence; Using the first correct candidate description paired with the first incorrect candidate description and other descriptions following it in the complete image description of the training image obtained by the sentence-by-sentence reasoning as the first correct subsequent description corresponding to the current sentence; The first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0018] Preferably, in the process supervision data generating unit, generating pairs of correct image descriptions and incorrect image descriptions as process supervision data comprises: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the complete image description obtained by reasoning as the first wrong full image description corresponding to the current sentence; Using the complete image description of the training image obtained by the sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct full image description corresponding to the current sentence; The first correct full-image description and the first incorrect full-image description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0019] Preferably, in the model optimization unit, optimizing the basic multimodal large model based on the training image and the process supervision data includes: Based on the training images and the process supervision data, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm.

[0020] Preferably, in the model optimization unit, the combined SFT algorithm and DPO algorithm are used to optimize the basic multimodal large model, including: Based on the correct image description in the process supervision data, using the SFT algorithm to determine the value of the first loss function; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, the DPO algorithm is used to determine the value of the second loss function, the weighted sum of the value of the first loss function and the value of the second loss function is calculated, and the parameters of the basic multimodal large model are updated based on the result of the weighted sum.

[0021] A computer-readable storage medium having computer instructions stored thereon, wherein the instructions, when executed by a processor, can implement any of the above-mentioned multimodal large model optimization methods.

[0022] An electronic device comprising at least a computer-readable storage medium and a processor; The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement any of the above-mentioned multimodal large model optimization methods.

[0023] As can be seen from the above technical solution, in this application, based on the image description output after the basic multimodal large model is inferred from the input image and the label data obtained by correcting the image description sentence by sentence, the basic multimodal large model is trained to obtain the process reward model, which is used to output information on whether each sentence in the image description is correct for the input image and its image description. Then, the basic multimodal large model is used to perform sentence-by-sentence reasoning on the training image for the image description, and for each current sentence obtained by reasoning, the process reward model is used to mark whether each candidate description of the current sentence is correct, so that the correct description and the wrong description at the same sentence description position can be obtained; and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description. Next, based on the paired correct candidate descriptions and wrong candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, paired correct image descriptions and wrong image descriptions are generated as process supervision data, thereby constructing a positive and negative sample pair of the image description, which, on the one hand, reflects the correctness of the image description, and on the other hand, ensures that the corresponding description position is the same. Finally, the basic multimodal large model is optimized based on the training image and the process supervision data. Since the process reward model and process supervision are used to optimize the multimodal large model, the reasoning efficiency can be effectively improved. At the same time, the process supervision data includes positive and negative sample pairs corresponding to the same description position, which further strengthens the relationship between the position and the positive and negative samples. This enables the multimodal large model to more effectively achieve hallucination suppression across the entire image and improve the accuracy of the model output. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of the basic process of the multimodal large model optimization method in this application; Figure 2 A schematic diagram of a specific process of optimizing a multi-modal large model in a specific embodiment of the present application; Figure 3 This is a block diagram of the implementation of the process reward model generation stage in the specific embodiment of this application; Figure 4 This is a block diagram of the implementation of the process supervision data generation stage in the specific embodiment of this application; Figure 5 This is a block diagram of the implementation of the model optimization stage in the specific embodiment of this application; Figure 6 It is a basic structural diagram of the multi-modal large model optimization device in this application; Figure 7 It is a schematic diagram of the basic structure of the electronic device in this application. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical means and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings.

[0026] The basic idea of ​​this application is to apply the process reward model and process supervision-based model optimization to the optimization of multimodal large models for image description to improve reasoning efficiency and enhance the hallucination suppression effect.

[0027] Figure 1 Schematic diagram of the basic process of the multi-modal large model optimization method in this application. Figure 1 As shown, the method includes: Step 101, based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, the basic multimodal large model is trained to obtain a process reward model.

[0028] This step is used to generate a process reward model. Among them, the basic multimodal large model used for image description can analyze the input image and output a text description of the image. The image description obtained specifically for the input image may contain multiple sentences. For the image description output after reasoning with the basic multimodal large model, the label data after sentence-by-sentence correction can be obtained, and the basic multimodal large model is trained based on the image description and the corresponding label data to generate a process reward model. The process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description. In other words, the process reward model is used to determine whether each sentence in the image description is correct.

[0029] Step 102, use the basic multimodal large model to perform sentence-by-sentence reasoning on the image description of the training image, and for each current sentence obtained by reasoning, use the process reward model to determine whether each candidate description of the current sentence is correct, and use the correct candidate description of the current sentence for the next sentence reasoning of the image description.

[0030] Steps 102 to 103 are used to generate process supervision data. Specifically, step 102 uses the basic multimodal large model to perform sentence-by-sentence reasoning on the image description of the training image, and step 103 organizes and generates process supervision data based on the reasoning results.

[0031] Among them, the image description usually includes multiple sentences. In the process of using the basic multimodal large model to process and obtain the image description, reasoning is usually performed sentence by sentence. Each time the current sentence is inferred, other descriptions before the current sentence can be used to infer and output one or more candidate descriptions of the current sentence, and then a candidate description is selected as the current sentence description. Based on the current sentence description and other previous descriptions, the next sentence is inferred, and so on, to obtain a complete image description of the training image.

[0032] The process reward model is introduced in the present application to participate in the sentence-by-sentence reasoning of the image description. Specifically, after each current sentence is inferred, one or more candidate descriptions of the current sentence are obtained, and the process reward model is further used to determine whether each candidate description of the current sentence is correct. The candidate description determined to be correct by the process reward model is called the correct candidate description; the candidate description determined to be correct by the process reward model is called the incorrect candidate description. The correct candidate description of the current sentence is used together with the preceding correct description of the current sentence for the next sentence reasoning of the image description. That is to say, in the sentence-by-sentence reasoning of the image description of the present application, the correct candidate description determined by the process reward model is selected each time for the next sentence reasoning, thereby obtaining a complete image description consisting entirely of correct descriptions.

[0033] Let me explain here that "previous correct description" refers to a set of correct descriptions obtained before a certain description is inferred. For example, if an image description includes 10 sentences in total and when inferring the 5th sentence, the previous correct description refers to the correct descriptions of the first 4 sentences.

[0034] Step 103, based on the pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, generate pairs of correct image descriptions and incorrect image descriptions as process supervision data.

[0035] For any current sentence of the image description, if the current sentence corresponds to an incorrect candidate description, then the position of the current sentence is also the position where the hallucination description may appear. If a current sentence corresponds to both a correct candidate description and an incorrect candidate description, then the current sentence has a pair of correct candidate descriptions and incorrect candidate descriptions, and the pair of descriptions corresponds to the same description position.

[0036] Based on the paired correct candidate descriptions and incorrect candidate descriptions of the current sentence, a paired correct image description and an incorrect image description are generated as a pair of positive and negative samples of process supervision data, and the positive and negative samples correspond to the same description position.

[0037] Step 104, optimizing the basic multimodal large model based on the training images and the process supervision data to obtain a multimodal large model that suppresses hallucinations.

[0038] This step optimizes the model based on process supervision data and training images. Specifically, various existing reinforcement learning methods can be used to optimize the model using supervision data. Since process supervision data includes positive samples and negative samples, the positive samples represent the correct description information that the model needs to strengthen, and the negative samples represent the hallucination description information that the model needs to suppress. At the same time, the positive and negative sample pairs also correspond to the same description position, thereby further strengthening the position information contained in the sample. Based on this model optimization, hallucination suppression can be effectively achieved in the entire image range.

[0039] So far, Figure 1 The basic process of the model optimization method in the present application is shown to be finished. Next, the specific implementation of the model optimization method in the present application is described through specific embodiments.

[0040] Figure 2 The schematic diagram of the specific process of the optimization method of the multimodal large model in the specific embodiment of this application is as follows. The whole process includes three stages: the process reward model generation stage (corresponding to the aforementioned Figure 1 Step 101 in the process), process supervision data generation phase (corresponding to the aforementioned Figure 1 Steps 102-103 in the process) and the model optimization phase (corresponding to the aforementioned Figure 1 Step 104 in the process). The following is a detailed description in stages.

[0041] Phase 1: Process reward model generation phase The first stage of treatment includes Figure 2 Steps 201 to 203 in the embodiment are as follows: Figure 3 shown.

[0042] Step 201, input the image into the basic multimodal large model, and generate a complete image description through reasoning of the basic multimodal large model; The processing of the basic multimodal large model can be performed in an existing manner, where an image is input and an output image description is indicated, and a complete image description for the entire image can be obtained after the large model is processed. The complete image description includes multiple description sentences.

[0043] Step 202: Obtain label data obtained by correcting each sentence in the image description.

[0044] The label data obtained by correcting each sentence in the image description includes: information on whether the description of the sentence is correct, and a correct description sentence obtained after correcting the wrong description sentence.

[0045] The specific correction process can be performed manually. This step is used to obtain the corresponding label data.

[0046] Step 203, based on the image description obtained in step 201 and the label data obtained in step 202, the basic multimodal large model is optimized and trained to generate a process reward model for determining whether the description sentence is correct.

[0047] For all the image descriptions obtained in step 201, they are marked as correct or not in step 202, and the wrong descriptions are modified, and the training data samples are organized based on these image descriptions. The correct descriptions are used as positive samples, and the wrong descriptions are used as negative samples. Usually, in addition to the correct description or wrong description of the current sentence, the sample also needs to include the input image of step 201 and the correct description of the preceding sequence of the current sentence in the image description.

[0048] The training samples are input into the multimodal large model for processing, and the binary classification information of whether the current sentence is correct or incorrect is output. Then, it is compared with the standard data, the value of the loss function is calculated, and the model parameters are adjusted according to the value of the loss function to generate a process reward model. The process reward model is used to determine whether each current description sentence is correct.

[0049] Phase 2: Process Supervision Data Generation Phase The second stage of processing is divided into two parts: multimodal large model reasoning and organizing process supervision data based on reasoning results, including Figure 2 Steps 204 to 206 in the embodiment are as follows: Figure 4 shown.

[0050] Step 204: Use the basic multimodal large model to perform sentence-by-sentence reasoning on the image description of the training image.

[0051] To distinguish the input images of the basic multimodal large model in the first stage from the input images of the basic multimodal large model in the second stage, the input images of the basic multimodal large model in the second stage are called training images.

[0052] In the process of step 204, the image description is inferred sentence by sentence for the training image. The specific process is described by taking the inference of a current description sentence as an example, and the current description sentence is referred to as the current sentence hereinafter.

[0053] First, the training image and the correct preceding description of the current description sentence are input into the basic multimodal large model, and the basic multimodal large model is instructed to output the current sentence of the image description through a text prompt. Usually, the current sentence output by the model includes one or more candidate descriptions.

[0054] Then, the training image, the previous correct description and all the candidate descriptions of the current sentence are input into the process reward model. The process reward model determines whether each candidate description of the current sentence is correct and marks it.

[0055] Finally, the candidate descriptions marked as correct are input into the basic multimodal large model for reasoning about the next description; the candidate descriptions marked as incorrect are saved.

[0056] Through the above sentence-by-sentence reasoning, the correct candidate description is used each time for the reasoning of the next description. After completing the complete reasoning of the training image, the obtained complete image description is composed entirely of correct descriptions.

[0057] Step 205, input the incorrect candidate descriptions obtained in the sentence-by-sentence reasoning into the basic multimodal large model, and obtain the subsequent description of the training image through reasoning.

[0058] Step 205 is used to complete another image description reasoning process. This process is performed on the incorrect candidate descriptions generated in the reasoning process of step 204. The incorrect candidate descriptions, the preceding correct descriptions and the training images are input into the basic multimodal large model, and the large model is instructed to output the image descriptions through text prompts. The basic multimodal large model uses the existing method to perform reasoning to obtain all subsequent descriptions of the training images. It should be noted here that the incorrect candidate descriptions based on which the image description reasoning in step 205 is performed can be selective, and specifically can be, for each current sentence, selecting the incorrect candidate descriptions from the paired correct candidate descriptions and incorrect candidate descriptions of the current sentence. The meaning of the paired correct candidate descriptions and incorrect candidate descriptions is introduced below.

[0059] In this application, in order to organize process supervision data, the correct candidate description and the wrong candidate description of the current sentence will be paired in the sentence-by-sentence reasoning of step 204. In the simplest way, when pairing, a correct candidate description and an wrong candidate description can be selected as a pair of correct candidate descriptions and wrong candidate descriptions, that is, only a pair of correct candidate descriptions and wrong candidate descriptions are formed, for example, the correct candidate description with the highest confidence and the wrong candidate description with the lowest confidence are selected as a pair of correct candidate descriptions and wrong candidate descriptions, and other candidate descriptions do not participate in subsequent processing. Alternatively, in order to enrich the process supervision data, any correct candidate description and any wrong candidate description can be paired, and all possible paired combinations are used as pairs of correct candidate descriptions and wrong candidate descriptions. For example, the current sentence has M correct candidate descriptions and N wrong candidate descriptions. For each correct candidate description, it can be paired with any wrong candidate description, and there can be M×N pairs of correct candidate descriptions and wrong candidate descriptions. In the above pairing process, the paired correct candidate descriptions and wrong candidate descriptions correspond to the same current description sentence, that is, to the same description position.

[0060] Based on the above pairing process, for any current sentence, there may or may not be a pair of correct candidate descriptions and incorrect candidate descriptions. For a current sentence that does not exist, there is no need to perform the image reasoning process of step 205. For a current sentence that exists, if its pair of correct candidate descriptions and incorrect candidate descriptions is only one pair, then only the incorrect candidate description in the pair can be selected to perform the image reasoning process of step 205. If the current sentence has multiple pairs of correct candidate descriptions and incorrect candidate descriptions, the image reasoning process of step 205 can be performed on the incorrect candidate descriptions in each pair in turn.

[0061] For the image reasoning process of step 205, all subsequent descriptions of the training image starting from the current sentence can be obtained. Since the reasoning process here is based on an incorrect candidate description, the subsequent description obtained is also inaccurate. This incorrect candidate description and the subsequent description obtained based on the incorrect candidate description are combined together to form an incorrect subsequent description; the previous correct description and the incorrect subsequent description can form a complete full-image image description. Of course, this full-image image description is also inaccurate. The full-image image description obtained based on the incorrect candidate description is referred to as the incorrect full-image description below. The incorrect subsequent description and the incorrect full-image description here are both based on the incorrect candidate description of the current sentence, so the incorrect subsequent description and the incorrect full-image description here are called the incorrect subsequent description and the incorrect full-image description of the current sentence. Since each current sentence may infer multiple candidate descriptions, for one incorrect candidate description, multiple incorrect subsequent descriptions and incorrect full-image descriptions may be obtained through the image reasoning process.

[0062] Step 206 , based on the pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, determine the pairs of correct image descriptions and incorrect image descriptions as process supervision data.

[0063] The above has introduced the paired correct candidate descriptions and incorrect candidate descriptions, and introduced the incorrect subsequent descriptions and incorrect full-image descriptions obtained by image reasoning based on the incorrect candidate descriptions through step 205. Step 206 introduces the organization of process supervision data.

[0064] In this embodiment, in order to make full use of the description position information, the process supervision data is organized according to the description position. Specifically, when organizing the process supervision data, a pair of correct image description and incorrect image description is determined based on the pair of correct candidate description and incorrect candidate description of the current sentence as a pair of positive and negative samples in the process supervision data.

[0065] The organization of specific process supervision data can be divided into three different granularities: 1. Description sentence granularity The process supervision data is organized at the sentence granularity, that is, the correct image description and the incorrect image description are at the sentence granularity level. Specifically, for each current sentence, the correct candidate description and the incorrect candidate description of the current sentence can be organized into a pair of positive and negative samples in the process supervision data; of course, the sample usually also includes a text prompt for indicating the generated image description and the preceding correct description. The pair of positive and negative samples obtained in this way corresponds to the same current sentence, and therefore corresponds to the same description position. At this granularity, the paired correct image description and the incorrect image description are also paired correct candidate descriptions and incorrect candidate descriptions.

[0066] 2. Granularity of Partial Image Description The process supervision data is organized at the granularity of the second half description after the current sentence, that is, the correct image description and the wrong image description are at the granularity level of the partial image description. Before introducing the organization of process supervision data at this granularity level, we first introduce the meaning of the correct subsequent description of the current sentence and the correct full image description.

[0067] For any current sentence, the correct candidate description in the paired correct candidate description and incorrect candidate description of the current sentence is used for the next sentence reasoning of the image description in step 204, and the subsequent reasoning of the training image is completed according to the sentence-by-sentence reasoning of step 204 to obtain a complete image description consisting entirely of correct descriptions; the aforementioned correct candidate description and other descriptions after it in the complete image description are combined together as the correct subsequent description corresponding to the current sentence, and the complete image description is used as the correct full image description corresponding to the current sentence. It can be seen from the above process that for the current sentence, the correct subsequent description and the correct full image description corresponding to the current sentence can be obtained through step 204, and the incorrect subsequent description and the incorrect full image description corresponding to the current sentence can be obtained through step 205. Among them, since there can be multiple candidate descriptions, there may be multiple correct full image descriptions, correct subsequent descriptions, incorrect full image descriptions, and incorrect subsequent descriptions corresponding to the current sentence.

[0068] When organizing process supervision data, for each current sentence, the correct subsequent description and incorrect subsequent description inferred from the paired correct candidate description and incorrect candidate description are organized into a pair of positive and negative samples of process supervision data. Of course, the sample usually also includes a text prompt for indicating the generated image description and the preceding correct description. The pair of positive and negative samples obtained in this way corresponds to the same current sentence, and therefore to the same description position. At this granularity, the paired correct image description and incorrect image description are also paired correct subsequent description and incorrect subsequent description.

[0069] 3. Granularity of Complete Image Description The process supervision data is organized at the granularity of the complete image description, that is, the correct image description and the incorrect image description are at the granularity level of the complete image description. When organizing the process supervision data, for each current sentence, the correct full image description and the incorrect full image description inferred by the paired correct candidate description and the incorrect candidate description are organized into a pair of positive and negative samples of the process supervision data. Of course, the sample usually also includes a text prompt for indicating the generated image description. The pair of positive and negative samples obtained in this way corresponds to the same current sentence, and therefore to the same description position. At this granularity, the paired correct image description and the incorrect image description are also paired correct full image description and incorrect full image description.

[0070] As mentioned above, the process supervision data can be organized from three different granularities, which are summarized in Table 1. In the specific implementation process, the above three different granularities can be used to organize the process supervision data at the same time, or one or two granularities can be selected to organize the process supervision data according to needs.

[0071]

[0072] The third stage: model optimization stage The third stage of treatment includes Figure 2 Step 207 in the embodiment is as follows: Figure 5 shown.

[0073] Step 207, based on the training images and process supervision data, the basic multimodal large model is optimized to obtain a multimodal large model that suppresses hallucinations.

[0074] In this embodiment, based on the training images and process supervision data, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm.

[0075] Specifically, based on the correct image description in the process supervision data, the SFT algorithm is used to determine the value SFT (Chosen) of the first loss function; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, the DPO algorithm is used to determine the value of the second loss function DPO (Chosen, Rejected); Calculate the weighted sum of the first loss function value SFT (Chosen) and the second loss function value DPO (Chosen, Rejected), and use it as the final loss function value Loss to update the parameters of the basic multimodal large model.

[0076] In the above manner, the DPO algorithm is based on the correct image description and the incorrect image description corresponding to the same description position, and contrastive learning is used to optimize the model to a model that meets the correct expectations. The SFT algorithm is used to optimize the model performance based on the correct image description, thereby effectively achieving hallucination suppression at all positions in the entire image.

[0077] So far, Figure 2 The method flow in the specific embodiment of the present application shown ends.

[0078] Through the above-mentioned optimization method of the multimodal large model in this application, it is possible to train the generation process reward model according to the needs of image description, and construct sentence-by-sentence process supervision data, partial process supervision data and full process supervision data. The process supervision data is used to effectively suppress image description hallucinations and improve the accuracy of model output.

[0079] The above is a specific implementation of the multi-modal large model optimization method in this application. This application also provides an optimization device for a multi-modal large model, which can be used to implement the above optimization method of this application. Figure 6 The basic structural diagram of the optimization device for the multi-modal large model provided in this application is as follows: Figure 6 As shown, the device includes: a process reward model generating unit, a process supervision data generating unit and a model optimization unit.

[0080] The process reward model generation unit is used to train the basic multimodal large model based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, so as to obtain the process reward model; wherein the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; A process supervision data generating unit is used to perform sentence-by-sentence reasoning of image descriptions for training images using a basic multimodal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, a process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description; and is also used to determine pairs of correct image descriptions and incorrect image descriptions based on pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning as process supervision data; The model optimization unit is used to optimize the basic multimodal large model based on the training images and process supervision data to obtain a multimodal large model for image description that suppresses hallucinations.

[0081] Optionally, the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and a correct current sentence obtained after correcting the wrong current sentence.

[0082] In the process reward model generation unit, when training the basic multimodal large model to obtain the process reward model, the input data includes: positive samples consisting of the input image, the previous correct description in the image description, and the correct current sentence, and negative samples consisting of the input image, the previous correct description in the image description, and the incorrect current sentence; the output data includes: information on whether the current sentence is correct.

[0083] Optionally, the correct candidate description and the wrong candidate description of the current sentence are a pair, including: a correct candidate description and an wrong candidate description selected from all candidate descriptions of the current sentence; or, The correct candidate descriptions and incorrect candidate descriptions of the current sentence are multiple pairs, including: all paired combinations of any correct candidate description and any incorrect candidate description among all candidate descriptions of the current sentence.

[0084] Optionally, in the process supervision data generation unit, determining a pair of correct image descriptions and incorrect image descriptions comprises: The pairs of correct candidate descriptions and incorrect candidate descriptions corresponding to the same sentence are taken as pairs of correct image descriptions and incorrect image descriptions, and organized into a pair of positive and negative samples in the process supervision data.

[0085] Optionally, in the process supervision data generation unit, determining a pair of correct image descriptions and incorrect image descriptions comprises: For each current sentence, determine the first wrong candidate description in the pair of correct candidate descriptions and wrong candidate descriptions of the current sentence, complete the subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the first wrong candidate description and other descriptions after it as the first wrong subsequent description corresponding to the current sentence; The first correct candidate description paired with the first wrong candidate description and other descriptions following it in the complete image description of the training image obtained by sentence-by-sentence reasoning are used as the first correct subsequent description corresponding to the current sentence; The first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0086] Optionally, in the process supervision data generating unit, generating pairs of correct image descriptions and incorrect image descriptions as process supervision data includes: For each current sentence, determine the first wrong candidate description in the pair of correct candidate descriptions and wrong candidate descriptions of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the complete image description obtained by reasoning as the first wrong full image description corresponding to the current sentence; The complete image description of the training image obtained by sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description is used as the first correct full image description corresponding to the current sentence; The first correct full-image description and the first incorrect full-image description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

[0087] Optionally, in the model optimization unit, the basic multimodal large model is optimized based on the training image and the process supervision data, including: Based on the training images and process supervision data, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm.

[0088] Optionally, in the model optimization unit, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm, including: Based on the correct image description in the process supervision data, the value of the first loss function is determined using the SFT algorithm; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, the DPO algorithm is used to determine the value of the second loss function, and the weighted sum of the values ​​of the first loss function and the second loss function is calculated. The parameters of the basic multimodal large model are updated based on the result of the weighted sum.

[0089] The present application also provides a computer-readable storage medium that stores instructions, which, when executed by a processor, can execute the steps in the optimization method for implementing a multimodal large model as described above. In practical applications, the computer-readable medium can be included in each device / apparatus / system in the above-mentioned embodiments, or it can exist independently without being assembled into the device / apparatus / system. Instructions are stored in a computer-readable storage medium, and the instructions stored therein can execute the steps in the optimization method for a multimodal large model as described above when executed by a processor.

[0090] According to the embodiments disclosed in the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof, but it is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.

[0091] Figure 7This application also provides an electronic device. Figure 7 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically: The electronic device may include a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When executing the program in the memory 702, a multimodal large model optimization method may be implemented.

[0092] Specifically, in practical applications, the electronic device may further include components such as a power supply 703, an input / output unit 704, etc. Those skilled in the art may understand that Figure 7 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Processor 701 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the server and processes data by running or executing software programs and / or modules stored in memory 702 and calling data stored in memory 702, thereby monitoring the electronic device as a whole.

[0093] The memory 702 can be used to store software programs and modules, that is, the above-mentioned computer-readable storage medium. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0094] The electronic device also includes a power supply 703 for supplying power to various components, which can be logically connected to the processor 701 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 703 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.

[0095] The electronic device may also include an input and output unit 704, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, and optical signal inputs related to user settings and function control. The input unit output 704 may also be used to display information input by the user or information provided to the user and various graphical user interfaces, which may be composed of graphics, text, icons, videos, and any combination thereof.

[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for optimizing a multimodal large model, characterized in that: include: Based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, the basic multimodal large model is trained to obtain a process reward model; wherein the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; The basic multimodal large model is used to perform sentence-by-sentence reasoning of image description on the training image; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description; Based on the pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, pairs of correct image descriptions and incorrect image descriptions are determined as process supervision data; The basic multimodal large model is optimized based on the training images and the process supervision data to obtain a multimodal large model that suppresses hallucinations.

2. The method according to claim 1, characterized in that The label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and a correct current sentence obtained after correcting an incorrect current sentence; When training the basic multimodal large model to obtain the process reward model, the input data includes: a positive sample consisting of the input image, the preceding correct description in the image description, and a correct current sentence, and a negative sample consisting of the input image, the preceding correct description in the image description, and an incorrect current sentence; the output data includes: information on whether the current sentence is correct.

3. The method according to claim 1, characterized in that For any current sentence, the pair of correct candidate description and incorrect candidate description is a pair, including: a correct candidate description and an incorrect candidate description selected from all candidate descriptions of the current sentence; or, For any current sentence, the pairs of correct candidate descriptions and incorrect candidate descriptions are multiple pairs, including: all paired combinations of any correct candidate descriptions and any incorrect candidate descriptions in all candidate descriptions of the current sentence.

4. The method according to claim 1, characterized in that: The determining of paired correct image descriptions and incorrect image descriptions comprises: The pairs of correct candidate descriptions and incorrect candidate descriptions corresponding to the same sentence are used as the pairs of correct image descriptions and incorrect image descriptions, and organized into a pair of positive and negative samples in the process supervision data.

5. The method according to claim 1, characterized in that The determining of paired correct image descriptions and incorrect image descriptions comprises: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the first wrong candidate description and other descriptions following it as the first wrong subsequent description corresponding to the current sentence; Using the first correct candidate description paired with the first incorrect candidate description and other descriptions following it in the complete image description of the training image obtained by the sentence-by-sentence reasoning as the first correct subsequent description corresponding to the current sentence; The first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

6. The method according to claim 1, 4 or 5, characterized in that: The generating of paired correct image descriptions and incorrect image descriptions as process supervision data includes: For each current sentence, determine a first wrong candidate description in a pair of a correct candidate description and an wrong candidate description of the current sentence, complete subsequent reasoning of the training image based on the preceding correct description of the current sentence and the first wrong candidate description, and use the complete image description obtained by reasoning as the first wrong full image description corresponding to the current sentence; Using the complete image description of the training image obtained by the sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct full image description corresponding to the current sentence; The first correct full-image description and the first incorrect full-image description corresponding to the same sentence are taken as a pair of correct image description and incorrect image description, and organized into a pair of positive and negative samples in the process supervision data.

7. The method according to claim 1, characterized in that The optimizing the basic multimodal large model based on the training image and the process supervision data comprises: Based on the training images and the process supervision data, the basic multimodal large model is optimized by combining the SFT algorithm and the DPO algorithm.

8. The method according to claim 7, characterized in that The combined SFT algorithm and DPO algorithm are used to optimize the basic multi-modal large model, including: Based on the correct image description in the process supervision data, using the SFT algorithm to determine the value of the first loss function; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, the DPO algorithm is used to determine the value of the second loss function, the weighted sum of the value of the first loss function and the value of the second loss function is calculated, and the parameters of the basic multimodal large model are updated based on the result of the weighted sum.

9. An optimization device for a multimodal large model for image description, characterized in that: include: a process reward model generation unit, a process supervision data generation unit, and a model optimization unit; The process reward model generation unit is used to train the basic multimodal large model based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting the image description sentence by sentence, so as to obtain the process reward model; wherein the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; The process supervision data generating unit is used to perform sentence-by-sentence reasoning of image descriptions on training images using the basic multimodal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence reasoning of the image description; and is also used to determine pairs of correct image descriptions and incorrect image descriptions based on pairs of correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning as process supervision data; The model optimization unit is used to optimize the basic multimodal large model based on the training image and the process supervision data to obtain a multimodal large model for image description that suppresses hallucinations.

10. An electronic device, characterized in that: The electronic device includes at least a computer-readable storage medium and also includes a processor; The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement the optimization method of the multimodal large model described in any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Method and system for solving illusion problem of large legal language model

    CN117744802A

  • Visual evidence-based video description object illusion correction method

    CN118887582A

  • Replay generation method and device, electronic equipment and medium

    CN119128108A

  • Training set construction method and device and electronic equipment

    CN119441470A

  • Task execution method and device based on large model, storage medium and program product

    CN119576719A

Cited By

  • Multi-modal large model illusion suppression contrast description data generation method and device

    CN121686147A