An Optimization Method, Device, and Electronic Device for a Multimodal Large Model

Through sentence-by-sentence reasoning and process-supervised data optimization, the hallucination suppression problem in image description is solved, and more efficient and accurate image description is achieved.

CN119940557BActive Publication Date: 2025-07-18HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510436082.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing multimodal large models have the problem of poor hallucination suppression in image description, especially the low efficiency of contrast decoding inference and unstable output, and the supervision reward supervision signal is unclear.

Method used

By conducting sentence-by-sentence reasoning of image description based on the basic multimodal large model, the process reward model is used to determine whether the candidate description of the current sentence is correct, and a pair of correct and incorrect image descriptions are generated as process supervision data, and the model is optimized by SFT and DPO algorithms to improve the hallucination suppression effect.

Benefits of technology

It improves the inference efficiency and accuracy of image description, effectively suppresses hallucinatory content, and improves the accuracy of model output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940557B_ABST
    Figure CN119940557B_ABST
Patent Text Reader

Abstract

The present application discloses an optimization method, device, and electronic device for a multi-modal large model. The method includes: training a basic multi-modal large model to obtain a process reward model; performing sentence-by-sentence inference on the training images for image description using the basic multi-modal large model; for each current sentence obtained by the inference, using the process reward model to determine whether each candidate description of the current sentence is correct, and using the correct candidate description of the current sentence for the next sentence inference of the image description; based on the paired correct candidate descriptions and incorrect candidate descriptions in the descriptions of each sentence obtained by the sentence-by-sentence inference, determining paired correct image descriptions and incorrect image descriptions as process supervision data; and optimizing the basic multi-modal large model based on the training images and the process supervision data to obtain a multi-modal large model that suppresses hallucinations. Applying the present application can effectively improve the performance of hallucination suppression during image description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to neural network technology, and particularly to an optimization method, device, and electronic device for multi-modal large models. Background Art

[0002] Image description has always been an important research field in the visual direction, capable of giving a text description of an image. After the rise of multi-modal large models, this field has made great progress: the detail and richness of image descriptions have been greatly improved, and image descriptions are no longer just simple descriptions of the main objects in the image. On the other hand, when multi-modal large models are applied to image descriptions, more hallucination content is introduced into the descriptions, resulting in the phenomenon of "saying" more, being right more, but also being wrong more. To suppress the problem of image description hallucinations, the current mainstream technologies include methods such as contrastive decoding and outcome-supervised reward models to optimize multi-modal large models, thereby achieving the suppression of image description hallucinations. Currently, contrastive decoding has low inference efficiency and unstable output; the supervision signal of outcome-supervised reward is not clear, and the hallucination suppression effect is not good. Summary of the Invention

[0003] This application provides an optimization method, device, and electronic device for multi-modal large models, which can effectively improve the performance of hallucination suppression when performing image descriptions.

[0004] To achieve the above object, this application adopts the following technical solutions:

[0005] An optimization method for a multi-modal large model, comprising:

[0006] Training the basic multi-modal large model based on the image description output after reasoning on the input image by the basic multi-modal large model and the label data obtained by correcting each sentence of the image description to obtain a process reward model; wherein, the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description;

[0007] Performing sentence-by-sentence reasoning on the training image for image description by using the basic multi-modal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, using the process reward model to determine whether each candidate description of the current sentence is correct, and using the correct candidate description of the current sentence for the next sentence reasoning of the image description;

[0008] Based on the paired correct candidate descriptions and incorrect candidate descriptions in the descriptions of each sentence obtained by sentence-by-sentence reasoning, determining the paired correct image descriptions and incorrect image descriptions as process supervision data;

[0009] Optimize the basic multimodal large model based on the training images and the process supervision data to obtain a multimodal large model that suppresses hallucinations.

[0010] Preferably, the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and the correct current sentence corresponding to the corrected wrong current sentence;

[0011] When training the basic multimodal large model to obtain a process reward model, the input data includes: positive samples composed of the input image, the previous correct descriptions in the image description, and the correct current sentence, and negative samples composed of the input image, the previous correct descriptions in the image description, and the wrong current sentence; the output data includes: information on whether the current sentence is correct.

[0012] Preferably, for any current sentence, the paired correct candidate description and wrong candidate description are a pair, including: one correct candidate description and one wrong candidate description selected from all the candidate descriptions of the any current sentence; or,

[0013] For any current sentence, the paired correct candidate descriptions and wrong candidate descriptions are multiple pairs, including: all paired combinations composed of any correct candidate description and any wrong candidate description among all the candidate descriptions of the any current sentence.

[0014] Preferably, the determination of the paired correct image description and wrong image description includes:

[0015] Use the paired correct candidate description and wrong candidate description corresponding to the same sentence as the paired correct image description and wrong image description, and organize them into a pair of positive and negative samples in the process supervision data.

[0016] Preferably, the determination of the paired correct image description and wrong image description includes:

[0017] For each current sentence, determine the first wrong candidate description among the paired correct candidate description and wrong candidate description of the current sentence, complete the subsequent inference of the training image based on the previous correct description of the current sentence and the first wrong candidate description, and use the first wrong candidate description and the other descriptions after it as the first wrong subsequent description corresponding to the current sentence;

[0018] Use the other descriptions after the first correct candidate description in the complete image description of the training image obtained by the sentence-by-sentence inference of the first correct candidate description paired with the first wrong candidate description as the first correct subsequent description corresponding to the current sentence;

[0019] Use the first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence as a pair of correct and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

[0020] Preferably, generating the pair of correct and incorrect image descriptions as process supervision data includes:

[0021] For each current sentence, determine the first incorrect candidate description among the correct and incorrect candidate descriptions paired with the current sentence. Based on the previous correct description of the current sentence and the first incorrect candidate description, complete the subsequent reasoning of the training image, and use the obtained complete image description as the first incorrect full-image description corresponding to the current sentence;

[0022] Use the complete image description of the training image obtained by the sentence-by-sentence reasoning with the first correct candidate description paired with the first incorrect candidate description as the first correct full-image description corresponding to the current sentence;

[0023] Use the first correct full-image description and the first incorrect full-image description corresponding to the same sentence as a pair of correct and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

[0024] Preferably, optimizing the basic multi-modal large model based on the training image and the process supervision data includes:

[0025] Based on the training image and the process supervision data, jointly optimize the basic multi-modal large model using the SFT algorithm and the DPO algorithm.

[0026] Preferably, jointly optimizing the basic multi-modal large model using the SFT algorithm and the DPO algorithm includes:

[0027] Based on the correct image description in the process supervision data, use the SFT algorithm to determine the value of the first loss function;

[0028] Based on the paired correct and incorrect image descriptions in the process supervision data, use the DPO algorithm to determine the value of the second loss function, calculate the weighted sum of the value of the first loss function and the value of the second loss function, and update the parameters of the basic multi-modal large model based on the result of the weighted sum.

[0029] An optimization device for a multi-modal large model for image description includes: a process reward model generation unit, a process supervision data generation unit, and a model optimization unit;

[0030] The process reward model generation unit is used to train the basic multi-modal large model based on the image description output after reasoning on the input image by the basic multi-modal large model and the label data obtained by correcting each sentence of the image description, so as to obtain a process reward model; wherein, the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description;

[0031] The process supervision data generation unit is used to perform sentence-by-sentence reasoning on the training image using the basic multi-modal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, use the process reward model to determine whether each candidate description of the current sentence is correct, and use the correct candidate description of the current sentence for the next sentence reasoning of the image description; it is also used to determine paired correct image descriptions and incorrect image descriptions based on the paired correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by sentence-by-sentence reasoning, as process supervision data;

[0032] The model optimization unit is used to optimize the basic multi-modal large model based on the training image and the process supervision data to obtain a multi-modal large model for image description that suppresses hallucinations.

[0033] Preferably, the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and the correct current sentence corresponding to the corrected incorrect current sentence.

[0034] In the process reward model generation unit, when training the basic multi-modal large model to obtain a process reward model, the input data includes: a positive sample composed of the input image, the previous correct description in the image description, and the correct current sentence, and a negative sample composed of the input image, the previous correct description in the image description, and the incorrect current sentence; the output data includes: information on whether the current sentence is correct.

[0035] Preferably, the paired correct candidate description and incorrect candidate description of the current sentence are a pair, including: one correct candidate description and one incorrect candidate description selected from all candidate descriptions of the current sentence; or,

[0036] The paired correct candidate description and incorrect candidate description of the current sentence are multiple pairs, including: all paired combinations composed of any correct candidate description and any incorrect candidate description among all candidate descriptions of the current sentence.

[0037] Preferably, in the process supervision data generation unit, the determination of paired correct image descriptions and incorrect image descriptions includes:

[0038] Use the paired correct candidate descriptions and incorrect candidate descriptions corresponding to the same sentence as the paired correct image descriptions and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

[0039] Preferably, in the process supervision data generation unit, the determination of the paired correct image descriptions and incorrect image descriptions includes:

[0040] For each current sentence, determine the first incorrect candidate description among the paired correct candidate descriptions and incorrect candidate descriptions of the current sentence. Complete the subsequent reasoning of the training image based on the previous correct description of the current sentence and the first incorrect candidate description, and use the first incorrect candidate description and the other descriptions after it as the first incorrect subsequent description corresponding to the current sentence;

[0041] Use the descriptions of the first correct candidate description and the other descriptions after it in the complete image description of the training image obtained through the sentence-by-sentence reasoning for the first correct candidate description paired with the first incorrect candidate description as the first correct subsequent description corresponding to the current sentence;

[0042] Use the first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence as the paired correct image descriptions and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

[0043] Preferably, in the process supervision data generation unit, the generation of the paired correct image descriptions and incorrect image descriptions as process supervision data includes:

[0044] For each current sentence, determine the first incorrect candidate description among the paired correct candidate descriptions and incorrect candidate descriptions of the current sentence. Complete the subsequent reasoning of the training image based on the previous correct description of the current sentence and the first incorrect candidate description, and use the obtained complete image description as the first incorrect full-image description corresponding to the current sentence;

[0045] Use the complete image description of the training image obtained through the sentence-by-sentence reasoning for the first correct candidate description paired with the first incorrect candidate description as the first correct full-image description corresponding to the current sentence;

[0046] Use the first correct full-image description and the first incorrect full-image description corresponding to the same sentence as the paired correct image descriptions and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

[0047] Preferably, in the model optimization unit, the optimization of the basic multi-modal large model based on the training image and the process supervision data includes:

[0048] Based on the training images and the process supervision data, the SFT algorithm and the DPO algorithm are jointly used to optimize the basic multi-modal large model.

[0049] Preferably, in the model optimization unit, the joint use of the SFT algorithm and the DPO algorithm to optimize the basic multi-modal large model includes:

[0050] Based on the correct image descriptions in the process supervision data, the value of the first loss function is determined using the SFT algorithm;

[0051] Based on the paired correct and incorrect image descriptions in the process supervision data, the value of the second loss function is determined using the DPO algorithm, the weighted sum of the value of the first loss function and the value of the second loss function is calculated, and the parameters of the basic multi-modal large model are updated based on the result of the weighted sum.

[0052] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the optimization method of the multi-modal large model described in any one of the above can be implemented.

[0053] An electronic device, which at least includes a computer-readable storage medium and also includes a processor;

[0054] The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement the optimization method of the multi-modal large model described in any one of the above.

[0055] As can be seen from the above technical solutions, in this application, based on the image descriptions output after inferring the input image by the basic multimodal large model and the labeled data obtained by correcting each sentence of the image description one by one, the basic multimodal large model is trained to obtain a process reward model, which is used to output whether each sentence in the image description is correct for the input image and its image description. Then, the basic multimodal large model is used to perform sentence-by-sentence inference on the training image, and for each current sentence obtained by the inference, the process reward model is used to mark whether each candidate description of the current sentence is correct, so that the correct description and the incorrect description at the same description position can be obtained; and the correct candidate description of the current sentence is used for the inference of the next sentence of the image description. Next, based on the paired correct candidate descriptions and incorrect candidate descriptions in each sentence description obtained by the sentence-by-sentence inference, paired correct image descriptions and incorrect image descriptions are generated as process supervision data, thereby constructing positive and negative sample pairs of the image description. On the one hand, the positive and negative sample pairs reflect whether the image description is correct, and on the other hand, they also ensure the same description position. Finally, the basic multimodal large model is optimized based on the training image and the process supervision data. Since the process reward model and the process supervision method are used to optimize the multimodal large model, the inference efficiency can be effectively improved. At the same time, the process supervision data includes positive and negative sample pairs corresponding to the same description position, further strengthening the relationship between the position and the positive and negative samples, and enabling the multimodal large model to more effectively suppress hallucinations within the full image range and improve the accuracy of the model output. Description of the Drawings

[0056] Figure 1 is a schematic diagram of the basic process of the multimodal large model optimization method in this application;

[0057] Figure 2 is a schematic diagram of the specific process of the multimodal large model optimization method in a specific embodiment of this application;

[0058] Figure 3 is a block diagram of the implementation of the process reward model generation stage in a specific embodiment of this application;

[0059] Figure 4 is a block diagram of the implementation of the process supervision data generation stage in a specific embodiment of this application;

[0060] Figure 5 is a block diagram of the implementation of the model optimization stage in a specific embodiment of this application;

[0061] Figure 6 is a schematic diagram of the basic structure of the multimodal large model optimization device in this application;

[0062] Figure 7 is a schematic diagram of the basic structure of the electronic device in this application. Detailed Embodiments

[0063] To make the purpose, technical means, and advantages of this application clearer, the following provides a more detailed description of this application in conjunction with the accompanying drawings.

[0064] The basic idea of this application is to apply the process reward model and the model optimization based on process supervision to the optimization of the multi-modal large model for image description, so as to improve the inference efficiency and the effect of hallucination suppression.

[0065] Figure 1 It is a schematic diagram of the basic process of the multi-modal large model optimization method in this application. As Figure 1 shown, this method includes:

[0066] Step 101: Based on the image description output after reasoning on the input image by the basic multi-modal large model and the label data obtained by correcting each sentence of the image description, train the basic multi-modal large model to obtain a process reward model.

[0067] This step is used to generate a process reward model. Among them, the basic multi-modal large model for image description can analyze the input image and output a text description of the image. Specifically, the image description obtained for the input image may contain multiple sentences. For the image description output after reasoning by the basic multi-modal large model, the label data corrected sentence by sentence can be obtained, and the basic multi-modal large model is trained based on the image description and the corresponding label data to generate a process reward model. This process reward model is used to output whether each sentence in the image description is correct for the input image and its image description. That is to say, the process reward model is used to judge whether each sentence in the image description is correct.

[0068] Step 102: Use the basic multi-modal large model to perform sentence-by-sentence reasoning on the training image for image description, and for each current sentence obtained by reasoning, use the process reward model to determine whether each candidate description of the current sentence is correct, and use the correct candidate description of the current sentence for the next sentence reasoning of the image description.

[0069] Steps 102 - 103 are used to generate process supervision data. Specifically, step 102 uses the basic multi-modal large model to perform sentence-by-sentence reasoning on the training image for image description, and step 103 organizes and generates process supervision data based on the reasoning result.

[0070] Among them, the image description usually includes multiple sentences. In the process of using the basic multi-modal large model to process and obtain the image description, reasoning is usually carried out sentence by sentence. When reasoning about the current sentence each time, one or more candidate descriptions of the current sentence can be inferred and output by using other descriptions before the current sentence. Then, one of the candidate descriptions is selected as the description of the current sentence. Based on this current sentence description and other previous descriptions, the reasoning of the next sentence is carried out, and so on, to obtain the complete image description of the training image.

[0071] In this application, a process reward model is introduced to participate in the sentence-by-sentence reasoning of the image description. Specifically, after inferring each current sentence, one or more candidate descriptions of the current sentence are obtained. Further, the process reward model is used to determine whether each candidate description of the current sentence is correct. For the candidate description determined by the process reward model to be correct, it is called a correct candidate description; for the candidate description determined by the process reward model to be incorrect, it is called an incorrect candidate description. The correct candidate description of the current sentence and the previous correct descriptions of the current sentence are used together for the reasoning of the next sentence of the image description. That is to say, in the sentence-by-sentence reasoning of the image description in this application, each time the correct candidate description determined by the process reward model is selected for the reasoning of the next sentence, and thus a complete image description composed entirely of correct descriptions can be obtained.

[0072] Here, it should be noted that the "previous correct descriptions" refer to a set of correct descriptions obtained before the reasoning of a certain sentence description. For example, if the image description includes 10 sentences in total, when reasoning about the 5th sentence, the previous correct descriptions refer to the correct descriptions of the first 4 sentences.

[0073] Step 103: Based on the paired correct candidate descriptions and incorrect candidate descriptions in the descriptions of each sentence obtained by sentence-by-sentence reasoning, generate paired correct image descriptions and incorrect image descriptions as process supervision data.

[0074] For any current sentence of the image description, if there is an incorrect candidate description corresponding to the current sentence, then the position of the current sentence is also the position where hallucinated descriptions may appear. If a certain current sentence corresponds to both a correct candidate description and an incorrect candidate description, then there are paired correct candidate descriptions and incorrect candidate descriptions for the current sentence, and these two descriptions correspond to the same description position.

[0075] Based on the paired correct candidate descriptions and incorrect candidate descriptions of the above current sentence, generate paired correct image descriptions and incorrect image descriptions as a pair of positive and negative samples of the process supervision data, and the positive and negative samples correspond to the same description position.

[0076] Step 104: Optimize the basic multi-modal large model based on the training image and the process supervision data to obtain a multi-modal large model that suppresses hallucinations.

[0077] This step optimizes the model based on process supervision data and training images. Specifically, various existing reinforcement learning methods can be adopted to optimize the model using the supervision data. Since the process supervision data includes positive samples and negative samples, where the positive samples represent the correct description information that the model needs to strengthen, and the negative samples represent the hallucination description information that the model needs to suppress. At the same time, the positive and negative sample pairs also correspond to the same description position, thereby further strengthening the position information contained in the samples. The model optimization carried out based on this can effectively achieve hallucination suppression in the full-image range.

[0078] So far, Figure 1 the basic process of the model optimization method in the present application shown ends. Next, the specific implementation of the model optimization method in the present application will be described through specific embodiments.

[0079] Figure 2 is a schematic diagram of the specific process of the optimization method for the multi-modal large model in a specific embodiment of the present application. The entire process includes three stages, namely: the process reward model generation stage (corresponding to step 101 in the aforementioned Figure 1 process), the process supervision data generation stage (corresponding to steps 102-103 in the aforementioned Figure 1 process), and the model optimization stage (corresponding to step 104 in the aforementioned Figure 1 process). The following will be described in detail by stage.

[0080] The first stage: the process reward model generation stage

[0081] The processing in the first stage includes Figure 2 steps 201-203 in, and its specific implementation block diagram is as Figure 3 shown.

[0082] Step 201, input the image into the basic multi-modal large model, and generate a complete image description through the inference of the basic multi-modal large model;

[0083] The processing of the basic multi-modal large model can be carried out in an existing manner. Input the image and indicate the output of the image description, then through the processing of the large model, a complete image description for the entire image can be obtained. This complete image description consists of multiple description sentences.

[0084] Step 202, obtain the label data obtained by correcting each sentence in the image description.

[0085] The label data obtained by correcting each sentence in the image description includes: the information on whether the sentence description is correct, and the correct description sentence corresponding to the corrected wrong description sentence.

[0086] The specific correction process can be carried out manually. This step is used to obtain the corresponding label data.

[0087] Step 203: Based on the image description obtained in step 201 and the label data obtained in step 202, optimize and train the basic multi-modal large model to generate a process reward model for judging whether a description sentence is correct.

[0088] For all the image descriptions obtained in step 201, it is marked whether they are correct through step 202, and the incorrect descriptions are modified. Based on these image descriptions, training data samples are organized. Among them, the correct descriptions are used as positive samples, and the incorrect descriptions are used as negative samples. Usually, in addition to including the correct or incorrect description of the current sentence in the sample, it is also necessary to include the input image in step 201 and the previous correct description of the current sentence in the image description.

[0089] Input the training samples into the multi-modal large model for processing, output the binary classification information indicating whether the current sentence is correct or incorrect, then compare it with the standard data, calculate the value of the loss function, and then adjust the model parameters according to the value of the loss function to generate a process reward model. This process reward model is used to judge whether each current description sentence is correct.

[0090] The second stage: the process supervision data generation stage

[0091] The processing in the second stage is divided into two parts: multi-modal large model inference and organizing process supervision data based on the inference results, specifically including Figure 2 Steps 204 - 206 in it, and its specific implementation block diagram is as Figure 4 shown.

[0092] Step 204: Use the basic multi-modal large model to perform sentence-by-sentence inference on the training image for image description.

[0093] To distinguish it from the input image of the basic multi-modal large model in the first stage, the input image of the basic multi-modal large model in the second stage is called the training image.

[0094] In the processing of step 204, perform sentence-by-sentence inference on the training image for image description. Taking the inference of a certain current description sentence as an example to illustrate the specific processing, hereinafter the current description sentence will be simply referred to as the current sentence.

[0095] First, input the training image and the previous correct description of the current sentence into the basic multi-modal large model, and at the same time, through text prompts, instruct the basic multi-modal large model to output the current sentence of the image description. Usually, the current sentence output by the model includes one or more candidate descriptions.

[0096] Then, input the training image, the previous correct description, and all the candidate descriptions of the output current sentence into the process reward model. The process reward model judges whether each candidate description of the current sentence is correct and makes a mark.

[0097] Finally, input the candidate descriptions marked as correct into the basic multi-modal large model for the inference of the next sentence description; save the candidate descriptions marked as incorrect.

[0098] Through the above sentence-by-sentence inference, using the correct candidate descriptions for the inference of the next sentence each time, after completing the full inference of the training image, the resulting complete image description consists entirely of correct descriptions.

[0099] Step 205: Input the incorrect candidate descriptions obtained in the sentence-by-sentence inference into the basic multi-modal large model, and obtain the subsequent descriptions of the training image through inference.

[0100] Step 205 is used to complete another image description inference process. This process is carried out for the incorrect candidate descriptions generated in the inference process of Step 204. Input the incorrect candidate descriptions, the previous correct descriptions, and the training image into the basic multi-modal large model, and instruct the large model to output the image description through text prompts. The basic multi-modal large model performs inference in the existing manner to obtain all subsequent descriptions of the training image. It should be noted here that the incorrect candidate descriptions based on which the image description inference is carried out in Step 205 can be selective. Specifically, for each current sentence, the incorrect candidate description among the paired correct candidate description and incorrect candidate description of the current sentence can be selected. The meanings of the paired correct candidate description and incorrect candidate description are introduced below.

[0101] In this application, for the purpose of organizing process supervision data, the correct candidate descriptions and incorrect candidate descriptions of the current sentence will be paired during the sentence-by-sentence inference in Step 204. Most simply, when pairing, one correct candidate description and one incorrect candidate description can be selected as the paired correct candidate description and incorrect candidate description, that is, only one pair of correct candidate description and incorrect candidate description is formed. For example, select the correct candidate description with the highest confidence and the incorrect candidate description with the lowest confidence as the paired correct candidate description and incorrect candidate description, and other candidate descriptions do not participate in subsequent processing. Or, to enrich the process supervision data, any correct candidate description and any incorrect candidate description can be paired, and all possible paired combinations are used as the paired correct candidate description and incorrect candidate description. For example, if there are M correct candidate descriptions and N incorrect candidate descriptions for the current sentence, for each correct candidate description, it can be paired with any one incorrect candidate description, then there can be M×N paired correct candidate descriptions and incorrect candidate descriptions. In the above pairing process, the paired correct candidate description and incorrect candidate description correspond to the same current description sentence, that is, they correspond to the same description position.

[0102] Based on the above pairing process, for any current sentence, there may or may not be paired correct candidate descriptions and incorrect candidate descriptions. For a current sentence without such pairs, the image reasoning process in step 205 is not required. For a current sentence with pairs, if there is only one pair of correct and incorrect candidate descriptions, only the incorrect candidate description in this pair needs to be selected to execute the image reasoning process in step 205. If there are multiple pairs of correct and incorrect candidate descriptions for the current sentence, the incorrect candidate descriptions in each pair can be sequentially used to execute the image reasoning process in step 205.

[0103] For the image reasoning process in step 205, all subsequent descriptions of the training image starting from the current sentence can be obtained. Since this reasoning process is based on an incorrect candidate description, the subsequent descriptions obtained are also inaccurate. Combining this incorrect candidate description and the subsequent descriptions obtained based on it is called an incorrect subsequent description. The previous correct description and the incorrect subsequent description can form a complete full-image description, and of course, this full-image description is also inaccurate. Hereinafter, this full-image description obtained based on the incorrect candidate description is called an incorrect full-image description. The incorrect subsequent description and the incorrect full-image description here are both based on the incorrect candidate description of the current sentence, so the incorrect subsequent description and the incorrect full-image description here are called the incorrect subsequent description and the incorrect full-image description of the current sentence. Since multiple candidate descriptions may be inferred for each current sentence, for an incorrect candidate description, multiple incorrect subsequent descriptions and incorrect full-image descriptions may be obtained through the image reasoning process.

[0104] Step 206: Based on the paired correct candidate descriptions and incorrect candidate descriptions in the descriptions obtained by sentence-by-sentence reasoning, determine the paired correct image descriptions and incorrect image descriptions as process supervision data.

[0105] The paired correct candidate descriptions and incorrect candidate descriptions have been introduced above, and the incorrect subsequent descriptions and incorrect full-image descriptions obtained by image reasoning from the incorrect candidate descriptions have been introduced through step 205. Step 206 introduces the organization of the process supervision data.

[0106] In this embodiment, to make more full use of the description position information, the process supervision data is organized based on the description position. Specifically, when organizing the process supervision data, the paired correct image descriptions and incorrect image descriptions are determined based on the paired correct candidate descriptions and incorrect candidate descriptions of the current sentence as a pair of positive and negative samples in the process supervision data.

[0107] The organization of the specific process supervision data can be divided into three different granularities:

[0108] I. Description sentence granularity

[0109] Organize the process supervision data at the sentence granularity, that is, the correct image descriptions and incorrect image descriptions are at the sentence granularity level. Specifically, for each current sentence, pair the correct candidate description and the incorrect candidate description of this current sentence to form a pair of positive and negative samples in the process supervision data; of course, the sample usually also includes the text prompt for indicating the generation of the image description and the previous correct description. The pair of positive and negative samples obtained in this way corresponds to the same current sentence, and thus corresponds to the same description position. At this granularity, the paired correct image descriptions and incorrect image descriptions are the paired correct candidate descriptions and incorrect candidate descriptions.

[0110] II. Granularity of Partial Image Descriptions

[0111] Organize the process supervision data at the granularity of the latter half description after the current sentence, that is, the correct image descriptions and incorrect image descriptions are at the granularity level of partial image descriptions. Before introducing the organization of the process supervision data at this granularity level, first introduce the meanings of the correct subsequent description and the correct full-image description of the current sentence.

[0112] For any current sentence, use the correct candidate description in the paired correct candidate description and incorrect candidate description of this current sentence for the next-sentence inference of the image description in step 204, and complete the subsequent inference of the training image according to the sentence-by-sentence inference in step 204 to obtain a complete image description composed entirely of correct descriptions; combine the aforementioned correct candidate description and the other descriptions after it in the complete image description as the correct subsequent description corresponding to the current sentence, and use the complete image description as the correct full-image description corresponding to the current sentence. It can be seen from the above process that for the current sentence, the correct subsequent description and the correct full-image description corresponding to the current sentence can be obtained through step 204, and the incorrect subsequent description and the incorrect full-image description corresponding to the current sentence can be obtained through step 205. Among them, since there can be multiple candidate descriptions, the correct full-image description, the correct subsequent description, the incorrect full-image description, and the incorrect subsequent description corresponding to the current sentence may all be multiple.

[0113] When organizing the process supervision data, for each current sentence, organize the correct subsequent description and the incorrect subsequent description inferred from the paired correct candidate description and incorrect candidate description into a pair of positive and negative samples in the process supervision data. Of course, the sample usually also includes the text prompt for indicating the generation of the image description and the previous correct description. The pair of positive and negative samples obtained in this way corresponds to the same current sentence, and thus corresponds to the same description position. At this granularity, the paired correct image descriptions and incorrect image descriptions are the paired correct subsequent descriptions and incorrect subsequent descriptions.

[0114] III. Granularity of Complete Image Descriptions

[0115] Organize the process supervision data with the complete image description as the granularity, that is to say, the correct image description and the wrong image description are at the granularity level of the complete image description. When organizing the process supervision data, for each current sentence, the correct full-image description and the wrong full-image description inferred from the paired correct candidate description and wrong candidate description are organized into a pair of positive and negative samples of the process supervision data. Of course, the samples usually also include the text prompts used to indicate the generation of the image description. The resulting pair of positive and negative samples corresponds to the same current sentence and thus the same description position. At this granularity, the paired correct image description and wrong image description are the paired correct full-image description and wrong full-image description.

[0116] As described above, the process supervision data can be organized from three different granularities, which are summarized in Table 1. In the specific implementation process, the above three different granularities can be used simultaneously to form the process supervision data, or one or two of the granularities can be selected according to needs for organizing the process supervision data.

[0117]

[0118] The third stage: the model optimization stage

[0119] The processing in the third stage specifically includes Figure 2 step 207 in [reference], and its specific implementation block diagram is as shown in Figure 5 shown.

[0120] Step 207: Optimize the basic multi-modal large model based on the training images and the process supervision data to obtain a multi-modal large model that suppresses hallucinations.

[0121] In this embodiment, based on the training images and the process supervision data, the basic multi-modal large model is optimized by combining the SFT algorithm and the DPO algorithm.

[0122] Specifically, based on the correct image descriptions in the process supervision data, use the SFT algorithm to determine the value of the first loss function SFT(Chosen);

[0123] Based on the paired correct image descriptions and wrong image descriptions in the process supervision data, use the DPO algorithm to determine the value of the second loss function DPO(Chosen, Rejected);

[0124] Calculate the weighted sum of the value of the first loss function SFT(Chosen) and the value of the second loss function DPO(Chosen, Rejected), and use it as the value of the final loss function Loss to update the parameters of the basic multi-modal large model.

[0125] In the above manner, based on the correct and incorrect image descriptions corresponding to the same description position through the DPO algorithm, the model is optimized to meet the correct expectations using contrastive learning, and the model performance is optimized based on the correct image description through the SFT algorithm. Thus, it is possible to effectively achieve hallucination suppression at all positions of the entire image.

[0126] So far, Figure 2 the method flow in the specific embodiment of the present application ends as shown.

[0127] Through the optimization method of the multimodal large model in the present application above, it is possible to train a process reward model for the needs of image description, construct sentence-by-sentence process supervision data, partial process supervision data, and full-process supervision data, and effectively suppress image description hallucinations using the process supervision data to improve the accuracy of the model output.

[0128] The above is the specific implementation of the optimization method of the multimodal large model in the present application. The present application also provides an optimization device for the multimodal large model, which can be used to implement the optimization method of the present application above. Figure 6 It is a schematic diagram of the basic structure of the optimization device for the multimodal large model provided by the present application. As Figure 6 shown, the device includes: a process reward model generation unit, a process supervision data generation unit, and a model optimization unit.

[0129] Among them, the process reward model generation unit is used to train the basic multimodal large model based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting each sentence of the image description, to obtain a process reward model; wherein, the process reward model is used to output information indicating whether each sentence in the image description is correct for the input image and its image description;

[0130] The process supervision data generation unit is used to perform sentence-by-sentence inference on the training image using the basic multimodal large model; wherein, in the sentence-by-sentence inference, for each current sentence obtained by the inference, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the inference of the next sentence of the image description; it is also used to determine paired correct and incorrect image descriptions based on the paired correct and incorrect candidate descriptions in the descriptions of each sentence obtained by the sentence-by-sentence inference, as the process supervision data;

[0131] The model optimization unit is used to optimize the basic multimodal large model based on the training image and the process supervision data to obtain a multimodal large model for image description that suppresses hallucinations.

[0132] Optionally, the labeled data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and the correct current sentence corresponding to the corrected incorrect current sentence.

[0133] In the process reward model generation unit, when training the base multi-modal large model to obtain the process reward model, the input data includes: positive samples composed of the input image, the previous correct descriptions in the image description, and the correct current sentence, and negative samples composed of the input image, the previous correct descriptions in the image description, and the incorrect current sentence; the output data includes: information on whether the current sentence is correct.

[0134] Optionally, a correct candidate description and an incorrect candidate description paired for the current sentence are a pair, including: a correct candidate description and an incorrect candidate description selected from all candidate descriptions of the current sentence; or,

[0135] The correct candidate descriptions and incorrect candidate descriptions paired for the current sentence are multiple pairs, including: all paired combinations composed of any correct candidate description and any incorrect candidate description among all candidate descriptions of the current sentence.

[0136] Optionally, in the process supervision data generation unit, determining the paired correct image descriptions and incorrect image descriptions includes:

[0137] Taking the paired correct candidate description and incorrect candidate description corresponding to the same sentence as the paired correct image description and incorrect image description, and organizing them into a pair of positive and negative samples in the process supervision data.

[0138] Optionally, in the process supervision data generation unit, determining the paired correct image descriptions and incorrect image descriptions includes:

[0139] For each current sentence, determining the first incorrect candidate description among the correct candidate descriptions and incorrect candidate descriptions paired for the current sentence, completing the subsequent reasoning of the training image based on the previous correct description of the current sentence and the first incorrect candidate description, and taking the first incorrect candidate description and the other descriptions after it as the first incorrect subsequent description corresponding to the current sentence;

[0140] Taking the descriptions after the first correct candidate description in the complete image description of the training image obtained by sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct subsequent description corresponding to the current sentence;

[0141] Taking the first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence as the paired correct image description and incorrect image description, and organizing them into a pair of positive and negative samples in the process supervision data.

[0142] Optionally, in the process supervision data generation unit, paired correct image descriptions and incorrect image descriptions are generated as process supervision data, including:

[0143] For each current sentence, determine the first incorrect candidate description among the paired correct candidate description and incorrect candidate description of the current sentence, complete the subsequent reasoning of the training image based on the previous correct description of the current sentence and the first incorrect candidate description, and use the complete image description obtained by the reasoning as the first incorrect full-image description corresponding to the current sentence;

[0144] Use the complete image description of the training image obtained by sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct full-image description corresponding to the current sentence;

[0145] Use the first correct full-image description and the first incorrect full-image description corresponding to the same sentence as the paired correct image description and incorrect image description, and organize them into a pair of positive and negative samples in the process supervision data.

[0146] Optionally, in the model optimization unit, optimize the basic multi-modal large model based on the training image and the process supervision data, including:

[0147] Optimize the basic multi-modal large model by jointly using the SFT algorithm and the DPO algorithm based on the training image and the process supervision data.

[0148] Optionally, in the model optimization unit, jointly use the SFT algorithm and the DPO algorithm to optimize the basic multi-modal large model, including:

[0149] Based on the correct image description in the process supervision data, use the SFT algorithm to determine the value of the first loss function;

[0150] Based on the paired correct image description and incorrect image description in the process supervision data, use the DPO algorithm to determine the value of the second loss function, calculate the weighted sum of the value of the first loss function and the value of the second loss function, and update the parameters of the basic multi-modal large model based on the result of the weighted sum.

[0151] This application also provides a computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed by a processor, can execute the steps in the method for optimizing a multi-modal large model as described above. In practical applications, the computer-readable medium can be included in each device / device / system of the above embodiments, or can exist separately without being assembled into the device / device / system. Among them, instructions are stored in the computer-readable storage medium, and the stored instructions can execute the steps in the method for optimizing a multi-modal large model as described above when executed by a processor.

[0152] According to the embodiments disclosed in the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above, but it is not used to limit the scope of protection of the present application. In the embodiments disclosed in the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device.

[0153] Figure 7 The present application also provides an electronic device. As Figure 7 shown, it shows a schematic structural diagram of the electronic device involved in the embodiments of the present application. Specifically:

[0154] The electronic device may include a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When executing the program in the memory 702, an optimization method for a multi-modal large model can be implemented.

[0155] Specifically, in actual applications, the electronic device may further include components such as a power supply 703 and an input / output unit 704. Those skilled in the art can understand that Figure 7 the structure of the electronic device shown in

[0156] does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0157] The memory 702 can be used to store software programs and modules, that is, the above-mentioned computer-readable storage medium. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the server, etc. In addition, the memory 702 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0158] The electronic device further includes a power supply 703 for powering each component, which can be logically connected to the processor 701 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 703 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0159] The electronic device may further include an input / output unit 704. The input / output unit 704 can be used to receive input digital or character information, and generate keyboard, mouse, joystick, and optical signal inputs related to user settings and function controls. The input / output unit 704 can also be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof.

[0160] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An optimization method for a multi-modal large model, characterized in that, Including: Training the basic multi-modal large model to obtain a process reward model based on the image description output after reasoning on the input image by the basic multi-modal large model and the label data obtained by correcting each sentence of the image description; wherein, the process reward model is used to output information indicating whether each sentence in the image description is correct for the input image and its image description; Performing sentence-by-sentence reasoning on the training image using the basic multi-modal large model; wherein, in the sentence-by-sentence reasoning, for each current sentence obtained by reasoning, using the process reward model to determine whether each candidate description of the current sentence is correct, and using the correct candidate description of the current sentence for the reasoning of the next sentence of the image description; Based on the paired correct candidate descriptions and incorrect candidate descriptions in the descriptions of each sentence obtained by the sentence-by-sentence reasoning, determining the paired correct image descriptions and incorrect image descriptions as the process supervision data; Optimizing the basic multi-modal large model based on the training image and the process supervision data to obtain a multi-modal large model that suppresses hallucinations; Wherein, the label data obtained by correcting each current sentence of the image description includes: information indicating whether the current sentence is correct, and the correct current sentence corresponding to the corrected incorrect current sentence; when training the basic multi-modal large model to obtain the process reward model, the input data includes: a positive sample composed of the input image, the previous correct description in the image description, and the correct current sentence, and a negative sample composed of the input image, the previous correct description in the image description, and the incorrect current sentence; the output data includes: information indicating whether the current sentence is correct.

2. The method according to claim 1, wherein For any current sentence, the paired correct candidate description and incorrect candidate description are a pair, including: a correct candidate description and an incorrect candidate description selected from all candidate descriptions of the any current sentence; or, For any current sentence, the paired correct candidate descriptions and incorrect candidate descriptions are multiple pairs, including: all paired combinations formed by any correct candidate description and any incorrect candidate description among all candidate descriptions of the any current sentence.

3. The method according to claim 1, wherein The determining of the paired correct image descriptions and incorrect image descriptions includes: Regarding the paired correct candidate description and incorrect candidate description corresponding to the same sentence as the paired correct image description and incorrect image description, and organizing them into a pair of positive and negative samples in the process supervision data.

4. The method according to claim 1, wherein The determining of the paired correct image descriptions and incorrect image descriptions includes: For each current sentence, determining the first incorrect candidate description among the paired correct candidate description and incorrect candidate description of the current sentence, completing the subsequent reasoning of the training image based on the previous correct description of the current sentence and the first incorrect candidate description, and using the first incorrect candidate description and other descriptions after it as the first incorrect subsequent description corresponding to the current sentence; Using the other descriptions after the first correct candidate description in the complete image description of the training image obtained by the sentence-by-sentence reasoning with the first correct candidate description paired with the first incorrect candidate description as the first correct subsequent description corresponding to the current sentence; The first correct subsequent description and the first incorrect subsequent description corresponding to the same sentence are used as a pair of correct image descriptions and incorrect image descriptions, and are organized into a pair of positive and negative samples in the process supervision data.

5. The method according to claim 1, 3 or 4, characterized in that, The determination of a pair of correct image descriptions and incorrect image descriptions as process supervision data includes: For each current sentence, determine the first incorrect candidate description among the correct candidate descriptions and incorrect candidate descriptions paired with the current sentence. Based on the previous correct description of the current sentence and the first incorrect candidate description, complete the subsequent reasoning of the training image, and use the complete image description obtained by the reasoning as the first incorrect full-image description corresponding to the current sentence; Use the complete image description of the training image obtained by the sentence-by-sentence reasoning of the first correct candidate description paired with the first incorrect candidate description as the first correct full-image description corresponding to the current sentence; Use the first correct full-image description and the first incorrect full-image description corresponding to the same sentence as a pair of correct image descriptions and incorrect image descriptions, and organize them into a pair of positive and negative samples in the process supervision data.

6. The method according to claim 1, wherein The optimization of the basic multimodal large model based on the training image and the process supervision data includes: Based on the training image and the process supervision data, jointly optimize the basic multimodal large model using the SFT algorithm and the DPO algorithm.

7. The method according to claim 6, characterized in that, The joint optimization of the basic multimodal large model using the SFT algorithm and the DPO algorithm includes: Based on the correct image descriptions in the process supervision data, use the SFT algorithm to determine the value of the first loss function; Based on the paired correct image descriptions and incorrect image descriptions in the process supervision data, use the DPO algorithm to determine the value of the second loss function, calculate the weighted sum of the value of the first loss function and the value of the second loss function, and update the parameters of the basic multimodal large model based on the result of the weighted sum.

8. An optimization device for a multi-modal large model for image description, characterized in that, Including: A process reward model generation unit, a process supervision data generation unit, and a model optimization unit; The process reward model generation unit is used to train the basic multimodal large model based on the image description output after the basic multimodal large model infers the input image and the label data obtained by correcting each sentence of the image description sentence by sentence, to obtain a process reward model; wherein, the process reward model is used to output information on whether each sentence in the image description is correct for the input image and its image description; the label data obtained by correcting each current sentence of the image description includes: information on whether the current sentence is correct, and the correct current sentence corresponding to the corrected incorrect current sentence; when training the basic multimodal large model to obtain a process reward model, the input data includes: a positive sample composed of the input image, the previous correct description in the image description, and the correct current sentence, and a negative sample composed of the input image, the previous correct description in the image description, and the incorrect current sentence; the output data includes: information on whether the current sentence is correct. The process supervision data generation unit is used to perform sentence-by-sentence inference on the training image using the basic multi-modal large model; wherein, in the sentence-by-sentence inference, for each current sentence obtained by inference, the process reward model is used to determine whether each candidate description of the current sentence is correct, and the correct candidate description of the current sentence is used for the next sentence inference of the image description; it is also used to determine paired correct image descriptions and incorrect image descriptions based on the paired correct candidate descriptions and incorrect candidate descriptions in the sentence descriptions obtained by sentence-by-sentence inference, as process supervision data. The model optimization unit is used to optimize the basic multi-modal large model based on the training image and the process supervision data to obtain a multi-modal large model for image description that suppresses hallucinations.

9. An electronic device, characterized in that, The electronic device includes at least a computer-readable storage medium and also includes a processor. The processor is used to read executable instructions from the computer-readable storage medium and execute the instructions to implement the optimization method of the multi-modal large model according to any one of claims 1 to 7 above.

Citation Information

Patent Citations

  • Visual evidence-based video description object illusion correction method

    CN118887582A