Visual question and answer multi-modal large model building method and device

By using a training data set containing images, complex prompt words and best answers, training a multimodal large model of visual Q&A has solved the problem that models in the prior art have difficulty understanding human intent, and achieving higher visual Q&A accuracy.

CN120012832AActive Publication Date: 2025-05-16HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510506137.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing multimodal big models have difficulty understanding the intention of human questions in visual questions and answers, resulting in the unsatisfactory answers.

Method used

By obtaining training data sets containing images, complex prompt words and best answers, training visual Q&A multimodal large models, and adjusting model parameters using direct preference optimization algorithms, so that the model can better understand the intent of the question and provide more accurate answers.

Benefits of technology

Improved the accuracy of visual question-and-answer, allowing the model to better understand the intent of asking questions and provide answers that meet the requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012832A_ABST
    Figure CN120012832A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual question and answer multi-mode large model establishing method and device. The method comprises the following steps: A1, acquiring a first training data set, wherein each piece of training data comprises at least one training image, a complex prompt word and an optimal answer; the complex cue word comprises a question and further comprises at least one of a background text and a constraint instruction; a2, extracting a piece of training data from the first training data set, inputting images and complex cue words in the piece of training data into a visual question and answer multi-modal large model to be trained, and outputting a predicted answer by the visual question and answer multi-modal large model; calculating a loss value according to the predicted answer and the optimal answer in the training data; adjusting parameters of the visual question and answer multi-modal large model by adopting the loss value; and returning to the step A2 until a training ending condition is met. According to the embodiment of the invention, the accuracy of visual question-answering is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal information processing, and in particular to a method and device for establishing a large multimodal model for visual question answering. Background Art

[0002] At present, multimodal big models refer to models that combine multimodal information such as text, images, videos, and audio for training. Multimodal big models have been applied in many scenarios, such as visual question answering, including: intelligent customer service applications for online shopping, specific behavior analysis in specific scenarios, and identification of illegal operations in specific scenarios. In these scenarios, images or videos are input into the multimodal big model together with questions, and the multimodal big model gives answers.

[0003] The current scheme for visual question answering using multimodal large models is as follows: a self-evolution method is designed with instruction fine-tuning data as the medium, and the multimodal large model is asked two questions in which the image is visible and the image is invisible. The former is considered to get a better answer than the latter, and direct preference optimization learning is performed accordingly. This scheme focuses more on the matching degree between the answer and the image content, and does not have the ability to follow instructions, that is, it fails to understand the human intention contained in the question, so the answer it gives may not be an ideal answer. Summary of the invention

[0004] One embodiment of the present invention proposes a visual question answering method and apparatus to improve the accuracy of visual question answering; another embodiment of the present invention proposes a non-transitory computer-readable storage medium to improve the accuracy of visual question answering.

[0005] The technical solution of the embodiment of the present invention is achieved as follows: A method for establishing a large multimodal model for visual question answering, the method comprising: A1. Obtain a first training data set, each piece of training data includes: at least one training image, a complex prompt word, and a best answer; wherein the complex prompt word includes a question, and also includes at least one of background text and constraint instructions; the background text is: background information that needs to be combined or referenced with respect to the training image when answering the question; the constraint instruction is: the description requirement of the questioner for the answer; A2. Extract a piece of training data from the first training data set, input the image and complex prompt words in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the best answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model; Return to step A2 until the training end condition is reached. The visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

[0006] Each piece of training data in the first training data set described in step A1 further includes: a non-optimal answer; Furthermore, the step A2 calculates the loss value according to the predicted answer and the best answer in the training data, including: The direct preference optimization algorithm is used to compare the predicted answer with the best answer and the non-best answer, and the loss value is calculated.

[0007] The non-optimal answer is: an intermediate answer or a worst answer.

[0008] The step A1 further comprises: B1. Obtain a first training data set, each piece of training data includes: at least one training image and a complex prompt word; and obtain an answer pool, the answer pool includes a plurality of answers corresponding to each complex prompt word in the first training data set; B2. Extract a piece of training data from the first training data set in sequence, and extract an answer corresponding to the complex prompt word in the training data from the answer pool in sequence; when an answer is extracted, the image, complex prompt word and the answer in the training data are sequentially input into the trained reward model, and the reward model outputs a reward value; when the reward value of each answer to the complex prompt word in the training data is obtained, the answer with the highest reward value is taken as the best answer, and a non-best answer is selected at the same time, and the best answer and the non-best answer are put into the training data of the first training data set; Return to step B2 until the best answer and the non-best answer are obtained for each training data in the first training data set.

[0009] The reward model is trained in the following way: C1. Extract one piece of training data from the first training data set in sequence, input the complex prompt words in the training data into multiple pre-selected answer multimodal large models respectively or input the same pre-selected answer multimodal large model trained by kernel sampling with different temperature coefficients in sequence to obtain multiple answers; provide the obtained multiple answers to the operator, and the operator annotates the score of each answer; put the complex prompt word and the multiple answers and their annotated scores into the answer pool; Return to step C1, and execute step C2 until step C1 is executed for all complex prompt words in the first training data set; C2, extracting one piece of training data from the first training data set in sequence, and taking out multiple answers and their annotation scores corresponding to the complex prompt words in the training data from the answer pool; C3, extracting one answer from multiple answers to the complex prompt word in sequence, inputting the image, complex prompt word and extracted answer in the training data into the reward model to be trained in sequence, the reward model outputs a prediction score, and uses a preset loss function to calculate the prediction score and the labeled score of the answer to obtain a loss value, and adjusts the parameters of the reward model according to the loss value; Return to step C3 until step C3 is executed for each answer to the complex prompt word, then return to step C2 until the training end condition is reached, and then a trained reward model is obtained.

[0010] Before step A1, the method further comprises: D1. Obtain a second training data set, where each piece of training data includes: at least one training image, one question, and one target answer; D2. Extract a piece of training data from the second training data set, input the image and question in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the target answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model; Return to step D2 until the training end condition is met, and the visual question answering multimodal large model is the pre-trained visual question answering multimodal large model; Furthermore, step A2 of inputting the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained is: inputting the image and complex prompt words in the training data into the pre-trained visual question answering multimodal large model.

[0011] When the complex prompt word includes a question and background text, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of question and background text, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text; Randomly select a capability type from the question capability type pool; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; The predicted question and the predicted background text are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

[0012] When the complex prompt word includes a question and a constraint instruction, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of questions and constraint instructions, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a capability type from the question capability type pool; input the text description of the image and the capability type into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; Randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the question into a trained large language model that can be used for constraint instruction synthesis, and the model sequentially outputs one or more predicted constraint instructions; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms; The predicted question and the predicted one or more constraint instructions are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

[0013] When the complex prompt word includes a question, background text, and constraint instructions, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of questions, background text and constraint instructions, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text; Randomly select a capability type from the question capability type pool; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; Randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the predicted question into a trained large language model that can be used for constraint instruction synthesis, and the model sequentially outputs one or more predicted constraint instructions; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms; The predicted question, the predicted background text and the predicted one or more constraint instructions are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

[0014] After the model outputs the predicted background text and before the randomly selecting a capability type from the problem capability type pool, the method further includes: The text description of the image and the predicted background text are input into a trained large language model that can be used for background text verification. The model outputs a prompt as to whether the predicted background text matches the text description of the image. If the prompt matches, the action of randomly selecting a capability type from the question capability type pool is executed; otherwise, the action of randomly selecting a background text example under a background type from the background text pool is returned.

[0015] After the model outputs the predicted problem and before the one or more constraint instruction examples are randomly selected from the constraint instruction pool, the method further includes: The text description of the image, the predicted background text and the predicted question are input into a trained large language model that can be used for question verification. The model outputs a prompt as to whether the predicted question matches the text description of the image and the predicted background text. If the prompt matches, the action of randomly selecting one or more constraint instruction examples from the constraint instruction pool is executed; otherwise, the action of randomly selecting a capability type from the question capability type pool is returned.

[0016] When the model outputs a predicted constraint instruction, it further includes: The predicted constraint instruction and the predicted question are input into a large language model that has been trained and can be used for constraint instruction verification. The model outputs a prompt as to whether the predicted constraint instruction and the predicted question match. If the prompt matches, the subsequent process continues; otherwise, the process returns to the action of randomly selecting a constraint instruction example from the constraint instruction pool.

[0017] The visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used, further comprising: An image of any scene and complex prompt words are input into the visual question answering multimodal large model, and the model outputs an answer.

[0018] A device for establishing a large multimodal model for visual question answering, the device comprising: The first training data set acquisition module is used to: acquire the first training data set, each piece of training data includes: at least one training image, a complex prompt word and a best answer; wherein the complex prompt word includes a question and at least one of background text and constraint instructions; the background text is: the background information that needs to be combined or referenced with respect to the training image when answering the question; the constraint instruction is: the description requirement of the questioner for the answer; The training module is used to: extract a piece of training data from the first training data set, input the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the best answer in the training data; use the loss value to adjust the parameters of the visual question answering multimodal large model until the training end condition is met, and the visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

[0019] A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, causes the processor to execute any of the above-described methods for establishing a multimodal large model for visual question answering.

[0020] In the above embodiment, the training data not only includes questions, but also includes at least one of background text and constraint instructions, so that the predicted answers output by the visual question answering multimodal large model are more consistent with the content of the image and more consistent with the questioner's questioning intention, and the answers also meet the description requirements of the questioner, thereby improving the instruction following ability of the visual question answering multimodal large model, that is, improving the accuracy of visual question answering. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0022] Figure 1 A flowchart of a method for establishing a large multimodal model for visual question answering provided by an embodiment of the present invention; Figure 2 A flow chart of a method for obtaining the best answer and non-best answer in training data provided by an embodiment of the present invention; Figure 3 A flowchart of a method for training a reward model provided by an embodiment of the present invention; Figure 4 A flowchart of a method for pre-training a large multimodal model for visual question answering provided by an embodiment of the present invention; Figure 5 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions and background texts provided by an embodiment of the present invention; Figure 6 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions and constraint instructions provided by an embodiment of the present invention; Figure 7 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions, background text and constraint instructions provided by an embodiment of the present invention; Figure 8 A schematic diagram of the structure of a device for establishing a large multimodal model for visual question answering provided by an embodiment of the present invention; Fig. 9 This is an image of a kitchen in an application example of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] It can be seen that the existing scheme of visual question answering through multimodal large models, when answering, focuses more on the matching degree between the image content and the answer, but fails to correctly understand the intention of the human question contained in the question. Moreover, after observation and analysis, the inventor found that: in many visual question answering scenarios, it is not only necessary to ask questions only about images or videos, but there will also be other additional requirements, such as: in addition to images or videos, it is also necessary to ask questions in combination with background text and constraints. For example: in the application scenario of illegal operation identification in a specific scene, if the specific scene is the back kitchen, the illegal operation identification is: the identification of operations that do not meet the hygiene standards. At this time, in addition to providing the back kitchen picture, the back kitchen operation hygiene standard text will also be provided. At this time, the existing multimodal large model for visual question answering cannot combine the back kitchen picture and the back kitchen operation hygiene standard text to give a correct answer to the question asked. For example, in a specific behavior analysis scenario in a specific scenario, a picture of an intersection is provided and the answer is required to analyze the road conditions, but the answer is limited to three sentences. The existing multimodal model for visual question answering is difficult to accurately analyze the problem using three sentences as required, or it may fail to complete the answer well due to excessive adherence to constraints.

[0026] Figure 1 Flow chart of a method for establishing a large multi-modal model for visual question answering provided by an embodiment of the present invention. Figure 1 As shown, the specific steps are as follows: Step 101: Obtain a first training data set, each piece of training data includes: at least one training image, a complex prompt word and a best answer, wherein the complex prompt word includes a question and at least one of background text and constraint instructions.

[0027] That is, there are three types of complex prompt words: 1. Question + background text, 2. Question + constraint instructions, 3. Question + background text + constraint instructions. When a complex prompt word contains background text, there is usually only one background text; when a complex prompt word contains constraint instructions, the number of constraint instructions is usually 1 to 3.

[0028] The question is: the specific problem that the questioner expects the visual question answering multimodal large model to solve, which directly reflects the needs of the questioner.

[0029] Background text refers to the background information that needs to be combined or referenced with respect to the training image when answering questions. Background text is uniquely defined background information that is unlikely to exist in the training corpus and knowledge of the large multimodal model of visual question answering. Background text can include a safety regulation, a hygiene specification, or a product manual.

[0030] Constraint instructions are the description requirements of the answer by the questioner, such as format constraints, expression constraints, grammar constraints, and / or conditional constraints. For example, constraint instructions can be one or any combination of the following: 1) The answer must start with "Summary:" and end with "Conclusion." 2) The answer cannot contain any spaces; 3) Answer using JSON format, including three keys: "title", "content" and "date"; 4) The answer must be more than 150 Chinese characters and divided into 3 paragraphs.

[0031] Step 102: extract a piece of training data from the first training data set, input the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate the loss value based on the predicted answer and the best answer in the training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model.

[0032] In this step 102, extracting a piece of training data from the first training data set may be: extracting a piece of training data from the first training data set sequentially or randomly.

[0033] Step 103: Return to step 102 until the training end condition is reached. The visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

[0034] After that, for any image in any scene, the image and complex prompt words are input into the visual question answering multimodal large model, and the model outputs the corresponding answer. Among them, the complex prompt words can be questions + background text, or questions + constraint instructions, or questions + background text + constraint instructions. Among them, the image can be: a product image in an online shopping scene, and the background text can be an introduction to the product; or, the image is an image in a traffic scene, and the background text is a traffic behavior specification; or, the image is a kitchen image, and the background text is the kitchen operation specification; or, the image is the operation screen of the operator of the production line, and the background text is the text related to the production line operation specification; or, the image is a picture of an area suspected of having safety hazards, and the background text is a list of safety inspection items for the factory; or, the image is a water surface picture, and the background text is a description of the types of floating objects that need to be reported and the types of floating objects that do not need to be paid attention to; and so on.

[0035] In the above embodiment, the training data not only includes questions, but also includes at least one of background text and constraint instructions, so that the predicted answers output by the visual question answering multimodal large model are more consistent with the content of the image and more consistent with the questioner's question intention, and the answers also meet the description requirements of the questioner, thereby improving the instruction following ability of the visual question answering multimodal large model, that is, improving the accuracy of visual question answering.

[0036] In order to further improve the instruction following ability of the large multimodal model of visual question answering, the following optimization scheme is proposed: Each piece of training data in the first training data set in step 101 further includes: a non-optimal answer; Furthermore, in step 102, a loss value is calculated based on the predicted answer and the best answer in the training data, including: using a DPO (Direct Preference Optimization) algorithm to compare the predicted answer with the best answer and the non-best answer respectively to calculate the loss value.

[0037] Among them, the DPO algorithm is a mature algorithm, and its specific calculation process is not described in detail in the embodiment of the present invention.

[0038] Among them, the non-best answer can be an intermediate answer or a worst answer.

[0039] In the above embodiment, by adding the best answer and the non-best answer to the training data, the predicted answer output by the visual question answering multimodal large model can be more and more inclined to the best answer and further away from the non-best answer during the training process, thereby further improving the instruction following ability of the visual question answering multimodal large model, that is, further improving the accuracy of visual question answering.

[0040] The following is a process of obtaining the best answer and the non-best answer in each training data in the first training data set. Obviously, the obtaining process is performed before step 101.

[0041] Figure 2 A flow chart of a method for obtaining the best answer and non-best answer in training data provided by an embodiment of the present invention, such as Figure 2 As shown, the specific steps are as follows: Step 201: Obtain a first training data set, each piece of training data includes: at least one training image and a complex prompt word; and obtain an answer pool, the answer pool includes a plurality of answers corresponding to each complex prompt word in the first training data set.

[0042] Step 201 is performed before step 101 , and the training data in the first training data set obtained at this time only contains training images and complex prompt words.

[0043] Step 202: extract one piece of training data from the first training data set obtained in step 201 in sequence, and extract one answer corresponding to the complex prompt word in the training data from the answer pool in sequence; when an answer is extracted, input the image, complex prompt word and the answer in the training data into the trained reward model in sequence, and the reward model outputs a reward value; when the reward value of each answer to the complex prompt word in the training data is obtained, take the answer with the highest reward value as the best answer, and select a non-best answer at the same time, and put the best answer and the non-best answer into the training data of the first training data set.

[0044] Among them, the non-optimal answer can be the answer with a middle reward value (i.e., the middle answer) or the answer with the lowest reward value (i.e., the worst answer).

[0045] Through step 202, the best answer and the non-best answer are obtained for a piece of training data in the first training data set.

[0046] Step 203: Return to step 202 until the best answer and the non-best answer are obtained for each training data in the first training data set.

[0047] In the above embodiment, the reward value is calculated for each of the multiple answers corresponding to each complex prompt word through the trained reward model, wherein the higher the reward value, the more accurate the answer is, and the answer with the highest reward value is taken as the best answer. At the same time, a non-best answer is selected, and the best answer and the non-best answer are put into the training data of the first training data set, thereby achieving the acquisition of the best answer and the non-best answer in the first training data set, laying a good foundation for using the first training data set for DPO training.

[0048] The training process of the reward model in step 202 is given below: Figure 3 Flow chart of a method for training a reward model provided by an embodiment of the present invention. Figure 3 As shown, the specific steps are as follows: Step 301: extract one piece of training data from the first training data set in sequence, input the complex prompt words in the training data into multiple pre-selected answer multimodal large models respectively, or input the same pre-selected answer multimodal large model trained using kernel sampling with different temperature coefficients in sequence to obtain multiple answers; provide the obtained multiple answers to the operator, and the operator annotates the score of each answer; put the complex prompt words and the multiple answers and their annotated scores into the answer pool.

[0049] Since the questions and constraints in the complex prompt words in the first training data set are various, it is very difficult to manually annotate the answers (including the best answer and the non-best answer) for each complex prompt word. Even some complex prompt words are difficult to annotate manually, so the synthesis of answers needs to be completed using a multimodal large model. In order to increase the diversity of answers and ultimately improve the accuracy of the visual question answering multimodal large model, multiple pre-selected answer multimodal large models can be used to generate different answers for the same complex prompt word, or the same pre-selected answer multimodal large model can be used but kernel sampling with different temperature coefficients to generate different answers.

[0050] Here, multiple pre-selected answer multimodal large models are used, or the same pre-selected answer multimodal large model is used but kernel sampling with different temperature coefficients is used to generate multiple answers for the same complex prompt word. The purpose is to increase the diversity of answers, so as to ultimately improve the instruction following ability of the visual question answering multimodal large model. This is because: different pre-selected answer multimodal large models, or the same multimodal large model using kernel sampling with different temperature coefficients, have different focuses when synthesizing answers. For example, a question that model A cannot answer may be answered by model B; and the answer synthesized by one model may not be correct, but using multiple models or kernel sampling with different temperature coefficients will increase the probability of obtaining an accurate answer.

[0051] Among them, the multimodal large models of the pre-selected answers have been pre-trained and have been verified to be models with better answering effects. The training data set contains a large number of complex prompt words and corresponding annotated answers.

[0052] Step 302: Return to step 301 until step 301 is executed for all complex prompt words in the first training data set, and then execute step 303.

[0053] Step 303: extract one piece of training data from the first training data set in sequence, and simultaneously extract multiple answers and their annotation scores corresponding to the complex prompt words in the training data from the answer pool.

[0054] Step 304: extract one answer from the multiple answers to the complex prompt word in turn, input the image, complex prompt word and the extracted answer in the training data into the reward model to be trained in turn, the reward model outputs a prediction score, and uses a preset loss function to calculate the prediction score and the labeled score of the answer to obtain a loss value, and adjust the parameters of the reward model according to the loss value.

[0055] Step 305: Return to step 304 until step 304 is executed for each answer to the complex prompt word, then return to step 303 until the training end condition is reached, at which time a trained reward model is obtained.

[0056] In practical applications, before using the first training data set to train the visual question answering multimodal large model, especially before using the DPO method to train the visual question answering multimodal large model, in order to improve the accuracy of the trained visual question answering multimodal large model, the visual question answering multimodal large model can be pre-trained first, and then the visual question answering multimodal large model obtained by pre-training is further trained using the second training data set to obtain the visual question answering multimodal large model used in the end. In the pre-training, complex prompt words are not used, but only simpler questions are used for training.

[0057] Figure 4 This is a flow chart of a method for pre-training a large multi-modal model for visual question answering provided by an embodiment of the present invention. Obviously, the pre-training process of the large multi-modal model for visual question answering is performed before step 101. Figure 4 As shown, the specific steps are as follows: Step 401: Obtain a second training data set, where each piece of training data includes: at least one training image, a question, and a target answer.

[0058] Step 402: extract a piece of training data from the second training data set, input the image and question in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate the loss value based on the predicted answer and the target answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model.

[0059] Step 403: Return to step 402 until the training end condition is reached. At this time, the visual question answering multimodal large model is a pre-trained visual question answering multimodal large model.

[0060] Thereafter, in step 102, the image and complex prompt words in the training data are input into the visual question answering multimodal large model to be trained: the image and complex prompt words in the training data are input into the pre-trained visual question answering multimodal large model.

[0061] Since the complex prompt words in the second training data set are diverse, specifically, the questions, background texts and constraint instructions in the complex prompt words are diverse, it is unrealistic to use manual annotation for complex prompt words. The following gives the corresponding synthesis methods according to the three composition methods of complex prompt words: Figure 5 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions and background text is provided for one embodiment of the present invention. Figure 5 As shown, the specific steps are as follows: Step 501: extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image. If the complex prompt word defined in the extracted piece of training data is composed of a question and background text, then perform the following steps: Step 501 is performed before step 101, and the purpose of this embodiment is to synthesize complex prompt words, so the first training data set at this time only includes training images.

[0062] Step 502: Input the image in the training data into the trained image-text description multimodal large model, and the model outputs the text description of the image.

[0063] The text description of an image includes: scenes, objects, and relationships between objects in the image. The relationships between objects include: spatial relationships and / or identity relationships between objects. Sufficient detailed information described in the text description of an image can be used as synthetic material for the question.

[0064] Among them, the image-text description multimodal large model has been pre-trained, and the training data set used contains a large number of images and corresponding annotated text descriptions, among which the annotated text descriptions include: text descriptions of the image scene, targets, and the relationship between targets.

[0065] Step 503: Randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text. The background text pool contains background text examples under each background type, and one background type contains one or more background text examples.

[0066] In practical applications, we first count the background types that the background text may cover, and then generate a multimodal large model through pre-trained background text examples to produce several background text examples for each background type. The background text examples under each background type constitute the background text pool. Background types include: knowledge-based background, standard-defined background, process-based background, social and application-based background, document-based background, and restricted background. The descriptions of each background type are as follows: 1) Knowledge background: It can be knowledge content in any field related to the image.

[0067] 2) Normative and standard definitional context: This can be the definition of behavioral norms, visual feature descriptions of behavior, management standards, counterfactual definitions, or counterfactual rules, etc.

[0068] 3) Process background: Give the process steps needed to solve certain problems and achieve certain goals.

[0069] 4) Social and applied background: information related to social production, life, or industry.

[0070] 5) Document background: such as the ingredient list or nutritional composition table of the food in the picture, generating the manual, usage document, or requirement document of the equipment or product in the picture.

[0071] 6) Restrictive background: Give some restrictions in the background to limit the scope or options of the large multimodal model of visual question answering to solve the problem.

[0072] The number of background text examples in each background type does not need to be large, about 10 will be enough, mainly serving as example samples for reference by the large language model that can be used for background text synthesis when generating background text.

[0073] In an optional embodiment, in order to facilitate the large language model that can be used for background text synthesis to identify which information in the input information is the text description of the image, which information is the background type, and which information is the background text example, a background text synthesis prompt word template is predefined. The background text synthesis prompt word template defines the format of the background text synthesis prompt word, and the format specifies the location of the text description of the image, the background type, and the background text example. For example, the background text synthesis prompt word template can be: "Based on the following text description of the image: '...', the following background type: '...' and the following background text example: '...' Synthesis problem", where the first place... should be filled in with the text description of the image, the second place... should be filled in with the selected background type, and the third place... should be filled in with the selected background text example. Furthermore, in this step 503, the text description of the image, the background type and the background text example are input into a large language model that has been trained and can be used for background text synthesis, specifically including: inserting the text description of the image, the background type and the background text example into a predefined background text synthesis prompt word template to obtain background text synthesis prompt words, and inputting the background text synthesis prompt words into the large language model that has been trained and can be used for background text synthesis.

[0074] Among them, the large language model that can be used for background text synthesis has been pre-trained, and the training data set used at least contains: a large number of image text descriptions and corresponding background types and background text examples, as well as corresponding annotated background text.

[0075] Step 504: randomly select a capability type from the question capability type pool; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs the predicted question.

[0076] In an optional embodiment, in order to facilitate the large language model that can be used for question synthesis to identify which information in the input information is the text description of the image, which information is the ability type, and which information is the background text, a question synthesis prompt word template is predefined. The question synthesis prompt word template defines the format of the question synthesis prompt word, and the format specifies the position of the text description of the image, the ability type, and the background text. For example, the question synthesis prompt word template can be: "Based on the following image text description: '...' and the following background text: '...', give a synthetic question under the ability type: '...'", wherein the first place... should be filled with the text description of the image, the second place... should be filled with the predicted background text, and the third place... should be filled with the selected ability type. In addition, in this step 504, the text description of the image, the ability type, and the predicted background text are input together into the trained large language model that can be used for question synthesis, specifically including: inserting the text description of the image, the ability type, and the predicted background text into the predefined question synthesis prompt word template to obtain the question synthesis prompt word, and inputting the question synthesis prompt word into the trained large language model that can be used for question synthesis.

[0077] The problem capability type pool includes all capability types covered by all problems. Capability types can be composed of two layers. The first layer is the capability categories, which include but are not limited to: recognition, conversion, analysis, positioning, reasoning, evaluation, security and other types; the second layer is the capability subcategories, each capability category is subdivided into one or more capability subcategories, such as: the recognition category can include: image description, style recognition, animal recognition, food recognition, behavior recognition, person recognition, license plate recognition and other capability subcategories.

[0078] After selecting a capability type from the question capability type pool for the question to be synthesized, the large language model that can be used for question synthesis will use this capability type as the capability type that the question needs to reflect or revolve around. For example, if the selected capability type is: behavior recognition subcategory under the recognition category, then for a classroom image, the text description of the image contains: scene: classroom, target: table and students, relationship between targets: students are on the left side of the table, then the large language model that can be used for question synthesis may form a question like "What behavior is the student on the left side of the table doing?" based on the text description of the image, the behavior recognition subcategory under the recognition category, and the background text.

[0079] Among them, the large language model that can be used for question synthesis has been pre-trained, and the training data set used contains at least: a large number of image text descriptions and corresponding ability types and background texts, as well as corresponding annotation questions.

[0080] Step 505: Use the predicted question and the predicted background text as complex prompt words for the piece of training data, and put the complex prompt words into the piece of training data in the first training data set.

[0081] After the background text and the question are synthesized, a complex prompt word can be synthesized according to a predefined complex prompt word template. The template includes but is not limited to using two line breaks to concatenate the question and the background text.

[0082] Step 506: Return to step 501 until all the training data consisting of question and background text for the complex prompt words defined in the first training data set have been synthesized into complex prompt words.

[0083] In the above embodiment, questions are synthesized based on the text description of the image and the background text, so that the questions are consistent with both the content of the image and the background text.

[0084] In an optional embodiment, in order to improve the accuracy of the synthesized background text, in step 503, after the model outputs the predicted background text and before step 504, it further includes: inputting the text description of the image and the predicted background text into a trained large language model that can be used for background text verification, and the model outputs a prompt as to whether the predicted background text and the text description of the image match; if the prompt matches, execute step 504; otherwise, return to step 503, that is, reselect the background type and background text example to re-predict a new background text until the predicted background text matches the text description of the image.

[0085] In an optional embodiment, the text description of the image and the predicted background text are input into a trained large language model that can be used for background text verification, specifically including: inserting the text description of the image and the predicted background text into a predefined background text verification prompt word template to obtain background text verification prompt words, and inputting the background text verification prompt words into a trained large language model that can be used for background text verification.

[0086] Among them, the background text verification prompt word template defines the format of the background text verification prompt word, such as: the background text verification prompt word template can be: "Judge whether the following background text: '...' and the following image text description: '...' match", where the first... should be filled in with the predicted background text, and the second... should be filled in with the text description of the image.

[0087] Among them, the large language model that can be used for background text verification has been pre-trained, and the training data set used at least contains: a large number of image text descriptions and corresponding background texts, and corresponding matching annotations.

[0088] In an optional embodiment, in order to improve the accuracy of the synthesized question, in step 504, after the model outputs the predicted question and before step 505, it further includes: inputting the text description of the image, the background text predicted in step 503 and the question predicted in step 504 into a trained large language model that can be used for question verification, and the model outputs a prompt as to whether the predicted question matches the text description of the image and the predicted background text. If the prompt matches, execute step 505; otherwise, return to step 504, that is, reselect the capability type to re-predict a new question until the predicted question matches the text description of the image and the predicted background text.

[0089] In an optional embodiment, the text description of the image, the background text predicted in step 503, and the question predicted in step 504 are input into a trained large language model that can be used for question verification, specifically including: inserting the text description of the image, the background text predicted in step 503, and the question predicted in step 504 into a predefined question verification prompt word template to obtain question verification prompt words, and inputting the question verification prompt words into a trained large language model that can be used for question verification.

[0090] Among them, the question verification prompt word template defines the format of the question verification prompt word, such as: the question verification prompt word template can be: "Judge whether the following question: '...' matches the following text description of the image: '...' and the following background text: '...'?", wherein the first place... should be filled in with the question predicted in step 504, the second place... should be filled in with the text description of the image, and the third place... should be filled in with the background text predicted in step 503.

[0091] Among them, the large language model that can be used for question verification has been pre-trained, and the training data set used at least contains: a large number of image text descriptions and corresponding background texts and questions, as well as corresponding matching annotations.

[0092] Figure 6 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions and constraint instructions is provided for one embodiment of the present invention. Figure 6 As shown, the specific steps are as follows: Step 601: extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; if the complex prompt word defined in the extracted piece of training data is composed of a question and a constraint instruction, then executing the following steps: Step 601 is performed before step 101, and the purpose of this embodiment is to synthesize complex prompt words, so the first training data set at this time only includes training images.

[0093] Step 602: Input the image in the piece of training data into the trained image-text description multimodal large model, and the model outputs the text description of the image.

[0094] Step 603: randomly select a capability type from a question capability type pool; input the text description of the image and the capability type into a trained large language model that can be used for question synthesis, and the model outputs a predicted question; wherein the question capability type pool contains all capability types covered by all questions.

[0095] Among them, the large language model that can be used for question synthesis has been pre-trained, and the training data set used contains at least: a large number of image text descriptions and corresponding ability types, as well as corresponding annotation questions.

[0096] Step 604: randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the question into a trained large language model that can be used for constraint instruction synthesis, and the model sequentially outputs one or more predicted constraint instructions; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms.

[0097] When synthesizing multiple constraint instructions in the same complex prompt word, the selected constraint instruction examples are different from each other to avoid conflicts between the synthesized constraint instructions.

[0098] In actual applications, the number of constraint instructions contained in each complex prompt word is predefined, and the number of constraint instructions is generally 1 to 3.

[0099] The constraint instruction examples in the constraint instruction pool can be manually written or generated by the model. The constraint forms include: format constraints, expression constraints, grammar constraints, and / or conditional constraints. For example, the constraint instruction examples under format constraints can be as follows: 1) The answer must start with "Summary:" and end with "Conclusion." 2) The answer cannot contain any spaces; 3) Answer using JSON format, including three keys: "title", "content" and "date"; 4) The answer must be more than 150 Chinese characters and divided into 3 paragraphs.

[0100] Individual constraint instructions generally do not contain specific questions, but are only requirements for the format of the answer and / or the expression style. Constraint instructions cannot exist independently of the question, especially for multimodal tasks. For example, if the question is to recognize text in an image, if the constraint instruction constrains the number of words, it will conflict with the question, so the synthesis of constraint instructions requires the participation of the question. In addition, if there are multiple constraint instructions, the constraint instructions need to be orthogonal to each other, and different constraint instructions need to come from different constraint forms to avoid conflicts between constraint instructions.

[0101] Constraint instructions are the requirements of how users expect the visual question answering multimodal large model to answer questions. They most directly reflect the instruction-following ability of the visual question answering multimodal large model and directly affect the user experience.

[0102] Among them, the large language model that can be used for constraint instruction synthesis has been pre-trained, and the training data set used at least includes: a large number of constraint instructions and corresponding questions, and corresponding labeled constraint instructions.

[0103] Step 605: Use the predicted question and the predicted one or more constraint instructions as complex prompt words for the piece of training data, and put the complex prompt words into the piece of training data in the first training data set.

[0104] After the question and all the constraint instructions are synthesized, the synthesized question and all the constraint instructions can be synthesized into a complex prompt word according to a preset complex prompt word template. The complex prompt word template includes but is not limited to: using two line breaks to splice the question and each constraint instruction.

[0105] Step 606: Return to step 601 until all the training data consisting of questions and constraint instructions for the complex prompt words defined in the first training data set have been synthesized into complex prompt words.

[0106] In the above embodiment, a question is synthesized based on the text description of the image so that the question conforms to the content of the image, and then a constraint instruction is synthesized based on the question so that the constraint instruction does not conflict with the question, thereby avoiding the situation where the final synthesized complex prompt word cannot answer the question, thereby improving the practicality of the synthesized complex prompt word.

[0107] In an optional embodiment, in order to improve the accuracy of the synthesized questions, in step 603, after the model outputs the predicted question and before step 604, it further includes: inputting the text description of the image and the predicted question into a trained large language model that can be used for question verification, and the model outputs a prompt as to whether the predicted question matches the text description of the image. If the prompt matches, execute step 604; otherwise, return to step 603, that is, reselect the capability type to re-predict a new question until the predicted question matches the text description of the image.

[0108] In an optional embodiment, the text description of the image and the predicted question are input into a trained large language model that can be used for question verification, specifically including: inserting the text description of the image and the predicted question into a predefined question verification prompt word template to obtain question verification prompt words, and inputting the question verification prompt words into a trained large language model that can be used for question verification.

[0109] Among them, the question verification prompt word template defines the format of the question verification prompt word, such as: the question verification prompt word template can be: "Judge whether the following question: '...' matches the text description of the following image: '...'?", where the first... should be filled in with the predicted question, and the second... should be filled in with the text description of the image.

[0110] Among them, the large language model that can be used for question verification has been pre-trained, and the training data set used contains at least: a large number of image text descriptions and corresponding questions, as well as corresponding matching annotations.

[0111] In an optional embodiment, in order to improve the accuracy of the synthesized constraint instructions, in step 604, after the model outputs a predicted constraint instruction, it further includes: inputting the predicted constraint instruction and the question predicted in step 603 into a trained large language model that can be used for constraint instruction verification, and the model outputs a prompt as to whether the predicted constraint instruction and the predicted question match; if the prompt matches, the subsequent process continues (i.e., continuing to predict the next constraint instruction for the current complex prompt word, or executing step 605 if all constraint instructions have been predicted for the current complex prompt word); otherwise, returning to step 604, i.e., reselecting an instruction constraint example to re-predict a new constraint instruction until the predicted constraint instruction and the predicted question match.

[0112] In an optional embodiment, the predicted constraint instruction and the problem predicted in step 603 are input into a trained large language model that can be used for constraint instruction verification, specifically including: inserting the predicted constraint instruction and the problem predicted in step 603 into a pre-defined constraint instruction verification prompt word template to obtain the constraint instruction verification prompt word, and inputting the constraint instruction verification prompt word into the trained large language model that can be used for constraint instruction verification.

[0113] Among them, the constraint instruction verification prompt word template defines the format of the constraint instruction verification prompt word, such as: the constraint instruction verification prompt word template can be: "Judge whether the following constraint instruction: '...' and the following question: '...' match", among which the first place... should be filled in with the constraint instruction predicted in step 604, and the second place... should be filled in with the question predicted in step 603.

[0114] Among them, the large language model that can be used for constraint instruction verification has been pre-trained, and the training data set used at least includes: a large number of constraint instructions and corresponding questions, and corresponding matching annotations.

[0115] Figure 7 A flowchart of a method for synthesizing complex prompt words when the complex prompt words include questions, background text, and constraint instructions is provided for one embodiment of the present invention. Figure 7 As shown, the specific steps are as follows: Step 701: extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; if the complex prompt word defined in the extracted piece of training data is composed of a question, background text, and constraint instructions, then executing the following steps: Step 701 is performed before step 101, and the purpose of this embodiment is to synthesize complex prompt words, so the first training data set at this time only includes training images.

[0116] Step 702: Input the image in the training data into the trained image-text description multimodal large model, and the model outputs the text description of the image.

[0117] Step 703: randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text.

[0118] Step 704: randomly select a capability type from a pool of question capability types; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs a predicted question; wherein the pool of question capability types contains all capability types covered by all questions.

[0119] Step 705: randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the predicted question into a trained large language model that can be used for constraint instruction synthesis, and the model outputs one or more predicted constraint instructions in turn; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms.

[0120] Step 706: Use the predicted question, the predicted background text, and the predicted one or more constraint instructions as complex prompt words for the piece of training data, and put the complex prompt words into the piece of training data in the first training data set.

[0121] After the background text, question and all constraint instructions are synthesized, the synthesized background text, question and all constraint instructions can be synthesized into a complex prompt word according to a preset complex prompt word template. The complex prompt word template includes but is not limited to: using two line breaks to splice the question, background text and each constraint instruction.

[0122] Step 707: Return to step 701 until all the training data consisting of questions, background texts and constraint instructions for the complex prompt words defined in the first training data set have been synthesized into complex prompt words.

[0123] In an optional embodiment, in order to improve the accuracy of the synthesized background text, in step 703, after the model outputs the predicted background text and before step 704, it further includes: inputting the text description of the image and the predicted background text into a trained large language model that can be used for background text verification, and the model outputs a prompt as to whether the predicted background text and the text description of the image match; if the prompt matches, execute step 704; otherwise, return to step 703, that is, reselect the background type and background text example to re-predict a new background text until the predicted background text matches the text description of the image.

[0124] In an optional embodiment, the text description of the image and the predicted background text are input into a trained large language model that can be used for background text verification, specifically including: inserting the text description of the image and the predicted background text into a predefined background text verification prompt word template to obtain background text verification prompt words, and inputting the background text verification prompt words into a trained large language model that can be used for background text verification.

[0125] Among them, the background text verification prompt word template defines the format of the background text verification prompt word, such as: the background text verification prompt word template can be: "Judge whether the following background text: '...' and the following image text description: '...' match", where the first... should be filled in with the predicted background text, and the second... should be filled in with the text description of the image.

[0126] In an optional embodiment, in order to improve the accuracy of the synthesized question, in step 704, after the model outputs the predicted question and before step 705, it further includes: inputting the text description of the image, the background text predicted in step 703 and the question predicted in step 704 into a trained large language model that can be used for question verification, and the model outputs a prompt as to whether the predicted question matches the text description of the image and the predicted background text. If the prompt matches, execute step 705; otherwise, return to step 704, that is, reselect the capability type to re-predict a new question until the predicted question matches the text description of the image and the predicted background text.

[0127] In an optional embodiment, the text description of the image, the background text predicted in step 703, and the question predicted in step 704 are input into a trained large language model that can be used for question verification, specifically including: inserting the text description of the image, the background text predicted in step 703, and the question predicted in step 704 into a predefined question verification prompt word template to obtain question verification prompt words, and inputting the question verification prompt words into a trained large language model that can be used for question verification.

[0128] Among them, the question verification prompt word template defines the format of the question verification prompt word, such as: the question verification prompt word template can be: "Judge whether the following question: '...' matches the following text description of the image: '...' and the following background text: '...'?", wherein the first place... should be filled in with the question predicted in step 704, the second place... should be filled in with the text description of the image, and the third place... should be filled in with the background text predicted in step 703.

[0129] In an optional embodiment, in order to improve the accuracy of the synthesized constraint instructions, in step 705, after the model outputs a predicted constraint instruction, it further includes: inputting the predicted constraint instruction and the question predicted in step 704 into a trained large language model that can be used for constraint instruction verification, and the model outputs a prompt as to whether the predicted constraint instruction and the predicted question match; if the prompt matches, the subsequent process continues (i.e., continuing to predict the next constraint instruction for the current complex prompt word, or executing step 706 if all constraint instructions have been predicted for the current complex prompt word); otherwise, returning to step 705, i.e., reselecting an instruction constraint example to re-predict a new constraint instruction until the predicted constraint instruction and the predicted question match.

[0130] In an optional embodiment, the predicted constraint instruction and the problem predicted in step 704 are input into a trained large language model that can be used for constraint instruction verification, specifically including: inserting the predicted constraint instruction and the problem predicted in step 704 into a pre-defined constraint instruction verification prompt word template to obtain the constraint instruction verification prompt word, and inputting the constraint instruction verification prompt word into the trained large language model that can be used for constraint instruction verification.

[0131] Among them, the constraint instruction verification prompt word template defines the format of the constraint instruction verification prompt word, such as: the constraint instruction verification prompt word template can be: "Judge whether the following constraint instruction: '...' and the following question: '...' match", among which the first place... should be filled in with the constraint instruction predicted in step 705, and the second place... should be filled in with the question predicted in step 704.

[0132] It should be noted that the various language models involved in the embodiments of the present invention include: a large language model that can be used for background text synthesis, a large language model that can be used for question synthesis, a large language model that can be used for constraint instruction synthesis, a large language model that can be used for background text verification, a large language model that can be used for question verification, and a large language model that can be used for constraint instruction verification. These six large language models can be trained independently and use the training data required by each during training; these six large language models can also use the same large language model, that is, the large language model has six functions at the same time: background text synthesis, question synthesis, constraint instruction synthesis, background text verification, question verification and constraint instruction verification, and the training data required by these six functions can be used during training.

[0133] It should be noted that the reason why the embodiment of the present invention adopts a large language model to generate the problem with the text description of the image as input instead of adopting a multimodal large model to directly generate the problem with the image as input is as follows: On the one hand, the instruction-following ability of the large multimodal model with the same parameter scale is weaker than that of the large language model. In the tasks of background text synthesis, question synthesis and constraint instruction synthesis, the large multimodal model cannot accurately understand the intention of the task, resulting in poor quality of the generated background text, questions and constraint instructions.

[0134] On the other hand, using images as input will make the synthesized background text, questions, and constraint instructions pay more attention to the subject in the image, and pay less attention to the details or small targets in the image, reducing the diversity and complexity of the synthesized background text, questions, and constraint instructions. For example: the subject of an image is a person, but the distant view in the image also contains a puppy running towards the person; then, if the image is used as input to synthesize the question, the synthesized question may only focus on the person, but not the puppy; but if the text description of the image is used as input to synthesize the question, then the targets in the text description of the image will include both the person and the puppy, as well as the relationship between the person and the puppy: the puppy is behind the person, so the synthesized question may contain information about the puppy.

[0135] It should be noted that the training data in the first training data set and the second training data set in the embodiment of the present invention need to be sufficient to improve the applicability and accuracy of the final visual answer multimodal large model.

[0136] Figure 8 This is a schematic diagram of the structure of a device for building a large multi-modal model for visual question answering provided by an embodiment of the present invention. Figure 8 As shown, it mainly includes: a first training data set acquisition module 81 and a training module 82, wherein: The first training data set acquisition module 81 is used to: acquire the first training data set, each training data includes: at least one training image, a complex prompt word and a best answer; wherein the complex prompt word includes a question and at least one of background text and constraint instructions; the background text is: the background information that needs to be combined or referenced with respect to the training image when answering the question; the constraint instruction is: the description requirements of the questioner for the answer.

[0137] The training module 82 is used to: extract a piece of training data from the first training data set, input the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate the loss value according to the predicted answer and the best answer in the training data; use the loss value to adjust the parameters of the visual question answering multimodal large model until the training end condition is met, and the visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

[0138] In an optional embodiment, each training data of the first training data set acquired by the first training data set acquisition module 81 further includes: non-optimal answers; and the training module 82 calculates the loss value based on the predicted answer and the best answer in the training data, including: using a direct preference optimization algorithm to compare the predicted answer with the best answer and the non-optimal answer, respectively, to calculate the loss value.

[0139] In an optional embodiment, the first training data set acquisition module 81 acquires the first training data set, each piece of training data includes: at least one training image, a complex prompt word and a best answer, and is further used to: B1. Obtain a first training data set, each piece of training data includes: at least one training image and a complex prompt word; and obtain an answer pool, the answer pool includes a plurality of answers corresponding to each complex prompt word in the first training data set; B2, extracting one piece of training data from the first training data set in step B1, and extracting one answer corresponding to the complex prompt word in the training data from the answer pool in step B1; when an answer is extracted, the image, complex prompt word and the answer in the training data are input into the trained reward model in sequence, and the reward model outputs a reward value; when the reward value of each answer to the complex prompt word in the training data is obtained, the answer with the highest reward value is taken as the best answer, and a non-best answer is selected at the same time, and the best answer and the non-best answer are put into the training data in the first training data set; Return to step B2 until the best answer and the non-best answer are obtained for each training data in the first training data set.

[0140] In an optional embodiment, the above device further includes a reward model training module, which is used to: C1. Extract one piece of training data from the first training data set in sequence, input the complex prompt words in the training data into multiple pre-selected answer multimodal large models respectively or input the same pre-selected answer multimodal large model trained by kernel sampling with different temperature coefficients in sequence to obtain multiple answers; provide the obtained multiple answers to the operator, and the operator annotates the score of each answer; put the complex prompt word and the multiple answers and their annotated scores into the answer pool; Return to step C1, and execute step C2 until step C1 is executed for all complex prompt words in the first training data set; C2, extracting one piece of training data from the first training data set in sequence, and taking out multiple answers and their annotation scores corresponding to the complex prompt words in the training data from the answer pool; C3, extracting one answer from multiple answers to the complex prompt word in sequence, inputting the image, complex prompt word and extracted answer in the training data into the reward model to be trained in sequence, the reward model outputs a prediction score, and uses a preset loss function to calculate the prediction score and the labeled score of the answer to obtain a loss value, and adjusts the parameters of the reward model according to the loss value; Return to step C3 until step C3 is executed for each answer to the complex prompt word, then return to step C2 until the training end condition is reached, and then a trained reward model is obtained.

[0141] In an optional embodiment, the above device further includes a pre-training module for: D1. Obtain a second training data set, where each piece of training data includes: at least one training image, one question, and one target answer; D2. Extract a piece of training data from the second training data set, input the image and question in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the target answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model; Return to step D2 until the training end condition is met, and the visual question answering multimodal large model is the pre-trained visual question answering multimodal large model; Furthermore, the training module 82 inputs the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained by: inputting the image and complex prompt words in the training data into the pre-trained visual question answering multimodal large model.

[0142] In an optional embodiment, the above device further includes a first complex prompt word synthesis module, which is used to: A training data is extracted from the first training data set in sequence, wherein each training data in the first training data set includes at least one training image; if the complex prompt word defined in the extracted training data is composed of a question and a background text, the image in the training data is input into a trained image-text description multimodal large model, and the model outputs a text description of the image; a background text example under a background type is randomly selected from a background text pool; the text description of the image, the background type and the background text example are input into a trained large language model that can be used for background text synthesis, and the model outputs a predicted background text; a capability type is randomly selected from a question capability type pool; the text description of the image, the capability type and the predicted background text are input into a trained large language model that can be used for question synthesis, and the model outputs a predicted question; wherein the question capability type pool contains all capability types covered by all questions; the predicted question and the predicted background text are used as the complex prompt word of the training data, and the complex prompt word is put into the training data of the first training data set.

[0143] In an optional embodiment, the above device further includes a second complex prompt word synthesis module, which is used to: A training data is extracted from the first training data set in sequence, wherein each training data in the first training data set includes at least one training image; if the complex prompt word defined in the extracted training data is composed of a question and a constraint instruction, the image in the training data is input into a trained image-text description multimodal large model, and the model outputs a text description of the image; a capability type is randomly selected from a problem capability type pool; the text description of the image and the capability type are input into a trained large language model that can be used for question synthesis, and the model outputs a predicted question; wherein the problem capability type pool contains all capability types covered by all questions; one or more constraint instruction examples are randomly selected from a constraint instruction pool, and each selected constraint instruction example is input into a trained large language model that can be used for constraint instruction synthesis together with the question, and the model outputs one or more predicted constraint instructions in sequence; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms; the predicted question and the predicted one or more constraint instructions are used as the complex prompt word of the training data, and the complex prompt word is put into the training data of the first training data set.

[0144] In an optional embodiment, the above device further includes a third complex prompt word synthesis module, which is used to: Extract one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; if the complex prompt word defined in the extracted training data is composed of a question, background text and a constraint instruction, input the image in the training data into a trained image-text description multimodal large model, and the model outputs a text description of the image; randomly select a background text example under a background type from a background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs a predicted background text; randomly select an ability type from a question ability type pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs a predicted background text; The predicted background text is input into a trained large language model that can be used for question synthesis, and the model outputs a predicted question; wherein the question ability type pool includes all ability types covered by all questions; one or more constraint instruction examples are randomly selected from the constraint instruction pool, and each selected constraint instruction example is respectively input into a trained large language model that can be used for constraint instruction synthesis together with the predicted question, and the model sequentially outputs the predicted one or more constraint instructions; wherein the constraint instruction pool includes constraint instruction examples under different constraint forms; the predicted question, the predicted background text and the predicted one or more constraint instructions are used as complex prompt words for the training data, and the complex prompt words are put into the training data of the first training data set.

[0145] In an optional embodiment, after the model outputs the predicted background text and before randomly selecting a capability type from the question capability type pool, the third complex prompt word synthesis module further includes: inputting the text description of the image and the predicted background text into a trained large language model that can be used for background text verification, and the model outputs a prompt as to whether the predicted background text and the text description of the image match; if the prompt matches, executing the action of randomly selecting a capability type from the question capability type pool; otherwise, returning to the action of randomly selecting a background text example under a background type from the background text pool.

[0146] In an optional embodiment, after the model outputs the predicted question and before randomly selecting one or more constraint instruction examples from the constraint instruction pool, the third complex prompt word synthesis module further includes: inputting the text description of the image, the predicted background text and the predicted question into a trained large language model that can be used for question verification, and the model outputs a prompt as to whether the predicted question matches the text description of the image and the predicted background text; if the prompt matches, executing the action of randomly selecting one or more constraint instruction examples from the constraint instruction pool; otherwise, returning to the action of randomly selecting a capability type from the question capability type pool.

[0147] In an optional embodiment, after the model outputs a predicted constraint instruction, the third complex prompt word synthesis module further includes: inputting the predicted constraint instruction and the predicted question into a trained large language model that can be used to verify the constraint instruction, and the model outputs a prompt as to whether the predicted constraint instruction and the predicted question match, and if the prompt matches, continuing the subsequent process; otherwise, returning to the action of randomly selecting a constraint instruction example from the constraint instruction pool.

[0148] In an optional embodiment, the above-mentioned device further includes an answer module, which is used to: input the image and complex prompt words of any scene into the visual question answering multimodal large model to be finally used, and the model outputs an answer.

[0149] An application example of the present invention is given below: Fig. 9 It is an image of a kitchen. The compound prompt words given by the questioner are as follows: The background text is: "Commercial Kitchen Hygiene Standards": The countertops must be kept clean and no food debris should be left; scissors must be washed and hung back on the rack after use, and cannot be placed alone on the countertop; oil stains on the wall must be cleaned in time, and no yellow or brown oil spots should appear; ... (The content of the standard is long, so the rest is omitted); The question is: Based on the above specifications, please determine whether there are any violations in the diagram; The binding instructions are: only explain the violations. If there are no violations, directly answer "there are no violations."

[0150] After the image and the complex prompt words consisting of background text, questions and constraint instructions are input into the visual question answering multimodal large model obtained by using the embodiment of the present invention, the answer output by the model is as follows: There is a violation. It is clearly visible in the picture that a pair of red scissors is placed alone on the countertop, which violates Article 2 of the Commercial Kitchen Hygiene Standards: "Scissors must be washed and hung back on the rack after use, and cannot be placed alone on the countertop."

[0151] It can be seen that the answer is consistent with the image content, the background text and the constraint instructions, and fully responds to the questioner's intention to ask the question.

[0152] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores instructions that, when executed by a processor, enable the processor to execute the method for establishing a large multimodal model for visual question answering as described in any of the aforementioned embodiments.

[0153] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways, and all of these combinations and / or combinations fall within the scope disclosed in the present application.

[0154] Specific embodiments are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core ideas of the present invention and is not used to limit the present application. For those skilled in the art, changes can be made in the specific implementation methods and application scopes according to the ideas, spirits and principles of the present invention, and any modifications, equivalent substitutions, improvements, etc. made therein should be included in the scope of protection of this application.

Claims

1. A method for establishing a large multimodal model for visual question answering, characterized in that: The method includes: A1. Obtain a first training data set, each piece of training data includes: at least one training image, a complex prompt word, and a best answer; wherein the complex prompt word includes a question, and also includes at least one of background text and constraint instructions; the background text is: background information that needs to be combined or referenced with respect to the training image when answering the question; the constraint instruction is: the description requirement of the questioner for the answer; A2. Extract a piece of training data from the first training data set, input the image and complex prompt words in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the best answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model; Return to step A2 until the training end condition is reached. The visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

2. The method according to claim 1, characterized in that Each piece of training data in the first training data set described in step A1 further includes: a non-optimal answer; Furthermore, the step A2 calculates the loss value according to the predicted answer and the best answer in the training data, including: The direct preference optimization algorithm is used to compare the predicted answer with the best answer and the non-best answer, and the loss value is calculated.

3. The method according to claim 2, characterized in that The non-optimal answer is: an intermediate answer or a worst answer.

4. The method according to claim 2 or 3, characterized in that: The step A1 further comprises: B1. Obtain a first training data set, each piece of training data includes: at least one training image and a complex prompt word; and obtain an answer pool, the answer pool includes a plurality of answers corresponding to each complex prompt word in the first training data set; B2. Extract a piece of training data from the first training data set in sequence, and extract an answer corresponding to the complex prompt word in the training data from the answer pool in sequence; when an answer is extracted, the image, complex prompt word and the answer in the training data are sequentially input into the trained reward model, and the reward model outputs a reward value; when the reward value of each answer to the complex prompt word in the training data is obtained, the answer with the highest reward value is taken as the best answer, and a non-best answer is selected at the same time, and the best answer and the non-best answer are put into the training data of the first training data set; Return to step B2 until the best answer and the non-best answer are obtained for each training data in the first training data set.

5. The method according to claim 4, characterized in that The reward model is trained in the following way: C1. Extract one piece of training data from the first training data set in sequence, input the complex prompt words in the training data into multiple pre-selected answer multimodal large models respectively, or input the same pre-selected answer multimodal large model trained by kernel sampling with different temperature coefficients in sequence, to obtain multiple answers; provide the obtained multiple answers to the operator, and the operator marks the score of each answer; Putting the complex prompt word, the multiple answers and their marked scores into an answer pool; Return to step C1, and execute step C2 until step C1 is executed for all complex prompt words in the first training data set; C2, extracting one piece of training data from the first training data set in sequence, and taking out multiple answers and their annotation scores corresponding to the complex prompt words in the training data from the answer pool; C3, extracting one answer from multiple answers to the complex prompt word in sequence, inputting the image, complex prompt word and extracted answer in the training data into the reward model to be trained in sequence, the reward model outputs a prediction score, and uses a preset loss function to calculate the prediction score and the labeled score of the answer to obtain a loss value, and adjusts the parameters of the reward model according to the loss value; Return to step C3 until step C3 is executed for each answer to the complex prompt word, then return to step C2 until the training end condition is reached, and then a trained reward model is obtained.

6. The method according to any one of claims 1 to 5, characterized in that: Before step A1, the method further comprises: D1. Obtain a second training data set, where each piece of training data includes: at least one training image, one question, and one target answer; D2. Extract a piece of training data from the second training data set, input the image and question in the piece of training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the target answer in the piece of training data; and use the loss value to adjust the parameters of the visual question answering multimodal large model; Return to step D2 until the training end condition is met, and the visual question answering multimodal large model is the pre-trained visual question answering multimodal large model; Furthermore, step A2 of inputting the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained is: inputting the image and complex prompt words in the training data into the pre-trained visual question answering multimodal large model.

7. The method according to claim 1, characterized in that When the complex prompt word includes a question and background text, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of question and background text, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text; Randomly select a capability type from the question capability type pool; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; The predicted question and the predicted background text are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

8. The method according to claim 1, characterized in that When the complex prompt word includes a question and a constraint instruction, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of questions and constraint instructions, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a capability type from the question capability type pool; input the text description of the image and the capability type into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; Randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the question into a trained large language model that can be used for constraint instruction synthesis, and the model sequentially outputs one or more predicted constraint instructions; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms; The predicted question and the predicted one or more constraint instructions are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

9. The method according to claim 1, characterized in that: When the complex prompt word includes a question, background text, and constraint instructions, the complex prompt word is obtained in the following manner: Extracting one piece of training data from the first training data set in sequence, wherein each piece of training data in the first training data set includes at least one training image; If the complex prompt word defined in the extracted training data is composed of questions, background text and constraint instructions, the image in the training data is input into the trained image-text description multimodal large model, and the model outputs the text description of the image; Randomly select a background text example under a background type from the background text pool; input the text description of the image, the background type and the background text example into a trained large language model that can be used for background text synthesis, and the model outputs the predicted background text; Randomly select a capability type from the question capability type pool; input the text description of the image, the capability type and the predicted background text into a trained large language model that can be used for question synthesis, and the model outputs the predicted question; wherein the question capability type pool contains all capability types covered by all questions; Randomly select one or more constraint instruction examples from the constraint instruction pool, input each selected constraint instruction example together with the predicted question into a trained large language model that can be used for constraint instruction synthesis, and the model sequentially outputs one or more predicted constraint instructions; wherein the constraint instruction pool contains constraint instruction examples under different constraint forms; The predicted question, the predicted background text and the predicted one or more constraint instructions are used as complex prompt words for the piece of training data, and the complex prompt words are put into the piece of training data in the first training data set.

10. The method according to claim 9, characterized in that After the model outputs the predicted background text and before the randomly selecting a capability type from the problem capability type pool, the method further includes: The text description of the image and the predicted background text are input into a trained large language model that can be used for background text verification. The model outputs a prompt as to whether the predicted background text matches the text description of the image. If the prompt matches, the action of randomly selecting a capability type from the question capability type pool is executed; otherwise, the action of randomly selecting a background text example under a background type from the background text pool is returned.

11. The method according to claim 9, characterized in that After the model outputs the predicted problem and before the one or more constraint instruction examples are randomly selected from the constraint instruction pool, the method further includes: The text description of the image, the predicted background text and the predicted question are input into a trained large language model that can be used for question verification. The model outputs a prompt as to whether the predicted question matches the text description of the image and the predicted background text. If the prompt matches, the action of randomly selecting one or more constraint instruction examples from the constraint instruction pool is executed; otherwise, the action of randomly selecting a capability type from the question capability type pool is returned.

12. The method according to claim 9, characterized in that When the model outputs a predicted constraint instruction, it further includes: The predicted constraint instruction and the predicted question are input into a large language model that has been trained and can be used for constraint instruction verification. The model outputs a prompt as to whether the predicted constraint instruction and the predicted question match. If the prompt matches, the subsequent process continues; otherwise, the process returns to the action of randomly selecting a constraint instruction example from the constraint instruction pool.

13. The method according to claim 1, characterized in that The visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used, further comprising: An image of any scene and complex prompt words are input into the visual question answering multimodal large model, and the model outputs an answer.

14. A device for building a large multimodal model for visual question answering, characterized in that: The device includes: The first training data set acquisition module is used to: acquire the first training data set, each piece of training data includes: at least one training image, a complex prompt word and a best answer; wherein the complex prompt word includes a question and at least one of background text and constraint instructions; the background text is: the background information that needs to be combined or referenced with respect to the training image when answering the question; the constraint instruction is: the description requirement of the questioner for the answer; The training module is used to: extract a piece of training data from the first training data set, input the image and complex prompt words in the training data into the visual question answering multimodal large model to be trained, and the visual question answering multimodal large model outputs a predicted answer; calculate a loss value based on the predicted answer and the best answer in the training data; use the loss value to adjust the parameters of the visual question answering multimodal large model until the training end condition is met, and the visual question answering multimodal large model at this time is the visual question answering multimodal large model finally used.

15. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores instructions, which, when executed by a processor, cause the processor to execute the method for establishing a visual question answering multimodal large model as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Question and answer pair construction method and device, electronic equipment and storage medium

    CN117688160A

  • Image annotation data generation method and device, image annotation data model training method and device, equipment and medium

    CN118097694A

  • Question and answer model training method, text processing method and reward model training method

    CN118350463A

  • Visual question-answering model training method, visual question-answering method and visual question-answering system

    CN118798373A

  • Visual question and answer method and device, equipment, storage medium and product

    CN118916471A