Method for training description generation model, description generation method, device and electronic device
Through a description generation model training method, the interaction between text prompts and description generation model is optimized, and the accuracy of image area description text generation is solved, and efficient and accurate description text generation is achieved.
Patent Information
- Application Number
- CN202510285822.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art is difficult to accurately generate description text of image areas, which affects application effects.
A description generation model training method is provided, and the parameters of the description generation model are optimized by using the first text prompt and the current description generation model, processing each first sample image, generating a description text, and differentially adjusting the matching result with the second sample description text.
It realizes the accurate generation of description text of image areas, which improves the accuracy and efficiency of description generation models.
Smart Images

Figure CN119810593B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and particularly to a method for training a description generation model, a description generation method, an apparatus, and an electronic device. Background Art
[0002] With the development of image processing technology, in some application scenarios (such as image search by text or image retrieval, etc.), it is often necessary to generate a description text for a specified image region in an image. The accuracy of generating the description text for the image region has a significant impact on the application effect. Therefore, how to accurately generate the description text for the image region has become an urgent problem to be solved. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a method for training a description generation model, a description generation method, an apparatus, and an electronic device to accurately generate a description text for a specified image region in an image. The specific technical solutions are as follows:
[0004] In the first aspect of the embodiments of the present application, a method for training a description generation model is provided. The method includes:
[0005] Using a first text prompt and the current description generation model, process each current first sample image to obtain a description text for describing the specified image region in the first sample image, as the current first sample description text; wherein, the first text prompt is used to indicate generating a description text for the specified image region in the input image.
[0006] Determine, from the current first sample description texts, a description text that matches the specified image region in the corresponding first sample image as the current second sample description text.
[0007] Input the first sample image corresponding to each current second sample description text and the first text prompt into the current description generation model to obtain a description text for the specified image region in the first sample image, as the current first predicted description text.
[0008] Based on the difference between the current first predicted description text and the current second sample description text, adjust the parameters of the current description generation model to obtain a new description generation model.
[0009] Optionally, the determining, from the current first sample description texts, a description text that matches the specified image region in the corresponding first sample image as the current second sample description text includes:
[0010] Input each current first sample image, the first sample description text corresponding to the first sample image, and a second text prompt into a pre-trained reward model to obtain a matching result of the specified image region in the first sample image and the corresponding first sample description text; wherein, the second text prompt is used to indicate whether the specified image region in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, the specified image region in the second sample image, and a description text that matches the specified image region in the second sample image; a negative sample pair includes: a third sample image, the specified image region in the third sample image, and a description text that does not match the specified image region in the third sample image;
[0011] In the case where the obtained matching result indicates that the specified image region in the first sample image matches the corresponding first sample description text, determine the first sample description text corresponding to the first sample image as the corresponding second sample description text.
[0012] Optionally, the reward model includes: a first large language model, a first adapter, a first image encoder, and a first token parser; the first image encoder is used to extract the features of the input image; the first adapter is used to convert the image features output by the first image encoder into a token form; the first token parser is used to convert the input text into a token form.
[0013] Optionally, the description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the reward model is trained through the following steps:
[0014] For each sample pair, input the sample image and the specified image region in the sample pair into the first image encoder in the reward model with an initial structure to obtain the features of the sample image and the features of the specified image region in the sample pair, and splice the obtained features;
[0015] Input the obtained splicing result into the first adapter in the reward model with an initial structure to obtain the token form of the splicing result;
[0016] Input the second text prompt and the description text in the sample pair into the first token parser in the reward model with an initial structure to obtain the token form of the second text prompt and the token form of the description text in the sample pair;
[0017] Input the tokenized form of the splicing result, the tokenized form of the second text prompt, and the tokenized form of the descriptive text in the sample pair into the first large language model in the reward model of the initial structure to obtain the predicted matching result between the specified image region and the descriptive text in the sample pair;
[0018] Based on the difference between the obtained predicted matching result and the true matching result of the sample pair, adjust the parameters of the reward model of the initial structure until the preset convergence condition is reached to obtain the trained reward model.
[0019] Optionally, the process of using the first text prompt and the current description generation model to process each current first sample image to obtain the descriptive text for describing the specified image region in the first sample image as the current first sample descriptive text includes:
[0020] For each current first sample image, when the proportion of the specified image region in the first sample image is less than the preset threshold, expand the specified image region in the first sample image to obtain the image to be processed;
[0021] Input the obtained image to be processed and the first text prompt into the current description generation model to obtain the descriptive text for describing the specified image region in the first sample image as the current first sample descriptive text;
[0022] When the proportion of the specified image region in the first sample image is not less than the preset threshold, input the first sample image and the first text prompt into the current description generation model to obtain the descriptive text for describing the specified image region in the first sample image as the current first sample descriptive text.
[0023] Optionally, the process of using the first text prompt and the current description generation model to process each current first sample image to obtain the descriptive text for describing the specified image region in the first sample image as the current first sample descriptive text includes:
[0024] Use the first text prompt and the current description generation model to process the first sample image according to different sampling parameters to obtain multiple descriptive texts for the specified image region in the first sample image as the current first sample descriptive text.
[0025] Optionally, after adjusting the parameters of the current description generation model based on the difference between the current first predicted descriptive text and the current second sample descriptive text to obtain a new description generation model, the method further includes:
[0026] Obtain a fourth sample image; among the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image region in the corresponding fourth sample image;
[0027] Use the first text prompt and the current description generation model to process each obtained fourth sample image according to different sampling parameters to obtain multiple description texts for the specified image region in the fourth sample image;
[0028] From the obtained description texts, determine at least one positive sample description text that matches the specified image region in the fourth sample image, and at least one negative sample description text that does not match the specified image region in the fourth sample image;
[0029] Input the fourth sample image and the first text prompt into the current description generation model to obtain the description text for the specified image region in the fourth sample image as the current second predicted description text;
[0030] Based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text, adjust the parameters of the current description generation model to obtain a new description generation model.
[0031] Optionally, the first text prompt is further used to indicate that the current description generation model generates description texts conforming to the first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included;
[0032] After adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, the method further includes:
[0033] Use the first text prompt and the current description generation model to process each current first sample image to obtain the description text in the first specified description style for the specified image region in the first sample image as the current third sample description text;
[0034] Adjust the description style of the current third sample description text from the first specified description style to the second specified description style to obtain the current fourth sample description text;
[0035] From the current fourth sample description text, determine the description text that matches the specified image region in the corresponding first sample image to obtain the current fifth sample description text;
[0036] Input the first sample image and the third text prompt corresponding to each current fifth sample description text into the current description generation model, and obtain the description text in the second specified description style for the specified image region in the first sample image as the current third predicted description text;
[0037] Based on the difference between the current third predicted description text and the current fifth sample description text, adjust the parameters of the current description generation model to obtain a new description generation model.
[0038] Optionally, after adjusting the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model, the method further includes:
[0039] Input the current first sample image and the third text prompt into the current description generation model, and obtain the description text in the second specified description style for the specified image region in the first sample image as the current fourth sample description text;
[0040] Return to execute the step of determining the description text that matches the specified image region in the corresponding first sample image from the current fourth sample description text to obtain the current fifth sample description text until a first preset condition is reached to obtain a new description generation model.
[0041] Optionally, after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, the method further includes:
[0042] Return to execute the step of using the first text prompt and the current description generation model to process each current first sample image to obtain the description text for describing the specified image region in the first sample image as the current first sample description text until a second preset condition is reached to obtain a new description generation model.
[0043] Optionally, the description generation model includes: a second large language model, a second adapter, a second image encoder, and a second token parser; the second image encoder is used to extract the features of the input image; the second adapter is used to convert the image features output by the second image encoder into a token form; the second token parser is used to convert the input text into a token form.
[0044] In a second aspect of the embodiments of the present application, a description generation method is further provided, and the method includes:
[0045] Obtain the image to be used and the text prompt to be used;
[0046] Input the to-be-utilized image and the to-be-utilized text prompt into a pre-trained description generation model to obtain a description text of the image region indicated by the to-be-utilized text prompt in the to-be-utilized image; wherein, the description generation model is trained based on any one of the above-described description generation model training methods.
[0047] Optionally, the position of the image region indicated by the to-be-utilized text prompt in the to-be-utilized image is obtained through the following steps:
[0048] Display the to-be-utilized image;
[0049] In response to a user's box selection operation on the displayed to-be-utilized image, determine the position indicated by the box selection operation in the to-be-utilized image to obtain the position of the image region indicated by the to-be-utilized text prompt in the to-be-utilized image.
[0050] Optionally, the description style of the obtained description text is the description style indicated by the to-be-utilized text prompt; a description style is used to indicate the language of the description text and / or the number of statements included.
[0051] In the third aspect of the embodiments of the present application, a description generation model training device is further provided, and the device includes:
[0052] A first sample description text generation module, configured to use a first text prompt and the current description generation model to process each current first sample image to obtain a description text for describing the specified image region in the first sample image as the current first sample description text; wherein, the first text prompt is used to indicate the generation of a description text of the specified image region in the input image.
[0053] A second sample description text determination module, configured to determine, from the current first sample description texts, a description text that matches the specified image region in the corresponding first sample image as the current second sample description text.
[0054] A first predicted description text generation module, configured to input the first sample image corresponding to each current second sample description text and the first text prompt into the current description generation model to obtain a description text of the specified image region in the first sample image as the current first predicted description text.
[0055] A first parameter adjustment module, configured to adjust the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model.
[0056] Optionally, the second sample description text determination module is specifically configured to input each current first sample image, the first sample description text corresponding to the first sample image, and a second text prompt into a pre-trained reward model to obtain a matching result between a specified image region in the first sample image and the corresponding first sample description text; wherein, the second text prompt is used to indicate whether the specified image region in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, a specified image region in the second sample image, and a description text that matches the specified image region in the second sample image; a negative sample pair includes: a third sample image, a specified image region in the third sample image, and a description text that does not match the specified image region in the third sample image; in the case where the obtained matching result indicates that the specified image region in the first sample image matches the corresponding first sample description text, the first sample description text corresponding to the first sample image is determined as the corresponding second sample description text.
[0057] Optionally, the reward model includes: a first large language model, a first adapter, a first image encoder, and a first token parser; the first image encoder is used to extract features of the input image; the first adapter is used to convert the image features output by the first image encoder into a token form; the first token parser is used to convert the input text into a token form.
[0058] Optionally, the description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the apparatus further includes:
[0059] A reward model training module, for each sample pair, inputting the sample image and the specified image region in the sample pair into the first image encoder in the reward model with an initial structure to obtain the features of the sample image and the features of the specified image region in the sample pair, and splicing the obtained features; inputting the obtained splicing result into the first adapter in the reward model with an initial structure to obtain the token form of the splicing result; inputting the second text prompt and the description text in the sample pair into the first token parser in the reward model with an initial structure to obtain the token form of the second text prompt and the token form of the description text in the sample pair; inputting the token form of the splicing result, the token form of the second text prompt, and the token form of the description text in the sample pair into the first large language model in the reward model with an initial structure to obtain a predicted matching result between the specified image region and the description text in the sample pair; adjusting the parameters of the reward model with an initial structure based on the difference between the obtained predicted matching result and the true matching result of the sample pair until a preset convergence condition is reached to obtain a trained reward model.
[0060] Optionally, the first sample description text generation module is specifically configured to, for each current first sample image, when the proportion of the specified image area in the first sample image is less than a preset threshold, expand the specified image area in the first sample image to obtain an image to be processed; input the obtained image to be processed and the first text prompt into the current description generation model to obtain a description text for describing the specified image area in the first sample image, as the current first sample description text; when the proportion of the specified image area in the first sample image is not less than the preset threshold, input the first sample image and the first text prompt into the current description generation model to obtain a description text for describing the specified image area in the first sample image, as the current first sample description text.
[0061] Optionally, the first sample description text generation module is specifically configured to use the first text prompt and the current description generation model to process the first sample image according to different sampling parameters to obtain multiple description texts for the specified image area in the first sample image, as the current first sample description text.
[0062] Optionally, the apparatus further includes:
[0063] A fourth sample image acquisition module, configured to acquire a fourth sample image after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model; among the description texts generated by the current description generation model for each fourth sample image, there is a description text that does not match the specified image area in the corresponding fourth sample image;
[0064] A description text inference module, configured to use the first text prompt and the current description generation model to process each acquired fourth sample image according to different sampling parameters to obtain multiple description texts for the specified image area in the fourth sample image;
[0065] A positive and negative sample description text determination module, configured to determine, from the obtained description texts, at least one positive sample description text that matches the specified image area in the fourth sample image, and at least one negative sample description text that does not match the specified image area in the fourth sample image;
[0066] A second predicted description text generation module, configured to input the fourth sample image and the first text prompt into the current description generation model to obtain a description text for the specified image area in the fourth sample image, as the current second predicted description text;
[0067] The second parameter adjustment module is used to adjust the parameters of the current description generation model based on the differences between the current second predicted description text and the determined positive sample description text, and between the current second predicted description text and the determined negative sample description text, to obtain a new description generation model.
[0068] Optionally, the first text prompt is further used to indicate that the current description generation model generates a description text that conforms to the first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included.
[0069] The apparatus further includes:
[0070] The third sample description text generation module is used to, after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, use the first text prompt and the current description generation model to process each current first sample image, to obtain the description text in the first specified description style of the specified image area in the first sample image as the current third sample description text.
[0071] The fourth sample description text generation module is used to adjust the description style of the current third sample description text from the first specified description style to the second specified description style to obtain the current fourth sample description text.
[0072] The fifth sample description text determination module determines, from the current fourth sample description text, the description text that matches the specified image area in the corresponding first sample image to obtain the current fifth sample description text.
[0073] The third predicted description text generation module is used to input the first sample image corresponding to each current fifth sample description text and the third text prompt into the current description generation model, to obtain the description text in the second specified description style of the specified image area in the first sample image as the current third predicted description text.
[0074] The third parameter adjustment model is used to adjust the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model.
[0075] Optionally, the apparatus further includes:
[0076] The fourth sample description text update module is used to adjust the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text. After obtaining a new description generation model, input each current first sample image and the third text prompt into the current description generation model to obtain the description text of the second specified description style in the specified image area of the first sample image, which is used as the current fourth sample description text; trigger the fifth sample description text determination module until a first preset condition is reached to obtain a new description generation model.
[0077] Optionally, the device further includes:
[0078] The trigger module is used to trigger the first sample description text generation module until a second preset condition is reached to obtain a new description generation model after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model.
[0079] Optionally, the description generation model includes: a second large language model, a second adapter, a second image encoder, and a second token parser; the second image encoder is used to extract the features of the input image; the second adapter is used to convert the image features output by the second image encoder into a token form; the second token parser is used to convert the input text into a token form.
[0080] In the fourth aspect of the embodiments of the present application, a description generation device is further provided, and the device includes:
[0081] The data acquisition module is used to acquire the image to be utilized and the text prompt to be utilized;
[0082] The description text generation module is used to input the image to be utilized and the text prompt to be utilized into a pre-trained description generation model to obtain the description text of the image area indicated by the text prompt to be utilized in the image to be utilized; wherein, the description generation model is trained based on the description generation model training method described in any of the above.
[0083] Optionally, the position of the image area indicated by the text prompt to be utilized in the image to be utilized is obtained through the following steps:
[0084] Display the image to be utilized;
[0085] In response to the user's box selection operation on the displayed image to be utilized, determine the position indicated by the box selection operation in the image to be utilized to obtain the position of the image area indicated by the text prompt to be utilized in the image to be utilized.
[0086] Optionally, the description style of the obtained description text is the description style indicated by the text prompt to be utilized; a description style is used to indicate the language of the description text and / or the number of statements included.
[0087] The embodiment of the present application also provides an electronic device, including:
[0088] A memory for storing a computer program;
[0089] A processor, when executing the program stored on the memory, implements the description generation model training method or the description generation method described in any one of the above.
[0090] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the description generation model training method or the description generation method described in any one of the above is implemented.
[0091] The embodiment of the present application also provides a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the description generation model training method or the description generation method described in any one of the above.
[0092] Advantageous effects of the embodiment of the present application:
[0093] A description generation model training method provided by the embodiment of the present application can use a first text prompt and the current description generation model to process each current first sample image, obtain a description text for describing a specified image area in the first sample image as the current first sample description text; wherein, the first text prompt is used to indicate generating a description text for a specified image area in the input image; determine, from the current first sample description texts, a description text that matches the specified image area in the corresponding first sample image as the current second sample description text; input the first sample image and the first text prompt corresponding to each current second sample description text into the current description generation model, obtain a description text for the specified image area in the first sample image as the current first predicted description text; and adjust the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model.
[0094] Based on the above processing, the description text of the specified image region in the sample image (i.e., the first sample description text) can be generated using the current latest description generation model. Due to the influence of the accuracy of the description generation model, the first sample description text inferred by the description generation model may not match the specified image region in the corresponding first sample image. Therefore, the description text that matches the specified image region in the corresponding first sample image (i.e., the second sample description text) can be further screened from the obtained first sample description text. Furthermore, each second sample description text can be used as a ground truth label, and the first sample image corresponding to each second sample description text can be used as a training sample to train the current description generation model to obtain a new description generation model. In this way, the description generation ability of the existing description generation model can be used to generate training samples, and the description generation ability of the existing description generation model can be further optimized to obtain a new description generation model. That is, the probability that the description generation model generates a description text that matches the specified image region in the input image is further increased. That is, the accuracy of the description text of the image region generated by the description generation model can be improved. Subsequently, using the trained description generation model, the description text of the image region can be accurately generated.
[0095] Of course, it is not necessary for any product or method implementing this application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0097] Figure 1 It is the first flowchart of the description generation model training method provided by the embodiment of the present application;
[0098] Figure 2 It is the second flowchart of the description generation model training method provided by the embodiment of the present application;
[0099] Figure 3 It is a flowchart of training the description generation model based on the reward model provided by the embodiment of the present application;
[0100] Figure 4a It is a schematic diagram of the image input to the reward model provided by the embodiment of the present application;
[0101] Figure 4b It is another schematic diagram of the image input to the reward model provided by the embodiment of the present application;
[0102] Figure 5 Schematic diagram of a process for training a reward model provided by an embodiment of the present application;
[0103] Figure 6 The first schematic diagram of the process for directly optimizing the preference of a description generation model provided by an embodiment of the present application;
[0104] Figure 7 The second schematic diagram of the process for directly optimizing the preference of a description generation model provided by an embodiment of the present application;
[0105] Figure 8 Schematic diagram of the process for performing one iteration on a description generation model provided by an embodiment of the present application;
[0106] Figure 9 Schematic diagram of the process for training the ability of a description generation model to generate description texts in a specified description style provided by an embodiment of the present application;
[0107] Figure 10 Schematic diagram of an image provided by an embodiment of the present application, including multiple image regions that need to generate description texts in different description styles;
[0108] Figure 11 Schematic diagram of the reuse of the ability of different description styles provided by an embodiment of the present application;
[0109] Figure 12 Schematic diagram of the process of a description generation method provided by an embodiment of the present application;
[0110] Figure 13 Schematic diagram of the process of another description generation method provided by an embodiment of the present application;
[0111] Figure 14 Schematic diagram of the structure of a description generation model training device provided by an embodiment of the present application;
[0112] Figure 15 Schematic diagram of the structure of a description generation device provided by an embodiment of the present application;
[0113] Figure 16 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0114] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0115] With the development of image processing technology, in some application scenarios (e.g., image search by text or image retrieval, etc.), it is often necessary to generate descriptive text for a specified image region in an image. The accuracy of generating the descriptive text for the image region has a significant impact on the application effect.
[0116] To accurately generate the descriptive text for the image region, an embodiment of this application provides a method for training a description generation model. Refer to Figure 1 , Figure 1 which is the first schematic flowchart of the method for training a description generation model provided by an embodiment of this application. The method for training a description generation model includes:
[0117] Step S101: Using the first text prompt and the current description generation model, process each current first sample image to obtain descriptive text for the specified image region in the first sample image, as the current first sample descriptive text.
[0118] Among them, the first text prompt is used to indicate generating descriptive text for the specified image region in the input image.
[0119] Step S102: Determine, from the current first sample descriptive texts, the descriptive text that matches the specified image region in the corresponding first sample image, as the current second sample descriptive text.
[0120] Step S103: Input the first sample image and the first text prompt corresponding to each current second sample descriptive text into the current description generation model to obtain descriptive text for the specified image region in the first sample image, as the current first predicted descriptive text.
[0121] Step S104: Based on the difference between the current first predicted descriptive text and the current second sample descriptive text, adjust the parameters of the current description generation model to obtain a new description generation model.
[0122] Based on the above processing, the description text of the specified image region in the sample image (i.e., the first sample description text) can be generated using the current latest description generation model. Due to the influence of the accuracy of the description generation model, the first sample description text inferred by the description generation model may not match the specified image region in the corresponding first sample image. Therefore, the description text that matches the specified image region in the corresponding first sample image (i.e., the second sample description text) can be further screened from the obtained first sample description text. Furthermore, each second sample description text can be used as a ground truth label, and the first sample image corresponding to each second sample description text can be used as a training sample to train the current description generation model to obtain a new description generation model. In this way, the description generation ability of the existing description generation model can be used to generate training samples, and the description generation ability of the existing description generation model can be further optimized to obtain a new description generation model. That is, to further increase the probability that the description generation model generates description text that matches the specified image region in the input image. This can also improve the accuracy of the description text of the image region generated by the description generation model. Subsequently, using the trained description generation model, the description text of the image region can be accurately generated.
[0123] Regarding step S101, the description generation model can be a multimodal large language model, which can include an existing pre-trained large language model. For example, the description generation model can include: a large language model (i.e., the second large language model in this application), an adapter (i.e., the second adapter in this application), an image encoder (i.e., the second image encoder in this application), and a token parser (i.e., the second token parser in this application). The second image encoder is used to extract the features of the input image; the second adapter is used to convert the image features output by the second image encoder into token form; the second token parser is used to convert the input text into token form. For example, the adapter in the embodiments of this application can be a neural network, and the image encoder can be a convolutional neural network. The token parser can also be called a Tokenizer. For example, the token parser can be a WordPiece Tokenizer (a character-based tokenizer), or it can also be a BPETokenizer (Byte-Pair Encoding Tokenizer, a byte-level encoding-based tokenizer).
[0124] In this way, the description generation model can also process the input image and text prompt to ensure that the description generation model can be used to process the input image in combination with the first text prompt to generate the description text of the specified image region in the input image. The description text of the specified image region in an image can be used to describe the image content of the specified image region in the image. For example, the image can be a photo of a forest, which includes trees and a yellow squirrel holding a pinecone. The specified image region in the image can be the image region where the squirrel is located. Correspondingly, the description text of the specified image region in the image can be "a yellow squirrel holding a pinecone in its hand". The positions of the specified image regions in different images can be the same, or they can also be different.
[0125] After determining the structure of the description generation model, the description generation model with the initial structure can be trained through multiple pre-obtained image-text pairs to obtain the first version of the description generation model. An image-text pair can include: a sample image and the description text of the specified image region in the sample image. For example, the image-text pair can come from a public dataset, or the sample image can be collected by an image acquisition device, and the description text of the specified image region in the sample image can be determined manually to obtain the image-text pair. The current description generation model can be the current latest version of the description generation model. When it is necessary to update the version of the description generation model, the current latest description generation model can be trained through step S101 and subsequent steps S102 - step S104 to obtain a new version of the description generation model.
[0126] The first sample image can be obtained in advance. For example, the first sample image can come from a public dataset, or the first sample image can also be collected by an image acquisition device and the specified image region in the collected image can be labeled manually. The specified image region in each first sample image can be any image region in the first sample image. For example, the specified image region in an image can be the foreground part of the image and can include at least one main object. For example, the main object can be an animal or a vehicle.
[0127] The first text prompt may instruct the description generation model to generate descriptive text for a specified image region in the input image. In one implementation, the first text prompt may include information about the location of the specified image region in the input image (e.g., the bounding box coordinates of the specified image region). For example, the first text prompt may be "Please generate descriptive text for the image region within the rectangle with the top-left vertex coordinates (a, b) and the bottom-right vertex coordinates (c, d) in the input image." In this implementation, for different images, when the location of the specified image region in the image is different, the information about the location of the specified image region included in the first text prompt is also different. That is, for different images, the first text prompt may be different.
[0128] In another implementation, the bounding box of the specified image region in the input image has been annotated, and the first text prompt may instruct the description generation model to generate descriptive text for the image region (i.e., the specified image region) within the annotated bounding box in the input image. For example, the first text prompt may be "Please generate descriptive text for the image region within the annotated bounding box in the input image."
[0129] By processing each first sample image using the first text prompt and the current latest version of the description generation model, descriptive text for the specified image region in the first sample image can be obtained as the current first sample descriptive text. The descriptive text for a specified image region in an image can be referred to as the descriptive text corresponding to that image. For each first sample image, one or more first sample descriptive texts corresponding to the first sample image can be obtained using the current latest version of the description generation model.
[0130] In one embodiment, step S101 includes: processing the first sample image according to different sampling parameters using the first text prompt and the current description generation model to obtain multiple descriptive texts for the specified image region in the first sample image as the current first sample descriptive text.
[0131] In the embodiments of the present application, during the process of processing each first sample image to obtain the first sample description text, the sampling parameters of the description generation model can be adjusted. For example, the sampling parameters can include at least one of Top-k, Top-p, and Temperature parameter. According to Top-k sampling, during the process of generating each word, the description generation model can be based on the probabilities of all possible words, and can randomly select one from the top k words with the highest probabilities as the generated word. According to Top-p sampling, during the process of generating each word, the description generation model can be based on the probabilities of all possible words, and select words in order from high to low probability until the sum of the probabilities of the selected words is not less than p, and can randomly select one from the selected words as the generated word. Correspondingly, the first text prompt and the description generation model can be used to process the first sample image according to different sampling parameters to obtain multiple description texts of the specified image area in the first sample image as the first sample description text.
[0132] Based on the above processing, for each first sample image, multiple first sample description texts corresponding to the first sample image can be obtained. In this way, the diversity of the first sample description text can be improved, and further ensure that the description text (i.e., the second sample description text) matching the specified image area in the corresponding first sample image can be determined from the obtained first sample description text, so as to obtain the sample data for optimizing the description generation ability of the description generation model. Ensure that the description generation ability of the description generation model can be further optimized (which can also be called supervised fine-tuning) to obtain a new description generation model. In this way, the accuracy of the description text of the image area generated by the description generation model can be improved. Subsequently, using the new description generation model, the description text of the image area can be accurately generated.
[0133] For step S102, the description text matching the specified image area in the corresponding first sample image can be determined from the current first sample descriptions as the current second sample description text. That a specified image area in an image matches a description text can mean that the content described by the description text conforms to the image content of the specified image area in the image. A first sample image may not have a corresponding second sample description text, or there may be one or more corresponding second sample description texts. In this way, the description text (i.e., the second sample description text) matching the specified image area in the corresponding first sample image can be screened out from the current first sample descriptions, and this screening process can be called rejection sampling. Subsequently, the current description generation model can be trained according to the screened second sample description text and the corresponding first sample image to optimize the description generation ability of the description generation model.
[0134] For example, for each first sample description, it can be determined manually according to preset preferences whether the first sample description text matches the specified image region in the corresponding first sample image. For example, the preset preferences can include dimensions such as description text accuracy, information richness, description style, etc., which can be determined according to needs and are not specifically limited. Alternatively, for each first sample description, a reward model that has been pre-trained to determine whether the specified image region in the input image matches the description text can be used to determine whether the first sample description text matches the specified image region in the corresponding first sample image. The specific manner of determining the second sample description text according to the reward model can refer to the relevant descriptions in steps S1021 - S1022 of the subsequent embodiments.
[0135] For steps S103 and S104, the first sample image and the first text prompt corresponding to each second sample description text can be input into the current description generation model. Correspondingly, the current description generation model can generate the description text of the specified image region in the input first sample image according to the indication of the first text prompt, that is, obtain the first predicted description text. Furthermore, the difference between the current first predicted description text and the current second sample description text can be calculated. For example, the similarity between the current first predicted description text and the current second sample description text can be calculated to represent the difference between the two. According to the difference between the current first predicted description text and the current second sample description text, the parameters of the current description generation model can be adjusted to obtain a new description generation model. For example, the parameters of the current description generation model can be adjusted until the first convergence condition is reached to obtain a new description generation model. For example, the first convergence condition can be that the number of times of adjusting the parameters of the description generation model reaches the first specified number, or it can also be that the difference between the current first predicted description text and the current second sample description text is less than the first preset difference.
[0136] Through the above steps S101 - S104, it is also possible to generate training samples using the description generation ability of the latest version of the description generation model, and use the obtained training samples to further optimize the description generation ability of the latest version of the description generation model to obtain a new description generation model. This process can be called one round of iteration. In this application, the number of rounds of iteration for the description generation model can be set according to needs and is not specifically limited. That is to say, in this application, the description generation model can be iterated one round or multiple rounds according to needs.
[0137] In one embodiment, after step S104, the description generation model training method further includes: returning to execute step S101 until the second preset condition is reached to obtain a new description generation model.
[0138] In the embodiments of the present application, after completing one round of iteration and updating the description generation model to obtain the latest description generation model, a new round of iteration can be continued. That is, the latest description generation model can be continuously used to process each current first sample image to generate new first sample description texts. The current first sample image can be the same as the first sample image used in the previous round of iteration, or new first sample images can be obtained. And new second sample description texts can be screened from the obtained new first sample description texts. Inputting the first sample image and the first text prompt corresponding to each new second sample description text into the new description generation model can obtain new first predicted description texts. Correspondingly, based on the difference between the new first predicted description texts and the new second sample description texts, the parameters of the new description generation model can be adjusted until the specified convergence condition is reached to obtain the latest description generation model. Furthermore, the latest description generation model can be continuously used to process each current first sample image to generate new first sample description texts. Repeat the above process until the second preset condition is reached to obtain a new description generation model.
[0139] The second preset condition can be that the number of times of adjusting the parameters of the description generation model reaches a second preset number of times, and the second preset number of times can be set according to needs without specific limitation.
[0140] Alternatively, after completing one round of iteration, the current latest description generation model can be used to process a second preset number of first sample images to obtain description texts of specified image regions in each first sample image. The second preset number can be set according to needs without specific limitation. And it can be determined whether the proportion of the description texts that match the specified image regions in the corresponding first sample images among the obtained description texts is greater than a second ratio. If the proportion of the description texts that match the specified image regions in the corresponding first sample images among the obtained description texts is greater than the second ratio, it is determined that the second preset condition is reached. The second ratio can be set according to needs without specific limitation. For each obtained description text, it can be determined whether the description text matches the specified image region in the corresponding image manually according to preset preferences. Or, a pre-trained model with the ability to determine whether a specified image region in an input image and a description text match can be used to determine whether the description text matches the specified image region in the corresponding image.
[0141] Based on the above processing, multiple rounds of iteration can be performed on the description generation model, which can also be called the self-loop iterative optimization of the description generation model. In this way, the ability of the description generation model to generate description text that matches the specified image region in the input image can be iteratively optimized, and thus the probability of the description generation model generating description text that matches the specified image region in the input image can be further increased. The accuracy of the description text of the image region generated by the description generation model can also be improved. Subsequently, using the trained description generation model, the description text of the image region can be accurately generated.
[0142] In one embodiment, step S101 includes:
[0143] Step 1: For each current first sample image, when the proportion of the specified image region in the first sample image is less than a preset threshold, expand the specified image region in the first sample image to obtain a to-be-processed image.
[0144] Step 2: Input the obtained to-be-processed image and the first text prompt into the current description generation model to obtain the description text for describing the specified image region in the first sample image, which is used as the current first sample description text.
[0145] Step 3: When the proportion of the specified image region in the first sample image is not less than the preset threshold, input the first sample image and the first text prompt into the current description generation model to obtain the description text for describing the specified image region in the first sample image, which is used as the current first sample description text.
[0146] In the embodiment of the present application, for each current first sample image, when the proportion of the specified image region in the first sample image is less than the preset threshold, it means that the specified image region in the first sample image is relatively small, and the specified image region in the first sample image can be expanded to obtain a to-be-processed image. For example, the preset threshold can be 30% or 25%. The specified image region can be expanded based on the center of the specified image region in the first sample image, and the expanded specified image region can be cropped from the first sample image to obtain the to-be-processed image. For example, based on the center of the specified image region in the first sample image, the bounding box of the specified image region can be expanded outward. For example, the length and width of the bounding box of the specified image region can be expanded by 10% or 12% of the original length and width respectively. Furthermore, the obtained to-be-processed image and the first text prompt can be input into the current description generation model to obtain the first sample description text.
[0147] When the specified image region in a first sample image is small, the image content of the specified image region in the first sample image is also less, resulting in low sensitivity of the description generation model to the features of the specified image region, which will also affect the accuracy of the feature recognition of the specified image region by the description generation model, resulting in low accuracy of the first sample description text generated by the description generation model. Through the above method, when the specified image region in a first sample image is small, the specified image region in the first sample image can be expanded, and a description text can be generated based on the expanded image. In this way, the problem that the first sample description text generated by the description generation model has low accuracy due to the small specified image region in the first sample image can be avoided. That is, the perception ability of the description generation model to the small specified image region (which can also be called a small target) in the input image can be improved. The accuracy of the first sample description text of the small specified image region generated by the description generation model can also be improved, and further the accuracy of the description text used to describe the small target in the training data used in each round of iteration process can be improved. Correspondingly, the inference performance of the description generation model for small targets can be enhanced.
[0148] For each current first sample image, when the proportion of the specified image region in the first sample image is not less than a preset threshold, it indicates that the specified image region in the first sample image is large (which can also be called a large target). The description generation model has high sensitivity to the features of the specified image region and high accuracy in feature recognition of the specified image region. In this case, the first sample image and the first text prompt can be directly input into the current description generation model to obtain a description text for describing the specified image region in the first sample image, as the current first sample description text. In this way, when the specified image region in a first sample image is large, the first sample image can be directly used as the input of the description generation model to generate the first sample description text. The context information of the specified image region can be kept complete, and the ability of the description generation model to generate description texts containing context information can be improved.
[0149] Subsequently, in the process of training the description generation model using the filtered second sample description texts and the corresponding first sample images, for both large targets and small targets, the complete first sample image is used as the input of the description generation model to increase the difficulty of the training samples, and thus the robustness of the trained description generation model can be improved.
[0150] Based on the above processing, while enhancing the perception ability of the description generation model to small targets, the description generation model can obtain more image content in the image other than the specified image region, so as to fully retain the context information of the specified image region and encourage the description generation model to generate a description containing the relationship between the subject in the specified image region of the input image and the information of the whole image.
[0151] In one embodiment, referring to Figure 2 , Figure 2 which is the second schematic flowchart of the description generation model training method provided by the embodiment of the present application. Step S102 includes:
[0152] Step S1021: Input each current first sample image, the first sample description text corresponding to the first sample image, and the second text prompt into a pre-trained reward model to obtain the matching result between the specified image region in the first sample image and the corresponding first sample description text.
[0153] Among them, the second text prompt is used to indicate whether to determine whether the specified image region in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, the specified image region in the second sample image, and the description text that matches the specified image region in the second sample image; a negative sample pair includes: a third sample image, the specified image region in the third sample image, and the description text that does not match the specified image region in the third sample image.
[0154] Step S1022: In the case where the obtained matching result indicates that the specified image region in the first sample image matches the corresponding first sample description text, determine the first sample description text corresponding to the first sample image as the corresponding second sample description text.
[0155] In the embodiment of the present application, positive sample pairs and negative sample pairs can be obtained in advance, and based on the obtained positive sample pairs and negative sample pairs, the reward model with an initial structure is trained so that the reward model learns the ability to determine whether the input image matches the input description text. The positive sample pairs and negative sample pairs can come from a public dataset, or can also be obtained by collecting images through an image acquisition device and manually annotating them. The second text prompt is used to indicate whether to determine whether the specified image region in the input image matches the input description text. In one implementation, the second text prompt may include information about the position of the specified image region in the input image (e.g., the bounding box coordinates of the specified image region). For example, the second text prompt may be "Please determine whether the image region in the rectangle with the upper left vertex coordinates (a, b) and the lower right vertex coordinates (c, d) in the input image matches the input description text". In another implementation, the bounding box of the specified image region in the input image has been annotated. For example, the second text prompt may be "Please determine whether the image region in the annotated bounding box in the input image matches the input description text".
[0156] Using the trained reward model in combination with the second text prompt, each first sample image and the first sample description text corresponding to the first sample image are processed to obtain a matching result between the specified image region in the first sample image and the corresponding first sample description text. The obtained matching result can indicate whether the specified image region in the first sample image matches the corresponding first sample description text.
[0157] The reward model may include an existing pre-trained large language model, and the type of large language model included in the reward model may be the same as that of the description generation model. For example, the reward model may include: a multimodal large language model (i.e., the first multimodal large language model in this application), an adapter (i.e., the first adapter in this application), an image encoder (i.e., the first image encoder in this application), and a token parser (i.e., the first token parser in this application). The first image encoder is used to extract the features of the input image; the first adapter is used to convert the image features output by the first image encoder into token form; the first token parser is used to convert the input text into token form. In this way, the reward model can also process the input image and text prompt to ensure that the reward model can be used to process the input image and description text in combination with the second text prompt to determine whether the specified image region in the input image matches the input description text.
[0158] For each first sample image, in the case where the obtained matching result indicates that the specified image region in the first sample image matches the corresponding first sample description text, the first sample description text corresponding to the first sample image can be determined as the corresponding second sample description text.
[0159] Based on the above processing, the reward model can be used to screen the second sample description text from the first sample description text, that is, the pre-trained reward model can be used to implement rejection sampling to improve the accuracy of the training data and obtain high-quality training data. Subsequently, the current description generation model can be trained according to the screened second sample description text and the corresponding first sample image to optimize the description generation ability of the description generation model.
[0160] In one embodiment, refer to Figure 3 , Figure 3 FIG. is a schematic flowchart of training a description generation model based on a reward model provided by an embodiment of the present application. The description generation model 31 (which may also be referred to as a region description model) may include: a large language model (i.e., the second large language model in this application), an adapter (i.e., the second adapter in this application), an image encoder (i.e., the second image encoder in this application), and a token parser (i.e., the second token parser in this application). Figure 3The multiple small trapezoids in it represent Tokens obtained through the adapter and the token parser.
[0161] For each image (i.e., the first sample image in the above embodiment), it can be determined whether the image is a small target, that is, it is determined whether the proportion of the specified image area in the first sample image is less than a preset threshold. If so, for the specified image area in the first sample image, a central cropped image can be obtained, and the obtained central cropped image and the bounding box coordinates of the specified image area in the image can be input into the previous version of the region description model (i.e., the current description generation model in the above embodiment) to generate the first sample description text, that is, steps one and two in the above embodiment. If not, the image and the bounding box coordinates of the specified image area in the image can be input into the previous version of the region description model to generate the first sample description text, that is, step three in the above embodiment.
[0162] And each first sample image, the first sample description text corresponding to the first sample image, and the bounding box coordinates of the specified image area in the first sample image can be input into the pre-trained reward model, that is, step S1021 in the above embodiment. That is, rejection sampling can be implemented through the reward model to determine the positive sample images that are not rejected by rejection sampling, that is, the first sample images corresponding to the selected second sample description texts in the above embodiment. The bounding box of the specified image area in the positive sample image can be called the positive sample bounding box, and the description text corresponding to the positive sample image (i.e., the second sample description text in the above embodiment) can be called the positive sample description.
[0163] For each current positive sample image, the positive sample image can be input into the image encoder in the current description generation model 31, and the corresponding positive sample description and the text prompt containing bounding box information (i.e., the first text prompt in the above embodiment) can be input into the token parser in the current description generation model 31 to infer the region description (i.e., the first predicted description text in the above embodiment) through the current description generation model 31. Specifically, reference can be made to the relevant description of step S103 in the above embodiment. Subsequently, based on the difference between the current first predicted description text and the current second sample description text, the parameters of the current description generation model 31 can be adjusted to obtain a new description generation model.
[0164] In one embodiment, the description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the reward model is trained through the following steps:
[0165] Step 1: For each sample pair, input the sample image and the specified image region in the sample pair into the first image encoder in the reward model with the initial structure, obtain the features of the sample image and the features of the specified image region in the sample pair, and splice the obtained features.
[0166] Step 2: Input the obtained splicing result into the first adapter in the reward model with the initial structure to obtain the tokenized form of the splicing result.
[0167] Step 3: Input the second text prompt and the description text in the sample pair into the first token parser in the reward model with the initial structure to obtain the tokenized form of the second text prompt and the tokenized form of the description text in the sample pair.
[0168] Step 4: Input the tokenized form of the splicing result, the tokenized form of the second text prompt, and the tokenized form of the description text in the sample pair into the first large language model in the reward model with the initial structure to obtain the predicted matching result between the specified image region and the description text in the sample pair.
[0169] Step 5: Based on the difference between the obtained predicted matching result and the true matching result of the sample pair, adjust the parameters of the reward model with the initial structure until the preset convergence condition is reached to obtain the trained reward model.
[0170] In the embodiment of the present application, the sample image to be utilized can be obtained in advance. The sample image to be utilized can be the same as the first sample image in the above embodiment, or it can also be different. For each sample image to be utilized, the latest description generation model can be used to process the sample image to be utilized in combination with the first text prompt to obtain the description text of the specified image region in the sample image to be utilized. Specifically, reference can be made to the relevant description of obtaining the first sample description text in the above embodiment. Furthermore, for each obtained description text, it can be determined whether the description text matches the specified image region in the corresponding sample image to be utilized manually according to the preset preference. Or, for each obtained description text, a pre-trained model with the ability to determine whether the specified image region and the description text in the input image match can be used to determine whether the description text matches the specified image region in the corresponding sample image to be utilized.
[0171] When it is determined that a description text matches a specified image region in a corresponding sample image to be utilized, the description text, the corresponding sample image to be utilized, and the specified image region in the sample image to be utilized can be used as a positive sample pair. In this case, the corresponding sample image to be utilized is the second sample image. Conversely, the description text, the corresponding sample image to be utilized, and the specified image region in the sample image to be utilized can be used as a negative sample pair. In this case, the corresponding sample image to be utilized is the third sample image.
[0172] For each obtained sample pair (including positive and negative sample pairs), the sample image in the sample pair and the specified image region in the sample image can be input into the first image encoder in the reward model with an initial structure to obtain the features of the sample image in the sample pair and the features of the specified image region in the sample image, and the obtained features can be concatenated. The obtained concatenated result is input into the first adapter in the reward model with an initial structure to obtain the concatenated result in a marked form. Also, the second text prompt and the description text in the sample pair can be input into the first token parser in the reward model with an initial structure to obtain the second text prompt in a marked form and the description text in the sample pair. The concatenated result in a marked form, the second text prompt, and the description text in the sample pair are input into the first large language model in the reward model with an initial structure. Correspondingly, the first large language model can infer the predicted matching result between the specified image region and the description text in the sample image of the sample pair.
[0173] The second text prompt can instruct the reward model to obtain an output result based on the chain of thought. That is, the first large language model in the reward model can output a conclusion on whether the specified image region in the sample image of the sample pair matches the description text, and can also output the analysis process for determining whether the two match. That is, the output result of the reward model can include: the conclusion on whether the specified image region in the input image matches the input description text, and the analysis process for determining whether the two match.
[0174] As Figure 4a shown, Figure 4a is a schematic diagram of an image input into the reward model provided by an embodiment of the present application. Figure 4a There is a little boy a and a little girl b sitting on a green swing. The little girl is smiling and wearing a pink hat. The dashed box 41 indicates the specified image region. The description text corresponding to the input into the reward model can be "a person sitting on a swing and smiling". The analysis process for the reward model to determine whether the two match can be "the object in this region is a girl wearing a pink hat. She is sitting on a swing and seems to be smiling". The conclusion for the reward model to determine whether the two match can be "yes". As Figure 4b shown, Figure 4bA schematic diagram of another input reward model image provided by an embodiment of this application. Figure 4b There is a little boy a and a little girl b sitting on a green swing. The dashed box 42 indicates the specified image area. The descriptive text corresponding to the input reward model can be "the children on the swing", and the analysis process for the reward model to determine whether they match can be "the object in this area is a green swing. There are two children sitting on the swing. The description 'the children on the swing' refers to the children sitting on the swing, not the swing itself", and the conclusion of the reward model to determine whether they match can be "no".
[0175] Furthermore, based on the difference between the obtained predicted matching result and the true matching result of this sample pair, the parameters of the reward model with the initial structure can be adjusted until the preset convergence condition is reached, and the trained reward model is obtained. If this sample pair is a positive sample pair, the true matching result of this sample pair can indicate that the specified image area in the sample image of this sample pair matches the descriptive text. If this sample pair is a negative sample pair, the true matching result of this sample pair can indicate that the specified image area in the sample image of this sample pair does not match the descriptive text. For example, the preset convergence condition can be that the number of times of adjusting the parameters of the reward model reaches a specified number, or it can also be that the difference between the obtained predicted matching result and the true matching result of this sample pair is less than a second preset difference.
[0176] Based on the above processing, sample pairs for training the reward model can be obtained using the latest description generation model. Training the reward model according to the obtained sample pairs can enable the reward model to learn the ability to determine whether the descriptive text generated by the description generation model matches the specified image area in the input image. Furthermore, using the trained reward model can accurately screen out the second sample descriptive text from the first sample descriptive text to ensure that the current description generation model can be trained according to the screened second sample descriptive text and the corresponding first sample image to optimize the description generation ability of the description generation model. Subsequently, using the trained description generation model, the descriptive text of the image area can be accurately generated.
[0177] In one embodiment, refer to Figure 5 , Figure 5 A schematic flowchart of a process for training a reward model provided by an embodiment of this application. The reward model 51 can include: a large language model (i.e., the first large language model in this application), an adapter (i.e., the first adapter in this application), an image encoder (i.e., the first image encoder in this application), and a token parser (i.e., the first token parser in this application). Figure 5 The multiple small trapezoids in it represent the tokens obtained through the adapter and the token parser.
[0178] For each sample image to be utilized (i.e., the complete image in Figure 5 ), the region description model of the previous version (i.e., the current description generation model) can be used, combined with the first text prompt, to process the sample image to be utilized, and obtain the description text of the specified image region in the sample image to be utilized. Specifically, reference can be made to the relevant description of obtaining the first sample description text in the above embodiments. Furthermore, for each obtained description text, it can be manually determined whether the description text conforms to human preferences to determine whether the description text matches the specified image region in the corresponding sample image to be utilized. If so, the description text (which can also be referred to as a positive sample description), the corresponding sample image to be utilized, and the specified image region in the sample image to be utilized (i.e., the bounding box region cropped image) can be used as a positive sample pair. If not, the description text (which can also be referred to as a negative sample description), the corresponding sample image to be utilized, and the specified image region in the sample image to be utilized can be used as a negative sample pair.
[0179] For each obtained sample pair (including positive and negative sample pairs), the sample image and the specified image region in the sample image in the sample pair can be input into the image encoder in the reward model with an initial structure, and the text prompt containing bounding box information (i.e., the second text prompt in the above embodiments) and the description text in the sample pair can be input into the token parser in the reward model with an initial structure, so as to infer the thought chain analysis process and conclusion through the reward model 51, that is, the analysis process of determining whether the two match in the form of a thought chain, and the conclusion of whether the specified image region in the input image matches the input description text. Furthermore, based on the difference between the obtained predicted matching result and the true matching result of the sample pair, the parameters of the reward model with an initial structure can be adjusted until the preset convergence condition is reached, and the trained reward model is obtained.
[0180] In one embodiment, refer to Figure 6 , Figure 6 which is the first schematic flowchart of directly optimizing the preference of the description generation model provided by the embodiments of the present application. After step S104, the description generation model training method further includes:
[0181] Step S601: Obtain a fourth sample image.
[0182] Among them, in the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image regions in the corresponding fourth sample images.
[0183] Step S602: Using the first text prompt and the current description generation model, process each obtained fourth sample image according to different sampling parameters to obtain multiple description texts for the specified image region in the fourth sample image.
[0184] Step S603: From the obtained description texts, determine at least one positive sample description text that matches the specified image region in the fourth sample image, and at least one negative sample description text that does not match the specified image region in the fourth sample image.
[0185] Step S604: Input the fourth sample image and the first text prompt into the current description generation model to obtain the description text for the specified image region in the fourth sample image as the current second predicted description text.
[0186] Step S605: Based on the difference between the current second predicted description text and the determined positive sample description texts, and the difference between the current second predicted description text and the determined negative sample description texts, adjust the parameters of the current description generation model to obtain a new description generation model.
[0187] In the embodiments of the present application, among the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image region in the corresponding fourth sample image, which means that the inference performance of the current latest description generation model for the scene to which the fourth sample image belongs is unstable. The scene to which the fourth sample image belongs can represent the scene to which the specified object included in the specified image region of the fourth sample image belongs. For example, the specified object can be a pedestrian, and the scene to which the fourth sample image belongs is the scene where the pedestrian is located, such as, an intersection.
[0188] In one implementation, during the process of the description generation model generating description texts, the scene with unstable inference performance of the description generation model can be determined through manual observation. In another implementation, the description generation model can be processed by a pre-trained reward model for the generated description texts and the corresponding images to determine whether the generated description texts match the specified image regions in the corresponding images. For each scene, it can be determined whether the inference performance of the description generation model for the scene is stable according to the number of occurrences of the situation where the generated description texts do not match the specified image regions in the corresponding images belonging to the scene. For example, in the case where the proportion of the description texts that do not match the specified image regions in the corresponding fourth sample images among the description texts generated by the description generation model for each fourth sample image is greater than a preset ratio, it can be determined that the inference performance of the description generation model for the scene to which the fourth sample image belongs is unstable.
[0189] After determining a scenario where the inference performance of the description generation model is unstable, sample images belonging to this scenario (i.e., the fourth sample images) can be obtained. Using the first text prompt and the current description generation model, each obtained fourth sample image is processed according to different sampling parameters, and multiple description texts for the specified image region in the fourth sample image can be obtained. Furthermore, at least one positive sample description text that matches the specified image region in the fourth sample image and at least one negative sample description text that does not match the specified image region in the fourth sample image can be determined from the obtained description texts. In one implementation, the positive sample description text and the negative sample description text can be selected from the obtained description texts manually according to preferences. In another implementation, a pre-trained reward model can be used to determine the matching result between the obtained description text and the specified image region in the fourth sample image. And the positive sample description text and the negative sample description text are selected from the obtained description texts according to the obtained matching result. When the matching result obtained by the reward model indicates that the obtained description text matches the specified image region in the fourth sample image, the positive sample description text and the negative sample description text can be selected from the obtained description texts manually according to preferences. For example, the preference can indicate a preference for description texts with richer description features. Then, description texts with relatively richer description features can be selected from the obtained description texts as the positive sample description text, and description texts with relatively less rich description features can be selected from the obtained description texts as the negative sample description text.
[0190] Input the fourth sample image and the first text prompt into the current description generation model to obtain the description text for the specified image region in the fourth sample image as the current second predicted description text. Furthermore, based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text, the parameters of the current description generation model can be adjusted to obtain a new description generation model. For example, the parameters of the current description generation model can be adjusted until the second convergence condition is reached to obtain a new description generation model. For example, the second convergence condition can be that the number of times of adjusting the parameters of the description generation model reaches the second specified number of times, or it can also be that the difference between the current second predicted description text and the determined positive sample description text is less than the first difference, and the difference between the current second predicted description text and the determined negative sample description text is greater than the second difference.
[0191] Based on the above processing, for scenarios where the inference performance of the latest description generation model is unstable, direct preference optimization can be performed to specifically optimize the description generation ability of the description generation model for the description text of the specified image region in the images belonging to such scenarios. In this way, the preference of the description generation model for the positive sample description text can be increased, and the preference for the negative sample description text can be reduced. Furthermore, the probability that the description generation model generates a description text that matches the specified image region in the input image can be further increased. Also, the accuracy of the description text of the image region generated by the description generation model can be improved. Subsequently, using the trained description generation model, the description text of the image region can be accurately generated.
[0192] In one embodiment, referring to Figure 7 , Figure 7 is the second flow diagram for directly optimizing the preference of the description generation model provided by the embodiments of the present application. The description generation model 71 may include: a large language model (i.e., the second large language model in the present application), an adapter (i.e., the second adapter in the present application), an image encoder (i.e., the second image encoder in the present application), and a token parser (i.e., the second token parser in the present application). Figure 5 The multiple small trapezoids in
[0193] represent tokens obtained through the adapter and the token parser. Figure 7 For each fourth sample image (i.e., the complete image in
[0194] Input the fourth sample image and the text prompt containing bounding box information (i.e., the first text prompt in the above embodiment) into the current description generation model 71, and the description text of the specified image region in the fourth sample image can be obtained as the current second predicted description text. Furthermore, based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text, the parameters of the current description generation model can be adjusted to obtain a new description generation model. This can improve the preference of the description generation model for positive sample description texts and reduce the preference for negative sample description texts, such that the preference for positive sample description texts is much greater than that for negative sample description texts.
[0195] In one embodiment, refer to Figure 8 , Figure 8 which is a schematic flowchart of one iteration of the description generation model provided by the embodiments of the present application. For each first sample image, the current latest description generation model (which can also be referred to as an image region description model) can be used to perform K inferences under different sampling parameters to obtain K sets of results (i.e., Figure 8 the same data list K sets of results in
[0196] ). That is, K first sample description texts of the specified image region in the first sample image are obtained. First, rule-based data cleaning can be performed on the obtained first sample description texts. For example, description texts containing error symbols and / or having too long a length are removed. Then, rejection sampling is performed using a pre-trained reward model to obtain supervised fine-tuning training data (i.e., second sample description texts). And the latest description generation model can be used to inject human preferences, that is, positive sample pairs and negative sample pairs for training the reward model are determined manually according to preferences. The reward model can be trained with supervised fine-tuning according to the obtained sample pairs.
[0197] In one embodiment, the ability of the description generation model to generate description texts in different description styles can also be trained. Refer to Figure 9 , Figure 9 which is a schematic flowchart of training the ability of the description generation model provided by the embodiments of the present application to generate description texts in a specified description style. The first text prompt is also used to indicate that the current description generation model generates description texts conforming to the first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included.
[0198] After step S104, the method for training the description generation model further includes:
[0199] Step S901: Using the first text prompt and the current description generation model, process each current first sample image to obtain a description text in the first specified description style for the specified image region in the first sample image, as the current third sample description text.
[0200] Step S902: Adjust the description style of the current third sample description text from the first specified description style to the second specified description style to obtain the current fourth sample description text.
[0201] Step S903: Determine, from the current fourth sample description text, the description text that matches the specified image region in the corresponding first sample image to obtain the current fifth sample description text.
[0202] Step S904: Input the first sample image corresponding to each current fifth sample description text and the third text prompt into the current description generation model to obtain a description text in the second specified description style for the specified image region in the first sample image, as the current third predicted description text.
[0203] Step S905: Based on the difference between the current third predicted description text and the current fifth sample description text, adjust the parameters of the current description generation model to obtain a new description generation model.
[0204] In the embodiments of the present application, the ability of the description generation model to generate description texts in different description styles can be trained. A description style can be used to indicate the language of the description text and / or the number of statements included. According to the number of statements included, it can be determined whether the description text is a long description or a short description. For example, the number of statements in a long description is not less than a preset statement number threshold. The number of statements in a short description is not greater than the preset statement number threshold. For example, the number of statements can be the number of full stops in the description text. The preset statement number threshold can be set as needed and is not specifically limited. For example, the preset statement number threshold can be 1. In this case, the number of statements in a long description can be greater than or equal to 1, and the number of statements in a short description is 1. In this way, it can be avoided that for an image region with relatively simple image content, setting the number of statements in a long description to be greater than 1 results in an error in the generated description text, and thus the accuracy of the description text generated by the description generation model can be further ensured.
[0205] For example, the description style can include Chinese long description, Chinese short description, English long description, English short description. See Figure 10 , Figure 10A schematic diagram of an image including multiple image regions for which description texts of different description styles need to be generated is provided in an embodiment of the present application. The description texts of different styles corresponding to each bounding box can be displayed near the bounding box. Specifically, there is a gray stone ( Figure 10 The corresponding Chinese short description can be "a small gray stone", the English short description can be "A gray rock on the ground", the English long description can be "A gray rock is visible, which is part of a larger pile of rocks. The rock's surface is rough and uneven", and the Chinese long description can be "a gray stone with a rough surface and irregular shape".
[0206] The image region in the second bounding box 12 contains a person riding a motorcycle wearing a white, blue, and orange jacket ( Figure 10 The corresponding Chinese short description could be "A person wearing a white, blue, and orange jacketriding a motorcycle", and the English short description could be "A person wearing a white, blue, and orange jacketriding a motorcycle", and the English long description could be "A person riding a motorcycle is visible. The rider is wearing a white and blue jacket, black pants, and a redand white helmet. The motorcycle is red and black. The rider is traveling ona dirt road surrounded by greenery and hills. The rider is moving away from the camera, and there are no other vehicles or people in sight", and the Chinese long description could be "A person riding a motorcycle, wearing a blue and white jacket, black pants, and a red and white helmet. The rider's posture shows that he is in control of the motorcycle, and the rear wheel of the motorcycle kicks up a cloud of dust".
[0207] There is a motorcycle exhaust pipe within the image region of the third bounding box 13. The corresponding short Chinese description can be "Motorcycle exhaust pipe", the short English description can be "A motorcycle exhaust pipe on a motorcycle", the long English description can be "A motorcycle exhaust pipe is visible, which is a component of the motorcycle's engine system. The exhaust pipe is cylindrical and appears to be made of metal, with a reflective surface. It is attached to the engine and extends towards the rear of the motorcycle", and the long Chinese description can be "The motorcycle exhaust pipe is part of the motorcycle exhaust system, which allows exhaust gas to be discharged from the engine. The exhaust pipe is cylindrical, made of metal on the surface, and there is a silencer at the end to reduce noise".
[0208] During one iteration of the description generation model, the first text prompt can be used to instruct the description generation model to generate descriptive text that conforms to the first specified description style. The first specified description style can be set as needed and is not specifically limited. By using the first text prompt and the current description generation model to process each current first sample image, the descriptive text of the specified image region in the first sample image can be obtained as the current third sample description text. The new description style that the current description generation model needs to learn can be called the second specified description style.
[0209] The description style of the current third sample description text can be adjusted from the first specified description style to the second specified description style to obtain the current fourth sample description text. For example, if the first specified description style is the short Chinese description and the second specified description style is the short English description, then the current third sample description text can be translated to obtain the current fourth sample description text. Or, if the first specified description style is the short Chinese description and the second specified description style is the long Chinese description, then the current third sample description text can be continued to write to obtain the current fourth sample description text. Both translation and continued writing can use existing translation tools or continued writing tools. For example, it can be implemented using a large language model with translation or continued writing functions.
[0210] From the current fourth sample description text, determine the description text that matches the specified image region in the corresponding first sample image to obtain the current fifth sample description text. This process can refer to the relevant description of screening the second sample description text from the first sample description text in the above embodiment. The third text prompt can instruct the description generation model to generate the description text of the second specified description style for the specified image region in the input image. For example, the third text prompt can be "Please generate the description text for the image region within the labeled bounding box in the input image, and the description style of the generated description text is a short English description."
[0211] Input the first sample image corresponding to each current fifth sample description text and the third text prompt into the current description generation model, and the description text of the second specified description style for the specified image region in the first sample image can be obtained as the current third predicted description text. Based on the difference between the current third predicted description text and the current fifth sample description text, adjust the parameters of the current description generation model to obtain a new description generation model. Specifically, it can refer to the relevant description of step S104 in the above embodiment.
[0212] Based on the above processing, the description generation model can be enabled to learn the ability to generate description texts of different description styles. Subsequently, the user can select the required description style from the description styles that the description generation model can provide and indicate the selected description style in the text prompt, so that the description generation model can generate the description text of the description style selected by the user. And in the process of training the description generation model's ability to generate description texts of a new description style, the description generation model that has been trained to have the ability to generate description texts of other description styles can be used to generate training samples. In this way, the ability of the trained description generation model to generate description texts of other description styles can be utilized to obtain high-quality training data for the new description style. That is, the optimization results of the description generation model for other description styles can be reused to obtain high-quality training data. Thus, the efficient optimization of the ability to generate description texts of different description styles can be achieved.
[0213] In one embodiment, after step S905, the description generation model training method further includes: inputting each current first sample image and the third text prompt into the current description generation model, obtaining the description text of the second specified description style for the specified image region in the first sample image as the current fourth sample description text; returning to execute step S903 until a first preset condition is reached to obtain a new description generation model.
[0214] In the embodiment of the present application, through the above steps S901 - S905, one round of iteration of the ability of the description generation model to generate description texts in the second specified description style can be achieved. Multiple rounds of iteration of the ability of the description generation model to generate description texts in the second specified description style can be performed. After step S905, each current first sample image and the third text prompt can be input into the currently latest description generation model to obtain the description text in the second specified description style of the specified image region in the first sample image, as the new fourth sample description text. Furthermore, new fifth sample description texts can be screened from the new fourth sample description texts. And the description generation model can be trained using the new fifth sample description texts, the corresponding first sample images, and the third text prompt. Thus, a new round of iteration of the ability of the description generation model to generate description texts in the second specified description style can be achieved.
[0215] The first preset condition can be that the number of times of adjusting the parameters of the description generation model reaches the first preset number of times, and the first preset number of times can be set as needed without specific limitation.
[0216] Alternatively, after one round of iteration is completed, the currently latest description generation model can be used to process the first preset number of first sample images to obtain the description texts of the specified image regions in each first sample image. The first preset number can be set as needed without specific limitation, and the first preset number can be the same as or different from the second preset number. And it can be determined whether the proportion of the description texts that match the specified image regions in the corresponding first sample images among the obtained description texts is greater than the first ratio. If the proportion of the description texts that match the specified image regions in the corresponding first sample images among the obtained description texts is greater than the first ratio, it is determined that the first preset condition is met. The first ratio can be set as needed without specific limitation, and the first ratio can be the same as or different from the second ratio. For each obtained description text, it can be determined whether the description text matches the specified image region in the corresponding image manually according to the preset preference. Or, a pre-trained model with the ability to determine whether the specified image region and the description text in the input image match can be used to determine whether the description text matches the specified image region in the corresponding image.
[0217] Based on the above processing, for each description style, multiple rounds of iteration of the description generation model can be performed. In this way, the accuracy of the description generation model in generating description texts in the specified description style can be further improved. Subsequently, using the trained description generation model, the description text of the image region can be accurately generated.
[0218] In one embodiment, refer to Figure 11 , Figure 11A schematic diagram of the ability reuse with different description styles provided by an embodiment of the present application. For each description style, one or more rounds of iteration can be performed on the description generation model, that is, in-task loop optimization is carried out. In the process of training the ability of the description generation model to generate description texts in a new description style, a trained description generation model with the ability to generate description texts in other description styles can be used to generate training samples. For example, a trained description generation model with the ability to generate description texts in the short description style of language A can be used to process the first sample image to generate description texts in the short description style of language A. And the description texts in the short description style of language B can be translated to obtain the training data in the short description style of language B. Or, the description texts in the long description style of language A can also be continued to obtain the training data in the long description style of language A. A trained description generation model with the ability to generate description texts in the long description style of language A can also be used to process the first sample image to generate description texts in the long description style of language A, and the description texts in the long description style of language B can be translated to obtain the training data in the long description style of language B. In this way, the ability of the trained description generation model to generate description texts in other description styles can be used to obtain high-quality training data in a new description style, realizing the mapping between the abilities to generate different description styles.
[0219] Based on the same inventive concept, an embodiment of the present application further provides a description generation method. Refer to Figure 12 , Figure 12 A flowchart of a description generation method provided by an embodiment of the present application. The method includes:
[0220] Step S1201: Obtain the image to be utilized and the text prompt to be utilized.
[0221] Step S1202: Input the image to be utilized and the text prompt to be utilized into a pre-trained description generation model to obtain the description text of the image region indicated by the text prompt to be utilized in the image to be utilized.
[0222] Wherein, the description generation model is trained based on the description generation model training method described in any one of the above.
[0223] Based on the above processing, during the training of the description generation model, the description generation ability of the description generation model obtained from the previous training can be utilized to generate training samples, and further optimize the description generation ability of the existing description generation model to obtain a new description generation model. That is, further increase the probability that the description generation model generates a description text that matches the specified image region in the input image. Thus, the accuracy of the description text of the image region generated by the description generation model can be improved. Correspondingly, using the trained description generation model, the description text of the image region can be accurately generated.
[0224] In one embodiment, the position of the image region indicated by the text prompt to be utilized in the image to be utilized is obtained through the following steps:
[0225] Display the image to be utilized;
[0226] In response to the user's box selection operation on the displayed image to be utilized, determine the position indicated by the box selection operation in the image to be utilized, and obtain the position of the image region indicated by the text prompt to be utilized in the image to be utilized.
[0227] In the embodiments of the present application, the image to be utilized can be displayed so that the user can browse the image to be utilized, and through the box selection operation, the user can box select the image region for which a description text needs to be generated in the displayed image to be utilized. Correspondingly, in response to the user's box selection operation on the displayed image to be utilized, the position indicated by the box selection operation in the image to be utilized can be determined, and the position of the image region indicated by the text prompt to be utilized in the image to be utilized can be obtained. In this way, it can be ensured that the description generation model can determine the position of the image region for which a description text needs to be generated in the image to be utilized, so as to accurately generate the description text of the image region indicated by the text prompt to be utilized in the image to be utilized.
[0228] In one embodiment, the description style of the obtained description text is the description style indicated by the text prompt to be utilized; a description style is used to indicate the language of the description text and / or the number of statements included.
[0229] In the embodiments of the present application, the user can select the description style indicated by the text prompt to be utilized according to needs. Correspondingly, the text prompt to be utilized can indicate that the description generation model generates a description text that conforms to the description style selected by the user. In this way, it can be ensured that the description generation model can generate a description text that meets the user's requirements.
[0230] In one embodiment, refer to Figure 13 , Figure 13It is a schematic flowchart of another description generation method provided by an embodiment of this application. The trained description generation model 131 may include: a large language model (i.e., the second large language model in this application), an adapter (i.e., the second adapter in this application), an image encoder (i.e., the second image encoder in this application), and a token parser (i.e., the second token parser in this application). Figure 3 The multiple small trapezoids in Figure 3 represent tokens obtained through the adapter and the token parser.
[0231] An image (i.e., the image to be utilized in the above embodiment) can be input into the image encoder in the current description generation model 31, and a text prompt containing bounding box information (i.e., the text prompt to be utilized in the above embodiment) can be input into the token parser in the description generation model 131, so as to infer, through the description generation model 131, a description text (which can also be referred to as a region description) of the image region indicated by the text prompt to be utilized in the image to be utilized. The text prompt to be utilized may contain bounding box coordinates. The user can perform a box selection operation on the displayed image to be utilized. The position indicated by the box selection operation is the region of the bounding box drawn by the user. According to the region of the bounding box drawn by the user, the bounding box coordinates can be parsed. Correspondingly, the parsed bounding box coordinates can be added to the text prompt to be utilized. And the user can select the description style (i.e., the specified description style) indicated by the text prompt to be utilized according to needs. Correspondingly, the text prompt to be utilized may also contain the specified description style, so as to indicate the description generation model to generate a description text that conforms to the specified description style.
[0232] It can be seen that the description generation method in this application can generate description texts based on an end-to-end multi-modal large model, and the description generation model in this application supports the function of generating descriptions of any region of an image in multiple languages and styles. Moreover, the user can select an image region and a description style according to needs, and the interactivity is strong.
[0233] Based on the same inventive concept, an embodiment of this application also provides a description generation model training device. Refer to Figure 14 , Figure 14 It is a schematic structural diagram of a description generation model training device provided by an embodiment of this application. The device includes:
[0234] A first sample description text generation module 1401, configured to use a first text prompt and the current description generation model to process each current first sample image, and obtain a description text for describing a specified image region in the first sample image as the current first sample description text; wherein, the first text prompt is used to indicate generating a description text of a specified image region in the input image.
[0235] The second sample description text determination module 1402 is configured to determine, from the current first sample description texts, the description text that matches the specified image region in the corresponding first sample image as the current second sample description text;
[0236] The first predicted description text generation module 1403 is configured to input the first sample image corresponding to each current second sample description text and the first text prompt into the current description generation model, and obtain the description text of the specified image region in the first sample image as the current first predicted description text;
[0237] The first parameter adjustment module 1404 is configured to adjust the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text, and obtain a new description generation model.
[0238] Based on the description generation model training device provided in the embodiments of the present application, the description text (i.e., the first sample description text) of the specified image region in the sample image can be generated by using the current latest description generation model. Due to the influence of the accuracy of the description generation model, the first sample description text inferred by the description generation model may not match the specified image region in the corresponding first sample image. Therefore, the description text that matches the specified image region in the corresponding first sample image (i.e., the second sample description text) can be further screened from the obtained first sample description texts. Furthermore, each second sample description text can be used as a true value label, and the first sample image corresponding to each second sample description text can be used as a training sample to train the current description generation model to obtain a new description generation model. In this way, the description generation ability of the existing description generation model can be used to generate training samples, and the description generation ability of the existing description generation model can be further optimized to obtain a new description generation model. That is, the probability that the description generation model generates a description text that matches the specified image region in the input image is further increased. That is, the accuracy of the description text of the image region generated by the description generation model can be improved. Subsequently, by using the trained description generation model, the description text of the image region can be accurately generated.
[0239] In one embodiment, the second sample description text determination module is specifically configured to input each current first sample image, the first sample description text corresponding to the first sample image, and a second text prompt into a pre-trained reward model to obtain a matching result between the specified image region in the first sample image and the corresponding first sample description text; wherein, the second text prompt is used to indicate whether the specified image region in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, the specified image region in the second sample image, and a description text that matches the specified image region in the second sample image; a negative sample pair includes: a third sample image, the specified image region in the third sample image, and a description text that does not match the specified image region in the third sample image; when the obtained matching result indicates that the specified image region in the first sample image matches the corresponding first sample description text, the first sample description text corresponding to the first sample image is determined as the corresponding second sample description text.
[0240] In one embodiment, the reward model includes: a first large language model, a first adapter, a first image encoder, and a first token parser; the first image encoder is used to extract the features of the input image; the first adapter is used to convert the image features output by the first image encoder into a token form; the first token parser is used to convert the input text into a token form.
[0241] In one embodiment, the description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the device further includes:
[0242] The reward model training module is used to, for each sample pair, input the sample image and the specified image region in the sample pair into the first image encoder in the reward model with the initial structure, obtain the features of the sample image and the features of the specified image region in the sample pair, and splice the obtained features; input the obtained splicing result into the first adapter in the reward model with the initial structure to obtain the tokenized form of the splicing result; input the second text prompt and the description text in the sample pair into the first token parser in the reward model with the initial structure to obtain the tokenized form of the second text prompt and the tokenized form of the description text in the sample pair; input the tokenized form of the splicing result, the tokenized form of the second text prompt, and the tokenized form of the description text in the sample pair into the first large language model in the reward model with the initial structure to obtain the predicted matching result between the specified image region and the description text in the sample pair; and adjust the parameters of the reward model with the initial structure based on the difference between the obtained predicted matching result and the true matching result of the sample pair until the preset convergence condition is reached, so as to obtain the trained reward model.
[0243] In one embodiment, the first sample description text generation module 1401 is specifically configured to, for each current first sample image, in the case where the proportion of the specified image region in the first sample image is less than the preset threshold, expand the specified image region in the first sample image to obtain an image to be processed; input the obtained image to be processed and the first text prompt into the current description generation model to obtain the description text for describing the specified image region in the first sample image, as the current first sample description text; in the case where the proportion of the specified image region in the first sample image is not less than the preset threshold, input the first sample image and the first text prompt into the current description generation model to obtain the description text for describing the specified image region in the first sample image, as the current first sample description text.
[0244] In one embodiment, the first sample description text generation module 1401 is specifically configured to use the first text prompt and the current description generation model to process the first sample image according to different sampling parameters to obtain multiple description texts for the specified image region in the first sample image, as the current first sample description text.
[0245] In one embodiment, the device further includes:
[0246] The fourth sample image acquisition module is used to obtain a fourth sample image after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model; among the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image region in the corresponding fourth sample image.
[0247] The description text inference module is used to process each obtained fourth sample image according to different sampling parameters by using the first text prompt and the current description generation model to obtain multiple description texts for the specified image region in the fourth sample image.
[0248] The positive and negative sample description text determination module is used to determine at least one positive sample description text that matches the specified image region in the fourth sample image and at least one negative sample description text that does not match the specified image region in the fourth sample image from the obtained description texts.
[0249] The second predicted description text generation module is used to input the fourth sample image and the first text prompt into the current description generation model to obtain the description text for the specified image region in the fourth sample image as the current second predicted description text.
[0250] The second parameter adjustment module is used to adjust the parameters of the current description generation model based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text to obtain a new description generation model.
[0251] In one embodiment, the first text prompt is further used to indicate that the current description generation model generates description texts conforming to the first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included.
[0252] The device further includes:
[0253] The third sample description text generation module is used to, after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, use the first text prompt and the current description generation model to process each current first sample image to obtain the description text in the first specified description style for the specified image region in the first sample image as the current third sample description text.
[0254] The fourth sample description text generation module is used to adjust the description style of the current third sample description text from the first specified description style to the second specified description style to obtain the current fourth sample description text;
[0255] The fifth sample description text determination module determines, from the current fourth sample description text, the description text that matches the specified image region in the corresponding first sample image to obtain the current fifth sample description text;
[0256] The third predicted description text generation module is used to input the first sample image and the third text prompt corresponding to each current fifth sample description text into the current description generation model to obtain the description text in the second specified description style of the specified image region in the first sample image as the current third predicted description text;
[0257] The third parameter adjustment model is used to adjust the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model.
[0258] In one embodiment, the device further includes:
[0259] The fourth sample description text update module is used to, after adjusting the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model, input each current first sample image and the third text prompt into the current description generation model to obtain the description text in the second specified description style of the specified image region in the first sample image as the current fourth sample description text; trigger the fifth sample description text determination module until a first preset condition is met to obtain a new description generation model.
[0260] In one embodiment, the device further includes:
[0261] The triggering module triggers the first sample description text generation module 1401 after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model until a second preset condition is met to obtain a new description generation model.
[0262] In one embodiment, the description generation model includes: a second large language model, a second adapter, a second image encoder, and a second token parser; the second image encoder is used to extract the features of the input image; the second adapter is used to convert the image features output by the second image encoder into a token form; the second token parser is used to convert the input text into a token form.
[0263] Based on the same inventive concept, an embodiment of the present application further provides a description generation device, see Figure 15 , Figure 15 which is a schematic structural diagram of a description generation device provided by an embodiment of the present application. The device includes:
[0264] A data acquisition module 1501, configured to acquire an image to be utilized and a text prompt to be utilized;
[0265] A description text generation module 1502, configured to input the image to be utilized and the text prompt to be utilized into a pre-trained description generation model, and obtain a description text of an image area indicated by the text prompt to be utilized in the image to be utilized; wherein, the description generation model is trained by using the description generation model training method according to any one of the above embodiments.
[0266] Based on the description generation device provided by the embodiment of the present application, during the training process of the description generation model, the description generation ability of the description generation model obtained in the previous training can be utilized to generate training samples, and further optimize the description generation ability of the existing description generation model to obtain a new description generation model. That is, further improve the probability that the description generation model generates a description text matching a specified image area in the input image. Thus, the accuracy of the description text of the image area generated by the description generation model can be improved. Correspondingly, by using the trained description generation model, the description text of the image area can be accurately generated.
[0267] In one embodiment, the position of the image area indicated by the text prompt to be utilized in the image to be utilized is obtained through the following steps:
[0268] Display the image to be utilized;
[0269] In response to a user's box selection operation on the displayed image to be utilized, determine the position indicated by the box selection operation in the image to be utilized, and obtain the position of the image area indicated by the text prompt to be utilized in the image to be utilized.
[0270] In one embodiment, the description style of the obtained description text is the description style indicated by the text prompt to be utilized; a description style is used to indicate the language of the description text and / or the number of statements included.
[0271] An embodiment of the present application further provides an electronic device, as Figure 16 shown, including:
[0272] A memory 1601, configured to store a computer program;
[0273] The processor 1602, when executing the program stored in the memory 1601, implements any of the above-described generation model training methods or any of the above-described generation methods.
[0274] And the above electronic device may further include a communication bus and / or a communication interface, and the processor 1602, the communication interface, and the memory 1601 complete communication with each other through the communication bus.
[0275] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0276] The communication interface is used for communication between the above electronic device and other devices.
[0277] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0278] The above processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0279] In another embodiment provided by the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-described generation model training methods or any of the above-described generation methods are implemented.
[0280] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the described generation model training methods or any one of the described generation methods in the above embodiments.
[0281] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0282] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device including the element.
[0283] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the description generation method, device, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0284] The foregoing is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A description generation model training method, characterized in that: The method comprises: Using the first text prompt and the current description generation model, each current first sample image is processed to obtain a description text for describing a specified image area in the first sample image as the current first sample description text; wherein the first text prompt is used to instruct the generation of a description text for the specified image area in the input image; Determine, from each of the current first sample description texts, a description text that matches the designated image area in the corresponding first sample image as the current second sample description text; Inputting the first sample image and the first text prompt corresponding to each current second sample description text into the current description generation model to obtain the description text of the specified image area in the first sample image as the current first predicted description text; Based on the difference between the current first predicted description text and the current second sample description text, adjusting the parameters of the current description generation model to obtain a new description generation model; The first text prompt is also used to indicate that: the current description generation model generates a description text that conforms to a first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included; After adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, the method further includes: Using the first text prompt and the current description generation model, each current first sample image is processed to obtain a description text of the first specified description style for a specified image area in the first sample image as the current third sample description text; Adjusting the description style of the current third sample description text from the first specified description style to the second specified description style to obtain a current fourth sample description text; Determine, from the current fourth sample description text, a description text that matches the designated image area in the corresponding first sample image, to obtain the current fifth sample description text; Inputting the first sample image and the third text prompt corresponding to each fifth sample description text into the current description generation model, obtaining the description text of the second specified description style of the specified image area in the first sample image as the current third predicted description text; Based on the difference between the current third predicted description text and the current fifth sample description text, the parameters of the current description generation model are adjusted to obtain a new description generation model.
2. The method according to claim 1, characterized in that The step of determining, from the current first sample description texts, a description text that matches a designated image area in the corresponding first sample image as the current second sample description text comprises: Input each current first sample image, the first sample description text corresponding to the first sample image, and the second text prompt into a pre-trained reward model to obtain a matching result between a specified image area in the first sample image and the corresponding first sample description text; wherein the second text prompt is used to indicate whether the specified image area in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, a specified image area in the second sample image, and a description text that matches the specified image area in the second sample image; a negative sample pair includes: a third sample image, a specified image area in the third sample image, and a description text that does not match the specified image area in the third sample image; When the obtained matching result indicates that the designated image area in the first sample image matches the corresponding first sample description text, the first sample description text corresponding to the first sample image is determined as the corresponding second sample description text.
3. The method according to claim 2, characterized in that The reward model includes: a first large language model, a first adapter, a first image encoder and a first tag parser; the first image encoder is used to extract features of an input image; the first adapter is used to convert the image features output by the first image encoder into a tag form; the first tag parser is used to convert the input text into a tag form.
4. The method according to claim 3, characterized in that The description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the reward model is trained by the following steps: For each sample pair, the sample image and the specified image region in the sample pair are input into the first image encoder in the reward model of the initial structure to obtain the features of the sample image and the features of the specified image region in the sample pair, and the obtained features are spliced; Input the obtained splicing result into the first adapter in the reward model of the initial structure to obtain a marked form of the splicing result; Inputting the second text prompt and the description text in the sample pair into the first token parser in the reward model of the initial structure to obtain the token form of the second text prompt and the token form of the description text in the sample pair; Inputting the marked form of the splicing result, the marked form of the second text prompt, and the marked form of the description text in the sample pair into the first language model in the reward model of the initial structure, and obtaining a predicted matching result between the specified image area in the sample pair and the description text; Based on the difference between the predicted matching result and the actual matching result of the sample pair, the parameters of the reward model of the initial structure are adjusted until the preset convergence condition is reached to obtain a trained reward model.
5. The method according to claim 1, characterized in that The method of using the first text prompt and the current description generation model to process each current first sample image to obtain a description text for describing a specified image area in the first sample image as the current first sample description text includes: For each current first sample image, when the proportion of the designated image area in the first sample image is less than a preset threshold, the designated image area in the first sample image is expanded to obtain an image to be processed; Inputting the obtained image to be processed and the first text prompt into the current description generation model to obtain a description text for describing the specified image area in the first sample image as the current first sample description text; When the proportion of the specified image area in the first sample image is not less than the preset threshold, the first sample image and the first text prompt are input into the current description generation model to obtain a description text for describing the specified image area in the first sample image as the current first sample description text.
6. The method according to claim 1, characterized in that The method of using the first text prompt and the current description generation model to process each current first sample image to obtain a description text for describing a specified image area in the first sample image as the current first sample description text includes: The first sample image is processed according to different sampling parameters using the first text prompt and the current description generation model to obtain a plurality of description texts of a designated image area in the first sample image as the current first sample description text.
7. The method according to claim 1, characterized in that After adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, the method further includes: Acquire a fourth sample image; wherein, among the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image area in the corresponding fourth sample image; Using the first text prompt and the current description generation model, each fourth sample image obtained is processed according to different sampling parameters to obtain a plurality of description texts for a specified image area in the fourth sample image; Determine, from the obtained description texts, at least one positive sample description text that matches the specified image region in the fourth sample image, and at least one negative sample description text that does not match the specified image region in the fourth sample image; Inputting the fourth sample image and the first text prompt into a current description generation model to obtain a description text of a specified image area in the fourth sample image as a current second predicted description text; Based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text, the parameters of the current description generation model are adjusted to obtain a new description generation model.
8. The method according to claim 1, characterized in that After adjusting the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model, the method further includes: Input each current first sample image and the third text prompt into the current description generation model to obtain a description text of the second specified description style for a specified image area in the first sample image as the current fourth sample description text; Return to the step of determining, from the current fourth sample description text, a description text that matches the designated image area in the corresponding first sample image to obtain the current fifth sample description text, until the first preset condition is met to obtain a new description generation model.
9. The method according to any one of claims 1 to 8, characterized in that: The description generation model includes: a second language model, a second adapter, a second image encoder and a second token parser; the second image encoder is used to extract features of an input image; the second adapter is used to convert image features output by the second image encoder into a token form; the second token parser is used to convert input text into a token form.
10. A description generation method, characterized in that: The method comprises: Obtaining images to be used and text prompts to be used; The image to be used and the text prompt to be used are input into a pre-trained description generation model to obtain a description text of the image area indicated by the text prompt to be used in the image to be used; wherein the description generation model is trained based on any one of the methods described in claims 1-9.
11. The method according to claim 10, characterized in that The position of the image area indicated by the to-be-used text prompt in the to-be-used image is obtained by the following steps: Displaying the image to be used; In response to a user's frame selection operation on the displayed image to be used, a position indicated by the frame selection operation in the image to be used is determined, and a position of an image area indicated by the text prompt to be used in the image to be used is obtained.
12. A description generation model training device, characterized in that: The device comprises: A first sample description text generation module is used to process each current first sample image using the first text prompt and the current description generation model to obtain a description text for describing a specified image area in the first sample image as the current first sample description text; wherein the first text prompt is used to indicate the generation of a description text for the specified image area in the input image; A second sample description text determination module is used to determine, from the current first sample description texts, a description text that matches the designated image area in the corresponding first sample image as the current second sample description text; A first predicted description text generation module, used for inputting the first sample image and the first text prompt corresponding to each current second sample description text into the current description generation model, and obtaining the description text of the specified image area in the first sample image as the current first predicted description text; A first parameter adjustment module, used to adjust the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model; The first text prompt is also used to indicate that: the current description generation model generates a description text that conforms to a first specified description style; a description style is used to indicate the language of the description text and / or the number of sentences included; The device also includes: A third sample description text generation module is used for adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, and then using the first text prompt and the current description generation model to process each current first sample image to obtain a description text of the first specified description style for a specified image area in the first sample image as the current third sample description text; a fourth sample description text generating module, configured to adjust the description style of the current third sample description text from the first specified description style to the second specified description style, to obtain a current fourth sample description text; A fifth sample description text determination module determines, from the current fourth sample description text, a description text that matches the designated image area in the corresponding first sample image, to obtain the current fifth sample description text; A third predicted description text generation module is used to input the first sample image and the third text prompt corresponding to each fifth sample description text into the current description generation model, and obtain the description text of the second specified description style of the specified image area in the first sample image as the current third predicted description text; The third parameter adjustment model is used to adjust the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text to obtain a new description generation model.
13. The device according to claim 12, characterized in that The second sample description text determination module is specifically used to input each current first sample image, the first sample description text corresponding to the first sample image, and the second text prompt into a pre-trained reward model to obtain a matching result between a specified image area in the first sample image and the corresponding first sample description text; wherein the second text prompt is used to indicate whether the specified image area in the input image matches the input description text; the reward model is pre-trained based on positive sample pairs and negative sample pairs; a positive sample pair includes: a second sample image, a specified image area in the second sample image, and a description text that matches the specified image area in the second sample image; a negative sample pair includes: a third sample image, a specified image area in the third sample image, and a description text that does not match the specified image area in the third sample image; when the matching result obtained indicates that the specified image area in the first sample image matches the corresponding first sample description text, the first sample description text corresponding to the first sample image is determined as the corresponding second sample description text; and / or, The reward model includes: a first large language model, a first adapter, a first image encoder and a first token parser; the first image encoder is used to extract features of an input image; the first adapter is used to convert the image features output by the first image encoder into a token form; the first token parser is used to convert the input text into a token form; and / or, The description text in each sample pair is obtained by processing the sample image in the sample pair using the first text prompt and the current description generation model; the device also includes: A reward model training module is used to input, for each sample pair, the sample image and the designated image area in the sample pair into the first image encoder in the reward model of the initial structure, obtain the features of the sample image and the designated image area in the sample pair, and splice the obtained features; input the obtained splicing result into the first adapter in the reward model of the initial structure to obtain the marked form of the splicing result; input the second text prompt and the description text in the sample pair into the first mark parser in the reward model of the initial structure to obtain the marked form of the second text prompt and the marked form of the description text in the sample pair; input the marked form of the splicing result, the marked form of the second text prompt and the marked form of the description text in the sample pair into the first large language model in the reward model of the initial structure to obtain the predicted matching result between the designated image area in the sample pair and the description text; based on the difference between the obtained predicted matching result and the actual matching result of the sample pair, adjust the parameters of the reward model of the initial structure until the preset convergence condition is reached to obtain a trained reward model; and / or, The first sample description text generation module is specifically used to, for each current first sample image, when the proportion of the specified image area in the first sample image is less than a preset threshold, expand the specified image area in the first sample image to obtain an image to be processed; input the obtained image to be processed and the first text prompt into the current description generation model to obtain a description text for describing the specified image area in the first sample image as the current first sample description text; when the proportion of the specified image area in the first sample image is not less than the preset threshold, input the first sample image and the first text prompt into the current description generation model to obtain a description text for describing the specified image area in the first sample image as the current first sample description text; and / or, The first sample description text generation module is specifically used to use the first text prompt and the current description generation model to process the first sample image according to different sampling parameters to obtain multiple description texts of the specified image area in the first sample image as the current first sample description text; and / or, The device also includes: A fourth sample image acquisition module, configured to acquire a fourth sample image after adjusting the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model; wherein, among the description texts generated by the current description generation model for each fourth sample image, there are description texts that do not match the specified image area in the corresponding fourth sample image; A description text inference module, used to process each fourth sample image acquired according to different sampling parameters by using the first text prompt and the current description generation model, to obtain a plurality of description texts of a specified image area in the fourth sample image; a positive and negative sample description text determination module, used to determine, from the obtained description texts, at least one positive sample description text that matches the specified image area in the fourth sample image, and at least one negative sample description text that does not match the specified image area in the fourth sample image; A second predicted description text generation module, used for inputting the fourth sample image and the first text prompt into a current description generation model to obtain a description text of a specified image area in the fourth sample image as a current second predicted description text; A second parameter adjustment module, used to adjust the parameters of the current description generation model based on the difference between the current second predicted description text and the determined positive sample description text, and the difference between the current second predicted description text and the determined negative sample description text, to obtain a new description generation model; and / or, The device also includes: a fourth sample description text updating module, for adjusting the parameters of the current description generation model based on the difference between the current third predicted description text and the current fifth sample description text, and after obtaining a new description generation model, inputting each current first sample image and the third text prompt into the current description generation model, and obtaining a description text of the second specified description style of the specified image area in the first sample image as the current fourth sample description text; and triggering the fifth sample description text determination module until the first preset condition is met, and a new description generation model is obtained; and / or, The device also includes: a triggering module, which adjusts the parameters of the current description generation model based on the difference between the current first predicted description text and the current second sample description text to obtain a new description generation model, and then triggers the first sample description text generation module until a second preset condition is met to obtain a new description generation model; and / or, The description generation model includes: a second language model, a second adapter, a second image encoder and a second token parser; the second image encoder is used to extract features of an input image; the second adapter is used to convert image features output by the second image encoder into a token form; the second token parser is used to convert input text into a token form.
14. A description generating device, characterized in that: The device comprises: A data acquisition module, used to acquire images to be used and text prompts to be used; A description text generation module is used to input the image to be used and the text prompt to be used into a pre-trained description generation model to obtain a description text of the image area indicated by the text prompt to be used in the image to be used; wherein the description generation model is trained based on any one of the methods described in claims 1-9.
15. The device according to claim 14, characterized in that The position of the image area indicated by the to-be-used text prompt in the to-be-used image is obtained by the following steps: Displaying the image to be used; In response to a user's frame selection operation on the displayed image to be used, determining a position indicated by the frame selection operation in the image to be used, and obtaining a position of an image area indicated by the text prompt to be used in the image to be used; and / or, The description style of the obtained description text is the description style indicated by the text prompt to be used; a description style is used to indicate the language of the description text and / or the number of sentences included.
Citation Information
Patent Citations
Large language model alignment fine tuning method and system based on enhanced rejection sampling training
CN117852616A
Method and device for generating wafer defect description based on multi-modal fusion
CN118037706A
Text generation model training method, text generation method and answer generation method
CN118364060A
Target positioning method and related equipment thereof
CN118568289A