An image captioning and question answering method based on multi-task training
By adopting a learnable dynamic weight allocation method in image description and question-answer tasks, combining large-scale data sets and multimodal feature extraction, the independence of image description and question-answer tasks in the prior art is solved, efficient multi-task training and testing in the same model is realized, and the comprehensive performance of image description and question-answer is improved.
Patent Information
- Application Number
- CN202310662783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-06
AI Technical Summary
Existing image description and image question and answer models usually treat two tasks as independent tasks, failing to effectively combine in the same model, resulting in a single use of data sets during training and testing, limiting the overall performance of the model.
Using a learningable dynamic weight allocation method, the BLIP pre-trained model is trained under the Pytorch framework, and the visual feature extraction module and text feature extraction module are used to combine CC3M, CC12M, SBU and LAION data sets to build a multi-task model, and distinguish image description and Q&A data through feature symbols, and dynamically adjust the loss function weight to realize multi-task training of image description and Q&A.
Under the same model, the image description and Q&A were significantly improved, the overall performance of the model was improved, and more image details were captured and high-quality Q&A pairs were generated.
Smart Images

Figure CN116561368B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image description and image question answering in computer vision and computer language, and particularly to an image description and question answering method based on multi-task training that uses a learnable dynamic allocation of weights to achieve multiple tasks within the same framework. Background Art
[0002] Image description and image question answering are two fundamental problems in computer vision and computer language. The image description task is to generate a descriptive text based on a given image; the image question answering task is to generate a corresponding answer based on a given image and question. With the development of large-scale image-text pre-training [1][2][3] computer vision and computer language tasks have been significantly improved.
[0003] Early work mainly focused on using pictures for image classification and object mining to identify image information. Now, with the increase in datasets and the use of large-scale network datasets, deep learning has achieved remarkable results. In image description and image question answering, both computer vision and computer language processing are involved, and the key is to find the connection between image information and text information.
[0004] Currently, with the development of neural networks, both image description and image question answering have made great progress. Most image description and image question answering models are based on the sequence-to-sequence framework. Previous work mostly used artificially annotated datasets such as COCO Caption [4] , Flick30k [5] , etc. RNN networks were trained to become models that could accurately predict text based on a given picture. Now, with the growth of computing power, large-scale network data has begun to be used in the pre-training of images and texts, such as Conceptual Captions [6] , Conceptual 12M [6] , SBU Caption [7] , LAION [8] .
[0005] However, existing models treat image description and image question answering as two independent tasks. When training and testing the image description task, only the image description dataset is simply used, and similarly, only the image question answering dataset is simply used when training and testing the image question answering task.
[0006] References:
[0007] [1] Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022. Blip: Boot-strapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086.
[0008] [2] Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 8748–8763. PMLR
[0009] [3] Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432.
[0010] [4] Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer
[0011] [5]Plummer,B.A.;Wang,L.;Cervantes,C.M.;Caicedo,J.C.;Hockenmaier,J.;and Lazebnik,S.2015.Flickr30k entities:Collecting region-to-phrase correspondences for richer image-to-sentence models.In Proceedings of the IEEE international conference on computer vision,2641–2649
[0012] [6]Changpinyo,S.;Sharma,P.;Ding,N.;and Soricut,R.2021.Conceptual 12m:Pushing web-scale image-text pre-training to recognize long-tail visual concepts.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition,3558–3568
[0013] [7]Ordonez,V.;Kulkarni,G.;and Berg,T.2011.Im2text:Describing images using 1million captioned photographs.Advances in neural information processing systems,24 Summary of the Invention
[0014] An object of the present invention is to provide an image description and question-answering method based on multi-task training that uses a learnable dynamic weight allocation method to achieve multiple tasks in the same framework in view of the above technical problems existing in the prior art.
[0015] The image description and question-answering method based on multi-task training according to the present invention includes the following steps:
[0016] 1) Pre-trained model: Use the pre-trained model of BLIP. This model uses the Pytorch framework and is trained on an 8-GPU system. Use vision transformer as the visual feature extraction module and BERT as the text feature extraction model. Load a large-scale image-text dataset to pre-train the model.
[0017] In step 1), the image-text dataset can adopt CC3M dataset, CC12M dataset, SBU dataset, and LAION dataset.
[0018] 2) Preprocess the image caption data and image question-answering data. Use identifiers to distinguish the input data. The two tasks of image captioning and image question-answering are implemented in the same framework.
[0019] (1) Image caption data: Add the corresponding feature symbol [cap] at the beginning of the input image caption data text so that the model can judge the input text and the corresponding person according to the symbol.
[0020] (2) Image question-answering data: For the image question-answering data, splice the image question and the answer into a question-answer pair, and add the [QA] symbol before the input question-answer text to indicate that the input is a question-answer pair.
[0021] 3) Build an image captioning and image question-answering multi-task model: The multi-task model has a total of three tasks, including the two tasks of model image captioning and image question-answering. Image question-answering includes two subtasks: image question answering and image question generation.
[0022] Load the pre-trained model in step 1), and use the preprocessed image caption data and image question-answering data in step 2) to fine-tune the model in the downstream task. Use multi-modal image-text feature extraction. The image becomes a visual vector through the visual feature extraction module and is used as the attention input to the text feature extraction module. The text uses the LM loss function to enable the model to learn the output content of the text.
[0023] 4) Calculate the corresponding loss function and its weight according to the tasks of model training:
[0024] Input the preprocessed image caption data and image question-answering data into the model. In the same batch, the model obtains both the image caption data and the image question-answering data at the same time, and distinguishes them through the feature symbols, so as to complete the multi-tasks of image captioning and image question-answering.
[0025] 5) Optimize the loss function and its corresponding weight, dynamically learn and allocate the loss weights of multiple tasks, and adjust the optimal learning method.
[0026] The learnable weights are as follows:
[0027]
[0028] Add the loss functions of the three tasks of model image description, image question answering, and image question generation through learnable weights, dynamically optimize the loss functions of the three tasks, and enable the multi-task model to achieve a mutually promoting effect; when the theta parameter (θ) increases, the corresponding weight L decreases, and conversely, when the theta parameter decreases, the corresponding weight L increases. The Theta parameter can be used as a regularization term for the multi-task loss function to prevent the weights of the loss functions of each task from deviating too much.
[0029] 6) Testing phase: Perform image description or image question answering according to requirements. When the user needs an image description, an image and the corresponding image description symbol need to be input, and the model outputs a description of the image; when image question answering is required, an image, a question, and the image question answering symbol are input, and the model outputs an answer; or an image and the image question answering symbol are input, and the model outputs the question and answer corresponding to the image.
[0030] The multi-task model framework of the present invention sequentially includes:
[0031] 1) Hybrid text learning module, which is used to distinguish different types of text by introducing text identifiers.
[0032] 2) Text encoding module, which is used to encode the input text into a format recognizable by the model.
[0033] 3) Multi-task dynamic loss weight allocation module, which is used to dynamically adjust the multi-task loss weights and improve task performance.
[0034] The multi-task model of the present invention has achieved significant improvements in both image description and image question answering, and the indicators have increased. The present invention can be implemented under the same model and share a common data set. In addition, the multi-task model of the present invention also has the function of outputting questions and corresponding answers according to the input image. The question-answer pairs generated by the present invention can capture more details of the image and have high quality. Description of the Drawings
[0035] Figure 1 It is the overall framework diagram of the model;
[0036] Figure 2 It is the comparison between the question-answer pairs generated by the model and the question-answer pairs in the data set. Detailed Embodiments
[0037] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following embodiments will further illustrate the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0038] AsFigure 1 , the general steps of the embodiments of the present invention are as follows:
[0039] I. Load a large-scale image-text pre-training model
[0040] Use the pre-training model of BLIP. This model uses the Pytorch framework and is trained on 8 GPUs. Use vision transformer as the visual feature extraction module and BERT as the text feature extraction model. The datasets used are CC3M, CC12M, SBU, and LAION, a total of 129M image-text pairs.
[0041] II. Preprocess the image caption data and image question-answer data
[0042] 1. Add the feature symbols corresponding to the image caption data.
[0043] For the image caption task, add the [cap] symbol at the beginning of the input text so that the model can judge the input text and the corresponding person according to the symbol.
[0044] 2. Concatenate the question-answer pairs of the image question-answer data and add the corresponding feature symbols.
[0045] For the image question-answer data, concatenate the question and the answer to form a question-answer pair, and add the [QA] symbol before the input text to indicate that the input is a question-answer pair.
[0046] III. Train the multi-task model
[0047] 1. Build an image caption and image question-answer multi-task model.
[0048] After loading the pre-trained model, use the image caption data and image question-answer data to fine-tune the model in the downstream task. Use multi-modal image-text feature extraction. Among them, the image becomes a visual vector through the visual feature extraction module and is used as the attention input to the text feature extraction module. Then, the LM loss function is used for the text, so that the model can learn the output content of the text.
[0049] 2. Calculate the corresponding loss function and its weight according to the tasks of model training.
[0050] Input the preprocessed image caption data and image question-answer data into the model. In the same batch, the model obtains the image caption data and image question-answer data at the same time and distinguishes them through the feature symbols, so as to complete the multi-tasks of image caption and image question-answer.
[0051] 3. Optimize the loss function and its corresponding weight at the same time.
[0052] Compared with the traditional multi-task learning method of simply adding weighted task losses, a method of learnable weights is proposed:
[0053]
[0054] The loss function can be as shown in the above formula. This model includes two tasks: image captioning and image question answering. Among them, image question answering includes two subtasks: image question answering and image question generation, making it a multi-task model with a total of three tasks. By adding the loss functions of the three tasks through learnable weights, the loss functions of the three tasks are dynamically optimized, enabling the multi-task model to achieve a mutually promoting effect. When the theta parameter increases, the corresponding weight L decreases, and conversely, when the theta parameter decreases, the corresponding weight L increases. The Theta parameter can be used as a regularization term for the multi-task loss function, ensuring that the weights of the loss functions of each task do not deviate too much.
[0055] IV. Testing the Multi-Task Model
[0056] In the testing phase, image captioning or image question answering can be performed according to requirements. When the user needs image captioning, they need to input an image and the corresponding image captioning symbol, and the model can output a description of the image; when image question answering is needed, an image and a question can be input along with the image question answering symbol, and the model can output the answer; or when only an image and the image question answering symbol are input, the model can output the question and answer corresponding to the image.
[0057] V. Experimental Tests
[0058] Three different experiments are conducted in the present invention to verify the performance of the introduced method, simultaneously visualize the ability of the present invention to generate question-answer pairs, and compare them with the question-answer pairs in the dataset. The present invention uses pytorch as the deep learning framework, trains on 8-card V100 GPUs, pre-trains the image feature network on ImageNet, and initializes the text feature network through BERT. The pre-trained BLIP model is loaded as the initialization model and fine-tuned on downstream tasks. The datasets used are COCO Caption and VQA2.0. In terms of data augmentation, random cropping is used to crop the images into images of size 384*384.
[0059] 1. Effect on Image Captioning
[0060] Image captioning is one of the fundamental tasks in computer vision and computer language, and the task is to provide a text description of the input image. The test uses the test set of COCO Caption for testing. Table 1 shows the performance of the model in image captioning. It can be seen that the model of the present invention has a significant advantage in image captioning.
[0061] Table 1. Performance in Image Captioning
[0062]
[0063] 2. Effect in Image Question Answering
[0064] Image question answering is one of the basic tasks in computer vision and computer language. The task is to answer the input image and question. The test uses the test set of VQA2.0 for testing. Table 2 shows the performance of the model in image question answering. It can be seen that the model of the present invention has great advantages in image question answering.
[0065] Table 2. Performance of the Model in Image Question Answering
[0066]
[0067] 3. Ablation Experiment
[0068] An ablation experiment on the multi-task loss weight allocation module is provided for image captioning and image question answering. It can be seen from the experiment that the image captioning and image question answering metrics are significantly improved in multi-task training. When using the multi-task dynamic loss weight allocation module, the performance of image captioning is greatly improved, and the performance of image question answering remains basically the same.
[0069] Table 3. Ablation Experiment of the Multi-task Dynamic Loss Weight Allocation Module
[0070]
[0071] 4. Comparison between the Generated Question-Answer Pairs and the Dataset
[0072] Figure 2 A comparison between the question-answer pairs generated by the present invention for the input pictures and the VQA2.0 dataset is provided. It can be seen that the question-answer pairs generated by the present invention can capture more picture details and have high quality.
[0073] Image captioning and image question answering are two basic tasks in the fields of computer vision and computer language. These two tasks are trained and tested using similar models. The main task of training the model is to train and learn the association between images and texts. Although these two tasks use similar models, they cannot be implemented in the same model. Existing methods can only implement one function in one model. Explore the connection between the two tasks and simultaneously implement the training and testing of these two tasks in a sequence to sequence framework. Experiments show that under the framework of the present invention, the image captioning and image question answering tasks can be implemented in the same model and promote each other, and the performance of these two tasks can be greatly improved.
Claims
1. An image captioning and question answering method based on multi-task training, characterized in that It includes the following steps: 1) Pre-train the model. Use the pre-trained model of BLIP and load a large-scale image-text dataset to pre-train the model. 2) Preprocess the image caption data and image question-answering data. Use identifiers to distinguish the input data. The two tasks of image captioning and image question-answering are implemented in the same framework. The preprocessing of the image caption data and image question-answering data includes: (1) Image caption data: Add the corresponding feature symbol [cap] at the beginning of the input image caption data text, so that the model can judge the input text and the corresponding person according to the symbol. (2) Image question-answering data: For the image question-answering data, splice the image question and the answer into a question-answer pair, and add the [QA] symbol before the input question-answer text to indicate that the input is a question-answer pair. 3) Construct an image captioning and image question-answering multi-task model. The multi-task model includes two tasks: image captioning and image question-answering. Image question-answering includes two sub-tasks: image question answering and image question generation. Load the pre-trained model in step 1), and use the preprocessed image caption data and image question-answering data in step 2) to fine-tune the model in the downstream task. Use multi-modal image-text feature extraction. The image becomes a visual vector after passing through the visual feature extraction module and is used as the attention input to the text feature extraction module. The text uses the LM loss function so that the model can learn the output content of the text. 4) Calculate the corresponding loss function and its weight according to the tasks of model training. 5) Optimize the loss function and its corresponding weight, dynamically learn to allocate the loss weights of multiple tasks, and adjust the optimal learning method. The learnable weights are as follows: Add the loss functions of the three tasks of model image captioning, image question answering, and image question generation through the learnable weights, and dynamically optimize the loss functions of the three tasks to make the multi-task model achieve a mutually promoting effect. When the θ parameter increases, the corresponding weight L decreases. On the contrary, when the θ parameter decreases, the corresponding weight L increases. The θ parameter is used as the regularization term of the multi-task loss function to prevent the loss function weights of each task from deviating too much. 6) Use the trained multi-task model to perform image captioning or image question-answering according to user needs. The framework of the multi-task model sequentially includes: a hybrid text learning module, a text encoding module, and a multi-task dynamic loss weight allocation module. The hybrid text learning module is used to distinguish different types of text by introducing text identifiers. The text encoding module is used to encode the input text into a format recognizable by the model. The multi-task dynamic loss weight allocation module is used to dynamically adjust the multi-task loss weights and improve the task performance.
2. The method for image description and question answering based on multi-task training according to claim 1, characterized in that In step 1), the pre-trained model uses the Pytorch framework and is trained on an 8-GPU. Use vision transformer as the visual feature extraction module and BERT as the text feature extraction model.
3. The method for image description and question answering based on multi-task training according to claim 1, characterized in that In step 1), the image-text dataset includes the CC3M dataset, CC12M dataset, SBU dataset, and LAION dataset.
4. The image description and question answering method based on multi-task training according to claim 1, characterized in that In step 4), calculate the corresponding loss function and its weight according to the tasks of model training: input the preprocessed image description data and image question-and-answer data into the multi-task model. In the same batch, the model simultaneously obtains the image description data and the image question-and-answer data, and differentiates them through feature symbols, so as to complete the multi-tasks of image description and image question-and-answer of the image.
5. The method for image description and question answering based on multi-task training according to claim 1, characterized in that In step 6), use the trained multi-task model to perform image description or image question-and-answer according to the user's needs. According to the needs, perform image description or image question-and-answer. When the user needs image description, an image and the corresponding image description symbol need to be input, and the model outputs the description of the corresponding image. When image question-and-answer is needed, input an image and a question along with the image question-and-answer symbol, and the model outputs the answer; or input an image and the image question-and-answer symbol, and the model outputs the question and answer corresponding to the image.
Citation Information
Patent Citations
Image analysis method and device, storage medium, electronic equipment and product
CN117009567A