Image question and answer method and device and medium
By segmenting the image into image blocks related to multiple feature points, combining the intermediate language model and the large language model, the problem of difficulty in capturing image information of VQA model is solved, and high-quality image question-and-answer function is realized.
Patent Information
- Application Number
- CN202510174585.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
The existing VQA model is difficult to effectively capture the global and local information of the image, resulting in the loss of image information and is unable to support answering user questions.
By segmenting the Q&A image into multiple image blocks based on feature points, filtering image blocks with high correlation with user problems, capturing global and local information of the image using the intermediate language model, and obtaining answers to user questions based on the overview through the large language model.
It realizes effective capture of global and local information of images, enhances the ability of natural language to overview visual images, and supports high-quality answers to user questions.
Smart Images

Figure CN120126166A_ABST
Abstract
Description
Technical Field
[0001] This application is at least related to the field of computer technology, and particularly relates to an image question answering method, device, and medium. Background Art
[0002] Visual Question Answering (VQA) involves technologies in multiple fields such as image processing, computer vision, and natural language processing. It requires the model to be able to answer natural language questions based on the content of the image.
[0003] However, images and texts are two different data streams in VQA, and their probability distributions vary greatly. Simple splicing and adding of main elements may not be sufficient to fuse the features of the two modalities. Moreover, existing VQA models often lack the grasp of image details, resulting in the loss of image information and being unable to support answering user questions. Summary of the Invention
[0004] In view of the above deficiencies, this application provides an image question answering method, device, and medium to solve the following technical problems: how to capture the global and local information of the image and effectively fuse visual and language information to support answering user questions.
[0005] In a first aspect, this application provides an image question answering method, which includes:
[0006] Segmenting the question answering image into multiple first question answering image blocks based on feature points;
[0007] Obtaining several second question answering image blocks in the first question answering image blocks with a relevance greater than a first threshold to the user question;
[0008] Obtaining a first overview of the question answering image and a second overview of the several second question answering image blocks based on an intermediate language model;
[0009] Obtaining an answer to the user question based on a large language model according to the first overview and the second overview.
[0010] Further, segmenting the question answering image into multiple first question answering image blocks based on feature points includes:
[0011] Obtaining the feature points of the question answering image, specifically including: extracting first feature points of the question answering image based on deep learning feature point detection SuperPoint or Scale-Invariant Feature Transform (SIFT), and screening second feature points from the first feature points using Non-Maximum Suppression (NMS);
[0012] Segment the Q&A image into multiple first Q&A image blocks according to the feature points, specifically including: segmenting multiple third Q&A image blocks with multiple sizes and / or multiple ratios from the Q&A image centered on each second feature point, and screening the first Q&A image blocks from the third Q&A image blocks based on the intersection over union (IoU) algorithm.
[0013] Further, obtain several second Q&A image blocks in the first Q&A image blocks with a relevance to the user question greater than a first threshold, specifically including:
[0014] Use the text encoder of the contrastive language-image pre-training (CLIP) model to obtain the first text encoding t of the user question;
[0015] Use the image encoder of the CLIP model to obtain the first visual encoding v of each first Q&A image block i , where i is the number of the first Q&A image block;
[0016] Calculate the first similarity score between each first Q&A image block and the user question
[0017] Obtain all s i > τ 3 of the first Q&A image blocks as the second Q&A image blocks with a relevance to the user question greater than the first threshold τ 3 of the user question.
[0018] Further, obtain the first overview of the Q&A image and the second overview of the several second Q&A image blocks based on an intermediate language model, specifically including:
[0019] Based on the intermediate language model, respectively obtain multiple third overviews of the Q&A image and multiple fourth overviews of each second Q&A image block, obtain a fifth overview with a relevance to the user question greater than a second threshold in the third overviews and a sixth overview with a relevance to the user question greater than a third threshold in the fourth overviews, and obtain the first overview according to the fifth overview and the second overview according to the sixth overview;
[0020] Among them, the intermediate language model is an image captioning (Image Caption) model trained based on a semi-supervised method. The semi-supervised method includes: segmenting multiple training images, obtaining manually annotated caption overviews for some of the segmented training image blocks, combining the Image Caption model with a large language model to obtain pseudo-annotated caption overviews for another part of the training image blocks, supervising the training of the Image Caption model based on the annotated caption overviews, and self-supervised training of the Image Caption model based on the pseudo-annotated caption overviews.
[0021] Further, based on the intermediate language model, obtain multiple third overviews of the Q&A image and multiple fourth overviews of each second Q&A image block respectively, obtain a fifth overview with a relevance to the user question greater than a second threshold in the third overviews and a sixth overview with a relevance to the user question greater than a third threshold in the fourth overviews, and obtain the first overview according to the fifth overview and the second overview according to the sixth overview, which specifically includes:
[0022] Obtain the attention over attention neural network AoANet model on top of the pre-trained first attention;
[0023] Use the first AoANet model multiple times to output multiple third overviews of the Q&A image, obtain the second similarity score between each third overview and the user question, obtain a fifth overview with a second similarity score greater than the second threshold from the third overviews, and obtain the first overview of the Q&A image according to the fifth overview;
[0024] Use the first AoANet model multiple times to output multiple fourth overviews of each of the several second Q&A image blocks, obtain the third similarity score between each fourth overview and the first overview and / or the user question, obtain a sixth overview with a third similarity score greater than the third threshold from the fourth overviews, and obtain the second overview of each of the several second Q&A image blocks according to the sixth overview.
[0025] Further, the method further includes pre-training the first AoANet model using a partially labeled dataset, which specifically includes:
[0026] Obtain training images;
[0027] Segment the training images into multiple training image blocks according to feature points;
[0028] Obtain the annotation overviews of the training images and partial training image blocks of each training image;
[0029] Perform the first training on the second AoANet model using the annotation overviews to obtain the third AoANet model;
[0030] Use the third AoANet model and the large language model to obtain the first pseudo-annotation overviews of the unannotated training image blocks;
[0031] Assign different weights to the annotation overviews and the first pseudo-annotation overviews;
[0032] Perform the second training on the third AoANet model using the annotation overviews and the first pseudo-annotation overviews with different weights to obtain the first AoANet model.
[0033] Further, where:
[0034] The training images include data images from different scenarios;
[0035] Segmenting the training images includes: extracting the top-k feature points in each training image, and segmenting each training image into k training image patches based on the top-k feature points;
[0036] The partial training image patches are n training image patches randomly selected from the k training image patches, where n < k / 3;
[0037] Cross-entropy loss is used as the loss function in the first training and the second training;
[0038] The first pseudo-label summary is obtained by using the third AoANet model to output multiple second pseudo-label summaries for each unlabeled training image patch, and inputting the multiple second pseudo-label summaries into a large language model for summarization;
[0039] Assigning different weights to the labeled summary and the first pseudo-label summary includes: assigning a first weight to the labeled summary and a second weight to the first pseudo-label summary, where the first weight is greater than 10 times the second weight.
[0040] Furthermore, obtaining the answer to the user's question based on the large language model according to the first summary and the second summary specifically includes:
[0041] Obtaining the size of the Q&A image and the position coordinates of each of the several second Q&A image patches in the Q&A image, as well as example Q&A pairs regarding the first summary and the second summary;
[0042] Writing the size and the first summary into the first blank of the Prompt template, writing each position coordinate and the corresponding second summary into the second blank of the Prompt template, and writing the example Q&A pairs into the third blank of the Prompt template to obtain the prompt word Prompt;
[0043] Inputting the Prompt into the large language model so that the large language model outputs the answer to the user's question according to the Prompt.
[0044] In a second aspect, the present application provides an image Q&A device, and the device includes:
[0045] A segmentation module, configured to segment a Q&A image into multiple first Q&A image patches based on feature points;
[0046] A screening module, connected to the segmentation module, configured to obtain several second Q&A image patches in the first Q&A image patches with a relevance to the user's question greater than a first threshold;
[0047] An overview module, connected to the screening module, is used to obtain a first overview of the Q&A image and a second overview of the several second Q&A image blocks based on an intermediate language model;
[0048] A Q&A module, connected to the overview module, is used to obtain an answer to the user's question based on a large language model according to the first overview and the second overview.
[0049] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the above-mentioned image Q&A method is implemented.
[0050] The present application provides an image Q&A method, device and medium. The image is segmented based on feature points, image blocks with high relevance to the user's question are screened, and the global information and local information of the image are captured by using an intermediate language model according to the overall image and the highly relevant local image blocks, enhancing the overview of the intermediate natural language of the visual image. Finally, a large language model is used to obtain a high-quality answer to the user's question according to the overview. Description of the Drawings
[0051] Figure 1 is a flowchart of an image Q&A method according to an embodiment of the present application;
[0052] Figure 2 is a schematic structural diagram of an image Q&A device according to an embodiment of the present application;
[0053] Figure 3 is a flowchart of another image Q&A method according to an embodiment of the present application. Detailed Embodiments
[0054] To enable those skilled in the art to better understand the technical solutions of the present application, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0055] It can be understood that the specific embodiments and drawings described herein are only used to explain the present application, rather than limiting the present application.
[0056] It can be understood that, without conflict, the various embodiments and features in the embodiments of the present application can be combined with each other.
[0057] It can be understood that for the convenience of description, only the parts related to the present application are shown in the drawings of the present application, and the parts unrelated to the present application are not shown in the drawings.
[0058] It can be understood that each module and unit involved in the embodiments of the present application may correspond to only one entity structure, or may be composed of multiple entity structures. Alternatively, multiple modules and units may also be integrated into one entity structure.
[0059] It is understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present application may occur in an order different from that marked in the accompanying drawings.
[0060] It is understood that in the flowcharts and block diagrams of the present application, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present application are shown. Among them, each block in the flowchart or block diagram may represent a module, unit, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart may be implemented by a hardware-based device for implementing the specified function, or by a combination of hardware and computer instructions.
[0061] It is understood that the modules and units involved in the embodiments of the present application may be implemented in software or in hardware. For example, the modules and units may be located in the processor.
[0062] Embodiment 1:
[0063] As Figure 1 shown, the present application provides an image question-answering method, and the method includes:
[0064] S1. Segment the question-answering image into multiple first question-answering image blocks based on feature points;
[0065] S2. Obtain several second question-answering image blocks in the first question-answering image blocks with a relevance to the user's question greater than a first threshold;
[0066] S3. Obtain a first overview of the question-answering image and a second overview of the several second question-answering image blocks based on an intermediate language model;
[0067] S4. Obtain an answer to the user's question based on a large language model according to the first overview and the second overview.
[0068] In this embodiment, the method segments the image based on feature points, screens the image blocks with high relevance to the user's question, captures the global information and local information of the image by using the intermediate language model according to the overall image and the local image blocks with high relevance, enhances the overview of the intermediate natural language for the visual image, and finally uses the large language model to obtain a high-quality answer to the user's question according to the overview. As Figure 1 shown, the method is correspondingly applied to such as Figure 2The device shown. It can be understood that steps S1 - S4 do not need to be strictly in the order of their execution. For example, the image can be pre - segmented, the first summary and the second summary can be pre - generated based on an intermediate language model. The second summary can be first generated for all the first question - answering image blocks, and after executing S2, the second summary corresponding to the second question - answering image block can be directly obtained.
[0069] Specifically, this embodiment provides a method for image question - answering based on feature - point detection, and a corresponding system can also be provided. It mainly considers capturing the global and local information of an image based on an intermediate - language model with a small amount of computing power to support user questions. In terms of mastering the local information of the image, the image is segmented based on feature points and the segmented image blocks are screened to improve the understanding of the image details; in terms of connecting the image with text data, the visual image is first transformed into a natural - language description and then input into an LLM (Large Language Model). There is no need to fine - tune the large model (abbreviation of the large - language model). With limited computing power, the intermediate natural language's summary of the visual image is enhanced to fully support answering user questions and obtain high - quality answers to user questions.
[0070] More specifically, the improvement of VQA model performance depends on the popularity of large model pre-training models and the use of large-scale VQA datasets, requiring the model to be able to effectively process and combine visual and language information to improve multi-modal feature fusion capabilities. For example, by pre-training with a large-scale dataset, the visual question answering (VQA) model can learn rich visual and language representations and then be fine-tuned on specific VQA tasks to improve performance. However, the alignment between the two modalities of images and texts often depends on high-quality pre-aligned datasets. In a formal usage environment, to implement VQA, a large amount of computing resources are required to train and optimize the model, and there are also high requirements for the quality of the dataset. There are multiple difficulties in fine-tuning large models: the computing power resources are limited and insufficient to support the training and fine-tuning of multi-modal large models; it is very difficult to obtain high-quality formal VQA data, and the cost of manually aligning questions and images and providing answer examples is extremely high. Research on VQA problems usually mainly focuses on model pre-training and VQA models based on intermediate languages. Multi-modal pre-training often requires adding additional visual modules and visual-text alignment modules on the basis of large models. Due to the extremely large number of parameters in the LLM, usually only a part of the model is selected for training, such as the visual encoder and the cross-attention mechanism, to fit the differences between modalities, which still requires high-quality datasets and long training times. If the LLM, visual encoder, and alignment module are jointly trained, it may even reduce the performance of the LLM and produce unpredictable results. Another approach is to directly describe visual information using natural language. Models based on intermediate languages often need to first convert visual images into natural language descriptions and input them into the LLM. However, this approach lacks the grasp of image details, resulting in the loss of image information and being unable to support answering user questions. Therefore, it is necessary to consider how to capture the global and local information of images with limited computing power and a lack of high-quality datasets to support user questions.
[0071] The method and system provided in this embodiment do not require fine-tuning of the large model. In the case of limited computing power and a small amount of high-quality data, it can enhance the natural language in the middle to summarize visual images, and can solve various problems faced in the VQA field, including: the complexity of multi-modal pre-training, the possible performance degradation caused by joint training, the information loss based on the intermediate language model, the dependence on high-quality data sets, and computing power limitations, etc. Specifically, it includes: 1) Challenges in multi-modal pre-training: In multi-modal pre-training, additional visual modules and visual-text alignment modules need to be added on the basis of a large model, which not only increases the complexity of the model but also raises the demand for computing power. In addition, due to the huge number of parameters in large language models (LLMs), usually only a part of the model is selected for training, such as the visual encoder and cross-attention mechanism, to adapt to the differences between modalities, which limits the generalization ability of the model. 2) Possible performance degradation caused by joint training: When the LLM is jointly trained with the visual encoder and alignment module, it may reduce the performance of the LLM and produce unpredictable results, which poses a challenge to the stability and reliability of the model. 3) Limitations of the VQA model based on intermediate language: The model based on intermediate language needs to first convert the visual image into a natural language description and then input it into the LLM. This way may lead to the loss of image details and cannot fully support answering users' questions because it lacks the grasp of image details. 4) Dependence on high-quality data sets: Existing VQA models usually rely on large-scale, fully annotated data sets for training. In the zero-shot setting, due to the lack of specific training data, it is difficult for the model to perform effective reasoning and generalization. 5) Computing power limitations: In the case of limited resources, how to capture the global and local information of the image to support users' questions, especially in the case of lacking high-quality data sets and computing power, how to effectively train the VQA model is an urgent problem to be solved.
[0072] This embodiment provides a system and method for image question answering based on feature point detection. Without fine-tuning the large model and with a small amount of high-quality data, it can perform image question answering and overcome the above-mentioned problems. The method is specifically as Figure 3 shown.
[0073] In one embodiment, S1: Segment the question-and-answer image into multiple first question-and-answer image blocks based on feature points, including:
[0074] Obtain the feature points of the question-and-answer image, specifically including: Extract the first feature points of the question-and-answer image based on the deep learning feature point detection SuperPoint or Scale-Invariant Feature Transform (SIFT), and use Non-Maximum Suppression (NMS) to screen the second feature points from the first feature points;
[0075] Segment the Q&A image into multiple first Q&A image blocks according to the feature points, specifically including: segmenting multiple third Q&A image blocks with multiple sizes and / or multiple ratios from the Q&A image centered on each second feature point, and screening the first Q&A image blocks from the third Q&A image blocks based on the Intersection over Union (IoU) algorithm.
[0076] In this embodiment, as Figure 3 shown, the method includes:
[0077] S01. Input an image and a question, that is, input an image (Q&A image) for answering the user's question and the user's question into the system described in this embodiment. The user's question is proposed by the user for the Q&A image;
[0078] S02. Obtain image feature points, calculate the feature points in the image, and obtain a feature point map based on non-maximum suppression (NMS, Non-Maximum Suppression); specifically including: extracting the feature points (first feature points) of the image based on the neural network SuperPoint (a method for feature point detection and descriptor generation based on deep learning) or the feature point extraction method SIFT (Scale-invariant feature transform); adopting the non-maximum suppression method to extract the feature points (second feature points) with the highest probability among the above feature points (first feature points) to obtain a feature point set P;
[0079] S03. Segment the image based on the image feature points, and use the information of the feature point map to obtain image blocks with different ratios and different sizes centered on the feature points; specifically including: taking the position of each point in the feature point set P as the center, and taking the size of τ 1 % of the original image (the whole Q&A image) as the benchmark to obtain three types of anchor sets A of 1:1, 2:1, and 1:2, where τ 1 is defaulted to 5; based on the confidence ranking of the feature points, and based on the IoU (Intersection over Union, a commonly used evaluation index in object detection to measure the overlap degree between the predicted bounding box and the true bounding box) algorithm, screen the anchor boxes not exceeding the threshold τ 2 and obtain a set I p of image blocks based on them, where τ 2 is set to 0.65;
[0080] It is understandable that SuperPoint, SIFT, NMS, and IoU can all be specifically implemented with reference to the prior art. In this embodiment, these technologies are mainly used to extract and screen image feature points and image patches; there are many other existing and mature technologies for extracting image feature points, and this application is not limited to using only the two methods of SuperPoint or SIFT.
[0081] In one embodiment, S2, obtaining a plurality of second question-and-answer image patches in the first question-and-answer image patch with a relevance greater than a first threshold to the user's question specifically includes:
[0082] Using the text encoder of the contrastive language-image pre-training CLIP model to obtain the first text encoding t of the user's question;
[0083] Using the image encoder of the CLIP model to obtain the first visual encoding v of each first question-and-answer image patch i , where i is the number of the first question-and-answer image patch;
[0084] Calculating the first similarity score between each first question-and-answer image patch and the user's question
[0085] Obtaining all s i >τ 3 of the first question-and-answer image patches as the second question-and-answer image patches with a relevance greater than the first threshold τ 3 to the user's question.
[0086] In this embodiment, as Figure 3 shown, the method further includes: S04, obtaining image patches with high relevance to the question, calculating the degree of relevance between the image patches finally obtained in step S03 and the text question in sequence, and setting a threshold to obtain a set of question-related image patches; specifically including:
[0087] Using the text encoder in CLIP (Contrastive Language-Image Pre-Training, a multi-modal pre-training neural network that is pre-trained with a large amount of paired image and text data to learn the alignment relationship between images and texts) to encode the question Query to obtain the text encoding t of the question = Encoder T (Query);
[0088] Using the image encoder in CLIP to encode the image patches finally obtained in step S03 to obtain the visual encoding v of the image patches i = Encoder I (I i ), where I i represents the image set Ip the i-th image patch in i is the corresponding visual encoding;
[0089] Calculate the similarity score between the question and the images, using the cosine similarity distance for calculation. The cosine similarity calculation formula is:
[0090]
[0091] Select the image patches whose similarity scores with the question exceed the threshold τ 3 to obtain a new set of image patches I c , where τ 3 is set to 0.6.
[0092] In one embodiment, S3. Obtain the first overview of the Q&A image and the second overviews of the several second Q&A image patches based on the intermediate language model, which specifically includes:
[0093] Based on the intermediate language model, obtain multiple third overviews of the Q&A image and multiple fourth overviews of each second Q&A image patch respectively. Obtain a fifth overview whose relevance to the user question in the third overviews is greater than a second threshold and a sixth overview whose relevance to the user question in the fourth overviews is greater than a third threshold. Obtain the first overview according to the fifth overview and obtain the second overview according to the sixth overview;
[0094] Among them, the intermediate language model is an image captioning Image Caption model trained based on a semi-supervised method. The semi-supervised method includes: segmenting multiple training images, obtaining manually annotated caption overviews for some of the segmented training image patches, combining the Image Caption model with a large language model to obtain pseudo-annotated caption overviews for another part of the training image patches, supervising the training of the Image Caption model based on the annotated caption overviews, and self-supervising the training of the Image Caption model based on the pseudo-annotated caption overviews.
[0095] In this embodiment, as Figure 3As shown, the method further includes: S05. Calculate the overview of the image patches, train an Image Caption model based on a semi-supervised manner to generate the overview information (Caption) of the image patches obtained in step S04, as well as the overview information of the overall Q&A image; in order to improve the accuracy of the overview information for answering user questions, generate the overviews of the same image multiple times through the Image Caption model, and filter out the overview information with a high degree of relevance to the question, and synthesize a Caption that is relevant to the question, diverse and clean; use a partially labeled dataset, train the Image Caption model based on a semi-supervised manner, construct it as a fine-tuning dataset, and draw on the overview ability of the large model, and combine the self-supervised manner to improve the generalization ability of the model.
[0096] In one embodiment, based on the intermediate language model, obtain multiple third overviews of the Q&A image and multiple fourth overviews of each second Q&A image patch respectively, obtain a fifth overview with a degree of relevance to the user question greater than a second threshold in the third overviews and a sixth overview with a degree of relevance to the user question greater than a third threshold in the fourth overviews, obtain the first overview according to the fifth overview and obtain the second overview according to the sixth overview, which specifically includes:
[0097] Obtain the Attention over Attention Network (AoANet) model above the first attention obtained through pre-training;
[0098] Use the first AoANet model multiple times to output multiple third overviews of the Q&A image, obtain the second similarity score between each third overview and the user question, obtain a fifth overview with a second similarity score greater than the second threshold from the third overviews, and obtain the first overview of the Q&A image according to the fifth overview;
[0099] Use the first AoANet model multiple times to output multiple fourth overviews of each of the several second Q&A image patches, obtain the third similarity score between each fourth overview and the first overview and / or the user question, obtain a sixth overview with a third similarity score greater than the third threshold from the fourth overviews, and obtain the second overview of each of the several second Q&A image patches according to the sixth overview.
[0100] In this embodiment, as Figure 3 shown in step S05, specifically includes:
[0101] Fine-tune the AoANet (Attention on Attention Network, a neural network architecture for Image Super-Resolution (ISR)) model using a partially labeled dataset, so that the model can capture the content of small-sized images;
[0102] Use the partially fine-tuned AoANet model to infer the overview information of the original image. This process is repeated M times to obtain a stable set C of Caption information. a ;
[0103] Again, use the text encoder and visual encoder in CLIP to calculate the similarity between the overview information of the original image and the original image, as well as the similarity between the overview information and the question, so as to further obtain a synthetic caption C of the complete image that is relevant to the question, diverse, and clean. ac , and the threshold τ 4 is set to 0.75;
[0104] Use the partially fine-tuned AoANet model to infer the overview information of the image patches related to the question, and filter them according to the similarity with the complete image. The threshold τ 5 is set to 0.35, so as to obtain a set C of Caption information for the image patches. p .
[0105] In one embodiment, the method further includes pre-training a first AoANet model using a partially labeled dataset, specifically including:
[0106] Obtain training images;
[0107] Segment the training images into multiple training image patches according to feature points;
[0108] Obtain the annotation overviews of the training images and partial training image patches of each training image;
[0109] Perform the first training on the second AoANet model using the annotation overviews to obtain a third AoANet model;
[0110] Use the third AoANet model and a large language model to obtain the first pseudo-annotation overviews of the training image patches without annotations;
[0111] Assign different weights to the annotation overviews and the first pseudo-annotation overviews;
[0112] Perform the second training on the third AoANet model using the annotation overviews and the first pseudo-annotation overviews with different weights to obtain the first AoANet model.
[0113] In this embodiment, the Image Caption model is trained based on a semi-supervised manner. The Image Caption model specifically uses the AoANet model, and the training process specifically includes:
[0114] Construct a fine-tuning dataset, collect data images from different scenarios, use the SuperPoint network to extract the top-k feature points significantly present in the images, and segment the images based on the feature points to obtain image patches, where k is set to 25. Manually annotate a part of the images, which includes the caption of the original image and the caption information of three randomly selected image patches (annotation overview);
[0115] Fine-tuning of the AoANet model (including the first training and the second training). Fine-tune the AoANet model in a supervised manner on a partially labeled dataset (the first training), and fine-tune the AoANet model in a self-supervised manner on another unlabeled dataset (the second training) to reduce the dependence on labeled data. The AoANet model uses cross-entropy loss as the loss function, and the calculation process is as follows:
[0116]
[0117] Where T represents the text length of the caption, y represents the word vector corresponding to the word in the caption output by the model, and θ represents the parameters of the AoANet model, represents the true sequence of the target caption;
[0118] In the training stage, first use the labeled data for training for 100 epochs (which means the training process of passing through all samples in the training dataset once and only once), the training weight is 1.0, sample once every 10 epochs, and sample three times in total (the first training). Use these three models and their weights to label the unlabeled data, obtain three image overviews, and feed these image overviews into the large model for summarization to obtain the pseudo-labels of the images (the first pseudo-annotation overview); then, use both the labeled data and the pseudo-labeled data for model training (the second training), the weight of the labeled data is 1.0, the weight of the pseudo-labeled data is 0.05, the labeled data constrains the model to be correctly optimized, and the pseudo-labeled data improves the generalization ability of the model. After that, re-label the unlabeled data every 10 epochs.
[0119] In one embodiment, where:
[0120] The training images include data images from different scenarios;
[0121] Segmenting the training images includes: extracting the top-k feature points in each training image, and segmenting each training image into k training image patches based on the top-k feature points;
[0122] The part of the training image patches are n training image patches randomly selected from the k training image patches, where n < k / 3;
[0123] The cross-entropy loss is used as the loss function in the first training and the second training;
[0124] The first pseudo-label summary is obtained by using the third AoANet model to output multiple second pseudo-label summaries for each unlabeled training image patch and inputting the multiple second pseudo-label summaries into a large language model for summarization;
[0125] Assigning different weights to the annotation summary and the first pseudo-label summary includes: assigning a first weight to the annotation summary and a second weight to the first pseudo-label summary, where the first weight is greater than 10 times the second weight.
[0126] In this embodiment, k is set to 25, n is set to 3, the first weight is set to 1.0, and the second weight is set to 0.05. The specific data can be adjusted according to the actual situation. The principle is to reduce the demand for labeled data while ensuring the model effect. The method of splitting the training images is similar to the method of splitting the Q&A images and corresponds to each other.
[0127] In one embodiment, S4. Based on the large language model, obtain the answer to the user's question according to the first summary and the second summary, specifically including:
[0128] Obtain the size of the Q&A image and the position coordinates of each of the several second Q&A image patches in the Q&A image, as well as example Q&A pairs regarding the first summary and the second summary;
[0129] Write the size and the first summary into the first part vacancy of the Prompt template, write each position coordinate and the corresponding second summary into the second part vacancy of the Prompt template, and write the example Q&A pairs into the third part vacancy of the Prompt template to obtain the prompt word Prompt;
[0130] Input the Prompt into the large language model so that the large language model outputs the answer to the user's question according to the Prompt.
[0131] In this embodiment, as Figure 3 shown, the method further includes:
[0132] S06. QA (Question-Answer) pair generation. For a group of overview information (Caption) of an image, generate related questions for the words or phrases that appear to obtain possible question-answer pairs. For example, use a large model to generate the corresponding context for the image description sets C ac and C p generate the corresponding context, and use the Caption information as the answer and use the large model to infer possible questions;
[0133] S07. Summarize the answers to the questions of the large model, splice the overview information (Caption) and Q&A pairs of the image, and the questions asked by the user, and feed them into a large model to obtain the final answer, specifically including: integrating the inference instruction, the Caption information C that has been obtained ac and C p , and the generated Q&A pair examples to construct a complete large model prompt, and feed the integrated prompt into the large model to obtain the answer to the question;
[0134] The prompt for this step contains four parts: instruction, overall overview of the image, overview of image patches, and Q&A pairs. First, the setting is "Please infer the answer to the question based on the context", the overall overview of the image is set to "The image size is ([weight], [height]) and mainly describes [captions]", and the overview of each image patch is set to "([x, y, x 1 , y 1 ) position contains [caption]", and the format of a single QA example is "Question: [question] Answer: [answer]", and they are concatenated; where, weight represents the width of the image, height represents the height of the image, (x, y) represents the coordinates of the upper left corner of the image patch, (x 1 , y 1 ) represents the coordinates of the lower right corner of the image patch.
[0135] The method described in this embodiment ultimately aims to implement an image understanding and image Q&A system. By integrating key modules such as feature point detection, image understanding, and Q&A generation, this system can effectively understand and answer user's image-based questions in the absence of high-quality datasets and multimodal large models; the feature point detection module is responsible for detecting key feature points from the input image and obtaining the corresponding anchor boxes; the image understanding module is responsible for understanding the image content and associating the user's questions with the image content; based on the results of the image understanding module, the Q&A generation module is responsible for generating answers to the user's questions.
[0136] The method mainly includes two parts: the inference process of image Q&A based on feature matching and the training of the image overview model based on semi-supervised learning.
[0137] The first part mainly includes: using the information of the feature point map, obtaining image patches of different scales and sizes centered on the feature points, and extracting the image patches related to the feature points to provide context information for the positioning and understanding of the problem; obtaining a set of image patches related to the problem, and screening the image patches directly related to the problem to prepare for generating an accurate image overview and question answer; obtaining the overview information (Caption) of the image patches, which involves generating a text description of the image patches to provide a text basis for understanding the image content and answering related questions; combining with a large model to obtain possible question-answer pairs, using the capabilities of the large model to generate possible question-answer pairs related to the problem to provide examples for determining the final answer; splicing the overview information (Caption) of the image, the question-answer pairs, and the question asked by the user, and feeding them into a large model to obtain the final answer. This step synthesizes all the information and obtains an accurate answer to the user's question through the deep learning capabilities of the large model.
[0138] The second part mainly adopts a semi-supervised method to train the model on a limited-labeled dataset. On the basis of ensuring the inference ability of the model, it improves the generalization ability of the model, works under limited computing power, combines supervised and self-supervised learning to enhance the model generalization ability on a small amount of labeled data, and pays more attention to the plug-and-play of the model (intermediate language description) and the compatibility with different large models, so that an effective image question-answer function can be realized even in a resource-constrained environment, which is more flexible, can be combined with different large models, has a lower dependence on high-quality datasets, and is more suitable for resource-constrained or fast-deployment scenarios.
[0139] Based on the above two parts, through natural language processing technology, the visual image information is combined with the large model. By extracting image features and inputting them into the language generation model to output a description, it can not only capture global and local information, but also jointly generate answers with the model, showing stronger versatility and generalization ability.
[0140] Embodiment 2:
[0141] As Figure 2 shown, the present application provides an image question-answer device, and the device includes:
[0142] A segmentation module 1, configured to segment the question-answer image into multiple first question-answer image patches based on feature points;
[0143] A screening module 2, connected to the segmentation module 1, configured to obtain several second question-answer image patches in the first question-answer image patches whose relevance to the user's question is greater than a first threshold;
[0144] The overview module 3, connected to the screening module 2, is used to obtain the first overview of the Q&A image and the second overview of the several second Q&A image blocks based on the intermediate language model;
[0145] The Q&A module 4, connected to the overview module 3, is used to obtain the answer to the user's question based on the large language model according to the first overview and the second overview.
[0146] In one embodiment, the segmentation module 1 includes:
[0147] The feature unit is used to obtain the feature points of the Q&A image, specifically: based on the deep learning feature point detection SuperPoint or the scale-invariant feature transform SIFT, extract the first feature points of the Q&A image, and use non-maximum suppression NMS to screen the second feature points from the first feature points;
[0148] The segmentation unit, connected to the feature unit, is used to segment the Q&A image into multiple first Q&A image blocks according to the feature points, specifically: take each second feature point as the center to segment multiple third Q&A image blocks with multiple sizes and / or multiple ratios from the Q&A image, and screen the first Q&A image blocks from the third Q&A image blocks based on the intersection over union IoU algorithm.
[0149] In one embodiment, the screening module 2 specifically includes:
[0150] The text encoding unit is used to obtain the first text encoding t of the user's question by using the text encoder of the contrastive language-image pre-training CLIP model;
[0151] The visual encoding unit is used to obtain the first visual encoding v of each first Q&A image block by using the image encoder of the CLIP model i , where i is the number of the first Q&A image block;
[0152] The similarity calculation unit, connected to the text encoding unit and the visual encoding unit, is used to calculate the first similarity score between each first Q&A image block and the user's question
[0153] The comparison unit, connected to the similarity calculation unit, is used to obtain all s i >τ 3 of the first Q&A image blocks as the second Q&A image blocks whose relevance to the user's question is greater than the first threshold τ 3 ;
[0154] In one embodiment, the overview module 3 is specifically used for:
[0155] Based on the intermediate language model, obtain multiple third overviews of the Q&A image and multiple fourth overviews of each second Q&A image block respectively. Obtain a fifth overview with a relevance greater than a second threshold to the user question in the third overviews and a sixth overview with a relevance greater than a third threshold to the user question in the fourth overviews. Obtain a first overview according to the fifth overview and a second overview according to the sixth overview;
[0156] Among them, the intermediate language model is an Image Caption model of an attention neural network AoANet above the first attention obtained by training in a semi-supervised manner. The semi-supervised manner includes: segmenting multiple training images, obtaining manually annotated annotated overviews for some of the segmented training image blocks, combining the Image Caption model with a large language model to obtain pseudo-annotated overviews for another part of the training image blocks, supervising the training of the Image Caption model based on the annotated overviews, and self-supervising the training of the Image Caption model based on the pseudo-annotated overviews.
[0157] In one embodiment, the overview module 3 specifically includes:
[0158] An intermediate language model unit, configured to obtain an attention neural network AoANet model above the first attention obtained by pre-training;
[0159] A first overview unit, connected to the intermediate language model unit, configured to output multiple third overviews of the Q&A image by using the first AoANet model multiple times, obtain a second similarity score between each third overview and the user question, obtain a fifth overview with a second similarity score greater than a second threshold from the third overviews, and obtain a first overview of the Q&A image according to the fifth overview;
[0160] A second overview unit, connected to the intermediate language model unit, configured to output multiple fourth overviews of each of the several second Q&A image blocks by using the first AoANet model multiple times, obtain a third similarity score between each fourth overview and the first overview and / or the user question, obtain a sixth overview with a third similarity score greater than a third threshold from the fourth overviews, and obtain a second overview of each of the several second Q&A image blocks according to the sixth overview.
[0161] In one embodiment, the device further includes a training module, configured to pre-train a first AoANet model by using a partially labeled data set, specifically including:
[0162] A training data unit, configured to obtain training images;
[0163] A training segmentation unit, connected to the training data unit, configured to segment the training images into multiple training image blocks according to feature points;
[0164] An artificial annotation unit, connected to the training segmentation unit, for obtaining training images and an annotation overview of partial training image patches of each training image;
[0165] A first training unit, connected to the artificial annotation unit, for performing a first training on the second AoANet model using the annotation overview to obtain a third AoANet model;
[0166] A pseudo-annotation unit, connected to the first training unit, for obtaining a first pseudo-annotation overview of the training image patches without annotations by using the third AoANet model and a large language model;
[0167] A weight assignment unit, connected to the pseudo-annotation unit, for assigning different weights to the annotation overview and the first pseudo-annotation overview;
[0168] A second training unit, connected to the weight assignment unit, for performing a second training on the third AoANet model using the annotation overview and the first pseudo-annotation overview with different weights to obtain a first AoANet model.
[0169] In one embodiment, wherein:
[0170] The training images include data images from different scenarios;
[0171] Segmenting the training images includes: extracting the top-k feature points in each training image and segmenting each training image into k training image patches based on the top-k feature points;
[0172] The partial training image patches are n training image patches randomly selected from the k training image patches, where n < k / 3;
[0173] Cross-entropy loss is used as the loss function in the first training and the second training;
[0174] The first pseudo-annotation overview is obtained by using the third AoANet model to output multiple second pseudo-annotation overviews of each training image patch without annotations and inputting the multiple second pseudo-annotation overviews into the large language model for summarization;
[0175] Assigning different weights to the annotation overview and the first pseudo-annotation overview includes: assigning a first weight to the annotation overview and a second weight to the first pseudo-annotation overview, and the first weight is greater than 10 times the second weight.
[0176] In one embodiment, the Q&A module 4 specifically includes:
[0177] A filling information unit, for obtaining the size of the Q&A image and the position coordinates of each of the several second Q&A image patches in the Q&A image, as well as example Q&A pairs regarding the first overview and the second overview;
[0178] A filling unit, connected to the filling information unit, is configured to write the size and the first overview into the first part vacancy of the Prompt template, write each of the position coordinates and the corresponding second overview into the second part vacancy of the Prompt template, and write the example Q&A pairs into the third part vacancy of the Prompt template to obtain the prompt word Prompt;
[0179] A Q&A unit, connected to the filling unit, is configured to input the Prompt into a large language model so that the large language model outputs an answer to the user's question according to the Prompt.
[0180] Embodiment 3:
[0181] Embodiment 3 of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it implements the image Q&A method as described in Embodiment 1 or implements the image Q&A device as described in Embodiment 2.
[0182] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program units, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0183] In addition, the present application can also provide a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the image Q&A method as described in Embodiment 1. The computer device can be the image Q&A device as described in Embodiment 2.
[0184] Wherein, the memory is connected to the processor. The memory can adopt flash memory or read-only memory or other memories, and the processor can adopt a central processing unit or a single-chip microcomputer.
[0185] Embodiments 1-3 of the present application provide an image question answering method, apparatus, and medium. The image is segmented based on feature points, image patches with high relevance to the user's question are screened, and the global information and local information of the image are captured using an intermediate language model according to the overall image and the locally highly relevant image patches, enhancing the overview of the visual image by the intermediate natural language. Finally, a large language model is used to obtain a high-quality answer to the user's question based on the overview.
[0186] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principle of the present application. However, the present application is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present application, and these modifications and improvements are also considered within the protection scope of the present application.
Claims
1. An image question answering method, characterized in that: The method comprises: Segmenting the question-answer image into a plurality of first question-answer image blocks based on the feature points; Acquire a plurality of second question-and-answer image blocks in the first question-and-answer image block whose relevance to the user question is greater than a first threshold; Acquire a first overview of the question and answer image and a second overview of the plurality of second question and answer image blocks based on an intermediate language model; An answer to the user question is obtained according to the first summary and the second summary based on a large language model.
2. The method according to claim 1, characterized in that The question-answer image is segmented into a plurality of first question-answer image blocks based on the feature points, including: Acquiring feature points of the question-answering image, specifically including: extracting first feature points of the question-answering image based on deep learning feature point detection SuperPoint or scale-invariant feature transformation SIFT, and using non-maximum suppression NMS to select second feature points from the first feature points; The question and answer image is segmented into multiple first question and answer image blocks according to the feature points, specifically including: segmenting third question and answer image blocks of multiple sizes and / or multiple ratios from the question and answer image with each second feature point as the center, and selecting the first question and answer image blocks from the third question and answer image blocks based on the intersection over union (IoU) algorithm.
3. The method according to claim 2, characterized in that Acquiring a plurality of second question-and-answer image blocks in the first question-and-answer image block whose relevance to the user question is greater than a first threshold value, specifically comprising: Use the text encoder of the contrastive language-image pre-trained CLIP model to obtain the first text encoding t of the user question; The image encoder of the CLIP model is used to obtain the first visual encoding v of each first question-answer image block. i , i is the number of the first question-answer image block; Calculate the first similarity score between each first question-answer image block and the user question Get all s i >τ3 is used as the second question-and-answer image block whose relevance to the user question is greater than the first threshold τ3.
4. The method according to any one of claims 1 to 3, characterized in that: Acquiring a first overview of the question-answer image and a second overview of the plurality of second question-answer image blocks based on the intermediate language model specifically includes: Based on the intermediate language model, a plurality of third overviews of the question-and-answer image and a plurality of fourth overviews of each second question-and-answer image block are respectively obtained, a fifth overview whose relevance to the user question is greater than a second threshold value among the third overviews and a sixth overview whose relevance to the user question is greater than a third threshold value among the fourth overviews are obtained, a first overview is obtained according to the fifth overview, and a second overview is obtained according to the sixth overview; Among them, the intermediate language model is an image overview Image Caption model obtained through semi-supervised training. The semi-supervised method includes: segmenting multiple training images, and obtaining manually annotated annotation overviews for some training image blocks obtained through segmentation, combining the Image Caption model with the large language model to obtain pseudo-annotated overviews for another part of the training image blocks, training the Image Caption model based on the annotation overview supervision, and training the Image Caption model based on the pseudo-annotated overview self-supervision.
5. The method according to claim 4, characterized in that Based on the intermediate language model, a plurality of third overviews of the question-answer image and a plurality of fourth overviews of each second question-answer image block are respectively obtained, a fifth overview whose relevance to the user question is greater than a second threshold value in the third overviews and a sixth overview whose relevance to the user question is greater than a third threshold value in the fourth overviews are obtained, a first overview is obtained according to the fifth overview, and a second overview is obtained according to the sixth overview, specifically comprising: Get the attention neural network AoANet model based on the first attention obtained in pre-training; Outputting multiple third overviews of the question-answer image using the first AoANet model multiple times, obtaining a second similarity score between each third overview and the user question, obtaining a fifth overview whose second similarity score is greater than a second threshold from the third overviews, and obtaining a first overview of the question-answer image according to the fifth overview; The first AoANet model is used multiple times to output multiple fourth overviews of each of the plurality of second question-answer image blocks, a third similarity score between each fourth overview and the first overview and / or the user question is obtained, a sixth overview having a third similarity score greater than a third threshold is obtained from the fourth overviews, and a second overview of each of the plurality of second question-answer image blocks is obtained based on the sixth overview.
6. The method according to claim 5, characterized in that The method further includes using a partially labeled data set for pre-training to obtain a first AoANet model, specifically comprising: Get training images; Segment the training image into multiple training image blocks according to feature points; Obtaining an overview of the annotations of the training images and partial training image patches of each training image; Performing a first training on the second AoANet model using the labeled overview to obtain a third AoANet model; Obtaining a first pseudo-annotated overview of unannotated training image patches using a third AoANet model and a large language model; Assigning different weights to the annotation overview and the first pseudo-annotation overview; The third AoANet model is trained for a second time using the annotation overviews and the first pseudo-annotation overviews with different weights to obtain the first AoANet model.
7. The method according to claim 6, characterized in that in: The training images include data images from different scenes; Segmenting the training images includes: extracting top-k feature points in each training image, and segmenting each training image into k training image blocks based on the top-k feature points; The partial training image blocks are n training image blocks randomly selected from k training image blocks, where n<k / 3; Cross entropy loss was used as the loss function in the first and second training; The first pseudo-annotated summary is obtained by outputting a plurality of second pseudo-annotated summaries of each unannotated training image block using the third AoANet model, and inputting the plurality of second pseudo-annotated summaries into the large language model for summarization; Assigning different weights to the annotation summary and the first pseudo annotation summary includes: assigning a first weight to the annotation summary and assigning a second weight to the first pseudo annotation summary, wherein the first weight is greater than 10 times the second weight.
8. The method according to any one of claims 1 to 3, characterized in that: Acquiring an answer to the user's question based on the first summary and the second summary based on the large language model specifically includes: Acquire the size of the question-and-answer image, the position coordinates of each of the plurality of second question-and-answer image blocks in the question-and-answer image, and example question-and-answer pairs about the first overview and the second overview; Writing the size and the first overview into a first blank of a Prompt template, writing each of the position coordinates and the corresponding second overview into a second blank of the Prompt template, and writing the example question-answer pair into a third blank of the Prompt template to obtain a prompt word Prompt; The Prompt is input into the large language model so that the large language model outputs the answer to the user's question according to the Prompt.
9. An image question-answering device, characterized in that: The device comprises: A segmentation module, used to segment the question-answer image into a plurality of first question-answer image blocks based on feature points; A screening module, connected to the segmentation module, for obtaining a plurality of second question-answer image blocks in the first question-answer image block whose relevance to the user question is greater than a first threshold; an overview module, connected to the screening module, for obtaining a first overview of the question-answer image and a second overview of the plurality of second question-answer image blocks based on an intermediate language model; The question-answering module is connected to the overview module and is used to obtain an answer to the user's question according to the first overview and the second overview based on a large language model.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image question answering method according to any one of claims 1 to 8 is implemented.