An Image-Text Visual Question Answering Method, System and Storage Medium
Through the multi-level inter-module information fusion network model, the semantic association between images, problems and OCR markers is deeply explored, and the problem of ignoring multi-mode interaction in the prior art is solved, achieving higher visual question-and-answer accuracy in image text.
Patent Information
- Application Number
- CN202111368159.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-18
AI Technical Summary
The prior art ignores different types of interactions between multiple modes in text visual question-and-answer task, and only focuses on specific image areas and OCR marking information guided by input problems, making it difficult to effectively understand and reason about the relationship between text in the image.
A multi-level inter-module information fusion network model is adopted, and by introducing cross-mode and in-mode interaction modules, nine possible relationships between image regions, problems and OCR markers are established, and the cross-mode and in-mode relationships are established using the proportional dot product attention method to deeply explore the semantic relationships between multiple modalities.
Effectively filter out multiple significant target areas related to answers from images, questions and OCR markers, improve the accuracy of visual question-and-answers in image text, and provide a general multimodal information interaction modeling method for multimodal applications.
Smart Images

Figure CN114092707B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and particularly to an image-text visual question answering method, system and storage medium. Background Art
[0002] Visual Question and Answering (VQA) is a complex multi-modal task aiming to automatically answer text questions related to the content of a given image, and it requires understanding both visual images and natural language questions simultaneously. Great progress has been made in the research in this field in recent years, and it has become one of the most active research fields in artificial intelligence. Its development generally has two directions: one is to focus on the relationship between modalities, such as co-attention; the other is to focus on the relationship within modalities, such as the BERT (Bidirectional Encoder Representation from Transformers) model for processing NLP (Natural Language Processing). However, most existing VQA tasks ignore a type of question that involves understanding and reasoning about text in images. Some research has proposed using text visual question answering tasks to solve this problem.
[0003] The text visual question answering task requires understanding the visual scene, question and text in the image simultaneously to infer the answer. Most of its models introduce an Optical Character Recognition (OCR) component to read the text in the image. For example, LoRRA adopts unidirectional attention on image regions and OCR tokens conditioned on the question to infer the answer. The disadvantage of these models is that they consider less the different types of interactions between multiple modalities, and only focus on and learn specific image regions and the OCR token information guided by the input question. Summary of the Invention
[0004] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides an image-text visual question answering method, system and storage medium. The image-text visual question answering method, system and storage medium of the present invention establish cross-modal and intra-modal interaction modules for multiple modalities of the text visual question answering task, and stack intra-modal and inter-modal information fusion modules to construct a multi-layer multi-intra-modal and inter-modal information fusion model. This model can obtain nine possible relationships among image regions, questions and OCR tokens, and finally predict the final answer according to the average value of the interaction features of all modules containing different levels of interaction information.
[0005] To achieve the above object, the present invention adopts the following technical solutions.
[0006] In some embodiments, an image-text visual question answering method is provided, which acquires a target image object and a target question object;
[0007] Extract image visual features from the target image object to obtain image visual features;
[0008] Extract image text features from the target image object to obtain image text features;
[0009] Extract question text features from the target question object to obtain question text features;
[0010] Transform the image visual features, image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension;
[0011] Fuse the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features that encode cross-modal and intra-modal relationships;
[0012] Input the image visual features, image text features, and question text features that encode cross-modal and intra-modal relationships into an answer generation module to obtain a target answer.
[0013] In some embodiments, the step of transforming the image visual features, image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension includes: using a linear transformation layer to transform the image visual features, image text features, and the question text features into the same feature space, where the linear transformation layer is used to input feature representations extracted by different encoders, convert them into the same feature dimension, and output.
[0014] In some embodiments, the step of fusing the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features that encode cross-modal and intra-modal relationships includes: inputting the image visual features, image text features, and question text features of the same dimension into a multi-layer intra-modal and inter-modal information fusion network to obtain image visual features, image text features, and question text features that encode cross-modal and intra-modal relationships; the multi-layer intra-modal and inter-modal information fusion network includes a cross-modal interaction module and an intra-modal interaction module, and the cross-modal interaction module and the intra-modal interaction module form an intra-modal and inter-modal information fusion module; wherein, the cross-modal interaction module is used to obtain the correlation between different modalities; the intra-modal interaction module is used to obtain the relationship between instances within each modality and provide supplementary information for the cross-modal interaction module.
[0015] In some embodiments, inputting the image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships into the answer generation module to obtain the target answer includes:
[0016] According to the output features of each intra-modal and inter-modal information fusion module, the answer generation module generates a predicted score for an answer from the output of each intra-modal and inter-modal information fusion module, and finally selects the candidate answer corresponding to the highest score among the average values of the predicted scores as the target answer.
[0017] In some embodiments, an image-text visual question answering system is further provided, including:
[0018] An interaction module for obtaining a target image object and a target question object and displaying the target answer; a feature extraction module for
[0019] extracting image visual features from the target image object to obtain image visual features;
[0020] extracting image text features from the target image object to obtain image text features;
[0021] extracting question text features from the target question object to obtain question text features;
[0022] An intra-modal and inter-modal information fusion module for transforming the image visual features, image text features, and question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension; and for fusing the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships;
[0023] An answer generation module for obtaining the target answer.
[0024] In some embodiments, a storage medium is further provided, where the storage medium stores computer instructions for causing a computer to execute the image-text visual question answering method as described in any one of the above.
[0025] In some embodiments, an electronic device is further provided, including:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image-text visual question answering method according to any one of the above.
[0029] The image-text visual question answering method provided by the embodiments of the present invention adopts a multi-modal feature extraction module for images and texts, an intra-modal and inter-modal information fusion module, and an answer generation module, deeply mines the semantic associations between multiple modalities, and proposes a multi-level intra-modal and inter-modal information fusion network model. By introducing spatial information as the supervision information for the attention mechanism, it not only mines the semantic connections within each modality, but also models the deep information interaction among the three modalities of images, questions, and OCR tags (texts in images). Thus, it can effectively screen out multiple significant target regions related to the answer from these three modalities to improve the accuracy of image-text visual question answering, and at the same time provides a general multi-modal information interaction modeling method for other multi-modal applications involving three or more modalities. Using the method described in the present invention for the image-text visual question answering task, the steps are simple, the efficiency is high, and the accuracy is high.
[0030] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least as follows: The present invention proposes cross-modal and intra-modal interaction modules for multiple (three or more) modalities, and the scaled dot-product attention (SDA) method is used to establish cross-modal and intra-modal relationships, and it can effectively screen out multiple significant target regions related to the answer from at least these three modalities (the three modalities of images, questions, and OCR tags), improving the accuracy of image-text visual question answering. The present invention constructs a multi-level intra-modal and inter-modal information fusion module for the text visual question answering task by stacking interaction blocks composed of intra-modal and inter-modal information fusion modules, which can model the multi-level interaction between multiple modes. The present invention predicts the final answer in a complementary manner by using the interaction features of all blocks containing different levels of interaction information, and the accuracy is high. The present invention uses the latest text visual question answering dataset for verification and conducts extensive ablation studies on the proposed method. The results show that the method and model performance proposed by the present invention are superior to the existing state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic flowchart of an image-text visual question answering method provided by some embodiments of the present invention.
[0032] Figure 2 is a schematic diagram of an application scenario of an image-text visual question answering method provided by some embodiments of the present invention.
[0033] Figure 3 is a schematic flowchart of an image-text visual question answering method provided by some embodiments of the present invention.
[0034] Figure 4 It is a schematic diagram of an answer generation module and an answer prediction process provided by some embodiments of the present invention.
[0035] Figure 5 It is a schematic diagram of the structure of an image - text visual question - answering system provided by some embodiments of the present invention.
[0036] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing the image - text visual question - answering method of the embodiments of the present invention provided by some embodiments of the present invention.
[0037] Figure 7A and Figure 7B It is a schematic diagram of visual question - answering for text in the figure in the prior art. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0039] The pictures in visual question - answering contain many scenes. For scenes such as stores, roads, and clothes, there are often text messages. When the question is related to this text, traditional visual question - answering methods will not be able to give the correct answer. Figure 7A and Figure 7B shows an example of visual question - answering for text in the figure in the prior art.
[0040] To solve the visual question answering task in this scenario, the present invention first uses an external OCR character recognition system to extract the text information in the image, and then uses the information of the three channels of the text question, the image, and the OCR recognized text to infer the answer. To explore the relationships among the object objects in the image, the semantic words of the question, and the OCR recognized text, the present invention proposes an intra-modal and inter-modal information fusion method for multi-modal data. The present invention also introduces the spatial position relationship between the object objects in the image and the OCR recognized text into the cross-modal interaction process, greatly improving the accuracy of relationship learning, thereby improving the accuracy of answer prediction. To adapt to questions of different complexities, the present invention proposes a multi-modal multi-level intra-modal and inter-modal information fusion method based on the intra-modal and inter-modal information fusion method, which can extract multi-modal fusion features of different levels of interaction. The image visual question answering model with text constructed based on this method mainly includes four parts: 1) Feature extraction: Extracting image object features based on Faster R-CNN, extracting question semantic information based on the LSTM network, and extracting OCR text features based on FastText. 2) Information interaction: Using the multi-modal multi-level intra-modal and inter-modal information fusion method to perform multi-level interaction on the three modal features to obtain multi-modal fusion features of multiple levels. 3) Answer prediction: Counting an answer candidate list from the training dataset, and then identifying the OCR text from the current picture, both of which are used as the answer space, and then using the fusion features obtained by the multi-level information interaction module to predict the answer.
[0041] In some embodiments of the present invention, as Figure 1 shown is a schematic flowchart of an image text visual question answering method provided by some embodiments of the present invention. This method can be executed by an image text visual question answering model and / or an image text visual question answering system. This system can be implemented in software and / or hardware, and is generally integrated in an electronic device. The electronic device can be a computer device or a server device, etc. In some embodiments, the image text visual question answering model includes a multi-level intra-modal and inter-modal information fusion model. Correspondingly, as Figures 1-3 shown, this method includes the following operations:
[0042] S110, obtaining a target image object and a target question object.
[0043] Among them, the target image object can be an image type with visual features, that is, an image. The so-called visual features can also be features that can be directly recognized from the image. Exemplarily, as Figure 2As shown, visual features can be features that can be intuitively obtained, such as the size, outline, and position of an object in an image. The target image object is generally an image type with text, and also includes image text information (Text in images), which is generally in text form. The target question object (Question) is the question configured by the target image object, which is generally in text form. The target image object and the target question object can be manually input by the user, or automatically obtained by the image text visual question answering system. Figure 2 Take this as an example, where the image object is Figure 2 Further, a target question object (Question) may be set for the image object, and the target question object has a target answer (Answer). For example, the target answer (Answer) for the target question object (Question) "what is the number and letter of the plane?" is "N328KF". Of course, one or more questions may be set for an image, and each question has a target answer, which is not limited in the embodiments of the present disclosure.
[0044] S120, extracting image visual features of the target image object to obtain image visual features.
[0045] The step of extracting image visual features from the target image object to obtain image visual features includes: using a Faster R-CNN object detection model to extract region-based visual features of the target image object, and retrieving a bounding box to obtain spatial information.
[0046] In order to infer the answer based on the information in the target image object, the multi-level intra-modal and inter-modal information fusion model first uses existing data sets (such as the ImageNet data set and the Visual Genome data set) to pre-train the Faster R-CNN (Faster Region-based Convolutional Neural Networks) object detection model, and then uses the Faster R-CNN object detection model to extract object-level image visual features. The Faster R-CNN object detection model can extract object features in the target image object. This embodiment uses the Faster R-CNN object detection model to extract the region-based visual features of the target image object and retrieve the bounding box to obtain spatial information.
[0047] In addition, during the training phase of the multi-level intra-module and inter-module information fusion model, only the parameters of the last fully connected layer fc7 of Faster R-CNN are fine-tuned, while other model parameters remain unchanged. The final object-level visual representation is X v ∈R N×2048 , where N = 100 indicates that the top 100 object region features with the highest confidence are selected for each image. In addition, the multi-level intra-module and inter-module information fusion model also extracts the bounding box B v ∈R N×4 corresponding to each object region as spatial information. The image visual feature extraction process can be expressed as:
[0048] (X v ,B v ) = Faster R-CNN(I), (1)
[0049] where I represents the input target image object, and B v i represents the spatial information of the i-th object region, including: the coordinates of the upper left and lower right corners of the bounding box of the object region in the image.
[0050] S130, perform image text feature extraction on the target image object to obtain image text features.
[0051] Among them, the performing image text feature extraction on the target image object to obtain image text features includes: inputting the image into an OCR system, obtaining OCR tags with bounding boxes, and extracting FastText vectors through the FastText model to obtain the representation and position information of the image text.
[0052] To obtain the feature information of the text in the image, the multi-level intra-module and inter-module information fusion model uses an external OCR text recognition system to extract the text in the image. Finally, L OCR tags are obtained for each image, and the value of L depends on the amount of text contained in each image. In addition, the multi-level intra-module and inter-module information fusion model also collects the position coordinates of the bounding boxes of each OCR tag in the image. For the l-th tag (where l ∈ [1,..., L]), the multi-level intra-module and inter-module information fusion model uses a pre-trained FastText model to extract its features containing sub-word information, and finally obtains a 300-dimensional feature vector X o l , to obtain the representation and position information of the image text. The extraction process of the image text features on the target image object can be expressed as:
[0053] (token,B o ) = OCR(I),
[0054] Xo = FastText(token), (2)
[0055] Among them, X o ∈R L×300 and B o ∈R L×4 respectively represent the image text features on the target image object and the coordinate positions of the bounding boxes of each character in the image, and token represents the OCR mark.
[0056] S140. Extract the problem text features from the target problem object to obtain the problem text features.
[0057] Among them, first, the length of the target problem object is aligned by using cropping or padding operations, then each word in the target problem object is encoded into a feature vector sequence through Glove word vectors, and then the sequence information is encoded through an LSTM network to obtain the problem text features.
[0058] To effectively obtain the problem text features, the multi-level intra-module and inter-module information fusion model uses a single-layer LSTM network to encode the problem statement. Specifically, the multi-level intra-module and inter-module information fusion model first aligns the lengths of all problems to M by using cropping or padding operations, and then encodes each word in the problem into a 300-dimensional feature vector through Glove word vectors. In this way, the problem is transformed into a feature vector sequence E = [e 1 ,..., e M . This feature vector sequence is then fed into the LSTM network for sequence information encoding to obtain the problem text features.
[0059] LSTM is an improvement of RNN. It attempts to solve the problems of gradient vanishing and gradient explosion in RNN. It uses a gating mechanism to determine which historical information needs to be retained, improving the long-term memory ability. After LSTM encodes all words, all intermediate outputs are used as the feature of the entire sentence. To effectively train the LSTM network, a residual structure links the word vectors and the sentence features. The formula for the problem text feature extraction process is as follows:
[0060] hidden m = LSTM(E), X q = [hidden m , E], (3)
[0061] Among them, hidden m ∈R (M-300)×1024 represents the output feature of the LSTM network, and X q ∈R M×1024 represents the problem text feature.
[0062] S150, transform the image visual features, image text features, and the problem text features into the same feature space to obtain image visual features, image text features, and problem text features of the same dimension.
[0063] After obtaining the features of three modalities (Text VQA task includes three modalities: image, text, and OCR tokens) in steps S120 - S140, the multi-level intra-modal and inter-modal information fusion model uses a linear transformation layer to transform the features of different dimensions of each modality into the same feature space, that is, transform the image visual features, image text features, and the problem text features into the same feature space using a linear transformation layer. The linear transformation layer is used to input the feature representations extracted by different encoders, convert them into the same feature dimension, and output. Specifically, the features of different dimensions of each modality are transformed into the same feature space as shown in formula (4):
[0064] X v 0 = FC(X v , θ v ),
[0065] X q 0 = FC(X q , θ q ),
[0066] X o 0 = FC(X o , θ o ), (4)
[0067] where X v 0 ∈R N×d , X q 0 ∈R M×d and X o 0 ∈R L×d are features with a dimension of d after the transformation of each modality. X v , X q , X o represent the extracted image visual features, problem text features, and image text features respectively, and θ v , θ q , θ o are the parameters of the corresponding fully connected layer FC.
[0068] S160, fuse the visual features, image text features, and question text features of the images in the same dimension to obtain visual features, image text features, and question text features that encode cross-modal and intra-modal relationships.
[0069] Specifically, input the visual features, image text features, and question text features of the same dimension into a multi-layer intra-modal and inter-modal information fusion network to obtain visual features, image text features, and question text features that encode cross-modal and intra-modal relationships. The multi-layer intra-modal and inter-modal information fusion network includes a cross-modal interaction module and an intra-modal interaction module, and the cross-modal interaction module and the intra-modal interaction module form an intra-modal and inter-modal information fusion module; among them, the cross-modal interaction module is used to obtain the correlation between different modalities; the intra-modal interaction module is used to obtain the relationship between instances within each modality and provide supplementary information for the cross-modal interaction module.
[0070] The multi-level intra-modal and inter-modal information fusion model uses the intra-modal and inter-modal information fusion module to fully model the interaction and fusion between multi-modal features. The intra-modal and inter-modal information fusion module first transmits the features X v 0 , X q 0 and X o 0 to the cross-modal interaction module. The cross-modal interaction module will learn the cross-modal relationships between the three modalities and update the features of the three modalities based on the SDA (or SDAG) mechanism, so that the output features of each modality will contain relevant information of other modalities.
[0071] The following specifically explains how to select the SDA and SDAG mechanisms in the cross-modal interaction module. The SDA mechanism is mainly used in cases where no additional information is required to guide relationship learning. No additional guiding information is needed in the relationship learning between the text modality and the image modality and between the text modality and the OCR token modality. Therefore, the cross-modal interaction module uses the SDA mechanism for relationship learning in their cross-modal interactions, while in the cross-modal interaction between the image modality and the OCR token modality, their spatial relationships are needed to assist in the learning of relevant weights. Therefore, the SDAG mechanism is used to learn the relevant weights.
[0072] Taking the question "What is the letter on the plane’s tail?" as an example, the positional relationship between the plane object and the text object in the image is crucial for correctly answering the question. Therefore, spatial information is introduced to fine-tune the cross-modal relevant weights. For this purpose, the SDAG mechanism uses the bounding box B v of the object in the image and the bounding box B oLearn the spatial relationship as the guiding information G. Specifically, in order to obtain rich spatial location information, the SDAG mechanism calculates the object bounding box B v and the text object bounding box B o 's center positions (C v ∈R N×2 and C o ∈R L×2 ) and sizes (i.e., width S v ∈R N×2 and height S o ∈R L×2 ) as well as their intersection over union (IOU), IOU ∈ R N×L . Then the SDAG mechanism concatenates these spatial information and passes it to a two-layer fully connected neural network with a sigmoid activation function to learn the spatial correlation weights between each object and the OCR token object. The SDAG mechanism applies the spatial correlation weight matrix G to the cross-modal interaction process between the image modality and the OCR tokens. The SDAG mechanism uses the representation of semantic information and spatial relationship information to infer the correlation weights. The introduction of spatial relationship information can reduce the correlation weights between two objects that are far apart or non-intersecting in the image, so as to learn more accurate correlation weights.
[0073] After obtaining the output from the cross-modal interaction module, the intra-modal and inter-modal information fusion module uses the intra-modal interaction module to learn the intra-modal relationships and uses this relationship to update the features of each modality. The learning of all intra-modal relationships is carried out through the SDA mechanism. Similar to the cross-modal interaction module, the intra-modal and inter-modal information fusion module adds a residual structure with layer regularization to the output of the intra-modal interaction module.
[0074] The intra-modal and inter-modal information fusion module is composed of a cross-modal interaction module and an intra-modal interaction module in series, and can perform cross-modal and intra-modal interactions between multi-modal data. To complete more complex interactions (such as the transfer relationship X v →X q →X o ), in some embodiments of the present invention, the multi-layer intra-modal and inter-modal information fusion module stacks multiple units with one intra-modal and inter-modal information fusion module as a unit to obtain high-level semantic relationship information. After performing the intra-modal and inter-modal information fusion T times, the multi-layer intra-modal and inter-modal information fusion module will output the image visual features, question text features, and image text features encoding cross-modal and intra-modal relationships, that is, X v T ∈R N×d , X q T ∈R M×d and Xo T ∈R L×d 。
[0075] In some embodiments, specifically, the cross-modal interaction module is used to capture the correlation between modalities, use SDA or SDAG to capture information related to other modalities with respect to this modality, and transmit this information to update this modality to complete the cross-modal interaction of this modality. When there is no guidance information, SDA is used for modeling, connecting the information flow from other modalities with the original features, and using a fully connected layer to convert the connected features into the output features of this modality. Among them, the SDA mechanism uses three groups of matrices as inputs, assumed to be the query matrix q, the key matrix k, and the value matrix v respectively, where q comes from modality one, and k and v come from the same modality two. The SDA mechanism uses the matrix product qk T to obtain the correlation distribution matrix between the two modalities, and then performs a weighted sum on the value matrix according to the correlation distribution to obtain the correlation information flow IF from modality two to modality one 1←2 . The SDA method is shown in formula (5)
[0076]
[0077] Among them, and respectively contain n q , n k and n v feature vectors, each feature vector having a dimension of d, is the correlation weight matrix. The result of the inner product is proportional to the vector dimension. When calculating the correlation weight matrix M, the inner product will be divided by the square root of the dimension d to normalize the weight value. The non-linear function softmax is applied to each row of the correlation weight matrix M to make the correlation weight values between 0 and 1, and the sum of the weights of each row is 1
[0078] The SDA mechanism of this embodiment can perform relational modeling of multi-modal data, and it learns the correlation weight matrix M through semantic features. Since the information contained in semantic features is relatively one-sided (for example, image visual features only contain the appearance information of visual objects), the learned correlation weights may not be accurate enough. Taking the sample in Figure 2 as an example, if the question is "What is the text on the airplane wing?", it is difficult to accurately learn that the OCR mark on the wing has a greater correlation weight with the airplane only through the appearance information of the airplane and the semantic information of the text marked by OCR, and this correlation weight is largely determined by the spatial position relationship between the airplane and the OCR mark
[0079] In some embodiments of the present invention, spatial relationships are introduced to assist in relevant weight learning, that is, the SDAG mechanism, which is the structure of the improved SDA mechanism. When calculating relevant weights, the SDAG mechanism incorporates external guidance information to assist in relationship learning. The external guidance information can be spatial position relationship information or other forms of information. When there is relevant information, SDAG is used for modeling, and spatial guidance information is utilized to calibrate and learn the relevant weights of regional positions and OCR tags to obtain output features. The SDAG method is shown in formula (6).
[0080]
[0081] Among them, ⊙ is element-wise multiplication, and the matrix G is external guidance information, which can be a spatial relationship matrix or other signals, and can be learned through a neural network module or manually set by humans.
[0082] The overall cross-modal interaction module (Cross-Modal Interaction, CMI) is shown in formula (7).
[0083]
[0084] Where is the original X information and information related to other modalities, features with a dimension of d after transformation of each modality.
[0085] The cross-modal interaction module in some embodiments of the present invention can connect the information flow from other modalities with the original features, and use a fully connected layer to convert the connected features into output features. This output feature contains key information from other features, and this interaction process can be easily extended to cases with more modalities.
[0086] In one embodiment, the cross-modal interaction module calculates the center positions and sizes of the object bounding box and the text object bounding box, as well as their intersection over union, so as to obtain richer spatial information, and applies the guidance information matrix to the interaction between the visual region and the OCR tag features to learn the spatial correlation weights between each object and the OCR tag object.
[0087] In some embodiments, the intra-modal interaction module is used to reveal the relationships between instances within each modality, and provide supplementary and important information for cross-modal interaction. Each feature is converted into query, key, and value features, and the features are input into different SDAs or SDAGs. The original features and information flow of each mode are added, and passed through a linear layer to obtain output features. The intra-modal interaction module of this embodiment uses residual connections to incorporate the information flow into the original features, adds the original features and the information flow, and passes through a linear layer to obtain updated features. This process can also be easily extended to more modes.
[0088] By using the cross-modal interaction module of the embodiments of the present invention, cross-modal relationships can be learned, and relevant information of all other modalities can be transmitted to one modality. For example, for the image modality, Figure 2 the aircraft area in can utilize the cross-modal interaction module to focus on the word "plane" in "Question" (the target question object) and "SpaceShipOne" and "N328KF" in "Text in images" (OCR tags).
[0089] The intra-modal interaction module (IMI) simulates intra-modal relationships (e.g., region-to-region, word-to-word). For example, Figure 2 the "number" and "letter" in "Question" (the target question object) in are the keys to understanding the semantics of the question and should be modeled. The intra-modal relationships within each modality are complementary to the cross-modal relationships.
[0090] In some embodiments of the present invention, the intra-modal and inter-modal interactions between multiple modalities are modeled by a scaled dot product attention model. By combining the cross-modal and intra-modal interaction modules, a module, namely the intra-modal and inter-modal information fusion module, is formed. The intra-modal and inter-modal information fusion module can simulate the complete interaction between multiple modalities (three or more). Based on the designed intra-modal and inter-modal information fusion module, the present invention proposes a multi-level complete interaction method for text VQA. The intra-modal and inter-modal information fusion module learns the potential relationships between image regions, question words, and OCR tags. By stacking multiple layers of intra-modal and inter-modal information fusion modules, their relationships are encoded in a multi-layer manner. In this way, different levels of relationships can be considered more comprehensively.
[0091] The multi-modal and multi-level relevant information flow fusion method proposed in the embodiments of the present invention deeply explores the internal relationships within a single modality and between multiple modalities in multi-modal data, extracts high-level multi-modal semantic association information, and overcomes the adverse effects of the semantic gap. This method uses the intra-modal and inter-modal relevant information flow extraction method to obtain the association relationships between multiple modalities and within a single modality in multi-modal data, and performs feature fusion according to the association relationships. Then, the multi-modal and multi-level relevant information flow fusion method performs multi-level intra-modal and inter-modal relevant information flow fusion to continuously and deeply capture the complex association relationships between multi-modal data in a multi-layer progressive manner, and extracts high-level semantic information. A large number of comparative and ablation experiment results on the TextVQA (Text Visual Question Answering) dataset show that the visual question answering model based on the multi-modal and multi-level relevant information flow fusion method has a 5.42% improvement in prediction accuracy compared to the current best model.
[0092] S170. Input the image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships into the answer generation module to obtain the target answer.
[0093] Generally, the questions in the TextVQA task are related to the text in the image. Therefore, the OCR tags should contain the answers corresponding to the questions. However, due to problems such as detection errors or missed detections in the OCR text recognition system, or questions given by question annotators that are not related to the text in the image, the answers to some questions cannot be found from the OCR tags. To solve this problem, the answer generation module uses a list of answers statistically obtained from the training dataset and the OCR tags recognized in the current image (i.e., the target image object) together as the answer space. Therefore, the answer to the question (i.e., the target question object) can come from the list of answers or the text recognized in the current image. The length of the list of answers in the answer generation module is a. First, in order to infer the correct answer, the answer generation module retains the correspondence between the row vectors of the image text features X o T ∈R L×d and the L OCR tags, and transforms the i-th OCR tag feature X o T,i ∈R d (where i ∈ [1,..., L]) into the prediction score that the i-th OCR tag is the answer through a multi-layer perceptron network Then, the answer generation module aggregates the image visual features and question text features through mean pooling operation, and fuses the image visual features and question text features by element-wise multiplication to obtain the multi-modal fusion features; subsequently, the answer generation module sends the fusion features into a multi-layer perceptron network to obtain the prediction score y voca for each answer in the list of answers. The answer generation module finally selects and y voca and takes the answer corresponding to the maximum prediction score as the predicted answer to the question.
[0094] The T stacked intra-modal and inter-modal information fusion modules in the multi-layer intra-modal and inter-modal information fusion module perform T times of intra-modal and inter-modal information fusion on the multi-modal data, gradually adding relevant information to the features of each modality. Different complexity answers may require different numbers of interactions. In some embodiments of the present invention, to effectively utilize the multi-layer interaction features, the multi-level feature joint prediction (MFJP) method is adopted to use the answer generation module to generate an answer prediction score for the output result of each intra-modal and inter-modal information fusion module. The t-th answer prediction score is denoted as The multi-level feature joint prediction method calculates the average value y of these scoresf Finally, take y f The candidate answer corresponding to the term with the highest score in is used as the final answer, i.e., the target answer. The multi-level feature joint prediction method takes into account the contributions of features at different abstraction levels to the answer.
[0095] Such as Figure 4 As shown, it is a schematic flowchart of the multi-level feature joint prediction method of the present invention for predicting answers, including the following steps:
[0096] (1) The i-th OCR token feature is converted by a classifier into a prediction score for the i-th OCR token
[0097] (2) The image visual features and the question text features are fused through a mean pooling operation
[0098] (3) The fused features are passed through a multi-layer perceptron network to generate a prediction score y voca ;
[0099] (4) Select and y voca The one with the highest score in is used as the score of the predicted answer
[0100] (5) Take the average value y of each score f , and the candidate answer corresponding to the highest score y final is used as the target answer.
[0101] In the embodiment of the present invention, for answer prediction, the output features of each layer are used, rather than only using the last layer to generate multiple candidate answer scores. Finally, the average value of these scores is taken to generate the final answer. The prediction using the multi-layer architecture can utilize multi-layer relationships, rather than just the relationships of the highest-level layer.
[0102] Another embodiment of the present invention also provides an evaluation model, and the latest image-text visual question answering dataset is selected.
[0103] This dataset has 28,408 images, which contain text from the Open Images dataset. Each image in the dataset contains 1-2 questions, and it is necessary to read the text in the image to answer each image, with a total of 45,336 questions. Then, for each question in this dataset, 10 manually proposed answers are collected, and the score of the model is statistically voted through these 10 answers.
[0104] The dataset is divided into three parts: the training dataset, the validation dataset, and the test dataset. The training dataset contains 34,602 questions, the validation dataset contains 5,000 questions, and the test dataset contains 5,734 questions. Since the human labels of the test dataset are not publicly released, the prediction results of the test dataset need to be submitted to a remote evaluation server to obtain the test scores.
[0105] The experimental setup is as follows: The dimensions of the visual features and word features in the questions are 2,048 and 1,024 respectively. N = 100 object regions from Faster R-CNN are used. The question length is fixed to M = 14 by truncation or padding. The OCR tokens are encoded as FastText features with a dimension of 300. The OCR tokens with a fixed length of L = 50 are used by discarding the extra tokens or adding zero vectors. The answers that appear more than 8 times in the training dataset are retained as answer vocabulary, resulting in a = 843 candidate answers. Among them, the hidden feature dimension is set to d = 512. All fully connected layers use the same dropout rate of 0.25. All models are trained using the Adamax optimizer, with a batch size of 64, gradient clipping of 0.25, and the learning rate set to 1.5 - 3, implemented using PyTorch. All ablation studies are conducted on the validation dataset, and the training dataset and the validation dataset are combined and tested on the test dataset without any additional datasets.
[0106] The present invention is compared with two state-of-the-art baseline models (variants of LoRRA), and the results show that the present invention is significantly superior to the other two models. Specifically, the performance on the validation dataset and the test dataset is 6.58 and 5.42 percentage points better than the baseline respectively.
[0107] Figure 5 It is a structural diagram of an image-text visual question answering system provided by an embodiment of the present invention. This embodiment is applicable to the situation of using a visual question answering model to process visual question answering tasks including image-text types. The device is implemented by software and / or hardware and is specifically configured in an electronic device. The electronic device can be a computer device or a server device, etc.
[0108] An image-text visual question answering system 1000, characterized by comprising:
[0109] A human-computer interaction module 100, configured to obtain a target image object and a target question object, and display a target answer;
[0110] A feature extraction module 200, configured to
[0111] Extract image visual features from the target image object to obtain image visual features;
[0112] Extract image text features from the target image object to obtain image text features;
[0113] Extract question text features from the target question object to obtain question text features;
[0114] The intra - and inter - module information fusion module 300 is configured to transform the image visual features, the image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension; and fuse the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features encoding cross - modal and intra - modal relationships;
[0115] The answer generation module 400 is configured to obtain the target answer.
[0116] The image - text visual question - answering system in this embodiment can execute the image - text visual question - answering method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the image - text visual question - answering method provided in any embodiment of the present disclosure.
[0117] In an embodiment of the present invention, an electronic device and a storage medium are further provided.
[0118] The storage medium provided in some embodiments of the present invention stores computer instructions, and the computer instructions are used to cause a computer to execute the image - text visual question - answering method as described in any of the above embodiments.
[0119] Some embodiments of the present invention further provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image - text visual question - answering method as described in any of the above embodiments.
[0120] Figure 6 The block diagram of an electronic device 600 that can be used to implement the embodiments of the present invention is shown. The electronic device 600 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed in the present invention.
[0121] The electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0122] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0123] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the image-text visual question answering method. For example, in some embodiments, the image-text visual question answering method can be implemented as a computer software program, which is tangibly contained in a machine-readable storage medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the above-described image-text visual question answering method can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the image-text visual question answering method in any other appropriate manner (e.g., by means of firmware).
[0124] The various embodiments of the systems and techniques described above in this invention can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0126] In the context of the present disclosure, a storage medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The storage medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127] To provide interaction with a user, the above-described image-text visual question answering system and image-text visual question answering method can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0128] The image-text visual question answering system and image-text visual question answering method described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with the implementations of the systems and technologies described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0129] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system, or a server combined with blockchain.
[0130] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An image-text visual question answering method, characterized in that, the method includes: obtaining a target image object and a target question object; extracting image visual features from the target image object to obtain image visual features; extracting image text features from the target image object to obtain image text features; extracting question text features from the target question object to obtain question text features; transforming the image visual features, image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension; fusing the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships; inputting the image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships into an answer generation module to obtain a target answer; the fusing the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships includes: inputting the image visual features, image text features, and question text features of the same dimension into a multi-layer intra-modal and inter-modal information fusion network to obtain image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships; the multi-layer intra-modal and inter-modal information fusion network includes a cross-modal interaction module and an intra-modal interaction module, and the cross-modal interaction module and the intra-modal interaction module form an intra-modal and inter-modal information fusion module; wherein, the cross-modal interaction module is used to obtain the correlation between different modalities; the intra-modal interaction module is used to obtain the relationship between instances within each modality and provide supplementary information for the cross-modal interaction module; The intra-modal and inter-modal information fusion module first transfers the features of multiple modalities to the cross-modal interaction module, and the cross-modal interaction module learns the cross-modal relationships between the three modalities and updates the features of the three modalities based on the SDA or SDAG mechanism, so that the output features of each modality contain the relevant information of other modalities; The cross-modal interaction module uses the SDA mechanism for relationship learning in the relationship learning between the text modality and the image modality and between the text modality and the OCR token modality, while using the SDAG mechanism to learn the relevant weights in the cross-modal interaction between the image modality and the OCR token modality; Using the SDAG mechanism to calculate the center position and size of the object object bounding box and the text object bounding box, as well as the intersection over union between the object object bounding box and the text object bounding box, so as to obtain richer spatial information to generate a guidance information matrix, and applying the guidance information matrix to the interaction between the visual region and the OCR token features to learn the spatial correlation weights between each object object and the OCR token object; the inputting the image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships into an answer generation module to obtain a target answer includes: The multi-layer feature joint prediction method uses the answer generation module to generate an answer prediction score for the output results of each layer of the in-mold and inter-mold information fusion module; the t-th answer prediction score is expressed as The multi-layer feature joint prediction method calculates the average value y of these scores f , and finally takes y f The candidate answer corresponding to the item with the highest score in is used as the final answer, that is, the target answer; The multi-level feature joint prediction method takes into account the contributions of features at different abstraction levels to the answer; The multi-level feature joint prediction method predicts the answer, including the following steps: The i-th OCR token feature is converted by a classifier into a prediction score regarding the i-th OCR token Fuse image visual features through mean pooling operation and question text features Integrate the above two features through element-wise multiplication to obtain fused features; Generate the predicted score y from the fused features through a multi-layer perceptron network voca ; Select and y voca The one with the highest score among them is used as the score of the predicted answer Take each score and calculate the average value y f , and use the candidate answer corresponding to the highest score y final as the target answer.
2. According to the image-text visual question answering method described in claim 1, wherein, The extracting of the image visual features of the target image object to obtain image visual features includes: using the Faster R-CNN object detection model to extract the region-based visual features of the target image object, and retrieving the bounding box to obtain spatial information.
3. According to the image-text visual question answering method described in claim 1, wherein, The extracting of the image text features of the target image object to obtain image text features includes: inputting the image into an OCR system, obtaining OCR tags with bounding boxes, and extracting FastText vectors to obtain the representation and position information of the image text.
4. According to the image-text visual question answering method described in claim 1, wherein, The extracting of the question text features of the target question object to obtain question text features includes: aligning the length of the target question object by using cropping or padding operations, then encoding each word in the target question object into a sequence of feature vectors through Glove word vectors, and then encoding the sequence information through an LSTM network, thereby obtaining the question text features.
5. According to the image-text visual question answering method described in claim 1, wherein, The transforming of the image visual features, image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension includes: using a linear transformation layer to transform the image visual features, image text features, and the question text features into the same feature space, and the linear transformation layer is used to input the feature representations extracted by different encoders, convert them into the same feature dimension and output.
6. An image-text visual question answering system, wherein, including: An interaction module, used to obtain the target image object and the target question object, and display the target answer; A feature extraction module, used to extract the image visual features of the target image object to obtain image visual features; extract the image text features of the target image object to obtain image text features; extract the question text features of the target question object to obtain question text features; An intra-modal and inter-modal information fusion module, used to transform the image visual features, image text features, and the question text features into the same feature space to obtain image visual features, image text features, and question text features of the same dimension; and fuse the image visual features, image text features, and question text features of the same dimension to obtain image visual features, image text features, and question text features encoding cross-modal and intra-modal relationships; An answer generation module, used to obtain the target answer; Fusing the visual features, image text features, and question text features of the same dimension to obtain visual features, image text features, and question text features encoding cross-modal and intra-modal relationships, including: inputting the visual features, image text features, and question text features of the same dimension into a multi-layer intra-modal and inter-modal information fusion network to obtain visual features, image text features, and question text features encoding cross-modal and intra-modal relationships; the multi-layer intra-modal and inter-modal information fusion network includes a cross-modal interaction module and an intra-modal interaction module, and the cross-modal interaction module and the intra-modal interaction module form an intra-modal and inter-modal information fusion module; wherein, the cross-modal interaction module is used to obtain the correlation between different modalities; the intra-modal interaction module is used to obtain the relationship between instances within each modality and provide supplementary information for the cross-modal interaction module; The intra-modal and inter-modal information fusion module first transfers the features of multiple modalities to the cross-modal interaction module, and the cross-modal interaction module learns the cross-modal relationships among the three modalities and updates the features of the three modalities based on the SDA or SDAG mechanism, so that the output features of each modality contain the relevant information of other modalities; The cross-modal interaction module uses the SDA mechanism for relationship learning in the relationship learning between the text modality and the image modality and between the text modality and the OCR token modality, while using the SDAG mechanism to learn the relevant weights in the cross-modal interaction between the image modality and the OCR token modality; Using the SDAG mechanism to calculate the center position and size of the object bounding box and the text bounding box, as well as the intersection over union between the object bounding box and the text bounding box, so as to obtain richer spatial information, generate a guidance information matrix, and apply the guidance information matrix to the interaction between the visual region and the OCR token features to learn the spatial correlation weights between each object and OCR token pair; The answer generation module is further configured to use a multi-layer feature joint prediction method to generate an answer prediction score for the output result of each layer of the in-module and inter-module information fusion module; the t-th answer prediction score is expressed as The multi-layer feature joint prediction method calculates the average value y of these scores f , and finally takes y f The candidate answer corresponding to the term with the highest score in is used as the final answer, that is, the target answer; The multi-layer feature joint prediction method considers the contribution of features at different abstraction levels to the answer; The multi-layer feature joint prediction method predicts the answer, including the following steps: The i-th OCR marker feature is converted by a classifier into a prediction score for the i-th OCR marker Fusing image visual features through mean pooling operation and question text features Integrating the above two features through element-wise multiplication to obtain fused features; Generate a predicted score y from the fused features through a multi-layer perceptron network voca ; Select and y voca The one with the highest score among them is used as the score of the predicted answer Take each score to calculate the average value y f , and select the candidate answer corresponding to the highest score y final as the target answer.
7. A storage medium, characterized in that, the storage medium stores computer instructions for causing a computer to execute the image text visual question answering method according to any one of claims 1-5.
8. An electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image text visual question answering method according to any one of claims 1-5.