A multi-task learning model combining image-text matching and visual reasoning, a visual common sense reasoning method, and a computer device
Through a multi-task learning model combining graphic and text matching and visual inference, the problem of cross-modal relationship understanding in visual common sense inference tasks is solved, and more efficient joint understanding of visual content and text semantics is achieved, and the model performance is improved.
Patent Information
- Application Number
- CN202210718706.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-23
AI Technical Summary
It is difficult for visual common sense reasoning tasks to fully understand the diverse visual content, semantic rich language expressions, and complex cross-modal relationships in images, resulting in insufficient model performance.
A multi-task learning model with joint graphic and text matching and visual inference is adopted to extract visual and text features through pre-trained models, combine multi-class cross-entropy loss functions and contrast learning loss functions, and jointly optimize training of visual common sense inference and graphic and text matching to achieve two-way promotion of graphic and text matching and visual inference.
The model's joint understanding of diversified visual content and advanced text semantics is improved, the model's perception ability and cross-modal feature alignment ability are enhanced, thereby improving the performance of visual common sense reasoning.
Smart Images

Figure CN114996502B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimedia computing, and specifically relates to a multi-task learning model combining image-text matching and visual reasoning, a visual common sense reasoning method and a computer device. Background Art
[0002] With the rapid growth of multimodal data in social networks, many challenging tasks have been paid attention to and studied in order to effectively analyze heterogeneous modal data. Visual Commonsense Reasoning (VCR) and Image Text Matching are two of these tasks, and are currently hot topics of research at home and abroad. Visual Commonsense Reasoning means that given a question about an image, the visual common sense reasoning model needs to provide not only the correct answer, but also reasonable reasons to prove the answer. Image-text matching means that given an image and a text description, the model needs to calculate whether the data of these two modalities are similar. In recent years, with the development of deep learning technology, visual common sense reasoning tasks and image-text matching models have made fruitful progress. However, the visual common sense reasoning task is still a challenging problem because it requires a comprehensive understanding of the diverse visual content in the image, semantically rich language expressions, and complex cross-modal relationships. Image-text matching tasks have achieved relatively better research results. The present invention hopes to use image-text matching to improve the performance of the visual common sense reasoning model.
[0003] In order to solve the above challenges, current methods resort to the overall attention mechanism or explore Transformer-based models with large-scale pre-training, but few researchers have combined visual reasoning tasks with image-text matching tasks. Since image-text matching also requires comprehensive and fine-grained learning of image features, the present invention believes that image-text matching can promote visual reasoning tasks. Therefore, it is very important for visual common sense reasoning tasks to have a more comprehensive understanding and learning of images and texts to obtain more discriminative feature representations.
[0004] In order to obtain fine-grained information in the visual modality and the language modality, it is proposed to use a multi-level form to perform feature representation in an all-round way; the image-text matching task requires a high degree of alignment between the visual modality and the language modality, and the VCR task also needs to achieve a high degree of alignment between the two modalities to mine deep semantic information, so the present invention proposes a model to achieve a multi-task learning framework that promotes the mutual promotion of image-text matching and visual common sense reasoning. Therefore, the present invention designs a multi-task learning method that combines image-text matching and visual reasoning to improve the feature learning and understanding reasoning ability of the visual common sense reasoning model, thereby improving the overall performance of the model. Summary of the invention
[0005] Regarding the above problems, the present invention mainly focuses on the visual common sense reasoning task. The purpose of the present invention is to use the image-text matching module to enhance the expressiveness of the visual common sense reasoning module. By combining the question and the response into a full-text form and inputting it into the image-text matching module, a joint modeling of higher-level text semantics and complex cross-modal relationships is obtained, so as to learn more discriminative feature representations and obtain a robust and high-performance visual common sense reasoning model. The technical solution for implementing the present invention is as follows:
[0006] A multi-task learning model that combines image-text matching and visual reasoning, which is obtained by the following steps:
[0007] Step S1: Use a pre-trained model to extract the features of the original image and text, and obtain a joint representation of the visual modality and the text modality;
[0008] Step S2: Use a multi-class cross-entropy loss function to optimize and train the visual common sense reasoning;
[0009] Step S3: Process the original dataset of visual common sense reasoning so that it can be used in the image-text matching module;
[0010] Step S4: Extract the pixel-level features of the image as global features, and extract the region-level features of the image as local features;
[0011] Step S5: Use a contrastive learning loss function to optimize and train the image-text matching;
[0012] Step S6: Realize the two-way promotion of image-text matching and visual common sense reasoning through parameter sharing, integrate all the above parts into a unified framework to obtain a multi-task learning model, and conduct overall training of the multi-task learning model.
[0013] All the above processes are described in detail in the specific implementation part.
[0014] The method for the present invention to perform visual common sense reasoning using the above multi-task learning model is as follows:
[0015] For any set of images, questions, and one of the candidate answers, first use the feature extraction methods in steps S1 and S4 to extract the features of the image and text, and obtain their cross-modal joint representation. Then, according to step S2, make the model calculate the probability that the current candidate answer is the correct answer. Then, according to S4, extract the local features and global features of the image, and align the extracted global features and local features locally and globally between the image and the text according to the method in step S5, and calculate the similarity between the image and the text. When the similarity is the largest, the visual common sense reasoning result is obtained.
[0016] Based on the above model and method, the present invention also proposes a computer device, which internally stores the execution instruction code or stored program code of the multi-task learning model that combines image-text matching and visual reasoning, or the execution instruction code or stored program code of the visual common sense reasoning method described above.
[0017] Advantages of the present invention:
[0018] (1) The present invention proposes a multi-task learning model that combines image-text matching and visual reasoning, improving the model's reasoning ability for jointly understanding diverse visual content and advanced text semantics.
[0019] (2) The present invention introduces the image-text matching task into the visual common sense reasoning task, enhancing the model's perception ability and helping the model more effectively align the two modal features.
[0020] (3) The present invention conducts joint training through image-text matching and visual reasoning to promote each other bidirectionally, further improving the performance of visual common sense reasoning. Description of the Drawings
[0021] Figure 1 is a framework diagram of the multi-task learning model of the present invention based on a combination of image-text matching and visual reasoning. The model uses joint learning of visual common sense reasoning and image-text matching to obtain a richer and more comprehensive feature representation. The model is obtained through the following steps: Detailed Embodiments
[0022] The present invention will be further described below with reference to the accompanying drawings.
[0023] Figure 1 is a framework diagram of the multi-task learning model of the present invention that combines image-text matching and visual reasoning. The model uses joint learning of visual common sense reasoning and image-text matching to obtain a richer and more comprehensive feature representation. The model is obtained through the following steps:
[0024] Step S1: Use a pre-trained model to extract the features of the original image and text;
[0025] The step S1 further includes the following steps:
[0026] Step S1-1: For each question in the training data, and its corresponding image and four options, extract its question feature image feature and the four option features where D q , D o , D r represents the dimension of the feature. In the embodiment, the image feature can be extracted by ResNet101 and processed by splicing to obtain a 512-dimensional visual feature (i.e., D o= 512), the problem features and option features can be extracted by BERT and concatenated to obtain 512-dimensional text features.
[0027] Step S1-2: For the text features q (or r) and image features o obtained according to Step S1-1, use the joint encoder f(·; θ) to concatenate the embedding representation of each word in the sentence with its corresponding local image representation, and then transform the concatenated feature representation through a long short-term memory network (LSTM). The output of each unit of the LSTM is pooled to obtain the final joint embedding representation f((o,q); θ), f((o,r); θ), where θ is a parameter during the training process.
[0028] Step S2, as Figure 1 shown, use the multi-class cross-entropy loss function to optimize and train the visual common sense reasoning;
[0029] The said Step S2 further includes the following steps:
[0030] Step S2-1: Send the two joint embedding representations obtained in Step S1-2 into a multi-layer perceptron MLP for probability score calculation, and then normalize the score using the softmax function. Specifically as follows:
[0031]
[0032] Here represents the result of normalization, w o and w q are two mapping matrices that can stabilize the training. The MLP used consists of two fully connected layers.
[0033] Step S2-2: Use a cross-entropy loss function to constrain the visual common sense reasoning based on the fused features and option features. The definition of the loss function is as follows:
[0034]
[0035] Here f(·) is the classification function, y i is the true result of option r i and L 1 is the classification loss function for basic visual common sense reasoning.
[0036] Step S3, as Figure 1 shown in the text form conversion part of, preprocess the original visual common sense reasoning dataset so that it can be used for image-text matching;
[0037] The said Step S3 further includes the following steps:
[0038] Step S3-1: Extract the initial questions and correct response sentences from the visual common sense reasoning dataset file. Connect the questions and correct responses to obtain the "full text" subtitle description denoted as c, and save it in a text file in the form of one line representing one text description, thus constituting the text description required by the image-text matching module.
[0039] Step S3-2: After completing S3-1, since in traditional image-text matching tasks, one image corresponds to five correct descriptions, while in the visual common sense reasoning dataset, some images correspond to two questions and some images correspond to three questions. To achieve a one-to-one correspondence between the images and the text description index numbers, the images are copied so that one image only corresponds to one correct text description. Therefore, it is necessary to extract the index numbers corresponding to the images and the corresponding descriptions respectively from the original visual common sense reasoning module dataset, and store them as the labels of the positive samples in a json file. For the descriptions with the same index number corresponding to the current image as the positive samples, the rest are negative samples.
[0040] Step S4, as Figure 1 shown in the image-text matching module part, extract the pixel-level features of the image as global features, and extract the region-level features of the image as local features;
[0041] The said Step S4 further includes the following steps:
[0042] Step S4-1: In the image-text matching part, first extract the pixel-level features of the image. For the pixel-level features, the present invention adjusts the CNN backbone network to increase the resolution of the input image to 512×512. Process it with two different CNNs: FasterRCNN pre-trained on ImageNet, bottom-up attention mechanism (BUTD) and ResNeXT-101(32×8d) pre-trained on Instagram (WSL), and the dimension of the joint embedding space is 1024. Use the pre-extracted object features as region features (BUTD feature). At the same time, use BiGRU or BERT-base as the text feature extractor to achieve the global alignment of the whole image and the text description. The specific feature calculation formula is as follows:
[0043]
[0044] where x is the image input into the ConvNet network, t is the option or question input into the SeqModel model, and the visual feature set has convolutional local representations, φ nis a spatial pixel-level feature vector from the feature map and object proposals. N represents the number of candidate bounding boxes for extracting image objects; the text feature set represents the contextualized word token feature sequence extracted from the sequence model, where M is the number of words, d 1 and d 2 are the feature dimensions.
[0045] Then, the output visual feature set and the text feature set are aggregated through the visual and text aggregators f visual (·) and f text (·) to further encode the overall vision and text and embed them as follows:
[0046] and
[0047] is the overall feature representation of the image, u is the overall feature representation of the text, and d 3 represents the dimension after mapping to the same embedding space.
[0048] Step S4-2: In the image-text matching part, extract the regional features of the image. Use the Faster R-CNN object detection model to extract the ROI features as local features to achieve local alignment between the key objects in the image and the key words in the text description;
[0049] Step S5, as shown in the image-text matching module part of Figure 1 , optimize and train the image-text matching module through the contrastive learning loss function;
[0050] The said Step S5 further includes the following steps:
[0051] Step S5-1: After extracting the pixel-level features of the image through Step S4-1, except for the description corresponding to the current image as the positive sample, the rest are negative samples. Assume the id of the correct corresponding description is i, then the features of the positive and negative samples are c i and {c 1 , c 2 ,..., c i-1 , c i+1 ,..., c n} in Step S3-1. Based on these features, construct the contrastive learning between the entire image and the entire sentence, model the relationship between different modalities, and enhance the understanding of language semantics. The specific contrastive loss function is as follows:
[0052]
[0053] Here, s(·) is the similarity metric function, τ is the temperature parameter, and τ is 0.2.
[0054] Step S5-2: After extracting the regional features of the image through Step S4-2, except for the description corresponding to the current image being the positive sample, the rest are negative samples. Assuming the id of the correct corresponding description is j, the features of the positive and negative samples are c in Step S3-1 j and {c 1 , c 2 ,..., c j-1 , c j+1 ,..., c n}. Based on these features, a contrastive learning between image regions and words is constructed to model the relationship between different modalities and enhance the understanding of language semantics. The specific contrastive loss function is as follows:
[0055]
[0056] Here, s(·) is the similarity metric function, τ is the temperature parameter, and τ is 0.2.
[0057] Step S6, as shown in the joint training part in Figure 1 , realizes the two-way promotion of image-text matching and visual common sense reasoning through parameter sharing, integrates all the above parts into a unified framework, and conducts the overall training of the multi-task learning model.
[0058] The said Step S6 further includes the following steps:
[0059] The integration of the unified framework results in a multi-task learning model, that is, optimizing the following loss function:
[0060] L = L 1 + λ 1 L g_sim + λ 2 L l_sim
[0061] Here, λ 1 , λ 2 are the balancing parameters, λ 1 = 0.6, λ 2 = 0.4, L 1 is the loss function of visual common sense reasoning, L g_sim is the contrastive loss function between the whole image and the whole sentence, and L l_sim is the contrastive loss function between image regions and words.
[0062] The method for the present invention to perform visual common sense reasoning using the above multi-task learning model is as follows:
[0063] For any set of images and questions, for one of the candidate answers, first, the feature extraction methods in steps s1 and s4 are used to extract the features of the images and text, and obtain their cross-modal joint representations. Then, according to step s2, the model calculates the probability that the current candidate answer is the correct answer. Then, according to s4, the local features and global features of the images are extracted, and the extracted global features and local features are subjected to local alignment and global alignment of the images and text according to the method in step S5, and the similarity between the images and text is calculated. Finally, the visual common sense reasoning result is obtained according to the cross-entropy loss function and the triplet ranking loss function.
[0064] Based on the above model and method, the present invention also proposes a computer device, which internally stores the execution instruction code of the multi-task learning model for joint graphic and text matching and visual reasoning or the stored program code, or the execution instruction code of the visual common sense reasoning method or the stored program code.
[0065] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent manners or changes that do not depart from the technology created by the present invention should be included in the protection scope of the present invention.
Claims
1. A construction method of a multi-task learning model that combines image-text matching and visual reasoning, characterized in that, it includes the following steps: S1: Extract the features of the original image and text, and obtain the joint representation of the visual modality and the text modality; S2: Use the multi-class cross-entropy loss function to optimize and train visual common sense reasoning; S3: Process the original dataset of visual common sense reasoning so that it can be used for image-text matching; S4: Extract the pixel-level features of the image as global features, and extract the region-level features of the image as local features; S5: Use the contrastive learning loss function to optimize and train image-text matching; S6: Realize the two-way promotion of image-text matching and visual common sense reasoning through parameter sharing, and fuse the above processes to obtain a multi-task learning model; The specific implementation of S6 includes: The multi-task learning model is obtained through the following fusion method: L = L 1 + λ 1 L g_sim + λ 2 L l_sim The λ here 1 , λ 2 is the balancing parameter, L 1 is the loss function for visual common sense reasoning, L g_sim is the contrastive loss function between the entire image and the entire sentence, L l_sim is the contrastive loss function between the image region and the word.
2. The construction method of a multi-task learning model that combines image-text matching and visual reasoning according to claim 1, characterized in that, the specific implementation of S1 includes: The step S1 further includes the following steps: S1-1: For each question in the training data, along with its corresponding image and four options, extract its question features Image features and option features Here, D q , D o , D r represents the dimension of the features; the image features are extracted and concatenated through ResNet101 to obtain 512-dimensional visual features, i.e., D o = 512. The question and option features are extracted and concatenated through BERT to obtain 512-dimensional text features; S1-2: Combine the problem feature q and option feature r obtained in step S1-1 i with the image feature o respectively. Use the joint encoder f(·; θ) to connect the embedding representation of each word in the sentence with its corresponding local image representation. Then, transform the concatenated feature representation through a long short-term memory network (LSTM). Pool the output of each unit of the LSTM to obtain the final joint representations f((o, q); θ) and f((o, r); θ).
3. The construction method of a multi-task learning model that combines image-text matching and visual reasoning according to claim 2, characterized in that, the specific implementation of S2 includes: S2-1: Send the joint embedding representation obtained in step S1-2 into a multi-layer perceptron MLP for score calculation, and then normalize the score using the softmax function, specifically as follows: The w here o and w q are two mapping matrices; S2-2: Use the cross-entropy loss function to constrain the visual common sense reasoning based on the fused features and option features, and the loss function is defined as follows: Here, f(·) is the classification function, and y i is the true result of option r i , and L 1 is the classification loss for basic visual common sense reasoning.
4. The construction method of a multi-task learning model that combines image-text matching and visual reasoning according to claim 1, characterized in that, the specific implementation of S3 includes: S3-1: Extract the initial unprocessed questions and correct response sentences from the visual common sense reasoning dataset file, connect the questions and correct responses to obtain a "full text" subtitle description denoted as c, and save it in a text file in the form of one line representing one text description, thus constituting the text description required for image-text matching; S3-2: To achieve a one-to-one correspondence between the required image and text description indexes, copy the image so that one picture only corresponds to one correct text description. Therefore, extract the id number of each image and the id number of each description from the original visual common sense reasoning dataset, and store them as the labels of positive samples in a json file. For the descriptions with the same index number corresponding to the current image as positive samples, the rest are negative samples.
5. The construction method of a multi-task learning model that combines image-text matching and visual reasoning according to claim 4, characterized in that, the specific implementation of S4 includes: S4-1: In the image-text matching part, first extract the pixel-level features of the image. For the pixel-level features, by adjusting the CNN backbone network, the resolution of the input image is increased to 512×512, and it is processed by two different CNNs: FasterRCNN is pre-trained on ImageNet, and uses the bottom-up attention mechanism and ResNeXT-101(32×8d) pre-trained on Instagram. The dimension of the joint embedding space is set to 1024; use the pre-extracted object features as region features; at the same time, use BiGRU or BERT-base as the text feature extractor to achieve the global alignment of the entire image and the text description. The specific feature calculation formula is as follows: Among them, the visual feature set has convolutional local representations, φ n is the spatial pixel-level feature vector from the feature mapping function and the target extraction box; the text feature represents the contextualized word token feature sequence extracted from the sequence model, where M is the number of words, d 1 and d 2 are the feature dimensions; Then the output visual features and text features are aggregated through the visual aggregator f visual (·) and the text aggregator f text (·) to further encode the overall visual and text embeddings as follows: and Step S4-2: In the image-text matching part, then extract the region features of the image; use the object detection model faster RCNN to extract the ROI features as local features to achieve the local alignment of the key objects in the image and the key words in the text description.
6. A method for constructing a multi-task learning model that combines image-text matching and visual reasoning according to claim 5, characterized in that, The specific implementation of S5 includes: S5-1: After extracting the pixel-level features of the image through step S4-1, only the text description that is consistent with the id number of this image is the positive sample, and the rest are negative samples. Assuming that the id of the correct corresponding description is i, the features of the positive and negative samples are c in step S3-1 respectively i and {c 1 , c 2 ,..., c i-1 , c i+1 ,..., c n}. Based on these features, contrastive learning between the entire image and the entire sentence is constructed to model the relationship between different modalities and enhance the understanding of language semantics. The specific contrastive loss function is as follows: Here, s(·) is the similarity metric function, and τ is the temperature parameter; S5-2: After extracting the regional features of the image through step S4-2, except for the description corresponding to the current image being a positive sample, the rest are negative samples. Assuming that the correct corresponding description corresponds to id j, the features of the positive and negative samples are c in step S3-1 respectively. j and {c 1 , c 2 ,..., c j-1 , c j+1 ,..., c n}. Based on these features, construct contrastive learning between image regions and words, model the relationships between different modalities, and enhance the understanding of language semantics. The specific contrastive loss function is as follows: Here, s(·) is the similarity metric function, and τ is the temperature parameter.
7. A visual common sense reasoning method for the model obtained by the method for constructing a multi-task learning model that combines image-text matching and visual reasoning according to any one of claims 1-6, characterized in that, For any set of images and questions, and one of the candidate answers, first use steps s1 and s4 to extract the features of the image and the text, and obtain their cross-modal joint representation. Then, according to step s2, make the model calculate the probability that the current candidate answer is the correct answer. Then, according to s4, extract the local features and global features of the image, and align the extracted global features and local features with the text for local and global alignment according to the method of step S5, and calculate the similarity between the image and the text. The result is obtained when the similarity is the largest.
8. A computer device, characterized in that, The computer device has the execution instruction code of the method for constructing a multi-task learning model that combines image-text matching and visual reasoning according to any one of claims 1-6 or the stored program code, or the execution instruction code of the visual common sense reasoning according to claim 7 or the stored program code.
Citation Information
Patent Citations
Cross-modal image-text matching method and device and computer readable storage medium
CN112905827A
Visual positioning method, device, equipment and medium
CN114511472A