Semantic disambiguation-based image-text retrieval method, storage medium and device
By using Faster-RCNN, Bi-GRU, Bert and GAT networks to process ambiguity words in the graphic search model, identifying the context information and semantic annotations of ambiguity words, the comprehension error caused by the context differences between ambiguity words in the prior art is solved, and the accuracy and performance of graphic search is improved.
Patent Information
- Application Number
- CN202510282112.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-08
AI Technical Summary
Existing graphic search techniques are difficult to accurately deal with the semantic differences of ambiguity in different contexts, resulting in the model being unable to accurately understand the meaning of text and pictures, especially when the text data is short and the semantic information is limited, the meaning of words and even the entire sentence is misunderstood.
The pre-trained Faster-RCNN model is used to extract image area features, the Bi-GRU and Bert models encode text data, use the GAT network to update the ambiguity representation, and calculate the meaning similarity between ambiguity words through the relevant layers, generate weight distribution, embed new ambiguity word representations, align with image area features, and adjust model parameters through loss function to achieve accurate recognition and matching of ambiguity words.
By identifying the context information of ambiguity and semantic annotations of knowledge sources, the accuracy and performance of graphic and text retrieval is improved, the lack of training data and data distribution dependence is compensated, better ambiguity perception fusion is achieved, and retrieval performance is improved.
Smart Images

Figure CN120277230A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and artificial intelligence, and particularly relates to a method, a storage medium and a device for image-text retrieval based on semantic disambiguation. Background Art
[0002] With the rapid development of the self-media era, multimodal data such as images, texts, and videos have shown explosive growth. The diverse content forms help humans perceive and understand the surrounding world more comprehensively, because people can easily achieve the alignment and complementarity of different forms of information, thereby obtaining knowledge more accurately and systematically. In the field of artificial intelligence cross-modal research, its core goal is to simulate the functions of the human brain and achieve semantic alignment and complementarity between different forms of information. As a basic task of cross-modal understanding, cross-modal retrieval aims to use one type of data as a query condition to retrieve another type of data. Among them, cross-modal retrieval between images and texts is the most common research direction.
[0003] With the rise of deep learning, researchers have widely applied its ideas to various tasks in natural language processing (NLP), and the field of image-text retrieval is no exception. Since the beginning of the 21st century, the modeling methods of deep learning have gradually become the mainstream for solving image-text retrieval problems. Because deep learning can automatically learn and extract features from raw data, it significantly reduces the cost of manually defining features. Most existing image-text retrieval studies mainly focus on joint space modeling or exploring the correspondence between image and text segments. In recent years, the research focus has shifted to using attention mechanisms to extract features of different modal data and calculate their similarities. Image-text matching methods can be divided into two categories: global alignment-based and local alignment-based. Global alignment-based methods map the entire image and the whole sentence to the same joint semantic space, while local alignment-based methods infer the global image-text similarity by aligning visual objects with text words, making the matching method more fine-grained and interpretable. Different from previous methods, this study proposes a novel cross-modal image-text retrieval method that explicitly integrates disambiguated word knowledge into the text and pictures of the matching model to more accurately understand the meanings of the text and pictures, thereby improving the matching performance.
[0004] Word sense disambiguation aims to identify the specific meaning of a word in a particular context. Current deep learning-based image-text retrieval studies usually use a unified word vector to represent the occurrences of a word at different positions in the text, which may cause the model to fail to accurately capture the semantic differences of disambiguated words in different contexts, thereby misinterpreting the meaning of the word or even the entire sentence. Especially in image-text retrieval tasks, the text data is usually relatively short and the semantic information is relatively limited, and the misrepresentation of disambiguated words will have a more significant impact on the overall understanding of the text. Summary of the Invention
[0005] To solve the problems existing in the above prior art, the present invention proposes a text-image retrieval method based on semantic disambiguation, which includes: obtaining the text words to be retrieved; inputting the text words to be retrieved into the trained text-image retrieval model to obtain a retrieval result;
[0006] Training the text-image retrieval model includes: obtaining an original data set, where the data in the original data set includes text data and image data; using a pre-trained Faster-RCNN model to extract the regional features of the image data; using a Bi-GRU model and a Bert model to encode the text data; using a GAT network to update the ambiguous words in the encoded text; using a relevant layer to calculate the semantic similarity between each updated ambiguous word, and generating a weight distribution of perceived meaning according to the semantic similarity; embedding each ambiguous word into the original text according to the weight distribution to obtain a new word representation of the ambiguous word; aligning the new word representation of the ambiguous word with the image regional features; matching the aligned words and images; calculating the loss function of the model according to the matching result, adjusting the parameters of the model, and completing the training of the model when the loss function converges.
[0007] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, any one of the above text-image retrieval methods based on semantic disambiguation is implemented.
[0008] To achieve the above object, the present invention also provides a text-image retrieval device based on semantic disambiguation, including a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the text-image retrieval device based on semantic disambiguation executes any one of the above text-image retrieval methods based on semantic disambiguation.
[0009] The beneficial effects of the present invention:
[0010] The present invention proposes a novel adaptive text disambiguation model based on knowledge sources. The model detects ambiguous words in the text with the help of semantic information of ambiguous words listed in public knowledge sources (such as WordNet), and identifies the precise semantics of each ambiguous word in the current context through the context information of each detected ambiguous word and the semantic annotation provided in the knowledge source. The invention can make up for the problems of lack of training data and strong dependence on data distribution in existing disambiguation algorithms. The present invention proposes a novel semantic joint representation space for ambiguous words and its semantic vector representation. The present invention constructs a joint feature space composed of ambiguous word semantic vectors and original text vectors. The present invention proposes a new multi-task cross-modal image-text retrieval model that integrates weakly supervised ambiguity perception. The fusion strategy can make full use of the complementary characteristics of disambiguation and retrieval tasks, so that ambiguity perception can be better integrated into the retrieval task, thereby improving retrieval performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a flow chart of a model module in an embodiment of the present invention;
[0012] Figure 2 This is a result diagram of a comparative experiment with other image-text retrieval algorithms after integrating text-ambiguous word perception in an embodiment of the present invention;
[0013] Figure 3 This is a result diagram of a comparative experiment with other image-text retrieval algorithms after integrating text-image-ambiguous word perception in an embodiment of the present invention;
[0014] Figure 4 is an ablation experiment diagram in an embodiment of the present invention;
[0015] Figure 5 This is a visual case study diagram for verifying the effectiveness of the experiment in the embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] A text-image retrieval method based on semantic disambiguation, the method comprising: obtaining a text word to be retrieved; inputting the text word to be retrieved into a trained text-image retrieval model to obtain a retrieval result; training the text-image retrieval model comprising: obtaining an original data set, the data in the original data set including text data and image data; extracting regional features of the image data using a pre-trained Faster-RCNN model; encoding the text data using a Bi-GRU model and a Bert model; updating disambiguating words in the encoded text using a GAT network; calculating the semantic similarity between each updated disambiguating word using a relevant layer, and generating a weight distribution of perceptual meaning according to the semantic similarity; embedding each disambiguating word into the original text according to the weight distribution to obtain a new representation of the disambiguating word; aligning the new representation of the disambiguating word with the image regional features; matching the aligned words and images; calculating a loss function of the model according to the matching result, adjusting the parameters of the model, and completing the training of the model when the loss function converges.
[0018] In this embodiment, the text-image retrieval method based on semantic disambiguation comprises the following steps:
[0019] Step 1, distillation process. Pre-train the semantic perception of text disambiguating words to obtain true labels of disambiguating words.
[0020] In this embodiment, nouns, verbs, adjectives, and adverbs listed in WordNet with more than one meaning are identified as disambiguating words, and these words may exhibit different meanings in different contexts. For each detected disambiguating word W i in the adopted data set, the present invention extracts four possible meanings {S i,1 , S i,2 , S i,3 , S i,4} and their corresponding annotations {g i,1 , …, g i,4} from WordNet, where the annotations are descriptive sentences.
[0021] The present invention mainly focuses on transfer learning, so existing and easily obtainable pre-trained models are used. In the process of optimizing the word sense disambiguation model and the text-image matching task, the pre-trained word sense disambiguation model is used as a teacher model to guide the word sense disambiguation distillation process of the present invention. And the semantic perception of the disambiguating word obtained under the pre-trained teacher model is used as the true label y ij of the disambiguating word in the present invention.
[0022] Step 2, initial representations of image I and text T. The present invention uses a pre-trained Faster-RCNN model to extract image regional features denoted as where r iDenote the features of the i-th region, and N = 36 represents the total number of image feature regions. The present invention uses Bi-GRU and Bert as the encoders of the text T.
[0023] Step 3: Representation of text ambiguous words. Use GAT to update the representation of ambiguous words. In the l-th iteration of training, update the sense-aware representation of the ambiguous word from to The calculation formula is:
[0024]
[0025] where, N(g i,k ) represents the set of words in g i,k . Then, using the same mechanism, based on the context words in the original text, update the ambiguous word as follows. The calculation formula is:
[0026]
[0027] where, is the sense embedding of the ambiguous word updated by GAT, is the word extracted from the original text in the (l - 1)-th layer of the GAT network, ω i is the i-th word in the word set, g i,k is the four sense interpretation functions corresponding to the ambiguous word, and N(ω i ) is the adjacent words of ω i .
[0028] Step 4: Calculation of the text-related layer. Calculate the similarity between the i-th ambiguous word and each sense of the ambiguous word through the inner product with the related layer. The weight distribution of all perceived senses is normalized by the Softmax function. The calculation formula is:
[0029]
[0030] q i,k = Softmax(p i,k ) = [q i,1 , …, q i,k
[0031] where, represents the sense embedding of the ambiguous word updated by GAT, and w i represents the representation of the ambiguous word updated by GAT. Then update w i to the sense-aware embedding of the ambiguous word with the highest predicted probability for subsequent update of the ambiguous word embedding.
[0032] Step 5: Integration of perceived meaning of ambiguous words. For each ambiguous word involved in the text, the original text embedding and the updated embedding of the ambiguous word are integrated to obtain a new semantic representation of the ambiguous word. This semantic representation integrates the original text semantic information and the semantic information of the ambiguous word. The calculation formula is:
[0033]
[0034] Step 6: Align ambiguous words with images and calculate image correlation layer.
[0035] In order to align ambiguous words with images. The invention uses GAT to obtain the features of ambiguous words and uses the image features obtained by Faster-RCNN for alignment operations. Specifically, the ambiguous word interpretation in each image-text pair is locally and globally aligned with the image features, and the local similarity probability distribution is obtained by calculating the similarity between each ambiguous word in the text and the local features of the image. Then, the KL loss function is used to close the local similarity probability distribution and the ambiguous word label (obtained by distillation of the pre-trained model). The similarity probability distribution is calculated as follows:
[0036]
[0037]
[0038] in, is the meaning embedding of the ambiguous word updated by GAT, is the local feature of the image, There are four explanations corresponding to ambiguous words. is the normalized similarity value, r i,k Similarity values calculated for local features and ambiguous words.
[0039] The KL loss function is:
[0040]
[0041] Among them, N represents the number of local features in the image-text pair, M represents the number of ambiguous words in each image-text pair, and y ij is the true label of the ambiguous word, is the predicted probability distribution of the ambiguous words output in the relevant layer.
[0042] In order to ensure that local alignment does not lose global information, the global features are obtained by fusing the image features through cat, and then the global similarity probability distribution is calculated by the global features and the ambiguous word interpretation in the text. The KL loss function is used to bring the global probability distribution and the ambiguous word label (obtained by distillation of the pre-trained model) closer. The calculation formula is as follows:
[0043]
[0044] Step 7, Matching of images and text. Calculate the cosine similarity of the images and text obtained in the above steps to measure the matching degree between them.
[0045] Step 8, Calculation of the loss function. Use three loss functions to train our sense perception model, namely the loss function for image-text matching, the loss function for word sense disambiguation, and the KL alignment loss function.
[0046] Step 8.1, Image-text matching loss function. Adopt Triplet Loss to measure the similarity between images and text in the embedding space learned by the model. The core idea is to train the model by comparing the similarity among three samples, which usually include a query sample (image or text) and two candidate samples (positive sample and negative sample). The goal of Triplet Loss is to make the distance between the positive sample and the anchor as small as possible, while the distance between the negative sample and the anchor as large as possible to promote correct matching. By minimizing the Triplet Loss, the model can learn better image-text embedding representations, thereby improving the accuracy and performance of image-text retrieval. The calculation formula of Triplet Loss is:
[0047]
[0048] where α is the margin parameter, S(I,T) is the similarity score function, I is the image, T is the text, is the text in the negative sample, is the image in the negative sample.
[0049] Step 8.2, Calculation of the disambiguation loss function. Adopt Cross-Entropy Loss as the calculation formula for the disambiguation loss of ambiguous words, and calculate the loss by comparing the difference between the predicted probability of each category by the model and the actual label. By minimizing the cross-entropy loss, the model can better learn the correct classification boundary and improve the accuracy of the model to find the correct meaning of ambiguous words. The calculation formula is:
[0050]
[0051] where N represents the number of image-text pairs, M represents the number of ambiguous words in each image-text pair. The supervision signal of WSD comes from the pre-trained WSD, y ij represents the true label of the ambiguous word, q ij represents the predicted probability distribution of the ambiguous word output in the relevant layer.
[0052] Step 8.3, Calculation of the overall loss function. In multi-task training, the calculation formula of the overall loss function is:
[0053] L = λL WSD + L match + L KL
[0054] where λ is a trade-off parameter.
[0055] A specific implementation of a graphic and text retrieval method based on semantic disambiguation, as Figure 1 shown, includes the following steps:
[0056] Step 1: Pretraining of the true labels of ambiguous words (i.e., the distillation process);
[0057] Step 2: Embedding representation of the original graphic and text;
[0058] Step 3: Using GAT to update the representations of the text and ambiguous words;
[0059] Step 4: The text-related layer predicts the probability distribution of the meanings of ambiguous words and updates the representations of ambiguous words;
[0060] Step 5: The image-related layer calculates the image and the probability distribution of the meanings of ambiguous words;
[0061] Step 5: Integrate the text-aware embedding and the original text embedding;
[0062] Step 6: Calculate the similarity score between the picture and the text for matching;
[0063] Step 7: Use three loss functions to supervise the training.
[0064] Implementation settings: Based on the above steps, the present invention first selects two graphic and text data sets for implementation, namely Flickr 30k and MS-COCO. Both the Flickr 30k data set and the MS-COCO data set are commonly used data sets for graphic and text retrieval tasks, which contain rich images, and each image is accompanied by a corresponding text description. The evaluation metrics of the present invention include image-to-text retrieval and text-to-image retrieval. In addition, we set the embedding dimensions of the image and the text to 2048 dimensions, the batch size to 64, and use the Adam optimizer.
[0065] Comparative experiment implementation: To verify that the present invention is superior to other similar inventions, the present invention performs graphic and text retrieval tasks on two data sets, Flickr30k and MS-COCO, respectively. The implementation results show that the present invention is superior to other similar inventions, as Figure 2 , Figure 3 shown.
[0066] Ablation experiment implementation: Next, an ablation experiment is carried out to check the influence of different components in the invention on the invention effect. As Figure 4As shown, compared with the baseline, the sense perception of the ambiguous words predicted by directly adding the pre-trained model slightly improves the performance of image-text retrieval; adding only the word sense disambiguation module in the present invention further improves the performance of image-text retrieval; adding both the word sense disambiguation module and the disambiguation loss function in the present invention achieves a greater improvement in the performance of image-text retrieval. This implementation proves that both the word sense disambiguation module and the disambiguation loss function of the present invention have a positive effect on image-text retrieval.
[0067] The present invention performs the process of text-to-image retrieval on the Flickr30k dataset. As Figure 5 , for the first sentence, the ambiguity-aware model disambiguates the word "apple", enhancing its meaning as a fruit while weakening its meaning as an electronic brand. This helps the matching model accurately retrieve the correct image. In the second sentence, the main ambiguous word is "nail". The ambiguity-aware model introduces disambiguation during the matching process, emphasizing the meaning of "nail" as a metal fastener and minimizing its meaning as a fingernail or other interpretations, thus improving the retrieval performance.
[0068] In an embodiment of the present invention, the present invention further includes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above-mentioned image-text retrieval methods based on semantic disambiguation.
[0069] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk, or optical disc and other various media that can store program codes.
[0070] An image-text retrieval device based on semantic disambiguation includes a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the image-text retrieval device based on semantic disambiguation executes any one of the above-mentioned image-text retrieval methods based on semantic disambiguation.
[0071] Specifically, the memory includes: ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc and other various media that can store program codes.
[0072] Preferably, the processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0073] The above-mentioned embodiments have further elaborated on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A graphic and text retrieval method based on semantic disambiguation, characterized in that include: Get the text words to be searched; Input the text words to be retrieved into the trained image-text retrieval model to obtain the retrieval results; Training the image-text retrieval model includes: obtaining the original data set, which includes text data and image data; using the pre-trained Faster-RCNN model to extract the regional features of the image data; using the Bi-GRU model and the Bert model to encode the text data; using the GAT network to update the ambiguous words in the encoded text; using the relevant layer to calculate the meaning similarity between each updated ambiguous word, and generating the weight distribution of the perceived meaning according to the meaning similarity; embedding each ambiguous word into the original text according to the weight distribution to obtain a new ambiguous word representation; aligning the new ambiguous word representation with the image regional features; Match the aligned words and images; calculate the model's loss function based on the matching results, adjust the model's parameters, and complete the model training when the loss function converges.
2. The method for graphic and text retrieval based on semantic disambiguation according to claim 1, wherein Using the GAT network to update the ambiguous words in the encoded text includes: Among them, is the meaning embedding of the ambiguous word updated by GAT, is the word extracted from the original text in the (l - 1)-th layer of the GAT network, ω i is the i-th word in the word set, N(g i,k ) is the word set in g i,k , g i,k is the four meaning interpretation functions corresponding to the ambiguous word, N(ω i ) is the adjacent word of ω i .
3. A method for image-text retrieval based on semantic disambiguation according to claim 1, characterized in that, Calculating the updated meaning similarity between each ambiguous word using the relevant layer includes: calculating the similarity between the i-th ambiguous word and each meaning of the ambiguous word through the inner product with the relevant layer, and normalizing the similarity using the Softmax function.
4. A method for image and text retrieval based on semantic disambiguation according to claim 1, characterized in that Embedding each ambiguous word into the original text according to the weight distribution includes: obtaining the original word vector w0 and the weighted meaning vector w of the weight distribution i , adding the original word vector w0 and the weighted meaning vector w of the weight distribution i , and using the sum as the final vector representation of the ambiguous word; embedding the final vector representation of the ambiguous word into the original text.
5. A method for image-text retrieval based on semantic disambiguation according to claim 1, characterized in that, Aligning the new ambiguous word representation with the image region features includes: locally and globally aligning the ambiguous word interpretation in each image-text pair with the image features, obtaining the local similarity probability distribution by calculating the similarity between each ambiguous word in the text and the local features of the image; using the KL loss function to bring the local similarity probability distribution and the ambiguous word label closer; obtaining the global features by fusing the image features through cat; calculating the global similarity probability distribution by combining the global features and the ambiguous word interpretation in the text; and using the KL loss function to bring the global probability distribution and the ambiguous word label closer.
6. The method for graphic and text retrieval based on semantic disambiguation according to claim 5, characterized in that The similarity probability distribution calculation formula is: Among them, is the sense embedding of the ambiguous word updated by GAT, is the local feature of the picture, are the four interpretations corresponding to the ambiguous word, is the similarity value after normalization, r i,k is the similarity value calculated from the local feature and the ambiguous word.
7. A method for image and text retrieval based on semantic disambiguation according to claim 5, characterized in that The KL loss function is: Among them, N represents the number of local features in the text-image pair, M represents the number of ambiguous words in each text-image pair, and y ij is the true label of the ambiguous word, is the predicted probability distribution of the ambiguous word output in the relevant layer.
8. A method for image and text retrieval based on semantic disambiguation according to claim 1, characterized in that The loss functions of the model include: using Triplet Loss to measure the similarity between images and texts in the embedding space learned by the model; using Cross-Entropy Loss as the loss calculation formula for disambiguating ambiguous words; and constructing the model loss function based on Triplet Loss, Cross-Entropy Loss, and KL Loss.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the image-text retrieval method based on semantic disambiguation as described in any one of claims 1 to 8.
10. An image-text retrieval device based on semantic disambiguation, characterized in that, It includes a processor and a memory; the memory is used to store computer programs; the processor is connected to the memory and is used to execute the computer programs stored in the memory, so that the image and text retrieval device based on semantic disambiguation executes the image and text retrieval method based on semantic disambiguation as described in any one of claims 1 to 8.