Method and apparatus for judging image-text pairs
By introducing phrase-level supervision mechanism and mask self-attention mechanism, the problem of difficult semantic mismatch at the phrase-level in the prior art image text retrieval is solved, and more efficient graphic and text retrieval performance and model interpretability are achieved.
Patent Information
- Application Number
- CN202210615255.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-01
AI Technical Summary
Existing image text retrieval methods mainly rely on sentence-level attention mechanisms, making it difficult to identify semantic mismatch between images and sentences on a fine-grained basis, especially at the phrase level.
A phrase-level supervision mechanism is introduced, and a phrase-level semantic label is generated, and a mask self-attention mechanism is used to establish a relationship model between modals and within modals, and the global, local and phrase-level matching degrees are calculated to improve the accuracy of the judgment of image text pairs.
Through the phrase-level supervision mechanism, the model can more accurately identify mismatched sentences, improve the performance of graphic and text retrieval, and enhance the interpretability and credibility of the model.
Smart Images

Figure CN115017356B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly relates to a method and device for judging image-text pairs. Background Art
[0002] Vision and language are two important aspects for humans to understand the world. To bridge vision and language, researchers have increasingly focused on multi-modal tasks. Image-text retrieval is one of the fundamental topics, aiming to query matching text (image) based on an image (text). Researchers extract features from text pairs to calculate an estimated matching degree, thereby measuring similarity. The model is optimized through triplet loss, making the modal features of correct image-text pairs superior to those of incorrect image-text pairs.
[0003] Figure 1 An example is shown, including a query image, some matching sentences, and non-matching sentences. In terms of matching scores, the model (Faghri et al. 2018) cannot distinguish between matching and non-matching sentences. A careful observation of this example reveals that most of the non-matching sentences actually match the picture in terms of scores, and only a few phrases have inconsistent semantics (two dogs, baseball field, etc.). Therefore, research shows that semantic mismatch between images and sentences often occurs at a finer grain. Summary of the Invention
[0004] The inventors of this application have shown through research that existing image-text retrieval studies mainly rely on sentence-level supervision to distinguish sentences that match or do not match a query image. However, semantic mismatch between images and sentences often occurs at a finer grain, i.e., the phrase level. In this paper, this application explores introducing additional phrase-level supervision to better identify mis-matched units in the text.
[0005] Embodiments of this application provide a method and device for judging image-text pairs to improve the effectiveness in terms of the overall performance of image-text retrieval.
[0006] Embodiments of this application provide a method for judging image-text pairs, including the following steps:
[0007] Generate phrase-level semantic labels based on the sentence-level semantic labels of the picture;
[0008] Establish an inter-modal relationship model and an intra-modal relationship model;
[0009] Calculate the image-text matching degree according to global pairing, local pairing, and phrase pairing, where the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated by the phrase-level semantic tags based on local matching.
[0010] Preferably, in the step of "generating phrase-level semantic tags according to the sentence-level semantic tags of the image", a parser is used to extract entity words, adjective plus entity words, and verb triples from the sentence-level semantic tags.
[0011] Preferably, in the step of "generating phrase-level semantic tags according to the sentence-level semantic tags of the image", the annotated text in the image library is used as the sentence-level semantic tags.
[0012] Preferably, in the step of "establishing the inter-modal relationship model and the intra-modal relationship model", a masked self-attention mechanism is adopted to establish the inter-modal relationship model; where the masked self-attention mechanism includes:
[0013] (1) At the visual end, no attention operation is performed between all region nodes and the global sentence node.
[0014] (2) At the language end, no attention operation is performed between all phrase, word nodes and the global image node.
[0015] (3) No attention operation is performed between each phrase node and any other word not included in the phrase itself.
[0016] Preferably, in the step of "establishing the inter-modal relationship model and the intra-modal relationship model", word embeddings, phrase embeddings, and global sentence embedding vectors are used as the input at the text end; the initial image vector and the global image node are used as the input at the image end.
[0017] Preferably, in the step of "calculating the image-text matching degree according to global pairing, local pairing, and phrase pairing, where the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated by the phrase-level semantic tags based on local matching", the global pairing represents the similarity between the overall image and the overall text.
[0018] Preferably, in the step of "calculating the image-text matching degree according to global pairing, local pairing, and phrase pairing, where the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated based on phrase-level semantic tags on the basis of the local pairing", the local pairing represents the similarity between the overall image and the local text and the similarity between the local image and the overall text.
[0019] Preferably, in the step of "calculating the image-text matching degree according to global pairing, local pairing, and phrase pairing, where the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated based on phrase-level semantic tags on the basis of the local pairing", the phrase pairing is generated by knowing the unmatched words and phrases at the text end and generating by multiplying coefficients.
[0020] An embodiment of the present application discloses a determination device based on an image-text pair, including:
[0021] One or more processors;
[0022] A memory; and
[0023] One or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to the above methods.
[0024] An embodiment of the present application discloses a computer-readable storage medium storing one or more programs, where the one or more programs include instructions,
[0025] When the instructions are executed by a computing device, the computing device is caused to execute any one of the methods according to the above methods.
[0026] In the embodiment of the present application, by introducing phrase nodes to expand the phrases input to the self-attention encoder and maintaining the hierarchical structure relationship between words and phrases during the encoding process, better multi-granularity semantic modeling is achieved. More importantly, previous work focused on finding better models for better cross-modal representation learning, and the loss function was limited to always being a sentence-level triplet loss. In the work of the present application, the present application provides a fine-grained supervision signal at the phrase level, rather than only providing a sentence-level matching (mismatching) signal, so as to guide the model to make decisions more based on irrelevant local parts and distinguish mismatching sentences. This method not only helps the model obtain better retrieval performance, but also is more interpretable and credible. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0028] Figure 1 is an example of querying an image. The score of this example is generated by VSE++ [9]. This example shows a set of mismatched sentences, a set of matched sentences and their corresponding text scene graphs, as well as phrase-level semantic tags. Among them, the underlined text segments represent mismatches at the phrase level.
[0029] Figure 2 is the overall framework of the model Semantic Structure-Aware Multimodal Transformer (SSAMT) proposed in the present application. Among them, the blank circles in the mask matrix M indicate that the query nodes in this column do not pay attention to the corresponding key nodes in this row.
[0030] Figure 3 Shows the comparison results of Recall@K (R@K) for cross-modal retrieval on MS-COCO 1K and Flickr30K.
[0031] Figure 4 Shows the comparison results of Recall@K (R@K) for cross-modal retrieval on MS-COCO. Detailed implementation manners
[0032] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0033] Construct multi-granularity semantic tags according to the query image, which are respectively sentence-level semantic tags and phrase-level semantic tags. For sentence-level semantic tags, the present application uses the entire sentence as the tag. For phrase-level semantic tags, the present application constructs the text scene graph of the sentence and extracts entities and various forms of triples from the graph as tags. Based on these multi-granularity semantic tags, the matching model desired by the present application is expected to be able to identify fine-grained mismatched semantic units while distinguishing incorrect sentences.
[0034] To perform cross-modal representation learning by leveraging both sentence-level and phrase-level information simultaneously, this application proposes a Semantic Structure Aware Multimodal Transformer (SSAMT) to model multi-granularity semantics in vision and language. On the language side, this application concatenates a sentence and its phrases as the input, while on the vision side, it uses an image and the object features within it for concatenation as the input.
[0035] This application uses the masked self-attention mechanism to model different granularity semantic units of the two modalities, and proposes a new attention mechanism for intra-modal and inter-modal interactions. The model learns the modalities of vision (image and region) and language (sentence and phrase) in multiple dimensions (global and local).
[0036] For optimization, this application uses global matching and local matching to calculate the similarity of image-text pairs. Among them, global matching calculates the overall matching score of an image and text, and local matching calculates similarity from a fine-grained perspective, including region-to-text and phrase-to-image.
[0037] In addition, for phrases extracted from mismatched sentences, this application proposes phrase-matching to guide the model to increase the scores between matching image-phrase pairs and reduce the scores between those mismatched image-phrase pairs. Experimental results based on MS-COCO (Lin et al., 2014) and Flickr30K (Plummer et al., 2015) show that, compared with some state-of-the-art methods, the performance of the model in this application is competitive. Further analysis shows that SSAMT can provide better interpretability by locating the mismatched phrases in mismatched sentences.
[0038] The overall framework of SSAMT in the embodiments of this application is as Figure 2 shown. It includes three main components, namely a multi-granularity semantic tagging system, a cross-modal feature learning model with multi-granularity semantics, and a multi-granularity matching loss.
[0039] The multi-granularity semantic tagging system automatically collects semantic tags from the annotated sentences of the query image. Cross-modal feature learning with multi-granularity semantics captures semantics of different granularities in the two modalities. The multi-scale loss is used to calculate the similarity between an image and the corresponding sentence. This application takes an image I i and a sentence T j as an example to calculate their matching scores.
[0040] The multi-granularity semantic tagging system can obtain the sentence-level semantic tags of an image and generate phrase-level semantic tags based on the sentence-level semantic tags of the image.
[0041] Each image in the vision and language dataset has multiple annotated sentences, for example, five in MS-COCO (Lin et al. 2014) and Flickr30K (Plummer et al. 2015). These sentences describe the multi-granularity semantics of the image, and the multi-granularity semantics include various objects, relationships, and scenes. This application proposes to use them to automatically configure the corresponding phrase-level semantic tags.
[0042] In practice, this application uses the scene graph parser of SPICE (Anderson et al. 2016) based on text scene graphs to mine object-relationship-object triples, object-attribute pairs, and object entities (entity words, adjectives plus entity words, and verb triples) from the above descriptive sentences (sentence-level semantic tags), where SPICE complies with SGAE1 (Yang et al. 2019b). For example, in Figure 2 , the retrieved phrases include "dog catches frisbee" and "yellow frisbee".
[0043] In addition, the words of each sentence are also collected as a supplement to the above phrases. The phrases and words are regarded as the phrase-level semantic tags L i of the image I i .
[0044] With the knowledge of semantic tags, this application can determine whether each phrase or word in sentence T j matches I i through a phrase matching method. The phrase matching method is as follows: If each phrase or word in T j appears in the phrase-level semantic tag L i or is included in a certain tag in L i , this application considers it positive, otherwise, it is negative. For example, if "black dog" is in L i , then "dog" in T j is a positive example, while "dogs" is a negative example.
[0045] To initialize the embedding of T j , this application prepares different strategies for words, phrases, and sentences.
[0046] (1) For word embeddings, this application uses the standard embedding layer of Devlin et al., which consists of context embeddings, position embeddings, and segment embeddings.
[0047] (2) For phrase embeddings, this application also maps each of them to a dense vector. Considering that there are three types of phrases as described above through the scene graph parser, including object attributes and relationships, this application adopts three phrase fragment embedding vectors for them. In this process, this application does not directly add semantically related embedding information, but will compensate for it later using the masked self-attention encoder.
[0048] (3) For sentences, this application uses the dense vector C T as an initialization to establish a global sentence node to capture the overall representation of the sentence in the subsequent modeling below.
[0049] In summary, the text embedding of this application is concatenated by three parts: word embeddings phrase embeddings and the global sentence embedding C T .
[0050] For the image I i , this application uses a pre-trained object detector to extract region features and repair them during training. To adapt to the hidden size of the encoder of this application, this application adds a fully connected layer to project each region feature to the same size and obtain the initial image vector After setting the global sentence node, this application also sets a global image node with C I to capture the overall semantics of the image.
[0051] To strengthen the interaction of multi-granularity semantics between images and texts, this application simultaneously adopts an inter-modal relationship and an intra-modal relationship model, and proposes a masked attention mechanism within SSAMT to learn multi-granularity semantics with an inherent structure.
[0052] The inter-modal relationship model aims to establish an interaction between the two modalities. This application uses the encoder of the self-attention encoder (Vaswani et al. 2017) as the backbone. In the following equations, this application concatenates the above-mentioned text embedding and image vector as the model input.
[0053]
[0054] In the original setting of the self-attention encoder, there are no different granularities or structures, and each element interacts with other elements in an unconstrained manner for attention.
[0055] Existing self-attention encoders are used for sequences where each element participates in other elements without constraints. In this application, in addition to word and region nodes in existing cross-modal self-attention encoders, this application also adopts phrase nodes and holistic semantic nodes from images and sentences respectively for modeling, where phrase nodes represent words in a phrase and holistic semantic nodes represent modality-related global nodes.
[0056] If the attention mechanisms of these two types of nodes H 0 focus on any nodes, they cannot learn the expected representations. To encode these nodes according to the corresponding structural constraints in a specific part of H 0 , this application uses masked attention to satisfy the structural constraints. In implementation, the masking matrix is all initialized to 0, which means that by default each node can process any other node. This application resets the values at specific positions to -∞ as the following requirements. An example of M is Figure 2 shown as follows.
[0057] (1) The global sentence node C T does not focus on the region nodes in it, and vice versa.
[0058] (2) The global image node C I does not focus on the phrase and word nodes in it, and vice versa.
[0059] (3) Each phrase node does not focus on any other words that the phrase itself does not contain, and vice versa. For example, in Figure 2 , the phrase node of P1 (dog catch frisbee) does not focus on the word node W3 (jump) that is not in the phrase. This application adds M to the following attention function, uses it to replace the original self-attention (SAN) in the transformer, and forms a new masked attention mechanism (mask transformer).
[0060]
[0061] After modeling the inter-modal relationships, this application obtains a series of outputs as shown in the following equations.
[0062]
[0063] where and are the global representations of the image and text corresponding to the global nodes C I and C T of the image and sentence. is the representation of the region, They are the representations of phrases and words, which are the local representations of images and sentences respectively.
[0064] The intra-modal relationship model is used to encode images and texts respectively as a supplement to the inter-modal relationship modeling, where the inputs of images and texts are and This application takes the outputs of C I and C T as the intra-modal global representations of images and sentences, denoted as and
[0065] During the training process, this application has a correct (matched) image-text pair (I i , T i ), an incorrect (unmatched) image I k and an incorrect (unmatched) sentence T j This application uses the triplet loss TriL α to train the model of this application. In the following TriL α (u, V, W), α is a scalar used to control the distance between the cosine scores of u and the positive sample V and the cosine scores of the negative sample W. The loss is to make each v ∈ V closer to u and push each w ∈ W away from u. Based on the multi-granularity semantic labels, this application uses three matching scores to measure the similarity of these image-text pairs, including global pairing, local pairing, and phrase pairing.
[0066]
[0067] For global pairing, both the intra-modal relationship model and the inter-modal relationship model have an impact on the global representations of images and sentences. Therefore, this application uses and to calculate the correct image-text pair (I i , T i ) and the incorrect image-text pair (I i , T j ). The corresponding loss equations are as follows:
[0068]
[0069]
[0070] For local pairing, the present application uses local matching based on the inter-modal relationship model to enhance fine-grained cross-modal matching. Local pairing has two parts: (1) Region-to-Sentence: the matching of each region with a sentence. (2) Phrase-to-Image: the similarity of each phrase (word) with an image. The present application uses the loss in the equation to make the local matching score of the correct image-text pair greater than that of the incorrect image-text pair. The specific equation is as follows:
[0071]
[0072] For phrase pairing, the present application divides each phrase or word of T j into matching
[0073] and non-matching according to the aforementioned phrase matching method. The present application repeats the above procedure for non-matching pairs (I k , T i ) to obtain and Considering that the matching part is the key to separating non-matching image-text pairs, the present application proposes to further push away the non-matching and matching parts in non-matching sentences. It can also be interpreted as punishing the non-matching part, which is to guide the matching model to make more decisions based on them.
[0074]
[0075] Based on these three types of matching methods and the corresponding losses, the present application obtains the overall loss as the following equation, which uses hyperparameters λ0, λ1, λ2, and λ3 to balance these losses.
[0076]
[0077] Previous image-text retrieval models usually use the most difficult image (text) in the batch data as the negative sample (text), which requires batch calculation of the matching scores of all pairwise image-text combinations. This is costly in inter-modal relationship modeling. Therefore, the present application samples negative instances through the intra-modal matching score to reduce the computational cost.
[0078] Specifically, during the inference process, the present application uses the following score(I i , T j ) for ranking. Where μ1 and μ2 are hyperparameters.
[0079]
[0080] Compared with these self-attention encoder-based models, the model of the present application extends the phrases input to the self-attention encoder by introducing phrase nodes and maintains the local structure of words during the encoding process to achieve better multi-granularity semantic modeling. More importantly, previous work focused on finding better models for better cross-modal representation learning, and the attention mechanism was always the sentence-level triplet loss. In the work of the present application, the present application provides fine-grained attention at the phrase level instead of only providing sentence-level matching (mismatching) signals, and the attention mechanism guides the model to distinguish mismatching sentences, more based on irrelevant local parts. This method not only helps the model obtain better retrieval performance, but also is more interpretable and credible.
[0081] Specific implementation effects:
[0082] The present application evaluates the model proposed in the present application on MS-COCO (Lin et al. 2014) and Flickr30K (Plummer et al. 2015). Each image in MS-COCO is accompanied by 5 manually annotated captions. The present application divides the dataset into a training set, a validation set, and a test set, with 113, 287 / 5, 000 / 5, 000 images respectively (Karpathy and Fei-Fei 2015). For MS-COCO 1K, the test set is further divided into 5 splits, and the reported performance is the average of 5 folds of 1K test images (Faghri et al. 2018). Flickr30K (Plummer et al. 2015) contains 31,000 images collected from the Flickr website. Each picture contains 5 descriptive sentences. The present application uses the same split as Karpathy and Fei-Fei (2015) for the training, validation, and test sets, with 1,000 images for validation, 1,000 images for testing, and the rest for training.
[0083] The present application compares the model of the present application with some classic and state-of-the-art methods, including VSE++ (Faghri et al. 2018), CAMP (Wang et al. 2019b), SCAN (Lee et al. 2018), SGM (Wang et al. 2018). 2020), VSRN (Li et al. 2019), BFAN (Liu et al. 2019), MMCA (Wei et al. 2020), GSMN (Liu et al. 2020). The results on MS-COCO 1K and Flickr30K are as Figure 3 shown, and the results on MS-COCO are as Figure 4As shown. In this application, it can be seen that the proposed SSAMT of this application is superior to all existing methods. The best R@1 of the image is 78.2% for text-to-image retrieval on MS-COCO 1K, and R@1 = 62.7%. For MS-COCO, the proposed method maintains its advantage, with an improvement of more than 3% in R@1 for image-to-text retrieval. In Flickr30K, the model of this application achieves the best performance, with an image-to-text R@1 of 75.4%.
[0084] In summary, in this paper, in order to make full use of the mismatched sentences at the phrase level and the sentence level, this application explores the construction of multi-granularity semantic tags, where the semantic tags at the phrase level are automatically constructed from the image-related sentences by extracting the phrases of object entities, object-attribute pairs, and object-relation-object triples.
[0085] This application concatenates the sentence and its phrases at the language end, and concatenates the image and its objects at the visual end, then proposes a masked self-attention mechanism for joint cross-modal modeling with multi-granularity semantics, and uses a multi-scale matching loss to capture the image-to-text matching and region-to-sentence / phrase-to-image matching.
[0086] Based on the phrase-image matching, this application uses semantic tags to determine the non-corresponding relationship between phrases and images, and adjusts the scores between image-phrase pairs. The experimental results show the effectiveness of the model of this application on MS-COCO and Flickr30K. Further analysis shows that semantic tags improve the efficiency of data utilization and guide the model to distinguish mismatched sentences based on more mismatched parts.
[0087] The embodiment of this application also includes a judgment device based on image-text pairs, including: one or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.
[0088] The embodiment of this application also includes a computer-readable storage medium storing one or more programs, where the one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the above methods.
[0089] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0090] It should be noted that the systems, devices, modules or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, in this specification, when describing the above devices, they are divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0091] In addition, in this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where the circumstances permit, the reference elements or components or steps (etc.) should not be construed as being limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.
[0092] In this embodiment, the above storage medium includes but is not limited to Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is used for the interface of network connection communication.
[0093] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained in contrast to other embodiments and will not be elaborated here.
[0094] Although different specific embodiments are mentioned in the content of the present application, the present application is not limited to the situations described by industry standards or embodiments. Some industry standards or implementation schemes slightly modified on the basis of the implementation described by the custom method or embodiment can also achieve the same, equivalent or similar, or predictable implementation effects after deformation as the above embodiments. The embodiments of applying these modified or deformed data acquisition, processing, output, judgment methods, etc. still fall within the scope of the optional implementation schemes of the present application.
[0095] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it may be executed in the method order shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device including the said elements.
[0096] The devices or modules etc. illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by the combination of multiple sub-modules, etc. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0097] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and the structures within the hardware component.
[0098] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform particular tasks or implement particular abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0099] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0100] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0101] Although the present application has been depicted through embodiments, those of ordinary skill in the art know that the present application has many variations and changes without departing from the spirit of the present application. It is hoped that the appended embodiments include these variations and changes without departing from the present application.
Claims
1. A method for judging image-text pairs, characterized in that, it includes the following steps: Generate phrase-level semantic tags according to the sentence-level semantic tags of the picture; Build an inter-modal relationship model and an intra-modal relationship model, and adopt a masked self-attention mechanism to build the inter-modal relationship model; wherein, the masked self-attention mechanism includes: (1) At the visual end, all regional nodes and global sentence nodes do not perform attention operations on each other; (2) At the language end, no attention operations are performed between all phrase, word nodes and global image nodes; (3) No attention operation is performed between each phrase node and any other word not included in the phrase itself; Calculate the picture-text matching degree according to global pairing, local pairing and phrase pairing, wherein the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated on the basis of local matching by the phrase-level semantic tags.
2. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "generating phrase-level semantic tags according to the sentence-level semantic tags of the picture", a parser is used to extract entity words, adjective-plus-entity words and verb triples from the sentence-level semantic tags.
3. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "generating phrase-level semantic tags according to the sentence-level semantic tags of the picture", the annotated text in the picture library is used as the sentence-level semantic tags.
4. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "building an inter-modal relationship model and an intra-modal relationship model", word embeddings, phrase embeddings and global sentence embedding vectors are used as inputs at the text end; The initial image vector and global image nodes are used as inputs at the image end.
5. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "calculating the picture-text matching degree according to global pairing, local pairing and phrase pairing, wherein the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated on the basis of local matching according to the phrase-level semantic tags", the global pairing represents the similarity between the overall picture and the overall text.
6. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "calculating the picture-text matching degree according to global pairing, local pairing and phrase pairing, wherein the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated on the basis of local pairing according to the phrase-level semantic tags", the local pairing represents the similarity between the overall picture and the local text and the similarity between the local picture and the overall text.
7. The method for judging image-text pairs according to claim 1, characterized in that, In the step of "calculating the image-text matching degree according to global pairing, local pairing, and phrase pairing, where the global pairing is generated by the inter-modal relationship model and the intra-modal relationship model, the local pairing is generated by the inter-modal relationship model, and the phrase pairing is generated based on phrase-level semantic tags on the basis of the local pairing", the phrase pairing is generated by knowing the unmatched words and phrases at the text end and multiplying by coefficients.
8. An apparatus for judging based on image-text pairs, comprising: one or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to claims 1-7.
9. A computer-readable storage medium storing one or more programs, the one or more programs including instructions, the instructions, when executed by a computing device, cause the computing device to execute any one of the methods according to claims 1-7.
Citation Information
Patent Citations
Cross-modal image-text matching method and device and computer readable storage medium
CN112905827A
Training method of image-text matching model, BI-directional search method, and relevant apparatus
US20200019807A1