Image finger expression understanding method based on zero sample
By combining the pre-trained CLIP model with the graph network and using self-attention strategy for multimodal aggregation, the problem of poor performance of the zero-sample image reference representation understanding (REC) method in complex expression and global reasoning tasks is solved, and efficient matching and recognition of unseen target categories is achieved.
Patent Information
- Application Number
- CN202510224460.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
Existing zero-sample images refer to Representative Understanding (REC) methods perform poorly when dealing with complex expressions and tasks requiring global reasoning and fail to effectively match the visual features of candidates.
The pre-trained CLIP model is used to integrate it into a two-stage algorithm based on graph networks. The graph network structure combined with self-attention strategy is used to perform multimodal aggregation of visual objects and text expressions, and the corresponding category labels of candidate objects whose categories are not covered in the training data are inferred.
Effectively solve the challenges in zero-sample REC tasks, especially when dealing with unseen target categories, perform better than the current state-of-the-art methods, and improve performance and matching accuracy through the combination of feature alignment and graph coordinated attention modules.
Smart Images

Figure CN120146204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of zero-shot image-text matching and pattern recognition, and particularly to a zero-shot based method for understanding image referential expressions. Background Art
[0002] Most of the existing referential expression comprehension (REC) methods rely on supervised learning. However, due to its inherent challenges, there is less research on unsupervised or zero-shot REC in the existing literature. Existing zero-shot REC methods often focus on local content and perform poorly in dealing with complex expressions and tasks that require global reasoning. On the other hand, Wang et al. proposed an unsupervised REC method that does not require training or paired phrase localization annotation. This method locates the target object by querying the semantic similarity between the expression and the predicted concept labels of candidate objects from various visual detectors. However, this method does not utilize the visual features of candidate objects for matching. Therefore, when the detected concepts are not sufficient to distinguish the appearance of the target object from other similar objects in the image, its performance is poor. Summary of the Invention
[0003] The present invention aims to solve the problem of unseen objects in the inference stage and focuses on the zero-shot REC task by utilizing the pre-trained CLIP model. The CLIP model performs excellently in zero-shot image classification tasks because it has been contrastively pre-trained using 400 million image-text pairs. To better utilize the capabilities of the CLIP model, it is adapted to the REC task through the training data of the target dataset. Specifically, the present invention proposes a zero-shot based method for understanding image referential expressions, which integrates the contrastive language-image pre-training (CLIP) model into a two-stage algorithm based on a graph network. The graph network structure can utilize the local and global context information of the detected objects and can work flexibly with additional modules. Through the general feature space learned by CLIP, the graph network method of the present invention can infer the corresponding class labels of candidate objects whose classes are not covered in the training data.
[0004] To achieve the above object, the present invention provides the following solution:
[0005] A zero-shot based method for understanding image referential expressions, comprising:
[0006] Obtain a target image and a target expression, input the target image and the target expression into the REC model, and obtain the localization result of the target object; the REC model is trained using a training set and is obtained by training in combination with a loss function related to the self-attention strategy and a total loss function; the training set includes: original target images and target categories and candidate category labels unrelated to the expression; the target image includes: an appearance map and a category map; the appearance map is the appearance feature of the candidate region and the nouns matched in the expression, and the nodes in the category map consider the category semantics of the candidate region and the nouns matched in the expression;
[0007] In the process of obtaining the localization result of the target object, use the language parsing module in the REC model to parse the target expression, and in combination with the self-attention strategy, obtain a vector sequence containing learnable vectors. Based on the graph co-attention module, perform appearance encoding and semantic encoding on the target image respectively, and in combination with the vector sequence, obtain the node features and node weights, edge features and edge weights of the target image. Through the language parsing module, perform multi-step reasoning on the node features and node weights, edge features and edge weights of the target image, and further use the matching module to calculate the matching scores between the composite nodes in the appearance map and the category map, and take the object corresponding to the node with the highest matching score as the localization result of the target object.
[0008] Optionally, obtaining the vector sequence containing learnable vectors includes:
[0009] Use the language parsing module to parse the target expression, embed each word in the target expression into a vector through a non-linear mapping function, obtain an embedded vector set, and encode the embedded vectors in the embedded vector set to generate an initial vector sequence;
[0010] Based on the self-attention strategy, determine the weights of the nouns in the target expression to generate a weight group:
[0011]
[0012] where, α n,l is the weight group, h l is the l-th initial vector sequence, L is L weights, W n is the learnable vector, and T is the transpose;
[0013] Convert the weight group into a learnable vector and apply it to the initial vector sequence to obtain the vector sequence containing learnable vectors:
[0014]
[0015] where, τ nis a vector sequence containing learnable vectors, h l is the l-th initial vector sequence, α n,l is the weight group.
[0016] Optionally, obtaining the node features and node weights of the target image includes:
[0017] Step 1.1, obtain the node features of the appearance graph:
[0018]
[0019] Among them, is the node feature, MLP a is the multi-layer perceptron, o i is the visual feature of the i-th candidate region, b i is the box feature of the i-th candidate region, ∈ a is the trainable bias term, is the feature of the th noun in the expression;
[0020] Step 1.2, calculate the matching degree between the candidate region and the nouns matching in the expression, and obtain the best matching noun:
[0021]
[0022] Among them, is the best matching noun, is the matching degree, W a,1 、W a,2 and W τ are trainable parameter matrices, N is N nouns, τ n is a vector sequence containing learnable vectors;
[0023] Step 1.3, according to the best matching noun, assign the node weight:
[0024]
[0025] Among them, is the result of calculating the similarity between the i-th target and the th noun, is the result of calculating the similarity between the k-th target and the th noun;
[0026] Step 1.4, replace the appearance graph with the category graph, that is, replace the visual features in the appearance graph with text features, and repeat steps 1.1 - 1.3 until the node features and node weights of the category graph are obtained.
[0027] Optionally, obtaining the edge features and edge weights of the target image includes:
[0028] Step 2.1, obtaining the edge features of the appearance graph:
[0029]
[0030] Wherein, is a trainable parameter matrix;
[0031] Step 2.2, calculating the correlation degree between the edge features and the expression:
[0032]
[0033] Wherein, W q and are trainable parameter matrices, and q is the representation of the entire expression;
[0034] Step 2.3, according to the correlation degree, aggregating the node features and normalizing the weights of the edges connected to the node to obtain the edge weights:
[0035]
[0036] Wherein, 1 [i≠j] is an indicator function, and K is the target in K images;
[0037] Step 2.4, replacing the appearance graph with a category graph, that is, replacing the visual features in the appearance graph with text features, and repeating steps 2.1 - 2.3 until the edge features and edge weights of the category graph are obtained.
[0038] Optionally, performing multi-step reasoning on the node features and node weights, and edge features and edge weights of the target image includes:
[0039]
[0040] Wherein, W t is a trainable parameter matrix, θ t is a trainable parameter vector, and are respectively the node features aggregated from the connected neighbor nodes and the self-loop transformation features of the i-th node in the t - 1 inference steps.
[0041] Optionally, obtaining the node features aggregated from the connected neighbor nodes and the self-loop transformation features of the i-th node in the t - 1 inference steps includes:
[0042]
[0043] Wherein, and are two trainable parameter matrices, and is a trainable parameter vector.
[0044] Optionally, calculating the matching scores between composite nodes in the appearance graph and the category graph includes:
[0045]
[0046] where * ∈ {a, c}, a is the appearance graph, c is the category graph, is the feature of the i-th node at the T-th iteration, and are two trainable parameter matrices, ||·|| represents the L 2 norm of the vector.
[0047] Optionally, the expression of the loss function related to the self-attention strategy is:
[0048]
[0049] where W e is a trainable parameter matrix, L e is the expression loss function, represents the word index corresponding to the p n -th noun in the expression, τ n is the expression of the n-th noun.
[0050] Optionally, the expression of the total loss function is:
[0051]
[0052] where, and are the scores of the GT nodes of the true labels in G a and G c , is the k-th score in the appearance graph, is the k-th score in the category graph.
[0053] The beneficial effects of the present invention are:
[0054] 1) By utilizing the CLIP model pre-trained on a large number of image-text pairs, the present invention effectively solves the challenging zero-shot Referring Expression Comprehension (REC) problem. It adopts a graph-based multimodal aggregation strategy to adapt the CLIP model to the downstream REC task, effectively aggregating the CLIP features of visual objects and text expressions. This strategy ensures that the expressions obtained from language processing are synchronized with the CLIP features, effectively bridging the gap between domains.
[0055] 2) When the CLIPREC of the present invention processes the zero-shot REC task of unseen target categories, it performs more excellently compared to the current state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a schematic diagram of a zero-shot image referential expression understanding method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0060] This embodiment aims to solve the problem of unseen objects in the inference stage and focuses on the zero-shot REC task by using the pre-trained CLIP model. The CLIP model performs excellently in the zero-shot image classification task because it has been contrastively pre-trained with 400 million pairs of image-text data. In order to better utilize the capabilities of the CLIP model, this embodiment adapts it to the REC task through the training data of the target dataset. Specifically, this embodiment proposes a novel graph network-based zero-shot REC domain adaptation network, called CLIPREC, which integrates the contrastive language-image pre-training (CLIP) model into a two-stage algorithm based on the graph network. The graph network structure can utilize the local and global context information of the detected objects and can work flexibly with additional modules. Through the general feature space learned by CLIP, the graph network method of this embodiment can infer the corresponding class labels of candidate objects whose classes are not covered in the training data.
[0061] The CLIP model is mainly used to extract image-text matching features, rather than performing reasoning or exploring the structure and semantics of a given image and expression. Therefore, it is not optimized for the REC task. Before using the CLIP model for zero-shot matching of unseen objects, this embodiment proposes a Feature Alignment (FA) module to synchronize the CLIP features with the expressions generated by the language parser for representation. In this embodiment, the FA module projects all features into a common subspace by using a trainable multi-layer perceptron (MLP) on the CLIP features, thereby achieving feature alignment.
[0062] The proposed domain adaptation network includes a graph co-attention module, which consists of two directed graphs: one representing the objects in a given image and the other representing the class labels of the objects. This module uses CLIP features to process unseen object categories and performs information aggregation based on a graph convolutional network (GCN). In addition, after the image encoder and text encoder of the pre-trained CLIP model, this embodiment adds two trainable multi-layer perceptrons (MLPs) for feature alignment. In this way, the visual and text features generated by the CLIP model can be better aligned with the expressions output by the language parser, thereby further improving performance. Finally, by calculating the similarity scores between the expressions and the graph nodes, the matching pair with the highest score is selected to complete the localization task of the target object.
[0063] As Figure 1 shown, this embodiment discloses a zero-shot based method for understanding image referential expressions, including: obtaining a target image and a target expression, inputting the target image and the target expression into the REC model to obtain the localization result of the target object; the REC model is trained using a training set and obtained by training in combination with a loss function related to the self-attention strategy and a total loss function; the target image includes: an appearance map and a category map; the appearance map is the appearance features of the candidate region and the nouns matched in the expression, and the nodes in the category map consider the category semantics of the candidate region and the nouns matched in the expression; in the process of obtaining the localization result of the target object, the language parsing module in the REC model is used to parse the target expression, and in combination with the self-attention strategy, a vector sequence containing learnable vectors is obtained, and the target image is respectively subjected to appearance encoding and semantic encoding based on the graph co-attention module, and in combination with the vector sequence, the node features and node weights, edge features and edge weights of the target image are obtained, and the node features and node weights, edge features and edge weights of the target image are subjected to multi-step reasoning by the language parsing module, and further the matching scores between the composite nodes in the appearance map and the category map are calculated by the matching module, and the object corresponding to the node with the highest matching score is used as the localization result of the target object.
[0064] Specifically:
[0065] In this embodiment, in view of the above problems, a zero-shot image referring expression understanding method is proposed. This method can effectively identify objects that have not appeared in the training samples. The model proposed in this embodiment consists of four parts: a language parsing module, a graph collaborative attention module, a multi-step reasoning module, and a matching module.
[0066] Figure 1 The proposed CLIPREC framework is used for zero-shot object-referring expression comprehension (REC). First, the expression is characterized by N + 1 representations, namely and q, where represents the representations of N objects in the expression, and q represents the representation of the entire expression. Then, the pre-trained CLIP model is used to infer the class labels of the target candidate objects, and based on the N + 1 expression representations, the CLIP features of the adapted target candidate objects, and their class labels, two attention graphs G a and G c are constructed. In each graph, the node weights represent the relationship between the candidate target and the expression, and the edge weights represent the relationship between two connected nodes and the expression. Finally, multi-step reasoning based on GCN is performed on the two attention graphs, and the target is located in the image by matching the target object specified by the given expression. The larger nodes in the graph represent the higher weights assigned in each graph. The parameters in the green and yellow modules are trainable and fixed, respectively.
[0067] Furthermore, obtaining the vector sequence containing learnable vectors includes:
[0068] Using the language parsing module to parse the target expression, embedding each word in the target expression into a vector through a non-linear mapping function, obtaining a set of embedding vectors, and encoding the embedding vectors in the set of embedding vectors to generate an initial vector sequence;
[0069] Based on the self-attention strategy, determining the weights of the nouns in the target expression to generate a weight set:
[0070]
[0071] where α n,e is the weight set, h e is the l-th initial vector sequence, L is the number of L weights, and W n is the learnable vector;
[0072] Converting the weight set into learnable vectors and applying them to the initial vector sequence to obtain a vector sequence containing learnable vectors:
[0073]
[0074] where τn is a vector sequence containing learnable vectors, h l is the l-th initial vector sequence, α n,l is the weight group.
[0075] Specifically:
[0076] Language interpreter:
[0077] Due to the self-attention based bidirectional long short-term memory network (Bi-LSTM), in this embodiment, a Bi-LSTM language parser combined with a self-attention strategy is adopted to obtain more accurate attention. Given an expression Q containing L words, each word is embedded into a vector through a non-linear mapping function to obtain an embedded vector set {e 1 , e 2 ,..., e L}. The Bi-LSTM is used to encode these embedded vectors to generate a vector sequence H = {h 1 , h 2 ,..., h L}. Among them, h l is the concatenation of the forward and backward outputs of the l-th word in the Bi-LSTM. At the same time, the overall representation of the expression is represented by the feature vector q, which is the concatenation of the final hidden states of the forward and backward LSTMs.
[0078] Let the number of nouns (entities) in the expression be N. In this embodiment, the text embedding representation of N nouns is where represents the word index corresponding to the p n -th noun in the expression, n = 1, 2,..., N. For the self-attention strategy, in this embodiment, N groups of weights are derived for each noun to emphasize the words related to the target noun in the expression. Specifically, for the n-th noun, its corresponding weight group contains L weights to highlight the words related to this noun. In this embodiment, a learnable vector W n is applied to H = {h 1 , h 2 ,..., h L}, and the weight group α n is obtained through the softmax function, that is:
[0079]
[0080] Through the weight group α n , the expression representation of the n-th noun can be expressed as:
[0081]
[0082] By repeating this process for N nouns, this embodiment obtains their representations The loss function related to the self-attention strategy is defined as:
[0083]
[0084] where W e is a trainable parameter matrix. L e is the expression loss function, which aims to help each noun (entity) find other relevant words (descriptions) in the expression. It models the relationship between the nth noun and other words in the expression through the soft attention weights In other words, among the weight group α n of the nth noun, the words with higher relevance to this noun will be assigned larger weights.
[0085] Furthermore, obtaining the node features and node weights of the target image includes:
[0086] Step 1.1: Obtain the node features of the appearance graph:
[0087]
[0088] where is the node feature, MLP a is a multi-layer perceptron, o i is the visual feature of the i-th candidate region, b i is the box feature of the i-th candidate region, ∈ a is a trainable bias term, is the feature of the th noun in the expression;
[0089] Step 1.2: Calculate the matching degree between the candidate region and the nouns that match in the expression, and obtain the best-matching noun:
[0090]
[0091] where is the best-matching noun, is the matching degree, W a,1 、W a,2 and W τ are trainable parameter matrices, N is N nouns, and τ n is a vector sequence containing learnable vectors;
[0092] Step 1.3: Assign node weights according to the best-matching noun:
[0093]
[0094] Among them, is the best matching noun;
[0095] Step 1.4: Replace the appearance graph with a category graph, that is, replace the visual features in the appearance graph with text features, and repeat Steps 1.1 - 1.3 until the node features and node weights of the category graph are obtained.
[0096] Furthermore, obtaining the edge features and edge weights of the target image includes:
[0097] Step 2.1: Obtain the edge features of the appearance graph:
[0098]
[0099] Among them, is a trainable parameter matrix;
[0100] Step 2.2: Calculate the correlation degree between the edge features and the expression:
[0101]
[0102] Among them, W q and are trainable parameter matrices, and q is the representation of the entire expression;
[0103] Step 2.3: According to the correlation degree, aggregate the node features and normalize the weights of the edges connected to the node to obtain the edge weights:
[0104]
[0105] Among them, 1 [i≠j] is an indicator function;
[0106] Step 2.4: Replace the appearance graph with a category graph, that is, replace the visual features in the appearance graph with text features, and repeat Steps 2.1 - 2.3 until the edge features and edge weights of the category graph are obtained.
[0107] Specifically:
[0108] Graph Co - attention Module:
[0109] Once the feature representation of the expression is obtained, this embodiment enters the next stage, constructing a graph structure for encoding the relationship between the objects in the image and the entities in the expression for subsequent inference and matching processes. For this purpose, this embodiment constructs two complete graphs: the appearance graph G a =(V a , ε a ) and the category graph G c =(V c , ε c), respectively for appearance encoding and semantic encoding. The appearance graph G a is constructed by referring to the appearance and expression information of the object candidate regions. The category graph G c is constructed in a similar way, but replaces the appearance information of the object candidate regions with their category labels, where the label of each candidate region is obtained by leveraging the zero-shot image-text matching ability of the CLIP model. Specifically, assume that in this embodiment, there is an image I, in which K object candidate regions are detected, and a predefined label set containing all possible category labels (including base classes and unseen classes). The image encoder and text encoder of the CLIP model are respectively used to extract the features of each candidate region in the image and the features of each category label in the label set. Then, by calculating the cosine similarity between each candidate region and the category label features, cross-modal matching can be achieved, so as to find the target label that best matches each candidate region. Similar to the graph-based REC method, the present invention also maintains an appearance graph G a =(V a , ε a ) and a category graph G c =(V c , ε c ), representing the visual appearance and category labels of K object candidate regions respectively. However, the node features and as well as the edge features and are all developed based on the CLIP model to achieve zero-shot REC.
[0110] The specific content is as follows:
[0111] Node features and node weights. The appearance graph G a and the category graph G c both contain K nodes, where K is the number of object candidate regions. The nodes in the appearance graph G a are encoded as the appearance features of the corresponding candidate regions and the matching nouns in the expression, while the nodes in the category graph G c consider the category semantics of the candidate regions and the matching nouns. This embodiment first describes the node features of the appearance graph G a =(V a , ε a ). For the i-th node in the appearance graph G a , its feature is calculated by a multi-layer perceptron (MLP):
[0112]
[0113] where the multi-layer perceptron MLP consists of fully connected layers, ε ais a trainable bias term. These parameters are used to adapt the connected feature vectors for inference and matching. Among the features input to MLP(·), the first part o i is the visual feature of the i-th object candidate region, extracted by the CLIP visual encoder. The second part b i = W b [x i , y i , w i , h i , w i h i is the box feature of the i-th candidate region, where W b is a trainable parameter matrix, (x i , y i ) are the normalized coordinates of the center of the candidate region, and w i , h i , w i h i are the normalized width, height, and area respectively. The third part is the feature of the -th noun in the expression. These noun features are calculated by Equation (2). Among the N nouns in the expression, the -th noun has the best match with the i-th candidate region. The following describes how to find the noun with the best match for the i-th candidate region. For the i-th candidate region in the image and the
[0114]
[0115] where W a,1 , W a,2 and W τ are trainable parameter matrices. Then, the noun with the best match for the i-th candidate region is determined by:
[0116]
[0117] where N represents the number of nouns in the expression. By repeating the process of Equation (2) for each candidate region i, the node features of the appearance graph can be calculated. To emphasize the nodes that match the nouns in the expression better, in this embodiment, an additional weight is assigned to each node Specifically, the node weight is defined as:
[0118]
[0119] In equations (4), (6), and (7), in this embodiment, the appearance graph G a the node features of each node i the index of the best-matching noun and weight For each node i in the category graph G c its node features the best-matching noun index and weight are calculated in the same way, but the visual feature o i is replaced with the text feature ζ obtained by the CLIP text encoder for the i-th object candidate region i .
[0120] Edge features and edge weights. The appearance graph G a and the category graph G c are both fully connected graphs initially. In this embodiment, the appearance graph G a =(V a , ε a ) edge features The method of this embodiment considers all features of the two nodes connected when deriving the edge features, that is:
[0121]
[0122] where is a trainable parameter matrix. This embodiment needs to emphasize the edges related to the expression. Therefore, the relevance degree of the edge to the expression q is estimated by the following formula:
[0123]
[0124] where W q and are trainable parameter matrices. When aggregating features from the i-th node, the weights of the edges connected to this node are normalized and set to:
[0125]
[0126] where 1 [i≠j] is the indicator function. The previous method, the REC method, shows that in the two-stage method, the output of the first-stage process contains many irrelevant object candidate boxes, which introduces a large amount of noise. By setting the edge weight to 0 (if its original value is less than the preset threshold), this adverse effect can be alleviated. In this work, this embodiment empirically sets the threshold to
[0127] For the appearance graph G a , the edge features have been described in (8), and the edge weights have been described in (10). For the category graph G c , the features and weights of each edge can be calculated by the same method, but all the involved visual features are replaced with the corresponding text features extracted using the CLIP text encoder.
[0128] Further, the multi-step reasoning on the node features and node weights, edge features and edge weights of the target image includes:
[0129]
[0130] where, W t is a trainable parameter matrix, θ t is a trainable parameter vector, and are the node features aggregated from the connected neighbor nodes and the self-loop transformation features of the i-th node in the (t - 1)-th inference step, respectively.
[0131] Further, obtaining the node features aggregated from the connected neighbor nodes and the self-loop transformation features of the i-th node in the (t - 1)-th inference step includes:
[0132]
[0133] where, and are two trainable parameter matrices, and are trainable parameter vectors.
[0134] Further, calculating the matching scores between the composite nodes in the appearance graph and the category graph includes:
[0135]
[0136] where, * ∈ {a, c}, a is the appearance graph, c is the category graph, is the feature of the i-th node at the T-th iteration, and are two trainable parameter matrices, ||·|| represents the L 2 norm of the vector.
[0137] Specifically: The multi-step reasoning module:
[0138] There are context relationships between entities in complex expressions. These relationships can be exploited to more precisely locate the target objects described in the expressions. With the help of a graph-based structure, these relationships can be effectively explored through multiple message passing steps between nodes in the graph. Specifically, once the graph structure is constructed, this embodiment performs multi-step reasoning through a T-layer Graph Convolutional Network (GCN), where T represents the number of reasoning steps for information aggregation on the appearance graph G a and the category graph G c . Below, this embodiment first describes the reasoning details on the appearance graph G a . In the t-th reasoning step, the node features, node weights, edge features, and edge weights of the appearance graph G a are respectively denoted as and . At the beginning, i.e., t = 0, this embodiment initializes these features to the results calculated in (4), (7), (8), and (10). That is and . This embodiment adopts a message passing strategy similar to GCN on G a , where in the t-th reasoning step, the feature of the i-th node is calculated as follows:
[0139]
[0140] where, W t is a trainable parameter matrix, θ t is a trainable parameter vector, and are respectively the node features aggregated from connected neighbor nodes and the self-loop transformation features of the i-th node in the (t - 1)-th reasoning step. Specifically:
[0141]
[0142] where, and are two trainable parameter matrices, and are trainable parameter vectors. By performing the forward propagation of the graph convolutional layer, each composite node in the appearance graph G a aggregates information from other nodes, and nodes more relevant to the expression will be assigned greater weights. Accordingly, the node weights and edge weights of the appearance graph G a (i.e., and ) are also updated to consider the expression representation q, as follows:
[0143]
[0144] Among them, and are trainable parameter matrices. and respectively represent the relevance scores of the nodes and edges of G in the t-th inference step a . When performing feature aggregation for the i-th node in the (t + 1)-th inference step, the weights of this node and the weights of the edges connected to it are further normalized by the scores of all nodes and edges, similar to (7) and (10). The above process is repeated T times for the T-step inference strategy. The calculation methods of the node features, node weights, edge features, and edge weights of the category graph G c are the same as those of the appearance graph G a , except that all the involved visual features are replaced with the corresponding text features extracted using the CLIP text encoder.
[0145] Matching module:
[0146] After T-step inference, this embodiment can calculate the matching scores between the expression and the composite nodes in the two graphs (G a and G c ) to identify the target object specified by the expression. In the graph-based representation, the nodes with higher matching scores are more relevant to the expression.
[0147] In this work, the formula for calculating the matching score of the i-th composite node in the graph G * is as follows:
[0148]
[0149] where * ∈ {a, c}, is the feature of the i-th node at the T-th iteration, and are two trainable parameter matrices. ||·|| represents the L 2 norm of the vector. The scores of the ground truth (GT) nodes in G a and G c are respectively denoted as and and should be maximized during the training phase. To consider both attention graphs simultaneously, this embodiment defines the loss function of this work as follows:
[0150]
[0151]
[0152] Network optimization: During the training process, this embodiment first uses the L in (3) eTo train the language parser proposed in this embodiment. Then, this embodiment fixes the training parameters of the language parser and learns the remaining modules of CLIPREC through L in (15).
[0153] Matching: In the inference stage, for each test sample, this embodiment constructs two attention maps through the same process as above. Even in the face of the correspondence between unseen target categories and candidate category labels, this embodiment can still perform matching through the features extracted by the image and text encoders of the pre-trained CLIP model. Then, these two maps are input into the multi-step inference module for information aggregation. The similarity between each composite node and the expression is calculated through (14) to obtain the matching score. The object corresponding to the node with the highest score will be retrieved as the target, thus completing the (zero-shot) reference expression comprehension (REC) task.
[0154] Algorithm steps:
[0155] The entire model is divided into two parts: training and testing.
[0156] Model training: Training preparation stage; Obtain the language features of the expression, the visual features of the image target, and the target label features. Set the number of training rounds R. Initialize the network model parameters.
[0157] Model training stage: Input the training data. Obtain the features of the expression through the language interpreter, H = {h 1 , h 2 ,..., h L} and the noun features Guided by the expression feature H = {h 1 , h 2 ,..., h L}, train the language interpreter through equations (1)-(3) to obtain the best representation of the noun. Generate the appearance map and label map through equations (4)-(10) and calculate the node weights and edge weights of the two maps. Guided by the expression noun features, perform multi-step inference on the two constructed multi-modal maps through equations (11)-(13). In each step of inference, the nodes of the map receive information from the nodes connected to it. Calculate the total loss through equations (14)-(15) and update the model parameters. Determine whether the maximum number of training rounds R has been reached. If so, enter step 8; otherwise, return to step 2. Save the model parameters and end the training.
[0158] Model testing:
[0159] Testing preparation stage: Prepare the data and obtain the language features of the expression, the visual features of the image target, and the target label features. Load the trained network model parameters.
[0160] Model testing phase: Input test data. Construct an appearance graph and a label graph through the model. Through a multi-step inference strategy, after T-step inference strategy, each node of the appearance graph and the label graph will obtain a score. The scores of the corresponding nodes of the appearance graph and the label graph are added together. Find the position of the picture corresponding to the node with the highest score. Output the prediction result.
[0161] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A zero-sample based method for understanding image referential expressions, characterized in that: include: Acquire a target image and a target expression, input the target image and the target expression into a REC model, and obtain a positioning result of the target object; The REC model is trained using a training set and is obtained by combining a loss function related to a self-attention strategy and a total loss function; The training set includes: the original target image and the target category and candidate category labels that are not related to the expression; the target image includes: an appearance map and a category map; the appearance map is the appearance features of the candidate area and the nouns matched in the expression, and the nodes in the category map consider the category semantics of the candidate area and the nouns matched in the expression; In the process of obtaining the positioning result of the target object, the target expression is parsed by using the language parsing module in the REC model, and combined with the self-attention strategy, a vector sequence containing learnable vectors is obtained, and the target image is respectively appearance encoded and semantically encoded based on the graph collaborative attention module, and the node features and node weights, edge features and edge weights of the target image are obtained in combination with the vector sequence. The node features and node weights, edge features and edge weights of the target image are multi-step inferenced through the language parsing module, and the matching score between the appearance graph and the composite node in the category graph is further calculated by the matching module, and the object corresponding to the node with the highest matching score is taken as the positioning result of the target object.
2. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: Obtaining the vector sequence including the learnable vectors includes: Parsing the target expression using the language parsing module, embedding each word in the target expression into a vector through a nonlinear mapping function, obtaining an embedding vector set, encoding the embedding vectors in the embedding vector set, and generating an initial vector sequence; Based on the self-attention strategy, the weights of the nouns in the target expression are determined to generate a weight group: in, For weight reorganization, For the initial vector sequence, L is L weights, W n is the learnable vector a is the appearance map, c is the category map, and T is the transpose; The weight group is converted into a learnable vector and applied to the initial vector sequence to obtain the vector sequence containing the learnable vector: Among them, τ n is a vector sequence containing learnable vectors, For the An initial vector sequence, Reorganization for weights.
3. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: Acquiring the node features and node weights of the target image includes: Step 1.1, obtaining node features of the appearance graph: in, is the node feature, MLP a is a multi-layer perceptron, o i is the visual feature of the i-th candidate region, b i is the box feature of the i-th candidate region, ∈ a is a trainable bias term, To express Characteristics of a noun; Step 1.2: Calculate the degree of match between the candidate region and the noun matched in the expression to obtain the best matching noun: in, is the best matching noun, is the matching degree, W a,1 , W a,2 and W τ is a trainable parameter matrix, N is the number of nouns, τ n is a vector sequence containing learnable vectors; Step 1.3: assign a weight to the node according to the best matching noun: in, For the i-th target and the The result of similarity calculation of nouns is: For the kth target and the The result of similarity calculation of nouns; Step 1.4, replacing the appearance graph with a category graph, that is, replacing the visual features in the appearance graph with text features, and repeating steps 1.1-1.3 until the node features and node weights of the category graph are obtained.
4. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: Acquiring edge features and edge weights of the target image includes: Step 2.1, obtaining edge features of the appearance graph: in, is the trainable parameter matrix; Step 2.2: Calculate the correlation between the edge feature and the expression: in, W q and is the trainable parameter matrix, and q is the representation of the entire expression; Step 2.3: According to the correlation degree, aggregate node features, and normalize the weights of the edges connected to the node to obtain the edge weights: Among them, 1 [i≠j] is the indicator function, K is the target in K images; Step 2.4, replacing the appearance graph with a category graph, that is, replacing the visual features in the appearance graph with text features, and repeating steps 2.1-2.3 until the edge features and edge weights of the category graph are obtained.
5. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: The multi-step reasoning of the node features and node weights, edge features and edge weights of the target image includes: Among them, W t is the trainable parameter matrix, θ t is a trainable parameter vector, and They are the node features aggregated from connected neighbor nodes in the t-1th reasoning step and the self-loop transformation features of the i-node.
6. The zero-sample-based image reference expression understanding method according to claim 5, characterized in that: Obtaining node features aggregated from connected neighboring nodes and self-loop transformation features of the i-node in the t-1-th reasoning step includes: in, and are two trainable parameter matrices, and is a trainable parameter vector.
7. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: Calculating the matching score between the composite node in the appearance graph and the class graph includes: Among them, *∈{a,c}, a is the appearance map, c is the category map, is the feature of the i-th node at the T-th iteration, and are two trainable parameter matrices, and ||·|| represents the L2 norm of the vector.
8. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: The loss function expression related to the self-attention strategy is: Among them, W e is the trainable parameter matrix, L e is the expression loss function, Indicates the pth n The word index corresponding to the noun, τ n is the expression for the nth noun.
9. The zero-sample based image reference expression understanding method according to claim 1, characterized in that: The total loss function expression is: in, and G a and G c The fraction of the true label GT nodes in , is the kth score in the appearance graph, is the kth score in the category graph.