Image-text cross-modal retrieval model, method and computer device based on local shared semantic center
By adopting a method based on local shared semantic center in cross-modal retrieval of image text, the problem of large amount of local alignment calculation and fine-grained semantic correspondence in the prior art is solved, and efficient image text retrieval and deep semantic understanding are achieved.
Patent Information
- Application Number
- CN202210718696.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-23
AI Technical Summary
Existing cross-modal retrieval methods for image text are computationally expensive in the process of local alignment, making it difficult to deeply explore the fine-grained semantic correspondence between images and text.
Using a method based on local shared semantic center, by learning the trainable image text shared semantic center, the fine-grained alignment of images and text is achieved, reducing the direct interaction of local features, thereby reducing the computational cost.
Effectively understand the semantic correspondence between images and text in a deeper way, improve retrieval efficiency, reduce the computational burden caused by local alignment, and provide auxiliary information through global alignment to improve retrieval accuracy.
Smart Images

Figure CN114969423B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image-text cross-modal retrieval, and in particular relates to an image-text cross-modal retrieval model, method and computer device based on a local shared semantic center. Background Art
[0002] Image-text cross-modal retrieval aims to use data in one modality to retrieve data in another modality that has the same semantics as the data. It is an important research direction in the fields of machine vision, natural language processing, and multimodal learning, and has become a research hotspot at home and abroad. In recent years, with the development of deep learning technology, image-text cross-modal retrieval has achieved excellent results. However, this task still faces huge challenges, because it not only requires a deep understanding of the semantic knowledge of images and texts, but also needs to cross the modality gap and obtain the semantic correspondence between different modalities.
[0003] To address the above challenges, current methods focus more on fine-grained correspondence between images and texts, highlight important semantic knowledge through local alignment, and learn images and texts more comprehensively. However, current methods ignore the heavy computational burden brought by local alignment. Therefore, while fully understanding images and texts, it is very important to reduce the interaction scale of local features for cross-modal image-text retrieval.
[0004] Recently, cluster learning methods have achieved great success in optimizing the common semantic representation of features. However, most of the current cluster feature learning focuses on global representation, thus ignoring fine-grained local information and failing to cope well with the challenge of cross-modal retrieval of images and texts. Therefore, the present invention designs a cluster center shared by images and texts, and adopts a soft assignment strategy to achieve fine-grained alignment between images and texts, thereby deeply understanding the semantic correspondence between images and texts and improving retrieval efficiency. Summary of the invention
[0005] The purpose of the present invention is to use the semantic center shared by trainable images and texts to represent the semantic commonality of local features of images and texts, and to achieve fine-grained alignment of images and texts through the semantic center, thereby mining deep-level image semantics and text semantics, avoiding direct interaction between local features of images and texts, and thus reducing the scale of calculations. It is also proposed to use global alignment as a supplement to local alignment, to achieve cross-modal semantic correspondence of images and texts from multiple angles, and to more comprehensively summarize semantic information. The technical solution for implementing the present invention is as follows:
[0006] A cross-modal image-text retrieval model based on local shared semantic centers is obtained through the following steps:
[0007] S1, extracts the regional features of the image and the word-level features of the text respectively, and then obtains the image features and text features for local alignment and global alignment respectively through two layers of independent mapping.
[0008] S2, clustering the image features and text features in step S1 to obtain k initialized shared semantic centers;
[0009] S3, calculating the similarity between the image text features in step S1 and the shared semantic center in step S2, and using the similarity to aggregate the image features into k image semantic representations corresponding to the shared semantic center, and aggregate the text features into k text semantic representations corresponding to the shared semantic center;
[0010] S4, modeling the pooling operation of the regional features of the image and the word-level features of the text in step 1 to obtain a global image representation and a global text representation;
[0011] S5, using the image semantic representation and text semantic representation with the same shared semantic center in step S3 to calculate the local similarity of the image and text, using the image global representation and text global representation in step S4 to calculate the global similarity of the image and text, the overall similarity of the image and text is represented by the weighted sum of the local similarity and the global similarity, and the modeling is completed.
[0012] S6, uses the overall similarity to train the image-text cross-modal retrieval model, and uses the trained model to perform real-time image-text cross-modal retrieval.
[0013] As a preferred technical solution, the specific process of image text feature extraction in step S1 includes:
[0014] Step 51-1, use the pre-trained Faster-RCNN to extract the regional features of the image, and pass the extracted regional features through two independent multi-layer perceptrons to map them respectively to obtain two sets of image features and
[0015] Step S1-2, the input text sentence is divided into words, and then filled with 0 to a fixed word length, the divided and filled text is sent to the pre-trained Bert to obtain word-level feature representation, and then two independent multi-layer perceptrons are used to map and obtain two sets of text features respectively. and
[0016] As a preferred technical solution, the specific process of initializing the semantic center in step S2 includes:
[0017] Step S2-1, randomly sampling image features and text features in the training data set;
[0018] Step S2-2: Perform K-means clustering on the randomly sampled image features and text features to obtain k initialized cluster centers. And k<<n;
[0019] Step S2-3, define the initialized cluster center C as a trainable shared semantic center and train it together with the model.
[0020] As a preferred technical solution, the specific process of obtaining the image-text alignment semantic representation in step S3 includes:
[0021] Step S3-1, for the image feature V in step S1-1 l and the text feature T in step S1-2 l Calculate the cosine distance with the shared semantic center C in step S2-3 respectively to obtain the similarity matrix between the image and the shared semantic center, and between the text and the shared semantic center, perform softmax operation on the similarity matrix, and obtain a normalized similarity matrix;
[0022] Step S3-2, using the value of the normalized similarity matrix in step S3-1 as the image feature V in step S1-1 l The weight of the text feature T in step S1-2 l The weights of the intra-modal features are accumulated as the image features and text features corresponding to the semantic center. Since there are k semantic centers in step S2-3, the number of image features and text features aligned according to the semantic centers is k.
[0023] As a preferred technical solution, the specific process of obtaining the global representation of the image text in step S4 includes:
[0024] Step S4-1, for the image feature V in step S1-1 g and the text feature T in step S1-2 g Perform different pooling operations, such as maximum pooling, second value pooling, minimum pooling, etc., to obtain pooled image features and text features;
[0025] Step 54-2, use bi-GRU to model the pooling features of the image and text respectively, find the coefficients required for the optimal pooling, and then obtain the global features of the image and text according to the optimal pooling strategy.
[0026] As a preferred technical solution, the specific process of calculating the image-text similarity and model training in step S5 includes:
[0027] S5-1, the fine-grained knowledge in the image and text has been aligned through step S3-2, and the local similarity between the image and the text corresponding to a certain semantic center is represented by the cosine distance of the image feature and the text feature aligned with the same semantic center in step S3-2. The sum of the local similarities of all semantic center alignments is calculated as the local similarity between the image and the text;
[0028] S5-2, the global similarity between the image and the text is represented by the cosine similarity of the global features of the image and the text in step S4-2.
[0029] S5-3,Finally, the overall similarity is expressed as the weighted sum of local similarity and global similarity. According to the overall similarity, ternary ranking loss is used for training.
[0030] As a preferred technical solution, the method process of image-text cross-modal retrieval in step S6 includes:
[0031] For any set of image-text pairs, first use the feature extraction method of step S1 to extract the features of the image and text, then extract the local features and global features of the image and the local features and global features of the text according to steps S3 and S4, and use the extracted global features and local features to perform local and global alignment of the image and text according to the method of step S5, and calculate the similarity between the image and text to obtain the retrieval result.
[0032] A computer device having built-in instructions or programs for the image-text cross-modal retrieval model based on local shared semantic centers, or instructions or programs for the image-text cross-modal retrieval method.
[0033] Beneficial effects of the present invention:
[0034] (1) The present invention solves the problem that the traditional local alignment of image-text cross-modal retrieval has large computational complexity and cannot deeply explore the fine-grained relationship between image and text.
[0035] (2) The image-text cross-modal retrieval method based on local shared semantic centers proposed in the present invention learns a set of trainable semantic centers shared by images and texts, so that the local features of images and texts can be indirectly aligned through the semantic centers, thereby deeply exploring the semantic relationship between images and texts and reducing the interaction cost brought by local alignment.
[0036] (2) The present invention applies soft assignment to the cluster matching problem. Since soft assignment makes the weight coefficient smooth and differentiable, the cluster center can be trained end-to-end with the model, thereby generating a reliable shared semantic center.
[0037] (3) Based on local alignment, the present invention uses the alignment of global features as auxiliary information to promote semantic matching between images and texts, and understands the relationship between images and texts from both local and global perspectives, thereby improving computational retrieval performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flowchart of image-text cross-modal retrieval based on local shared semantic centers. DETAILED DESCRIPTION
[0039] The present invention first extracts the regional features of the image and the word-level features of the text, and after two layers of independent mapping, obtains the image features and text features for local alignment and global alignment respectively. A clustering method is used to obtain an initialized cluster center group, and the cluster center is set as a trainable shared semantic center, which is updated as the network is trained. The image text features are aligned to the corresponding shared semantic center by the cosine distance between the image text features and the shared semantic center, thereby obtaining the same number of image local features and text local features as the shared semantic center. The global features of the image and the global features of the text are calculated by a method of modeling the image text feature pooling operation. The local features of the image text are used for local alignment, and the global features of the image text are used for global alignment, and finally the multi-angle image text similarity is obtained, and the model is trained using the ternary sorting loss.
[0040] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] Figure 1 This is a flowchart of the image-text cross-modal retrieval method based on local shared semantic centers proposed by the present invention. The present invention first extracts features from the image text, then defines a set of trainable shared semantic centers, calculates the local features of the image-text alignment through the relationship between the image text and the shared semantic centers, models the global features of the image text by the pooling method of the image-text features, calculates the overall image-text similarity using local alignment and global alignment, and finally uses the ternary ranking loss for training, specifically including the following steps:
[0042] S1, image-text feature extraction: extract the regional features of the image and the word-level features of the text respectively, and then obtain the image features and text features for local alignment and global alignment respectively through two layers of independent mapping.
[0043] Specific implementation: Use the pre-trained Faster-RCNN to extract the regional features of the image, and use the extracted regional features as two independent multi-layer perceptrons MLP Vl and MLP Vg Input, mapping to obtain two sets of image features and
[0044] The input text sentence is then segmented into words, padded with 0 to a fixed word length, and the segmented and padded text is fed into the pre-trained Bert to obtain word-level feature representations. The extracted word-level feature representations are then used as two independent layers of multi-layer perceptrons (MLPs). Tl and MLP Tg Input, respectively mapped to obtain two sets of text features and
[0045] S2, initialize shared semantic centers: perform K-Means clustering on the image features and text features in step S1 to obtain k initialized shared semantic centers.
[0046] Specific implementation: First, randomly sample the image features and text features in the training data set to obtain several untrained image features and text features, and then perform K-means clustering on the randomly sampled image features and text features to obtain k initialized cluster centers. And k<<n, then the initialized cluster center C is defined as a trainable shared semantic center, and the parameters of the shared semantic center are updated as the network is trained.
[0047] S3, learning the aligned semantic representation of image text: calculating the similarity between the image text features in step S1 and the shared semantic center in step S2, and using the similarity to aggregate the image features into k image semantic representations corresponding to the shared semantic center, and aggregate the text features into k text semantic representations corresponding to the shared semantic center.
[0048] Specific implementation: For the image feature V in step S1 l and text features T l The cosine distance is calculated with the shared semantic center C in step S2 to obtain the similarity matrix between the image and the shared semantic center, and between the text and the shared semantic center. A softmax operation is performed on the similarity matrix to obtain a normalized similarity matrix.
[0049] Then the value of the normalized similarity matrix in step S3 is used as the image feature V in step S1 l The weight and text feature T l The weights of the intra-modal features are accumulated as the image features and text features corresponding to the semantic center. Since there are k semantic centers in step S2, the number of image features and text features aligned according to the semantic centers is k.
[0050] S4, learning global representation of image and text: use bi-GRU to model the regional features of the image and the word-level features of the text in step 1 to obtain the optimal global representation of the image and the global representation of the text.
[0051] Specific implementation: For the image feature V in step S1 g and text features T g Different pooling operations are performed, such as maximum pooling, second value pooling, minimum pooling, etc., to obtain pooled image features and text features.
[0052] The bi-GRU is further used to model the pooling features of the image and text respectively, and the coefficients required for the optimal pooling are obtained, and then the global features of the image and text are obtained according to the optimal pooling strategy.
[0053] S5, image-text similarity calculation: the local similarity of the image and text is calculated using the image semantic representation and text semantic representation with the same shared semantic center in step S3, and the global similarity of the image and text is calculated using the image global representation and text global representation in step S4. The overall similarity between the image and text is represented by the weighted sum of the local similarity and the global similarity.
[0054] Specific implementation: The fine-grained knowledge in the image and text has been aligned through step S3. The local similarity between the image and the text corresponding to a certain semantic center is represented by the cosine distance of the image features and text features aligned with the same semantic center in step S3. The sum of the local similarities of all semantic center alignments is calculated as the local similarity between the image and the text.
[0055] The global similarity between the image and the text is represented by the cosine similarity of the global features of the image and the text in step S4. The overall similarity between the image and the text is calculated by the weighted sum of the local similarity and the global similarity, and finally the training is performed using the ternary ranking loss based on the overall similarity.
[0056] The present invention is described below by means of specific embodiments. The implementation of the present invention includes the model establishment and training process and the image text retrieval process, which are described in detail below.
[0057] 1. The model establishment and training process includes the following:
[0058] 1.1 Feature extraction process of image text
[0059] The regional features of the image and the word-level features of the text are extracted using the pre-trained Faster R-CNN and pre-trained Bert, respectively. In order to align the image and text from a local and global perspective, both the image features and the text features are extracted using two independent multi-layer perceptrons.
[0060] 1.1.1 Image Feature Extraction
[0061] Given an image I, use the pre-trained Faster R-CNN to detect region r in the image i , and extract each region r i The characteristic f i Then two independent multi-layer perceptrons are used to transform the regional features f i Mapped separately and
[0062]
[0063]
[0064] In formula (1) and (2), MLP Vl 、MLP Vg Represents two independent multi-layer perceptrons, which respectively obtain image features for local alignment and global alignment, expressed as and
[0065] 1.1.2 Feature Extraction of Text
[0066] Given a text S, first use the word segmentation tool to divide the text into multiple independent words and fill the words with 0 to a fixed length. i Input into the pre-trained Bert to obtain word-level text features z i Then two independent multi-layer perceptrons are used to transform the word-level features z of the text i Mapped separately and
[0067] z i =Bert(s i )#(3)
[0068]
[0069]
[0070] In formula (4) and (5), MLP Tl 、MLP Tg Represents two independent multi-layer perceptrons, and the text features for local alignment and global alignment are represented as and
[0071] 1.2 Initialization of the semantic center
[0072] First, the image features V used for local alignment in the training dataset are l and text features T for local alignment l Perform random sampling to obtain several untrained image features and text features, and then perform K-means clustering on the randomly sampled image features and text features to obtain k initialized cluster centers And k<<n, then the initialized cluster center C is defined as a trainable shared semantic center, and the parameters of the shared semantic center are updated as the network is trained.
[0073] 1.3 Aligned Semantic Representation of Image and Text
[0074] According to the semantic commonalities between the image text and the shared semantics, semantically aligned image context features and text context features are obtained. Since the local features of the image and the local features of the text are aligned based on the shared semantic center, the local similarity between the image and the text can be represented by the context features of the image and the text under the same shared semantic center.
[0075] 1.3.1 Obtaining aligned semantic representation of images
[0076] In order to obtain the image context features aligned with the shared semantic center, the cosine similarity between the image features and the shared semantic center is calculated:
[0077]
[0078] In formula (6) represents the transpose of the i-th shared semantic center, represents the jth image feature used for local alignment, Represents the cosine similarity between the i-th shared semantic center and the j-th image feature used for local alignment. The cosine similarity matrix is operated with softamx to obtain the normalized similarity matrix:
[0079]
[0080] In formula (7), λ represents the temperature coefficient, Represents the normalized cosine similarity, As The weight of the corresponding semantic center c is calculated i Local features of the image:
[0081]
[0082] In formula (8), Refers to the i-th shared semantic center c iThe image context features of the image are used to obtain the shared semantically aligned image features.
[0083] 1.3.2 Obtaining aligned semantic representation of text
[0084] As in step 1.3.1, in order to obtain the text context features aligned with the shared semantic center, calculate the cosine similarity between the text features and the shared semantic center:
[0085]
[0086] In formula (9) represents the transpose of the i-th shared semantic center, represents the jth text feature used for local alignment, Represents the cosine similarity between the i-th shared semantic center and the j-th text feature used for local alignment. The cosine similarity matrix is operated with softamx to obtain the normalized similarity matrix:
[0087]
[0088] In formula (10), λ represents the temperature coefficient, Represents the normalized cosine similarity, As The weight of the corresponding semantic center c is calculated i Text context features:
[0089]
[0090] In formula (11), Refers to the i-th shared semantic center c i The text context features of
[0091] 1.4 Global Representation of Image Text
[0092] Global alignment of image and text provides more general and comprehensive semantic information for understanding the shared semantics of image and text than local alignment. Therefore, semantic alignment from a global perspective can be regarded as auxiliary information for image and text alignment.
[0093] 1.4.1 Extracting global features of images
[0094] Perform multiple pooling on the image features used for global alignment in step 1.1.1 to obtain a global representation of multiple images:
[0095]
[0096] In formula (12), Represents the pooling result of image features, max i Represents the characteristics Perform the i-th value pooling, for example, when i=1, max 1 It means to perform maximum pooling on image features. The result of maximum pooling. In order to find the best pooling strategy, bi-GRU is used to model all pooling results to approximate maximum pooling, second value pooling, average pooling or more complex pooling results:
[0097]
[0098] In formula (13) Represents the positional encoding of image features, Represents the output of the bi-GRU corresponding to the position encoding, with a dimension of Each position encoding corresponds to the output of bi-GRU They are all d-dimensional features, and the fully connected layer is used to map their dimensions to Then use softmax to normalize:
[0099]
[0100] In formula (14), w v represents the weight matrix of the fully connected layer, b v Represents the bias of the fully connected layer. Since the output dimension of the fully connected layer is So w v The dimension is b v The dimension is Represents the weight coefficient corresponding to the i-th value pooling result. The global features of the image are represented by the weighted sum of the pooling results:
[0101]
[0102] 1.4.2 Extracting a global representation of text
[0103] As in step 1.4.1, perform multiple pooling on the text features used for global alignment in step 1.1.1 to obtain a global representation of multiple texts:
[0104]
[0105] In formula (18), Represents the pooling result of text features, max i Represents the characteristics Perform the i-th value pooling, for example, when i=1, max 1It means to perform maximum pooling on text features. The result of maximum pooling. In order to find the best pooling strategy, bi-GRU is used to model all pooling results to approximate maximum pooling, second value pooling, average pooling or more complex pooling results:
[0106]
[0107] In formula (17) Represents the positional encoding of text features, Represents the output of the bi-GRU corresponding to the position encoding, with a dimension of Use a fully connected layer to map its dimensions to Then use softmax to normalize:
[0108]
[0109] In formula (18), w t represents the weight matrix of the fully connected layer, b t Represents the bias of the fully connected layer. Since the output dimension of the fully connected layer is So w t The dimension is b t The dimension is Represents the weight coefficient corresponding to the pooling result of the i-th value. The global feature of the text is represented by the weighted sum of the pooling results:
[0110]
[0111] 1.5 Image-text similarity calculation
[0112] Since the local features of image and text have been aligned by the shared semantic center, the local similarity between image and text can be calculated by the image context features and text context features under the same shared semantics; the global similarity between image and text is calculated by the global features of image and text as auxiliary information to improve retrieval accuracy.
[0113] 1.5.1 Local Similarity of Image and Text
[0114] Steps 1.3.1 and 1.3.2 extract the image context features and text context features that are aligned with shared semantics, respectively. The local similarity is represented by the cosine similarity of the image and text context features:
[0115]
[0116] In formula (20) Indicates that in the shared semantic center ci Contextual features of the image below Contextual features of the text The cosine similarity between them. The sum of the similarities of all aligned semantic centers is taken as the local similarity between the image and the text:
[0117]
[0118] 1.5.2 Global Similarity of Image and Text
[0119] Steps 1.4.1 and 1.4.2 extract the global representations of the image and text respectively. The global similarity is represented by the cosine similarity of the global representations of the image and text:
[0120]
[0121] In formula (22), R g (v, t) represents the global feature g of the image v and the global feature g of the text t The cosine similarity between .
[0122] 1.5.3 Overall similarity of image and text
[0123] According to steps 1.5.1 and 1.5.2, the local similarity and global similarity between the image and the text have been obtained. The overall similarity between the image and the text is determined by the local similarity and the global similarity:
[0124] R(v, t) = β 1 R l (v,t)+β 2 R g (v, t)#(23)
[0125] In formula (23), β 1 and β 2 is a hyperparameter that determines the local and global ratios. In practice, β 1 Set to 0.2, and β 2 Setting it to 1 can achieve better results. According to the obtained similarity, the ternary ranking loss is used for training:
[0126]
[0127] In formula (24), Δ is a hyperparameter. In practice, setting Δ to 0.15 can achieve better results. (v, t) represents the dataset The positive sample pairs in represents the hardest negative sample of v, It means that under the condition t′≠t, when When , R(v, t′) is maximized. represents the hardest negative sample of t, It means that under the condition v′≠v, when When , R(v, t′) is maximized. [x]+≡max(0, x), the distance between positive sample pairs is shortened by using the ternary ranking loss; where t′ and v′ are intermediate variables.
[0128] 2. Image-text cross-modal retrieval process
[0129] After the model is fully trained on the training set, for any image to be tested, the similarity between the image and all the texts in the test library is calculated using formula (23), and the text with the greatest similarity is retrieved as the retrieval result; given a piece of text to be tested, the similarity between the text and all the images in the test library is calculated using formula (23), and the image with the greatest similarity is retrieved as the retrieval result.
[0130] In summary, the present invention discloses a method for cross-modal retrieval of images and texts based on local shared semantic centers. The method performs cross-modal semantic alignment of images and texts from the perspectives of local alignment and global alignment. For local alignment, a semantic center shared by images and texts is trained, which describes the semantic commonalities of local features of images and texts. Therefore, the local features of images and texts can be aligned according to the same semantic center. This local alignment method ignores the complex direct interaction of local features, reduces the amount of local alignment calculations while mining fine-grained semantic information; for global alignment, images and texts are represented using global features, which can obtain more comprehensive semantic knowledge, and can be used as auxiliary information to improve the accuracy of cross-modal retrieval. Therefore, the present invention solves the problems of redundant calculation and low recall rate of local alignment in cross-modal retrieval of images and texts.
[0131] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. All equivalent methods or changes that do not deviate from the technical creation of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing an image-text cross-modal retrieval model based on local shared semantic centers. It is characterized in that The method comprises the following steps: S1, extracts the regional features of the image and the word-level features of the text respectively, and then obtains the image features and text features for local alignment and global alignment respectively through two layers of independent mapping; S2, cluster the image features and text features in S1 to obtain k initialized shared semantic centers; S3, obtaining image-text aligned semantic representation: calculating the similarity between the image-text features in S1 and the shared semantic center in step S2, and using the similarity to aggregate the image features into k image-aligned semantic representations corresponding to the shared semantic center, namely, the image context features, and aggregate the text features into k text-aligned semantic representations corresponding to the shared semantic center, namely, the text context features; S4, modeling the pooling operation of the regional features of the image and the word-level features of the text in step 1 to obtain a global image representation and a global text representation; S5, using the image semantic representation and text semantic representation with the same shared semantic center in step S3 to calculate the local similarity of the image and text, using the image global representation and text global representation in step S4 to calculate the global similarity of the image and text, the overall similarity of the image and text is represented by the weighted sum of the local similarity and the global similarity, and the modeling is completed; The specific implementation of S1 includes: S1.1 Image feature extraction Given an image I, use the pre-trained Faster R-CNN to detect region r in the image i , and extract each region r i The characteristic f i , and then use two independent multi-layer perceptrons to transform the image region features f i Mapped separately and In formula (1) and (2), MLP Vl 、MLP Vg Represents two independent multi-layer perceptrons, which respectively obtain image features for local alignment and global alignment, expressed as and S1.2 Feature extraction of text Given a text S, first use the word segmentation tool to divide the text into multiple independent words, and fill the words with 0 to a fixed length. Input the fixed-length word sequence into the pre-trained Bert to obtain word-level text features, and then use two independent multi-layer perceptrons to convert the word-level features of the text into i Mapped separately and z i =Bert(s i )(3) In formula (3), Bert represents the pre-trained Bert network, s i Represents the original input text, z i Represents the word-level features of the text extracted by Bert. In formulas (4) and (5), MLP Tl 、MLP Tg Represents two independent multi-layer perceptrons, and the text features for local alignment and global alignment are represented as and The specific implementation of S2 includes: S2.1 Image features V used for local alignment in the training dataset l and text features T for local alignment l Perform random sampling to obtain several untrained image features and text features. S2.2 Perform K-means clustering on randomly sampled image features and text features to obtain k initialized cluster centers k<<n and k<<n, S2.3 defines the initialized cluster center C as a trainable shared semantic center, and the parameters of the shared semantic center are updated as the network is trained.
2. According to the method for constructing an image-text cross-modal retrieval model based on local shared semantic center according to claim 1, It is characterized in that The specific implementation of S3 includes: S3.1 Obtaining aligned semantic representation of images In order to obtain the image context features aligned with the shared semantic center, the cosine similarity between the image features and the shared semantic center is calculated: In formula (6) represents the transpose of the i-th shared semantic center, represents the jth image feature used for local alignment, Represents the cosine similarity between the i-th shared semantic center and the j-th image feature used for local alignment. The cosine similarity matrix is operated with softamx to obtain the normalized similarity matrix: In formula (7), λ represents the temperature coefficient, Represents the normalized cosine similarity, As The weight of the corresponding semantic center c is calculated i Local features of the image: In formula (8), Refers to the i-th shared semantic center c i The image context features of the image are used to obtain the shared semantically aligned image features. S3.2 Obtaining aligned semantic representation of text As in step S3.1, in order to obtain the text context features aligned with the shared semantic center, the cosine similarity between the text features and the shared semantic center is calculated: In formula (9) represents the transpose of the i-th shared semantic center, represents the jth text feature used for local alignment, Represents the cosine similarity between the i-th shared semantic center and the j-th text feature used for local alignment. The cosine similarity matrix is operated with softamx to obtain a normalized similarity matrix: In formula (10), λ represents the temperature coefficient, Represents the normalized cosine similarity, As The weight of the corresponding semantic center c is calculated i Text context features: In formula (11), Refers to the i-th shared semantic center c i The text context features of 3. According to claim 1, a method for constructing an image-text cross-modal retrieval model based on local shared semantic center, It is characterized in that The specific implementation of S4 includes: S4.1 Extracting global features of images Perform multiple pooling on the image features used for global alignment in step 1 to obtain a global representation of multiple images: In formula (12), max i Represents the characteristics Perform the i-th value pooling, Use bi-GRU to model all pooling results to approximate different pooling results: In formula (13) Represents the positional encoding of image features, Represents the output of the bi-GRU corresponding to the position encoding, with a dimension of Use a fully connected layer to map its dimensions to Then use softmax to normalize: In formula (14), w v represents the weight matrix of the fully connected layer, b v Represents the bias of the fully connected layer. Since the output dimension of the fully connected layer is So w v The dimension is b v The dimension is The global features of the image are represented by the weighted sum of the pooling results: S4.2 Extracting global representation of text As in step S4.1, the text features used for global alignment in step 1 are multi-pooled to obtain global representations of multiple texts: In formula (18), max i Represents the characteristics Perform the i-th value pooling, Use bi-GRU to model all pooling results to approximate the pooling results of different pooling: In formula (17) Represents the positional encoding of image features, Represents the output of the bi-GRU corresponding to the position encoding, with a dimension of Use a fully connected layer to map its dimensions to Then use softmax to normalize: In formula (18), w t represents the weight matrix of the fully connected layer, b t Represents the bias of the fully connected layer. Since the output dimension of the fully connected layer is So w t The dimension is b t The dimension is Represents the weight coefficient corresponding to the pooling result of the i-th value. The global feature of the text is represented by the weighted sum of the pooling results:
4. According to claim 3, a method for constructing an image-text cross-modal retrieval model based on local shared semantic center, It is characterized in that The specific implementation of S5 includes: S5.1 Local similarity between image and text The local similarity is represented by the cosine similarity of the image-text context features: In formula (20) Indicates that in the shared semantic center c i Contextual features of the image below Contextual features of the text The cosine similarity between the two images is taken as the sum of the similarities of all aligned semantic centers as the local similarity between the image and the text: S5.2 Global Similarity of Image and Text The global similarity is represented by the cosine similarity between the global representation of the image and the global representation of the text: In formula (22), R g (v,t) represents the global feature g of the image v and the global feature g of the text t The cosine similarity between ; S5.3 Overall similarity of image and text According to the local similarity and global similarity between the image and text obtained in steps S5.1 and S5.2, the overall similarity between the image and the text is determined by the local similarity and the global similarity: R(v,t)=β 1 R l (v,t)+β 2 R g (v,t)(23) In formula (23), β 1 and β 2 is a hyperparameter that determines the local-global ratio.
5. A method for constructing an image-text cross-modal retrieval model based on local shared semantic center according to any one of claims 1 to 4, It is characterized in that The overall similarity is used to train the image-text cross-modal retrieval model; the details are as follows: Based on the obtained overall similarity, the ternary ranking loss is used for training: In formula (24), Δ is a hyperparameter, and (v, t) represents the dataset The positive sample pairs in represents the hardest negative sample of v, represents the hardest negative sample of t, [x] + ≡max(0,x), using the ternary ranking loss to shorten the distance between positive sample pairs.
6. The image-text cross-modal retrieval method based on the local shared semantic center image-text cross-modal retrieval model according to any one of claims 1 to 4, It is characterized in that For any image to be tested, it is input into the model constructed by any one of claims 1-4, the overall similarity between the image and all texts in the model test library is calculated, and the text with the greatest similarity is retrieved as the retrieval result; for any section of text to be tested, the similarity between the text and all images in the test library is calculated, and the image with the greatest similarity is retrieved as the retrieval result.
7. A computer device, It is characterized in that The computer device has built-in instructions or programs for the image-text cross-modal retrieval model based on local shared semantic center constructed according to any one of claims 1 to 4, or instructions or programs for the image-text cross-modal retrieval method according to claim 6.
Citation Information
Patent Citations
Cross-modal image text retrieval method based on credibility self-adaptive matching network
CN111026894A
Cross-modal image text retrieval method of hybrid fusion model
CN112784092A