A method for image text retrieval based on multi-level network

By building a multi-level network in image text retrieval, training subnets at the global level, relationship level and digital level respectively, and fusing their similarity, the problems of low efficiency and low accuracy of image text retrieval in the prior art are solved, and more efficient and accurate retrieval effects are achieved.

CN114357148BActive Publication Date: 2025-06-06ZHEJIANG LAB +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111619401.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-06-06
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

Existing image text retrieval methods are inefficient and have low accuracy, and cannot effectively capture fine-grained semantic information of images and text, especially ignore the alignment of digital-level information.

Method used

Using image text retrieval method based on multi-level networks, global-level subnets, relationship-level subnets and digital-level subnets are constructed, and these subnets are trained to obtain the global level similarity, relationship-level similarity and digital-level similarity of images and text, and fuse them to generate multi-level overall similarity.

Benefits of technology

By aligning global information, fine-grained relationship information and digital information, the accuracy and efficiency of image text retrieval is significantly improved, the alignment process is simplified, and the difficulty of model training is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114357148B_ABST
    Figure CN114357148B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image text retrieval technology, and in particular to an image text retrieval method based on a multi-level network, comprising: constructing an image text retrieval model having a global level subnetwork, a relation level subnetwork, and a digital level subnetwork; constructing a training data set for training the image text retrieval model, wherein the training data set includes image text pairs; inputting the image text pairs in the training data set into the global level subnetwork, the relation level subnetwork, and the digital level subnetwork of the image text retrieval model respectively, so as to generate corresponding global level similarity, relation level similarity, and digital level similarity respectively, and train the corresponding subnetworks separately; performing image text retrieval based on the trained image text retrieval model. The image text retrieval method in the present invention can improve the retrieval efficiency and retrieval accuracy of image texts, thereby improving the effect of image text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image text retrieval, and in particular to an image text retrieval method based on a multi-level network. Background Art

[0002] Image text retrieval means retrieving the relevant text description sentences in the text library given a search image, or retrieving the corresponding image in the picture library given a text description. It has important applications in many fields, such as large image and video websites, where users input query text and use text image retrieval technology to retrieve images or videos related to the query text, thereby achieving rapid indexing of multimedia data, improving multimedia data management efficiency, and enhancing user experience.

[0003] Existing methods measure the similarity between images and texts by mapping them into a common space, but such methods cannot capture the fine-grained semantic information in images and sentences. To this end, the Chinese patent with publication number CN109255047A discloses "An image-text mutual retrieval method based on complementary semantic alignment and symmetric retrieval", which includes: extracting deep visual features using a model fused with a target-based convolutional neural network and a scene distribution-based convolutional neural network; encoding the text using a long short-term memory network to extract the corresponding semantic feature representation; using two mapping matrices to map the visual features and text features to the same cross-modal embedding space respectively; using the k-nearest neighbor method to retrieve in the cross-modal embedding space to obtain the initial retrieval list; using the mutual nearest neighbor method to symmetric the neighbor relationship of the bidirectional retrieval, reordering the initial retrieval list to obtain the final retrieval ranking list.

[0004] The above-mentioned existing image-text mutual retrieval methods use the interactive information after cross-processing of images and texts to more accurately mine image semantic information and text semantic information. However, the existing methods that integrate another form of contextual information through the cross-attention mechanism to obtain relational information mostly require the execution of image-based attention mechanism alignment and text-based attention mechanism alignment. However, this attention mechanism-based alignment method is very time-consuming, which leads to low efficiency of image-text retrieval. At the same time, the existing methods ignore the numerical information of the image text. For example, the numerical level information such as "four" and "Three" is not aligned, so that the images retrieved by the text "four people are jumping from the top ofstairs" and "Three people are jumping from the top of stairs" are the same, that is, the accuracy of image-text retrieval is not high.

[0005] Therefore, how to design a method that can improve the efficiency and accuracy of image text retrieval is a technical problem that needs to be solved urgently. Summary of the invention

[0006] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is: how to provide an image text retrieval method based on a multi-level network to improve the retrieval efficiency and retrieval accuracy of image text, thereby improving the effect of image text retrieval.

[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0008] An image text retrieval method based on a multi-level network comprises the following steps:

[0009] S1: Construct a graph-text retrieval model with global level subnetwork, relation level subnetwork and number level subnetwork;

[0010] S2: Construct a training dataset for training the image-text retrieval model, which includes image-text pairs;

[0011] S3: Input the image-text pairs in the training dataset into the global level sub-network, the relation level sub-network and the number level sub-network of the image-text retrieval model respectively to generate the corresponding global level similarity, relation level similarity and number level similarity respectively, and then train the corresponding sub-networks separately based on the global level similarity, relation level similarity and number level similarity respectively;

[0012] S4: Perform image text retrieval based on the trained image-text retrieval model.

[0013] Preferably, in step S2, the image-text pairs in the training data set are preprocessed and feature extracted to obtain the image global features and image local features of the corresponding image and the text global features and text local features of the corresponding text.

[0014] Preferably, in step S3, the global level sub-network is trained by the following steps:

[0015] S301: Inputting the image global features and text global features of the image-text pair into the global level sub-network;

[0016] S302: Calculating corresponding global level similarity based on image global features and text global features;

[0017] S303: Calculate the corresponding global level triplet loss based on the global level similarity, and train the global level sub-network through the global level triplet loss.

[0018] Preferably, by formula Sg (v,t)=sim g (g v ,g t ) Calculate the global level similarity;

[0019] The objective function of the global level triplet loss is

[0020] Where: S g (v, t) represents the global level similarity of the image-text pair (v, t); g v Represents the global features of the image; g t Represents the global features of the text; L global represents the global level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter; N represents the number of image-text pairs.

[0021] Preferably, in step S3, the relationship level sub-network is trained by the following steps:

[0022] S311: inputting the image local features and text local features of the image-text pair into the relation level sub-network;

[0023] S312: The local features of the image are passed through a Transformer encoder to aggregate the relationship information between each image region and generate corresponding image relationship features;

[0024] S313: The local features of the text are passed through the Bert module to integrate the relationship information between words and generate corresponding text relationship features;

[0025] S314: Calculating corresponding relationship level similarities based on image relationship features and text relationship features;

[0026] S315: Calculate the corresponding relationship-level triplet loss based on the relationship-level similarity, and train the relationship-level sub-network through the relationship-level triplet loss.

[0027] Preferably, the Transformer encoder includes a multi-head self-attention mechanism layer for calculating attention multiple times, and a fully connected forward feedback layer for obtaining rich semantic feature representation; the Transformer encoder can aggregate the relationship information between each image region to generate corresponding image relationship features, and make each feature of the image relationship feature contain the semantic information of the image region and the relationship information with other regions.

[0028] Preferably, by formula Calculate relation level similarity;

[0029] The objective function of the relation-level triple loss is L relation =max(0,mS r (v + ,t - )+S r (v + ,t - ));

[0030] Where: S r (v, t) represents the relation-level similarity between the image-text pair (v, t); s ij r Indicates the similarity between the i-th feature in the image relation feature and the j-th feature in the text relation feature; L relation represents the relation-level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

[0031] Preferably, in step S3, the digital level sub-network is trained by the following steps:

[0032] S321: Inputting the image local features and text local features of the image-text pair into the digital level sub-network;

[0033] S322: Calculate the similarity between each region of the image based on the local features of the image to obtain the corresponding image region similarity matrix; then select similar regions with similarity greater than γ from the image region similarity matrix to form a similar region set; finally, convert the quantity information of the similar region set into a vector through the Bert module and fuse it with the local features of the image in proportion to generate the corresponding image digital features;

[0034] S323: Calculate the similarity between each word in the text based on the local features of the text to obtain a corresponding text word similarity matrix; then select similar words with a similarity greater than γ from the text word similarity matrix to form a similar word set; finally, convert the quantity information of the similar word set into a vector through the Bert module and fuse it with the local features of the text in proportion to generate a corresponding text digital feature;

[0035] S324: Calculating corresponding digital level similarities based on the image digital features and the text digital features;

[0036] S325: Calculate the corresponding digit-level triplet loss based on the digit-level similarity, and train the digit-level sub-network through the digit-level triplet loss.

[0037] Preferably, by formula S ij v =sim(l i v ,l j v ) Calculate the image region similarity matrix;

[0038] The set of similar regions is represented as

[0039] The digital features of the image are represented as in,

[0040] By formula S ij t =sim(l i t ,l j t ) Calculate the text word similarity matrix;

[0041] The set of similar words is represented as

[0042] The text numeric feature is represented as in,

[0043] By formula Calculate the number level similarity;

[0044] The objective function of the digit-level triplet loss is L digit =max(0,mS r (v + ,t - )+S r (v + ,t - ));

[0045] Where: S ij v Represents the image region similarity matrix; l i v Represents the local feature L of the image v The i-th feature in D v Represents the digital features of the image; num v represents the number of features in the similar region set V; S ij t Represents the text word similarity matrix; l i tRepresents the local feature of the text L t The i-th feature in D t Indicates the numeric feature of text; num t represents the number of features in the similar word set T; S d (v, t) represents the numerical similarity between the image-text pair (v, t); s ij d Indicates the similarity between the i-th feature in the image digital feature and the j-th feature in the text digital feature; L digit represents the digital level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

[0046] Preferably, in step S4, when performing image-text retrieval, the global-level similarity, relationship-level similarity and number-level similarity output by the global-level subnetwork, relationship-level subnetwork and number-level subnetwork of the image-text retrieval model are fused to generate corresponding multi-level overall similarities, and the retrieval results are scored and sorted based on the multi-level overall similarities;

[0047] Among them, through the formula S overall =S r +αS d +βS g Calculate multi-level overall similarity;

[0048] Where: S overall represents the multi-level overall similarity, S r Represents the relationship level similarity; S d Indicates the numerical level similarity; S g represents the global level similarity; α and β represent the weighted hyperparameters, which are used to adjust the proportion of semantic information at each level of the network.

[0049] Compared with the prior art, the image text retrieval method in the present invention has the following beneficial effects:

[0050] The present invention trains the global level subnetwork, relationship level subnetwork and digital level subnetwork of the image-text retrieval model respectively through training data sets, so that the image-text retrieval model can obtain the global level similarity, relationship level similarity and digital level similarity of images and texts respectively, that is, it can capture the global information, fine-grained relationship information and digital information of images and texts and align them, thereby improving the accuracy of image-text retrieval.

[0051] The way in which the present invention aligns global information, fine-grained relational information and digital information is simpler and less time-consuming than the existing attention mechanism-based alignment, thereby improving the efficiency of image text retrieval.

[0052] The present invention separately trains the global level sub-network, the relationship level sub-network and the digital level sub-network, which can ensure the training effect of each sub-network and effectively reduce the training difficulty of the model, thereby improving the training effect of the image-text retrieval model.

[0053] The present invention scores and sorts the search results by fusing global level similarity, relationship level similarity and digital level similarity to generate multi-level overall similarity, which can ensure the accuracy and effectiveness of the search results output by the graphic and text retrieval model. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0055] Figure 1 It is the logic block diagram of the image text retrieval method;

[0056] Figure 2 This is the network structure diagram of the image and text retrieval model. DETAILED DESCRIPTION

[0057] The following is a further detailed description through specific implementation methods:

[0058] Example:

[0059] This embodiment discloses an image text retrieval method based on a multi-level network.

[0060] like Figure 1 As shown, the image text retrieval method based on a multi-level network includes the following steps:

[0061] S1: Construct a graph-text retrieval model with a global level subnetwork, a relation level subnetwork (local / relation level subnetwork), and a digital level subnetwork;

[0062] S2: Construct a training dataset for training the image-text retrieval model, which includes image-text pairs;

[0063] S3: Combination Figure 2As shown, the image-text pairs in the training data set are respectively input into the global level sub-network, the relation level sub-network and the digital level sub-network of the image-text retrieval model to generate the corresponding global level similarity, relation level similarity and digital level similarity respectively, and then the corresponding sub-networks are separately trained based on the global level similarity, relation level similarity and digital level similarity respectively;

[0064] S4: Perform image text retrieval based on the trained image-text retrieval model. When performing image text retrieval, the global level similarity, relationship level similarity, and number level similarity output by the global level subnetwork, relationship level subnetwork, and number level subnetwork of the image-text retrieval model are integrated to generate the corresponding multi-level overall similarity, and the retrieval results are scored and sorted based on the multi-level overall similarity;

[0065] Among them, through the formula S overall =S r +αS d +βS g Calculate multi-level overall similarity;

[0066] Where: S overall represents the multi-level overall similarity, S r Represents the relationship level similarity; S d Indicates the numerical level similarity; S g represents the global level similarity; α and β represent the weighted hyperparameters, which are used to adjust the proportion of semantic information at each level of the network.

[0067] The present invention trains the global level subnetwork, relation level subnetwork and digital level subnetwork of the image-text retrieval model respectively through the training data set, so that the image-text retrieval model can obtain the global level similarity, relation level similarity and digital level similarity of the image and text respectively, that is, it can capture the global information, fine-grained relation information and digital information of the image and text and align them, so as to improve the accuracy of image-text retrieval. At the same time, the alignment method of the present invention for global information, fine-grained relation information and digital information is simpler and less time-consuming than the existing alignment based on the attention mechanism, so as to improve the efficiency of image-text retrieval. In addition, the method of the present invention for separately training the global level subnetwork, relation level subnetwork and digital level subnetwork can ensure the training effect of each subnetwork, and can effectively reduce the training difficulty of the model, so as to improve the training effect of the image-text retrieval model. Finally, the present invention scores and sorts the retrieval results by fusing the global level similarity, relation level similarity and digital level similarity to generate multi-level overall similarity, so as to ensure the accuracy and effectiveness of the retrieval results output by the image-text retrieval model.

[0068] In the specific implementation process, a Web platform is built based on a high-performance server, and the image-text retrieval model is used as the interface for its back-end call to realize the function of mutual retrieval between images and texts. Users upload images or enter texts to retrieve related data of another modality, and the retrieval results are returned. Among them, the computer equipment used in the hardware part is a service computer based on AMD processor and NVIDIA GeForce GTX 1080Ti GPU, and the required image and text data can be transmitted to its storage system through the network. The codes are all implemented in Python language.

[0069] The training datasets include MS-COCO and Flickr30K. Among them, the MS-COCO dataset is one of the most commonly used datasets in image and sentence retrieval tasks. It contains 123,287 pictures, each with 5 text annotations, 5,000 pictures are used for verification and testing, and the remaining pictures are used for training. The Flickr30k dataset contains 31,783 pictures, each with 5 text annotations, 1,000 pictures are used for verification and testing, and the remaining pictures are used for training.

[0070] In the specific implementation process, the image-text pairs in the training data set are preprocessed and feature extracted to obtain the image global features and image local features of the corresponding image and the text global features and text local features of the corresponding text.

[0071] Specifically, global feature extraction is performed on the image: the image data is input into the Resnet101 model pre-trained on the Visual Genome dataset, the output result of its pooling layer Pool5 is taken, and it is mapped to a 1024-dimensional feature representation through a fully connected neural network. The global feature representation of the image is g v .

[0072] Extract local features of the image: Use the object detection model Faster-RCNN pre-trained on the Visual Genome dataset to extract the salient areas of the image, take the top 36 target areas with the highest scores, and then use the pre-trained Resnet-101 model to extract the features of each area. Take the output of the pooling layer pool5 as the feature of each area, and map it to a 1024-dimensional feature representation through a fully connected neural network. The local feature representation of the image is:

[0073] Extract global and local features from the text: Use the WordPiece model to segment each text data, and map the word to a feature representation with 768 dimensions. The global feature of the text is represented as g t , the local feature of the text is expressed as

[0074] The present invention performs preprocessing and feature extraction on image-text pairs and obtains the image global features and image local features of the image and the text global features and text local features of the text, so that the global level sub-network, the relationship level sub-network and the digital level sub-network can be effectively trained separately based on the image global features, the image local features, the text global features and the text local features, thereby improving the training effect of the image-text retrieval model.

[0075] In the specific implementation process, the global level sub-network is trained through the following steps:

[0076] S301: Inputting the image global features and text global features of the image-text pair into the global level sub-network;

[0077] S302: Calculating corresponding global level similarity based on image global features and text global features;

[0078] S303: Calculate the corresponding global level triplet loss based on the global level similarity, and train the global level sub-network through the global level triplet loss. Obtain N image-text pairs from the training data set, and minimize the global level triplet loss function to make the feature representations of similar image-text pairs close, thereby achieving alignment of the global features of the image and text.

[0079] Specifically, through the formula S g (v,t)=sim g (g v ,g t ) Calculate the global level similarity;

[0080] The objective function of the global level triplet loss is

[0081] Where: S g (v, t) represents the global level similarity of the image-text pair (v, t); g v Represents the global features of the image; g t Represents the global features of the text; L global represents the global level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter; N represents the number of image-text pairs.

[0082] The present invention trains the global level sub-network through the above steps, so that the global level similarity output by the global level sub-network is close to the feature representation, which can realize the alignment of the global features of the image and text, and then can effectively capture the global information of the image and text to help improve the accuracy of image text retrieval.

[0083] In the specific implementation process, the relationship level sub-network is trained through the following steps:

[0084] S311: inputting the image local features and text local features of the image-text pair into the relation level sub-network;

[0085] S312: The local features of the image are passed through a Transformer encoder to aggregate the relationship information between each image region and generate corresponding image relationship features;

[0086] S313: The local features of the text are passed through the Bert module to integrate the relationship information between words and generate corresponding text relationship features;

[0087] S314: Calculating corresponding relationship level similarities based on image relationship features and text relationship features;

[0088] S315: Calculate the corresponding relation-level triplet loss based on the relation-level similarity, and train the relation-level sub-network through the relation-level triplet loss. By minimizing the relation-level triplet loss function, the image and text can be aligned at the relational feature level, thereby capturing the fine-grained relational semantic information of the image and text.

[0089] Specifically, through the formula Calculate relation level similarity;

[0090] Image relation features are expressed as Based on local image features calculate;

[0091] The text relation feature is represented as Based on local features of text calculate;

[0092] The objective function of the relation-level triple loss is L relation =max(0,mS r (v + ,t - )+S r (v + ,t - ));

[0093] Where: S r (v, t) represents the relation-level similarity between the image-text pair (v, t); s ijr Indicates the similarity between the i-th feature in the image relation feature and the j-th feature in the text relation feature; L relation represents the relation-level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

[0094] Among them, the Transformer encoder includes a multi-head self-attention mechanism layer for multiple calculations of attention, and a fully connected forward feedback layer for obtaining rich semantic feature representations; the Transformer encoder can aggregate the relationship information between each image region to generate corresponding image relationship features, and make each feature of the image relationship feature contain the semantic information of the image region and the relationship information with other regions.

[0095] In the multi-head self-attention mechanism layer, since the attention is calculated h times, it is called a multi-head attention mechanism. It is obtained by mapping the query value Q, key value K, and real value V h times through different mapping methods.

[0096] Specifically, given a set X = {x 1 ,x 2 ,...,x m},in as well as Given a set X, we can obtain the query value Q by mapping the matrix. X =XW i Q , key value K X =XW i K And the true value V X =XW i V , where the weight matrix Then, the attention weights are weighted and summed to obtain:

[0097]

[0098] The attention values ​​of each head are concatenated to obtain:

[0099] head i =Attention(XW i Q ,XW i K ,XW i V );

[0100] in, h represents the number of heads.

[0101] In order to obtain a more semantically rich feature representation, the position information of the image area is integrated into the feature representation through a fully connected feedforward network layer. The formula is described as follows:

[0102] FFN(x)=ReLu(xW 1 +b 1 )W 2 +b 2 ;

[0103] in,

[0104] The present invention trains the relationship level sub-network through the above steps, so that the relationship level similarity output by the relationship level sub-network can be aligned at the relationship feature level, thereby being able to capture the fine-grained relationship semantic information of images and texts to help improve the accuracy of image text retrieval.

[0105] In the specific implementation process, the digital level sub-network is trained through the following steps:

[0106] S321: Inputting the image local features and text local features of the image-text pair into the digital level sub-network;

[0107] S322: Calculate the similarity between each region of the image based on the local features of the image to obtain the corresponding image region similarity matrix; then select similar regions with a similarity greater than γ (set as needed) from the image region similarity matrix to form a similar region set; finally, convert the quantity information of the similar region set into a vector through the Bert module and fuse it with the local features of the image in proportion to generate the corresponding image digital features;

[0108] S323: Calculate the similarity between each word in the text based on the local features of the text to obtain a corresponding text word similarity matrix; then select similar words with a similarity greater than γ (set as needed) from the text word similarity matrix to form a similar word set; finally, convert the quantity information of the similar word set into a vector through the Bert module and fuse it with the local features of the text in proportion to generate the corresponding text digital features;

[0109] S324: Calculating corresponding digital level similarities based on the image digital features and the text digital features;

[0110] S325: Calculate the corresponding digit-level triplet loss based on the digit-level similarity, and train the digit-level sub-network through the digit-level triplet loss. By minimizing the digit-level triplet loss function, it is possible to align images and texts at the digit level, thereby capturing fine-grained digit information by minimizing the loss.

[0111] Specifically, through the formula S ij v =sim(l i v ,l j v ) Calculate the image region similarity matrix;

[0112] The set of similar regions is represented as

[0113] The digital features of the image are represented as in,

[0114] By formula S ij t =sim(l i t ,l j t ) Calculate the text word similarity matrix;

[0115] The set of similar words is represented as

[0116] The text numeric feature is represented as in,

[0117] By formula Calculate the number level similarity;

[0118] The objective function of the digit-level triplet loss is L digit =max(0,mS r (v + ,t - )+S r (v + ,t - ));

[0119] Where: S ij v Represents the image region similarity matrix; l i v Represents the local feature L of the image v The i-th feature in D v Represents the digital features of the image; num v represents the number of features in the similar region set V; S ijt Represents the text word similarity matrix; l i t Represents the local feature of the text L t The i-th feature in D t Indicates the numeric feature of text; num t represents the number of features in the similar word set T; S d (v, t) represents the numerical similarity between the image-text pair (v, t); s ij d Indicates the similarity between the i-th feature in the image digital feature and the j-th feature in the text digital feature; L digit represents the digital level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

[0120] The present invention trains the digital level sub-network through the above steps, so that the digital level similarity output by the digital level sub-network can align the image and text at the digital level, and then can effectively capture fine-grained digital information to help improve the accuracy of image text retrieval.

[0121] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described with reference to the preferred embodiments of the present invention, it should be understood by those skilled in the art that various changes can be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims. At the same time, common sense such as the well-known specific structures and characteristics in the embodiments are not described in detail here. Finally, the scope of protection claimed by the present invention shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A multi-level network-based image text retrieval method, It is characterized in that The following steps are involved: S1: Construct a graph-text retrieval model with global level subnetwork, relation level subnetwork and number level subnetwork; S2: Construct a training dataset for training the image-text retrieval model, which includes image-text pairs; In step S2, the image-text pairs in the training data set are preprocessed and feature extracted to obtain the image global features and image local features of the corresponding image and the text global features and text local features of the corresponding text; S3: Input the image-text pairs in the training dataset into the global level sub-network, the relation level sub-network and the number level sub-network of the image-text retrieval model respectively to generate the corresponding global level similarity, relation level similarity and number level similarity respectively, and then train the corresponding sub-networks separately based on the global level similarity, relation level similarity and number level similarity respectively; In step S3, the digit level sub-network is trained by the following steps: S321: Inputting the image local features and text local features of the image-text pair into the digital level sub-network; S322: Calculate the similarity between each region of the image based on the local features of the image to obtain the corresponding image region similarity matrix; then select similar regions with similarity greater than γ from the image region similarity matrix to form a similar region set; finally, convert the quantity information of the similar region set into a vector through the Bert module and fuse it with the local features of the image in proportion to generate the corresponding image digital features; S323: Calculate the similarity between each word in the text based on the local features of the text to obtain a corresponding text word similarity matrix; then select similar words with a similarity greater than γ from the text word similarity matrix to form a similar word set; finally, convert the quantity information of the similar word set into a vector through the Bert module and fuse it with the local features of the text in proportion to generate a corresponding text digital feature; S324: Calculating corresponding digital level similarities based on the image digital features and the text digital features; S325: Calculate the corresponding digit-level triplet loss based on the digit-level similarity, and train the digit-level sub-network through the digit-level triplet loss; S4: Perform image text retrieval based on the trained image-text retrieval model.

2. The image text retrieval method based on a multi-level network as claimed in claim 1, It is characterized in that In step S3, the global level sub-network is trained by the following steps: S301: Inputting the image global features and text global features of the image-text pair into the global level sub-network; S302: Calculating corresponding global level similarity based on image global features and text global features; S303: Calculate the corresponding global level triplet loss based on the global level similarity, and train the global level sub-network through the global level triplet loss.

3. The image text retrieval method based on a multi-level network as claimed in claim 2, Features: By formula S g (v,t)=sim g (g v ,g t ) Calculate the global level similarity; The objective function of the global level triplet loss is Where: S g (v, t) represents the global level similarity of the image-text pair (v, t); g v Represents the global features of the image; g t Represents the global features of the text; L global represents the global level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter; N represents the number of image-text pairs.

4. The image text retrieval method based on a multi-level network as claimed in claim 1, It is characterized in that In step S3, the relationship level sub-network is trained by the following steps: S311: inputting the image local features and text local features of the image-text pair into the relation level sub-network; S312: The local features of the image are passed through a Transformer encoder to aggregate the relationship information between each image region and generate corresponding image relationship features; S313: The local features of the text are passed through the Bert module to integrate the relationship information between words and generate corresponding text relationship features; S314: Calculating corresponding relationship level similarities based on image relationship features and text relationship features; S315: Calculate the corresponding relationship-level triplet loss based on the relationship-level similarity, and train the relationship-level sub-network through the relationship-level triplet loss.

5. The image text retrieval method based on a multi-level network as claimed in claim 4, Features: The Transformer encoder includes a multi-head self-attention mechanism layer for multiple calculations of attention, and a fully connected forward feedback layer for obtaining rich semantic feature representations; the Transformer encoder can aggregate the relationship information between each image region to generate corresponding image relationship features, and make each feature of the image relationship feature contain the semantic information of the image region and the relationship information with other regions.

6. The image text retrieval method based on a multi-level network as claimed in claim 4, Features: By formula Calculate relation level similarity; The objective function of the relation-level triple loss is L relation =max(0,mS r (v + ,t - )+S r (v + ,t - )); Where: S r (v, t) represents the relation-level similarity between the image-text pair (v, t); s ij r Indicates the similarity between the i-th feature in the image relation feature and the j-th feature in the text relation feature; L relation represents the relation-level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

7. The image text retrieval method based on a multi-level network as claimed in claim 1, Features: By formula S ij v =sim(l i v ,l j v ) Calculate the image region similarity matrix; The set of similar regions is represented as The digital features of the image are represented as in, By formula S ij t =sim(l i t ,l j t ) Calculate the text word similarity matrix; The set of similar words is represented as The text numeric feature is represented as in, By formula Calculate number level similarity; The objective function of the digit-level triplet loss is L digit =max(0,mS r (v + ,t - )+S r (v + ,t - )); Where: S ij v Represents the image region similarity matrix; l i v Represents the local feature L of the image v The i-th feature in D v Represents the digital features of the image; num v represents the number of features in the similar region set V; S ij t Represents the text word similarity matrix; l i t Represents the local feature of the text L t The i-th feature in D t Indicates the numeric feature of the text; num t represents the number of features in the similar word set T; S d (v, t) represents the numerical similarity between the image-text pair (v, t); s ij d Indicates the similarity between the i-th feature in the image digital feature and the j-th feature in the text digital feature; L digit represents the digital level triplet loss; (v + ,t - ) represents a matching image-text pair; (v + ,t - ) represents an unmatched image-text pair, that is, an image-text pair with the smallest similarity; m represents the margin hyperparameter.

8. The image text retrieval method based on a multi-level network as claimed in claim 1, Features: In step S4, when performing image-text retrieval, the global-level similarity, relation-level similarity, and number-level similarity output by the global-level subnetwork, relation-level subnetwork, and number-level subnetwork of the image-text retrieval model are fused to generate corresponding multi-level overall similarities, and the retrieval results are scored and sorted based on the multi-level overall similarities; Among them, through the formula S overall =S r +αS d +βS g Calculate multi-level overall similarity; Where: S overall Represents the multi-level overall similarity, S r Represents the relationship level similarity; S d Indicates the numerical level similarity; S g represents the global level similarity; α and β represent the weighted hyperparameters, which are used to adjust the proportion of semantic information at each level of the network.

Citation Information

Patent Citations

  • Image-text mutual retrieval method based on complementary semantic alignment and symmetric retrieval

    CN109255047A

  • Two-stage network image text cross-media retrieval method

    CN110059217A

  • Cross-modal image text retrieval method based on credibility self-adaptive matching network

    CN111026894A