A Cross-modal Hashing Retrieval Feature Fusion Method, System and Storage Medium
By designing a cross-modal hash retrieval model, using tensor fusion and adversarial training technology, the semantic features of data of different modalities are effectively fused, and the problem of low cross-modal retrieval accuracy in the existing technology is solved, and more efficient feature fusion and retrieval accuracy are achieved.
Patent Information
- Application Number
- CN202311063089.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-08-22
AI Technical Summary
The prior art is difficult to effectively fuse the semantic features of data of different modalities in cross-modal retrieval, resulting in low retrieval accuracy and large differences in features between different modalities, making it difficult to directly adopt commonly used feature fusion methods.
The data pair composed of images and text is trained to design a cross-modal hash retrieval model, including semantic feature extraction module, tensor fusion module and adversarial feature fusion module. The semantic features of different modal data are fused through tensor fusion operators, and the feature fusion effect is enhanced by adversarial training.
The accuracy of cross-modal hash retrieval is improved, the semantic interval between different modal data is reduced, the computational amount of the model is reduced, and the capability of feature discriminator is enhanced.
Smart Images

Figure CN116992054B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer multimodality, and more specifically, relates to a cross-modal hashing retrieval feature fusion method, system, and storage medium. Background Technique
[0002] Cross-modal retrieval refers to using data of one modality to query and obtain data semantically related to other multiple modalities in the database. For example, a user can submit a photo on a website, and the website will return related text descriptions, pictures, videos, sounds, etc. Due to the advantages of hash codes and deep learning, the combination of deep learning and hash codes is mainly applied in cross-modal retrieval now.
[0003] In cross-modal retrieval based on the combination of deep learning and hash codes, when performing feature fusion, it is required to fuse the features of different modality data into new features with as little loss of the features of the modality itself as possible, so as to minimize the semantic gap between different modality data. Currently, relatively effective feature fusion methods include attention mechanism, graph neural network, and adversarial learning. Among them, in the method of using adversarial learning for feature fusion, the semantic features of the original data are not directly extracted in the adversarial learning process, but adversarial learning is directly used for feature fusion, which causes the network to not fully extract the semantic information of the data, and thus the retrieval accuracy is not high; at the same time, due to the large feature differences between different modalities, common feature fusion methods cannot be directly used.
[0004] Moreover, in the existing technology, models that only adopt the adversarial idea generally set a discriminator for each modality to judge whether the feature belongs to that modality, which will lead to a weakening of the interaction result between that modality and other modality data, resulting in a large semantic gap between different modality data. For the method of adopting a complete generative adversarial network model, due to the need for an additional generator, the computational cost is too large and it is difficult to be extended to the usage environments of more modalities. If the feature encoding network is directly used as the generator, there will be a problem of mode collapse in adversarial learning. Summary of the Invention
[0005] Aiming at the defects and improvement requirements of the existing technology, the present invention provides a cross-modal hashing retrieval feature fusion method, system, and storage medium, and its purpose is to improve the cross-modal hashing retrieval accuracy.
[0006] To achieve the above object, according to the first aspect of the present invention, a cross-modal hashing retrieval feature fusion method is provided, including:
[0007] Training stage: Using a data pair composed of images and texts as a training sample data set to train a cross-modal hashing retrieval model, and the label is the corresponding category annotation in the images and texts;
[0008] Among them, the cross-modal hashing retrieval model includes:
[0009] A semantic feature extraction module, configured to extract semantic features of different-modal data in training samples and predict hash codes corresponding to the corresponding modalities;
[0010] A tensor fusion module, configured to fuse the semantic features of different-modal data by using a tensor fusion operator, and use the hash codes corresponding to the fused semantic features to semantically guide the hash codes predicted in the semantic feature extraction module; the tensor fusion operator is: Γ = ((Γ c ×1T i )×2T t )×3T o , Γ c is a multi-modal weight matrix, used to guide the weight allocation in the process of fusing semantic features of different modalities; T i , T t , T o are weight matrices for a single modality or task respectively; × k The operator represents the projection calculation of the semantic feature corresponding to the k-th modality and the weight matrix;
[0011] An adversarial feature fusion module, configured to perform adversarial training with the semantic features of different-modal data to semantically guide the semantic features of different-modal data in the semantic feature extraction module;
[0012] Application stage: Input the text and pictures to be retrieved into the trained semantic feature extraction module to obtain hash codes corresponding to the corresponding modalities, and obtain retrieval results according to the hash codes.
[0013] Further, in the semantic feature extraction module, the semantic features of the different-modal data include picture semantic feature f i , text semantic feature f t , and label semantic feature f l ;
[0014] In the tensor fusion module, fusing the semantic features of different-modal data by using the tensor fusion operator includes:
[0015] After the picture semantic feature f i and the text semantic feature f t are respectively subjected to a layer of non-linear mapping, the corresponding picture feature tensor (f i ) T T i and text feature tensor (f t ) T T t are obtained;
[0016] Adopt the multi-modal weight matrix Γ c For the picture feature tensor (f i ) T T i and the text feature tensor (f t ) T T t perform fusion and weight assignment to obtain the fused tensor
[0017] The fused tensor and the label semantic feature f l are respectively passed through a layer of non-linear mapping to obtain the corresponding picture-text fusion feature tensor and the label feature tensor (f l ) T T t ;
[0018] Adopt the multi-modal weight matrix Γ c For the picture-text fusion feature tensor and the label feature tensor (f l ) T T t perform fusion and weight assignment to obtain the re-fused tensor z:
[0019] The tensor z is passed through a layer of non-linear mapping again to obtain the fused semantic feature f f : f f = z T T o .
[0020] Furthermore, each time tensor fusion is performed, it also includes reducing the rank of the multi-modal weight matrix Γ c , including:
[0021] The k-th result z in the current tensor fusion result k is expressed as: where and are two feature tensors obtained after passing two different modal semantic features to be fused through a layer of non-linear mapping, Γ c [:,:,k] is the multi-modal weight matrix for the k-th result in tensor fusion, and the value of k is 1, 2,..., n, where n is the total number of data pairs in the dataset;
[0022] Each multi-modal weight matrix is expressed as the sum of R first-order matrices: is the outer product operation, m r and nr They are the first-order weight matrices corresponding to the weight matrices of the image and text modality data respectively; among them, the rank of the multi-modal weight matrix corresponding to each data pair is a constant R;
[0023] Define R matrices And Then the tensor after fusing the semantic features of two different modalities is: Among them, (A)*(B) represents the intelligent product of element A and element B; l i , l t , l o respectively represent the feature lengths corresponding to the image data, text data, and the fused data after Tucker decomposition.
[0024] Furthermore, the semantic feature extraction module includes:
[0025] A semantic feature extraction layer for extracting the semantic features of different modality data in the training samples;
[0026] A first label prediction layer for predicting the labels corresponding to the semantic features of the different modality data;
[0027] A first hash code generation layer for predicting the hash codes corresponding to the semantic features of the different modality data;
[0028] In the training stage, the loss function of the semantic feature extraction module is:
[0029]
[0030] Among them, represents the hash code of the j-th multi-modal data pair predicted by the first hash code generation layer, * represents different modalities of the data; represents the hash code corresponding to the label in the i-th data pair, H * is the hash code predicted by the first hash code generation layer, H l represents the hash code corresponding to the sample label; is the predicted label, L is the label corresponding to the sample; S ij is the similarity matrix of the data pair samples; β * , γ * , δ * are hyperparameters to be learned respectively; n represents the total number of data pairs in the dataset.
[0031] Furthermore, the adversarial feature fusion module includes a feature discriminator, and the feature discriminator is also used for adversarial training with the fused semantic features output by the tensor fusion module;
[0032] The adversarial feature fusion module further includes a hash code discriminator, which is used to perform adversarial training with the hash code predicted by the semantic feature extraction module.
[0033] Further, the tensor fusion module is also used to fuse the doped image semantic feature f i and the text semantic feature f t to obtain a doped and fused feature; the feature discriminator is also used to perform adversarial training on the doped and fused feature;
[0034] wherein, the doping and fusion of the image semantic feature f i and the text semantic feature f t include:
[0035] Multiply the image semantic feature f i by the first coefficient σ to obtain the feature f i σ; multiply the text semantic feature f t by the second coefficient ζ to obtain the feature f t ζ; when the text semantic feature f i is doped into the image semantic feature f t , σ >> ζ; when the image semantic feature f t is doped into the text semantic feature f i , ζ >> σ;
[0036] Input the feature f i σ and the feature f t ζ into the tensor fusion module for feature fusion.
[0037] Further, in the training stage, it also includes gradient penalty on the feature discriminator, including: regarding the label semantic feature f l as the real data of adversarial learning, regarding the image semantic feature f i and the text semantic feature f t as the generated data of adversarial learning, and performing weighted averaging on the distributions of the real data and the generated data to obtain critical data;
[0038] Calculate the gradient of the critical data in the feature discriminator, and punish the gradient through the loss function of gradient penalty; wherein, the loss function L penaf of the gradient penalty is: represents the critical data, represents the gradient calculation of the critical data backpropagated by the feature discriminator, D f represents the feature discriminator.
[0039] Further, in the semantic feature extraction module, a loss function for semantic guidance is applied to the predicted hash code in the semantic feature extraction module using the hash code corresponding to the fused semantic feature is:
[0040]
[0041] where H f is the hash code corresponding to the fused semantic feature, and H * is the predicted hash code in the semantic feature extraction module;
[0042] In the adversarial feature fusion module, the loss function for semantic guidance of the semantic features of different modal data in the semantic feature extraction module is:
[0043]
[0044]
[0045] where represents the loss of guiding the semantic features of the image with the adversarial feature fusion module, represents the loss function of guiding the semantic features of the text with the adversarial feature fusion module; Λ f* represents the output predicted by the feature discriminator, and Λ h* is the output predicted by the hash code discriminator. Taking i, t, and l respectively, it correspondingly represents that the semantic features input to the discriminator come from the image, text, and label modal data.
[0046] According to the second aspect of the present invention, a cross-modal hashing retrieval feature fusion system is provided, including a computer-readable storage medium and a processor;
[0047] The computer-readable storage medium is used to store executable instructions;
[0048] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of the first aspects.
[0049] According to the third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to any one of the first aspects is implemented.
[0050] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0051] (1) In the cross-modal hashing retrieval feature fusion method of the present invention, the semantic features of different modalities are fused through the designed tensor fusion operator. In this fusion operator, the core tensor Γ cis the multimodal weight matrix corresponding to the tensor fusion module in the neural network. As an interaction matrix, it can guide the weight distribution during the fusion of different modal data features, enabling the fused result to reflect the potential distribution of the original features. Using the fused features to guide the semantic feature extraction module can make full use of the semantic information of the original semantic features, reduce the semantic gap between different modal data, and improve the retrieval accuracy.
[0052] (2) Further, the present invention designs a method for reducing the rank of the multimodal weight matrix Γ c by approximating the core tensor Γ c through outer product multiplication and summation, rather than directly multiplying multi-dimensional data, which greatly reduces the number of parameters and the computational complexity of the cross-modal hashing retrieval model. At the same time, by reducing the rank of the multimodal weight matrix Γ c , it can also reduce the noise of the semantic features of different modalities and further improve the retrieval accuracy.
[0053] (3) Further, the loss function of the semantic feature extraction module of the present invention fully applies the semantic information of the labels and uses the semantic information contained in the labels as a guide, enabling the hash codes generated by the labels to accurately reflect the relationship between different data pairs. The first term of the loss function realizes the diffusion of semantic information by multiplying the hash code generated by the label by the hash code predicted by the first hash code generation layer. After the diffusion of semantic information, it is constrained by a similarity matrix, that is, the similarity of the hash codes generated by related data pairs is increased, and the similarity of the hash codes of the data pairs with "0" in the similarity matrix (indicating that there is no relationship between the data pairs) is decreased. The second term is constrained by the labels, enabling the feature encoding neural network (i.e., the semantic feature extraction module) to learn the mathematical distribution of the class semantic information contained in the data. The last term is guided by the hash code generated by the label, making the hash codes generated by each multimodal data pair approach in a unified and correct direction. Constrained by these three losses, the feature encoding networks corresponding to each modality can more accurately learn the semantic distribution contained in the multimodal data.
[0054] (4) Further, the feature discriminator of the present invention can improve the ability of the feature discriminator through adversarial training with the fused semantic features output by the tensor fusion module;
[0055] At the same time, the adversarial feature fusion module also includes a hash code discriminator, which conducts adversarial training on the predicted hash code and the hash code of the corresponding modal data, thereby enhancing the adversarial effect of the adversarial feature fusion module.
[0056] (5) Further, by fusing the doped picture semantic features and text semantic features and then inputting them into the feature discriminator for adversarial training, the interactivity between the currently output modal data of the discriminator and other modal data can be enhanced, and the semantic gap between different modal data can be further reduced.
[0057] (6) Further, considering that different multi-modal data have different distributions, by regarding the label semantic feature f l output by the semantic feature extraction layer as the real data in adversarial learning, and regarding the picture semantic feature f i and text semantic feature f t as the generated data in adversarial learning to generate critical data, when performing gradient penalty on the discriminator based on the designed critical data, the characteristics of multiple modalities can be considered simultaneously. When the gradient is backpropagated, the updated direction will not deviate too much from a certain modality, reducing the fluctuation of gradient update. It can also prevent the gradient from concentrating in one interval, improving the stability of the discriminator. When the adversarial feature fusion module performs adversarial training, the semantic feature extraction module can be directly used as the generator, thereby reducing the computational amount of the model. At the same time, it also avoids the problem of mode collapse in adversarial learning when directly using the feature encoding network as the generator.
[0058] (7) Further, through the designed loss function for guiding the semantic feature extraction module based on tensor fusion, the hash codes generated by each modality can be approximated towards the hash code generated after tensor fusion, enabling a relationship between the hash codes generated between different modalities. Based on the loss function for guiding the semantic feature extraction module in adversarial feature fusion, during adversarial training, the feature encoding networks of text and pictures are regarded as the generators in the generative adversarial network. The learning objective of the generator is to be able to deceive the discriminator and make the discriminator make wrong judgments. Since the objective of the discriminator is to determine which modality the data comes from, the objective of the generator is to make the generated results approximate the characteristics of other modalities as much as possible, thereby achieving the effect of feature fusion. Finally, the two competing opponents, the generator and the discriminator, will gradually become more and more powerful, and the degree of feature fusion will also become deeper. By guiding the semantic feature extraction module through the tensor fusion module and the adversarial feature fusion module, the semantic gap between different modal data can be broken. Brief Description of the Drawings
[0059] Figure 1 is a framework diagram of the cross-modal hashing retrieval feature fusion method in an embodiment of the present invention.
[0060] Figure 2 is a precision simulation curve graph when the retrieval method provided in an embodiment of the present invention and the commonly used cross-modal retrieval method in the prior art perform picture retrieval of text.
[0061] Figure 3 This is the precision simulation curve graph when the retrieval method provided by the embodiment of the present invention and the commonly used cross-modal retrieval method in the prior art perform text retrieval of pictures. Specific implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0063] As Figure 1 shown, the cross-modal hashing retrieval feature fusion method of the present invention includes: a training stage and an application stage;
[0064] The training stage includes: using a data pair composed of images and texts as a training sample data set to train a cross-modal hashing retrieval model, and the label is the corresponding category annotation in the images and texts; in the embodiment of the present invention, in the image and text databases used for cross-modal retrieval, the images and texts are labeled with category labels to obtain a standard training data set with labels;
[0065] Among them, the cross-modal hashing retrieval model includes: a semantic feature extraction module, a tensor fusion module, and an adversarial feature fusion module; in the embodiment of the present invention, the tensor fusion module and the adversarial feature fusion module constitute a semantic feature fusion module.
[0066] The semantic feature extraction module is used to extract the semantic features of different modal data in the training samples, including the image semantic feature f i , the text semantic feature f t , and the label semantic feature f l , and predict the hash code corresponding to the modality;
[0067] The tensor fusion module is used to perform feature fusion on the semantic features of different modal data by using the constructed tensor fusion operator to obtain the fused semantic features; and use the hash code corresponding to the fused semantic features to semantically guide the hash code predicted in the semantic feature extraction module; among them, the tensor fusion operator Γ is: Γ = ((Γ c ×1T i )×2T t )×3T o , where Γ c is the core tensor, which is the multi-modal weight matrix in the neural network corresponding to the tensor fusion module and is used to guide the weight distribution in the process of fusing the semantic features of different modalities; T i ,T t ,T oare three factor matrices, which are weight matrices for single modality or task respectively; the core tensor Γ c and three factor matrices T i , T t , T o are obtained by training the tensor fusion module; during the training process of the semantic feature extraction module, the corresponding network parameters are adjusted backward by minimizing the loss function of the tensor fusion module, and the core tensor Γ c and three factor matrices T i , T t , T o ; where d i , d t , d o represent the dimensions of the image semantic feature, the text semantic feature, and the fused feature respectively, and l i , l t , l o represent the feature lengths corresponding to different modalities of data (image data, text data, and fused data) after Tucker decomposition for different modalities of data samples. In the embodiments of the present invention, for the convenience of simplifying the calculation, it is set that d i = d t = d o , l i = l t = l o ; × k The operator represents the projection calculation of the semantic feature tensor corresponding to the k-th modality and the weight matrices (Γ c , T i , T t , T o );
[0068] After fusing the image semantic feature f i and the text semantic feature f t using the above-mentioned tensor fusion operator Γ, the result of fusing the image semantic feature and the text semantic feature is obtained Then, the result of fusing the image semantic feature and the text semantic feature and the label semantic feature f l are fused using the above-mentioned tensor fusion operator Γ to obtain the fused semantic feature.
[0069] The adversarial feature fusion module is used to perform adversarial training with the semantic features of different modalities of data to provide semantic guidance for the semantic features of different modalities in the semantic feature extraction module;
[0070] Application stage: Execute the cross-modal hashing retrieval task using the trained cross-modal hashing retrieval model; specifically, input the text and pictures to be retrieved into the trained semantic feature extraction module after semantic guidance to obtain the hash codes of the corresponding modalities, and obtain the retrieval result based on the hash codes.
[0071] In the embodiment of the present invention, the semantic feature extraction module serves as a data encoder for different modality data that combines global information and local information. The semantic feature extraction module is semantically guided by the tensor fusion module and the adversarial feature fusion module to reduce the semantic gap.
[0072] Specifically, the semantic feature extraction module includes: a semantic feature extraction layer, a first label prediction layer, and a first hash code generation layer;
[0073] The semantic feature extraction layer is used to extract the semantic features of different modality data in the training samples; among them, the different modality data includes picture data, text data, and label data.
[0074] The first label prediction layer is used to predict the corresponding label That is, the corresponding semantic information; during the training process, the first label prediction layer is trained with the goal of minimizing the feature loss between the predicted label and the actual label, which reflects the loss training of local semantic information. In the embodiment of the present invention, the corresponding loss function adopts the mean square error loss function.
[0075] The first hash code generation layer is used to predict the hash codes of the corresponding modality data according to the semantic features of different modality data; during the training process, the first hash code generation layer is trained with the goal of minimizing the feature loss between the predicted hash code and the hash code generated by the actual label, and minimizing the loss between the similarity matrix corresponding to the predicted hash code and the similarity matrix of the data samples. Among them, the loss function between the predicted hash code and the hash code generated by the actual label reflects the loss training of local semantic information, and the loss function between the similarity matrix corresponding to the predicted hash code and the similarity matrix of the data samples reflects the loss training of global semantic information. Among them, the calculation of the similarity matrix of the data samples includes:
[0076] Preprocess the data pairs in the dataset: adjust the picture data in the data pairs to the same size, and represent the text data in the data pairs as vectors using the bag-of-words model;
[0077] Generate the similarity matrix of the data samples according to the number of data pairs sharing the same label among the preprocessed data pairs, that is, this similarity matrix can reflect the similarity between the data pairs. The greater the similarity, the closer the connection between the data pairs.
[0078] Denote the loss function of the first label prediction layer is:
[0079]
[0080] where L is the label corresponding to the multimodal data, is the predicted label, * represents different modalities of the data. In the embodiments of the present invention, i, t, and l are taken respectively, corresponding to image data, text data, and label data. In other embodiments, other data modalities may also be included; that is, are respectively when corresponding to the label corresponding to the image data, the label corresponding to the text data, and the label corresponding to the category annotation data, ||·|| F represents the Frobenius norm. β l represents the hyperparameter to be learned.
[0081] The global loss function in the first hash code generation layer is:
[0082]
[0083]
[0084] where represents the global loss in the first hash code generation layer, δ l ||H * -H l || F represents the local loss in the first hash code generation layer; n represents the total number of data pairs in the dataset, and i and j respectively represent the i-th and j-th data pairs; are respectively the hash codes corresponding to the labels in the i-th and j-th data pairs, S ij is the similarity matrix of the data pair samples; α l , γ l , δ l respectively represent the hyperparameters to be learned; H * is the hash code corresponding to the data of different modalities predicted by the first hash code generation layer, H l represents the hash code corresponding to the actual label. In this loss function, for the global loss, the first term constrains through the similarity matrix of the data samples, making the distance between the shared labeled data pairs closer, while the distance between unrelated data pairs becomes farther. The purpose of adding the logarithm is to avoid overfitting of the model and reduce the occurrence of dead neurons due to the comparison of mostly 0s with 0s (indicating no relationship between data pairs).
[0085] During the training of the cross-modal hashing retrieval model, the loss function of the first label prediction layer and the loss function corresponding to the first hash code generation layer are weighted and summed as the total loss function of the semantic feature extraction module. And backpropagate to calculate the gradient, and update the network parameters corresponding to the semantic feature extraction module, so that the semantic feature extraction module can extract the global and local information of image, text and label data. In the embodiment of the present invention, an adaptive moment estimator is used to update the network parameters corresponding to the semantic feature extraction module, where the corresponding network parameters include the core tensor Γ c and three factor matrices T i , T t , T o .
[0086] Specifically, the total loss function of the semantic feature extraction module is:
[0087]
[0088] where β * , γ * , δ * are respectively hyperparameters in the semantic feature extraction module; represents the hash code corresponding to the j-th multi-modal data pair predicted by the first hash code generation layer; * represents different modalities of data. The total loss function of the semantic feature extraction module designed in the present invention fully applies the semantic information of the label, and uses the semantic information contained in the label as a guide, so that the hash code generated by the label can accurately reflect the relationship between different data pairs. The first term of the total loss function uses the method of multiplying the hash code generated by the label by the hash code predicted by the first hash code generation layer to realize the diffusion of semantic information; after the diffusion of semantic information, it is constrained by the similarity matrix, that is, the similarity of the hash codes generated by the related data pairs becomes larger, and the similarity of the hash codes of the data pairs with "0" in the similarity matrix (indicating that there is no relationship between the data pairs) becomes smaller. The second term is constrained by the label, so that the feature encoding neural network (that is, the semantic feature extraction module) can learn the mathematical distribution of the class semantic information contained in the data. The last term is guided by the hash code generated by the label, so that the hash code generated by each multi-modal data pair will approach in a unified and correct direction. Constrained by the three losses in the above formula, the feature encoding networks corresponding to each modality can more accurately learn the semantic distribution contained in the multi-modal data. Therefore, the hash codes generated by the non-linear mapping of the multi-modal data through the feature encoding network can roughly reflect the relationship between different data pairs within the modality.
[0089] Specifically, the tensor fusion module includes: a semantic feature fusion layer, a second label prediction layer, and a second hash code generation layer;
[0090] The semantic feature fusion layer is used to perform feature fusion on the semantic features of different modality data to obtain the fused semantic features;
[0091] The second label prediction layer is used to predict the corresponding label according to the fused semantic features; similar to the training process of the first label prediction layer, during the training process, the predicted label is used to calculate the loss with the actual label, generating a label loss function based on tensor fusion, and the second label prediction layer is trained for loss with the goal of minimizing this loss.
[0092] The second hash code generation layer is used to predict the hash code of the corresponding modality data according to the fused semantic features, and use the predicted hash code of the corresponding modality data to semantically guide the hash code predicted in the semantic feature extraction module; similar to the training process of the first hash code generation layer, during the training process, the predicted hash code is multiplied by the hash code corresponding to the label in other data pairs, and then guided by a similarity matrix, generating a global information loss function based on tensor fusion.
[0093] The label loss function based on tensor fusion and the global information loss function based on tensor fusion are weighted and summed to obtain the total loss function of the tensor fusion module. Backpropagation is used to calculate the gradient and update the network parameters corresponding to the tensor fusion module; in the embodiments of the present invention, the adaptive moment estimation optimizer is used to update the network parameters corresponding to the tensor fusion module.
[0094] Specifically, the total loss function of the tensor fusion module is:
[0095]
[0096] where represents the hash code of the j-th data pair predicted by the second hash code generation layer; α f 、γ f 、β f respectively represent hyperparameters to be learned; L f represents the label predicted by the second label prediction layer.
[0097] Specifically, by constructing a loss function between the hash code of the corresponding modality data predicted by the second hash code generation layer and the hash code of the corresponding modality data predicted by the first hash code generation layer to minimize this loss function With the goal of using the hash code of the predicted corresponding modal data to semantically guide the hash code predicted in the semantic feature extraction module. Specifically, based on tensor fusion to guide the loss function of the semantic feature extraction module is:
[0098]
[0099] where H f is the hash code of the corresponding modal data predicted by the second hash code generation layer, that is, the hash code corresponding to the fused semantic features, and H * is the hash code of the corresponding modal data predicted by the first hash code generation layer. Through this loss function the hash codes generated by each modality are approximated towards the hash code generated after tensor fusion, which can make the hash codes generated between different modalities related, thereby breaking the semantic gap between different modal data.
[0100] Specifically, in the semantic feature fusion layer, a tensor fusion operator is used to fuse the image semantic feature f i , the text semantic feature f t and the semantic feature f l of the label, including:
[0101] After the image semantic feature f i and the text semantic feature f t are respectively subjected to a layer of non-linear mapping and projected into a unified embedding space, the corresponding image feature tensor (f i ) T T i and the text feature tensor (f t ) T T t are obtained;
[0102] Use the core tensor Γ c to perform a fusion and weight assignment on the image feature tensor (f i ) T T i and the text feature tensor (f t ) T T t to obtain the tensor after fusing the image semantic feature and the text semantic feature That is, represents the result of fusing the image and text features, which can reflect the potential distribution of the original features;
[0103] The above realizes the fusion of the image semantic feature and the text semantic feature, and again uses the above method to realize the result after fusing the image and text features and the label semantic feature f lThe fusion, that is:
[0104] The result after fusing the image text features and the label semantic feature f l After passing through a non-linear mapping layer respectively, they are projected into a unified embedding space to obtain the corresponding image text fusion feature tensor and the label feature tensor (f l ) T T t ;
[0105] Use the core tensor Γ c to perform one fusion and weight assignment on the image text feature tensor and the label feature tensor (f l ) T T t to obtain the tensor z after fusing the semantic feature after image text fusion and the semantic feature of the label, that is, z represents the result after fusing the image text label features;
[0106] Then convert the tensor z after the second fusion into a set feature dimension to obtain the feature fusion result of the required dimension; specifically, after passing the tensor z after the second fusion through another non-linear mapping layer, convert it into the set feature dimension to obtain the result f after fusing the image text semantic feature and the label semantic feature under the required dimension f , f f = z T T o .
[0107] Through the tensor fusion operator designed by the present invention, feature fusion is performed on the multi-modalities corresponding to the image semantic feature, the text semantic feature, and the label semantic feature. In this fusion operator, the core tensor is the multi-modal weight matrix in the neural network corresponding to the tensor fusion module. As an interaction matrix, it can guide the weight assignment during the fusion process of different modality data features, so that the fused result can reflect the potential distribution of the original features; using the fused features to guide the semantic feature extraction module can make full use of the semantic information of the original semantic features, reduce the semantic gap between different modality data, and improve the retrieval accuracy.
[0108] Furthermore, in order to reduce the computational complexity of the cross-modal hashing retrieval model and denoise the semantic features of different modalities at the same time, the present invention also designs a cross-modal rank reduction method to perform rank reduction on the core tensor Γ c , specifically including:
[0109] When performing tensor fusion, denote and After the semantic features of two different modalities to be fused are subjected to a layer of non - linear mapping, two feature tensors are obtained, and then there is z k represents the fusion result corresponding to the k - th data pair in the current tensor fusion result, where the value of k is 1, 2, …, n, and n is the total number of data pairs in the dataset; Γ c [:,:,k] is the multi - modal weight matrix for the k - th result during tensor fusion; in the embodiments of the present invention, when fusing the picture semantic features and the text semantic features, when the result after fusing the picture - text features is further fused with the label semantic features,
[0110] Let the rank of the multi - modal weight matrix Γ c corresponding to each data pair be the same constant R, and represent each multi - modal weight matrix as the sum of R first - order matrices: where, is the outer - product operation, m r and n r are the first - order weight matrices of the weight matrices corresponding to the picture and text modality data respectively;
[0111] Define R matrices and Then the tensor after fusing the semantic features of two different modalities is: where, (A)*(B) represents the intelligent product of element A and element B.
[0112] Through this cross - modal rank - reduction method, the core tensor Γ c is approximated by means of outer - product multiplication and summation, greatly reducing the number of parameters.
[0113] Specifically, the adversarial feature fusion module includes: a feature discriminator;
[0114] The feature discriminator is used to perform adversarial training with the semantic features of different modality data output by the semantic feature extraction layer to obtain a first predicted triple; among them, the first predicted triple is the probability that the semantic feature of the corresponding modality belongs to the picture, label, and text; during the training process, the first predicted triple is compared with the set label triple to generate a loss function based on the feature discriminator label With the goal of minimizing this loss function, the feature discriminator is trained. Among them, the set label triple is used to indicate which modality the data feature comes from, which is equivalent to the label in the discriminator task. That is, the meaning of "discrimination" is to judge which modality the input semantic feature comes from. In the embodiments of the present invention, the output of the feature discriminator is is the special discriminator D fThe parameters normalize the output result of the discriminator, enabling it to be extended to scenarios of multiple modalities, that is, *in addition to being able to represent images, text, or labels, it can also be other data modalities.
[0115] To enhance the adversarial effect, the adversarial feature fusion module further includes a hash code discriminator for adversarial training with the hash code of the corresponding modality data predicted by the first hash code generation layer; similar to the feature discriminator, during training, the hash code discriminator is trained with the goal of minimizing the loss between the triplet predicted by the hash code discriminator and the set label triplet, generating a loss function based on the hash code discriminator label In the embodiment of the present invention, the output of the hash code discriminator is: Wherein, represents the parameters of the hash code discriminator D H of.
[0116] Specifically, the loss function based on the feature discriminator label is:
[0117]
[0118] Wherein, Λ f* is the first predicted triplet output by the feature discriminator; * represents different modalities of the features input to the discriminator, taking i, t, and l respectively, corresponding to the semantic features input from the image, text, and label modality data; B i , B t , B l respectively represent the set label triplets.
[0119] The loss function based on the hash code discriminator label is:
[0120]
[0121] Wherein, Λ h* is the predicted triplet output by the hash code discriminator.
[0122] In the embodiment of the present invention, three one-hot vectors B * ={B i , B t , B l} are preset according to experience as the set label triplets to guide the learning of the feature and hash code discriminators, so as to strengthen the intra-modal characteristics of features from different modalities, that is, the output result of the correct neuron will be much larger than that of the incorrect neuron; as an alternative implementation, B i =[1.05, 0.05, 0.05] T , B l= [0.05, 1.05, 0.05] T , B t = [0.05, 0.05, 1.05] T respectively represent that the features come from images, labels, and texts; among them, the way of not using "1" and "0" as elements in each label triple is because not setting absolute values allows the discriminator to continuously learn and avoid overfitting.
[0123] To improve the ability of the discriminator, the feature discriminator is also used for adversarial training with the fused semantic features output by the tensor fusion module; that is, the fused semantic features output by the tensor fusion module are input into the feature discriminator for adversarial training to improve the discrimination ability of the discriminator.
[0124] To enhance the interaction between the current output modal data of the discriminator and other modal data, and further reduce the semantic gap between different modal data, the tensor fusion module is also used to fuse the doped image semantic features and text semantic features to obtain the doped and fused features; the feature discriminator is also used to perform adversarial training on the doped and fused features to obtain the second predicted triple; during the training process, the second predicted triple is compared with the set label triple to generate a loss function based on the discriminator's fused features With the goal of minimizing this loss function, the feature discriminator is trained. Among them, the loss function Λ fit represents the second predicted triple obtained after doping the text semantic features in the image semantic features, performing feature fusion, and then inputting them into the feature discriminator; Λ fti represents the second predicted triple obtained after doping the image semantic features in the text semantic features, performing feature fusion, and then inputting them into the feature discriminator.
[0125] Specifically, the doping of image and text semantic features includes: multiplying the image semantic feature f i by the first coefficient σ to obtain the feature f i σ; multiplying the text semantic feature f t by the second coefficient ζ to obtain the feature f t ζ; inputting the feature f i σ and the feature f t ζ into the tensor fusion module for feature fusion; among them, when doping the text semantic features in the image semantic features, σ >> ζ; when doping the image semantic features in the text semantic features, ζ >> σ.
[0126] As a further design of the present invention, in order to improve the stability of the adversarial feature fusion module, enable the semantic feature extraction module to be directly used as the generator of the adversarial network, reduce the computational cost of the model, and solve the mode collapse problem of adversarial learning, the present invention designs a gradient penalty method to constrain the gradient in the feature discriminator, specifically including:
[0127] Regarding the label semantic feature f output by the semantic feature extraction layer l as the real data in the original adversarial learning concept, regarding the image semantic feature f i and the text semantic feature f t as the generated data in the original adversarial learning concept, and performing weighted averaging on the distributions of the label semantic feature f l , the image semantic feature f i and the text semantic feature f t data to obtain the critical data;
[0128] Calculating the gradient of the critical data in the feature discriminator, and punishing the gradient through the loss function of gradient penalty; among them, the loss function L of gradient penalty penaf is: represents the data sampled from the image feature, text feature and label feature, that is, the critical data, represents the gradient calculation of the feature discriminator for the backpropagation of the critical data, that is, calculating the gradient of the backpropagation of the result of the feature discriminator judging the sampled data, D f represents the feature discriminator.
[0129] In the embodiment of the present invention, considering that different multi-modal data have different distributions, by using the label semantic feature f output by the semantic feature extraction layer l as the real data in adversarial learning, and using the image semantic feature f i and the text semantic feature f t as the generated data in adversarial learning to generate critical data, when performing gradient penalty on the discriminator based on the designed critical data, the characteristics of multiple modalities can be considered simultaneously, so that when the gradient is backpropagated, the updated direction will not deviate too much from a certain modality, reducing the fluctuation of gradient update, and also avoiding the gradient being concentrated in one interval, improving the stability of the discriminator, enabling the adversarial feature fusion module to directly use the semantic feature extraction module as the generator during adversarial training, thereby reducing the computational amount of the model; it also avoids the mode collapse problem of adversarial learning when directly using the feature encoding network as the generator.
[0130] In the embodiment of the present invention, the loss function based on the discriminator tag, the loss function based on the discriminator fusion feature, and the loss function based on the gradient penalty are summed up to obtain the total loss function of the adversarial feature fusion module. Backpropagation is used to calculate the gradient and update the network parameters corresponding to the adversarial feature fusion module. In the embodiment of the present invention, the RMSProp optimizer is used to update the network parameters corresponding to the adversarial feature fusion module.
[0131] Specifically, the total loss function of the adversarial feature fusion module for feature training is:
[0132]
[0133] where λ is the network hyperparameter to be learned.
[0134] Specifically, the adversarial feature fusion module is used to semantically guide the semantic feature extraction module:
[0135] The fused semantic features output by the tensor fusion module are input into the feature discriminator to obtain the third predicted triple. The third predicted triple is compared with the triples that do not belong to this modality to generate the loss function loss for guiding the semantic feature extraction module based on the adversarial feature fusion. G* is:
[0136]
[0137]
[0138] where represents the loss function for guiding the image encoding network based on the adversarial feature fusion, that is, the loss of guiding the image semantic features with the adversarial feature fusion module. represents the loss function for guiding the text encoding network with the adversarial feature fusion module; Λ f* represents the output predicted by the feature discriminator. * represents the semantic feature modality input into the discriminator, taking i, t, l, which respectively represent that the input semantic features come from the image, text, and label modality data; Λ h* is the output predicted by the hash code discriminator.
[0139] It can be seen that in the present invention, the feature encoding networks of text and images are directly regarded as the generator in the generative adversarial network. The learning objective of the generator is to be able to deceive the discriminator so that the discriminator makes a wrong judgment. Since the objective of the discriminator is to judge which modality the data comes from, the objective of the generator is to make the generated result approximate the characteristics of other modalities as much as possible, so as to achieve the effect of feature fusion. Finally, the two competing opponents, the generator and the discriminator, will gradually become more and more powerful, and the degree of feature fusion will also become deeper and deeper.
[0140] Finally, the total loss function L when using the tensor fusion module and the adversarial feature fusion module to guide the semantic feature extraction module * is:
[0141]
[0142] Using this guidance loss function to perform backpropagation on the semantic feature extraction module to achieve semantic guidance for the semantic feature extraction module, so as to reduce the semantic gap between different modality data.
[0143] Specifically, in the application stage, the text and images to be retrieved are input into the semantic feature extraction module after training and semantic guidance. The first hash code generation layer outputs the corresponding hash code, and the hash code is saved as the database to be retrieved; the text or image to be queried is input into the semantic feature extraction module after training and semantic guidance to generate the corresponding hash code, and the Hamming distance is calculated between the generated hash code and the hash codes in the database, and the first m data with relatively small Hamming distances are selected as the corresponding image or text retrieval results; where m is a preset value obtained by those skilled in the art through empirical analysis. That is to say, the present invention can implement the two-way retrieval tasks of retrieving images from text and retrieving text from images.
[0144] The retrieval accuracy of the method in the MirFlickr25K dataset using the embodiment of the present invention and the retrieval accuracies of other methods in the prior art are as Figure 2 and Figure 3 shown. The abscissa in the figure represents the recall rate, and the ordinate represents the retrieval accuracy, which is measured using the MAPTOPK metric. MAPTOPK refers to the retrieval accuracy of the first K data related to the query data in the database. Among them, the ATCMR curve is the accuracy curve of the method of the present invention, and DCMH, CHN, PRDH, SSAH, MLCAH, and AGAH are all commonly used retrieval methods in the art. I->T represents retrieving text from images, and T->I represents retrieving images from text. It can be seen that when retrieving text from images, the retrieval accuracy of the method of the present invention is significantly improved compared with other methods in the prior art, and when retrieving images from text, the retrieval accuracy of the method of the present invention can also be higher than the retrieval accuracies of other methods, and is more stable than other methods.
[0145] The present invention also provides a cross-modal hashing retrieval feature fusion system, including a computer-readable storage medium and a processor; the computer-readable storage medium is used for storing executable instructions; the processor is used for reading the executable instructions stored in the computer-readable storage medium and executing the steps corresponding to the cross-modal hashing retrieval feature fusion method in the above embodiments.
[0146] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the cross-modal hashing retrieval feature fusion method in the above embodiments is implemented.
[0147] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A cross-modal hashing retrieval feature fusion method, characterized in that, Including: Training stage: Using a data pair composed of images and texts as a training sample data set to train a cross-modal hashing retrieval model, and the label is the corresponding category annotation in the images and texts; Among them, the cross-modal hashing retrieval model includes: A semantic feature extraction module, which is used to extract the semantic features of different-modal data in the training samples and predict the hash codes corresponding to the corresponding modalities; A tensor fusion module, which is used to fuse the semantic features of different modal data by using a tensor fusion operator, and use the hash code corresponding to the fused semantic features to semantically guide the hash code predicted in the semantic feature extraction module; the tensor fusion operator is: Γ = ((Γ c ×1T i )×2T t )×3T o , Γ c is a multi-modal weight matrix, which is used to guide the weight assignment in the process of fusing the semantic features of different modalities; T i , T t , T o are respectively weight matrices for a single modality or task; the × k operator represents the projection calculation of the semantic feature corresponding to the k-th modality and the weight matrix. An adversarial feature fusion module, which is used to perform adversarial training with the semantic features of different-modal data to semantically guide the semantic features of different-modal data in the semantic feature extraction module; Application stage: Inputting the text and pictures to be retrieved into the trained semantic feature extraction module to obtain the hash codes corresponding to the corresponding modalities, and obtaining the retrieval result according to the hash codes; In the semantic feature extraction module, the semantic features of the different modality data include the picture semantic feature f i , the text semantic feature f t , and the label semantic feature f l ; In the tensor fusion module, the semantic features of different-modal data are fused by using the tensor fusion operator, including: The picture semantic feature f i and the text semantic feature f t After passing through a non-linear mapping layer respectively, the corresponding picture feature tensor (f i ) T T i and the text feature tensor (f t ) T T t ; Adopt a multi-modal weight matrix Γ c For the picture feature tensor (f i ) T T i and the text feature tensor (f t ) T T t Perform fusion and weight assignment to obtain the fused tensor The fused tensor and the label semantic feature f l After passing through a non-linear mapping respectively, the corresponding image-text fusion feature tensor and the label feature tensor (f l ) T T t ; Adopt a multi-modal weight matrix Γ c For the picture-text fusion feature tensor And the label feature tensor (f l ) T T t Perform fusion and weight assignment to obtain the re-fused tensor z: After passing the tensor z through another layer of non-linear mapping, the fused semantic feature f is obtained f : f f = z T T o .
2. The method according to claim 1, characterized in that, When performing tensor fusion each time, it also includes performing rank reduction on the multi-modal weight matrix Γ c which includes: Denote the k-th result z in the current subtensor fusion result as: k as follows: where and After two semantic features of different modalities to be fused respectively undergo a layer of non - linear mapping, two feature tensors, Γ, are obtained. c [:,:,k] is the multi - modal weight matrix for the k - th result during tensor fusion, where the value of k is 1, 2, …, n, and n is the total number of data pairs in the dataset. Each multimodal weight matrix is represented as the sum of R first-order matrices: is the outer product operation, m r and n r are the first-order weight matrices of the weight matrices corresponding to the image and text modality data respectively; where the rank of the multimodal weight matrix corresponding to each data pair is a constant R; Define R matrices and Then the tensor after the semantic feature fusion of two different modalities is as follows: where, (A)*(B) represents the intelligent product of element A and element B; respectively represent the corresponding feature lengths after the Tucker decomposition of the picture data, text data, and the fused data.
3. The method according to claim 1 or 2, characterized in that, The semantic feature extraction module includes: A semantic feature extraction layer, which is used to extract the semantic features of different-modal data in the training samples; A first label prediction layer, which is used to predict the labels corresponding to the semantic features of the different-modal data; A first hash code generation layer, which is used to predict the hash codes corresponding to the semantic features of the different-modal data; In the training stage, the loss function of the semantic feature extraction module is: Among them, represents the hash code of the j-th multi-modal data pair predicted by the first hash code generation layer, * represents different modalities of the data; represents the hash code corresponding to the label in the i-th data pair, H * is the hash code predicted by the first hash code generation layer, H l represents the hash code corresponding to the sample label; is the predicted label, L is the label corresponding to the sample; S ij is the similarity matrix of the data pair samples; β * , γ * , δ * are hyperparameters to be learned respectively; n represents the total number of data pairs in the dataset.
4. The method according to claim 1, characterized in that, The adversarial feature fusion module includes a feature discriminator, and the feature discriminator is also used to perform adversarial training with the fused semantic features output by the tensor fusion module; The adversarial feature fusion module also includes a hash code discriminator, and the hash code discriminator is used to perform adversarial training with the hash codes predicted by the semantic feature extraction module.
5. The method according to claim 4, characterized in that, The tensor fusion module is further configured to fuse the doped picture semantic feature f i and the text semantic feature f t to obtain a doped and fused feature; the feature discriminator is further configured to perform adversarial training on the doped and fused feature; Among them, the picture semantic feature f i and the text semantic feature f t are doped and fused, including: Multiply the picture semantic feature f i by the first coefficient σ to obtain the feature f i σ; multiply the text semantic feature f t by the second coefficient ζ to obtain the feature f t ζ; when the picture semantic feature f i is doped with the text semantic feature f t , σ >> ζ; when the text semantic feature f t is doped with the picture semantic feature f i , ζ >> σ; Input the feature f i σ and the feature f t ζ into the tensor fusion module for feature fusion.
6. The method according to claim 4, characterized in that, In the training phase, it also includes performing gradient penalty on the feature discriminator, including: regarding the label semantic feature f l as the real data for adversarial learning, regarding the image semantic feature f i and the text semantic feature f t as the generated data for adversarial learning, performing weighted average on the distributions of the real data and the generated data to obtain the critical data; Calculate the gradient of the critical data in the feature discriminator, and penalize the gradient through a loss function with gradient penalty; where the loss function L of the gradient penalty penaf is: represents the critical data, represents the gradient calculation of the critical data backpropagated by the feature discriminator, D f represents the feature discriminator.
7. The method according to claim 4, characterized in that, In the semantic feature extraction module, a loss function for semantically guiding the predicted hash code in the semantic feature extraction module with the hash code corresponding to the fused semantic feature is as follows: Among them, H f is the hash code corresponding to the fused semantic feature, and H * is the hash code predicted in the semantic feature extraction module; In the adversarial feature fusion module, the loss function for semantically guiding the semantic features of different-modal data in the semantic feature extraction module is: Among them, represents the loss of guiding the semantic features of the image with the adversarial feature fusion module, represents the loss function of guiding the semantic features of the text with the adversarial feature fusion module; Λ f* represents the output of the prediction by the feature discriminator, Λ h* is the output of the prediction by the hash code discriminator * Taking i, t, and l respectively, they correspondingly represent that the semantic features input to the discriminator come from the image, text, and label modality data respectively.
8. A cross-modal hashing retrieval feature fusion system, characterized in that, Including a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Cross-modal image-text mutual search method based on tensor fusion and reordering
CN110442741A
Image-text retrieval method and system based on attention mechanism and gating mechanism
CN112966135A