Global semantic guided cross-modal hash retrieval method
By introducing a tag-based modal specific feature enhancement module and a global semantic structure capture module in the cross-modal hash retrieval method, the problem of ignoring the structural relationship between data points and multi-source semantic association in the existing technology is solved, and efficient cross-modal retrieval and improvement of hash code quality is achieved.
Patent Information
- Application Number
- CN202510253930.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-23
AI Technical Summary
When using semantic labels to construct similarity matrix to guide hash function learning, the existing cross-modal hash retrieval method ignores the inherent structural relationships and multi-source semantic associations between data points, resulting in poor performance of hash codes.
A cross-modal hash retrieval method for global semantic guidance is proposed. Through the modal specific feature enhancement module and the global semantic structure capture module based on labels, consistent multi-source semantic associations and global semantic structure relationships are mined, refined global semantic enhancement cross-modal consistency features are generated, and comparative learning is carried out for semantic structure guidance.
This method can not only efficiently explore the consistency information between different modal data, but also accurately capture the global semantic structural relationship between samples and generate hash codes with excellent recognition and quality.
Smart Images

Figure SMS_256 
Figure QLYQS_14 
Figure QLYQS_35
Abstract
Description
Technical Field
[0001] The invention relates to a global semantics-guided cross-modal hash retrieval method, belonging to the technical field of cross-modal retrieval of multimedia retrieval. Background Art
[0002] With the surge in the amount of multimedia data from various sources, the demand for cross-modal retrieval methods is growing, and traditional unimodal retrieval methods can no longer meet the demand. Cross-modal retrieval is an important and common task in the field of multimedia analysis, which aims to find relevant data in different media types, such as searching text with images or searching videos with text. To improve the efficiency of these tasks, hashing technology has become a research focus due to its efficient storage and fast computation. Cross-modal hashing (CMH) aims to effectively map high-dimensional data points of different media types into a common compact binary encoding space by preserving the data structure or semantic relationship, thereby facilitating fast retrieval from large multimedia datasets.
[0003] Currently, most supervised CMH methods use similarity matrices constructed from semantic labels to guide hash function learning. Thanks to semantic guidance, these methods can generate more accurate hash codes. However, they still have limitations: they mainly rely on class labels as supervisory information, such as using label similarity matrices to generate shared hash codes, but ignore the inherent structural relationship between data points, and lack exploration of multi-source semantic associations (image-label, text-label, and cross-modal associations). In addition, although some supervised CMH methods have designed effective hash functions to learn semantically preserved hash codes, they are often limited to modality-specific features and may miss key cross-modal consistency information, resulting in poor hash code performance. Summary of the invention
[0004] The purpose of the present invention is to overcome the deficiencies of the above-mentioned prior art and provide a global semantics-guided cross-modal hash retrieval method.
[0005] The technical solution provided by the present invention is as follows: a global semantics-guided cross-modal hash retrieval method, characterized in that it comprises the following steps: 1) Step S1, using the MS-COCO dataset, establish an image dataset, a text dataset, and a dataset of labels corresponding to the image text, and divide the three datasets into a training set and a test set; 2) Step S2, constructing a tag-based modality-specific feature enhancement module, aligning image features, text features and tag features through three loss functions, and mining consistent multi-source semantic associations; 3) Step S3, construct a global semantic structure capture module to model the global semantic structure relationship between samples; generate global semantic enhanced cross-modal consistency features through this module; 4) Step S4, using the refined global semantics to enhance cross-modal consistency features and fuse them with image features and text features respectively, to obtain image features and text features that preserve global semantics; and performing semantic structure-guided contrastive learning; 5) Step S5, constructing the total objective function of the global semantics-guided hashing method on the training set, implementing the method using Pytorch, and selecting SGD as the optimization algorithm; generating hash codes for images and texts; 6) Step S6, performing cross-modal retrieval based on the hash codes of the images and the hash codes of the texts in the test set.
[0006] Furthermore, in step S1: The training set is represented as ,contain samples; The representation of a sample is ,in, Represents an image, Represents text, Represents the label; AlexNet convolutional neural network is used to extract the features of the images in the training set to obtain image features, which are represented as ; Use the BOW algorithm and fully connected neural network to extract the features of the text in the training set and obtain the text features, which are expressed as ; Use the BOW algorithm and fully connected neural network to extract the features of the labels in the training set and obtain the label features, which are expressed as .
[0007] Furthermore, the step S2 comprises the following steps: Step S21, introduce image-label loss function and the text-label loss function , respectively align the image features and text features with the corresponding label features; The definition is as follows: ; in, For the The image features of samples, For the The text features of samples, For the The label features of the samples; is the label similarity matrix, if ,but ,otherwise ;if ,but ,otherwise Indicates The labels of samples, Indicates The labels of samples, Indicates The labels of samples, and Represents a transpose operation; Step S22, introducing image-text loss function , align image features and text features: ; Among them, if ,but ,otherwise Represents a transpose operation; In step S23, the tag-based modality-specific feature enhancement loss is expressed as: ; in, is the balance coefficient.
[0008] Furthermore, the step S3 comprises the following steps: Step S31: The outputs of the three feature extraction networks, i.e., image features , label features and text features Splicing into comprehensive features ;in, , , is the dimension size; Step S32, using a linear mapping matrix With comprehensive characteristics Perform matrix multiplication to promote cross-modal fusion of all modalities and obtain initial cross-modal consistency features : = ; in, express The row vector of , , express row vector; using two different linear mapping matrices , With comprehensive characteristics Perform matrix multiplication to obtain query features and key features : ; in, Used to generate query features , Used to generate key features , ; Step S33, according to the self-attention mechanism, the global semantic structure relationship matrix between all samples is defined as: ; Using the global semantic structure relationship matrix Initial cross-modal consistency features Enhanced to obtain global semantic enhanced cross-modal consistency features , the specific steps are as follows: ; in, Representation Matrix Middle The vector representation of the data of samples is Indicates The samples and The similarity between samples is removed by MLP. The redundancy of , and the refined global semantics to enhance cross-modal consistency features , as shown in the following formula: ; Among them, MLP is a multi-layer perceptron.
[0009] Furthermore, the step S4 comprises the following steps: Step S41: Fusing refined global semantics to enhance cross-modal consistency features and image features , and obtain the image features that preserve global semantics ; Fusion refined global semantics to enhance cross-modal consistency features and text features Get text features that preserve global semantics and The formula is: ; in, Indicates The global semantics preserved image features of samples, Indicates The global semantics of the samples are preserved in the text features; Step S42, define the loss function of contrastive learning guided by semantic structure, and the contrastive learning loss with the image as the anchor point is: ; in, Representation and From the same sample The global semantics of the text features preserved; Indicates The global semantics of the samples are preserved in the text features; The contrastive learning loss with text as anchor is: ; in, Representation and From the same sample Global semantically preserved image features; Indicates The global semantics of the image features preserved by samples; In the above formula represents the cosine similarity, is the temperature parameter; is the activation function; The semantic structure guided contrastive learning loss is defined as: .
[0010] Furthermore, the step S5 comprises the following steps: Step S51, constructing the overall objective function: ; in, is a hyperparameter used as a balancing factor; Step S52: Hash code of the image and the hash code of the text The image features preserved by global semantics and the text features preserved by global semantics are generated by the following formula: ; in, is a symbolic function.
[0011] Furthermore, in the step S6: the images of all samples in the test set are sent to the AlexNet convolutional neural network to obtain image features; the texts of all samples are sent to the fully connected neural network through the BOW algorithm to obtain text features; the image features and the text features are used to obtain the hash codes of the images and the texts in the test set according to the method of steps S1 to S5 of the training process; cross-modal retrieval is performed, and based on the hash codes, data of another modality related to the data to be queried in the test set is queried in the training set.
[0012] The beneficial effects of the present invention are as follows: the present invention introduces a tag-based modality-specific feature enhancement module, and aligns the features of the image and the tag, the features of the text and the tag, and the features of the image and the text through the three loss functions of image-tag loss, text-tag loss and image-text loss, respectively, in order to deeply mine the consistent multi-source semantic association information. In addition, in order to accurately capture the global semantic structure relationship, the present invention proposes a global semantic structure capture module (GSSA) based on the transformer attention mechanism, which captures the global semantic structure relationship through the similarity within and between samples, and generates refined global semantic enhancement cross-modal consistency features. Finally, the present invention integrates modality-specific features and refined global semantic enhancement cross-modal consistency features, and uses the global semantic structure relationship to guide contrastive hash learning to optimize the hash network to obtain a hash code with global semantics preservation. Compared with many existing technologies, the present invention can not only efficiently explore the consistency information between different modal data, but also accurately capture the global semantic structure relationship between samples, and the resulting hash representation has excellent recognition. In addition, by guiding the contrastive hash learning and optimizing the hash network through the global semantic structure relationship, the hash code quality of the present invention is significantly improved. DETAILED DESCRIPTION
[0013] The specific implementation modes of the present invention are described in detail below: Although the present invention specifies two modalities, image and text, the algorithm can be easily extended to other modalities. For the convenience of description, the present invention only considers two modalities, image and text.
[0014] A global semantics-guided cross-modal hash retrieval method comprises the following steps: 1) Step S1, using the MS-COCO dataset, establish an image dataset, a text dataset, and a dataset of labels corresponding to the image text, and divide the three datasets into a training set and a test set; The training set is represented as ,contain samples; The representation of a sample is ,in, Represents an image, Represents text, Represents the label. The features of the images in the training set are extracted using the AlexNet convolutional neural network to obtain the image features, which are represented as ; Use the BOW algorithm and fully connected neural network to extract the features of the text in the training set and obtain the text features, which are expressed as ; Use the BOW algorithm and fully connected neural network to extract the features of the labels in the training set and obtain the label features, which are expressed as ; 2) Step S2, construct a label-based modality-specific feature enhancement module, align image features, text features and label features through three loss functions, and mine consistent multi-source semantic associations.
[0015] Specifically include the following steps: Step S21, usually, image features and text features are semantically aligned with their corresponding label features. We introduce the image-label loss function and the text-label loss function , respectively aligning the image features and text features with the corresponding label features.
[0016] The definition is as follows: ; in, For the The image features of samples, For the The text features of samples, For the Label features of samples; is the label similarity matrix, if ,but ,otherwise ;if ,but ,otherwise . Indicates The labels of samples, Indicates The labels of samples, Indicates The labels of samples, and Represents a transpose operation.
[0017] Step S22: Subsequently, in order to align image features and text features to reduce the heterogeneous differences between images and texts, an image-text loss function is introduced. : ; Among them, if ,but ,otherwise . Represents a transpose operation.
[0018] In step S23, the tag-based modality-specific feature enhancement loss is expressed as: ; in, is the balance coefficient, and its value is 0.10.
[0019] 3) Step S3: Construct a global semantic structure capture module, which effectively models the global semantic structure relationship between samples. Through this module, global semantic enhanced cross-modal consistency features are generated;
[0020] It includes the following steps: Step S31: First, concatenate the outputs of the three feature extraction networks, namely the image features , label features and text features into a comprehensive feature . Among them, , , is the dimension size.
[0021] Step S32: Use a linear mapping matrix to perform matrix multiplication with the comprehensive feature to promote cross-modal fusion of all modalities and obtain the initial cross-modal consistency feature : = ; Among them, represents the row vector of , , , represents the row vector of ; Similarly, use two different linear mapping matrices , to perform matrix multiplication with the comprehensive feature to obtain the query feature and the key feature : ; Among them, is used to generate the query feature , is used to generate the key feature , .
[0022] In the self-attention mechanism, the query feature and the transpose of the key feature perform a dot product operation to calculate the attention weight to capture the global structural relationship between features. , , serve as the parameter matrices in the Transformer self-attention layer and are automatically learned by training the Transformer model.
[0023] Step S33, according to the self-attention mechanism, the global semantic structure relationship matrix between all samples is defined as: ; Then, we use the global semantic structure relationship matrix Initial cross-modal consistency features Enhanced to obtain global semantic enhanced cross-modal consistency features , this feature benefits from other samples with strong semantic structure correlation and can effectively maintain consistent semantic structure relationship. The specific steps are as follows: ; in, Representation Matrix Middle The vector representation of the data of samples is Indicates The samples and The similarity relationship between samples. By all modes In order to alleviate this situation, the output is processed by MLP to reduce redundancy, and after removing the redundancy, a refined global semantic enhancement cross-modal consistency feature is obtained. , as shown in the following formula: ; Among them, MLP is a multi-layer perceptron.
[0024] 4) Step S4: Fusion of refined global semantics to enhance cross-modal consistency features and image features , and obtain the image features that preserve global semantics ; Fusion refined global semantics to enhance cross-modal consistency features and text features , get the text features that preserve global semantics .right and Conduct contrastive learning guided by semantic structure.
[0025] It comprises the following steps: Step S41: Fusing refined global semantics to enhance cross-modal consistency features and image features , and obtain the image features that preserve global semantics ; Fusion refined global semantics to enhance cross-modal consistency features and text features Get text features that preserve global semantics . and The formula is expressed as: ; in, Indicates The global semantics preserved image features of samples, Indicates The global semantics preserved text features of samples.
[0026] Step S42, define the loss function of contrastive learning guided by semantic structure, and the contrastive learning loss with the image as the anchor point is: ; in, Representation and From the same sample The global semantics of the text features are preserved. Indicates The global semantics preserved text features of samples.
[0027] The contrastive learning loss with text as anchor is: ; in, Representation and From the same sample The global semantics of the image are preserved. Indicates The global semantics preserved image features of samples.
[0028] In the above formula represents the cosine similarity, is the temperature parameter, which is taken as 0.50 here. It is an activation function whose purpose is to relax the feature vector to reduce the information loss caused by quantizing image features and text features into image hash codes and text hash codes.
[0029] Finally, the semantic structure guided contrastive learning loss is defined as: ; 5) Step S5, construct the overall objective function of the global semantics-guided hashing method on the training set, implement the method using Pytorch, and select SGD as the optimization algorithm; generate the hash code of the image and the hash code of the text ; It comprises the following steps: Step S51, constructing the overall objective function: ; in, is a hyperparameter used as a balance coefficient and takes a value of 0.10.
[0030] Step S52: Hash code of the image and the hash code of the text The image features preserved by global semantics and the text features preserved by global semantics can be generated by the following formula: ; in, is a symbolic function.
[0031] 6) Step S6, send the images of all samples in the test set to the AlexNet convolutional neural network to obtain image features; send the texts of all samples to the fully connected neural network through the BOW algorithm to obtain text features; the image features and text features are used to obtain the hash codes of the images and texts in the test set according to the method of steps S1 to S5 of the training process. Perform cross-modal retrieval, and query the data of another modality related to the data to be queried in the test set in the training set based on the hash code.
[0032] This example is trained and validated on the MS-COCO dataset, which contains a total of 123,287 pairs of images and texts. After filtering out sample pairs without label information, we used 122,218 sample pairs containing 80 categories. We randomly selected 7762 pairs as the query set, and the remaining pairs were used for the retrieval set, in which we randomly selected 18,000 pairs of images and texts for training. We used the AlexNet convolutional neural network to extract 4096-dimensional image features and the BOW algorithm (bag of words model) to extract 2000-dimensional text features. The mean average precision (mAP@50) is used as the performance evaluation criterion, where 50 means that the mAP value is calculated by the first 50 returned samples. This scheme is compared with DSPH (Y. Huo, Q. Qin, J. Dai, L. Wang, W. Zhang, L. Huang, and C. Wang. DeepSemantic-Aware Proxy Hashing for Multi-Label Cross-Modal Retrieval. IEEETransactions on Circuits and Systems for Video Technology, 34(1):576–589,2024.), and the accuracy of 16-bit, 32-bit, 64-bit and three code lengths on image retrieval text and text retrieval image tasks is shown in Table 1.
[0033] The MSCOCO dataset is used for verification, and the retrieval accuracy is shown in Table 1.
[0034] Table 1 mAp scores of different bits on the MS-COCO dataset
[0035] This paper proposes a new cross-modal hash retrieval method, Global Semantic Guided Hashing (GCSGH), to perform cross-modal retrieval. This hash learning framework has several key features. First, the proposed tag-based modality-specific feature enhancement module effectively explores consistent and rich multi-source semantic association information between three different source data: images, tags, and texts. Secondly, the designed semantic structure capture (GSSA) module captures the global semantic structural relations between samples and modalities, and generates refined global semantic enhanced cross-modal consistency features. We further merge these features with modality-specific features and incorporate them into the contrastive hash learning process guided by global semantic structural relations. Therefore, the generated hash code can excellently preserve the global semantic structural relations and modality consistency information. Experimental results on benchmark datasets show that GCSGH outperforms the existing best baselines in cross-modal retrieval tasks.
[0036] It should be understood that the parts not elaborated in detail in this specification belong to the prior art. The above embodiments are only descriptions of the preferred implementation methods of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made by ordinary engineers and technicians in this field to the technical solution of the present invention should fall within the protection scope determined by the claims of the present invention.
Claims
1. A global semantics-guided cross-modal hash retrieval method, characterized in that: It includes the following steps: 1) Step S1, using the MS-COCO dataset, establish an image dataset, a text dataset, and a dataset of labels corresponding to the image text, and divide the three datasets into a training set and a test set; 2) Step S2, constructing a tag-based modality-specific feature enhancement module, aligning image features, text features and tag features through three loss functions, and mining consistent multi-source semantic associations; 3) Step S3, construct a global semantic structure capture module to model the global semantic structure relationship between samples; generate global semantic enhanced cross-modal consistency features through this module; 4) Step S4, using the refined global semantics to enhance cross-modal consistency features and fuse them with image features and text features respectively, to obtain image features and text features that preserve global semantics; and performing semantic structure-guided contrastive learning; 5) Step S5, constructing the total objective function of the global semantics-guided hashing method on the training set, implementing the method using Pytorch, and selecting SGD as the optimization algorithm; generating hash codes for images and texts; 6) Step S6, performing cross-modal retrieval based on the hash codes of the images and the hash codes of the texts in the test set.
2. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: In the step S1: The training set is represented as ,contain samples; The representation of a sample is ,in, Represents an image, Represents text, Represents the label; AlexNet convolutional neural network is used to extract the features of the images in the training set to obtain image features, which are represented as ; Use the BOW algorithm and fully connected neural network to extract the features of the text in the training set and obtain the text features, which are expressed as ; Use the BOW algorithm and fully connected neural network to extract the features of the labels in the training set and obtain the label features, which are expressed as .
3. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: The step S2 comprises the following steps: Step S21, introduce image-label loss function and the text-label loss function , respectively align the image features and text features with the corresponding label features; The definition is as follows: ; in, For the The image features of samples, For the The text features of samples, For the The label features of the samples; is the label similarity matrix, if ,but ,otherwise ;if ,but ,otherwise Indicates The labels of samples, Indicates The labels of samples, Indicates The labels of samples, and Represents a transpose operation; Step S22, introducing image-text loss function , align image features and text features: ; Among them, if ,but ,otherwise Represents a transpose operation; In step S23, the tag-based modality-specific feature enhancement loss is expressed as: ; in, is the balance coefficient.
4. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: The step S3 comprises the following steps: Step S31: The outputs of the three feature extraction networks, i.e., image features , label features and text features Splicing into comprehensive features ;in, , , is the dimension size; Step S32, using a linear mapping matrix With comprehensive characteristics Perform matrix multiplication to promote cross-modal fusion of all modalities and obtain initial cross-modal consistency features : = ; in, express The row vector of , , express row vector; using two different linear mapping matrices , With comprehensive characteristics Perform matrix multiplication to obtain query features and key features : ; in, Used to generate query features , Used to generate key features , ; Step S33, according to the self-attention mechanism, the global semantic structure relationship matrix between all samples is defined as: ; Using the global semantic structure relationship matrix Initial cross-modal consistency features Enhanced to obtain global semantic enhanced cross-modal consistency features , the specific steps are as follows: ; in, Representation Matrix Middle The vector representation of the data of samples is Indicates The samples and The similarity between samples is removed by MLP. The redundancy of , and the refined global semantics to enhance cross-modal consistency features , as shown in the following formula: ; Among them, MLP is a multi-layer perceptron.
5. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: The step S4 comprises the following steps: Step S41: Fusing refined global semantics to enhance cross-modal consistency features and image features , and obtain the image features that preserve global semantics ; Fusion refined global semantics to enhance cross-modal consistency features and text features Get text features that preserve global semantics and The formula is: ; in, Indicates The global semantics preserved image features of samples, Indicates The global semantics of the samples are preserved in the text features; Step S42, define the loss function of contrastive learning guided by semantic structure, and the contrastive learning loss with the image as the anchor point is: ; in, Representation and From the same sample The global semantics of the text features preserved; Indicates The global semantics of the samples are preserved in the text features; The contrastive learning loss with text as anchor is: ; in, Representation and From the same sample Global semantically preserved image features; Indicates The global semantics of the image features preserved by samples; In the above formula represents the cosine similarity, is the temperature parameter; is the activation function; The semantic structure guided contrastive learning loss is defined as: 。 6. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: The step S5 comprises the following steps: Step S51, constructing the overall objective function: ; in, is a hyperparameter used as a balancing factor; Step S52: Hash code of the image and the hash code of the text The image features preserved by global semantics and the text features preserved by global semantics are generated by the following formula: ; in, is a symbolic function.
7. A global semantics-guided cross-modal hash retrieval method as claimed in claim 1, characterized in that: In the step S6: sending the images of all samples in the test set to the AlexNet convolutional neural network to obtain image features; The text of all samples is sent to the fully connected neural network through the BOW algorithm to obtain text features; the image features and text features are used to obtain the hash code of the image and the hash code of the text in the test set according to the method of steps S1 to S5 in the training process; cross-modal retrieval is performed, and based on the hash code, the data of another modality related to the data to be queried in the test set is queried in the training set.