Cross-modal hash retrieval method based on label-driven semantic perception learning

By constructing a tag-driven semantic perception learning method, the problems of information loss and redundancy in cross-modal retrieval are solved, and the semantic information of cross-modal samples is enhanced by using teacher models and attention mechanisms, achieving more efficient cross-modal retrieval accuracy.

CN120353989APending Publication Date: 2025-07-22HEBEI UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510492565.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the existing cross-modal search methods process different modal data, there are problems of information loss and redundancy, which leads to inaccurate search results and makes it difficult to effectively use label information for semantic understanding.

Method used

A tag-driven semantic perception learning method is constructed, initial feature extraction is performed through CLIP, combined with feature extraction, information supplementation, filtering and hash modules of the teacher model, the semantic information of cross-modal samples is enhanced by using GCN and attention mechanisms, and the semanticity of the hash code is improved by maintaining the loss through quantitative differences.

Benefits of technology

It improves the accuracy and efficiency of cross-modal retrieval, significantly improves the retrieval performance, especially showing higher average accuracy on large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353989A_ABST
    Figure CN120353989A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-modal hash retrieval method based on label-driven semantic perception learning. The method comprises the following steps: S1, acquiring a data set; s2, constructing a CLIP, and inputting samples in the data set into the CLIP for initial feature extraction; s3, constructing a teacher model, wherein the teacher model comprises a feature extraction module, GLoVe, an information supplement module, an information filtering module and a feature hash module; s4, constructing a student model, wherein the student model comprises an image MLP and a text MLP; s5, constructing target loss functions corresponding to the teacher model and the student model; s6, parameters of the teacher model are updated according to the target loss function, and training of the student model is guided after training of the teacher model is completed; and S7, generating respective corresponding hash codes for the data in the database and the query data by using the trained student model, and outputting a retrieval result according to the distance between the hash codes. The information filtering module reduces the information redundancy of the sample, and the completeness of the semantic content of the sample is improved through the information supplementing module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-modal retrieval method, specifically a cross-modal hashing retrieval method based on label-driven semantic perception learning. Background Art

[0002] With the rapid development of the Internet, the amount of multi-modal data is increasing day by day. Since the implementation of cross-modal retrieval is beneficial to the unified management and analysis of data in different modalities on the Internet, it has gradually become a hot issue. The common representation features of different modalities generated by traditional methods have a high dimensionality, resulting in a high cost for storing data features and a low retrieval efficiency. To solve these problems, the cross-modal hashing retrieval method maps the high-dimensional common representation features into low-dimensional binary hash codes, and quickly calculates the similarity between samples through the Hamming distance, effectively improving the retrieval efficiency.

[0003] In recent years, compared with shallow cross-modal hashing, since deep cross-modal hashing methods can learn hash codes with powerful representation capabilities, they have gradually become the mainstream of research. According to whether label information is used, the current deep cross-modal hashing methods can be roughly divided into supervised deep cross-modal hashing and unsupervised deep cross-modal hashing. Unsupervised deep cross-modal hashing focuses on the feature structure and distribution of the data itself, and learns the hash function by exploring the potential structure and relationship between modalities. Since no label data is required to participate in the training, unsupervised deep cross-modal hashing is more suitable for large-scale data retrieval. However, due to the lack of label information, the hash codes generated by these methods contain less semantic information and it is difficult to obtain good retrieval performance. In contrast, supervised deep cross-modal hashing methods use label information as supervision information to guide the learning process of hash codes, and can more accurately capture the features and relationships of the data, thereby generating cross-modal hash codes containing rich semantic information.

[0004] Although the current cross-modal hashing methods have all achieved certain results, the implementation of these methods is based on the ideal situation, that is, the data in different modalities in the data pair can completely match the label information. However, due to the inherent characteristics of different modality data and the noise that may be introduced by manual annotation, this assumption is usually difficult to hold. As Figure 1 shown, the images and texts in the dataset cannot completely correspond to the labels one by one. Specifically, it can be summarized into two situations: 1. Information missing, that is, some label information is not included in a certain modality data. For example, in Figure 1 (a), the label information of "dining table" is not described in the corresponding text, and in Figure 1 (b), "wine glass", "bottle" and "oven" are missing in the corresponding text. 2. Information redundancy, that is, a certain modality data contains information of non-semantic labels. For example, inFigure 1 In (c), the trees in the image and the word "mountain" in the text do not appear in the corresponding tags. Therefore, simply aligning the image-text pairs in the dataset directly may lead to deviations in semantic understanding during the cross-modal retrieval process, inaccurate retrieval results, and thus affect the retrieval performance. Summary of the Invention

[0005] The objective of the present invention is to provide a cross-modal hashing retrieval method based on label-driven semantic perception learning to solve the problem of inaccuracy in existing cross-modal retrieval technologies.

[0006] The objective of the present invention is achieved as follows:

[0007] A cross-modal hashing retrieval method based on label-driven semantic perception learning includes the following steps:

[0008] S1. Obtain a cross-modal retrieval dataset; the dataset includes image samples, text samples, and labels corresponding to each sample respectively; and the display content of an image sample is the expression content of a corresponding text sample.

[0009] S2. Construct Contrastive Language-Image Pre-Training (CLIP), and input the samples in the dataset into CLIP for initial feature extraction.

[0010] S3. Construct a teacher model, and the teacher model includes:

[0011] A feature extraction module for extracting features from the data output by CLIP to obtain image representations and text representations.

[0012] Global Vectors for Word Representation (GLoVe) for performing initial feature extraction on the labels corresponding to the samples to obtain label representations.

[0013] An information supplementation module for respectively supplementing information for the data output by the feature extraction module and the data output by GLoVe.

[0014] An information filtering module for filtering the data output by the information supplementation module; and

[0015] A feature hashing module for mapping the data output by the information filtering module into teacher hash codes.

[0016] S4. Build a student model, where the student model includes an image Multilayer Perceptron (MLP) and a text MLP, and both the image MLP and the text MLP include three fully connected layers; the student model is used to map the data output by CLIP into student hash codes;

[0017] S5. Construct the target loss function of the teacher model according to the teacher hash code, and construct the target loss function of the student model according to the teacher hash code and the student hash code;

[0018] S6. According to the target loss function corresponding to the teacher model, train the parameters corresponding to the teacher model; after the teacher model is trained, according to the trained teacher model and the target loss function corresponding to the student model, train the parameters corresponding to the student model;

[0019] S7. Use the trained student model to generate respective corresponding hash codes for the data in the database and the query data, and output the retrieval results according to the distance between the hash codes.

[0020] Further, the feature extraction module includes a text feature extraction module and an image feature extraction module, and both the text feature extraction module and the image feature extraction module include two fully connected layers.

[0021] Further, the information filtering module includes at least two attention modules, and each attention module includes a first fully connected layer, a second fully connected layer, a third fully connected layer, and a softmax function;

[0022] The data processing method of the attention module is as follows:

[0023] The input is divided into three paths corresponding to three fully connected layers. After the outputs of the first fully connected layer and the second fully connected layer pass through the softmax function, they are calculated with the output of the third fully connected layer to obtain the attention value of the input data.

[0024] Further, the feature hashing module includes an image hashing layer, a text hashing layer, and a label hashing layer, and each of the image hashing layer, the text hashing layer, and the label hashing layer includes a connection layer.

[0025] Further, the information supplement module includes a graph convolutional network (GCN);

[0026] The specific method for the information supplement module to supplement the data output by the feature extraction module and GLoVe in step S3 is as follows:

[0027] Sa3-1. Build a sample-label graph based on the data output by the feature extraction module and GLoVe. The vertices of the sample-label graph include image representations, text representations, and label representations;

[0028] Sa3-2. Calculate the adjacency matrix according to the data of the vertices of the sample-label graph;

[0029] Sa3-3. Input the adjacency matrix and the vertex feature matrix into the GCN to obtain the finally output features.

[0030] Furthermore, the specific way to calculate the adjacency matrix is as follows:

[0031] In the sample-label graph, no edges are set between data of the same modality, and edges are set between corresponding image and text pairs; Determine the edges between the text representation and the label representation according to the similarity between the text representation and the label representation, and determine the edges between the image representation and the label representation according to the similarity between the image representation and the label representation; Determine the edges between label representations by using the feature similarity between labels. The obtained adjacency matrix A of the sample-label graph is:

[0032]

[0033] where A *2* represents the edge of "pointing to", *∈{v,t,l}, A v2v , A t2t , A v2t and A t2v are all identity matrices

[0034] Furthermore, the target loss function corresponding to the teacher model is:

[0035]

[0036] where is the semantic contrast loss, is the classification loss, and α1 is a hyperparameter;

[0037] The target loss function corresponding to the student model is:

[0038]

[0039] where is the teacher-student hash code alignment loss, is the quantization difference preservation loss, is the quantization loss, and β1 and β2 are hyperparameters.

[0040] Furthermore, the teacher hash code output by the teacher model includes a teacher image hash code and a teacher text hash code;

[0041] The calculation method of the semantic contrast loss is as follows:

[0042] Calculate the intra-modal semantic contrast loss of the image according to the teacher image hash code and the semantic weight; calculate the intra-modal semantic contrast loss of the image according to the teacher text hash code and the semantic weight; calculate the inter-modal semantic contrast loss according to the teacher image hash code, the teacher text hash code and the semantic weight; calculate the semantic contrast loss according to the intra-modal semantic contrast loss of the image, the intra-modal semantic contrast loss of the image and the inter-modal semantic contrast loss.

[0043] Furthermore, the student hash code output by the student model includes a student image hash code and a student text hash code; the teacher hash code output by the teacher model includes a teacher image hash code and a teacher text hash code;

[0044] The quantization difference preservation loss is:

[0045]

[0046] wherein, is the semantic weight, y i is the label of the i-th sample, y j is the label of the j-th sample, is the i-th student image hash code, is the j-th student image hash code, is the i-th student text hash code, is the j-th student text hash code, wherein, is the teacher image hash code, is the teacher text hash code.

[0047] Furthermore, step S7 includes the following steps:

[0048] S7-1. Input the data in the database into CLIP and the trained student model in sequence to obtain the hash codes of the data in the database;

[0049] S7-2. Input the query data into CLIP and the trained student model to obtain the hash code of the query data;

[0050] S7-3. Compare the hash code of the query data with the hash codes of the data in the database that are in different modalities from the query data to obtain the distance between the hash codes, and output the data in the database within the preset distance as the retrieval result.

[0051] The present invention proposes a cross-modal hashing retrieval method based on label-driven semantic-aware learning (LSLCHR). First, by constructing an information supplementation module, the information of category labels is supplemented into cross-modal samples to enhance the semantic information of cross-modal samples. Second, an information filtering module is constructed, and an attention mechanism guided by multi-labels is designed to filter out the non-semantic information of cross-modal samples. In addition, a quantization difference preservation loss is proposed to keep the difference between sample features consistent with the difference between their corresponding hash codes, effectively reducing the change in sample similarity caused by the binary quantization process of hash codes. Experimental results on three widely used datasets show that this method can achieve good retrieval performance.

[0052] The present invention mainly includes an information supplementation module and an information filtering module. The information supplementation module constructs a graph composed of cross-modal samples and category labels, and supplements the information of category labels into cross-modal samples through GCN to solve the problem of information loss and improve the completeness of the semantic content of cross-modal samples. The information filtering module uses multi-labels as a guide and uses the attention mechanism to extract the semantic features of cross-modal samples, filter out non-semantic features, and further align cross-modal features to solve the problem of information redundancy. In addition, to improve the semantics of binary hash codes, the present invention proposes a quantization difference preservation loss, which keeps the difference between sample features consistent with the difference between their corresponding hash codes, thereby improving the semantics of hash codes and effectively reducing the information loss during the quantization process of hash codes. Applying the student model established by the above method to cross-modal retrieval improves the accuracy of cross-modal retrieval.

[0053] Compared with other cross-modal retrieval methods on the MIRFlickr-25K, MS COCO, and NUS-WIDE datasets, the present invention has a higher mean average precision, and the present invention has achieved significant performance improvement compared with other baseline methods. Description of the Drawings

[0054] Figure 1 are images, texts, and their corresponding labels.

[0055] Figure 2 is the flow chart of the present invention.

[0056] Figure 3 is the data processing schematic diagram of the attention module. Detailed Implementation Manner

[0057] The present invention will be further described in detail below with reference to the drawings.

[0058] As Figure 2 shown, the cross-modal hashing retrieval method based on label-driven semantic perception learning of the present invention includes the following steps:

[0059] S1. Obtain a data set.

[0060] Among them, the data set includes image samples, text samples, and labels corresponding to each sample respectively; and the display content of an image sample is the expression content of a corresponding text sample, that is, they have the same semantics.

[0061] The training set O used in the present invention includes images, texts, and their corresponding labels, and the texts and images correspond to each other. where o i represents the i-th instance, and n is the number of instances in the data set. An instance where represents the i-th image, represents the i-th text. Each instance has a corresponding label vector, o i The corresponding label vector is y i where y i =(y i1 , y i2 ,..., y ic ), when o i belongs to the j-th category, y ij =1, otherwise, y ij =0. The purpose of cross-modal hashing retrieval is to learn two hashing functions and for mapping images and texts into hash codes B x ∈{-1, +1} n×k , B y ∈{-1, +1} n×k , and k is the length of the hash code.

[0062] S2. Construct CLIP, and input the samples in the data set into CLIP for initial feature extraction.

[0063] First, use CLIP as a feature extractor to extract 512-dimensional initial image features and initial text features

[0064]

[0065] S3. Construct a teacher model.

[0066] Among them, the teacher model includes a feature extraction module, GLoVe, an information supplement module, an information filtering module, and a feature hashing module.

[0067] The feature extraction module includes a text feature extraction module and an image feature extraction module. Both the text feature extraction module, the image feature extraction module, and GLoVe include two fully connected layers.

[0068] The feature extraction module extracts features from the data output by CLIP. The initial image features are input into the image feature extraction module to obtain the image representation f i v The initial text features are input into the text feature extraction module to extract the corresponding text representation f i t :

[0069]

[0070] where d is the dimension, and θ v are the network parameters of the fully connected layer for extracting the image representation, and θ t are the network parameters of the fully connected layer for extracting the text representation.

[0071] GLoVe extracts initial features from the labels corresponding to the samples to obtain the initial features e of each class label i and maps them to the label representation l through a fully connected layer i :

[0072]

[0073] where θ l are the network parameters of the fully connected layer for extracting the label representation.

[0074] Finally, after extracting the features of the data in the dataset, an image representation matrix a text representation matrix and a label representation matrix

[0075] can be obtained. The activation function of the fully connected layer is tanh(·).

[0076] The information supplementation module includes GCN.

[0077] The teacher model constructs a sample-label graph G. Among them, V is the vertex set of graph G, which is composed of the image representation, the text representation, and the label representation. The feature matrix of the vertex set is denoted as E is the edge set of graph G.

[0078] The adjacency matrix of G is A:

[0079]

[0080] where A *2* represents the edge from * to *, and * ∈ {v, t, l}.

[0081] To highlight the relationship between the labels and data of different modalities, no edges are set between data of the same modality in G, that is, A v2v and A t2t are both identity matrices For data of different modalities, the present invention only sets edges between corresponding image-text pairs, that is, A v2t and A t2v are also both identity matrices

[0082] To describe the relationship between the data and the labels, the present invention constructs A using the feature similarity between the data and the labels v2l and A t2l :

[0083]

[0084] where σ is the threshold value

[0085] The present invention takes the edges with feature similarity less than σ as noise edges and removes them. Since the present invention constructs the relationship between the data and the labels using feature similarity, therefore, A l2v =(A v2l ) T , A l2t =(A t2l ) T .

[0086] Similarly, the present invention constructs A using the feature similarity between the labels l2l :

[0087]

[0088] Finally, the adjacency matrix A of the graph G is:

[0089]

[0090] Input the adjacency matrix A of the sample-label graph G and the vertex feature matrix F into the GCN to supplement the information of the multi-modal data features and generate the label feature H containing the multi-modal data information (l+1) :

[0091]

[0092] where D is the degree matrix of A, D ii =∑ j A ij , W (l) is the transformation matrix of the l-th layer, d (l) is the input dimension of the l-th layer of the GCN, d(l+1) is the output dimension of the (l + 1)-th layer of the GCN, is the activation function, and the total number of layers of the GCN is set to g. The initial input of the GCN is H (0) = F, and the finally output feature matrix is Among them, is the supplementary image representation, is the supplementary text representation, is the supplementary label representation, and d' is the representation dimension output by the information supplementation module.

[0093] Among them, the activation function of the GCN is relu(·).

[0094] The supplementary image representation and supplementary text representation generated by the information supplementation module already include relatively complete semantic information, but still contain non-semantic redundant information that affects the retrieval performance, such as background information in images, grammatical structures in texts, etc. Some methods use the masking mechanism to decompose data features into modality-common features and modality-private features. However, these methods do not consider the semantic information of the data when decomposing the features.

[0095] The information filtering module uses the multi-label representation as a guide and combines the self-attention mechanism to filter out the redundant information irrelevant to semantics in the supplementary image representation and supplementary text representation.

[0096] Multi-head attention is added to the information filtering module. The information filtering module includes at least two attention modules. The inputs of the information filtering module enter all the attention modules respectively, and each data obtains multiple attention values Among them, M is the total number of attention heads. And the outputs of all the attention modules are concatenated and mapped to obtain the filtered representation of the sample is:

[0097]

[0098] Among them, △ ∈ {v, t}, is the linear mapping matrix.

[0099] As Figure 3 shown, each attention module includes a first fully connected layer, a second fully connected layer, a third fully connected layer, and a softmax function;

[0100] The data processing method of the attention module is:

[0101] The input is divided into three paths corresponding to the three fully connected layers. The outputs of the first fully connected layer and the second fully connected layer are calculated with the output of the third fully connected layer after passing through the softmax function to obtain the attention value of the input data.

[0102] The input data of each attention module is processed as follows: First, the supplementary representations of different modalities are respectively obtained through different linear mappings

[0103]

[0104] where, W △q 、W △k and are linear mapping matrices.

[0105] The traditional self-attention mechanism only relies on the sample representation itself for calculation to obtain the self-attention value, without considering the semantics of attention. To solve this problem, the multi-labels of the samples are used as semantic supervision information to generate more accurate attention values

[0106]

[0107] where, is the self-attention value corresponding to the i-th sample, is the semantic attention value corresponding to the i-th sample, is the multi-label representation corresponding to the i-th sample, and are linear mapping matrices, is the weight matrix.

[0108] Each data needs to go through the above calculations. For data of different modalities, filtered image representations and filtered text representations are obtained.

[0109] All the parameters in the information filtering module are denoted as θ a .

[0110] After information supplementation and information filtering of the samples, the feature hashing module maps the data output by the information filtering module to teacher hash codes.

[0111] The feature hashing module includes an image hashing layer, a text hashing layer, and a label hashing layer. The image hashing layer, the text hashing layer, and the label hashing layer all include a connection layer.

[0112] The filtered image representation is mapped to a teacher image hash code through the image hashing layer, the filtered text representation is mapped to a teacher text hash code through the text hashing layer, and the supplementary label representation is mapped to a label hash code through the label hashing layer:

[0113]

[0114] where, θh = {θ hv , θ ht , θ hl} are the learnable parameters of the feature hashing module of the teacher model.

[0115] During the training process, tanh(·) needs to be used as the activation function of the hashing layer to avoid the vanishing gradient problem caused by using sign(·).

[0116] S4. Construct the student model.

[0117] Among them, the student model includes an image MLP and a text MLP, and both the image MLP and the text MLP include three fully connected layers; the student model is used to map the data output by CLIP into student hash codes.

[0118] The student hash codes include student image hash codes and student text hash codes.

[0119] In order to learn a hashing function with information supplementation and information filtering functions, the present invention introduces a lightweight student model. The student model learns under the guidance of the teacher model, so as to distill the knowledge of the information supplementation module and the information filtering module in the teacher model into the student model. The student model is only composed of MLP, including an image MLP and a text MLP. Both the image MLP and the text MLP include three fully connected layers. The image MLP is used to map the initial image features extracted by CLIP into student image hash codes, and the text MLP is used to map the initial text features extracted by CLIP into student text hash codes:

[0120]

[0121] Among them, θ sv and θ st are the learnable parameters of the corresponding hashing function.

[0122] Similarly, during the training process, the hashing function uses tanh(·) as the activation function.

[0123] S5. Construct the target loss function of the teacher model according to the teacher hash codes, and construct the target loss function of the student model according to the teacher hash codes and the student hash codes.

[0124] In order to transfer the knowledge of the teacher module to the student module, a teacher-student hash code alignment loss is constructed to align the sample hash codes of the student model with the sample hash codes of the teacher model. The teacher-student hash code alignment loss is:

[0125]

[0126] To enable the image hash code and text hash code of the teacher model to maintain in-modal semantic similarity, an in-modal semantic contrast loss is constructed, which includes the in-modal semantic contrast loss of images and the in-modal semantic contrast loss of text. Among them, the in-modal semantic contrast loss of images is:

[0127]

[0128] where is the semantic weight, and is one of the items in

[0129] Samples with more shared labels should have a greater semantic weight. By minimizing the distances between semantically similar hash codes within the image modality can be made closer, while making semantically dissimilar image hash codes move away from each other.

[0130] The in-modal semantic contrast loss of text is

[0131]

[0132] By minimizing the distances between semantically similar hash codes within the text modality can be made closer, while making semantically dissimilar text hash codes move away from each other.

[0133] To maintain cross-modal semantic similarity, a cross-modal semantic contrast loss is also constructed and

[0134]

[0135] is the loss obtained by comparing all texts with the image, and

[0136] is the loss obtained by comparing all images with the text.

[0137]

[0138] Finally, by combining the in-modal and cross-modal semantic contrast losses, the semantic contrast loss

[0139]

[0140] where The predicted score indicating that the i-th image belongs to the j-th label, The predicted score indicating that the i-th text belongs to the j-th label.

[0141] By using the true label information to guide the training of the predicted labels of the samples, the classification loss can be obtained

[0142]

[0143] Among them, is the predicted label of the i-th image hash code, is the predicted label of the i-th text hash code.

[0144] To optimize the generation process of the student hash code, a quantization loss is introduced into the student model:

[0145]

[0146] Among them,

[0147] Minimizing this loss can reduce the information loss caused in the process of binarizing the hash code.

[0148] Currently, most methods only minimize the quantization loss to minimize the gap between different modality hash codes and their corresponding binary hash codes, while ignoring the relationship between the differences of different sample hash codes and the differences of the corresponding binary hash codes. To solve this problem, the present invention proposes a quantization difference preservation loss for maintaining the consistency between the differences of different sample hash codes and the differences of the corresponding binary hash codes.

[0149] Quantization difference preservation loss is:

[0150]

[0151] Among them, the first two terms are used to maintain the consistency between the differences of the student sample hash codes within the modality and the differences of the binary hash codes, and the third term is used to maintain the consistency between the differences of the student sample hash codes and the differences of the binary hash codes between modalities. q ij is used to make the samples with more shared labels have a greater weight in the quantization difference preservation loss.

[0152] The objective loss function of the teacher model is:

[0153]

[0154] Among them, α1 is used to measure the importance of the classification loss in the objective loss function

[0155] The objective loss function of the student model is:

[0156]

[0157] Among them, β1 and β2 are respectively used to measure the importance of the quantization difference preservation loss and the quantization loss in the objective loss function.

[0158] S6. Train the parameters corresponding to the teacher model according to the objective loss function corresponding to the teacher model; after the teacher model is trained, train the parameters corresponding to the student model according to the trained teacher model and the objective loss function corresponding to the student model.

[0159] The network parameters include: θ v 、θ t 、θ l 、 θ a 、θ h 、θ sv and θ st . The hyperparameters include σ, α1, β1 and β2.

[0160] Divide the dataset obtained in step S1 into a training set and a test set. The training set is used during training, and the test set is used to verify the performance after training is completed. Before training starts, initialize the network parameters and hyperparameters, and update the network parameters of the teacher model and the student model respectively through the loss function and using the BP algorithm.

[0161] Iteratively update the parameters of the teacher model according to the objective loss function of the teacher model. After the teacher model is trained, use the objective loss function of the student model and the trained teacher model to guide the training of the student model. After the student model is trained, obtain the trained student model.

[0162] S7. Use the trained student model to generate respective hash codes for the data in the database and the query data, and output the retrieval result according to the distance between the hash codes.

[0163] S7-1. Input the data in the database into CLIP and the trained student model in sequence to obtain the hash codes of the data in the database.

[0164] The data in the database includes a large amount of different-modal data (such as text and images), which are the objects to be retrieved. Input the data in the database into CLIP to obtain the initial features of the data in the database, and then input the initial features of the data in the database into the trained student model to obtain the hash code corresponding to each data in the database.

[0165] S7-2. Input the query data into CLIP and the trained student model to obtain the hash code of the query data.

[0166] When a query is needed, the query data is input into CLIP to obtain the initial features of the query data, and the initial features of the query data are input into the trained student model to obtain the hash code of the query data.

[0167] S7-3. Compare the hash code of the query data with the hash codes of the data in the database that have different modalities from the query data to obtain the distance between the hash codes, and output the data in the database within the preset distance as the retrieval result.

[0168] The query data is of one modality, and the required retrieval result should be data of another modality. That is, it is necessary to compare the hash code of the query data with the hash codes of the data in the database that have different modalities from the query data to obtain the distance between the hash codes, output the data within the preset distance as the retrieval result, or arrange the data in the database that have different modalities from the query data according to the size of the distance, and take the k data with the closest distance as the retrieval result.

[0169] S8. Method evaluation.

[0170] To verify the effectiveness of the cross-modal retrieval method of the present invention, the present invention conducts experiments on three commonly used cross-modal datasets, namely MIRFlickr-25K, MS COCO, and NUS-WIDE.

[0171] MIRFlickr-25K is a multi-label dataset with 24 categories, containing a total of 25,000 image-text pairs. After deleting the image-text pairs without labels, 20,015 image-text pairs are obtained. Randomly select 2,000 pairs from them as the query set, and the remaining 18,015 pairs as the retrieval set. In addition, randomly select 10,000 pairs from the retrieval set as the training set.

[0172] MS COCO is a multi-label dataset with 80 categories, including 123,287 image-text pairs. After deleting the image-text pairs without labels, 122,218 image-text pairs are obtained, and randomly select 5,000 pairs from them as the query set, and the remaining 117,218 pairs as the retrieval set. Then select 10,000 pairs from the retrieval set as the training set.

[0173] NUS-WIDE is a multi-label dataset with 81 categories, including 269,648 image-text pairs. Extract the subset belonging to 21 of the most common categories, totaling 190,421 image-text pairs. Randomly select 2,100 pairs from them as the query set, and the remaining 190,321 pairs as the retrieval set, and randomly select 10,000 pairs from the retrieval set as the training set.

[0174] The present invention is implemented under the PyTorch framework and uses the Adam optimizer to update network parameters. The learning rate on the three datasets is set to 0.00005, and the hyperparameters are set to α1 = 1, β1 = 0.001, and β2 = 0.01. The threshold parameter σ of the sample-label graph is set to 0.2. The batch-sizes for the MIRFlickr-25K, MS-COCO, and NUS-WIDE datasets are set to 128, 1024, and 128 respectively, and the number of epochs during training is set to 20, 100, and 30 respectively. The activation function of the information supplement module GCN is relu(·), and the activation functions of the fully connected layers are all tanh(·). The number of attention heads M of the information filtering module is set to 12. In the experiment, the length K of the hash code is set to 16 bits, 32 bits, and 64 bits respectively.

[0175] The present invention is compared with Deep cross-modal hashing (DCMH), Self-supervised adversarial hashing networks (SSAH), Deep joint-semantics reconstructing hashing (DJRSH), Adversary guided asymmetric hashing (AGAH), Deep cross-modal hashing with hashing functions and unified hashcodes jointly learning (DCHUC), Modality-invariant asymmetric networks (MIAN), Efficient hierarchical message aggregation hashing (HMAH). LSLCHR is the present invention. Cross-modal retrieval includes two retrieval tasks, namely image-to-text (I2T) and text-to-image (T2I). The present invention uses the mean Average Precision (mAP) to evaluate the retrieval performance of the method of the present invention and the comparative methods in these two retrieval tasks. mAP is a widely used evaluation metric for evaluating retrieval performance. For a query instance and the returned R retrieval results, the average precision (AP) can be expressed as:

[0176]

[0177] Where: n is the actual number of similar samples of the query instance in the retrieval set, and P(r) refers to the precision of the first r retrieval samples. When the r-th sample is similar to the query, δ(r) = 1; otherwise, δ(r) = 0. The present invention sets R as the size of the retrieval set, and finally, the mAP of all query instances in the query set is averaged to obtain the mAP.

[0178] Table 1: mAP Results of Different Models on Three Datasets

[0179]

[0180]

[0181] Table 1 shows the mAP results of all methods on three datasets, and the best results are shown in bold. According to the results in Table 1, it can be observed that:

[0182] Compared with other methods, the present invention can obtain the best retrieval performance in most cases. Specifically, the average mAP values of the present invention are 1.06%, 2.88%, and 1.89% higher than those of the sub-optimal methods in the three datasets of MIRFlickr-25K, MS COCO, and NUS-WIDE, respectively.

[0183] Among the comparison methods, DCMH, SSAH, AGAH, DCHUC, MIAN, HMAH, and SCH are all supervised methods. Among them, DCMH, AGAH, DCHUC, MIAN, and HMAH use discrete label information as the supervision information in the target loss function; SSAH uses continuous label features as the supervision information to optimize the generation of sample hash codes; SCH divides the sample relationships into three categories according to the labels: completely different labels, shared partial labels, and completely identical labels, and applies different triplet loss constraints to these three types of relationships. These methods all use label information in the network model to improve the semanticity of hash codes, but the utilization of label information is not sufficient. In contrast, the present invention takes into account the role of labels in information supplementation, information filtering, and the target loss function to fully enhance the semanticity of cross-modal samples. Therefore, the present invention can obtain the highest average mAP value in the three datasets.

[0184] It can be found that compared with other methods, the performance advantage of the present invention is the most obvious in the MS COCO dataset. This is mainly because the MS COCO data contains 81 category labels, far exceeding the number of categories in the MIRFlickr-25K and NUS-WIDE datasets. The present invention makes full use of the category labels as a guide in the process of information supplementation and information filtering. More labels mean that more abundant semantic information and semantic relevance can be incorporated, and thus, cross-modal hash codes of higher quality can be generated.

[0185] As the length of the hash code increases, the performance of most methods shows an upward trend. This phenomenon is mainly due to the fact that longer hash codes can contain richer semantic information, thereby improving the accuracy of retrieval to a certain extent. For example, the mAP values of the DCMH and DJRSH methods in the three datasets increase with the increase of the hash code length. However, there are also some special cases. For example, in the "I2T" task of the MIRFlickr-25K dataset, when the hash code length of AGAH is 16 bits, the mAP value is 0.7248, and when the hash code length is 64 bits, the mAP value drops to 0.7195. In the "T2I" task of the MS COCO dataset, when the hash code length of DCHUC is 32 bits, the mAP value is 0.5269, and when the hash code length is 64 bits, the mAP value drops to 0.5185. In the above cases, as the hash code length increases, the retrieval performance decreases instead. This may be because longer hash codes also contain more redundant information, which may impair the performance of the retrieval task in some cases.

[0186] Compared with HMAH, the present invention shows better performance in different retrieval tasks. Specifically, both HMAH and the present invention use GCN to fuse the features of cross-modal samples. Among them, HMAH constructs a graph containing cross-modal samples, and aggregates each other's features through GCN to reduce heterogeneous differences. Unlike HMAN, the graph constructed by the present invention contains not only cross-modal samples but also category labels. Since cross-modal samples aggregate the information of category labels, the semantics of cross-modal sample features can be enhanced. It can be found that the average mAP value of the present invention in the MIRFlickr-25K, MS COCO and NUS-WIDE datasets is 4.27%, 2.88% and 8.53% higher than that of HMAH respectively. The retrieval method of the present invention is more advantageous.

Claims

1. A cross-modal hashing retrieval method based on label-driven semantic perception learning, characterized in that, It includes the following steps: S1. Obtain a cross-modal retrieval dataset; the dataset includes image samples, text samples, and labels corresponding to each sample respectively; and the display content of an image sample is the expression content of a corresponding text sample; S2. Construct CLIP, and input the samples in the dataset into CLIP for initial feature extraction; S3. Construct a teacher model, and the teacher model includes: A feature extraction module for extracting features from the data output by CLIP to obtain an image representation and a text representation; GLoVe for performing initial feature extraction on the labels corresponding to the samples to obtain label representations; An information supplementation module for respectively supplementing information to the data output by the feature extraction module and the data output by GLoVe; An information filtering module for filtering the data output by the information supplementation module; and A feature hashing module for mapping the data output by the information filtering module into a teacher hash code; S4. Construct a student model, and the student model includes an image MLP and a text MLP, and both the image MLP and the text MLP include three fully connected layers; the student model is used to map the data output by CLIP into a student hash code; S5. Construct a target loss function for the teacher model according to the teacher hash code, and construct a target loss function for the student model according to the teacher hash code and the student hash code; S6. Train the parameters corresponding to the teacher model according to the target loss function corresponding to the teacher model; after the teacher model is trained, train the parameters corresponding to the student model according to the target loss function corresponding to the trained teacher model and the student model; S7. Use the trained student model to generate respective hash codes for the data in the database and the query data, and output a retrieval result according to the distance between the hash codes.

2. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, wherein The feature extraction module includes a text feature extraction module and an image feature extraction module, and both the text feature extraction module and the image feature extraction module include two fully connected layers.

3. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, characterized in that The information filtering module includes at least two attention modules, and each attention module includes a first fully connected layer, a second fully connected layer, a third fully connected layer, and a softmax function; The data processing method of the attention module is: The input is divided into three paths corresponding to three fully connected layers, and the outputs of the first fully connected layer and the second fully connected layer are calculated after passing through the softmax function and the output of the third fully connected layer to obtain the attention value of the input data.

4. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, characterized in that The feature hashing module includes an image hashing layer, a text hashing layer, and a label hashing layer, and the image hashing layer, the text hashing layer, and the label hashing layer all include a connection layer.

5. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, characterized in that The information supplementation module includes GCN; The specific method for the information supplementation module in step S3 to supplement information to the data output by the feature extraction module and GLoVe is: Sa3-1. Establish a sample-label graph according to the data output by the feature extraction module and GLoVe, and the vertices of the sample-label graph include an image representation, a text representation, and a label representation; Sa3-2. Calculate the adjacency matrix according to the data of the vertices of the sample-label graph; Input the adjacency matrix and the vertex feature matrix into the GCN to obtain the finally output features.

6. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 5, characterized in that, The specific method for calculating the adjacency matrix is as follows: No edges are set between data of the same modality in the sample-label graph, and edges are set between corresponding image and text pairs; determine the edges between the text representation and the label representation according to the similarity between the text representation and the label representation, and determine the edges between the image representation and the label representation according to the similarity between the image representation and the label representation; utilize the feature similarity between labels to determine the edges between label representations, and obtain the adjacency matrix A of the sample-label graph as: Among them, A *2* represents the edge of "*pointing to*", *∈{v,t,l}, A v2v , A t2t , A v2t and A t2v are all identity matrices 7. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, characterized in that The target loss function corresponding to the teacher model is: Among them, is the semantic contrast loss, is the classification loss, and α1 is a hyperparameter; The target loss function corresponding to the student model is: Among them, is the teacher-student hash code alignment loss, is the quantization difference preservation loss, is the quantization loss, and β1 and β2 are hyperparameters.

8. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 7, characterized in that, The teacher hash codes output by the teacher model include teacher image hash codes and teacher text hash codes; The calculation method of the semantic contrast loss is: Calculate the intra-modal semantic contrast loss of the image according to the teacher image hash code and the semantic weight; calculate the intra-modal semantic contrast loss of the image according to the teacher text hash code and the semantic weight; calculate the inter-modal semantic contrast loss according to the teacher image hash code, the teacher text hash code and the semantic weight; calculate the semantic contrast loss according to the intra-modal semantic contrast loss of the image, the intra-modal semantic contrast loss of the image and the inter-modal semantic contrast loss.

9. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 7, wherein The student hash codes output by the student model include student image hash codes and student text hash codes; the teacher hash codes output by the teacher model include teacher image hash codes and teacher text hash codes; The quantization difference retention loss is as follows: Among them, is the semantic weight, y i is the label of the i-th sample, y j is the label of the j-th sample, is the image hash code of the i-th student, is the image hash code of the j-th student, is the text hash code of the i-th student, is the text hash code of the j-th student, Among them, is the image hash code of the teacher, is the text hash code of the teacher.

10. The cross-modal hashing retrieval method based on label-driven semantic perception learning according to claim 1, characterized in that Step S7 includes the following steps: S7-1. Input the data in the database into CLIP and the trained student model in sequence to obtain the hash codes of the data in the database; S7-2. Input the query data into CLIP and the trained student model to obtain the hash code of the query data; S7-3. Compare the hash code of the query data with the hash codes of the data of different modalities from the query data in the database to obtain the distance between the hash codes, and output the data within the preset distance in the database as the retrieval result.

Citation Information

Cited By

  • Image-text cross-modal Hash retrieval model training method based on discriminative generation sampling

    CN120832432A

  • Text-image cross-modal hash retrieval model training method based on discriminative generative sampling

    CN120832432B