Cross-modal retrieval method based on comparative learning and balanced hash coding
Through comparative learning and balanced hash coding methods, the problems of modal heterogeneity and hash code imbalance in cross-modal retrieval are solved, and efficient and accurate cross-modal retrieval is achieved.
Patent Information
- Application Number
- CN202510410535.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
In the existing cross-modal search methods, the cross-modal heterogeneity problem makes it difficult to directly compare data of different modalities, and the hash code quality is unbalanced, which affects the search accuracy and efficiency.
The method based on contrast learning and balanced hash encoding is adopted, and the semantics of different modalities are aligned through contrast learning, and the hash code distribution is optimized by combining the quantized loss of the optimal transmission. The modal specific and shared hash codes are used for cross-modal searching, and the semantic index is used for coarse screening and hash code fine-grained matching.
The accuracy and efficiency of cross-modal retrieval is improved, and hash collision is reduced through semantic alignment and hash code equality optimization, and the retrieval accuracy and efficiency are improved.
Smart Images

Figure CN120296109A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal retrieval, and in particular to a cross-modal retrieval method based on contrast learning and balanced hash coding. Background Art
[0002] In cross-modal retrieval tasks, the computational complexity in the high-dimensional feature space is high, the storage cost is large, and it is affected by the "curse of dimensionality". Therefore, cross-modal hashing retrieval technology has received extensive attention. This method maps high-dimensional features to low-dimensional binary hash codes to achieve efficient retrieval. Compared with traditional methods, hashing retrieval performs similarity matching in the Hamming space, has high computational efficiency and low storage cost, and can effectively optimize the retrieval performance of multi-modal data. Existing cross-modal hashing methods are mainly divided into supervised and unsupervised methods. The supervised method uses label information to learn the hash function to improve the semantic consistency of the hash code. The unsupervised method does not require labels and learns the hash function through the features of the data itself.
[0003] However, there are still some problems in existing methods. First, the cross-modal heterogeneity problem makes data of different modalities have different feature representation methods, and this difference makes it difficult to directly compare data of different modalities, thus affecting the retrieval accuracy. Second, the quality problem of hash codes is also an important challenge. Traditional hash methods are prone to generating unbalanced hash codes, and some bit positions may contain redundant information, resulting in a decline in the discriminative ability of the hash representation, thus affecting the overall retrieval performance. Summary of the Invention
[0004] The purpose of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a cross-modal retrieval method based on contrast learning and balanced hash coding, which uses contrast learning to achieve semantic alignment of different modalities, and optimizes the distribution characteristics of hash codes through a quantization loss based on optimal transport to improve the retrieval performance.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: a cross-modal retrieval method based on contrast learning and balanced hash coding, including the following steps:
[0006] S1: Extract modality-specific features and modality-shared features from the input multi-modal data respectively. The modality-specific features are used to maintain the unique information of each modality, and the modality-shared features retain the information of all modalities. Align the modality-specific features of different modalities through contrast learning, and use the quantization loss based on optimal transport to optimize the modality-specific features and modality-shared features, so as to optimize the balance of the hash codes obtained according to the modality-specific features and modality-shared features;
[0007] S2: Perform a binarization operation on the modality-specific features and modality-shared features extracted in step S1 to generate a modality-specific hash code and a modality-shared hash code;
[0008] S3: Based on the modality-shared hash code extracted in step S2, perform semantic clustering through the K-means algorithm to generate a semantic index;
[0009] S4: Coarsely screen the candidate samples in the cross-modal retrieval stage through the semantic index obtained in step S3, and then perform fine-grained comparison through the modality-specific hash code and modality-shared hash code generated in step S2 to achieve efficient cross-modal retrieval.
[0010] Furthermore, the specific operation steps of step S1 are as follows:
[0011] S11: Extract modality-specific features and modality-shared features from multi-modal data; for a multi-modal data set containing images and texts, where the image modality and text modality are denoted as X1 and X2 respectively, and each modality contains N samples, denoted as X i′ , where X i′ represents the image modality or text modality, i′ = 1, 2, d i′ is the feature dimension of the image modality or text modality; for each modality, an independent feature extraction network is constructed respectively, where the ResNet18 network is used for the image modality to extract visual features, and a two-layer fully connected network is used for the text modality to process the topic distribution vector of the text to obtain the text representation; in the feature extraction network of each modality, the final feature output passes through two different fully connected layers to generate the modality-specific features and modality-shared features of each modality respectively; the modality-specific features are calculated by the first fully connected layer, representing the unique information of the modality:
[0012]
[0013] In the formula, represents the modality-specific feature of the image modality or text modality, F i′ (X i′ ) represents the feature extraction network of the image modality or text modality, is the fully connected layer of the image modality or text modality, is 's parameter; the modality-shared features are calculated by the second fully connected layer, representing the common information between different modalities:
[0014]
[0015] In the formula, represents the modality-shared feature of the image modality or text modality, It is a fully connected layer for generating shared features. is a parameter of
[0016] S12: Semantically align the modality-specific features extracted in step S11 through contrastive learning to address the cross-modal heterogeneous gap problem. Contrastive learning is achieved by constructing a specific contrastive loss function, and the expression of the contrastive loss function is:
[0017]
[0018] In the formula, is the contrastive loss, n is the number of samples, and respectively represent the specific features of the i-th sample in two modalities, represents the specific feature of the j-th sample in the text modality, sim(·) is the cosine similarity, and let the letter a represent let the letter b represent or is defined as τ is the temperature parameter, used to control the smoothness of positive and negative samples in contrastive learning, and exp(·) is the exponential function;
[0019] S13: To optimize the distribution characteristics of the generated hash codes, a quantization loss based on optimal transport is introduced. First, define the hash slice Wasserstein distance to measure the difference between the actual distribution and the target distribution:
[0020]
[0021] In the formula, m is the length of the hash code, D1 and D2 are two distributions, D 1l,: and D 2l,: are one-dimensional samples of D1 and D2 on dimension l, and w is the one-dimensional Wasserstein distance;
[0022] The quantization loss based on optimal transport has the expression:
[0023]
[0024] In the formula, is the quantization loss based on optimal transport, and let x represent or represents the sign function, U is the uniform discrete distribution. In the uniform discrete distribution, the value of each bit takes -1 or 1 with equal probability, and different bits are independent of each other. It is the hashed sliced Wasserstein distance; by minimizing the quantization loss based on optimal transport, the generated hash codes can maintain good balance; the balance of hash codes means that the probabilities of each bit of the hash code being 1 and -1 are the same, and each bit of the hash code should be as uncorrelated as possible, that is, different bits capture different features. The balance improves the information capacity of the hash code and prevents information redundancy;
[0025] S14: The final loss function is:
[0026]
[0027] In the formula, λ1 and λ2 are hyperparameters; by optimizing the above loss function to optimize the network parameters.
[0028] Furthermore, in step S2, the formula for calculating the modality-specific hash code is:
[0029]
[0030] In the formula, represents the modality-specific hash code of the image modality or the text modality;
[0031] The formula for calculating the modality-shared hash code is:
[0032]
[0033] In the formula, B c represents the modality-shared hash code.
[0034] Furthermore, in step S3, clustering is performed through the following optimization objective:
[0035] min H,C ||B c -CH|| 2
[0036] s.t.C∈{-1,1} m×k ,
[0037] H∈{0,1} k×n ,1 T H = 1 T
[0038] In the formula, ‖·‖ represents the 2-norm, m is the hash code length, C∈{-1,1} m×k represents the cluster center matrix, where k is the number of cluster categories, and each column represents a cluster center; H∈{0,1} k×n represents the semantic index matrix, and the elements in this semantic index matrix are represented by H hl denoted as, H hl= 1 indicates that the l-th sample belongs to the h-th cluster, otherwise it is 0; H ensures that each sample is assigned to a specific semantic category.
[0039] Furthermore, in step S4, by combining the modality-specific hash codes and modality-shared hash codes in step S2 and the semantic index in step S3, efficient cross-modal retrieval is achieved; a two-stage matching strategy is adopted, that is, first perform a rough screening based on the semantic index, and then calculate the fine-grained similarity through the hash codes. The specific operation steps are as follows:
[0040] S41: Perform a rough screening of candidate samples for the query sample according to the semantic index generated in step S3, that is, select those samples with the same semantic index as the query sample as the candidate sample set;
[0041] S42: In the candidate sample set, combine the modality-specific hash codes and modality-shared hash codes generated in step S2, and achieve fine-grained similarity matching by calculating the Hamming distance between the hash codes, so as to determine the final retrieval result.
[0042] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0043] 1. Cross-modal semantic alignment: Introduce a contrastive learning loss in the cross-modal retrieval task, align different modality data in the shared semantic space, and improve the retrieval accuracy.
[0044] 2. Optimize the quality of hash codes: Utilize the quantization loss based on optimal transport to improve the balance of hash codes, and then improve the retrieval performance.
[0045] 3. Dual-hash structure design: Combine modality-specific hash codes and modality-shared hash codes, which not only retain modality features but also enhance semantic consistency, reduce hash conflicts, and improve the retrieval accuracy.
[0046] 4. Efficient retrieval strategy: Use the semantic index to screen candidate samples, and then combine the hash code calculation to calculate the fine-grained similarity, improving the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow framework diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0049] As Figure 1 shown, this embodiment discloses a cross-modal retrieval method based on contrastive learning and balanced hash coding, and the specific situation is as follows:
[0050] S1: Extract modality-specific features and modality-shared features from the input multi-modal data respectively. The modality-specific features are used to preserve the unique information of each modality, and the modality-shared features retain the information of all modalities. By contrastive learning, align the modality-specific features of different modalities to solve the cross-modal heterogeneous gap problem and improve the retrieval accuracy. Optimize the modality-specific features and modality-shared features using the quantization loss based on optimal transport, so as to optimize the balance of the hash codes obtained from the modality-specific features and modality-shared features. The specific operation steps are as follows:
[0051] S11: Extract modality-specific features and modality-shared features from multi-modal data; for a multi-modal data set containing images and texts, where the image modality and text modality are denoted as X1 and X2 respectively, and each modality contains N samples, denoted as X i′ , where X i′ represents the image modality or the text modality, i′ = 1, 2, d i′ is the feature dimension of the image modality or the text modality; for each modality, construct an independent feature extraction network respectively. Among them, the ResNet18 network is used for the image modality to extract visual features, and a two-layer fully connected network is used for the text modality to process the topic distribution vector of the text to obtain the text representation; in the feature extraction network of each modality, the final feature output passes through two different fully connected layers, and the modality-specific features and modality-shared features of each modality are generated respectively; the modality-specific features are calculated by the first fully connected layer, representing the unique information of the modality:
[0052]
[0053] In the formula, represents the modality-specific feature of the image modality or the text modality, F i′ (X i′ ) represents the feature extraction network of the image modality or the text modality, is the fully connected layer of the image modality or the text modality, is parameters of; the modality-shared features are calculated by the second fully connected layer, representing the common information between different modalities:
[0054]
[0055] In the formula, represents the modality-shared feature of the image modality or the text modality, is the fully connected layer used to generate the shared features, is parameters of;
[0056] S12: Semantically align the modality-specific features extracted in step S11 through contrastive learning to address the cross-modal heterogeneous gap problem. Contrastive learning is achieved by constructing a specific contrastive loss function, and the expression of the contrastive loss function is as follows:
[0057]
[0058] In the formula, is the contrastive loss, n is the number of samples, and respectively represent the modality-specific features of the i-th sample in two modalities, represents the modality-specific feature of the j-th sample in the text modality, sim(·) is the cosine similarity, and is represented by the letter a is represented by the letter b or is defined as τ is the temperature parameter, used to control the smoothness of positive and negative samples in contrastive learning, and exp(·) is the exponential function;
[0059] S13: To optimize the distribution characteristics of the generated hash codes, a quantization loss based on optimal transport is introduced. First, define the hash slice Wasserstein distance to measure the difference between the actual distribution and the target distribution:
[0060]
[0061] In the formula, m is the length of the hash code, D1 and D2 are two distributions, D 1l,: and D 2l,: are one-dimensional samples of D1 and D2 in dimension l, and w is the one-dimensional Wasserstein distance;
[0062] The quantization loss based on optimal transport has the following expression:
[0063]
[0064] In the formula, is the quantization loss based on optimal transport, and is represented by x or represents the sign function, U is the uniform discrete distribution. In the uniform discrete distribution, the value of each bit takes -1 or 1 with equal probability, and different bits are independent of each other. is the hash slice Wasserstein distance; by minimizing the quantization loss based on optimal transport, the generated hash codes can maintain good balance; the balance of hash codes means that the probabilities of each bit of the hash code being 1 and -1 are the same, and each bit of the hash code should be as uncorrelated as possible, that is, different bits capture different features. Balance improves the information capacity of hash codes and prevents information redundancy;
[0065] S14: The final loss function is:
[0066]
[0067] where λ1 and λ2 are hyperparameters; by optimizing the above loss function to optimize the network parameters.
[0068] S2: For the modality-specific features and modality-shared features extracted in step S1, generate modality-specific hash codes and modality-shared hash codes through binarization operations;
[0069] The formula for calculating the modality-specific hash code is:
[0070]
[0071] where represents the modality-specific hash code of the image modality or the text modality;
[0072] The formula for calculating the modality-shared hash code is:
[0073]
[0074] where B c represents the modality-shared hash code.
[0075] S3: Based on the modality-shared hash codes extracted in step S2, perform semantic clustering through the K-means algorithm to generate semantic indexes;
[0076] Cluster through the following optimization objective:
[0077] min H,C ||B c -CH|| 2
[0078] s.t. C ∈ {-1, 1} m×k ,
[0079] H ∈ {0, 1} k×n , 1 T H = 1 T
[0080] where ||·|| represents the 2-norm, m is the hash code length, C ∈ {-1, 1} m×k represents the cluster center matrix, where k is the number of cluster categories, and each column represents a cluster center; H ∈ {0, 1} k×n represents the semantic index matrix, and the elements in this semantic index matrix are represented by H hl denoted by, H hl= 1 indicates that the l-th sample belongs to the h-th cluster, otherwise it is 0; H ensures that each sample is assigned to a specific semantic category.
[0081] S4: Combine the modality-specific hash codes and modality-shared hash codes in step S2 and the semantic index in step S3 to achieve efficient cross-modal retrieval; adopt a two-stage matching strategy, that is, first perform a rough screening based on the semantic index, and then calculate the fine-grained similarity through the hash codes. The specific operation steps are as follows:
[0082] S41: Perform a rough screening of candidate samples for the query sample according to the semantic index generated in step S3, that is, select those samples with the same semantic index as the query sample as the candidate sample set;
[0083] S42: In the candidate sample set, combine the modality-specific hash codes and modality-shared hash codes generated in step S2, and achieve fine-grained similarity matching by calculating the Hamming distance between the hash codes, so as to determine the final retrieval result.
[0084] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A cross-modal retrieval method based on contrastive learning and balanced hash coding, characterized in that Including the following steps: S1: Extract modality-specific features and modality-shared features from the input multi-modal data respectively. The modality-specific features are used to preserve the unique information of each modality, and the modality-shared features retain the information of all modalities. Align the modality-specific features of different modalities through contrastive learning, and optimize the modality-specific features and modality-shared features using the quantization loss based on optimal transport, so as to optimize the balance of the hash codes obtained according to the modality-specific features and modality-shared features; S2: For the modality-specific features and modality-shared features extracted in step S1, generate modality-specific hash codes and modality-shared hash codes through binarization operations; S3: Generate semantic indexes through semantic clustering of the modality-shared hash codes extracted in step S2 by the K-means algorithm; S4: Coarsely screen candidate samples in the cross-modal retrieval stage through the semantic indexes obtained in step S3, and then perform fine-grained comparison through the modality-specific hash codes and modality-shared hash codes generated in step S2 to achieve efficient cross-modal retrieval.
2. The cross-modal retrieval method based on contrastive learning and balanced hash coding according to claim 1, characterized in that The specific operation steps of step S1 are as follows: S11: Extract modality-specific features and modality-shared features from multimodal data; for a multimodal dataset containing images and texts, where the image modality and the text modality are denoted as X1 and X2 respectively, each modality contains N samples, denoted as X i′ , where X i′ represents the image modality or the text modality, i′ = 1, 2, d i′ is the feature dimension of the image modality or the text modality; for each modality, an independent feature extraction network is constructed respectively, where the ResNet18 network is used for the image modality to extract visual features, and a two-layer fully-connected network is used for the text modality to process the topic distribution vector of the text to obtain the text representation; in the feature extraction network of each modality, the final feature output passes through two different fully-connected layers to generate the modality-specific features and modality-shared features of each modality respectively; the modality-specific features are calculated by the first fully-connected layer, representing the unique information of that modality: wherein, represents modality-specific features of an image modality or a text modality, and F i′ (X i′ ) represents a feature extraction network for an image modality or a text modality, is a fully connected layer for an image modality or a text modality, is a parameter of; the modality-shared features are calculated by the second fully connected layer, representing the common information between different modalities: In the formula, represents the modality-shared feature of the image modality or the text modality, is a fully connected layer for generating the shared feature, is the parameter of; S12: Semantically align the modality-specific features extracted in step S11 through contrastive learning to solve the cross-modal heterogeneous gap problem. Contrastive learning is achieved by constructing a specific contrastive loss function. The expression of the contrastive loss function is: Wherein, is the contrastive loss, n is the number of samples, and respectively represent the specific features of the i-th sample in two modalities, represents the specific feature of the j-th sample in the text modality, sim(·) is the cosine similarity, represented by the letter a represented by the letter b or is defined as is the temperature parameter, used to control the smoothness of positive and negative samples in contrastive learning, exp(·) is the exponential function; S13: To optimize the distribution characteristics of the generated hash codes, a quantization loss based on optimal transport is introduced. First, the hash slice Wasserstein distance is defined to measure the difference between the actual distribution and the target distribution: where m is the length of the hash code, D1 and D2 are two distributions, D 1l,: and D 2l,: are the one-dimensional samples of D1 and D2 in dimension l, is the one-dimensional Wasserstein distance; The quantization loss based on optimal transport, the expression is: wherein, is the quantization loss based on optimal transport, where x represents or denotes the sign function, U is a uniform discrete distribution, where the value of each bit in the uniform discrete distribution takes -1 or 1 with equal probability, and different bits are independent of each other, is the hash slice Wasserstein distance; by minimizing the quantization loss based on optimal transport, the generated hash code can maintain good balance; the balance of the hash code means that the probabilities of each bit of the hash code being 1 and -1 are the same, and each bit of the hash code should be as uncorrelated as possible, that is, different bits capture different features, and the balance improves the information capacity of the hash code and prevents information redundancy; The final loss function is: where λ1 and λ2 are hyperparameters; by optimizing the above loss function to optimize the network parameters.
3. The cross-modal retrieval method based on contrastive learning and balanced hashing coding according to claim 2, wherein, In step S2, the formula for calculating the modality-specific hash code is: In the formula, represents a modality-specific hash code for an image modality or a text modality; The formula for calculating the modality-shared hash code is: In the formula, B c represents the modal shared hash code.
4. The cross-modal retrieval method based on contrastive learning and balanced hashing coding according to claim 3, characterized in that In step S3, cluster through the following optimization objective: Where, ‖·‖ represents the 2-norm, m is the length of the hash code, and C ∈ {-1, 1} m×k represents the cluster center matrix, where k is the number of clustering categories, and each column represents a cluster center; H ∈ {0, 1} k×n represents the semantic index matrix, and the elements in this semantic index matrix are represented by H hl represents, H hl = 1 indicates that the l-th sample belongs to the h-th cluster, otherwise it is 0; H ensures that each sample is assigned to a specific semantic category.
5. The cross-modal retrieval method based on contrastive learning and balanced hash coding according to claim 4, characterized in that, In step S4, combine the modality-specific hash code and modality-shared hash code in step S2 and the semantic index in step S3 to achieve efficient cross-modal retrieval; adopt a two-stage matching strategy, that is, first coarsely screen based on the semantic index, and then calculate the fine-grained similarity through the hash code. The specific operation steps are as follows: S41: Coarsely screen candidate samples for the query sample according to the semantic index generated in step S3, that is, select those samples with the same semantic index as the query sample as the candidate sample set; S42: In the candidate sample set, combine the modality-specific hash code and modality-shared hash code generated in step S2, and achieve fine-grained similarity matching by calculating the Hamming distance between the hash codes, so as to determine the final retrieval result.
Citation Information
Cited By
Multi-modal data approximate query method and system based on hybrid block chain
CN121560962A