Unsupervised cross-modal hash retrieval method based on multistage interaction

Through the unsupervised cross-modal hash retrieval method of multi-level interaction, the problems of insufficient information interaction and incomplete semantic information mining are solved, and the cross-modal retrieval performance is improved, especially in the nonlinear structural capture of multi-modal data.

CN120296192APending Publication Date: 2025-07-11KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510359164.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing unsupervised cross-modal hashing method has limitations in insufficient information interaction and incomplete semantic information mining, resulting in poor cross-modal retrieval performance, especially in the nonlinear structural capture of multimodal data.

Method used

Unsupervised cross-modal hash retrieval method based on multi-level interaction is adopted, and the cross-modal public hash learning module, a hash learning module for specific modals and a semantic similarity maintenance module is used to deeply interact from coarse to fine-grained, and hash codes for specific modals are learned, and information interaction is enhanced by self-attention mechanism and dual contrast learning.

Benefits of technology

It significantly improves cross-modal retrieval performance, improves the discriminantity of specific modal hash codes, and ensures that similar samples learn similar hash codes through instance similarity constraints, which improves the retrieval accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296192A_ABST
    Figure CN120296192A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised cross-modal Hash retrieval method based on multistage interaction, which belongs to the technical field of cross-modal Hash retrieval, and mainly comprises the following steps of: preprocessing cross-modal data; an unsupervised cross-modal Hash retrieval network based on multistage interaction is constructed; network training: inputting the image-text training samples into the constructed network in batches for network training; and modal retrieval: inputting the image-text query sample set and the retrieval sample set into the trained multi-level interaction-based unsupervised cross-modal Hash retrieval network, respectively generating corresponding Hash codes, and obtaining a query result by calculating the Hamming distance between the Hash codes of the query sample and the retrieval sample. The minimum Hamming distance is the final query result. According to the method, high-dimensional multi-modal features can be compressed into compact binary codes, the cross-modal retrieval efficiency is remarkably improved, and the method can be used for real-time image search and cross-modal recommendation systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-modal hashing retrieval method related to image texts, and belongs to the technical field of cross-modal hashing retrieval. Background Art

[0002] With the rapid development of digital media and information society, multi-modal data has grown exponentially, covering various forms such as texts, videos, and images. Cross-modal retrieval plays a crucial role in studying the correlations between various multi-modal data. In recent years, cross-modal hashing technology has attracted wide attention in the field of cross-modal retrieval due to its high retrieval speed and low storage cost.

[0003] Cross-modal hashing retrieval methods can be roughly divided into two categories: supervised and unsupervised according to whether real labels are used. Among them, supervised hashing methods achieve relatively ideal retrieval performance by using label information. However, manual annotation of semantic labels is not only time-consuming but also costly. Therefore, unsupervised hashing learning methods have attracted wide attention and provided effective solutions to the label annotation problem. In the field of unsupervised cross-modal hashing retrieval, deeply mining semantic information plays a key role in learning hash codes. Some shallow unsupervised techniques map the features of different modalities into a common latent space through matrix factorization to construct the correlation between heterogeneous data. However, such shallow methods have obvious limitations in capturing the non-linear structure of multi-modal data, resulting in poor performance. Therefore, deep hashing methods have attracted the attention of researchers due to their strong feature representation ability. Many studies have attempted to construct a multi-modal similarity matrix to explore the semantic correlation between different modalities. These methods mainly rely on the features of the original data to calculate similarity. However, compared with instance-level similarity, feature similarity fails to fully mine the internal relationship between different modalities. To address the above limitations, some studies have attempted to use the correlation of instance neighborhoods to construct a similarity matrix between data instances. Specifically, to construct a high-quality similarity matrix, Yu et al. used graph-neighbor coherence (GC) to reveal the potential relationship between an instance and its neighborhood. Zhang et al. proposed a high-order similarity measurement method from a non-local perspective. In addition, graph convolutional networks (GCNs) are used to aggregate the semantic information within and between modalities to improve cross-modal retrieval performance. However, most existing methods are limited to single-level information interaction (such as inter-modal correlation or instance-level correlation), resulting in problems such as insufficient information interaction and incomplete mining of semantic information.

[0004] Research shows that information interaction can effectively bridge the semantic gap between different modalities. In cross-modal retrieval tasks, due to the heterogeneous feature representations of different modalities such as images and texts, direct similarity calculation often has difficulty capturing the underlying semantic alignment. Information interaction, by realizing the sharing and fusion of cross-modal semantic information, generates more consistent representations in the shared feature space, improves the accuracy of cross-modal matching, and thus enhances the retrieval performance.

[0005] Based on this background, in order to achieve fast and accurate cross-modal retrieval, the present invention proposes a simple and effective unsupervised cross-modal hashing retrieval method. Summary of the Invention

[0006] The present invention proposes an unsupervised cross-modal hashing retrieval method based on multi-level interaction. Through multi-level interaction from coarse-grained to fine-grained, deep interaction is carried out between common information and specific information, between modalities and between instances, so as to learn the hash codes of specific modalities and effectively improve the cross-modal retrieval performance.

[0007] The technical solution of the present invention is: an unsupervised cross-modal hashing retrieval method based on multi-level interaction. The specific steps of the unsupervised cross-modal hashing retrieval method based on multi-level interaction are as follows:

[0008] Step1: Prepare several general image-text datasets for network training;

[0009] Step2: Preprocess the image-text. Crop and adjust the resolution of the images and standardize them. Tokenize the texts, construct a vocabulary to support vector representation of the texts, and divide the image-text pairs into non-overlapping training sample sets, query sample sets, and retrieval sample sets;

[0010] Step3: Construct an unsupervised cross-modal hashing retrieval network based on multi-level interaction. The entire network includes a cross-modal common hashing learning module, a specific modality hashing learning module, and a semantic similarity preservation module;

[0011] The unsupervised cross-modal hashing retrieval network based on multi-level interaction first uses the cross-modal common hashing learning module to learn common hash codes and maintain modality invariance; then, in the specific modality hashing learning module, through the guidance of the common hash codes and the combination of dual contrast learning, the discriminability of the specific modality hash codes is enhanced, and thus the interaction between common information and specific information and between modalities is realized at the coarse-grained level;

[0012] In the fine-grained interaction stage, through the semantic similarity preservation module, the learning of the common hash codes and the specific modality hash codes is constrained by the instance similarity to promote the information interaction between instances, so as to ensure that similar samples learn similar hash codes;

[0013] Step 4: Train the unsupervised cross-modal hashing retrieval network with multi-level interaction using the image-text training sample set, and use the image-text query sample set to verify the trained network after each batch of training is completed to check the status and convergence of the current method;

[0014] Step 5: Input the image-text query sample set and retrieval sample set into the trained unsupervised cross-modal hashing retrieval network based on multi-level interaction, generate the corresponding hash codes respectively, obtain the query results by calculating the Hamming distance between the hash codes of the query samples and retrieval samples, and the one with the smallest Hamming distance is the final query result.

[0015] Furthermore, the specific steps of Step 3 are as follows:

[0016] Step 3.1: Build a cross-modal common hashing learning module;

[0017] The cross-modal common hashing learning module first extracts the feature representations of images and texts respectively using specific-modal feature extractors; secondly, promotes the interaction of intra-modal information by introducing the self-attention mechanism to capture the potential semantic associations within the modality; finally, to further enhance the semantic consistency between modalities, a reconstruction loss function is introduced to maintain modality invariance, thereby effectively learning the common hash codes between different modalities;

[0018] The specific-modal feature extractor consists of an image feature extraction network based on the ResNet-18 architecture and a text feature extraction network based on the BoW model;

[0019] Step 3.2: Build a specific-modal hashing learning module;

[0020] The specific-modal hashing learning module includes a cross-modal preliminary interaction module and dual contrast learning; specifically, first uses the self-attention mechanism for preliminary modality interaction, secondly, strengthens the interaction through dual contrast learning; finally, optimizes the specific-modal hash codes under the guidance of the common hash codes to achieve the interaction of common information and specific information at the coarse-grained level;

[0021] Step 3.3: Build a semantic similarity preservation module;

[0022] The semantic similarity preservation module first constructs a specific-modal similarity matrix, fuses the multi-modal similarity modeling instance relationships, and uses this to guide the learning of the common hash codes and specific-modal hash codes, thereby enhancing the interaction between instances at the fine-grained level.

[0023] Furthermore, Step 4 includes Step 4.1: For the part of optimizing the common hash code:

[0024] The specific steps of Step 4.1 are as follows:

[0025] Step 4.1.1. Calculate the error loss using the reconstruction error loss L r Keeping the modal invariance, the loss function is:

[0026]

[0027] where ||·|| F is the Frobenius norm of the matrix, F I ′ and F T ′ represents the reconstructed features of image and text modalities, respectively, and F I and F T They are the feature representations of images and texts respectively;

[0028] Step 4.1.2. Using discrete loss L dis Let's learn the public hash code B:

[0029]

[0030] Where, sign(·) is the signum function;

[0031] Step 4.1.3, at the fine-grained interaction level, use the adjacency-related enhancement loss to keep the instance similarity in the public hash code B as much as possible, expressed as:

[0032]

[0033] Among them, cos(·,·) represents the cosine similarity of two vectors, S is the instance similarity matrix, and They are used to maintain instance similarity in images, texts and joint representations respectively.

[0034] The cross-modal common hash learning module is updated using a stochastic gradient descent algorithm, and the hash code reconstruction error loss, discrete loss, and adjacency-related enhancement loss are integrated into the same framework to learn the hash code calculation loss value, which can be expressed as:

[0035] L cmch =L r +L ace +L dis .

[0036] Furthermore, the Step 4 includes Step 4.2, optimizing the hash code part of the specific mode, and the specific steps of Step 4.2 are as follows:

[0037] Step4.2.1. To further enhance modal interaction and improve the representation ability, the specific-modal hashing learning module introduces dual contrastive learning. The first contrastive learning focuses on the representation level:

[0038]

[0039] where and represent the i-th elements of Z I and Z T respectively, τ represents the temperature coefficient, and <·,·> represents the inner product; it should be noted that and are truly aligned sample pairs; therefore, the first contrastive loss is mathematically expressed as follows:

[0040]

[0041] The second contrastive learning focuses on the specific-modal hash code level, and its objective function is expressed as:

[0042]

[0043] where and represent the i-th elements of the specific-modal continuous hash codes and respectively, and are mutually aligned, and these specific-modal continuous hash codes and are generated by a hash layer based on Z I and Z T as inputs, where the hash layer is composed of a multi-layer perceptron MLP. The second contrastive loss is shown as follows:

[0044]

[0045] Step4.2.2. To achieve coarse-grained interaction between common information and specific information, a common hash code B is introduced to guide the learning process of the specific-modal continuous hash codes and . Its loss function is defined as follows:

[0046]

[0047] Step4.2.3. At the fine-grained interaction level, to maintain the instance similarity between the specific-modal continuous hash codes and , the following structural consistency loss is proposed:

[0048]

[0049] From another perspective, the loss function maintains the numerical consistency between different-modal hash codes and the common-modal hash code, while the structural consistency loss maintains the structural consistency within and between modalities of the specific-modal hash codes;

[0050] Therefore, the final loss is as follows:

[0051]

[0052] Update the hash code learning module of the specific modality using the stochastic gradient descent algorithm. By integrating the public double contrast loss and the structural consistency loss into a unified learning framework, therefore, the total objective function of the hash learning module of the specific modality is as follows:

[0053] L msh = L ch + λL cf + L f

[0054] where λ is used to measure the importance of each loss function.

[0055] The present invention also provides an unsupervised cross-modal hashing retrieval system based on multi-level interaction, including: a module for executing the unsupervised cross-modal hashing retrieval method based on multi-level interaction described above.

[0056] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the unsupervised cross-modal hashing retrieval method based on multi-level interaction is implemented.

[0057] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the unsupervised cross-modal hashing retrieval method based on multi-level interaction is implemented.

[0058] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the unsupervised cross-modal hashing retrieval method based on multi-level interaction is implemented.

[0059] The beneficial effects of the present invention are as follows:

[0060] 1. The present invention proposes a new unsupervised cross-modal hashing framework. Through multi-level interactions from coarse-grained to fine-grained, deep interactions are carried out between common information and specific information, between modalities and between instances, thereby learning the hash codes of specific modalities and effectively improving the cross-modal retrieval performance.

[0061] 2. The present invention proposes an inter-modal interaction strategy. First, the self-attention mechanism is used for cross-modal preliminary interaction, and then double contrast learning is carried out to achieve deeper information interaction. This progressive interaction significantly improves the discriminability of the specific modal hash codes.

[0062] 3. In the fine-grained interaction stage, the method proposed by the present invention uses instance similarity to constrain the learning of the common hash code and the specific modal hash code to promote information interaction between instances, so as to ensure that similar samples learn similar hash codes. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is the flow chart of this method;

[0064] Figure 2 is the comparison experimental graph of the top-N curves of the method of the present invention and other cross-modal hashing retrieval methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The embodiments and effects of the present invention will be further described below with reference to the accompanying drawings.

[0066] Refer to Figure 1 , the implementation steps of this example include the following.

[0067] Step1. Prepare several general image-text data sets for network training;

[0068] Step2. Preprocess the image-text. Crop and adjust the resolution of the image and standardize it. Tokenize the text, construct a vocabulary to support the vector representation of the text, and divide the image-text pairs into non-overlapping training sample sets, query sample sets, and retrieval sample sets;

[0069] Step3. Construct an unsupervised cross-modal hashing retrieval network based on multi-level interaction. The entire network consists of a cross-modal common hash learning module, a specific modal hash learning module, and a semantic similarity preservation module; the unsupervised cross-modal hashing retrieval network based on multi-level interaction first uses the cross-modal common hash learning module to learn the common hash code and maintain modal invariance. Then, in the specific modal hash learning module, through the guidance of the common hash code and the combination of double contrast learning, the discriminability of the specific modal hash code is enhanced, and thus the interaction between common information, specific information, and modalities is realized at the coarse-grained level. In the fine-grained interaction stage, through the semantic similarity preservation module, the learning of the common hash code and the specific modal hash code is constrained by instance similarity to promote information interaction between instances, so as to ensure that similar samples learn similar hash codes.

[0070] Step 4: Train the unsupervised cross-modal hashing retrieval network with multi-level interactions using the image-text training sample set, and after each batch of training is completed, use the image-text query sample set to verify the trained network to check the status and convergence of the current method;

[0071] Step 5: Input the image-text query sample set and retrieval sample set into the trained unsupervised cross-modal hashing retrieval network based on multi-level interactions, generate the corresponding hash codes respectively, obtain the query results by calculating the Hamming distance between the hash codes of the query samples and retrieval samples, and the one with the smallest Hamming distance is the final query result.

[0072] Furthermore, the specific steps of the said Step 3 are as follows:

[0073] Step 3.1: Build a cross-modal common hashing learning module;

[0074] This module consists of two feature extraction networks, two encoders, two decoders and three self-attention modules: First, adopt the ResNet-18 network architecture to extract the deep features of the image modality, and at the same time use the BoW model to extract the semantic features of the text modality. Secondly, introduce the self-attention mechanism to promote the interaction of information within the modality to capture the potential semantic associations within the modality. Finally, to further enhance the semantic consistency between modalities, introduce a reconstruction loss function to maintain modality invariance, so as to effectively learn the common hash codes between different modalities.

[0075] Furthermore, the specific steps of the said Step 3.1 are as follows:

[0076] Step 3.1.1: Learn the deep representation while maintaining modality invariance:

[0077] To learn the deep representation while maintaining modality invariance, first use modality-specific feature extractors to extract the semantic features of images and texts respectively. These features are denoted as and where d I and d T represent the dimensions of the image and text features respectively. Then, learn the shallow semantic representations of images and texts through two independent encoders, namely and where d I ′ and d′ T represent the dimensions of J I and J T respectively, and its mathematical expression can be represented as:

[0078] J I =Encoder I (F I )

[0079] J T = Encoder T (F T )

[0080] In addition, to capture deeper semantic relationships between instances, first, the shallow representation J I and J T are respectively input into their own self-attention modules, and then through the modality-specific fully connected layers to obtain their deep latent semantic representations H I and H T . Therefore, the mathematical formulation of this process is as follows:

[0081] RS2(·)= Norm(Drop(FFN(·))+(·))

[0082] (q * ) i = J * W * Q ,(k * ) j = J * W * K ,(v * ) j = J * W * V

[0083]

[0084] H * = FNN((J * ) i )

[0085] where * ∈ {I, T} represents the image or text modality. Norm(·) and Drop(·) represent layer normalization and Dropout function respectively. FFN(@) is a feed-forward network. (v * ) j and (k * ) j are respectively the j-th elements of the value matrix and the key matrix. (q * ) i is the i-th element of the query matrix. and are trainable parameters. d k* and d v* are respectively the dimensions of the key matrix and the value matrix. d * ' is J *The dimension. FNN(·) represents the fully connected layer of a specific modality, which plays the role of reducing the representation dimension. (J * ′) i is the i-th element of the high-level representation of a specific modality, while is the i-th element of the intermediate-level representation of a specific modality. (·) represents a placeholder, which is the input of a function.

[0086] Finally, the deep representations H I and H T are input into the decoder. The mathematical expression is as follows:

[0087] F I ′ = Decoder I (H I )

[0088] F T ′ = Decoder T (H T )

[0089] The proposed UCHRMI method adopts the following reconstruction error loss to maintain modality invariance:

[0090]

[0091] where ||·|| F is the Frobenius norm of the matrix. F I ′ and F T ′ represent the reconstructed features of the image and text modalities respectively.

[0092] Step3.1.2, Learning of common hash codes;

[0093] To learn the common hash code B, the UCHRMI method first concatenates J I and J T to generate the common shallow representation J S . Then, J S is input into the self-attention layer, and the deep latent joint representation H S is generated through the fully connected layer. The mathematical description of this process is as follows:

[0094] q i = J S W Q , k j = J S W K , v j = J S W V

[0095]

[0096] RS2(·) = Norm(Drop(FFN(·))+(·))

[0097]

[0098] H S = FNN((J S ′) i )

[0099] where d k and d v are the dimensions of the key matrix and the value matrix respectively, and d s = d I ′ + d′ T represents the dimension of J S . and represent training parameters. (J S ′) i is the i-th element of the high-level common representation, and

[0100] is the i-th element of the middle-level common representation.

[0101]

[0102] where sign(·) is the signum function.

[0103] At the fine-grained interaction level, in order to maintain the instance similarity in the common hash code B as much as possible, this cross-modal retrieval network proposes an adjacency correlation enhancement loss, denoted as:

[0104]

[0105] where cos(·,·) represents the cosine similarity between two vectors, and S is the instance similarity matrix.

[0106] and are used to maintain the instance similarity in the image, text, and joint representation respectively.

[0107] Step3.2. Build a modality-specific hash learning module;

[0108] This module first uses the self-attention mechanism for preliminary cross-modal interaction. Then, it uses dual contrast learning for deeper cross-modal interaction. Finally, it guides the learning of modality-specific hash codes through the common hash code, realizing the interaction between common information and specific information at the coarse-grained level.

[0109] Furthermore, first concatenate the features and to generate the multi-modal feature where d = d I + d T . The multi-modal fusion Transformer encoder takes the multi-modal feature F as the input and uses the self-attention mechanism to achieve cross-modal interaction. The output of the Transformer encoder is denoted as Z = [Z I , Z T .

[0110] To further enhance cross-modal interaction and improve the representation ability, dual contrast learning is introduced in this model. The first contrast learning focuses on the representation level:

[0111]

[0112] where and represent the i-th elements of Z I and Z T respectively. τ represents the temperature coefficient, and <·,·> represents the inner product. It should be noted that and are truly aligned sample pairs. Therefore, the first contrast loss is mathematically expressed as follows:

[0113]

[0114] The second contrast learning focuses on the modality-specific hash code level, and its objective function can be expressed as:

[0115]

[0116] where and represent the i-th elements of and respectively. and are aligned with each other respectively. These modality-specific continuous hash codes and are generated through a hash layer based on Z I and Z T as the input, where the hash layer is composed of a multi-layer perceptron MLP. The second contrast loss is shown as follows:

[0117]

[0118] To achieve the coarse-grained interaction between public information and specific information, the learning process of the continuous hash code of a specific modality is guided by utilizing the public hash code B, and the loss function is defined as follows:

[0119]

[0120] At the fine-grained interaction level, the structural consistent loss is used to maintain the instance similarity between the continuous hash codes of specific modalities and :

[0121]

[0122] From another perspective, maintains the numerical consistency between the hash codes of different modalities and the public modality hash code, while maintains the structural consistency within and between the specific modality hash codes.

[0123] Therefore, the final loss is as follows:

[0124]

[0125] Step 3.3. Build a semantic similarity preservation module;

[0126] The semantic similarity preservation module integrates the similarities of multiple modalities to model instance similarity, and uses this similarity to constrain the learning of the public hash code and the specific modality hash code, realizing the fine-grained interaction between instances.

[0127] Furthermore, the module first constructs an instance similarity matrix with m pairs of instance samples. Specifically, the initial similarity matrix S v and S t is constructed by weighted summation, and its formula is as follows: r :

[0128] S r =γ1S I +(1 - γ1)S T

[0129] S I =cos(F I ,F I )

[0130] ST = cos(F T , F T )

[0131] where S I and S T respectively represent the cosine similarity between the image and text modality instances. γ1 weighs the importance of these two matrices.

[0132] It should be noted that the cross-modal retrieval task mainly focuses on the instances with the most significant and top-ranked semantic relationships, while instance pairs with smaller similarity values are more likely to introduce noise, which has an adverse impact on hash code learning. To reduce the interference of such noise, this module assigns the k smallest element values in each row of the similarity matrix S r to -1, thus generating a denoised similarity matrix

[0133]

[0134] where, represents the element in the i-th row and j-th column of, (S r ) ij represents the similarity between the i-th sample and the j-th sample, and e i (k) represents the set composed of the first k minimum values in the i-th row of S r .

[0135] In addition, by introducing a non-linear operation to further enhance the adjacency correlation between more similar sample pairs, the enhanced similarity matrix can be expressed as follows:

[0136]

[0137] where represents the identity matrix, which is used to enhance the similarity of the instance itself while expanding the similarity distance from other instances. is a matrix of all 1s. Sigmoid(·) is used as an activation function to compress the input value to a range close to 0 or 1, thus significantly expanding the similarity distance between similar and dissimilar instances, making it more prominent.

[0138] To more effectively mine the adjacency semantic relationship between instances, the denoised similarity matrix and the enhanced similarity matrix are combined to generate the final multi-modal joint semantic similarity matrix, which is mathematically expressed as follows:

[0139]

[0140] where γ2 is a weighing parameter used to balance the denoised similarity matrix and the enhanced similarity matrix Se The importance between.

[0141] Among them, the hyperparameter γ2 is used to adjust the denoising similarity matrix With the enhanced similarity matrix S e The weight distribution between them is to balance the contribution of the two.

[0142] Furthermore, the specific steps of Step 4 are as follows:

[0143] This image-text retrieval method includes two parts: a public hash code and a hash code of a specific mode.

[0144] Step 4.1: Optimize the public hash code part:

[0145] Step 4.1.1. Calculate the error loss using the reconstruction error loss L r Keeping the modal invariance, the loss function is:

[0146]

[0147] in‖·‖ F is the Frobenius norm of the matrix, F I ′ and F T ′ represents the reconstructed features of image and text modalities, respectively, and F I and F T They are the feature representations of images and texts respectively;

[0148] Step 4.1.2. Using discrete loss L dis Let's learn the public hash code B:

[0149]

[0150] Where, sign(·) is the signum function;

[0151] Step 4.1.3, at the fine-grained interaction level, use the adjacency-related enhancement loss to keep the instance similarity in the public hash code B as much as possible, expressed as:

[0152]

[0153] Among them, cos(·,·) represents the cosine similarity of two vectors, S is the instance similarity matrix, and They are used to maintain instance similarity in images, texts and joint representations respectively.

[0154] Update the common hash code learning module using the stochastic gradient descent algorithm. By integrating the reconstruction error loss, discrete loss, and adjacent correlation enhancement loss into a unified learning framework, the total objective function of the cross-modal common hash learning module is expressed as follows:

[0155] L cmch = L r + L ace + L dis

[0156] Step4.2: Optimize the hash code part of a specific modality:

[0157] Step4.2.1. To further enhance modality interaction and improve the representation ability, the hash learning module of a specific modality introduces dual contrast learning. The first contrast learning focuses on the representation level:

[0158]

[0159] where and represent the i-th elements of Z I and Z T respectively, τ denotes the temperature coefficient, and <·,·> denotes the inner product; it should be noted that and are truly aligned sample pairs; therefore, the first contrast loss is mathematically expressed as follows:

[0160]

[0161] The second contrast learning focuses on the hash code level of a specific modality, and its objective function is expressed as:

[0162]

[0163] where and represent the i-th elements of the continuous hash codes and of a specific modality respectively, and are aligned with each other. These continuous hash codes and of a specific modality are generated by a hash layer based on Z I and Z T as inputs, where the hash layer is composed of a multi-layer perceptron MLP. The second contrast loss is shown as follows:

[0164]

[0165] Step4.2.2. To achieve coarse-grained interaction between public information and specific information, a public hash code B is introduced to guide the learning process of the continuous hash code of a specific modality. The loss function is defined as follows: and The learning process of is as follows:

[0166]

[0167] Step4.2.3. At the fine-grained interaction level, to maintain the instance similarity between the continuous hash codes of a specific modality and the following structural consistency loss is proposed:

[0168]

[0169] From another perspective, the loss function maintains the numerical consistency between the hash codes of different modalities and the public modality hash code, while the structural consistency loss maintains the structural consistency within and between the specific modality hash codes;

[0170] Therefore, the final loss is as follows:

[0171]

[0172] Use the stochastic gradient descent algorithm to update the hash code learning module of the specific modality. By integrating the public dual contrast loss and the structural consistency loss into a unified learning framework. Therefore, the total objective function of the hash learning module of the specific modality is as follows:

[0173] L msh = L ch + λL cf ++ L f

[0174] where λ is used to measure the importance of each loss function;

[0175] The present invention also provides an unsupervised cross-modal hashing retrieval system based on multi-level interaction, including: a module for executing the unsupervised cross-modal hashing retrieval method based on multi-level interaction.

[0176] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the unsupervised cross-modal hashing retrieval method based on multi-level interaction.

[0177] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the unsupervised cross-modal hashing retrieval method based on multi-level interaction is implemented.

[0178] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the unsupervised cross-modal hashing retrieval method based on multi-level interaction is implemented.

[0179] Further, in this embodiment, the hyperparameter λ of the network is set to 0.5, and the remaining hyperparameters are set as follows: for the MIRFlickr dataset, γ1 = 0.6, γ2 = 0.6 and k = 3000; the Adam optimizer is used for model training, and the batch size is set to 128. The batch size is 64, and 150 rounds of iterations are trained.

[0180] The hardware platform for the simulation experiment in this embodiment is: an NVIDIA GeForce RTX3090 GPU with a model of 24GB of memory; the Ubuntu 20.04.5 LTS operating system is adopted, and the configured virtual environment includes: Python3.9.16, Pytorch1.13.1, cuda11.6, etc.

[0181] We send the cross-modal query sample set into the above-mentioned trained unsupervised cross-modal retrieval (UCHRMI) network based on multi-level interaction to calculate the general evaluation index: the mean average precision (mAP). The larger this index is, the better the classification effect is.

[0182] To evaluate the effectiveness of this method, we retrieve the cross-modal dataset MIR Flickr using 11 advanced unsupervised cross-modal hashing methods: CVH, CMFH, STMH, JIMFH, DRMFH, DJSRH, JDSH, DGCPN, CIRH, UCCH, UCMFH.

[0183] The CVH method mentioned above refers to the unsupervised cross-modal hashing retrieval method proposed by Kumar S et al. in "Learning hash functions for cross-view similarity search", abbreviated as CVH;

[0184] The CMFH method mentioned above refers to the unsupervised cross-modal hashing retrieval method proposed by Ding G et al. in "Collective matrix factorization hashing for multimodal data", abbreviated as CMFH;

[0185] The so-called STMH method refers to the unsupervised cross-modal hashing retrieval method proposed by Di Wang X G et al. in "Semantic topic multimodal hashing for cross-media retrieval", abbreviated as STMH;

[0186] The so-called JIMFH method refers to the unsupervised cross-modal hashing retrieval method proposed by Wang D et al. in "Joint and individual matrix factorization hashing for large-scale cross-modal retrieval", abbreviated as JIMFH;

[0187] The so-called DRMFH method refers to the unsupervised cross-modal hashing retrieval method proposed by Yao T et al. in "Discrete robust matrix factorization hashing for large-scale cross-media retrieval", abbreviated as DRMFH;

[0188] The so-called DJSRH method refers to the unsupervised cross-modal hashing retrieval method proposed by Su S et al. in "Deep joint-semantics reconstructing hashing for large-scale unsupervised cross-modal retrieval", abbreviated as DJSRH;

[0189] The so-called JDSH method refers to the unsupervised cross-modal hashing retrieval method proposed by Liu S et al. in "Joint-modal distribution-based similarity hashing for large-scale unsupervised deep cross-modal retrieval", abbreviated as JDSH;

[0190] The so-called DGCPN method refers to the unsupervised cross-modal hashing retrieval method proposed by Yu J et al. in "Deep graph-neighbor coherence preserving network for unsupervised cross-modal hashing", abbreviated as DGCPN;

[0191] The CIRH method mentioned above refers to the unsupervised cross-modal hashing retrieval method proposed by Zhu L et al. in "Work together: Correlation-identity reconstruction hashing for unsupervised cross-modal retrieval", abbreviated as CIRH;

[0192] The UCCH method mentioned above refers to the unsupervised cross-modal hashing retrieval method proposed by Hu P et al. in "Unsupervised contrastive cross-modal hashing", abbreviated as UCCH;

[0193] The UCMFH method mentioned above refers to the unsupervised cross-modal hashing retrieval method proposed by Xia X et al. in "When CLIP meets cross-modal hashing retrieval: A new strong baseline", abbreviated as UCMFH;

[0194] Table 1 mAP values of all methods on the MIR Flickr dataset for the image-to-text (I2T) task

[0195]

[0196] Table 2 mAP values of all methods on the MIR Flickr dataset for the image-to-text (I2T) task

[0197]

[0198] It can be seen from Table 1 and Table 2 that the mAP values of the UCHRMI method proposed in the present invention in the two query tasks in the unsupervised cross-modal retrieval scenario of the MIR Flickr dataset are higher than those of other comparison methods. Further proving that the UCHRMI method proposed in the present invention is more powerful in unsupervised cross-modal retrieval.

[0199] Figure 2 This is a comparison experimental graph of the top-N curves of the method of the present invention and other unsupervised cross-modal hashing retrieval methods with different hash code lengths;

[0200] It can be clearly seen from the figure that the present method surpasses all comparison methods in retrieval accuracy by retrieving more correct instances, and still demonstrates a powerful retrieval ability compared to other methods.

[0201] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. An unsupervised cross-modal hashing retrieval method based on multi-level interaction, characterized in that: It includes the following steps: Step 1. Prepare several general image-text datasets for network training; Step 2. Preprocess the image-text. Crop the image, adjust the resolution and standardize it. Tokenize the text, construct a vocabulary to support the vector representation of the text, and divide the image-text pairs into non-overlapping training sample sets, query sample sets, and retrieval sample sets; Step 3. Construct an unsupervised cross-modal hashing retrieval network based on multi-level interaction. The entire network includes a cross-modal common hashing learning module, a modality-specific hashing learning module, and a semantic similarity preservation module; The unsupervised cross-modal hashing retrieval network based on multi-level interaction first uses the cross-modal common hashing learning module to learn the common hash code and maintain modality invariance; then, in the modality-specific hashing learning module, through the guidance of the common hash code and the combination of dual contrast learning, enhance the discriminability of the modality-specific hash code, and then achieve the interaction between common information, specific information, and modalities at the coarse-grained level; In the fine-grained interaction stage, through the semantic similarity preservation module, use the instance similarity to constrain the learning of the common hash code and the modality-specific hash code to promote the information interaction between instances, so as to ensure that similar samples learn similar hash codes; Step 4. Train the unsupervised cross-modal hashing retrieval network based on multi-level interaction with the image-text training sample set, and verify the trained network with the image-text query sample set after each batch of training is completed to check the status and convergence of the current method; Step 5. Input the query sample set and retrieval sample set of the image-text into the trained unsupervised cross-modal hashing retrieval network based on multi-level interaction, generate the corresponding hash codes respectively, and obtain the query results by calculating the Hamming distance between the query sample and the retrieval sample hash codes. The one with the smallest Hamming distance is the final query result.

2. The unsupervised cross-modal hashing retrieval method based on multi-level interaction according to claim 1, wherein: The specific steps of Step 3 are as follows: Step 3.

1. Build a cross-modal common hashing learning module; The cross-modal common hashing learning module first uses modality-specific feature extractors to extract the feature representations of images and texts respectively; secondly, introduce a self-attention mechanism to promote the interaction of intra-modal information to capture the potential semantic associations within the modality; finally, to further enhance the semantic consistency between modalities, introduce a reconstruction loss function to maintain modality invariance, so as to effectively learn the common hash code between different modalities; The modality-specific feature extractor consists of an image feature extraction network based on the ResNet-18 architecture and a text feature extraction network based on the BoW model; Step 3.

2. Build a modality-specific hashing learning module; The modality-specific hashing learning module includes a cross-modal preliminary interaction module and dual contrast learning; specifically, first use the self-attention mechanism for preliminary modality interaction, secondly, strengthen the interaction through dual contrast learning; finally, optimize the modality-specific hash code under the guidance of the common hash code to achieve the interaction between common information and specific information at the coarse-grained level; Step 3.

3. Build a semantic similarity preservation module; The semantic similarity preservation module first constructs a modality-specific similarity matrix, and then fuses the multimodal similarity modeling instance relationship to guide the learning of public hash codes and modality-specific hash codes, thereby enhancing the interaction between instances at a fine-grained level.

3. The unsupervised cross-modal hashing retrieval method based on multi-level interaction according to claim 1, wherein: The Step 4 includes Step 4.1, optimizing the public hash code part: The specific steps of Step 4.1 are as follows: Step4.1.

1. Calculate the error loss and utilize the reconstruction error loss L r Maintain modal invariance, and the loss function is as follows: where ||·|| F is the Frobenius norm of the matrix, F I ′ and F T ′ represent the reconstructed features of the image and text modalities respectively, F I and F T are the feature representations of the image and text respectively; Step4.1.

2. Use the discrete loss L dis to learn the common hash code B: Where, sign(·) is the signum function; Step 4.1.3, at the fine-grained interaction level, use the adjacency-related enhancement loss to keep the instance similarity in the public hash code B as much as possible, expressed as: where cos(·,·) represents the cosine similarity between two vectors, and S is the instance similarity matrix, and are used to maintain instance similarities in the image, text, and joint representations, respectively. The cross-modal common hash learning module is updated using a stochastic gradient descent algorithm, and the hash code reconstruction error loss, discrete loss, and adjacency-related enhancement loss are integrated into the same framework to learn the hash code calculation loss value, which can be expressed as: L cmch = L r + L ace + L dis .

4. The unsupervised cross-modal hashing retrieval method based on multi-level interaction according to claim 1, characterized in that: The Step 4 includes Step 4.2, optimizing the hash code part of the specific mode, and the specific steps of Step 4.2 are as follows: Step 4.2.

1. In order to further enhance modal interaction and improve representation capabilities, the modality-specific hash learning module introduces double contrastive learning. The first contrastive learning focuses on the representation level: Among them and represent the i-th elements of Z I and Z T respectively, τ represents the temperature coefficient, <·,·> represents the inner product; it should be noted that and are respectively the truly aligned sample pairs; therefore, the mathematical expression of the first contrastive loss is as follows: The second contrastive learning focuses on the hash code level of a specific modality, and its objective function is expressed as: where and represent the consecutive hash codes of a specific modality, respectively and are the i-th elements of and are respectively aligned with each other, and these consecutive hash codes of specific modalities and are generated by a hash layer based on Z I and Z T as inputs, where the hash layer is composed of a multi-layer perceptron MLP, and the second contrastive loss is shown as follows: Step4.2.

2. To achieve coarse-grained interaction between public information and specific information, a public hash code B is introduced to guide the learning process of the continuous hash code of a specific modality, and its loss function is defined as follows: and as follows: Step4.2.

3. At the fine-grained interaction level, to maintain the instance similarity between and , the following structural consistency loss is proposed: From another perspective, the loss function maintains the numerical consistency between different modal hash codes and the common modal hash code, while the structural consistency loss maintains the structural consistency within and between modalities of the specific modal hash code; Therefore, the final loss is as follows: The stochastic gradient descent algorithm is used to update the modality-specific hash code learning module. By integrating the public dual contrast loss and structural consistency loss into a unified learning framework, the overall objective function of the modality-specific hash learning module is as follows: L msh = L ch + λL cf + L f Among them, λ is used to measure the importance of each loss function.

5. Unsupervised cross-modal hashing retrieval system based on multi-level interaction, characterized in that include: A module for executing the unsupervised cross-modal hash retrieval method based on multi-level interaction as described in any one of claims 1 to 4.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the unsupervised cross-modal hash retrieval method based on multi-level interaction as described in any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the unsupervised cross-modal hash retrieval method based on multi-level interaction as described in any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the unsupervised cross-modal hash retrieval method based on multi-level interaction as described in any one of claims 1 to 4 is implemented.

Citation Information

Cited By

  • Unsupervised cross-modal retrieval method based on memory preserving hash

    CN121502053A