Asymmetric cross-modal hash learning method for potential structure information
By jointly exploring the clustering structure and semantic similarity of multimodal samples and designing a discrete optimization algorithm to generate hash codes, the problems of insufficient discriminability and high computational overhead in existing methods are solved, and the accuracy and efficiency of cross-modal retrieval are improved.
Patent Information
- Application Number
- CN202510868511.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-26
AI Technical Summary
Existing cross-modal hashing methods ignore the potential clustering structure information of samples when learning hash codes, resulting in insufficient discriminability, high computational overhead, and significant quantization errors, which affects retrieval performance.
By jointly exploring the clustering structure, sample similarity and semantic similarity of multimodal samples, an auxiliary variable decomposition objective function is introduced, and an effective discrete optimization algorithm is designed to generate more discriminative hash codes.
The retrieval performance of cross-modal retrieval is improved, quantization errors are avoided, computational complexity is reduced, and the generated hash code can more accurately reflect the potential structural information of multimodal data.
Smart Images

Figure CN120705605A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal retrieval technology, and in particular to an asymmetric cross-modal hash learning method for potential structural information. Background Art
[0002] Currently, with the rapid development of information technology, the amount of data in various forms has exploded, posing a significant challenge to information retrieval tasks. In recent years, cross-modal retrieval has attracted widespread attention in the fields of computer vision and multimedia. Its goal is to retrieve semantically related samples from another modality given a query sample from one modality, for example, using text to search for images related to its content, and vice versa. As an approximate nearest neighbor search method for multimodal data, cross-modal hashing (CMH) has been proven to be one of the most effective methods for various cross-modal retrieval tasks. The core idea of the CMH method is to transform heterogeneous multimodal data into a Hamming space, where efficient similarity search can be performed by performing an XOR operation on binary codes.
[0003] However, due to the differences in dimensionality and distribution of data from different modalities, there is a semantic gap between samples from different modalities, which leads to insufficient discriminability of the hash codes learned by the CMH method, greatly reducing the retrieval performance. Summary of the Invention
[0004] Based on this, it is necessary to provide an asymmetric cross-modal hash learning method for potential structural information to address the above technical problems. This method can obtain more discriminative hash codes, thereby improving the retrieval performance of cross-modal retrieval.
[0005] The present invention adopts the following technical solutions: The present invention provides an asymmetric cross-modal hashing learning method for potential structural information, comprising: Establish an objective function; the objective function aims to jointly explore and analyze the multimodal sample clustering structure, sample similarity information, and semantic similarity information; Auxiliary variables are introduced to separate the variables of the objective function, and the calculation formulas of the projection matrix of different modes, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables are obtained; Based on the multimodal training samples and label matrix, as well as the calculation formulas for the projection matrix of each modality, the calculation formula for the cluster center of the projection space, the calculation formula for the cluster indicator matrix, the calculation formula for the mapping matrix, the calculation formula for the discrete hash code, and the calculation formula for the auxiliary variables, the projection matrix, the cluster center of the projection space, the cluster indicator matrix, the mapping matrix, the discrete hash code, and the auxiliary variables of each modality training sample are iteratively updated; the update of each variable in the iterative process depends on the current optimal solution of other variables; Based on the discrete hash code and multimodal training samples obtained in the last iteration, the hash function of each modality is determined; the hash function is used to perform similarity retrieval on the query sample in the multimodal data.
[0006] Optionally, the objective function is: ; in, represents the number of modes, Indicates the The training samples of each modality, Indicates the The projection matrix of the modal, represents the cluster indicator matrix, Indicates the The cluster centers of the projection space of the modes, represents a discrete hash code, represents the mapping matrix, The Laplacian matrix representing the consistent similarity graph of multiple modalities, 、 and represents the balance parameter, S represents the pairwise semantic similarity matrix, , represents the identity matrix, Y represents the label matrix corresponding to the multimodal training samples, is used The norm is obtained by normalizing each row vector of Y.
[0007] Optionally, The projection matrix of the modal The calculation formula is: ; in, is a diagonal matrix, and the diagonal elements are calculated as ; No. The projection space of the modal The calculation formula of the cluster center is: ; in, and It is through The singular value decomposition of The calculation formula of the cluster indicator matrix V is: ; in, , the optimal solution is given by Given, where and Depend on The compact singular value decomposition of is obtained; , µ is the relaxation parameter that ensures A is positive definite, , δ >0 is the penalty factor, V=Z and Z>0; Mapping Matrix The calculation formula is: ; in, , the optimal solution is given by Given, and Depend on The compact singular value decomposition of is obtained; The calculation formula of discrete hash code B is: ; in, is the element-wise sign function; Auxiliary variables The calculation formula is: .
[0008] Optionally, before iteratively updating the projection matrix of each modality, the cluster center of the projection space, the cluster indicator matrix, the mapping matrix, the discrete hash code, and the auxiliary variables, the method includes: Based on the multimodal training samples, a consistent similarity graph of multiple modalities is constructed; Based on the consistent similarity graphs of multiple modalities, the cluster indicator matrix and discrete hash code are randomly initialized using the standard normal distribution, and the diagonal matrix is initialized using the identity matrix. .
[0009] Optionally, q A modal hash function for: ; in, is the regularization parameter.
[0010] Optionally, the method further includes: Determine a hash code of the query sample based on a hash function of the query sample and the modality corresponding to the query sample; Similarity retrieval in multimodal data via hash codes.
[0011] Optionally, query the sample's hash code The calculation formula is: ; in, Indicates the q A sample query of the modality.
[0012] The present invention provides an asymmetric cross-modal hash learning device for latent structural information, comprising: The building module is used to establish the objective function; the objective function aims to jointly explore and analyze the sample clustering structure, sample similarity information and semantic similarity information of multiple modalities; Decomposition module, used to introduce auxiliary variables to separate the variables of the objective function, and obtain the calculation formulas of the projection matrix of different modes, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables; The training module is used to iteratively update the projection matrix, cluster center of projection space, cluster indicator matrix, mapping matrix, discrete hash code and auxiliary variables of each modal training sample based on the multimodal training samples and label matrix, as well as the calculation formula of the projection matrix of each modality, the calculation formula of the cluster center of projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables; the update of each variable in the iterative process depends on the current optimal solution of other variables; the hash function of each modality is determined based on the discrete hash code and multimodal training samples obtained in the last iteration; the hash function is used to perform similarity retrieval on the query sample in the multimodal data.
[0013] The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned asymmetric cross-modal hash learning method for potential structural information.
[0014] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the asymmetric cross-modal hash learning method for potential structural information is implemented.
[0015] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: The present invention jointly explores the multi-level potential structural information of multimodal samples, including clustering structure, sample similarity and semantic similarity, so that they can promote each other, avoiding the limitation of only considering sample similarity information in the existing technology, and then introduces auxiliary variables to decompose the objective function, decomposing the objective function into multiple independent sub-problems. Each sub-problem corresponds to the optimization of a variable, and constructs a discrete hash code directly from the learned projection matrix, cluster center of the projection space, cluster indicator matrix, and mapping matrix. By alternating and iteratively updating variables such as the projection matrix, cluster center, and discrete hash code, the global optimal solution is gradually approached. This not only retains the sample similarity and potential clustering structure information in the discrete hash code, but also avoids the quantization error problem, and can obtain accurate discriminant hash codes, thereby improving the retrieval effect of cross-modal retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0017] Figure 1 A schematic diagram of the process flow of an asymmetric cross-modal hash learning method for latent structural information provided by the present invention; Figure 2 The overall framework diagram of an asymmetric cross-modal hashing learning method for latent structural information provided by the present invention; Figure 3 Schematic diagram of the PR curve for comparing baseline methods on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets; Figure 4 Schematic diagram of the Top-k accuracy curves of the baseline methods compared on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets; Figure 5 Schematic diagram of the visualization results of LSOAH and ASFOH on MIRFlickr; Figure 6 Schematic diagram of the change of the objective function value of LSOAH with the number of iterations on all data sets; Figure 7 Schematic diagram of the change of MAP value relative to corresponding parameters on three data sets under 64 hash bits for the method provided by the present invention; Figure 8 A schematic diagram of a computer device for implementing an asymmetric cross-modal hashing learning method for latent structural information provided by the present invention. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] A fundamental challenge facing CMH methods is how to explore the latent structure of cross-modal samples while eliminating data heterogeneity. Existing CMH methods are primarily categorized as unsupervised and supervised, depending on whether or not labels are used. Unsupervised methods can learn hash codes without supervised information. Their core approach is to maintain sample distance relationships by leveraging the topological structure of samples. Supervised methods learn semantically-aware hash codes by exploiting label information, achieving superior performance compared to unsupervised methods. In recent years, deep neural networks, with their powerful data representation learning capabilities, have greatly facilitated the development of deep cross-modal hashing methods. Although deep cross-modal hashing methods have demonstrated excellent performance in cross-modal retrieval, even with the help of high-performance devices such as Graphics Processing Units (GPUs) or Tensor Processing Units (TPUs), they still require large training datasets and lengthy training times. Therefore, this paper aims to propose an efficient cross-modal hashing method that fully exploits the latent structure of heterogeneous samples and generates more discriminative hash codes without relying on expensive hardware. As far as the applicant knows, this is still a problem that needs to be solved urgently when facing large-scale heterogeneous data scenarios.
[0020] Although existing supervised CMH methods have good performance, the following problems still need to be solved: (1) Most existing CMH methods use sample similarity or semantic similarity to guide the learning of hash codes, ignoring the potential clustering structure information of samples, resulting in insufficient discriminability of the learned hash codes. (2) Most existing CMH methods use complex fusion mechanisms to fuse predefined similarity graphs in different modalities, making the retrieval performance highly dependent on the quality of the predefined graphs, sensitive to noise, and computationally expensive. (3) To facilitate optimization, some methods relax the binary constraints of hash codes to continuous values, resulting in significant quantization errors and suboptimal solutions.
[0021] Two commonly used methods in the prior art are described below.
[0022] Unsupervised methods based on sample similarity learning: Unsupervised CMH aims to generate hash codes by exploring the intrinsic structural information of samples to maintain the similarity structure of the original samples. The basic components of some typical unsupervised CMH methods are mainly developed from collaborative matrix factorization. For example, Latent Semantic Sparse Hashing (LSSH) uses matrix factorization and sparse coding to learn the latent space and then combines the learned latent features to generate hash codes. Collaborative Matrix Factorization Hashing (CMFH) adopts matrix factorization technology to learn shared representations and generates hash codes by quantizing the shared representations. Building on this, some methods introduce graph techniques to model sample similarity and promote hash code learning. Fusion Similarity Hashing (FSH) simultaneously learns the fusion similarity and hash codes of samples, thereby maintaining the fusion similarity in the hash codes. Adaptive Structural Similarity Preserving for Unsupervised Cross Modal Hashing (ASSPH) uses adaptive correlation expansion and constraints to adaptively learn joint semantic relationships between samples and embed them into hash code learning. Unsupervised Cross-Modal Hashing with Modality-Interaction (UCHM) directly learns cross-modal similarity from raw data to guide deep hashing networks in learning hash codes. Mining Similarity Relationships for Unsupervised Cross-Modal Hashing (MSRH) jointly learns to fuse similarity graphs and feature cross-reconstruction in a unified framework, bridging the semantic gap and preserving potential cross-modal semantic similarity. These methods primarily focus on exploring sample similarities within and between modalities to guide hash code learning. They can utilize high-level semantic information to process heterogeneous data and obtain hash codes. However, they only consider sample similarity and ignore the potential clustering structure and fusion similarity of multimodal data that contain complementary information, which often contains the key discriminative and heterogeneous correlation information for cross-modal retrieval performance.
[0023] Supervised methods based on label semantics: In contrast, supervised CMH can utilize given supervisory information (e.g., labels) to achieve better retrieval performance. This is mainly because label semantic information can effectively help the model capture high-level semantic relationships between heterogeneous data and overcome the semantic gap problem between modalities. Based on the semantic form used, supervised CMH can be roughly divided into two categories: label-based hashing and hashing methods based on label pairwise similarity. The former directly uses binary labels to supervise hash code learning. For example, discrete cross-modal hashing (Learning Discriminative Binary Codes for Large-scale Cross-modal Retrieval, DCH) directly learns a linear mapping from samples to hash codes to predict category information in a classification framework. Fast Discriminative Discrete Hashing for Large-Scale Cross-Modal Retrieval (FDDH) simultaneously regresses the target hash code to its relaxed label and uses an orthogonal transformation scheme to map the embedded sample into a semantic subspace. A Scalable Discrete Matrix Factorization Hashing Framework for Cross-Modal Retrieval (SCRATCH) combines collective matrix factorization and label semantic embedding to learn shared semantic representations and generate hash codes from them. Exploiting Subspace Relation in Semantic Labels for Cross-Modal Hashing (SRLCH) jointly learns a mapping from label space and nonlinear embedded samples to Hamming space, obtaining a unified form of hash codes and hash functions. Efficient discrete cross-modal hashing with semantic correlations and similarity preserving (EDCH) transforms heterogeneous features into a common latent representation and then learns a bilinear feature-label mapping to generate hash codes. Supervised Adaptive Similarity Consistent Latent Representation Hashing (SCLRH) learns a similarity-consistent representation for multimodal data and then generates hash codes based on the consistent representation and labels. The latter usually uses labels to construct pairwise semantic matrices to measure the cross-modal similarity between samples.For example, Semantics-Preserving Hashing for Cross-View Retrieval (SePH) converts semantic similarity into a probability distribution and transfers it to the hash code by minimizing the KL divergence. Semi-Relaxation Supervised Hashing for Cross-Modal Retrieval (SRSH) learns hash codes and hash functions through a relaxation strategy guided by a predefined semantic similarity matrix. Supervised Matrix Factorization Hashing for Cross-Modal Retrieval (SMFH) maintains the local structure consistency of samples from different modalities through graph regularization and generates hash codes that satisfy semantic consistency through collaborative matrix factorization. Adaptive Label Correlation Based Asymmetric Discrete Hashing for Cross-Modal Retrieval (ALECH) learns latent features and hash codes through an asymmetric discrete strategy guided by label correlation. These methods have achieved good performance. However, DCH, FDDH, SCRATCH, SRLCH, and ALECH utilize label semantics to learn hash codes, which may ignore sample semantic information that contains potential separation information. SRSH and SMFH use relaxed optimization schemes and ignore quantization losses. DCH jointly learns hash functions and hash codes, which increases the optimization complexity. SCLRH uses an adaptive pairwise similarity graph, which is difficult to handle large-scale datasets. Furthermore, pre-defining the similarity matrix using label information alone does not accurately reflect the true similarity relationship between image-text pairs in the same category. Recently, several deep hashing methods (such as Deep Cross-Modal Hashing (DCMH), Self-Supervised Adversarial Hashing (SSAH), Deep Multiscale Fusion Hashing (DMFH), and Deep Cross-Modal Hashing Based on Semantic Consistent Rankin (DCH-SCR)) have been proposed, which combine feature and hash code learning into an end-to-end learning framework.Although deep hashing methods often show excellent performance due to the powerful data representation learning ability of DNNs, most of them are limited by high computational cost and exhaustive search for optimal parameters. Therefore, designing an interpretable objective function with an efficient discrete optimization strategy is crucial for learning discriminative compact binary hash codes.
[0024] To address these issues, this paper proposes a latent structure-oriented asymmetric cross-modal hashing learning method (LSOAH). This method combines sample clustering, sample similarity, and semantic similarity exploration with hash code learning. By integrating clustering, sample similarity, and semantic similarity exploration into hash code learning, we can directly obtain the most discriminative hash codes for cross-modal retrieval. The main contributions of this paper are as follows.
[0025] (1) A joint learning framework is proposed, which integrates the exploration of multi-level latent structural information of samples (i.e., clustering structure, sample similarity, and semantic similarity) and the learning of hash codes into a unified framework. The consistent structural information between different modalities is embedded in the hash codes, thereby eliminating the heterogeneity gap between multimodal samples.
[0026] (2) By performing orthogonal basis decomposition on the projected samples, a modal consistency cluster indicator matrix is obtained, which is used to guide the learning of hash codes that retain the sample clustering information as much as possible.
[0027] (3) By fusing sample similarity graphs from different modalities through the Hadamard product and associating them with hash codes in an asymmetric strategy, hash codes that can perceive sample similarity and semantic similarity can be directly generated.
[0028] (4) An efficient discrete optimization algorithm is designed to solve the model, avoiding large quantization errors. Extensive experiments on multiple benchmark datasets show that the LSOAH method outperforms the existing CMH method.
[0029] The technical solutions provided by various embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0030] Figure 1 The figure is a flow chart of an asymmetric cross-modal hash learning method for latent structural information in the present invention, which specifically includes the following steps: S101, establish an objective function; the objective function aims to jointly explore and analyze the sample clustering structure, sample similarity information and semantic similarity information of multiple modalities.
[0031] First, let’s define the symbols involved in this invention. In this article, matrices and vectors are represented by bold capital letters and lowercase letters respectively. , , ,and Respectively represent the first i Row, No. j Column and ( i , j ) elements. If M is a square matrix, represents the transpose of the matrix M, and Tr(M) represents the trace of M. and Represents the Frobenius norm of the matrix and the vector norm. Indicates M The norm, I, 1, and 0 represent the identity matrix, the all-one vector, and the all-zero vector, respectively. Table 1 lists other symbols and definitions used in this paper.
[0032] Table 1 The present invention first analyzes from three aspects: potential clustering structure exploration, sample similarity embedding and semantic similarity preservation.
[0033] (1) Exploration of potential clustering structures To eliminate the heterogeneity differences between different modalities, CMFH and its variants decompose heterogeneous data into the product of modality-specific factors and a common latent representation, and generate hash codes by quantizing the latent representation. However, CMFH has two problems that need to be solved: 1) It does not consider the potential separable information in the feature distribution, that is, the clustering structure of the samples; 2) The generation of hash codes requires rounding or relaxing the quantization operation of the latent representation, which causes significant quantization errors and greatly reduces the retrieval performance. To solve the first problem, the projected heterogeneous data can be orthogonally decomposed to generate a multimodal-specific basis matrix and a modality-consistent clustering indicator matrix. The process can be expressed as follows:
[0034] , (1) in, It is q The projection matrix of the modal samples, represents the modal consistency cluster indicator matrix, It is q The specific basis matrix of the mode, γ is the equilibrium parameter. The orthogonality constraint ensures that Each column of is independent, so in the projection space middle, It can be viewed as a set of cluster centers. Orthogonality and non-negativity constraints Ensure that each row of V has a non-zero element, which means that each sample can only be associated with one cluster through the cluster indicator matrix V.
[0035] From the perspective of multimodal data information mining, the model has the following characteristics: First, the sample information may have inherent differences between different modalities, and the cluster centers in different modalities may be different. Therefore, the orthogonal basis matrix is designed to be modality-specific to capture the unique information of multiple modalities. Secondly, different modalities describe the same sample object in different ways, and they should share a consistent clustering structure. Therefore, a shared cluster indicator matrix V is introduced to explore the consistency information in heterogeneous data. In addition, Regularized replace Orthogonal decomposition helps to more flexibly utilize key features to explore the potential clustering structure of samples and reduce the negative impact of redundant and noisy features.
[0036] (2) Sample Similarity Embedding Samples from different modalities have complementary information, and fusion similarity is the key information to capture the geometric structure of cross-modal samples. Inspired by the idea of GSF, this paper first constructs multiple sample similarity graphs for multimodality, and then calculates the fusion similarity graph to embed the intrinsic structure shared by different modalities into the hash code. q modal n samples ,set up is the pairwise sample distance, q modal k -NN graph The structure is as follows
[0037] (2) In order to mine the complementary shared similarity structure information between different modalities, the present invention uses Hadamard product to extract common edges from multiple sample similarity graphs and converts multiple graphs into Fusion into a consistency graph G. The details are as follows:
[0038] (3) Where Π represents the Hadamard product of the sequence.
[0039] In order to explore the geometric structure between heterogeneous samples and make full use of the complementary information between different modalities, the fusion similarity graph G is embedded in the hash code learning. Intuitively, the distance between the clusters to which different samples belong is The smaller the value, the corresponding similarity value The bigger the value, the smaller the value, and vice versa:
[0040] (4) Where L represents the Laplace matrix of the graph G, and the calculation formula is L=D−G. The calculation formula of the diagonal matrix D is Formula (4) is used to pull similar samples closer and push different samples further apart in the shared space V. Finally, V is used to guide the learning of B, thereby maintaining the fusion similarity between samples in the hash code.
[0041] (3) Preserving semantic similarity In order to better utilize label information, semantic similarity is introduced to establish semantic connections between heterogeneous samples. A typical mechanism is to minimize the difference between pairwise label distance and pairwise hash code distance. The learning problem is defined as follows:
[0042] (5) Where S represents the pairwise semantic similarity matrix. In order to reduce the computational complexity and align the similarity between labels and hash codes, S is defined as ,in, is used The norm is obtained by normalizing each row vector of Y. is the identity matrix. To simplify the calculation, it is not necessary to explicitly calculate S during optimization, but its expression can be used directly.
[0043] However, formula (5) requires solving a symmetric inner product difference problem, which is an NP-hard problem. To solve this problem, consider replacing one B in the inner product with other variables. As mentioned above, embedding the clustering structure into the hash code can significantly improve cross-modal retrieval performance. Therefore, the present invention uses V to construct B and uses it to replace one B in the inner product, that is:
[0044] (6) in, is the mapping matrix, α and β is a balancing parameter. The orthogonality constraint on P aims to preserve clustering information while reducing redundancy to a certain extent. Obviously, the hash code B and the mapping P are learned jointly, which enables B to preserve clustering structure information and effectively avoid the quantization loss problem faced by the CMFH method.
[0045] By combining formulas (1), (4) and (6), the overall objective function of LSOAH is derived as follows: (7) in, represents the number of modes, Indicates the The training samples of each modality, Indicates the The projection matrix of the modal, represents the cluster indicator matrix, Indicates the The cluster centers of the projection space of the modes, represents a discrete hash code, represents the mapping matrix, represents the Laplacian matrix of the multimodal consistency graph, 、 and represents the balance parameter, S represents the pairwise semantic similarity matrix, , Y represents the label matrix corresponding to the training samples of multiple modalities, is used The norm is obtained by normalizing each row vector of Y.
[0046] It should be noted that the objective function is the object of optimization in the training process. Generally speaking, it is to minimize the objective function through multiple iterations. The values of each variable corresponding to the minimum objective function are the final model parameter values.
[0047] S102, introduce auxiliary variables to perform variable separation on the objective function, and obtain the calculation formulas of the projection matrix of different modes, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables.
[0048] Specifically, in order to facilitate optimization, the objective function (Formula (7)) is first reformulated into the following equivalent form: ; (8) An auxiliary variable Z is introduced to separate the non-negative constraint of V, ensuring that V=Z and Z>0. By incorporating the equality constraint V=Z into the objective function through the penalty term, we obtain:
[0049] ; (9) in, δ >0 is the penalty factor. If the penalty factor δ If is large enough, the local optimal solution of formula (8) can be obtained by minimizing formula (9). Considering that formula (9) is jointly non-convex, the alternating multiplier optimization method is used to solve it alternately, as shown below.
[0050] No. q The projection matrix of the modal Steps: With other variables fixed, The sub-problem is (10) have ,in is a diagonal matrix, and the diagonal entries are calculated as By setting , get the The projection matrix of the modal The calculation formula is:
[0051] (11) in, is a diagonal matrix, and the diagonal elements are calculated as .
[0052] No. q The cluster centers of the projected space of the modes Steps: When other variables are fixed, you can update by solving the following problems : (12) Solve formula (12) and get The cluster centers of the projected space of the modes The calculation formula is: (13) in, and It is through Formula (12) is an orthogonal Pluck problem, so it is updated at each iteration. This will result in a monotonically decreasing target value.
[0053] Cluster indicator matrix V step: When other variables are fixed, the sub-problem about V is: (14) in, Formula (14) is a second-order Stiefel manifold problem.
[0054] To solve formula (14), it needs to be transformed into: (15) in, , µ is the relaxation parameter that ensures that A is positive definite. Then, the Lagrangian function of formula (15) is obtained:
[0055] (16) where Λ is the Lagrange multiplier. By setting ,get:
[0056] (17) Inspired by the power iteration method, based on formula (17), the optimization problem formula (15) is transformed into solving the following problem, that is, the optimization problem of the cluster indicator matrix V is: (18) in, , the optimal solution is given by Given, where and Depend on The compact singular value decomposition of is obtained; , µ is the relaxation parameter that ensures A is positive definite, , δ >0 is the penalty factor, V=Z and Z>0.
[0057] Mapping matrix P step: When other variables are fixed, the sub-problem about P is: (19) By constraints The optimization problem formula (19) can be simplified to: (20) in, Similar to formula (18), the optimal solution is given by Given, where and Depend on The compact singular value decomposition of .
[0058] Discrete Hash Code B Step: By removing items that are not related to B, the subproblem about B can be formulated as: (twenty one) Formula (21) can be equivalently transformed into the following problem, namely formula (22).
[0059] (twenty two) because By setting , the calculation formula of the discrete hash code matrix B is:
[0060] (twenty three) in, is the element-wise sign function.
[0061] Auxiliary variable Z step: When other variables are fixed, the sub-problems about Z are as follows: (twenty four) The subproblem related to Z is easy to solve. Auxiliary variables The calculation formula is:
[0062] (25) S103, based on the multimodal training samples and label matrix, as well as the calculation formula of the projection matrix of each modality, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables, iteratively update the projection matrix of each modality training sample, the cluster center of the projection space, the cluster indicator matrix, the mapping matrix, the discrete hash code and the auxiliary variables; the update of each variable in the iterative process depends on the current optimal solution of other variables.
[0063] In one embodiment, before iteratively updating the projection matrix, cluster centers of the projection space, cluster indicator matrix, mapping matrix, discrete hash code, and auxiliary variables of each modality, a consistent similarity graph of multiple modalities is constructed based on the multimodal training samples; based on the consistent similarity graph of multiple modalities, the cluster indicator matrix and discrete hash code are randomly initialized using a standard normal distribution, and the diagonal matrix is initialized using the identity matrix. The specific implementation method is the same as the specific calculation method of the above embodiment, and is not limited in this embodiment.
[0064] S104, determining the hash function of each modality based on the discrete hash code obtained in the last iteration and the multimodal training sample; the hash function is used to perform similarity retrieval on the query sample in the multimodal data.
[0065] After obtaining the optimal discrete hash code B using the above process, a modality-specific hash function is learned for predicting samples. This can be solved by a binary classification problem. Typical methods include but are not limited to SVM, DNN, and linear regression, which can be used to learn classifiers. In order to balance retrieval speed and performance, the present invention uses linear regression to learn hash functions. In addition, graph regularization is introduced to further improve generalization ability. Specifically, q The hash function of a modality can be obtained through the following model:
[0066] (26) in, η >0 is the regularization parameter. By applying formula (26) to the variable The partial derivative of is set to zero, and the optimal , that is, q A modal hash function Defined as:
[0067] (27) In one embodiment, after obtaining the hash functions of each modality, the hash code of the query sample can be determined based on the query sample and the hash function of the modality corresponding to the query sample; and similarity retrieval can be performed in the multimodal data using the hash code.
[0068] For the q Query samples of modalities , its hash code b can be generated as follows: (28) The above examples explain the proposed method (LSOAH) from three aspects: model construction, optimization process and theoretical analysis. The model has four aspects: latent clustering structure exploration, sample similarity embedding, semantic similarity preservation and hash function learning. The overall framework of LSOAH is as follows: Figure 2 The optimization process gives the update process of each variable, and the theoretical analysis shows the convergence and complexity of the algorithm.
[0069] The specific theoretical analysis includes: the convergence and computational complexity of the LSOAH method.
[0070] (1) Convergence analysis: 1. For the optimization problem, i.e., formula (8), ADMM is used to iteratively calculate the optimal solution of variables Θ, U, V, P, B and auxiliary variable Z. The proof of LSOAH convergence is given by the following Theorem 1.
[0071] Theorem 1: The objective function of optimizing LSOAH is as t The iteration of is monotonically decreasing, that is , in Indicates the first t The objective function value at the iteration 。
[0072] In one embodiment, the proof process of Theorem 1 is given, which theoretically proves that the objective function is monotonically decreasing after updating each variable in each iteration. The specific proof process is as follows:
[0073] For convenience, the objective function of formula (9) is expressed as: (29) renew :According to formula (29), only The sub-problems are: (30) According to the update rule of formula (11), The following fixed The solution to the problem is as follows: (31) This shows that (32) For any nonzero row vector f, , the inequality is given by: (33) Applying the above inequality to and ,get: (34) Formula (34) can be rewritten in matrix form as follows: (35) Combining formula (32) and (A35), we get: (36) This means, (37) renew :because It is updated according to formula (13), and we get: (38) Formula (38) can be regarded as an orthogonal programming problem, so the update of each iteration is will make the value of the objective function monotonically decreasing, which means (39) Update V: According to the update rule of V, that is, formula (18), we have (40) This shows that (41) based on , formula (41) can be rewritten as: (42) Since the matrix A is positive definite, we have (43) According to formulas (42), (43) and , it can be inferred that (44) This shows (45) Update P: According to the update rule of P, that is, formula (20), we have: (46) This shows that: (47) According to formula (47), we have: (48) Update B: According to the update rule of B, that is, formula (22), we have: (49) This shows (50) According to formula (50), it can be inferred that: (51) This shows that: (52) Update Z: Similarly, according to the update formula of Z, we have: (53) This shows that: (54) According to formulas (37), (39), (45), (48), (52) and (54), we can conclude that: (55) This shows that the objective function (Formula (29)) decreases monotonically with each iteration. In addition, since the function (Formula (29)) is convex with respect to each variable, the update algorithm converges.
[0074] (2) Complexity analysis t Represents the number of iterations. In the hash code generation phase, for each iteration, update , the time overhead of V and P is Solving for B and Z involves only some basic operations such as matrix / vector addition and matrix / vector multiplication, which are very fast compared to operations such as matrix inversion and eigendecomposition, so their computational complexity can be ignored. During the hash function learning phase, the computation of the projection parameter W is approximately .because , the total complexity of LSOAH is In summary, the complexity of LSOAH is n It is linear.
[0075] In a specific embodiment, to comprehensively evaluate LSOAH, we conducted extensive experiments on three public datasets: MIRFlickr, NUS-WIDE, and IAPR-TC12. MIRFlickr contains 25,000 samples consisting of image-text pairs. Each sample belongs to at least one of 24 categories and is represented by a 512-D GIST visual feature vector and a 1386-D bag-of-words text vector, respectively. We selected samples with at least 20 labels, resulting in a total of 20,015 samples. The dataset was randomly split into 18,015 / 2,000 training (retrieval) and query sets. NUS-WIDE contains 269,648 samples, with 186,577 samples with the top 10 most frequent labels selected from 81 categories. Each image is represented by a 500-D bag-of-visual-words feature vector, and the text is described by a 1,000-D bag-of-words feature vector. The NUS-WIDE dataset was split into 184,577 / 2,000 training (retrieval) and query sets. IAPR-TC12 contains 20,000 samples, each annotated with at least one of 255 labels. Each image and text is represented by a 512-D GIST feature vector and a 2912-D bag-of-words vector, respectively. The dataset is randomly split into 18,000 / 2,000 training (retrieval) / query sets.
[0076] The effectiveness of the proposed method is verified by two retrieval tasks: 1) image2text, which is to retrieve text using image queries; 2) text2image, which is to retrieve images using text queries. In the experiment, three widely used evaluation metrics are used, namely, Average Precision (MAP), Precision-Recall (PR), and Top- k Precision (P@ k ) to evaluate the quality of retrieval results. MAP comprehensively considers the average precision of all test samples. The larger the MAP, the better the retrieval results, indicating that the samples relevant to the query are ranked higher. The PR curve analyzes the trade-off between precision and recall under different thresholds. Generally, the larger the area under the curve, the better the overall performance. k Indicates the search sequence before k The proportion of similar samples in the results, P@ k The trend of the curve reflects the accuracy performance.
[0077] Comparison Methods and Experimental Details: This paper compares LSOAH with ten cross-modal hashing methods, including two unsupervised methods CMFH and FSH, and eight supervised methods, namely DCH, SMFH, SCRATCH, SRLCH, ALECH, FDDH, SCLRH, and ASFOH. All comparison methods use the code and parameters given in their papers for experimental settings. About the hyperparameters of LSOAH α 、 β 、 λ and γ Grid search technology is used to select candidate sets In addition, for the proposed LSOAH, the fixed δ = To ensure the orthogonality of the cluster indicator matrix, the regularization parameter η and the maximum number of iterations τ According to experience, they are set to and 15. All experiments were performed on MATLAB R2017b using Windows 10 system on a workstation equipped with Intel(R) Core(TM) i7-8700CPU@3.2GHz and 16GB RAM.
[0078] Results and Discussion: Experimental results for two retrieval tasks on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets (I→T represents querying related text using an image, and T→I represents querying related images using text) are given in Tables 2, 3, and 4, respectively. The hash code lengths on each dataset are set to 8 bits, 16 bits, 32 bits, 64 bits, and 128 bits, respectively. The best results are shown in bold. Table 2 shows the MAP results of LSOAH and all baselines on MIRFlickr, Table 3 shows the MAP results of LSOAH and all baselines on NUS-WIDE, and Table 4 shows the MAP results of LSOAH and all baselines on IAPR-TC12.
[0079] Table 2 Table 3 Table 4 The results in the table above demonstrate that, for both cross-modal information retrieval tasks, LSOAH significantly outperforms ten compared CMH methods across all three datasets, regardless of code length. Specifically, on the MIRFlickr, NUS-WIDE, and IAPR-TC12 benchmark datasets, the proposed LSOAH achieves average performance improvements over the second-best CMH method, ASFOH, by 0.77%, 0.91%, and 1.62% in image-to-text retrieval, and by 0.73%, 1.18%, and 1.78% in text-to-image search, respectively. 2) The MAP scores of most hashing methods increase with increasing code length. The proposed LSOAH significantly improves retrieval performance with longer code lengths. This phenomenon is attributed to the fact that longer hash codes can encode more semantic information. 3) Another observation is that the average precision on the text-to-image retrieval task is higher than that on the image-to-text retrieval task. This phenomenon may be due to the fact that text conveys more precise semantic information than images and is less susceptible to noise and outliers. From these observations, it can be seen that the asymmetric framework that incorporates the latent structure of the samples gives LSOAH a significant advantage.
[0080] In addition to the MAP evaluation criteria, we also compared the PR curves and P@ k Accuracy curve, the results are as follows Figure 3 and Figure 4 As shown, Figure 3 The relationship curves of the precision and recall of the baseline methods on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets are shown below. Figure 4 Top-k accuracy curves of the baseline methods compared on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets. Figure 3 It can be seen that the area under the PR curve corresponding to the LSOAH method proposed in the present invention is the largest, which shows that LSOAH can maintain a performance balance between precision and recall in image2text and text2image retrieval tasks on all data. In LSOAH, hash code learning simultaneously explores and utilizes the multi-level potential structure of heterogeneous samples, including clustering structure, sample similarity and semantic similarity, which makes the heterogeneous samples highly aligned with the semantic information in the Hamming space. In this way, LSOAH not only eliminates the semantic gap between different modalities, but also makes full use of the complementary supervision information between different modalities to generate the most discriminative hash codes. However, other comparison methods ignore the clustering structure, resulting in insufficient mining of sample separation information, especially CMFH does not even consider semantic similarity, so its area under the PR curve is very small. Similarly, Figure 4It shows that for different k value settings, the method of the present invention has better comprehensive retrieval performance than other comparison methods, which is consistent with the results reflected by the MAP value and PR curve.
[0081] Combined with Tables 2, 4 and Figure 3 、 Figure 4 Based on the results, the following conclusions can be drawn: (1) The performance of unsupervised CMH methods (such as FSH and CMFH) is weaker than that of supervised CMH methods. (2) CMH methods that utilize sample similarity and semantic similarity (such as ASFOH, FDDH, and ALECH) are better than methods that only use label information (such as SCRATCH, SRLCH, and DCH) in terms of retrieval performance. (3) Compared with methods optimized using discrete relaxation strategies (i.e., SMFH and CMFH), CMH methods using discrete optimization strategies (i.e., FDDH, ALECH, SCRATCH, SRLCH, SCLRH, and DCH) show better retrieval performance. (4) Compared with existing retrieval baseline methods, LSOAH utilizes the multi-level latent structure of samples and achieves better retrieval results.
[0082] Retrieval result visualization: To verify the effectiveness of LSOAH from multiple perspectives, we further visualized the retrieval results of LSOAH and its best competitor ASFOH on MIRFlickr, as shown in Figure 2. Figure 5 As shown, Figure 5 Visualization results of LSOAH and ASFOH on MIRFlickr. Bold labels are manually marked as relevant. In the experiment, the code length is set to 64 bits. Figure 5 In , the first column shows the true label of the query sample, the second column shows the query sample consisting of two image-text pairs, and the last column shows the corresponding retrieval results for the image2text and text2image tasks. For each image query, all labels that appear in the top 10 retrieved texts are sorted according to their frequency of occurrence. For each text query, the top 10 retrieved images are shown, which are sorted according to the Hamming distance between the corresponding hash codes. Figure 5Visualization results show that for a given query sample, the proportion of relevant samples in LSOAH's retrieval results is much greater than that in ASFOH's retrieval results in both the image2text and text2image tasks. Specifically, in the image2text task, LSOAH retrieves nearly all key tags relevant to the retrieved image, and these tags are more semantically relevant to the image. In contrast, the tags retrieved by ASFOH have weaker semantic relevance to the images, and some highly relevant tags are not retrieved, such as metal related to the first image, green related to the second image, deer, and other key tags. On the other hand, in the text2image task, the top images in LSOAH's retrieval results closely match the semantic information of the corresponding text query, while the images retrieved by ASFOH are semantically confused with the query text. For example, the horse in the second image and the duck in the fourth image do not match the dog in the text query. This is primarily because LSOAH simultaneously leverages clustering structure, sample similarity, and semantic similarity to generate more discriminative hash codes, thereby improving retrieval quality.
[0083] Convergence and effectiveness analysis: The above examples prove that the learning process of LSOAH is convergent. In order to analyze the LSOAH optimization algorithm more comprehensively, the present invention verifies the convergence of the learning algorithm through experiments. Specifically, Figure 6 The figure shows how the objective function value of LSOAH changes with the number of iterations on all datasets. In order to clearly show the results, the objective function value is normalized to the range of [0,1] by dividing by the maximum value in the experiment. Figure 6The results show that the optimization algorithm converges after five iterations, verifying its convergence. The above analysis shows that the computational complexity of LSOAH increases linearly with the size of the training set, making it scalable to large-scale datasets. To further verify the time performance of the proposed method, Table 5 reports the training time (in seconds) of all methods on the MIRFlickr, NUS-WIDE, and IAPR-TC12 datasets. Table 5 shows that the training time of CMFH, FSH, and DCH increases significantly with increasing code length, while the training time of SMFH, SCROCH, ALECH, FDDH, ASFOH, and LSOAH increases only slightly with increasing code length. This is because the latter methods directly generate hash codes in a single step through a discrete optimization strategy. FDDH is the fastest due to its closed-form solution for the hash code. FSH is the slowest, followed by SCLRH, primarily because they incorporate graph similarity into hash code learning. Compared to the most competitive ASFOH, LSOAH offers superior performance and training speed. This is because ASFOH requires expensive kernel matrix calculations and singular value decomposition operations in its solution, while LSOAH constrains the geometric structure of samples through efficient graph regularization. Although the time cost of LSOAH is higher than some baseline methods, the performance is significantly improved.
[0084] Table 5 Parameter sensitivity analysis: The objective function of the LSOAH proposed in this paper contains four hyperparameters α 、 β 、 λ and γ . Specifically, α and β The weights of the terms are kept for semantic similarity, λ Used to adjust the contribution of sample similarity embedding, γ is the weight used to control the sparse regularization term of the projection matrix. Considering the influence of each hyperparameter on the contribution ratio between different modules of LSOAH, an alternating strategy is adopted. Adjust one parameter within a certain range while keeping other parameters fixed at their optimal values. Figure 7The figure shows the changes of MAP values of the method of the present invention on three datasets relative to the corresponding parameters under 64 hash bits, specifically including: parameter sensitivity analysis of α, β, λ and γ on the three datasets, (a) represents MAP and α of 64-bit I2T task, (b) represents MAP and β of 64-bit I2T task, (c) represents MAP and λ of 64-bit I2T task, (d) represents MAP and γ of 64-bit I2T task, (e) represents MAP and α of 64-bit T2I task, (f) represents MAP and β of 64-bit T2I task, (g) represents MAP and λ of 64-bit T2I task, and (h) represents MAP and γ of 64-bit T2I task. For the MIRFlickr dataset, the MAP value is relatively stable within a large range of changes in these parameters, that is, On the NUS-WIDE dataset, LSOAH achieves stable retrieval performance in a wider range of parameters, i.e., In addition, when LSOAH achieves the best retrieval performance on the IAPR-TC12 dataset. It can be observed that α and β It has a significant impact on the retrieval performance of all three datasets. α and β Less than MAP drops rapidly when α and β exist and When the pressure is between 0 and 1, MAP reaches the highest value and tends to be stable. α and β The main reason for the consistency of MAP is that they control the contribution of semantic similarity to hash code generation. This further illustrates the necessity of semantic similarity preservation for LSOAH to achieve the best retrieval performance. On the other hand, when λ and γ Take a very small value (i.e. less than ), the MAP score is very low, which means that it is necessary to explore the clustering structure in LSOAH.
[0085] Ablation Experiments: To analyze LSOAH more deeply, ablation experiments are conducted on the modules in LSOAH to analyze their contribution to the overall performance. The degenerate variants of LSOAH include LSOAH1, LSOAH2 and LSOAH3. α is set to zero, i.e., discarding the semantic similarity embedding module from LSOAH. LSOAH2 eliminates the constraint To ignore the potential sample clustering structure exploration. LSOAH3 by the parameter λSet to zero to abandon sample similarity embedding. The ablation experiment results of the above variants are shown in Table 6, where the MAP of all variants is lower than that of LSOAH, indicating that each module has different contributions to the overall performance.
[0086] Table 6 Specifically, the following conclusions can be drawn based on the experimental results: (1) From Table 6, we can calculate that for the I→T task, the MAP scores of LSOAH1 on the three datasets are 1.6%, 1.3%, and 1.36% lower than those of LSOAH on average. For the T→I task, the MAP scores of LSOAH1 on the three datasets are 1.45%, 1.61%, and 1.78% lower than those of LSOAH on average. These results show that semantic similarity embedding promotes the learning of discriminative hash codes. This is because semantic similarity embedding allows the hash codes to still maintain the local structure of the original label space.
[0087] (2) Similarly, for the I→T task, the MAP scores of LSOAH2 on the three datasets are 3.71%, 3.01%, and 2.89% lower than those of LSOAH, respectively. For the T→I task, the MAP scores of LSOAH2 on the three datasets are 3.41%, 3.54%, and 3.44% lower than those of LSOAH, respectively. These results indicate that exploring the latent clustering structure can significantly improve the retrieval performance of hash code learning. This is because exploring the latent clustering structure of samples can capture the separation information of the samples, thereby obtaining high-level semantic information and generating more discriminative hash codes.
[0088] (3) Referring to Table 6, LSOAH3 omits sample similarity and does not fully explore the local structure between samples, resulting in retrieval performance significantly weaker than LSOAH.
[0089] Comparison with Deep Hashing: To validate the effectiveness of LSOAH, we compared it with five leading deep cross-modal hashing methods on the MIRFlickr dataset: 1) DCMH, 2) SSAH, 3) DMFH, and 4) DCH-SCR. To ensure a fair comparison, we replaced the shallow features of the image modality used in previous experiments with 4096-D vectors extracted using the VGG-16 network. The code lengths ranged from 16 to 64, with the best results shown in bold. Table 7 reports the MAP results for LSOAH and all deep baselines.
[0090] Table 7 As can be seen from Table 7, LSOAH consistently performs well in all comparison experiments. This is because LSOAH not only makes full use of deep features, but also exploits the clustering structure and manifold similarity of samples.
[0091] In the present invention, LSOAH is proposed for cross-modal information retrieval, focusing on how to jointly utilize the clustering structure and similarity structure of samples to guide the learning of hash codes for cross-modal retrieval. Specifically, in order to explore the discriminant clustering structure, the projected data is decomposed into a modality-specific basis matrix and a modality-consistent clustering indicator matrix. At the same time, an asymmetric mechanism is used to construct a hash code from the clustering indicator matrix, and the clustering structure information and semantic similarity are embedded in the hash code. In addition, graph learning and fusion mechanisms are used to maintain consistent similarity between multimodal data and hash codes. By seamlessly modeling clustering and similarity structure learning and semantic similarity exploration in a joint objective function, the present invention designs an efficient optimization algorithm and theoretically proves its convergence. A large number of experiments on multiple benchmark datasets have confirmed the superiority of this method.
[0092] This paper proposes an asymmetric cross-modal hashing learning method (LSOAH) for latent structural information for cross-modal retrieval. Specifically, LSOAH learns a common representation of multimodal data through orthogonal decomposition, where the samples of each specific modality are projected and decomposed into a modality-specific basis matrix and a modality-consistent cluster indicator matrix, and an asymmetric mechanism is used to associate the cluster indicator matrix with a hash code. Furthermore, LSOAH exploits the Hadamard product of sample similarity graphs in different modalities to explore the consistent similarity of samples and embeds it into the common representation. Finally, a unified optimization objective function is proposed to simultaneously explore sample clustering structure, sample similarity and semantic similarity, as well as the learning of hash codes, and an alternating optimization algorithm is developed based on this, which has theoretically proven convergence. Experimental results on three benchmark datasets confirm the effectiveness of the proposed LSOAH in cross-modal retrieval.
[0093] When applying the asymmetric cross-modal hash learning method for potential structural information provided by the present invention, it is not necessary to Figure 1 The steps are executed in the order shown. The specific execution order of the steps can be determined according to needs, and the present invention does not limit this.
[0094] The above is an asymmetric cross-modal hash learning method for latent structural information provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding asymmetric cross-modal hash learning device for latent structural information, which includes: The building module is used to establish the objective function; the objective function aims to jointly explore and analyze the sample clustering structure, sample similarity information and semantic similarity information of multiple modalities; A decomposition module is used to introduce auxiliary variables to decompose the objective function, and obtain the calculation formulas of the projection matrix of different modes, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code, and the calculation formula of the auxiliary variables; The training module is used to iteratively update the projection matrix, cluster center of projection space, cluster indicator matrix, mapping matrix, discrete hash code and auxiliary variables of each modal training sample based on the multimodal training samples and label matrix, as well as the calculation formula of the projection matrix of each modality, the calculation formula of the cluster center of projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables; the update of each variable in the iterative process depends on the current optimal solution of other variables; the hash function of each modality is determined based on the discrete hash code and multimodal training samples obtained in the last iteration; the hash function is used to perform similarity retrieval on the query sample in the multimodal data.
[0095] For the specific limitations of the asymmetric cross-modal hash learning device for potential structural information, please refer to the limitations of the asymmetric cross-modal hash learning method for potential structural information above, which will not be repeated here. The various modules in the above-mentioned asymmetric cross-modal hash learning device for potential structural information can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0096] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 The proposed asymmetric cross-modal hashing learning method for latent structural information.
[0097] The present invention also provides Figure 8 The structural diagram of the computer equipment shown in FIG. Figure 8 As shown in the figure, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The proposed asymmetric cross-modal hashing learning method for latent structural information.
[0098] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0099] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. An asymmetric cross-modal hashing learning method for latent structural information, characterized by: include: Establish the objective function; The objective function aims to jointly explore and analyze the multimodal sample clustering structure, sample similarity information, and semantic similarity information; Auxiliary variables are introduced to separate the variables of the objective function, and the calculation formulas of the projection matrix of different modes, the calculation formula of the cluster center of the projection space, the calculation formula of the cluster indicator matrix, the calculation formula of the mapping matrix, the calculation formula of the discrete hash code and the calculation formula of the auxiliary variables are obtained; Based on the multimodal training samples and label matrix, as well as the calculation formulas for the projection matrix of each modality, the calculation formula for the cluster center of the projection space, the calculation formula for the cluster indicator matrix, the calculation formula for the mapping matrix, the calculation formula for the discrete hash code, and the calculation formula for the auxiliary variables, the projection matrix, the cluster center of the projection space, the cluster indicator matrix, the mapping matrix, the discrete hash code, and the auxiliary variables of each modality training sample are iteratively updated; the update of each variable in the iterative process depends on the current optimal solution of other variables; Determine the hash function of each modality based on the discrete hash code and multimodal training samples obtained in the last iteration; The hash function is used to perform similarity retrieval on query samples in multimodal data.
2. The method according to claim 1, characterized in that The objective function is: ; in, represents the number of modes, Indicates the The training samples of each modality, Indicates the The sample projection matrix of the modalities, represents the cluster indicator matrix, Indicates the The cluster centers of the projection space of the modes, represents a discrete hash code, represents the mapping matrix, The Laplacian matrix representing the consistent similarity graph of multiple modalities, 、 and represents the balance parameter, S represents the pairwise semantic similarity matrix, , represents the identity matrix, Y represents the label matrix corresponding to the multimodal training samples, is used The norm is obtained by normalizing each row vector of Y.
3. The method according to claim 2, characterized in that No. The projection matrix of the modal The calculation formula is: ; in, is a diagonal matrix, and the diagonal elements are calculated as ; No. The cluster centers of the projected space of the modes The calculation formula is: ; in, and It is through The singular value decomposition of The calculation formula of the cluster indicator matrix V is: ; in, , the optimal solution is given by Given, where and Depend on The compact singular value decomposition of is obtained; , µ is the relaxation parameter that ensures A is positive definite, , δ >0 is the penalty factor, V=Z and Z>0; Mapping Matrix The calculation formula is: ; in, , the optimal solution is given by Given, and Depend on The compact singular value decomposition of is obtained; The calculation formula of discrete hash code B is: ; in, is the element-wise sign function; Auxiliary variables The calculation formula is: 。 4. The method according to claim 3, characterized in that Before iteratively updating the projection matrix of each modality, the cluster center of the projection space, the cluster indicator matrix, the mapping matrix, the discrete hash code, and the auxiliary variables, the method includes: Based on the multimodal training samples, a consistent similarity graph of multiple modalities is constructed; Based on the consistent similarity graphs of multiple modalities, the cluster indicator matrix and discrete hash code are randomly initialized using the standard normal distribution, and the diagonal matrix is initialized using the identity matrix. .
5. The method according to claim 3, characterized in that No. q A modal hash function for: ; in, is the regularization parameter.
6. The method according to claim 1, wherein The method further comprises: Determine a hash code of the query sample based on a hash function of the query sample and the modality corresponding to the query sample; Similarity retrieval in multimodal data via hash codes.
7. The method according to claim 6, characterized in that Query the hash code of the sample The calculation formula is: ; in, Indicates the q A sample query of the modality.
Citation Information
Cited By
Medical image-text hashing retrieval method based on nonlinear projection and prototype-like alignment
CN122470774A
Medical image-text hashing retrieval method based on nonlinear projection and prototype-like alignment
CN122470774B