Unsupervised cross-modal hash retrieval method, system and equipment based on implicit features
By constructing multimodal feature matrix and pseudo-label matrix, the hash code and projection matrix are optimized, and the problem of insufficient utilization of supervision information in cross-modal retrieval is solved, and the search accuracy is improved. It is suitable for unsupervised cross-modal hash retrieval.
Patent Information
- Application Number
- CN202510864919.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing cross-modal hash retrieval method has insufficient utilization of supervision information, resulting in low retrieval accuracy, especially when the label missing rate is high, the search accuracy is significantly reduced, and the hierarchy and semantic representation of semantic labels are not fully mined.
By extracting modal features from the original data of different modalities, building a multimodal feature matrix, performing pair-by-two typical correlation analysis, building a common latent representation matrix, and generating a target hash code matrix and projection matrix through pseudo-label matrix and alternating iteration optimization, making full use of pseudo-label information for cross-modal hash retrieval.
The accuracy of cross-modal retrieval is improved, the problem of insufficient utilization of supervision information in traditional methods is solved, and efficient cross-modal data retrieval is achieved in the case of no labels or missing labels.
Smart Images

Figure CN120371996A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cross-modal retrieval, and in particular, to an unsupervised cross-modal hashing retrieval method, system and device based on implicit features. Background Art
[0002] With the deep integration of the Internet, multimedia technology and artificial intelligence, the data form is undergoing a paradigm shift from single-modal to multi-modal collaborative expression of images, texts, audios and videos. The exponential growth of such multi-modal data poses dual challenges to cross-modal retrieval systems: on the one hand, it is necessary to solve the semantic gap problem between heterogeneous modalities such as text retrieval of images and video understanding of audios, and on the other hand, it is necessary to meet the engineering requirements of real-time retrieval of data in the order of terabytes.
[0003] Traditional multi-modal hashing methods achieve cross-modal retrieval by aligning feature spaces or constructing shared semantic representations in low-dimensional spaces, which have greater advantages in training efficiency and retrieval speed. These methods are based on the modeling method of handcrafted features and highly rely on supervised information. However, there are "supervision-semantic double gaps" in the utilization of supervised information by these methods: on the one hand, they highly rely on labeled data, and the retrieval accuracy will decrease significantly when the label missing rate increases. On the other hand, the hierarchical structure and semantic representation of semantic labels are not fully mined, and the generated hash codes do not make full use of the effective semantic information of the labels.
[0004] Therefore, the existing cross-modal hashing retrieval methods have relatively low accuracy for cross-modal retrieval. Summary of the Invention
[0005] The present application aims to propose an unsupervised cross-modal hashing retrieval method, system and device based on implicit features, which can improve the accuracy of cross-modal retrieval.
[0006] In a first aspect, an embodiment of the present application provides an unsupervised cross-modal hashing retrieval method based on implicit features, and the method includes: Extract modal features from raw data of different modalities to construct a multi-modal feature matrix, where the raw data of different modalities includes text data, image data, video data, and audio data; Perform pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain multiple canonical variable matrices of different modalities, and construct a common latent representation matrix according to the multiple canonical variable matrices of different modalities; Cluster the features in the common latent representation matrix to obtain multiple cluster centers, and calculate the distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; Calculate the pseudo-label value for each sample according to the distance matrix, and construct a pseudo-label matrix according to the pseudo-label value of each sample; Decompose the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtain a target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix; Initialize a consensus representation matrix according to each modality feature matrix in the multi-modal feature matrix, and determine the projection matrix of each modality feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modality feature matrix; Binarize the consensus representation matrix into a hash code matrix, and construct a target loss function according to the target sample latent feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; Determine a target hash code matrix and a target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on the target modality data according to the target hash code matrix and the target projection matrix.
[0007] Compared with the prior art, the first aspect of the present application has the following beneficial effects: This method extracts modal features from raw data of different modalities to construct a multi-modal feature matrix. The raw data of different modalities includes text data, image data, video data, and audio data. Perform pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain multiple canonical variable matrices of different modalities, and construct a common latent representation matrix based on the multiple canonical variable matrices of different modalities. Cluster the features in the common latent representation matrix to obtain multiple cluster centers, and calculate the distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix. According to the distance matrix, calculate the pseudo-label value of each sample, and construct a pseudo-label matrix based on the pseudo-label value of each sample. Decompose the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtain the target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix. Initialize a consensus representation matrix according to each modal feature matrix in the multi-modal feature matrix, and determine the projection matrices of each modal feature matrix and the target sample latent feature matrix respectively based on the target sample latent feature matrix, the consensus representation matrix, and each modal feature matrix. Binarize the consensus representation matrix into a hash code matrix, and construct a target loss function based on the target sample latent feature matrix, each modal feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix. Determine the target hash code matrix and the target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on the target modal data based on the target hash code matrix and the target projection matrix. In this way, by calculating the pseudo-label value of each sample to construct a pseudo-label matrix, and later constructing a target loss function based on the pseudo-label matrix to obtain an optimized hash code matrix (i.e., the target hash code matrix) and an optimized projection matrix (i.e., the target projection matrix), cross-modal hashing retrieval of the target modal data is realized, that is, by generating a pseudo-label matrix and making full use of pseudo-label information to deeply mine the semantic consistency between various modal data, the problem of the "supervision-semantic double gap" existing in traditional multi-modal hashing when using supervision information is solved, thereby improving the accuracy of cross-modal retrieval.
[0008] In some embodiments, calculating the pseudo-label value of each sample according to the distance matrix includes: ; where represents the pseudo-label value of each sample, represents the distance matrix in the th row and represents the distance matrix in the th Elements of the column Indicates the total number of cluster centers Indicates the hyperparameter for controlling the distance decay rate
[0009] In some embodiments, obtaining the target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix includes: Obtain the score of each sample, and construct an iterative optimization loss function according to the score of each sample, the pseudo-label latent feature matrix, and the sample latent feature matrix Fix the pseudo-label latent feature matrix and optimize the sample latent feature matrix Fix the sample latent feature matrix and optimize the pseudo-label latent feature matrix Until the iterative optimization loss function converges after alternate iterative optimization to obtain the target sample latent feature matrix
[0010] In some embodiments, determining the projection matrix of each modality feature matrix and the target sample latent feature matrix according to the target sample latent feature matrix, the consensus representation matrix, and each modality feature matrix includes: Determine the projection matrix of each modality feature matrix according to the consensus representation matrix and each modality feature matrix: ; Determine the projection matrix of the target sample latent feature matrix according to the consensus representation matrix and the target sample latent feature matrix: ; Wherein Represents the projection matrix of the th modality feature matrix And Represent hyperparameters Represents the consensus representation matrix Represents the th modality feature matrix Represents the transpose Represents the identity matrix Represents the projection matrix of the target sample latent feature matrix Represents the target sample latent feature matrix
[0011] In some embodiments, constructing the target loss function according to the target sample latent feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix includes: ; ; Among them, represents the projection matrix of the -th modal feature matrix, represents the projection matrix of the target sample latent feature matrix, represents the consensus representation matrix, represents the hash code matrix, represents the number of modalities, is used to balance the contribution of each feature matrix in constructing the consensus representation matrix, represents the -th modal feature matrix, represents the F2 norm, represents the target sample latent feature matrix, represents the pseudo-label matrix, represents the number of bits of the hash code, represents the number of samples, represents the consensus representation matrix and -dimensional all vector product, represents -dimensional vector, represents the identity matrix, , , and represent hyperparameters.
[0012] In some embodiments, the determining the target hash code matrix and the target projection matrix according to the target loss function includes: When optimizing the projection matrix, fix other parameters in the target loss function, and extract the functional formula related to the projection matrix from the target loss function to construct a first optimization function; When optimizing the consensus representation matrix, fix other parameters in the target loss function, and extract the functional formula related to the consensus representation matrix from the target loss function to construct a second optimization function; When optimizing the hash code matrix, fix other parameters in the target loss function, and extract the functional formula related to the hash code matrix from the target loss function to construct a third optimization function; By minimizing the first optimization function, the second optimization function, and the third optimization function, minimize the target loss function to determine the target hash code matrix and the target projection matrix.
[0013] In some embodiments, the cross-modal hashing retrieval of the target modal data according to the target hash code matrix and the target projection matrix includes: Obtain a target projection matrix corresponding to the target modal data; Use the target hash code matrix as a hash code database; Extract the target feature vector of the target modal data; Calculate the hash code of the target modal data according to the target projection matrix corresponding to the target modal data and the target feature vector; Calculate the Hamming distance between the hash code of the target modal data and the hash codes of all data in the hash code database, and perform cross-modal hash retrieval on the target modal data according to the Hamming distance.
[0014] In a second aspect, an embodiment of the present application further provides an unsupervised cross-modal hash retrieval system based on implicit features. The system includes: A feature matrix construction unit, configured to extract modal features from raw data of different modalities to construct a multi-modal feature matrix. The raw data of different modalities includes text data, image data, video data, and audio data; A representation matrix construction unit, configured to perform pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain multiple canonical variable matrices of different modalities, and construct a common latent representation matrix according to the multiple canonical variable matrices of different modalities; A distance matrix construction unit, configured to cluster the features in the common latent representation matrix to obtain multiple cluster centers, and calculate the distance between each cluster center and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; A pseudo-label matrix construction unit, configured to calculate the pseudo-label value of each sample according to the distance matrix, and construct a pseudo-label matrix according to the pseudo-label value of each sample; An alternating iteration optimization unit, configured to decompose the pseudo-label matrix into the product of a pseudo-label implicit feature matrix and a sample implicit feature matrix, and obtain a target sample implicit feature matrix by alternately iteratively optimizing the pseudo-label implicit feature matrix and the sample implicit feature matrix; A projection matrix determination unit, configured to initialize a consensus representation matrix according to each modal feature matrix in the multi-modal feature matrix, and determine the projection matrix of each modal feature matrix and the target sample implicit feature matrix respectively according to the target sample implicit feature matrix, the consensus representation matrix, and each modal feature matrix; A loss function construction unit, configured to binarize the consensus representation matrix into a hash code matrix, and construct a target loss function according to the target sample latent feature matrix, each type of modal feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; A cross-modal hashing retrieval unit, configured to determine a target hash code matrix and a target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on target modal data according to the target hash code matrix and the target projection matrix.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute an unsupervised cross-modal hashing retrieval method based on latent features as described above.
[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions for enabling a computer to execute an unsupervised cross-modal hashing retrieval method based on latent features as described above.
[0017] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the above first aspect, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and / or additional aspects and advantages of the present application will become apparent and easier to understand from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of an unsupervised cross-modal hashing retrieval method based on latent features provided by the present application; Figure 2 is a schematic overall framework diagram of the best embodiment of an unsupervised cross-modal hashing retrieval method based on latent features provided by the present application; Figure 3 is a schematic structural diagram of an embodiment of an unsupervised cross-modal hashing retrieval system based on latent features provided by the present application; Figure 4 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.
[0020] In the description of the present application, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0021] In the description of the present application, it should be understood that for the orientation description, such as up, down, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application.
[0022] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0023] Traditional multi-modal hashing methods achieve cross-modal retrieval by aligning feature spaces or constructing shared semantic representations in a low-dimensional space, and have greater advantages in training efficiency and retrieval speed. These methods are based on the modeling method of handcrafted features and highly rely on supervised information. However, there is a "supervision-semantic double gap" in the utilization of supervised information by these methods: on the one hand, they highly rely on labeled data, and the retrieval accuracy will significantly decrease when the label missing rate increases. On the other hand, the hierarchical structure and semantic representation of semantic labels are not fully mined, and the generated hash codes do not make full use of the effective semantic information of the labels. Therefore, the existing cross-modal hashing retrieval methods have relatively low accuracy for cross-modal retrieval.
[0024] To solve the problem that the existing cross-modal hashing retrieval methods have relatively low accuracy for cross-modal retrieval, the present application proposes an unsupervised cross-modal hashing retrieval method, system and device based on implicit features.
[0025] Refer to Figure 1 , the flowchart of the unsupervised cross-modal hashing retrieval method based on implicit features provided by the embodiments of the present application. The unsupervised cross-modal hashing retrieval method based on implicit features is applied to an electronic device, and the electronic device can be a server or a mobile terminal, etc. As Figure 1 shown, the unsupervised cross-modal hashing retrieval method based on implicit features may include the following steps: Step S100: Extract modal features from raw data of different modalities to construct a multi-modal feature matrix. The raw data of different modalities includes text data, image data, video data, and audio data. Step S200: Perform pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain multiple canonical variable matrices of different modalities, and construct a common latent representation matrix based on the multiple canonical variable matrices of different modalities. Step S300: Cluster the features in the common latent representation matrix to obtain multiple cluster centers, and calculate the distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix. Step S400: Calculate the pseudo-label values of each sample according to the distance matrix, and construct a pseudo-label matrix according to the pseudo-label values of each sample. Step S500: Decompose the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtain the target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix. Step S600: Initialize a consensus representation matrix according to each modal feature matrix in the multi-modal feature matrix, and determine the projection matrices of each modal feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modal feature matrix. Step S700: Binarize the consensus representation matrix into a hash code matrix, and construct a target loss function according to the target sample latent feature matrix, each modal feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix. Step S800: Determine the target hash code matrix and the target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on the target modal data according to the target hash code matrix and the target projection matrix.
[0026] In this embodiment, by extracting modal features from raw data of different modalities to construct a multi-modal feature matrix, the raw data of different modalities includes text data, image data, video data, and audio data; performing pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain a variety of canonical variable matrices of different modalities, and constructing a common latent representation matrix according to the variety of canonical variable matrices of different modalities; clustering the features in the common latent representation matrix to obtain multiple cluster centers, and calculating the distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; calculating the pseudo-label value of each sample according to the distance matrix, and constructing a pseudo-label matrix according to the pseudo-label value of each sample; decomposing the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtaining a target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix; initializing a consensus representation matrix according to each modal feature matrix in the multi-modal feature matrix, and determining the projection matrices of each modal feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modal feature matrix; binarizing the consensus representation matrix into a hash code matrix, and constructing a target loss function according to the target sample latent feature matrix, each modal feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; determining a target hash code matrix and a target projection matrix according to the target loss function, and performing cross-modal hashing retrieval on the target modal data according to the target hash code matrix and the target projection matrix. In this way, by calculating the pseudo-label value of each sample to construct a pseudo-label matrix, and later constructing a target loss function according to the pseudo-label matrix to obtain an optimized hash code matrix (i.e., the target hash code matrix) and an optimized projection matrix (i.e., the target projection matrix), so as to realize cross-modal hashing retrieval of the target modal data, that is, by generating a pseudo-label matrix and making full use of pseudo-label information to deeply mine the semantic consistency between various modal data, the problem of the "supervision-semantic double gap" existing in traditional multi-modal hashing when using supervision information is solved, thereby improving the accuracy of cross-modal retrieval.
[0027] In some embodiments, calculating the pseudo-label value of each sample according to the distance matrix includes: ; where represents the pseudo-label value of each sample, represents the distance matrix in the th row and represents the distance matrix in the th The elements of the column Indicates the total number of cluster centers Indicates the hyperparameter used to control the distance decay rate
[0028] In some embodiments, by alternately iteratively optimizing the pseudo-label implicit feature matrix and the sample implicit feature matrix, a target sample implicit feature matrix is obtained, including: Obtain the score of each sample, and construct an iterative optimization loss function according to the score of each sample, the pseudo-label implicit feature matrix, and the sample implicit feature matrix Fix the pseudo-label implicit feature matrix and optimize the sample implicit feature matrix Fix the sample implicit feature matrix and optimize the pseudo-label implicit feature matrix Until the iterative optimization converges to the iterative optimization loss function, a target sample implicit feature matrix is obtained
[0029] In this embodiment, by obtaining the score of each sample, constructing an iterative optimization loss function according to the score of each sample, the pseudo-label implicit feature matrix, and the sample implicit feature matrix; fixing the pseudo-label implicit feature matrix and optimizing the sample implicit feature matrix; fixing the sample implicit feature matrix and optimizing the pseudo-label implicit feature matrix; until the iterative optimization converges to the iterative optimization loss function, a target sample implicit feature matrix is obtained. In this way, by alternately iteratively optimizing to obtain the target sample implicit feature matrix, a relatively accurate target sample implicit feature matrix can be obtained, laying a good data foundation for guiding the learning of the consensus representation and the generation of the hash code matrix in the later stage
[0030] In some embodiments, according to the target sample implicit feature matrix, the consensus representation matrix, and each modal feature matrix, the projection matrix of each modal feature matrix and the target sample implicit feature matrix is determined, including: According to the consensus representation matrix and each modal feature matrix, determine the projection matrix of each modal feature matrix ; According to the target sample implicit feature matrix and the consensus representation matrix, determine the projection matrix of the target sample implicit feature matrix ; Wherein Represents the projection matrix of the th modal feature matrix And Represent hyperparameters Represents the consensus representation matrix Represents the th modal feature matrix Represents the transpose Represents the identity matrix The projection matrix representing the projection of the target sample implicit feature matrix, represents the target sample implicit feature matrix.
[0031] In some embodiments, according to the target sample implicit feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix, a target loss function is constructed, including: ; ; wherein, represents the th projection matrix of the modality feature matrix, represents the projection matrix of the target sample implicit feature matrix, represents the consensus representation matrix, represents the hash code matrix, represents the number of modalities, represents a parameter used to balance the contribution of each feature matrix in constructing the consensus representation matrix, represents the th modality feature matrix, represents the F2 norm, represents the target sample implicit feature matrix, represents the pseudo-label matrix, represents the number of bits of the hash code, represents the number of samples, represents the consensus representation matrix and dimensional all vector product, represents dimensional vector, represents the identity matrix, , , and represent hyperparameters.
[0032] In this embodiment, according to the target sample implicit feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix, a target loss function is constructed. In this way, by comprehensively constructing the target loss function with multiple parameters and optimizing the target loss function later, more accurate results can be obtained.
[0033] In some embodiments, according to the target loss function, a target hash code matrix and a target projection matrix are determined, including: When optimizing the projection matrix, fix the other parameters in the target loss function, and extract the functional formula related to the projection matrix from the target loss function to construct the first optimization function; When optimizing the consensus representation matrix, fix other parameters in the objective loss function, extract the functional related to the consensus representation matrix from the objective loss function, and construct a second optimization function; When optimizing the hash code matrix, fix other parameters in the objective loss function, extract the functional related to the hash code matrix from the objective loss function, and construct a third optimization function; By minimizing the first optimization function, the second optimization function, and the third optimization function, minimize the objective loss function to determine the target hash code matrix and the target projection matrix.
[0034] In this embodiment, by minimizing the first optimization function, the second optimization function, and the third optimization function, the objective loss function is minimized to determine the target hash code matrix and the target projection matrix. In this way, the obtained target hash code matrix and target projection matrix are more accurate, thus laying a good data foundation for later cross-modal data retrieval. And this embodiment does not use label information, but deeply mines the semantic consistency between various modal data by generating a pseudo-label matrix and making full use of pseudo-label information, solving the "supervision-semantic double gap" problem existing in traditional multi-modal hashing when using supervision information. This embodiment is applicable to the efficient retrieval of large-scale multi-modal data.
[0035] In some embodiments, according to the target hash code matrix and the target projection matrix, perform cross-modal hashing retrieval on the target modal data, including: Obtain the target projection matrix corresponding to the target modal data; Use the target hash code matrix as the hash code database; Extract the target feature vector of the target modal data; According to the target projection matrix and the target feature vector corresponding to the target modal data, calculate the hash code of the target modal data; Calculate the Hamming distance between the hash code of the target modal data and the hash codes of all data in the hash code database, and perform cross-modal hashing retrieval on the target modal data according to the Hamming distance.
[0036] In this embodiment, by using the obtained accurate target hash code matrix and target projection matrix to perform cross-modal hashing retrieval on the target modal data, the accuracy of cross-modal retrieval can be improved.
[0037] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: With the deep integration of the Internet, multimedia technology, and artificial intelligence, the data form is undergoing a paradigm shift from a single modality to a multi-modal collaborative expression of images, text, audio, and video. This exponential growth of multi-modal data poses a dual challenge to cross-modal retrieval systems: on the one hand, it is necessary to address the semantic gap problem between heterogeneous modalities such as text retrieving images and video understanding audio, and on the other hand, it is necessary to meet the engineering requirements of real-time retrieval of data in the terabyte scale.
[0038] Due to the great advantages of hash technology in quickly querying and efficiently storing large amounts of multi-modal data, it provides a feasible solution for the application of large-scale cross-modal retrieval. The core of cross-modal hashing lies in constructing a unified Hamming space mapping model to measure the comparability of different modality data through binary hash coding. Many cross-modal hashing methods have been proposed, and their main idea is to project cross-modal data into the Hamming space of binary hash codes while preserving the inherent semantic similarity.
[0039] In recent years, while achieving performance breakthroughs, the cross-modal hashing methods based on deep learning also face severe engineering implementation bottlenecks in terms of the huge time and computing power consumed. This computationally intensive characteristic leads to two major challenges in industrial-level and terabyte-scale data scenarios: 1) The gradient synchronization overhead in distributed training increases exponentially; 2) It is difficult to meet the real-time response in edge computing scenarios. Therefore, this method is restricted in the use of ultra-large datasets and tasks with high real-time requirements.
[0040] Traditional multi-modal hashing methods achieve cross-modal retrieval by aligning feature spaces or constructing shared semantic representations in a low-dimensional space, and have greater advantages in training efficiency and retrieval speed. These methods are based on the modeling method of handcrafted features and highly rely on supervision information. However, there are "supervision-semantic double gaps" in the utilization of this supervision information: on the one hand, they highly rely on labeled data, and the retrieval accuracy will decrease significantly when the label missing rate increases. On the other hand, the hierarchical structure and semantic representation of semantic labels are not fully mined, and the generated hash codes do not make full use of the effective semantic information of the labels, which will also lead to a decrease in retrieval accuracy.
[0041] This embodiment relates to an unsupervised cross-modal hashing learning method based on generating sample implicit features with pseudo-labels, which is applied to the field of multi-modal data retrieval. It aims to solve the problem of insufficient retrieval accuracy in multi-modal data retrieval due to the lack of label information or insufficient utilization of label information. This method is applicable to multi-modal data retrieval with paired unlabeled information. Refer to Figure 2 , the technical solution of this embodiment specifically includes the following steps: Step S1. Extract the multimodal feature matrix. There are significant differences in the structure and form of the original data of different modalities. Through feature extraction, this heterogeneous data can be converted into vectors that are convenient for unified processing, so as to eliminate the form differences and provide a basis for subsequent cross-modal alignment.
[0042] Specifically, in the multimodal retrieval task, the original data forms of different modalities vary greatly (for example, text is discrete symbols and images are pixel matrices), and direct comparison or fusion is almost impossible. To solve the heterogeneity problem of different-modal data, it is necessary to first extract the multimodal feature matrix so that it can be effectively compared and calculated in a unified space. For data of different modalities, an appropriate feature extraction model needs to be selected according to the modal features and task requirements. The multimodal hashing retrieval method in this embodiment has greater advantages in training efficiency and retrieval speed compared with the traditional multimodal hashing retrieval method. Therefore, a traditional feature extraction method with higher computational efficiency is preferably used. Next, taking the extraction of text features by the bag-of-words model and the extraction of image features by the convolutional neural network as examples: 1. Extract text features by the bag-of-words model.
[0043] The bag-of-words model quickly captures the global keyword information of the text by counting word frequencies and is suitable for efficiently processing high-dimensional sparse text data. Specifically: A. Segment the original text, remove stop words, and convert it into a standardized word unit; B. Extract unique words from all documents, construct a vocabulary, and assign a unique index to each word; C. For each document, mark whether the word appears in the document (1 or 0), so as to map it into a vector with the length of the vocabulary.
[0044] 2. Extract image features by the convolutional neural network.
[0045] The convolutional neural network automatically learns local features through convolutional layers and abstracts high-level semantic information layer by layer. Specifically: A. Use methods such as size normalization and channel standardization to convert the original image into a standardized format suitable for model processing; B. Use convolutional kernels to extract local spatial features, and then introduce non-linearity through non-linear activation functions to enhance the model's expression ability; C. Through the pooling layer, reduce the spatial dimension of the feature map to reduce the computational amount; D. Repeat the stacking of "convolution-activation-pooling" modules to extract more complex features layer by layer.
[0046] Step S2. Learning common latent representations. Apply canonical correlation analysis to the multimodal feature matrix to find the weight matrix and canonical variable matrix of each modality, so as to maximize the correlation between the canonical variables after various modal transformations, and then obtain the common latent representation through the canonical variable matrix.
[0047] Specifically, canonical correlation analysis is applied to the multimodal feature matrix to find the weight matrix of each mode and the canonical variate matrix , so that the correlation between the typical variables after various modal transformations is maximized. If there are only two modal data, canonical correlation analysis can be directly applied; for more than two modal data, multiple groups of data are sequentially subjected to pairwise canonical correlation analysis, and the weight matrix is gradually optimized. Taking two modes as an example, the canonical variable matrices obtained from their characteristic matrices are recorded as and , only the dimensional features with cumulative correlation coefficients exceeding 10% are retained, still recorded as and , the asymmetric fusion of the two is the common potential representation matrix : ; ; in, Represents the weight coefficient.
[0048] Step S3. Construct a pseudo label matrix. Apply a clustering method to the features in the common latent representation matrix to obtain several cluster centers. Treat each cluster center as a pseudo label, and treat the similarity between each sample in the common latent representation matrix and the cluster center as the value of the sample under this pseudo label. Calculate the similarity between the cluster center matrix and all column vector pairs of the common latent representation matrix to obtain a pseudo label matrix.
[0049] Specifically, for the common latent representation matrix The features in the clustering method are clustered, such as the k-means clustering method. The clustering method in this embodiment can adopt a clustering method known to those skilled in the art, and this embodiment does not specifically limit it. Center points (cluster centers), denoted as , treating each cluster center as a pseudo label. Calculate the Cluster centers and No. samples (i.e., the common latent representation matrix The The square of the distance between columns: ; True label values often have marks for each sample under only a few labels, and no marks under other labels. To make the pseudo-labels closer to the nature of the true labels and also to enhance the guiding role of the nearest neighbor pseudo-labels, the distance matrix (distance matrix contains multiple elements) is sparsified, that is, only the elements with the smallest values in each column are retained, and the remaining elements are all set to . Finally, according to the distance matrix the pseudo-label matrix (distance matrix contains multiple elements) is calculated, and to balance the influence of the pseudo-labels on each sample, the pseudo-label values of each sample are normalized: ; where, and are the elements in the th row and th column and the elements in the th row and th column of the distance matrix respectively, is a hyperparameter used to control the distance attenuation rate.
[0050] Step S4. Obtaining the sample latent features. The pseudo-label matrix can be regarded as the product of the pseudo-label factor matrix and the sample factor matrix. Using the alternating least squares method, that is, the method of alternately optimizing these two product matrices, the optimal pseudo-label factor matrix and sample factor matrix are obtained, and the sample factor matrix is regarded as the sample latent feature matrix and applied to subsequent consensus representation learning.
[0051] Specifically, the pseudo-label matrix can be regarded as the product of the pseudo-label factor matrix (i.e., the pseudo-label latent feature matrix) and the sample factor matrix (i.e., the sample latent feature matrix), that is: ; where, , , ( ), represents the number of pseudo-labels (i.e., the number of clustering centers), represents the number of samples in a certain modality, represents the latent factor dimension. Using the alternating least squares method, the problem is decomposed by alternately fixing one matrix and optimizing the other matrix: Fix the pseudo-label factor matrix , and optimize the sample factor matrix : For each sample , only consider the set of pseudo-labels that have scores (pseudo-label values) for this sample . Define , the factor matrix of the pseudo-labels that have scores for the sample ; , that is, the known score vector of the sample . The part of the loss function with respect to is: ; The closed-form solution is: ; where represents the hyperparameter that controls the strength of the regularization term, represents the identity matrix of dimension represents the transpose, represents the factor matrix of the samples scored by the pseudo-label (i.e., the submatrix of ), represents the factor matrix of the pseudo-labels that score the sample (i.e., the submatrix of ).
[0052] Fix the sample factor matrix , and optimize the pseudo-label factor matrix : Similarly, we can get: ; Use the alternating least squares method to alternate and iterate until convergence.
[0053] Regard the optimized sample factor matrix as the sample latent feature matrix and apply it to subsequent consensus representation learning. Step S5. Learn the consensus representation. Use the feature matrix and the sample latent feature matrix to learn the consensus representation matrix and the corresponding projection matrix. Specifically, the feature matrix and the sample latent feature matrix are multiplied by the corresponding projection matrices, and the results should be as close as possible to the consensus representation matrix. To balance the contributions in constructing the consensus representation matrix, the loss function of each feature matrix and the consensus representation matrix will be multiplied by the corresponding hyperparameter.
[0054] Specifically, the consensus representation matrix and the corresponding projection matrix are learned using each modal feature matrix and the sample latent feature matrix. Specifically, the feature matrix and the sample latent feature matrix are multiplied by the corresponding projection matrix, and the result should be as close as possible to the consensus representation matrix. Among them, a modal feature matrix obtains an initialized consensus representation matrix, and then the projection matrix corresponding to the sample latent feature matrix is calculated based on the sample latent feature matrix and the consensus representation matrix, and the projection matrix corresponding to each modal feature matrix is calculated based on each modal feature matrix and the consensus representation matrix, and then the projection matrix and the consensus representation matrix are gradually iteratively optimized in step S8.
[0055] Denote as the consensus representation matrix, denote as the projection matrix of the th modal feature matrix, denote as the projection matrix of the sample latent feature matrix, where is the number of modalities, and the optimization objective can be written as the formula: ; is used to balance the contribution of each feature matrix in constructing the consensus representation matrix, and the regularization terms of are to limit the complexity of the projection matrix and improve the generalization of the model,
[0056] Step S6. Learning the hash code. Intuitively, the hash code matrix is the ultimate learning goal of each sample feature matrix, while the consensus representation matrix is a continuous value estimate of the hash code matrix. Therefore, the simplest approach is to directly binarize the consensus representation matrix to obtain the hash code matrix. However, considering that this will cause a certain loss of precision, the norm of the difference between the two is added to the total loss function.
[0057] Specifically, intuitively, the hash code matrix is the ultimate learning goal of each sample feature matrix, while the consensus representation matrix is a continuous value estimate of the hash code matrix. Therefore, the simplest approach is to directly binarize the consensus representation matrix to obtain the hash code matrix, that is: ; where represents the sign function, which outputs when the input is positive, and otherwise.
[0058] However, considering that this will cause a certain loss of precision, the norm of the difference between the two is optimized instead: .
[0059] Step S7. Utilization of pseudo-label information. To utilize the label information of the pseudo-label matrix to guide the generation of the consensus representation matrix and the hash code matrix, the product of the consensus representation matrix and the hash code matrix should be as close as possible to (the number of bits of the hash code) times the inner product of the pseudo-label matrix.
[0060] Specifically, in Step S5, this embodiment uses the sample implicit feature matrix decomposed from the pseudo-label matrix to guide the learning of the consensus representation, which is one aspect of the utilization of pseudo-label information. On the other hand, the label information of the pseudo-label matrix is directly used to guide the generation of the consensus representation matrix and the hash code matrix. If , where represents the number of bits of the hash code, the optimization objective is written as the formula: ; Intuitively, if is larger, it indicates that the similarity between the th sample and the th sample is higher, and the inner product of the th column sample of the consensus representation matrix and the th column sample of the hash code matrix should be larger.
[0061] Step S8. Optimization of the main loss function and each variable.
[0062] Combining steps S5 to S7, the overall optimization objective, i.e., the main loss function, is as follows: ; ; where, are all hyperparameters. From the perspective of optimizing each variable, the methods of optimizing and are the same. Denote as , is denoted as . The above formula can be simplified as: ; ; where, represents the hyperparameter for balancing the contributions of each modality when the modality feature matrix is used to learn the consensus representation matrix .
[0063] Optimizing : Fix other variables and only consider the functional related to . The optimization problem is transformed into: ; Let the first derivative be 0, and we get: ; Solving it gives: ; When , .
[0064] Optimizing : Fix other variables and only consider the functional related to . The optimization problem is transformed into: ; ; It is transformed into the following form: ; ; This is an orthogonal Procrustes problem with zero-mean constraint. Denote , . To obtain the optimal , perform eigenvalue decomposition on to get: ; Among them, is a diagonal matrix composed of positive eigenvalues, is the matrix 's rank. are the corresponding eigenvectors, contains the remaining eigenvectors corresponding to the eigenvalue 0, obtained through Schmidt orthogonalization . Denote , is a random orthogonal matrix. Finally, 's optimal solution is: ; Optimize : Fix other variables and only consider the functional related to . The optimization problem is transformed into: ; ; transformed into the following form: ; ; Binarization can be solved: ; Step S9. Retrieval of new data.
[0065] In step S8, the hash codes of the training data (as the hash code database) and the projection matrices of each modality have been obtained. For the new data (i.e., target modality data) that needs to be retrieved, use the following steps to retrieve relevant data in the database: Use the corresponding method provided in step S1 to extract the feature vector of the new data ; If the new data belongs to the th modality, obtain the corresponding projection matrix , and use to obtain the hash code of the new data; Calculate the Hamming distance between the hash code of the new data and the hash codes of all data in the database in sequence. For any two hash codes and , the Hamming distance calculation formula is: ; Among them, is the hash code length, is the indicator function. If then it is , otherwise it is .
[0066] Return the other modality data related to the Hamming distance of the new data in ascending order of the Hamming distance of the new data as the retrieval result.
[0067] Compared with the prior art, the technical solution of this embodiment has the following advantages: This embodiment is a type of unsupervised cross-modal hashing learning that only utilizes the pairwise correspondence relationship of multi-modal data without using label information. By generating a pseudo-label matrix and making full use of the pseudo-label information to deeply mine the semantic consistency between various modality data, it can improve the accuracy of cross-modal hashing retrieval and solve the "supervision-semantic double gap" problem existing in traditional multi-modal hashing when using supervision information, that is, it solves the problems that traditional methods highly rely on supervision information and do not fully mine supervision information. The technical solution of this embodiment is applicable to the efficient retrieval of large-scale multi-modal data.
[0068] Referring to Figure 3 , the embodiment of the present application also provides an unsupervised cross-modal hashing retrieval system based on latent features. The system includes a feature matrix construction unit 100, a representation matrix construction unit 200, a distance matrix construction unit 300, a pseudo-label matrix construction unit 400, an alternating iterative optimization unit 500, a projection matrix determination unit 600, a loss function construction unit 700, and a cross-modal hashing retrieval unit 800, where: The feature matrix construction unit 100 is used to extract modality features from the original data of different modalities and construct a multi-modal feature matrix. The original data of different modalities includes text data, image data, video data, and audio data; The representation matrix construction unit 200 is used to perform pairwise canonical correlation analysis on every two modality feature matrices in the multi-modal feature matrix to obtain multiple different modality canonical variable matrices, and construct a common latent representation matrix according to the multiple different modality canonical variable matrices; The distance matrix construction unit 300 is used to cluster the features in the common latent representation matrix to obtain multiple cluster centers, and calculate the distance between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; The pseudo-label matrix construction unit 400 is used to calculate the pseudo-label value of each sample according to the distance matrix, and construct a pseudo-label matrix according to the pseudo-label value of each sample; The alternating iterative optimization unit 500 is used to decompose the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtain the target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix; A projection matrix determination unit 600, configured to initialize a consensus representation matrix according to each modality feature matrix in the multi-modal feature matrix, and determine the projection matrices of each modality feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modality feature matrix; A loss function construction unit 700, configured to binarize the consensus representation matrix into a hash code matrix, and construct a target loss function according to the target sample latent feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; A cross-modal hashing retrieval unit 800, configured to determine a target hash code matrix and a target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on the target modality data according to the target hash code matrix and the target projection matrix.
[0069] It should be noted that since an unsupervised cross-modal hashing retrieval system based on latent features in this embodiment and the above-mentioned unsupervised cross-modal hashing retrieval method based on latent features are based on the same inventive concept, the corresponding content in the method embodiments is equally applicable to this system embodiment and will not be elaborated here.
[0070] Refer to Figure 4 , this application embodiment also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned unsupervised cross-modal hashing retrieval method based on latent features of the present disclosure.
[0071] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.
[0072] The electronic device of the embodiment of the present application will be introduced in detail below.
[0073] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the unsupervised cross-modal hashing retrieval method based on implicit features in the embodiments of the present disclosure.
[0074] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices. It can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between the various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0075] The embodiments of the present disclosure also provide a storage medium. This storage medium is a computer-readable storage medium, and this computer-readable storage medium stores computer-executable instructions. These computer-executable instructions are used to cause a computer to execute the above-mentioned unsupervised cross-modal hashing retrieval method based on implicit features.
[0076] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0077] The embodiments described in the embodiments of the present disclosure are to more clearly illustrate the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided in the embodiments of the present disclosure. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present disclosure are equally applicable to similar technical problems.
[0078] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0079] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0080] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0081] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0082] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0083] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0084] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0085] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0086] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs. The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the gist of the present application.
[0087] The above has described the embodiments of the present application in detail with reference to the drawings, but the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art to which the present application pertains, various changes can be made without departing from the gist of the present application.
Claims
1. An unsupervised cross-modal hashing retrieval method based on implicit features, characterized in that The method includes: extracting modal features from raw data of different modalities to construct a multi-modal feature matrix, where the raw data of different modalities includes text data, image data, video data, and audio data; performing pairwise canonical correlation analysis on every two modal feature matrices in the multi-modal feature matrix to obtain multiple canonical variable matrices of different modalities, and constructing a common latent representation matrix according to the multiple canonical variable matrices of different modalities; clustering the features in the common latent representation matrix to obtain multiple cluster centers, and calculating the distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; calculating the pseudo-label values of each sample according to the distance matrix, and constructing a pseudo-label matrix according to the pseudo-label values of each sample; decomposing the pseudo-label matrix into the product of a pseudo-label latent feature matrix and a sample latent feature matrix, and obtaining a target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix; initializing a consensus representation matrix according to each modal feature matrix in the multi-modal feature matrix, and determining the projection matrices of each modal feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modal feature matrix; binarizing the consensus representation matrix into a hash code matrix, and constructing a target loss function according to the target sample latent feature matrix, each modal feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; determining a target hash code matrix and a target projection matrix according to the target loss function, and performing cross-modal hashing retrieval on target modal data according to the target hash code matrix and the target projection matrix.
2. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, wherein The calculating the pseudo-label values of each sample according to the distance matrix includes: ; wherein, represents the pseudo-label value of each sample, represents the distance matrix the -th row and -th column element in it, represents the distance matrix the -th row and -th column element in it, represents the total number of clustering centers, represents the hyperparameter used to control the distance decay rate.
3. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, wherein The obtaining a target sample latent feature matrix by alternately iteratively optimizing the pseudo-label latent feature matrix and the sample latent feature matrix includes: obtaining the score of each sample, and constructing an iterative optimization loss function according to the score of each sample, the pseudo-label latent feature matrix, and the sample latent feature matrix; fixing the pseudo-label latent feature matrix and optimizing the sample latent feature matrix; fixing the sample latent feature matrix and optimizing the pseudo-label latent feature matrix; until the iterative optimization loss function converges after alternate iterative optimization to obtain a target sample latent feature matrix.
4. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, wherein The determining the projection matrices of each modal feature matrix and the target sample latent feature matrix respectively according to the target sample latent feature matrix, the consensus representation matrix, and each modal feature matrix includes: determining the projection matrix of each modal feature matrix according to the consensus representation matrix and each modal feature matrix: ; determining the projection matrix of the target sample latent feature matrix according to the consensus representation matrix and the target sample latent feature matrix: ; Among them, represents the projection matrix of the th modal feature matrix, and represent hyperparameters, represents the consensus representation matrix, represents the th modal feature matrix, represents the transpose, represents the identity matrix, represents the projection matrix of the target sample latent feature matrix, represents the target sample latent feature matrix.
5. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, wherein Constructing a target loss function based on the target sample implicit feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix, includes: ; ; Among them, represents the projection matrix of the th modal feature matrix, represents the projection matrix of the target sample latent feature matrix, represents the consensus representation matrix, represents the hash code matrix, represents the number of modalities, is used to balance the contribution of each feature matrix in constructing the consensus representation matrix, represents the th modal feature matrix, represents the F2 norm, represents the target sample latent feature matrix, represents the pseudo-label matrix, represents the number of bits of the hash code, represents the number of samples, represents the consensus representation matrix and dimensional all vector product, represents dimensional vector, represents the identity matrix, , , and represent hyperparameters.
6. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, wherein Determining a target hash code matrix and a target projection matrix according to the target loss function, includes: When optimizing the projection matrix, fixing other parameters in the target loss function, and extracting a functional related to the projection matrix from the target loss function to construct a first optimization function; When optimizing the consensus representation matrix, fixing other parameters in the target loss function, and extracting a functional related to the consensus representation matrix from the target loss function to construct a second optimization function; When optimizing the hash code matrix, fixing other parameters in the target loss function, and extracting a functional related to the hash code matrix from the target loss function to construct a third optimization function; By minimizing the first optimization function, the second optimization function, and the third optimization function, minimizing the target loss function to determine the target hash code matrix and the target projection matrix.
7. The unsupervised cross-modal hashing retrieval method based on implicit features according to claim 1, characterized in that Performing cross-modal hashing retrieval on target modality data according to the target hash code matrix and the target projection matrix, includes: Obtaining a target projection matrix corresponding to the target modality data; Using the target hash code matrix as a hash code database; Extracting a target feature vector of the target modality data; Calculating a hash code of the target modality data according to the target projection matrix corresponding to the target modality data and the target feature vector; Calculating a Hamming distance between the hash code of the target modality data and the hash codes of all data in the hash code database, and performing cross-modal hashing retrieval on the target modality data according to the Hamming distance.
8. An unsupervised cross-modal hashing retrieval system based on implicit features, characterized in that, The system includes: A feature matrix construction unit, configured to extract modality features from raw data of different modalities to construct a multi-modal feature matrix, where the raw data of different modalities includes text data, image data, video data, and audio data; A representation matrix construction unit, configured to perform pairwise canonical correlation analysis on every two modality feature matrices in the multi-modal feature matrix to obtain multiple typical variable matrices of different modalities, and construct a common latent representation matrix according to the multiple typical variable matrices of different modalities; A distance matrix construction unit, configured to cluster features in the common latent representation matrix to obtain multiple cluster centers, and calculate distances between the cluster centers and each sample in the common latent representation matrix to construct a distance matrix, where the sample is a column in the common latent representation matrix; A pseudo-label matrix construction unit, configured to calculate pseudo-label values of each sample according to the distance matrix, and construct a pseudo-label matrix according to the pseudo-label values of each sample; An alternating iteration optimization unit, configured to decompose the pseudo-label matrix into a product of a pseudo-label implicit feature matrix and a sample implicit feature matrix, and obtain a target sample implicit feature matrix by alternately iteratively optimizing the pseudo-label implicit feature matrix and the sample implicit feature matrix; A projection matrix determination unit, configured to initialize a consensus representation matrix according to each modality feature matrix in the multi-modal feature matrix, and determine projection matrices for each modality feature matrix and the target sample latent feature matrix respectively according to each modality feature matrix, the target sample latent feature matrix, and the consensus representation matrix; A loss function construction unit, configured to binarize the consensus representation matrix into a hash code matrix, and construct a target loss function according to the target sample latent feature matrix, each modality feature matrix, the consensus representation matrix, the projection matrix, the hash code matrix, and the pseudo-label matrix; A cross-modal hashing retrieval unit, configured to determine a target hash code matrix and a target projection matrix according to the target loss function, and perform cross-modal hashing retrieval on target modality data according to the target hash code matrix and the target projection matrix.
9. An electronic device, characterized in that, Comprising at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the unsupervised cross-modal hashing retrieval method based on latent features according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the unsupervised cross-modal hashing retrieval method based on latent features according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unsupervised cross-modal hash retrieval method and system based on virtual label regression
CN110674323A
Depth cross-modal hash image retrieval method based on joint semantic matrix
CN113177132A
Multi-modal retrieval method and system based on weak supervision hash learning
CN114329109A
Cross-modal remote sensing image-text retrieval method based on single-mode feature modeling
CN117932101A
Unsupervised cross-modal retrieval method and system based on hypergraph convolution, medium and equipment
CN118916497A