Online semi-supervised collective matrix factorization hash method and system
Through the online semi-supervised collective matrix decomposition hashing method, multimodal features and pseudo-label generation, combined with graph regularization and stream clustering, the problems of low accuracy of unsupervised methods and high cost of supervised methods are solved, and efficient cross-modal retrieval in dynamic data environments are achieved.
Patent Information
- Application Number
- CN202510408807.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-15
AI Technical Summary
Existing unsupervised methods have lower accuracy when learning hash functions, while supervised methods are costly on large-scale data sets and cannot effectively update the hash model in a dynamic data environment, resulting in poor retrieval performance.
The online semi-supervised collective matrix decomposition hashing method is adopted. By obtaining multimodal features, the hash bucket index table and clustering center are constructed, pseudo-labels are generated, stream clustering is updated, and hash code learning is performed in combination with graph regularization constraints. The similarity sparse graph construction and stream clustering methods are used to obtain the parameters of the hash model.
While reducing space and time costs, it effectively constructs similarity sparse graphs to generate robust pseudo-labels, improves the accuracy and efficiency of cross-modal retrieval, and is suitable for dynamic data environments.
Smart Images

Figure CN120494123A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to an online semi-supervised collective matrix decomposition hashing method and system. Background Art
[0002] Unsupervised methods rely on spatial relationships to learn hash functions, but due to the lack of label information, the accuracy is low. Supervised methods, on the other hand, use labeled training data to achieve higher accuracy, but are costly and impractical for large-scale datasets.
[0003] Although the above methods can effectively update the hash model in a dynamic data environment, the supervised method requires that all data be labeled, which is unrealistic in a real data environment; and the unsupervised method cannot utilize label information, resulting in poor retrieval performance. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to provide an online semi-supervised collective matrix decomposition hashing method and system.
[0005] The technical solution adopted by the present invention is:
[0006] On the one hand, an embodiment of the present invention provides an online semi-supervised collective matrix decomposition hashing method, wherein the online semi-supervised collective matrix decomposition hashing method comprises the following steps:
[0007] Acquiring multimodal features; the multimodal features include image features and text features;
[0008] Obtaining a hash bucket index table and cluster centers based on the multimodal features;
[0009] Obtaining a similarity sparse graph according to the hash bucket index table and the cluster center;
[0010] Obtaining pseudo labels according to the similarity sparse graph;
[0011] Perform streaming clustering updates based on the pseudo labels to obtain an enhanced label matrix;
[0012] Performing collective matrix decomposition according to the enhanced label matrix to obtain an objective function;
[0013] Obtaining parameters of a hash model according to the objective function;
[0014] According to the parameters of the hash model, cross-modal retrieval data is obtained.
[0015] Furthermore, the obtaining of multimodal features comprises the following steps:
[0016] Acquire image features; the image features include labeled image data and unlabeled image data;
[0017] Acquire text features; the text features include labeled text data and unlabeled text data;
[0018] The image features and the text features are spliced together to obtain multimodal features.
[0019] Furthermore, the step of obtaining a hash bucket index table and a cluster center based on the multimodal features includes the following steps:
[0020] Based on the multimodal features, the feature space is preliminarily divided using unsupervised hashing, and a hash code of each multimodal feature is calculated;
[0021] The multimodal features with the same hash code are classified into the same hash bucket to obtain a hash bucket index table; the hash bucket index table records the hash bucket to which each multimodal feature belongs;
[0022] According to the hash bucket index table, the multimodal features are randomly selected as cluster centers.
[0023] Furthermore, obtaining a similarity sparse graph according to the hash bucket index table and the cluster center includes the following steps:
[0024] Calculating the distance between each of the multimodal features and the cluster center, and selecting training samples;
[0025] Splicing the cluster center and the training sample to obtain a splicing feature matrix;
[0026] The similarities between the splicing features in the same hash bucket are calculated according to the hash bucket index table, and a similarity sparse graph is obtained by using the similarity calculation.
[0027] Furthermore, the pseudo labels are obtained according to the similarity sparse graph, and the formula used includes:
[0028] constructing an asymmetric normalized Laplace matrix according to the similarity sparse graph;
[0029] Calculation is performed according to the asymmetric normalized Laplace matrix to generate a pseudo label.
[0030] Furthermore, performing streaming clustering update based on the pseudo labels to obtain an enhanced label matrix includes the following steps:
[0031] Perform streaming clustering updates based on the pseudo labels to obtain final pseudo labels;
[0032] Combining the final pseudo labels with the true labels to obtain an enhanced label matrix;
[0033] The step of updating the streaming clustering has multiple rounds of calculations. The pseudo labels are used as pseudo labels in round 0. In the streaming clustering update calculation in round t, the following steps are included:
[0034] Selecting several unused multimodal features as sample features for the tth round;
[0035] Obtain the cluster center of round t based on the sample features of round t and the pseudo labels of round t-1;
[0036] Selecting several unused multimodal features as training samples for the tth round;
[0037] Splicing the cluster centers of the t-th round and the training samples of the t-th round to obtain a t-th round splicing feature matrix;
[0038] Calculate the similarity between the t-th round splicing feature matrices in the same hash bucket, and use the similarity calculation to obtain a t-th round similarity sparse graph;
[0039] According to the similarity sparse graph of the t-th round, an asymmetric normalized Laplace matrix is constructed, the pseudo labels of the t-th round are obtained, and the streaming clustering update calculation of the t-th round is completed;
[0040] Wherein, t is an integer, t is greater than or equal to 1, and if the pseudo label of the tth round meets the setting conditions, the pseudo label of the tth round is used as the final pseudo label.
[0041] Furthermore, performing collective matrix decomposition according to the enhanced label matrix to obtain the objective function includes the following steps:
[0042] Performing collective matrix decomposition based on the enhanced label matrix to obtain a latent semantic basis matrix and a feature basis matrix; the latent semantic basis matrix includes a latent semantic basis matrix of the image modality and a latent semantic basis matrix of the text modality; the feature basis matrix includes an image feature basis matrix and a text feature basis matrix;
[0043] Setting an auxiliary matrix; the auxiliary matrix includes a first auxiliary matrix, a second auxiliary matrix, and a third auxiliary matrix;
[0044] Constructing an objective function based on the auxiliary matrix, the latent semantic basis matrix, and the feature basis matrix;
[0045] The formula used in the objective function includes:
[0046]
[0047] in, is the objective function; i is the weight parameter that balances the image and text modalities is the regularization term; M (t) is the correlation matrix between the two modes; Represents the image feature matrix; represents the text feature matrix; L (t) is the enhanced label matrix; T1 (t) a latent semantic basis matrix representing the image modality; a latent semantic basis matrix representing the text modality; represents the image feature basic matrix; Representing the text feature basic matrix; represents the first auxiliary matrix; represents the second auxiliary matrix; represents the third auxiliary matrix; represents the eigenvalue matrix; μ and γ are balance parameters that control the contribution of the corresponding terms.
[0048] Furthermore, obtaining the parameters of the hash model according to the objective function includes the following steps:
[0049] According to the objective function and the multimodal features, using iterative updating to obtain an updated latent semantic basis matrix of the image modality, an updated latent semantic basis matrix of the text modality, an updated basis matrix of the image features, and an updated basis matrix of the text features;
[0050] Parameters of a hash model are obtained according to the latent semantic basis matrix of the updated image modality, the latent semantic basis matrix of the updated text modality, the basis matrix of the updated image features, and the basis matrix of the updated text features.
[0051] On the other hand, an embodiment of the present invention also provides an online semi-supervised collective matrix decomposition hashing system, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the online semi-supervised collective matrix decomposition hashing method as described above.
[0052] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the online semi-supervised collective matrix decomposition hashing method as described above.
[0053] The embodiments of the present application include at least the following beneficial effects: The present application provides an online semi-supervised collective matrix decomposition hashing method and system. The present invention can obtain multimodal features; multimodal features include image features and text features; based on the multimodal features, a hash bucket index table and cluster centers are obtained; based on the hash bucket index table and cluster centers, a similarity sparse graph is obtained; based on the similarity sparse graph, pseudo labels are obtained; based on the pseudo labels, streaming clustering updates are performed to obtain an enhanced label matrix; based on the enhanced label matrix, collective matrix decomposition is performed to obtain an objective function; based on the objective function, the parameters of the hash model are obtained; based on the parameters of the hash model, cross-modal retrieval data is obtained. The present invention can use graph regularization constraints to perform hash code learning to retain the spatial structure information of unlabeled data, and also uses a streaming clustering method to effectively construct a similarity sparse graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 1 is a flow chart of an online semi-supervised collective matrix decomposition hashing method provided by an embodiment of the present invention;
[0055] Figure 2 Schematic diagram of implementing nearest neighbor image retrieval using OsCMFH provided by an embodiment of the present invention;
[0056] Figure 3 FIG. 4 is a schematic diagram of the OsCMFH algorithm flow provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0058] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0059] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0061] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0062] 1) OsCMFH, Online Collaborative Matrix Factorization Hashing;
[0063] 2) Implicit restart Arnoldi algorithm, an iterative method based on Krylov subspace for computing eigenvalues of sparse matrices;
[0064] 3) Krylov subspace, a search space for iterative methods;
[0065] 4) Hessenberg, Heisenberg matrix.
[0066] Considering that real-world data is often partially labeled, this paper urgently requires semi-supervised methods for cross-modal retrieval tasks. Therefore, this paper proposes a semi-supervised online cross-modal hashing method. This method generates pseudo-labels for unlabeled data through Laplace normalization and hash-based similarity sparse graph construction. Furthermore, it utilizes graph regularization constraints for hash code learning to preserve the spatial structure of the unlabeled data. Furthermore, this method employs streaming clustering to efficiently construct the similarity sparse graph.
[0067] The online semi-supervised collective matrix factorization hashing (OsCMFH) method of an embodiment of the present invention uses collective matrix factorization to learn the discriminative hash code of streaming data. OsCMFH adopts a new hash-based similarity sparse graph construction method to reduce space and time costs. Subsequently, asymmetric Laplace normalization is applied to the similarity sparse graph to generate pseudo labels. In addition, the method constructs a non-normalized Laplace eigenvector matrix of the similarity sparse graph and combines the pseudo labels to guide the learning process of the hash code. After each round of learning, the cluster center of the similarity sparse graph is updated using a streaming clustering method for use in the next round of similarity sparse graph construction.
[0068] The embodiments of the present invention are further described below with reference to the accompanying drawings.
[0069] On the one hand, the embodiment of the present invention provides an online semi-supervised collective matrix decomposition hashing method, referring to Figure 1 ,an online semi-supervised collective matrix factorization hashing method includes the following steps:
[0070] S100, obtaining multimodal features; the multimodal features include image features and text features;
[0071] S200, obtaining a hash bucket index table and cluster centers based on multimodal features;
[0072] S300, obtaining a similarity sparse graph according to the hash bucket index table and the cluster center;
[0073] S400, obtaining pseudo labels according to the similarity sparse graph;
[0074] S500: Perform streaming clustering update based on the pseudo labels to obtain an enhanced label matrix;
[0075] S600, performing collective matrix decomposition according to the enhanced label matrix to obtain the target function;
[0076] S700, obtaining parameters of a hash model according to the objective function;
[0077] S800: Obtain cross-modal retrieval data according to the parameters of the hash model.
[0078] The step S100 of acquiring multimodal features disclosed in the embodiment of the present invention includes the following steps:
[0079] S110, obtaining image features; image features include labeled image data and unlabeled image data;
[0080] S120, obtaining text features; text features include labeled text data and unlabeled text data;
[0081] S130: Concatenate the image features and the text features to obtain multimodal features.
[0082] S200 disclosed in the embodiment of the present invention obtains a hash bucket index table and a cluster center based on multimodal features, including the following steps:
[0083] S210, using unsupervised hashing to preliminarily divide the feature space according to the multimodal features, and calculating the hash code of each multimodal feature;
[0084] S220: Multimodal features with the same hash code are grouped into the same hash bucket to obtain a hash bucket index table; the hash bucket index table records the hash bucket to which each multimodal feature belongs;
[0085] S230: Randomly select multimodal features as cluster centers according to the hash bucket index table.
[0086] S300 disclosed in the embodiment of the present invention obtains a similarity sparse graph based on the hash bucket index table and the cluster center, including the following steps:
[0087] S310, calculating the distance between each multimodal feature and the cluster center, and selecting training samples;
[0088] S320, concatenating the cluster centers and the training samples to obtain a concatenated feature matrix;
[0089] S330 , calculating the similarity between the splicing features in the same hash bucket according to the hash bucket index table, and obtaining a similarity sparse graph using the similarity calculation.
[0090] S400 disclosed in the embodiment of the present invention obtains pseudo labels based on the similarity sparse graph, and the formula used includes:
[0091] S410, constructing an asymmetric normalized Laplace matrix according to the similarity sparse graph;
[0092] S420: Calculate according to the asymmetric normalized Laplace matrix to generate a pseudo label.
[0093] S500 disclosed in the embodiment of the present invention performs streaming clustering update based on pseudo labels to obtain an enhanced label matrix, including the following steps:
[0094] S510: Perform streaming clustering update based on the pseudo labels to obtain final pseudo labels;
[0095] S520, combining the final pseudo label and the true label to obtain an enhanced label matrix;
[0096] S530, the step of stream clustering update has multiple rounds of calculations, and the pseudo labels are used as the pseudo labels of round 0. In the stream clustering update calculation of round t, the following steps are included:
[0097] S531. Select several unused multimodal features as sample features for the tth round;
[0098] S532. Obtain the cluster center of round t based on the sample features of round t and the pseudo labels of round t-1;
[0099] S533, selecting several unused multimodal features as training samples for the tth round;
[0100] S534, performing splicing based on the cluster centers of the t-th round and the training samples of the t-th round to obtain the t-th round splicing feature matrix;
[0101] S535, calculating the similarity between the t-th round splicing feature matrices in the same hash bucket, and obtaining the t-th round similarity sparse graph by using the similarity calculation;
[0102] S536: Based on the similarity sparse graph of the t-th round, construct an asymmetric normalized Laplace matrix, obtain the pseudo label of the t-th round, and complete the streaming clustering update calculation of the t-th round;
[0103] S537, where t is an integer greater than or equal to 1. If the pseudo-label of the tth round meets the set conditions, the pseudo-label of the tth round is used as the final pseudo-label.
[0104] As an optional embodiment, the calculation data and corresponding symbols used in the present invention include:
[0105] At time t, a new image-text data block arrives with Indicates that Represent the feature matrices of the image and text modules respectively. In the calculation process of the similarity sparse graph, the features of the two modalities need to be spliced into To calculate the similarity. Each feature matrix contains labeled and unlabeled data, that is, in and Represent the feature matrices of unlabeled data and labeled data respectively, and The same process is also performed. N1 represents the number of samples in the newly arrived feature matrix, and d1 and d2 represent the dimensions of the feature matrices of the two modalities respectively. u and N l Represents the number of unlabeled and labeled samples. For labeled data, there is a true label matrix Where c is the number of categories.
[0106] The feature matrix of the previous time step is expressed as The number of samples is expressed as N2. The total number of samples is N=N1+N2.
[0107] The hash code of all samples is represented by B∈{0,1} N×k , where k is the length of the hash code. q query dataset of samples or Their hash codes are determined by the hash function and Generate, where and They represent the mapping matrices of the two modes respectively. sgn(·) represents the sign function.
[0108] Pseudo label generation according to an embodiment of the present invention:
[0109] First, the feature space is initially partitioned, and samples are assigned to different hash buckets. Then, within each bucket, the similarity between labeled and unlabeled samples is calculated. A similarity sparse graph is constructed for each unlabeled sample using the most similar labeled samples. This graph undergoes asymmetric Laplace normalization and label propagation, and is used to construct a non-normalized graph Laplace matrix. The feature matrix of this matrix participates in and guides the hash code learning process.
[0110] In addition, at the end of each round of learning, the features are clustered, and a streaming clustering method is used to meet the needs of online hashing. The cluster centers are integrated into the similarity sparse graph construction of the next round. This embodiment avoids high storage and computing costs by constructing a similarity sparse graph instead of a dense adjacency matrix. The similarity sparse graph reduces the need to store all pairwise similarities, greatly improves storage efficiency, reduces computational complexity, and naturally adapts to efficient algorithms. At the same time, streaming clustering introduces historical data to construct a similarity sparse graph while avoiding redundant calculations. At the end of this round of training, assuming there are n cluster centers, for the i-th sample that is not selected as a cluster center in the feature matrix, it is compared with the cluster center matrix of the previous round. The distance between the nearest cluster centers in is calculated as follows:
[0111]
[0112] in Represents the nearest cluster center. Samples are assigned to the cluster formed by their nearest cluster center. In the first round, the initial cluster center will be randomly selected from the existing samples. Subsequently, the cluster center is dynamically updated based on the distance between each sample and its nearest cluster center, which is calculated as follows:
[0113]
[0114] in, Indicates that the last round was allocated to C j The sample set of clusters formed, S (t) Represents the sample set of the current round. The label of each cluster center is calculated by weighting the labels of the samples in the cluster, in a similar way. The cluster center of the previous round is spliced with the training samples to form a new feature matrix To construct a sparse graph (similarity sparse graph), first, a simple unsupervised hashing method is used in the first round to initially divide the feature space. The hash function is as follows:
[0115]
[0116] in, is the random mapping matrix and b is the bias term. The similarity between unlabeled samples and labeled samples is calculated as follows:
[0117]
[0118] To construct a sparse graph, only the similarity between samples in the same hash bucket is calculated, as follows:
[0119]
[0120] Then, for each unlabeled sample, keep N s The most similar samples are found and the similarity is normalized. The construction process of the similarity sparse graph is as follows:
[0121]
[0122] To avoid the probability transfer function's over-reliance on prior knowledge and existing labels, this method uses a normalized Laplacian matrix for label propagation. This method captures the spatial structure of the data, allowing for more efficient label propagation using the spatial information of unlabeled data. Furthermore, the method exhibits good adaptability when processing high-dimensional sparse data. Its calculation is as follows:
[0123]
[0124] Among them, D (t) represents the degree matrix, is the asymmetric normalized Laplace matrix. Label matrix Initialize it to a zero matrix and use the true label matrix Replace the part corresponding to the labeled sample. Then, the label is propagated to the unlabeled sample through the graph structure, which is calculated as follows:
[0125]
[0126] Among them, α is a parameter that controls the degree of dependence on the initial label matrix. After multiple iterations until convergence, the true label is gradually propagated to the unlabeled data, and the probability of each sample corresponding to the label is stored in the matrix. Finally, the label with the highest probability is selected as the pseudo label. Specifically, if multiple labels need to be generated, the label with a probability greater than a certain threshold can be assigned to the sample as a pseudo label. The label matrix of some unlabeled samples is used as its pseudo label, and these pseudo labels are combined with the true labels to form the final label matrix L (t) .
[0127] S600 disclosed in the embodiment of the present invention performs collective matrix decomposition based on the enhanced label matrix to obtain the objective function, including the following steps:
[0128] S610: Perform collective matrix decomposition based on the enhanced label matrix to obtain a latent semantic basis matrix and a feature basis matrix; the latent semantic basis matrix includes a latent semantic basis matrix of the image modality and a latent semantic basis matrix of the text modality; the feature basis matrix includes an image feature basis matrix and a text feature basis matrix;
[0129] S620, setting an auxiliary matrix; the auxiliary matrix includes a first auxiliary matrix, a second auxiliary matrix, and a third auxiliary matrix;
[0130] S630, constructing an objective function based on the auxiliary matrix, the latent semantic basis matrix, and the feature basis matrix;
[0131] The formula used for the objective function includes:
[0132]
[0133] in, is the objective function; i is the weight parameter that balances the image and text modalities is the regularization term; M (t) is the correlation matrix between the two modes; Represents the image feature matrix; represents the text feature matrix; L (t) is the enhanced label matrix; T1 (t) The latent semantic basis matrix representing the image modality; The latent semantic basis matrix representing the text modality; Represents the image feature basis matrix; Represents the basic matrix of text features; represents the first auxiliary matrix; represents the second auxiliary matrix; represents the third auxiliary matrix; represents the eigenvalue matrix; μ and γ are balance parameters that control the contribution of the corresponding terms.
[0134] As an optional implementation, the present invention takes into account that in cross-modal retrieval, multimodal data contains shared features and modality-specific features. Most methods focus on the shared latent semantic representation and ignore the modality-specific features. The present invention adopts matrix decomposition to capture the shared latent semantic representation at the same time. and specific latent semantic representations and k1 and k2 represent the dimensions of the shared latent semantic representation matrix and the specific latent semantic matrix, respectively, k = k1 + k2. In addition, the auxiliary matrix is introduced To combine the label matrix, and
[0135] In addition, the sparse similarity graph of samples can reflect the similarity relationship between samples and is very suitable for guiding hash code learning. It can ensure that similar samples in the sparse graph have similar hash codes in the Hamming space. In order to utilize this sparse graph, we first construct a non-normalized graph Laplacian matrix. This matrix combines node degree information to alleviate the degree imbalance problem and reveals the global structure and partition characteristics of the graph through spectral decomposition. It is calculated as follows:
[0136]
[0137] Subsequently, this method uses the implicit restarted Arnoldi algorithm to extract the eigenvalues of the graph Laplacian matrix. By iteratively constructing an orthogonal basis of the Krylov subspace, the eigenvalue problem of the graph Laplacian matrix is transformed into a smaller-scale Hessenberg eigenvalue problem. The extracted eigenvalue matrix is denoted as Therefore, the objective function of the embodiment of the present invention is as follows:
[0138]
[0139] in, is a regularization term used to avoid overfitting. and denote the latent semantic basis matrices of image and text modalities respectively, and and is the basic matrix of image and text features. i is the weight parameter that balances the image and text modalities M (t) is the correlation matrix between the two modes. μ and γ are the balance parameters that control the contribution of the corresponding terms. By applying the constraints Specific representations of images and text are associated. Therefore, the latent semantic representation of the training data can be expressed as:
[0140]
[0141] When generating hash codes for new data, the hash codes for old data are also updated. Therefore, the total training data for online learning includes new data and accumulated old data. Similarly, the objective function for old data is It can be constructed by replacing all parameters in the objective function with the corresponding parameters of the previous round. Therefore, the overall objective function is as follows:
[0142]
[0143] S700 disclosed in the embodiment of the present invention obtains the parameters of the hash model according to the objective function, including the following steps:
[0144] S710, using iterative updating to obtain an updated latent semantic basis matrix of the image modality, an updated latent semantic basis matrix of the text modality, an updated basis matrix of the image features, and an updated basis matrix of the text features according to the objective function and the multimodal features;
[0145] S720. Obtain parameters of the hash model according to the updated latent semantic basis matrix of the image modality, the updated latent semantic basis matrix of the text modality, the updated basis matrix of the image features, and the updated basis matrix of the text features.
[0146] As an optional implementation manner, the parameter solving step of the embodiment of the present invention includes:
[0147] The present invention takes into account that the objective function is non-convex for all parameters, so direct optimization is challenging. However, when other parameters are fixed, the objective function is convex for any given parameter. Therefore, an iterative update scheme can be used to solve the objective function. The specific optimization process is as follows:
[0148] 1. Update the latent semantic basis matrix T1 of the image modality (t) :
[0149]
[0150]
[0151] in, and is the historical intermediate matrix of the previous iteration.
[0152] 2. Update the latent semantic basis matrix of text modality
[0153] in, is the historical intermediate matrix of the previous iteration.
[0154] 3. Update the basic matrix of image features
[0155]
[0156] in, and is the historical intermediate matrix of the previous iteration.
[0157] 4. Update the basic matrix of text features
[0158]
[0159] in, and is the historical intermediate matrix of the previous iteration.
[0160] 5. Update the correlation matrix M between the two modalities (t) :
[0161]
[0162] in, and is the historical intermediate matrix of the previous iteration.
[0163] 6. Update the auxiliary matrix
[0164]
[0165] 7. Update the auxiliary matrix
[0166]
[0167] 8. Update the auxiliary matrix
[0168]
[0169] The above process will be repeated until convergence. Then, according to the Q of formula (11) (t) Update and can be done through sgn(Q (t) ) directly obtains the hash code of the training data. The hash model of the previous round can also be updated online based on the results of the current round without accessing the old data. Assume and They represent the updated parameters of the previous round of hash model, which can be calculated as follows:
[0170]
[0171] For the query sample, the hash function can be obtained through the linear regression model. The original image and text features can be linearly projected into the latent semantic space, and the hash function is as follows:
[0172]
[0173] By using the same method as before, P1 (t)The update formula can be obtained as follows:
[0174]
[0175] for The update formula is the same as P1 (t) Similarly, the hash code of the query sample can be obtained through the projection matrix P1 (t) and Directly obtain, in the previous calculation, I is the identity matrix.
[0176] As an optional implementation, refer to Figure 2 Nearest neighbor image retrieval and Figure 3 The OsCMFH algorithm flow chart of the embodiment of the present invention implements the nearest neighbor image retrieval method, wherein the maximum number of iterations is 10, the hash code length is set to 64, α is set to 0.5, and N c is 20, N s =10, λ, μ, and γ are 0.1, 0.01, and 0.01, respectively. This implementation implements the nearest neighbor image retrieval function, requiring the retrieval of the 100 samples in the database that are most relevant to the current query sample. At time t, OsCMFH first uses a hash-based similarity sparse graph construction method to construct a similarity sparse graph of unlabeled samples, and uses a Laplace normalized label propagation algorithm to generate pseudo labels. Subsequently, collective matrix decomposition is used to generate hash codes for multimodal data. Next, all mapped results are stored in the database, and the Hamming distance between samples is calculated using the data stored in the database to evaluate the similarity between samples. Finally, the 100 candidate samples with the smallest asymmetric distance to the query sample are returned as the retrieval results.
[0177] The present invention proposes an online semi-supervised collective matrix factorization hashing (OsCMFH) method. OsCMFH adopts a novel hash-based similarity sparse graph construction method, which effectively reduces space and time costs. Subsequently, asymmetric Laplace normalization is applied to the similarity sparse graph to generate pseudo-labels, and these pseudo-labels are used to guide the learning process of hash codes. In addition, the method introduces a streaming clustering method to update the cluster centers after each round of learning and integrate them into the sparse graph construction of the next round. Experimental results show that OsCMFH performs better than other comparison methods in cross-modal retrieval tasks under semi-supervised non-stationary data environments.
[0178] The key points of the present invention are:
[0179] 1. The proposed OsCMFH method is a brand-new hashing method. Although it refers to the collective matrix factorization method commonly used in cross-modal hashing, the semi-supervised process, hash code objective function, and streaming clustering method are brand-new.
[0180] 2. We propose an online semi-supervised cross-modal hashing method, OsCMFH, which utilizes hashing-based sparse graph construction and asymmetric Laplacian matrix with label propagation to generate robust pseudo-labels from unlabeled data.
[0181] 3. OsCMFH combines a graph regularization term and a denormalized graph Laplacian matrix and updates parameters online via collective matrix factorization, thereby enhancing the learning of discriminative hash codes.
[0182] 4. A large number of experiments were conducted in various dynamic data environments. The experimental results show that OsCMFH has significantly better retrieval performance than the comparison methods.
[0183] On the other hand, an embodiment of the present invention also provides an online semi-supervised collective matrix decomposition hashing system, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the above-mentioned online semi-supervised collective matrix decomposition hashing method.
[0184] The processor and the memory can be connected via a bus or other means. The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0185] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned online semi-supervised collective matrix decomposition hashing method.
[0186] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0187] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. An online semi-supervised collective matrix factorization hashing method, characterized by: The online semi-supervised collective matrix decomposition hashing method comprises the following steps: Acquiring multimodal features; the multimodal features include image features and text features; Obtaining a hash bucket index table and cluster centers based on the multimodal features; Obtaining a similarity sparse graph according to the hash bucket index table and the cluster center; Obtaining pseudo labels according to the similarity sparse graph; Perform streaming clustering updates based on the pseudo labels to obtain an enhanced label matrix; Performing collective matrix decomposition according to the enhanced label matrix to obtain an objective function; Obtaining parameters of a hash model according to the objective function; According to the parameters of the hash model, cross-modal retrieval data is obtained.
2. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The obtaining of multimodal features comprises the following steps: Acquire image features; the image features include labeled image data and unlabeled image data; Acquire text features; the text features include labeled text data and unlabeled text data; The image features and the text features are spliced together to obtain multimodal features.
3. The online semi-supervised collective matrix decomposition hashing method according to claim 1, characterized in that The step of obtaining a hash bucket index table and a cluster center based on the multimodal features includes the following steps: Based on the multimodal features, the feature space is preliminarily divided using unsupervised hashing, and a hash code of each multimodal feature is calculated; The multimodal features with the same hash code are classified into the same hash bucket to obtain a hash bucket index table; the hash bucket index table records the hash bucket to which each multimodal feature belongs; According to the hash bucket index table, the multimodal features are randomly selected as cluster centers.
4. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The step of obtaining a similarity sparse graph according to the hash bucket index table and the cluster center includes the following steps: Calculating the distance between each of the multimodal features and the cluster center, and selecting training samples; Splicing the cluster center and the training sample to obtain a splicing feature matrix; The similarities between the splicing features in the same hash bucket are calculated according to the hash bucket index table, and a similarity sparse graph is obtained by using the similarity calculation.
5. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The pseudo labels are obtained according to the similarity sparse graph, and the formula used includes: constructing an asymmetric normalized Laplace matrix according to the similarity sparse graph; Calculation is performed according to the asymmetric normalized Laplace matrix to generate a pseudo label.
6. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The step of performing streaming clustering update based on the pseudo labels to obtain an enhanced label matrix includes the following steps: Perform streaming clustering updates based on the pseudo labels to obtain final pseudo labels; Combining the final pseudo labels with the true labels to obtain an enhanced label matrix; The step of updating the streaming clustering has multiple rounds of calculations. The pseudo labels are used as pseudo labels in round 0. In the streaming clustering update calculation in round t, the following steps are included: Selecting several unused multimodal features as sample features for the tth round; Obtain the cluster center of round t based on the sample features of round t and the pseudo labels of round t-1; Selecting several unused multimodal features as training samples for the tth round; Splicing the cluster centers of the t-th round and the training samples of the t-th round to obtain a t-th round splicing feature matrix; Calculate the similarity between the t-th round splicing feature matrices in the same hash bucket, and use the similarity calculation to obtain a t-th round similarity sparse graph; According to the similarity sparse graph of the t-th round, an asymmetric normalized Laplace matrix is constructed, the pseudo labels of the t-th round are obtained, and the streaming clustering update calculation of the t-th round is completed; Wherein, t is an integer, t is greater than or equal to 1, and if the pseudo label of the tth round meets the setting conditions, the pseudo label of the tth round is used as the final pseudo label.
7. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The method of performing collective matrix decomposition according to the enhanced label matrix to obtain the objective function includes the following steps: Performing collective matrix decomposition based on the enhanced label matrix to obtain a latent semantic basis matrix and a feature basis matrix; the latent semantic basis matrix includes a latent semantic basis matrix of the image modality and a latent semantic basis matrix of the text modality; the feature basis matrix includes an image feature basis matrix and a text feature basis matrix; Setting an auxiliary matrix; the auxiliary matrix includes a first auxiliary matrix, a second auxiliary matrix, and a third auxiliary matrix; Constructing an objective function based on the auxiliary matrix, the latent semantic basis matrix, and the feature basis matrix; The formula used in the objective function includes: in, is the objective function; i is the weight parameter that balances the image and text modalities is the regularization term; M (t) is the correlation matrix between the two modes; Represents the image feature matrix; represents the text feature matrix; L (t) is the enhanced label matrix; T1 (t) a latent semantic basis matrix representing the image modality; a latent semantic basis matrix representing the text modality; represents the image feature basic matrix; Representing the text feature basic matrix; represents the first auxiliary matrix; represents the second auxiliary matrix; represents the third auxiliary matrix; represents the eigenvalue matrix; μ and γ are balance parameters that control the contribution of the corresponding terms.
8. The online semi-supervised collective matrix factorization hashing method according to claim 1, characterized in that The step of obtaining the parameters of the hash model according to the objective function comprises the following steps: According to the objective function and the multimodal features, using iterative updating to obtain an updated latent semantic basis matrix of the image modality, an updated latent semantic basis matrix of the text modality, an updated basis matrix of the image features, and an updated basis matrix of the text features; Parameters of a hash model are obtained according to the latent semantic basis matrix of the updated image modality, the latent semantic basis matrix of the updated text modality, the basis matrix of the updated image features, and the basis matrix of the updated text features.
9. An online semi-supervised collective matrix factorization hashing system, characterized by: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the online semi-supervised collective matrix decomposition hashing method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the online semi-supervised collective matrix decomposition hashing method according to any one of claims 1 to 8.