Multi-view clustering method for large data
The feature matrix is processed through low-pass filtering and graph learning modules, combining sparse and rank constraints, and the long running time and memory overflow problems of deep learning methods on large data sets is solved, achieving efficient and accurate multi-view clustering.
Patent Information
- Application Number
- CN202411944950.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-09
AI Technical Summary
The deep learning method of existing multi-view clustering methods has a long run time, easy memory overflow, and it is difficult to effectively process large data sets.
A multi-view clustering method is proposed. The feature matrix is processed through low-pass filtering and graph learning modules, the anchor matrix is constructed, and the clustering results are obtained through sparse constraints and rank constraints, avoiding the computational complexity and memory overflow problems of deep learning methods.
This method can effectively reduce calculation costs, reduce noise and missing, directly obtain clustering labels, avoid errors and randomness in post-processing, and improve the accuracy and robustness of clustering.
Smart Images

Figure CN119961703A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph clustering of large data, and in particular to a multi-view clustering method for large data. Background Art
[0002] Graph clustering is a data mining technology that aims to group nodes in graph structured data and cluster similar nodes. The goal is to maintain the integrity and high dimensionality of image features while achieving fast and effective grouping. It is often used in movie recommendations and medical image analysis. Multi-view clustering uses richer graph information on the basis of graph clustering to seek consistent clustering results.
[0003] Single-view clustering is a method of partitioning data based on a single feature set or view. In the real world, data often has information from multiple sources, which may present different characteristics and structures in different feature spaces. Therefore, single-view clustering methods cannot fully explore the intrinsic structure and characteristics of the data. Multi-view clustering can more comprehensively describe the characteristics and structure of the data by integrating information from multiple views, thereby improving the accuracy and robustness of clustering.
[0004] The key steps of multi-view clustering methods usually include: data preprocessing, building a graph similarity matrix, building a Laplacian matrix, fusing multi-view information, solving eigenvectors and clustering. With the development of deep learning, deep learning is the main graph embedding method used in multi-view clustering. Deep learning methods require high computational complexity and space complexity. In particular, real-world data sets often have a large number of nodes, which may lead to problems such as long running time and memory overflow in deep learning methods. Summary of the invention
[0005] The purpose of the present invention is to provide a multi-view clustering method for large data to solve the problems of long running time and easy memory overflow in the deep learning method of the existing multi-view clustering method.
[0006] To achieve the above object, the present invention provides a basic solution: a multi-view clustering method for large data, comprising the following steps:
[0007] S1. Input multiple feature matrices, adjacency matrices, anchor point numbers, and cluster numbers composed of multiple view data into the database, and normalize the feature matrices;
[0008] S2. Filter the feature matrix by low-pass filtering to obtain a filtered feature matrix;
[0009] S3. The filtered feature matrix is used to construct an anchor matrix through the graph learning module;
[0010] S4. Split the anchor matrix into a shared anchor matrix and an independent feature anchor matrix through sparse constraints;
[0011] S5. construct an augmented matrix using a shared anchor matrix, and concatenate the augmented matrices to construct a normalized Laplace matrix;
[0012] S6. The clustering result is obtained by performing rank constraints on the normalized Laplace matrix and connectivity constraints on the shared anchor matrix.
[0013] The principle and beneficial effect of the present invention are that multi-view clustering can more comprehensively describe the characteristics and structure of data by integrating the information of multiple views, thereby improving the accuracy and robustness of clustering. When in use, the feature matrix, adjacency matrix, number of anchor points and number of clustering clusters of the multi-view data are first input, and then the feature matrix is filtered through a filter to obtain a feature matrix, which can reduce the noise and missingness in the multi-view data. Then, the feature matrix is constructed into an anchor graph through a graph learning module. The graph learning module solves the problem of complex calculation and too many parameters in deep learning, and effectively reduces the calculation cost. Then, a rank constraint is designed for the Laplacian matrix of the anchor graph, and the anchor graph is divided into a shared anchor graph and an independent feature anchor graph. The clustering label can be directly obtained without post-processing after the graph is embedded, which can avoid the errors generated in the post-processing and the influence of the randomness of the traditional clustering method. Finally, a connectivity constraint is imposed on the shared anchor graph, and the clustering result can be generated directly from the shared anchor graph without referencing other clustering methods to the obtained shared graph, and no post-processing process is required.
[0014] Solution 2 is the preferred solution of the basic solution. In step S1, the formula for normalizing the adjacency matrix is:
[0015]
[0016] in:
[0017] A V represents the normalized adjacency matrix;
[0018] in represents the input adjacency matrix, represents the element in the i-th row and j-th column of the v-th view of the adjacency matrix. If there is an edge between the i-th row and the j-th column, then If there is no edge between the i-th row and the j-th column, then
[0019] D V =diag(d v1 ,...,d vN ) is the degree matrix of the adjacency matrix, and diag represents the diagonal matrix;
[0020] I represents the identity matrix; the normalized adjacency matrix is a method for processing graph data to solve the degree deviation and gradient vanishing problems. After normalization, the contribution of each node is balanced and the gradient calculation is more stable.
[0021] Solution 3 is the preferred basic solution. In step S2, the formula for obtaining the filtered feature matrix is:
[0022] in:
[0023] X represents the input feature matrix;
[0024] G represents low-pass filtering. The formula for calculating low-pass filtering is: G f =Up(Λ)U -1 , where U represents the orthogonal eigenvector, p represents the frequency response function, Λ represents the increasing eigenvalue, and f represents the graph signal, because p(λ q ) is decreasing and non-negative on [0, 2], and the low-pass filter formula is derived as follows: And because the eigenvalues λ of all normalized Laplacian matrices q All are within the interval [0,2], so the formula for calculating the low-pass filter is derived as:
[0025] Where Lnorm represents the normalized Laplace matrix; a low-pass filter is an electronic filter whose main function is to allow signals below the cutoff frequency to pass through while suppressing or attenuating signals above the cutoff frequency. It is often used to smooth images, remove noise, and retain low-frequency details of images, thereby blurring the image and reducing noise and missing information in multi-view data.
[0026] Solution 4 is the preferred solution of Solution 3. In step S3, the anchor matrix construction formula is:
[0027] Finally, we get v S V and an S∈R m*N ,
[0028] in:
[0029] S V represents the anchor matrix in the vth view of the original data, and S represents the shared anchor matrix of the original data;
[0030] Among them B V represents the anchor matrix of the v-th view, represents the element of the mth anchor point in the anchor point matrix, and B V yes part of , T represents transpose;
[0031] a(S V , S) represents the regularization term;
[0032] The formula of the graph learning module is:
[0033] where Z∈R N*N represents the similarity matrix, represents the reconstruction loss, represents the regularization term, which constrains Z to ensure that the learned graph is non-negative and normalized.
[0034] Anchor graphs are a teaching and learning tool that helps students deeply understand the learning content through visual presentation, supports them to complete tasks independently, and reinforces the procedures and routines in the classroom. In image processing, anchors provide a reference frame for each target, helping to better estimate the size, shape and position of the target, thereby simplifying the target detection task and improving detection efficiency and accuracy. In this method, anchor graphs can directly obtain cluster labels without post-processing after graph embedding, which can avoid the errors generated in post-processing and the influence of randomness of traditional clustering methods.
[0035] Solution 5 is the preferred solution of Solution 4. In step S4, the sparse constraint is to introduce a weight parameter for each view to determine the importance of different views. By learning the anchor matrix on the vth view, the anchor matrix is decomposed into a shared anchor matrix and an independent feature anchor matrix. The decomposition formula is:
[0036]
[0037] in:
[0038] i represents the i-th vector in the corresponding matrix;
[0039] B V Represents the anchor matrix of the vth view;
[0040] S V represents the anchor matrix, because the shared anchor matrix and the independent feature anchor matrix satisfy so Where C represents the shared anchor matrix, H V represents the independent feature anchor matrix, F represents the norm, and λ is the weight.
[0041] Through B V , C and H V Continuously update so that When the minimum value is reached, B V , C and H Vis the optimal value; sparse constraint refers to finding a solution that is as sparse as possible under certain conditions. In this method, it is used for anchor image denoising, compression and reconstruction to improve image quality.
[0042] Solution 6: This is the preferred solution of Solution 5. In step S5, the formula for constructing an augmented matrix using a shared anchor matrix is:
[0043] Where Ga represents the augmented matrix, C represents the shared anchor matrix, and C T represents the transposed matrix of the shared anchor matrix;
[0044] The augmented matrices are concatenated to construct a normalized Laplace matrix, and the construction formula is: Where L Ga represents the normalized Laplace matrix of Ga, D V represents the degree matrix of the augmented matrix, and I represents the identity matrix.
[0045] Solution 7: This is the preferred solution of Solution 6. In step S6, the augmented matrix is composed of shared anchor matrices, so they have the same number of connected components. Ga The rank constraint imposed is equivalent to the connectivity constraint imposed on C. The clustering result is obtained from the number of clusters in the rank-constrained clustering. The formula for the rank constraint is: rank(L Ga )=N+m-cl, where m represents the number of anchor points, cl represents the number of clusters, and N represents the number of nodes. The connectivity constraint requires that there is no situation in the space that can be divided into two or more non-empty, non-intersecting sets. The clustering result can be generated directly from the shared anchor graph without referencing other clustering methods on the obtained shared graph.
[0046] Scheme 8: This is the preferred option of Scheme 5, obtaining C and H V In the formula, the variable H V , B V and C adopt the strategy of alternating update. Alternating update means that when one variable is updated, the values of other variables are fixed by default:
[0047] Fixed H V and C, for B V The updates are as follows:
[0048] in Tr(B vT Q V ) represents the trace of the matrix, by Q V Perform singular value decomposition to calculate the optimal B V ;
[0049] Fixed B V and C, for H VThe updates are as follows: in H V After updating Derivation, get each The closed-form solution of
[0050] Fixed B V and H V When , the update of C is as follows: in R∈R (N+m)*cl , R represents the indicator matrix, and R T R=I cl , I cl Represents the identity matrix with length and width cl;
[0051] γ represents a sufficiently large constant;
[0052] R and C are updated alternately until C has exactly cl connected components:
[0053] When C is fixed, the update for R is as follows: Based on Ga is composed of C and The augmented matrix composed of
[0054] Where D represents the degree matrix of Ga, R represents the singular matrix, and R U represents the singular left matrix, R V represents the singular right matrix, R U ∈R N*cl , R V ∈R N*cl , and D U ∈R N*N and D V ∈R m*m It is decomposed from R and D;
[0055] When the optimal R is obtained and R is fixed, the update of C is as follows:
[0056]
[0057] Where t represents the number of updates, c represents the elements of C, d represents the elements of D, w represents the elements of W, and r represents the elements of R. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 The present invention is a flowchart of a multi-view clustering method for large data. DETAILED DESCRIPTION
[0059] The present invention is further described in detail below through specific embodiments:
[0060] Example
[0061] like Figure 1 A multi-view clustering method for large data, comprising the following steps:
[0062] S1. Input multiple feature matrices, adjacency matrices, anchor point numbers, and cluster numbers composed of multiple view data into the database, and normalize the feature matrices. The formula for the normalized adjacency matrix is:
[0063]
[0064] in:
[0065] A V represents the normalized adjacency matrix;
[0066] in represents the input adjacency matrix, represents the element in the i-th row and j-th column of the v-th view of the adjacency matrix. If there is an edge between the i-th row and the j-th column, then If there is no edge between the i-th row and the j-th column, then
[0067] D V =diag(d v1 ,...,d vN ) is the degree matrix of the adjacency matrix, and diag represents the diagonal matrix;
[0068] I represents the identity matrix;
[0069] S2. Filter the feature matrix through low-pass filtering, and the formula of the filtered feature matrix is:
[0070] in:
[0071] X represents the input feature matrix;
[0072] G represents low-pass filtering. The formula for calculating low-pass filtering is: G f =Up(Λ)U -1 , where U represents the orthogonal eigenvector, And because the eigenvalues λ of all normalized Laplacian matrices q All are within the interval [0,2], so the formula for calculating the low-pass filter is derived as:
[0073] Where Lnorm represents the normalized Laplace matrix;
[0074] S3. The filtered feature matrix is used to construct an anchor matrix through the graph learning module. The anchor matrix construction formula is:
[0075] Finally, we get v S V and an S∈R m*N ,
[0076] in:
[0077] S V represents the anchor matrix in the vth view of the original data, and S represents the shared anchor matrix of the original data;
[0078] Among them B V represents the anchor matrix of the v-th view, represents the element of the mth anchor point in the anchor point matrix, and B V yes part of , T represents transpose;
[0079] a(S V , S) represents the regularization term;
[0080] S4. Split the anchor matrix into a shared anchor matrix and an independent feature anchor matrix through sparse constraints. The sparse constraint introduces a weight parameter for each view to determine the importance of different views. By learning the anchor matrix on the vth view, the anchor matrix is decomposed into a shared anchor matrix and an independent feature anchor matrix. The decomposition formula is:
[0081]
[0082] in:
[0083] i represents the i-th vector in the corresponding matrix;
[0084] B V Represents the anchor matrix of the vth view;
[0085] S V represents the anchor matrix, because the shared anchor matrix and the independent feature anchor matrix satisfy so
[0086] Where C represents the shared anchor matrix, H V represents the independent feature anchor matrix, F represents the norm, and λ is the weight.
[0087] Through B V , C and H V Continuously update so that When the minimum value is reached, B V , C and H V is the optimal value;
[0088] Variable H V, B V and C adopt the strategy of alternating update. Alternating update means that when one variable is updated, the values of other variables are fixed by default:
[0089] Fixed H V and C, for B V The updates are as follows:
[0090] in Tr(B vT Q V ) represents the trace of the matrix, by Q V Perform singular value decomposition to calculate the optimal B V ;
[0091] Fixed B V and C, for H V The updates are as follows: in H V After updating Derivation, get each The closed-form solution of
[0092] Fixed B V and H V When , the update of C is as follows: in R∈R (N+m)*cl , R represents the indicator matrix, and R T R=I cl , I cl Represents the identity matrix with length and width cl;
[0093] γ represents a sufficiently large constant;
[0094] R and C are updated alternately until C has exactly cl connected components:
[0095] When C is fixed, the update for R is as follows: Based on Ga is composed of C and The augmented matrix composed of
[0096] Where D represents the degree matrix of Ga, R represents the singular matrix, and R U represents the singular left matrix, R V represents the singular right matrix, R U ∈R N*cl , R V ∈R N*cl , and D U ∈R N*N and D V ∈R m*m It is decomposed from R and D;
[0097] When the optimal R is obtained and R is fixed, the update of C is as follows:
[0098]
[0099] Where t represents the number of updates, c represents the elements of C, d represents the elements of D, w represents the elements of W, and r represents the elements of R.
[0100] S5. Use the shared anchor matrix to construct the augmented matrix, concatenate the augmented matrix to construct the normalized Laplace matrix, and use the shared anchor matrix to construct the augmented matrix formula:
[0101] Where Ga represents the augmented matrix, C represents the shared anchor matrix, and C T represents the transposed matrix of the shared anchor matrix;
[0102] The augmented matrices are concatenated to construct a normalized Laplace matrix, and the construction formula is: Where L Ga represents the normalized Laplace matrix of Ga, D V represents the degree matrix of the augmented matrix, I represents the identity matrix;
[0103] S6. By rank-constraining the normalized Laplacian matrix and connectivity-constraining the shared anchor matrix, the clustering result is obtained. The augmented matrix is composed of the shared anchor matrices, so they have the same number of connected components. Ga The rank constraint imposed is equivalent to the connectivity constraint imposed on C. The clustering result is obtained from the number of clusters in the rank-constrained clustering. The formula for the rank constraint is: rank(L Ga )=N+m-cl, where m represents the number of anchor points, cl represents the number of clusters, and N represents the number of nodes.
[0104] The implementation method of this embodiment is as follows: when in use, multiple feature matrices, adjacency matrices, anchor point numbers and clustering cluster numbers composed of multiple view data are input into the database, and the feature matrix is normalized, and then the feature matrix is filtered by low-pass filtering to obtain a filtered feature matrix, which can reduce the noise and missingness in the multi-view data. Then, the feature matrix is constructed into an anchor matrix through a graph learning module. The graph learning module solves the problem of complex calculation and too many parameters in deep learning, and effectively reduces the calculation cost. Then, a sparse constraint is designed for the Laplacian matrix of the anchor matrix, and the anchor matrix is divided into a shared anchor matrix and an independent feature anchor matrix. The sparse constraint can denoise, compress and reconstruct the anchor matrix, improve the image quality, and directly obtain the clustering label without post-processing after graph embedding. This can avoid the errors generated in the post-processing and the influence of the randomness of the traditional clustering method. Finally, a connectivity constraint is imposed on the shared anchor matrix, and the clustering result can be generated directly from the shared anchor matrix without referencing other clustering methods for the obtained shared anchor matrix, and no post-processing process is required.
[0105] In order to verify the effectiveness of this method, this method is compared with 13 existing algorithms. Four indicators are used, including clustering accuracy (ACC), normalized mean similarity (NMI), purity (Purity) and weighted harmonic mean (F-score) of clustering results. The Dermatology dataset contains diagnostic information of six skin diseases; ACM is a paper network; Caltech101-20 comes from the image dataset Caltech101, which contains images of 101 different objects; DBLP is an author information network; YTF10, YTF20 and YTF50 are sub-datasets extracted from the video dataset YouTube Faces for comparison. The comparison results are shown in Table 1:
[0106] Table 1 Clustering effects of various algorithms on seven data sets
[0107]
[0108]
[0109]
[0110] According to these results, BDMVLC performs well on most datasets, compared with the four algorithms that also use anchor graphs. LMVSC learns the anchor graphs of each view separately and then fuses them, while this algorithm combines the anchor graphs of each view with the shared anchor graph in a unified framework, which has better performance. Compared with SMVSC, MSGL, and FPMVS that directly generate shared anchor graphs, BDMVLC filters out view-specific noise while learning the shared anchor graph. In addition, by imposing rank constraints on the anchor graph, this algorithm can directly generate clustering results without errors in post-processing and will not be affected by the randomness of traditional clustering methods.
[0111] The operation time of the multi-view clustering method on each data set is compared. Since some algorithms cannot run on large data sets due to memory reasons, only the 7 algorithms that can run on all data sets are compared here. The specific results are shown in Table 2:
[0112] Table 2 Running time of various algorithms on seven data sets (s)
[0113]
[0114] As can be seen from the table above, BDMVLC has excellent operating efficiency. Although the efficiency of BMVC and FastMICE is higher than ours, their clustering effect is poor. Taking all factors into consideration, BDMVLC is an algorithm with high efficiency, good effect and strong versatility.
[0115] The above is only an embodiment of the present invention, and the common knowledge such as the known specific structure and characteristics in the scheme is not described in detail here. It should be pointed out that for those skilled in the art, several deformations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A multi-view clustering method for large data, characterized in that The following steps are involved: S1. Input multiple feature matrices, adjacency matrices, anchor point numbers, and cluster numbers composed of multiple view data into the database, and normalize the feature matrices; S2. Filter the feature matrix by low-pass filtering to obtain a filtered feature matrix; S3. The filtered feature matrix is used to construct an anchor matrix through the graph learning module; S4. Split the anchor matrix into a shared anchor matrix and an independent feature anchor matrix through sparse constraints; S5. construct an augmented matrix using a shared anchor matrix, and concatenate the augmented matrices to construct a normalized Laplace matrix; S6. The clustering result is obtained by performing rank constraints on the normalized Laplace matrix and connectivity constraints on the shared anchor matrix.
2. A multi-view clustering method for large data according to claim 1, characterized in that ,In step S1, the formula for normalizing the adjacency matrix is: in: A V represents the normalized adjacency matrix; in represents the input adjacency matrix, represents the element in the i-th row and j-th column of the v-th view of the adjacency matrix. If there is an edge between the i-th row and the j-th column, then If there is no edge between the i-th row and the j-th column, then D V =diag(d v1 ,...,d vN ) is the degree matrix of the adjacency matrix, and diag represents the diagonal matrix; I represents the identity matrix.
3. A multi-view clustering method for large data according to claim 1, characterized in that: In step S2, the formula for obtaining the filtered feature matrix is: in: X represents the input feature matrix; G represents low-pass filtering. The formula for calculating low-pass filtering is: G f =Up(Λ)U -1 , where U represents the orthogonal eigenvector, And because the eigenvalues λ of all normalized Laplacian matrices q All are within the interval [0,2], so the formula for calculating the low-pass filter is derived as: Where Lnorm represents the normalized Laplace matrix.
4. A multi-view clustering method for large data according to claim 3, characterized in that: In step S3, the anchor matrix construction formula is: Finally, we get v S V and an S∈R m*N , in: S V represents the anchor matrix in the vth view of the original data, and S represents the shared anchor matrix of the original data; Where BV represents the anchor matrix of the vth view, represents the element of the mth anchor point in the anchor point matrix, and BV is part of , T represents transpose; a(S V , S) represents the regularization term.
5. A multi-view clustering method for large data according to claim 4, characterized in that In step S4, the sparse constraint is to introduce a weight parameter for each view to determine the importance of different views. By learning the anchor matrix on the vth view, the anchor matrix is decomposed into a shared anchor matrix and an independent feature anchor matrix. The decomposition formula is: in: i represents the i-th vector in the corresponding matrix; B V Represents the anchor matrix of the vth view; S V represents the anchor matrix, because the shared anchor matrix and the independent feature anchor matrix satisfy so Where C represents the shared anchor matrix, H V represents the independent feature anchor matrix, F represents the norm, and λ is the weight; Through B V , C and H V Continuously update so that When the minimum value is reached, B V , C and H V is the optimal value.
6. A multi-view clustering method for large data according to claim 5, characterized in that: In step S5, the augmented matrix is constructed using the shared anchor matrix: Where Ga represents the augmented matrix, C represents the shared anchor matrix, and C T represents the transposed matrix of the shared anchor matrix; The augmented matrices are concatenated to construct a normalized Laplace matrix, and the construction formula is: Where L Ga represents the normalized Laplace matrix of Ga, D V represents the degree matrix of the augmented matrix, and I represents the identity matrix.
7. A multi-view clustering method for large data according to claim 6, characterized in that: In step S6, the augmented matrix is composed of shared anchor matrices, so they have the same number of connected components, so for L Ga The rank constraint imposed is equivalent to the connectivity constraint imposed on C. The clustering result is obtained from the number of clusters constrained by the rank constraint. The formula for the rank constraint is: rank(L Ga )=N+m-cl, where m represents the number of anchor points, cl represents the number of clusters, and N represents the number of nodes.
8. A multi-view clustering method for large data according to claim 5, characterized in that: Obtain C and H V In the formula, the variable H V , B V and C adopt the strategy of alternating update. Alternating update means that when one variable is updated, the values of other variables are fixed by default: Fixed H V and C, for B V The updates are as follows: in Tr(B vT Q V ) represents the trace of the matrix, by Q V Perform singular value decomposition to calculate the optimal B V ; Fixed B V and C, for H V The updates are as follows: in H V After updating Derivation, get each The closed-form solution of Fixed B V and H V When , the update of C is as follows: in R∈R (N+m)*cl , R represents the indicator matrix, and R T R=I cl , I cl Represents the identity matrix with length and width cl; γ represents a sufficiently large constant; R and C are updated alternately until C has exactly cl connected components: When C is fixed, the update for R is as follows: Based on Ga is composed of C and The augmented matrix composed of Where D represents the degree matrix of Ga, R represents the singular matrix, and R U represents the singular left matrix, R V represents the singular right matrix, R U ∈R N *cl , R V ∈R N*cl , and D U ∈R N*N and D V ∈R m*m It is decomposed from R and D; When the optimal R is obtained and R is fixed, the update of C is as follows: Where t represents the number of updates, c represents the elements of C, d represents the elements of D, w represents the elements of W, and r represents the elements of R.