Deep clustering method for multi-view self-representation and clustering joint optimization
By pre-training the autoencoder network and self-representation learning, combined with the joint optimization of shared self-representation and unique self-representation of the view, the problems of insufficient utilization of complementarity between views and high-dimensional noise sensitivity in the traditional multi-view clustering method are solved, and efficient and robust multi-view clustering is achieved.
Patent Information
- Application Number
- CN202510476941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional multi-view clustering methods are difficult to effectively utilize the complementarity between views, and are inefficient in processing super-large samples and are sensitive to high-dimensional noise, resulting in inadequate clustering results.
View features are extracted through pre-trained autoencoder network, combined with joint optimization of shared self-representation and unique self-representation of view, low-dimensional embedding is generated using clustering, and self-representation is optimized by two-part graph clustering algorithm, and gradually iterative updates are updated to improve clustering performance and computing efficiency.
It significantly improves the performance and computing efficiency of multi-view clustering, reduces the computational complexity in super-large sample scenarios, and enhances the robustness and accuracy of clustering results.
Smart Images

Figure CN120372339A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning and data mining, and in particular to a deep clustering method for jointly optimizing multi-view self-representation and clustering. Background Art
[0002] Multi-view clustering is a key research issue in the fields of data mining and machine learning. Its goal is to integrate the feature information of multiple heterogeneous views, mine the internal structure of the data, and achieve accurate grouping of samples. With the wide application of multi-source data in fields such as image processing, text analysis, and sensor networks, the importance of multi-view clustering has become increasingly prominent. However, actual multi-view data usually presents significant complexity: the feature distributions of different views are highly heterogeneous, high-dimensional features are redundant, and the sample size may be extremely large. These characteristics pose great challenges to traditional clustering methods.
[0003] Traditional multi-view clustering methods often have obvious deficiencies when dealing with such data. Many methods tend to directly splice multi-view features or only rely on the information of a single view, resulting in the difficulty of effectively utilizing the complementarity and diversity between views. At the same time, traditional techniques based on matrix factorization or subspace learning usually assume strong consistency between views, while ignoring the non-linear associations between heterogeneous views and the influence of high-dimensional noise, resulting in less robust clustering results. In addition, traditional clustering methods rely on the construction of a full-sample similarity matrix, and the computational complexity increases quadratically or even cubically in the case of extremely large samples, with a huge time overhead and difficulty in meeting real-time requirements. Some existing self-representation learning methods have tried to alleviate some problems through sparse representation, but most are limited to shallow feature extraction, difficult to deeply capture the complex interaction patterns between views, and sensitive to noise, easily reducing the clustering performance due to the interference of redundant features. These defects limit the application effect of traditional methods in actual multi-view data analysis, and there is an urgent need for more effective solutions. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems that traditional multi-view clustering methods are difficult to effectively utilize the complementarity between views, have low efficiency in processing extremely large samples, and are sensitive to high-dimensional noise. A deep clustering method for jointly optimizing multi-view self-representation and clustering is provided. By pre-training an autoencoder network to extract view features, combining the joint optimization of shared self-representation and view-unique self-representation, using clustering to generate low-dimensional embeddings, and finally achieving efficient multi-view clustering. This method generates anchor points in the initialization stage, balances self-representation, consistency, diversity, and reconstruction loss through the objective function, and significantly improves the clustering performance and computational efficiency.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: A deep clustering method for jointly optimizing multi-view self-representation and clustering, comprising the following steps:
[0006] S1: Obtain multiple views, and take each data point in each view as a sample. Use the VDA algorithm to select the most representative points in each view as anchor points;
[0007] S2: Use the anchor points selected in step S1 and all samples of multiple views to construct and pre-train an autoencoder network that only contains reconstruction loss;
[0008] S3: After completing pre-training and obtaining the initial weights of the autoencoder network, introduce a self-representation module on the basis of this network to simultaneously learn shared self-representation and view-unique self-representation; Use the pre-trained autoencoder network to encode all samples and anchor points to obtain the representation in the latent space; For each view, construct a similarity matrix between samples and anchor points respectively, and decompose it into shared self-representation and view-unique self-representation; The shared self-representation is responsible for extracting the consistent structure between views, while the view-unique self-representation extracts the diversity information between views; The shared self-representation and view-unique self-representation learned by this self-representation module will play a key role in the subsequent clustering steps;
[0009] S4: Use the shared self-representation and view-unique self-representation obtained in step S3 to construct a comprehensive self-representation required for bipartite graph clustering, construct a bipartite graph affinity matrix according to the comprehensive self-representation, and use the bipartite graph clustering algorithm to obtain the clustering results of samples and anchor points;
[0010] S5: Use the clustering results of samples and anchor points obtained in step S4 to construct a distance representation of samples and anchor points, and feedback the distance representation to the self-representation module to iteratively correct the shared self-representation and view-unique self-representation. Through multiple iterations, the clustering results and the self-representation module promote each other, and gradually converge the shared self-representation and view-unique self-representation;
[0011] S6: After the iteration is completed, combine the finally obtained shared self-representation and view-unique self-representation to obtain a comprehensive self-representation, and use the comprehensive self-representation for spectral clustering to obtain the final clustering result.
[0012] Furthermore, the specific operation steps of step S1 are as follows:
[0013] S11: Obtain multiple views from different sources, and each view is represented by a matrix where X v represents the matrix representation of the v-th view, m represents the number of views, n represents the number of samples, represents the set of real numbers, and d v represents the feature dimension of the v-th view; In this matrix, each row corresponds to a sample, and each column corresponds to a feature dimension;
[0014] S12: Use the VDA algorithm to select the most representative anchor points from each view. The specific process is as follows:
[0015] First, splice each view in the feature dimension to form:
[0016]
[0017] In the formula, d represents the sum of all view dimensions, X represents the matrix representation after splicing all view dimensions, each row of X represents a sample, and each column corresponds to the feature dimension after splicing all views;
[0018] Next, calculate the sample variance for each sample of X to form a vector Q:
[0019] Q = [u1, u2,..., u n
[0020] In the formula, u n represents the variance of the nth sample;
[0021] After normalizing Q, iteratively select the sample corresponding to the maximum value as the anchor point;
[0022] Repeat the above process until a sufficient number of anchor points are obtained, and combine these anchor points into:
[0023]
[0024] In the formula, t is the number of anchor points, A v represents the set of anchor points for the vth view, represents the feature vector of the tth anchor point under the vth view; in this way, the most representative set of anchor points A v is obtained in each view.
[0025] Furthermore, the specific operation steps of step S2 are as follows:
[0026] S21: Build and initialize an autoencoder network to capture the non-linear features of samples in multiple views; the autoencoder network includes an encoder and a decoder; let W e represent the encoder weights, W d represent the decoder weights, and the hidden layer size is h; for the vth view, the input feature dimension is d v ; in the initialization stage, use the random normal distribution to assign initial values to W e and W d ;
[0027] S22: After completing the construction of the autoencoder network, input the matrix and the set of anchor points into the autoencoder network for forward propagation and backward propagation; the encoder passes through:
[0028]
[0029] Obtain the latent representation of the samples and its anchor latent representation where ReLU is the activation function, and W e maps the input from dimension d v to dimension h; subsequently, the decoder passes through:
[0030]
[0031] In the formula, represents the reconstructed sample, and represents the reconstructed anchor;
[0032] restore the latent representations of the samples and anchors back to the input dimension respectively;
[0033] By minimizing the reconstruction loss L rec :
[0034]
[0035] In the formula, ||·|| F represents the Frobenius norm, and thus W e and W d can be updated until convergence; finally, the weights W e and W d of the pre-trained autoencoder network can be obtained.
[0036] Furthermore, the specific operation steps of step S3 are as follows:
[0037] S31: After obtaining the weights W e and W d of the autoencoder network through pre-training, encode all samples and anchors through the autoencoder network to obtain the latent representations of the samples and anchors where represents the latent representation of the samples of the v-th view, and represents the latent representation of the anchors of the v-th view;
[0038] S32: Use the latent representations of the samples and anchors to construct the similarity matrix of the samples and anchors, and decompose the similarity matrix of the samples and anchors into two parts: one part is the shared self-representation for capturing the consistency information between views and the other part is the view-unique self-representation for capturing the diversity information between views The self-representation module obtains the latent representation Z of the samples by combining the anchor latent representation with two self-representation matrices v :
[0039] Z v ≈H v (U + S v )T
[0040] where U+S v acts as the coefficient matrix, linearly combining the anchor point latent representations to obtain the sample latent representations; the loss function L of the self-representation module self is as follows:
[0041] L self =||Z v -H v (U+S v ) T ||
[0042] This self-representation module takes into account the consistency and diversity of multiple views at the subspace level, providing richer and more accurate similarity information for clustering;
[0043] S33: To ensure the block structure of the obtained shared self-representation and view-unique self-representations and the diversity of the view-unique self-representations, appropriate regularization terms need to be introduced into the above loss function, including: ||U|| 2,1 and ∑ v ||S v || 2,1 The norm constraint is used to ensure row sparsity, promoting the block structure of the shared self-representation and view-unique self-representations, ||S v ⊙S w ||0 is used to encourage the view-unique self-representations to be as dissimilar as possible at the element level, where v≠w ensures different views; since the zero norm is difficult to optimize in the actual operation process, it is further relaxed to the one norm, obtaining the following formula:
[0044] ||S v ⊙S w ||1
[0045] Combining the above regularization terms with the reconstruction loss of the autoencoder network in the previous stage, the designed loss L of the shared and view-unique self-representation module can be obtained self_all , which is specifically as follows:
[0046]
[0047] By jointly iteratively updating the variables W e , W d , U, S v , it is possible to fully exploit the multi-view diversity while maintaining the multi-view Figure 1 consistency, thereby constructing a more robust self-representation module.
[0048] Furthermore, the specific operation steps of step S4 are:
[0049] S41: Based on the shared self-representation and view-unique self-representation learned in step S3, construct a comprehensive self-representation B for bipartite graph clustering. The definition of the comprehensive self-representation is as follows:
[0050]
[0051] In the formula, represents taking the average over the number of views; Integrate the similarities between samples and anchors. The similarity between the i-th sample and the j-th anchor is denoted as B ij ; Through the above comprehensive self-representation B, not only the shared similarity information brought by the shared self-representation U is considered, but also the view-unique similarity information of each view's unique self-representation S v is incorporated;
[0052] S42: After obtaining the comprehensive self-representation B in step S41, construct a bipartite graph affinity matrix P in the following way:
[0053]
[0054] In the formula, 0 indicates that the corresponding block is a zero matrix. Use the bipartite graph clustering algorithm. Define the degree matrix D as a diagonal matrix, and the element in the i-th row and j-th column represents the sum of the edge weights between the i-th node and the remaining nodes; Construct the normalized Laplacian matrix according to the degree matrix D
[0055]
[0056] S43: After obtaining the normalized Laplacian matrix , use the SVD decomposition algorithm to perform eigen-decomposition on , and take the eigenvectors corresponding to the first c smallest non-zero eigenvalues as the low-dimensional representations of samples and anchors, denoted as matrix F:
[0057]
[0058] Among them, the first n rows of matrix F represent the low-dimensional representations of samples and the last t rows represent the low-dimensional representations of anchors After obtaining the low-dimensional representation F1 of samples, use the K-means clustering algorithm to cluster the n samples to obtain preliminary clustering labels;
[0059] Through the above step S4, the multi-view shared self-representation and view-unique self-representation are combined to form a comprehensive self-representation, and the comprehensive self-representation is closely combined with bipartite graph clustering to obtain the low-dimensional representations and clustering results of samples and anchors, which are used to guide the update and correction of the shared self-representation and view-unique self-representation in step S5.
[0060] Further, the specific operation steps of step S5 are as follows:
[0061] S51: According to the clustering result obtained in step S4, construct the distance representation between the sample and the anchor point using the low-dimensional representation F1 of the sample and the low-dimensional representation F2 of the anchor point, that is, the distance matrix Θ. Let represent the coordinates of the i-th sample in the low-dimensional space, represent the coordinates of the j-th anchor point in the low-dimensional space. Then, the distance matrix Θ between the i-th sample and the j-th anchor point is ij defined as:
[0062]
[0063] In the formula, represents the square of the Euclidean distance; when Θ ij is less than the preset threshold, it means that the i-th sample and the j-th anchor point are closer in this low-dimensional space, and vice versa;
[0064] S52: After obtaining the distance representation between the sample and the anchor point, feedback this distance representation to the self-representation module to correct the shared self-representation and the view-unique self-representation. Specifically: introduce the ||Θ⊙B||0 zero-norm constraint to align the distance matrix with the comprehensive self-representation. In the actual operation process, since the zero-norm is difficult to optimize, it is further relaxed to the one-norm, which is specifically expressed as:
[0065] L refine = ||Θ⊙B||1
[0066] In the formula, ⊙ is the Hadamard product, ||·||1 represents the sum of the absolute values of the elements. When the distance is large, the corresponding B should be close to zero, and when the distance is small, the corresponding B should not be zero; this can make the comprehensive self-representation B match the distance matrix Θ between the sample and the anchor point;
[0067] S53: Combining all the losses mentioned above, the total loss function L is:
[0068] L = L rec + αL self-all + βL refine
[0069] In the formula, both α and β are hyperparameters, and different hyperparameters are selected according to different datasets during the experimental setup to complete the clustering algorithm;
[0070] S54: Iterate the process of steps S31 - S53, so that the clustering result, the shared self-representation, and the view-unique self-representation are alternately updated, promoting each other and gradually converging. Finally, output the converged shared self-representation U and view-unique self-representation S v .
[0071] Furthermore, the specific operation steps of step S6 are as follows:
[0072] S61: Based on the shared self-representation U and the view-unique self-representation S output in step S54 v construct a comprehensive self-representation B for bipartite graph clustering, and construct a new bipartite graph affinity matrix P' according to step S42. Subsequently, calculate the Laplacian matrix of P'
[0073] S62: On the basis of what is obtained in step S61 use the sample low-dimensional representation F1 obtained in step S43, and then use the low-dimensional representation to perform K-means clustering on n samples to obtain a clustering label vector Y ∈ R n , where the i-th element represents the cluster index to which the i-th sample belongs. Thus, the final clustering result for multiple views is obtained.
[0074] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0075] 1. By pre-training an autoencoder network and self-representation learning, the sensitivity of shallow methods to high-dimensional noise and non-linear features is overcome, and the utilization efficiency of complementarity and diversity between views is improved.
[0076] 2. The anchor points and the optimization of the objective function significantly reduce the computational complexity in the scenario of extremely large samples. Compared with the quadratic or cubic time overhead of traditional spectral clustering, the method of the present invention is more efficient. At the same time, considering the Figure 1 consistency and diversity constraints enhances the learning of self-representations.
[0077] 3. Jointly optimize the graph representation and clustering to enhance the robustness and accuracy of the clustering results.
[0078] In summary, the present invention effectively solves the limitations of traditional methods in multi-view data processing, performs excellently in scenarios with strong heterogeneity and large sample sizes, provides an efficient and robust solution for multi-view clustering tasks, and has significant theoretical significance and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a flow framework diagram of the method of the present invention.
[0080] Figure 2 is a flow chart of the anchor point selection algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0081] The present invention will be described in further detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0082] As Figure 1 and Figure 2As shown in the figure, this embodiment discloses a deep clustering method for joint optimization of multi-view self-representation and clustering, and the specific situation is as follows:
[0083] S1: Obtain multiple views, and take each data point in each view as a sample. Use the VDA algorithm to select the most representative points in each view as anchor points. The specific operation steps are as follows:
[0084] S11: Obtain multiple views from different sources. Each view is represented by a matrix where X v represents the matrix representation of the v-th view, m represents the number of views, n represents the number of samples, represents the set of real numbers, and d v represents the feature dimension of the v-th view; in this matrix, each row corresponds to a sample, and each column corresponds to a feature dimension;
[0085] S12: Use the VDA algorithm to select the most representative anchor points from each view. The specific process is as follows:
[0086] First, concatenate each view in the feature dimension to get:
[0087]
[0088] where d represents the sum of all view dimensions, X represents the matrix representation after concatenating all view dimensions, each row of X represents a sample, and each column corresponds to the feature dimension after concatenating all views;
[0089] Next, calculate the sample variance for each sample in X to form a vector Q:
[0090] Q = [u1, u2,..., u n
[0091] where u n represents the variance of the n-th sample;
[0092] After normalizing Q, iteratively select the sample corresponding to the maximum value as the anchor point;
[0093] Repeat the above process until enough anchor points are obtained. Combine these anchor points into:
[0094]
[0095] where t is the number of anchor points, A v represents the set of anchor points for the v-th view, represents the feature vector of the t-th anchor point under the v-th view; in this way, the set of the most representative anchor points A v is obtained in each view.
[0096] S2: Using the anchor points selected in step S1 and all samples of multiple views, construct and pre-train an autoencoder network that only contains the reconstruction loss. The specific operation steps are as follows:
[0097] S21: Build and initialize the autoencoder network to capture the non-linear features of samples in multiple views; the autoencoder network includes an encoder and a decoder; let W e represent the encoder weights, and W d represent the decoder weights. The hidden layer size is h; for the v-th view, the input feature dimension is d v ; in the initialization stage, use the random normal distribution to assign initial values to W e and W d ;
[0098] S22: After completing the construction of the autoencoder network, input the matrix and the anchor point set into the autoencoder network for forward propagation and backward propagation; the encoder passes through:
[0099]
[0100] to obtain the sample latent representation and its anchor point latent representation where ReLU is the activation function, and W e maps the input from d v dimensions to h dimensions; then the decoder passes through:
[0101]
[0102] In the formula, represents the reconstructed sample, represents the reconstructed anchor point;
[0103] Restore the latent representations of the sample and the anchor point back to the input dimension respectively;
[0104] By minimizing the reconstruction loss L rec :
[0105]
[0106] In the formula, ||·|| F represents the Frobenius norm, and then W e and W d can be updated until convergence; finally, the weights W e and W d of the pre-trained autoencoder network can be obtained.
[0107] S3: After completing pre-training and obtaining the initial weights of the autoencoder network, introduce a self-representation module on the basis of this network to simultaneously learn shared self-representations and view-unique self-representations; use the pre-trained autoencoder network to encode all samples and anchors to obtain representations in the latent space; for each view, construct a similarity matrix between samples and anchors respectively, and decompose it into shared self-representations for extracting the consistent structure between views and view-unique self-representations for extracting the diverse information between views; the shared self-representations and view-unique self-representations learned by this self-representation module will play a key role in the subsequent clustering steps; the specific operation steps are as follows:
[0108] S31: After completing pre-training to obtain the weights W of the autoencoder network e and W d then, use the autoencoder network to encode all samples and anchors to obtain the latent representations of samples and anchors where represents the latent representation of samples in the v-th view, represents the latent representation of anchors in the v-th view;
[0109] S32: Use the latent representations of samples and anchors to construct a similarity matrix between samples and anchors, and decompose the similarity matrix of samples and anchors into two parts: one part is the shared self-representation for capturing the consistent information between views the other part is the view-unique self-representation for capturing the diverse information between views The self-representation module obtains the sample latent representation Z by combining the anchor latent representation with two self-representation matrices v :
[0110] Z v ≈H v (U + S v ) T
[0111] where U + S v acts as a coefficient matrix to linearly combine the anchor latent representation to obtain the sample latent representation; the loss function L self of the self-representation module is as follows:
[0112] L self = ||Z v - H v (U + S v ) T ||
[0113] This self-representation module takes into account the consistency and diversity of multiple views at the subspace level, providing richer and more accurate similarity information for clustering;
[0114] S33: To ensure the block structure of the obtained shared self-representation and view-unique self-representation, as well as the diversity of the view-unique self-representation, appropriate regularization terms need to be introduced into the above loss function, including: ||U|| 2,1 and ∑ v ||S v || 2,1 The norm constraint is used to ensure row sparsity and promote the block structure of the shared self-representation and view-unique self-representation. ||S v ⊙S w ||0 is used to encourage the view-unique self-representation to be as dissimilar as possible at the element level, where v≠w ensures different views; since the zero norm is difficult to optimize in the actual operation process, it is further relaxed to the one norm, and the following formula is obtained:
[0115] ||S v ⊙S w ||1
[0116] Combining the above regularization terms with the reconstruction loss of the autoencoder network in the previous stage, the loss L of the designed shared and view-unique self-representation module can be obtained self_all , as follows:
[0117]
[0118] Jointly iteratively updating the variables W e , W d , U, S v can fully exploit the multi-view diversity while maintaining multi-view Figure 1 consistency, thereby constructing a more robust self-representation module.
[0119] S4: Using the shared self-representation and view-unique self-representation obtained in step S3, construct the comprehensive self-representation required for bipartite graph clustering, construct the bipartite graph affinity matrix based on the comprehensive self-representation, and obtain the clustering results of samples and anchors using the bipartite graph clustering algorithm. The specific operation steps are as follows:
[0120] S41: Construct a comprehensive self-representation B for bipartite graph clustering based on the shared self-representation and view-unique self-representation learned in step S3. The definition of the comprehensive self-representation is:
[0121]
[0122] In the formula, represents taking the average over the number of views; Integrate the similarity correlations between samples and anchors. The similarity between the i-th sample and the j-th anchor is represented as B ij; Through the above comprehensive self-representation B, not only the shared similarity information brought by the shared self-representation U is considered, but also the view-unique similarity information of each view is integrated v of the view-unique similarity information;
[0123] S42: After obtaining the comprehensive self-representation B in step S41, construct a bipartite graph affinity matrix P in the following way:
[0124]
[0125] where 0 indicates that the corresponding block is a zero matrix, Using the bipartite graph clustering algorithm, define the degree matrix D as a diagonal matrix, and the element in the i-th row and j-th column represents the sum of the edge weights of the i-th node and the remaining nodes; construct a normalized Laplacian matrix according to the degree matrix D
[0126]
[0127] S43: After obtaining the normalized Laplacian matrix , use the SVD decomposition algorithm to perform eigenvalue decomposition on , and take the eigenvectors corresponding to the first c smallest non-zero eigenvalues as the low-dimensional representations of the samples and anchors, denoted as matrix F:
[0128]
[0129] where the first n rows of matrix F represent the low-dimensional representations of the samples and the last t rows represent the low-dimensional representations of the anchors After obtaining the low-dimensional representation F1 of the samples, use the K-means clustering algorithm to cluster the n samples to obtain the preliminary clustering labels;
[0130] Through the above step S4, the multi-view shared self-representation and view-unique self-representation are combined to form a comprehensive self-representation, and the comprehensive self-representation is closely combined with bipartite graph clustering to obtain the low-dimensional representations and clustering results of the samples and anchors, which are used to guide the update and correction of the shared self-representation and view-unique self-representation in step S5.
[0131] S5: Use the clustering results of the samples and anchors obtained in step S4 to construct the distance representations of the samples and anchors, and feedback the distance representations to the self-representation module to iteratively correct the shared self-representation and view-unique self-representation. Through the mutual promotion of the clustering results and the self-representation module through multiple iterations, the shared self-representation and view-unique self-representation are gradually converged. The specific operation steps are as follows:
[0132] S51: According to the clustering results obtained in step S4, construct the distance representation of the samples and anchors, that is, the distance matrix Θ, by using the low-dimensional representation F1 of the samples and the low-dimensional representation F2 of the anchors. Let Denote the coordinates of the \(i\)-th sample in the low-dimensional space, and denote the coordinates of the \(j\)-th anchor point in the low-dimensional space. Then, the distance matrix \(\Theta\) between the \(i\)-th sample and the \(j\)-th anchor point is ij defined as:
[0133]
[0134] where, denotes the square of the Euclidean distance; when \(\Theta\) ij is less than the preset threshold, it means that the \(i\)-th sample and the \(j\)-th anchor point are closer in this low-dimensional space, and vice versa;
[0135] S52: After obtaining the distance representation between the sample and the anchor point, feedback this distance representation to the self-representation module to correct the shared self-representation and the view-unique self-representation. Specifically: introduce the \(l_0\)-norm constraint \(\|\Theta\odot B\|_0\) to align the distance matrix with the comprehensive self-representation. In the actual operation process, since the \(l_0\)-norm is difficult to optimize, it is further relaxed to the \(l_1\)-norm, which is specifically expressed as:
[0136] L refine =\(\|\Theta\odot B\|_1\)
[0137] where, \(\odot\) is the Hadamard product, \(\|\cdot\|_1\) represents the sum of the absolute values of the elements. When the distance is large, the corresponding \(B\) should be close to zero, and when the distance is small, the corresponding \(B\) should not be zero; this can make the comprehensive self-representation \(B\) match the distance matrix \(\Theta\) between the sample and the anchor point;
[0138] S53: Combining all the losses mentioned above, the total loss function \(L\) is:
[0139] L = L rec +\(\alpha L\) self-all +\(\beta L\) refine
[0140] where, both \(\alpha\) and \(\beta\) are hyperparameters, and different hyperparameters are selected according to different datasets during the experimental setup to complete the clustering algorithm;
[0141] S54: Iterate the process of S31 - S53, so that the clustering results, the shared self-representation, and the view-unique self-representation are alternately updated, promoting each other and gradually converging. Finally, output the converged shared self-representation \(U\) and view-unique self-representation \(S\) v .
[0142] S6: After the iteration is completed, combine the finally obtained shared self-representation and view-unique self-representation to obtain the comprehensive self-representation, and use the comprehensive self-representation for spectral clustering to obtain the final clustering result. The specific operation steps are:
[0143] S61: Based on the shared self-representation U and view-unique self-representation S output in step S54 v Construct a comprehensive self-representation B for bipartite graph clustering, and construct a new bipartite graph affinity matrix P' according to step S42. Subsequently, calculate the Laplacian matrix of P'
[0144] S62: Obtained in step S61 On this basis, use the sample low-dimensional representation F1 obtained in step S43. Then use the low-dimensional representation to perform K-means clustering on n samples to obtain a clustering label vector Y ∈ R n , where the i-th element represents the cluster index to which the i-th sample belongs. Thus, the final clustering result for multi-views is obtained.
[0145] The above embodiments are preferred embodiments of the present invention. However, the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A deep clustering method for joint optimization of multi-view self-representation and clustering, characterized in that, It includes the following steps: S1: Obtain multiple views, and take each data point in each view as a sample. Use the VDA algorithm to select the most representative points in each view as anchor points; S2: Use the anchor points selected in step S1 and all samples of multiple views to construct and pre-train an autoencoder network that only contains reconstruction loss; S3: After completing pre-training and obtaining the initial weights of the autoencoder network, introduce a self-representation module on the basis of this network to simultaneously learn shared self-representation and view-unique self-representation; use the pre-trained autoencoder network to encode all samples and anchor points to obtain the representation in the latent space; for each view, construct the similarity matrix between samples and anchor points respectively, and decompose it into shared self-representation and view-unique self-representation; the shared self-representation is responsible for extracting the consistent structure between views, and the view-unique self-representation extracts the diversity information between views; the shared self-representation and view-unique self-representation learned by this self-representation module will play a key role in the subsequent clustering steps; S4: Use the shared self-representation and view-unique self-representation obtained in step S3 to construct the comprehensive self-representation required for bipartite graph clustering, construct the bipartite graph affinity matrix according to the comprehensive self-representation, and use the bipartite graph clustering algorithm to obtain the clustering results of samples and anchor points; S5: Use the clustering results of samples and anchor points obtained in step S4 to construct the distance representation of samples and anchor points, feedback the distance representation to the self-representation module, iteratively correct the shared self-representation and view-unique self-representation, and promote each other between the clustering results and the self-representation module through multiple iterations, gradually converging the shared self-representation and view-unique self-representation; S6: After the iteration is completed, combine the finally obtained shared self-representation and view-unique self-representation to obtain the comprehensive self-representation, and use the comprehensive self-representation for spectral clustering to obtain the final clustering result.
2. The deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 1, wherein The specific operation steps of step S1 are as follows: S11: Obtain multiple views from different sources, each view represented by a matrix v = [1, 2,..., m], where X v represents the matrix representation of the v-th view, m represents the number of views, n represents the number of samples, represents the set of real numbers, and d v represents the feature dimension of the v-th view; in this matrix, each row corresponds to a sample and each column corresponds to a feature dimension; S12: Use the VDA algorithm to select the most representative anchor points from each view. The specific process is as follows: First, splice each view in the feature dimension into: where d represents the sum of all view dimensions, X represents the matrix representation after splicing all view dimensions, each row of X represents a sample, and each column corresponds to the feature dimension after splicing all views; Next, calculate the sample variance for each sample of X to form a vector Q: Q = [u1, u2,..., u n where u n represents the variance of the nth sample; After normalizing Q, iteratively select the sample corresponding to the maximum value as the anchor point; Repeat the above process until enough anchor points are obtained, and combine these anchor points into: where \(t\) is the number of anchor points, \(A\) v represents the set of anchor points of the \(v\)-th view, and represents the feature vector of the \(t\)-th anchor point under the \(v\)-th view; thus, the most representative set of anchor points \(A\) is obtained in each view v .
3. A deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 2, characterized in that The specific operation steps of step S2 are as follows: S21: Build and initialize an autoencoder network to capture the non - linear features of samples in multiple views; the autoencoder network includes an encoder and a decoder; let W e represent the encoder weights, and W d represent the decoder weights, with the hidden layer size being h; for the v - th view, the input feature dimension is d v ; in the initialization stage, use a random normal distribution to initialize the values of W e and W d ; S22: After completing the construction of the autoencoder network, input the matrix and the anchor point set into the autoencoder network for forward propagation and backward propagation; the encoder passes through: Obtain the latent representation of the sample and its anchor latent representation where ReLU is the activation function, and W e maps the input from dimension d v to dimension h; subsequently, the decoder passes through: In the formula, represents the reconstructed sample, represents the reconstructed anchor point; Restore the latent representations of samples and anchor points back to the input dimension respectively; By minimizing the reconstruction loss L rec : where ||·|| F denotes the Frobenius norm, and thus W e and W d can be updated until convergence; finally, the weights W e and W d of the pre-trained autoencoder network can be obtained.
4. A deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 3, characterized in that The specific operation steps of step S3 are as follows: S31: After obtaining the weights W of the autoencoder network through pre-training e and W d all samples and anchors are encoded through the autoencoder network to obtain the latent representations of the samples and anchors where represents the latent representation of the samples of the v-th view, represents the latent representation of the anchors of the v-th view; S32: Construct a similarity matrix of samples and anchors using the sample and anchor latent representations, and decompose the similarity matrix of samples and anchors into two parts: one part is the shared self-representation for capturing inter-view consistency information and the other part is the view-unique self-representation for capturing inter-view diversity information The self-representation module obtains the sample latent representation Z by combining the anchor latent representation with two self-representation matrices v : Z v ≈H v (U+S v ) T where U+S v acts as a coefficient matrix to linearly combine the anchor point latent representations to obtain the sample latent representations; the loss function L of the self-representation module self is as follows: L self = ||Z v - H v (U + S v ) T || This self-representation module takes into account the consistency and diversity of multiple views at the subspace level, providing richer and more accurate similarity information for clustering; S33: To ensure the block structure of the obtained shared self-representation and view-unique self-representations, as well as the diversity of view-unique self-representations, appropriate regularization terms need to be introduced into the above loss function, including: ||U|| 2,1 and ∑ v ||S v || 2,1 The norm constraint is used to ensure row sparsity and promote the block structure of shared self-representations and view-unique self-representations, ||S v ⊙S w ||0 is used to encourage view-unique self-representations to be as dissimilar as possible at the element level, where v≠w ensures different views; since the zero norm is difficult to optimize in actual operations, it is further relaxed to the first norm, resulting in the following formula: ||S v ⊙S w ||1 Combining the above regular terms with the reconstruction loss of the autoencoder network in the previous stage, the loss L of the designed shared and view-unique self-representation module can be obtained, which is specifically as follows: self_all , as follows: For variable W e and W d , U, S v Performing joint iterative updates can fully exploit multi-view diversity while maintaining multi-view consistency, thereby constructing a more robust self-representation module.
5. A deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 4, characterized in that The specific operation steps of step S4 are as follows: S41: Construct a comprehensive self-representation B for bipartite graph clustering based on the shared self-representation and view-unique self-representation learned in step S3. The definition of the comprehensive self-representation is: In the formula, represents the average over the number of views; Integrate the similarity associations between the samples and the anchors. The similarity between the i-th sample and the j-th anchor is denoted as B ij ; Through the above comprehensive self-representation B, not only the shared similarity information brought by the shared self-representation U is considered, but also the view-unique similarity information of each view's unique self-representation S v is integrated; S42: After obtaining the comprehensive self-representation B in step S41, construct the bipartite graph affinity matrix P in the following way: In the formula, 0 indicates that the corresponding block is a zero matrix. Using the bipartite graph clustering algorithm, define the degree matrix D as a diagonal matrix, where the element in the i-th row and j-th column represents the sum of the edge weights between the i-th node and the remaining nodes. Construct the normalized Laplacian matrix according to the degree matrix D S43: After obtaining the normalized Laplacian matrix perform eigen-decomposition on it using the SVD decomposition algorithm and take the eigenvectors corresponding to the first c smallest non-zero eigenvalues as the low-dimensional representations of the samples and the anchors, denoted as matrix F: Among them, the first n rows of the matrix F represent the low-dimensional representation of the samples The last t rows represent the low-dimensional representation of the anchor points After obtaining the low-dimensional representation F1 of the samples, use the K-means clustering algorithm to cluster the n samples to obtain the preliminary clustering labels; Through the above step S4, the multi-view shared self-representation and the view-unique self-representation are combined to form a comprehensive self-representation, and the comprehensive self-representation is closely combined with the bipartite graph clustering to obtain the low-dimensional representations and clustering results of the samples and the anchors, which are used to guide the update and correction of the shared self-representation and the view-unique self-representation in step S5.
6. A deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 5, characterized in that The specific operation steps of step S5 are as follows: S51: According to the clustering result obtained in step S4, a distance representation between the sample and the anchor point is constructed from the low-dimensional representation F1 of the sample and the low-dimensional representation F2 of the anchor point, that is, the distance matrix Θ. Let represent the coordinates of the i-th sample in the low-dimensional space, represent the coordinates of the j-th anchor point in the low-dimensional space. Then, the distance matrix Θ between the i-th sample and the j-th anchor point is ij defined as: In the formula, represents the square of the Euclidean distance; when Θ ij is less than the preset threshold, it means that the i-th sample is closer to the j-th anchor point in this low-dimensional space, and vice versa, it is farther. S52: After obtaining the distance representation of the samples and the anchors, feedback this distance representation to the self-representation module to correct the shared self-representation and the view-unique self-representation. Specifically: introduce the ||Θ⊙B||0 zero-norm constraint to align the distance matrix with the comprehensive self-representation. In the actual operation process, since the zero-norm is difficult to optimize, it is further relaxed to the one-norm, which is specifically expressed as: L refine = ||Θ⊙B||1 In the formula, ⊙ is the Hadamard product, ||·||1 represents the sum of the absolute values of the elements. When the distance is large, the corresponding B should be close to zero; when the distance is small, the corresponding B should not be zero. This can make the comprehensive self-representation B match the distance matrix Θ of the samples and the anchors. S53: Combining all the losses mentioned above, the total loss function L is obtained as: L = L rec + αL self-all + βL refine In the formula, both α and β are hyperparameters, and different hyperparameters are selected according to different data sets during the experimental setting to complete the clustering algorithm. S54: Iterate the process of steps S31 - S53 so that the clustering results, the shared self - representation, and the view - unique self - representation are alternately updated, promoting each other and gradually converging. Finally, output the converged shared self - representation U and view - unique self - representation S v 。 7. A deep clustering method for joint optimization of multi-view self-representation and clustering according to claim 6, characterized in that The specific operation steps of step S6 are as follows: S61: Based on the shared self-representation U and the view-unique self-representation S output in step S54 v Construct a comprehensive self-representation B for bipartite graph clustering, and construct a new bipartite graph affinity matrix P' according to step S42, and then calculate the Laplacian matrix of P' S62: Based on what is obtained in step S61 , the sample low-dimensional representation F1 obtained in step S43 is used. Then, K-means clustering is performed on the n samples using the low-dimensional representation to obtain the clustering label vector Y ∈ R n , where the i-th element represents the cluster index to which the i-th sample belongs. Thus, the final clustering result for the multi-view is obtained.
Citation Information
Cited By
Multi-dimensional data processing method and system corresponding to power load prediction
CN121233943A