Adaptive graph clustering method and device based on robust symmetric non-negative matrix factorization
Through the adaptive graph clustering method of robust symmetric non-negative matrix factorization, combined with integrated clustering and multi-round iterative optimization, the initialization sensitivity and noise sensitivity problems of the symmetric non-negative matrix factorization algorithm are solved, the clustering performance and adaptability are improved, and more accurate clustering results are achieved.
Patent Information
- Application Number
- CN202210695708.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-06-20
AI Technical Summary
Existing symmetric non-negative matrix factorization algorithms are mathematically non-convex optimization problems. They are sensitive to initialization and are sensitive to noise and outliers, resulting in poor clustering performance. They also lack local geometric information and have difficulty reconstructing the data space in a low-dimensional space.
An adaptive graph clustering method based on robust symmetric non-negative matrix factorization is adopted. By integrating clustering and multi-round iterative optimization, combined with L2,p norm and three Laplace matrix construction methods, the robustness and adaptability are enhanced and the clustering results are optimized.
The robustness and adaptability of the clustering algorithm are improved, the accuracy and scalability of the clustering results are improved, and it can better handle high-dimensional data containing noise and outliers.
Smart Images

Figure CN115062707B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to an adaptive graph clustering method and device based on robust symmetric non-negative matrix decomposition. Background Art
[0002] Non-negative Matrix Factorization (NMF) is a commonly used data representation method and clustering technique and is widely used in data mining and machine learning. When the data distribution has a linear structure, NMF can achieve good clustering performance. However, NMF cannot use the nonlinear structure of the input data to produce clustering results. Symmetric Non-negative Matrix Factorization (SNMF) is a special type of constrained NMF, which can be regarded as a graph clustering algorithm. It can decompose nonlinear data and directly generate clustering indicators. Compared with spectral clustering (SC), symmetric non-negative matrix factorization clustering performs better. At the same time, SNMF is more flexible in the choice of similarity measurement between samples. The graph-based SNMF model contains a basic manifold structure.
[0003] However, while SNMF has achieved good performance in most clustering tasks, it still has some limitations. One major issue is that SNMF is mathematically formulated as a non-convex optimization problem and is sensitive to variable initialization. The quality of the initialization matrix will seriously affect its clustering performance. Another major issue is that the objective function defined by the Frobenius norm (i.e., the L2 loss function) is sensitive to noise and outliers. However, in practical applications, most real data contains noise and outliers. The square of the Frobenius norm reconstruction error can cause the solution to deviate from the true value when measuring outliers, reducing the robustness of the algorithm. In addition, a single graph has defects when representing the manifold: its topological information is incomplete, and the local geometric structure is insufficient to reconstruct the data space in a low-dimensional space. Combining local and global geometric information can provide a more comprehensive understanding of the original data space. Moreover, when dealing with high-dimensional data containing noise and outliers, robustness, sensitivity to variable initialization, and adaptive graph learning are not considered simultaneously, resulting in poor clustering results.
[0004] To address the above issues, no effective solutions have been proposed so far. Summary of the Invention
[0005] The purpose of this application is to provide an adaptive graph clustering method and device based on robust symmetric non-negative matrix decomposition, which can improve the algorithm clustering performance and adaptability to data with different structures, thereby improving the accuracy of clustering results.
[0006] The present application provides an adaptive graph clustering method and apparatus based on robust symmetric non-negative matrix factorization, which is implemented as follows:
[0007] An adaptive graph clustering method based on robust symmetric non-negative matrix factorization, comprising:
[0008] Acquire a target data set, wherein the target data set includes a plurality of sample vectors;
[0009] Recalling a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, clustering multiple sample vectors in the target data set, and obtaining multiple groups of clustering results for the multiple sample vectors;
[0010] The multiple clustering results are processed through integrated clustering to obtain the optimal clustering result of the target data set.
[0011] In one embodiment, a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization is called, and multiple sample vectors in the target data set are clustered to obtain multiple groups of clustering results for the multiple sample vectors, including:
[0012] Generate an empirical similarity matrix based on the target dataset;
[0013] Repeat the following operations to obtain multiple cluster partition matrices: generate a set of random non-negative matrices based on the empirical similarity matrix; obtain a cluster partition matrix corresponding to the random non-negative matrix based on the random non-negative matrix;
[0014] Obtaining a similarity matrix after reconstructing the target data set according to the multiple clustering matrices;
[0015] The similarity matrix after reconstructing the target data set is decomposed by the robust symmetric non-negative matrix in the objective function to obtain multiple groups of clustering results.
[0016] In one embodiment, the plurality of clustering results are processed by ensemble clustering to obtain an optimal clustering result for the target data set, including:
[0017] assigning a weight to each of the plurality of clustering results by automatic weighting;
[0018] Multiple rounds of alternating iterative optimization are performed to obtain the optimal clustering result of the target data set.
[0019] In one embodiment, performing multiple rounds of alternating iterative optimization to obtain the optimal clustering result of the target data set includes:
[0020] Perform a preset number of iterative optimization rounds on each parameter in the objective function in the following manner to obtain the optimal clustering result of the target data set:
[0021] Selecting a parameter from the objective function as a target parameter for current iterative optimization;
[0022] The preset updating rule is called to iteratively optimize the target parameter. During the iterative optimization of the target parameter, other parameters except the target parameter are fixed to remain unchanged.
[0023] In one embodiment, the objective function is:
[0024]
[0025]
[0026] Where W is the reconstructed similarity matrix that measures the similarity between the i-th and j-th samples of X. is the target dataset, n is the number of samples, m is the feature dimension of the target dataset, is the clustering result of the target data set, H r The position of the maximum value in each row indicates the sample X i The cluster members, c represents the number of clusters, α r is the weight vector The rth element of is used to balance the contribution of each group of clustering results. Represents a full 1 vector, used to ensure that each α r are all effective weights, λ≥0 is the regularization parameter, For use in H r The low-dimensional representation H ri and H rj Keep X i and X j Local information with similar features, L represents the Laplacian matrix, ||·|| 2,p L represents the matrix 2,p norm, 0<p≤1, tr(·) means finding the trace of the matrix, and T means finding the transpose of the matrix.
[0027] In one embodiment, the Laplacian matrix constructed in the objective function is at least one of the following:
[0028] Laplacian matrix:
[0029]
[0030] Normalized Laplacian matrix:
[0031]
[0032] Shifted Laplace matrix:
[0033]
[0034] Where D is the diagonal matrix,
[0035] In one embodiment, the preset update rule is:
[0036] The Laplace matrix is In the case of H r The update formula is:
[0037]
[0038] The Laplace matrix is In the case of H r The update formula is:
[0039]
[0040] The Laplace matrix is In the case of H r The update formula is:
[0041]
[0042]
[0043] in, is a diagonal matrix;
[0044]
[0045]
[0046]
[0047] ε1,ε2>0 is used to avoid the situation where the denominator of the diagonal elements of the diagonal matrix is zero. I is the identity matrix, ||·|| F represents the Frobenius norm of the matrix, and ||·||2 represents the L2 norm of the vector.
[0048] An adaptive graph clustering device based on robust symmetric non-negative matrix factorization, comprising:
[0049] An acquisition module, configured to acquire a target data set, wherein the target data set includes a plurality of sample vectors;
[0050] A clustering module is used to call a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix decomposition, cluster multiple sample vectors in the target data set, and obtain multiple groups of clustering results for the multiple sample vectors;
[0051] The processing module is used to process the multiple clustering results through integrated clustering to obtain the optimal clustering result of the target data set.
[0052] A terminal device comprises a processor and a memory for storing instructions executable by the processor, wherein the steps of the above method are implemented when the processor executes the instructions.
[0053] A computer-readable storage medium stores a computer program / instruction thereon, which implements the steps of the above method when executed by a processor.
[0054] The adaptive graph clustering method and device based on robust symmetric non-negative matrix decomposition provided by the present application take into account that for existing clustering algorithms, the intrinsic structure of the data is generally explored from the original feature space. Then, if only global geometric information is used, it is difficult to reconstruct the data space in a low-dimensional space and it is impossible to fully understand the original data characteristics. To this end, in this example, a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix decomposition is used to better reflect the local geometric structure of the data, thereby enhancing the robustness of the algorithm, and combining the idea of integrated clustering with robust symmetric non-negative matrix decomposition to improve the algorithm clustering performance and adaptability to data with different structures, thereby improving the accuracy of the clustering results. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0056] Figure 1 This is a method flow chart of an embodiment of an adaptive graph clustering method based on robust symmetric non-negative matrix factorization provided by the present application;
[0057] Figure 2 This is a method flow chart of another embodiment of the adaptive graph clustering method based on robust symmetric non-negative matrix factorization provided by the present application;
[0058] Figure 3 This is a hardware structure block diagram of an electronic device for an adaptive graph clustering method based on robust symmetric non-negative matrix decomposition provided by the present application;
[0059] Figure 4 This is a schematic diagram of the module structure of an embodiment of an adaptive graph clustering device based on robust symmetric non-negative matrix decomposition provided by the present application. DETAILED DESCRIPTION
[0060] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0061] Considering that the existing clustering algorithms generally explore the intrinsic structure of data from the original feature space and utilize the global or local geometric features of the data, it is difficult to reconstruct the data space in a low-dimensional space and fully understand the original data features with only global geometric information. 2,p The (0<p1) norm and graph regularization under three Laplace matrix construction methods are used to better reflect the local geometric structure of the data, enhance the robustness of the algorithm, improve the algorithm clustering performance and adaptability to data with different structures, and combine the integrated clustering idea with robust symmetric non-negative matrix factorization to improve the accuracy of the clustering results.
[0062] Figure 1 It is a method flow chart of an embodiment of the adaptive graph clustering method based on robust symmetric non-negative matrix decomposition provided by the present application. Although the present application provides the method operation steps or device structure as shown in the following embodiments or drawings, more or fewer operation steps or module units may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of the present application and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be connected according to the method or module structure shown in the embodiment or drawings for sequential execution or parallel execution (for example, a parallel processor or multi-threaded processing environment, or even a distributed processing environment).
[0063] Specifically, such as Figure 1 As shown, the above-mentioned adaptive graph clustering method based on robust symmetric non-negative matrix factorization may include the following steps:
[0064] Step 101: Acquire a target data set, wherein the target data set includes multiple sample vectors;
[0065] Furthermore, the target dataset also carries noise and outliers;
[0066] For example, you can choose a database As the input data matrix, that is, the target data set, the data matrix includes n samples, and each column in the data matrix is a sample vector, that is, x1 is a sample vector.
[0067] Step 102: Retrieve a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, perform clustering on a plurality of sample vectors in the target data set, and obtain a plurality of clustering results of the plurality of sample vectors;
[0068] Furthermore, during the clustering process, robustness to noise and outliers can be maintained;
[0069] That is, we can first establish the similarity matrix P = X based on experience T X, and randomly generate a set of non-negative matrices H r , the above P, H r The input is fed into the preset clustering algorithm, and both local and global graph structures are considered. Regularization parameters can be introduced into the preset clustering algorithm to obtain more precise data features, thereby obtaining a more accurate cluster partition matrix, thereby improving the clustering effect.
[0070] Step 103: Process the multiple clustering results through ensemble clustering to obtain the optimal clustering result of the target data set.
[0071] Specifically, the above step 103 can be to regard the robust SNMF decomposition matrices with various random initializations as different clustering results, adopt an automatic weighting strategy to assign appropriate weights to different clustering results, and use the idea of ensemble clustering to obtain the clustering results of each sample vector in the target data set.
[0072] In the above example, considering that the existing clustering methods do not consider robustness, sensitivity of variable initialization and adaptive graph learning at the same time, in this example, a method is provided through L 2,p The (0<p≤1) norm is used to better reflect the geometric structure of the data, enhancing the algorithm's robustness and adaptability to data with different structures. This approach combines ensemble clustering with symmetric non-negative matrix factorization, thereby improving the scalability of clustering and the final clustering effect. This approach addresses the technical issues of low clustering performance and poor universality inherent in existing clustering methods, effectively improving clustering performance.
[0073] Specifically, the above step 102: calling a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, clustering multiple sample vectors in the target data set, and obtaining multiple groups of clustering results for the multiple sample vectors may include:
[0074] S1: Generate an empirical similarity matrix based on the target dataset;
[0075] S2: Repeat the following operations to obtain multiple cluster partition matrices: generate a set of random non-negative matrices based on the empirical similarity matrix; and obtain the cluster partition matrices corresponding to the random non-negative matrices based on the random non-negative matrices;
[0076] S3: Obtaining a similarity matrix after reconstructing the target data set according to the multiple clustering matrices;
[0077] S4: Decomposing the similarity matrix after reconstructing the target data set through the robust symmetric non-negative matrix in the objective function to obtain multiple groups of clustering results.
[0078] Accordingly, the above step 103: processing the multiple clustering results by integrated clustering to obtain the optimal clustering result of the target data set may include:
[0079] S1: assigning a weight to each of the multiple clustering results by automatic weighting;
[0080] S2: Perform multiple rounds of alternating iterative optimization to obtain the optimal clustering result of the target data set.
[0081] Specifically, performing multiple rounds of alternating iterative optimization to obtain the optimal clustering result of the target data set may include: performing preset rounds of iterative optimization on each parameter in the objective function in the following manner to obtain the optimal clustering result of the target data set: selecting a parameter from the objective function as the target parameter of the current iterative optimization; calling a preset update rule to iteratively optimize the target parameter, and in the process of iteratively optimizing the target parameter, fixing other parameters except the target parameter to remain unchanged.
[0082] The above objective function can be expressed as:
[0083]
[0084]
[0085] Where W is the reconstructed similarity matrix that measures the similarity between the i-th and j-th samples of X. is the target dataset, n is the number of samples, m is the feature dimension of the target dataset, is the clustering result of the target data set, H r The position of the maximum value in each row indicates the sample X i The cluster members, c represents the number of clusters, α r is the weight vector The rth element of is used to balance the contribution of each group of clustering results. Represents a full 1 vector, used to ensure that each α r are all effective weights, λ≥0 is the regularization parameter, For use in H r The low-dimensional representation H ri and H rj Keep X i and X j Local information with similar features, L represents the Laplacian matrix, ||·|| 2,p L represents the matrix 2,p norm, 0<p≤1, tr(·) means finding the trace of the matrix, and T means finding the transpose of the matrix.
[0086] Among them, α r is the weight vector The rth element of , which balances the contribution of each cluster partition, specifically, There are weight vectors in the first term of the objective function, which can be understood as the weighting of the contribution of each group of cluster divisions. The more accurate the division, the higher the weight, and the more vague the division, the lower the weight. represents a vector of all 1s, constraining α T 1=1 avoids the trivial solution of α (i.e., α=0), and α≥0 ensures that every α r are all effective weights. λ≥0 is the regularization parameter, Can be in H r The low-dimensional representation H ri and H rj Keep X i and X j It can also capture local information with similar features, thereby greatly improving its representation ability and clustering performance.
[0087] By introducing L into the objective function 2,p The (0<p≤1) norm achieves the technical effect of effectively improving clustering performance by selecting different p to maintain the robustness, flexibility and universality of the method.
[0088] That is, we can first generate a set of random non-negative matrices (k represents the number of elements in the set), and k cluster partition matrices (or cluster membership matrices) are obtained by robust symmetric non-negative matrix decomposition. (when When M r,ij =1, otherwise, Mr,ij =0;H r,ij 、M r,ij They are H r and M r Then, we construct a reconstruction similarity matrix that measures the similarity between the i-th and j-th samples of X. A new set of better cluster partition matrices is generated under multiple initializations. The process is repeated until the termination criterion or the maximum number of iterations is reached.
[0089] The objective function described above considers both the effects of noise and outliers and the local geometric structure of the target dataset (data features selected based on local geometry are always superior to those selected based on global geometry). This establishes a joint framework that simultaneously considers robustness, sensitivity to variable initialization, and adaptive graph learning. This approach, using an alternating iterative approach to optimize the objective function, achieves joint optimization while significantly reducing computational time.
[0090] That is, introducing L 2,p The (0<p≤1) norm suppresses the influence of noise and outliers. Without the help of any additional information, the sensitivity of robust SNMF to the initialization data is used to gradually improve the clustering performance. It is used to deal with clustering problems with noise and outlier interference. Combined with graph regularization under three Laplace matrix construction methods, the local invariant properties of the manifold are used to fully maintain the local geometric structure of adjacent data.
[0091] Furthermore, the alternating iteration method can be used to update the iteration, that is, the above objective function is iteratively updated multiple times to obtain a set of cluster partition matrices M r , and directly obtain the cluster label of each sample.
[0092] Specifically, calling the preset objective function of the adaptive graph clustering method based on robust symmetric non-negative matrix factorization and clustering the multiple sample vectors may include:
[0093] S1: Obtain a target data set, where the target data set includes multiple sample vectors with noise and outliers;
[0094] S2: calling a preset objective function of a graph clustering method based on adaptive robust symmetric non-negative matrix factorization to obtain different clustering results of the multiple sample vectors while maintaining robustness to noise and outliers;
[0095] S3: Using the ensemble clustering idea, the optimal clustering result of each sample vector in the target data set is adaptively obtained.
[0096] When performing multiple rounds of alternating iterative optimization to obtain the clustering results of the multiple sample vectors, a parameter of the objective function can be selected as the parameter for the current iterative optimization; the preset update rule is called to iteratively optimize the current parameter, and in the process of iteratively optimizing the current parameter, other parameters are fixed unchanged; after iteratively optimizing the parameters in the objective function for the preset rounds, the optimal parameters are used to solve and obtain a high-quality reconstructed similarity matrix of the multiple sample points. For example, the preset rounds can be set to 10 rounds, then 10 rounds of optimization iterations are performed, the parameters after 10 rounds of optimization iterations are used as the optimal parameters, and the objective function is solved with the optimal parameters to obtain a reconstructed similarity matrix as the final reconstructed similarity matrix.
[0097] When the alternating iterative method is used to update and solve the above non-convex objective function, the above objective function can be rewritten as:
[0098]
[0099]
[0100] The above question is equivalent to the following question:
[0101]
[0102]
[0103] The Laplacian matrix constructed in the objective function can be at least one of the following:
[0104] Laplacian matrix:
[0105]
[0106] Normalized Laplacian matrix:
[0107]
[0108] Shifted Laplace matrix:
[0109]
[0110] Where D is the diagonal matrix,
[0111] Correspondingly, the preset update rule is:
[0112] 1) When the Laplace matrix is In the case of H r The update formula is:
[0113]
[0114] The Laplace matrix is In the case of H r The update formula is:
[0115]
[0116] The Laplace matrix is In the case of H r The update formula is:
[0117]
[0118] 2)
[0119] in, is a diagonal matrix, ε1,ε2>0 avoids the situation where the denominator of the diagonal elements of the diagonal matrix is zero. I is the identity matrix, ||·|| F represents the Frobenius norm of the matrix, and ||·||2 represents the L2 norm of the vector.
[0120] 3)
[0121] Because the reconstruction error decreases sharply at the beginning, manifold regularization becomes less effective during iterations, so λ needs to be dynamically calculated. The above formula dynamically (proportionally) balances the reconstruction error and the manifold regularization term. This dynamic update accelerates objective function convergence, especially for large datasets. The constant β represents the ratio of the manifold regularization term to the reconstruction error and can be determined experimentally.
[0122] It is directly interpreted as the clustering result of the target data set, that is, H r The position of the maximum value of each row can indicate the sample X i The cluster members of c are cluster numbers. Therefore, in the objective function of the adaptive graph clustering, dimensionality reduction can be achieved by using symmetric non-negative matrix decomposition, and the category information can be reflected in the matrix H r middle.
[0123] The specific optimization and update process of the adaptive graph clustering algorithm based on robust symmetric non-negative matrix factorization can be implemented as follows:
[0124] Input: empirical similarity matrix P, hyperparameters k, λ, ε=10 -3 ;
[0125] Output: a set of clustering results
[0126] S1: Initialize a set of random non-negative matrices W = P;
[0127] S2: According to the above The formula for updating H r ;
[0128] S3: According to the above α r The formula to update α r ;
[0129] S4: dynamically obtain λ according to the above λ formula;
[0130] S5: According to the above formula Update W;
[0131] S6: According to the above formula Update G r ;
[0132] S7: Through the above formula Returns a set of clustering results M r . Among them H r,ij 、M r,ij They are H r and M r The element in the i-th j-th column of .
[0133] That is, in the above example, without any additional information, L is introduced 2,p (0<p≤1) norm, and the sensitivity of robust symmetric non-negative matrix factorization to the initialization data is used to obtain different clustering results; combined with graph regularization under three Laplace matrix construction methods, the local invariant properties of the manifold are used to fully maintain the local geometric structure of adjacent data; using the ensemble clustering idea, the optimal clustering result of each sample vector in the target data set is adaptively obtained, and the alternating iteration method is used for iterative update to optimize the solution.
[0134] The above method is described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating the present application and does not constitute an improper limitation to the present application.
[0135] In this example, a set of random non-negative moments is first generated. Robust symmetric non-negative matrix factorization is then used to obtain k cluster partition matrices (or cluster membership matrices). A reconstructed similarity matrix is then constructed to measure the similarity between samples in the target dataset. A new set of better cluster partition matrices is then generated with multiple initializations. This process is repeated until a termination criterion or the maximum number of iterations is reached.
[0136] Specifically, you can follow the Figure 2 The steps shown achieve:
[0137] Step S201: Obtain sample data as the target data set:
[0138] For example, you can choose a database As the input data matrix, that is, the target data set, the data matrix includes n samples, and each column in the data matrix is a sample vector, that is, x1 is a sample vector.
[0139] Step S202: Obtain different clustering results of the multiple sample vectors:
[0140] We can first establish the similarity matrix P = X based on experience T X, and randomly generate a set of non-negative matrices H r , the above P, H r The input is fed into the preset clustering algorithm, and both local and global graph structures are considered. Regularization parameters can be introduced into the preset clustering algorithm to obtain more precise data features, thereby obtaining a more accurate cluster partition matrix, thereby improving the clustering effect.
[0141] Among them, the objective function of the above algorithm can be expressed as:
[0142]
[0143]
[0144] in, is the target dataset, It is directly interpreted as the clustering result of the target data set, that is, H r The position of the maximum value of each row can indicate the sample X i The cluster members of , c is the number of clusters. r yes The rth element of , the weight vector balances the contribution of each partition, represents a vector of all 1s, constraining α T 1=1 avoids the trivial solution of α (i.e., α=0), and α≥0 ensures that every α r are all effective weights. λ≥0 is the regularization parameter, Can be in H r The low-dimensional representation H ri and H rj Keep X i and X j It can also capture local information with similar features, thereby greatly improving its representation ability and clustering performance.
[0145] L is selected as: ① Laplace matrix ②Normalized Laplacian ③Shifted Laplacian D is the diagonal matrix, Since the Laplacian matrix inherently contains noise information, its best low-rank approximation mainly encodes the noise, while the normalized Laplacian matrix and the shifted Laplacian matrix mainly encode the clustering information. 2,p L represents the matrix 2,p norm (0<p≤1), tr(·) represents the trace of the matrix, and T represents the transpose of the matrix.
[0146] Furthermore, in this example, the optimization solution is updated by the alternating iteration method. Specifically, in the process of updating the optimization solution, other parameters are fixed and then the desired parameters are solved.
[0147] In the process of updating the optimization solution, the following update rules can be followed:
[0148] 1) H under three Laplace matrix construction methods r The update formulas are
[0149] ①
[0150] ②
[0151] ③
[0152] 2)
[0153] in, is a diagonal matrix, ε1,ε2>0 avoids the situation where the denominator of the diagonal elements of the diagonal matrix is zero. I is the identity matrix, ||·|| F represents the Frobenius norm of the matrix, and ||·||2 represents the L2 norm of the vector.
[0154] 3)
[0155] Because the reconstruction error decreases sharply at the beginning, this means that the manifold regularization is less effective during the iterations, so λ needs to be dynamically calculated. The above formula dynamically (proportionally) balances the reconstruction error and the manifold regularization term. Dynamically updating λ also accelerates objective function convergence, especially for large datasets. The constant β represents the ratio of the manifold regularization term to the reconstruction error and can be obtained through specific experiments.
[0156] S203: Clustering.
[0157] Specifically, the robust SNMF decomposition matrices with various random initializations are regarded as different clustering results, and an automatic weighting strategy is used to assign appropriate weights to different clustering results. The idea of ensemble clustering is used to obtain the clustering label H of each sample vector in the target dataset. r .
[0158] The clustering accuracy (ACC) of adaptive graph clustering based on robust symmetric non-negative matrix factorization under different p values provided in this example will change. That is, the best ACC on different datasets changes with the change of p. There is no universally optimal p value for all datasets. A better solution is to keep p flexible.
[0159] The method embodiments provided in the above embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on an electronic device as an example, Figure 3 This is a hardware structure diagram of an electronic device for an adaptive graph clustering method based on robust symmetric non-negative matrix decomposition provided by this application. Figure 3 As shown, the electronic device 10 may include one or more (only one is shown in the figure) processors 02 (the processor 02 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 04 for storing data, and a transmission module 06 for communication functions. It will be understood by those skilled in the art that Figure 3 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3 More or fewer components than shown, or with Figure 3 Different configurations shown.
[0160] The memory 04 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the adaptive graph clustering method based on robust symmetric non-negative matrix decomposition in the embodiment of the present application. The processor 02 executes various functional applications and data processing by running the software programs and modules stored in the memory 04, that is, the adaptive graph clustering method based on robust symmetric non-negative matrix decomposition of the above-mentioned application is realized. The memory 04 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 04 may further include a memory remotely arranged relative to the processor 02, and these remote memories can be connected to the electronic device 10 via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0161] The transmission module 06 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of the electronic device 10. In one embodiment, the transmission module 06 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 06 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0162] At the software level, the above-mentioned adaptive graph clustering device based on robust symmetric non-negative matrix factorization can be Figure 4 Shown, including:
[0163] An acquisition module 401 is configured to acquire a target data set, wherein the target data set includes a plurality of sample vectors;
[0164] A clustering module 402 is configured to call a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, perform clustering on a plurality of sample vectors in the target data set, and obtain a plurality of clustering results for the plurality of sample vectors;
[0165] The processing module 403 is configured to process the multiple clustering results through integrated clustering to obtain an optimal clustering result of the target data set.
[0166] In one embodiment, the clustering module 402 can specifically generate an empirical similarity matrix based on the target data set; repeatedly perform the following operations to obtain multiple cluster partition matrices: generate a set of random non-negative matrices based on the empirical similarity matrix; obtain a cluster partition matrix corresponding to the random non-negative matrix based on the random non-negative matrix; obtain a similarity matrix after reconstructing the target data set based on the multiple cluster partition matrices; decompose the similarity matrix after reconstructing the target data set through the robust symmetric non-negative matrix in the objective function to obtain multiple groups of clustering results.
[0167] In one embodiment, the processing module 403 may specifically assign a weight to each of the multiple clustering results by automatic weighting; and perform multiple rounds of alternating iterative optimization and solving to obtain the optimal clustering result of the target data set.
[0168] In one embodiment, performing multiple rounds of alternating iterative optimization to obtain the optimal clustering result of the target data set may include: performing preset rounds of iterative optimization on each parameter in the objective function in the following manner to obtain the optimal clustering result of the target data set: selecting a parameter from the objective function as the target parameter for the current iterative optimization; calling a preset update rule to iteratively optimize the target parameter, and in the process of iteratively optimizing the target parameter, fixing other parameters except the target parameter to remain unchanged.
[0169] In one embodiment, the objective function may be:
[0170]
[0171]
[0172] Where W is the reconstructed similarity matrix that measures the similarity between the i-th and j-th samples of X. is the target dataset, n is the number of samples, m is the feature dimension of the target dataset, is the clustering result of the target data set, H r The position of the maximum value in each row indicates the sample X i The cluster members, c represents the number of clusters, α r is the weight vector The rth element of is used to balance the contribution of each group of clustering results. Represents a full 1 vector, used to ensure that each α r are all effective weights, λ≥0 is the regularization parameter, For use in H r The low-dimensional representation H ri and H rj Keep X i and X j Local information with similar features, L represents the Laplacian matrix, ||·|| 2,p L represents the matrix 2,p norm, 0<p≤1, tr(·) means finding the trace of the matrix, and T means finding the transpose of the matrix.
[0173] In one embodiment, the Laplacian matrix constructed in the objective function may be at least one of the following:
[0174] Laplacian matrix:
[0175]
[0176] Normalized Laplacian matrix:
[0177]
[0178] Shifted Laplace matrix:
[0179]
[0180] Where D is the diagonal matrix,
[0181] In one embodiment, the preset update rule may be:
[0182] The Laplace matrix is In the case of H r The update formula is:
[0183]
[0184] The Laplace matrix is In the case of H r The update formula is:
[0185]
[0186] The Laplace matrix is In the case of H r The update formula is:
[0187]
[0188]
[0189] in, is a diagonal matrix;
[0190]
[0191]
[0192]
[0193] ε1,ε2>0 is used to avoid the situation where the denominator of the diagonal elements of the diagonal matrix is zero. I is the identity matrix, ||·|| F represents the Frobenius norm of the matrix, and ||·||2 represents the L2 norm of the vector.
[0194] The embodiments of the present application also provide a specific implementation of an electronic device that can implement all the steps in the adaptive graph clustering method based on robust symmetric non-negative matrix decomposition in the above embodiment. The electronic device specifically includes the following contents: a processor, a memory, a communication interface, and a bus; wherein the processor, the memory, and the communication interface communicate with each other through the bus; the processor is used to call a computer program in the memory, and when the processor executes the computer program, all the steps in the adaptive graph clustering method based on robust symmetric non-negative matrix decomposition in the above embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0195] Step 1: Obtain a target data set, wherein the target data set includes multiple sample vectors;
[0196] Step 2: Retrieve a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, perform clustering on multiple sample vectors in the target data set, and obtain multiple groups of clustering results for the multiple sample vectors;
[0197] Step 3: Process the multiple clustering results through ensemble clustering to obtain the optimal clustering result of the target data set.
[0198] As can be seen from the above description, the embodiment of the present application takes into account that for existing clustering algorithms, the intrinsic structure of the data is generally explored from the original feature space. Then, if only global geometric information is used, it is difficult to reconstruct the data space in a low-dimensional space and it is impossible to fully understand the original data characteristics. To this end, in this example, a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix decomposition is used to better reflect the local geometric structure of the data, thereby enhancing the robustness of the algorithm, and combining the integrated clustering idea with robust symmetric non-negative matrix decomposition to improve the algorithm clustering performance and adaptability to data with different structures, thereby improving the accuracy of the clustering results.
[0199] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the adaptive graph clustering method based on robust symmetric non-negative matrix factorization in the above-mentioned embodiment. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the adaptive graph clustering method based on robust symmetric non-negative matrix factorization in the above-mentioned embodiment. For example, when the processor executes the computer program, the following steps are implemented:
[0200] Step 1: Obtain a target data set, wherein the target data set includes multiple sample vectors;
[0201] Step 2: Retrieve a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, perform clustering on multiple sample vectors in the target data set, and obtain multiple groups of clustering results for the multiple sample vectors;
[0202] Step 3: Process the multiple clustering results through ensemble clustering to obtain the optimal clustering result of the target data set.
[0203] As can be seen from the above description, the embodiment of the present application takes into account that for existing clustering algorithms, the intrinsic structure of the data is generally explored from the original feature space. Then, if only global geometric information is used, it is difficult to reconstruct the data space in a low-dimensional space and it is impossible to fully understand the original data characteristics. To this end, in this example, a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix decomposition is used to better reflect the local geometric structure of the data, thereby enhancing the robustness of the algorithm, and combining the integrated clustering idea with robust symmetric non-negative matrix decomposition to improve the algorithm clustering performance and adaptability to data with different structures, thereby improving the accuracy of the clustering results.
[0204] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the hardware + program embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0205] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0206] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment).
[0207] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0208] Although the present specification embodiment provides the method operation steps as described in the embodiment or flow chart, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiment is only one way in the order of execution of many steps and does not represent a unique execution order. When the device or terminal product in practice is executed, it can be performed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (such as a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements not only include those elements, but also include other elements not clearly listed, or also include elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements.
[0209] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing the embodiments of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules that implement the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0210] Those skilled in the art will also appreciate that, in addition to implementing the controller in pure computer-readable program code, it is entirely possible to implement the same functionality by logically programming the method steps in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered structures within the hardware component. Alternatively, the devices for implementing various functions can be considered both software modules implementing the method and structures within the hardware component.
[0211] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0212] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0214] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0215] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0216] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0217] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0218] Embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.
[0219] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the embodiments in this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.
[0220] The above description is merely an example of the embodiments of this specification and is not intended to limit the embodiments of this specification. For those skilled in the art, various modifications and variations of the embodiments of this specification are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of this specification shall be included within the scope of the claims of the embodiments of this specification.
Claims
1. An adaptive graph clustering method based on robust symmetric non-negative matrix factorization, characterized by: include: Acquire a target data set, wherein the target data set includes a plurality of sample vectors; Recalling a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, clustering multiple sample vectors in the target data set, and obtaining multiple groups of clustering results for the multiple sample vectors; Processing the multiple clustering results through integrated clustering to obtain the optimal clustering result of the target data set; Wherein, the objective function is: Among them, \(W\) is the reconstructed similarity matrix that measures the similarity between the \(i\)-th and \(j\)-th samples of \(X\). is the target data set, \(n\) is the number of samples, and \(m\) is the feature dimension of the target data set. is the clustering result of the target data set, and the position of the maximum value in each row of \(H\) j , i , 2,p , 2,p indicates the clustering members of the sample \(X\). i \(c\) represents the number of clusters, and \(\alpha\) r is the \(r\)-th element of the weight vector used to balance the contributions of each group of clustering results. represents the all-1 vector, used to ensure that each \(\alpha\) r is an effective weight, \(\lambda\geq0\) is the regularization parameter, used to retain the local information of \(X\) with similar features in the low-dimensional representation \(H\) r of \(H\) and \(H\) ri and \(H\) [[ID=2 The Laplace matrix constructed in the objective function is at least one of the following: Laplacian matrix: Normalized Laplacian matrix: Shifted Laplace matrix: Where D is the diagonal matrix, 2. The method according to claim 1, characterized in that Recalling a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix factorization, clustering multiple sample vectors in the target data set, and obtaining multiple groups of clustering results for the multiple sample vectors, including: Generate an empirical similarity matrix based on the target dataset; Repeat the following operations to obtain multiple cluster partition matrices: generate a set of random non-negative matrices based on the empirical similarity matrix; obtain a cluster partition matrix corresponding to the random non-negative matrix based on the random non-negative matrix; Obtaining a similarity matrix after reconstructing the target data set according to the multiple clustering matrices; The similarity matrix after reconstructing the target data set is decomposed by the robust symmetric non-negative matrix in the objective function to obtain multiple groups of clustering results.
3. The method according to claim 2, characterized in that The multiple clustering results are processed through ensemble clustering to obtain the optimal clustering result of the target data set, including: assigning a weight to each of the plurality of clustering results by automatic weighting; Multiple rounds of alternating iterative optimization are performed to obtain the optimal clustering result of the target data set.
4. The method according to claim 3, characterized in that Perform multiple rounds of alternating iterative optimization to obtain the optimal clustering result of the target data set, including: Perform a preset number of iterative optimization rounds on each parameter in the objective function in the following manner to obtain the optimal clustering result of the target data set: Selecting a parameter from the objective function as a target parameter for current iterative optimization; The preset updating rule is called to iteratively optimize the target parameter. During the iterative optimization of the target parameter, other parameters except the target parameter are fixed to remain unchanged.
5. The method according to claim 1, wherein the preset update rule is: The Laplace matrix is In the case of H r The update formula is: The Laplace matrix is In the case of H r The update formula is: The Laplace matrix is In the case of H r The update formula is: in, is a diagonal matrix; ε1,ε2>0 is used to avoid the situation where the denominator of the diagonal elements of the diagonal matrix is zero. I is the identity matrix, ||·|| F represents the Frobenius norm of the matrix, and ||·||2 represents the L2 norm of the vector.
6. An adaptive graph clustering device based on robust symmetric non-negative matrix factorization, characterized in that: include: An acquisition module, configured to acquire a target data set, wherein the target data set includes a plurality of sample vectors; A clustering module is used to call a preset objective function of adaptive graph clustering based on robust symmetric non-negative matrix decomposition, cluster multiple sample vectors in the target data set, and obtain multiple groups of clustering results for the multiple sample vectors; A processing module, configured to process the plurality of clustering results through integrated clustering to obtain an optimal clustering result of the target data set; Wherein, the objective function is: Where W is the reconstructed similarity matrix that measures the similarity between the i-th and j-th samples of X. is the target dataset, n is the number of samples, m is the feature dimension of the target dataset, is the clustering result of the target data set, H r The position of the maximum value in each row indicates the sample X i The cluster members, c represents the number of clusters, α r is the weight vector The rth element of is used to balance the contribution of each group of clustering results. Represents a full 1 vector, used to ensure that each α r are all effective weights, λ≥0 is the regularization parameter, For use in H r The low-dimensional representation H ri and H rj Keep X i and X j Local information with similar features, L represents the Laplacian matrix, ||·|| 2,p L represents the matrix 2,p norm, 0<p≤1, tr(·) means finding the trace of the matrix, and T means finding the transpose of the matrix; The Laplace matrix constructed in the objective function is at least one of the following: Laplacian matrix: Normalized Laplacian matrix: Shifted Laplace matrix: Where D is the diagonal matrix, 7. A terminal device comprising a processor and a memory for storing processor-executable instructions, wherein the processor implements the steps of the method according to any one of claims 1 to 5 when executing the instructions.
8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Microbial data clustering method based on robust symmetric non-negative matrix factorization
CN113723537A
Robust local and global regularization non-negative matrix factorization clustering method
CN114254703A