Multi-view clustering method and system for rapid graph filtering augmentation

By constructing a two-part graph and thermonuclear diffusion mechanism, designing low-pass filters, smooth enhancement and joint optimization of multi-view clustering, the problem of low-quality base partitioning in multi-view clustering is solved, and more efficient and accurate clustering results are achieved.

CN120296452APending Publication Date: 2025-07-11SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510347829.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing multi-view clustering method combines information from different perspectives, there is low-quality base partition instability and noise sensitivity, resulting in poor clustering effect, especially when there is a lot of noise, it is easy to fall into the local optimal solution, affecting the overall effect.

Method used

By constructing a two-part graph and designing a low-pass filter based on the thermonuclear diffusion mechanism, linear combination is performed, a consensus graph filter is generated, the original base partition is smoothed and enhanced, and joint optimization is performed on the smooth base partition and the original base partition, and combining K-means clustering, the final multi-view clustering results are generated.

Benefits of technology

It significantly improves the quality and clustering effect of the consensus graph filter, suppresses excessive smoothing, maintains linear computing complexity, can effectively process large-scale data sets, and improves the accuracy and stability of clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296452A_ABST
    Figure CN120296452A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view clustering method and system for rapid graph filtering augmentation, and the method comprises the steps: carrying out the partitioning of an original basis of each view, generating an anchor point matrix through K-means clustering, and constructing a bipartite graph based on the probability adjacent weight between a sample and an anchor point; constructing low-pass filters based on a bipartite graph, and performing linear combination on the plurality of low-pass filters to obtain a consensus graph filter; performing smooth enhancement processing on the original base partition of each view based on a consensus graph filter to obtain a smooth base partition; respectively carrying out clustering on the smooth base partition and the original base partition; constructing a joint optimization framework and performing optimization; and combining the clustering results of the smooth base partition and the original base partition based on the optimized joint optimization framework to generate a final multi-view clustering result. According to the method, the precision and efficiency of multi-view clustering are remarkably improved, and the method has the capability of efficiently processing large-scale data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-view clustering, and particularly relates to a multi-view clustering method and system with fast graph filtering augmentation. Background Art

[0002] With the rapid development of multimedia data, multi-view data such as images, videos, and audios are becoming increasingly common in practical applications. These data from different modalities provide complementary perspectives, promoting more comprehensive data understanding and information extraction. Multi-view Clustering (MVC), as an important data analysis method, reveals potential patterns and structures by fusing data from multiple perspectives and is widely used in fields such as image classification, medical diagnosis, and social network analysis. However, due to the heterogeneity and noise between different perspectives, how to effectively fuse information from each perspective has become the core challenge in multi-view clustering.

[0003] Existing MVC methods can be divided into two categories: early fusion and late fusion according to the different information fusion stages. In recent years, late fusion methods have attracted much attention due to their high computational efficiency. Their strategy is to first independently generate base partitions on each view and then fuse them on the partitions to obtain a consistent clustering result. However, the main challenge faced by late fusion methods in practical applications is the dependence on low-quality base partitions. For example, the experimental results of the COIL100 dataset show that the average clustering accuracy of the base partitions is only 36.19%, while the variance is as high as 16.82%. The root cause of this problem lies in: firstly, the representation ability of the original data is limited, and traditional feature representation methods are difficult to fully capture the potential clustering structure; secondly, when generating base partitions based on a single view, the information from different perspectives is not effectively fused, resulting in the instability of the base partitions under noisy views; finally, the clustering algorithm is sensitive to data and is prone to falling into local optimal solutions when there is a lot of noise, exacerbating the instability of the base partitions. These factors jointly lead to low-quality base partitions, seriously affecting the overall effect of late fusion methods.

[0004] To solve the above problems and achieve linear computational complexity, it is urgent to propose a multi-view clustering method and system with fast graph filtering augmentation. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes a multi-view clustering method and system with fast graph filtering augmentation to solve the problems existing in the above prior art.

[0006] To achieve the above object, the present invention provides a multi-view clustering method with fast graph filtering augmentation, including the following steps:

[0007] For the original base partition of each view, an anchor matrix is generated by K-means clustering, and a bipartite graph is constructed based on the probability proximity weights between samples and anchors;

[0008] Based on the heat kernel diffusion mechanism, a low-pass filter based on the bipartite graph is constructed, and a linear combination of several low-pass filters is performed to obtain a consensus graph filter for multi-view clustering;

[0009] Based on the consensus graph filter, the original base partition of each view is smoothed and enhanced to obtain a smoothed base partition;

[0010] K-means clustering is performed on the smoothed base partition and the original base partition respectively;

[0011] A joint optimization framework is constructed, and the joint optimization framework is optimized based on the block coordinate rotation method;

[0012] Based on the optimized joint optimization framework, the clustering results of the smoothed base partition and the original base partition are combined to generate the final multi-view clustering result.

[0013] Optionally, the process of constructing a low-pass filter based on the bipartite graph based on the heat kernel diffusion mechanism includes:

[0014] Introduce a diagonal matrix into the bipartite graph to obtain a sample-anchor similarity graph; perform column normalization on the sample-anchor similarity graph to obtain a sample-anchor-sample similarity graph; based on the heat kernel diffusion mechanism, perform weighted combination on each order of the sample-anchor-sample similarity graph to obtain the corresponding low-pass filter.

[0015] Optionally, the formula of the low-pass filter is as follows:

[0016]

[0017] where t represents the order of diffusion, is the weight coefficient of heat kernel diffusion, S v is the sample-anchor-sample similarity graph, η is a non-negative parameter controlling the attenuation rate, L v is the normalized Laplacian matrix constructed for the v-th view, G v is the low-pass filter with respect to L v of.

[0018] Optionally, the formula of the consensus graph filter is as follows:

[0019]

[0020] where β v is the weight of the v-th graph filter.

[0021] Optionally, the objective function of the joint optimization framework is to minimize the weighted sum of the clustering errors on the smoothed basis partition and the original basis partition. The optimization variables include the clustering indicator matrix, the clustering centroid matrix, the view weight, and the graph filter weight.

[0022] Optionally, the formula of the objective function of the joint optimization framework is as follows:

[0023]

[0024] where α p is a non - negative view weight coefficient, β v is a non - negative graph filter weight coefficient, λ is a trade - off hyperparameter between the smoothed basis partition clustering and the original basis partition clustering, G v is a low - pass filter with respect to L v H p is the original basis partition, Y is the consensus clustering indicator matrix, and C p is the clustering centroid matrix learned from the p - th basis partition.

[0025] The present invention also provides a fast graph - filtering - augmented multi - view clustering system for implementing the above - mentioned method, including: a bipartite graph construction module, a filter construction module, a smoothing processing module, a joint clustering module, and an iterative update module;

[0026] The bipartite graph construction module is used to generate an anchor matrix for the original basis partition of each view through K - means clustering, and construct a bipartite graph based on the probability proximity weights between samples and anchors;

[0027] The filter construction module is used to construct a low - pass filter based on the bipartite graph based on the heat - kernel diffusion mechanism, and perform a linear combination of several low - pass filters to obtain a consensus graph filter for multi - view clustering;

[0028] The smoothing processing module is used to perform smoothing enhancement processing on the original basis partition of each view based on the consensus graph filter to obtain a smoothed basis partition;

[0029] The joint clustering module is used to perform K - means clustering on the smoothed basis partition and the original basis partition respectively, construct a joint optimization framework, and combine the clustering results of the smoothed basis partition and the original basis partition based on the joint optimization framework to generate a multi - view clustering result;

[0030] The iterative update module is used to iteratively update the joint optimization framework based on the block - coordinate rotation method, and further generate the final multi - view clustering result.

[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the method.

[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.

[0033] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.

[0034] Compared with the prior art, the present invention has the following advantages and technical effects:

[0035] The present invention proposes a fast graph filtering augmented multi-view clustering method and system. By constructing a bipartite graph and designing a graph filter induced by high-order diffusion based on the heat kernel, the linear combination of multiple graph filters is learned under the multi-view clustering framework, which significantly improves the quality of the consensus graph filter and the clustering effect, while maintaining linear complexity and being able to effectively process large-scale data sets. In addition, the present invention proposes a joint optimization framework, and by performing joint clustering on the original base partition and the smoothed partition enhanced by the graph filter, the advantages of the graph filter are fully utilized, and the degradation of the clustering performance caused by the over-smoothing problem of the graph filter is effectively suppressed. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0037] Figure 1 It is a schematic flowchart of the fast graph filtering augmented multi-view clustering method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0039] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0040] Embodiment 1

[0041] Graph filters are key tools in graph signal processing, used to process signals on graph structures, transmit information through the adjacency relationships between nodes, and achieve functions such as signal smoothing, noise reduction, or feature extraction. According to spectral characteristics, graph filters can be divided into low-pass, band-pass, and high-pass filters: low-pass filters are used to smooth signals and remove noise, band-pass filters extract specific frequency components, and high-pass filters highlight significant structural features. Graph filters are widely applied in fields such as image processing, graph neural networks, data denoising, and social network analysis. Especially when dealing with data with complex topological structures, they can effectively capture local and global features of the data.

[0042] In the prior art, various methods have been explored to apply graph filters to multi-view clustering. Although graph filters play an important role in multi-view clustering, there are still the following key problems to be solved:

[0043] Disconnection between design and task: The design of graph filters and the multi-view clustering process are usually carried out in stages, resulting in the upstream graph filter design not being optimally adapted to the downstream clustering task, which limits the improvement of the overall performance.

[0044] Over-smoothing problem: High-order graph filters are prone to causing over-smoothing phenomena, suppressing important structural information in the graph, and thus affecting the accuracy of clustering results.

[0045] High computational complexity: The construction complexity of graph filters is relatively high, and the complexity of data filtering is also high, which may bring significant computational overhead on large-scale data sets and limit their practical applications.

[0046] To solve the above problems, this embodiment proposes a multi-view clustering method and system with fast graph filter augmentation, which optimizes the design of graph filters, makes them closely combined with the clustering task, avoids over-smoothing, and reduces the computational complexity to improve the efficiency and performance of multi-view clustering.

[0047] As a specific implementation manner, the method includes the following steps:

[0048] For the original base partition of each view, generate an anchor matrix through K-means clustering, and construct a bipartite graph based on the probability proximity weights between samples and anchors;

[0049] Based on the heat kernel diffusion mechanism, construct a low-pass filter based on the bipartite graph, and perform a linear combination of several low-pass filters to obtain a consensus graph filter for multi-view clustering;

[0050] Based on the consensus graph filter, perform smoothing enhancement processing on the original base partition of each view to obtain a smoothed base partition;

[0051] Perform K-means clustering on the smoothed base partition and the original base partition respectively;

[0052] Construct a joint optimization framework and optimize the joint optimization framework based on the block coordinate rotation method;

[0053] Combine the clustering results of the smoothed basis partition and the original basis partition based on the optimized joint optimization framework to generate the final multi-view clustering result.

[0054] As a specific implementation manner, this embodiment proposes a multi-bipartite graph-induced multi-view consensus graph filter to smooth and enhance the basis partition, and combines its low-rank structure to achieve efficient linear filtering operations; then proposes a fast graph filtering augmented multi-view clustering method to improve the post-fusion clustering performance and prevent over-smoothing degradation through the joint optimization of the original basis partition clustering, the smoothed basis partition clustering, and the consensus graph filter coefficient learning. The overall framework is as Figure 1 shown.

[0055] Implementable, for the original basis partition of each view, generate an anchor matrix through K-means clustering, and construct a bipartite graph based on the probability proximity weight between the samples and the anchors. The process includes:

[0056] To characterize the proximity relationship on the samples and maintain linear computational complexity, this embodiment uses the K-means algorithm to select the anchors and characterizes the relationship between the samples and the anchors through the probability proximity weight. Specifically, first in the basis partition where n is the number of samples, d is the dimension of the basis partition, run the K-means algorithm to cluster the samples into m classes and use the cluster centers as the anchor matrix where m is the number of anchors. Then construct the bipartite graph corresponding to the samples and the anchors through the probability weights on the neighbors The calculation formula is as follows:

[0057]

[0058] where, h v,i and o v,i respectively represent the i-th row of the basis partition H v and the anchor matrix O v In this embodiment, the neighborhood size k = 5 is set.

[0059] Implementable, based on the heat kernel diffusion mechanism, the process of constructing a bipartite graph-based low-pass filter includes:

[0060] Although the above bipartite graph is often used for large-scale clustering, it has certain limitations. First, the selection of the number of anchor points and the anchor point set cannot guarantee optimality for downstream tasks. Second, the non-negative weights of the bipartite graph only characterize the first-order similarity between samples and a small number of anchor points, which limits its expression of the global structure and usually leads to misclassification of in-cluster samples outside the neighborhood as negative samples and misclassification of out-of-cluster samples within the neighborhood as positive samples. Therefore, in this embodiment, the similarity of sample-anchor-sample is further utilized to capture the indirect association relationship between samples.

[0061] Specifically, given the sample-anchor bipartite graph Z obtained from the v-th view v , this embodiment first introduces a diagonal matrix whose j-th diagonal element is Next, the column-normalized sample-anchor similarity graph can be obtained. Then, the sample-anchor-sample similarity graph S v can be derived, expressed as where It can be verified that S v is a doubly stochastic matrix, that is, Although these similarity graphs can extract local manifold structures, they do not consider high-order neighborhood information, resulting in the direct neglect of the similarity between distant samples and thus loss of global information. Therefore, in this embodiment, high-order connections are introduced through graph diffusion, and the heat kernel diffusion mechanism is used to weight and combine similarity graphs of each order to obtain the corresponding graph filter as follows:

[0062]

[0063] where t represents the order of diffusion, is the weight coefficient of heat kernel diffusion, η is a non-negative parameter controlling the decay rate, and L v is the normalized Laplacian matrix constructed from the v-th view It is not difficult to see that G v is a low-pass filter with respect to L v Since G v only considers the information of the v-th view, this embodiment proposes to learn a consensus graph filter for multi-view clustering through the linear combination of multiple low-pass filters:

[0064]

[0065] where β v is the weight of the v-th graph filter, and its optimal solution is automatically determined by the joint optimization of the downstream clustering task. This consensus filter adaptively fuses the high-order connections of each view in the multi-view clustering task, thereby achieving more effective data denoising and enhancement.

[0066] It can be implemented to perform smoothing enhancement on the original base partition of each view based on the consensus graph filter to obtain a smoothed base partition. The process includes:

[0067] Insufficient base partition quality is the main bottleneck of post-fusion multi-view clustering. A smoother graph signal can often present a clearer cluster structure. Therefore, in this embodiment, the above-mentioned multi-view consensus graph filter is used to smooth and enhance the original base partition H p to obtain a smoothed partition

[0068]

[0069] Although the above filtering can effectively enhance the base partition, the complexity of direct calculation is relatively high. Specifically, the complexity of calculating is The complexity of calculating exp(-ηL v ) is And the complexity of calculating G v H p is To reduce the computational overhead and improve the usability of the consensus graph filter in multi-view clustering, this embodiment proposes a high-order graph filtering method with linear complexity.

[0070] Specifically, in this embodiment, by exploring the low-rank structure of the bipartite graph, the sample-side graph filter is transformed into a heat kernel diffusion process on the anchor points, and the computational complexity of the graph filter is reduced from to In addition, in this embodiment, instead of calculating the filter G v , directly calculate G v filtered on H p , further reducing the complexity from to The specific process of this linear graph filtering calculation is as follows:

[0071]

[0072] In the above formula, the complexity of calculating is The complexity of inverting is The complexity of heat kernel diffusion on the anchor point side is The complexity of calculating other matrix multiplications from right to left in turn is Therefore, the total time complexity of the above formula is Through the above acceleration optimization, the time complexity of this embodiment is successfully reduced from to linear Greatly improves the computational efficiency of filtering enhancement.

[0073] To further improve the computational efficiency in the iterative optimization process, in this embodiment, the smoothed basis partitions obtained by applying the low-pass filter to each basis partition under each view are pre-computed before the iteration starts:

[0074]

[0075] In each iteration, instead of repeatedly calculating the graph filtering operation, the weighted sum of each row of Equation (6) is calculated using the filter weights learned by the model. By performing preprocessing at the beginning of the calculation, this strategy effectively reduces the redundant calculations in the iterative process, thus significantly improving the overall iterative efficiency of the algorithm.

[0076] Implementably, K-means clustering is respectively performed on the smoothed basis partitions and the original basis partitions; a joint optimization framework is constructed, and the joint optimization framework is optimized based on the block coordinate rotation method; based on the optimized joint optimization framework, the clustering results of the smoothed basis partitions and the original basis partitions are combined to generate the final multi-view clustering result. The process includes:

[0077] K-means is a classic clustering method. In this embodiment, orthogonal K-means clustering is constructed on each smoothed basis partition and the information in each smoothed basis partition is fused through a shared clustering indicator matrix. The specific optimization problem is formulated as:

[0078]

[0079] where α p is a non-negative view weight coefficient used to measure the clustering ability of each view; β v is a non-negative graph filter weight coefficient used to measure the contribution of each graph filter to the consensus graph filter; is the clustering centroid matrix learned for the p-th basis partition. By imposing the constraint the correlation between the centroids is eliminated to better distinguish clusters; is an element of the consensus clustering indicator matrix, where c is the number of clusters, and y ij ∈ {0, 1}.

[0080] Although graph filtering enhancement can effectively improve the clustering quality, the potential over-smoothing of graph filters may lead to the over-fusion of local structures that are originally important for discrimination, thereby weakening the discrimination ability of clustering. To solve this problem, this embodiment proposes a joint optimization framework, and the objective function of the joint optimization framework is to minimize the weighted sum of clustering errors on the smoothed base partition and the original base partition. By combining the multi-view K-means clustering of the base partition after graph filtering enhancement with the multi-view K-means clustering of the original base partition, it not only retains the advantages of graph filtering in enhancing the data structure expression but also effectively suppresses the negative impact of over-smoothing on the clustering accuracy through the joint optimization strategy.

[0081] Specifically, the objective function of the joint optimization framework can be expressed as:

[0082]

[0083] where λ is a trade-off hyperparameter between the clustering of the smoothed base partition and the clustering of the original base partition.

[0084] According to Equation (8), it can be concluded that: (1) From the perspective of the relationship between graph filter learning and clustering, the objective function of the joint optimization framework in this embodiment realizes the seamless integration of graph filter learning and clustering tasks by using the learned optimal consensus graph filter to enhance downstream multi-view clustering and using the optimized clustering results to guide the learning of the consensus graph filter coefficients. (2) From the perspective of the original base partition and the smoothed base partition, this embodiment performs joint clustering on the base partitions before and after filtering, fuses the original features and smoothed features of the base partition to avoid the over-smoothing effect caused by graph filtering, and uniformly characterizes the cluster structure by sharing the cluster centroids of the base partitions before and after filtering. (3) From the perspective of multi-view clustering and graph filter weights, this embodiment decouples and models the clustering weights and graph filter weights, which not only improves the characterization and interpretation ability of each group of weights but also simplifies the subsequent optimization process. It can be seen that the method of this embodiment makes full use of the collaborative enhancement of multi-graph filters, the fusion clustering of multi-views, and the joint optimization of multi-tasks, thereby achieving a better multi-view clustering effect.

[0085] Furthermore, the optimization solution of the joint optimization framework:

[0086] The objective function in formula (8) includes a total of four types of variables, namely Y, and This embodiment designs a block coordinate rotation method to alternately update all variables of problem (8).

[0087] (1) Iteratively update

[0088] After fixing other variables, the sub-problem of optimizing C p can be expressed as:

[0089]

[0090] Among them, Let the singular value decomposition of B p be expressed as Then the optimal solution of C p can be expressed as:

[0091]

[0092] (2) Iteratively update Y:

[0093] After fixing other variables, combining tr(Y T Y) = n, the sub-problem of optimizing Y can be formulated as:

[0094]

[0095] Among them, Then the solution of Y is:

[0096]

[0097] (3) Iteratively update

[0098] After fixing other variables, the optimization problem of can be formulated as:

[0099]

[0100] Among them, According to the Cauchy - Schwarz Inequality, we can get: When the equal sign holds. Therefore, the closed - form solution of α p is:

[0101]

[0102] (4) Iteratively update β:

[0103] After fixing other variables, the optimization problem of β can be formulated as:

[0104]

[0105] Among them, its elements are its elements This is a standard quadratic programming problem, and the optimal solution can be obtained using existing tools.

[0106] Implementable. The optimization algorithm proposed in this implementation has convergence guarantee. Specifically, the optimization problem in Equation (8) is solved by the block coordinate descent method. In this process, each sub-problem decomposed is a convex optimization problem and the global optimal solution can be obtained. Therefore, as each iteration proceeds, the objective function value will monotonically decrease. Since the objective function has a lower bound and is monotonically decreasing, the convergence of the algorithm can be ensured.

[0107] The computational complexity of the algorithm can be divided into multiple parts. The computational complexities of generating anchor points and constructing the bipartite graph are respectively and The complexity of the high-order graph filtering pre-computation is In the iterative optimization part, the computational complexity of updating is where t represents the iteration number of the algorithm. The computational complexity of updating Y is The computational complexity of updating α is The computational complexity of updating β is Since t, m, c, d << n, the overall computational complexity of the entire algorithm can be simplified to This indicates that the method proposed in this embodiment can be efficiently executed on large-scale data sets.

[0108] Theoretical analysis:

[0109] In this embodiment, the effects of the constructed high-order graph filter on the base partition quality and subsequent clustering results are analyzed from two aspects: graph filter and spectral graph theory.

[0110] Analysis from the perspective of graph filter:

[0111] The eigenvalues of the graph Laplacian matrix can reflect its structural characteristics: smaller eigenvalues correspond to the overall characteristics in the graph, such as cluster structures, while larger eigenvalues capture more subtle details and noises. Therefore, in order to achieve better clustering results, a low-pass graph filter can be used to weaken the larger eigenvalues while retaining the smaller eigenvalues. Given a graph Laplacian matrix L with eigenvalues λ1 ≤ λ2 ≤ … ≤ λ n . Let represent a certain transformation applied to the Laplacian matrix, which makes the eigenvalues become where is the transformation function. Definition of the low-pass graph filter:

[0112] Definition 1 (Low-pass graph filter): For a graph filter If there exists an integer 1 ≤ K < n and coefficients ζ, satisfying:

[0113]

[0114] Then is regarded as a (K, ζ) low-pass graph filter, where ζ is called the low-pass coefficient.

[0115] Theorem 1: The filter constructed by Equation (2) is a low-pass filter.

[0116] Proof: For perform singular value decomposition P v = QΣN T , and let its singular values be 1 ≥ σ1 ≥ σ2 ≥ … ≥ σ m ≥ 0. Therefore, the corresponding eigenvalues of the Laplacian matrix are denoted as {λ1, λ2, …, λ n} and satisfy 0 ≤ λ1 ≤ λ2 ≤ … ≤ λ m = … = 1. Denote the eigenvalues of the graph filter constructed by Equation (2) as In this embodiment, its low-pass coefficient ζ in Definition 1 is calculated as follows:

[0117]

[0118] Given the non-zero singular value σ v of P K > σ K+1 > 0, the corresponding eigenvalue of L v is λ K < λ K+1 . Therefore, there exists an integer 1 ≤ K < n such that 0 < λ K < λ K+1 , and the corresponding coefficients satisfy:

[0119]

[0120] Therefore, Equation (2) is a low-pass filter. As mentioned before, it can enhance small eigenvalues and suppress large eigenvalues to enhance the overall representation while suppressing noise, thereby revealing a clearer clustering structure.

[0121] Analyzed from the perspective of spectral graph theory:

[0122] In the post-fusion multi-view clustering task, the base partition H p can be regarded as a graph signal. If the base partition H p has a clear clustering structure, it should follow the clustering and manifold assumptions, that is, the data in the same class should be close to each other. Smooth graph signals tend to follow the clustering and manifold assumptions. The smoothness of graph signals can be described by Definition 2.

[0123] Definition 2 (Smoothness of Graph Signals): Given a normalized similarity graph whose corresponding normalized graph Laplacian matrix is \(L = I - S\), for any graph signal its smoothness is defined as:

[0124]

[0125] Next, in this embodiment, it will be proved in Theorem 2 that applying the low-pass filter of Equation (2) to \(H\) p to obtain \(G\) v \(H\) p results in a smoother graph signal.

[0126] Theorem 2: Applying the graph filter constructed by Equation (2) to \(H\) p to obtain the graph signal \(G\) v \(H\) p is smoother than the original graph signal \(H\) p .

[0127] Proof: In this embodiment, we first consider the \(i\)-th signal (i.e., the \(i\)-th column of the basis partition \(H\) p ). The analysis of other signals is similar. Next, this embodiment will prove that

[0128] the above inequality holds because Therefore, \(G\) v \(H\) p is a smoother graph signal.

[0129] According to Theorem 2, by applying the low-pass filter of Equation (2) to the basis partitions under each view, this embodiment can obtain a smoother basis partition, thereby obtaining a clearer clustering structure, which follows the clustering and manifold assumptions.

[0130] Applying the consensus filter of Equation (3) to the basis partition \(H\) p can be regarded as applying the low-pass filters (Equation (6)) constructed under different views to \(H\) p one by one, generating \(V\) smoothed basis partitions after being processed by each low-pass filter. Then, by weighted combining these smoothed results, the consensus filter-enhanced Through this process, not only can the signal be effectively smoothed and the noise influence be reduced, but also the diverse information between different views can be captured, so that the basis partition enhanced by the consensus filter has a clearer cluster structure.

[0131] Experiment:

[0132] This experiment aims to comprehensively evaluate the effectiveness of the algorithm (FGFMVC) in this embodiment. Two representative base partition generation processes are adopted in this embodiment to evaluate the performance of all methods. The FGFMVC is compared with 10 advanced post-fusion multi-view algorithms on 9 widely used real-world datasets in this embodiment, and the performance of each method is reported on multiple evaluation metrics. In addition, parameter sensitivity analysis and ablation experiments are conducted in this embodiment to further illustrate the feasibility of the joint clustering strategy for the original base partition and the smoothed partition enhanced by graph filtering.

[0133] Benchmark datasets:

[0134] Table 1 lists the basic information of the 9 benchmark datasets involved, including WebACE, K1B, MouseBladder, Zeisel, Macosko, COIL100, CITECBMC, TDT2, and MouseRetina. The number of samples and the number of clusters in these datasets are between 2340 and 27499 and between 6 and 100 respectively, covering various types from news articles, protein sequences to images and videos.

[0135] Comparison methods:

[0136] The method proposed in this embodiment is compared with 10 advanced multi-view post-fusion clustering methods, including AWP, LFMVC, OPLF, ALMVC, LFLKA, MMLMVC, ERMKC, sLGm, HKLMVC, and RIWLF. In addition, the baseline AvgH is also compared in this embodiment, that is, running K-means on each base partition and reporting its average performance.

[0137] Generation of base partitions:

[0138] Two common strategies are adopted in this embodiment to generate base partitions to evaluate the performance of the algorithm under different input environments.

[0139] Construction of the original kernel base partition:

[0140] Same as LFMVC, OPLF, and ALMVC, 12 original kernel matrices are constructed for each dataset here and the base partitions are generated through Kernel K-means. These kernel matrices include 7 Gaussian kernels, 1 linear kernel, and 4 polynomial kernels.

[0141] Construction of the local kernel base partition:

[0142] Same as LFLKA and sLGm, local structures are further mined on the basis of the 12 original kernel matrices here to obtain the corresponding local kernels, and the base partitions are also generated on the local kernels by means of Kernel K-means.

[0143] Table 1

[0144]

[0145] Experimental setup:

[0146] All comparison methods were executed following the parameter search and setting strategies recommended in their original literature. To mitigate the impact of randomness, each method was run 10 times, and the average results were reported. The number of clusters for all datasets was set to the true number of classes. FGFMVC contains two hyperparameters, the number of anchor points m and the trade-off parameter λ. In this embodiment, m and λ were selected from [4c, 6c, 8c, 10c] and [0, 0.1, …, 1] respectively through grid search, where c is the number of clusters. At the same time, the dimension d of the base partition was fixed to the number of clusters c, the heat kernel diffusion coefficient η was fixed to 9, and the number of neighbors k for bipartite graph construction was fixed to 5. Two commonly used clustering evaluation metrics were used in this experiment, including clustering accuracy (Accuracy, ACC) and adjusted Rand index (Adjusted Rand Index, ARI), to evaluate the clustering performance. All experiments were conducted in an environment equipped with an AMD Ryzen 7 5700G CPU (3.8 GHz), 64 GB of RAM, and MATLAB 2022b (64-bit).

[0147] Experimental results:

[0148] The clustering performance results of the FGFMVC algorithm and 10 comparison algorithms on 9 datasets under two sets of base partition settings are shown in Tables 2 and 3. The following conclusions can be drawn from this embodiment:

[0149] (1) The clustering results of FGFMVC under both types of base partitions are generally better than other methods. Tables 2 and 3 show that FGFMVC performs excellently in terms of both the ACC and ARI metrics, regardless of whether the input is the original kernel base partition or the local kernel base partition. Generally speaking, compared with the sub-optimal method, the average ACC of FGFMVC on all original kernel base partitions of the data is increased by 28%, and the average ARI is increased by 53%; on the local kernel base partition, the average ACC is increased by 16%, and the average ARI is increased by 24%. These results fully verify the superior performance of FGFMVC in fusing clustering under different base partitions.

[0150] (2) The enhancement effect of FGFMVC on the original kernel base partition is more significant, and the performance improvement amplitude is larger compared with the local kernel base partition. This may be because the local kernel has sparsified the data by extracting local structure information during construction, removing some noise and redundancy, making the base partition clearer and the gain space of graph filtering smaller. While the original kernel base partition contains more potential information and noise, and the graph filtering has a more obvious improvement effect on it in FGFMVC.

[0151] (3) Benefiting from the design of joint clustering, FGFMVC is significantly superior to other graph filtering enhancement methods in terms of performance. For example, HKLMVC also uses graph filtering but fails to effectively alleviate the over-smoothing degradation phenomenon and still lags behind FGFMVC. Through joint clustering on the original base partition and the filtered smooth base partition, FGFMVC not only exploits the advantages of graph filtering but also suppresses over-smoothing, thereby further improving the clustering performance.

[0152] Table 2

[0153]

[0154]

[0155] Table 3

[0156]

[0157] Running time comparison:

[0158] Comparing the 10-time average running times of all methods on each dataset under the original kernel base partition, the results of the local kernel base partition are similar. The results show that the running efficiency of FGFMVC is significantly better than that of MMLMVC, ERMKC, sLGm, and HKLMVC. Although the latter three methods achieve good clustering results on multiple datasets, their complexities are and This indicates that FGFMVC can not only obtain superior clustering performance but also achieve efficient computation, without sacrificing computational complexity for improving clustering accuracy. This further highlights the practicality of FGFMVC in the post-fusion multi-view clustering task.

[0159] Parameter sensitivity analysis:

[0160] FGFMVC contains two hyperparameters: the number of anchor points m and the trade-off parameter λ. In this embodiment, the parameter sensitivity of the algorithm is evaluated in the grid of parameters m = [4c, 6c, 8c, 10c] and λ = [0, 0.1, 0.2,..., 1]. Taking four datasets such as Zeisel, Macosko, COIL100, and MouseRetina as examples, the ACC results of FGFMVC under different parameter combinations in the two base partition settings of the original kernel and the local kernel are compared respectively.

[0161] It can be concluded that FGFMVC exhibits good stability under different parameter settings. Especially when the local kernel-based partition is used as the input, the stability of the parameters is more obvious. When the parameter value is large, FGFMVC can usually obtain better clustering results. However, on most datasets, when only using the smoothed basis partition enhanced by the graph filter, i.e., λ = 1, it does not achieve the optimal performance. This may be because the over-smoothed graph filter eliminates some key local structures, reducing the distinguishability of the cluster structures. In contrast, FGFMVC combines multi-view K-means on the smoothed basis partition and multi-view K-means on the original basis partition through joint clustering. This not only retains the data enhancement advantage of the graph filter but also effectively suppresses the negative impact of over-smoothing. According to the experimental results, the range of the parameter λ in the FGFMVC method is recommended to be set as 0.6 - 0.9.

[0162] Convergence analysis:

[0163] In this embodiment, taking four datasets such as Zeisel, Macosko, COIL100, and MouseRetina as examples, the change curves of the objective function value and ACC of FGFMVC on the original kernel-based partition with the number of iterations are obtained. It can be concluded that the objective function value of FGFMVC strictly monotonically decreases with the increase of the number of iterations, which verifies the convergence of the algorithm. In addition, this method can quickly converge within about 10 iterations on different datasets. At the same time, the ACC value increases with the decrease of the objective function value, further verifying the effectiveness of the FGFMVC method.

[0164] Ablation experiment:

[0165] In this embodiment, three ablation methods are designed to verify the effectiveness of the graph filter enhancement strategy in the FGFMVC algorithm. First, AvgH is used as the baseline, representing the result of single partition without graph filter enhancement. Second, the AvgH-GF method is adopted, where graph filter enhancement is applied to each basis partition, i.e., Equation (5), and then K-means is run and the average result is reported. Finally, the NGFMVC method is adopted, i.e., λ = 0, which is a multi-view clustering method without using the graph filter enhancement strategy. Tables 4 and 5 respectively show the ACC ablation experiment results of all datasets under the two settings of the original kernel and the local kernel.

[0166] By comparing AvgH and AvgH-GF, it can be found that the smoothed base partition enhanced by the graph filter significantly improves the average clustering performance, which verifies the enhancement effect of the graph filter on the single-view base partition. By comparing AvgH and NGFMVC, it can be concluded that without introducing the graph filter enhancement strategy, NGFMVC significantly outperforms the single-view result by integrating the comprehensive information of multiple views. By comparing AvgH-GF, NGFMVC, and FGFMVC, it can be found that FGFMVC further improves the clustering performance based on AvgH-GF and NGFMVC. The performance advantage of FGFMVC stems from its in-depth exploration of the high-order correlations of the base partition, enhancing the base partition smoothing through the consensus graph filter, and combining the joint clustering mechanisms before and after smoothing to effectively overcome the over-smoothing problem of the graph filter.

[0167] Table 4

[0168]

[0169]

[0170] Table 5

[0171]

[0172] Conclusion:

[0173] This embodiment proposes a multi-view clustering method based on fast graph filter augmentation, aiming to solve the problem of low-quality base partitions and achieve linear computational complexity. The method introduces the learning and fast filtering mechanism of the multi-view optimal consensus graph filter, induces the high-order graph filter based on the bipartite graph heat kernel diffusion, and realizes the linear combination of multiple graph filters in the multi-view clustering framework. This method significantly improves the quality of the consensus graph filter and the clustering effect, while maintaining high-efficiency linear complexity. In addition, this embodiment also proposes a new multi-view clustering framework, which gives full play to the advantages of the graph filter by performing joint clustering on the base partition and the smoothed partition enhanced by the graph filter, and effectively suppresses the negative impact of the over-smoothing phenomenon on the clustering performance. Under the construction of two common base partitions, the experimental results of multiple datasets verify the effectiveness and superiority of this method in multi-view clustering, and the ablation experiment verifies the effectiveness of the graph filter and the rationality of the joint clustering.

[0174] Embodiment 2

[0175] This embodiment also provides a multi-view clustering system with fast graph filter augmentation for implementing the above method, including: a bipartite graph construction module, a filter construction module, a smoothing processing module, a joint clustering module, and an iterative update module;

[0176] The bipartite graph construction module is used to generate an anchor matrix for the original base partition of each view through K-means clustering, and construct a bipartite graph based on the probability proximity weights between samples and anchors;

[0177] The filter construction module is used to construct a low-pass filter based on the bipartite graph based on the heat kernel diffusion mechanism, and perform a linear combination of several low-pass filters to obtain a consensus graph filter for multi-view clustering;

[0178] The smoothing processing module is used to perform smoothing enhancement processing on the original base partition of each view based on the consensus graph filter to obtain a smoothed base partition;

[0179] The joint clustering module is used to perform K-means clustering on the smoothed base partition and the original base partition respectively, construct a joint optimization framework, and combine the clustering results of the smoothed base partition and the original base partition based on the joint optimization framework to generate a multi-view clustering result;

[0180] The iterative update module is used to iteratively update the joint optimization framework based on the block coordinate rotation method, and further generate the final multi-view clustering result.

[0181] Embodiment 3

[0182] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method.

[0183] Embodiment 4

[0184] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method are implemented.

[0185] Embodiment 5

[0186] This embodiment also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method are implemented.

[0187] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A fast graph filtering augmented multi-view clustering method, characterized in that It includes the following steps: For the original base partition of each view, generate an anchor matrix through K-means clustering, and construct a bipartite graph based on the probability proximity weights between samples and anchors; Based on the heat kernel diffusion mechanism, construct a low-pass filter based on the bipartite graph, and perform a linear combination of several low-pass filters to obtain a consensus graph filter for multi-view clustering; Based on the consensus graph filter, perform smoothing enhancement processing on the original base partition of each view to obtain a smoothed base partition; Perform K-means clustering on the smoothed base partition and the original base partition respectively; Construct a joint optimization framework, and optimize the joint optimization framework based on the block coordinate rotation method; Based on the optimized joint optimization framework, combine the clustering results of the smoothed base partition and the original base partition to generate the final multi-view clustering result.

2. The method according to claim 1, wherein: The process of constructing a low-pass filter based on the bipartite graph based on the heat kernel diffusion mechanism includes: Introduce a diagonal matrix into the bipartite graph to obtain a sample-anchor similarity graph; perform column normalization on the sample-anchor similarity graph to obtain a sample-anchor-sample similarity graph; based on the heat kernel diffusion mechanism, perform weighted combination on each order of sample-anchor-sample similarity graph to obtain the corresponding low-pass filter.

3. The method according to claim 2, wherein: The formula of the low-pass filter is as follows: where t represents the order of diffusion, is the weight coefficient of the heat kernel diffusion, S v is the sample-anchor-sample similarity graph, η is a non-negative parameter controlling the decay rate, L v is the normalized Laplacian matrix constructed from the v-th view, G v is the low-pass filter with respect to L v of.

4. The method according to claim 3, wherein: The formula of the consensus graph filter is as follows: where β v is the weight of the v-th graph filter.

5. The method according to claim 1, wherein: The objective function of the joint optimization framework is to minimize the weighted sum of clustering errors on the smoothed base partition and the original base partition, and the optimization variables include a clustering indicator matrix, a clustering centroid matrix, view weights, and graph filter weights.

6. The method according to claim 5, wherein: The formula of the objective function of the joint optimization framework is as follows: where α p is a non - negative view weight coefficient, β v is a non - negative graph filter weight coefficient, λ is a trade - off hyperparameter between the smoothed base partition clustering and the original base partition clustering, G v is a low - pass filter with respect to L v , H p is the original base partition, Y is the consensus clustering indicator matrix, and C p is the clustering centroid matrix learned from the p - th base partition.

7. A multi-view clustering system with fast graph filtering augmentation, characterized in that The apparatus for implementing the method according to any one of claims 1-6 includes: a bipartite graph construction module, a filter construction module, a smoothing processing module, a joint clustering module, and an iterative update module; The bipartite graph construction module is used to generate an anchor matrix through K-means clustering for the original base partition of each view, and construct a bipartite graph based on the probability proximity weights between samples and anchors; The filter construction module is used to construct a low-pass filter based on the bipartite graph based on the heat kernel diffusion mechanism, and perform a linear combination of several low-pass filters to obtain a consensus graph filter for multi-view clustering; The smoothing processing module is used to perform smoothing enhancement processing on the original base partition of each view based on the consensus graph filter to obtain a smoothed base partition; The joint clustering module is used to perform K-means clustering on the smoothed base partition and the original base partition respectively, construct a joint optimization framework, and based on the joint optimization framework, combine the clustering results of the smoothed base partition and the original base partition to generate a multi-view clustering result; The iterative update module is used to iteratively update the joint optimization framework based on the block coordinate rotation method, and further generate the final multi-view clustering result.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.