Multi-view collaborative fusion method and system suitable for multi-view data clustering

Through the multi-graph collaborative fusion method, the sub-graph similarity matrix is ​​adaptively learned and fused into a unified similarity graph, which solves the problems of information redundancy and loss in multi-view data clustering, and achieves high-precision and robust clustering effect.

CN119992271APending Publication Date: 2025-05-13LINGBO ZHIXIN (CHANGZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510097244.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Multi-view data clustering can easily lead to the redundancy of modal overlap and the loss of modal exclusive information during data fusion.

Method used

The multi-graph collaborative fusion method is adopted to construct the sub-graph similarity matrix through adaptive learning. The fused sub-similar graph is a unified similarity graph, and spectral clustering is performed on the basis of the unified similarity graph. Finally, the K-means algorithm is used to complete the clustering task.

Benefits of technology

It effectively avoids information redundancy and information loss, improves clustering accuracy and robustness, can capture similarity information between samples more accurately, and provides a flexible optimization method for clustering results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992271A_ABST
    Figure CN119992271A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view collaborative fusion method and system suitable for multi-view data clustering, and relates to the technical field of computer data clustering. The method comprises the following steps: constructing a sub-graph similarity matrix by adopting adaptive learning; reserving sample similarity information, and fusing a plurality of sub-similar graphs by adopting adaptive learning to form a unified similar graph; and on the basis of unifying the similar graph, extracting low-dimensional embedding representation of the samples and finishing final clustering. According to the MCBMF, mutual reinforcement learning of similar graph construction and unified graph fusion is adopted, closer data capture is carried out, samples are initialized in multiple distance measurement modes, the similarity and importance relation between each view and the internal samples of the view is accurately measured, and therefore distribution characteristics of complex data are comprehensively captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer data clustering, and in particular to a multi-graph collaborative fusion method and system suitable for multi-view data clustering. Background Art

[0002] Multi-view data clustering belongs to the research scope of multi-task learning, and is also a clustering algorithm. The purpose of this type of algorithm is to make full use of the common information and private information of different views of the same object and the relationship between them to enhance the performance of the model. Specifically, it makes full use of the private information of different views (the difference between the view and other views). This approach will make the effective information between different views complementary and integrated, and obtain more comprehensive and comprehensive information. The combined use of multi-view data can effectively make up for the missing information of a single view and reduce the impact of the noise of a single view on the model, thereby retaining the most valuable information parts of different views and improving the performance of the model. The algorithm finally makes full use of the information of a single object at different angles to complete the clustering task of the object with high accuracy. For example: to identify the label of a website, by using the text content, hyperlinks, and viewing crowd of this website; in the problem of image recognition, the color, texture, shape, etc. of images from different perspectives are used for analysis and clustering.

[0003] Many solutions have been proposed in recent years around this problem, which can be divided into five categories:

[0004] 1. Co-training algorithm;

[0005] This type of algorithm aims to train two classifiers, maximize their interaction, and iterate continuously so that the two classifiers eventually reach the same clustering view of the samples. The specific steps are as follows:

[0006] 1) Separate the labeled data and unlabeled data into two different datasets (views), and train the first classifier with the labeled data.

[0007] 2) A portion of the samples with higher prediction confidence obtained by training the first classifier (as knowledge) is added to the second view as labeled samples, and these newly added samples are used to train the second classifier.

[0008] 3) Use the second classifier to train the unlabeled data, and add a portion of samples with higher confidence (as knowledge) to the first view as labeled samples, and use these newly added samples to train the first classifier.

[0009] Repeat the above content to increase the number of labeled samples and the two classifiers to exchange information continuously, and finally converge to the consistency of sample categories (consistent with the category of samples).

[0010] 2. Multiple kernel learning;

[0011] The purpose of multi-core learning is to effectively improve the search space capacity of existing kernel functions (linear kernel functions, polynomial kernels, and Gaussian kernels) and effectively improve the generalization ability of the model. Since the kernels in multi-core learning can correspond to multiple views, multi-core learning is widely used in the processing of multi-view data. The specific steps of multi-core learning are as follows:

[0012] 1) Different views can be processed by different kernel functions to obtain the kernel matrix.

[0013] 2) Learn the weights of different kernel matrices through various methods such as Bayesian Information Criterion or cross-validation.

[0014] 3) Fusing the kernel functions on different views and the corresponding kernel matrices through linear or nonlinear superposition.

[0015] Finally, a clustering algorithm is performed on the fused kernel matrix.

[0016] 3. Multi-view subspace clustering;

[0017] Multi-view subspace clustering is a unified clustering method based on multiple views of data. The main strategy of this method is to embed sample data into the subspace, reduce the amount of data, and remove weak noise information. Then the data of different subspaces are unified into one subspace, and finally, based on the formed unified subspace sample data, the similarity between different samples is measured to perform clustering tasks. The main process of the multi-view subspace algorithm is as follows:

[0018] 1) For each view, initialize the subspace and learn the low-order embedding form of the sample in the subspace to obtain a low-dimensional representation of the sample. Common methods include principal component analysis and autoencoder.

[0019] 2) Fusion: Fusion the sample data of each subspace into a new subspace and calculate the similarity of samples in the fused subspace.

[0020] 3) Clustering: clustering algorithm is used to perform clustering tasks on similar graphs obtained by fusion subspace calculation with lower complexity.

[0021] 4. Multi-task multi-view clustering;

[0022] Multi-task multi-view clustering effectively utilizes the consistency and complementarity of different views. The algorithm executes multiple tasks together, extracts the features of related tasks from different views, and combines multiple tasks to learn the same and shared views of different tasks on the same sample (that is, based on the connection between different tasks, it emphasizes the influence between tasks). Use clustering algorithms to divide tasks into different clusters, and finally perform clustering tasks on different sets and then integrate the results. The specific steps of multi-task multi-view clustering are as follows:

[0023] 1) Learning one or more feature tasks from a single view.

[0024] 2) Combine multiple feature tasks and views to learn multi-task representations.

[0025] 3) Divide multiple tasks into different clusters through clustering algorithms.

[0026] 4) Integrate the results of each cluster to get the final clustering result.

[0027] 5. Multi-view graph clustering;

[0028] Multi-view graph clustering is different from multi-kernel function. The main task of multi-kernel function is to learn a better kernel function and continuously improve the learned kernel function and the corresponding weights. Multi-view graph clustering, on the other hand, learns a similarity graph from each view. The key is to fuse the similarity graphs formed by these views into a unified graph, and perform graph cutting on the unified graph (commonly spectral clustering), and then complete the clustering task. It is worth mentioning that there are usually two research directions for similar clustering tasks on a unified graph: the first is to perform the clustering task when cutting the graph, and the clustering task is completed after the graph is cut. The other is to perform an independent clustering task after the entire graph is cut. The detailed steps of multi-view graph clustering are shown below:

[0029] 1) Generate a sub-similarity graph through each view according to a pre-defined similarity measurement rule.

[0030] 2) Learn the comprehensive similarity graph by fusing the features of each sub-similarity graph (graph fusion).

[0031] 3) Cut the comprehensive similarity graph and cut off some weakly connected edges to preprocess the subsequent clustering tasks.

[0032] 4) Perform clustering tasks to complete the cluster division of samples.

[0033] At present, some studies have improved the multi-view clustering method and proposed Multiview Consensus Graph Clustering, which uses the structure of the graph to improve the clustering performance. The main three innovations are:

[0034] Consensus Graph Learning: The authors propose to learn a consensus graph that minimizes the inconsistency between different views. They avoid relying on a predefined similarity matrix by imposing a rank constraint on the Laplacian matrix to ensure that the graph structure has connected components corresponding to the number of clusters.

[0035] New cost function: A new inconsistency cost function is used to regularize the multi-view graph structure into a consensus. The Laplacian matrix rank is constrained so that it automatically forms kkk connected components (i.e., kkk clusters).

[0036] Obtain clustering results: After the consensus graph is learned, no further post-processing (such as k-means clustering) is required.

[0037] This multi-view Figure 1 MCGC has certain advantages in overcoming the limitations of traditional single-view data clustering methods by simultaneously utilizing the collaborative information of multiple views to improve clustering performance. However, MCGC also has the following shortcomings:

[0038] Independence problem in the model: In the MCGC method, the construction and fusion of the similarity graph of each view are separate, which may result in the information of each view not being able to fully reflect its internal relationship.

[0039] Lack of post-processing: One of the characteristics of MCGC is that it does not require post-processing, such as K-means clustering, but directly uses strongly connected components to obtain clustering results. However, this may lack flexibility in some scenarios and cannot optimize the initial clustering results.

[0040] Limitations of loss function design: Although MCGC weighs the difference between the similarity graph of each view and the shared graph through a new loss function, the design of the loss function may not be sufficient to fully capture the distribution characteristics of complex data.

[0041] The present invention belongs to a multi-view graph clustering method, and proposes a graph clustering method based on multi-graph fusion (MCBMF, Multi-view data clustering based on multi-graph fusion), which is different from the algorithms used in previous studies to initialize sample similarity, including Euclidean distance, Hamiltonian distance, Gaussian kernel function, etc. MCBMF adopts an adaptive similarity strategy to learn the similarity relationship between samples in each subgraph. By continuously iteratively updating and refining the similarity value, the algorithm can more accurately capture the similarity information between different samples. In addition, previous multi-view clustering algorithms based on similarity graphs usually consider the importance of different views, because the importance of different view data is not consistent. However, such algorithms ignore the different importance of samples. MCBMF improves this defect and considers the weights of samples in different views. After adaptively learning the subgraph similarity graph and fusing the sub-similar graphs into a unified graph, MCBMF uses a classic spectral clustering algorithm to learn the embedding results of samples in the subspace from the unified graph. Finally, the K-means algorithm is used in the sample subspace to complete the final clustering task and obtain the clustering information of the sample. At the same time, the present invention further overcomes the above-mentioned multi-view Figure 1 Problems with Multi-Content Graph Clustering (MCGC). Summary of the invention

[0042] The purpose of the present invention is to provide a multi-graph collaborative fusion method and system suitable for multi-view data clustering, so as to solve the problems of information redundancy of modal overlap and modal exclusive information loss caused by multi-view data clustering (MvC) in the data fusion process.

[0043] To achieve the above object, the present invention adopts the following technical solutions:

[0044] In a first aspect, the present invention proposes a multi-graph collaborative fusion method suitable for multi-view data clustering, which comprises at least the following steps:

[0045] S1, generate sub-similarity graph; construct sub-graph similarity matrix using adaptive learning;

[0046] S2, merging sub-similar graphs into a unified similarity graph; retaining sample similarity information, and using adaptive learning to fuse multiple sub-similar graphs to form a unified similarity graph;

[0047] S3. Obtain the clustering information of the samples; based on the unified similarity graph, extract the low-dimensional embedding representation of the samples and complete the final clustering.

[0048] Preferably, the S1 is as follows:

[0049] First, the objective function is established to maintain the structure of the graph. Second, the geometric similarity is maintained, and normalization is used to change the direction characteristics without changing the length of the subsequent similarity vector. After that, the sparsity of the subgraph similarity matrix is ​​guaranteed, and the l_{F} norm is added for sparsity constraint.

[0050] The obtained subgraph similarity matrix S is as follows:

[0051]

[0052] in, S (v) ∈R N×N , d v represents the number of samples of the vth subview, N represents the dimension of the sample, and V is the number of sub-images; and Respectively represent the i-th and j-th columns in the v-th sub-graph data; Represents the data in the i-th column of the similarity graph, Represents the similarity between the i-th and j-th samples in the v-th subgraph; 1 N represents a full-1 column vector of length N, and β is a hyperparameter used to characterize the weights of different learning modules.

[0053] Preferably, S2 is specifically as follows:

[0054] Based on the linear weighted superposition method, a unified global similarity graph matrix U is generated. In this process, the sample similarity information is retained by adding an adaptive weight linear construction regularization term. The details are as follows:

[0055]

[0056] Among them, the matrix U∈R N×N is the learned unified similarity graph, U i Represented as the i-th column data of the unified similarity graph; matrix W∈R V×N is the adaptive weight matrix, Represents the i-th column of the v-th subgraph weight matrix.

[0057] Preferably, S3 is specifically as follows:

[0058] The RatioCut graph cutting strategy in spectral clustering is adopted. On the basis of the unified similarity graph, the graph is cut according to the edge weight information, and the Laplace matrix is ​​used to ensure its symmetry; the final matrix is ​​as follows:

[0059]

[0060] Among them, L U represents the Laplace matrix; λ represents the hyperparameter; D UThe matrix is ​​initialized as

[0061] Furthermore, the final matrix uses an alternating iterative method to solve the variables therein one by one; specifically: other matrices are fixed as constants by default, and one of the matrices is regarded as a variable to solve the update rule.

[0062] Furthermore, the solution process is as follows:

[0063] Fix W, U and H and update S^{(v)} as follows:

[0064]

[0065] Fix S^{(v)}, U and H, and update W as follows:

[0066]

[0067] Fix S^{(v)}, W and H, and update U as follows:

[0068]

[0069] Fix S^{(v)}, W and U, and update H as follows:

[0070]

[0071] s H T H=I.

[0072] The second aspect of the present invention proposes a multi-graph collaborative fusion system (i.e., MCBMF model) suitable for multi-view data clustering, including:

[0073] The local structure preservation module of sub-view data adopts adaptive learning of sub-graph similarity matrix to generate sub-similarity graph;

[0074] The multi-graph fusion module based on adaptive learning adopts linear weighted superposition and adds adaptive weight linear construction regularization terms to fuse multiple sub-similar graphs into a unified similar graph;

[0075] The subspace embedding module based on spectral clustering adopts the RatioCut graph cutting strategy in spectral clustering to learn the final representation of the sample in the subspace and obtain the low-dimensional embedding representation of the sample.

[0076] Compared with the prior art, the present invention has the following beneficial effects:

[0077] (1) To address the shortcoming of MCGC, that is, the construction and fusion processes of the similarity graph of each view are separate problems, the present invention adopts the mutual reinforcement learning of similarity graph construction and unified graph fusion to achieve closer data capture.

[0078] (2) In the present invention, MCBMF uses a variety of distance metrics to initialize samples and accurately measures the similarity and importance relationship between each view and its internal samples, thereby fully capturing the distribution characteristics of complex data.

[0079] (3) In the present invention, MCBMF adopts a more flexible clustering method, which provides the possibility of subsequent data processing to optimize the clustering results. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 The structural block diagrams of different multi-view data clustering algorithms proposed in the background technology ((a) is a collaborative training algorithm, (b) is a multi-core learning algorithm, (c) is a multi-view subspace clustering algorithm, and (d) is a multi-task multi-view clustering algorithm);

[0081] Figure 2 It is a structural block diagram of a multi-graph collaborative fusion system suitable for multi-view data clustering in the present invention;

[0082] Figure 3 It is a bar graph of ACC evaluation results in the present invention;

[0083] Figure 4 Schematic diagram of the objective function value convergence curve of MCBMF on four different data sets in the present invention;

[0084] Figure 5 Schematic diagram of unified graph visualization of MCBMF on four different data sets in the present invention;

[0085] Figure 6 Schematic diagram of the k value parameter test of MCBMF on four different data sets in the present invention. DETAILED DESCRIPTION

[0086] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0087] Embodiment 1:

[0088] This paper proposes a graph clustering scheme based on multi-view fusion, MCBMF, which aims to solve the problem of data fusion and data clustering under multi-modality and ensure the accuracy of data clustering. The flowchart of MCBMF is as follows: Figure 2 shown.

[0089] The model is composed of three process modules, namely, the sub-view data local structure preservation module, which is used to adaptively learn the sub-graph similarity matrix; the multi-graph fusion module based on adaptive learning, which is used to fuse multiple sub-graphs to form a fusion graph; and the subspace embedding module based on spectral clustering, which is used to learn the final representation of the sample in the subspace. The details are as follows:

[0090] Step 1: Design a local structure preservation module for sub-view data; construct a sub-similarity graph of multi-view data through adaptive learning to more accurately capture the similarity between samples.

[0091] In previous studies, multi-graph learning based on similarity strategy often constructs a similarity matrix through a fixed rule. For example, the Hammond distance, Euclidean distance, and Gaussian kernel are used to construct similarity. In the present invention, an adaptive learning method is used to study the similarity matrix of the constructed subgraph. Its objective function is as follows:

[0092]

[0093] Among them, among them, S (v) ∈R N×N , d v Represents the number of samples of the vth subview, N represents the dimension of the sample, and V is the number of subimages. and They represent the i-th and j-th columns in the v-th sub-graph data respectively. Represents the data in the i-th column of the similarity graph, Represents the similarity between the i-th and,th samples in the v-th subgraph. N represents a column vector of length N that is all ones.

[0094] Formula (1) uses the recently popular popular learning to maintain the structure of the graph. Under this conditional constraint, the distance between samples in the high-dimensional space is measured. If the distance is close, then the similarity between them is high, otherwise, the similarity between them is low. At the same time, in order to maintain the similarity in the geometric sense, normalization is used so that when the subsequent similar vector changes, only the directional characteristics are changed without changing the length. Secondly, according to a priori, the sparse property has good robustness to noise and missingness. In order to further ensure the robustness of the algorithm, that is, to ensure the sparsity of the S^{(v)} matrix, the l_{2,1} norm is added for sparsity constraints.

[0095]

[0096] Among them, β is a hyperparameter used to characterize the weights of different learning modules.

[0097] In the subsequent optimization process, l_{2,1} makes it difficult to solve, so the model is relaxed and replaced by l_{F} constraint. In order to ensure its performance, the present invention adds a non-zero number constraint in the process of sub-similarity matrix optimization. Then formula (2) is simplified as follows:

[0098]

[0099] Step 2: Design a multi-graph fusion module based on adaptive learning; generate a unified global similarity graph matrix U based on the linear weighted superposition method.

[0100] Linear construction is very common in multi-graph learning of similarity strategies. Such algorithms usually equip each view with a separate weight, so as to obtain the information of the fusion graph by weighted sum. MCBMF not only considers the importance of different views from the perspective of views, but also considers the importance of different samples from the perspective of samples. For example: assuming that the current view data is an image, and the current task is car recognition, then the importance of the set of samples composed of pixels in the area where the car logo is located must be much higher than the sample set of other irrelevant backgrounds. Therefore, the importance of different samples is also different. Adding an adaptive weight linear construction regularization term and simplifying it, we get the following loss function:

[0101]

[0102] Among them, the matrix U∈R N×N is the learned fusion graph, U i Represented as the i-th column data of the fusion graph. Matrix W∈R V×N is the adaptive weight matrix, Represents the i-th column of the v-th subgraph weight matrix.

[0103] Regarding the setting of hyperparameters, the objective loss function only contains the beta parameter. As mentioned above, the S matrix should have a non-negative property, and the similarity matrix does not have a value less than 0. However, in the actual optimization update process, the iterative form of the S matrix may result in the generation of negative numbers. In order to maintain non-negativity, it is updated to (S_{ij})_+. However, doing so will result in the loss of similarity information between the two samples. In the case of ideal hyperparameters, it is possible to update to only retain appropriate sample similarity information, but in reality, the existence of hyperparameters will lead to too much redundant similarity information, or too little similarity information will be retained. In order to retain the appropriate information just right, only one hyperparameter is retained, which is used as a variable to control the learning of appropriate similarity information during the S^{(v)} learning process. So far, the present invention has completed the modeling of the learning process from the subgraph similarity matrix to the unified graph. The final matrix U is a unified graph that comprehensively considers the importance of samples and the importance of views.

[0104] Step 3: Design a subspace embedding module based on spectral clustering; based on the unified similarity graph, extract the low-dimensional embedding representation of the sample and complete the final clustering.

[0105] On the basis of the unified graph, it is necessary to cut the graph according to the edge weight information, cut off the redundant weakly connected (small edge weight) edges, leave the effective and determined edge weight information, and maintain the accuracy of the model. The present invention adopts the very popular RatioCut graph cutting strategy in spectral clustering. Due to a series of reasons such as update order and weight, the matrix U may not be guaranteed to be symmetric during the learning process. Since the undirected graph must be a symmetric graph as a bidirectional edge structure, the Laplace matrix will definitely ensure its symmetry, so the present invention constructs the Laplace matrix in the following way:

[0106]

[0107] This ensures that the constructed Laplace matrix is ​​symmetric. In summary, the final loss function model is:

[0108]

[0109] Among them, L U represents the Laplace matrix; λ represents the hyperparameter, D U The matrix is ​​initialized as

[0110] It is difficult to solve equation (6) directly, so the present invention uses an alternating iterative approach to solve the variables one by one.

[0111] Specifically, the other matrices are fixed as constants by default, while one of the matrices is regarded as a variable, and a series of methods such as the Lagrange multiplier method (ALM) are used to solve the update rule.

[0112] 1) Fix W, U and H and update S^{(v)} as follows:

[0113]

[0114] 2) Fix S^{(v)}, U and H, and update W as follows:

[0115]

[0116] 3) Fix S^{(v)}, W and H, and update U as follows:

[0117]

[0118] 4) Fix S^{(v)}, W and U, and update H as follows:

[0119]

[0120] s H T H=I (10)

[0121] The overall algorithm flow pseudo code is as follows:

[0122] Algorithm: Alternating iterative algorithm based on multi-graph fusion;

[0123] Input: Data matrix and hyperparameters λ>0 and σ>0;

[0124] Output: Sample clustering results:

[0125] Initialization: random matrices U, W, H satisfying the constraints of equation (6) and S calculated by equation (7) (V) ;

[0126] When not converged, execute:

[0127] 1: Update S according to formula (7) (V) ;

[0128] 2: Update W according to formula (8);

[0129] 3: Update U according to formula (9);

[0130] 4: According to L U Update H by the feature vectors corresponding to the c smallest features.

[0131] Perform clustering on the matrix using the K-means algorithm and return the sample clustering results.

[0132] Experimental verification:

[0133] 1) Select a dataset;

[0134] Evaluation experiments of multi-view clustering based on multi-image fusion are conducted on four public datasets1, as shown in Table 1. The detailed introduction of the datasets is as follows:

[0135] 100leaves: A dataset of leaf images. This dataset comes from the UCI database and consists of 1,600 samples of 100 plant species. It contains three independent views and is used as a dataset for multi-image learning tasks.

[0136] 3-sources: Multi-view text dataset. This dataset consists of news articles from three online news services, providing a total of 948 news article datasets, where each news service source corresponding to the article is considered as a view. This dataset selects 169 articles from these 948 articles to form the current dataset.

[0137] Hdigit: Handwritten digit image dataset. This dataset is composed of two sources: the MNIST handwritten digit dataset and the USPS handwritten digit dataset, so this dataset corresponds to two views. It contains a total of 10,000 samples.

[0138] NGs: Multi-view text dataset. It is a subset of the 20 newsgroups dataset. NGs consists of 500 newsgroup documents. Each raw document is preprocessed using three different methods (giving three viewpoints, views) and annotated with one of five topic labels.

[0139] Table 1 Dataset description

[0140] Dataset Number of samples Number of views Number of categories <![CDATA[d1]]> <![CDATA[d2]]> <![CDATA[d3]]> 100leaves 1600 3 100 64 64 64 3-sources 169 3 6 3560 3631 3068 Hdigit 10000 2 10 784 256 0 NGs 500 3 5 2000 2000 2000

[0141] 2) Comparison algorithm:

[0142] In order to verify the feasibility of the proposed algorithm, this paper uses 10 multi-view clustering algorithms for comparative analysis. The relevant algorithm introduction is as follows:

[0143] (1) MKC (Multi-view K-means Clustering) (Xiao et al. 2013): The main idea of ​​the MKC algorithm is to integrate the differences of multiple perspectives into one optimization objective to achieve comprehensive clustering.

[0144] (2) Multi-view clustering via Non-negative Matrix Factorization (Liu et al. 2013): MultiNMF decomposes the view matrix into the product of non-negative matrices, selects the feature selection matrices in it and cross-constructs a new feature selection matrix. The clustering task is performed on the constructed matrix.

[0145] (3) MSC (Multi-view Spectral Clustering) (Xia et al. 2014): The MSC algorithm is a spectral clustering method based on multiple perspectives. The algorithm combines similarity information from multiple perspectives, comprehensively represents the intrinsic structure of the data through a unified framework, and achieves clustering.

[0146] (4) ASMV (Adaptive Structure-based Multi-view clustering) (Zhan et al. 2018): ASMV is a multi-view clustering method based on adaptive structure. The algorithm uses the correlation between data from different perspectives, transforms the data into a universal low-dimensional space, and realizes clustering.

[0147] (5) MGL (Multiple Graph Learning) (Nie, Jing, et al. 2016): The core idea of ​​MGL is based on the concept of mutual learning. The algorithm first decomposes the input data (such as an image) into multiple parts and constructs the corresponding graph structure. For each graph, MGL uses methods such as graph convolution and graph attention mechanism to learn the relationship and feature representation between nodes. Between different graphs, a mutual learning mechanism is used to transfer information and optimize the model.

[0148] (6) CoregSC (Co-regularized Spectral Clustering) (Kumar et al. 2011b): CoregSC is an algorithm for clustering analysis proposed by Canyi Lu et al. in 2016. The algorithm is based on two basic concepts: spectral clustering and co-regularization, and aims to solve the problem of inaccurate clustering results when there are noise and outliers in the data set.

[0149] (7) AMGL (A framework for multiview clustering and semi-supervised classification) (Nimrod and Ron nd): AMGL is a framework for multiview clustering and semi-supervised classification that can automatically weigh the contributions of features from different views without manually setting hyperparameters. The algorithm is based on the idea of ​​graph neural network (GNN) and subspace aggregation and can effectively process datasets with multiple views.

[0150] (8) MCGC (Multiview consensus graph clustering) (Zhan et al. 2019): In the MCGC algorithm, the construction of the consistency graph and the subgraph decomposition process are key steps. In order to construct the consistency graph, MCGC adopts online learning and dynamic adjustment methods to ensure that the graph can better reflect the relationship between samples. At the same time, in order to perform subgraph decomposition, MCGC adopts a method based on sparse subspace aggregation, emphasizing feature fusion while retaining data structure information.

[0151] (9) NEMO (Cancer subtyping by integration of partial multi-omic data) (Wang et al. 2014): NEMO is a novel multi-omics clustering algorithm that can perform clustering on partial data without data interpolation. It uses a neighborhood-based approach to construct relationships between samples and uses some feature selection strategies to fuse multiple omics datasets for clustering.

[0152] (10) SNF (Similarity network fusion for aggregating data types on a genomic scale) (Nie, Li, et al. 2016): The SNF method is a multi-group data integration method based on network fusion, which aims to effectively integrate data from different types. This method uses similarity networks as a bridge between data sets, fuses multiple data sets into an overall network, and partitions the network through a spectral clustering algorithm to obtain interactive information between partitions.

[0153] 3) Evaluation indicators:

[0154] The MCBMF evaluation indicators include accuracy rate ACC, standardized mutual information NMI, and adjusted Rand coefficient ARI.

[0155] 4) Experimental analysis:

[0156] The ACC evaluation values ​​of the clustering results can be seen in Table 2. In order to highlight the algorithms with excellent performance, the corresponding values ​​of the algorithms that achieved the best and second best performance on each data set are bolded and visualized as a bar chart ( Figure 3 ).

[0157] Table 2 ACC comparison test results

[0158]

[0159] As shown in Table 2, the MCBMF algorithm shows good performance on all data sets. Among them, the SNF algorithm achieved the most outstanding performance in 100leaves, and achieved very high evaluations with MCBMF on the large-scale Hdigit data set. However, the SNF algorithm has a certain gap with MCBMF on the other two data sets, indicating that the robustness of the SNF algorithm is not as good as MCBMF when adapting to different types and distributions of data sets. It also shows that MCBMF has strong robustness in addition to good performance. Among them, the graph-based algorithms include ASMV, MGL, SNF and MCG C, MCBMF. Except for MGL, other algorithms have very good results. It shows that graph algorithms are very suitable for multi-view clustering tasks. At the same time, MCBMF has achieved better performance because it fully mines the information in the graph. In addition, the effect of MGL on different data sets fluctuates greatly because the algorithm is highly sensitive to the k value in K-means. MultiN MF produced very poor results on the NGs data set because the algorithm is not suitable for processing data with negative values.

[0160] The NMI and ARI evaluation values ​​of the clustering results can be seen in Tables 3 and 4.

[0161] Table 3 NMI comparison test results

[0162]

[0163] Table 4 ARI comparison experimental results

[0164]

[0165]

[0166] As shown in Tables 3 and 4, under the NMI evaluation index, except for the 100leaves dataset, MCB MF has achieved the highest performance on other datasets, especially on NGs. Since the data distribution and type of the NGs dataset are different from other datasets, some algorithms have poor robustness and cannot adapt well to this dataset, resulting in poor algorithm performance under ARI evaluation. For example: MultiNMF, MSC, ASMV, etc. However, MCBMF can also achieve an excellent performance of 95.54% on this dataset.

[0167] In summary, the experimental results in Table 2, Table 3 and Table 4 show that the MCBMF algorithm has excellent performance and robustness.

[0168] In addition to the normal data set, the present invention also adds an Average column, which is the average of the evaluation of the corresponding algorithm on the four data sets under the evaluation criteria. Through this value, the comprehensive performance of each algorithm can be more directly understood. The data shows that MCBMF has the highest Average value in all three evaluation criteria, followed by MGL and SNF, two graph algorithms.

[0169] Furthermore, the experimental design calculates the objective loss function value of MCBMF on different data sets in each iteration. When the difference between the objective function value of this round and the objective function value of the previous round is less than 10^{-11}, the iteration will be stopped. The objective function convergence graph of MCBMF for different data sets is shown in the figure below: Figure 4 shown.

[0170] In addition, the present invention visualizes the similarity between different samples. The unit with a similarity value of 0 is represented by white, and the other values ​​close to 0 are represented by light blue, and close to 1 are represented by dark blue. Figure 5 As shown in the figure, it can be seen that MCBMF has recognized regular matrix graphics on the 100leaves, Hdigit and NGs datasets. This shows that after fusing the unified graph, the MCBMF algorithm can accurately capture the similarity information between samples, which reflects the effectiveness of the MCBMF algorithm. It is worth mentioning that Figure 5 This is the result of unified graph visualization learned under the condition that MCBMF uses each sample node to ensure that the degree is 15 and initialize the sub-similar graph. The most appropriate degree setting will be discussed next.

[0171] Assume that the k value represents the degree, that is, the number of non-zero cells corresponding to a single sample. Non-zero cells represent that the connection between the two samples has been preserved. For example, if the cell value in the xth row and yth column in S(v) is non-zero, it means that there is an edge between sample x and sample y. A cell with a value of zero means there is no edge. A similarity graph between samples is constructed based on the position of non-zero values. The designed experiment tests the model accuracy when the k value changes, in order to find the most suitable k value as the model parameter of MCBMF under the corresponding data set, so as to achieve the best model effect. The experimental test of the k value is as follows Figure 6 As shown:

[0172] Therefore, in the 100leaves dataset, the k value used by MCBMF is 18; in the 3sources dataset, the k value of 45 can meet the best performance; similarly, the k value is initialized to 21 on the Hdigit and NGs datasets. Finally, MCBMF achieves excellent performance in terms of the ACC evaluation index compared with other MvC algorithms, as shown in Table 2.

[0173] In summary, MCBMF has good performance and excellent robustness, and supports more extensive data, which fully demonstrates the effectiveness of the algorithm.

[0174] The above description is only used to help understand the method of the present invention and its core essence, but the protection scope of the present invention is not limited thereto. For those skilled in the art in the art, equivalent replacement or change according to the technical solution and inventive concept of the present invention within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A multi-graph collaborative fusion method suitable for multi-view data clustering, characterized in that: At least the following steps are included: S1, generate sub-similarity graph; construct sub-graph similarity matrix using adaptive learning; S2, merging sub-similar graphs into a unified similarity graph; retaining sample similarity information, and using adaptive learning to fuse multiple sub-similar graphs to form a unified similarity graph; S3. Obtain the clustering information of the samples; based on the unified similarity graph, extract the low-dimensional embedding representation of the samples and complete the final clustering.

2. The multi-graph collaborative fusion method for multi-view data clustering according to claim 1, characterized in that: The S1 is specifically as follows: First, the objective function is established to maintain the structure of the graph. Second, the geometric similarity is maintained, and normalization is used to change the direction characteristics without changing the length of the subsequent similarity vector. After that, the sparsity of the subgraph similarity matrix is ​​guaranteed, and the l_{F} norm is added for sparsity constraint. The obtained subgraph similarity matrix S is as follows: in, S (v) ∈R N×N , N represents the dimension of the sample, and V is the number of subgraphs; and Respectively represent the i-th and j-th columns in the v-th sub-graph data; Represents the data in the i-th column of the similarity graph, Represents the similarity between the i-th and j-th samples in the v-th subgraph; 1 N represents a full-1 column vector of length N, and β is a hyperparameter used to characterize the weights of different learning modules.

3. The multi-graph collaborative fusion method suitable for multi-view data clustering according to claim 2, characterized in that: The S2 is specifically as follows: Based on the linear weighted superposition method, a unified global similarity graph matrix U is generated. In this process, the sample similarity information is retained by adding an adaptive weight linear construction regularization term. The details are as follows: Among them, the matrix U∈R N×N is the learned unified similarity graph, U i Represented as the i-th column data of the unified similarity graph; matrix W∈R V×N is the adaptive weight matrix, Represents the i-th column of the v-th subgraph weight matrix.

4. The multi-graph collaborative fusion method for multi-view data clustering according to claim 3, characterized in that: The S3 is as follows: The RatioCut graph cutting strategy in spectral clustering is adopted. On the basis of the unified similarity graph, the graph is cut according to the edge weight information, and the Laplace matrix is ​​used to ensure its symmetry; the final matrix is ​​as follows: Among them, L U represents the Laplace matrix; λ represents the hyperparameter; D U The matrix is ​​initialized as 5. The multi-graph collaborative fusion method for multi-view data clustering according to claim 4, characterized in that: The final matrix uses an alternating iterative method to solve the variables therein one by one; specifically: other matrices are fixed as constants by default, and one of the matrices is regarded as a variable to solve the update rule.

6. The multi-graph collaborative fusion method suitable for multi-view data clustering according to claim 5, characterized in that: The specific solution process is as follows: Fix W, U and H and update S^{(v)} as follows: Fix S^{(v)}, U and H, and update W as follows: Fix S^{(v)}, W and H, and update U as follows: Fix S^{(v)}, W and U, and update H as follows:

7. A multi-graph collaborative fusion system for multi-view data clustering applied in any of the methods of claims 1-6, characterized in that: include: The local structure preservation module of sub-view data adopts adaptive learning of sub-graph similarity matrix to generate sub-similarity graph; The multi-graph fusion module based on adaptive learning adopts linear weighted superposition and adds adaptive weight linear construction regularization terms to fuse multiple sub-similar graphs into a unified similar graph; The subspace embedding module based on spectral clustering adopts the RatioCut graph cutting strategy in spectral clustering to learn the final representation of the sample in the subspace and obtain the low-dimensional embedding representation of the sample.