A Granule-Based Multi-View Contrastive Clustering Method
Through the multi-view comparison clustering method based on the sphere, the sphere set is constructed and the sphere correlation matrix is established, which solves the shortcomings of the existing multi-view comparison learning method in balancing the consistency and complementarity of the view, and realizes effective multi-view data clustering and local structure retention.
Patent Information
- Application Number
- CN202411687496.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-11-25
AI Technical Summary
Existing multi-view comparison learning methods have shortcomings in balancing the consistency and complementarity of views, making them difficult to effectively apply on large-scale datasets, and the relationships across view clusters cannot be clearly measured.
Using the multi-view comparison clustering method based on particle spheres, we use the particle sphere set, generate the particle sphere overlap relationship matrix and particle sphere association matrix, and establish the particle sphere association within and across views to achieve multi-grain size comparison learning.
This method can effectively segment the sample set into coarse-grained particle spheres, preserve the local topology of the sample set, realize the effective clustering of multi-view data, and improve the accuracy and interpretability of clustering results.
Smart Images

Figure CN119474926B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of clustering methods, and particularly relates to a multi-view contrast clustering method based on granulospheres. Background Art
[0002] Multi-view data is collected from different sensors or feature extractors and usually exhibits heterogeneity. For example, a web page contains pictures, texts, videos, etc., which are regarded as different views and reflect part of the information of the web page from their respective perspectives. Multi-view clustering has received continuous attention in recent years, aiming to cluster multi-view data into multiple clusters in an unsupervised paradigm. The key challenge lies in how to balance the consistency and complementarity of multiple views so as to learn the most comprehensive consensus representation. Traditional multi-view clustering methods are mainly divided into three categories: subspace learning, graph learning, and multi-kernel learning. These methods involve matrix decomposition and fusion and have high computational complexity, making it difficult to apply them to large-scale data sets, which limits their practical applications.
[0003] In recent years, multi-view clustering methods based on deep learning have received extensive attention due to their excellent feature representation capabilities. Such methods expand the deep single-view clustering methods and select specific feature extractors according to the view characteristics for different views. For example, the method proposed by Deep canonical correlation analysis uses a deep neural network to project the data of two views into a common space, where the feature representations of the two views are highly correlated, making it a non-linear extension of canonical correlation analysis. The method proposed by Deep subspace clustering with sparsity prior is a deep subspace clustering method with a sparse prior. It projects the input data into a latent space, maintains the local structure by minimizing the reconstruction loss, and introduces sparse prior information into the latent representation learning to preserve the sparse reconstruction relationship of the entire data set. Duet Robust Deep Subspace Clustering models the collapse mode (such as in-view noise) from the perspectives of data reconstruction and latent self-expression using two regularization norms, and also proposes a smoothing method for non-differentiable norms to optimize the loss function using gradient-based methods. DIMC-net: Deep Incomplete Multi-view Clustering Network extracts high-level features of multiple views through view-specific autoencoders and introduces a fusion graph-based constraint to preserve the local geometric structure of the data. Deep multi-view spectral clustering via ensemble applies ensemble clustering to fuse similarity graphs from different views and uses graph autoencoders to learn a common spectral embedding. It designs a unified optimization framework to minimize the graph reconstruction loss, orthogonality loss, and graph contrast loss simultaneously. These methods combine deep learning with traditional multi-view learning ideas by introducing the constraints of traditional methods (such as neighborhood graph constraints and self-expression constraints) into the latent space of deep module projection. This allows the model to learn concise and comprehensive representations, maximizing the preservation of the structural information of the input data.
[0004] Existing multi-view contrast learning methods generally operate at two scales: instance level and cluster level. Instance-level methods construct positive and negative pairs based on sample correspondence, aiming to bring positive pairs closer and push negative pairs further apart in the latent space. Cluster-level methods focus on calculating the cluster assignments of samples under each view and maximize view consensus by reducing distribution differences, such as minimizing the KL divergence or maximizing the mutual information. However, these two types of methods either introduce false negative pairs, resulting in reduced model discriminability, or ignore local structures and cannot explicitly measure the relationships between cross-view clusters. Summary of the Invention
[0005] The present invention provides a multi-view contrast clustering method based on grain balls for the problems existing in the prior art.
[0006] The technical solution adopted by the present invention is: a multi-view contrast clustering method based on grain balls, including the following steps:
[0007] Step 1: Obtain multi-view data to form a data set;
[0008] Step 2: Construct a multi-view contrast clustering model, which includes a data processing module, a grain ball generation module, and an output module;
[0009] The data processing module is used to reduce the dimension of the multi-view data in the data set to obtain multi-view data features;
[0010] The grain ball generation module is used to construct a grain ball set according to each view corresponding to the multi-view data features, and construct an intra-view grain ball overlap relationship matrix and an inter-view grain ball association matrix;
[0011] Obtain a mask matrix according to the intra-view grain ball overlap relationship matrix and the inter-view grain ball association matrix, and obtain a fused feature through contrastive learning;
[0012] The output module is used to output a clustering result according to the fused feature;
[0013] Step 3: Train the multi-view contrast clustering model to obtain a trained multi-view contrast clustering model;
[0014] Step 4: Obtain the required clustering result according to the trained multi-view contrast clustering model.
[0015] Further, the process of constructing the grain ball set in step 2 is as follows:
[0016] Under each view, through the k-means clustering method, k clusters, that is, k grain balls, are obtained, and the k grain balls form a grain ball set; repeating the above process for each view can obtain the entire grain ball set.
[0017] Further, the center c i and radius r i The calculation method is as follows:
[0018]
[0019] In the formula: i is the serial number of the grain ball, n i is the number of samples in the grain ball, j is the serial number of the samples included in the grain ball, and x j is the j-th sample of the current grain ball.
[0020] Further, the element v in the overlap relationship matrix A Satisfy the following relationship:
[0021]
[0022] Where: v is the view number;
[0023] The calculation process of the distance between the granular balls is as follows:
[0024]
[0025] Where: is the distance between the i-th granular ball and the j-th granular ball in the v-th view, is the center of the i-th granular ball in the v-th view, is the center of the j-th granular ball in the v-th view.
[0026] Furthermore, the granular ball association matrix P (m,n) in the elements Satisfy:
[0027]
[0028] Where: m and n are both view numbers, t i is a certain granular ball in view m, t j is a certain granular ball in view n, t both is the number of samples in the intersection of the two granular balls, and τ is the threshold parameter;
[0029] Among them:
[0030] t both = length(Id both )
[0031] Where: Id both is the common sample set of the i-th granular ball in view m and the j-th granular ball in view n.
[0032] Furthermore, the mask matrix M is:
[0033]
[0034] Where: P (n,m) is the transpose of P (m,n) .
[0035] Furthermore, the loss function in the training process includes a contrast loss and an encoding loss;
[0036] L = L con + λL rec
[0037] Where: L is the loss function, L con is the contrast loss, Lrec is the encoding loss, and λ is a parameter.
[0038] Furthermore, the contrast loss function is as follows:
[0039]
[0040] In the formula: V is the number of views, and L (m,n) is the loss between views m and n.
[0041] Furthermore, the encoding loss is:
[0042]
[0043] In the formula: E v (·; θ v ) is the encoder, D v (·; φ v ) is the decoder, is, θ v is a parameter, φ v is a parameter, N is the number of samples within a view, V is the number of views, v is the view serial number, and i is the sample serial number.
[0044] A system for a multi-view contrast clustering method based on granular balls, comprising a data acquisition and processing module, a granular ball generation module, and an output module;
[0045] The data acquisition and processing module is used to acquire a multi-view data set, perform dimensionality reduction on the multi-view data in the data set, and obtain multi-view data features;
[0046] The granular ball generation module is used to construct a granular ball set according to each view corresponding to the multi-view data features, construct an intra-view granular ball overlap relationship matrix and an inter-view granular ball association matrix;
[0047] Obtain a mask matrix according to the intra-view granular ball overlap relationship matrix and the inter-view granular ball association matrix, and obtain a fused feature through contrast learning;
[0048] The output module is used to output a clustering result according to the fused feature.
[0049] The beneficial effects of the present invention are:
[0050] (1) The method of the present invention divides the sample set into coarse-grained granular balls and establishes associations between intra-view and cross-view granular balls; these associations are strengthened in the shared latent space, thereby realizing multi-granularity contrast learning;
[0051] (2) The granular balls in the present invention are located between instances and clusters, retaining the local topological structure of the sample set. Description of the Drawings
[0052] Figure 1 It is a schematic structural diagram of the multi-view contrast clustering model in the present invention.
[0053] Figure 2 It is the clustering result visualized by t-SNE on the MNIST-USPS dataset using the method of the present invention in the embodiment of the present invention.
[0054] Figure 3 It is the parameter analysis experiment of different p and d on Caltech101-20 ( Figure 3 a) and Cora ( Figure 3 b) in the embodiment of the present invention, where the vertical axis represents the clustering accuracy. Detailed implementation manners
[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0056] A multi-view contrast clustering method based on grain balls includes the following steps:
[0057] Step 1: Obtain multi-view data to form a dataset;
[0058] Given a multi-view dataset with N samples Each sample has V instances from different views. Let d v be the feature dimension of the v-th view, which is usually different for different views.
[0059] Step 2: Construct a multi-view contrast clustering model. The multi-view contrast clustering model includes a data processing module, a grain ball generation module, and an output module; the structure is as Figure 1 shown.
[0060] The data processing module is used to reduce the dimension of the multi-view data in the dataset to obtain multi-view data features;
[0061] For subsequent comparison and fusion, it is necessary to unify the cross-view feature dimensions, that is, project and transform different views to the same dimension through an encoder. A deep autoencoder is used as a representation learning framework to effectively extract basic low-dimensional embeddings from the original features.
[0062] For example, for the v-th view, E v (·; θ v ) is the encoder, D v (·; φ v ) is the decoder, θ v is the parameter, φ v is the parameter; set the output dimension of to d. Obtain the feature matrix and store it in the latent space. Subsequently, use this feature to construct grain balls.
[0063] The feature corresponding to the i-th instance of the v-th view Is expressed as:
[0064]
[0065] Where: Is X v The i-th instance of.
[0066] The granule ball generation module is used to construct a granule ball set according to each view corresponding to the multi-view data features, and construct the granule ball overlap relationship matrix within each view and the granule ball association matrix between views;
[0067] First, a granularity control parameter p is introduced, which can roughly reflect the generation granularity of the granule balls. Let k be the total number of granule balls generated under a single view. The sample number N and k satisfy:
[0068]
[0069] Then, k-means is run under a single view to divide into k clusters, and each cluster is regarded as a granule ball.
[0070] Then the i-th granule ball is expressed as GB i ; The center c i And radius r i The calculation method is as follows:
[0071]
[0072] Where: i is the serial number of the granule ball, n i Is the number of samples in the granule ball, j is the serial number of the samples included in the granule ball, and x j Is the j-th sample of the current granule ball.
[0073] For different views, the same granularity parameter p is set, and the corresponding k is also the same, that is, the same number of granule balls are divided in each view.
[0074] For the v-th view, H v Represents the feature after being projected by the encoder E v Using the above granule ball method, the granule ball set of the v-th view is obtained Let Represents the combination of the granule ball sets of all views, C v Represents the center matrix of the v-th view. R v Represents the radius matrix of the v-th view. For the i-th granule ball under the v-th view Its center is The radius is The calculation of the center and radius is gradient-preserving. Let D vThe granule distance matrix representing the v-th view. The distance between the i-th granule and the j-th granule is:
[0075]
[0076] In the formula: is the distance between the i-th granule and the j-th granule in the v-th view, is the center of the i-th granule in the v-th view, is the center of the j-th granule in the v-th view.
[0077] Based on D v the granule overlap relationship matrix A within the view can be obtained v : Each of its elements satisfies the following relationship:
[0078]
[0079] In the formula: v is the view serial number;
[0080] where A v is used as part of the contrast learning positive and negative sample pair mask, and the relationship between granules within the view is established through the matrix set
[0081] According to the granule overlap relationship matrix within the view and the granule association matrix between views, a mask matrix is obtained, and fused features are obtained through contrast learning;
[0082] Next, the association between cross-view instances is constructed:
[0083] Two balls in different views are regarded as neighbors in the latent space if each of them contains instances of the same sample from their respective views. However, this method lacks robustness. When p is relatively large, the granules also become larger, and due to randomness, two cross-view balls may contain very few common samples.
[0084] The present invention solves this problem by adopting the following method:
[0085] First, identify the common sample set between two balls according to the stored sample indexes:
[0086]
[0087] In the formula: and are two granules of views m and n, respectively containing t i and t j samples, and Id(·) is the sample index contained in the granule.
[0088] Calculate the number of samples in Id both :
[0089] t both = length(Id both )
[0090] The inter-view granulocyte correlation matrix P (m,n) The elements in Satisfy:
[0091]
[0092] Where: m and n are both view numbers, t i Is a certain granulocyte in view m, t j Is a certain granulocyte in view n, t both Is the number of samples in the intersection of the two granulocytes, and τ is the threshold parameter, representing the minimum proportion of the common samples required for two cross-view granulocytes.
[0093] The matrix set Reflects whether any two granulocytes within a view overlap, and the matrix set Indicates whether any two balls between views have sufficient intersections. The goal is that these related granulocyte pairs are as close as possible in the latent space, while the unrelated granulocyte pairs should be far apart. The center of the granulocyte is used to represent the entire granulocyte during the calculation process. For the sake of easy calculation, for any two views m and n, the combined center matrix is defined as:
[0094]
[0095] Then {A m , A n} and {P (m,n) , P (m,n)} are concatenated into a unified mask matrix:
[0096]
[0097] Where: P (n,m) Is the transpose of P (m,n) .
[0098] The unified mask matrix M ensures that the granulocytes within the same view and across different views are properly considered during the contrast learning process.
[0099] Step 3: Train the multi-view contrast clustering model to obtain the trained multi-view contrast clustering model;
[0100]
[0101] Where: Ω i Represents the set of granulocyte serial numbers associated with the i-th granulocyte, φ iDenote the set of granule numbers associated with the \(i\)-th granule, \(M\). ij Denote whether there is an association between the \(i\)-th granule and the \(j\)-th granule, \(M\). iz Denote whether there is an association between the \(i\)-th granule and the \(z\)-th granule; \(\cos(\cdot,\cdot)\) is the cosine similarity between two vectors.
[0102] Perform the same calculation process between any two views, and take the average of all losses as the final contrast loss. The contrast loss function is as follows:
[0103]
[0104] In the formula: \(V\) is the number of views, \(L\). (m,n) is the loss between views \(m\) and \(n\).
[0105] The encoding loss is:
[0106]
[0107] In the formula: \(E\). v (·; \(\theta\). v ) is the encoder, \(D\). v (·; \(\varphi\). v ) is the decoder, is, \(\theta\). v is the parameter, \(\varphi\). v is the parameter, \(N\) is the number of samples in the view, \(V\) is the number of views, \(v\) is the view number, and \(i\) is the sample number.
[0108] Combine the above two loss functions with the regularization parameter \(\lambda\) to obtain the total loss function, and this objective function can be represented by any gradient-based optimization method such as the Adam algorithm.
[0109] The loss function includes the contrast loss and the encoding loss;
[0110] \(L = L\). con + \(\lambda L\). rec
[0111] In the formula: \(L\) is the loss function, \(L\). con is the contrast loss, \(L\). rec is the encoding loss, and \(\lambda\) is the parameter.
[0112] By optimizing the objective function, i.e., the loss function, the multi-view features gradually have clear clustering boundaries in the latent space, showing good discriminability.
[0113] The output module is used to output the clustering result according to the fused features;
[0114] Feature matrix Obtain the fused matrix through contrastive learning The k-means algorithm is directly executed on the fusion matrix to obtain the clustering labels of each sample, that is, the clustering result.
[0115] Step 4: Obtain the required clustering result according to the trained multi-view contrast clustering model.
[0116] Embodiment
[0117] To illustrate the effect of the method of the present invention, seven multi-view benchmark datasets are used.
[0118] BBCSport includes 544 sports news articles in 5 subject areas, with 3183-dimensional MTX features and 3203-dimensional TERMS features, forming 2 views. Caltech101-20 contains a total of 101 classes, and 20 widely used classes and 2386 samples are selected.
[0119] Cora contains 4 views of content, inbound, outbound, and citations extracted from documents. Scene15 consists of 4568 natural scenes, divided into 15 groups. Each scene is described by three types of features: GIST, SIFT, and LBP. MNIST-USPS is a popular handwritten digit dataset containing 5000 samples with two different types of digit images. ALOI-100 consists of 10800 object images, and each image is described by 4 different features. NoisyMNIST uses the original images as views Figure 1 , and randomly selects intra-class images of Gaussian white noise as views Figure 2 . The specific datasets are shown in Table 1.
[0120] Table 1. Multi-view datasets
[0121]
[0122] The method of the present invention is compared with existing methods, and the existing methods include
[0123] Completer (Lin et al. 2021), MFLVC (Xu et al. 2022), DealMVC (Yang et al. 2023b), DMCE (Zhao, Yang, and Nie 2023), CSPAN (Jin et al. 2023b), ADPAC (Xu et al. 2023b), SURE (Yang et al. 2023a).
[0124] Commonly used metrics are used to evaluate the results: clustering accuracy ACC, normalized mutual information NMI, and purity PUR.
[0125] For each view, the encoder consists of several linear layers with ReLU activation functions between each pair of layers. Except for the BBCSport and Cora datasets, all other datasets use the same encoder structure with dimensions set to {d v , 2000, 500, 500, d}, where d v is the input feature dimension for each view. d is the projected feature dimension, which is the same for all views. After encoding, the input is normalized. The decoder reflects the encoder structure. For the Cora dataset, the same dimensions are used but without the activation function between layers, resulting in a linear projection. For BBCSport, given its sample size of 544, we use a single-layer linear projection with encoder dimensions set to {d v , d}. In the clustering stage, the projected features of each view are fused with equal weights and then the k-means algorithm is applied to obtain the clustering labels.
[0126] The implementation of the method of the present invention is performed using PyTorch 2.3 (Paszke et al. 2019) on a Windows 10 operating system supported by an NVIDIA GeForce GTX 1660 Ti GPU. The Adam optimizer with a learning rate of 0.0001 and a weight decay of 0 is adopted. The batch size is usually set to 256 or 1024, depending on the dataset size. Except for BCSport and Cora, the regularization parameter λ is usually set to 1 in most datasets and is adjusted to 0 due to differences in projection methods (e.g., linear or non-linear). The threshold parameter τ for all datasets is uniformly set to 0.1. The granularity parameter p significantly affects the experimental results and will be analyzed later. It should be emphasized that the granularity parameter reflects the average granularity, and the set of granular balls may contain balls smaller than or equal to this granularity.
[0127] The results are shown in Table 2, where the best results are in bold and the second-best results are underlined.
[0128] Table 2. Clustering results of different methods on 7 datasets
[0129]
[0130] As can be seen from Table 2, among the three given metrics, the method of the present invention achieves the best results or the second-best results in most cases. Even on the NoisyMNIST dataset, the proposed method ranks approximately third, with a very small gap from the top two methods. Taking the Cora dataset as an example, the method of the present invention reaches an accuracy of 65.44%, greatly exceeding the best comparative result of 49.07%. This proves the effectiveness and competitiveness of the proposed method.
[0131] Compared with classical multi-view contrastive learning methods (such as Completer, DealMVC, SURE), this method consistently achieves more favorable clustering results on most datasets. As a representative, SURE is an instance-level contrastive learning method that focuses on the false negative pair problem and has achieved the best or second-best results on the Scene-15, MNIST-USPS, ALOI-100, and NoisyMNIST datasets. It is worth noting that the method of the present invention does not show a significant performance degradation on these four datasets and is significantly better than other methods. In addition, on the remaining three datasets, the proposed method is significantly better than SURE, highlighting the effectiveness of contrastive learning at the granulocyte level.
[0132] The clustering results visualized by t-SNE on the MNIST-USPS dataset are as follows Figure 2 As shown, it can be seen that as the number of optimization rounds increases, the clustering structure becomes clearer. This indicates that the method of the present invention can effectively reveal the underlying cluster structure.
[0133] To further verify the effectiveness of the method of the present invention at the granulocyte level, ablation experiments were conducted on the Caltech101-20 dataset. The feature dimension d was set to 128. Based on this, three experimental settings were adopted. The first setting trained the model only according to the reconstruction loss. The second setting included the reconstruction loss and the instance-level contrastive loss (i.e., p = 1). The third setting contained the reconstruction loss and the granulocyte contrastive loss, and the granularity parameter p was set to 2 (the method of the present invention).
[0134] The experimental results are shown in Table 3, from which it can be seen that the instance-level contrastive method performs poorly on this dataset, while the granulocyte contrastive method has achieved significant improvement, indicating the feasibility and effectiveness of this method.
[0135] Table 3. Ablation experiment results on the Caltech101-20 dataset
[0136]
[0137] There are two important hyperparameters in the method of the present invention, the granularity parameter p and the projected feature dimension d; the former essentially reflects the average size of granulocytes, and when p is set to 1, it is equivalent to instance-level contrastive learning. The latter affects the amount of original feature information contained in the latent representation. If d is too small, important information may be lost, and if it is too large, it will increase the complexity of optimization and memory requirements.
[0138] Experiments were conducted on Caltech101-20 and Cora. The parameter p was varied in the range [1, 2, 3, 4, 8, 16], and the parameter d was varied in the range [8, 16, 32, 64, 128, 256]. The parameters for all datasets were selected from these ranges. The results are as Figure 3 shown. It can be seen from the figure that when p is set to 2, the method of the present invention performs well. As p increases, the performance gradually decreases because larger granular balls can no longer be effectively associated based only on the overlap and intersection sizes. When p becomes too large, the method of the present invention essentially degrades to a contrast method at the cluster level, where reducing the cluster assignment differences may be better.
[0139] Therefore, the parameter p is usually set to 1, 2, or 4. The parameter d has not much influence on the results and is usually set to 64.
[0140] The convergence of the proposed method on the Cora dataset was evaluated by tracking the loss value and the corresponding clustering performance as the number of rounds increased. The total loss value gradually decreased and converged within 100 rounds. These results demonstrate the strong convergence performance of the proposed method.
[0141] The present invention conducts contrastive learning at the granular ball level, avoiding directly using adjacent samples to construct negative pairs while retaining the local structure information of the sample set, and overcoming the disadvantages of instance-level and cluster-level methods. The granular ball construction method of the present invention is different from the classical method of continuously dividing the dataset until the minimum granularity is reached. It can directly divide the sample set into multiple granular balls according to the granularity parameter, avoiding the disadvantage that non-adjacent samples are divided into the same granular ball in the boundary region.
[0142] The present invention uses granular balls to model the local structure of the sample set, and establishes intra-view and inter-view granular ball connections based on the overlap and intersection sizes respectively. By making the connected granular balls closer in the latent space, the proposed model learns highly discriminative features.
Claims
1. A multi-view contrast clustering method based on granular sphere, characterized in that: The following steps are involved: Step 1: Acquire multi-view data to form a data set; the multi-view includes pictures, texts, and videos; Step 2: construct a multi-view comparison clustering model, which includes a data processing module, a particle sphere generation module and an output module; The data processing module is used to reduce the dimension of the multi-view data in the data set to obtain the multi-view data features; The particle sphere generation module is used to construct a particle sphere set according to each view corresponding to the multi-view data features, and to construct a particle sphere overlap relationship matrix within each view and a particle sphere correlation matrix between views; Overlapping Relationship Matrix A v Elements in The following relations are satisfied: Where: v is the view number; The distance between the particles is calculated as follows: Where: For the v In the view i The ball and j The distance between the spheres, For the v In the view i The center of the sphere, For the v In the view j The center of the sphere; Inter-view particle-sphere correlation matrix Elements in satisfy: Where: m and n They are all view numbers. t i For View m A ball in t j For View n A ball in is the number of samples in the intersection of two spheres, τ is the threshold parameter; in: Where: For View m Middle i Spheres and views n Middle j A public sample set of spheres; The mask matrix is obtained according to the particle-ball overlap relationship matrix within the view and the particle-ball correlation matrix between views, and the fusion feature is obtained through contrast learning; the mask matrix M is: Where: for The transpose of The output module is used to output clustering results based on fusion features; Step 3: Train the multi-view comparative clustering model to obtain a trained multi-view comparative clustering model; Multi-view comparative clustering model: Where: , , Indicates i There is a set of ball numbers associated with each ball. represents the set of ball numbers associated with the i-th ball, M ij Indicates i The ball and j Is there a correlation between the particles? M iz Indicates i The ball and z Are there any connections between the particles? is the cosine similarity between two vectors; Step 4: Obtain the desired clustering results based on the trained multi-view comparative clustering model.
2. The multi-view contrast clustering method based on particle sphere according to claim 1, characterized in that: The process of constructing the particle sphere set in step 2 is as follows: In each view, we use the k-means clustering method to obtain k clusters, i.e. k A ball, k The particles constitute a particle set; repeating the above process for each view can obtain the entire particle set.
3. The multi-view contrast clustering method based on particle sphere according to claim 2, characterized in that: The center of the sphere c i and radius r i The calculation method is as follows: Where: i is the serial number of the ball, n i is the number of samples in the sphere, j is the serial number of the sample contained in the sphere, x j is the number of the current ball j samples.
4. The multi-view contrast clustering method based on particle sphere according to claim 1, characterized in that: The loss function in the training process includes contrast loss and encoding loss; Where: L is the loss function, is the contrast loss function, is the encoding loss function, To balance the hyperparameters.
5. The multi-view contrast clustering method based on particle sphere according to claim 4, characterized in that: The contrast loss function is as follows: Where: V is the number of views, For View m and n The loss between.
6. The multi-view contrast clustering method based on particle sphere according to claim 4, characterized in that: The encoding loss function is: Where: For the encoder, For the decoder, For the i The sample v view instances, As parameters, As parameters, N is the number of samples in the view, V is the number of views, v is the view number, i is the sample serial number.
7. A system using the multi-view comparison clustering method based on particle spheres as claimed in any one of claims 1 to 6, characterized in that: It includes data acquisition and processing module, pellet generation module and output module; The data acquisition and processing module is used to acquire a multi-view data set, perform dimensionality reduction on the multi-view data in the data set, and obtain multi-view data features; The particle sphere generation module is used to construct a particle sphere set according to each view corresponding to the multi-view data features, and to construct a particle sphere overlap relationship matrix within each view and a particle sphere correlation matrix between views; The mask matrix is obtained according to the particle-ball overlap relationship matrix within the view and the particle-ball correlation matrix between views, and the fusion feature is obtained through contrast learning; The output module is used to output clustering results based on fusion features.
Citation Information
Patent Citations
Particle ball generation method, data classification method and classifier
CN116432065A
Gene clustering analysis method based on Gaussian distribution pellets
CN118298930A