A method to improve the accuracy of single-cell multi-view clustering
Through the technology of multi-view data fusion and integration, the problem that traditional single-cell clustering algorithms cannot fully reflect data complexity is solved, and higher clustering accuracy and stability are achieved.
Patent Information
- Application Number
- CN202410875275.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-02
AI Technical Summary
Traditional single-cell clustering algorithms are based on a single perspective only, and cannot fully reflect the complexity of the data, and it is easy to ignore the complementarity between data from multiple perspectives.
Through multi-view data fusion and integration, multi-view weight learning, cross-view consistency learning and multi-graph fusion technologies are adopted to build multi-view data feature space, optimize the map structure, and improve the accuracy of single-cell clustering.
It improves the accuracy and stability of single-cell multi-view clustering, enhances the generalization ability and robustness of the model, and can better mine hidden patterns and structures between data.
Smart Images

Figure CN118969103B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of single-cell RNA sequencing technology, and relates to the fields of data mining, artificial intelligence, algorithm optimization, and single-cell multi-omics, and in particular to a method for improving the accuracy of single-cell multi-view clustering. Background Art
[0002] Single-cell RNA sequencing technology can reveal the differences between different cell types and has made significant contributions to the study of dynamic changes in cells during development, disease and treatment. This technology can provide gene expression data at the single-cell level, helping researchers to gain a deeper understanding of changes and interactions within cells. By performing high-throughput sequencing on individual cells, scientists can reveal the characteristics of different cell subtypes, the expression patterns of cells in different states, and the interaction network between cells. This information is of great significance in revealing the mechanisms of disease occurrence and development, screening drug targets, and personalized treatment, and has brought revolutionary changes to biomedical research.
[0003] Cell feature extraction and clustering algorithm design are the main problems in single-cell clustering. Single-view clustering relies only on the information of a single view and cannot fully reveal the complexity of the data. Different views usually contain complementary information, which may be ignored or omitted in single-view clustering. Therefore, it is urgent to introduce an algorithm to improve the accuracy of single-cell multi-view clustering in single-cell clustering to improve the problems we are facing. In order to avoid the single feature space of single-cell RNA sequencing data being insufficient to fully characterize the function of cells, multi-view learning is used in this system to comprehensively characterize single-cell RNA sequencing data from different angles, capturing different aspects and characteristics of single-cell data, and being able to use this complementary information to provide more comprehensive cell type clustering. The accuracy of cell clustering directly affects the identification and interpretation of cell types, subtypes and their characteristics. Through accurate cell clustering, researchers can more accurately identify different cell clusters and further study their biological characteristics, functions and relationships under physiological and pathological conditions.
[0004] A method to improve the accuracy of single-cell multi-view clustering involves multi-view data fusion and integration, in which the key technologies involved include building multi-views, view weight learning, cross-view learning, multi-image fusion, etc. The details are as follows:
[0005] Constructing multiple views: k-nearest neighbors are usually used to construct multiple similarity graphs. The basic idea of k-nearest neighbors is to use neighboring sample data for prediction or classification. For a cell sample, by calculating the distance between the sample and all samples in the training set, the k training samples closest to the sample are found, and then the label information of these k samples is used for prediction or classification. The specific steps include calculating the distance between samples to determine the similarity connection relationship between them. For each sample, the k-nearest neighbor algorithm will find the k nearest neighbors to it, and these neighbors can be regarded as having a strong connection in the similarity graph.
[0006] Multi-view weight learning: Most of the existing multi-view unsupervised feature selection methods have the following problems: the sample similarity matrix, the weight matrix of different views, and the weight matrix of features are often pre-defined, which cannot effectively characterize the real structure between data and reflect the importance of different views and features, resulting in the failure to select useful features. To solve the above problems, AWMVPL learns the weights of different views through a self-weighting scheme when considering the correlation between views, and integrates the information of the learned view-specific proximity matrix into the view common clustering indicator matrix through an adaptive weighting scheme, which outputs the final clustering result. ALMUFS performs adaptive learning of view weights and feature weights based on multi-view clustering, which can simultaneously achieve feature selection and guarantee clustering performance, and adaptively learns the similarity matrix of samples under the Laplace rank constraint, and constructs a multi-view unsupervised feature selection method based on adaptive learning.
[0007] Cross-view learning: In multi-view clustering, since the data comes from multiple different views or feature spaces, these views may contain complementary information, but may also be noisy or inconsistent. Therefore, ensuring consistency between different views is crucial to improving the performance of clustering. Figure 1 Consistency learning is a common technique in the field of multi-view learning, which aims to enhance learning performance by leveraging the consistency between different views or modalities. Figure 1 Consistent learning can effectively combine the information of multiple views, thereby improving the generalization ability and robustness of the model. Gcfagg obtains consistent data representation of multiple views through cross-sample and cross-view feature aggregation, fully exploring the complementarity of similar samples. UPMGC-S effectively utilizes the structural information of each view to refine the cross-view correspondence. DCMSC realizes multi-view clustering by constraining the learned clustering label matrix and diversity and consistency in the data space. CGL unifies spectral embedding and low-order tensor learning into a unified optimization framework to jointly determine the spectral embedding matrix and tensor representation. GSF solves the problem of multi-view clustering by seamlessly integrating the graph structure of different views to make full use of the geometric properties of the underlying data structure.
[0008] Multi-graph fusion: Integrate graphs from different sources into a joint graph, integrate multiple data views and consider cross-view Figure 1 Consistency can improve the effectiveness and accuracy of graph clustering. The key challenge of multi-graph fusion clustering is to effectively integrate information from multiple graphs. This involves solving problems such as heterogeneity of graph structures or feature spaces, noise, and inconsistency between different views. To address these challenges, various techniques have been developed, including graph alignment and weighted graph fusion. Graph alignment techniques aim to align nodes or edges in different graphs to ensure that similar entities or relationships are represented consistently across views. Weighted graph fusion methods assign weights to each graph based on its relative importance or quality, and then combine them into a fused graph. Figure 1 Consistency learning can enforce consistency between different views by incorporating similarity measures or constraints into clustering objectives. Figure 1 Consistency, multi-graph fusion clustering can improve the effectiveness and accuracy of clustering analysis. It can help identify more complex data structures and patterns, better handle noisy or incomplete data, and provide more robust and reliable clustering results. With the advancement of graph neural networks and other deep learning techniques, it is possible to further improve the performance of multi-graph fusion clustering by leveraging these powerful representation learning methods.
[0009] Single-cell RNA sequencing technology can reveal the differences between different cell types and has made significant contributions to the study of dynamic changes in cells during development, disease and treatment. This technology can provide gene expression data at the single-cell level, helping researchers to gain a deeper understanding of changes and interactions within cells. By performing high-throughput sequencing on individual cells, scientists can reveal the characteristics of different cell subtypes, the expression patterns of cells in different states, and the interaction network between cells. This information is of great significance in revealing the mechanisms of disease occurrence and development, screening drug targets, and personalized treatment, and has brought revolutionary changes to biomedical research. Summary of the invention
[0010] The technical problem to be solved by this patent is that the traditional single-cell clustering algorithm is only based on the clustering results of a single perspective and cannot fully reflect the complexity of the data. At the same time, since there is usually complementarity between multi-perspective data, it is easy for the single-perspective clustering algorithm to ignore this. To this end, the purpose of this patented technology is to establish a data feature space from multiple perspectives and study the existing problems. Specifically, by performing weight learning and structure learning on the multi-perspective similarity graph, the initialized map is optimized to improve the clustering accuracy of single cells.
[0011] The present invention is achieved by the following measures:
[0012] This patent discloses a method for improving the accuracy of single-cell multi-view clustering, including the following steps:
[0013] S01: Input a set of data points , counting cells and cells Similarity information is used to generate multi-view data from single single-cell data, and a kernel function is applied to generate multi-view data {V m}, as shown in the following formula:
[0014]
[0015]
[0016]
[0017] Obtained by , and The multi-view data {V m}; m represents the number of views, T represents the transpose between x and y;
[0018] S02: Apply k-nearest neighbors to construct a multi-view similarity graph {S m} , the similarity connection relationship between samples is determined by calculating the distance between them, as shown in the following formula:
[0019]
[0020] in , k is the number of neighbors;
[0021] S03: Obtain a weighted graph by weighting different views: First, initialize and construct the weight w of each view m (w m =1 / m), through {S m} and w to obtain the initialized weighted graph U and initialize F1, as shown in the following formula:
[0022]
[0023] Where U is {S m The average similarity matrix of},
[0024] Solve the following formula to obtain F1:
[0025] ;
[0026] The F1 is the eigenvector obtained after eigendecomposition of U;
[0027] S04:
[0028] Optimization m}, w, U, as shown in the following formula:
[0029]
[0030]
[0031] ;
[0032] S05: According to {S m}Construct H: The specific H construction is shown in the following formula:
[0033]
[0034] According to H, the consistency structure diagram A is learned. The specific construction of A is shown in the following formula:
[0035]
[0036]
[0037] in and is the weight factor;
[0038] F2 is the eigenvector obtained after eigendecomposition of A;
[0039] S06: Optimize structure graph A and normalize it:
[0040] The multi-view structure graph H is generated using the following formula:
[0041]
[0042] Based on H, a similarity matrix A is learned, where A corresponds to the learned structure graph;
[0043] The objective function is as follows:
[0044]
[0045]
[0046] in and is the weight factor;
[0047] Before optimizing A, you need to first get H through S, and then get A by applying the objective function after getting H, and then perform the following optimization process.
[0048] (1) Fix F2 and update A, so the objective function becomes
[0049]
[0050] The system will iteratively learn Figure A according to the following formula:
[0051]
[0052] Update A j The post-normalization A uses the following formula:
[0053]
[0054] (2) Fix A and update F2. The specific formula is as follows:
[0055]
[0056] The optimal F2 is formed by arranging the eigenvectors of the first c smallest eigenvalues of the Laplacian matrix of A. Finally, the optimized structure diagram is obtained;
[0057] c is the number of clusters;
[0058] S07: Fusion of weighted graph and structure graph, and use of Construct a symmetric normalized matrix T,
[0059] The specific formula is as follows:
[0060]
[0061] S08: Apply spectral clustering: By normalizing the Laplacian matrix L and applying the k-means algorithm to the eigenvector F, the data points are finally divided into c clusters of different cell types.
[0062] The above method for improving the accuracy of single-cell multi-view clustering is preferably: in step S4, the method for optimizing the weighted graph is:
[0063] (1) Fix w, U and F1 and update {S m}:
[0064]
[0065] in
[0066] (2) Fixed {S m}, F1 and U, update w using the following formula:
[0067]
[0068] (3) Fixed w, {S m} and F1, update U using the following formula:
[0069]
[0070] Finally, the optimized weighted graph U is obtained.
[0071] The above method for improving the accuracy of single-cell multi-view clustering is preferably: in step S8, the specific method of spectral clustering is:
[0072] (1) Construct the degree matrix. For matrix T, the degree matrix D is a diagonal matrix whose diagonal elements are the degrees of each vertex. The degree here is defined as the sum of the weights of the edges adjacent to the vertex. Since T is symmetric, the degree is obtained by directly calculating the sum of each row. In order to avoid division by zero, a small constant eps is added. The calculation formula of the degree matrix is as follows:
[0073]
[0074] (2) The construction of the normalized Laplace matrix L, the specific formula is as follows:
[0075]
[0076] (3) Eigendecomposition: Perform eigendecomposition on the normalized Laplace matrix L to obtain the eigenvector matrix V,
[0077]
[0078] Where V is the eigenvector matrix, whose columns are the eigenvectors F of L; Λ is the eigenvalue diagonal matrix, whose diagonal elements are the eigenvalues of L;
[0079] (4) Normalization of eigenvectors. In order to obtain the normalized eigenvector F, the extracted eigenvector F is normalized row by row.
[0080]
[0081] Finally, the k-means algorithm is used to cluster the normalized feature vector F and finally divide it into c clusters.
[0082] The advantage of this patent is that it improves the accuracy of single-cell clustering through a variety of technologies. Multi-view data extraction feature technology can use the complementary information between different views to improve the accuracy and stability of clustering. Figure 1 Consistent learning technology can effectively combine information from multiple views and enhance the generalization and robustness of the model. Figure 1Consistent learning can better mine hidden patterns and structures between data and improve the interpretability of clustering results. In general, this patented technology performs well in processing multi-source data, multi-modal data or heterogeneous data, has a wide range of applications, and achieves effective fusion and consistency learning between data from different views, providing reliable analysis methods and tools for data mining, information retrieval and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Attached Figure 1-6 The clustering results of this patent and the comparative method are shown in the figure.
[0084] In the figure, the y-axis represents the method, namely the comparison method (GMC\GSF\FastMICE\MVCC\MVGL) and the method of this patent (OUR); the x-axis is 8 public single-cell data sets, and the legend on the right uses color to indicate the quality of the clustering results. The larger the value, the darker the color, the better. Therefore, the color of each box represents the high or low clustering result, and the numerical value is the value of the clustering result, so that the results of this patent's method and the comparison method can be compared. The last row is the result of this patent's method on all data sets, which is almost higher than other methods. DETAILED DESCRIPTION
[0085] In order to make the technical solutions and advantages of the present invention more clear, they will be described in detail below with reference to specific embodiments.
[0086] This embodiment discloses a method for improving the accuracy of single-cell multi-view clustering. Specifically, in a single-cell system, a method for comprehensively characterizing single-cell functions using multiple perspectives is implemented. The single-cell multi-view clustering system includes: constructing a multi-perspective data subsystem, a multi-view weighting subsystem, a multi-view structure subsystem, and a spectral clustering subsystem.
[0087] 1. Constructing a multi-view data subsystem First, we use multi-view learning to comprehensively represent single-cell RNA sequencing data from different perspectives. Specifically, we input a set of data points. , counting cells and cells The similarity information of is used to generate multi-view data from single single-cell data. The specific formula is as follows:
[0088]
[0089]
[0090]
[0091] This results in V (1) , V (2) and V (3) The multi-view data {Vm}.
[0092] Finally, we use k nearest neighbors to construct a multi-similarity graph. We determine the similarity connection between samples by calculating the distance between them. For each sample, k nearest neighbors will find the k nearest neighbors, which can be considered to have strong connections in the similarity graph. We use the following formula to initialize the matrix {S v}:
[0093]
[0094] in , k is the number of neighbors;
[0095] 2. The multi-view weighted subsystem obtains a weighted graph by adjusting the weights of multiple views. It is divided into the following two parts: initialization and optimization.
[0096] First, we need to initialize the weight of each view and weight it to get the weighted graph U. Initialize the weight of each view, that is, . U is initialized by concatenating {Sv} and by the following formula.
[0097]
[0098] Where U is {S m The average similarity matrix of}.
[0099] Solve the following formula to obtain F1:
[0100]
[0101] The following is the part about optimizing weighted graph.
[0102] (1) Fix w, U and F1 and update {S m}:
[0103]
[0104] in
[0105] (2) Fixed {S m}, F1 and U, update w using the following formula:
[0106]
[0107] (3) Fixed w, {S m} and F1, update U using the following formula:
[0108]
[0109] Finally, the optimized weighted graph U is obtained.
[0110] 3. The multi-view structure subsystem is to build a consistent graph structure between multiple views, which can provide a deeper understanding of the same underlying structure in multiple views and better explore the potential correlation relationship between views, which is crucial for multi-view clustering. Specifically, the multi-view structure graph H is generated using the following formula.
[0111]
[0112] Next, we will learn a similarity matrix A based on H, where A is the corresponding learned structure graph. The objective function is as follows:
[0113]
[0114]
[0115] in and is the weight factor.
[0116] The following is the optimization structure diagram.
[0117] (1) Fix F2 and update A, so the objective function becomes
[0118]
[0119] The system will iteratively learn Figure A according to the following formula:
[0120]
[0121] Update A j The post-normalization A uses the following formula:
[0122]
[0123] (2) Fix A and update F2. The specific formula is as follows:
[0124]
[0125] The optimal F2 is composed of the eigenvectors of the first c smallest eigenvalues of the Laplacian matrix of A. Finally, we get the optimized structure diagram.
[0126] 4. Implementation of the spectral clustering subsystem First, the weighted graph and the structure graph are combined, and the symmetric normalized matrix T is constructed. The specific formula is as follows:
[0127]
[0128]
[0129] The low-dimensional embedding representation of the graph is obtained through spectral decomposition, and then the k-means algorithm is used to cluster these embeddings to achieve spectral clustering. The details are as follows:
[0130] (5) Construct the degree matrix. For matrix T, the degree matrix D is a diagonal matrix whose diagonal elements are the degrees of each vertex. The degree here is defined as the sum of the weights of the adjacent edges of the vertex. Since T is symmetric, we can directly calculate the sum of each row to get the degree. In order to avoid division by zero, a small constant eps is added. The calculation formula of the degree matrix is as follows:
[0131]
[0132] (6) The construction of the normalized Laplace matrix L, the specific formula is as follows:
[0133]
[0134] (7) Eigendecomposition: Perform eigendecomposition on the normalized Laplace matrix L to obtain the eigenvector matrix V.
[0135]
[0136] where V is the eigenvector matrix whose columns are the eigenvectors F of L. Λ is the eigenvalue diagonal matrix whose diagonal elements are the eigenvalues of L.
[0137] (8) Normalization of eigenvectors: In order to obtain the normalized eigenvector F, the extracted eigenvector F is normalized row by row.
[0138]
[0139] (9) Finally, the k-means algorithm is used to cluster the normalized feature vector F and finally divide it into c clusters.
[0140] The following are examples of applications of single-cell multi-view clustering technology:
[0141] (1) Cancer research: Applying single-cell multi-view clustering technology to single-cell data in tumor tissue can help reveal the heterogeneity of tumor cells, identify specific subpopulations of cells, and explore the mechanisms of tumor development and metastasis.
[0142] (2) Immunology research: By analyzing the composition and function of individual immune cells through multi-view clustering technology, we can gain a deeper understanding of the functional differentiation, phenotypic transition and immune response process of immune cells, and provide a new perspective for the study of disease immune mechanisms.
[0143] (3) Drug screening: Using single-cell multi-view clustering technology to analyze single-cell data after drug treatment can help screen out drugs that have specific effects on specific cell subpopulations, providing important reference for drug development and personalized treatment.
[0144] (4) Neuroscience research: Applying multi-view clustering technology to analyze single-cell neuronal data can help reveal the distribution, functional characteristics, and interconnection patterns of different types of neurons in the brain, and promote our understanding of the complex neural networks of the brain.
[0145] The above examples of the application of single-cell multi-view clustering technology in different fields illustrate the potential application value of this technology in biomedical research, drug development, and disease diagnosis.
[0146] The following are the effect judgment indicators, detection methods, results, data analysis, etc. of the above method.
[0147] Judgment indicators: This paper uses 6 clustering indicators to judge the quality of the algorithm, namely, adjusted Rand index (ARI), normalized mutual information (NMI), F-score, clustering accuracy (ACC), purity, and precision. The larger the value of these clustering indicators, the better the clustering effect.
[0148] (1)
[0149] (2)
[0150] (3)
[0151] (4)
[0152] (5)
[0153] (6)
[0154] Detection method: Compare the clustering result label with the true label, and calculate the above 6 evaluation indicators to obtain the accuracy value of the clustering result.
[0155] Results: Attached Figure 1-6 The clustering results of this patent and its comparison method are shown.
[0156] In the figure, the y-axis represents the method, namely the comparison method (GMC\GSF\FastMICE\MVCC\MVGL) and the method of this patent (OUR); the x-axis is 8 public single-cell data sets, and the legend on the right uses color to indicate the quality of the clustering results. The larger the value, the darker the color, the better. Therefore, the color of each box represents the high or low clustering result, and the numerical value is the value of the clustering result, so that the results of this patent's method and the comparison method can be compared. The last row is the result of this patent's method on all data sets, which is almost higher than other methods. In the figure, 0 means that it is not implemented correctly.
[0157] Data analysis: The clustering results of our method on 8 single-cell datasets are significantly better than other multi-view clustering methods, and slightly lower than the MVCC method on the Pollen dataset. The values of the six evaluation indicators, including ARI, NMI, F-score, ACC, Purity, and Precision, are all around 0.8, and the best performance is achieved on the Usoskin dataset, Goolam dataset, and Yan dataset, indicating that this patent has significantly improved the accuracy of single-cell clustering.
Claims
1. A method for improving the accuracy of single-cell multi-view clustering, characterized in that The following steps are involved: S01: Input a set of data points , counting cells and cells Similarity information is used to generate multi-view data from single single-cell data, and a kernel function is applied to generate multi-view data {V m }, as shown in the following formula: Obtained by , and The multi-view data {V m }; m represents the number of views, T represents the transpose between x and y; S02: Apply k-nearest neighbors to construct a multi-view similarity graph {S m } , the similarity connection relationship between samples is determined by calculating the distance between them, as shown in the following formula: in , k is the number of neighbors; S03: Obtain a weighted graph by weighting different views: First, initialize and construct the weight w of each view m , w m =1 / m, through {S m } and w to obtain the initialized weighted graph U and initialize F1, as shown in the following formula: Where U is {S m The average similarity matrix of}, Solve the following formula to obtain F1: ; The F1 is the eigenvector obtained after eigendecomposition of U; S04: Optimization m }, w, U, as shown in the following formula: ; S05: According to {S m }Construct H: The specific H construction is shown in the following formula: According to H, the consistency structure diagram A is learned. The specific construction of A is shown in the following formula: in and is the weight factor; F2 is the eigenvector obtained after eigendecomposition of A; S06: Optimize structure graph A and normalize it: The multi-view structure graph H is generated using the following formula: Based on H, a similarity matrix A is learned, where A corresponds to the learned structure graph; The objective function is as follows: in and is the weight factor; Before optimizing A, you need to first get H through S, and then get A by applying the objective function after getting H, and then perform the following optimization process: (1) Fix F2 and update A, so the objective function becomes The system will iteratively learn Figure A according to the following formula: Update A j The post-normalization A uses the following formula: (2) Fix A and update F2. The specific formula is as follows: F2 is formed by arranging the eigenvectors of the first c smallest eigenvalues of the Laplacian matrix of A. Finally, the optimized structure diagram is obtained; c is the number of clusters; S07: Fusion of weighted graph and structure graph, and use Construct a symmetric normalized matrix T, The specific formula is as follows: S08: Apply spectral clustering: By normalizing the Laplacian matrix L and applying the k-means algorithm to the eigenvector F, the data points are finally divided into c clusters of different cell types.
2. The method for improving the accuracy of single-cell multi-view clustering according to claim 1, characterized in that: In step S04, the method for optimizing the weighted graph is: (1) Fix w, U and F1 and update {S m }: in (2) Fixed {S m }, F1 and U, update w using the following formula: (3) Fixed w, {S m } and F1, update U using the following formula: Finally, the optimized weighted graph U is obtained.
3. The method for improving the accuracy of single-cell multi-view clustering according to claim 1, characterized in that: In step S08, the specific method of spectral clustering is: (1) Construct the degree matrix. For the matrix T, the degree matrix D is a diagonal matrix whose diagonal elements are the degrees of each vertex. The degree is the sum of the weights of the edges adjacent to the vertex. Since T is symmetric, the degree is obtained by directly calculating the sum of each row. In order to avoid division by zero, a small constant eps is added. The calculation formula of the degree matrix is as follows: (2) The construction of the normalized Laplace matrix L, the specific formula is as follows: (3) Eigendecomposition: Perform eigendecomposition on the normalized Laplace matrix L to obtain the eigenvector matrix V, Where V is the eigenvector matrix, and its columns are the eigenvectors F of L; Λ is an eigenvalue diagonal matrix, and its diagonal elements are the eigenvalues of L; (4) Normalization of eigenvectors. In order to obtain the normalized eigenvector F, the extracted eigenvector F is normalized row by row. Finally, the k-means algorithm is used to cluster the normalized feature vector F and finally divide it into c clusters.
Citation Information
Patent Citations
Multi-view clustering method based on non-negative matrix factorization and diversity-consistency
CN108776812A
Single-cell RNA sequencing data clustering method based on low-rank characterization and improved spectral clustering
CN115223659A