A high-precision single-cell clustering method and system based on marker genes and ensemble learning

By using cell marker gene set and ensemble learning methods, combined with SNN-Cliq and SOM algorithms, the problem of insufficient feature processing in single-cell clustering is solved, and high-precision single-cell clustering is achieved.

CN115512772BActive Publication Date: 2025-06-27SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211159840.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-06-27
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

Existing single-cell clustering tools rely on simple unsupervised methods in feature processing, ignoring the research results of gene expression, resulting in low clustering accuracy and susceptibility to noise.

Method used

The cell marker gene set is used as prior knowledge for feature extraction, and combined with the single-cell clustering method SNN-Cliq and the self-organized mapping method SOM of deep learning, an integrated clustering model is constructed for feature processing and clustering.

Benefits of technology

Through the guidance of marker gene sets, the noise impact is reduced, the characteristics that characterize cell characteristics are extracted, the accuracy and robustness of clustering are improved, and the feature extraction and clustering performance is better than the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512772B_ABST
    Figure CN115512772B_ABST
Patent Text Reader

Abstract

The present invention relates to a high-precision single-cell clustering method and system based on marker genes and ensemble learning, including: Step 1: Feature extraction; using a feature extraction algorithm and a dimensionality reduction algorithm to reduce the dimensionality of the cell expression matrix and extract cell features; each element in the cell expression matrix corresponds to the expression of a gene / transcript in a given cell; Step 2: Inner-layer clustering; the expression matrix after feature extraction is used as input and applied to the inner-layer clustering method; the inner-layer clustering method includes the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning; Step 3: Calculate the consensus matrix; use the clustering-based similarity partitioning algorithm CSPA to calculate the consensus matrix C; Step 4: Consensus clustering; construct a graph c according to the consensus matrix C, where the nodes Node in the graph c represent cells, and the weight edge of the edge represents the probability that two nodes are in the same partition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a high-precision single-cell clustering method and system based on marker genes and ensemble learning, belonging to the technical field of data clustering. Background Art

[0002] The clustering of single cells is the most important part of single-cell RNA sequencing data analysis. Single-cell RNA sequencing data has problems such as noise and high sparsity, which pose great challenges to high-precision clustering algorithms. For single-cell clustering, the quality of feature selection has a significant impact on clustering accuracy. Current single-cell clustering tools mainly rely on some simple unsupervised feature selection methods for feature processing, while ignoring the guiding role of existing research results in feature extraction. For example, the feature processing part often uses simple measurement methods related to the statistical moments of gene expression, combined with classical data dimensionality reduction operations such as Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), or Uniform Manifold Approximation and Projection (UMAP), and finally uses clustering methods such as spectral clustering, hierarchical clustering, and K-means for clustering. Such a feature processing method is prone to losing features that characterize cell types. Therefore, it is necessary to construct a high-precision cell clustering algorithm to achieve accurate feature extraction and accurate cell grouping.

[0003] Cell marker genes, as genes specifically expressed in different cell populations, their expression patterns can effectively guide the cell grouping process. One possible reason why many single-cell clustering algorithms have low clustering accuracy and are easily affected by noise is that unsupervised feature extraction is not easy to identify the gene set with the largest differences between cell populations. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention uses a cell marker gene set as a prior knowledge set for feature extraction, and integrates two excellent clustering methods, including the single-cell clustering method SNN-Cliq and the self-organizing map (SOM) of deep learning, for single-cell clustering. The present invention proposes a high-precision single-cell clustering algorithm - SCMcluster (Single cell cluster using marker genes), which uses two integrated single-cell marker databases to apply it to feature extraction and constructs an integrated clustering model for further clustering after feature processing.

[0005] The technical solution of the present invention is as follows:

[0006] A high-precision single-cell clustering method based on marker genes and ensemble learning, comprising the following steps:

[0007] Step 1: Feature extraction; using a feature extraction algorithm and a dimensionality reduction algorithm to reduce the dimensionality of the cell expression matrix and extract cell features; each element in the cell expression matrix corresponds to the expression of a gene / transcript in a given cell, where columns correspond to cells and rows correspond to genes / transcripts;

[0008] Step 2: Inner clustering; the expression matrix after feature extraction is used as input and applied to the inner clustering method; the inner clustering method includes the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning; the expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning;

[0009] The expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq, including:

[0010] First, use the Euclidean distance to calculate the similarity matrix D corresponding to the expression matrix M;

[0011] Then, for the similarity matrix D, regard it as a weighted graph to construct a KNN graph;

[0012] Again, construct a shared neighbor graph according to the KNN graph;

[0013] Finally, by the strategy of finding quasi-cliques in the constructed shared neighbor graph, iterate continuously until the final subgraph is obtained; the parameters r_cutoff and merge_cutoff represent the nearest neighbor radius and merge threshold of each pair of Cliq respectively;

[0014] The expression matrix after feature extraction is used as input and applied to the self-organizing mapping method SOM of deep learning. The topological structure of SOM includes an input layer, a competitive layer, and an output layer; the input layer is used to receive and transfer the expression matrix after feature extraction; the competitive layer is used to analyze and compare the expression matrix, find patterns and classify; the output layer is used to output the clustering result;

[0015] Step 3: Calculate the consensus matrix; use the clustering-based similarity partitioning algorithm CSPA to calculate the consensus matrix C;

[0016] Step 4: Consensus clustering; construct a graph c according to the consensus matrix C. The nodes Node in the graph c represent cells, and the weight edge of the edge represents the probability that two nodes are in the same partition; the final output single-cell clustering result is the label obtained after consensus clustering.

[0017] More preferably, in step one, specifically: using a marker gene set to screen the columns of the cell expression matrix, extracting features that have a greater impact on cell types; and setting a variance threshold, further reducing the dimension through variance screening, and genes with a variance change lower than the variance threshold are screened out.

[0018] More preferably, in step two, the Euclidean distance calculation formula is as shown in formula (I):

[0019]

[0020] In formula (I), d(x, y) represents the distance between two cells, n represents the number of features; x and y respectively represent cell x and cell y, and x i , y i respectively represent the i-th expression values of cell x and cell y.

[0021] More preferably, in step two, for the similarity matrix D, regarding it as a weighted graph to construct a KNN graph, including: taking the nodes in the similarity matrix D as the nodes in the KNN graph, K is the number of the nearest neighbors, and the distance between two nodes is the Euclidean distance between these two nodes.

[0022] More preferably, in step two, constructing a shared neighbor graph according to the KNN graph, including: the nodes of the shared neighbor graph are cells, and the edges are defined according to whether there exists at least one pair of nodes with a common KNN; the weight w(x i , y i ) of the edge e(x i , x j ) is defined as the difference between k and the highest average rank in the KNN graph, and the calculation formula is as shown in formula (II):

[0023]

[0024] In formula (II), k is the size of the nearest neighbor list, rank(v, x i ) represents the position of node v in the nearest neighbor list NN(x i ), rank(v, x i ) represents the position of node v in the nearest neighbor list NN(x j ), and rank(v, x j ) represents the position of node v in the nearest neighbor list NN(x j ).

[0025] More preferably, r_cutoff = 0.7 and merge_cutoff = 0.5.

[0026] Further preferably, in step two, by the strategy of finding quasi-cliques in the constructed shared neighbor graph and continuously iterating until the final subgraph is obtained, it includes: First, use the greedy algorithm in the shared neighbor graph to find the maximum quasi-clique associated with each node. After finding all possible quasi-cliques, eliminate redundancy by deleting the quasi-cliques that are completely included in other quasi-cliques; then, identify clusters by merging quasi-cliques, and finally, assign nodes to unique clusters.

[0027] Preferably according to the present invention, in step two, the expression matrix after feature extraction is used as an input and applied to the self-organizing mapping method SOM of deep learning, specifically including:

[0028] First, randomly extract m input samples from the data set, that is, the expression matrix after feature extraction, as the initial weights. For the cell vector X and the weight vector W, perform normalization processing to obtain and Initialize the winning neighborhood r_t; the cell vector X refers to the vector composed of the gene expression values of cells, and the initialization of the weight vector W is performed by randomly selecting the cell vector X.

[0029] Then, for the normalized samples including and Calculate the dot product, and select the node with the largest dot product after calculation as the winning node, as shown in Equation (III):

[0030]

[0031] Finally, adjust the weights of the nodes in the winning neighborhood, that is, update the neurons in the topological neighborhood of the winning neuron using the inner star rule, as shown in Equation (IV):

[0032]

[0033] The finally obtained network weights approach the average value of each input vector; determine whether the learning rate η is lower than the threshold eps. When the learning rate decays to be lower than the threshold eps, the iteration ends.

[0034] Preferably according to the present invention, in step three, the element m of the consensus matrix C ij is defined as the probability that two cells are classified into the same class, and the definitions are shown in Equations (V) and (VI):

[0035] C = {m ij} n×n (V)

[0036]

[0037] Wherein, n represents the number of cells, M represents the number of clustering methods in the first inner layer, represents whether cells i, j are classified into the same class in the m-th clustering method of the first layer.

[0038] Further preferably, M = 2.

[0039] According to the present invention, preferably, in step four, a graph c is constructed based on the consensus matrix C, as shown in formulas (VII) and (VIII):

[0040] Node = n_of_C (VII)

[0041] edge = m ij (VIII)

[0042] Wherein, n represents the points in the consensus matrix D, that is, the cell numbers, and the nodes (Node) in the constructed graph c are in the same order as the nodes in the consensus matrix.

[0043] A high-precision single-cell clustering system based on marker genes and ensemble learning, comprising:

[0044] A feature extraction module, configured to: reduce the dimension of the cell expression matrix and extract cell features by using a feature extraction algorithm and a dimensionality reduction algorithm;

[0045] An inner layer clustering module, configured to: apply the expression matrix after feature extraction as an input to an inner layer clustering method; the inner layer clustering method includes a single-cell clustering method SNN-Cliq and a self-organizing mapping method SOM of deep learning; the expression matrix after feature extraction is respectively applied as an input to the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning;

[0046] A consensus matrix calculation module, configured to: calculate a consensus matrix C by using a clustering-based similarity partitioning algorithm CSPA;

[0047] A consensus clustering module, configured to: construct a graph c according to the consensus matrix C.

[0048] The beneficial effects of the present invention are:

[0049] 1. The marker gene set in the single-cell clustering algorithm - SCMcluster proposed by the present invention can reduce the influence of noise on single-cell data as prior knowledge and effectively extract the features characterizing cells.

[0050] 2. The integrated clustering model in the single-cell clustering algorithm - SCMcluster proposed by the present invention combines the advantages of different clustering methods, further improving the accuracy and robustness of clustering. Tests have proven that the single-cell clustering algorithm - SCMcluster proposed by the present invention is superior to existing methods in terms of feature extraction and clustering performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flow chart of the high-precision single-cell clustering method based on marker genes and ensemble learning of the present invention;

[0052] Figure 2 is a schematic diagram of applying the single-cell RNA sequencing data before and after feature processing of the present invention to eight different clustering methods and measuring the clustering results with five evaluation indicators;

[0053] Figure 3 is a schematic diagram of the comparison of the dimensionality reduction effects of SCMcluster of the present invention with PCA used in t-SNE and pcaReduce and UMAP;

[0054] Figure 4(a) is a schematic diagram of the performance comparison of the high-precision single-cell clustering method based on marker genes and ensemble learning of the present invention on the real dataset Muraro;

[0055] Figure 4(b) is a schematic diagram of the performance comparison of the high-precision single-cell clustering method based on marker genes and ensemble learning of the present invention on the real dataset Baron. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0057] Example 1

[0058] A high-precision single-cell clustering method based on marker genes and ensemble learning, as Figure 1 shown, includes the following steps:

[0059] SCMcluster takes the cell expression matrix M' as input, where the columns correspond to cells and the rows correspond to genes / transcripts; each element of M' corresponds to the expression of a gene / transcript in a given cell; SCMcluster is based on five basic steps ( Figure 1 ). For each parameter in these steps, the user can easily adjust it or set it to a reasonable default value.

[0060] Step 1: Feature extraction; Single-cell RNA sequencing data usually has high dimensionality and sparsity. Therefore, in the present invention, in order to reduce redundant data dimensions, enhance the representational ability of features, and improve the operation speed of clustering algorithms, feature extraction algorithms and dimensionality reduction algorithms are used to reduce the dimension of the cell expression matrix and extract cell features; each element in the cell expression matrix corresponds to the expression of a gene / transcript in a given cell, where columns correspond to cells and rows correspond to genes / transcripts;

[0061] For the feature screening part, only the constructed marker gene set is used, and feature extraction can be achieved by constructing a more sound marker gene scoring system and integrating a more comprehensive marker gene dataset.

[0062] Step 2: Inner clustering; The expression matrix after feature extraction is used as input and applied to the inner clustering method; the present invention integrates two advanced clustering methods, and the inner clustering method includes the single-cell clustering method SNN-Cliq and the self-organizing map method SOM of deep learning; the expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq and the self-organizing map method SOM of deep learning;

[0063] The expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq, including:

[0064] SNN-Cliq is a single-cell clustering algorithm based on subgraph partitioning, and this algorithm is suitable for sparse and large single-cell RNA datasets. First, the Euclidean distance is used to calculate the similarity matrix D corresponding to the expression matrix M;

[0065] Then, for the similarity matrix D, it is regarded as a weighted graph to construct a KNN (K-Nearest Neighbor) graph;

[0066] Again, a shared neighbor graph (SNN graph) is constructed based on the KNN graph;

[0067] Finally, by the strategy of finding quasi-cliques in the shared neighbor graph constructed by the single-cell clustering method SNN-Cliq, continuous iteration is carried out until the final subgraph is obtained; the parameters r_cutoff and merge_cutoff represent the nearest neighbor radius and merge threshold of each pair of Cliq respectively;

[0068] The expression matrix after feature extraction is used as input and applied to the self-organizing map method SOM of deep learning. The topological structure of SOM includes an input layer, a competition layer, and an output layer; the input layer is used to receive and transfer the expression matrix after feature extraction; the competition layer is used to analyze and compare the expression matrix, find patterns and classify; the output layer is used to output the clustering results;

[0069] The topological structure of the competitive layer is a one-dimensional linear structure. Since SOM can represent high-dimensional input data in a low-dimensional space and has the ability of dimensionality reduction, in order to preserve the global features of the data, the present invention removes the dimensionality reduction operation when using SOM.

[0070] In addition, in the design of SOM, an improved network can be used to achieve faster clustering.

[0071] Step 3: Calculate the consensus matrix; use the clustering-based similarity partitioning algorithm CSPA to calculate the consensus matrix C;

[0072] Step 4: Consensus clustering; in the clustering framework of the present invention, the final partitioning result is obtained by spectral clustering. According to the definition of the consensus matrix, the consensus matrix has a similar structure to the adjacency matrix of a graph, both of which are symmetric square matrices. Therefore, a graph c is constructed based on the consensus matrix C. The nodes Node in the graph c represent cells, and the weight edge of the edge represents the probability that two nodes (i.e., cells) are in the same partition. As a subgraph segmentation method, the spectral clustering algorithm shows good performance in finding subgraphs. The present invention uses the spectral clustering method on the weighted graph c constructed by the consensus matrix C to obtain the final cell clustering result.

[0073] The final output single-cell clustering result is the label obtained after performing consensus clustering (clustering of spectral clustering on the consensus matrix).

[0074] Embodiment 2

[0075] A high-precision single-cell clustering method based on marker genes and ensemble learning according to Embodiment 1, wherein the difference lies in:

[0076] In step one, specifically: use the marker gene set to screen the columns of the cell expression matrix, extract the features that have a greater impact on cell types; and set a variance threshold to further reduce the dimension through variance screening, and genes with a variance change lower than the variance threshold are screened out. The detailed steps include: First, construct the marker gene set. Use two relatively comprehensive public single-cell databases - the CellMarker database (http: / / biocc.hrbmu.edu.cn / CellMarker / download.jsp) and the PanglaoDB database (http: / / biocc.hrbmu.edu.cn / CellMarker / download.jsp), extract the Official gene symbol of different species in the PanglaoDB database and the geneSymbol from cancer cells in the CellMarker database as the marker gene set. Then, use the constructed marker gene set to screen the rows of the expression matrix and extract the covered genes as features. Finally, use the variance screening method to further reduce the dimension. Set the variance threshold, and genes with a variance change lower than the variance threshold are screened out. According to the above steps, the expression matrix after feature extraction is obtained.

[0077] In step two, the Euclidean distance calculation formula is shown in formula (I):

[0078]

[0079] In formula (I), d(x,y) represents the distance between two cells, n represents the number of features; x and y represent cell x and cell y respectively, and x i , y i respectively represent the i-th expression value of cell x and cell y.

[0080] In step two, for the similarity matrix D, regard it as a weighted graph to construct the KNN graph, including: take the nodes in the similarity matrix D as the nodes in the KNN graph, K is the number of the nearest neighbors, and the distance between two nodes is the Euclidean distance between these two nodes.

[0081] In step two, construct the shared neighbor graph according to the KNN graph, including: the nodes of the shared neighbor graph are cells, and the edges are defined according to whether there exists at least one pair of nodes (that is, a pair of cells) that have a common KNN; the weight w(x i , y i ) of the edge e(x i , x j ) is defined as the difference between k and the highest average rank in the KNN (K-Nearest Neighbor) graph, and the calculation formula is shown in formula (II):

[0082]

[0083] In formula (II), k is the size of the nearest neighbor list, and rank(v, x i ) represents the position of node v in the i nearest neighbor list NN(x i ), and rank(v, x j ) represents the position of node v in the j nearest neighbor list NN(x j ).

[0084] r_cutoff = 0.7, merge_cutoff = 0.5.

[0085] In step two, by searching for quasi-cliques in the constructed shared neighbor graph and iterating continuously until the final subgraph is obtained, it includes: First, use the greedy algorithm in the shared neighbor graph to find the maximum quasi-clique associated with each node. After finding all possible quasi-cliques, eliminate redundancy by deleting the quasi-cliques that are completely contained in other quasi-cliques; then, identify clusters by merging quasi-cliques, and finally, assign nodes to unique clusters.

[0086] In step two, the expression matrix after feature extraction is used as input and applied to the self-organizing mapping method SOM of deep learning, specifically including:

[0087] First, randomly select m input samples from the data set, that is, the expression matrix after feature extraction, as the initial weights. For the cell vector X and the weight vector W, perform normalization to obtain and Initialize the winning neighborhood r_t; the cell vector X refers to the vector composed of the gene expression values of cells, and the initialization of the weight vector W is performed by randomly selecting the cell vector X;

[0088] Then, calculate the dot product for the normalized samples including and and select the node with the largest calculated dot product as the winning node, as shown in formula (III):

[0089]

[0090] Finally, adjust the weights of the nodes within the winning neighborhood, that is, update the neurons within the topological neighborhood of the winning neuron using the inner star rule, as shown in formula (IV):

[0091]

[0092] The finally obtained network weights approach the average value of each input vector; it is determined whether the learning rate η is lower than the threshold eps, and when the learning rate decays to be lower than the threshold eps, the iteration ends. The threshold eps is the end point of the learning rate decay and can be determined according to actual needs. If not specified, it defaults to 0.

[0093] In step three, the element m of the consensus matrix C ij is defined as the probability that two cells are classified into the same class, and the definitions are shown in formulas (V) and (VI):

[0094] C = {m ij} n×n (V)

[0095]

[0096] where n represents the number of cells, M represents the number of clustering methods in the first inner layer, indicates whether cells i and j are classified into the same class in the m-th clustering method of the first layer.

[0097] M = 2.

[0098] In step four, a graph c is constructed according to the consensus matrix C, as shown in formulas (VII) and (VIII):

[0099] Node = n_of_C (VII)

[0100] edge = m ij (VIII)

[0101] where n represents the points in the consensus matrix D, that is, the cell numbers, and the nodes (Node) in the constructed graph c are in the same order as the nodes in the consensus matrix.

[0102] To comprehensively evaluate the performance of the high-precision single-cell clustering method based on marker genes and ensemble learning of the present invention, five commonly used clustering evaluation indicators are introduced: including the Rand Index (RI), Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), Adjusted Mutual Information (AMI), and Fowlkes and Mallows Index (FMI). In addition, a sparsity coefficient is constructed to quantify the sparsity degree of the matrix. The specific definitions are as follows:

[0103] Sparce_index = x ÷ num(M)

[0104] Among them, x represents the number of 0 elements in matrix M, and num(M) represents the total number of elements in the matrix.

[0105] Comparing SCMcluster with four traditional clustering algorithms and four widely used single-cell clustering algorithms (Table 1) verifies the superiority of the method, as shown in Table 1:

[0106] Table 1

[0107]

[0108] The single-cell RNA sequencing data before and after feature processing are applied to eight different clustering methods (four traditional clustering methods and four single-cell clustering methods), and five evaluation indicators are used to measure the correctness of the clustering results. Figure 2 This is a schematic diagram of the present invention applying the single-cell RNA sequencing data before and after feature processing to eight different clustering methods and using five evaluation indicators to measure the clustering results; considering that three single-cell clustering methods, SC3, Seurat, and pcaReduce, cover data dimensionality reduction processing internally and SOM itself can also be regarded as a dimensionality reduction algorithm, in order to verify the reliability, the variance screening step is removed when applying the above methods.

[0109] Some dimensionality reduction sub-steps widely used in single-cell clustering are also compared, including three methods: t-SNE, PCA used in pcaReduce, and UMAP. Figure 3 This is a schematic diagram of the comparison of the dimensionality reduction effects of SCMcluster of the present invention with t-SNE, PCA used in pcaReduce, and UMAP.

[0110] The method proposed in the present invention was compared with general single-cell clustering methods in two single-cell RNA datasets of human and mouse cells respectively, and the performance on cross-species datasets was compared. Figure 4(a) is a schematic diagram of the performance comparison of the high-precision single-cell clustering method based on marker genes and ensemble learning in the present invention on the real dataset Muraro; Figure 4(b) is a schematic diagram of the performance comparison of the high-precision single-cell clustering method based on marker genes and ensemble learning in the present invention on the real dataset Baron. The results show that SCMcluster performs better than six test methods on all benchmark datasets. Specifically, on the human dataset muraro, the ARI value is 8.5% higher than that of the second-ranked spectral clustering and 20.0% higher than that of the third-ranked SC3; on the mouse dataset, the performance of SCMcluster is far higher than that of other methods, and the five index values are 94.41%, 97.74%, 89.83%, 89.66%, and 95.99% respectively, while the results of the second-ranked SC3 are only 43.36%, 81.70%, 73.19%, 72.64%, and 58.18%. The analysis and comparison results show that the clustering model SCMcluster proposed in the present invention performs best in feature processing and clustering performance.

[0111] Example 3

[0112] A high-precision single-cell clustering system based on marker genes and ensemble learning, comprising:

[0113] A feature extraction module, configured to: reduce the dimension of the cell expression matrix and extract cell features by using a feature extraction algorithm and a dimensionality reduction algorithm;

[0114] An inner-layer clustering module, configured to: apply the expression matrix after feature extraction as input to an inner-layer clustering method; the inner-layer clustering method includes the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning; apply the expression matrix after feature extraction as input to the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning respectively;

[0115] A consensus matrix calculation module, configured to: calculate a consensus matrix C by using a clustering-based similarity partitioning algorithm CSPA;

[0116] A consensus clustering module, configured to: construct a graph c according to the consensus matrix C.

Claims

1. A high-precision single-cell clustering method based on marker genes and ensemble learning, characterized in that, The steps are as follows: Step 1: Feature extraction; Use a feature extraction algorithm and a dimensionality reduction algorithm to reduce the dimensionality of the cell expression matrix and extract cell features; each element in the cell expression matrix corresponds to the expression of a gene / transcript in a given cell, where the columns correspond to cells and the rows correspond to genes / transcripts; Step 2: Inner clustering; the expression matrix after feature extraction is used as input and applied to the inner clustering method; the inner clustering method includes the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning; the expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq and the self-organizing mapping method SOM of deep learning; The expression matrix after feature extraction is used as input and applied to the single-cell clustering method SNN-Cliq, including: First, calculate the similarity matrix D corresponding to the expression matrix M using the Euclidean distance; Then, for the similarity matrix D, regard it as a weighted graph to construct a KNN graph; Again, construct a shared neighbor graph according to the KNN graph; Finally, by the strategy of finding quasi-cliques in the constructed shared neighbor graph and continuously iterating until the final subgraph is obtained; the parameters r_cutoff and merge_cutoff represent the nearest neighbor radius and the merging threshold of each pair of Cliq respectively; The expression matrix after feature extraction is used as input and applied to the self-organizing mapping method SOM of deep learning. The topological structure of SOM includes an input layer, a competitive layer, and an output layer; the input layer is used to receive and transfer the expression matrix after feature extraction; the competitive layer is used to analyze and compare the expression matrix, find patterns and classify; the output layer is used to output the clustering results; Step 3: Calculate the consensus matrix; use the clustering-based similarity partitioning algorithm CSPA to calculate the consensus matrix C; Step 4: Consensus clustering; construct a graph c according to the consensus matrix C. The nodes Node in the graph c represent cells, and the weight edge of the edge represents the probability that the two nodes are in the same partition; the final output single-cell clustering result is the label obtained after consensus clustering; Among them, r_cutoff = 0.7 and merge_cutoff = 0.

5.

2. The high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, wherein In Step 1, specifically: use the marker gene set to screen the columns of the cell expression matrix, extract the features that have a greater impact on the cell type; and set a variance threshold, and further reduce the dimensionality through variance screening. Genes with a variance change lower than the variance threshold are screened out.

3. A high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, characterized in that, In Step 2, the Euclidean distance calculation formula is shown in Equation (I): In formula (I), d(x, y) represents the distance between two cells, n represents the number of features; x, y represent cell x and cell y, x respectively. i ,y i represent the i-th expression value of cell x and cell y respectively.

4. A high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, characterized in that In Step 2, for the similarity matrix D, regard it as a weighted graph to construct a KNN graph, including: taking the nodes in the similarity matrix D as the nodes in the KNN graph, K is the number of the nearest neighbors, and the distance between two nodes is the Euclidean distance between the two nodes.

5. A high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, characterized in that In step two, construct a shared neighbor graph based on the KNN graph, including: the nodes of the shared neighbor graph are cells, and the edges are defined according to whether there exists at least one pair of nodes with a common KNN; the weight w(x i , y i ) of the edge e(x i , x j ) is defined as the difference between k and the highest average rank in the KNN graph, and the calculation formula is shown in formula (II): In formula (II), k is the size of the nearest neighbor list, and rank(v, x i ) represents the position of node v in the i nearest neighbor list NN(x i ), and rank(v, x j ) represents the position of node v in the j nearest neighbor list NN(x j ).

6. The high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, wherein In step 2, by the strategy of finding quasi - cliques in the constructed shared neighbor graph, iterate continuously until the final sub - graph is obtained, including: First, use the greedy algorithm in the shared neighbor graph to find the maximum quasi - clique associated with each node. After finding all possible quasi - cliques, eliminate redundancy by deleting the quasi - cliques that are completely contained in other quasi - cliques; Then, identify clusters by merging quasi - cliques. Finally, assign nodes to unique clusters.

7. A high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, characterized in that In step 2, the expression matrix after feature extraction is used as input and applied to the self - organizing map method SOM of deep learning, specifically including: First, randomly select m input samples from the dataset, that is, the expression matrix after feature extraction, as the initial weights. For the cell vector X and the weight vector W, perform normalization processing to obtain and Initialize the winning neighborhood r_t; the cell vector X refers to the vector composed of the gene expression values of cells, and the initialization of the weight vector W is performed by randomly selecting the cell vector X; Then, for the normalized samples including and calculate the dot product, and select the node with the largest dot product after calculation as the winning node, as shown in Equation (III): Finally, adjust the weights of the nodes in the winning neighborhood, that is, update the neurons within the topological neighborhood of the winning neuron using the inner - star rule, as shown in Equation (IV): The finally obtained network weights approach the average value of each input vector; Determine whether the learning rate η is lower than the threshold eps. When the learning rate decays to be lower than the threshold eps, the iteration ends.

8. A high-precision single-cell clustering method based on marker genes and ensemble learning according to claim 1, characterized in that In step 3, the element m of the consensus matrix C ij is defined as the probability that two cells are classified into the same class, and the definition is shown in formulas (V) and (VI) as follows: C = {m ij} n×n (V) where n represents the number of cells, M represents the number of clustering methods in the first inner layer, indicates whether cells i and j are classified into the same class in the m-th clustering method of the first layer; Where M = 2.

9. A high-precision single-cell clustering method based on marker genes and ensemble learning according to any one of claims 1-8, characterized in that In step 4, construct graph c according to the consensus matrix C, as shown in Equation (VII) and Equation (VIII): Node=n_of_C(VII) edge=m ij (VIII) Where n represents the points in the consensus matrix C, that is, the cell numbers. The nodes Node in the constructed graph c are in the same order as the nodes in the consensus matrix.

10. A high-precision single-cell clustering system based on marker genes and ensemble learning, characterized in that, Including: A feature extraction module, configured to: use a feature extraction algorithm and a dimensionality reduction algorithm to reduce the dimension of the cell expression matrix and extract cell features; Each element in the cell expression matrix corresponds to the expression of a gene / transcript in a given cell, where the columns correspond to cells and the rows correspond to genes / transcripts; An inner - layer clustering module, configured to: use the expression matrix after feature extraction as input and apply it to the inner - layer clustering method; The inner - layer clustering method includes the single - cell clustering method SNN - Cliq and the self - organizing map method SOM of deep learning; The expression matrix after feature extraction is respectively used as input and applied to the single - cell clustering method SNN - Cliq and the self - organizing map method SOM of deep learning; A consensus matrix calculation module, configured to: calculate the consensus matrix C using the clustering - based similarity partitioning algorithm CSPA; A consensus clustering module, configured to: construct graph c according to the consensus matrix C; The nodes Node in graph c represent cells, and the weight edge of the edge represents the probability that two nodes are in the same partition; The finally output single - cell clustering result is the label obtained after consensus clustering.

Citation Information

Patent Citations

  • Method of identifying cell types based on single-cell RNA sequencing data

    CN110797089A

  • Large-scale single cell transcriptome data efficient clustering method

    CN113178233A