A method, device and computer readable storage medium for cell heterogeneity analysis of large-scale single-cell sequencing data
By combining hierarchical sampling and nuclear nonnegative matrix factorization with deep neural networks, the problems of high computational complexity and low recognition ability in cell heterogeneity analysis of large-scale single-cell sequencing data are solved, and more efficient and accurate cell type identification is achieved.
Patent Information
- Application Number
- CN202310784889.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-06-29
AI Technical Summary
Existing technologies suffer from high computational complexity and low recognition ability in the analysis of cell heterogeneity in large-scale single-cell sequencing data. Furthermore, traditional methods suffer from poor clustering results and high computational cost.
A representative subset is selected using hierarchical sampling, clustering is performed using kernel nonnegative matrix factorization, and the remaining subset is classified using a deep neural network. Finally, the clustering results of the overall data are obtained by merging the results.
It improves the cell type identification capability of large-scale single-cell sequencing data, reduces computational complexity, and enhances the accuracy and robustness of clustering and classification.
Smart Images

Figure CN116933072B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of biological information, and particularly relates to a cell heterogeneity analysis method, device and computer readable storage medium for large-scale single cell sequencing data. BACKGROUND
[0002] Single-cell RNA sequencing (scRNA-seq) technology provides unprecedented possibilities for intercellular heterogeneity analysis by analyzing gene expression in individual cells. Cell type recognition is a very critical link in the process of cell heterogeneity analysis, and is usually achieved by clustering. Due to technical noise and random transcription, single-cell sequencing data often has characteristics such as high dimensionality, high noise and high sparsity, which affect the recognition of cell types and further hinder downstream analyses such as gene differential expression and heterogeneity analysis. In recent years, with the rapid development of single-cell sequencing technology, a large amount of complex single-cell sequencing data has been generated, and heterogeneity analysis is restricted in many aspects, mainly in computational complexity and analysis accuracy, and special analysis techniques are needed to achieve accurate data interpretation.
[0003] Traditional methods reduce the dimensionality of single-cell sequencing data and use traditional machine learning clustering algorithms in low-dimensional space for heterogeneity analysis, such as k-means, hierarchical clustering, density clustering, and graph-based clustering. Considering the special properties of single-cell sequencing data, clustering methods designed for single-cell sequencing data have been proposed, such as SC3, Seurat, SIMLR, and Scanpy. SC3 performs consistent clustering on single-cell data through parallelization, and Seurat and scanpy are similar single-cell analysis frameworks. By default, they use Louvain algorithm and k-nearest neighbor method to segment cells with similar gene expression and achieve cell recognition. SIMLR clusters single-cell data based on kernel similarity learning. When the number of cells is small, the average performance of SC3 is better than other methods. The results of scanpy are similar to seruat. Large-scale single-cell sequencing data has high dimensionality, sparsity and complexity, which brings higher computational complexity and lower recognition ability to traditional heterogeneity analysis methods. The best-performing SC3 clustering method in the above methods has relatively high computational complexity when analyzing large-scale single-cell sequencing data.
[0004] In analyzing large-scale scRNA sequence data, some methods use integration or parallel computing to complete such a large amount of calculation. Sharp is to aggregate millions of cells by integrating random projection methods. Bigscale uses large sample size to estimate high-precision and comprehensive noise numerical model. Sharp improves the dimensionality reduction method. The model of Bigscale can quantify the distance between cells. Both methods are hierarchical clustering of data, but the accuracy of the results is still low. The SSCC framework proposed by Ren uses feature engineering and projection technology to improve the clustering framework through subsampling and classification. However, the SSSC method uses random subset selection for clustering, which can cause biased clustering results and affect the accuracy of subsequent classification. It uses traditional clustering and classification algorithms, which not only have large computational complexity, but also have poor clustering results. SUMMARY
[0005] The technical problem to be solved by the present application is how to analyze cell heterogeneity of large-scale single cell sequencing data and how to effectively identify cell types of large-scale single cell sequencing data.
[0006] To solve the above technical problems, the present application first provides a method for clustering and / or cell heterogeneity analysis of single cell sequencing data, which can include the following steps:
[0007] A1, representative subset obtaining: using hierarchical sampling method to extract a representative subset V from a single cell sequencing data set s , dividing the single cell sequencing data set into the representative subset V s and the remaining subset V r Two parts;
[0008] A2, clustering of the representative subset: using kernel non-negative matrix factorization method to cluster the representative subset to obtain the clustering label of the representative subset;
[0009] A3, obtaining classification label of the remaining subset: using the clustering label of the representative subset and the representative subset as a training set, using a deep neural network method for training to obtain a neural network classification model, and using the neural network classification model to obtain the classification label of the remaining subset;
[0010] A4, output of overall data clustering result: merging the clustering label of the representative subset and the classification label of the remaining subset to obtain the clustering result and / or cell heterogeneity analysis result of the single cell sequencing data.
[0011] The hierarchical sampling can be stratified random sampling.
[0012] In the above method, the representative subset obtaining of A1 can further include the following steps:
[0013] A1-1) Data dimension reduction: using principal component analysis method to reduce dimension of the single-cell sequencing data set, then using farthest point sampling method to obtain k representative cell sample center points, and dividing the single-cell sequencing data set based on the representative cell sample center points to obtain k sample modules (layers);
[0014] A1-2) Representative subset obtaining: using module optimization algorithm apri cot to select representative data of the sample module from the sample module, and combining the representative data of each sample module to obtain the representative subset;
[0015] A1-3) Representative subset and remaining subset ratio determination: using principal component analysis method to reduce dimension of the representative subset and the remaining subset of the single-cell sequencing data set after removing the representative subset, obtaining the reduced representative subset and the reduced remaining subset, calculating the KL divergence of the sample distribution of the reduced representative subset and the sample distribution of the reduced remaining subset, determining the proportion of the representative subset in the single-cell sequencing data set according to the value of the KL divergence, and dividing the single-cell sequencing data set into the representative subset V s and the remaining subset V r .
[0016] In A1-3), the minimum value of the KL divergence can be selected to determine the proportion of the representative subset in the single-cell sequencing data set.
[0017] In the above method, A2 the clustering of the representative subset can include the following steps:
[0018] Constructing a non-linear mapping for the representative subset, obtaining a kernel matrix induced after mapping, performing non-negative matrix factorization on the kernel matrix to obtain a coefficient matrix and a basis matrix, performing hierarchical clustering on the coefficient matrix to obtain a clustering label, calculating a similarity matrix of the clustering label to obtain a consistency similarity matrix, and performing clustering on the consistency similarity matrix to obtain a clustering label of the representative subset.
[0019] In the above method, A2 in the process of clustering the representative subset using kernel non-negative matrix factorization method, the determination method of the parameter rank k can be: performing singular value decomposition on the single-cell sequencing data set, and setting the parameter rank k when the square sum of the first singular value accounts for 98% of the square sum of all singular values.
[0020] In the above method, A3 the neural network classification model can be obtained by the method comprising the following steps:
[0021] The clustering label of the representative subset and the representative subset are used as a training set, and a unidirectional multilayer feedforward neural network is used to train the training set based on a back propagation algorithm of a stochastic gradient descent to obtain the neural network classification model.
[0022] To solve the above technical problems, the application also provides a device for single-cell sequencing data clustering and / or cell heterogeneity analysis, which can include the following modules:
[0023] B1, a representative subset obtaining module: used for obtaining a representative subset V from a single-cell sequencing data set using hierarchical sampling method s The single-cell sequencing data set is divided into a representative subset V s and a remaining subset V r Two parts;
[0024] B2, a clustering module of the representative subset: used for clustering the representative subset using kernel non-negative matrix factorization method to obtain a clustering label of the representative subset;
[0025] B3, a remaining subset classification label obtaining module: used for using the clustering label of the representative subset and the representative subset as a training set, training using a deep neural network method to obtain a neural network classification model, and obtaining the remaining subset classification label using the neural network classification model;
[0026] B4, an overall data clustering result output module: used for merging the clustering label of the representative subset and the remaining subset classification label to obtain a clustering result of the single-cell sequencing data and / or a cell heterogeneity analysis result.
[0027] In the above device, the representative subset obtaining module B1 can further include the following modules:
[0028] B1-1) data dimensionality reduction module: using principal component analysis method to reduce the dimensionality of the single-cell sequencing data set, then using farthest point sampling method to obtain k representative cell sample center points, and dividing the single-cell sequencing data set based on the representative cell sample center points to obtain k sample modules;
[0029] B1-2) representative subset obtaining module: using module optimization algorithm apricot to select representative data of the sample module from the sample module, and merging the representative data of each sample module to obtain the representative subset;
[0030] B1-3) represents the subset and the remaining subset ratio determination module: using principal component analysis method to reduce dimensionality of the representative subset and the remaining subset of the representative subset in the single cell sequencing data set, obtaining the reduced dimensionality of the representative subset and the reduced dimensionality of the remaining subset, calculating the KL divergence of the sample distribution of the reduced dimensionality of the representative subset and the sample distribution of the reduced dimensionality of the remaining subset, determining the proportion of the representative subset in the single cell sequencing data set according to the value of the KL divergence, and dividing the single cell sequencing data set into representative subset and remaining subset.
[0031] In the above device, the value of the minimum KL divergence in B1-3) can be selected to determine the proportion of the representative subset in the single cell sequencing data set.
[0032] In the above device, the clustering module of the representative subset in B2 can be realized by the method comprising the following steps:
[0033] The representative subset is constructed into a nonlinear mapping, a kernel matrix induced by the mapping is obtained, a non-negative matrix factorization is performed on the kernel matrix to obtain a coefficient matrix and a basis matrix, a hierarchical clustering is performed on the coefficient matrix to obtain a clustering label, a similarity matrix of the clustering label is calculated to obtain a consistency similarity matrix, and the clustering label of the representative subset is obtained by clustering the consistency similarity matrix.
[0034] In the B2 module of the above device, during the clustering of the representative subset using the kernel non-negative matrix factorization method, the determination method of the parameter rank k can be: singular value decomposition is performed on the single cell sequencing data set, and when the square sum of the first singular value accounts for 98% of the square sum of all singular values, the parameter rank k is set.
[0035] In the above device, the neural network classification model in B3 can be obtained by the method comprising the following steps:
[0036] The clustering label of the representative subset and the representative subset are used as a training set, a one-way multilayer feedforward neural network is used, a cross-entropy loss function is adopted, and a back propagation algorithm based on stochastic gradient descent is used to train the training set, and the neural network classification model is obtained.
[0037] In order to solve the above technical problems, the present application also provides a computer readable storage medium storing a computer program, wherein the computer program can make the computer execute the steps of the above method.
[0038] Any of the following applications of the above method and / or the above device and / or the above computer readable storage medium also belongs to the protection scope of the present application:
[0039] D1, in the preparation of the product for identifying the cell type of the sequencing data;
[0040] D2, use in preparing a product for predicting tumor cell heterogeneity;
[0041] D3, use in preparing a product for developing a tumor marker.
[0042] In order to solve the sampling bias problem in the prior art, the present application proposes a large-scale single cell sequencing data heterogeneity analysis method. In the case of uneven distribution of sequencing data, the most representative subset is selected for heterogeneity analysis according to the stratified sampling method. Based on the KL divergence, the optimal sampling division is determined, and the heterogeneity analysis ability of the representative subset is migrated to the remaining single cell data set, and the overall clustering classification algorithm is improved, and the type recognition ability of large-scale single cell sequencing data is strengthened.
[0043] Compared with the prior art, the effect of the present application lies in the selection of the optimal representative subset. Unlike the random sampling of the prior art, the present application uses stratified sampling method to select the most representative subset. In addition, the size of the sample quantity is not explicitly stated in the prior art, while the present application optimizes the optimal sample quantity based on the KL-divergence, and optimally divides the data of the original data clustering classification. The present application uses kernel non-negative matrix factorization for low-dimensional representation and clustering analysis. Compared with the traditional clustering method such as k-means in the prior art, KNMF can extract the non-linear complex features hidden in the original data in non-negative matrix factorization. For data classification, the neural network used in the present application performs deep learning on the samples, and the data set is more efficient and accurate in expression, and has stronger robustness, and has better performance under large data samples. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The framework diagram of the large-scale single cell sequencing data heterogeneity analysis method of the present application.
[0045] Figure 2 The flowchart of the present application for selecting the best representative set by stratified sampling.
[0046] Figure 3 The results of the clustering method KNMF and the existing five clustering methods on the representative subset.
[0047] Figure 4 The results of the classification method DNN used in the present application and the existing three clustering methods on the remaining subset.
[0048] Figure 5 The results of the KNMF-DNN method and other clustering methods on the whole large-scale single cell sequencing data.
[0049] Figure 6The KL divergence curves at different sampling rates and the clustering results of the KNMF-DNN framework at different sampling rates. The left figure is the KL divergence curve at different sampling rates; the right figure is the clustering result of the KNMF-DNN framework at different sampling rates. DETAILED DESCRIPTION
[0050] The application will be further described in conjunction with the specific embodiments. The examples given are only to illustrate the application, and are not intended to limit the scope of the application. The examples provided below can serve as a guide for further improvement by those skilled in the art, and do not constitute any limitation on the application.
[0051] The experimental methods in the following examples are all routine methods, unless otherwise specified, according to the techniques or conditions described in the literature in the art or according to the product instructions. The materials, reagents, etc. used in the following examples, unless otherwise specified, can be obtained commercially.
[0052] Example 1, Establishment of a method for analyzing cell heterogeneity based on large-scale single-cell sequencing data
[0053] The data set used in the embodiments of the application includes three single-cell sequencing data sets:
[0054] zeisel data set: containing 3005 cells, 19972 genes, 9 cell types (related literature: Zeisel, A, et al. Brain structure. Cell types in the mouse cortex and hippocampus revealed by single-cell RNA-seq. Science 347.6226 (2015): 1138);
[0055] Marques data set: containing 5053 cells, 23556 genes, 13 cell types (related literature: S. Marques, A. Zeisel, S. Codeluppi, D. Van Bruggen, A. Mendanha Falcao, L. Xiao, H. Li, M. Haring, H. Hochgerner, R. A. Romanov et al., “Oligodendrocyte heterogeneity in the mouse juvenile and adult central nervous system,” Science, vol. 352, no. 6291, pp. 1326-1329, 2016);
[0056] Covid-19 dataset: contains 8004 cells, 26924 genes, 8 cell types (relevant website: https: / / cells.ucsc.edu / ?ds=covid19-bronch-epi#).
[0057] The method established by the present application based on large-scale single cell sequencing data heterogeneity analysis is shown in the method framework diagram as Figure 1
[0058] 1. A method for selecting a representative subset V from a gene expression matrix using hierarchical sampling (stratified random sampling) s .
[0059] Given the gene expression matrix V of the single cell sequencing raw data, the raw data set V is:
[0060] V∈R n×p Formula I,
[0061] n represents the number of cells in the data set V, and the gene expression matrix V is an n-row p-column matrix, and p represents the number of genes in the data set, and the unit is individual. For example, for the zeisel data set, n represents 3005 cells, p represents 19972 genes, and the features in each column are the expression amounts of genes; R represents that the value of the matrix is a real number.
[0062] Select M cells from n cells as a representative subset.
[0063] 1.1 Dimensionality reduction and sample division module of the original data
[0064] The original data is subjected to principal component analysis method (PCA) dimensionality reduction; then 5 center point representative cell samples c1, c2,..., c5 are selected from the entire data set using farthest point sampling method (FPS).
[0065] The distance between other cell samples and the center point in the gene expression matrix of the original data is calculated (distance measurement: for point cloud data, the Euclidean distance, i.e. the straight line distance between two points in space, is generally used), and the other cell samples are assigned to the nearest center point cell sample, and the data is divided into 5 sample modules, each containing a center point cell sample c and corresponding assigned cell samples.
[0066] 1.2 Best representative subset V s Obtain
[0067] To summarize the massive dataset into a minimal redundant subset that still represents the original data, the apricot module optimization algorithm (related literature: Schreiber J, Bilmes J, Noble W S. apricot: Submodular selection for data summarization in Python[J].2019.DOI:10.48550 / arXiv.1906.03543) was used to optimize the submodules and select representative subsets from the large dataset. Based on the module optimization algorithm, representative data (including gene and cell type data corresponding to the cells) of M / 5 cells were selected from each sample module obtained in step 1.1. The original gene expression of the representative data was then restored based on the index, resulting in representative sample data selected from each of the 5 blocks.
[0068] The representative sample data selected from each of the five blocks are combined to form the final best representative subset V. s (Contains M cells and their corresponding data) Figure 1 The flowchart for selecting the best representative set by stratified sampling (using the representative subset) is as follows: Figure 2 As shown.
[0069] Using a hierarchical representative set selection method can ensure the representativeness of the selected dataset. In addition, it is necessary to determine the size of the representative subset. In this embodiment, the M values of the three datasets are as follows: for the COVID dataset, the sampling rate is 40% and M is 3200; for the Zeisel dataset, the sampling rate is 50% and M is 1500; and for the Marques dataset, the sampling rate is 50% and M is 2500.
[0070] 1.3 Partitioning the representative subset and the remaining subset V r
[0071] The representative subset V is determined by calculating the KL divergence between the sample distribution of the representative set and the distribution of the remaining sample data after removing the representative sample. s and the remaining subset V r ( Figure 1 The optimal partitioning is represented by the remaining subset.
[0072] For large-scale high-dimensional data, the PCA dimensionality reduction method is used to reduce the dimension of the representative subset and the remaining subset by 20, obtaining the dimensionality-reduced representative subset and the dimensionality-reduced remaining subset. The first 20 principal components are used to represent the original data.
[0073] The present application assumes that the probability distribution of the sequencing data is a multivariate Gaussian distribution, and the probability distribution functions of the representative sample and the remaining sample are p1(x|μ1,∑1) and p2(x|μ2,∑2) respectively:
[0074]
[0075]
[0076] In formula II, p1 is the probability distribution function of the representative sample, μ1 is the mean value of the representative sample, and ∑1 is the covariance of the representative sample;
[0077] In formula III, p2 is the probability distribution function of the remaining sample, μ2 is the mean value of the remaining sample, and ∑2 is the covariance of the remaining sample.
[0078] For different sampling rates, the distributions of the representative subset and the remaining subset are obtained, and the KL divergence is calculated, and the calculation formula is:
[0079]
[0080] When the sampling rates are 10%, 20%, 30%, 40%, 50%, 60%, 70% and 80% respectively, the distributions p1(x|μ1,∑1) and p2(x|μ2,∑2) of the representative subset V s and the remaining subset V r , and the sampling rate with the minimum KL divergence is the optimal proportion of the representative subset. In this embodiment, the optimal proportions of the three data set samples are: for the covid data set, the sampling rate is 40%; for the zeisel data set, the sampling rate is 50%; and for the Marques data set, the sampling rate is 50%.
[0081] 2. Clustering of the representative subset V s
[0082] The Kernel NMF (KNMF) method is used to cluster the representative subset V s selected in step 1.
[0083] For a high-dimensional nonlinear gene expression matrix V s =[v1,v2...v s ,], a nonlinear mapping is constructed, and the kernel matrix induced by the mapping is represented as:
[0084]
[0085]
[0086] In formula V, p represents the number of subset sample expression genes; i represents the ith cell, j represents the jth cell, and k represents the kth expression gene; v ik represents the expression amount of the kth gene in the ith cell, v jk represents the expression amount of the kth gene in the jth cell.
[0087] The non-negative matrix factorization is performed on the kernel matrix to obtain a coefficient matrix and a base matrix:
[0088] K s×s ≈(W φ ) s×k φ k×s (H φ ) φ In formula VII, W φ is the coefficient matrix, and H i is the base matrix.
[0090] The coefficient matrix can be regarded as a dimension-reduced representation of the original matrix. Uniform clustering is performed on the coefficient matrix W i , and the specific process is as follows:
[0091] The kernel non-negative matrix factorization is performed on the kernel matrix K for multiple times to obtain:
[0092] K=W i H i i=1, 2,..., m formula VIII
[0093] The hierarchical clustering is performed on W i to obtain clustering labels L i , and a similarity matrix S s of L s is calculated.
[0094] According to
[0095] a uniform similarity matrix S is obtained, the hierarchical clustering is performed on 1-S, and finally, the uniform clustering labels L are obtained, which are the clustering labels L s representing the subset V r .
[0096] When clustering using the KNMF clustering method, it is necessary to determine the rank parameter k. This invention performs singular value decomposition (SVD) on the original COVID-19 dataset. The rank parameter k is set when the sum of the squares of the first singular value accounts for 98% of the sum of squares of all singular values, i.e., 23. Regarding the selection of the parameter k for kernel nonnegative matrix factorization, referring to previous singular value decomposition methods, it is generally found that the sum of the first 10% or even 1% of singular values accounts for 99% of all singular values, and the sum of the squares of the first k singular values accounts for 98% of the sum of squares of all singular values. Therefore, we choose a k value that satisfies this condition as the rank parameter k for KNMF. Taking COVID-19 and Zeisel as examples, performing SVD on a representative subset matrix of the COVID-19 dataset, the sum of the squares of the first 23 singular values is 98% of the total sum of squares of singular values. Therefore, for the COVID-19 dataset, the parameter k is 23 when performing kernel nonnegative matrix factorization. After performing singular value decomposition on a representative subset of the Zeisel dataset, the sum of squares of the first 50 singular values accounts for 98.23% of the total sum of squares of singular values. Therefore, the parameter k is 50 in the Zeisel dataset.
[0097] 3. Obtaining the category tags of the remaining subset
[0098] The labels L obtained from clustering in step 2 and the representative subset V obtained in step 1 are used together. s The training set is used to train the system using a Deep Neural Network (DNN) method. The remaining data V is then used... r As a test set, it is classified by a trained neural network classification model.
[0099] A neural network classification model is obtained by training a unidirectional multi-layer feedforward neural network with hidden layers using the backpropagation (BP) algorithm based on stochastic gradient descent (SCG).
[0100] The step size is adjusted using the Levenberg-Marquardt method, assuming it is adjusted according to λ each time. k The sign adjustment a k The second-order term is calculated as follows:
[0101]
[0102] The step size is:
[0103]
[0104] in It is a set of non-zero weight vectors, where v is Rn The weight vector in the gene expression matrix space, E(v) is the global error function, and E′(v) is the gradient of the error function. It is a quadratic approximation of the error function. λ k The sign function is represented by a (if less than or equal to 0, it means the Hessian matrix of the error function is non-positive definite; if greater than 0, it means the Hessian matrix of the error function is positive definite), a k σk represents the step size, k represents the number of iterations, T represents the transpose, and σk represents a parameter greater than 0 and less than or equal to 1.
[0105] Cross-entropy is chosen as the loss function, and Softmax regression is used to transform the forward propagation results of the unidirectional multilayer feedforward neural network into a probability distribution. The output of softmax can be regarded as a probability distribution and used for multi-label classification.
[0106] Using a neural network classification model, the remaining subset V of the test set will be obtained. r Category tag L r .
[0107] 4. Clustering of the overall data V
[0108] Finally, in step 2, the representative subset V... s Cluster labels L obtained from clustering s and the remaining subset V in step 3 r Classify to obtain label L r Merging achieves clustering of the overall data V.
[0109] Example 2: Effectiveness analysis of a method for analyzing cellular heterogeneity based on large-scale single-cell sequencing data.
[0110] In this embodiment of the invention, KNMF-DNN (Kernel Nonnegative Matrix Factorization-Depth Neural Network) is used for cluster analysis on the Zeisel, Marques, and COVID-19 datasets. The clustering evaluation metrics, Normalized Mutual Information (NMI) and Adjusted Rand Index (ARI), are used to evaluate the effectiveness of the method for analyzing cellular heterogeneity based on large-scale single-cell sequencing data established in this invention.
[0111] This embodiment uses three comparisons to illustrate the advantages of KNMF-DNN.
[0112] 1. Comparative Analysis of KNMF and Other Clustering Algorithms
[0113] First, the KNMF clustering method, along with four other classic clustering algorithms, is used to analyze the representative subset V selected in step 1 of Example 1.s The four comparative clustering algorithms are: single-cell consensus clustering (SC3) with high accuracy, k-means clustering algorithm (K-means) which is the most classic and widely used clustering algorithm, hierarchical clustering (HC), and single-cell interpretation via multi-kernel learning (SIMLR) which is a stable multi-kernel learning framework.
[0114] Figure 3 The benchmark test results of the KNMF method used in the embodiments of the present application and the other five clustering methods on the above three data sets are shown. For the covid-19 data set, by calculating the KL divergence at different sampling rates, it is found that when the sampling rate is 40%, the KL divergence is the smallest. 3200 cells are selected to form a representative set using stratified sampling, and clustering is performed on the representative set. When clustering by the KNMF clustering method, the parameter rank k needs to be determined. The present application performs singular value decomposition on the original data, and when the square sum of the first singular value accounts for 98% of the square sum of all singular values, i.e., 23, the parameter rank k is set. The consistency parameter is set to 10 (the parameter m when performing multiple kernel non-negative matrix factorization on the gene expression matrix of the representative set and calculating the consensus matrix is the consistency parameter). The NMI score of the KNMF clustering is 0.87 (covid-19 in the left middle graph), and the ARI value is 0.82 (covid-19 in the right middle graph), which is higher than that of the other five clustering algorithms. SC3 ranks second with an NMI value of 0.76 (covid-19 in the left middle graph) and an ARI value of 0.77 (covid-19 in the right middle graph). The results of these two algorithms are obviously higher than those of the remaining four algorithms ( Figure 3 ). Figure 3 Figure 3 Figure 3 Similarly, for the zeisel data set, the representative set selected by stratified sampling has a size of 1500 cells. The parameter k is set in the same way as above, i.e., 50. The consistency parameter m is set to 10. The NMI score of the KNMF clustering of the present application is 0.80 (zeisel in the left middle graph), and the ARI value is 0.84 (zeisel in the right middle graph), which is higher than that of the other five clustering algorithms. SC3 clustering ranks second with an NMI value of 0.74 (zeisel in the left middle graph) and an ARI value of 0.77 (zeisel in the right middle graph). Figure 3
[0115] Figure 3 Figure 3 Figure 3
[0116] For the Marques dataset, 2500 cells were selected as the representative subset, the parameter k was set to 19, and the consistency parameter m was set to 10. The NMI score of KNMF clustering was 0.46 Figure 3 Marques in the middle left), and the ARI value was 0.29 Figure 4 Marques in the middle right), which was better than other clustering algorithms. It can be seen that the KNMF algorithm used in the present application is better than other clustering algorithms, which will ensure that the label obtained by clustering is more accurate, which will help the next step of training the neural network, and thus improve the classification effect.
[0117] 2. Analysis of comparison results of classification methods
[0118] Figure 4 The deep neural network (DNN) used for the method of the present application was compared with the classification effects of three classification methods (k-nearest neighbor, support vector machine, and random forest). For the clustering labels obtained by the KNMF method, the present application used DNN, k-nearest neighbor, support vector machine, and random forest to classify the remaining subset relative to the representative subset, and obtained the labels of the entire dataset.
[0119] In the covid-19 dataset, 3200 cells were used to predict the labels of the remaining 4804 cells. In deep neural network classification, the number of hidden layer neurons was set to 23, and the NMI and ARI values of the entire framework clustering were 0.79 Figure 4 represented by DNN of covid-19 in the middle left) and 0.72 Figure 4 represented by DNN of covid-19 in the middle right).
[0120] For the zeisel dataset, 1500 cells were used to predict the labels of the remaining 1505 cells. In deep neural network classification, the number of hidden layer neurons was set to 12, and the NMI and ARI values of the entire framework clustering were 0.81 Figure 4 represented by DNN of zeisel in the middle left) and 0.85 Figure 4 represented by DNN of zeisel in the middle right).
[0121] For the Marques dataset, 2500 cells were used to predict the labels of the remaining 2553 cells. In deep neural network classification, the number of hidden layer neurons was set to 40, and the NMI and ARI values of the entire framework clustering were 0.45 Figure 4 represented by DNN of Marques in the middle left) and 0.30 Figure 4 represented by DNN of Marques in the middle right). From Figure 5It can be clearly seen that the classification result of the DNN is more accurate than the other three classification algorithms.
[0122] 3. Comparison results of the KNMF-DNN method and other clustering methods on the entire large-scale data set
[0123] Figure 5 For the comparison results of the KNMF-DNN method and other clustering methods on the entire large-scale data set, the comparison methods are SC3, bigscale (related literature: Iacono G, Mereu E, Guillaumet-Adkins A, Corominas R, Cuscó I, Rodriguez-Esteban G, Gut M, Perez-Jurado LA, Gut I, Heyn H. bigSCale: an analytical framework for big-scale single-cell data. Genome Res. 2018 Jun; 28 (6): 878-890. doi: 10.1101 / gr.230771.117. Epub 2018 May 3), SIMLR, hierarchical clustering (HC), and KNMF method used on the entire data set. From Figure 5 it can be seen that the KNMF-DNN is obviously better than the other comparison methods (left and right in Figure 5 ). SC3 has too high computational complexity and runs out of memory when processing large-scale data such as the covid-19 data set and the Marques data set, and cannot obtain results. Directly performing KNMF clustering on the covid-19 data set also has too high complexity and cannot be calculated. For the Zeisel data set, the NMI value (left in Figure 5 zeisel in Figure 6 ) and the ARI value (right in zeisel in
[0124] ) of the KNMF clustering and SC3 are similar, ranking second. The other three clustering methods have poor accuracy on the three data sets. Therefore, it can be seen that the method framework for analyzing cell heterogeneity based on large-scale single-cell sequencing data established by the present application is more suitable for large-scale single-cell sequencing data than traditional methods, and greatly enhances the cell type recognition ability of the sequencing data.
[0124] 4. Result analysis of KL divergence determining the sampling rate
[0125] The KL divergence between the distribution of the screened sample data and the distribution of the remaining data can determine the optimal sampling rate, and the number of representative subsets obtained by comparing different divisions can verify the final results of the framework: taking the covid-19 dataset as an example, ensure that the steps and parameters of the framework remain unchanged, only the division ratio is different, and the KNMFDNN method is used to cluster the data for different representative subsets.
[0126] The KL divergence curves of the covid-19 dataset under different sampling rates and the clustering results of the KNMF-DNN framework under different sampling rates. From the results in the figure, it can be seen that when the sampling rate is 40%, the KL divergence between the distribution of the representative subset and the distribution of the remaining subset is the smallest. For the method of the present application, only the sampling rate is changed, and the rest of the process and parameters remain unchanged, and finally it is also confirmed that when the sampling rate is 40%, the overall clustering result is optimal. This means that it is feasible to use the KL divergence to determine the optimal proportion of the representative subset.
[0127] When using the KNMF method for clustering, the present application considers whether different k values will affect the results. For the same sample screened under the optimal proportion, the present application analyzes the variance under similar k values: taking the covid-19 dataset as an example, the p value of NMI is 0.0305, and the p value of ARI is 0.0315, both of which are less than 0.05. Therefore, for the same sample, the selection of k has a statistically significant effect on the results.
[0128] In addition, the present application also considers whether different samples screened under the optimal proportion will affect the results. Similarly, taking the covid-19 dataset as an example, five representative sample sets are selected at a proportion of 40%. Set the parameter k to 23 and analyze the variance of the results. Take α = 0.05, the p value of NMI is 0.0305, and the p value of ARI is 0.0315. Both are greater than 0.05. Therefore, under the optimal proportion, different samples have no significant statistical impact on the results.
[0129] The above describes the present application in detail. For those skilled in the art, without departing from the purpose and scope of the present application, and without unnecessary experiments, the present application can be implemented in a wider range under equivalent parameters, concentrations and conditions. Although the present application gives a special example, it should be understood that further improvements can be made to the present application. In summary, according to the principle of the present application, this application intends to include any changes, uses or improvements of the present application, including changes made by conventional techniques known in the art, which deviate from the scope disclosed in the present application.
Claims
1. A method of clustering and / or cell heterogeneity analysis of single cell sequencing data, characterized in that: The method comprises the following steps: A1. Representative subset obtaining: obtaining a representative subset from a single-cell sequencing dataset using stratified sampling dividing the single-cell sequencing dataset into the representative subset and a remaining subset two parts; A2, clustering of the representative subset: clustering the representative subset using a kernel non-negative matrix factorization method to obtain clustering labels of the representative subset; A3, obtaining of classification labels of the remaining subset: using the clustering labels of the representative subset and the representative subset as a training set, training using a deep neural network method to obtain a neural network classification model, and using the neural network classification model to obtain the classification labels of the remaining subset; A4, output of overall data clustering results: merging the clustering labels of the representative subset and the classification labels of the remaining subset to obtain clustering results and / or cell heterogeneity analysis results of the single-cell sequencing data; A1, obtaining of the representative subset comprises the following steps: A1-1) data dimension reduction: using a principal component analysis method to reduce the dimension of the single-cell sequencing data set, then using a farthest point sampling method to obtain k representative cell sample center points, and dividing the single-cell sequencing data set based on the representative cell sample center points to obtain k sample modules; A1-2) obtaining of the representative subset: using a module optimization algorithm apri cot to select representative data of the sample modules from the sample modules, and merging the representative data of each sample module to obtain the representative subset; A1-3) Representative subset and remaining subset ratio determination: using principal component analysis method to reduce dimensionality of the representative subset and the remaining subset of the single cell sequencing data set from which the representative subset is removed, obtaining the reduced dimensionality representative subset and the reduced dimensionality remaining subset, calculating the KL divergence of the sample distribution of the reduced dimensionality representative subset and the sample distribution of the reduced dimensionality remaining subset, determining the proportion of the representative subset in the single cell sequencing data set according to the value of the KL divergence, dividing the single cell sequencing data set into the representative subset and the remaining subset .
2. The method of claim 1, wherein: A2, clustering of the representative subset comprises the following steps: constructing a non-linear mapping of the representative subset to obtain a kernel matrix induced after mapping, performing non-negative matrix factorization on the kernel matrix to obtain a coefficient matrix and a basis matrix, performing hierarchical clustering on the coefficient matrix to obtain clustering labels, calculating a similarity matrix of the clustering labels to obtain a consistency similarity matrix, and performing clustering on the consistency similarity matrix to obtain clustering labels of the representative subset.
3. The method according to claim 1 or 2, characterized in that: A3, the neural network classification model is obtained by a method comprising the following steps: using the clustering labels of the representative subset and the representative subset as a training set, using a one-way multi-layer feedforward neural network, adopting a cross-entropy loss function, and training the training set based on a backpropagation algorithm of stochastic gradient descent to obtain the neural network classification model.
4. An apparatus for single-cell sequencing data clustering and / or cell heterogeneity analysis, characterized in that: The device comprises the following modules: B1, a representative subset obtaining module: configured to obtain a representative subset from the single-cell sequencing dataset using a hierarchical sampling method dividing the single-cell sequencing dataset into the representative subset and a remaining subset two parts; B2, clustering module of the representative subset: configured to cluster the representative subset using a kernel non-negative matrix factorization method to obtain clustering labels of the representative subset; B3, obtaining module of classification labels of the remaining subset: configured to use the clustering labels of the representative subset and the representative subset as a training set, train using a deep neural network method to obtain a neural network classification model, and use the neural network classification model to obtain the classification labels of the remaining subset; B4, overall data clustering result output module: configured to merge the clustering labels of the representative subset and the classification labels of the remaining subset to obtain clustering results and / or cell heterogeneity analysis results of the single-cell sequencing data; B1, the representative subset obtaining module comprises the following modules: B1-1) data dimension reduction module: using principal component analysis method to reduce dimension of the single-cell sequencing data set, then using furthest point sampling method to obtain k representative cell sample center points, and dividing the single-cell sequencing data set based on the representative cell sample center points to obtain k sample modules; B1-2) representative subset obtaining module: using module optimization algorithm apri cot to select representative data of the sample module from the sample module, and combining the representative data of each sample module to obtain the representative subset; B1-3) representative subset and remaining subset ratio determination module: using principal component analysis method to reduce dimension of the representative subset and the remaining subset of the single-cell sequencing data set after removing the representative subset, obtaining the reduced dimension representative subset and the reduced dimension remaining subset, calculating the KL divergence of the sample distribution of the reduced dimension representative subset and the sample distribution of the reduced dimension remaining subset, determining the proportion of the representative subset in the single-cell sequencing data set according to the value of the KL divergence, and dividing the single-cell sequencing data set into representative subset and remaining subset.
5. The apparatus of claim 4, wherein: B2 the clustering module of the representative subset is realized by a method comprising the following steps: constructing a non-linear mapping of the representative subset, obtaining a kernel matrix induced after mapping, performing non-negative matrix factorization on the kernel matrix to obtain a coefficient matrix and a basis matrix, performing hierarchical clustering on the coefficient matrix to obtain clustering labels, calculating the similarity matrix of the clustering labels to obtain a consistency similarity matrix, and performing clustering on the consistency similarity matrix to obtain the clustering labels of the representative subset.
6. The apparatus of claim 4 or 5, wherein: B3 the neural network classification model is obtained by a method comprising the following steps: using the clustering labels of the representative subset and the representative subset as a training set, using a one-way multi-layer feedforward neural network, adopting a cross-entropy loss function, and training the training set based on a backpropagation algorithm of stochastic gradient descent to obtain the neural network classification model.
7. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program enables a computer to perform the steps of the method according to any one of claims 1-3.
8. Any one of the following applications of the method according to any one of claims 1-3 and / or the device according to any one of claims 4-6 and / or the computer readable storage medium according to claim 7: D1, application in preparing a product for identifying cell types of sequencing data; D2, application in preparing a product for predicting tumor cell heterogeneity; D3, application in preparing a product for developing tumor markers.
Citation Information
Patent Citations
Big data clustering method based on decomposition and composition
CN104063518A
Kernel nonnegative matrix factorization-based dictionary learnt and sparse feature expressed face recognition method and system
CN106897685A
Cited By
Tumor immune microenvironment single-cell transcriptome analysis and cell communication analysis system
CN122551915A