A method for analyzing disease-specific cells in single-cell transcriptomes based on iterative EM clustering
Through the combination of iterative EM clustering and scoring algorithms, the problem of single-cell transcriptome data analysis in non-cancer diseases is solved, and the precise identification and distinction of disease-specific cells is achieved, which is suitable for single-cell transcriptome data analysis under various disease conditions.
Patent Information
- Application Number
- CN202310083812.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-02-08
AI Technical Summary
The prior art is difficult to identify and analyze small differential changes in single-cell transcriptome data in non-cancer diseases, find out cell types with significant changes, and effectively distinguish disease-specific cells.
A single-cell transcriptome disease-specific cell analysis method based on iterative EM clustering is proposed. By obtaining single-cell transcriptome data, using the clustering method and marker gene classification of graphs, combining scoring algorithms and iterative EM clustering algorithms, we identify and distinguish disease-specific cells.
The analysis of single-cell transcriptome data sets under all disease conditions is achieved, and it does not rely on relevant databases. It can identify slight differential changes in the cell population and accurately distinguish disease-specific cells, improving the accuracy and reliability of the analysis.
Smart Images

Figure CN116486920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for analyzing single-cell transcriptome data, and in particular to a single-cell transcriptome disease-specific cell analysis method based on iterative EM clustering. Background Art
[0002] Single-cell transcriptome sequencing (scRNA-seq) can be used to identify the unique transcriptome characteristics of each cell. Using this information, previously unanswerable questions can be answered, including identifying new cell types, simulating the dynamic changes of cells during development, and identifying different gene regulatory mechanisms between cell subtypes. In recent years, many disease research institutions have used scRNA-seq data to study diseases in the hope of more accurately characterizing the causes of disease at the cellular and molecular levels.
[0003] Most existing methods analyze patient scRNA-seq data by studying the changes in the rough overall aspects. For example, by identifying scRNA-seq data, we can explore whether the proportion of each type of cell in the patient tissue is different from that of healthy people, or further explore whether the proportion of cell subpopulations in each major cell type changes in patients. In exploring the changes of various types of cells under disease conditions, previous methods were mostly affected by traditional transcriptome (RNA-seq) analysis methods, and were mostly used to directly compare the differences between the patient group and the normal group. However, this method of comparing differences gradually compares certain cell groups or certain cell subpopulations due to the existence of scRNA-seq data characteristics. Obviously, this method of sticking to the overall difference comparison often loses a lot of information, and even adds a lot of noise to interfere with the analysis. Usually in the patient group cells, not all cells are cells with significant differential expression under the disease, referred to as disease-specific cells. Usually, there are also some inert cells in them, and these cells do not have significant differential expression under a certain disease condition. To this end, a method is needed to find such disease-specific cells and combine them with the relevant theories of single-cell transcriptomes for targeted analysis. The latest research has used iterative KMeans clustering to gradually find disease-specific cells in epithelial cells of gastric cancer patients. This method first uses the RNA-seq data in the TCGA gastric adenocarcinoma dataset to make differential genes, and obtains the malignant and non-malignant scores of the corresponding cells through the AddModuleScore scoring function of the Seurat package, and then iterates KMeans to obtain the final malignant cells and non-malignant cells under disease conditions, successfully conducting a detailed analysis of gastric cancer. However, cancer is a special disease with a related TCGA database, and cancer cells have not only many differentially expressed genes compared to normal cells, but also large differences in gene expression, so it is easy to divide cells into two categories through the above iterative algorithm.
[0004] However, not all diseases have a mature RNA-Seq database as a reference like cancer. Moreover, the pathogenic cell types of other diseases or the cell types with specific changes are not as clear as those of cancer. At the same time, the specific changes that occur in disease-specific cells may not be as strong as those in cancer cells. Therefore, it is of great significance to propose an algorithm that can be applied to scRNA-seq datasets under all disease conditions, does not rely on relevant databases, can identify subtle differential changes in cell populations, find cell types with significant changes, and can well distinguish disease-specific cells in such cell populations. Summary of the Invention
[0005] The present invention aims to propose a method for analyzing disease-specific cells in single-cell transcriptomes based on iterative EM clustering, which can study scRNA-seq datasets under all disease conditions, does not rely on relevant databases, can identify subtle differential changes in cell populations, find cell types with significant changes, and can well distinguish disease-specific cells in such cell populations.
[0006] The method for analyzing disease-specific cells in single-cell transcriptomes based on iterative EM clustering in the present invention includes the following steps:
[0007] Step 1: Obtain single-cell transcriptome datasets of healthy individuals and patients under a certain disease condition, use graph clustering methods to divide the single-cell transcriptome datasets into multiple clusters, and classify each cluster through marker genes to obtain cell populations of each type of cell;
[0008] Step 2: For a certain cell population of interest, obtain the up-regulated and down-regulated genes of this cell population under the current disease condition;
[0009] Step 3: According to the obtained up-regulated and down-regulated genes, use a scoring algorithm to assign an initial up-regulation score and an initial down-regulation score to each cell in the cell population of interest, and project the cells into a two-dimensional score space according to the scoring results;
[0010] Step 4: Cluster the cells in the score space into two categories using the EM clustering algorithm;
[0011] Step 5: Arbitrarily select one of the two clusters after clustering, obtain its up-regulated and down-regulated genes relative to the other cluster, and use the scoring algorithm to re-assign an up-regulation score and a down-regulation score to each cell in this cell population, and then project the results into a two-dimensional score space;
[0012] Step 6: Repeat Step 4 and Step 5 until the clustering results of two consecutive times are the same or reach the set number of iterations, and output the two types of cells after iterative clustering;
[0013] Step 7: Calculate the initial disease score set of cells according to the results of the initial up - regulation score and the initial down - regulation score as follows:
[0014]
[0015] Among them, is the initial disease score set of cells, is the initial up - regulation score set of cells, is the initial down - regulation score set of each cell;
[0016] Then, according to the clustering result after iteration, calculate the class score set of cells after iteration as follows:
[0017]
[0018] Among them, Score last is the class score set obtained based on a certain class of cells after iterative clustering, is the final up - regulation score set obtained according to the result of the last iteration, is the final down - regulation score set obtained according to the result of the last iteration;
[0019] Subsequently, re - sort in ascending order according to the initial disease scores of the cells therein, and then arrange the class scores of each cell in Score last in the order corresponding to the cells after rearrangement of , and calculate the Pearson correlation coefficient r between the rearranged and Score last ;
[0020] Finally, classify the clustered cells into disease - specific cells and control group normal cells according to one of the following strategies:
[0021] ① If r > 0, it means that one class of cells selected in Step 5 is disease - specific cells, and the other class is control group normal cells. If r < 0, the classification is opposite;
[0022] ② If the rearranged Score last generally shows an upward trend as the same as that after rearrangement of , it means that one class of cells selected in Step 5 is disease - specific cells, and the other class is control group normal cells. Otherwise, the classification is opposite.
[0023] Furthermore, Step 7 also includes judging whether the change result of the score after iteration is reasonable according to the absolute value of the calculated Pearson correlation coefficient r. If |r| > 0.6, it means the result is reasonable, otherwise it means the result is unreasonable;
[0024] If the change in the score after iteration is unreasonable, it is determined that the cell population has no significant change in the disease environment;
[0025] If the change is reasonable, the two types of cells output are classified as disease-specific cells and normal cells in the control group.
[0026] Furthermore, it also includes step 8, repeating steps 2 to 7 to analyze a new group of cells until all cell populations of interest are analyzed.
[0027] Furthermore, the specific steps of the scoring algorithm in step 3 are as follows:
[0028] First, randomly select M genes different from the feature genes to be scored as reference genes;
[0029] Then, use the Kmeans clustering algorithm to divide all cells into K categories according to the total expression of all genes in each cell; next, calculate the average value of the reference gene expression of cells in each category, and the calculation method is as follows:
[0030]
[0031] In the formula, is the average value of the reference gene expression of cells in the kth category, N k is the total number of cells in the kth category, M represents the number of selected reference genes, and expre ij represents the expression of the jth reference gene of the ith cell in the kth category;
[0032] Finally, divide the average value of the up-regulated or down-regulated gene expression in each cell by the average value of the reference gene expression of cells in this category to obtain the original up-regulated or down-regulated score of this cell, and perform maximum-minimum normalization on the original up-regulated or down-regulated scores of all cells to obtain the up-regulated or down-regulated score of this cell.
[0033] Furthermore, the calculation method of the up-regulated score of each cell is as follows:
[0034]
[0035] where cell i ∈Cell k
[0036] Let InitialScore up =[initialupscore1, initialupscore2, ……, initialupscore N
[0037]
[0038] Wherein is the average value of the up-regulated gene expression levels of the i-th cell, upexpre ij is the expression level of the j-th up-regulated gene of the i-th cell, M up represents the total number of up-regulated genes calculated by the differential gene expression algorithm, initialupscore i represents the initial up-score of the i-th cell, cell i represents the i-th cell, Cell k represents the k-th type of cells separated by the above Kmeans clustering algorithm, which is a cell set, InitialScore up represents the set of the original up-scores of all cells, N is the total number of cells, upscore i is the up-score of cell i after max-min normalization;
[0039] Similarly, the calculation formula for the down-score of each cell is as follows:
[0040]
[0041] where cell i ∈Cell k
[0042] Let InitialScore down =[initialdownscore1, initialdownscore2, ……, initialdownscore N
[0043]
[0044] Wherein is the average value of the down-regulated gene expression levels of the i-th cell, downexpre ij is the expression level of the j-th down-regulated gene of the i-th cell, M down represents the total number of down-regulated genes calculated by the differential gene expression algorithm, initialdownscore i represents the original down-score of cell i, InitialScore down represents the set of the original down-scores of all cells, downscore i is the down-score of cell i after max-min normalization;
[0045] Finally, through sorting, the up-score set and down-score set of cells are obtained:
[0046] Score up = [upscore1, upscore2, ……, upscore N
[0047] Score down = [downscore1, downscore2, ……, downscore N
[0048] In the formula, Score up is the up - regulation score set, and Score down is the down - regulation score set.
[0049] Furthermore, in step 2, the Wilcoxon rank - sum test is used to obtain up - regulated and down - regulated genes.
[0050] The beneficial effects of the present invention are as follows: A set of analysis methods applicable to scRNA - seq data sets under all disease conditions are proposed, and disease - specific cells therein can be found. The single - cell transcriptome disease - specific cell classification task is realized. By performing differential expression analysis and scoring and dimensionality reduction on the single - cell transcriptome data of the disease group and the control group, and then using the iterative EM clustering algorithm to classify cells, the differences in the disease group are projected onto each cell for subsequent analysis at a higher resolution. Since the Gaussian mixture model in the EM clustering algorithm used has good clustering effect on elliptical distribution data, it is more suitable for the data distribution in the score space, making the clustering results more accurate and reasonable.
[0051] In some embodiments, classifying and scoring according to the total cell expression using the improved scoring algorithm can avoid more noise interference in the case of fewer differential genes;
[0052] In some embodiments, the rationality judgment method after iteration judges whether the cells change significantly under disease conditions from the correlation of cell scores before and after iteration, making the judgment result more persuasive.
[0053] Compared with the methods proposed in the prior art, the method in the present invention is more perfect in structure and applicable to scRNA - seq data under various disease conditions, providing a new modular idea for the analysis of this type of data set. Brief Description of the Drawings
[0054] Figure 1 is the flowchart of the single - cell transcriptome disease - specific cell analysis method based on iterative EM clustering in the embodiments of the present invention.
[0055] Figure 2 It is a schematic diagram of the first distribution of the iterative clustering results in the score space in the embodiments of the present invention.
[0056] Figure 3 It is a schematic diagram of the second distribution of the iterative clustering results in the score space in the embodiments of the present invention. Detailed implementation manners
[0057] In order to make the implementation steps and advantages of the present invention clearer, the following will describe the specific implementation manners in combination with the legends.
[0058] The process framework of the single-cell transcriptome disease-specific cell analysis method based on iterative EM clustering in this embodiment is basically as Figure 1 shown, and the specific implementation steps are as follows:
[0059] After obtaining the scRNA-seq of patients and healthy people, it is first necessary to cluster them using the graph clustering method, and then classify them using the marker genes of related cells. In some embodiments, the graph clustering method usually consists of two steps: drawing a graph and identifying the graph. Drawing a graph usually consists of two steps: k-nearest neighbor (KNN) and shared nearest neighbor (SNN). This graph can be regarded as the result of calculating the similarity between cells. Identifying the graph is usually implemented by the community detection algorithm - Louvain algorithm, that is, selecting the group of cells with the highest similarity in the graph as a cell group. After the cell groups are divided, the marker genes of related cells in relevant databases such as CellMarker and PanglaoDB are used to classify the cells. Then, for the cells of interest, further subsequent analysis is carried out.
[0060] Each of the following steps is to perform a separate analysis on one of the selected cells of interest. First, it is necessary to obtain the differentially expressed genes between the diseased group and the healthy group in the analyzed cell group. Differentially expressed genes refer to some genes in a group of cells whose expression levels are significantly different statistically from those of another group of cells. Differentially expressed genes can be further divided into up-regulated and down-regulated genes, that is, significantly increased or decreased statistically. In some embodiments, the Wilcoxon rank sum test method is used to obtain the differentially expressed genes between the diseased group and the healthy group in the analyzed cell group. The Wilcoxon rank sum test is usually used to infer whether there is a difference in a certain index between two population distributions. Because of its fast speed and suitability for large sample datasets, it is mostly used for obtaining differentially expressed genes in scRNA-Seq data.
[0061] After obtaining the differentially expressed genes, according to the up-regulated genes and down-regulated genes of the diseased group relative to the healthy group among the differentially expressed genes, the up-score and down-score of each cell in the analyzed cell population need to be obtained using the scoring algorithm proposed in this embodiment. When scoring cells, the relative expression levels of the characteristic genes of each cell need to be considered, because the absolute expression levels of the obtained data may be interfered by many factors, such as different cell development cycles, or the capture differences of different cells by the instrument. The scoring algorithm proposed in this embodiment will avoid the interference of these factors to a certain extent, and the scoring algorithm is as follows:
[0062] First, randomly select M genes that are different from the characteristic genes to be scored as reference genes. Then, use the Kmeans clustering algorithm to divide all cells into K categories according to the total expression level of each cell. Next, calculate the average value of the reference gene expression levels of the cells in each category. The calculation method is as follows:
[0063]
[0064] In the formula, is the average value of the reference gene expression levels of the cells in the k-th category, N k is the total number of cells in the k-th category, M represents the number of selected reference genes, and expre ij represents the expression level of the j-th reference gene of the i-th cell in the k-th category.
[0065] Finally, divide the average value of the characteristic gene expression level in each cell by the average value of the reference gene expression levels of the cells in its category to obtain the original score of the cell. Normalize the original scores of all cells using the maximum-minimum normalization method to obtain the score of the cell. The score calculated using the up-regulated genes is the up-score, and the score calculated using the down-regulated genes is the down-score.
[0066] In this embodiment, the specific calculation formula for the up-score of each cell is as follows:
[0067]
[0068] where cell i ∈Cell k
[0069] Let InitialScore up =[initialupscore1, initialupscore2, ……, initialupscore N
[0070]
[0071] In the formula, is the average value of the up-regulated gene expression levels of the i-th cell, upexpre ij is the expression level of the j-th up-regulated gene in the i-th cell, M up represents the total number of up-regulated genes calculated by the differential gene expression algorithm, initialupscore i represents the initial up-score of the i-th cell, cell i represents the i-th cell, Cell k represents the k-th type of cells separated by the above Kmeans clustering algorithm, which is a set of cells, InitialScore up represents the set of the original up-scores of all cells, N is the total number of cells, upscore i is the up-score of cell i after max-min normalization;
[0072] Similarly, the calculation formula for the down-score of each cell is as follows:
[0073]
[0074] where cell i ∈Cell k
[0075] Let InitialScore down =[initialdownscore1, initialdownscore2, ……, initialdownscore N
[0076]
[0077] In the formula, is the average value of the down-regulated gene expression levels of the i-th cell, downexpre ij is the expression level of the j-th down-regulated gene in the i-th cell, M down represents the total number of down-regulated genes calculated by the differential gene expression algorithm, initialdownscore i represents the original down-score of cell i, InitialScore down represents the set of the original down-scores of all cells, downscore i is the down-score of cell i after max-min normalization.
[0078] Finally, after sorting, the up-score sets and down-score sets of each cell are obtained:
[0079] Score up = [upscore1, upscore2, ……, upscore N
[0080] Score down = [downscore1, downscore2, ……, downscore N ,
[0081] where Score up is the up - score set, and Score down is the down - score set.
[0082] After scoring, each cell needs to be projected into the corresponding two - dimensional score space according to the obtained up - scores and down - scores.
[0083] Subsequently, the iterative EM clustering algorithm is used to cluster the cells in the score space, and the results and data of the first clustering are saved. The iterative EM clustering algorithm includes two steps: clustering using the EM algorithm and re - obtaining the differentially expressed genes of the two classes after clustering and projecting the clustering results into the score space for the next use of the EM algorithm for clustering. Among them, any one of the classes of cells after clustering is selected, and its up - regulated and down - regulated genes relative to the other class are obtained, and each cell in this cell population is re - scored with the up - score and down - score using the scoring algorithm, and then the results are put into the two - dimensional score space; these two steps are continuously repeated during the iteration until the termination condition of the iteration is met, that is, the clustering results of the two consecutive times are the same, or the predetermined number of termination times is reached. According to the experimental process, after iterating to a certain extent, there may always be a very small number of cells that keep fluctuating between the categories of the two consecutive iterations. Since the number of these cells in the intermediate form is extremely small, it will not affect the overall iterative convergence. Therefore, after setting a suitable number of termination times, the iterative optimization of cell classification has been generally completed.
[0084] Then, it is necessary to determine whether there are significant changes in this type of cell population under this disease. If there are no obvious changes, it means that this type of cell population has no significance for analysis under this disease condition. If there are significant changes, it means that the role played by this type of cell population under this disease can be further analyzed, etc. The method for determining whether there are significant changes in this type of cell population under this disease is to judge whether the change between the first clustering result and the last iterative clustering result is reasonable. Because the first clustering result is obtained from the differentially expressed genes between the cells of patients and healthy people, it has strong reference value. If the result after iterative clustering is very different from the result of the first clustering, it indicates that the differential expression of genes in this cell between patients and healthy people is not significant, resulting in a large change in the clustering result during the subsequent iterative process because this difference is not utilized. Since the change in the clustering result after iteration is ultimately reflected in the scores calculated from the differentially expressed genes of the two categories before and after iteration, in this embodiment, the initial disease score and the category score after iteration are defined, and the Pearson correlation coefficient r between the two scores is obtained to judge the reasonableness of the iterative change. Statistical principles indicate that when |r| > 0.6, the two data have strong correlation. Therefore, 0.6 is set as the threshold to judge whether the clustering change before and after iteration is reasonable. In this example, the generation process of the Pearson correlation coefficient is as follows:
[0085] First, calculate the initial disease score set of cells according to the results of the initial up-regulation score and the initial down-regulation score as follows:
[0086]
[0087] Among them, is the initial disease score set of cells, is the initial up-regulation score set of cells, is the initial down-regulation score set of cells;
[0088] And so on, the category score set after each iteration is expressed as:
[0089]
[0090] Among them, n represents the number of iterations, Score n represents the category score set obtained based on a certain category after the nth iterative clustering. Similarly, is the up-regulation score set after the nth iteration, is the down-regulation score set after the nth iteration;
[0091] Then, according to the clustering result after the iteration is completed, calculate the category score set of each cell after the iteration is completed as follows:
[0092]
[0093] Among them, Score last is a set of class scores obtained based on a certain class of cells after the last iterative clustering, is the final up-regulated score set obtained according to the results of the last iteration, is the final down-regulated score set obtained according to the results of the last iteration;
[0094] Subsequently, is re-sorted in ascending order of the initial disease scores of each cell therein, and then the class scores of each cell in Score last are arranged according to the cell order corresponding to the re-arranged , and the Pearson correlation coefficient r between the re-arranged and Score last is calculated. Among them, and Score last are approximately normally distributed after testing, so they are directly set to be normally distributed in the calculation;
[0095] Finally, the cells in the cell population with significant changes need to be classified into disease-specific cells and control group normal cells. There are two classification strategies. One is to judge according to the Pearson correlation coefficient r obtained in the previous step. If r > 0, it means that the class selected in the last iteration for calculating differentially expressed genes is the same as the initially selected class (patient group), then the cells of this class are disease-specific cells, and the cells of the other class are control group normal cells. Otherwise, the classification is the opposite. The other is to sort the initial disease scores, and the class scores after iteration are sorted according to the order of the initial scores. If the overall change order of the class scores after iteration (decreasing from large to small or increasing from small to large, which can be visually seen by drawing a graph) is the same as the change order of the initial disease scores, it means that the class selected in the last iteration for calculating differentially expressed genes is the same as the initially selected class (patient group), then the cells of this class are disease-specific cells, and the cells of the other class are control group normal cells. Otherwise, the classification is the opposite.
[0096] After classification, the cell population can be further divided into subpopulations, and then combined with the disease-specific cells and control group normal cells classified in the previous step for further bioinformatics research. After the analysis of this cell population is completed, the above several steps can be repeated to analyze a new cell population of interest.
[0097] In this embodiment, computational example demonstrations were also carried out based on the single-nucleus transcriptome (snRNA-seq) dataset of the brains of patients with autism spectrum disorder collected by D Velmeshev et al. (Source: Velmeshev D, Schirmer L, Jung D, et al. Single-cell genomics identifies cell type–specific molecular changes in autism[J]. Science, 2019, 364(6441):685-689.). SnRNA-seq is very similar to scRNA-seq and is mainly used for sequencing single cells in brain tissue. This dataset has divided the original snRNA-Seq dataset into 17 cell clusters, which are fibrous astrocytes (AST-FB), protoplasmic astrocytes (AST-PP), oligodendrocyte precursor cells (OPC), Parvalbumin interneurons (IN-PV), somatostatin interneurons (IN-SST), SV2C interneurons (IN-SV2C), VIP interneurons (IN-VIP), layer 2 / 3 excitatory neurons (L2 / 3), layer 4 excitatory neurons (L4), layer 5 / 6 cortical projection neurons (L5 / 6), layer 5 / 6 corticocortical projection neurons (L5 / 6-CC), mature neurons (Neu-mat), NRGN-expressing neurons 1 (Neu-NRGN-I), NRGN-expressing neurons 2 (Neu-NRGN-II), microglia (Microglia), oligodendrocytes (Oligodendrocytes), and endothelial cells (Endothelial). This dataset consists of two parts: the prefrontal cortex (PFC) and the anterior cingulate cortex (ACC). In this example, the data in the ACC was used for experiments.
[0098] After being processed by the above single-cell transcriptome disease-specific cell analysis method based on iterative EM clustering, the Pearson correlation coefficients calculated for each cell cluster are shown in Table 1 below:
[0099] Table 1 Pearson correlation coefficients of each cell after processing
[0100]
[0101]
[0102] As can be seen from the above table, a total of 13 types of cells were calculated to possibly have differences and could be divided into two categories.
[0103] In summary, the distribution of these 13 types of cells after iterative clustering can be roughly divided into three categories. The first two categories are relatively special, and the third category is a mixture of the first two categories.
[0104] The first distribution is as shown in Figure 2 . Here, the vertical axis represents the gene expression upregulation score (Score.UP) of patients relative to healthy individuals, and the horizontal axis represents the gene expression downregulation score (Score.DOWN). All the cells in the cell population can be roughly divided into two categories. The first category is disease-specific cells, and the other category is normal cells in the control group. Obviously, disease-specific cells have a relatively high Score.UP and a relatively low Score.DOWN, and are located in the Figure 2 upper left corner, while normal cells in the control group are just the opposite, located in the Figure 2 lower right corner. As can be seen from Figure 2 , disease-specific cells are mainly composed of patient cells, with only a very small number of healthy individual cells, which may be due to noise. Normal cells in the control group are composed of both patient cells and healthy individual cells. Therefore, it can be speculated that there may be a new subset of cells in the patient's body, and this subset of cells only exists in patients and not in normal healthy individuals. At the same time, it can also indicate that the changes of these cells in patients are relatively large. In this example, the final iterative distribution of some subsets is roughly like this first type of distribution, such as Microglia and Neu-mat.
[0105] Another type of distribution is as shown in Figure 3 . Disease-specific cells are mainly composed of patient cells, but are also accompanied by many healthy individual cells, while normal cells in the control group are mainly composed of healthy individual cells. Therefore, it can be speculated that there are mainly two forms of these cells in normal individuals, one is the activated state, and the other is the resting state. When the patient is ill, the resting state cells may be transformed into the activated state due to environmental changes, and the extreme case is to be completely transformed into the activated state. This change is relatively smaller compared to the change in the first type. In this example, the final iterative distribution of some subsets is roughly like this second type of distribution, such as L4 and Neu-NRGN-I.
[0106] The third case is a mixture of the first two cases. There may be subsets of cells in the first case, or subsets of cells in the second case. Of course, it may also be a non-extreme state of the second case. Further bioinformatics experiments are needed to explore the changes of the cells. In this experiment, a total of 9 types of cells are in this state.
[0107] In summary, the method in this embodiment realizes the disease-specific cell classification task of single-cell transcriptome. By performing differential expression analysis and scoring and dimensionality reduction on the single-cell transcriptome data of the disease group and the control group, and then using the iterative EM clustering algorithm to classify the cells, the differences in the disease group are projected onto each cell for subsequent analysis at a higher resolution. Since the Gaussian mixture model in the EM clustering algorithm used has a good clustering effect on elliptical distribution data, it is more suitable for the data distribution in the score space, making the clustering results more accurate and reasonable. Using the improved scoring algorithm to classify and score according to the total cell expression can avoid more noise interference when the number of differential genes is small; the rationality judgment method after iteration judges whether the cells change significantly under disease conditions from the correlation of cell scores before and after iteration, making the judgment result more persuasive.
Claims
1. A method for analyzing disease-specific cells in single-cell transcriptomes based on iterative EM clustering, characterized in that, The following steps are involved: Step 1: Obtain single-cell transcriptome datasets of healthy people and patients with certain disease conditions, use graph clustering methods to divide the single-cell transcriptome datasets into multiple clusters, and classify each cluster by marker genes to obtain cell populations of each type of cell; Step 2: For a cell population of interest, obtain the up-regulated and down-regulated genes of the cell population under the current disease condition; Step 3: Based on the obtained up-regulated and down-regulated genes, use the scoring algorithm to score each cell in the cell population of interest with an initial up-regulated score and an initial down-regulated score, and project the cells into a two-dimensional score space based on the scoring results; Step 4: Cluster the cells in the score space into two categories using the EM clustering algorithm; Step 5: Randomly select a type of cell after clustering, obtain its up-regulated and down-regulated genes relative to the other type, and use the scoring algorithm to re-score the up-regulated score and down-regulated score for each cell in the cell group, and then put the results into a two-dimensional score space; Step 6: Repeat steps 4 and 5 until the clustering results are consistent, or the set number of iterations is reached, and output the two types of cells after iterative clustering; Step 7: Calculate the initial disease score set for each cell based on the results of the initial up-regulation score and the initial down-regulation score as follows: Among them, is the initial disease score set of the cells, is the initial up-regulated score set of the cells, is the initial down-regulated score set of the cells; Then, according to the clustering results after the iteration, the category score set of cells after the iteration is calculated as follows: Among them, Score last is a set of class scores obtained based on a certain class of cells after iterative clustering, is a set of final up-regulated scores obtained according to the results of the last iteration, is a set of final down-regulated scores obtained according to the results of the last iteration; Subsequently, is re-sorted in ascending order according to the initial disease scores of each cell, and then the class scores of each cell in Score last are arranged according to the cell order corresponding to the re-sorted , and the Pearson correlation coefficient r between the re-sorted and Score last is calculated; Finally, the clustered cells are classified into disease-specific cells and control normal cells according to one of the following strategies: ① If r>0, it means that one type of cells selected in step 5 is disease-specific cells, and the other type is normal cells in the control group. If r<0, the classification is opposite; ② If the rearranged Score last is generally the same as that after extraction and arrangement and maintains an upward trend, it indicates that the type of cells selected in step 5 is disease-specific cells, while the other type is normal control cells; otherwise, the classification is reversed.
2. The method according to claim 1, wherein Step 7 also includes judging whether the score change result after iteration is reasonable according to the absolute value of the calculated Pearson correlation coefficient r. If |r|>0.6, it means the result is reasonable, otherwise it means the result is unreasonable. If the score changes unreasonably after iteration, it is considered that the cell population has no significant changes in the disease environment; If the changes are reasonable, the two types of cells output are classified into disease-specific cells and control group normal cells.
3. The method according to claim 1, characterized in that, The method further comprises step 8, repeating steps 2 to 7 to analyze a new group of cells until the analysis of all the cell groups of interest is completed.
4. The method according to claim 1, characterized in that, The specific steps of the scoring algorithm in step 3 are as follows: First, M genes that are different from the up-regulated and down-regulated genes to be scored are randomly selected as reference genes; Then, the Kmeans clustering algorithm was used to classify all cells into K categories according to the total expression of all genes in each cell; Next, the average expression of the reference gene in each category is calculated as follows: In the formula, is the average value of the reference gene expression levels of the cells in the k-th category, N k is the total number of cells in the k-th category, M represents the number of selected reference genes, expre ij represents the expression level of the j-th reference gene of the i-th cell in the k-th category; Finally, the average expression level of up-regulated or down-regulated genes in each cell was divided by the average expression level of the reference gene in the category to obtain the original up-regulated or down-regulated score of the cell. The original up-regulated or down-regulated scores of all cells were normalized to the maximum and minimum to obtain the up-regulated or down-regulated score of the cell.
5. The method according to claim 4, characterized in that, The up-regulated score for each cell was calculated as follows: where cell i ∈ Cell k Set InitialScore up = [initialupscore1, initialupscore2, ……, initialupscore N In the formula, is the average value of the up-regulated gene expression levels of the i-th cell, upexpre ij is the expression level of the j-th up-regulated gene of the i-th cell, M up represents the total number of up-regulated genes calculated by the differential gene expression algorithm, initialupscore i represents the initial up-regulation score of the i-th cell, cell i represents the i-th cell, Cell k represents the k-th type of cells separated by the above Kmeans clustering algorithm, which is a cell set, InitialScore up represents the set of the original up-regulation scores of all cells, N is the total number of cells, upscore i is the up-regulation score of cell i after max-min normalization; Similarly, the down-regulation score of each cell is calculated as follows: where cell i ∈ Cell k Set InitialScore down = [initialdownscore1, initialdownscore2, ……, initialdownscore N In the formula, is the average value of the down-regulated gene expression levels of the i-th cell, downexpre ij is the expression level of the j-th down-regulated gene in the i-th cell, M down represents the total number of down-regulated genes calculated by the differential gene expression algorithm, initialdownscore i represents the original down-regulation score of cell i, InitialScore down represents the set of original down-regulation scores of all cells, downscore i is the down-regulation score of cell i after max-min normalization; Finally, through sorting, we get the up-regulated score set and down-regulated score set of cells: Score up = [upscore1, upscore2, ……, upscore N Score down = [downscore1, downscore2,..., downscore N where Score up is the up-regulated score set, and Score down is the down-regulated score set.
6. The method according to claim 1, characterized in that In step 2, the up-regulated and down-regulated genes are obtained by using the Wilcoxon rank sum test.
Citation Information
Patent Citations
Analysis method based on 10X unicell transcriptome sequencing data
CN109979538A
Tumor malignant cell gene prognosis risk model construction method
CN115424728A