Tumor cell identification method based on single-cell sequencing data
Through the tumor cell identification method based on single-cell sequencing data, identifying and using programmed cell death-related genes, the problem of difficult identification of tumor cells with unknown markers in the prior art is solved, and a more accurate and comprehensive identification of large-scale cell types is achieved.
Patent Information
- Application Number
- CN202410682103.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-05-29
AI Technical Summary
The prior art is difficult to accurately identify tumor cells with unknown markers, and it is impossible to use genes related to programmed cell death to identify tumor cells.
Through tumor cell identification methods based on single-cell sequencing data, programmed cell death-related genes specifically expressed in tumor cells were identified, and cell type identification was performed using multiple classification models and fuzzy comprehensive evaluation.
It improves the accuracy of tumor cell markers, enables more comprehensive and quantitative identification of cell types, and provides a more comprehensive perspective for tumor research.
Smart Images

Figure CN118486374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor cell detection, and in particular to a tumor cell identification method based on single-cell sequencing data. Background Art
[0002] Cancer (tumor) is caused by uncontrolled cell growth. During the development of cancer, normal cells mutate and transform into tumor cells. Tumor cells grow out of control and invade surrounding normal tissues. Eventually, disseminated tumor cells colonize in secondary sites, achieving cancer metastasis. Despite advances in surgical resection and treatment, cancer remains the leading cause of mortality worldwide. The regeneration and metastasis of tumor cells in cancer are at the core of cancer development and metastasis, making it necessary to identify tumor cells in cancer for cancer treatment.
[0003] Traditional tumor cell identification methods help identify tumor cells by marking tumor markers, but this method relies on known markers, such as cell surface markers, nuclear markers, etc., and cannot accurately identify tumor cells with unknown markers. Studies have shown that genomic variation leading to abnormal cell DNA copy number is an important driving factor for cancer, and tumor cells often carry a higher proportion of copy number variation (CNV) than normal cells. Therefore, by inferring copy number variation, tumor cells can be identified to a certain extent. In addition, single-cell sequencing technology can provide genomic information at the level of individual cells, making it possible to identify copy number variation at the cellular level and quickly infer the possibility of tumor cell formation, but this method cannot discover new markers, which further limits the research on tumor cell detection and development. Therefore, tumor cell identification methods based on single-cell sequencing technology remain to be developed. At present, many studies have shown that programmed cell death plays a huge role in the occurrence and development of tumors. However, there is no method to use genes related to programmed cell death in tumor cell identification. Summary of the invention
[0004] Based on the above problems, the present invention provides a method for identifying tumor cells based on single-cell sequencing data. This method uses genes related to programmed cell death as features to greatly improve the accuracy of tumor cell markers for tumor cell identification, and provides a reference for the application of single-cell sequencing technology in tumor cell detection.
[0005] The specific technical scheme of the present invention is:
[0006] A tumor cell identification method based on single-cell sequencing data, including a tumor cell marker acquisition module and a tumor cell identification module:
[0007] The tumor cell marker acquisition module is used to identify programmed cell death-related genes specifically expressed in tumor cells;
[0008] The tumor cell identification module is used to identify tumor cells; using the tumor cell marker genes obtained by the tumor marker acquisition module as features, multiple classification models are used to identify cell types, and fuzzy comprehensive evaluation is used to integrate the prediction results of multiple models to quantify the cell classification results.
[0009] Preferably, in the tumor cell marker acquisition module, the markers obtained are differential genes related to programmed cell death in tumor cells, and specifically include the following steps:
[0010] Step S11: Obtain a set of programmed cell death-related genes and tumor tissue single-cell sequencing data, perform quality control on the data, and use copyKAT software to infer cell copy number variation;
[0011] Step S12: extract the prediction content in the output result to obtain the cell type file cell_type containing the cell name and cell chromosome multiples;
[0012] Step S13: Read the single-cell gene expression matrix and integrate the cell_type file content according to the cell name;
[0013] Step S14: grouping based on chromosome multiples annotation, where cells with chromosome multiples labeled as diploid are classified as a normal cell group (Normal), and cells with chromosome multiples labeled as aneuploid are classified as a tumor cell group (Tumor);
[0014] Step S15: using the differential analysis software limma to compare the gene expression levels between the Normal and Tumor groups to determine the genes with significant differences between the different groups, and obtain the programmed death-related genes among them;
[0015] Step S16: Calculate the correlation between gene expression and cell type, and use the correlated genes as a marker candidate gene set;
[0016] It should be noted that copyKAT and limma software are only one way to implement the tumor cell enrichment module. It can be understood that the key to the tumor marker acquisition module is to obtain markers related to programmed cell death in tumor cells at single-cell resolution. As for the specific tumor cell and differential gene identification methods, it is not ruled out that other bioinformatics or experimental methods can also be used to achieve it.
[0017] Preferably, in step S15, the differential genes related to programmed cell death are obtained by taking the intersection of the differential genes and the programmed cell death gene set, thereby improving the biological information of the marker.
[0018] Preferably, in step S16, the mutual information and chi-square value between gene expression and cell type are calculated respectively, and the intersection of the genes in the top 30% of the mutual information value and the genes in the top 30% of the chi-square value is taken to obtain markers with higher comprehensive correlation, thereby improving the accuracy of the markers in tumor cell identification.
[0019] Preferably, the tumor identification module comprehensively considers the prediction results of multiple classification models to quantify the cell type, specifically comprising the following steps:
[0020] Step S21: Based on the genes in the marker candidate gene set, the expression matrix and cell type are read, the data is divided into a training set and a test set, and multiple classification models are trained on the training set;
[0021] Step S22: For each cell in the test set, a cell type identification matrix is constructed according to its type prediction probability under each classification model. The matrix is shown as follows:
[0022] Tumor cells Normal cells Model 1 P11 P12 Model 2 P21 P22 … … … Model Pn1 Pn2
[0023] The elements in the matrix represent the probability of a cell being classified as a tumor cell or a normal cell in the model.
[0024] Step S23: setting weights for each model to obtain a weight vector;
[0025] Step S24: performing weighted summation on the weight vector and the cell type identification matrix to calculate the membership scores of the cells to the two categories under multi-model integration;
[0026] Step S25: According to the maximum membership principle, determine the final classification result of the cells under the comprehensive combination of multiple models;
[0027] Preferably, the setting of each model weight in step S23 is achieved by normalizing the AUC value, and the calculation formula is as follows:
[0028]
[0029] Where AUCi refers to the area under the curve of model i on the training data set, and ∑AUC refers to the sum of the areas under the curves of all models on the training data set.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] The tumor cell marker acquisition module of the present invention realizes the accurate identification of markers related to programmed cell death in tumor cells at the single-cell level, providing ideas for the subsequent in-depth study of the pathogenesis of tumors and the development of related anti-tumor drugs. In addition, in the tumor cell identification module, the identification of tumor cells comprehensively considers multiple classification models and quantitatively evaluates the categories to which cells belong, which can comprehensively and quantitatively identify cell types and provide a more comprehensive perspective for tumor research. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is the technical roadmap of tumor identification based on single-cell sequencing data of the present invention.
[0033] Figure 2 This is a flow chart of the tumor cell marker acquisition module in Example 1 of the present invention.
[0034] Figure 3 This is a flow chart of the tumor cell identification module in embodiment 2 of the present invention. DETAILED DESCRIPTION
[0035] The present invention will be further described below in conjunction with the embodiments.
[0036] Refer to the technical route Figure 1 , operate the tumor cell marker acquisition module and the tumor cell identification module to realize tumor cell identification based on single-cell sequencing data.
[0037] Example 1: Tumor cell marker acquisition module process
[0038] Reference process Figure 2 The tumor cell marker acquisition module process includes obtaining single-cell sequencing data of glioblastoma tissue and then identifying programmed cell death-related genes that are differentially expressed in tumor cells; the specific steps are as follows:
[0039] Step S11: Obtain a set of programmed cell death-related genes, including 1965 genes in 18 modes. Obtain single-cell sequencing data of glioblastoma tissue, and obtain 3211 cells and 20809 genes after quality control of the data. Then, use copyKAT software to infer cell copy number variation;
[0040] Step S12: extract the prediction content in the output result to obtain the cell type file cell_type containing the cell name and cell chromosome multiples;
[0041] Step S13: integrating cell_type and cell expression matrix;
[0042] Step S14: grouping the matrix based on the chromosome multiples labeling, the chromosome multiples labeling diploid are classified as the normal cell group (Normal), and the chromosome multiples labeling aneuroploid are classified as the tumor cell group (Tumor), wherein the Normal group includes 1621 cells and 20809 genes, and the Tumor group includes 1551 cells and 20809 genes;
[0043] Step S15: using the differential analysis software limma to compare the gene expression levels between the Normal and Tumor groups, 2288 genes with significant expression differences in the Tumor group were obtained, and the intersection with the programmed cell death-related gene set was taken to obtain 1040 differentially expressed programmed cell death-related genes;
[0044] Step S16: Calculate the correlation between the expression levels of the 1040 genes and the cell types, and extract 255 markers with higher correlation.
[0045] Furthermore, the intersection of the top 30% of the mutual information values and the top 30% of the chi-square values was extracted to obtain 255 markers with high comprehensive correlation; the specific 255 markers and their corresponding programmed cell death patterns are shown in Table 1.
[0046] Table 1: 255 markers and their corresponding programmed cell death patterns
[0047]
[0048]
[0049]
[0050] Example 2: Tumor cell identification module process
[0051] Reference process Figure 3 The process of the tumor cell identification module is to use the genes obtained by the tumor marker acquisition module as characteristic genes, and then use multiple classification models to classify the cells, and use fuzzy comprehensive evaluation to quantify the cell classification results by integrating the prediction results of multiple models.
[0052] The tumor identification module considers the predictions of multiple classification models to quantify cell types. It includes the following steps:
[0053] Step S21: Based on the genes in the marker candidate gene set, the expression matrix and cell type are read, the data is divided into a training set and a test set, and then the random forest, support vector machine and XGBoost models are trained on the training set;
[0054] Step S22: For each cell in a completely independent test set, a cell type identification matrix is constructed according to its type prediction probability under each classification model; the specific type identification matrix constructed for a certain cell is shown in Table 2.
[0055] Table 2 Cell type identification matrix
[0056]
[0057]
[0058] The elements in the matrix represent the probability of a cell being classified as a tumor cell or a normal cell in the model.
[0059] Step S23: setting weights for each model to obtain a weight vector;
[0060] Step S24: performing weighted summation on the weight vector and the cell XX type identification matrix to obtain the membership scores of the cell XX to the two categories;
[0061] Step S25: According to the maximum membership principle, determine the final classification result of cell XX under the comprehensive combination of multiple models;
[0062] Specifically, the setting of each model weight in step S23 is achieved by normalizing the AUC value, and the calculation formula is as follows:
[0063]
[0064] Where AUCi refers to the area under the curve of model i on the training data set, and ∑AUC refers to the sum of the areas under the curves of all models on the training data set.
[0065] Specifically, the weights of the support vector machine, random forest, and XGBoost models are shown in Table 3:
[0066] Table 3 Weights of support vector machine, random forest, and XGBoost models
[0067] Support Vector Machine Random Forest XGBoost 0.273967552 0.363569322 0.362463127
[0068] Specifically, the calculation process of the membership score of cell XX to tumor cells and normal cells is as follows:
[0069]
[0070] That is, the membership scores for normal cells and tumor cells are 0.916 and 0.084, respectively.
[0071] Specifically, according to the maximum membership principle, the final classification result of the cells under the multi-model integration was determined to be normal cells, which is consistent with the prediction result of copyKAT.
[0072] In summary, the present invention uses advanced bioinformatics algorithms and high-throughput sequencing technology, and utilizes a tumor cell marker acquisition module to achieve accurate identification of markers related to programmed cell death in tumor cells at the single-cell level, thereby providing strong data support for the subsequent quantitative evaluation of the cell category. This module has established a complete quantitative evaluation system, which comprehensively scores the characteristic information of cells in multiple dimensions and gives a specific quantitative index to reflect the probability of the cell category. This quantitative evaluation method makes the cell identification results more comparable and explainable. In addition to traditional tumor cell types, this module can also identify other types of cells, such as normal cells.
Claims
1. A method for identifying tumor cells based on single-cell sequencing data, characterized in that: Including tumor cell marker acquisition module and tumor cell identification module: The tumor cell marker acquisition module is used to identify programmed cell death-related genes specifically expressed in tumor cells. The obtained markers are differential genes related to programmed cell death in tumor cells. Specifically, the module comprises the following steps: Step S11: Obtain a set of programmed cell death-related genes and tumor tissue single-cell sequencing data, perform quality control on the data, and use copyKAT software to infer cell copy number variation; Step S12: extract the predicted content in the output result to obtain a cell type file cell_type containing the cell name and cell chromosome multiples; Step S13: Read the single-cell gene expression matrix and integrate the cell_type file content according to the cell name; Step S14: grouping based on chromosome multiples annotation, where the chromosome multiples annotated as diploid are grouped as a normal cell group Normal, and the chromosome multiples annotated as aneuploid are grouped as a tumor cell group Tumor; Step S15: using the differential analysis software limma to compare the gene expression levels between the Normal and Tumor groups to determine the genes with significant differences between the different groups, and obtain the programmed cell death-related genes; Step S16: Calculate the correlation between gene expression and cell type, and use the correlated genes as a marker candidate gene set; The tumor cell identification module is used to identify tumor cells, using the tumor cell marker genes obtained by the tumor cell marker acquisition module as features, using multiple classification models to identify cell types, and using fuzzy comprehensive evaluation to integrate the prediction results of multiple models to quantify the cell classification results.
2. A method for identifying tumor cells based on single-cell sequencing data according to claim 1, characterized in that: In the step S15, the differential genes related to programmed cell death are obtained by taking the intersection of the differential genes and the programmed cell death gene set.
3. A method for identifying tumor cells based on single-cell sequencing data according to claim 1, characterized in that: In step S16, the mutual information and chi-square value between gene expression and cell type are calculated respectively, and the intersection of the top 30% of the mutual information value and the top 30% of the chi-square value is taken to obtain markers with high comprehensive correlation.
4. A method for identifying tumor cells based on single-cell sequencing data according to claim 1, characterized in that: The tumor cell identification module comprehensively considers the prediction results of multiple classification models to quantify the cell type, specifically including the following steps: Step S21: Based on the genes in the marker candidate gene set, the expression matrix and cell type are read, the data is divided into a training set and a test set, and multiple classification models are trained on the training set; Step S22: For each cell in the test set, a cell type identification matrix is constructed according to its type prediction probability under each classification model. The matrix is shown as follows: The rows in the matrix represent classifiers, the columns represent cell types, and the elements P in the matrix i1 represents the probability that a cell is classified as a tumor cell in classifier i, and the element P i2 represents the probability that a cell is classified as a normal cell in classifier i; Step S23: setting weights for each model to obtain a weight vector; Step S24: performing weighted summation on the weight vector and the cell type identification matrix to calculate the membership scores of the cells to the two categories under multi-model integration; Step S25: According to the maximum membership principle, the final classification result of the cells under the comprehensive combination of multiple models is determined.
5. A method for identifying tumor cells based on single-cell sequencing data according to claim 4, characterized in that: In step S23, the setting of each model weight is achieved by normalizing the AUC value, and the calculation formula is as follows: Among them, AUCi refers to the area under the curve of model i on the training data set, and ∑AUC refers to the sum of the areas under the curves of all models on the training data set.
Citation Information
Patent Citations
Liver cancer patient survival prediction model construction method based on cell death related genes
CN115019965A
Tumor single cell transcriptome sequencing data cell type identification method and system
CN117037907A