Precise identification method for benign and malignant cells based on multi-dimensional characteristics of single cell transcriptome

By using a multidimensional feature-based modeling method based on single-cell transcriptome data, the challenge of identifying benign and malignant cells in scenarios involving multiple patients, multiple samples, and early-stage tumors has been solved. This method enables stable identification and accurate determination of malignant cells, providing reliable tumor diagnostic support, especially in tumor margin regions and early lesion stages.

CN121983137APending Publication Date: 2026-05-05BEIJING INSTITUTE OF GENOMICS CHINESE ACADEMY OF SCIENCES (CHINA NATIONAL CENTER FOR BIOINFORMATION)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INSTITUTE OF GENOMICS CHINESE ACADEMY OF SCIENCES (CHINA NATIONAL CENTER FOR BIOINFORMATION)
Filing Date
2026-01-21
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably distinguish between benign and malignant tumor cells based on single-cell transcriptome data in scenarios involving multiple patients, multiple samples, and early-stage tumors. This is especially true in tumor margin areas, residual cell populations after treatment, early lesions, or the initial stages of tumor development. Current single-feature-based judgment strategies are unable to provide reliable conclusions and are prone to missing a small number of key malignant clonal cells.

Method used

A multidimensional feature-based modeling method based on single-cell transcriptome data was adopted, combining allele frequency, copy number variation, and gene expression characteristics. Data processing and analysis were performed using the inferCNV algorithm, Seurat tool, and Numbat tool to construct a unified modeling and quantitative evaluation framework, including single-cell transcriptome data preprocessing, copy number variation analysis, allele-specific copy number variation inference, and comprehensive judgment. Finally, a voting rule was used to accurately identify benign and malignant cells.

Benefits of technology

It improves the ability to identify malignant cells in early-stage tumors and cells in the transitional state between benign and malignant tumors, providing a more reliable molecular basis for precise tumor diagnosis and personalized treatment, and enhancing the accuracy and robustness of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983137A_ABST
    Figure CN121983137A_ABST
Patent Text Reader

Abstract

The invention discloses a method for accurately identifying benign and malignant cells based on multi-dimensional characteristics of a single cell transcriptome, and belongs to the technical fields of bioinformatics, tumor molecular biology and cell identification. On the basis of single cell transcriptome sequencing data of tumor tissues and para-carcinoma tissues, three types of information including copy number variation, allele specific copy number variation and tumor-related transcriptional characteristics are synthesized, and final benign and malignant identification is performed on each EpCAM positive epithelial cell through a multi-evidence voting strategy. The method disclosed by the invention has relatively high stability and accuracy in a multi-patient, multi-sample and early tumor scene, particularly improves the recognition capability of early lesion and malignant cells in benign and malignant boundary transition state cells, and provides a reliable technical means for precise diagnosis and individualized treatment of tumors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of single-cell transcriptomics, bioinformatics, tumor molecular diagnostics, and cell identification technology, specifically to a method for accurately identifying benign or malignant cells based on single-cell transcriptomics data and by integrating multidimensional characteristics such as alleles, copy number variations, and gene expression. Background Technology

[0002] The occurrence and progression of tumors are driven by multiple factors, including genetic mutations, genomic instability, abnormal epigenetic regulation, and microenvironmental stress. Tumor tissue typically consists of various cell types, such as tumor cells, reactive epithelial cells, stromal cells, and various immune cells. Different cell populations exhibit significant differences in proliferative activity, genomic instability, metabolic state, and invasive / metastatic capabilities. Tumor cells often present a continuum of morphological and molecular characteristics with normal or reactive cells from their origin tissues. Especially in the early stages of tumor development, the genomic aberration load and transcriptional signatures of malignant cells are relatively limited, and their expression patterns are highly similar to those of normal or reactive cells, making it more difficult to distinguish early malignant cells from benign cells using only traditional indicators. Accurately differentiating benign cells from malignant tumor cells in complex tissue contexts is a crucial foundation for tumor subtyping, efficacy evaluation, and the development of targeted therapy strategies.

[0003] Traditional methods for determining benign or malignant tumors primarily rely on histopathological morphological observation and the detection of a limited number of immunohistochemical markers, combined with molecular detection results such as some driver gene mutations or copy number variations. These methods are limited by factors such as observer experience, differences in sampling sites, and the limited coverage of markers. For cell populations with ambiguous boundaries, similar morphology, or those affected by treatment, inconsistencies in judgment or inability to reach definitive conclusions can easily arise. While somatic mutation and copy number variation analysis based on whole-cell samples can reflect the overall genomic burden of tumors, it cannot resolve the benign or malignant at the single-cell level.

[0004] With the rapid development of single-cell RNA sequencing (scRNA-seq) technology, gene expression profiles of various cell types in tumor tissues can be obtained at single-cell resolution, enabling direct identification of malignant tumor cells at the expression level and differentiation from normal or reactive cells. Current tumor cell identification strategies based on single-cell transcriptome data mainly include: expression scoring using known tumor-associated marker genes or gene sets; identifying cells with significant chromosomal instability based on inferred copy number variation (CNV) patterns; and, in some studies, attempting indirect inference of mutations or allele imbalances using single-cell transcriptome data. However, in early stages of tumor development or samples with a low proportion of tumor cells, the copy number variation load of malignant clones is often low and the magnitude of change is limited. Even when copy number variation inference is performed at the single-cell transcriptome level, the signal is not significant, making it difficult to accurately identify the molecular characteristics of a small number of early malignant cells. Overall, the above methods still mainly rely on single or a few feature dimensions, and their stability and generalization ability remain limited in multi-patient, multi-sample, and multi-platform data environments.

[0005] On the one hand, strategies relying solely on marker genes or expression signatures are easily influenced by the functional state, differentiation degree, and microenvironment of tumor cells. Some benign cells with active proliferation or under stress may also exhibit similar expression patterns, leading to false positives. Conversely, some malignant cells with minimal genomic alterations or in a quiescent state may be misclassified as benign, a problem more common in early lesions or mild atypical hyperplasia. On the other hand, analysis methods for inferring copy number variations based on transcriptome data are highly sensitive to sequencing depth, gene coverage, and reference cell selection. For tumor types with low copy number variation loads, highly confounded samples, or data with significant technical noise, the inference results are often unstable. Furthermore, existing methods generally lack systematic utilization of allele-level information such as allele expression imbalance, recessive deletions, and focal copy number variations, making it difficult to make joint judgments based on allele characteristics, copy number variation characteristics, and overall gene expression characteristics.

[0006] In real-world research and application scenarios involving multiple patients and centers, there is a lack of a method that can systematically integrate multidimensional features such as allele information, copy number variation patterns, and gene expression profiles based on single-cell transcriptome data to construct a unified scoring or classification model, and to provide a stable, interpretable, and cross-cohort-generalizable accurate identification of benign and malignant cells. This is especially true in tumor margin regions, post-treatment residual cell populations, early lesions, or the initial stages of tumor development, where a large number of cells are in a benign-malignant boundary or transitional state. Existing single-feature-based judgment strategies are even less likely to provide reliable conclusions and are prone to missing a small number of key malignant clones.

[0007] Therefore, there is an urgent need to develop an analytical method based on single-cell transcriptome data that integrates multidimensional features such as allele frequency, copy number variation, and gene expression to uniformly model and quantitatively assess the benign and malignant nature of cells. This would improve the accuracy, robustness, and generalizability of benign and malignant cell identification, especially enhancing the ability to identify malignant cells in early-stage tumors and transitional cells, thus providing more reliable molecular evidence for precise tumor diagnosis and personalized treatment. Summary of the Invention

[0008] The purpose of this invention is to address the shortcomings of existing technologies that rely on only a single feature dimension and struggle to reliably distinguish between benign and malignant tumor cells in multiple patients, multiple samples, and early-stage tumor scenarios. This invention provides a precise method for identifying benign and malignant cells based on multidimensional features of the single-cell transcriptome. This method uses single-cell transcriptome data as a foundation, integrating multidimensional features such as allele frequency characteristics, copy number variation characteristics, and gene expression to construct a unified modeling and quantitative evaluation framework. This enables stable identification of benign and malignant cells in samples from different sources, particularly improving the ability to identify malignant cells in early-stage tumors and cells in the transitional state between benign and malignant tumors.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] A method for precise identification of benign and malignant cells based on multidimensional features of single-cell transcriptomes includes the following steps:

[0011] 1) Preprocessing of single-cell transcriptome data

[0012] Single-cell transcriptome data from the target tumor samples were preprocessed, including normalization and screening for hypervariable genes. Principal component analysis (PCA) and UMAP (Uniform Manifold Approximation and Projection) were used for dimensionality reduction, and a clustering algorithm was employed to divide the cells into different clusters. Cell types were annotated based on known marker genes, primarily classifying them into EpCAM (Epithelial Cell Adhesion Molecule) positive epithelial cells, mesenchymal cells, and immune cells. EpCAM-positive epithelial cells were selected as the candidate cell set for determining benignity or malignancy, while immune cells and mesenchymal cells were used as normal reference cell populations for subsequent copy number variation inference.

[0013] 2) Inference of malignant cells based on copy number variation

[0014] The inferCNV algorithm is used to analyze the copy number variation of candidate cells in the target tumor sample. Immune cells are used as a normal reference cell population to infer the copy number variation (CNV) of candidate cells. The sample-specific CNV judgment threshold is determined based on the average CNV level of the mesenchymal cell cluster. Cells with an average CNV level exceeding the CNV judgment threshold are marked as malignant cells.

[0015] 3) Malignant cell inference based on tumor-related transcriptional signature scores

[0016] Based on the differences in transcriptome expression between tumor tissue and adjacent normal tissue bulk, malignant and non-malignant characteristic gene sets were constructed, and the malignancy score of each candidate cell was calculated using the Seurat tool. The characteristic gene sets were updated through multiple iterations to ultimately determine whether a cell was malignant or non-malignant. The termination condition for multiple iterations of the characteristic gene set was: the consistency ratio between the current and previous iterations of the characteristic gene set reached or exceeded a preset threshold (95%), or the number of iterations (10 iterations) reached a preset upper limit.

[0017] 4) Inference of malignant cells based on allele-specific copy number variations

[0018] Using the Numbat tool, allele information of target tumor samples is extracted to infer copy number variations. Key parameters are set for modeling. Based on tumor probability, genotype status, and clonal tags, malignant cells are identified.

[0019] 5) Comprehensive determination of final benign or malignant nature

[0020] Combining the inference results based on copy number variation, allele-specific copy number variation, and tumor-related transcriptional signature scores, a voting rule was used to make a final determination of benign or malignant characteristics of EpCAM-positive epithelial cells.

[0021] Step 1) above involves preprocessing and cell type clustering of single-cell transcriptome data, specifically including:

[0022] Single-cell transcriptome sequencing data of tumor tissue and / or adjacent normal tissue samples were obtained, and a Seurat object containing a UMI counting matrix was constructed. The object was preprocessed by normalization, high-variance gene screening, etc., and the cells were divided into multiple cell clusters by combining principal component analysis, UMAP and other dimensionality reduction methods and clustering algorithms.

[0023] Based on known marker genes, cell clusters are annotated to classify cells into at least EpCAM-positive epithelial cells, mesenchymal cells, and immune cells including B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells; preferably, each cell type contains at least two annotation genes.

[0024] Based on the cell type annotation results, EpCAM-positive epithelial cells were screened from all cells and only retained to construct a candidate cell set for determining benign or malignant cells. At the same time, immune cells were used as the normal reference cell population in subsequent copy number variation inference and their annotation information was retained.

[0025] Step 2) above, which infers malignant cells based on copy number variations, specifically includes:

[0026] First, for the target tumor sample, single-cell transcriptome objects with completed cell type annotation were selected. The original UMI count matrix of the whole genome and the corresponding cell type annotation information were extracted, and a gene sequence file containing gene chromosomal positions and gene sequences was prepared based on the reference genome annotation file. The original UMI count matrix of the whole genome and the cell type annotation information were organized into the input format required by the copy number variation inference algorithm inferCNV, generating cell annotation files and expression matrix files. In the cell type annotation, B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells were pre-specified as normal reference cell populations; the inferCNV analysis objects were created, and sex chromosomes and mitochondrial chromosomes were excluded during the modeling process to reduce artifacts.

[0027] The inferCNV process uses uniform preset parameters, including filtering for low-expression genes, enabling Hidden Markov Models for CNV state prediction, and combining graph clustering with the Leiden algorithm to further subdivide potential tumor cell subpopulations. This process infers the CNV profile of each candidate cell on each chromosome segment and the expression matrix generated by inferCNV. After the inferCNV analysis is completed, the results are added to the metadata of the Seurat object through an ensemble function. The normalized CNV index column starting with proportion_scaled_cnv_chr is extracted from the metadata. The average CNV level of each cell is obtained by averaging the CNV indexes of this type on each chromosome.

[0028] Using cell clusters or cell types as grouping units, violin plots and box plots of average CNV levels are drawn to observe the CNV distribution characteristics of each cell population. Based on a global initial threshold, and combined with the average CNV level distribution of each mesenchymal cell cluster in the same sample, a sample-specific CNV determination threshold is obtained. Each cell is binary classified according to this CNV determination threshold; cells with an average CNV level greater than the threshold are marked as malignant cells. Preferably, the global initial threshold is an average CNV level of 0.10, and the sample-specific CNV determination threshold is adjusted to a suitable level by observing the average CNV level distribution of each mesenchymal cell cluster in the same sample.

[0029] In addition, for samples with paired adjacent normal tissue, the above analysis process was repeated for single-cell transcriptome sequencing data of adjacent normal tissue. For EpCAM positive cell clusters with an average CNV level higher than the CNV determination threshold, they were regarded as false positive malignant cell clusters caused by factors such as technical noise, focal non-tumor amplification, or alignment error, and were removed from the previously defined malignant cell set. Finally, a set of malignant epithelial cells based on copy number variation and corrected for adjacent normal tissue was obtained.

[0030] Step 3) above, inference of malignant cells based on tumor-related transcriptional signature scores, specifically includes:

[0031] First, publicly available tissue block transcriptome datasets were collected for the tumor type under study, and differential expression analysis was performed between tumor tissues and corresponding normal tissues; a p-value less than a preset threshold (e.g., 0.001) was used to determine the expression level. Multiple change ( If the absolute value exceeds a preset threshold (e.g., 1), gene expression will be used. Genes are sorted from high to low fold change. The top 50 (e.g., 50) highly expressed genes from tumor tissue are selected to construct a malignant characteristic gene set, and the top 50 (e.g., 50) highly expressed genes from normal tissue are selected to construct a non-malignant characteristic gene set.

[0032] Subsequently, using the single-cell transcriptome data with annotation information obtained in step 1), the AddModuleScore function in the Seurat package is used to calculate the malignancy score and non-malignancy score for each EpCAM-positive epithelial cell based on the malignancy feature gene set and the non-malignancy feature gene set, respectively. The difference score between the malignancy score and the non-malignancy score is defined as the benign-malignancy difference score of the cell, and the difference score is normalized to the interval of 0 to 1. Based on the normalized benign-malignancy difference score, the k-means unsupervised clustering algorithm is used to perform binary classification on the EpCAM-positive epithelial cells to obtain the initial set of malignant epithelial cells and non-malignant epithelial cells.

[0033] Based on this, malignant and non-malignant epithelial cells obtained from the current round of classification were divided into two groups. Differential expression analysis was performed at the single-cell level. Genes upregulated in malignant epithelial cells were selected as a new set of malignant characteristic genes, and genes upregulated in non-malignant epithelial cells were selected as a new set of non-malignant characteristic genes. Based on the updated characteristic gene sets, the malignancy score and non-malignancy score of each cell were recalculated using the AddModuleScore function, and the difference between the two was used as the new benign and malignant differentiation score. The epithelial cells were again divided into "malignant" and "non-malignant" categories by k-means clustering, and the malignancy and non-malignancy labels of each cell were updated.

[0034] After each iteration, the consistency ratio between the classification result of the current iteration and the classification result of the previous iteration is calculated. When the consistency reaches or exceeds a preset threshold (e.g., 95%) or the number of iterations reaches a preset upper limit (e.g., 10 iterations), the iteration is terminated. The classification result obtained at the time of termination is used as the final benign or malignant determination result based on transcriptional features.

[0035] Step 4) above, which involves inferring malignant cells based on allele-specific copy number variations, has the following specific requirements:

[0036] First, for the target tumor sample to be analyzed, its single-cell transcriptome alignment results (BAM file), cell barcode file, population SNP reference variant file, and genetic map file are used as input. The Numbat pipeline pileup_and_phase is run to extract allele readings and perform haplotype phasing on heterozygous SNP sites across the whole genome, generating an allele count data file containing reference allele / variant allele reading information for each cell, which is used as input for subsequent Numbat modeling.

[0037] Secondly, the run_numbat function of Numbat is called to infer allele-specific copy number variations in the target tumor sample. The UMI count matrix of the candidate cells to be judged in the tumor sample is used as the input of count_mat expression level, and the UMI count matrix of B lymphocytes, T lymphocytes, natural killer cells and myeloid cells is used as the ref_internal normal reference expression intensity parameter.

[0038] By setting key modeling parameters and combining expression-level and allele-level signals, copy number variation inferences can be performed on tumor samples. Key modeling parameters need to be set: t is used as a numerical stability term to smooth zero counts, ensuring the numerical stability of log-likelihood calculations; gamma characterizes the excessive dispersion of allele counts, reasonably modeling allele noise at the single-cell level; min_cells limits the minimum number of cells participating in pseudo-bulk modeling, ensuring allele coverage and estimation accuracy; multi_allelic indicates whether to allow the identification of multi-allelic copy number variations, controlling the richness of the copy number state space; min_LLR serves as a threshold for filtering low-confidence CNV events by log-likelihood ratio, eliminating events with insufficient statistical support; max_entropy is an upper limit parameter for the complexity of the candidate copy number state space, used to suppress overly complex or highly uncertain CNV states, improving the robustness of clonal partitioning; call_clonal_loh enables the identification of clonal loss of heterozygosity (LOH) states, enhancing the ability to resolve allele-level anomalies; init_k specifies the initial number of clones, guiding the initial partitioning and subsequent optimization process of the tumor clonal structure. The specific values ​​of the above parameters need to be dynamically adjusted according to the stage of the tumor to be analyzed, the inferred copy number variation load, and the clonal complexity: for samples with low copy number variation load and suspected early-stage tumors, the state space complexity can be appropriately reduced or the initial number of clones can be reduced to avoid excessive division; for samples with highly complex clonal structures and suspected late-stage or multiclonal samples, the state complexity limit can be appropriately relaxed or the initial number of clones can be increased to more fully characterize the subclonal structure.

[0039] Finally, the tumor probability p_cnv_x at the expression level, the tumor probability p_cnv_y at the allele level, and the combined posterior tumor probability p_cnv for each cell are integrated into the Seurat object metadata of the tumor sample. Violin plots and box plots of the three are drawn with cell clusters as the grouping unit. EPCAM-positive epithelial cells with p_cnv higher than the preset threshold of 0.90 and also showing a high tumor probability in p_cnv_x and p_cnv_y are identified as malignant cells. The remaining EpCAM-positive epithelial cells are regarded as non-malignant or suspicious cells, forming a benign or malignant determination result based on allele-specific copy number variation.

[0040] Step 5 above) comprehensively determines the final benign or malignant nature, including:

[0041] For each EpCAM-positive epithelial cell, the judgment results based on copy number variation, allele-specific copy number variation, and tumor-related transcriptional feature scores are summarized to form a benign / malignant determination vector containing three pieces of evidence. A voting rule is set: when at least two pieces of evidence support that a cell is a malignant tumor cell, it is finally determined to be a malignant cell; otherwise, it is determined to be a benign or non-malignant cell. According to the above voting fusion strategy, each candidate EpCAM-positive epithelial cell is comprehensively judged, which is the final result of the accurate identification of benign and malignant cells in this invention.

[0042] Based on the aforementioned method for precise identification of benign and malignant cells, this invention provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method for precise identification of benign and malignant cells.

[0043] The present invention also provides a computer-readable storage medium having the computer program stored thereon.

[0044] The present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method for precise identification of benign and malignant cells of the present invention.

[0045] Compared with existing methods, the present invention has the following advantages:

[0046] 1) The multi-dimensional feature comprehensive modeling framework established in this invention can stably distinguish between early tumors and non-malignant cells, overcoming the shortcomings of traditional methods in cross-sample and early tumor identification.

[0047] 2) The accuracy and robustness of malignant cell identification were improved through comprehensive analysis based on copy number variation and allele-specific copy number variation.

[0048] 3) This invention combines multidimensional features (gene expression, copy number variation, allele specificity analysis) and transcriptional feature scoring, providing a more reliable means of identifying malignant cells, and is particularly suitable for the accurate determination of early tumors and transitional cells at the boundary between benign and malignant tumors.

[0049] 4) The identification of malignant cells based on tumor-related transcriptional features can further provide solid technical support for the discovery, precise diagnosis and treatment of tumor biomarkers. Attached Figure Description

[0050] Figure 1 A schematic diagram of the overall process of this invention for the precise identification of benign and malignant cells based on multidimensional features of tissue block transcriptome and single-cell transcriptome.

[0051] Figure 2In this embodiment of the invention, single-cell transcriptome cell type annotation UMAP diagram of lung adenocarcinoma tissue (A) and its adjacent normal tissue (B).

[0052] Figure 3 The distribution map of average CNV levels of different cell clusters inferred by inferCNV in this embodiment of the invention.

[0053] Figure 4 The UMAP distribution map of malignant epithelial cells in this embodiment of the invention is based on copy number variation and corrected for adjacent normal tissue.

[0054] Figure 5 In this embodiment of the invention, the distribution map of UMAP for malignant and non-malignant epithelial cells is distinguished based on tumor-related transcriptional feature scores.

[0055] Figure 6 In this embodiment of the invention, Numbat infers the probability distribution map in different cell clusters.

[0056] Figure 7 UMAP distribution map of malignant epithelial cells based on allele-specific copy number variation determination in this embodiment of the invention.

[0057] Figure 8 UMAP distribution map of the final benign / malignant cell determination results in this embodiment of the invention. Detailed Implementation

[0058] The technical solution of the present invention will be described in more detail below through embodiments. The parameters and specific implementation details are used to explain the feasibility and implementation effect of the present invention, and do not constitute a limitation of the present invention.

[0059] This embodiment uses single-cell transcriptome data from tumor tissue and adjacent tissue of a lung adenocarcinoma patient for demonstration. Figure 1 The flowchart below shows the method for precise identification of benign and malignant cells according to the present invention. The experimental steps and results are described in detail below:

[0060] 1) Preprocessing of single-cell transcriptome data and cell type clustering

[0061] First, the expression matrices of single-cell transcriptome sequencing data from tumor tissues and adjacent normal tissues of lung adenocarcinoma patients were obtained, normalized, and highly variable genes were screened. Next, dimensionality reduction was performed using PCA, and cluster analysis was conducted on the dimensionality-reduced data to group cells based on similarity. Finally, UMAP was used for non-linear dimensionality reduction to generate a two-dimensional cell type distribution map.

[0062] Based on known cell type marker genes, we performed type annotation on cell clusters in single-cell transcriptome data from tumor and adjacent normal samples. Specifically, EpCAM-positive epithelial cells were identified by EpCAM, KRT19, and PROM1; fibroblasts were identified by PDGFRA, FN1, and DCN; endothelial cells were identified by PECAM1, VWF, and CD34; B lymphocytes were identified by CD79A and MS4A1; T lymphocytes were labeled by CD3D and CD3E; natural killer cells were identified by NKG7 and GNLY; and myeloid cells were labeled by CD14 and CD68. Through expression analysis of these marker genes, we successfully classified cells into multiple types, including EpCAM-positive epithelial cells, mesenchymal cells, and B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells. Finally, we obtained the following results: Figure 2 The UMAP cluster diagram shown clearly illustrates the distribution of these different cell types.

[0063] Based on the aforementioned cell type annotation results, EpCAM-positive epithelial cells were screened from all cells to form a candidate cell set for determining benignity or malignancy. Meanwhile, immune cells were retained and used as a reference normal cell population for subsequent copy number variation inference.

[0064] 2) Inference of malignant cells based on copy number variation

[0065] First, prepare the input data for the inferCNV algorithm. Extract the original UMI count matrix (counts matrix) of all genes in the tumor sample and the corresponding cell type annotation information (meta.data$celltype). Prepare a reference genome annotation file containing the chromosomal location and gene order of the genes. Organize the original UMI count matrix and cell annotation information into the format required by the inferCNV algorithm. Generate two files: one is a cell annotation file containing the barcode of each cell and its corresponding cell type, and the other is an expression matrix file containing the expression data of each gene in different cells. In the cell type annotation, we pre-specify B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells as normal reference cell populations. During model construction, exclude the influence of sex chromosomes (chrX and chrY) and mitochondrial chromosomes (chrM) to reduce artifact interference.

[0066] Next, inferCNV analysis was performed. Preset parameters for inferCNV were used, including filtering for low-expression genes, predicting CNV states using a hidden Markov model, and further subdividing potential tumor subgroups using graph clustering (Leiden algorithm). The analysis results were added to the metadata of the Seurat object via an ensemble function. The average CNV level for each cell on each chromosome was calculated by extracting the normalized CNV index column starting with `proportion_scaled_cnv_chr` from the metadata.

[0067] Calculate the mean CNV value for each cell on each chromosome and normalize it so that the CNV value is within the range of 0-1. Plot violin and box plots of the mean CNV level by cell clusters, as follows: Figure 3 As shown, based on a global initial threshold of 0.10, a sample-specific CNV determination threshold of 0.08 was obtained by adjusting the threshold according to the average CNV level distribution of each mesenchymal cell cluster in the same sample. Based on this final threshold, each cell was classified into two categories, and cells with CNV levels greater than the threshold were marked as malignant cells.

[0068] For paired adjacent normal tissue samples, the above analysis procedure was repeated. In these adjacent normal tissue samples, any EpCAM-positive cell clusters with a mean CNV level above a threshold were considered false-positive malignant cell clusters caused by factors such as cell type characteristics or focal non-tumor amplification. These cell clusters were removed from the previously defined malignant cell set. After this correction process, the final malignant epithelial cells, based on copy number variation and corrected for adjacent normal tissue, were obtained, such as... Figure 4 As shown.

[0069] 3) Malignant cell inference based on tumor-related transcriptional signature scores

[0070] Tissue block transcriptome data from 348 tumors and adjacent normal samples published in TCGA and GEO were collected, and differential expression analysis was performed using DESeq2. After adjustment, p-values ​​less than 0.001 were considered statistically significant. Under the premise that the fold change is greater than 1, differentially expressed genes are classified as follows: The genes were sorted from highest to lowest expression level. The top 50 highly expressed genes were selected to obtain the malignant characteristic gene set of lung adenocarcinoma, and the bottom 50 highly expressed genes were selected to obtain the non-malignant characteristic gene set.

[0071] For each cell in the single-cell transcriptome data of lung adenocarcinoma, the AddModuleScore function of the Seurat package was used to calculate the malignancy score and non-malignancy score of each EpCAM-positive epithelial cell based on the malignancy feature gene set and the non-malignancy feature gene set, respectively. The difference score between benign and malignant cells was defined by subtracting the malignancy score from the non-malignancy score and then normalized to the 0-1 range by linear scaling.

[0072] Unsupervised binary classification of EpCAM-positive epithelial cells was performed using the k-means clustering algorithm (k=2). Cells with lower normalized differential scores were defined as the initial "malignant epithelial cells," and the other group as the initial "non-malignant epithelial cells," obtaining the first round of benign / malignant determination results. Using the malignant and non-malignant epithelial cells obtained in the previous round as two groups, differential expression analysis was performed at the single-cell level. Cells were ranked by average expression difference and statistical significance. The top 50 upregulated genes from malignant epithelial cells were selected to update the "malignant characteristic gene set," and the top 50 upregulated genes from non-malignant epithelial cells were selected to update the "non-malignant characteristic gene set." Based on the updated characteristic gene sets, the AddModuleScore algorithm was called again to calculate the malignancy and non-malignancy scores for each cell, updating the benign / malignancy differential scores and performing k-means clustering again to obtain a new round of benign / malignant labels. After each iteration, the consistency ratio between the current and previous cell labels was calculated, with a target ratio of 95% or higher. After two iterations, the final determination result based on tumor-related transcriptional feature scores was obtained, as shown below. Figure 5 As shown.

[0073] 4) Inference of malignant cells based on allele-specific copy number variations

[0074] For lung adenocarcinoma tumor samples, the single-cell transcriptome BAM file, cell barcode file, population SNP reference variant file, and genetic map file were used as input to run the Numbat accompanying pileup_and_phase script. Allele readings and haplotype phasing were performed on heterozygous SNP loci across the entire genome, resulting in an allele count file containing reference allele / variant allele readings for each cell. This file served as the allele-level input for subsequent Numbat modeling.

[0075] The `run_numbat` function of `Numbat` is called to infer allele-specific copy number variations in the target tumor sample. Specifically, the UMI count matrix of the tumor sample is used as the input `count_mat` to provide copy number signals at the gene expression level; the count matrices of B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells are used as the input `ref_internal` to estimate the expression baseline of each gene in the diploid state.

[0076] The following settings were made for key parameters of CNV detection and filtering when modeling lung adenocarcinoma samples: parameter t was set to... This approach strikes a balance between fragment resolution and noise robustness while also enabling the detection of subclonal events. The parameter `gamma` is set to 20 to adapt to 10×Genomics UMI data; `min_cells` is set to 50 to ensure allele coverage; `multi_allelic` is set to TRUE to allow the identification of multi-allelic copy number variation events; `min_LLR` is set to 5 to retain only events with sufficient confidence for phylogenetic inference; `max_entropy` is initially set to 0.5 to suppress overly complex or highly uncertain CNV states; `call_clonal_loh` is set to TRUE to enhance the resolution of allele-level anomalies; and `init_k` is initially set to 3 to guide the initial partitioning of tumor clonal structures and subsequent optimization processes.

[0077] The specific values ​​of the above parameters can be dynamically adjusted according to the stage of the tumor to be analyzed, the estimated copy number variation load, and the clonal complexity: For tumor samples with low copy number variation load, suspected early stage, or relatively simple clonal structure, the metastasis probability can be appropriately increased, and max_entropy can be reduced to 0.4. At the same time, min_cells can be set to 30-50 and init_k can be set to 2-3 to improve the judgment criteria and avoid excessive division of rare CNVs; For tumor samples with highly complex clonal structure, suspected late stage, or polyclonal, t can be reduced to Increase max_entropy to 0.6-0.8, set init_k to 4-5, increase min_cells to 100, and decrease min_LLR to 3-4 to relax the filtering restrictions on complex CNV events and more fully characterize potential subclonal structures.

[0078] The tumor probability at the expression level (p_cnv_x), the tumor probability at the allele level (p_cnv_y), and the combined posterior tumor probability (p_cnv) for each cell are integrated into the Seurat object `meta.data` of the lung adenocarcinoma sample. Violin plots and box plots of these three probabilities are then drawn, grouped by cell clusters. Figure 6 As shown in the figure, cells that meet the criteria of p_cnv>0.90 and show a high probability of tumor formation at both the expression and allele levels are classified as malignant epithelial cells, while the remaining EpCAM-positive cells are considered non-malignant or benign cells, thus obtaining the following results. Figure 7 The results shown are from UMAP, which determines benign or malignant tumors based on allele-specific copy number variations.

[0079] 5) The final benign or malignant nature is determined by a comprehensive assessment of the module results.

[0080] For each cell in the single-cell data of lung adenocarcinoma, the results based on copy number variation, tumor score, and allele-specific copy number are summarized. If at least two pieces of evidence support that the cell is a malignant tumor cell, it is classified as a malignant cell; otherwise, it is classified as a benign or non-malignant cell. The results are as follows: Figure 8 The final result of precise identification of malignant cells is shown.

[0081] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made in accordance with the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for precise identification of benign and malignant cells based on multidimensional features of single-cell transcriptomes, comprising the following steps: 1) Single-cell transcriptome data preprocessing: The single-cell transcriptome data of the target tumor samples were normalized and high-variance genes were screened. Combined with principal component analysis and UMAP dimensionality reduction, the cells were divided into different clusters using a clustering algorithm. Cell type annotation was performed based on known marker genes, mainly divided into EpCAM-positive epithelial cells, mesenchymal cells, and immune cells. EpCAM-positive epithelial cells were selected as the candidate cell set for determining benignity or malignancy, and immune cells and mesenchymal cells were used as normal reference cell populations for subsequent copy number variation inference. 2) Malignant cell inference based on copy number variation: Copy number variation analysis is performed on candidate cells of the target tumor sample using the inferCNV algorithm. Immune cells are used as a normal reference cell population to infer the copy number variation of candidate cells. The sample-specific CNV judgment threshold is determined based on the average CNV level of the mesenchymal cell cluster. Cells with an average CNV level exceeding the CNV judgment threshold are marked as malignant cells. 3) Malignant cell inference based on tumor-related transcriptional feature scores: Based on the differences in transcriptome expression between tumor tissue and adjacent normal tissue, malignant and non-malignant feature gene sets are constructed, and the malignancy score of each cell is calculated using the Seurat tool; through multiple rounds of iterative updates to the feature gene set, the final determination of malignant and non-malignant cells is obtained; among which, The termination condition for multiple rounds of iterative updates of the feature gene set is: the consistency ratio between the current round and the previous round of the feature gene set reaches or exceeds a preset threshold or the number of iteration rounds reaches a preset upper limit; 4) Malignant cell inference based on allele-specific copy number variation: The Numbat tool is used to extract allele information from the target tumor sample to infer copy number variation; key parameters are set for modeling, and malignant cells are identified based on tumor probability, genotype status and clonal tag information. 5) Comprehensive determination of final benign or malignant: Combining the inference results based on copy number variation, allele-specific copy number variation and tumor-related transcriptional feature scores, a voting rule is used to make a final determination of the benign or malignant nature of EpCAM-positive epithelial cells.

2. The method for precise identification of benign and malignant cells as described in claim 1, characterized in that, In step 1), for single-cell transcriptome sequencing data of tumor tissue and / or adjacent normal tissue samples, a Seurat object containing a UMI counting matrix is ​​constructed. This object is then normalized and high-variance gene screening is performed. Principal component analysis, UMAP dimensionality reduction, and clustering algorithms are used to divide the cells into multiple cell clusters. Based on known marker genes, each cell cluster is annotated with its type, classifying the cells into at least EpCAM-positive epithelial cells, mesenchymal cells, and immune cells including B lymphocytes, T lymphocytes, natural killer cells, and myeloid cells. Based on the cell type annotation results, EpCAM-positive epithelial cells are selected from all cells and only retained to construct a candidate cell set for determining benignity or malignancy. At the same time, immune cells are used as the normal reference cell population in subsequent copy number variation inference.

3. The method for precise identification of benign and malignant cells as described in claim 2, characterized in that, Step 2) Extract the whole genome UMI count matrix and cell type annotation information from the Seurat object, and construct a gene sequence file containing gene chromosome positions and gene sequences in conjunction with the reference genome annotation file. Organize the whole genome UMI count matrix and cell type annotation information into the cell annotation file and expression matrix file required by the inferCNV algorithm. In the cell type annotation, pre-specify immune cells as normal reference cells, construct the inferCNV object and exclude sex chromosomes and mitochondrial chromosomes. Run inferCNV with preset parameters to filter low-expression genes, predict CNV status using the hidden Markov model and Leiden graph clustering, estimate the CNV level of each candidate cell on each chromosome segment, write the inferCNV output results into the Seurat object metadata, extract the CNV index column starting with proportion_scaled_cnv_chr to calculate the average CNV level of each cell on each chromosome, and determine the sample-specific CNV judgment threshold based on the average CNV level distribution of mesenchymal cell clusters on the global initial threshold. Cells with an average CNV level exceeding the CNV judgment threshold are marked as malignant cells.

4. The method for precise identification of benign and malignant cells as described in claim 3, characterized in that, In step 2), the global initial threshold is an average CNV level of 0.

10. The sample-specific CNV determination threshold is adjusted by observing the distribution of the average CNV level of each mesenchymal cell cluster in the same sample. For samples with paired adjacent normal tissue, the inferCNV analysis process is repeated on the single-cell transcriptome data of adjacent normal tissue. For EpCAM positive cell clusters with an average CNV level higher than the CNV determination threshold, they are regarded as false positive malignant cell clusters and removed from the malignant cell set to obtain the malignant epithelial cell set corrected by adjacent normal tissue.

5. The method for precise identification of benign and malignant cells as described in claim 2, characterized in that, Step 3) Perform differential expression analysis based on the transcriptome data of tumor tissue and adjacent normal tissue. After adjustment, the p-value should be less than the preset threshold. When the absolute value of the fold change exceeds a preset threshold, differentially expressed genes are classified according to... The genes with the highest fold change were sorted from high to low. A malignant characteristic gene set was constructed from the top few highly expressed genes in tumor tissue, and a non-malignant characteristic gene set was constructed from the top few highly expressed genes in normal tissue. In the single-cell transcriptome data from step 1), the Seurat AddModuleScore function was used to calculate a malignant score and a non-malignant score for each EpCAM-positive epithelial cell based on both the malignant and non-malignant characteristic gene sets. The difference between the malignant and non-malignant scores was defined by subtracting the malignant score from the non-malignant score, and the difference score was normalized to the 0-1 range. Based on the normalized difference score, the k-means unsupervised clustering algorithm was used to cluster the EpCAM cells. CAM-positive epithelial cells are binary classified to obtain an initial set of malignant and non-malignant epithelial cells. Based on this, single-cell differential expression analysis is performed on the malignant and non-malignant epithelial cells obtained in the current round of classification as two groups. The sets of malignant and non-malignant characteristic genes are updated. AddModuleScore calculation and k-means clustering are repeated. After each iteration, the consistency ratio between the current round and the previous round of classification results is calculated. The iteration is terminated when the consistency ratio reaches or exceeds a preset threshold or the number of iterations reaches a preset upper limit. The classification result at the time of termination is used as the benign or malignant determination result based on tumor-related transcriptional features.

6. The method for precise identification of benign and malignant cells as described in claim 5, characterized in that, The parameter settings for differential expression analysis in step 3) include: Differential expression analysis of tissue block transcriptome data was performed using DESeq2. After Benjamini–Hochberg correction, an adjusted p-value of less than 0.001 was considered the minimum acceptable value. A change in multiple greater than 1 is used as the screening threshold, and then... Sort by the ratio from highest to lowest; The top 50 highly expressed genes from tumor tissues were selected to construct a malignant characteristic gene set, and the top 50 highly expressed genes from normal tissues were selected to construct a non-malignant characteristic gene set. When iteratively updating the feature gene set, differential expression analysis at the single-cell level was used. The top 50 genes upregulated in malignant epithelial cells were selected as the new malignant feature gene set, and the top 50 genes upregulated in non-malignant epithelial cells were selected as the new non-malignant feature gene set. The iterative process calculates the consistency ratio between the benign and malignant labels of the current round and the previous round after each round. It terminates when the consistency ratio reaches or exceeds 95% or when the number of iterations reaches 10 rounds.

7. The method for precise identification of benign and malignant cells as described in claim 2, characterized in that, Step 4) For the target tumor sample, use its single-cell transcriptome alignment result BAM file, cell barcode file, population SNP reference variant file, and genetic map file as input to run the Numbat pileup_and_phase workflow. This extracts allele readings and performs haplotype phasing on heterozygous SNP sites across the entire genome, generating an allele count file containing reference allele and variant allele readings for each cell. Then, call the Numbat run_numbat function, using the UMI count matrix of the candidate cells in the tumor sample as the count_mat input and the UMI count matrix of the immune cells as the ref_internal input, setting parameters including t, gamma, and min_. Key modeling parameters, including cells, multi_allelic, min_LLR, max_entropy, call_clonal_loh, and init_k, were used to construct a hidden Markov model that jointly evaluates the expression level and allele level signals. Allele-specific copy number variation was inferred from tumor samples to obtain the tumor probability p_cnv_x at the expression level, the tumor probability p_cnv_y at the allele level, and the joint posterior tumor probability p_cnv for each cell. These probabilities were written into the Seurat object metadata. Based on a set p_cnv threshold, the distributions of p_cnv_x and p_cnv_y were combined to classify EpCAM-positive epithelial cells into malignant and non-malignant cells.

8. The method for precise identification of benign and malignant cells as described in claim 7, characterized in that, The key modeling parameters mentioned in step 4) include: Parameter t takes This is used to strike a balance between CNV fragment resolution and noise robustness; The parameter gamma is set to 20 to characterize the excessive dispersion of allele counts; The parameter min_cells is set to 50, which limits the minimum number of cells required to participate in pseudo-bulk modeling. Setting the parameter multi_allelic to TRUE allows the recognition of multiallelic copy number mutation events; The parameter min_LLR is set to 5, which is used to filter low-confidence CNV events according to the log-likelihood ratio; The parameter max_entropy is initially set to 0.5 to suppress overly complex or highly uncertain CNV states; Setting the parameter call_clonal_loh to TRUE enables the recognition of clonal heterozygosity loss states. The parameter init_k is initially set to 3 to specify the initial number of clones; The above parameters are adjusted within a preset range based on tumor stage, estimated copy number variation load, and clonal complexity to balance the sensitivity and specificity of CNV detection. The criteria for determining malignant cells include: setting a threshold of 0.90 for the combined posterior tumor probability p_cnv. When an EpCAM-positive epithelial cell satisfies p_cnv > 0.90 and shows a high tumor probability in both the expression level tumor probability p_cnv_x and the allele level tumor probability p_cnv_y, it is determined to be a malignant cell supported by the allele-specific copy number variation module.

9. The method for precise identification of benign and malignant cells as described in claim 1, characterized in that, In step 5), for each EpCAM-positive epithelial cell, the results of step 2) based on copy number variation, step 3) based on tumor-related transcriptional feature score and step 4) based on allele-specific copy number variation are summarized to form a benign / malignant determination vector containing three pieces of evidence. When at least two pieces of evidence support that the cell is a malignant tumor cell, it is finally determined to be a malignant cell; otherwise, it is determined to be a benign or non-malignant cell.

10. A computer program product or computer-readable storage medium, the computer program product comprising a computer program, the computer-readable storage medium storing the computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for precise identification of benign and malignant cells as described in any one of claims 1 to 9.