Method, device, electronic equipment and storage medium for single-cell transcriptome cell type automatic annotation based on consensus voting

By combining consensus voting and a large language model, the accuracy of single-cell transcriptome cell type annotation and the identification of rare cell types were solved, achieving higher annotation accuracy and robustness, and making it suitable for automated annotation of multiple cell types.

CN120954520BActive Publication Date: 2026-05-29GUANGZHOU NAT LAB

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU NAT LAB
Filing Date
2025-06-16
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing single-cell transcriptome cell type annotation methods are insufficient in terms of accuracy and ability to annotate rare cell types, especially when dealing with complex or rare cell types, they are prone to errors and omissions.

Method used

We employ a consensus-based voting approach, combining multiple cell type annotation methods and a large language model. We use a consensus-based voting algorithm to annotate single-cell transcriptome data by cell type, and then utilize the rich prior knowledge of the large language model to make corrections, thereby improving the accuracy and robustness of the annotation.

Benefits of technology

It effectively identifies rare cell types, broadens the application scope of cell type annotation, improves the accuracy and robustness of cell type annotation, reduces the error of a single method, and enhances the interpretability and reliability of annotation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954520B_ABST
    Figure CN120954520B_ABST
Patent Text Reader

Abstract

The application provides a single-cell transcriptome cell type automatic annotation method and device based on consensus voting, electronic equipment and storage medium, relates to the field of medical biotechnology, and integrates a plurality of initial annotation results obtained by a plurality of cell type annotation methods by applying an ensemble learning strategy, so as to reduce errors that may exist in a single annotation method, and improve the accuracy and robustness of cell type annotation. In addition, the method uses single-cell transcriptome data of a sample, combines rich prior knowledge and strong reasoning ability of a large language model, and has the ability to discover rare cell types, so as to effectively identify rare cell types, widen the application range of cell type annotation, and improve the general tissue annotation capability of cell types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pharmaceutical biotechnology, and in particular to a method, apparatus, electronic device, and storage medium for automated annotation of single-cell transcriptome cell types based on consensus voting. Background Technology

[0002] Cell type annotation is a core component of single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data analysis, and its accuracy directly determines the scientific value of downstream research. When comparing differences in cell type composition among different samples or conducting cell type-specific functional analyses, subtle deviations in annotation quality can trigger a cascade of errors.

[0003] Traditional manual annotation relies on researchers' prior knowledge of cell markers and a deep understanding of sequencing technology principles, which inherently suffers from high subjectivity, significant time consumption, and difficulty in scaling. With scRNA-seq technology being incorporated into the foundational technology stack of various large-scale research projects, developing efficient and reliable automated annotation tools has become an urgent need in the field of computational biology.

[0004] Current mainstream automated annotation methods can be broadly categorized into supervised learning and unsupervised learning paradigms based on whether they rely on external reference datasets. Supervised learning methods achieve label transfer by establishing a mapping between the reference dataset and the query dataset. Typical implementations include the scClassify algorithm based on random forests and scANVI, which employs a deep learning framework. However, these methods face several technical bottlenecks: first, the reference dataset needs to be sufficiently comprehensive in terms of cell type coverage and cell number, which is particularly challenging when studying rare cell types; second, the reference set and the target set need to maintain a certain degree of similarity, but this condition is often difficult to meet in scRNA-seq research. In contrast, unsupervised learning methods integrate multi-dimensional bioinformatics evidence to identify endogenous markers and then annotate cell types. These methods typically use class-specific genes or cell gene expression patterns in the dataset to be annotated, combined with cell type markers from authoritative databases (such as CellMarker and MarkerDB) or relevant literature, and employ various algorithms to assess cell types in cell populations or cells. However, while unsupervised methods avoid dependence on reference datasets, they also have limitations: First, not all cell types have well-defined and universally applicable marker genes, especially for some rare or newly discovered cell types, where there is a lack of sufficient marker genes; second, even for the same cell type, expression characteristics may differ at different time points or in different tissue environments, thus affecting the accurate differentiation of cell types.

[0005] In summary, existing cell type annotation methods have at least the following shortcomings:

[0006] 1) Insufficient accuracy of annotation results: The annotation results have a high error rate, especially when dealing with complex or rare cell types, which may lead to inaccurate classification or missed detection, thus affecting the reliability and accuracy of the overall annotation results.

[0007] 2) Insufficient annotation capabilities for rare cell types: Existing cell type annotation methods generally perform well when dealing with common cell types, but they are weak in recognizing rare or under-studied cell types, making it easy for these cell types to be missed or misclassified. Summary of the Invention

[0008] This invention provides a method, apparatus, electronic device, and storage medium for automated annotation of single-cell transcriptome cell types based on consensus voting, in order to address the deficiencies in related technologies.

[0009] This invention provides a method for automated cell type annotation of single-cell transcriptome based on consensus voting, comprising:

[0010] Based on multiple cell type annotation methods, cell type annotation is performed on the single-cell transcriptome data of the samples, and cell type consensus annotation results are obtained based on consensus voting algorithm. Cells of the same type are clustered to obtain each cell population and the category marker gene of each cell population.

[0011] Based on the clinical information of the samples and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population, and the predicted annotation results of each cell population are obtained.

[0012] Based on the predicted annotation results of each cell population, the cell type consensus annotation results of each cell population are corrected to obtain the cell type annotation results of the sample.

[0013] Preferably, the multiple cell type annotation methods include single-cell basic model annotation methods based on generative artificial intelligence (AI) and non-generative training annotation methods.

[0014] Preferably, the annotation method for single-cell basic models based on generative AI includes annotation methods based on large language models.

[0015] Preferably, non-generative training annotation methods include scATOMIC and CellAssign.

[0016] Preferably, based on multiple cell type annotation methods, cell type annotation is performed on the single-cell transcriptome data of the sample, including:

[0017] Based on the single-cell transcriptome data, standardized single-cell gene expression data were determined, and quality control analysis was performed on the standardized single-cell gene expression data to obtain quality control analysis results data.

[0018] Based on the aforementioned multiple cell type annotation methods, cell type annotations were performed on the quality control analysis results data.

[0019] Preferably, based on the single-cell transcriptome data, standardized single-cell gene expression data are determined, including:

[0020] Upstream processing software is used to convert the single-cell transcriptome data into standardized single-cell gene expression data for downstream analysis.

[0021] Preferably, based on the predicted annotation results of each cell population, the cell type consensus annotation results of each cell population are corrected, including:

[0022] If the proportion of mixed types in the target cell population is greater than the preset proportion, then the consensus annotation results of the cell types of the target cell population are filled in based on the predicted annotation results of the target cell population.

[0023] Preferably, based on the predicted annotation results of the target cell population, the cell type consensus annotation results of the target cell population are filled in, including:

[0024] If the target cell type in the predicted annotation results of the target cell population is a standard cell type, and the average gene activity assessment score of the target cell type is greater than the average gene activity assessment score of other standard cell types in the target cell population, the consensus annotation results of the cell type of the target cell population are supplemented based on the predicted annotation results of the target cell population.

[0025] Preferably, based on the predicted annotation results of each cell population, the cell type consensus annotation results of each cell population are corrected to obtain the cell type annotation results of the sample, and then the process further includes:

[0026] If there are potentially uncertain cell types in the cell type annotation results, the uncertain cell types in the cell type annotation results are determined based on the potential uncertain cell types and the gene activity assessment scores of the remaining cell types in the cell type annotation results.

[0027] Based on the number of uncertain cell types in the cell type annotation results and the total number of cell types in the cell type annotation results, the uncertainty score of the cell type annotation results is calculated;

[0028] The potential uncertain cell types are determined based on the mathematical distribution of gene activity assessment scores for each cell.

[0029] Preferably, the lower quartile of the gene activity assessment score for each cell of the potentially uncertain cell type is less than the upper quartile of the gene activity assessment score for each cell of half of the remaining cell types in the cell type annotation results of any sample.

[0030] Preferably, the gene activity assessment score of each cell is calculated based on the average expression value of the cell marker gene of the cell population in each cell.

[0031] Preferably, the cell marker genes corresponding to each cell population are obtained based on a large language model.

[0032] Preferably, the uncertain cell types in the cell type annotation results are determined based on the potential uncertain cell types and the gene activity assessment scores of the remaining cell types in the cell type annotation results, including:

[0033] A T-test was performed based on the gene activity assessment scores of a predetermined number of cells of the potentially uncertain cell type and the gene activity assessment scores of a predetermined number of cells of each remaining cell type in the cell type annotation results.

[0034] Calculate the difference between the mean score of the potentially uncertain cell type and the mean score of each remaining cell type, and determine the uncertain cell type based on the results of the T-test and the difference between the mean scores.

[0035] Preferably, the results of the T-test include each p-value.

[0036] Preferably, the uncertain cell type is determined based on the results of the T-test and the difference between the mean scores, including:

[0037] If all p-values ​​are less than the preset probability and the difference between the mean scores is greater than 0, then the potentially uncertain cell type is determined to be the accurate cell type.

[0038] Otherwise, the potentially uncertain cell type is defined as an uncertain cell type.

[0039] Preferably, based on the potential uncertain cell types and the gene activity assessment scores of cells of each remaining cell type in the cell type annotation results, the uncertain cell types in the cell type annotation results are determined, and the process further includes:

[0040] Replace the uncertain cell types in the cell type annotation results with mixed types and store them.

[0041] Preferably, based on a consensus voting algorithm, the cell type consensus annotation results are obtained, including:

[0042] Convert the cell types in the initial annotation results obtained by the various cell type annotation methods into standard cell types;

[0043] For any cell in the sample, the neighboring cells of the cell are determined. Based on the initial annotation results of the neighboring cells under the same cell type annotation method, a majority vote is performed on the cell to obtain the internal voting result.

[0044] A joint majority vote is performed on the internal voting results under different cell type annotation methods to obtain the type consensus annotation result for any given cell.

[0045] Preferably, the neighboring cells of any cell include a specified number of cells that are highly similar to any cell.

[0046] Preferably, based on the initial annotation results of the neighboring cells under the same cell type annotation method, a majority vote is performed on any one of the cells to obtain the internal voting results, including:

[0047] Based on the initial annotation results of the specified number of neighboring cells under the same cell type annotation method, a majority vote is performed on any one of the cells to obtain the internal voting result.

[0048] Preferably, each cell population in the sample includes a different number of class marker genes at different thresholds;

[0049] Based on the clinical information of the samples and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population, resulting in predicted annotation results for each cell population, including:

[0050] Based on the clinical information of the samples and the category marker genes of each cell population at each quantity threshold, the large language model is applied to annotate each cell population in each sample by cell type, and the alternative annotation results corresponding to each quantity threshold are obtained.

[0051] The candidate annotation results corresponding to each quantity threshold are subject to majority voting to obtain the predicted annotation results for each cell population.

[0052] This invention also provides an automated cell type annotation device for single-cell transcriptome based on consensus voting, comprising:

[0053] The consensus annotation module is used to annotate the single-cell transcriptome data of the sample with cell type based on multiple cell type annotation methods, and obtain the cell type consensus annotation results based on the consensus voting algorithm. It then clusters cells of the same type to obtain each cell population and the category marker genes of each cell population.

[0054] The large language model prediction module is used to annotate the cell types of each cell population based on the clinical information of the sample and the category marker genes of each cell population, and to obtain the prediction annotation results of each cell population.

[0055] The annotation result completion module is used to correct the cell type consensus annotation results of each cell population based on the predicted annotation results of each cell population, so as to obtain the cell type annotation results of the sample.

[0056] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for automated annotation of single-cell transcriptome cell types based on consensus voting as described above.

[0057] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automated annotation of single-cell transcriptome cell types based on consensus voting as described above.

[0058] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for automated annotation of single-cell transcriptome cell types based on consensus voting as described above.

[0059] This invention provides a method, apparatus, electronic device, and storage medium for automated single-cell transcriptome cell type annotation based on consensus voting. The method utilizes single-cell transcriptome data from samples and combines it with the rich prior knowledge and powerful reasoning capabilities of a large language model, thereby enabling the discovery of rare cell types. This effectively identifies rare cell types, broadens the application scope of cell type annotation, and improves the general tissue annotation capability of cell types. Furthermore, this method employs an ensemble learning strategy, integrating multiple initial annotation results obtained from various cell type annotation methods, reducing the errors that may exist with single annotation methods, and improving the accuracy and robustness of cell type annotation. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is one of the flowcharts illustrating the method for automated annotation of single-cell transcriptome cell types based on consensus voting provided by the present invention.

[0062] Figure 2 This is a schematic diagram comparing the results of the automatic cell type annotation method provided by this invention with CellAssign in lung cancer data.

[0063] Figure 3 This is a schematic diagram illustrating the application of the automatic cell type annotation method provided by this invention in skin cancer data.

[0064] Figure 4 This is a schematic diagram illustrating the application of the automatic cell type annotation method provided by this invention in cross-sequencing platform data.

[0065] Figure 5 This is the second flowchart of the method for automated annotation of single-cell transcriptome cell types based on consensus voting provided by the present invention.

[0066] Figure 6 This is a schematic diagram of the device for automated annotation of single-cell transcriptome cell types based on consensus voting provided by the present invention.

[0067] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0069] Cell type classification in biology is a complex problem, especially given the continuous changes in cell states. Cell type definitions are often imprecise, and even experienced experts may disagree on the specific phenotypes of certain cells. The continuity of cell states, the randomness of sequencing processes, and the background knowledge of personnel annotating manually can all lead to discrepancies in cell annotations within the same dataset. Furthermore, the identification of cell subtypes and the redefinition of marker genes between different datasets can also cause differences in cell type identification. Therefore, although various automated cell type annotation methods exist, there is no absolutely optimal method due to differences in cell type granularity, experimental interference factors, and the technical dependence and sparsity of gene expression.

[0070] To address the shortcomings of existing cell type annotation methods, such as insufficient accuracy and limited annotation capabilities for rare cell types, this invention provides a method for automated single-cell transcriptome cell type annotation based on consensus voting.

[0071] Figure 1 This is a flowchart illustrating a method for automated annotation of single-cell transcriptome cell types based on consensus voting, as provided in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0072] S1. Based on multiple cell type annotation methods, cell type annotation is performed on the single-cell transcriptome data of the sample, and cell type consensus annotation results are obtained based on consensus voting algorithm. Cells of the same type are clustered to obtain each cell population and the category marker gene of each cell population.

[0073] S2, Based on the clinical information of the sample and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population to obtain the predicted annotation results of each cell population;

[0074] S3. Based on the predicted annotation results of each cell population, the cell type consensus annotation results of each cell population are corrected to obtain the cell type annotation results of the sample.

[0075] Specifically, the automated cell type annotation method for single-cell transcriptome based on consensus voting provided in this embodiment of the invention is executed by an automated cell type annotation device based on single-cell transcriptome consensus voting. This device can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0076] First, perform step S1 to obtain single-cell transcriptome data for the sample. The sample can include one or more samples, and the category can be set as needed, such as pan-cancer samples; no specific limitation is made here. The single-cell transcriptome data for each sample can be pan-cancer single-cell transcriptome data included in the Gene Expression Omnibus (GEO) database. Here, single-cell transcriptome data refers to single-cell ribonucleic acid sequencing (scRNA-seq) data.

[0077] By utilizing existing cell type annotation methods, cell type annotation can be performed on single-cell transcriptome data of a sample, resulting in multiple initial annotation results for each cell in each sample. The existing cell type annotation methods applied in this invention include single-cell basic model annotation methods based on generative artificial intelligence and non-generative training annotation methods.

[0078] Single-cell annotation methods based on generative artificial intelligence utilize generatively pre-trained biological large-scale language models, aiming to achieve efficient transfer learning across tasks through large-scale pre-training and fine-tuning. In addition, non-generative training annotation methods mainly include scATOMIC and CellAssign. scATOMIC is a single-cell annotation tool based on a random forest model, while CellAssign is a method based on a probabilistic graphical model.

[0079] By inputting the single-cell transcriptome data of the sample into scATOMIC for cell type annotation, an initial annotation result can be obtained for each cell in each sample, which can be labeled as the first initial annotation result.

[0080] By compiling a list of marker genes for cell types in GEO tissues and combining it with single-cell transcriptome data from the samples, cell type annotation can be performed using the CellAssign package. This yields an initial annotation result for each cell in each sample, which can then be labeled as a second initial annotation result.

[0081] The single-cell transcriptome data of the samples are screened for highly variable genes. Then, cell type annotation is performed using the highly variable genes combined with a single-cell basic model annotation method based on generative artificial intelligence. This yields an initial annotation result for each cell in each sample, which can be labeled as the third initial annotation result.

[0082] Subsequently, the Seurat package was used to analyze the quality control analysis results, and then the cells in each sample were clustered to obtain the cell populations in each sample and the category marker genes of each cell population in each sample.

[0083] The Seurat package analyzes single-cell transcriptome data from samples as follows:

[0084] Normalization: The LogNormalize method was used to normalize the single-cell transcriptome data of the samples.

[0085] Highly variable gene screening: Select 3000 (can be manually modified) highly variable genes for subsequent analysis.

[0086] Data scaling: All gene data are scaled using the ScaleData function in the Seurat package. This step standardizes the gene expression data so that the mean is 0 and the variance is 1, to facilitate subsequent principal component analysis (PCA) and cluster analysis.

[0087] Parameter definition: Parameters are automatically selected based on the number of cells in the sample.

[0088] Principal component analysis: Calculate the PCA results.

[0089] Cluster analysis: Cell clustering is performed based on PCA results to distinguish different cell populations (clusters).

[0090] Category marker gene screening: Highly expressed positive genes in each cluster are selected as category marker genes.

[0091] Meanwhile, by leveraging the analysis results of single-cell transcriptome data of the samples using the Seurat package, a consensus voting algorithm is employed to conduct consensus voting on the initial annotation results of multiple automatic annotation methods, thereby obtaining the cell type consensus annotation results for each sample.

[0092] At this point, each sample will have a summary file containing cell type annotations, including the following information:

[0093] barcode: cell barcode.

[0094] scATOMIC_cl: The first initial annotation result from scATOMIC.

[0095] CellAssign_cl: The second initial annotation result from CellAssign.

[0096] scLLM-broad: Third initial annotation result from the single-cell basic model annotation method.

[0097] cellredefine: Consensus voting results for the cell type of the sample.

[0098] In summary, each sample yielded a .rda file containing a summary of cell type annotations, category marker genes, and Seurat meta information for subsequent analysis.

[0099] Subsequently, step S2 is executed, using the clinical information of each sample and the category marker genes of each cell population in each sample, applying the Large Language Model (LLM) to annotate the cell types of each cell population in each sample, and obtaining the predicted annotation results of each cell population.

[0100] The clinical information for each sample may include the type of the target disease, whether it has metastasized, and the tissue from which the sample was taken. A large language model, or simply a large model, refers to a Natural Language Processing (NLP) model with a massive number of parameters. The number of model parameters and / or the complexity of the model structure exceed a preset threshold. During training, the model processes large-scale text data and possesses the ability to understand and generate natural language.

[0101] Clinical information for each sample and the class marker genes for each cell population within that sample are input into a large language model. Using the cell type annotation R package based on the large language model, cell type annotation is performed on each cell population in each sample, yielding predicted annotation results. Here, each cell population in each sample corresponds to one predicted annotation result. The cell types in the predicted annotation results can include standard cell types, heterogeneous types (other), and unclassified cell types from the cell ontology database. Here, the standard cell type is a defined cell type.

[0102] Finally, in step S3, the predicted annotation results of each cell population can be used to correct the consensus annotation results of cell types for each cell population. Specifically, it is determined whether the predicted annotation results of each cell population are consistent with the consensus annotation results of cell types. If they are consistent, no action is taken. If they are inconsistent, the consensus annotation results of cell types for that cell population can be replaced with the predicted annotation results, or the consensus annotation results of cell types for that cell population and the predicted annotation results can be combined to determine a comprehensive annotation result, and this comprehensive annotation result can be used to replace the consensus annotation results of cell types for that cell population.

[0103] The method for automated single-cell transcriptome cell type annotation based on consensus voting provided in this invention employs an ensemble learning strategy to integrate multiple initial annotation results obtained from various cell type annotation methods. This reduces the potential errors of individual annotation methods and improves the accuracy and robustness of cell type annotation. Furthermore, this method utilizes single-cell transcriptome data from samples and combines the rich prior knowledge and powerful reasoning capabilities of a large language model, thereby enabling the discovery of rare cell types. This effectively identifies rare cell types, broadens the application scope of cell type annotation, and enhances the general tissue annotation capabilities of cell types.

[0104] Building upon the above embodiments, to reduce the difficulty of subsequent analysis processes and improve the quality of single-cell transcriptome data, upstream processing software, such as sequencing platform-specific tools (e.g., Cell Ranger for 10x Genomics and STARsolo for Smart-seq2), can be used to perform quality control and preprocessing of the single-cell transcriptome data before cell type annotation. This generates standardized single-cell gene expression data for subsequent analysis. Here, the standardized single-cell gene expression data can be a standardized single-cell gene expression count matrix.

[0105] Subsequently, scanpy can be used to perform quality control analysis on standardized single-cell gene expression data to obtain quality control analysis results.

[0106] The quality control analysis process may include the following steps:

[0107] 1) Preliminary filtering steps for low-quality cells and low-expression genes: When filtering cells, retain cells with at least 200 detected genes to help remove some poor-quality cells or cells with high background noise, preventing them from affecting subsequent analysis. When filtering genes, retain only genes expressed in at least 3 cells to remove low-frequency genes not expressed in most cells. Low-frequency genes may be noise or expressed at very low levels, and have no biological significance.

[0108] 2) The step of calculating and using the proportion of mitochondrial genes as an indicator to eliminate potentially stressed or dying cells: An excessively high proportion of mitochondrial genes usually indicates that the cells may be under stress or have begun to die, which may interfere with subsequent analysis. Therefore, the expression level of mitochondrial genes is often used as one of the quality control standards in single-cell RNA sequencing (RNA-seq). Mitochondrial genes are labeled by checking if their names begin with "MT-". Then, the sc.pp.calculate_qc_metrics function is used to calculate the quality control indicators for each cell, including the proportion of mitochondrial genes. Finally, cells with a mitochondrial gene proportion exceeding 10% are filtered out.

[0109] 3) Steps to further clean up standardized single-cell gene expression data and remove abnormal cells through outlier detection: outlier detection is performed on the number of genes detected in each cell, and abnormal cells are identified and filtered by using a 3-fold median bias.

[0110] 4) Steps to detect and remove doublets using the Scrublet algorithm to avoid data contamination: Further, the Scrublet algorithm is used to detect doublets in the standardized single-cell gene expression data. Here, a doublet refers to two cells whose mRNA was incorrectly mixed together during sequencing.

[0111] Subsequently, by using multiple cell type annotation methods to annotate the quality control analysis results data by cell type, the accuracy and reliability of multiple initial annotation results for each cell in each sample can be improved.

[0112] Based on the above embodiments, the step of correcting the cell type consensus annotation results of each cell population based on the prediction annotation results of each cell population includes:

[0113] If the proportion of mixed types in the target cell population is greater than the preset proportion, then the consensus annotation results of the cell types of the target cell population are filled in based on the predicted annotation results of the target cell population.

[0114] Specifically, here, the target cell population refers to the cell population whose consensus annotation results show a higher proportion of heterogeneous types than a preset proportion. The preset proportion can be set as needed, for example, it can be set to 0.5, or it can be set to other values; no specific limitation is made here.

[0115] If a target cell population exists, the predicted annotation results of the target cell population can be used to fill in the consensus annotation results of the cell types of the target cell population.

[0116] In this embodiment of the invention, the consensus annotation results of cell types for the target cell population are filled in, which can reduce the workload of annotation result correction while ensuring the accuracy and reliability of the cell type annotation results of the sample.

[0117] Based on the above embodiments, the step of supplementing the cell type consensus annotation results of the target cell population based on the predicted annotation results of the target cell population includes:

[0118] If the target cell type in the predicted annotation results of the target cell population is a standard cell type, and the average gene activity assessment score of the target cell type is greater than the average gene activity assessment score of other standard cell types in the target cell population, the consensus annotation results of the cell type of the target cell population are supplemented based on the predicted annotation results of the target cell population.

[0119] Specifically, if a target cell population exists, preset conditions need to be introduced. The filling operation will only be performed when the preset conditions are met. This ensures the significance of the filling operation and further improves the accuracy of the cell type annotation results of the filled target cell population.

[0120] Here, the preset conditions may include that the target cell type in the prediction annotation results of the target cell population is a standard cell type in the cell ontology database, and that the average gene activity assessment score of the target cell type in the target cell population is greater than the average gene activity assessment score of all other standard cell types in the target cell population.

[0121] Therefore, before imputation, it is necessary to determine whether the target cell type in the predicted annotation results of the target cell population is a standard cell type, and whether the mean gene activity assessment score of the target cell type is greater than the mean gene activity assessment score of other standard cell types in the target cell population.

[0122] Each cell of each cell type has a gene activity assessment score, and the mean gene activity assessment score for each cell type is the average of the gene activity assessment scores of all cells of that cell type. Under the preset conditions, the mean gene activity assessment score of the target cell type in the target cell population must be greater than the mean gene activity assessment score of all other standard cell types in the target cell population. Other standard cell types in the target cell population refer to cell types that are not heterogeneous or undefined, excluding the target cell type.

[0123] Here, the cell type annotation results for each sample can be used, combined with the cell marker genes provided by the large language model, to calculate the cell gene activity score (module score) using the quality control analysis data. For example, for each cell type in any sample, each cell type can be input into the large language model. By querying the large language model, it can provide the cell marker genes for each cell type; for example, each cell type may include 5 cell marker genes. Subsequently, using the cell marker genes for each cell type, the mean gene activity assessment score for each cell type can be calculated, i.e., the average gene activity assessment score of all cells in each cell type, and box plot visualization can be performed.

[0124] Taking the gene activity assessment score (modulescore) of each cell of the target cell type in the target cell population as an example, it can be calculated using the AddModuleScore function of the Seurat package. The calculation process may include:

[0125] Identify the cell marker genes of the target cell population to form the target gene set;

[0126] Calculate the expression value of the target gene set, that is, the average expression value of the target gene set in a single cell;

[0127] Calculate the expression values ​​of the control gene set, which is the average expression value of randomly selected control genes in the same cells.

[0128] The difference between the expression values ​​of the target gene set and the expression values ​​of the control gene set is used as the gene activity assessment score for a single cell.

[0129] In this embodiment of the invention, in addition to limiting the cell type consensus annotation results to be filled for the target cell population, it is also necessary to limit the cell type and gene activity assessment score. This can further reduce the workload of annotation result correction while ensuring the accuracy and reliability of the cell type annotation results of the sample.

[0130] Because many existing cell annotation methods lack quantitative assessment of the uncertainty of annotation results, researchers cannot effectively judge the reliability of the annotation results, thus affecting the reliability of subsequent analyses.

[0131] Based on this, and building upon the above embodiments, the cell type consensus annotation results for each cell population are revised based on the predicted annotation results of each cell population to obtain the cell type annotation results for the sample. This process further includes:

[0132] If there are potentially uncertain cell types in the cell type annotation results, the uncertain cell types in the cell type annotation results are determined based on the potential uncertain cell types and the gene activity assessment scores of the remaining cell types in the cell type annotation results.

[0133] Based on the number of uncertain cell types in the cell type annotation results and the total number of cell types in the cell type annotation results, the uncertainty score of the cell type annotation results is calculated;

[0134] The potential uncertain cell types are determined based on the mathematical distribution of gene activity assessment scores for each cell.

[0135] Specifically, after obtaining the cell type annotation results for the samples, the uncertainty of the cell type annotation results for each sample can be quantitatively assessed, thereby enhancing the interpretability and practicality of the cell type annotation results for each sample and providing more reliable support for further biological analysis and clinical applications.

[0136] First, determine whether there are potentially uncertain cell types in the cell type annotation results of the sample. The criterion can be to determine whether there is a cell type in any sample, and whether the lower quartile of the gene activity assessment score of each cell of that type is lower than the upper quartile of the gene activity assessment score of each cell of half of the remaining cell types in the cell type annotation results of any sample. If it is lower, then the type is determined to be a potentially uncertain cell type.

[0137] In cases where potentially uncertain cell types exist in the cell type annotation results of any sample, the gene activity assessment scores of these potentially uncertain cell types, along with the gene activity assessment scores of the remaining cell types in the cell type annotation results of any sample, can be used to determine the uncertain cell types in the cell type annotation results. This process assesses the reliability of the annotation for potentially uncertain cell types, or in other words, whether the potentially uncertain cell types are accurate. If a potentially uncertain cell type is inaccurate, i.e., it exhibits uncertainty, then that potentially uncertain cell type is defined as an uncertain cell type.

[0138] Here, the gene activity assessment scores of cells with potential uncertainty can be compared with the gene activity assessment scores of cells of each remaining cell type. T-tests can be performed separately, and the results of the T-tests can be calculated, as well as the difference between the mean scores.

[0139] Each potentially uncertain cell type and each remaining cell type can have a corresponding T-test result and score mean difference. By using the results of the T-test and the score mean differences, we can determine whether the annotation of the potentially uncertain cell type is reliable, and thus define the uncertain cell type.

[0140] Subsequently, using the number of uncertain cell types and the total number of cell types in the cell type annotation results of any sample, the uncertainty score of the cell type annotation results for any sample is calculated. If the number of uncertain cell types in the cell type annotation results of any sample is n, and the total number of cell types in the cell type annotation results of any sample is N, then the uncertainty score of the cell type annotation results for any sample can be expressed as n / N.

[0141] In this embodiment of the invention, the uncertainty of cell type annotation results of samples is quantitatively evaluated, thereby enhancing the interpretability and practicality of cell type annotation results for each sample and providing more reliable support for further biological analysis and clinical applications.

[0142] Based on the above embodiments, determining the uncertain cell type in the cell type annotation results based on the gene activity assessment scores of the cells of the remaining cell types in the cell type annotation results, including:

[0143] A T-test was performed based on the gene activity assessment scores of a predetermined number of cells of the potentially uncertain cell type and the gene activity assessment scores of a predetermined number of cells of each remaining cell type in the cell type annotation results.

[0144] Calculate the difference between the mean score of the potentially uncertain cell type and the mean score of each remaining cell type, and determine the uncertain cell type based on the results of the T-test and the difference between the mean scores.

[0145] Specifically, in this embodiment of the invention, a preset number of cells can be selected from the potentially uncertain cell types and each remaining cell type. The gene activity scores of these cells are then used to perform a T-test, yielding the results of each T-test, and the mean difference of each score is calculated. The preset number can be set as needed. If the number of cells in the potentially uncertain cell types is greater than or equal to 50, the preset number can be set to 50. If the number of cells in the potentially uncertain cell types is less than 50, the preset number can be set to be equal to the number of cells in the uncertain cell types.

[0146] Each potentially uncertain cell type and each remaining cell type can have a corresponding T-test result and score mean difference. By using the results of the T-test and the score mean differences, we can determine whether the annotation of the potentially uncertain cell type is reliable, and thus define the uncertain cell type.

[0147] In this embodiment of the invention, by combining the results of the T-test and the difference in the mean scores, the uncertain cell types in the cell type annotation results can be quickly and accurately determined.

[0148] Based on the above embodiments, each result of the T-test includes each p-value of the T-test.

[0149] Based on the above embodiments, the uncertain cell type is determined based on the results of the T-test and the difference between the mean scores, including:

[0150] If all p-values ​​are less than the preset probability and the difference between the mean scores is greater than 0, then the annotation of the potentially uncertain cell type is determined to be credible and an accurate cell type.

[0151] Otherwise, the potentially uncertain cell type is defined as an uncertain cell type.

[0152] Specifically, when determining the certainty of a potentially uncertain cell type, if all p-values ​​are less than the preset probability and the difference between the mean scores is greater than 0, then the potentially uncertain cell type is determined to be an accurate cell type. The preset probability can be set as needed, for example, it can be set to 0.05, or it can be set to other values; no specific limitation is made here.

[0153] Otherwise, if there is a p-value greater than or equal to the preset probability, or a difference in the mean score less than or equal to 0, then the potentially uncertain cell type can be determined as an uncertain cell type.

[0154] In this embodiment of the invention, by using the difference between each p-value and the mean score, it is possible to quickly and accurately determine whether a potentially uncertain cell type is an accurate cell type.

[0155] Based on the above embodiments, the uncertain cell types in the cell type annotation results are determined based on the gene activity assessment scores of the cells of the remaining cell types in the cell type annotation results, and then the process further includes:

[0156] Replace the uncertain cell types in the cell type annotation results of any of the samples with mixed types, and save the results to the results file.

[0157] Specifically, after determining the reliability of the annotations for potentially uncertain cell types, the uncertain cell types in the annotation results for each sample can be identified and replaced with a heterogeneous type, i.e., "other." The cell type annotation results for each sample are then stored as a results file. This results file can be in CSV format for subsequent selective analysis.

[0158] Existing cell type annotation methods suffer from inconsistencies in labeling the same cell population due to differences in annotation standards, datasets, and algorithms. This inconsistency limits the comparability and accuracy of cell type annotation results.

[0159] Based on this, and building upon the above embodiments, a consensus annotation result for cell types is obtained using a consensus voting algorithm, including:

[0160] Convert the cell types in the initial annotation results obtained by the various cell type annotation methods into standard cell types;

[0161] For any cell in the sample, the neighboring cells of the cell are determined. Based on the initial annotation results of the neighboring cells under the same cell type annotation method, a majority vote is performed on the cell to obtain the internal voting result.

[0162] A joint majority vote is performed on the internal voting results under different cell type annotation methods to obtain the type consensus annotation result for any given cell.

[0163] Specifically, in the process of consensus voting on multiple initial annotation results, the quality control analysis results, first initial annotation results, second initial annotation results, and third initial annotation results of each sample can be read separately according to the sample label.

[0164] Then, the cell types in the initial annotation results corresponding to the non-generative training annotation method are converted into standard cell types in the cell ontology database, thereby maintaining the consistency and comparability of the cell type annotation results. It is understandable that the single-cell basic model annotation method based on generative AI, because it uses standard cell types as labels for training, can be considered to output prediction annotation results belonging to the standard cell types in the cell ontology database.

[0165] Subsequently, for any cell in each sample, the neighboring cells of that cell can be determined. The neighboring cells can be a specified number of cells with high similarity to that cell. The specified number can be set as needed, such as one or more, and is not limited here.

[0166] The `FindNeighbors` function in the `Seurat` package can be used to calculate the similarity between any cell and other cells. Then, the similarity and neighbor information between any given cell and other cells are extracted, and a specified number of neighboring cells with high similarity to any given cell are counted. If multiple numbers are specified, they can be set to 9, or other values.

[0167] By using the initial annotation results of neighboring cells under the same cell type annotation method, a majority vote can be performed on any cell to obtain the internal voting result. This majority vote is a self-vote among cell type annotation methods, belonging to the internal voting under the same cell type annotation method.

[0168] If the most frequent cell type accounts for more than 0.5% of the total votes, then that cell type is considered as the voting result for any cell under the same cell type annotation method; otherwise, it is defined as a mixed type. If a mixed type or uncertain cell type appears, then during the voting process, the total number of votes needs to be reduced to remove mixed or uncertain cell types from the total number of votes, and the cell type is redefined based on the voting results.

[0169] Subsequently, a joint majority vote is performed on the voting results of any cell under different cell type annotation methods to obtain the type consensus annotation result for that cell. If there are three different cell type annotation methods, each cell type annotation method corresponds to an internal voting result, and a joint majority vote is required on the three internal voting results to obtain the type consensus annotation result for that cell. For example, if the internal voting result for a cell under all three cell type annotation methods is "B cell", since the proportion of votes for "B cell" is 3 / 3 = 1, which is greater than 0.5, the cell can be defined as "B cell".

[0170] Additionally, if the internal voting result of a certain cell type annotation method is a mixed type, then during the majority vote, it is also necessary to perform a vote reduction operation on the total number of votes in order to remove the mixed type from the total number of votes.

[0171] Finally, the cell type consensus annotation results for each sample are constructed using the cell type consensus annotation results for each cell in each sample.

[0172] In this embodiment of the invention, a process is introduced to convert cell types in the initial annotation results corresponding to non-generative training annotation methods into standard cell types in a cell ontology database, which can maintain the consistency and comparability of cell type annotation results. Moreover, through self-voting of the same cell type annotation method and joint majority voting of different cell type annotation methods, the cell type annotation results of each sample can be made more accurate and reliable.

[0173] Based on the above embodiments, each cell population in the sample includes a different number of category marker genes; the process of annotating cell types for each cell population using a large language model based on the clinical information of the sample and the category marker genes of each cell population, and obtaining the predicted annotation results for each cell population, includes:

[0174] Based on the clinical information of the samples and the category marker genes of each cell population at each quantity threshold, the large language model is applied to annotate each cell population in each sample by cell type, and the alternative annotation results corresponding to each quantity threshold are obtained.

[0175] The candidate annotation results corresponding to each quantity threshold are subject to majority voting to obtain the predicted annotation results for each cell population.

[0176] Based on the above embodiments, during the consensus voting process of multiple initial annotation results, for any sample, tryCatch can be used to catch possible errors. If an error occurs, the error information will be written to the log, and that sample will be skipped, and the cell types of subsequent samples will be automatically annotated.

[0177] Based on the above embodiments, each cell population in each sample includes a different number of class marker genes at different thresholds;

[0178] Based on the clinical information of the samples and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population, resulting in predicted annotation results for each cell population, including:

[0179] Based on the clinical information of the samples and the category marker genes of each cell population at each quantity threshold, the large language model is applied to annotate each cell population in each sample by cell type, and the alternative annotation results corresponding to each quantity threshold are obtained.

[0180] The candidate annotation results corresponding to each quantity threshold are subject to majority voting to obtain the predicted annotation results for each cell population.

[0181] Specifically, each cell population in each sample can include different number thresholds of class marker genes. For example, different number thresholds can include 3, and can take values ​​of 10, 30, and 50 respectively.

[0182] Furthermore, when applying the class marker genes of each cell population in each sample, the clinical information of each sample and the class marker genes of each cell population in each sample at each quantity threshold can be input into the large language model. The large language model then performs cell type annotation on each cell population in each sample, obtaining candidate annotation results corresponding to each quantity threshold. If different quantity thresholds can include three, then each cell population in each sample can have three candidate annotation results.

[0183] Subsequently, a majority vote was conducted on the candidate annotation results corresponding to each quantity threshold to obtain the predicted annotation results for each cell population in each sample.

[0184] In this embodiment of the invention, by setting different thresholds for category marker genes, the reliability of the prediction annotation results output by the large language model can be guaranteed.

[0185] The performance of the automatic cell type annotation method provided in this embodiment of the invention is verified by using real data from cancer patients as single-cell transcriptome data of different samples.

[0186] (a) The real data from cancer patients came from 20 lung cancer patients, including 10 brain metastasis samples, 7 lymph node metastasis samples, and pleural effusion samples. The sequencing technology used for validation was the 10x Genomics platform.

[0187] First, standardized single-cell gene expression data from lung cancer patients underwent preprocessing, i.e., quality control analysis, as detailed in the above embodiments, and will not be repeated here. This preprocessing ensures the reliability of subsequent analyses by screening cells and genes to retain high-quality cells for later analysis.

[0188] Then, the automatic cell type annotation method provided in the embodiments of the present invention is used to annotate the cell types of lung cancer patients.

[0189] The main steps include:

[0190] 1) Annotate cells using multiple cell type annotation methods respectively;

[0191] 2) Use the Seurat package to read the matrix data after quality control analysis, identify neighboring cells, and integrate the initial annotation results of multiple cell type annotation methods by consensus voting;

[0192] 3) Using the category marker genes obtained from the Seurat package, cell type annotation is assisted by a large language model to obtain the predicted annotation results;

[0193] 4) Aggregate the predicted annotation results with the cell type annotation results of each sample obtained by consensus voting to finally obtain the type annotation results of each cell.

[0194] Figure 2 (A) is a comparison graph of the type annotation results (LLM-ConV) of each cell obtained by the automatic cell type annotation method provided by the present invention and the actual annotation results (Original) in real data of lung cancer patients; Figure 2 (B) is a comparison graph of the cell type annotation results (CellAssign) obtained by the CellAssign method with the actual annotation results (Original) in real data from lung cancer patients.

[0195] like Figure 2As shown in (A), the cell type annotation results (LLM-ConV) obtained by the automatic cell type annotation method provided in this embodiment of the invention are compared with the actual annotation results (Original) in the real data of lung cancer patients, and the two are highly consistent.

[0196] like Figure 2 As shown in (B), when the cell type annotation results (CellAssign) obtained by the CellAssign method are compared with the actual annotation results (Original) in the real data of lung cancer patients, there are significant differences in some cell types.

[0197] Therefore, compared with the CellAssign method, the automatic cell type annotation method provided in this embodiment of the invention performs better in terms of annotation accuracy. For example, when comparing the type annotation results of epithelial cells with the actual annotation results in real data from lung cancer patients, the consistency between the automatic cell type annotation method provided in this embodiment of the invention and the annotation results in real data from lung cancer patients reaches 0.9, while the similarity of the CellAssign method is only 0.7. More importantly, the CellAssign method incorrectly annotates oligodendrocytes as T / NK cell types, while the type annotation results of the automatic cell type annotation method provided in this embodiment of the invention successfully avoid this error.

[0198] Conclusion: The automatic cell type annotation method provided in this embodiment of the invention demonstrates superior accuracy in annotation. Comparison with actual annotation results from real data of lung cancer patients and with the performance of the CellAssign method shows that the automatic cell type annotation method provided in this embodiment of the invention offers more accurate and reliable cell type annotation results.

[0199] (ii) The real data from cancer patients came from 5 patients with basal cell carcinoma, and all samples were primary tumors. The sequencing technology used for validation was the 10x Genomics platform.

[0200] By downloading the raw data from the PRJNA753840 project and performing quantitative analysis using Cell Ranger, standardized single-cell gene expression data from basal cell carcinoma patients were ultimately obtained.

[0201] First, standardized single-cell gene expression data from basal cell carcinoma patients underwent preprocessing, i.e., quality control analysis, as detailed in the above embodiments, and will not be repeated here. This preprocessing ensures the reliability of subsequent analyses by screening cells and genes to retain high-quality cells for later analysis.

[0202] Then, the automatic cell type annotation method provided in the embodiments of the present invention is used to annotate the cell types of basal cell carcinoma patients.

[0203] Figure 3 (A) is a diagram showing the cell type annotation results obtained in the GSM5514166 sample using the automatic cell type annotation method provided by this invention; Figure 3 (B) is a gene activity assessment score map of the marker gene obtained by applying the automatic cell type annotation method provided by the present invention to the GSM5514166 sample.

[0204] like Figure 3 As shown in (A), a specific cell type, melanocyte, was successfully annotated in the GSM5514166 sample. This is a characteristic cell type of skin cancer. Figure 3 In (A), the coordinate axes UMAP_1 and UMAP_2 are abstract coordinate axes of the data after dimensionality reduction in two-dimensional space, representing the relative positional relationship of the samples in the high-dimensional structure.

[0205] It is worth noting that neither CellAssign nor scATOMIC could annotate this cell type, even though this cell type was also annotated in real data from basal cell carcinoma patients, which supports the cell type annotation results obtained by the automatic cell type annotation method provided in this embodiment of the invention.

[0206] Furthermore, the gene activity assessment score of the marker genes associated with Melanocyte cells in each cell was calculated using the AddModuleScore function. Figure 3 As shown in (B), the Melanocyte cells annotated by the automatic cell type annotation method provided in this embodiment of the invention have significantly higher gene activity assessment scores than other cell types. Figure 3 In (B), the horizontal axis represents the cell types annotated, and the vertical axis represents the score.

[0207] Conclusion: The automated cell type annotation method provided in this embodiment of the invention can be successfully applied to basal cell carcinoma data and effectively identify Melanocyte cells that are specifically associated with skin cancer. This result indicates that the method can not only effectively annotate common cell types, but also identify and annotate tissue-specific and even rare cell types, demonstrating broad applicability and flexibility.

[0208] (III) The real data of cancer patients came from 30 lung cancer patients, including 49 samples, of which 19 were primary tumors, 2 were of unknown type, and 28 were metastatic tumors. The sequencing technology used for validation was the Smart-seq platform.

[0209] By downloading the raw data from the PRJNA591860 project and performing quantitative analysis using htseq-count, standardized single-cell gene expression data from lung cancer patients were ultimately obtained.

[0210] First, gene IDs in the standardized single-cell gene expression data of lung cancer patients were converted into gene symbols. Then, the count value of each gene was standardized based on gene length. Following this, preprocessing, i.e., quality control analysis, was performed, as detailed in the above embodiment, and will not be repeated here. This preprocessing ensures the reliability of subsequent analyses by screening cells and genes to retain high-quality cells for later analysis. After the above processing, 9481 cells were ultimately selected for subsequent analysis.

[0211] Then, the automatic cell type annotation method provided in this embodiment of the invention is used to annotate the lung cancer data on the Smart-seq platform.

[0212] The main steps include:

[0213] 1) Annotate cells using multiple cell type annotation methods respectively;

[0214] 2) Use the Seurat package to read the matrix data after quality control analysis, identify neighboring cells, and integrate the initial annotation results of multiple cell type annotation methods by consensus voting;

[0215] 3) Using the category marker genes obtained from the Seurat package, cell type annotation is assisted by a large language model to obtain the predicted annotation results;

[0216] 4) Aggregate the predicted annotation results with the cell type annotation results of each sample obtained by consensus voting to finally obtain the type annotation results of each cell.

[0217] Figure 4 (A) is a graph showing the cell type annotation results of lung cancer data on the Smart-seq platform using the automatic cell type annotation method provided by this invention; Figure 4 (B) is a heatmap of marker gene expression results of cell type annotation of lung cancer data on the Smart-seq platform using the automatic cell type annotation method provided by the present invention. Figure 4 In (A), the coordinate axes UMAP_1 and UMAP_2 are abstract coordinate axes of the dimensionality-reduced data in two-dimensional space, representing the relative positional relationship of the lung cancer data of the Smart-seq platform in the high-dimensional structure.

[0218] like Figure 4As shown in (A), the annotated cells are clustered according to cell type, with cell groups of similar biological backgrounds being closer together, such as NK cells and T cells, consistent with existing biological prior knowledge. Furthermore, the reliability of the annotation results can be further verified by examining the expression of marker genes in each annotated cell type, such as... Figure 4 As shown in (B), each cell type highly expresses its specific marker gene, further supporting the accuracy of the type annotation results of the automatic cell type annotation method provided in the embodiments of the present invention. Figure 4 In (B), the horizontal axis represents marker genes, and the vertical axis represents cell types.

[0219] Conclusion: The automated cell type annotation method provided in this embodiment of the invention can be successfully applied to lung cancer single-cell transcriptome data generated by Smart-seq technology. Validation of the cell type annotation results shows that the annotated cell types conform to prior biological knowledge, indicating that this method is not only applicable to data from the 10x Genomics platform but also possesses robustness and broad applicability across sequencing platforms.

[0220] In summary, such as Figure 5 As shown in the embodiments of the present invention, the method for automated annotation of single-cell transcriptome cell types based on consensus voting can obtain quality control analysis results data by standardizing and performing quality control analysis on single-cell transcriptome data for the target disease.

[0221] Then, based on multiple cell type annotation methods, the cell types in the quality control analysis results data were annotated, resulting in multiple initial annotation results for each cell in each sample. Simultaneously, the Seurat package was used to analyze the quality control analysis results data, yielding the analysis results.

[0222] On the one hand, the analysis results can be used to conduct consensus voting on multiple initial annotation results, and a consensus result for cell type can be obtained for each sample.

[0223] On the other hand, by using the analysis results, we can combine the cell type library of the large language model to annotate the cell types of each cell population in each sample under different thresholds of category marker genes, and obtain the predicted annotation results.

[0224] Finally, the cell type annotation results obtained from consensus voting are aggregated with the predicted annotation results obtained from the large language model. That is, the cell type annotation results of the target cell population are filled in using the predicted annotation results of the target cell population. In addition, a judgment process based on preset conditions can be introduced before filling in the annotation results.

[0225] like Figure 6As shown, based on the above embodiments, this embodiment of the invention provides an automated cell type annotation device based on single-cell transcriptome consensus voting, comprising:

[0226] The consensus annotation module 61 is used to annotate the single-cell transcriptome data of the sample with cell types based on multiple cell type annotation methods, and obtain cell type consensus annotation results based on consensus voting algorithm, and cluster cells of the same type to obtain each cell population and the category marker gene of each cell population.

[0227] The large language model prediction module 62 is used to annotate the cell types of each cell population based on the clinical information of the sample and the category marker genes of each cell population, and to obtain the prediction annotation results of each cell population.

[0228] The annotation result correction module 63 is used to correct the cell type consensus annotation result of each cell population based on the predicted annotation result of each cell population, so as to obtain the cell type annotation result of the sample.

[0229] Based on the above embodiments, the annotation result correction module is specifically used for:

[0230] If the proportion of mixed types in the target cell population is greater than the preset proportion, then the consensus annotation results of the cell types of the target cell population are filled in based on the predicted annotation results of the target cell population.

[0231] Based on the above embodiments, the annotation result correction module is further used for:

[0232] If the target cell type in the predicted annotation results of the target cell population is a standard cell type, and the average gene activity assessment score of the target cell type is greater than the average gene activity assessment score of other standard cell types in the target cell population, the consensus annotation results of the cell type of the target cell population are supplemented based on the predicted annotation results of the target cell population.

[0233] Based on the above embodiments, a module for quantitative evaluation of annotation results is also included, used for:

[0234] If there are potentially uncertain cell types in the cell type annotation results, the uncertain cell types in the cell type annotation results are determined based on the potential uncertain cell types and the gene activity assessment scores of the remaining cell types in the cell type annotation results.

[0235] Based on the number of uncertain cell types in the cell type annotation results and the total number of cell types in the cell type annotation results, the uncertainty score of the cell type annotation results is calculated;

[0236] The potential uncertain cell types are determined based on the mathematical distribution of gene activity assessment scores for each cell.

[0237] Based on the above embodiments, the annotation result quantification and evaluation module is further used for:

[0238] A T-test was performed based on the gene activity assessment scores of a predetermined number of cells of the potentially uncertain cell type and the gene activity assessment scores of a predetermined number of cells of each remaining cell type in the cell type annotation results.

[0239] Calculate the difference between the mean score of the potentially uncertain cell type and the mean score of each remaining cell type, and determine the uncertain cell type based on the results of the T-test and the difference between the mean scores.

[0240] Based on the above embodiments, the consensus annotation module is specifically used for:

[0241] Convert the cell types in the initial annotation results obtained by the various cell type annotation methods into standard cell types;

[0242] For any cell in the sample, the neighboring cells of the cell are determined. Based on the initial annotation results of the neighboring cells under the same cell type annotation method, a majority vote is performed on the cell to obtain the internal voting result.

[0243] A joint majority vote is performed on the internal voting results under different cell type annotation methods to obtain the type consensus annotation result for any given cell.

[0244] Based on the above embodiments, each cell population includes category marker genes with different number thresholds; the large language model prediction module is specifically used for:

[0245] Based on the clinical information of the samples and the category marker genes of each cell population at each quantity threshold, the large language model is applied to annotate the cell types of each cell population to obtain the alternative annotation results corresponding to each quantity threshold.

[0246] The candidate annotation results corresponding to each quantity threshold are subject to majority voting to obtain the predicted annotation results for each cell population.

[0247] Specifically, the functions of each module in the automated cell type annotation device based on single-cell transcriptome consensus voting provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-mentioned method-like embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0248] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the consensus-voting-based automated annotation method for single-cell transcriptome cell types provided in the above embodiments.

[0249] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0250] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the method for automated annotation of single-cell transcriptome cell types based on consensus voting provided in the above embodiments.

[0251] In another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automated annotation of single-cell transcriptome cell types based on consensus voting provided in the above embodiments. This computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transient computer-readable storage medium, and is not specifically limited herein.

[0252] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0253] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0254] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automated cell type annotation of single-cell transcriptome based on consensus voting, characterized in that, include: Based on multiple cell type annotation methods, cell type annotation is performed on the single-cell transcriptome data of the samples, and cell type consensus annotation results are obtained based on consensus voting algorithm. Cells of the same type are clustered to obtain each cell population and the category marker gene of each cell population. Based on the clinical information of the samples and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population, resulting in predicted annotation results for each cell population; the clinical information includes the type of the target disease, whether it has metastasized, and the sampling tissue of the sample; the target disease corresponds to the sample; Based on the predicted annotation results of each cell population, the cell type consensus annotation results of each cell population are corrected to obtain the cell type annotation results of the sample. The process of correcting the cell type consensus annotation results for each cell population based on the predicted annotation results of each cell population includes: If the proportion of mixed types in the target cell population is greater than the preset proportion, then the cell type consensus annotation result of the target cell population is filled based on the predicted annotation result of the target cell population; The process of supplementing the cell type consensus annotation results of the target cell population based on the predicted annotation results of the target cell population includes: If the target cell type in the predicted annotation results of the target cell population is a standard cell type, and the average gene activity assessment score of the target cell type is greater than the average gene activity assessment score of other standard cell types in the target cell population, the consensus annotation results of the cell type of the target cell population are supplemented based on the predicted annotation results of the target cell population.

2. The method for automated cell type annotation of single-cell transcriptome based on consensus voting according to claim 1, characterized in that, The process involves correcting the cell type consensus annotation results for each cell population based on the predicted annotation results of each cell population to obtain the cell type annotation results for the sample, and then further includes: If there are potentially uncertain cell types in the cell type annotation results, the uncertain cell types in the cell type annotation results are determined based on the potential uncertain cell types and the gene activity assessment scores of the remaining cell types in the cell type annotation results. Based on the number of uncertain cell types in the cell type annotation results and the total number of cell types in the cell type annotation results, the uncertainty score of the cell type annotation results is calculated; The potential uncertain cell types are determined based on the mathematical distribution of gene activity assessment scores for each cell.

3. The method for automated cell type annotation of single-cell transcriptome based on consensus voting according to claim 2, characterized in that, The determination of uncertain cell types in the cell type annotation results based on the gene activity assessment scores of the cells of the remaining cell types in the cell type annotation results, including: A T-test was performed based on the gene activity assessment scores of a predetermined number of cells of the potentially uncertain cell type and the gene activity assessment scores of a predetermined number of cells of each remaining cell type in the cell type annotation results. Calculate the difference between the mean score of the potentially uncertain cell type and the mean score of each remaining cell type, and determine the uncertain cell type based on the results of the T-test and the difference between the mean scores.

4. The method for automated cell type annotation of single-cell transcriptome based on consensus voting according to any one of claims 1-3, characterized in that, The consensus annotation results for cell types obtained based on the consensus voting algorithm include: Convert the cell types in the initial annotation results obtained by the various cell type annotation methods into standard cell types; For any cell in the sample, the neighboring cells of the cell are determined. Based on the initial annotation results of the neighboring cells under the same cell type annotation method, a majority vote is performed on the cell to obtain the internal voting result. A joint majority vote is performed on the internal voting results under different cell type annotation methods to obtain the type consensus annotation result for any given cell.

5. The method for automated annotation of single-cell transcriptome cell types based on consensus voting according to any one of claims 1-3, characterized in that, Each cell population includes a different number of category marker genes at different thresholds; Based on the clinical information of the samples and the category marker genes of each cell population, a large language model is applied to annotate the cell types of each cell population, resulting in predicted annotation results for each cell population, including: Based on the clinical information of the samples and the category marker genes of each cell population at each quantity threshold, the large language model is applied to annotate the cell types of each cell population to obtain the alternative annotation results corresponding to each quantity threshold. The candidate annotation results corresponding to each quantity threshold are subject to majority voting to obtain the predicted annotation results for each cell population.

6. An automated cell type annotation device for single-cell transcriptome based on consensus voting, characterized in that, include: The consensus annotation module is used to annotate the single-cell transcriptome data of the sample with cell type based on multiple cell type annotation methods, and obtain the cell type consensus annotation results based on the consensus voting algorithm. It then clusters cells of the same type to obtain each cell population and the category marker genes of each cell population. The large language model prediction module is used to annotate the cell types of each cell population based on the clinical information of the sample and the category marker genes of each cell population, and to obtain the prediction annotation results of each cell population; the clinical information includes the type of the target disease, whether it has metastasized, and the sampling tissue of the sample; the target disease corresponds to the sample; The annotation result correction module is used to correct the cell type consensus annotation results of each cell population based on the predicted annotation results of each cell population, so as to obtain the cell type annotation results of the sample. The annotation result correction module is specifically used for: If the proportion of mixed types in the target cell population is greater than the preset proportion, then the cell type consensus annotation result of the target cell population is filled based on the predicted annotation result of the target cell population; The process of supplementing the cell type consensus annotation results of the target cell population based on the predicted annotation results of the target cell population includes: If the target cell type in the predicted annotation results of the target cell population is a standard cell type, and the average gene activity assessment score of the target cell type is greater than the average gene activity assessment score of other standard cell types in the target cell population, the consensus annotation results of the cell type of the target cell population are supplemented based on the predicted annotation results of the target cell population.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for automated annotation of single-cell transcriptome cell types based on consensus voting as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for automated annotation of single-cell transcriptome cell types based on consensus voting as described in any one of claims 1-5.