Automatic single cell annotation method and system
By integrating multiple algorithms and re-annotating and verifying, the problems of tediousness and misjudgment in the single-cell annotation process are solved, achieving highly accurate and reliable cell type identification and providing transparent and repeatable annotation results.
Patent Information
- Application Number
- CN202511666251.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-06
AI Technical Summary
Existing single-cell annotation methods are cumbersome and inefficient, and the manual annotation process is prone to errors, especially in the construction of disease-specific maps where it is difficult to match with public databases. Furthermore, the supervised methods do not take into account the differences in tissue compartments that lead to incorrect annotations.
A multi-algorithm integration and re-annotation verification method is adopted. Marker genes are extracted based on an annotated reference map, and combined with the pan-class score of lineage and tissue compartment division. Cell type probability scores are calculated using weighted geometric mean and probability normalization. Finally, the final label is generated through consistency discrimination and conflict resolution.
It significantly improves the accuracy, reliability, and traceability of single-cell annotation, reduces bias caused by human intervention, avoids misjudgment, and provides highly interpretable and reproducible annotation results.
Smart Images

Figure CN121483397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated cell type annotation technology, and in particular to an automated single-cell annotation method and system. Background Technology
[0002] Accurate identification of different cell types in complex tissue samples is a crucial prerequisite for elucidating the role of cell populations in various biological processes, including homeostasis and damage repair. Traditionally, cell sorting and microscopy techniques have been widely used to isolate cell types and identify them using marker genes. In recent years, single-cell transcriptome sequencing (scRNA-seq) has been established as a high-throughput method for routinely mapping different cell populations in tissue samples and studying various biological processes in disease and development. scRNA-seq technology provides detailed molecular profiles of various cell types and has become an essential technique for various human cell atlas projects.
[0003] Cell population identification typically relies on unsupervised cell clustering based on cellular transcriptome profiling, followed by analysis of differentially expressed marker genes between clusters. These marker genes are then manually examined using available information from literature or cell marker databases, ultimately assigning a cell type label to each cluster. However, this manual annotation process is tedious and inefficient because marker genes are often expressed in multiple cell clusters and correspond to various cell types. Furthermore, the expression of different marker genes varies within specific cell types. For example, the SCGB1A1 gene, a marker gene for Club cells, is also expressed in goblet cells, but at a lower abundance.
[0004] Furthermore, the cell annotation process becomes more complex, particularly in disease-specific atlas construction, due to difficulties in matching with public databases. For example, the SCGB3A2 gene is considered a marker gene for novel distal airway secretory cells, but SCGB3A2+ / SFTPC+ alveolar type 2 epithelial cells (AT2) can be detected in the lung tissue of patients with chronic obstructive pulmonary disease (COPD), easily leading to lineage origin errors during annotation. Another popular cell type assignment method utilizes a reference dataset—a set of previously annotated cell types in single-cell data—to train a classification algorithm and apply it to new single-cell datasets. However, this supervised approach does not consider cell type differences across tissue compartments, causing specific cell types to potentially appear incorrectly in datasets across different regions. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide an automated single-cell annotation method and system. Through innovative multi-algorithm integration and re-annotation verification methods, the accuracy, reliability, and traceability of single-cell annotation are significantly improved.
[0006] To achieve the above objectives, the present invention provides the following solution: An automated single-cell annotation method includes: Marker genes for each cell type were extracted based on an annotated reference map, a reference gene set was established, and weighting rules were set according to prior knowledge. The set of strong marker genes for confidence scoring was determined, a pan-class scoring system was established according to four lineages: epithelial, endothelial, immune, and matrix, and tissue compartments were divided to form a list of identifiable cell types. In at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, within the range of candidate cell types that match the broad class score, the probability score of each cell for each candidate cell type is calculated using weighted geometric mean and probability normalization based on the reference gene set and the weighting rules, and a confidence score is generated based on the set of strongly labeled genes, and the type corresponding to the highest probability is taken as the initial annotation label. Consistency judgment and conflict resolution are performed on the outputs of the four default marker gene identification and classification algorithms. The probability scores and confidence levels are used as decision parameters to generate integrated labels and quality identifiers between "determined label", "Undetermined", and "Lowquality" according to preset rules. For easily confused or transitional cells, in single cell type re-identification or multiple cell type re-identification modes, the probability score is recalculated and the integrated label is corrected within a limited range of candidate types according to preset conditions to obtain the final label. The final label is output and associated with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.
[0007] Preferably, the preset conditions include at least one of the following: The integration tag is Undetermined or Low quality, or; The integrated label is inconsistent with the general category score.
[0008] Preferably, the weighting rule is a high / medium / low stratification rule; wherein: marker genes that exist only in a single cell type or are retained only in the target cell type after tissue compartment filtration are set to high weight; marker genes that are limited to two cell types in the same tissue compartment and the same lineage are set to medium weight; and widely expressed genes that lack cell specificity are set to low weight.
[0009] Preferably, the confidence level is divided into 1-5 levels according to the number of genes that match the set of strongly marked genes, and is used as one of the decision parameters for initial annotation and integration decisions.
[0010] Preferably, the range of candidate cell types is determined by the intersection of the broad category score and the list of identifiable cell types, and cell types that are outside the list of identifiable cell types do not participate in the initial annotation and re-annotation.
[0011] Preferably, the integrated decision of the four marker gene identification and classification algorithms follows at least one of the following rules: (1) When the outputs of the four algorithms are consistent, or the outputs of the three algorithms are consistent, the cell type corresponding to the consistency is determined as the integration label; (2) When the output of two algorithms is the first cell type and the output of the other two algorithms is the second cell type, compare the confidence of the first cell type and the second cell type. The one with higher confidence is determined as the integration label; when the confidence of the two is equal, the integration label is marked as Undetermined. (3) When the four algorithms output four different cell types respectively, the integration label is marked as Low quality.
[0012] Preferably, when the optimal cell type generated by the semi-supervised mode and the optimal cell type generated by the unsupervised mode belong to different lineages, and the confidence level of both the semi-supervised mode and the unsupervised mode is greater than zero, the integration tag is marked as Low quality; wherein, the following pairings constitute exceptions: (1) The optimal cell type in the semi-supervised mode is epithelial and the optimal cell type in the unsupervised mode is matrix; or the optimal cell type in the unsupervised mode is epithelial and the optimal cell type in the semi-supervised mode is matrix. (2) The optimal cell type in the semi-supervised mode is endothelial and the optimal cell type in the unsupervised mode is matrix; or the optimal cell type in the unsupervised mode is endothelial and the optimal cell type in the semi-supervised mode is matrix. In exceptional pairings, the confidence levels of the semi-supervised mode and the unsupervised mode are compared: if the confidence level of the semi-supervised mode is higher than that of the unsupervised mode, the optimal cell type of the semi-supervised mode is identified as the integration tag; if the confidence level of the unsupervised mode is higher than that of the semi-supervised mode, the optimal cell type of the unsupervised mode is identified as the integration tag; if the confidence levels of the semi-supervised mode and the unsupervised mode are equal, the integration tag is marked as Undetermined.
[0013] Preferably, when the confidence levels of the four marker gene identification and classification algorithms are all zero in a certain mode while the confidence level of another mode is non-zero, the integrated decision result of the mode with non-zero confidence level is used as the integrated label.
[0014] Preferably, the defined candidate type range includes the integrated label and a predefined set of easily confused types; within the defined candidate type range, the probability score is recalculated using a weighted geometric mean and probability normalization method to correct the integrated label.
[0015] An automated single-cell annotation system, comprising: The reference gene set and weight rule construction unit is used to extract marker genes for each cell type based on the annotated reference map, establish the reference gene set, and set weight rules according to prior knowledge. Strong marker genes, lineage scoring, and tissue compartment division units are used to determine the set of strong marker genes used for confidence scoring. A pan-class scoring system is established according to four lineages: epithelial, endothelial, immune, and matrix. Tissue compartment division is performed to form a list of identifiable cell types. The initial annotation calculation unit is used to calculate the probability score of each cell for each candidate cell type in a range of candidate cell types that match the broad class score, in at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, based on the reference gene set and the weighting rules, using weighted geometric mean and probability normalization, and to generate confidence based on the set of strongly labeled genes, and take the type corresponding to the highest probability as the initial annotation label. The multi-algorithm integration decision unit is used to perform consistency judgment and conflict resolution on the outputs of the four default included marker gene identification and classification algorithms, and uses the probability score and confidence level as decision parameters to generate an integrated label and quality label between "determined label", "undetermined", and "low quality" according to preset rules. The re-annotation verification unit is used to recalculate the probability score and correct the integrated label within a limited range of candidate types and according to preset conditions for easily confused or transitional cells, under single cell type re-identification or multiple cell type re-identification modes, to obtain the final label. The results output and evidence association unit is used to output the final label and associate it with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.
[0016] The present invention discloses the following technical effects: (1) This invention extracts marker genes for each cell type based on an annotated reference map and combines lineage panclass scores and tissue compartment division to achieve more accurate cell type annotation. This reference gene set and weighting rules based on prior knowledge effectively solve the problems of inaccurate marker gene selection and unclear lineage division in the prior art, thereby significantly improving the accuracy and consistency of the annotation results.
[0017] (2) This invention performs consistency judgment and conflict resolution on the output results of four marker gene identification and classification algorithms, and uses probability scores and confidence levels as decision-making criteria to make flexible adjustments between "identified label", "undetermined" and "Lowquality". By integrating multiple algorithms, the errors that may be caused by a single algorithm are avoided, and the quality and reliability of cell annotation results are greatly improved.
[0018] (3) Unlike traditional methods that rely on manual verification, this invention recalculates probability scores and corrects integrated labels within the candidate type range through single-cell-type re-identification or multiple-cell-type re-identification modes, effectively eliminating the bias caused by manual intervention. This mechanism ensures the accuracy and consistency of the verification process, providing a reliable guarantee for the final annotation results.
[0019] (4) This invention, through a “re-annotation verification” mechanism, can recalculate probability scores and correct labels within a limited range of candidate types for easily confused or transitional cells, thus solving the problem of misjudging these cell types in traditional methods. This mechanism significantly improves the comprehensiveness and detail of cell annotation and avoids common annotation errors in existing technologies.
[0020] (5) This invention can output the final label and associate it with the corresponding probability score and expression evidence related to the reference gene set, providing comprehensive support for data traceability and rich evidence for subsequent reanalysis. This feature solves the problem of lack of transparency and data evidence support in the prior art, making the annotation results more interpretable and reproducible, and providing more reliable data support for future research and clinical applications. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the annotation of the scRNA-seq dataset 1 provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the annotation of the scRNA-seq dataset 2 provided in an embodiment of the present invention; Figure 4 The flowchart for the initial annotation of scIA-Suite provided in the embodiments of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The purpose of this invention is to provide an automated single-cell annotation method and system, which significantly improves the accuracy and consistency of single-cell annotation by combining multi-algorithm integrated decision-making with a precise re-annotation verification mechanism.
[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Figure 1 The method flowchart provided in the embodiments of the present invention is as follows: Figure 1 As shown, this invention provides an automated single-cell annotation method, comprising: Step 100: Extract marker genes for each cell type based on the annotated reference map, establish a reference gene set, and set weight rules according to prior knowledge; Step 200: Determine the set of strong marker genes for confidence scoring, establish a pan-class scoring system according to four lineages: epithelial, endothelial, immune, and matrix, and divide tissue compartments to form a list of identifiable cell types; Step 300: In at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, within the range of candidate cell types that match the broad class score, calculate the probability score of each cell for each candidate cell type using weighted geometric mean and probability normalization based on the reference gene set and weight rules, generate confidence based on the set of strongly labeled genes, and take the type corresponding to the highest probability as the initial annotation label. Step 400: Perform consistency judgment and conflict resolution on the outputs of the four default marker gene identification and classification algorithms, and generate an integrated label and quality label between "determined label", "undetermined", and "low quality" according to preset rules, using probability score and confidence level as decision parameters. Step 500: For easily confused or transitional cells, under single cell type re-identification or multiple cell type re-identification modes, within the limited candidate type range, the probability score is recalculated according to preset conditions and the integrated label is corrected to obtain the final label. Step 600: Output the final label and associate it with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.
[0027] Specifically, the technical approach of this embodiment is as follows: 1. The acquisition and preprocessing of the reference dataset are as follows: This invention utilizes scRNA-seq datasets from the following respiratory tract sources (e.g.) Figure 2 and Figure 3 (As shown) is used as an example. First, download the HLCA core map file in h5ad format (website), load it in R using the schard package (v) and convert it into a processing object suitable for the Seurat package (v) format. Use the NormalizeData and FindVariableFeatures functions to normalize the map and identify hypervariable genes, respectively. Given the existing hierarchical labels of the map, this embodiment first uses the 5th level label as the cell type reference. Then, based on the complexity of disease-specific cell transcriptomes and prior knowledge, adjustments are made, and finally 48 cell types are used as annotation references, namely: basal cells (resting, differentiated, and hillock-like), mucous gland cells (ducts, mucous, and serous), secretory cells (proximal and non-proximal goblet, Club), alveolar epithelial cells (AT2 and AT1), ciliated cells (multiciliated and intermediate), neuroendocrine cells, cluster cells, ionocytes, arterial endothelium, venous endothelium (pulmonary and systemic), capillary endothelium (transient) Typical and gastric cells), lymphoendothelial cells (mature and differentiated), macrophages (alveolar, monocyte-derived, and peripheral vascular), monocytes (classical and non-classical), dendritic cells (type 1, type 2, plasmacytoid, and migratory), neutrophils, mast cells, T cells (CD4 and CD8), NK cells, B cells, plasma cells, hematopoietic stem cells, fibroblasts (alveolar, adventitia, peribronchial, and subpleural), myofibroblasts, smooth muscle cells, pericytes, and mesothelial cells. Next, the FindAllMarkers function of the Seurat package, the cosg function of the COSG package (v), the getMarkers function of the genesorteR package (v), and the wilcoxauc function of the presto package (v) were used in sequence to identify the marker genes for each cell type (except neutrophils). Finally, the top 20 marker genes for each cell type (screening threshold: adjusted P < 0.05, log2(fold change) > 0.25) were selected as the reference gene set for subsequent analysis.
[0028] 2. Data preparation for initial annotation using the scIA-Suite tool: First, the cell types were categorized according to epithelial, endothelial, immune, and stromal lineages: 1) Epithelial cells: basal cells (resting type, transitional type and hillock-like), mucous gland cells (ducts, mucus and serous fluid), secretory cells (goblet cells, club cells), alveolar epithelial cells (AT2 and AT1), ciliated cells (multiciliated and intermediate cells), neuroendocrine cells, cluster cells and ionizing cells; 2) Endothelial cells: arterial endothelium, venous endothelium (pulmonary and systemic), capillary endothelium (classical and pneumocellular), and lymphoid endothelium (mature and differentiated). 3) Immune cells: macrophages (alveolar type, monocyte-derived type, and peripheral vascular type), monocytes (classical and non-classical), dendritic cells (type 1, type 2, plasma cell-like, and migratory type), neutrophils, mast cells, T cells (CD4 and CD8), NK cells, B cells, plasma cells, and hematopoietic stem cells; 4) Matrix cells: fibroblasts (alveolar type, adventitia type, peribronchial type, subpleural type), myofibroblasts, smooth muscle cells, pericytes, and mesothelial cells.
[0029] Secondly, the cell types that should be included in the tissue compartments are classified: 1) Tissue origin of proximal airway: mucous gland cells (ducts, mucus, and serous fluid), secretory cells (proximal goblet cells and Club cells), basal cells (resting, differentiated, and hillock-like), ciliated cells (multiciliated and intermediate), neuroendocrine cells, cluster cells, ionizing cells, arterial endothelium, venous endothelium (pulmonary and systemic), lymphoendothelium (mature and differentiated), monocytes (classical and non-classical), dendritic cells (type 1, type 2, plasmacytoid, and migrating), mast cells, T cells (CD4 and CD8), NK cells, B cells, plasma cells, fibroblasts (peribronchial type), and smooth muscle cells; 2) Tissue origins of intermediate airways: secretory cells (non-proximal goblet cells and Club cells), basal cells (resting, differentiated, and hillock-like), ciliated cells (multiciliated and intermediate), neuroendocrine cells, cluster cells, ionocytes, macrophages (alveolar, monocyte-derived), dendritic cells (type 1, type 2, plasmacytoid, and migrating), mast cells, B cells, T cells (CD4 and CD8), NK cells, and hematopoietic stem cells; 3) Distal airway and lung tissue origins: basal cells (resting, differentiated, and hillock-like), secretory cells (non-proximal goblet, Club), alveolar epithelial cells (AT2 and AT1), ciliated cells (multiciliated and intermediate), neuroendocrine cells, cluster cells, ionizing cells, arterial endothelium, venous endothelium (pulmonary and systemic), capillary endothelium (classical and air cell type), lymphoendothelium (mature and differentiated), macrophages (alveolar, monocyte-derived, and peripheral vascular type), monocytes (classical and non-classical), dendritic cells (type 1, type 2, plasmacytoid, and migratory), neutrophils, mast cells, T cells (CD4 and CD8), NK cells, B cells, plasma cells, hematopoietic stem cells, fibroblasts (alveolar, adventitia, peribronchial, and subpleural), myofibroblasts, smooth muscle cells, pericytes, and mesothelial cells. The dataset to be processed is filtered based on sample origin to remove cell types from the reference gene set that should not be identified.
[0030] Next, the marker genes selected by the four algorithms are weighted according to the following rules: 1) High-weighted genes: ① Existing only in a specific cell type. For example, ABCA3 has been identified as a specific gene marker for AT2; ② Existing in a specific cell type after removing cell types from different tissue compartments. For example, LTF was found only in goblet cells from distal airway tissue after excluding serous cells and ductal cells of mucous glands from proximal airway tissue; ③ Recognized in two cell types of different major categories. For example, LAMP3 was identified only in AT2 cells of epithelial cells and migratory dendritic cells of immune cells; ④ Specific gene markers provided by the HLCA atlas. For example, MFSD2A, C8orf4, and C11orf96 have been defined as specific marker genes for AT2.
[0031] 2) Medium-weighted genes: ① Present in two cell types of the same major category within the same tissue compartment. For example, HOPX is only recognized in AT2 and AT1; ② Recognized in more than two cell types, but after removing cell types from different tissue compartments, they are present in at most two cell types of the same major category. For example, when mucous gland cells from proximal airway tissue are excluded, TSPAN8 is only identified in Club and Goblet cells from two distal airway tissues; 3) Low-weighted genes: Not expressed cell-specifically. For example, KRT5 is identified as a marker gene for resting, differentiated, and hillock-like basal cells. Based on the above rules, the four algorithms are predefined into proximal, intermediate, and distal weighted gene sets according to different tissue compartments.
[0032] To improve annotation accuracy, during the initial annotation process, five strong marker genes are assigned to each included cell type for confidence level calculation. Scores from 1 to 5 correspond to the expression levels of these five strong marker genes for a given cell type. This weighting is an optional setting for the initial annotation process.
[0033] Finally, based on the reference gene set or the preprocessed weighted gene set, a weighted geometric mean and probability normalization calculation are used to define the cell type with the highest probability as the final prediction result.
[0034] 3. Initial annotation process for scIA-Suite tool: This tool initially offers three annotation modes: semi-supervised, unsupervised, and predefined module prediction. Specific usage is as follows: 1) Semi-supervised mode: In this mode, cell clustering information is pre-matched to a broad range of cell types, namely epithelial lineage (marker genes EPCAM, FXYD3, and ELF3), endothelial lineage (marker genes CLDN5, ECSCR, and CLEC14A), immune lineage (marker genes PTPRC, CD53, and CORO1A), and matrix lineage (marker genes COL1A2, DCN, and MFAP4). This allows for fully automated prediction of the dataset to be annotated within the eligible cell types. This mode can be used to eliminate annotation distortion caused by crosstalk between marker genes of different cell types.
[0035] 2) Unsupervised mode: In this mode, cell clustering information is not preset, and the dataset to be annotated is automatically predicted among cell types that meet the criteria. This mode can be used to discover potential low-quality (e.g., two-cell) or transitional cells.
[0036] 3) Predefined Module Mode: In this mode, the dataset to be annotated is pre-divided into lineage modules based on a broad class of cell types, and fully automated prediction is performed according to the cell types matched by the modules. This mode can be used in scenarios where the number of cells is not critical, and strict matching of broad class of cell types is required to avoid including potential double cells.
[0037] Ultimately, the output will include cell type labels, probability scores, gene expression levels, and other information. See the flowchart below. Figure 4 .
[0038] 4. Supplementing the annotation process for the scIA-Suite tool Supplementary comments apply to the output results of the first two modes using the default initial commenting function (four algorithms). The built-in processing rules for this workflow are as follows: 1) When the pan-macro cell type score only matches the semi-supervised mode (in the semi-supervised mode, four...) If the cell type predicted by the algorithm belongs to only one broad cell type and the score of that broad cell type is greater than 0, it will be annotated according to the following rules: ① Label the same cell type predicted by three or more algorithms. For example, if any three algorithms predict AT2, then label it as AT2.
[0039] ② If any two algorithms predict the same cell type, the remaining two algorithms will predict the same cell type. For the other cell type, the cell type with the higher confidence level is identified. When the confidence levels of the two algorithms are equal, it is marked as Undetermined. For example: ① Any two algorithms predict AT2 with a confidence level of 3; the other two algorithms predict AT1 with a confidence level of 2, and it is marked as AT2; ② Any two algorithms predict AT2 with a confidence level of 2; the other two algorithms predict AT1 with a confidence level of 2, and it is marked as Undetermined. ③ If any two algorithms predict the same cell type, and the remaining two algorithms predict two different cell types, then the cell type is labeled as "Low quality" when the confidence levels of both cell types are greater than 0, or when either cell type is equal to or greater than the confidence level of the former two algorithms for the same cell type. For example: ① If any two algorithms predict AT2 with a confidence level of 2, and the other two algorithms predict AT1 and Basal resting with confidence levels of 0 and 0 respectively, then the cell type is labeled as AT2; ② If any two algorithms predict AT2 with a confidence level of 0, and the other two algorithms predict AT1 and Basal resting with confidence levels of 0 and 0 respectively, then the cell type is labeled as "Low quality"; ③ If any two algorithms predict AT2 with a confidence level of 3, and the other two algorithms predict AT1 and Basal resting with confidence levels of 5 and 4 respectively, then the cell type is classified as "Low quality".
[0040] ④ If the four algorithms predict different cell types, they are labeled as Low quality. For example, if the four algorithms predict AT2, AT1, Club (non-nasal), and Basal resting respectively, they are labeled as Low quality. 2) When the general class score only matches the unsupervised mode (in the semi-supervised mode, the cell types predicted by the four algorithms belong to only one general class cell type, but the score of that general class cell type is equal to 0), the annotation is performed according to the following rules: ① Label the same cell type predicted by three or more algorithms. For example: any three If the algorithm predicts AT2, then it is labeled as AT2.
[0041] ② If any two algorithms predict the same cell type, and the remaining two algorithms uniformly predict another cell type belonging to the same major category, then the cell type with the higher confidence level is marked; if the confidence levels of the two algorithms are equal, it is marked as Undetermined. For example, AT2 and AT1 are both panepithelial cell types. If any two algorithms predict AT2 with a confidence level of 3, and the other two algorithms predict AT1 with a confidence level of 2, then it is marked as AT2; otherwise, it is marked as AT1. If any two algorithms predict AT2 with a confidence level of 2, and the other two algorithms predict AT1 with a confidence level of 2, then it is marked as Undetermined.
[0042] ③ If any two algorithms predict the same cell type, and the remaining two algorithms predict a different cell type from a different major category, then the cell type with a confidence level greater than 0 is marked (the confidence level of the other cell type must be equal to 0). If both confidence levels are greater than 0 or equal, it is marked as Low quality. In special cases, if the two cell types from different major categories are either panepithelial or panstromal, or either panendothelial or panstromal, then the cell type with the higher confidence level is marked; when both confidence levels are equal, it is marked as Undetermined. For example: if any two algorithms predict AT2 with a confidence level of 3, and the other two algorithms predict EC generalcapillary with a confidence level of 0, then it is marked as AT2; otherwise, it is marked as EC generalcapillary. If any two algorithms predict AT2 with a confidence level of 2, and the other two algorithms predict EC generalcapillary with a confidence level of 2, then it is marked as Low quality. In special cases, if any two algorithms predict AT2 with a confidence level of 2, and the other two algorithms predict Alveolar fibroblasts with a confidence level of 1, then it is labeled as AT2. In special cases, if any two algorithms predict AT2 with a confidence level of 1, and the other two algorithms predict Alveolar fibroblasts with a confidence level of 2, then it is labeled as Alveolar fibroblasts. In special cases, if any two algorithms predict AT2 with a confidence level of 2, and the other two algorithms predict Alveolar fibroblasts with a confidence level of 2, then it is labeled as Undetermined.
[0043] ④ If any two algorithms predict the same cell type, and the remaining two algorithms predict two different cell types, then the cell type is labeled as "Low quality" when the confidence scores for both cell types are greater than 0, or when either cell type is equal to or greater than the confidence scores for the same cell type predicted by the two algorithms. For example: If any two algorithms predict AT2 with a confidence score of 2, and the other two algorithms predict AT1 and Basal resting with confidence scores of 0 and 0 respectively, then the cell type is labeled as AT2; if any two algorithms predict AT2 with a confidence score of 0, and the other two algorithms predict AT1 and Basal resting with confidence scores of 0 and 0 respectively, then the cell type is labeled as "Low quality"; if any two algorithms predict AT2 with a confidence score of 3, and the other two algorithms predict AT1 and Basal resting with confidence scores of 5 and 4 respectively, then the cell type is labeled as "Low quality".
[0044] ⑤ If the four algorithms predict different cell types respectively, the cell type is marked as Low quality. For example, if the four algorithms predict AT2, AT1, Club (non-nasal), and Basal resting respectively, the cell type is marked as Low quality. 3) When the general class score does not perfectly match the semi-supervised mode (in the semi-supervised mode, the cell type predicted by the four algorithms is any one of the general class cell types, and there exists any other general class score greater than 0 besides that general class score being greater than 0) or when all general class scores are equal to 0, if the best cell type predicted by the four algorithms in the semi-supervised mode is exactly the same as the cell type predicted by the four algorithms in the unsupervised mode, then it is marked according to the annotation rules of the above semi-supervised mode; if the two cell types are not equal, it is marked according to the following rules: ① If the confidence level of the cell type predicted by the four algorithms in both semi-supervised and unsupervised modes is equal to 0, then it is marked as Low quality; ② If the confidence level of any algorithm predicting a cell type in semi-supervised mode is greater than 0, and the confidence level of any algorithm predicting a cell type in unsupervised mode is greater than 0, except in special cases, if the best cell type predicted in unsupervised mode and the best cell type predicted in semi-supervised mode are from different major categories, it is judged as Low quality; if they are from the same major category, the cell type is determined according to the rules of semi-supervised mode (first major point). Special cases are as follows: if the best cell types of the two modes are any cell type in panepithelial and panstromal, or any cell type in panendothelial and panstromal, then it is marked as the cell type with higher confidence; when the confidence levels of the two are equal, it is marked as Undetermined. For example: in a special case, if the best cell type predicted by any two algorithms is AT2 with a confidence level of 2, and the best cell type predicted by the other two algorithms is Alveolar fibroblasts with a confidence level of 1, then it is marked as AT2; otherwise, it is marked as Alveolar fibroblasts. In special cases, if the best cell type predicted by any two algorithms is AT2 with a confidence level of 2, and the best cell type predicted by the other two algorithms is Alveolar fibroblasts with a confidence level of 2, then it is marked as Undetermined.
[0045] ③ If the confidence scores of the cell types predicted by the four algorithms in the unsupervised mode are all equal to 0, while in the semi-supervised mode, any one of the four algorithms has a confidence score greater than 0, then the cell types are marked according to the rules of the semi-supervised mode.
[0046] ④ If the confidence scores of the cell types predicted by the four algorithms in the semi-supervised mode are all equal to 0, while in the unsupervised mode, any one of the four algorithms has a confidence score greater than 0, then the cell types are marked according to the rules of the unsupervised mode.
[0047] Ultimately, the dataset to be annotated can obtain high-precision cell labels after a supplementary annotation process.
[0048] 5. ScIA-Suite tool re-annotation process: This process includes two modes, specifically: 1) Single cell type re-identification: Suitable for labeling and validating a group of cells to check if they belong to another specific cell type. For example, annotations for monocytes and neutrophils are easily confused; this pattern can be used to identify potential neutrophils in the final annotation results. 2) Re-identification of multiple cell types: Suitable for distinguishing two or more difficult-to-distinguish cell types. For example, distinguishing AT2 cells from airway secretory cells.
[0049] Corresponding to the above method, this embodiment also provides an automated single-cell annotation system, including: The reference gene set and weight rule construction unit is used to extract marker genes for each cell type based on the annotated reference map, establish the reference gene set, and set weight rules according to prior knowledge. Strong marker genes, lineage scoring, and tissue compartment division units are used to determine the set of strong marker genes used for confidence scoring. A pan-class scoring system is established according to four lineages: epithelial, endothelial, immune, and matrix. Tissue compartment division is performed to form a list of identifiable cell types. The initial annotation calculation unit is used to calculate the probability score of each cell for each candidate cell type in a range of candidate cell types that match the broad class score, in at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, based on the reference gene set and the weighting rules, using weighted geometric mean and probability normalization, and to generate confidence based on the set of strongly labeled genes, and take the type corresponding to the highest probability as the initial annotation label. The multi-algorithm integration decision unit is used to perform consistency judgment and conflict resolution on the outputs of the four default included marker gene identification and classification algorithms, and uses the probability score and confidence level as decision parameters to generate an integrated label and quality label between "determined label", "undetermined", and "low quality" according to preset rules. The re-annotation verification unit is used to recalculate the probability score and correct the integrated label within a limited range of candidate types and according to preset conditions for easily confused or transitional cells, under single cell type re-identification or multiple cell type re-identification modes, to obtain the final label. The results output and evidence association unit is used to output the final label and associate it with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.
[0050] The beneficial effects of this invention are as follows: This invention addresses the issue that current scRNA-seq dataset annotation heavily relies on manual identification by researchers, a cumbersome and highly subjective process. Furthermore, under large-scale, multi-batch, multi-tissue conditions, classic annotation tools such as singleR and Seurat's TransferData function cannot accurately obtain high-quality cell tags. scIA-Suite provides a standardized pipeline that makes the annotation process objective, transparent, and reproducible.
[0051] The tool of this invention has a built-in high-quality list of marker genes (including weighted assignments), which supports rapid application to another dataset (such as different disease stages, different treatment conditions, or different species), accelerating the analysis process of new data and ensuring the consistency of annotation labels. Furthermore, reliable annotation can avoid misleading conclusions due to annotation errors.
[0052] The technical advantages of this invention are as follows: 1) Hierarchical annotation framework: Adopt the process of "general category scoring -> grouping -> detailed annotation within the group -> multi-algorithm integration" to minimize erroneous annotations; 2) Knowledge-driven weighting system: Integrating prior knowledge (the specificity of marker genes) into the automated annotation process in a quantitative and adjustable manner; 3) Integrated decision engine: Identifies cell types through a complex decision system, preserves transitional cells as much as possible, filters contradictory and low-quality data, further improving accuracy while reducing false positives; 4) Special case handling: Able to identify and handle recognized special cell type pairs, including AT2 / stromal cells and endothelial / stromal cells.
[0053] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0054] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An automated single-cell annotation method, characterized in that, include: Marker genes for each cell type were extracted based on an annotated reference map, a reference gene set was established, and weighting rules were set according to prior knowledge. The set of strong marker genes for confidence scoring was determined, a pan-class scoring system was established according to four lineages: epithelial, endothelial, immune, and matrix, and tissue compartments were divided to form a list of identifiable cell types. In at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, within the range of candidate cell types that match the broad class score, the probability score of each cell for each candidate cell type is calculated using weighted geometric mean and probability normalization based on the reference gene set and the weighting rules, and a confidence score is generated based on the set of strongly labeled genes, and the type corresponding to the highest probability is taken as the initial annotation label. Consistency judgment and conflict resolution are performed on the outputs of the four default marker gene identification and classification algorithms. The probability score and confidence level are used as the decision parameters to generate an integrated label and quality label between "determined label", "Undetermined", and "Lowquality" according to preset rules. For easily confused or transitional cells, in single cell type re-identification or multiple cell type re-identification modes, the probability score is recalculated and the integrated label is corrected within a limited range of candidate types according to preset conditions to obtain the final label. The final label is output and associated with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.
2. The automated single-cell annotation method according to claim 1, characterized in that, The preset conditions include at least one of the following: The integration tag is Undetermined or Low quality, or; The integrated label is inconsistent with the general category score.
3. The automated single-cell annotation method according to claim 1, characterized in that, The weighting rule is a high / medium / low stratification rule; wherein: marker genes that exist only in a single cell type or are retained only in the target cell type after tissue compartment filtration are set to high weight; marker genes that are limited to two cell types in the same tissue compartment and the same lineage are set to medium weight; and widely expressed genes that lack cell specificity are set to low weight.
4. The automated single-cell annotation method according to claim 1, characterized in that, The confidence level is divided into 1-5 levels based on the number of genes that match the set of strongly labeled genes, and is used as one of the decision-making parameters for initial annotation and integration.
5. The automated single-cell annotation method according to claim 1, characterized in that, The range of candidate cell types is determined by the intersection of the broad category score and the list of identifiable cell types. Cell types that are outside the list of identifiable cell types are not included in the initial annotation and re-annotation.
6. The automated single-cell annotation method according to claim 1, characterized in that, The integration decision of the four marker gene identification and classification algorithms follows at least one of the following rules: (1) When the outputs of the four algorithms are consistent, or the outputs of the three algorithms are consistent, the cell type corresponding to the consistency is determined as the integration label; (2) When the output of two algorithms is the first cell type and the output of the other two algorithms is the second cell type, compare the confidence of the first cell type and the second cell type. The one with higher confidence is determined as the integration label; when the confidence of the two is equal, the integration label is marked as Undetermined. (3) When the four algorithms output four different cell types respectively, the integration label is marked as Low quality.
7. The automated single-cell annotation method according to claim 1, characterized in that, When the optimal cell type generated by the semi-supervised mode and the optimal cell type generated by the unsupervised mode belong to different lineages, and the confidence level of both the semi-supervised mode and the unsupervised mode is greater than zero, the integration tag will be marked as "Low quality"; the following pairings constitute exceptions: (1) The optimal cell type in the semi-supervised mode is epithelial and the optimal cell type in the unsupervised mode is matrix; or the optimal cell type in the unsupervised mode is epithelial and the optimal cell type in the semi-supervised mode is matrix. (2) The optimal cell type in the semi-supervised mode is endothelial and the optimal cell type in the unsupervised mode is matrix; or the optimal cell type in the unsupervised mode is endothelial and the optimal cell type in the semi-supervised mode is matrix. In exceptional pairings, the confidence levels of the semi-supervised mode and the unsupervised mode are compared: if the confidence level of the semi-supervised mode is higher than that of the unsupervised mode, the optimal cell type of the semi-supervised mode is identified as the integration tag; if the confidence level of the unsupervised mode is higher than that of the semi-supervised mode, the optimal cell type of the unsupervised mode is identified as the integration tag; if the confidence levels of the semi-supervised mode and the unsupervised mode are equal, the integration tag is marked as Undetermined.
8. The automated single-cell annotation method according to claim 1, characterized in that, When the confidence levels of the four marker gene identification and classification algorithms are all zero in one mode, while the confidence level of another mode is non-zero, the integrated decision result of the mode with non-zero confidence level is used as the integrated label.
9. The automated single-cell annotation method according to claim 1, characterized in that, The defined candidate type range includes the integrated label and a predefined set of easily confused types; within the defined candidate type range, the probability score is recalculated using a weighted geometric mean and probability normalization method to correct the integrated label.
10. An automated single-cell annotation system, characterized in that, include: The reference gene set and weight rule construction unit is used to extract marker genes for each cell type based on the annotated reference map, establish the reference gene set, and set weight rules according to prior knowledge. Strong marker genes, lineage scoring, and tissue compartment division units are used to determine the set of strong marker genes used for confidence scoring. A pan-class scoring system is established according to four lineages: epithelial, endothelial, immune, and matrix. Tissue compartment division is performed to form a list of identifiable cell types. The initial annotation calculation unit is used to calculate the probability score of each cell for each candidate cell type in a range of candidate cell types that match the broad class score, in at least one of the semi-supervised mode, unsupervised mode, or predefined module mode, based on the reference gene set and the weighting rules, using weighted geometric mean and probability normalization, and to generate confidence based on the set of strongly labeled genes, and take the type corresponding to the highest probability as the initial annotation label. The multi-algorithm integration decision unit is used to perform consistency judgment and conflict resolution on the outputs of the four default included marker gene identification and classification algorithms, and uses the probability score and confidence level as decision parameters to generate an integrated label and quality label between "determined label", "Undetermined", and "Low quality" according to preset rules. The re-annotation verification unit is used to recalculate the probability score and correct the integrated label within a limited range of candidate types and according to preset conditions for easily confused or transitional cells, under single cell type re-identification or multiple cell type re-identification modes, to obtain the final label. The results output and evidence association unit is used to output the final label and associate it with the corresponding probability score and expression evidence related to the reference gene set for retrospective and reanalysis.