A single-cell transcriptome cell annotation method and system fusing a large language model
By integrating large language models and homology alignment technology, the accuracy and efficiency issues of single-cell transcriptome annotation methods have been addressed, realizing an automated and intelligent cell annotation workflow applicable to both model and non-model species, thus improving the accuracy and efficiency of annotation.
Patent Information
- Application Number
- CN202411755491.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Existing single-cell transcriptome annotation methods suffer from problems such as accuracy dependence on expert experience, time consumption, and low efficiency. In particular, the annotation results are inaccurate in non-model species, and software annotation tools are severely affected by the accuracy of the reference dataset.
By integrating large language models (such as Chat GPT-4) with homology alignment technology, and through data preprocessing, gene sequencing and screening, homology alignment and prompt word construction, combined with multiple question-and-answer sessions to determine cell type, an automated and intelligent cell annotation process is achieved.
It improves the accuracy and versatility of cell type annotation, can handle both model and non-model species, reduces manual intervention, shortens analysis time, and meets the needs of high-throughput data analysis.
Smart Images

Figure CN119601094B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics analysis technology, and particularly relates to single-cell transcriptome data processing and cell type annotation technology, specifically a single-cell transcriptome cell annotation method and system that integrates a large language model. Background Technology
[0002] The development of single-cell transcriptome sequencing technology is progressing rapidly, with continuously increasing throughput and decreasing costs, and it has been widely applied in basic biological research and clinical research. This technology can analyze gene expression profiles at the single-cell level, providing crucial data support for a deeper understanding of cellular heterogeneity, cellular function, and disease pathogenesis. For example, in tumor research, single-cell transcriptome sequencing helps reveal the gene expression characteristics of different cell types in the tumor microenvironment, providing a theoretical basis for precision tumor treatment.
[0003] However, cell type annotation, a fundamental step in single-cell transcriptome sequencing analysis, faces numerous limitations with existing methods. Manual annotation, the traditional approach, relies on experts comparing highly expressed genes in cell subpopulations with typical marker genes to determine cell type. This method's accuracy is highly dependent on expert experience; different experts may annotate the same data differently. Furthermore, manual annotation is extremely time-consuming and inefficient for large-scale single-cell transcriptome data, failing to meet the demands of high-throughput data analysis and thus hindering its standardized, large-scale application. Software annotation methods, such as singleR and cellassign, determine cell type by comparing the correlation between new sequencing data and a reference dataset. While these tools offer efficient and automated annotation, they also suffer from several problems. Their accuracy is severely constrained by the accuracy of the reference dataset; differences in reference datasets across different tissues or species can lead to inaccurate annotation results. For example, a reference set for liver cells cannot be directly applied to annotate small intestinal cells. Additionally, errors in the reference dataset directly impact annotation effectiveness. As the application scope of single-cell sequencing continues to expand, constructing high-quality reference sets for various tissue samples requires significant manpower and time.
[0004] Large language models (such as Chat GPT-4 and Zhipu Qingyan) have achieved significant success in the field of natural language processing. Their training process integrates a vast database of human scientific and technological papers, enabling them to understand biological knowledge. Recent studies have demonstrated the effectiveness of large language models in the biomedical field, such as in tasks like literature retrieval and question answering. However, their application in single-cell transcriptome annotation is still in the exploratory stage. Existing literature on cell marker genes primarily focuses on common model organisms such as humans and mice, with a severe lack of information on non-model organisms. Summary of the Invention
[0005] This invention aims to overcome the shortcomings of existing cell annotation methods, integrate large language models into the cell annotation analysis process, and improve the model's universality across different species through strategies such as homology annotation, thereby developing a more efficient, accurate, and universal single-cell transcriptome cell annotation method and system.
[0006] The technical solution of the present invention, which integrates a large language model with a single-cell transcriptome cell annotation method, is as follows.
[0007] Data preprocessing: The newly obtained single-cell transcriptome data were preprocessed, including data cleaning, quality control, and gene expression matrix construction. Advanced data quality control algorithms were used to remove low-quality data and outliers, and the data were normalized to ensure accuracy and comparability, providing a high-quality data foundation for subsequent analysis.
[0008] Identify and rank upregulated genes in each cell subpopulation: For each gene in a given subpopulation, use the Wilcoxon rank-sum test (a non-parametric test primarily used to compare the distributions of two independent samples) to compare its expression value with other cell subpopulations to identify significantly upregulated genes in each subpopulation. The specific criteria are: the gene is expressed in at least 25% of cells within the target subpopulation, a p-value ≤ 0.01, and a log2 fold change ≥ 0.36. Based on the differential analysis results, rank the upregulated genes in each subpopulation, prioritizing ranking by p-value from smallest to largest. When p-values are the same, further rank by fold change from largest to smallest. This helps to screen for genes with significant expression differences in cell subpopulations that may play a key role in cell type identification.
[0009] Homology alignment for non-model species (if the target cell is a non-model species): When the target cell for annotation is a non-model species (e.g., non-human, non-mouse, non-Arabidopsis), homology alignment is required to determine the gene symbol number (a gene identifier). If the target species is an animal, human or mouse is selected as the reference model species; if it is a plant, Arabidopsis is selected as the reference model species. Homology alignment is performed using BLASTP software (a software developed by NCBI for searching similarity in protein sequence databases), with a homology threshold set to E-value (expected value) < 1 × 10⁻⁶. -5With an Identity score ≥ 0.8, genes are sorted according to their E-value, and the gene with the lowest E-value is selected as the homologous gene of the source reference type species. Then, based on the symbol number of the reference type species' gene, the gene symbol number (a gene identifier) of the non-type species is defined. Through a precise homology comparison algorithm, homologous genes are accurately screened and the gene symbol number of the non-type species is correctly defined, so that the gene information of the non-type species can be effectively mapped to the gene system of the reference type species. This is crucial for mapping the gene information of the non-type species to the known gene system of the type species, and helps to use the knowledge resources of the type species to perform cell annotation of the non-type species.
[0010] Proprietary cue words are constructed by recording information such as the names of upregulated marker genes for species, tissues, and subpopulations as variables, defining input and output formats, and constructing cue words to be submitted to the large language model. For model species (human, mouse, Arabidopsis), 10 of the top-ranking genes in each subpopulation are selected to construct cue words; for non-model species, 20 of the top-ranking genes in each subpopulation are selected. The constructed cue words are accurately formatted, with complete variable information, and can clearly and accurately convey the gene expression characteristics of cell subpopulations to the large language model, guiding the large language model to make accurate cell type determinations. For example, taking human small cell lung cancer as an example, the constructed cue word format is: "You are a biologist, and youare very good at cell annotation."
[0011] Identify the cell types of human small cell lung cancer using the following markers, each row separately. Some may be a mixture of multiple cell types.
[0012] The following is the data to be analyzed. Each row represents one type of cells and the corresponding cell markers. Cell ID is before the semicolon, and gene markers are after the semicolon.
[0013] 0: BTBD17, LINC01896, CALML3, TRIM9, GRP, KLK12, NPW, KIF19, DLL3,ESPL1.
[0014] 1: H3C2, H1-5, CENPA, DEPDC1, DLGAP5, UBE2C, PBK, ASPM, NUF2, SGO1.
[0015] 2: CD8A, CD8B, GZMK, TRGC2, CCL5, IFNG, GZMA, GZMH, LAG3, APOBEC3G.
[0016] 3: IL7R, LTB, CD6, P2RY10, ITK, SPOCK2, KLRB1, CD3E, TRAC, CD2.
[0017] 4: S100A12, FCN1, IL1B, CD300E, FOLR3, EREG, PTX3, LILRA5, CLEC4E, S100A8.
[0018] 5: SERPINB4, SERPINB3, MMP10, AQP5, BPIFB1, SLC6A14, SCGB1A1, MSLN,KRT17, SCGB3A1.
[0019] 6: CCL18, LILRB5, APOE, APOC1, C1QB, C1QA, MARCO, C1QC, OTOA, TREM2.
[0020] ## Please output the cell types of each row. The format is as follows. Please output the result directly without further explanation.
[0021] 0 Monocytes
[0022] 1 Neutrophil
[0023] 2 B cells
[0024] Submit to the large language model and determine the cell type: Submit the constructed prompts to the large language model (such as Chat GPT-4) through the API (running the Chat GPT-4 application programming interface), and repeat the question-and-answer five times. Summarize the responses of the large language model, and determine the cell type and marker genes corresponding to each subgroup according to the concept of mode in statistics. For example, if three out of five question-and-answers answer that a certain subgroup is "B cells", then this subgroup is finally determined to be of the "B cells" type. By asking multiple times and statistically counting the majority answers, the reliability and stability of cell type determination are improved.
[0025] Furthermore, the present invention also proposes a single-cell transcriptome cell annotation system integrating a large language model, including a data preprocessing module, a gene sorting and screening module, a homology alignment module (for non-model species), a prompt construction module, a large language model interaction module, and a cell type determination module. Each module works together to achieve an automated and intelligent cell annotation process.
[0026] The single-cell transcriptome cell annotation system integrating a large language model of the present invention includes the following.
[0027] Data preprocessing module: Responsible for receiving single-cell transcriptome data and performing quality control on the data, including removing low-quality data, outliers, etc., to ensure the accuracy and reliability of the data. At the same time, perform standardization processing on the data to make the data of different samples comparable and provide a high-quality data basis for subsequent analysis steps.
[0028] Gene sequencing and screening module: This module acquires and sorts upregulated genes from each cell subpopulation, filters significantly upregulated genes according to set criteria, and then sorts them. It utilizes advanced statistical analysis algorithms to ensure the accuracy of gene screening and sorting, providing crucial gene information for subsequent homology alignment and suggestion word construction.
[0029] Homology Alignment Module (for Non-Model Species): This module performs homology alignment operations when processing non-model species data. Corresponding to the homology alignment steps for non-model species described above, efficient homology alignment is performed using tools such as BLASTP software (a software developed by NCBI for searching similarity in protein sequence databases). This determines the symbol number (a gene identifier) of genes in non-model species, achieving a mapping between genes in non-model species and genes in model species, thus expanding the applicability of cell annotation methods.
[0030] The cue word construction module constructs proprietary cue words according to predetermined rules, based on species and tissue information provided by the data preprocessing module and genes selected by the gene sequencing and screening module. This module can flexibly adapt to different species and data characteristics, generating accurate and effective cue words so that the large language model can accurately understand data features and task requirements.
[0031] Large Language Model Interaction Module: Responsible for interacting with large language models (such as Chat GPT-4), submitting prompts to the large language model and receiving its responses. Through an optimized API (using the Chat GPT-4 application programming interface), it ensures efficient communication with the large language model, enabling multiple question-and-answer sessions and result aggregation, providing a basis for accurate cell type determination.
[0032] Cell type determination module: Based on the responses obtained from the large language model interaction module, the cell type is determined according to the statistical concept of mode. This module performs statistical analysis on the output results of the large language model, integrates multiple question-and-answer results, and arrives at the most reliable cell type determination conclusion, improving the accuracy and stability of cell annotation.
[0033] The beneficial effects of the single-cell transcriptome annotation method and system that integrates a large language model in this invention are as follows.
[0034] 1. Improve the accuracy of annotations.
[0035] By optimizing the construction of prompt words and taking the mode of multiple question-and-answer sessions, this invention can more accurately determine cell types. Compared with traditional manual annotation, it reduces the subjective error of human judgment; compared with software annotation, it is not affected by the limitations of the reference dataset and can more accurately identify cell types, especially when dealing with complex or rare cell types.
[0036] 2. Enhance versatility.
[0037] By employing a homology alignment strategy, the method of this invention can process single-cell transcriptome data from both model and non-model species, effectively addressing the challenges faced by existing methods in cell annotation of non-model species. This significantly expands the applicability of cell annotation methods, making them valuable for application in biological research across different species.
[0038] 3. Achieve automation and intelligence.
[0039] The automated and intelligent cell annotation workflow and the efficient processing capabilities of the large language model enable this invention to rapidly process large-scale single-cell transcriptome data. Compared to time-consuming manual annotation, it reduces the degree of human intervention, significantly shortens the analysis time, and improves annotation efficiency, helping to process massive amounts of single-cell transcriptome data in a short time and meeting the needs of high-throughput data analysis. Attached Figure Description
[0040] Figure 1 The flowchart of a single-cell transcriptome annotation method integrating a large language model is shown.
[0041] Figure 2 The flowchart of a single-cell transcriptome annotation system integrating a large language model is shown. Detailed Implementation
[0042] The present invention will be described in detail below with reference to the embodiments.
[0043] Example 1: Single-cell annotation of mouse liver.
[0044] Data Preparation and Preprocessing: Single-cell transcriptome data of mouse liver were collected, including gene expression values. Data quality control was performed to remove low-quality cells and genes, ensuring data reliability. Genes with extremely low expression levels or not expressed in a large number of cells, as well as cell data potentially containing technical errors, were removed. Differential analysis was performed on the preprocessed data to calculate the expression differences of each gene in different cell subpopulations, providing a data foundation for subsequent gene screening and sequencing.
[0045] Gene screening and sequencing: Following the gene screening criteria described above, the Wilcoxon rank-sum test was used to identify significantly upregulated genes in each subpopulation of mouse liver cells. During the screening process, the expression of each gene in different subpopulations was rigorously compared to ensure that the screened genes exhibited significant subpopulation-specific expression differences. The screened upregulated genes were then sorted, primarily based on P-values from smallest to largest, with genes having the same P-value further sorted by fold change from largest to smallest. For a given subpopulation, a sorted list of upregulated genes was obtained, where the genes at the top of the list exhibited higher expression specificity and larger fold changes within that subpopulation.
[0046] Cue word construction: Based on the subpopulation gene sequencing results, the top 10 genes were extracted, and cue words were constructed. During the construction process, variables such as mouse species, liver tissue, and gene information of each subpopulation were accurately recorded, and cue words were generated according to a predetermined format. The constructed cue words are:
[0047] "You are a biologist, and you are very good at cell annotation."
[0048] Identify the cell types of mouse liver cells using the markers listed below, separately for each row. Some may be a mixture of multiple cell types.
[0049] The following is the data to be analyzed. Each row represents one type of cells and the corresponding cell markers. Cell ID is before the semicolon, and gene markers are after the semicolon.
[0050] 0: Gm46565, Lyve1, Bmp2, Oit3, Id4, Kdr, Stab2, Aass, Gm30938,4930578G10Rik
[0051] 1: Vsig4, Timd4, C6, Cd5l, Clec4f, Cd207, Folr2, C1qa, Mmp12, Spic
[0052] 2: Pax5, Fcmr, Ms4a1, Cr2, Ebf1, Chst3, Fcer2a, Cacna1i, Gm31243, Blk
[0053] 3: Chil3, F13a1, Arhgef37, Sirpd, Ace, F10, Tppp3, Gpr141, Gm56663, Nupr1
[0054] 4: Gja5, Nrg1, Adgrg6, Ednrb, Mecom, Vegfc, Sdc1, Gask1b, Jag1, Fbln2
[0055] 5: Nsg2, Lef1, Sidt1, Cd8b1, Tcf7, Fgf13, Gm44174, Il7r, Ms4a4b, Gm2682
[0056] 6: Izumo1r, Cd4, Ctla4, Tnfsf8, Icos, Cd5, Cd40lg, Slamf1, A430093F15Rik, Ikzf2
[0057] 7: Stfa2l1, S100a8, S100a9, Il36g, Retnlg, Cxcr2, Slc7a11, Trem1, Dhrs9, Il1r2
[0058] 8: Il4, Il12rb2, Cd40lg, Zbtb16, Kcnq5, Il2ra, Gramd3, Podnl1, Klrc2, Acoxl
[0059] 9: Gzmk, Pdcd1, Eomes, Ccl5, Cd8a, Gm44174, Trgc2, Cd8b1, Tigit, Cst7
[0060] 10: Gstm3, Bhmt, A2ml1, Rgn, Fabp2, Otc, Sult1d1, Akr1c6, Agmat,Glyat
[0061] 11: Rspo3, Wnt9b, Bmp4, Lhx6, Selp, Vwf, Plppr5, Tgfb2, Rnf165, Cdh13
[0062] 12: Cd209a, Mreg, Klri1, Il4i1, Clec10a, Flt3, Nectin1, Tnip3,Spint1, Tspan33
[0063] 13: Pkhd1, Fermt1, Sftpd, Gm57079, Hnf4g, Gm19935, Sytl5, Cd200l1,Ankrd1, Nipal2
[0064] 14: Bmp5, Pdgfra, Ednra, Pde3a, Tbx20, Tcf21, Nr1h5, Gdf2, Srpx2,Lamb1
[0065] 15: Nefl, Stmn3, Gap43, Stmn2, Tubb3, Sncg, Tubb2b, mt-Atp6, mt-Cytb,Map1b
[0066] 16: Klra6, Klra1, Klra9, Klra7, Rasgrf1, Itgad, Ccdc136, Gzmb, Cd160,Cd7
[0067] 17: Esco2, Mxd3, Pclaf, Pimreg, Ska1, Pbk, Cenpf, Ccna2, Sgo1, Shcbp1
[0068] 18: Zmat4, Ncr1, Klrb1a, Cd200r2, Klrb1b, Gm43647, Xcl1, Dscam,Klrb1c, Gm36723
[0069] 19: Iglv1, Jchain, Tnfrsf17, Igha, Derl3, 5830418P13Rik, Igkc, Ccr10,Bhlha15, Eef1akmt3
[0070] 20: Gm38410,
[0071] 21: Siglech, Gm21762, Klk1, Grm8, Cox6a2, Gm30605, Havcr1, Cd209d, Paqr5, Duxbl1
[0072] 22: Skap1, Nkg7, Itk, Ccl5, Rpl37a, Cxcr6, Il2rb, Thy1, Gm2682, Cd226
[0073] 23: Cma1, Klra4, Klra8, Klri2, Gzma, Adamts14, Gm43647, Prf1, Ncr1,Itga2
[0074] 24:6030498E09Rik, Fam20a, Nova2, Gmpr, Rgs7bp, Pde10a, Pakap, Ptch2, Klhl4, Gm57038
[0075] 25: Chst4, Bnc1, Muc16, Upk3b, Myl7, Msln, Upk1b, Gm36535, Lrrn4, Pkhd1l1
[0076] 26: Negr1, Fcrl5, Gm26620, Cd300e, 4931431B13Rik, Klra17, Adamdec1, Itgad, Fam89a, Sox5
[0077] 27: Ebf1, Iglc2, Cd79a, Ighd, Ms4a1, Bank1, Pax5, Iglc3, Blk, Aff3
[0078] 28: Ms4a2, Mcpt8, Cyp11a1, Fcer1a, Cd200r3, Cpa3, Il6, Gm32511, Alox15, Csrp3
[0079] ## Please output the cell types of each row. The format is asfollows. Please output the result directly without further explanation.(Chinese translation: ## Please output the cell types of each row. The format is as follows. Please output the result directly without further explanation.)
[0080] 0 Monocytes(Chinese translation: Single cells)
[0081] 1 Neutrophil(Chinese translation: Neutrophil)
[0082] 2 B cells”(Chinese translation: B cells)。
[0083] Submission to the large language model and result determination: Submit the prompts to the Chat GPT-4 model through the API, and repeat 5 independent questions. During the questioning process, ensure that each submitted prompt is accurate and independent of the previous questions to obtain diverse model responses. Aggregate the 5 responses of the model, and determine the cell type corresponding to each subgroup according to the concept of mode in statistics. If 3 out of 5 responses of a certain subgroup are determined to be "Monocytes", then this subgroup is finally annotated as "Monocytes". Organize and analyze the annotation results, and compare them with the manual annotation results. In the actual case of single-cell annotation of mouse liver, the results show that among 28 subgroups, 20 subgroups are completely consistent with the manual annotation results, and another 3 subgroups are manually annotated as unknown cells (unknow), but the large language model successfully returned the annotation results. This fully demonstrates the accuracy and effectiveness of the method of the present invention in single-cell annotation of mouse liver.
[0084] Comparison table of single-cell annotation of mouse liver:
[0085] Cell subset Manual annotation AI annotation result First output Second output Third output Fourth output Fifth output 0 Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell 1 Kupffer cell Kupffer cell Kupffer cell Kupffer cell Kupffer cell Kupffer cell Kupffer cell 2 B cell B cell B cell B cell B cell B cell B cell 3 Macrophage Macrophage Macrophage Macrophage Macrophage Neutrophil Macrophage 4 Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell Endothelial cell 5 T cell T cell T cell T cell T cell T cell T cell 6 T cell T cell T cell T cell T cell T cell T cell 7 Neutrophil Neutrophil Neutrophil Neutrophil Neutrophil Neutrophil Neutrophil 8 T cell T cell T cell T cell T cell T cell T cell 9 T cell T cell T cell T cell T cell T cell T cell 10 Hepatocytes Hepatocyte Hepatocyte Hepatocyte Hepatocyte Hepatocyte Hepatocyte 11 Endothelial cell Stellate cell Hepatic stellate cell Stellate cell Stellate cell Endothelial cell Stellate cell 12 Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell 13 Cholangiocytes Hepatocyte Hepatocyte Hepatocyte Hepatocyte Hepatocyte Hepatocyte 14 HSC Mesenchymal cell Mesenchymal stem cell Mesenchymal cell Mesenchymal cell Endothelial cell Mesenchymal cell 15 unkown Neuron Neuron Neuron Neuron Neuronal cell Neuron 16 NK cell NK cell Natural killer cell NK cell NK cell NK cell NK cell 17 Diving_cell Cell cycle cell Proliferating cell Cell cycle cell Cell cycle cell Cell divisioncycle Cell cycle cell 18 NK cell NK cell Natural killer cell NK cell NK cell NK cell NK cell 19 Plasma_B B cell B cell B cell B cell B cell B cell 20 Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell Dendritic cell 21 Plasmacytoid_Dendritic_cell Plasmacytoiddendritic cell Dendritic cell Plasmacytoiddendritic cell Plasmacytoiddendritic cell Dendritic cell Plasmacytoiddendritic cell 22 T cell T cell T cell T cell T cell T cell T cell 23 NK cell NK cell Natural killer cell NK cell NK cell NK cell NK cell 24 Endo Neuron Neuron Neuron Neuron Neuronal cell Neuron 25 unkown Epithelial cell Cholangiocyte Epithelial cell Epithelial cell Epithelial cell Epithelial cell 26 Macrophages B cell B cell B cell B cell NK cell B cell 27 unkown B cell B cell B cell B cell B cell B cell 28 Basophill Mast cell Mast cell Mast cell Mast cell Mast cell Mast cell
[0086] Example 2: Annotation of sheep dorsal muscle cells.
[0087] Data preparation and preprocessing (similar steps to Example 1): Collect single-cell transcriptome data of sheep dorsal muscle, and perform strict quality control and differential analysis. Since sheep is a non-model species, special attention should be paid to data normalization and outlier handling during data preprocessing to ensure the accuracy of subsequent homologous alignment.
[0088] Homology alignment and gene screening / sorting: Gene symbol numbers are obtained by homology alignment with human genes. Using BLASTP software, gene sequences are sorted according to a set homology threshold (E-value < 1 × 10⁻⁶). -5 (Identity ≥ 0.8) Genes from sheep back muscle cells were compared with human genes to screen for homologous genes and determine their symbol numbers. Based on the homologous genes, the upregulated genes in each subgroup were identified and sorted according to the screening criteria and sorting method described in the invention. For a specific subgroup of sheep back muscle cells, after homologous alignment and screening sorting, a set of representative upregulated genes and their corresponding sorting information were obtained.
[0089] Cue word construction: Cue words were constructed by extracting the top 20 genes from the subpopulation gene sequencing results (since this is a non-model species). When constructing the cue words, accurate information such as sheep species, back muscle tissue, and genetic information was entered to ensure that the cue words accurately reflect the characteristics of sheep back muscle cells. The constructed cue words are:
[0090] "You are a biologist, and you are very good at cell annotation."
[0091] Identify the cell types of sheep dorsal muscle using the following markers, one for each row. Some cells may be a mixture of multiple cell types.
[0092] The following is the data to be analyzed. Each row represents one type of cells and the corresponding cell markers. Cell ID is before the semicolon, and gene markers are after the semicolon.
[0093] 0: PACRG, CCDC85A, PRAG1, BTNL9, LNX1, LRRC36, ZNF366, APOL3, DACH1,CYYR1, LRRC3B, ADGRL4, SHANK3, RASGRF2, CA4, TSPAN13, ROBO4, ABLIM3, Stox2,VWF
[0094] 1: ncbi_102185059, UCP3, ncbi_108636275, PEBP4, MYBPC2, PFKM, PFKFB4,SH3RF2, DDIT4L, NOS1, MYH2, FBP2, ncbi_106503003, TMEM233, ASB4, Igfn1, MYH1,ATP2A1, AMPD1, ACTN3
[0095] 2: KCNJ8, OTOGL, SEMA5B, GPRIN3, EGFLAM, CNTN4, ADRA1D, ncbi_108633370, SCG3, ncbi_102187240, RERG, NDUFA4L2, CPM, TMEM176B, CDH6, PDGFRB,COX4I2, C8orf48, GUCY1B1, THBS4
[0096] 3: PAX7, CHRDL2, ST8SIA1, RYR3, MEGF10, ncbi_106503297, NCAM1,C1QTNF3, APOE, Asic2, Hs6st2, PPP1R14B, Adamts20, CDH4, CADM2, GPC3, TAGLN3,ncbi_102186022, ATRNL1, C3orf70
[0097] 4: PRG4, FMO2, NCAM2, CYP2F3, DPP4, ADAM22, CFD, BMPER, RIPOR3,CDH18, FBN2, ncbi_102190915, NOX4, ADAM19, HSD11B1, MYOC, KAZN, LUM, C3,GSTM1
[0098] 5: Myh7b, MYOZ2,
[0099] 6: SERPINF1, PCOLCE2, PI16, COL14A1, MFAP5, COL1A1, VIT, S100A4, NOVA1, COL1A2, TNXB, FN1, UST, DPT, PPL, ABI3BP, PDZRN4, MEDAG, Fbln2, LVRN
[0100] 7: ADRA1D, CPM, OTOGL, KCNJ8, ncbi_102187240, RGS5, APOLD1, TMEM176B, ncbi_108633370, NES, SEMA5B, STAC, RERG, FAM162B, GPRIN3, PDGFRB, NDUFA4L2, TRPC3, EGFLAM, ANO1
[0101] 8: ITK, CD2, THEMIS, CTSW, KLRK1, BCL11B, GRAP2, IL7R, KLRC1, GZMH, TRGC1, CD3E, CD3D, Klre1, ncbi_106502126, CD3G, ncbi_108637283, STAT4, Tox2, HCST
[0102] 9: AZGP1, ACVR1C, MAL2, COL22A1, ABCD2, ACSM1, KLB, PLIN1, APOL6,NNAT, PKP2, ADIPOQ, CHL1, SPEF2, ASGR1, CIDEC, ADIG, SLC22A4, NALCN, SLC1A3
[0103] 10: MYOCD, MYH11, ACTG2, KCNAB1, NRIP2, RCAN2, TGFB3, CNN1, LSAMP,DGKG, ITGA8, CKB, KCTD16, Sorbs2, SLC22A3, ANKRD6, COL4A5, RYR2, DSTN, LBH
[0104] 11: TNNC1, MYL2, CKM, TNNI1, TNNI2, MB, ENO3, MYOZ1, MYLPF, TNNC2, DES, ACTA1, SLN, MYH7, C1QTNF9, KLF2, ALDOA, CRYAB, PGAM2, ID1
[0105] 12: VSIG4, C4BPA, BLNK, ncbi_102188010, RAB3C, CD163, C1QA, F13A1,C1qc, C1QB, BRN, CSF1R, BTK, CTLA4, P2RY12, MRC1, ncbi_106503915, CD74, HLA-DRA, CD86
[0106] 13: TNNC1, CCN1, CRYAB, MB, TMEM176B, Klhl29, ACTA1, MYH1, CPM, ncbi_102187240, ATP2A1, FAM162B, Ppp1r27, Ntn4, CKM, Tnnt3, TNNI2, ADCY5, COX6A2, SEMA5B
[0107] 14: MMRN1, CCL21, RELN, SCN3A, PTN, C8orf34, SEMA3D, NR2F1, KCNIP4, TC2N, PIEZO2, TNFAIP8L3, DLGAP1, TLL1, Pard6g, LYPD6, RAB11FIP1, RALYL, PKHD1L1, Bst1
[0108] 15: CDH1, ARMC4, Clstn2, Map3k7cl, Sema3a, NCKAP5, EDA2R, CDKN1A,NR5A2, NRK, ART3, FST, ABI3BP, PARM1, Bmpr1b, GPC6, ARPP21, UNC45B, MYBPH, KCNQ5
[0109] 16: BMX, GJA5, ncbi_106502978, PRDM16, APOA1, DKK3, THSD7A, HMCN1, EFNA5, ADAMTS1, PLCG2, COL8A1, NEBL, ncbi_102181308, HEY1, CTNNAL1, PCSK5, ABCB1, LTBP1, . TSPAN1
[0110] 17: NELL1, PLVAP, SYT15, VCAM1, ADGRG6, Klhl14, IL1R1, HMCN1, KIF26B, CYP7B1, THSD7A, FABP4, NOS1AP, CTNNAL1, MMP16, CALCRL, Ehd4, Rgs6, AQP1, TSPAN12
[0111] 18: FMO2, CDH18, PDGFRA, PCSK6, PRG4, COL6A3, SHCBP1L, NOX4, DCN,CCDC80, SPRY2, GSN, NUPR1, BMPER, CYP2F3, ncbi_102189207, Col6a2, ncbi_102190915, LUM, RGMA
[0112] 19: ncbi_106503297, CHRDL2, NCAM1, RYR3, Asic2, PAX7, APOE, GREB1, ST8SIA1, MEGF10, CADM2, CDH4, TAGLN3, CDH6, PPP1R14B, TMEM176B, FAM162B, ANO1, Adamts20, THBS4
[0113] 20: TRPM3, ID4, LAMA3, Slit3, CNN1, LRRC17, Prrx1, MAMDC2, TRPC6,FST, SEMA5A, Sh3gl2, RIMS1, ADGRL3, PTGDS, NHS, IGFBP5, LSAMP, NCKAP5, Synpo2
[0114] 21: ST8SIA1, C1QTNF3, RYR3, CHRDL2, NCAM1, CDH4, SLC16A10, PAX7,Adamts20, ncbi_108636275, EDA2R, Hs6st2, FBXO32, BTC, LRRTM3, KCNN3, CHRNB1,ASB4, SLIT2, PFKM
[0115] 22: ncbi_108638533, ncbi_106503398, ncbi_106503383, ncbi_102168946,ncbi_102191280, ncbi_108637754, GADL1, ncbi_108638502, KLHL38, NR5A2, MYBPH,LRRC2, EDA2R, NEURL1, PLAAT1, SCN4A, ARPP21, ncbi_108635579, TMEM182, CPED1
[0116] 23: SLC28A3, SERPINA1, CSF3R, VNN2, MS4A8, SDS, S100A8, S100A12,NCF1, SELL, HCK, C1orf162, MCTP2, CCL23, C5AR1, PLEK, RIPOR2, ncbi_102175442,SYK, EMB
[0117] 24: CDH19, ncbi_102189639, IL1RAPL2, UGT8, Xkr4, ATP10B, Nrxn3,GRIK2, SORCS1, PLP1, PRIMA1, ERBB3, PTPRZ1, GRIK3, ADAM23, PPP2R2B, TMEM178B,NLGN1, SCN7A, CNTNAP2
[0118] 25: KRT8, KRT18, KRT19, CLDN20, CLDN1, ITGB4, ROR2, SBSPON, KLF5,ncbi_106503347, Il1rapl1, FZD1, Robo2, PDLIM2, SFRP5, EPHB2, ncbi_102172959,SMOC2, UNC5C, DOCK5
[0119] ## Please output the cell types of each row. The format is as follows. Please output the result directly without further explanation.
[0120] 0 Monocytes
[0121] 1 Neutrophil
[0122] 2 B cells”
[0123] Submit the large language model and result determination (similar steps as in Example 1): Submit the prompt to the Chat GPT-4 model for 5 independent questions and summarize the reply results. When processing the annotation of sheep dorsal muscle cells, since sheep is a non-model species, some gene expression patterns and biological characteristics different from those of model species may be encountered. However, the method of the present invention can still effectively perform cell type annotation through homologous alignment and the powerful learning ability of the large language model. Determine the cell type according to the concept of mode in statistics and compare it with the manual annotation results. In the case of sheep dorsal muscle cell annotation, the results show that among 25 subpopulations, 16 subpopulations are consistent with the manual annotation results. In addition, 4 subpopulations are manually annotated as unknown cells (unknow), while the large language model successfully gives the annotation results. This further proves the feasibility and advantages of the method of the present invention in the cell annotation of single-cell transcriptomes of non-model species.
[0124] Comparison table of sheep dorsal muscle cell annotation:
[0125] Cell subsets Manual annotation AI annotation result First output Second output Third output Fourth output Fifth output 0 Endothelialcells Endothelialcells Endothelialcells Endothelial Cells Endothelial Cells Endothelial cells Endothelial cells 1 Myofiber Myofiber Myofiber Muscle Cells Skeletal Muscle Cells Muscle fibers Myofibers 2 Smoothmuscle cells Smoothmuscle cells Smooth musclecells Fibroblasts Smooth Muscle Cells Smooth musclecells Pericytes 3 Satellitecells Satellitecells Satellite cells Satellite Cells Satellite Cells Satellite cells Satellite cells 4 Fibroblasts Fibroblasts Fibroblasts Fibroblasts Fibro-AdipogenicProgenitors (FAPs) Fibroblasts Fibro-adipogenicprogenitors (FAPs) 5 Myofiber Myofiber Myofiber Muscle Cells Cardiac Muscle Cells Muscle fibers Cardiomyocytes 6 Tenocyte Fibroblasts Fibroblasts Fibroblasts Fibroblasts Fibroblasts Tenocytes 7 Smoothmuscle cells Smoothmuscle cells Smooth musclecells Pericytes Vascular Smooth MuscleCells Smooth musclecells Pericytes 8 [[ID=八]]T cell T cells T cells T Cells T Cells T cells T cells 9 Unknow Adipocytes Adipocytes Adipocytes Adipocytes Adipocytes Adipocytes 10 SmoothMuscle Cells SmoothMuscle Cells Myofiber Smooth MuscleCells Myofibroblasts Smooth musclecells Smooth muscle cells 11 Endothelial Myofiber Myofiber Muscle Cells Myofibers Muscle fibers Skeletal muscle cells 12 B cell Macrophages Macrophages Macrophages Macrophages Macrophages Macrophages 13 Unknow Myofiber Myofiber Muscle Cells Myofibers Muscle fibers Myofibers 14 Unknow SchwannCells Fibroblasts Fibroblasts Schwann Cells Schwann cells Schwann cells 15 Myoblast Endothelialcells Endothelialcells Epithelial Cells Muscle ProgenitorCells Schwann cells Fibroblasts 16 Endothelialcells Endothelialcells Endothelialcells Smooth MuscleCells Pericytes Endothelial cells Vascular smooth musclecells 17 Endothelialcells Endothelialcells Endothelialcells Endothelial Cells Vascular EndothelialCells Endothelial cells Vascular endothelialcells 18 Fibroblasts Fibroblasts Fibroblasts Fibroblasts Fibro-AdipogenicProgenitors (FAPs) Fibroblasts Fibroblasts 19 Unknow Satellitecells Satellite cells Satellite Cells Satellite Cells Satellite cells Satellite cells 20 Pericytes Myofibroblasts Myofibroblasts Smooth MuscleCells Myofibroblasts Fibroblasts Myoendothelialprogenitors 21 Satellitecells Satellitecells Satellite cells Satellite Cells Satellite Cells Satellite cells Satellite cells 22 Myofiber Myofiber Myofiber Smooth MuscleCells Muscle ProgenitorCells Schwann cells Neuromuscular junctioncells 23 Neutrophils Neutrophils Neutrophils Neutrophils Neutrophils Neutrophils Neutrophils 24 Schwann cell Schwanncells Schwann cells Glial Cells Schwann Cells Schwann cells Oligodendrocytes 25 Fibroblast Epithelialcells Epithelialcells Epithelial Cells Epithelial Cells Epithelial cells Epithelial cells
Claims
1. A method for single-cell transcriptome cell annotation of a fusion large language model, characterized by, Comprising the following steps: S1, data preprocessing: Preprocessing of single-cell transcriptome data obtained by new sequencing, including data cleaning, quality control and construction of gene expression matrix; S2, obtaining and sorting up-regulated genes in each cell subpopulation: For each gene in a given subpopulation, compare its expression value with other cell subpopulations using Wilcoxon rank sum test to determine significantly up-regulated genes in each cell subpopulation, the specific criteria for the significantly up-regulated genes are: the gene is expressed in at least 25% of the cells within the target subpopulation, P value ≤ 0.01 and log2 fold change ≥ 0.36; Sort the up-regulated genes of each subpopulation based on the results of differential analysis, preferentially sort according to P value from small to large, when P value is the same, further sort according to fold change from large to small; S3, homologous alignment of non-model species: When the annotation target cell is a non-model species, homologous alignment is required to determine the gene symbol number, if the target species is an animal, select human or mouse as the reference model species, if the target species is a plant, select Arabidopsis thaliana as the reference model species; Homologous relationship between non-model species and reference model species genes was determined by using BLASTP software for homologous alignment, and the symbol number of the gene of the non-model species was defined based on the symbol number of the reference model species gene, with the threshold value of homology set as E-value < 1 x 10 -5 , Identity≥0.8, and the gene with the smallest E-value value was selected as the homologous gene derived from the reference model species according to the E-value ranking; S4, constructing a proprietary prompt word: Record the information of species, tissue, subpopulation up-regulated ranking gene name as a variable, and define the input and output formats, construct the prompt word for submitting to the large language model, for each subpopulation of the model species, select the top 10 genes in the ranking to construct the prompt word, for each subpopulation of the non-model species, select the top 20 genes in the ranking to construct the prompt word; S5, submitting the large language model and judging the cell type: Submit the constructed prompt word to the large language model through the application programming interface (API), repeat the question and answer multiple times, summarize the replies of the large language model, and determine the cell type and marker gene corresponding to each subpopulation according to the concept of mode in statistics.
2. A single-cell transcriptome cell annotation system of a fusion large language model, characterized by, Comprising: S1, data preprocessing module: Used for receiving single-cell transcriptome data, and performing quality control and standardization processing on the data; S2, gene sorting and screening module: Used to realize the function of step S2 in claim 1, obtain and sort up-regulated genes in each cell subpopulation, and screen and sort significantly up-regulated genes according to the set standard; For each gene in a given subpopulation, compare its expression value with other cell subpopulations using Wilcoxon rank sum test to determine significantly up-regulated genes in each cell subpopulation, the specific criteria for the significantly up-regulated genes are: the gene is expressed in at least 25% of the cells within the target subpopulation, P value ≤ 0.01 and log2 fold change ≥ 0.36; Sort the up-regulated genes of each subpopulation based on the results of differential analysis, preferentially sort according to P value from small to large, when P value is the same, further sort according to fold change from large to small; S3, homologous alignment module: When processing non-model species data, used to realize the function of step S3 in claim 1, determine the symbol number of non-model species genes, if the target species is an animal, select human or mouse as the reference model species, if the target species is a plant, select Arabidopsis thaliana as the reference model species; Homologous relationship between non-model species and reference model species genes is determined by using BLASTP software for homologous alignment, symbol number of the gene of the non-model species is defined based on the symbol number of the reference model species gene, parameter setting is that threshold value of homology is E-value < 1 x 10 -5 , Identity ≥ 0.8, and the gene with the minimum E-value value is selected as the homologous gene derived from the reference model species according to E-value ranking; S4, prompt word construction module: constructing a special prompt word according to the information provided in step S2 of claim 1 and the genes screened in step S3 of claim 1, and selecting 10 top-ranked genes for each subpopulation of model species to construct a prompt word, and selecting 20 top-ranked genes for each subpopulation of non-model species to construct a prompt word; S5, a large language model interaction module: responsible for interacting with a large language model, submitting a prompt word to the large language model and receiving its reply; Through an optimized API calling mode, efficient communication with the large language model is ensured, stable performance of multiple question-answering and accurate summarization of results are realized, and problems such as data transmission errors and communication interruptions are avoided to affect the determination of cell types; S6, a cell type determination module: According to the reply obtained in step S5 of claim 1, the cell type is determined according to the concept of mode in statistics, the results of each question-answering are accurately recorded and analyzed when the large language model replies, and a rigorous determination is made according to the concept of mode in statistics. For the case where the frequency is the same or similar, it has the ability to further analyze and judge, so as to improve the accuracy and reliability of the determination of cell types.
Citation Information
Patent Citations
Method for annotating cell identities based on single cell transcriptome clustering results
CN110060729A
Non-model species cell annotation method for bidirectional homologous comparison
CN115083523A
Cited By
Large language model driven single-cell double-score iterative annotation method
CN122392655A
Large language model driven single-cell double-score iterative annotation method
CN122392655B