Cell type annotation method and device based on plant single cell transcriptome data and readable storage medium thereof
By integrating multi-level methods of marker gene expression, cross-dataset correlation, differential gene enrichment and quasi-time sequence analysis, the problem of low annotation accuracy of plant single-cell transcriptome data in the prior art is solved, and efficient and systematic cell type annotation of a variety of plant species is achieved.
Patent Information
- Application Number
- CN202510841143.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology relies on artificial experience and limited basic data sets, resulting in low cell type annotation accuracy of plant single-cell transcriptome data, which is particularly difficult to apply to non-modal plant species, lacks a systematic multi-dimensional integration method, and is not standardized in the process.
By integrating marker gene expression analysis, cross-dataset correlation analysis, differential gene function enrichment analysis and quasi-time sequence trajectory analysis, a multi-level cell type annotation method is constructed, and a homologous gene alignment and standardized process is used to reduce artificial dependence, and annotate the dynamic development trajectory from the gene layer to the subpopulations.
It significantly improves the accuracy and universality of cell type annotation, can be applied to a variety of plant species, especially non-modal plants, reduces the impact of artificial judgment, and forms a reproducible systematic method.
Smart Images

Figure CN120356520A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, specifically to the analysis technology of plant single-cell transcriptome data, and particularly to a method for annotating plant cell types based on multi-dimensional data integration. Background Art
[0002] In the analysis of plant single-cell transcriptome data, cell type annotation is a key step in parsing cell heterogeneity. The current mainstream methods mainly rely on the knowledge and experience of analysts for manual annotation, lacking a systematic methodology. Although some automatic annotation software can assist in the analysis, it highly depends on high-quality basic data sets (such as standard single-cell expression atlases) and manual judgment, with limited universality. In addition, although the preprocessing processes such as quality control, normalization, dimensionality reduction, and clustering of plant single-cell transcriptome data have been relatively mature, the publicly available basic data sets of plant single-cell transcriptome are scarce. Especially for non-model plant species, the available marker genes and cross-species data resources are limited, resulting in low accuracy of cell type annotation, fragmented processes, and difficulty in meeting diverse research needs.
[0003] Therefore, there is an urgent need for a method, device, and readable storage medium for annotating cell types based on plant single-cell transcriptome data that can effectively solve the problems existing in the prior art, realize multi-level annotation from the gene level to the dynamic development trajectory of subpopulations, and significantly improve the accuracy and universality of annotation. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, and readable storage medium for annotating cell types based on plant single-cell transcriptome data, aiming at the problems existing in the current technology, such as relying on manual experience and limited basic data sets, lacking a systematic multi-dimensional integration method, resulting in low accuracy of cell type annotation for plant single-cell transcriptome data, irregular processes, and especially difficulty in being applicable to non-model plant species.
[0005] The core technology of the present invention mainly constructs a systematic method for annotating cell types of plant single-cell transcriptome data by integrating multi-dimensional information such as marker gene expression analysis, cross-dataset correlation analysis, differential gene function enrichment analysis, and pseudotime trajectory analysis, realizes multi-level annotation from the gene level to the dynamic development trajectory of subpopulations, and significantly improves the accuracy and universality of annotation.
[0006] In the first aspect, the present invention provides a method for annotating cell types based on plant single-cell transcriptome data, and the method includes the following steps: S1. For the plant single-cell transcriptome expression matrix that has been clustered, obtain cell type marker genes based on a database or literature, and initially annotate the types of cell clusters by visually analyzing the expression of cell type marker genes in each cell cluster; S2. Collect conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same or different species as known datasets, calculate the expression correlation between the data to be analyzed and the known datasets, and assist in annotating the cell cluster types based on the expression correlation results; S3. Identify the differentially expressed genes in each cell cluster, perform functional enrichment analysis on the differentially expressed genes, and infer the cell types based on the enrichment results; S4. For cell clusters that are preliminarily annotated as the same cell type or have heterogeneity, perform pseudotime analysis to construct the cell developmental trajectory, and annotate the cell subsets based on the cell developmental trajectory and the enrichment results of the differential genes in the trajectory; S5. Integrate the results of steps S1 - S4 to generate a comprehensive annotation of plant cell types.
[0007] Furthermore, in step S1, if the target species lacks cell type marker genes, the homologous genes are obtained by the following method: Obtain the amino acid sequences of the marker genes of the model species or related species from the genomic database, and obtain the homologous genes of the target species through the sequence alignment algorithm. The E - value threshold of the sequence alignment algorithm is ≤1×10 -5 .
[0008] Furthermore, in step S2, the expression correlation is the Pearson correlation coefficient, which is calculated by calculating the Pearson correlation coefficient of the column vectors of the gene expression matrices of the data to be analyzed and the known datasets; Perform batch effect removal on datasets from different sources before analysis.
[0009] Furthermore, in step S3, the screening conditions for differentially expressed genes are: adjusted P - value < 0.01 and fold change ≥ 2 - fold.
[0010] Furthermore, in step S4, the pseudotime analysis distinguishes the developmental stages or functional states of cell subsets by constructing a single - cell gene expression trajectory and combining the functional enrichment results of highly expressed genes at different stages in the trajectory.
[0011] Furthermore, the batch effect removal is performed using the ComBat function.
[0012] Furthermore, the plant includes model plants or non - model plants. The model plants include Arabidopsis thaliana and rice, and the non - model plants include soybean.
[0013] In the second aspect, the present invention provides a cell type annotation device based on plant single - cell transcriptome data, including: A marker gene expression pattern identification module, for the plant single - cell transcriptome expression matrix that has been clustered, based on the database or literature, obtain cell type marker genes, and preliminarily annotate the types of cell clusters by visually analyzing the expression of cell type marker genes in each cell cluster; An expression correlation module that collects conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same or different species as known datasets, calculates the expression correlation between the data to be analyzed and the known datasets, and assists in annotating cell cluster types based on the expression correlation results; A cell type inference module that identifies differentially expressed genes in each cell cluster, performs functional enrichment analysis on the differentially expressed genes, and infers cell types based on the enrichment results; A cell subpopulation annotation module that performs pseudotime analysis on cell clusters initially annotated as the same cell type or having heterogeneity to construct a cell development trajectory, and annotates cell subpopulations based on the cell development trajectory and the enrichment results of the differential genes in the trajectory; A comprehensive annotation generation module that integrates the results of the above modules to generate a comprehensive annotation of plant cell types.
[0014] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned cell type annotation method based on plant single-cell transcriptome data.
[0015] In a fourth aspect, the present invention provides a readable storage medium in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the cell type annotation method based on plant single-cell transcriptome data as described above.
[0016] The main contributions and innovations of the present invention are as follows: 1. Break through the limitations of a single method and achieve systematic annotation The prior art relies on manual experience or a single automatic tool. The present invention forms a multi-level annotation system by integrating gene expression analysis (marker genes / homologous genes), cross-dataset correlation analysis (intra-species / cross-species transcriptome data), functional enrichment analysis (GO pathways), and pseudotime dynamic trajectory analysis, avoiding the one-sidedness of a single method and significantly improving the comprehensiveness and accuracy of annotation.
[0017] 2. Alleviate the problem of lack of basic datasets and be applicable to non-model species In view of the current situation of insufficient plant single-cell basic data, the present invention uses data of model species (such as Arabidopsis thaliana, rice) to assist in annotating non-model species (such as soybean) through homologous gene alignment (such as BLASTP, E value ≤ 1×10 -5 ), breaking through the "data dependence" bottleneck and broadening the scope of application of the method.
[0018] 3. Dynamically analyze cell heterogeneity and achieve refined annotation Existing static clustering methods are difficult to distinguish the functions of cell subsets. By performing pseudotime analysis (such as Monocle2) to construct the cell development trajectory and combining the enrichment results of trajectory-differential genes (such as stress response and nutrient metabolism pathways), the functional states of subsets within the same cell cluster (such as early / late cells) can be accurately distinguished, filling the gap in the annotation of cell dynamic heterogeneity in the existing technology.
[0019] 4. Reduce manual dependence and improve annotation standardization By means of a standardized process (such as identification of marker gene expression patterns, calculation of Pearson correlation coefficients, screening thresholds for differential genes, and gene function enrichment analysis), the influence of the subjective judgment of analysts is reduced, forming a reproducible systematic method to solve the problems of fragmented manual annotation processes and unreliable results in the existing technology.
[0020] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects, and advantages of the present invention will become more apparent and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a UMAP visualization result of cell clustering and a marker gene annotation result diagram for annotating single-cell transcriptome data of rice seed embryos imbibed for 6 hours based on a known data set according to Embodiment 2 of the present invention; Figure 2 is a diagram showing the expression of marker genes of rice vascular cells in the data to be analyzed according to Embodiment 2 of the present invention; Figure 3 is a diagram showing the expression of marker genes of coleoptile cells in the data to be analyzed according to Embodiment 2 of the present invention; Figure 4 is a correlation heat map of single-cell transcriptome data (abscissa) of rice seed embryos imbibed for 6 hours and spatial transcriptome data (ordinate) treated in the same way according to Embodiment 2 of the present invention; Figure 5 is a correlation heat map of single-cell transcriptome data (abscissa) of rice seed embryos imbibed for 6 hours and single-cell transcriptome data (ordinate) of rice seed embryos imbibed for 24 hours according to Embodiment 2 of the present invention; Figure 6 is a UMAP visualization result of cell clustering and an annotation result diagram based on the correlation with known data according to Embodiment 2 of the present invention; Figure 7 is a dimensionality reduction diagram of pseudotime analysis according to Embodiment 2 of the present invention; Figure 8 It is a visualization diagram of differentially expressed genes in the pseudotime process according to Embodiment 2 of the present invention; Figure 9 It is a final annotation result diagram of single-cell transcriptome data of rice seed embryos imbibed for 6 hours according to Embodiment 2 of the present invention; Figure 10 It is a UMAP visualization result of cell clustering and a marker gene annotation result diagram for annotating single-nucleus transcriptome data of soybean cotyledon stage based on known cell type marker genes according to Embodiment 3 of the present invention; Figure 11 It is a diagram showing the expression of marker genes of soybean endosperm cells in the data to be analyzed according to Embodiment 3 of the present invention; Figure 12 It is a diagram showing the expression of marker genes of embryo cells in the data to be analyzed according to Embodiment 3 of the present invention; Figure 13 It is a correlation heatmap of single-nucleus transcriptome data (ordinate) of soybean cotyledon stage and transcriptome data of Arabidopsis seed development (abscissa) according to Embodiment 3 of the present invention; Figure 14 It is a UMAP visualization result diagram of cell clustering and an annotation result diagram based on correlation with known data according to Embodiment 3 of the present invention; Figure 15 It is a dimensionality reduction diagram of pseudotime analysis according to Embodiment 3 of the present invention; Figure 16 It is a visualization diagram of differentially expressed genes in the pseudotime process according to Embodiment 3 of the present invention; Figure 17 It is a final annotation result diagram of single-nucleus transcriptome data of soybean cotyledon stage according to Embodiment 3 of the present invention; Figure 18 It is a flowchart of a cell type annotation method based on plant single-cell transcriptome data according to Embodiment 1 of the present invention; Figure 19 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0022] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0023] It should be noted that: In other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0024] The prior art relies on manual experience and limited basic data sets, lacking a systematic multi-dimensional integration method, resulting in low accuracy of cell type annotation and non-standardized processes for plant single-cell transcriptome data, and it is particularly difficult to apply to non-model plant species. Model plants refer to plants with short growth cycles and small genomes. Arabidopsis thaliana is often selected as a model plant for biological experiments; non-model plants are plants other than model plants, such as soybeans.
[0025] Based on this, the present invention solves the problems existing in the prior art by integrating multi-dimensional information such as marker gene expression analysis, cross-dataset correlation analysis, differential gene function enrichment analysis, and pseudotime trajectory analysis.
[0026] Embodiment 1 The present invention aims to propose a cell type annotation method based on plant single-cell transcriptome data. Specifically, referring to Figure 18 , the method includes the following steps: S1. For the plant single-cell transcriptome expression matrix that has completed clustering, obtain cell type marker genes based on a database or literature, and initially annotate the types of cell clusters by visually analyzing the expression of cell type marker genes in each cell cluster. In this embodiment, the specific process steps are as follows: (1a) Preliminary analysis of single-cell transcriptome data: For the plant single-cell transcriptome expression matrix, preliminary analysis can be performed using software such as Seurat or Scanpy to obtain the results of cell dimensionality reduction and clustering, that is, cell cluster information.
[0027] (1b) Collection of cell type marker genes: Download databases such as PlantscRNAdb (such as the latest v4 version), scPlantDB (such as the latest v1.1 version), etc., or cell type marker genes of the corresponding species reported in existing literature.
[0028] If there are no cell type marker genes for the species under study, collect the homologous genes of marker genes from other species (such as model species Arabidopsis thaliana, rice, etc., or species closely related to the study species). Specifically, download the corresponding amino acid sequence file of the species from a genomic database such as Ensembl Plants, and use BLASTP software (such as version v2.5.0) for alignment, setting the E-value to 1×10 -5 And retain the best alignment result between the two species as the homologous gene between the two species.
[0029] (1c)Expression analysis of cell type marker genes: Use the DotPlot function or FeaturePlot function in Seurat (such as version v5.3.0) (or visualization functions such as dotplot and violin in Scanpy) to visualize the known cell type marker genes. If the gene expression is high, the cell type of this cell cluster is the corresponding cell type of these genes.
[0030] S2. Collect conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same or different species as known datasets, calculate the expression correlation between the data to be analyzed and the known datasets, and assist in annotating the cell cluster types based on the expression correlation results; In this embodiment, the specific process steps are as follows: (2a)Preparation of known datasets: Collect and download data related to the species and tissues under study from public databases such as NCBI and NGDC, including conventional transcriptome sequencing data of the same species and tissue (such as using techniques like laser microdissection to separate cell types), single-cell transcriptome sequencing data (such as based on platforms like 10X Genomics or BD Rhapsody), and spatial transcriptome sequencing data (such as based on platforms like 10X Visium or BGI Stereo-seq); conventional transcriptome sequencing data, single-cell transcriptome sequencing data, and spatial transcriptome sequencing data of different species and the same tissue, etc.
[0031] Among them, the data of different species can be associated by the homologous gene identification method described in step (1b). This step obtains the gene expression matrix of the known data, that is, the rows of the matrix are genes, and the columns are sample names / tissue names / cell type names, etc.
[0032] (2b)Calculation of the average gene expression value of the data to be analyzed: Obtain the average expression value of all genes in each cluster, such as using the AverageExpression function in the Seurat package. This step obtains the gene expression matrix of the data to be analyzed, that is, the rows of the matrix are genes, and the columns are cell clusters.
[0033] (2c)Correlation between the data to be analyzed and the known data: Calculate the Pearson correlation coefficient for each column (i.e., sample / tissue / cell type / cell cluster, etc.) in the expression matrix obtained in steps (2a) and (2b), such as using the cor function in R language. It should be noted that before calculating the Pearson correlation coefficient, it is necessary to perform batch effect removal on datasets from different sources, such as using the ComBat function in the R package sva (e.g., version v3.52.0).
[0034] S3. Identify the differentially expressed genes in each cell cluster, perform functional enrichment analysis on the differentially expressed genes, and infer the cell type based on the enrichment results; In this embodiment, the specific process steps are as follows: (3a)Identification of cell cluster differential genes: Identify the highly expressed genes in each cell cluster compared to other cells according to the cell clustering situation, such as using the FindAllMarkers function in the Seurat package for identification (other methods can also be used to identify differentially expressed genes), and screen genes with an adjusted P-value < 0.01 and a fold change of more than 2 times as candidate genes for subsequent enrichment analysis.
[0035] (3b)Differential gene enrichment analysis: Based on the differentially expressed genes identified in step (3a), perform functional enrichment analysis, i.e., gene ontology (GO) analysis. Specifically, use the enrichGO function in the R package clusterProfiler (e.g., version v4.12.6) to analyze which biological processes (BP), molecular functions (MF), and cellular components (CC) the highly expressed genes in each cell cluster are enriched in. Based on the enriched pathways and the biological characteristics of the samples, judge the cell type corresponding to the cell cluster. It should be noted that when using the enrichGO function, it is necessary to specify the annotation database corresponding to the species, and this data can be directly downloaded from the OrgDb topic on the Bioconductor website, or constructed using the amino acid sequences of the species.
[0036] S4. For cell clusters initially annotated as the same cell type or with heterogeneity, perform pseudotime analysis to construct the cell developmental trajectory, and annotate the cell subsets based on the cell developmental trajectory and the enrichment results of the differential genes in the trajectory; In this embodiment, the specific process steps are as follows: (4a)Determination of cell clusters for pseudotime analysis: Based on the analysis results of steps S1 to S3, each cell cluster can basically correspond to a specific cell type. For multiple cell clusters corresponding to the same cell type, or a single cell cluster possibly having multiple cell types, they can be separately extracted for pseudotime analysis. To extract the data set in the single-cell expression matrix, the subset function of the Seurat package can be used, or the expression values of the corresponding cell clusters can be extracted according to the cell clustering records in the metadata information.
[0037] (4b)Pseudotime analysis: Based on the cell clusters extracted in step (4a), the R package Monocle2 (v2.28.0) is used for pseudotime analysis. Specifically, the newCellDataSet function of the Monocle2 package is used to create an object for pseudotime analysis; the estimateSizeFactors function and the estimateDispersions function are used to normalize the data; the detectGenes function and the differentialGeneTest function are used to identify the differentially expressed genes during the pseudotime process; the reduceDimension function is used for data dimensionality reduction and the orderCells function is used to construct the pseudotime trajectory. For the differentially expressed genes identified during the pseudotime process, the plot_pseudotime_heatmap function of the Monocle2 package is used for visualization and gene classification (the classification criterion is genes highly expressed in different periods such as the early, middle, and late stages of pseudotime), and the number of classifications can be specified using the num_clusters parameter; at the same time, the cutree function in the stats package is used to extract the gene information of different classifications for gene enrichment analysis, and this analysis process refers to step S3. Finally, based on the pseudotime trajectory and the gene enrichment pathways highly expressed in different periods, more detailed subpopulation annotations are performed on the cell clusters, and at the same time, the cell type annotation results of the first three steps can also be verified.
[0038] S5. Integrate the results of steps S1 - S4 to generate a comprehensive annotation of plant cell types.
[0039] In this embodiment, different information obtained from steps S1 to S4 is integrated to comprehensively annotate the cell types of the plant to be studied. Specifically, the annotations at the gene level (step S1) and the dataset level (step S2) highly depend on the existing research status. For common plant species, corresponding data resources are available, while for special plant species, relevant information is lacking, and the cell types can only be indirectly annotated through the form of homologous gene alignment. Steps S1 and S2 provide a preliminary landscape of cell type annotation, while gene enrichment analysis (step S3) and pseudotime analysis (step S4) can perform more refined annotations on cell types and are more dependent on the biological background of the annotators.
[0040] Thus, the present invention integrates these 4 different annotation strategies to systematically annotate the single-cell transcriptome sequencing data of any plant (the present invention does not involve the pre-quality control, normalization, dimensionality reduction, clustering, etc. steps of single-cell transcriptome, and takes the single-cell transcriptome data after clustering as the object).
[0041] Embodiment 2 Based on the same concept, this embodiment proposes a cell type annotation method based on the single-cell transcriptome data of the embryo after 6 hours of rice seed imbibition. The specific steps are as follows: Step 1: Annotation based on known cell type marker genes Taking the single-cell transcriptome data of the embryo of rice seeds imbibed for 6 hours as an example for cell type annotation (data source: CNGBdb, number CNP0004152). After obtaining the expression matrix, perform preliminary data analysis, including quality control, normalization, dimensionality reduction, and clustering, to obtain the clustering information of each cell, that is, cell clusters ( Figure 1 , a total of 11 cell clusters are divided). At the same time, download the cell type marker genes of rice in the plant single-cell related databases PlantscRNAdb (such as version v4) and scPlantDB (such as version v1.1), and check the expression of these marker genes in the data to be analyzed ( Figure 2 and Figure 3 ). Due to the limitations of the marker genes included in these two databases, the results show that cell clusters 1, 3, 6, 9, and 11 may be cell types related to vascular tissue, and cell cluster 8 may be the coleoptile cell type. If only using the single method of marker genes, only the above-mentioned cell clusters 1, 3, 6, 8, 9, and 11 can be annotated (such as Figure 1 the blue font in), and the annotation is relatively rough.
[0042] Among them, Figure 1 represents the UMAP visualization result of cell clustering and the marker gene annotation result (blue font). Each point in the figure represents a cell. Figure 2Indicates the expression of marker genes for rice vascular cells in the data to be analyzed. Figure 3 Indicates the expression of marker genes for coleoptile cells in the data to be analyzed. The abscissa is the gene, and the ordinate is the cell cluster (corresponding to Figure 1 the 11 cell clusters in); the size of the circle represents the proportion of cells of this gene in the total number of cells in this cluster, and the depth of the color represents the average expression value of this gene.
[0043] Step 2: Annotate based on known datasets Download public datasets related to the process of rice seed germination. In this example, download the spatial transcriptome data during the process of rice seed germination (data source STOmics, number STT0000049) and the single-cell transcriptome sequencing data of rice seeds imbibed for 24 hours (data source CNGBdb, number CNP0004152). As Figure 4 、 Figure 5 shown, compare the expression of cell clusters in the data to be analyzed with the correlation of these two types of data respectively.
[0044] Results ( Figure 6 ) show that except for cell clusters 2 and 7 that cannot be annotated, other cell clusters can be annotated according to their correlation with other data. If only using the single method of correlation with known datasets, cell clusters 2 and 7 cannot be annotated.
[0045] Among them, the partial R code for the correlation between the data to be analyzed and the spatial transcriptome data ( Figure 4 ) is as follows: library(sva) exprSet11 = read.table("rice_stereo_6HAI_1_clusters_res1_avg.txt",sep="\t", header = TRUE) exprSet12 = read.table("rice_stereo_6HAI_2_clusters_res1_avg.txt",sep="\t", header = TRUE) exprSet13 = read.table("rice_stereo_6HAI_3_clusters_res1_avg.txt",sep="\t", header = TRUE) exprSet14 = read.table("rice_stereo_6HAI_4_clusters_res1_avg.txt",sep="\t", header = TRUE) exprSet = cbind(exprSet11, exprSet12[rownames(exprSet11),], exprSet13[rownames(exprSet11),], exprSet14[rownames(exprSet11),]) exprSet = na.omit(exprSet) exprSet2<- read.table('rice_6HAI_clusters_res1_avg_BD.txt', header =TRUE) exprSet2 = exprSet2[rowSums(exprSet2)>1,] exprSet2 = exprSet2[rownames(exprSet),] exprSet2 = na.omit(exprSet2) dat = cbind(exprSet[rownames(exprSet2),], exprSet2) colnames(dat) dat_temp = dat a = ComBat(scale(dat_temp), batch = c(rep('batch1',40), rep('batch2',11))) M3<- cor(a, method ='spearman') pheatmap::pheatmap(M3[1:40,41:51], display_numbers = TRUE) The partial R code for the correlation between the data to be analyzed and the single-cell transcriptome data of rice seeds at 24 hours of imbibition ( Figure 5 ) is as follows: library(sva) aa<- read.table('rice_24HAI_known_avg_BD.txt', header = TRUE) colnames(aa) = paste0('24HAI_',colnames(aa)) aa = aa[rowSums(aa)>1,] bb <- read.table('rice_6HAI_clusters_res1_avg_BD.txt', header = TRUE) colnames(bb) <- paste0('6HAI_', colnames(bb)) bb <- bb[rowSums(bb) > 1, ] dat <- cbind(aa, bb[rownames(aa), ]) dat <- na.omit(dat) colnames(dat) a <- ComBat(scale(dat), batch = c(rep('batch1', 14), rep('batch2', 11))) M3 <- cor(a, method ='spearman') pheatmap::pheatmap(M3[1:14, 15:25]) Figure 4 A heatmap showing the correlation between single-cell transcriptome data (abscissa) of rice seeds imbibed for 6 hours and spatial transcriptome data (ordinate) under the same treatment. Figure 5 A heatmap showing the correlation between single-cell transcriptome data (abscissa) of rice seeds imbibed for 6 hours and single-cell transcriptome data (ordinate) of rice seeds imbibed for 24 hours. Figure 4 and Figure 5 In [figure] and [figure], the abscissa represents different cell clusters to be analyzed, and the ordinate represents the average expression levels in different regions of the spatial transcriptome. Figure 6 UMAP visualization results of cell clustering and annotation results based on correlation with known data (blue font). Each point in the figure represents a cell.
[0046] Among them, Figure 4 and Figure 5 the letters represent abbreviations of cell types, PL: plumule; RA: radicle; CO: coleoptile; COL1: coleoptile epidermis; COL2: coleoptile cortex; LS: lateral scale; SC: scutellum; SCL1: scutellum epidermis; SCL2: scutellum cortex; SCL1-2: intermediate layer between scutellum cortex and epidermis; SCL3: scutellum vascular; EP-CR: epiblast and root sheath; UN: unknown cell type; BESCL2: cortex at both ends of scutellum. HAI represents the number of hours after seed imbibition.
[0047] Step 3: Annotation based on the results of differential gene enrichment in cell clusters Perform gene enrichment analysis on the highly expressed genes of the cell clusters for the data to be analyzed. On the one hand, verify the accuracy of the cell cluster annotation in step 2. On the other hand, it also provides information for the annotation of cell clusters 2 and 7 that could not be annotated in step 2. The genes highly expressed in cell clusters 2 and 7 mainly function in aleurone grains and are involved in the biosynthesis and metabolic processes of starch (Table 1). Therefore, cell clusters 2 and 7 may be aleurone layer cells. The annotation method of cell cluster differential gene enrichment analysis can only assist in verifying the preliminary annotation results of steps 1 and 2, as well as the annotation of individual unknown cell clusters. Since there are many gene enrichment results and different cell types may have similar biological processes, it is difficult to determine the cell type of a certain cell cluster from scratch without the preliminary annotation in the early stage.
[0048] Table 1 Enrichment of highly expressed genes in cell clusters 2 and 7
[0049] Step 4: Perform annotation based on the results of pseudotime analysis Pseudotime analysis is not only used for trajectory construction, but also verifies the previous annotation through the trajectory gene enrichment results, solving the problem that static clustering cannot distinguish the functions of subpopulations. In this example, cell clusters 2 and 7 are selected for pseudotime analysis to explore why these two clusters of cells, both being aleurone layer cells, have two different states ( Figure 7 ). The results of pseudotime analysis show that compared with cell cluster 7, cell cluster 2 is an earlier cell type. By performing enrichment analysis on the genes highly expressed in the early stage, it is found that these genes are more involved in biological processes such as stress response (Table 2), and may be aleurone layer cells that actively respond to relevant treatments during the preparation of single-cell suspensions; while cell cluster 7 is more involved in biological processes such as nutrient reservoir activity and may be aleurone layer cells that are working normally. The annotation method of pseudotime analysis can only assist in verifying the preliminary annotation results of steps 1, 2, and 3, as well as the annotation of cell subpopulations. Since there are many gene enrichment results and different cell types may have similar biological processes, it is difficult to perform cell subpopulation annotation and functional classification of subpopulations without the preliminary annotation in the early stage.
[0050] Table 2 Enrichment of highly expressed genes in different periods during the pseudotime analysis of cell clusters 2 and 7 (corresponding to Figure 8 )
[0051] Among them, Figure 7 represents the dimensionality reduction diagram of pseudotime analysis. Figure 7 The color in the upper part of Figure 7 expresses the early and late stages of pseudotime, with light colors representing early stages and dark colors representing late stages; the lower part of Figure 8 shows the distribution of cell cluster 2 (blue) and cell cluster 7 (red) on the pseudotime trajectory. Each point represents a cell.Visualize the differentially expressed genes in the pseudotime process. Red represents high expression, and blue represents low expression. Among them, the heatmap divides all differentially expressed genes into 3 clusters. Clusters 1 and 2 are highly expressed in the late stage, and cluster 3 is highly expressed in the early stage.
[0052] Step 5: Integrate information for annotation Integrate the above 4 steps. Finally, the annotation results of the single-cell transcriptome data of rice seed embryos at 6 hours of imbibition are as Figure 9 (Each point in the figure represents a cell, and the blue font is the annotation result).
[0053] It can be seen that the present invention can greatly annotate plant cell types. However, due to the limitations of existing knowledge, there may still be some cells that cannot be clearly annotated.
[0054] Example 3 Based on the same concept, this example presents the cell type annotation based on the single-nucleus transcriptome data of soybean cotyledon stage. The specific steps are as follows: Step 1: Annotate based on known cell type marker genes Take the single-nucleus transcriptome data of soybean cotyledon stage as an example for cell type annotation (data source GEO, number GSE270392). After obtaining the expression matrix, perform preliminary data analysis, including quality control, normalization, dimensionality reduction, and clustering, to obtain the clustering information of each cell, that is, cell clusters ( Figure 10 , a total of 18 cell clusters are divided). At the same time, download the cell type marker genes of soybean in the plant single-cell related databases PlantscRNAdb (such as version v4) and scPlantDB (such as version v1.1), and view the expression of these marker genes in the data to be analyzed (partially such as Figure 11 and Figure 12 ). The results show that thanks to the large number of soybean marker genes collected in the database, except for cell clusters 7 and 11, other cell clusters can be preliminarily annotated ( Figure 10 blue font in). However, if only using the single method of marker genes, the annotation is relatively rough and can only distinguish the major categories of cell types (such as cell clusters 8, 9, and 16 are all endosperm cells), and it is impossible to determine what types cell clusters 7 and 11 are.
[0055] Among them, Figure 10 represents the UMAP visualization result of cell clustering and the annotation result of marker genes (blue font). Each point in the figure represents a cell. Figure 11 and Figure 12 respectively represent the expression of marker genes of soybean endosperm cells and embryonic cells in the data to be analyzed. The abscissa is the gene, and the ordinate is the cell cluster (corresponding to Figure 10in 18 cell clusters); the size of the circle represents the proportion of cells expressing the gene in the total number of cells in the cluster, and the shade of the color represents the average expression value of the gene.
[0056] Step 2: Annotation based on known datasets Since soybean is a non-model species, there is relatively little publicly available data related to seed development. Therefore, download the dataset related to Arabidopsis seed development and use the method for identifying homologous genes described in step (1b) for association. In this example, download the conventional transcriptome data of Arabidopsis seed development (data source: GEO, accession number GSE12404), which uses laser microdissection technology to isolate different regions of the seed and perform sequencing. Compare the expression of cell clusters in the data to be analyzed with the correlation of Arabidopsis data respectively ( Figure 13 , and the R language code for reference is the same as that in the same step of Example 2). The results ( Figure 14 ) show that all cell clusters to be analyzed can be annotated according to their correlation with Arabidopsis data. Based on the annotation results of step 1, this step can not only annotate the unknown cell clusters 7 and 11, but also refine the annotation of cell clusters (such as annotating cell cluster 4 as chalazal seed coat) ( Figure 14 in red font).
[0057] Among them, Figure 13 represents the correlation heatmap of the single-nucleus transcriptome data (ordinate) of soybean cotyledon stage and the transcriptome data of Arabidopsis seed development (abscissa). The ordinate represents different cell clusters to be analyzed, and the abscissa represents the average expression level of different regions during Arabidopsis seed development. The letters in the figure represent the abbreviations of cell types, namely CZE: chalazal endosperm; MCE: micropylar endosperm; PEN: peripheral endosperm; EP: embryo body; SC: seed coat; CZSC: chalazal seed coat; CEN: cellularized endosperm. Figure 14 represents the UMAP visualization results of cell clustering and annotation. The blue font is the annotation result obtained based on step 1, and the red font is the annotation result based on the correlation with Arabidopsis data on the basis of step 1. Each point in the figure represents a cell.
[0058] Step 3: Annotation based on the enrichment results of differentially expressed genes in cell clusters Perform gene enrichment analysis on the highly expressed genes of the cell clusters for the data to be analyzed. On the one hand, verify the accuracy of the cell cluster annotation in step 2, and on the other hand, can also perform fine annotation on the cell clusters in step 2. For example, cell clusters 3 and 14 were annotated as embryonic cells in the previous stage. The enrichment results of the highly expressed genes in these two cell clusters are shown (Table 3). The genes related to cell cluster 14 are enriched in biological processes such as epidermal development and cutin biosynthesis. Therefore, cell cluster 14 can be subdivided into embryonic epidermal cells. The annotation method of cell cluster differential gene enrichment analysis can only assist in verifying the preliminary annotation results of steps 1 and 2, as well as the annotation of individual unknown cell clusters. Since there are many gene enrichment results and different cell types may have similar biological processes, it is difficult to determine the cell type of a certain cell cluster from scratch without preliminary annotation in the previous stage.
[0059] Table 3 Enrichment of highly expressed genes in cell cluster 14
[0060] Step 4: Perform annotation based on the results of pseudotime analysis Pseudotime analysis is not only used for trajectory construction, but also verifies the previous annotation through the trajectory gene enrichment results, solving the problem that static clustering cannot distinguish the functions of subpopulations. In this example, cell clusters 2, 4, 5, and 6 were selected for pseudotime analysis to explore why these cells, which are all seed coat cells, have different states.
[0061] The results of pseudotime analysis ( Figure 15 ) show that compared with cell clusters 4, 5, and 6, cell cluster 2 is an earlier cell type. And by performing enrichment analysis on the genes highly expressed in the early stage, it is found that these genes are more involved in biological processes such as cell wall growth (Table 4); the early-middle cell cluster 5 is more involved in biological processes such as carbohydrate transmembrane transport; the mid-late cell cluster 4 is involved in the osmotic signal-related pathway; the late cell cluster 6 is related to vascular transport (Table 6). The annotation method of pseudotime analysis can only assist in verifying the preliminary annotation results of steps 1, 2, and 3, as well as the annotation of cell subpopulations. Since there are many gene enrichment results and different cell types may have similar biological processes, it is difficult to perform cell subpopulation annotation and functional classification of subpopulations without preliminary annotation in the previous stage.
[0062] Table 4 Enrichment of highly expressed genes in different periods during the pseudotime analysis of cell clusters 2 and 7 (corresponding to Figure 16 )
[0063] Among them, Figure 15 represents the dimensionality reduction diagram of pseudotime analysis. Figure 15The color of the upper figure represents the early or late pseudotime, with light colors representing early stages and dark colors representing late stages; Figure 15 The lower figure shows the distribution of cell clusters 2 (blue), 4 (red), 5 (green), and 6 (yellow) on the pseudotime trajectory. Each point represents a cell. Figure 16 It represents the visualization of differentially expressed genes during the pseudotime process. Red represents high expression and blue represents low expression. Among them, the heatmap divides all differentially expressed genes into 4 clusters, and clusters 1, 2, 3, and 4 are highly expressed in the early and middle stages, late stage, early stage, and middle and late stages of pseudotime, respectively.
[0064] Step 5: Integrate information for annotation Integrate the above 4 steps, and finally the annotation results of the single-cell transcriptome data of soybean cotyledon stage are as Figure 17 (The blue font is the annotation result according to step 1, and the red font is to verify and optimize the annotation based on the analysis results of steps 2 - 4. Each point in the figure represents a cell).
[0065] It can be seen that the present invention can greatly annotate plant cell types. However, due to the limitation of existing knowledge, there may still be some cells that cannot be clearly annotated.
[0066] For the convenience of understanding, the software, databases, and functions mentioned in the present invention are explained by category as follows: Table 5 Software Tools
[0067] Table 6 Databases
[0068] Among them, the core functions are as follows: DotPlot, belonging to the tool Seurat, used for visualizing marker gene expression (dot plot); FeaturePlot, belonging to the tool Seurat, used for visualizing marker gene expression (gene expression bubble plot); dotplot / violin, belonging to the tool Scanpy, used for visualizing marker gene expression (dot plot / violin plot); AverageExpression, belonging to the tool Seurat, used for calculating the average gene expression value in cell clusters; cor, belonging to the R basic function, used for calculating the Pearson correlation coefficient (dataset correlation analysis); ComBat, belonging to the sva package, used for removing batch effects of different datasets; FindAllMarkers, belonging to the tool Seurat, used for identifying differentially expressed genes in cell clusters; enrichGO in clusterProfiler, a tool belonging to Gene Ontology (GO) for functional enrichment analysis; newCellDataSet in Monocle2, a tool for creating a pseudotime analysis object; estimateSizeFactors in Monocle2, a tool for normalizing pseudotime data; detectGenes in Monocle2, a tool for identifying differentially expressed genes during pseudotime; orderCells in Monocle2, a tool for constructing a cell developmental trajectory; plot_pseudotime_heatmap in Monocle2, a tool for visualizing the heatmap of differentially expressed genes in pseudotime; cutree in the stats package, a tool for extracting classification information of differentially expressed genes in pseudotime.
[0069] Among them, the description of key parameters is as follows: Homologous gene alignment: The E-value threshold of BLASTP is set to 1×10 -5 , and the best alignment result is retained.
[0070] Differentially expressed gene screening: Adjusted P-value < 0.01 and fold change ≥ 2.
[0071] Dataset batch removal: The ComBat function is used to eliminate technical biases in data from different sources.
[0072] Pseudotime analysis: The number of differentially expressed gene classifications is specified through the num_clusters parameter, and the gene cluster information is extracted in combination with the cutree function.
[0073] Example 4 Based on the same concept, the present invention also proposes a cell type annotation device based on plant single-cell transcriptome data, including: A marker gene expression pattern identification module, for the plant single-cell transcriptome expression matrix that has completed clustering, obtaining cell type marker genes based on a database or literature, and initially annotating the type of cell clusters by visually analyzing the expression of cell type marker genes in each cell cluster; An expression correlation module, collecting conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same species or different species as known datasets, calculating the expression correlation between the data to be analyzed and the known datasets, and assisting in annotating the cell cluster type based on the expression correlation results; The cell type inference module identifies the differentially expressed genes in each cell cluster, performs functional enrichment analysis on the differentially expressed genes, and infers the cell type based on the enrichment results; The cell subset annotation module performs pseudotime analysis on the cell clusters that are initially annotated as the same cell type or exhibit heterogeneity to construct the cell development trajectory, and annotates the cell subsets based on the cell development trajectory and the enrichment results of the differential genes in the trajectory; The comprehensive annotation generation module integrates the results of the above modules to generate a comprehensive annotation of plant cell types.
[0074] Example 5 This example also provides an electronic device. Refer to Figure 19 , which includes a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0075] Specifically, the above processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0076] Among them, the memory 404 may include a mass storage for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 404 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0077] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0078] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the cell type annotation methods based on plant single-cell transcriptome data in the above embodiments.
[0079] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0080] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0081] The input / output device 408 is used to input or output information.
[0082] Embodiment Six This embodiment also provides a readable storage medium. The readable storage medium stores a computer program, and the computer program includes program code for controlling a process to execute the process. The process includes the cell type annotation method based on plant single-cell transcriptome data according to Embodiment One.
[0083] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0084] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representations, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.
[0085] Embodiments of the present invention can be implemented by computer software, which is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow, as Figure 9 described in [reference], can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.
[0086] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0087] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for annotating cell types based on plant single-cell transcriptome data, characterized in that, Including the following steps: S1. For the plant single-cell transcriptome expression matrix that has completed clustering, obtain cell type marker genes based on databases or literature, and preliminarily annotate the types of cell clusters by visually analyzing the expression of the cell type marker genes in each cell cluster; S2. Collect conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same or different species as known datasets, calculate the expression correlation between the data to be analyzed and the known datasets, and assist in annotating the cell cluster types based on the expression correlation results; S3. Identify the differentially expressed genes in each cell cluster, perform functional enrichment analysis on the differentially expressed genes, and infer the cell types according to the enrichment results; S4. For cell clusters that are preliminarily annotated as the same cell type or have heterogeneity, perform pseudotime analysis to construct a cell developmental trajectory, and annotate cell subpopulations based on the cell developmental trajectory and the enrichment results of the differential genes in the trajectory; S5. Integrate the results of steps S1 - S4 to generate a comprehensive annotation of plant cell types.
2. The cell type annotation method based on plant single cell transcriptome data according to claim 1, characterized in that In step S1, if the target species lacks the cell type marker genes, homologous genes are obtained by the following methods: Obtain the amino acid sequences of marker genes of model species or related species from the genomic database, and obtain homologous genes of the target species through a sequence alignment algorithm, where the E-value threshold of the sequence alignment algorithm is ≤ 1×10 -5 .
3. The cell type annotation method based on plant single cell transcriptome data according to claim 1, wherein, In step S2, the expression correlation is the Pearson correlation coefficient, which is calculated by calculating the Pearson correlation coefficient of the column vectors of the gene expression matrices of the data to be analyzed and the known datasets; Perform batch effect removal on datasets from different sources before analysis.
4. A method for cell type annotation based on plant single cell transcriptome data according to claim 1, wherein In step S3, the screening conditions for the differentially expressed genes are: adjusted P value < 0.01 and fold change ≥ 2-fold.
5. The method for cell type annotation based on plant single cell transcriptome data according to claim 1, wherein In step S4, the pseudotime analysis distinguishes the developmental stages or functional states of cell subpopulations by constructing a single-cell gene expression trajectory and combining the functional enrichment results of the highly expressed genes at different stages in the trajectory.
6. The method for annotating cell types based on plant single-cell transcriptome data according to claim 3, wherein, The batch effect removal is performed using the ComBat function.
7. A method for annotating cell types based on plant single-cell transcriptome data according to any one of claims 1 to 6, characterized in that, The plant includes model plants or non-model plants. The model plants include Arabidopsis thaliana and rice, and the non-model plants include soybean.
8. A cell type annotation device based on plant single cell transcriptome data, characterized in that Including: A marker gene expression pattern identification module that, for the plant single-cell transcriptome expression matrix that has completed clustering, obtains cell type marker genes based on databases or literature, and preliminarily annotates the types of cell clusters by visually analyzing the expression of the cell type marker genes in each cell cluster; An expression correlation module that collects conventional transcriptome, single-cell transcriptome, or spatial transcriptome datasets of the same or different species as known datasets, calculates the expression correlation between the data to be analyzed and the known datasets, and assists in annotating the cell cluster types based on the expression correlation results; A cell type inference module that identifies the differentially expressed genes in each cell cluster, performs functional enrichment analysis on the differentially expressed genes, and infers the cell types according to the enrichment results; A cell subpopulation annotation module that, for cell clusters that are preliminarily annotated as the same cell type or have heterogeneity, performs pseudotime analysis to construct a cell developmental trajectory, and annotates cell subpopulations based on the cell developmental trajectory and the enrichment results of the differential genes in the trajectory; A comprehensive annotation generation module that integrates the results of the above modules to generate a comprehensive annotation of plant cell types.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the cell type annotation method based on plant single-cell transcriptome data according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, the computer program includes program code for controlling a process to execute the process, and the process includes the cell type annotation method based on plant single-cell transcriptome data according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cell subset annotation method based on single cell transcriptome sequencing
CN112700820A
Sequencing analysis method of single cell transcriptome based on ONT sequencing
CN119091964A
Method for identifying elncRNAs regulation target gene in single cell based on single cell ATAC-seq
CN119580822A
Experimental method based on depletion-type nk cells as lung cancer treatment target
CN119811489A
Method and system for constructing dynamic gene regulatory network, and computer device
WO2024077533A1
Cited By
Cell subset classification method based on plant single cell transcriptome marker gene
CN121483387A
Cell annotation method and device, electronic equipment and computer readable storage medium
CN121905305A