Method for efficiently screening plant cell type marker genes
Through single-cell transcriptome data analysis combined with weighted average calculation of FC, Diff and Avg indicators, marker genes with high specificity and abundance were screened, solving the problems of low success rate and high cost of verification in the prior art, and achieving accurate screening of marker genes for plant cell type.
Patent Information
- Application Number
- CN202510447818.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art has low verification success rate, inaccurate marker gene screening, and high experimental cost in plant cell type marker gene screening, mainly due to insufficient consideration of gene expression level and specificity.
Single-cell transcriptome data acquisition, analysis and differential gene screening methods were used, and weighted average calculation was performed in combination with three indicators of FC, Diff and Avg. Marker genes with high cell type specificity and expression abundance were screened out, and in situ hybridization was verified.
The accuracy of marker gene screening and verification success rate have been improved, reducing experimental costs.
Smart Images

Figure CN120356522A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a method for efficiently screening marker genes for plant cell types. Background Art
[0002] The identification of plant cell types is an important topic in botanical research. Especially in the fields of genomics and cell biology, accurately identifying different cell types in plants is of great significance for understanding plant physiological processes, developmental mechanisms, and disease prevention and control. In recent years, with the development of single-cell sequencing technology, researchers can identify different cell types by providing cell type marker genes through differential gene analysis. However, although these methods can provide clues for the screening of cell type marker genes, they often face the problem of low verification success rate in actual verification.
[0003] Traditional methods for cell type identification usually rely on screening significantly different genes from large-scale gene expression data, and then designing probes based on these genes for in situ hybridization verification. However, simply relying on differential gene screening does not consider the expression abundance and specificity of these genes in specific cell types, which may lead to insufficient expression of the selected genes or non-specific expression in the target cell types, thus affecting the success rate of screening cell type marker genes. The disadvantages of the existing technologies are as follows:
[0004] (1) Low verification success rate: Since the screened marker genes may not have sufficient cell type specificity and expression abundance, the success rate of in situ hybridization verification is low.
[0005] (2) Imprecise screening of marker genes: Existing differential gene screening methods may not fully consider the expression level and specificity of genes in the target cell types, resulting in insufficient expression or non-specificity of the selected marker genes, which are not sufficient to be used as marker genes for plant cell types.
[0006] (3) High experimental cost: Based on the low verification success rate, more probes need to be synthesized for in situ hybridization experiments, making the experimental cost high.
[0007] Therefore, there is an urgent need to develop a method for efficiently screening marker genes for plant cell types. Summary of the Invention
[0008] Aiming at the deficiencies of the existing technology, according to the first aspect of the technical solution of the present invention, the present invention provides a method for efficiently screening marker genes for plant cell types, and the method comprises the following steps: (1) obtaining single-cell transcriptome data; (2) analyzing single-cell transcriptome data; (3) analyzing differential genes of plant cell types; (4) screening marker genes for plant cell types; (5) verifying marker genes for plant cell types.
[0009] In some embodiments, the (1) acquisition of single-cell transcriptome data includes preparing a plant cell nucleus suspension, constructing a single-cell transcriptome library, and sequencing the single-cell library to obtain plant single-cell transcriptome data.
[0010] Optionally, the library construction and sequencing instruments are selected from the MGI DNBSEQ-TaiM 4 instrument of MGI, the DNBSEQ-T7 instrument of MGI, or the GEM-X instrument of 10x genomics.
[0011] In some embodiments, the (2) analysis of single-cell transcriptome data includes filtering, merging, normalizing, logarithmizing, dimensionality reduction, and clustering of the plant single-cell transcriptome sequencing data using bioinformatics software to obtain plant cell type clusters.
[0012] Optionally, the bioinformatics software includes DNBELab_C_Series (v2.1.3) and / or Seurat (v5.2.0).
[0013] In some embodiments, the (3) analysis of differentially expressed genes in plant cell types includes calculating genes with significant differential expression in cell types to obtain the fold change (FC), the proportion of cells in the current cell type in which the gene expression is detected (PCT1), the proportion of cells in other cell types in which the gene expression is detected (PCT2), and the adjusted p-value (Padj) obtained based on the Bonferroni correction using all features in the dataset. A list of differentially expressed genes is obtained according to Padj < 0.05.
[0014] In some embodiments, the (4) screening of marker genes for plant cell types includes (4.1) calculating the difference in the proportion of cells expressing the gene (Diff); (4.2) calculating the expression value (Avg) of the gene in the cell type; and (4.3) screening marker genes for the cell type. Optionally, Diff = PCT1 - PCT2.
[0015]
[0016] The calculation formula of the Avg is as follows:
[0017] where N c is the number of cells in population c, where c represents the cell type cluster, and X i,g is the expression level of gene g corresponding to cell i.
[0018] In some embodiments, the screening of cell type marker genes (4.3) includes obtaining three corresponding ranking values by sorting candidate genes in descending order according to FC, Diff, and Avg values. The weights of FC, Diff, and Avg are 0.4, 0.4, and 0.2 respectively. The weighted average is used to calculate the score of the gene. The scores are sorted in ascending order, and the genes with higher rankings are candidate cell type marker genes.
[0019] In some embodiments, the calculation formula for the score of the gene is as follows:
[0020]
[0021] Where Score(g) is the score of gene g;
[0022] N(FC) is the ranking of the FC value of gene g;
[0023] N(Diff) is the ranking of the Diff value of gene g;
[0024] N(Avg) is the ranking of the Avg value of gene g.
[0025] In some embodiments, the verification of the plant cell type marker gene is to perform in situ hybridization verification on the top N genes, optionally, where N is a positive integer.
[0026] Finally, the present invention provides the application of the above method in screening plant cell type marker genes or preparing a kit for screening plant cell type marker genes.
[0027] Compared with the prior art, the present invention has at least the following beneficial effects:
[0028] a) Accurate screening of marker genes: Fully considering the expression level of genes and their specificity in target cell types, and combining three indicators for weighted average calculation to accurately screen marker genes for plant cell types.
[0029] b) High verification success rate: Since the screened marker genes have sufficient cell type specificity and expression abundance, the success rate of in situ hybridization verification is high.
[0030] c) Low experimental cost: Based on the high verification success rate, accurate probes are synthesized for in situ hybridization experiments, resulting in low experimental costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart of the technical solution;
[0032] Figure 2 It is a cell type clustering diagram;
[0033] Figure 3 It is a figure for in situ hybridization and gene expression. Specific implementation manners
[0034] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0035] Example 1
[0036] A method for efficiently screening plant cell type marker genes, the flow chart of the method is as Figure 1 shown, and it mainly includes the following steps:
[0037] (1) Acquisition of single-cell transcriptome data;
[0038] (2) Analysis of single-cell transcriptome data;
[0039] (3) Analysis of differentially expressed genes in plant cell types;
[0040] (4) Screening of plant cell type marker genes;
[0041] (5) Verification of plant cell type marker genes.
[0042] Example 2
[0043] A method for efficiently screening plant cell type marker genes, the method includes the following detailed steps:
[0044] (1) Acquisition of single-cell transcriptome data: Prepare a plant cell nucleus suspension, construct a single-cell transcriptome library using the MGI DNBSEQ C-TaiM4 instrument, sequence the single-cell library using the MGI DNBSEQ-T7 instrument, and obtain plant single-cell transcriptome data;
[0045] (2) Analysis of single-cell transcriptome data: Use DNBSEQ_C_Series (v2.1.3) and Seurat (v5.2.0) software to filter, merge, normalize, logarithmize, reduce dimensions and cluster the plant single-cell transcriptome sequencing data to obtain plant cell type clusters.
[0046] (3) Analysis of differentially expressed genes in plant cell types: c) Use the FindAllMarkers module in Seurat (v5.2.0) to calculate the genes with significantly different expression in cell types, and obtain the fold change (FC), the proportion of cells in which the gene expression is detected in the current cell type (PCT1), the proportion of cells in which the gene expression is detected in other cell types (PCT2), and the adjusted p-value (Padj) obtained based on the Bonferroni correction using all features in the dataset. According to Padj less than 0.05, obtain the differential gene list.
[0047] (4) Screening of plant cell type marker genes:
[0048] (4.1) Calculate the difference in cell proportion of gene expression (Diff): Diff = PCT1 - PCT2;
[0049] (4.2) Calculate the expression value (Avg) of the gene in the cell type: For each gene g and cell population c, assume that cell i belongs to population c, where c represents the cell type cluster, and the expression level of this cell is X i,g , then the average expression level avg g (c) of this gene g in population c is calculated by the formula:
[0050]
[0051] where N c is the number of cells in population c.
[0052] X i,g is the expression level of cell i corresponding to gene g.
[0053] (4.3) Screen cell type marker genes: Candidate genes are sorted from largest to smallest according to the FC, Diff, and Avg values to obtain the corresponding three ranking values. The weights of FC, Diff, and Avg are 0.4, 0.4, and 0.2 respectively. The weighted average is used to calculate the score of the gene, and the scores are sorted from smallest to largest. The genes with higher rankings are candidate cell type marker genes. The formula for calculating the score (Score) of the gene is as follows:
[0054]
[0055] where Score(g) is the score of gene g.
[0056] N(FC) is the ranking of the FC value of gene g.
[0057] N(Diff) is the ranking of the Diff value of gene g.
[0058] N(Avg) is the ranking of the Avg value of gene g.
[0059] (5) Verification of plant cell type marker genes: In situ hybridization verification is performed on the top 5 genes.
[0060] Example 3
[0061] (1) Single-cell transcriptome data acquisition: Prepare a nuclear suspension using maize root tip tissue. Construct a single-cell transcriptome library using the MGI DNBSEQ-TaiM 4 instrument, and sequence the single-cell library using the MGI DNBSEQ-T7 instrument to obtain single-cell transcriptome data of maize root tip tissue;
[0062] (2) Single-cell transcriptome data analysis: Use DNBSEQ-C Series (v2.1.3) and Seurat
[0063] (v5.2.0) software to filter, merge, normalize, logarithmize, reduce dimensions, and cluster the plant single-cell transcriptome sequencing data to obtain cell type clusters of maize root tips (such as Figure 2 ).
[0064] (3) Analysis of differentially expressed genes in plant cell types: Use the FindAllMarkers module in Seurat (v5.2.0) to calculate genes with significantly differentially expressed cell types, obtaining the fold change (FC), the proportion of cells in the current cell type in which the gene expression is detected (PCT1), the proportion of cells in other cell types in which the gene expression is detected (PCT2), and the adjusted p-value (Padj) obtained based on the Bonferroni correction of all features in the dataset used. Obtain a list of differentially expressed genes according to Padj less than 0.05.
[0065] (4) Screening of marker genes for plant cell types:
[0066] (4.1) Calculate the difference in the proportion of cells expressing the gene: Diff = PCT1 - PCT2;
[0067] (4.2) Calculate the expression value (Avg) of the gene in the cell type: For each gene g and cell population c, assume that cell i belongs to population c, and the expression level of this cell is X i,g , then the average expression level avg g (c) of the gene g in population c is calculated by the formula:
[0068]
[0069] where N c is the number of cells in population c, where c represents the cell type cluster.
[0070] X i,g is the expression level of cell i corresponding to gene g.
[0071] (4.3) Screening for cell type marker genes: Candidate genes are sorted from largest to smallest according to the FC, Diff, and Avg values to obtain corresponding three ranking values. The weights of FC, Diff, and Avg are 0.4, 0.4, and 0.2 respectively. The weighted average is used to calculate the score of the gene. The scores are sorted from smallest to largest, and the genes with higher rankings are candidate cell type marker genes (see Table 1).
[0072] The calculation formula for the score (Score) of a gene is as follows:
[0073]
[0074] where Score(g) is the score of gene g.
[0075] N(FC) is the ranking of the FC value of gene g.
[0076] N(Diff) is the ranking of the Diff value of gene g.
[0077] N(Avg) is the ranking of the Avg value of gene g.
[0078] Table 1. Candidate cell type marker genes
[0079]
[0080] (5) Verification of plant cell type marker genes: In situ hybridization verification is performed on the screened genes (such as Figure 3 ).
[0081] In summary, the screening of the marker genes of the present invention is accurate: fully considering the expression level of the genes and their specificity in the target cell type, and combining three indicators for weighted average calculation to accurately screen the marker genes of plant cell types. The verification success rate is high: Since the screened marker genes have sufficient cell type specificity and expression abundance, the success rate of in situ hybridization verification is high. The experimental cost is low: Based on the high verification success rate, accurate probes are synthesized for in situ hybridization experiments, resulting in a low experimental cost. In the past, only one indicator of expression difference was used to provide marker genes. Among the previous 10 markers, only 2 markers were successfully verified. According to three indicators, 15 marker genes are provided, and 11 of them are successfully verified.
[0082] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for efficiently screening plant cell type marker genes, characterized in that The method comprises the following steps: (1) acquiring single-cell transcriptome data; (2) analyzing single-cell transcriptome data; (3) analyzing differential genes of plant cell types; (4) screening plant cell type marker genes; and (5) verifying plant cell type marker genes.
2. The method according to claim 1, wherein The (1) single-cell transcriptome data acquisition includes preparing a plant cell nuclear suspension, constructing a single-cell transcriptome library, and sequencing the single-cell library to obtain plant single-cell transcriptome data.
3. The method according to claim 1, characterized in that, The (2) single-cell transcriptome data analysis includes using bioinformatics software to filter, merge, standardize, logarithmize, reduce dimension and cluster the plant single-cell transcriptome sequencing data to obtain plant cell type clustering.
4. The method according to claim 1, wherein The (3) plant cell type differential gene analysis includes calculating genes with significant differential expression in cell types, obtaining the differential expression fold (FC), the proportion of cells in which the gene expression is detected in the current cell type (PCT1), the proportion of cells in which the gene expression is detected in other cell types (PCT2), and the corrected p-value (Padj) based on the Bonferroni correction using all the features in the data set, and obtaining a differential gene list based on Padj being less than 0.
05.
5. The method according to claim 1, characterized in that The (4) screening of plant cell type marker genes includes (4.1) calculating the cell ratio difference (Diff) of gene expression; and (4.2) calculating the expression value (Avg) of the gene in the cell type; (4.3) screening cell type marker genes; optionally In other words, Diff=PCT1-PCT2; the calculation formula of Avg is as follows: Among them, c represents the cell type cluster, and N c is the number of cells in population c, and X i,g is the expression level of gene g corresponding to cell i.
6. The method according to claim 1, characterized in that The (4.3) screening of cell type marker genes includes sorting candidate genes according to FC, Diff and Avg values from large to small to obtain corresponding three ranking values, the weights of FC, Diff and Avg are 0.4, 0.4 and 0.2 respectively, and the weighted average is used to calculate the score of the gene. The score is sorted from small to large, and the top ranked genes are candidate cell type marker genes.
7. The method according to claim 6, wherein The score calculation formula of the gene is as follows: Among them, Score(g) is the score of gene g; N(FC) is the ranking of the FC value of gene g; N(Diff) is the ranking of the Diff value of gene g; N(Avg) is the rank of the Avg value of gene g.
8. The method according to claim 1, wherein The plant cell type marker gene verification is performed by in situ hybridization verification of the top N genes.
9. The method according to claim 8, characterized in that, Said N is a positive integer.
10. Use of the method according to any one of claims 1 to 9 for screening plant cell type marker genes or for preparing a kit for screening plant cell type marker genes.