Method for determining the relationship between a plant growth gene and a developmental stage and cell type
By employing partial least squares decomposition and cluster analysis methods, combined with hypergeometric distribution tests, the problem of insufficient monitoring of the dynamic characteristics of key genes in existing technologies has been solved. This enables in-depth analysis of the relationship between genes and cell types during plant growth and development, improving the success rate and accuracy of experiments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI CHENSHAN BOTANICAL GARDEN
- Filing Date
- 2022-09-23
- Publication Date
- 2026-05-15
AI Technical Summary
The lack of effective methods in the current technology to monitor and systematically analyze the dynamic characteristics of key genes in plant growth and development has led to a large amount of useful information being ignored or buried, making it impossible to deeply understand the relationship between key genes and plant developmental stages and cell types.
The partial least squares method was used to decompose the single-cell dataset. Combined with cluster analysis and hypergeometric distribution test, the correlation between the expression of key genes and cell type and developmental state was represented by a sigma graph. Cluster analysis was performed using Seurat software.
It significantly improves the purposefulness and success rate of gene function experiments, enabling preliminary inferences about the role of key target genes in plant growth processes and providing a deeper understanding of the relationship between genes and developmental stages and cell types.
Smart Images

Figure CN115579067B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant developmental cell data research, and in particular to a method for determining the relationship between plant growth genes and developmental stages and cell types. Background Technology
[0002] Single-cell transcriptome sequencing technology can comprehensively depict the expression profile of individual cells in plant samples at the time of sampling. Current plant single-cell studies have identified the main expression characteristics of phloem, column cells, and QCs in Arabidopsis root tissue development1, and further revealed the complex process of Arabidopsis root cell differentiation and key transcription factors such as LAR3, ATHB-20, and GATA4 through pseudo-time series analysis. More plant single-cell related studies have emerged both domestically and internationally. These studies currently mainly focus on cell differentiation and tissue development in model plants such as Arabidopsis, tomato, and rice2-6, providing the possibility to study the details of cell differentiation and key transcription factors during plant growth and development.
[0003] However, current analyses of plant single-cell data largely focus on identifying marker genes and key transcription factors through differential analysis, followed by a series of wet experiments to verify these findings, such as gene knockout. Little is known about the relationships between key genes and plant development, or between cell types. Furthermore, the presence of both characteristic information and interference in plant developmental single-cell data means that current data mining remains superficial and lacks depth. There is currently no method for specifically monitoring and systematically analyzing the dynamic characteristics of certain key genes during plant growth and development, resulting in a significant amount of useful information being buried or ignored by other interfering signals. Fully mining this information will help understand the dynamic relationship between the expression of a key gene and the plant developmental stage, its role in different cell types, the regulatory relationships and activation pathways among key genes during plant growth and development, and the plant traits and functions regulated by key genes. This will provide crucial information for optimizing plant breeding, increasing yield, and improving stress resistance.
[0004] Therefore, those skilled in the art are dedicated to developing a method for determining the relationship between plant growth genes and developmental stages and cell types, in order to predict the role of target key genes in the plant growth process. Summary of the Invention
[0005] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is how to extract the relationship between developmental state and cell type related to the dynamic changes of upregulation or downregulation of specific key genes from single-cell data of different stages of plant growth and development (e.g., from the phloem tissue growth period, plant stem cell growth period, root cap growth period, xylem growth period, etc.).
[0006] To achieve the above objectives, the present invention provides a method for determining the relationship between plant growth genes and developmental stages and cell types, comprising the following steps:
[0007] Step 1: Divide the plant development single-cell dataset into two datasets, X, based on the expression and non-expression of key genes. 表达 Second dataset X 未表达 Extract the first dataset X using the following formula. 表达 and the second dataset X 未表达 The difference in characteristic components between them;
[0008]
[0009] Among them, A 表达 It is X 表达 The partial least squares basis matrix, B 未表达 It is X 未表达 The partial least squares basis matrix;
[0010] V 表达 It is X 表达 Compared to X 未表达 The increment, V 未表达 It is X 未表达 Compared to X 表达 The increment;
[0011] Step 2: Use cluster analysis methods, based on V 表达 and V 未表达 Cluster the cells;
[0012] Step 3: Calculate V 表达 The correlation p1 value with the labeled cell type; calculate V 未表达 The p2 value correlated with the labeled cell type;
[0013] Step 4: Calculate V 表达 The p3 value correlated with the developmental state of the labeled plant; calculate V 未表达 The p4 value correlated with the developmental state of the labeled plant;
[0014] Step 5: Connect the significant correlations between the p1, p2 and p3, p4 values obtained in steps 3 and 4 with lines.
[0015] Furthermore, in step 1, the partial least squares method is used to extract the first dataset X. 表达 and the second dataset X 未表达 The difference in characteristic components between them.
[0016] Furthermore, the clustering analysis method used in step 2 is the Seurat method.
[0017] Furthermore, the formula for calculating the p1 value in step 3 is as follows:
[0018]
[0019] In this case, n cells are selected from N cells in a certain cell subtype as the denominator, and the numerator is V. 表达 There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
[0020] Furthermore, the formula for calculating the p2 value in step 3 is as follows:
[0021]
[0022] In this case, n cells are selected from N cells in a certain cell subtype as the denominator, and the numerator is V. 未表达 There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
[0023] Furthermore, the formula for calculating the p3 value in step 4 is as follows:
[0024]
[0025] The denominator is n cells selected from N cells in a cell population at a certain developmental stage of a plant, and the numerator is V. 表达 Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
[0026] Furthermore, the formula for calculating the p4 value in step 4 is as follows:
[0027]
[0028] The denominator is n cells selected from N cells in a cell population at a certain developmental stage of a plant, and the numerator is V. 未表达 Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
[0029] Furthermore, in step 3, the values of p1 and p2 are calculated using the hypergeometric distribution test.
[0030] Furthermore, in step 4, the values of p3 and p4 are calculated using the hypergeometric distribution test.
[0031] Furthermore, in step 5, a xuan diagram is used to connect the lines.
[0032] Compared with existing technologies, the advantages of this invention are as follows: Traditional methods for studying the function of a key gene often involve experimental knockout, which is both time-consuming and has a high failure rate. This invention proposes a research method based on plant single-cell transcriptome data to study the correlation between key genes and developmental state and cell type. This method, based on statistical modeling analysis, can preliminarily predict the role of the target key gene in plant growth. This will significantly improve the purposefulness and success rate of gene function experiments.
[0033] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description
[0034] Figure 1 This is a flowchart of a preferred embodiment of the present invention;
[0035] Figure 2 This is a clustering result classified based on whether a single gene is expressed or not, according to a preferred embodiment of the present invention.
[0036] Figure 3 This is a preferred embodiment of the invention, showing the relationship between the classification results based on gene expression and cell type and state. Detailed Implementation
[0037] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0038] This method is applicable to the study of developmental status and cell type relationships related to the dynamic changes of upregulation or downregulation of specific key genes in single-cell data from different stages of plant growth and development (e.g., from the phloem growth stage, plant stem cell growth stage, root and crown growth stage, xylem growth stage, etc.).
[0039] Using single-cell datasets comparing the expression and non-expression of key genes in plant development, a supervised machine learning algorithm was used to extract only variables related to the dynamic process of key gene expression, eliminating features generated by other irrelevant biological processes.
[0040] Partial least squares (PLS) was used to decompose the covariance matrix between single-cell datasets showing key gene expression and non-expression. The results showed the state difference of the same cell when the key gene was expressed versus when it was not expressed. Cluster analysis was then performed based on this state difference.
[0041] The hypergeometric distribution test was used to calculate the p-value of the correlation between cell clustering results and cell type. A sigma plot was used to represent significant correlations with p-values as lines.
[0042] The hypergeometric distribution test was used to calculate the p-value of the correlation between cell clustering results and cell sample growth state attributes. A syntactic graph was used to represent significant correlations with p-values as lines.
[0043] like Figure 1 The diagram shown is a flowchart of the method of the present invention, as detailed below:
[0044] Step 1: Divide the plant developmental single-cell dataset into X groups based on the expression and non-expression of key genes. 表达 and X 未表达 The partial least squares method is used to extract the differential feature components between the two datasets.
[0045]
[0046] Among them, A 表达 It is X 表达 The partial least squares basis matrix, B 未表达 It is X 未表达 The partial least squares basis matrix;
[0047] V 表达 It is X 表达 Compared to X 未表达 The increment, V 未表达 It is X 未表达 Compared to X 表达 The increment;
[0048] Step 2: Use cluster analysis methods such as Seurat, based on V 表达 and V 未表达 Cluster the cells.
[0049] Step 3: Calculate V using the hypergeometric distribution test. 表达 The correlation p1 value with the labeled cell type; calculate V 未表达 The p2 value, which correlates with the labeled cell type, is calculated using the following formula:
[0050]
[0051] Where n cells are selected from N cells in a certain cell subtype as the denominator, and the numerator is V. 表达There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
[0052]
[0053] Where n cells are selected from N cells in a certain cell subtype as the denominator, and the numerators are V... 未表达 There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
[0054] Step 4: Calculate V using the hypergeometric distribution test. 表达 The p3 value correlated with the developmental state of the labeled plant; V was calculated using the hypergeometric distribution test. 未表达 The p4 value, which correlates with the developmental state of the labeled plant, is calculated using the following formula:
[0055]
[0056] The denominator is n cells selected from N cells in a cell population at a certain developmental stage of a plant, and the numerator is V. 表达 Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
[0057]
[0058] The denominator is n cells selected from N cells in a cell population at a certain developmental stage of a plant, and the numerator is V. 未表达 Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
[0059] Step 5: Use a xuan diagram to represent the significant correlation between the p1 and p2 values obtained in steps 3 and 4 by connecting lines.
[0060] In one embodiment of the present invention, using single-cell data 7 from tomato aboveground root testing, cells are divided into two groups based on gene expression and non-expression, and matrix X is used to represent these groups. 表达 and X 未表达 This means that, substituting into formula (1), we obtain the differential characteristic cell populations of all gene dynamic changes during the process of this gene changing from non-expression to expression, V. 表达 and V 未表达 Cluster analysis was performed using the Seurat8 software package, and the clusters were obtained as follows: Figure 2As shown in Table 1, the expression and non-expression of this gene are shown in three cases, developmental status: 0, 1, 3, 5, and the number of overlapping cells between five cell types: distal phloem parenchyma, phloem parenchyma, root cap, stem cells, and transitional cells. The p-value is calculated for all cell numbers in Table 1 using formula (2). A threshold is set for the p-value (generally 0.05 or 0.01, which can be modified according to the specific data and common p-value threshold selection rules). Cases with p-values less than the threshold in Table 1 are linked with gray lines, while cases with p-values greater than the threshold are not linked. Figure 3 As shown.
[0061] Table 1 includes the following information: gene expression and non-expression, developmental status: 0, 1, 3, 5, and five cell types: distal phloem parenchyma, phloem parenchyma, root cap, stem cells, and the values of the overlap between transitional cells, p1 and p2.
[0062]
[0063] Table 2 includes the following information: Gene expression and non-expression were calculated using the hypergeometric distribution test, developmental state: 0, 1, 3, 5, and the values of p3 and p4 for the overlap between five cell types: distal phloem parenchyma, phloem parenchyma, root cap, stem cells, and transitional cells.
[0064]
[0065] Traditional methods for studying the function of key genes often involve experimental knockout. This is both time-consuming and has a high failure rate. This invention proposes a research method based on plant single-cell transcriptome data to study the correlation between key genes and developmental state and cell type. This method is entirely based on statistical modeling analysis and can preliminarily predict the role of target key genes in plant growth. This will significantly improve the purposefulness and success rate of gene function experiments.
[0066] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for determining the relationship between plant growth genes and developmental stages and cell types, characterized in that, Includes the following steps: Step 1: Divide the plant development single-cell dataset into two datasets, X, based on the expression and non-expression of key genes. 表达 Second dataset X 未表达 Extract the first dataset X using the following formula. 表达 and the second dataset X 未表达 The difference in characteristic components between them; (1) Among them, A 表达 It is X 表达 The partial least squares basis matrix, B 未表达 It is X 未表达 The partial least squares basis matrix; V 表达 It is X 表达 Compared to X 未表达 The increment, V 未表达 It is X 未表达 Compared to X 表达 The increment; Step 2: Use cluster analysis methods based on... and Cluster the cells; Step 3: Calculate using the hypergeometric distribution test. The correlation p1 value with the labeled cell type was calculated using the hypergeometric distribution test. The p2 value correlated with the labeled cell type; Step 4: Calculate using the hypergeometric distribution test. The p3 value correlated with the developmental state of the labeled plant was calculated using the hypergeometric distribution test. The p4 value correlated with the developmental state of the labeled plant; Step 5: Connect the significant correlations between the p1, p2 and p3, p4 values obtained in steps 3 and 4 with lines.
2. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, In step 1, the partial least squares method is used to extract the first dataset X. 表达 and the second dataset X 未表达 The difference in characteristic components between them.
3. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, The clustering analysis method used in step 2 is the Seurat method.
4. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, The formula for calculating the value of p1 in step 3 is as follows: In this case, n cells are selected from N cells in a certain cell subtype as the denominator, and the numerator is... There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
5. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, The formula for calculating the p2 value in step 3 is as follows: In this case, n cells are selected from N cells in a certain cell subtype as the denominator, and the numerator is... There are a total of M cells, of which i fall into this cell subtype and ni do not. m is the number of cells that fall into this cell subtype, and i is a value from m to M.
6. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, The formula for calculating the p3 value in step 4 is as follows: Where n cells are selected from N cells in a cell population at a certain developmental stage of a plant as the denominator, and the numerator is... Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
7. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, The formula for calculating the p4 value in step 4 is as follows: Where n cells are selected from N cells in a cell population at a certain developmental stage of a plant as the denominator, and the numerator is... Of the total number of M cells, i fall into this developmental state and ni do not. m is the number of cells falling into this cell subtype, and i is a value from m to M.
8. The method for determining the relationship between plant growth genes and developmental stages and cell types as described in claim 1, characterized in that, Step 5 uses the Xuan Diagram connection.