Method for improving stress resistance of secondary metabolites of tea tree

By integrating multi-omics data and using a dynamic Bayesian network model, the problem of identifying the causal regulatory relationships in the process of tea tree stress resistance and secondary metabolite synthesis was solved, enabling precise regulation and efficient breeding of tea trees and providing a direct breeding application solution.

CN121687184BActive Publication Date: 2026-07-03四川省农业科学院茶叶研究所
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies rely on single-gene data when analyzing the stress resistance and secondary metabolite synthesis process of tea plants. This leads to fragmented and static regulatory network models, which cannot reveal the dynamic causal regulatory relationships between genes, metabolites and phenotypes, thus limiting the targeting and efficiency of molecular breeding.

Method used

By collecting multi-omics time-series data of tea trees under abiotic stress, a spatiotemporal dynamic data matrix integrating genome, transcriptome, metabolome and phenome was constructed. Causal inference was performed using a dynamic Bayesian network model, a multidimensional quantitative evaluation algorithm was designed to screen key regulatory genes, and in vivo computer simulation was conducted to generate gene editing or molecular marker-assisted selection breeding programs.

Benefits of technology

It enables precise control of tea tree stress resistance and secondary metabolite content, shortens the breeding cycle, reduces trial and error costs, provides a direct breeding application solution, solves the problems of information bias and subjectivity in existing technologies, and has industrial application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121687184B_ABST
    Figure CN121687184B_ABST
Patent Text Reader

Abstract

This invention relates to the field of plant biotechnology and discloses a method for enhancing the stress resistance of tea plant secondary metabolites. The method includes: collecting multi-omics time-series data of tea plants under drought stress to construct a spatiotemporally aligned multidimensional data matrix; using dynamic Bayesian networks to model the dynamic causal regulatory relationships between genes, metabolites, and phenotypes; designing a multidimensional evaluation algorithm integrating network topology and biological effects to screen key regulatory genes; quantitatively predicting their effects on target metabolites such as catechins through in vivo computer perturbation simulations; and finally generating an optimal gene editing or molecular marker-assisted breeding scheme. This invention achieves a breakthrough from association analysis to causal inference, improving the accuracy and efficiency of synergistic improvement of tea plant stress resistance and secondary metabolites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of plant biotechnology, specifically relating to a method for enhancing the stress resistance of secondary metabolites in tea trees. Background Technology

[0002] As an important economic crop, the quality and resistance of tea plants are highly dependent on the accumulation level of secondary metabolites.

[0003] These compounds not only determine the flavor and health value of tea, but also play a crucial defensive role in plants' response to abiotic and biotic stresses such as drought, low temperatures, and pests and diseases. In recent years, with the intensification of climate change and the increasing complexity of planting environments, enhancing the ability of tea trees to synthesize secondary metabolites under adverse conditions has become one of the core breeding objectives.

[0004] However, this trait is regulated by multiple genes and interacts deeply with environmental factors, exhibiting highly complex quantitative genetic characteristics that are difficult to improve through traditional phenotypic selection.

[0005] Analyzing the molecular mechanisms of secondary metabolism and stress resistance in tea plants using omics technologies has become a research hotspot. Genomics can reveal the basis of genetic variation, transcriptomics reflects the dynamics of gene expression under stress response, and metabolomics directly characterizes the abundance changes of functional molecules. While a single omics approach can provide local information, the lack of cross-level data integration makes it impossible to accurately depict the causal regulatory pathways between genes, transcription, and metabolism, leading to ambiguity in the identification of key regulatory nodes and limiting the precise implementation of molecular design breeding.

[0006] In current technologies, most tea tree breeding still relies on field phenotypic observation and experience-based judgment, which is time-consuming and inefficient. Even with the introduction of molecular marker-assisted selection, it often focuses on single genes or simple traits, making it difficult to achieve the combined goals of high metabolite content and strong stress resistance. At the same time, the lack of ability to systematically model multi-omics data makes parental selection lack predictive basis, and the phenotypic expression of hybrid offspring is uncontrollable.

[0007] Especially when faced with synergistic optimization requirements such as high EGCG content and drought resistance, existing methods are unable to efficiently mine the optimal gene combination from massive germplasm resources, which seriously restricts the progress of precision tea breeding.

[0008] Therefore, there is an urgent need for an intelligent method that integrates multi-omics information, can construct regulatory networks, and supports breeding decisions, so as to realize the transformation of the breeding paradigm from experience-driven to data-driven. Summary of the Invention

[0009] The technical problem to be solved by this invention is that existing technologies, when analyzing the complex biological process of tea tree stress resistance and secondary metabolite synthesis, rely heavily on single-atomistic data. This results in fragmented, static, and mostly correlational descriptions of the constructed regulatory network models, which cannot reveal the deep-seated and dynamic causal regulatory relationships between different molecular levels. This severely restricts the formulation of targeted and efficient molecular breeding strategies.

[0010] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for improving the stress resistance of tea tree secondary metabolites.

[0011] This invention provides a method for enhancing the stress resistance of secondary metabolites in tea trees. By collecting multi-omics time-series data of tea trees under abiotic stress conditions, a spatiotemporal dynamic data matrix integrating genomic, transcriptomic, metabolomic, and phenotypic information is constructed.

[0012] A dynamic Bayesian network model was used to make causal inferences on this high-dimensional time series data, and a multi-level network that can reflect the dynamic causal regulatory relationship between genes, metabolites and phenotypes was constructed.

[0013] Design a multidimensional quantitative evaluation algorithm that integrates network topology attributes and biological effect weights to identify and rank key regulatory hub genes in the network;

[0014] Based on the constructed causal network model, in vivo computer simulations were performed on the upregulation or downregulation of the selected key genes to quantitatively predict their impact on the synthesis flux of downstream target secondary metabolites.

[0015] Ultimately, based on the simulation prediction results, the optimal gene editing or molecular marker-assisted selection breeding scheme is generated to achieve precise regulation of the synergistic enhancement of tea tree stress resistance and the content of specific secondary metabolites.

[0016] This invention provides a method for enhancing the stress resistance of tea plant secondary metabolites, comprising:

[0017] A multi-omics time-series dataset of tea plants under specified abiotic stress conditions was obtained. The multi-omics time-series dataset includes genomic sequence data, transcriptomic data, metabolomic data, and phenotypic data collected at different time points.

[0018] Data standardization and spatiotemporal alignment are performed on the multi-omics time series datasets to construct a unified multi-dimensional time series data matrix;

[0019] Based on the multidimensional time-series data matrix, a dynamic Bayesian network algorithm is used to construct a tea tree stress resistance regulation network model that reflects the causal relationship between genes, metabolites and target traits.

[0020] Based on the preset multidimensional quantitative evaluation index system, the key regulatory gene set was screened and sorted in the tea tree stress resistance regulation network model;

[0021] Using the tea tree stress resistance regulatory network model, in vivo computer perturbation simulations were performed on each gene in the key regulatory gene set to quantitatively predict the effect of changes in their expression levels on the content of target secondary metabolites.

[0022] Based on the quantitative prediction results of the in vivo computer perturbation simulation, the target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites were identified.

[0023] Furthermore, the acquisition of the tea tree multi-omics time-series dataset under specified abiotic stress conditions specifically includes:

[0024] Tea tree populations with consistent genetic backgrounds were selected and divided into a control group and a treatment group. The treatment group was subjected to a 15% polyethylene glycol 6000 solution to simulate drought stress.

[0025] Tender leaf samples from two groups of tea trees were collected simultaneously at six time points: 0 hours, 6 hours, 12 hours, 24 hours, 48 ​​hours, and 72 hours after the stress treatment.

[0026] Whole-genome resequencing technology was used to obtain genomic sequence data and single nucleotide polymorphism (SNP) site information of tea plant populations;

[0027] Gene expression profiles of samples at all time points were obtained using whole transcriptome sequencing technology;

[0028] Ultra-high performance liquid chromatography-tandem mass spectrometry was used to detect non-targeted metabolites in samples at all time points and obtain metabolite abundance data.

[0029] The relative water content, proline content, malondialdehyde content, and catechin content of leaves were measured and recorded simultaneously at all time points as phenotypic data.

[0030] Furthermore, the step of performing data standardization and spatiotemporal alignment on the multi-omics time-series dataset to construct a unified multi-dimensional time-series data matrix specifically includes:

[0031] Transcriptome gene expression data were standardized using the number of reads per million transcripts;

[0032] Metabolomics data were standardized using a combination of internal standard normalization and total ion current normalization.

[0033] All time-series data, including transcriptomic, metabolomic, and phenotypic data, are processed using a cubic spline interpolation algorithm to fit them from discrete sampling time points into continuous time function curves. Then, at a uniform, preset time resolution, i.e., every hour, resampling is performed to generate data sequences with perfectly aligned time points.

[0034] Integrate all aligned omics data into OK A multidimensional time-series data matrix of columns, in which This represents the total number of observed variables, including all genes, metabolites, and phenotypic indicators. This represents the total number of time points after unified resampling.

[0035] Furthermore, the construction of a tea tree stress resistance regulation network model reflecting the causal relationship between genes, metabolites, and target traits based on the multidimensional time-series data matrix and using a dynamic Bayesian network algorithm specifically includes:

[0036] Using the multidimensional time-series data matrix as input, a two-time-slice dynamic Bayesian network structure is initialized. This structure includes the current time slice. With the next time film +1;

[0037] A hill-climbing algorithm based on Bayesian information criterion scoring is used as a structure learning strategy to search for the optimal network topology in the variable space. The network topology is a directed acyclic graph, where nodes represent genes or metabolites and directed edges represent direct causal effects between variables at time lag units.

[0038] The maximum likelihood estimation algorithm is used to learn the conditional probability distribution parameters of each node in the network and quantify the causal strength of directed edges.

[0039] Through an iterative learning process, a fully defined tea tree stress resistance regulatory network model is finally output, which accurately describes the effects of any gene or metabolite over time. How the state probabilistically determines the state of other genes or metabolites over time The +1 state.

[0040] Furthermore, the step of screening and ranking the key regulatory gene set in the tea tree stress resistance regulation network model according to the preset multidimensional quantitative evaluation index system specifically includes:

[0041] The key regulatory index is defined, and its calculation formula is as follows:

[0042] ;

[0043] , , The preset weighting coefficients satisfy... In this embodiment, we take =0.4, =0.4, =0.2.

[0044] To standardize network centrality, it is calculated as follows: first calculate the node centrality... out of degree That is, from Count the number of directed edges sent; then calculate the number of nodes. betweenness centrality That is, all shortest paths in the network pass through The proportion; then and A weighted average is calculated with each weight being 0.5 to obtain the original centrality score. Finally, the score is standardized using the Z-score method to make its mean 0 and standard deviation 1.

[0045] Furthermore, the step of using the tea tree stress resistance regulatory network model to perform in vivo computer perturbation simulations on each gene in the key regulatory gene set, and quantitatively predicting the impact of changes in their expression levels on the content of target secondary metabolites, specifically includes:

[0046] Genes in the key regulatory gene set were selected sequentially as perturbation targets;

[0047] In the tea tree stress resistance regulation network model, the node state in the initial time slice is forcibly set to a high expression level, which is twice the original average expression level, or a low expression level, which is 0.5 times the original average expression level.

[0048] The forward sampling logical stear sampling algorithm is adopted. Based on the conditional probability distribution learned in the network, the state probability distribution of all other nodes in the network is recursively calculated from the initial time slice to the last time slice.

[0049] Extract the target secondary metabolites, namely catechins and epigallocatechin gallate, which characterize stress resistance, and calculate the mean predicted concentrations at the last time slice. Compare the mean predicted concentrations with the mean baseline concentrations without perturbation to calculate the net percentage effect of the change in expression level on the content of the target secondary metabolites.

[0050] This perturbation simulation process is repeated for all genes in the key regulatory gene set.

[0051] Furthermore, the determination of target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites based on the quantitative prediction results of the in vivo computer perturbation simulation specifically includes:

[0052] The in vivo computer perturbation simulation results of all key regulatory genes were compiled to form a comprehensive effect matrix including gene identifiers, perturbation types, percentage effects on catechin content, and percentage effects on epigallocatechin gallate content. A regulatory benefit threshold was set, i.e., the increase in the total content of target secondary metabolites must be greater than 20%. Genes in the comprehensive effect matrix that meet the regulatory benefit threshold and show positive regulatory effects on both target metabolites were screened. The screened genes were used as the final breeding target genes, and the corresponding regulatory strategies were determined according to the perturbation type that produced the largest positive effect in the simulation, i.e., high expression or low expression, i.e., genetic improvement through gene overexpression technology or gene silencing technology.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] 1. This invention constructs a dynamic and comprehensive molecular change map by collecting multi-omics time-series data and performing strict spatiotemporal alignment. It overcomes the information bias and limitations caused by the analysis of a single omics or static time point in the existing technology, and can capture the complete dynamic chain of gene expression, metabolite accumulation and phenotypic changes in tea trees during the stress response process.

[0055] 2. This invention innovatively introduces dynamic Bayesian networks for causal inference, constructing a directional causal regulatory network from high-dimensional complex data. Compared with the correlation analysis methods commonly used in existing technologies, this invention can reveal the direct regulatory relationships and information flow between molecules, providing a theoretical basis for accurately locating key upstream regulatory factors and achieving a cognitive breakthrough from "correlation" to "causation".

[0056] 3. This invention designs a multidimensional quantitative evaluation system that integrates network topology and biological significance to objectively and accurately screen key regulatory genes. This avoids the subjectivity and one-sidedness of traditional methods that rely on experience or single indicators for screening, and improves the accuracy of target gene identification.

[0057] 4. This invention utilizes the constructed causal network model to perform in vivo computer perturbation simulation, achieving efficient prediction of gene function. It can quantitatively assess the potential impact of different gene editing strategies on target traits before time-consuming and laborious wet experiments are required, shortening the breeding research and development cycle, reducing trial and error costs, and providing a powerful prediction and decision support tool for molecular design breeding.

[0058] 5. The final output of this invention is a set of target genes with clear regulatory strategies that have been quantitatively predicted and verified. This provides a direct and operable technical solution for breeding practice, solves the problem that existing research results are difficult to directly transform into breeding applications, and has industrial application value. Attached Figure Description

[0059] Figure 1 This is a schematic diagram of the overall technical solution architecture of a method for improving the stress resistance of secondary metabolites in tea trees proposed in this invention;

[0060] Figure 2 This is a schematic diagram illustrating the core principle framework of constructing a causal network for regulating the stress resistance of tea trees based on a dynamic Bayesian network in this invention.

[0061] Figure 3 This is a flowchart illustrating the logical process framework for multi-omics time-series data acquisition, standardization, and spatiotemporal alignment in this invention.

[0062] Figure 4 This is a flowchart illustrating the logical framework of the multidimensional quantitative evaluation and ranking algorithm for key regulatory genes in this invention.

[0063] Figure 5 This is a logical flowchart of the in vivo computer perturbation simulation and effect prediction of key genes based on the causal network model in this invention.

[0064] Figure 6 This is a schematic diagram of the multi-level interaction relationships and data flow generated by the terminal breeding target gene screening and regulation strategy in this invention. Detailed Implementation

[0065] Please refer to Figures 1 to 6 This invention provides a method for enhancing the stress resistance of tea plant secondary metabolites. Its core lies in the integration of multi-omics time-series data, the construction of dynamic causal networks, the quantitative evaluation of key regulatory genes, and in vivo computer perturbation simulation, ultimately generating target genes and their regulatory strategies that can be directly used for molecular breeding. The following will describe this embodiment in detail according to the method steps S1 to S6 explicitly listed in the invention description.

[0066] The method includes the following steps:

[0067] S1, Obtain a multi-omics time-series dataset of tea trees under specified abiotic stress conditions;

[0068] S2, perform data standardization and spatiotemporal alignment processing on the multi-omics time series dataset to construct a unified multi-dimensional time series data matrix;

[0069] S3. Based on the multidimensional time-series data matrix, a tea tree stress resistance regulation network model reflecting the causal relationship between genes, metabolites and target traits is constructed using a dynamic Bayesian network algorithm.

[0070] S4. Based on the preset multidimensional quantitative evaluation index system, the key regulatory gene set is screened and sorted in the tea tree stress resistance regulation network model.

[0071] S5. Using the tea tree stress resistance regulation network model, in vivo computer perturbation simulation is performed on each gene in the key regulatory gene set to quantitatively predict the effect of its expression level changes on the content of target secondary metabolites.

[0072] S6. Based on the quantitative prediction results of the in vivo computer perturbation simulation, the target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites are identified.

[0073] In step S1, a multi-omics time-series dataset of tea plants under specified abiotic stress conditions is obtained. The specific implementation process of this step is as follows:

[0074] A population of tea tree clones with highly consistent genetic backgrounds was selected as the experimental material and randomly divided into a control group and a treatment group.

[0075] The treatment group was given a 15% polyethylene glycol 6000 solution to simulate drought stress, while the control group maintained a normal water supply.

[0076] At six time points—0, 6, 12, 24, 48, and 72 hours after the stress treatment was initiated—the second to fourth functional tender leaves from the top of the two tea plantations were collected simultaneously as samples.

[0077] All samples were immediately flash-frozen in liquid nitrogen and stored in an ultra-low temperature freezer at -80 degrees Celsius until use.

[0078] Subsequently, whole-genome resequencing technology was used to sequence the entire tea plant population, obtaining complete genome sequence information and single nucleotide polymorphism site maps for each individual.

[0079] RNA was extracted and library constructed from all samples at the above 6 time points using whole transcriptome sequencing technology to obtain whole genome expression profiles for each sample at the corresponding time points, with a data output of more than 30 gigabit base pairs per sample.

[0080] Untargeted metabolomics analysis was performed on all samples using ultra-high performance liquid chromatography-tandem mass spectrometry to detect and quantify all detectable small molecule metabolites in the samples, and to obtain a metabolite abundance data matrix.

[0081] At the same time, physiological and biochemical indicators were measured for samples at each time point, including relative leaf water content, free proline content, malondialdehyde content, and the content of the main components of catechins, including catechins, epicatechin, epigallocatechin, epigallocatechin gallate, etc. These indicators together constitute phenotypic data.

[0082] This completes the collection of a multi-omics time-series dataset covering four dimensions: genome, transcriptome, metabolome, and phenome.

[0083] In step S2, the multi-omics time series dataset is subjected to data standardization and spatiotemporal alignment to construct a unified multi-dimensional time series data matrix.

[0084] The specific implementation process for this step is as follows:

[0085] First, the transcriptome data were standardized using the reads per million transcripts method to eliminate systematic bias caused by differences in sequencing depth between different samples.

[0086] Secondly, the metabolomics data were double-standardized. First, internal standard compounds were used for correction to eliminate instrument response fluctuations, and then the total ion current normalization method was used to eliminate differences in sample loading.

[0087] For phenotypic data, since they are absolute measurements, only unit standardization and outlier removal are required. After completing the internal standardization within each omics, the spatiotemporal alignment phase begins.

[0088] Since the original sampling was conducted at only 6 discrete time points, while subsequent dynamic modeling requires continuous and high-resolution time series, cubic spline interpolation was applied to fit all time-series variables, including each gene in the transcriptome, each metabolite in the metabolome, and each index in the phenotype.

[0089] The algorithm uses the observations at the original six time points as control points to generate a smooth, continuous function curve that accurately describes the dynamic trend of the variable throughout the entire 72-hour stress period.

[0090] Subsequently, resampling is performed at a uniform preset time resolution. In this embodiment, the time resolution is set to data points per hour, that is, 73 time points are generated from 0 hours to 72 hours.

[0091] By performing this resampling operation on all variables, we ensured that the transcriptomic, metabolomic, and phenotypic data were fully aligned over time.

[0092] Finally, all aligned variable data were integrated into... OK Multidimensional time series data matrix of columns , where N represents the total number of observed variables, which is equal to the sum of the number of all detected genes, the number of all identified metabolites, and the number of all phenotypic indicators; This represents the total number of time points after unified resampling, in this embodiment. It equals 73.

[0093] The matrix Each row corresponds to a unique biological variable, and each column corresponds to a specific time point; matrix elements Indicates the first The variable in the first... Standardized values ​​at each time point.

[0094] In step S3, based on the multidimensional time-series data matrix, a dynamic Bayesian network algorithm is used to construct a tea tree stress resistance regulation network model that reflects the causal relationship between genes, metabolites and target traits.

[0095] The specific implementation process of this step is as follows: the multidimensional time series data matrix D constructed in step S2 is used as input data.

[0096] Initialize a standard two-time-slice dynamic Bayesian network structure, which consists of two contiguous time slices, labeled as follows: and +1, where T represents the current time. +1 represents the next moment. The set of nodes in the network. From the matrix All row variables constitute, i.e. It includes all genes, metabolites, and phenotypic indicators.

[0097] The learning objective of the network is to determine the set of directed edges between nodes. This allows the resulting directed acyclic graph to optimally interpret the observed time-series data.

[0098] A scoring function based on the Bayesian information criterion is used to evaluate the merits of different network structures. This scoring function strikes a balance between model fit and model complexity, preventing overfitting.

[0099] In the network structure search phase, a hill-climbing algorithm is used as an optimization strategy. Starting from an initial empty network or a random network, it iteratively adds, deletes, or reverses single directed edges, gradually moving towards the desired structure. The network structure with the higher score moves until a local optimum is reached.

[0100] After learning, network topology Each directed edge is determined. → Indicates the time of variable u The state of variable v in time The +1 state has a direct causal effect.

[0101] After the structure is determined, the maximum likelihood estimation algorithm is used to learn the conditional probability distribution parameters of each node in the network.

[0102] For continuous variables, assuming they follow a Gaussian distribution, their conditional probability distribution is composed of a linear combination of parent nodes plus Gaussian noise. Maximum likelihood estimation is then transformed into solving the coefficients of a multiple linear regression.

[0103] The absolute values ​​of these regression coefficients quantify the causal strength of the corresponding directed edges. The final output tea tree stress resistance regulation network model is a complete probabilistic graphical model, which not only defines the causal direction between variables but also quantifies the strength of causal influences, and can be used for subsequent simulations and predictions.

[0104] In step S4, based on a preset multidimensional quantitative evaluation index system, a set of key regulatory genes is screened and sorted in the tea tree stress resistance regulation network model. The specific implementation process of this step is as follows: First, for each gene node in the network... Calculate its key regulatory index The formula for calculating this index is:

[0105] ;

[0106] , , The preset weighting coefficients satisfy... In this embodiment, we take =0.4, =0.4, =0.2.

[0107] To standardize network centrality, it is calculated as follows: first calculate the node centrality... out of degree That is, from Count the number of directed edges sent; then calculate the number of nodes. betweenness centrality That is, all shortest paths in the network pass through The proportion; then and A weighted average is calculated with each weight being 0.5 to obtain the original centrality score. Finally, the score is standardized using the Z-score method to make its mean 0 and standard deviation 1.

[0108] To standardize the strength of causal effects, the calculation method is as follows: extract from nodes The absolute values ​​of the conditional probability distribution parameters (i.e., regression coefficients) corresponding to all directed edges are taken and summed to obtain the original causal effect strength, which is then standardized using the Z-score method.

[0109] The standardized path length to the target phenotypic trait is calculated as follows: In the network, with the node representing the content of the target secondary metabolite as the endpoint, the path length from the node is calculated. The weighted shortest path distance to the endpoint is calculated, with the path weight being the reciprocal of the causal strength of the corresponding edge. The shorter the distance, the more direct the control path. The original path length is also standardized using the Z-score method.

[0110] Substitute the above three standardized indicators into The formula calculates the key regulatory index for each gene node.

[0111] Then all gene nodes were arranged according to The values ​​are sorted in descending order from highest to lowest.

[0112] Finally, the top 5% of genes were selected as the set of key regulatory genes. .

[0113] The genes in this set are considered to occupy a central hub position in the network and have the strongest comprehensive regulatory potential for downstream target traits.

[0114] In step S5, the tea tree stress resistance regulatory network model is used to perform in vivo computer perturbation simulation on each gene in the key regulatory gene set to quantitatively predict the effect of changes in their expression levels on the content of target secondary metabolites.

[0115] The specific implementation process of this step is as follows: traverse every gene in the key regulatory gene set KG. .

[0116] For each gene Two independent perturbation simulations were performed: the first was for high expression perturbation, and the second was for low expression perturbation.

[0117] In high expression perturbation, gene k is placed in the initial time slice The node state =0 is forcibly set to twice its mean expression at all time points in the original data; in low expression perturbation, it is set to 0.5 times the original mean.

[0118] After setting the initial perturbation state, the forward sampling logical stearic sampling algorithm is used to recursively simulate the network state. This algorithm starts from... Starting with 0, calculate based on the conditional probability distribution already learned in the network. The probability distribution of the states of all other nodes at time =1 is used to sample the specific instance state; then, using The sampled state at T=1 is taken as input, and the state at T=2 is calculated and sampled. This process is repeated time-sliced ​​forward until the last time slice. =72.

[0119] This process is repeated 1000 times to obtain a sufficient sample size for statistical inference.

[0120] After the simulation was completed, the target secondary metabolite nodes, namely catechins and epigallocatechin gallate, were extracted. The mean of 1000 predicted concentration samples at time 72 is calculated as the predicted concentration under perturbation.

[0121] At the same time, running a baseline simulation without any disturbance also yields the results for the target metabolite in... =72 is the baseline predicted average concentration.

[0122] The predicted disturbance concentration is compared with the baseline concentration to calculate the net percentage impact, using the formula: (disturbance concentration - baseline concentration) / baseline concentration × 100%.

[0123] right This process was repeated for all genes to obtain a complete perturbation effect dataset, which recorded the quantitative effects of each key gene on the two target metabolites under high / low expression perturbation.

[0124] In step S6, based on the quantitative prediction results of the in vivo computer perturbation simulation, the target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites are determined.

[0125] The specific implementation process for this step is as follows:

[0126] First, organize all the perturbation simulation results obtained in step S5 and construct the comprehensive effect matrix. .

[0127] Each row of the matrix corresponds to a perturbation type (high or low expression) of a key regulatory gene, and the columns include the gene identifier, perturbation type, percentage effect on catechin content, and percentage effect on epigallocatechin gallate content.

[0128] Secondly, a threshold for the regulatory effect is set. In this embodiment, the threshold is that the effect of increasing the total content of the target secondary metabolites must be greater than 20%.

[0129] The target total content is defined as the weighted sum of the contents of catechins and epigallocatechin gallate, with the weights set according to their importance in tea quality.

[0130] Next, filter the rows in the combined effects matrix EM that meet the following two conditions:

[0131] First, its effect on increasing the total content of the target is greater than 20%;

[0132] Second, its effects on both catechin and epigallocatechin gallate metabolites are positive, meaning the percentage effect is greater than 0.

[0133] Genes that meet the criteria are the final breeding target genes.

[0134] Finally, for each selected target gene, the regulatory strategy is determined based on the type of perturbation that produces the greatest positive effect: if high expression perturbation brings the greatest benefit, the regulatory strategy is to genetically modify it through gene overexpression technology; if low expression perturbation brings the greatest benefit, the regulatory strategy is to knock down its expression through gene silencing technology or CRISPR-Cas9 gene editing technology.

[0135] This generates a breeding program containing a list of target genes and their corresponding precise regulatory strategies, which can be directly used in subsequent molecular design breeding practices.

[0136] As part of the system of this invention, a dedicated data processing and analysis system was constructed to support the implementation of the above method.

[0137] The system is deployed on a high-performance computing cluster and includes a data acquisition module, a data preprocessing module, a network modeling module, a gene evaluation module, a perturbation simulation module, and a strategy generation module.

[0138] The data acquisition module is responsible for receiving and storing raw data streams from sequencers, mass spectrometers, and physiological and biochemical testing equipment.

[0139] The data preprocessing module executes the standardization and spatiotemporal alignment algorithm in step S2, which integrates a cubic spline interpolation engine and a data matrix builder.

[0140] The network modeling module implements the dynamic Bayesian network learning algorithm in step S3, including a BIC score calculator, a hill climber, and a maximum likelihood parameter estimator.

[0141] The gene evaluation module executes the multidimensional quantitative evaluation process in step S4, and has built-in network centrality calculator, causal effect strength calculator and shortest path algorithm.

[0142] The perturbation simulation module is responsible for the in vivo computer simulation of step S5. Its core is the forward sampling logic stear sampling engine, which can efficiently perform large-scale Monte Carlo simulations.

[0143] The strategy generation module executes the screening and decision-making logic of step S6, and automatically outputs the final list of breeding target genes and regulatory instructions based on preset thresholds.

[0144] All modules communicate via a unified data bus and share a central database, ensuring data flow consistency and automated processing. This system provides a solid hardware and software foundation for the efficient and accurate implementation of the method of this invention.

[0145] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish an entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0146] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for enhancing the stress resistance of tea plant secondary metabolites, characterized in that, include: A multi-omics time-series dataset of tea plants under specified abiotic stress conditions was obtained. The multi-omics time-series dataset includes genomic sequence data, transcriptomic data, metabolomic data, and phenotypic data collected at different time points. Data standardization and spatiotemporal alignment are performed on the multi-omics time series datasets to construct a unified multi-dimensional time series data matrix; Based on the multidimensional time-series data matrix, a dynamic Bayesian network algorithm is used to construct a tea tree stress resistance regulation network model that reflects the causal relationship between genes, metabolites and target traits. Based on the preset multidimensional quantitative evaluation index system, the key regulatory gene set was screened and sorted in the tea tree stress resistance regulation network model; Using the aforementioned tea plant stress resistance regulatory network model, in vivo computer perturbation simulations were performed on each gene in the key regulatory gene set to quantitatively predict the impact of changes in their expression levels on the content of target secondary metabolites, including: Genes in the key regulatory gene set were selected sequentially as perturbation targets; In the tea tree stress resistance regulation network model, the node state in the initial time slice is forcibly set to a high expression level, which is twice the original average expression level, or a low expression level, which is 0.5 times the original average expression level. The forward sampling logical stear sampling algorithm is adopted. Based on the conditional probability distribution learned in the network, the state probability distribution of all other nodes in the network is recursively calculated from the initial time slice to the last time slice. Extract the target secondary metabolites, namely catechins and epigallocatechin gallate, which characterize stress resistance, and the mean predicted concentrations at the last time step. The predicted mean concentration was compared with the baseline mean concentration without perturbation to calculate the net percentage effect of the change in expression level on the content of the target secondary metabolite. This perturbation simulation process is repeated for all genes in the key regulatory gene set; Based on the quantitative prediction results of the in vivo computer perturbation simulation, the target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites were identified.

2. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 1, characterized in that, The multi-omics time-series datasets are subjected to data standardization and spatiotemporal alignment to construct a unified multidimensional time-series data matrix, including: Transcriptome gene expression data were standardized using the number of reads per million transcripts; Metabolomics data were standardized using a combination of internal standard normalization and total ion current normalization. All time-series data, including transcriptomic, metabolomic, and phenotypic data, are processed using a cubic spline interpolation algorithm to fit them from discrete sampling time points into continuous time function curves. Then, at a uniform, preset time resolution, i.e., every hour, resampling is performed to generate data sequences with perfectly aligned time points. Integrate all aligned omics data into OK A multidimensional time-series data matrix of columns, in which This represents the total number of observed variables, including all genes, metabolites, and phenotypic indicators. This represents the total number of time points after unified resampling.

3. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 1, characterized in that, Based on the aforementioned multidimensional time-series data matrix, a dynamic Bayesian network algorithm is used to construct a tea tree stress resistance regulation network model reflecting the causal relationship between genes, metabolites, and target traits, including: Using the multidimensional time-series data matrix as input, a two-time-slice dynamic Bayesian network structure is initialized. This structure includes the current time slice. With the next time film +1; A hill-climbing algorithm based on Bayesian information criterion scoring is used as a structure learning strategy to search for the optimal network topology in the variable space. The network topology is a directed acyclic graph, where nodes represent genes or metabolites and directed edges represent direct causal effects between variables at time lag units. The maximum likelihood estimation algorithm is used to learn the conditional probability distribution parameters of each node in the network and quantify the causal strength of directed edges. Through an iterative learning process, a fully defined tea tree stress resistance regulation network model is finally output. This model accurately describes how the state of any gene or metabolite at time T probabilistically determines the state of other genes or metabolites at time T+1.

4. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 1, characterized in that, Based on a pre-defined multidimensional quantitative evaluation index system, a set of key regulatory genes was screened and ranked in the tea tree stress resistance regulation network model, including: The key regulatory index is defined, and its calculation formula is as follows: ; , , The preset weighting coefficients satisfy... ; To standardize network centrality, the calculation method is as follows: first calculate the node centrality. out of degree That is, from Count the number of directed edges sent; then calculate the number of nodes. betweenness centrality That is, all shortest paths in the network pass through The proportion; then and A weighted average is calculated with each weight being 0.5 to obtain the original centrality score. Finally, the score is standardized using the Z-score method to make its mean 0 and standard deviation 1. To standardize the causal effect strength, the calculation method is as follows: extract from the node The absolute values ​​of the conditional probability distribution parameters corresponding to all directed edges are taken and summed to obtain the original causal effect strength, which is then standardized using the Z-score method. To determine the standardized path length to the target phenotypic trait, the calculation method is as follows: In the network, with the node representing the content of the target secondary metabolite as the endpoint, calculate the path length from the node... The weighted shortest path distance to the destination is calculated, with the path weight being the reciprocal of the causal strength of the corresponding edge. The shorter the distance, the more direct the control path. The original path length is also standardized using the Z-score method.

5. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 1, characterized in that, Based on the quantitative prediction results of the in vivo computer perturbation simulation, target genes and their regulatory strategies for enhancing the stress resistance of tea plant secondary metabolites were identified, including: The in vivo computer perturbation simulation results of all key regulatory genes were compiled to form a comprehensive effect matrix that includes gene identifiers, perturbation types, percentage effects on catechin content, and percentage effects on epigallocatechin gallate content. Set thresholds for regulatory effectiveness; Genes that meet the regulatory benefit threshold and exhibit positive regulatory effects on both target metabolites in the comprehensive effect matrix are selected. The selected genes are used as the final breeding target genes, and the corresponding regulatory strategies are determined based on the type of perturbation that produces the greatest positive effect in the simulation, i.e., high expression or low expression, i.e., genetic improvement through gene overexpression technology or gene silencing technology.

6. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 1, characterized in that, The total content of the target secondary metabolites is the weighted sum of the contents of catechins and epigallocatechin gallate, with the weights set according to their importance in tea quality.

7. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 4, characterized in that, The standardized network centrality is standardized by using the Z-score method to standardize the original centrality score, which is the weighted average of the out-degree and betweenness centrality scores.

8. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 4, characterized in that, The standardized causal effect strength is standardized by the Z-score method, which is the sum of the absolute values ​​of the conditional probability parameters of all directed edges originating from the gene node.

9. The method for enhancing the stress resistance of tea tree secondary metabolites according to claim 4, characterized in that, The standardized path length to the target phenotypic trait is standardized by the Z-score method using the weighted shortest path distance, where the edge weight of the weighted shortest path distance is the reciprocal of the corresponding causal strength.

Citation Information

Patent Citations

  • Predicting effects of gene regulatory sequences on endophenotypes using machine learning

    CA3258766A1

  • Multi-character collaborative screening and breeding method for dipsacus asper germplasm resources

    CN120555642A