Wheat quality trait gene positioning method and system based on correlation analysis
By applying standardized perturbation sequences to wheat quality traits, collecting dynamic molecular phenotypes and performing parameterization, the problem that static endpoint phenotypes in existing technologies cannot reveal the dynamic environmental genetic basis of wheat quality traits is solved. This enables precise genetic mapping of trait stability and environmental adaptability, supporting early breeding selection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KELAN AGRICULTURAL TECHNOLOGY (HENAN) CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing genome-wide association studies rely on static endpoint phenotypes, which cannot reveal the genetic basis of wheat quality traits under dynamic environments, resulting in limited explanatory power of genetic mapping and difficulty in ensuring environmental adaptability.
By applying standardized perturbation sequences, dynamic molecular phenotypes are collected, parameterized to obtain dynamic response spectrum parameters, and genome-wide association analysis is performed. Combined with causal mediation analysis and cross-scale prediction models, genetic loci controlling trait plasticity and environmental adaptability are located.
This study reveals and locates the complex dynamic processes controlling wheat quality traits, providing deeper biological insights and enabling precise genetic mapping of trait stability and environmental adaptability, thus supporting early breeding selection.
Smart Images

Figure CN121884931A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biotechnology and crop genetics and breeding technology, and in particular to a method and system for locating wheat quality trait genes based on association analysis. Background Technology
[0002] Wheat is one of the world's most important food crops, and the quality traits of its grains, such as protein content, gluten strength, and flavor, directly determine its final processing use and economic value. Improving wheat quality traits, especially its performance stability across different years and geographical environments, through genetic improvement has long been a core goal in wheat breeding. Genome-wide association analysis (GWAS), as an important modern molecular breeding technique, statistically correlates genetic variations across the entire genome with target traits, and has become a key tool for elucidating the genetic basis of complex traits and discovering superior gene resources.
[0003] Currently, in the genetic analysis of wheat quality traits, conventional genome-wide association studies (GWAS) typically rely on linking genotype data with endpoint phenotypes measured after natural growth in the field and post-harvest. These endpoint phenotypes are mostly static observations, such as final protein content, wet gluten content, or sedimentation value.
[0004] However, this analytical approach based on static endpoint phenotypes inherently limits its view of trait formation as a "black box." It merely captures the final result of the interaction between genetic and environmental factors throughout the reproductive period, completely ignoring the dynamic information of trait responses to environmental changes during development. Therefore, this method struggles to distinguish the genetic differences behind the same endpoint phenotype achieved by two distinct developmental pathways, thus losing a significant amount of valuable genetic information related to process regulation.
[0005] Furthermore, the field endpoint phenotype is the final cumulative result of the interaction between various complex, random, and uncontrollable environmental factors (such as drought, heat waves, and strong light) experienced by the plant throughout its growth period and a specific genotype. This highly confounded interaction effect makes it extremely difficult to accurately isolate the specific gene effects controlling the environmental adaptability or stability of traits from the final phenotypic values. Consequently, the located loci often have limited explanatory power, and their effectiveness under different environments is difficult to guarantee. Summary of the Invention
[0006] In view of the above problems, this invention is proposed to provide a method and system for wheat quality trait gene mapping based on association analysis to overcome or at least partially solve the above problems. It can solve the problem that existing gene mapping methods rely on static endpoint phenotypes, thus failing to reveal and locate the genetic basis of the association between the dynamic processes controlling crop response to environmental changes and trait stability. It can also reveal and locate genetic loci that control complex dynamic processes such as trait plasticity, environmental adaptability and stress resilience, providing a more profound biological insight.
[0007] Specifically, according to one aspect of the present invention, a method for mapping wheat quality trait genes based on association analysis includes the following steps:
[0008] S1. Obtain whole-genome typing data of the associated population;
[0009] S2. Apply a standardized perturbation sequence to the associated population;
[0010] S3. After applying the standardized perturbation sequence, biological samples of individuals in the associated population are collected at multiple preset time points, and the concentration of key molecules in the samples is detected to obtain dynamic molecular phenotypes.
[0011] S4. The dynamic molecular phenotype is parameterized to obtain a set of quantitative dynamic response spectrum parameters;
[0012] S5. Using the dynamic response spectrum parameters as a phenotype, perform genome-wide association analysis with the whole-genome typing data to locate gene loci associated with the dynamic response spectrum parameters.
[0013] Optionally, in S2, applying the normalized perturbation sequence specifically includes:
[0014] The grouping of individuals within the associated group includes at least a single perturbation group and a sequence perturbation group;
[0015] Apply a single challenge coercion to the single disturbance group;
[0016] A stress sequence comprising a pre-processing stress and the challenge stress is applied to the perturbation group of sequences.
[0017] Optionally, the parameterization process in S4 further includes:
[0018] Stress memory parameters are extracted by comparing the dynamic molecular phenotypes of the sequence perturbation group and the single perturbation group; the stress memory parameters include response intensity gain ( The calculation method is as follows:
[0019] ;
[0020] in, The peak response intensity of the sequence perturbation group. The peak response intensity of the single disturbance group.
[0021] Optionally, the dynamic response spectrum parameters in S4 include: peak response intensity ( Peak response time ), response recovery rate ( ) and total response ( ).
[0022] Optionally, the method further includes the following steps:
[0023] Obtain the endpoint quality traits of the associated population in a field environment and calculate the trait stability parameters;
[0024] Causal mediation analysis was performed on the gene loci located in S5 to examine the mediating effect of the dynamic response spectrum parameters in the genetic pathway of "gene locus-trait stability parameters".
[0025] Optionally, the causal mediation analysis includes calculating indirect effects ( The calculation method is as follows:
[0026] ;
[0027] Wherein, a represents the effect of the gene locus on the dynamic response spectrum parameter, and b represents the effect of the dynamic response spectrum parameter on the trait stability parameter after controlling the gene locus.
[0028] Optionally, the trait stability parameter is the coefficient of variation of the endpoint quality trait under multiple environments ( ). ).
[0029] Optionally, the method further includes:
[0030] Based on the results of the causal mediation analysis, a cross-scale prediction model is constructed with the dynamic response spectrum parameters as feature variables and the trait stability parameters as response variables.
[0031] Optionally, the key molecules in S3 are metabolites or transcripts selected from metabolic pathways related to wheat quality.
[0032] The present invention also provides a localization system based on any of the above-mentioned gene localization methods, comprising:
[0033] The data acquisition module is used to acquire whole-genome typing data of related populations;
[0034] A perturbation execution module is used to apply a standardized perturbation sequence to the associated population;
[0035] The phenotypic acquisition module is used to collect biological samples of individuals in the associated population at multiple preset time points after applying the standardized perturbation sequence, and to detect the concentration of key molecules in the samples in order to obtain dynamic molecular phenotypes.
[0036] The parameterization module is used to parameterize the dynamic molecular phenotype to obtain a set of quantitative dynamic response spectrum parameters.
[0037] The association analysis module is used to perform genome-wide association analysis with the dynamic response spectrum parameters as phenotypes and the whole genome typing data to locate gene loci associated with the dynamic response spectrum parameters.
[0038] This invention discloses a method and system for mapping wheat quality trait genes based on association analysis. By applying standardized perturbation sequences and collecting time-series dynamic molecular phenotypes, this invention transforms the research object of association analysis from a static endpoint trait to a dynamic response process. Compared to traditional methods that can only link genotypes to the final result, this invention can capture and quantify the dynamic characteristics of genes in response to external stimuli, such as speed, intensity, and resilience. This allows for the revelation and localization of genetic loci controlling complex dynamic processes such as trait plasticity, environmental adaptability, and stress resilience, providing a more profound biological insight.
[0039] Furthermore, in the wheat quality trait gene mapping method and system based on association analysis of the present invention, by parameterizing the dynamic molecular phenotype, this method reduces the dimensionality of high-dimensional and complex response curves to low-dimensional parameters that can be directly used for association analysis. This not only solves the technical problem that time-series data is difficult to use directly for genome-wide association analysis, but also makes the genetic mapping target more precise and can directly deconstruct the genetic basis that controls specific dynamic response links.
[0040] Furthermore, the wheat quality trait gene mapping method and system based on association analysis of the present invention establishes a verifiable and quantitative causal bridge between the microscopic molecular dynamic response measured in the laboratory and the macroscopic trait stability observed in the field by introducing causal mediation analysis and cross-scale prediction models. This not only clarifies the mechanism of action of the mapped genes, that is, verifies that genes regulate the final field trait performance by influencing specific dynamic response processes, but also enables the construction of a model based on rapid laboratory measurement results to predict the stability of breeding materials in variable field environments, providing a new technical approach for achieving precise and efficient early breeding selection.
[0041] The above and other objects, advantages and features of the present invention will become more apparent to those skilled in the art from the following detailed description of specific embodiments of the invention in conjunction with the accompanying drawings. Attached Figure Description
[0042] The following sections will describe some specific embodiments of the invention in detail by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0043] Figure 1 This is a flowchart illustrating a method for locating wheat quality trait genes based on association analysis according to an embodiment of the present invention.
[0044] Figure 2 This is a system architecture diagram of a wheat quality trait gene localization system based on association analysis according to an embodiment of the present invention. Detailed Implementation
[0045] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0046] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0047] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may represent singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to such processes, methods, products, or apparatus.
[0048] Figure 1 This is a flowchart of a wheat quality trait gene mapping method based on association analysis according to an embodiment of the present invention, as shown below. Figure 1 As shown, this embodiment of the invention provides a molecular design breeding method for increasing the content of functional components in wheat, comprising the following steps:
[0049] S1. Obtain whole-genome typing data of the associated population.
[0050] S2. Apply a standardized perturbation sequence to the associated group.
[0051] S3. After applying the standardized perturbation sequence, biological samples of individuals in the associated population are collected at multiple preset time points, and the concentration of key molecules in the samples is detected to obtain dynamic molecular phenotypes.
[0052] S4. The dynamic molecular phenotype is parameterized to obtain a set of quantitative dynamic response spectrum parameters.
[0053] S5. Using the dynamic response spectrum parameters as a phenotype, perform genome-wide association analysis with the whole-genome typing data to locate gene loci associated with the dynamic response spectrum parameters.
[0054] In this embodiment, by applying standardized perturbation sequences and collecting time-series dynamic molecular phenotypes, the research object of association analysis is transformed from a static endpoint trait to a dynamic response process. Compared with traditional methods that can only associate genotype with the final result, this invention can capture and quantify the dynamic characteristics of genes in the process of responding to external stimuli, such as speed, intensity, and resilience. This allows for the revelation and localization of genetic loci controlling complex dynamic processes such as trait plasticity, environmental adaptability, and stress resilience, providing a more profound biological insight.
[0055] In addition, this embodiment parameterizes dynamic molecular phenotypes, which reduces high-dimensional and complex response curves to low-dimensional parameters that can be directly used for association analysis. This not only solves the technical problem that time-series data is difficult to use directly for genome-wide association analysis, but also makes the genetic mapping target more precise and can directly deconstruct the genetic basis that controls specific dynamic response links.
[0056] In some embodiments of the present invention, step S1 provides a genetic basis for subsequent association analysis based on perturbation response. This step aims to construct a high-density, high-quality whole-genome genetic variation map, assigning each individual in the population an accurate digital genotype identity, and is the cornerstone for connecting genotype with subsequent dynamic response phenotype.
[0057] This step includes:
[0058] S11. Construction and preparation of related groups.
[0059] To ensure that subsequent analyses capture a broad range of genetic variations related to quality traits, this embodiment uses an association analysis population comprising N wheat germplasm resources. The purpose of constructing this population is to maximize genetic diversity; therefore, its sources preferably include core germplasm from different ecological regions of the world, local varieties that have undergone long-term local domestication, intermediate breeding materials carrying superior genes, and commercially validated varieties in production.
[0060] Before conducting formal experiments, to ensure the homozygosity and stability of the genetic background and reduce the interference of heterozygosity on subsequent analysis, all individual materials within the population underwent necessary purification treatment. Preferably, each individual underwent at least two generations of self-pollination to obtain homozygous inbred lines, thereby providing uniform and reproducible genetic material for subsequent genotyping and phenotypic identification.
[0061] S12, Implementation of whole-genome typing technology.
[0062] To obtain genetic variation information covering the entire genome, this embodiment performs high-throughput genotyping on all N individuals in the associated population.
[0063] Preferably, whole-genome resequencing technology is used. This technology is chosen because it can provide the most comprehensive genome-wide variation information. It can not only obtain high-density single nucleotide polymorphism markers, but also simultaneously detect multiple types of genetic variations such as insertions / deletions, structural variations, and copy number variations, providing the most complete data foundation for subsequent in-depth mining of genetic loci associated with traits.
[0064] In other preferred embodiments, high-density SNP chips or genotyping sequencing technologies can be used as alternatives or supplementary solutions, depending on cost and throughput requirements.
[0065] S13, Bioinformatics Analysis Process.
[0066] Regardless of the sequencing technology used, the raw data must be processed through a rigorous bioinformatics workflow to generate a high-quality genotype dataset for association analysis.
[0067] First, the raw reads obtained from sequencing undergo quality assessment and preprocessing to remove adapter sequences and low-quality reads. Then, the preprocessed, high-quality sequencing reads are aligned to the wheat reference genome. This alignment process aims to precisely locate each sequencing fragment to its original position on the genome.
[0068] After the alignment is completed, standardized bioinformatics tools are used to detect genetic variations. This process includes steps such as local realignment and base quality recorrection to identify SNP sites in all individuals in the population relative to the reference genome.
[0069] S14. Quality control of genotype data.
[0070] To ensure the accuracy and reliability of subsequent association analysis, the initial detected SNP marker set undergoes rigorous quality control and filtering. This step is crucial for eliminating false positive associations and improving statistical power. The filtering criteria primarily include:
[0071] Minimal allele frequency filtering: This step removes allele loci that occur at extremely low frequencies in the entire population. The aim is to eliminate rare variants for which there is insufficient statistical information for association testing; these rare variants may also originate from sequencing or alignment errors.
[0072] Marker missing rate filtering: SNP loci that failed to be successfully genotyped on too many individuals in the population are removed. This step aims to ensure that each marker used for analysis has sufficient data integrity.
[0073] Individual heterozygosity filtering: Individuals with an abnormally high proportion of heterozygous loci are removed. Since the materials used in this embodiment are purified, an abnormally high heterozygosity rate may indicate sample contamination or that the material itself is not completely homozygous. Removing such individuals can ensure the genetic purity of the population.
[0074] S15. Generate a high-quality, high-density whole-genome SNP marker matrix.
[0075] After step S14 above, a high-quality, high-density whole-genome SNP marker matrix is finally generated.
[0076] In this embodiment, the final generated matrix provides an accurate digital representation of the genotype information of each individual in the population at tens of thousands or even millions of genetic loci, forming a reliable genetic data foundation necessary for subsequent steps such as dynamic response spectrum parameterization, genome-wide association analysis, and cross-scale causal network inference.
[0077] In some embodiments of the present invention, step S2 is one of the core elements that distinguishes the present invention from traditional association analysis methods. Its purpose is to transform the passive observation of naturally occurring and uncontrollable environmental influences in the field into the application of active, repeatable, and precisely quantified physical or chemical perturbations under laboratory conditions, thereby stimulating and capturing the dynamic response process of gene networks in response to external stimuli.
[0078] This step includes:
[0079] S21. Preparation of the experimental system and plant culture.
[0080] To ensure the standardization of the perturbation application and the reproducibility of the experimental results, this step is performed in one or more plant cultivation systems with precisely controllable environmental conditions. Preferably, the system is an artificial climate chamber or an automated greenhouse, capable of precisely regulating key environmental factors such as photoperiod, light intensity, ambient temperature, relative humidity, and nutrient supply.
[0081] Before the experiment began, each germplasm in the associated population was propagated and potted in a controlled system. All plants were placed under uniform and optimized standard growing conditions until they reached the target growth stage for perturbation. This was to eliminate experimental errors introduced by inconsistencies in initial growth environment and developmental stage, ensuring that all individuals were in similar physiological states when subjected to perturbation. Preferably, the target growth stage for perturbation was selected as a period that plays a decisive role in the formation of final grain quality, such as the grain-filling stage of wheat.
[0082] S22. Experimental design and execution of perturbation sequences.
[0083] To systematically analyze the response patterns of gene networks under different stress histories, this embodiment grouped the cultured associated plant populations and applied a carefully designed perturbation sequence scheme. This scheme included at least the following three treatment groups:
[0084] Control group: A subset of plants served as the control group. Throughout the experimental period, this group was maintained under optimal, stress-free standard growth conditions. The purpose of setting up a control group was to provide a baseline state, against which all subsequent molecular phenotypic changes in perturbation treatment groups were compared, thereby accurately identifying the true response induced by the perturbation.
[0085] Single perturbation group: Another group of plants is used as a single perturbation group. During a pre-set challenge time ( A standardized, acute challenge stress pulse was applied to this group of plants. Standardization means that the type, intensity, and duration of the stress pulse were predefined and precisely controlled. For example, a heat shock pulse can be defined as rapidly increasing the ambient temperature from the normal growth temperature by a preset temperature difference (…). ), and maintain a standard duration ( (The temperature is then restored to normal.) The purpose of setting up a single perturbation group is to detect and quantify the primary response patterns of genetic systems to single, sudden environmental stimuli.
[0086] Sequence perturbation group: The third group of plants was used as the sequence perturbation group. This group was designed to probe the stress plasticity and memory effects of gene networks. Specifically, at an early pretreatment time ( The plants in this group were first subjected to a mild, non-lethal pretreatment stress. The intensity and type of this pretreatment stress could be the same as or different from the challenge stress. After the plants had undergone pretreatment and a recovery period, they were subjected to a challenge stress at the exact same time point as the single perturbation group. Apply the exact same challenge stress pulse to the perturbation group of the sequence.
[0087] By setting up a set of perturbations and comparing their subsequent responses with a single perturbation set, this invention can quantify the molecular basis of system state changes induced by pretreatment stress, i.e., the so-called "stress memory" or "cross tolerance." This design of probing the intrinsic properties of a system through a "pretreatment-challenge" sequence constitutes a key technical feature of the method of this invention.
[0088] In this embodiment, by applying the aforementioned grouping and perturbation sequences, the previously observed as a whole population is transformed into multiple subsets under different, controllable stress states. These plants, treated with different perturbation histories, have their internal molecular regulatory networks effectively activated, providing ideal and informative biological materials for subsequent high-throughput dynamic molecular phenotypic acquisition.
[0089] In some embodiments of the present invention, the fundamental purpose of step S3 is to capture and record the molecular-level response that changes continuously over time caused by a controllable perturbation, thereby transforming the phenotype under study from a static, endpoint observation into a high-dimensional dynamic dataset containing rich process information.
[0090] This step includes:
[0091] S31. Formulation and execution of time-series sampling scheme.
[0092] To fully depict the entire dynamic process of gene regulatory networks from stimulation to response, peak, and eventual recovery to a steady state, this embodiment has developed and implemented a sophisticated time-sequential sampling scheme.
[0093] Specifically, for each treatment group among the control group (CK), single perturbation group (SP), and sequential perturbation group (SQ) set in the aforementioned steps, after the start of the challenge stress (defined as T_0), at a series of pre-set time points... Dynamic sampling was conducted. These time points were chosen to cover key phases of the response process, including, for example, the initial response period, the rapid growth period, the peak period, and the recovery period.
[0094] The biological samples taken are preferably tissues directly related to the physiological processes of wheat grain quality formation. For example, the flag leaf, a key source organ for photosynthetic product supply during the grain-filling stage, can be selected, or the developing grain, the direct site of quality trait formation, can be selected.
[0095] To ensure that the collected samples accurately reflect the molecular state at the moment of sampling, the sampling process must be completed rapidly. The collected samples are immediately flash-frozen in liquid nitrogen and then transferred to an ultra-low temperature environment (e.g., -80°C) for storage. This operation aims to instantly terminate all biochemical reactions within the sample, thereby preserving the original abundance information of various molecules at each time point to the maximum extent, providing a high-quality sample guarantee for subsequent accurate molecular detection.
[0096] S32. Selection of key molecules and high-throughput detection.
[0097] This invention does not involve the indiscriminate detection of all molecules in a sample, but rather purposefully focuses on a predetermined set of key molecules. The selection of these key molecules is based on their known biological associations with the target trait (i.e., wheat quality). Preferably, this set of key molecules may be key amino acids involved in the synthesis and accumulation pathways of grain proteins, carbohydrate precursors involved in the starch synthesis pathway, or relevant phenolic and lipid metabolites affecting dough processing characteristics and flavor. In other embodiments, this set of key molecules may also be transcripts of key functional genes regulating these metabolic pathways.
[0098] To achieve high-throughput and high-precision quantification of the concentrations of these key molecules, this embodiment preferably employs targeted metabolomics technology. For example, liquid chromatography-tandem mass spectrometry (LC-MS / MS) can be used to perform absolute or relative quantitative analysis of a predetermined group of metabolites in the sample extract. This technology features high selectivity and high sensitivity, enabling accurate determination of the concentration of target compounds in complex biological matrices.
[0099] S33, the generation of dynamic molecular phenotypes.
[0100] Through the above sampling and detection process, for each individual in the associated population, under each treatment group (CK, SP, SQ), and for each key molecule being detected, this embodiment can obtain an ordered dataset consisting of its concentration values at a series of time points.
[0101] This dataset represents the dynamic molecular phenotype defined in this invention. Specifically, for any individual or molecule, its dynamic molecular phenotype can be represented as a function of concentration over time or a discrete-time series. ,in For the sampling time point, This represents the molecular concentration measured at this time point. This dynamic molecular phenotype provides richer biological information compared to traditional single-time-point phenotypic values. It not only includes the absolute level of molecular abundance, but more importantly, it records the speed, intensity, and recovery pattern of the molecule's response to external perturbations.
[0102] In this embodiment, the output of this step is a large, multi-dimensional, dynamic molecular phenotypic database. This database provides the raw data input for subsequent parameterization steps and is a key link in the information transformation from biological processes to quantitative traits that can be used for genetic analysis.
[0103] In some embodiments of the present invention, step S4 is crucial for transforming high-dimensional time-series data into low-dimensional, biologically significant quantitative traits. Its fundamental purpose is to abstract and refine the core dynamic characteristics contained in each complex response curve into a set of numerical values that can be directly used for genome-wide association analysis. In other words, this step uses mathematical methods to extract indicators describing the speed, strength, and duration of the response process, thereby creatively transforming the object of association analysis from endpoint state values into a quantitative description of the system's dynamic behavior.
[0104] Specifically, the following phenotypes are observed:
[0105] S41, Primary phenotype: Extraction of kinetic parameters.
[0106] First, this embodiment focuses on the dynamic molecular phenotype, i.e., the response curve, of each molecule in each individual obtained in a single perturbation group (SP). The data is processed to extract a set of parameters describing its fundamental dynamic characteristics. These parameters together constitute the primary phenotype of this invention.
[0107] Preferably, this process is implemented through function fitting or direct feature extraction algorithms. Specifically, the extracted dynamic parameters may include the following:
[0108] Peak response intensity ( This parameter quantifies the maximum change in molecular concentration that can occur in response to stress. It reflects the ability or flux limit of a gene regulatory network to produce relevant molecules in response to a stimulus. It is calculated as the difference between the maximum value of the response curve over the entire observation period and the initial value before the response begins.
[0109] ;
[0110] in, The response curves for a single disturbance group are shown. The moment when the challenge of coercion begins.
[0111] Peak response time ( This parameter quantifies the time required for the system to reach its maximum response intensity. It reflects the agility or rate of the molecular response. The smaller the value, the faster the system responds to the stimulus. It is defined as the time point at which the response curve reaches its maximum value.
[0112] ;
[0113] Response recovery rate ( This parameter quantifies the ability or rate at which the molecular concentration of a system recovers to its baseline steady-state level after reaching a peak response. It reflects the resilience of the system or the efficiency of its negative feedback regulation mechanism. Preferably, this parameter can be obtained by fitting an exponential decay function (e.g., ...) to the decay phase of the response curve. To obtain, among which This is the attenuation coefficient.
[0114] Total response ( This parameter is obtained by calculating the area under the response curve, representing the total cumulative change in molecule concentration throughout the entire response and recovery process. It comprehensively reflects the strength and duration of the response and can be used to assess the overall biological effect of the molecule throughout the stress event. It is calculated by integrating the baseline-corrected response curve.
[0115] ;
[0116] in, This is the time point at which the observation ended.
[0117] S42, Secondary phenotype: Derivation of stress memory parameters.
[0118] Secondly, to further explore the plasticity and memory effects of gene networks, this embodiment derives a set of higher-order quantitative parameters describing stress memory by comparing the kinetic parameters of the sequence perturbation group (SQ) and the single perturbation group (SP). These parameters constitute the secondary phenotype of this invention.
[0119] Response intensity gain ( This parameter quantifies whether preprocessing stress enhances or weakens the system's response to subsequent challenge stress. It is obtained by calculating the difference in peak response intensity between a sequence of perturbations and a single perturbation.
[0120] ;
[0121] in, denoted as the peak response intensity of the sequence perturbation group.
[0122] Response rate gain ( This parameter quantifies whether preprocessing stress alters the agility of the system response. It is obtained by calculating the difference in peak response times between the two processing groups.
[0123] ;
[0124] in, is the peak response time of the sequence perturbation group.
[0125] Recovery ability bonus ( This parameter is used to quantify the impact of pretreatment stress on the system's resilience. It is obtained by calculating the difference in response recovery rates between the two treatment groups.
[0126] ;
[0127] in, is the response recovery rate of the sequence perturbation group.
[0128] In this embodiment, the raw time-series data of each individual and each molecule were successfully transformed into a set of quantitative phenotypic vectors containing multiple kinetic and stress memory parameters. This novel phenotypic set not only significantly reduces the data dimensionality, but more importantly, each parameter has a clear biological interpretation. This provides more informative and directly relevant inputs to subsequent genome-wide association studies, thus laying the foundation for locating gene loci that control these complex dynamic characteristics.
[0129] This embodiment parameterizes dynamic molecular phenotypes, reducing high-dimensional and complex response curves to low-dimensional parameters that can be directly used for association analysis. This not only solves the technical problem that time-series data is difficult to use directly for genome-wide association analysis, but also makes the genetic mapping target more precise and can directly deconstruct the genetic basis that controls specific dynamic response links.
[0130] Although existing studies have explored the molecular responses of specific genes to single stresses under controlled laboratory conditions, these studies are often disconnected from the macroscopic phenotypic performance of crops in variable field environments, lacking a bridge connecting microscopic molecular dynamics with macroscopic agronomic phenotypic stability. Current technologies generally lack an effective method to actively stimulate and systematically quantify the dynamic response characteristics of gene networks (such as response rate, resilience, stress memory, and plasticity) and use them as direct breeding targets, which limits the in-depth understanding and utilization of the genetic mechanisms of crop environmental adaptation.
[0131] In some embodiments of this invention, although existing studies have explored the molecular responses of specific genes to single stresses under controlled laboratory conditions, these studies are often disconnected from the macroscopic phenotypic performance of crops in variable field environments, lacking a bridge connecting microscopic molecular dynamics and macroscopic agronomic phenotypic stability. Existing technologies generally lack an effective method to actively stimulate and systematically quantify the dynamic response characteristics of gene networks (such as response rate, resilience, stress memory, and plasticity) and use them as direct breeding targets, which limits the in-depth understanding and utilization of the genetic mechanisms of crop environmental adaptation.
[0132] Therefore, in step S5, the purpose of this step is to establish a statistical association between the whole genome typing data obtained in the preceding steps and the newly created dynamic response spectrum parameters, so as to identify and locate those genetic loci (genes or regulatory elements) that control specific dynamic response characteristics across the whole genome.
[0133] This step includes:
[0134] S51. Preparation of input data for association analysis.
[0135] The execution of this step relies on two sets of core data: the first is the high-quality whole-genome SNP marker matrix constructed in the previous step, which represents the genetic background information (genotype) of each individual in the population; the second is the dynamic response spectrum parameters obtained by parameterization in the previous step, which represents the dynamic response characteristics (phenotype) of each individual on a specific molecule.
[0136] Specifically, this method performs a genome-wide association analysis independently for each dynamic response spectrum parameter. For example, it analyzes the peak response intensity of all individuals in the associated population to a specific molecule. The set of values is used as a phenotypic vector. The peak response time of all individuals was correlated with the SNP genotype matrix of the entire population. Similarly, the peak response time of all individuals was correlated with the peak response time of the entire population. Value set, response recovery rate ) value set, and even stress memory parameters (such as Each set of values is used as an independent phenotypic vector for genome-wide association analysis.
[0137] S52, Correction of population genetic structure.
[0138] To avoid spurious associations caused by confounding factors such as group structure and implicit kinship between individuals, this embodiment requires the evaluation and correction of these factors before conducting association analysis. This is a crucial prerequisite for ensuring the reliability of the final localization results.
[0139] Preferably, using a quality-controlled whole-genome SNP dataset, the genetic structure of the population is inferred through principal component analysis (PCA) or model-based clustering methods to generate a population structure matrix (Q matrix). Simultaneously, the genetic similarity between any two individuals in the population is calculated using all SNP markers to construct a kinship matrix (K matrix). These two matrices will be included as covariates in subsequent association analysis models to effectively control for false positives caused by population stratification and kinship.
[0140] S53. Implementation of the genome-wide association analysis model.
[0141] To control for the aforementioned confounding factors and examine the significance of the association between each SNP marker and the target dynamic response spectrum parameters, this embodiment preferably employs a mixed linear model for genome-wide association analysis (MLM). MLM can simultaneously treat population structure as a fixed effect and kinship as a random effect.
[0142] The mathematical expression for this model is as follows:
[0143] ;
[0144] The definitions and functions of each parameter are as follows:
[0145] It is The phenotypic vector, whose elements are the phenotype vectors of the associated population. For a specific dynamic response spectrum parameter (e.g.), an individual... The observed values.
[0146] It is The fixed effects design matrix. This matrix contains at least two parts: a The genotype vector of the target SNP marker to be tested, and a Other covariates, such as the Q-matrix used to correct population structure.
[0147] It is The fixed effects parameter vector, which includes the effect values of the SNP to be tested and the effect values of other covariates, is one of the core parameters that the model needs to estimate.
[0148] It is The random effects design matrix is usually an identity matrix.
[0149] It is This vector represents the random genetic background effects, used to capture polygenic cumulative effects determined by kinship that are not included in fixed effects. This vector is assumed to follow a multivariate normal distribution, i.e. ,in The aforementioned kinship matrix is the result of the calculation. It represents additive genetic variance.
[0150] It is The random residual vector, assumed to follow a normal distribution. ,in It is the identity matrix. This represents the residual variance.
[0151] In other implementations, to further improve computational efficiency or statistical power, derivative models of MLM or other advanced multi-site association analysis models may be used.
[0152] S54. Identification of associated sites.
[0153] By sequentially performing the model fitting and testing on each SNP marker on the genome, this method obtains the p-value of the association between each SNP and a specific dynamic response spectrum parameter. To control the false positive rate caused by multiple testing, a strict significance threshold needs to be set. SNP loci with p-values below this threshold are identified as gene loci significantly associated with that dynamic response spectrum parameter.
[0154] S15. Generate a high-quality, high-density whole-genome SNP marker matrix.
[0155] After step S14 above, a high-quality, high-density whole-genome SNP marker matrix is finally generated.
[0156] In this embodiment, the output of this step is a series of lists, each corresponding to a dynamic response spectrum parameter, and detailing the SNP loci significantly associated with this parameter across the entire genome, along with their locations, effect values, and significance levels. These located loci are key candidate gene regions controlling the dynamic response behavior of crops under specific perturbations, providing precise targets for subsequent functional validation and molecular breeding applications.
[0157] In this embodiment, the present invention establishes a verifiable and quantitative causal bridge between the microscopic molecular dynamic response measured in the laboratory and the macroscopic trait stability observed in the field by introducing causal mediation analysis and cross-scale prediction models. This not only clarifies the mechanism of action of the located genes, that is, verifies that genes regulate the final field trait performance by influencing specific dynamic response processes, but also enables the construction of a model based on rapid laboratory measurement results to predict the stability of breeding materials in variable field environments, providing a new technical approach for achieving precise and efficient early breeding selection.
[0158] Please see the appendix Figure 2 The present invention also provides a system based on any of the above-mentioned association analysis-based methods for mapping wheat quality trait genes, comprising:
[0159] The data acquisition module is used to acquire whole-genome typing data of related populations;
[0160] The perturbation execution module is used to apply a standardized perturbation sequence to the associated population;
[0161] The phenotypic acquisition module is used to collect biological samples from individuals in an associated population at multiple preset time points after applying a standardized perturbation sequence, and to detect the concentration of key molecules in the samples in order to obtain dynamic molecular phenotypes.
[0162] The parameterization module is used to parameterize dynamic molecular phenotypes to obtain a set of quantitative dynamic response spectrum parameters.
[0163] The association analysis module is used to perform genome-wide association analysis with dynamic response spectrum parameters as phenotypes and whole-genome typing data to locate gene loci associated with dynamic response spectrum parameters.
[0164] This invention transforms the research object of association analysis from a static endpoint trait to a dynamic response process by applying standardized perturbation sequences and collecting time-sequential dynamic molecular phenotypes. Compared to traditional methods that can only link genotypes to the final result, this invention can capture and quantify the dynamic characteristics of genes in response to external stimuli, such as speed, intensity, and resilience. This allows for the revelation and localization of genetic loci controlling complex dynamic processes such as trait plasticity, environmental adaptability, and stress resilience, providing a more profound biological insight.
[0165] Furthermore, this invention parameterizes dynamic molecular phenotypes, reducing high-dimensional and complex response curves to low-dimensional parameters that can be directly used for association analysis. This not only solves the technical challenge of using time-series data directly for genome-wide association analysis, but also makes genetic mapping more precise and can directly deconstruct the genetic basis controlling specific dynamic response links.
[0166] Finally, by introducing causal mediation analysis and cross-scale prediction models, this invention establishes a verifiable and quantitative causal bridge between the microscopic molecular dynamic response measured in the laboratory and the macroscopic trait stability observed in the field. This not only elucidates the mechanism of action of the located genes, i.e., verifies that genes regulate the final field trait performance by influencing specific dynamic response processes, but also enables the construction of models based on rapid laboratory measurements to predict the stability of breeding materials in variable field environments, providing a new technical approach for achieving precise and efficient early breeding selection.
[0167] Therefore, those skilled in the art should recognize that although numerous exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications conforming to the principles of the present invention can be directly determined or derived from the disclosure of the present invention without departing from the spirit and scope of the invention. Thus, the scope of the present invention should be understood and construed as covering all such other variations or modifications.
Claims
1. A method for mapping wheat quality trait genes based on association analysis, characterized in that, Includes the following steps: S1. Obtain whole-genome typing data of the associated population; S2. Apply a standardized perturbation sequence to the associated population; S3. After applying the standardized perturbation sequence, biological samples of individuals in the associated population are collected at multiple preset time points, and the concentration of key molecules in the samples is detected to obtain dynamic molecular phenotypes. S4. The dynamic molecular phenotype is parameterized to obtain a set of quantitative dynamic response spectrum parameters; S5. Using the dynamic response spectrum parameters as a phenotype, perform genome-wide association analysis with the whole-genome typing data to locate gene loci associated with the dynamic response spectrum parameters.
2. The method for mapping wheat quality trait genes based on association analysis according to claim 1, characterized in that, In S2, applying the standardized perturbation sequence specifically includes: The grouping of individuals within the associated group includes at least a single perturbation group and a sequence perturbation group; Apply a single challenge coercion to the single disturbance group; A stress sequence comprising a pre-processing stress and the challenge stress is applied to the perturbation group of sequences.
3. The method for mapping wheat quality trait genes based on association analysis according to claim 2, characterized in that, The parameterization process in S4 also includes: Stress memory parameters are extracted by comparing the dynamic molecular phenotypes of the sequence perturbation group and the single perturbation group; the stress memory parameters include response intensity gain ( The calculation method is as follows: ; in, The peak response intensity of the sequence perturbation group. The peak response intensity of the single disturbance group.
4. The method for mapping wheat quality trait genes based on association analysis according to claim 1, characterized in that, The dynamic response spectrum parameters in S4 include: peak response intensity ( Peak response time ), response recovery rate ( ) and total response ( ).
5. The method for mapping wheat quality trait genes based on association analysis according to claim 1, characterized in that, The method also includes the following steps: Obtain the endpoint quality traits of the associated population in a field environment and calculate the trait stability parameters; Causal mediation analysis was performed on the gene loci located in S5 to examine the mediating effect of the dynamic response spectrum parameters in the "gene locus-trait stability parameter" genetic pathway.
6. The method for mapping wheat quality trait genes based on association analysis according to claim 5, characterized in that, The causal mediation analysis includes calculating indirect effects ( The calculation method is as follows: ; Wherein, a represents the effect of the gene locus on the dynamic response spectrum parameter, and b represents the effect of the dynamic response spectrum parameter on the trait stability parameter after controlling the gene locus.
7. The method for mapping wheat quality trait genes based on association analysis according to claim 5, characterized in that, The stability parameter of the trait is the coefficient of variation of the endpoint quality trait under multiple environments. ).
8. The method for mapping wheat quality trait genes based on association analysis according to claim 5, characterized in that, The method also includes: Based on the results of the causal mediation analysis, a cross-scale prediction model is constructed with the dynamic response spectrum parameters as feature variables and the trait stability parameters as response variables.
9. The method for mapping wheat quality trait genes based on association analysis according to claim 1, characterized in that, The key molecules in S3 are metabolites or transcripts selected from metabolic pathways related to wheat quality.
10. A positioning system based on the gene positioning method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire whole-genome typing data of related populations; A perturbation execution module is used to apply a standardized perturbation sequence to the associated population; The phenotypic acquisition module is used to collect biological samples of individuals in the associated population at multiple preset time points after applying the standardized perturbation sequence, and to detect the concentration of key molecules in the samples in order to obtain dynamic molecular phenotypes. The parameterization module is used to parameterize the dynamic molecular phenotype to obtain a set of quantitative dynamic response spectrum parameters. The association analysis module is used to perform genome-wide association analysis with the dynamic response spectrum parameters as phenotypes and the whole genome typing data to locate gene loci associated with the dynamic response spectrum parameters.