A method for constructing a whole growth period simulation model of corn based on genomic information

By employing a three-step strategy based on gene information and utilizing whole-genome resequencing data and field trial data, a high-precision maize growth period simulation model was constructed. This solved the problems of time-consuming parameter calibration and "different parameters with the same effect" in large-scale breeding, and achieved efficient and accurate prediction of maize growth period.

CN121415861BActive Publication Date: 2026-04-17BEIJING CIIC INT INST OF BIOLOGICAL AGRI +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CIIC INT INST OF BIOLOGICAL AGRI
Filing Date
2025-12-26
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing maize breeding models are time-consuming and labor-intensive to calibrate parameters in large-scale breeding, suffer from the problem of "different parameters with the same effect", fail to fully utilize population genetic structure information, and are difficult to achieve efficient, physiologically reasonable parameter estimation and accurate prediction.

Method used

A three-step strategy based on gene information was adopted to construct a high-precision maize growth period simulation model using whole-genome resequencing data and field trial data. Parameter inversion and optimization were performed using population genetic structure analysis, and parameter modeling was carried out in conjunction with the APSIM model.

Benefits of technology

It significantly improves parameter optimization efficiency and prediction accuracy, takes into account individual differences and population genetic consistency, has the potential to diagnose heterogeneity within genetic subpopulations, and is suitable for high-throughput parameterized modeling of large-scale breeding programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415861B_ABST
    Figure CN121415861B_ABST
Patent Text Reader

Abstract

This invention relates to the field of smart agriculture and crop modeling technology, specifically disclosing a method for constructing a maize full-growth-period simulation model based on genetic information. The method first acquires whole-genome resequencing data and multi-growth-stage phenotypic data of a target maize population over several consecutive years at the same ecological location. Based on the resequencing data, the population is divided into several subpopulations with consistent genetic backgrounds. Subsequently, using a strategy of "single-family parameter inversion – subpopulation template construction – secondary optimization and calibration of single genotypes under subpopulation template constraints," the genetic subpopulation characteristics based on the genome are correlated with the APSIM crop model to construct a high-precision maize growth period prediction model. This invention overcomes the shortcomings of traditional crop models, such as low efficiency in parameter calibration for large-scale breeding materials and susceptibility to local optima. It is particularly suitable for efficient and accurate growth period prediction and adaptability evaluation of a large number of breeding materials within a unified ecological zone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart agriculture and crop modeling technology, and in particular to a parametric modeling method that integrates population genetics and process-oriented crop models, for the accurate and efficient simulation and prediction of the growth period of large-scale maize breeding materials. Background Technology

[0002] The growth period is one of the core agronomic traits that determines the regional adaptability and final yield of maize varieties. Accurate prediction of the growth period is of great significance for variety zoning, sowing date determination and high-yield cultivation measures (Tang et al., 2025; Wang et al., 2023). At present, crop models based on physiological and ecological processes (such as APSIM) are the main tools for predicting crop growth and development (He et al., 2024). However, in modern maize breeding practice, the application potential of traditional models has the following shortcomings: (1) The parameter calibration mode is crude. Traditional APSIM applications usually perform "variety-by-variety" parameter calibration for a single variety or a small number of varieties, which is time-consuming and labor-intensive. When there are hundreds or thousands of materials to be evaluated in the breeding program, this "one case, one calibration" mode obviously cannot meet the needs of high-throughput breeding. (2) The problem of "different parameters with the same effect": In the parameter inversion process of traditional crop models, there are usually complex correlations and compensation relationships between model parameters. Especially in large-scale parameter inversion, the model is prone to getting trapped in local optima, that is, although different parameter combinations produce similar fitting effects, they may not be physiologically reasonable. This situation leads to a lack of interpretability of the model and seriously affects the model's generalization ability and predictive stability. (3) The genetic structure information of the population is not fully utilized. With the development of high-throughput sequencing technology, whole-genome resequencing can provide high-density, whole-genome-wide variation data at a relatively low cost. These data provide new opportunities for maize breeding, allowing for a more accurate understanding of the population's genetic structure and revealing the differences between different genotypes. In recent years, in order to enhance the genetic realism of crop models, gene-based crop models have gradually become a research hotspot. Related studies have mainly explored two strategies: first, by directly embedding major genes into the model as parameters controlling vernalization requirements or photoperiod sensitivity, thereby simulating the staghorn or flowering time responses of different genotypes (He et al., 2024); second, by using the effects of major or minor QTLs to calibrate model parameters, thereby enhancing the model's ability to distinguish differences between genotypes (Bustos-Kortset et al., 2019; Hu et al., 2022; Kadam et al., 2019). Although these methods have demonstrated the feasibility of genetic parameterization, they generally rely on a small number of loci, making it difficult to reflect the essential characteristic of quantitative traits being regulated by the cumulative regulation of multiple genes. Furthermore, most existing implementations are only applicable to single traits or limited environments, rarely taking into account G×E interactions, phenotypic plasticity, and long-term adaptation. In summary, current technologies lack a parameterized modeling method for maize growth period that can efficiently and physiologically reasonable estimate parameters for hundreds or thousands of genotypes by using whole-genome information to characterize the genetic structure of materials at the population scale, while ensuring that the parameters can distinguish individual differences and maintain the consistency of genetic structure. Summary of the Invention

[0003] (a) Purpose of the invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a method for constructing a simulation model of the entire growth period of maize based on genetic information, so as to achieve rapid and accurate prediction of the growth period of large-scale breeding populations.

[0005] (II) Technical Solution

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The present invention provides a method for constructing a maize full growth period simulation model based on gene information. Based on field trial data and whole genome resequencing data from multiple consecutive years at the same ecological site, a high-precision maize growth period prediction model is constructed through population structure analysis and APSIM model parameterization modeling. The core of this method is a three-step strategy: Single-family parameter inversion: In the training year, each family is treated as an independent individual, and its specific reproductive period model parameters are inverted to obtain the initial parameter vector corresponding to each maize genotype; Subpopulation template aggregation: Based on the genetic subpopulations resolved by whole-genome resequencing, the initial parameter vectors of multiple maize genotypes belonging to the same genetic subpopulation are aggregated using statistical means, variances, and standard deviations to construct the parameter constraint template corresponding to that genetic subpopulation; Secondary optimization and calibration under subpopulation template constraints: Using the subpopulation parameter constraint templates obtained in the above steps as prior constraints based on population structure, secondary optimization and calibration are performed on the reproductive period model parameters of each maize genotype to simultaneously reduce the deviation between the simulated and measured reproductive period values ​​of that genotype under multiple environmental conditions and the deviation of the genotype parameters relative to the parameter constraint template of its genetic subpopulation, thus obtaining the final reproductive period model parameters corresponding to each maize genotype. In the above process, this invention predefines a set of key reproductive period sensitivity parameters in the APSIM model and sets reasonable biological value ranges for each parameter to ensure that the inversion results can both fit the observed data and have physiological rationality.

[0007] (III) Beneficial Effects

[0008] This invention introduces "population genetic structure" as a constraint variable for APSIM reproductive period parameters into the modeling process for the first time. This makes the model parameters no longer simple statistical fit quantities, but "individual-level parameters" that are offset around the center of physiological characteristics of genetic subpopulations. This effectively solves the key technical problems of traditional crop models at the population scale, such as "different parameters with the same effect, uninterpretable parameters, and unstable cross-year prediction performance".

[0009] Compared with the prior art, the present invention has the following beneficial effects:

[0010] (1) Significantly improves parameter optimization efficiency and prediction accuracy. By first inverting parameters at the single-family scale, then constructing parameter constraint templates at the genetic subpopulation scale, and performing secondary optimization and calibration on each genotype under template constraints, this invention can provide a search starting point close to the optimal solution for global optimization, greatly reduce the risk of getting trapped in local optima, accelerate the convergence speed, and improve the prediction accuracy of the final model.

[0011] (2) Balancing individual differences with population genetic consistency. The final output is a set of independent parameters for each genotype. At the same time, the deviation of parameters is controlled by template constraints, so that individual parameters can reflect the differences between materials without deviating from the physiological characteristics of their respective subgroups.

[0012] (3) Possesses the potential for diagnosing heterogeneity within genetic subpopulations. By performing statistical analysis of the mean, variance, and standard deviation of the initial and final genotype parameter vectors within each genetic subpopulation, this invention can identify genotypes with abnormal parameters or materials with high heterozygosity in genetic background during model inference, providing auxiliary decision-making basis for material screening, subpopulation division verification, and outlier detection during the breeding process.

[0013] (4) Make full use of whole-genome resequencing data. This invention transforms high-density resequencing genome data into parameter constraint templates with clear mechanistic significance, opening up the chain of "genome information – population structure – physiological parameters – phenotypic prediction", and significantly enhancing the mechanistic and interpretable nature of the model.

[0014] (5) Applicable to the high-throughput parameterized modeling requirements of large-scale breeding programs. This method can independently output the final reproductive period parameter vector for each genotype in the population. It has the characteristics of parallelization and scalability, and can process the reproductive period calibration and prediction tasks of thousands of materials in a unified ecological zone. It provides an efficient modeling tool for the multi-material, multi-year, and multi-environment phenotypic evaluation commonly found in modern breeding projects. Attached Figure Description

[0015] Figure 1 This is a flowchart of the overall method of the present invention, illustrating the sequence and relationship of steps including data acquisition, population structure analysis, single-family parameter inversion, subpopulation template construction, and genotype secondary optimization and calibration under template constraints.

[0016] Figure 2 is a statistical analysis chart of meteorological data collected in Pinggu District, Beijing, for two years (2021–2022) during the data collection period of this invention.

[0017] Figure 3 is a genetic structure diagram of the research population based on resequencing data in this invention.

[0018] Figure 4 shows the comparison results of the modeling method of the present invention and the traditional overall parameter optimization method in the simulation accuracy of the growing season in Pinggu District, Beijing for two years (2021–2022). The figure shows the scatter plot of the model prediction value and the field measured value and the relevant analysis results. Detailed Implementation

[0019] The method of this invention generally includes steps such as data acquisition, population structure analysis, single-family parameter inversion, subpopulation template construction, and secondary genotype optimization and calibration under template constraints. The specific process, sequence, and relationships are as follows: Figure 1 As shown.

[0020] Example 1: Data Preparation

[0021] Study area and materials: Pinggu District of Beijing was selected as a unified ecological site, and about 1223 maize inbred lines were planted at the site for two consecutive years. The experiment adopted a randomized block design.

[0022] Phenotypic data collection: During each growing season, key growth stages such as seedling emergence, tasseling, and flowering were recorded for each inbred line.

[0023] Genotypic data acquisition: Whole-genome resequencing was performed on all 1223 maize inbred lines, with an average sequencing depth of approximately 20X. After sequencing data quality control, alignment, and variant detection, a high-quality SNP dataset covering the entire genome was obtained.

[0024] Environmental and Management Data: Daily meteorological data, including maximum temperature (Tmax), minimum temperature (Tmin), precipitation, and solar radiation, were collected from the experimental station. Field management practices such as sowing date, irrigation, and fertilization were also recorded to drive the crop model. (Meteorological data reference) Figure 2 .

[0025] Example 2: Group Structure Analysis

[0026] The resequencing genomic data obtained in Example 1 were analyzed for population structure using ADMIXTURE software, and the optimal number of subpopulations K was determined by cross-validation error. Based on the K values, the 1223 self-crossed H1 individuals were divided into 7 subpopulations with clear genetic backgrounds, denoted as K1-K7. The subpopulation classification is as follows: Figure 3 As shown.

[0027] Example 3: Three-step parametric modeling

[0028] Step 3.1: Single-family parameter inversion

[0029] The field-measured growth period data from the first year in Pinggu District, Beijing, were used as the training set. APSIM growth period simulation models were independently constructed for each of the 1223 maize inbred lines in the population, and the following key sensitivity parameters for growth period were inverted: shoot_lag, shoot_rate, tt_emerg_to_endjuv, tt_endjuv_to_init, tt_flag_to_flower, photoperiod_crit1, photoperiod_crit2, and photoperiod_slope. The initial default values ​​and search ranges for these parameters were set according to the parameter ranges shown in Table 1.

[0030]

[0031] (1) Construction of objective function: For each maize inbred line family g, the following inversion objective function is constructed:

[0032]

[0033]

[0034] The objective function aims to minimize the sum of squared errors between the simulated and measured reproductive periods.

[0035] (2) Optimization Algorithm: In this embodiment, the differential evolution algorithm is used to optimize and invert the above objective function. The specific process includes: randomly initializing multiple candidate parameter individuals in the parameter space; continuously updating the population through mutation, crossover and selection operations; iteratively searching to gradually converge the objective function value to the minimum; when the objective function converges or reaches the maximum number of iterations, outputting the optimal parameter vector θ of the family. g Through high-performance parallel computing, parameter inversion was performed independently for each of the 1223 families, ultimately obtaining 1223 sets of initial parameter vectors.

[0036] Step 3.2: Subgroup Template Aggregation

[0037] All 1223 initial parameter vectors obtained in step 3.1 are assigned to K genetic subgroups C1–C7 according to the population structure classification results obtained in Example 2. For any genetic subgroup, the set of all family parameter vectors belonging to that subgroup is denoted as {θ_g|g∈Ck}. The mean of the statistics of this set on each parameter dimension is calculated, and the parameter constraint template vector Template_{Ck} for that subgroup is constructed.

[0038]

[0039]

[0040] Step 3.3: Secondary optimization calibration under subgroup template constraints

[0041] Based on the K subpopulation parameter constraint template obtained in step 3.2, this embodiment further incorporates all reproductive period observation data from the first and second years to perform constrained secondary optimization calibration on the parameter vector of each genotype. For the inbred family g belonging to the k-th genetic subpopulation Ck, the present invention constructs the following constrained objective function, and the final secondary optimized parameter θ is obtained after correction. g *

[0042]

[0043]

[0044] Example 4: Overall Predictive Performance Verification

[0045] To verify the effectiveness of the method of this invention, this embodiment selects field-measured growth period data from the Pinggu District of Beijing in the second year as an independent validation set to evaluate the generalization performance of the "template-constrained quadratic optimization model" constructed in Example 3. First, the genetic subpopulation to which the validation material belongs is determined. For any maize genotype g in the validation set, its genetic subpopulation Ck is determined based on its whole-genome resequencing data using the ADMIXTURE model output from Example 2. Second, the quadratic optimization parameter vector θ obtained in Example 3 for this subpopulation is called. g * , serving as input parameters for the APSIM model, drive the APSIM model to make independent predictions; finally, prediction performance evaluation metrics are calculated for all 1223 genotypes using the following two metrics:

[0046] (1) Root Mean Square Error (RMSE)

[0047]

[0048] (2) Coefficient of determination (R²)

[0049]

[0050]

[0051] Example 5: Comparison of Predictive Performance between Traditional Methods and the Method of the Present Invention

[0052] To further verify the stability and generalization ability of the population structure-driven parameterization method of the present invention in cross-year prediction, this embodiment uses field measured growth period data from Pinggu District, Beijing for two consecutive years to compare the performance of the two modeling methods: (1) Traditional overall parameter optimization method in the traditional method of overall optimization without considering population genetic structure: calibration stage (first year in Pinggu): R 2 The figure was 0.991, and the RMSE was 3.01 days; Independent assessment phase spanning the year (Pinggu's second year): R 2 The accuracy was 0.915, and the RMSE was 3.07 days. Although the goodness of fit was high within the calibration year, the prediction accuracy decreased significantly across years, indicating that the traditional method lacks the ability to model genetic differences between materials and has weak generalization performance.

[0053] (2) The predictive performance of the method of the present invention is as follows in the same year: Calibration phase (first year in Pinggu): R 2 The RMSE was 0.998, and the RMSE was 1.29 days; Independent assessment phase (Pinggu Year 2): R 2 The mean squared error was 0.984, and the RMSE was 1.29 days. Compared with traditional methods, the method of this invention significantly improves the linearity and error control of fertility period prediction, and maintains highly stable predictive performance, especially in cross-year assessments.

[0054] (3) Results analysis Figure 4 It can be seen that the prediction accuracy of the method of this invention is significantly higher than that of the traditional method: the RMSE of the reproductive period is reduced from about 3.0 days in the traditional method to about 1.3 days, and the error is reduced by more than 50%. The method of this invention has significantly stronger cross-year migration ability: the performance of the traditional method deteriorates significantly in cross-environment and cross-year predictions, while the method of this invention maintains consistent and high prediction accuracy between two years. Population structure information provides effective genetic constraints: the genotype-level final parameters obtained under the constraint of subpopulation templates enable the APSIM model to explain the genetic consistency and differences of materials within the population, thereby improving the robustness and interpretability of the model.

[0055] In summary, the "parametric modeling method for maize growth period that integrates population structure" provided by this invention is significantly superior to traditional modeling methods in terms of calibration accuracy, cross-year stability, and characterization of population genetic differences, and has stronger applicability and promotion value.

[0056] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. For those skilled in the art, various equivalent substitutions or modifications can be made to the above embodiments without departing from the spirit and substance of the present invention, and all such substitutions or modifications should fall within the protection scope of the present invention.

[0057] References:

[0058] Bustos-Korts, D., Malosetti, M., Chenu, K., Chapman, S., Boer, MP, Zheng, B., & van Eeuwijk, FA (2019). From QTLs to Adaptation Landscapes: Using Genotype-To-Phenotype Models to Characterize G×E Over Time. Frontiers in Plant Science , 10 https: / / doi.org / 10.3389 / fpls.2019.01540;

[0059] He, Y., Xiong, W., Hu, P., Huang, D., Feurtado, JA, Zhang, T.,Hao, C., DePauw, R., Zheng, B., Hoogenboom, G., Dixon, LE, Wang, H., & Challinor, AJ (2024). Climate change enhances stability of wheat-flowering-date. Science of The Total Environment , 917 , 170305; https: / / doi.org / 10.1016 / j.scitotenv.2024.170305

[0060] Hu, P., Chapman, SC, Sukumaran, S., Reynolds, M.,&Zheng, B. (2022). Phenological optimization of late reproductive phase for raising wheat yield potential in irrigated mega-environments. Journal of ExperimentalBotany , 73 (12), 4236–4249. https: / / doi.org / 10.1093 / jxb / erac144;

[0061] Kadam, N. N., Jagadish, S. V. K., Struik, P. C., van der Linden, C.G.,&Yin, X. (2019). Incorporating genome-wide association into eco-physiological simulation to identify markers for improving riceyields. Journal of Experimental Botany , 70 (9), 2575–2586. https: / / doi.org / 10.1093 / jxb / erz120;

[0062] Tang, C., Wang, C., Zhang, Z., Cao, Y., Bulut, M., Xiao, Y., Li, X.,Xiong, T., Yan, J.,&Guo, T. (2025). Redefining agroecological zones in Chinato mitigate climate change impacts on maize production. Molecular Plant , 18 (11), 1799–1802. https: / / doi.org / 10.1016 / j.molp.2025.09.005;

[0063] Wang, W., Guo, W., Le, L., Yu, J., Wu, Y., Li, D., Wang, Y., Wang,H., Lu, X., Qiao, H., Gu, X., Tian, J., Zhang, C.,&Pu, L. (2023). Integrationof high-throughput phenotyping, GWAS, and predictive models reveals thegenetic architecture of plant height in maize. Molecular Plant ,16 (2), 354–373.https: / / doi.org / 10.1016 / j.molp.2022.11.016。

Claims

1. A method for constructing a whole growth period simulation model of maize based on genetic information, characterized by, Includes the following steps: (1) Obtain whole genome resequencing data of the target maize population, as well as phenotypic data of the growth period, including seedling emergence, tasseling and flowering, recorded in the same ecological location and in multiple consecutive years, and obtain meteorological data and field management information of the corresponding locations; (2) Based on the whole genome resequencing data of the target maize population, the genetic structure of the maize population is analyzed, and the target maize population is divided into several genetic subpopulations with consistent genetic background; the analysis of the genetic structure of the maize population is carried out by using the ADMIXTURE population structure analysis method, which maximizes the correlation of genetic structure between populations and divides the maize population into multiple genetic subpopulations. (3) Based on the APSIM crop model, the reproductive period model parameters are inverted for each maize genotype in the population to obtain the initial parameter vector corresponding to each maize genotype. The reproductive period model parameters include shoot_lag, shoot_rate, tt_emerg_to_endjuv, tt_endjuv_to_init, photoperiod_crit1, photoperiod_crit2 and photoperiod_slope. The initial default values ​​of the above parameters are shown in Table 1. The reproductive period model parameter inversion for each maize genotype in the population adopts the differential evolution algorithm, with the optimization objective being to minimize the sum of squared errors between the simulated and measured values ​​of the reproductive period of the genotype under multiple environmental conditions. The specific process includes: randomly initializing multiple candidate parameter individuals in the parameter space; continuously updating the population through mutation, crossover and selection operations; iteratively searching to gradually converge the objective function value to the minimum; when the objective function converges or reaches the maximum number of iterations, outputting the optimal parameter vector θg of the family; and using high-performance parallel computing, performing the parameter inversion for 1223 maize genotypes. Each family independently completed parameter inversion, ultimately obtaining 1223 sets of initial parameter vectors; (4) Based on the genetic subgroup, the initial parameter vectors of multiple maize genotypes belonging to the same genetic subgroup are aggregated with statistical mean, variance and standard deviation to construct the parameter constraint template corresponding to the genetic subgroup; (5) Under the constraint of the parameter constraint template, the growth period model parameters of each maize genotype are optimized and calibrated twice to simultaneously reduce the deviation between the simulated and measured values ​​of the growth period of the genotype under multiple environmental conditions and the deviation of the genotype parameters from the parameter constraint template of its genetic subgroup, so as to obtain the final growth period model parameters corresponding to each maize genotype. (6) The APSIM crop model is driven by the model parameters of the final growth period corresponding to each maize genotype to realize the prediction of the growth period of each maize genotype under multiple environmental conditions; the input data of the APSIM crop model includes daily maximum temperature, daily minimum temperature, rainfall and solar radiation.

2. The method of claim 1, wherein, The construction of the parameter constraint template corresponding to the genetic subgroup refers to, for the genetic subgroup described in claim 1, denoting the set of all family parameter vectors belonging to the subgroup as {θ_g| g ∈ Ck}, calculating the mean of the statistics of the set on each parameter dimension, and constructing the parameter constraint template vector Template_{Ck} for the subgroup:

3. The method of claim 1, wherein, The secondary optimization calibration is performed by constructing a constraint objective function that includes a fitting error term and a parameter deviation penalty term. The fitting error term characterizes the difference between the simulated and measured reproductive periods of the genotype under different environmental and reproductive stage conditions. The parameter deviation penalty term constrains the degree of deviation between the genotype parameter vector to be optimized and the parameter constraint template of its genetic subpopulation. By optimizing the constraint objective function, the final secondary optimized model parameters θg* are obtained.

Citation Information

Patent Citations

  • Quantitative prediction method, system and device based on crop growth period phenotype and regional adaptability thereof

    CN118171785A

  • Corn genetic performance prediction method fusing multiple environmental factors

    CN120954498A