Methods and systems to enhance seed production
By employing learned prediction models to synchronize flowering in hybrid seed production through genotype and environment interaction analysis, the method enhances yield and reduces costs by optimizing planting dates and management, addressing asynchronous flowering issues.
Patent Information
- Application Number
- PCT/US2025/031176
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-28
- Publication Date
- 2025-12-04
AI Technical Summary
Current hybrid seed production methods face challenges in ensuring synchrony of flowering between male and female inbred parents due to environmental variation, leading to asynchronous flowering and reduced yield and quality, as well as high production costs.
Utilizing learned prediction models trained on genotype by environment interactions to predict phenotypes for each parent, determining optimal planting dates and management decisions based on location-specific data, and updating these models dynamically during the growing season to enhance synchronization and yield.
The method increases hybrid seed yield by at least 1 bu/ac to 50 bu/ac compared to traditional methods, improving yield and reducing costs by optimizing planting dates and field management.
Smart Images

Figure US2025031176_04122025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS TO ENHANCE SEED PRODUCTIONBACKGROUND
[0001] The production of commercially relevant quantities of hybrid seed often involves planting of a hybrid seed production field with rows comprising seed of a male inbred parent of a hybrid plant and rows comprising seed of a female inbred parent of the hybrid plant.
[0002] One challenge of hybrid seed production is to ensure synchrony of flowering of the two inbred parents. Most current hybrid seed production field planting methods rely on using an average growing degree unit for each of the inbred parents determined over several locations. Thus, environmental variation at specific locations may result in asynchronous flowering which can affect seed production yield and quality.
[0003] Accordingly, there is a need to develop new hybrid seed production methods to boost yield and seed quality and reduce hybrid seed production cost.SUMMARY
[0004] Provided are methods for increasing yield in a hybrid seed production field comprising inputting genotypic information from a female parent of a hybrid plant and location-specific environmental data for a hybrid seed production field into a first learned prediction model, the first learned prediction model trained to learn a relationship of genotype (that includes genetic variations introduced by breeding, site-specific genome editing, transgenes, and a combination thereof) by environment interactions or genotype by environment by management interactions to predict a phenotype for at least one crop growth event of the female parent in the hybrid seed production field, inputting genotypic information from a male parent of the hybrid plant and the location-specific environmental data for the hybrid seed production field into a second learned prediction model, the second learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for the at least one crop growth event for the male parent in the hybrid seed production field, generating by the first learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the female parent, generating by the second learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the male parent, and determining,based on the prediction of the phenotype for the at least one crop growth event for the female parent and the male parent in the hybrid seed production field, an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, or any combination thereof. In certain embodiments, the method further comprises planting the female parent and male parent in the hybrid seed production field using the determined optimal planting dates. In certain embodiments, the method further comprises (i) updating the location-specific environmental data in the first learned prediction model, the second learned prediction model, or both at least one time during the growing season, (ii) generating by the first learned prediction model, the second learned prediction model, or both comprising the updated location-specific environmental data a prediction of the phenotype for at least one crop growth event in the hybrid seed production field for the female parent, the male parent, or both, and (iii) determining, based on the prediction of the phenotype for the at least one crop growth event for the female parent, the male parent, or both a management decision.
[0005] Also provided is a method for increasing yield in a hybrid seed production field comprising generating a dynamic crop growth model (CGM) for a first parent of a hybrid plant, the dynamic CGM generated by inputting genotypic information from the first parent of the hybrid plant and location-specific environmental data for a hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent in the hybrid seed production field, generating a dynamic CGM for a second parent of the hybrid plant, the dynamic CGM generated by inputting genotypic information from the second parent of the hybrid plant and location-specific environmental data for the hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the second parent in the hybrid seed production field, determining an optimal planting date for the first parent and an optimal planting date for the second parent in the hybrid seed production field based on the dynamic CGM generated for the first parent and the second parent of the hybrid plant, planting the first parent on the determined optimal planting date and planting the second parent on the determined optimal planting date, updating the dynamic CGM for the first parent, the second parent or both the first and second parent at least one time by inputting observed location-specific environmental data into the learned prediction model, and determining at least one field or crop management decision based on the updated dynamic CGM for the first parent, the updated CGM for the second parent or the updated CGM for both the first and second parent.
[0006] Further provided is a method of allocating fields for hybrid seed production comprising (a) inputting genotypic information from a first parent and a second parent for each of a plurality of hybrid products and location-specific environmental data for a plurality of hybrid seed production fields into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields and (b) determining the optimal hybrid seed production field for at least one of the plurality of hybrid products based on the simulated plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields. In certain embodiments, after (a) and before (b) the method further comprises inputting the simulated plant growth for the first parent and the second parent for one or more of the plurality of hybrid products in the plurality of hybrid seed production fields into at least one additional learned prediction model, the at least one additional learned prediction model trained to learn a relationship between hybrid production performance in an environment and production costs and wherein the optimal hybrid seed production field for the at least one of the plurality of hybrid products is based on both the simulated plant growth and the output of the at least one additional learned prediction model.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The disclosure can be more fully understood from the following detailed description and the accompanying drawings.
[0008] Fig. 1 provides a schematic providing a description of the CGM-WGP infrastructure to predict phenotypes for in-field in-season crop growth development.
[0009] Figs. 2A and 2B provide graphs depicting the difference between the days-to-silk and days-to-shed (simulated and observed values are represented by a black dot and a grey triangle, respectively) for two different hybrids (Hybl (Fig. 2A), Hyb2 (Fig. 2B)) in three fieldsconducted at different years. The vertical black line represents the mean of the simulated difference across the three fields.
[0010] Fig. 3 provides a schematic providing a description of the Maize Flowering Synchrony crop growth model. The three physiological parameters (tin, coblf, and ebRl determined the two phenotypes’ days-to-silk and days-to-shed. The scheme at the top represents the relationship between leaf appearance and thermal time. The scheme at the bottom represents the ear biomass development over time.
[0011] Fig. 4 provides graphs depicting the difference genotype specific simulated days-to-silk and simulated days-to-shed for a set of 50 female and male inbreds (x-axis and y-axis ) based on 2018, 2020 and 2021 weather data. Female and male inbreds are ranked by increasing value of days-to-silk and days-to-shed using the location 2018 and the order is kept the same for 2020 and 2021. The hybrids Hybl and Hyb2 are represented by the small squares.
[0012] Fig. 5 is a block diagram illustrating an exemplary computer system including a server and a computing device according to an embodiment as disclosed herein.DETAILED DESCRIPTION
[0013] The present disclosure provides methods for increasing yield in a seed production field, such as, for example, a hybrid seed yield in a seed production field. The methods involve using learned prediction models trained to learn a relationship between genotype by environment interactions or genotype by environment by management interactions to predict one or more phenotypes for plants in a seed production field. The methods also involve using the predicted phenotypes to determine planting dates in the field and / or field management decisions.
[0014] As used herein, a “seed production field” “production field” and the like refers to a field used to produce a commercially relevant quantity of seed. In certain embodiments, the field is a hybrid seed production field, also referred to herein as a hybrid production field, in which a commercially relevant quantity of hybrid seed is produced. In certain embodiments, the field is a variety seed production field, also referred to herein as a variety production field, in which a commercially relevant quantity of variety seed (e.g., soybean) is produced. As used herein, a commercially relevant quantity refers to an amount of seed that is needed for a variety or hybrid to be sold for commercial use. The amount of seed considered to be a commercially relevant amount is not particularly limited and may vary based upon the phenotypic characteristics of theseed or plant grown therefrom and economic factors. In certain embodiments of the methods described herein, the commercially relevant quantity of seed comprises at least 10, 25, 50, 75, 100, 250, 500, 1000, 5000, 10,000, or 100,000 units. A unit of seed, as used herein, refers to about 80,000 maize seeds, 150,000 sunflower seeds, 140,000 soybean seeds, 50 pounds of seed sorghum and canola.
[0015] As used herein, “increased yield” “increasing yield” or the like refers to any detectable gain in the amount of seed harvested per unit of land and may include reference to bushels per acre (bu / ac) or kilograms per hectare of a crop at harvest, as adjusted for seed moisture, when compared to an appropriate control (e.g., field planted using BLUP models). The increased yield of the methods described herein is not particularly limited. In certain embodiments of the methods described herein, the yield of a field or average yield across multiple fields planted using the prediction models described herein is increased by at least 1 bu / ac, 2 bu / ac, 3 bu / ac, 4 bu / ac, 5 bu / ac, 10 bu / ac, 15 bu / ac, 20 bu / ac, 25 bu / ac, or 50 bu / ac as compared to an appropriate control (e.g., yield of field planted using BLUP models or average yield across multiple fields using BLUP models).
[0016] Accordingly, one aspect of the disclosure provides methods for increasing yield in a hybrid seed production field comprising inputting genotypic information from a female parent of a hybrid plant and location-specific environmental data for a hybrid seed production field into a first learned prediction model, the first learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for at least one crop growth event of the female parent in the hybrid seed production field, inputting genotypic information from a male parent of the hybrid plant and the location-specific environmental data for the hybrid seed production field into a second learned prediction model, the second learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for the at least one crop growth event for the male parent in the hybrid seed production field, generating by the first learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the female parent, generating by the second learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the male parent, and determining, based on the prediction of the phenotype for the at least one crop growth event for the femaleparent and the male parent in the hybrid seed production field, an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, or any combination thereof. In certain embodiments, the first learned prediction model and the second learned prediction model are the same model. In certain embodiments, the first learned prediction model and the second learned prediction model are two different learned prediction models (e.g., trained using different data sets or different statistical models). In certain embodiments, the method further comprises planting the female parent and male parent in the hybrid seed production field using the determined optimal planting dates, thereby increasing hybrid seed yield as compared to a control (e.g., yield from a field in which the first parent and second parent are planted using a GDU average across all locations). In certain embodiments, the method further comprises (i) updating the location-specific environmental data in the first learned prediction model, the second learned prediction model, or both at least one time during the growing season, (ii) generating by the first learned prediction model, the second learned prediction model, or both comprising the updated location-specific environmental data a prediction of the phenotype for at least one crop growth event in the hybrid seed production field for the female parent, the male parent, or both, and (iii) determining, based on the prediction of the phenotype for the at least one crop growth event for the female parent, the male parent, or both a management decision.
[0017] The setup of the hybrid seed production field for use in the methods described herein is not particularly limited and may be determined based on, for example, the hybrid to be produced or field conditions. For example, the planting of a hybrid seed production field often involves the planting of rows comprising seed of the male parent of the hybrid plant and rows comprising seed of the female parent of the hybrid plant. Recent advances, however, allow for planting of seed of the female parent of the hybrid plant and artificially pollinating the female parent plant grown from the female seed with pollen captured from the male plant. In certain embodiments, the hybrid seed production field is planted in a pattern comprising a row of seed from the male parent followed by 2, 3, 4, 5, or 6 rows of seed from the female parent.
[0018] As used herein a “learned prediction model” refers to a model that has been trained using data (e.g., historical data) to predict likely outcomes. A learned prediction model is usually not a fixed model and is validated or revised regularly to incorporate changes in the underlying data. The type of learned prediction model for use in the methods described herein is not particularlylimited and may be any type of learned prediction model that can be trained to learn a relationship of genotype by environment interactions and / or genotype by environment by management interactions to predict a phenotype for at least one crop growth event or to generate a crop growth model (e.g., dynamic crop growth model). Similarly, the method of training the learned prediction model is not particularly limited and may be any method known in the art, or described herein, to train a model to learn a relationship of genotype by environment interactions and / or genotype by environment by management interactions to predict a phenotype.
[0019] In certain embodiments, the learned prediction model comprises a machine learning model. Any suitable machine learning model may be used in the in the methods and systems described herein. Types of models include without limitation statistical models, such as probability models, regression models, and those involving deep learning, such as supervised, self-supervised, and unsupervised models, or combinations thereof. In certain embodiments, the machine learning model is a classification model, a regression model, a clustering model, a dimensionality reduction model, a distribution model, for example, a multivariate or univariate Gaussian distribution model, or a deep learning model. In certain embodiments, the deep learning model is part of an ensemble model. In certain embodiments, the deep learning model is an ensemble model comprising two or more models. In certain embodiments, the deep learning model is a supervised learning model. The supervised learning model may be a classification or regression model. The machine learning models include support vector machines, neural networks, such as SVM-DA (Support Vector machines) or ANN (Artificial Neural Networks), or deep learning algorithms and the like.
[0020] In certain embodiments, the machine learning model for use in the methods described herein is an artificial neural network (ANN). ANNs are configured to synthesize or learn from a plurality of inputs to produce an output. One or more variables in the algorithms can have weights that are applied to each equation and optimized as the neural network is trained. Based on the amount of training information, the deep learning models or networks get better at producing more helpful outputs. In certain embodiments, the ANN includes a plurality of input factors that may be used to train predicted phenotypic information. These factors include, but are not limited to, QTLs, SNPs, haplotypes, yield, and other historical agronomic or breeding phenotypic components.
[0021] In certain embodiments, the learned prediction model comprises a linear statistical model. As used herein, a “linear statistical model” is a model that provides a linear relationship between an independent variable (e.g., genotype) and a dependent variable (e.g., phenotype) to predict the outcome of future events. The type of linear statistical model is not particularly limited and may be any linear statistical model known in the art. In certain embodiments, the linear statistical model comprises a logistic generalized linear prediction model.
[0022] In certain embodiments, the machine learning model for use in the methods described herein is a Bayesian genomic prediction model (e g., BayesA, BayesB, or BayesC). In certain embodiments, the training dataset for the Bayesian genomic prediction model comprises phenotypes at multiple locations for the trait or traits of interest and marker genotype scores from a plurality of material representative of the germplasm of interest. In certain embodiments, the training dataset is used to estimate the marker effects for the trait of interest in the population at a given environment using the prediction model BayesA, BayesB, and / or BayesC. As used herein, a "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a plant or particular plant material or cell. In some instances, this characteristic is visible to the human eye, such as seed or plant size, or can be measured by biochemical techniques, such as detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process, e.g. by measuring uptake of carbon dioxide, or by the observation of the expression level of a gene or genes, e.g., by employing Northern blot analysis, RT-PCR, microarray gene expression assays, or reporter gene expression systems, or by agricultural observations such as stress tolerance, yield, or pathogen tolerance.
[0023] In certain embodiments, the learned prediction model comprises a crop growth model (CGM). As used herein a “crop growth model (CGM)” refers to a model used to predict phenotypes for a plant across multiple environmental conditions by simulating the growth of a plant in a given environment. One type of CGM referred to herein as a “dynamic crop growth model” refers to a crop growth model in which the data in the model is updated during the plant growth and the simulated growth of the plant is adjusted based on the updated data. For example, in a dynamic crop growth model the environmental conditions may be updated to include the actual weather data and / or forecasted weather data after planting such that the simulated growth uses the updated data.
[0024] In certain embodiments of the CGMs described herein, the CGM or dynamic CGM comprises the use of whole genome prediction (WGP) such that the CGM or dynamic CGM for use in the methods described herein comprises a WGP-CGM. As used herein “genome prediction”, “whole genome prediction” and the like refers to prediction methods in which polymorphic markers densely distributed over the genome are used regardless of their association with the trait of interest. Examples of whole genome prediction can be found in, for example, Meuwissen et al. (Genetics, 157(4): 1819-1829 (2001)) and de los Campos (Genetics, 193(2): 327-345 (2013)). The GEBVs are often calculated but not limited to the sum of all additive effects estimated with a genomic prediction model. As used herein, “genomic estimated breeding values” (GEBVs) refer to a measurable degree to which one or more, polymorphic markers, haplotypes and / or genotypes heritably affect the expression of a phenotype associated with a trait. In certain embodiments of the WGP-CGM described herein, the WGP predicts physiological parameters related to the crop growth event and the CGM uses the predicted physiological parameters to simulate the growth phenotype. For example, The WGP prediction in certain embodiments uses WGP to predict total leaf number, minimum ear biomass, or leaf appearance rate, or any combination thereof and the prediction of the trait is used in the CGM to predict flowering time.
[0025] For example, provided herein is a method for increasing yield in a hybrid seed production field comprising generating a dynamic crop growth model (CGM) (e.g., dynamic WGP-CGM) for a first parent of a hybrid plant, the dynamic CGM generated by inputting genotypic information from the first parent of the hybrid plant and location-specific environmental data for a hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent in the hybrid seed production field, generating a dynamic CGM for a second parent of the hybrid plant, the dynamic CGM generated by inputting genotypic information from the second parent of the hybrid plant and location-specific environmental data for the hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the second parent in the hybrid seed production field, determining an optimal planting date for the first parent and an optimal planting date for the second parent in the hybridseed production field based on the dynamic CGM generated for the first parent and the second parent of the hybrid plant. In certain embodiments, the method further comprises planting the first parent on the determined optimal planting date and planting the second parent on the determined optimal planting date, updating the dynamic CGM for the first parent, the second parent or both the first and second parent at least one time during a growing season by inputting observed, forecasted, or observed and forecasted location-specific environmental data into the learned prediction model, and determining at least one field or crop management decision based on the updated dynamic CGM for the first parent, the updated CGM for the second parent or the updated CGM for both the first and second parent.
[0026] As used herein, a growing season refers to the time between planting the seed to harvesting the resulting seed from a plant grown from the planted seed. In certain embodiments, the dynamic CGM is updated on a selected day after planting and / or a selected day before a predicted crop growth event. For example, in certain embodiments, the dynamic crop growth model is updated 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, or 150 days after planting and / or 5, 10, 15, 20, 25, 30, or 35 days before a predicted crop growth event such as, for example, flowering time, silking, shedding, or harvest. In certain embodiments, the dynamic CGM is updated daily, every 2 days, every 3 days, every 4 days, every 5 days, every 6 days, weekly, every two weeks, every three weeks, monthly, every 2 months, every 3 months, every 4 months, every 5 months, or every 6 months.
[0027] "Genomic information" or “genotypic profile” as used herein generally refers to a set of information about the entire genome of a given plant or group of plants (genome-wide), or it can encompass a specific subset of the genome of a given plant or group of plants, or any combination thereof in a given plant or group (e.g., population) of plants. In certain embodiments of the methods described herein, the genotypic information (e.g., genotypic information of a first or a second parent) includes information regarding the presence or absence in the genome of a specific set of mutations, single nucleotide polymorphisms (SNPs), insertion of bases, deletion of bases, genotypic markers, other sequence information, or any combination thereof.
[0028] The method of genotyping the plants in the methods described herein is not particularly limited and may be any method known in the art, or described herein, that can determine the genotype of the plant at one or more positions within the genome of the plant. In certain embodiments, plants are genotyped using molecular biological assays such as, for example, PCR,DNA sequencing, restriction fragment length polymorphism identification or whole genome sequencing. The plants can be genotyped individually or from a pooled DNA sample.
[0029] Location-specific environmental data of the methods described herein includes, but is not limited to, information for or relating to geographical locations such as latitude and longitude information, land features such as elevation, site topography, historical and forecasted climate conditions e.g. historical weather conditions and forecast weather conditions, including but not limited to wind direction, wind velocity, cloud cover, humidity, relative humidity, sunrise, sunset, temperature, precipitation, water vapor, vapor pressure deficit, snow depth, barometric pressure, season, heat index, visibility, dew point, air quality, storms, solar radiation. In certain embodiments, environmental data may be obtained from remotely-sensed imagery, including visual light, infrared, near-infrared, multi -spectral, and hyperspectral imaging bands, in some aspects, these bands may be combined into vegetative indices, including but not limited to the normalized difference vegetation index (NDVI), enhanced vegetation index (EVI), weighted difference vegetation index (WDVI), or normalized difference water index (NDWI); soil type or soil substrate type e.g. sand, loam, clay, soil conditions e.g. aeration level, temperature, (ground) water level, soil moisture, humidity level, pH, composition such as organic matter, degree of compaction, ground soil organic carbon estimates or capacity, soil toxicities, soil nutrients, inputs or applied products such as fertilizers, herbicides, insecticides, seed treatments, seed- or soil-applied agricultural biologicals soil drainage; (crop) plant conditions e.g. plant population density, planting date, nutrient application, plant height, evapotranspiration rate, gross primary productivity (GPP), growing and harvesting season, growth cycle, stage of development, such as plant phenological stage - including flowering and grain filling, green-up, dry down, and senescence, seed type, vegetation, crop variety, chemical, physical and nutritional requirements; biotic stresses such as plant disease resistance level, including but not limited to Northern Leaf Blight (NLB) and Goss's Wilt (GOSWLT, plant herbicide tolerance level, physical injuries e.g. from pathogens, herbicides, storms, plant stress; abiotic stresses such as early and late root lodging, stalk lodging, brittlesnap, willowing, management decisions such as row count, irrigation, irrigation location, and tillage; previous and / or subsequent crops in a crop rotation system, crop production methods, e.g. crops grown in the open field, in a growth chamber, or in a greenhouse; disease and pest events, weeds, or any combinations thereof.
[0030] In certain embodiments, the location-specific environmental data comprises one or more of soil type, historical temperature such as, for example, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 year average; and forecast temperature and / or moisture, such as, for example, 3-day, 5-day, 6-day, 7-day, 8-day, 9-day, 10-day, or 15-day forecast. The location-specific environmental data may be taken at a regional level, a town level, a field level, a sub-field level, a plot level, a row level, or any combination thereof.
[0031] The location-specific environmental data may be obtained in any suitable manner or using any technique, including without limitation imaging devices, cameras, and sensors. The data may be from any number of sources, including but not limited to an aerial source, including but not limited to one or more satellites, including high resolution satellites, airplanes, helicopters, balloons and UAV platforms; a ground-based source including but not limited to one or more trucks, tractors, rovers, or other vehicles or objects such as weather stations that are land-bound, or a mobile-source including but not limited to one or more hand-held or mobile devices, or sources of data that do not fall into any of the other categories.
[0032] The location-specific environmental data used in the methods and systems herein may be obtained from one or more sources such as an external database, an internal database, a private data source, a public data source, or combinations thereof.
[0033] Non-limiting examples include public weather information from public databases and sources such as the National Oceanic and Atmospheric (NOAA), National Weather Service (NWS), Meteorological Simulation Data Ingest System (MADIS), Esri Open Data Portal, National Aeronautics and Space Administration (NASA), European Space Agency, Natural Resources Conservation Service (NRCS-SSURGO), United States Geological Survey (USGS), and Parameter-elevation Regressions on Independent Slopes Model (PRISM); public soil information from public databases and sources such as the United States Department of Agriculture (USDA) SSURGO soil survey database.
[0034] The location-specific environmental data may be obtained using an environmental monitoring device. In certain embodiments, the environmental monitoring device may comprise a mechanized and / or electronic device that may be configured to determine the weather conditions at a location. For example, the environmental monitoring device may comprise one or more sensors (e.g., thermometers, barometers, humidity sensors, rain gauges, anemometers, and / or aerial cameras) that may be used to determine environmental conditions comprisingtemperature (e.g., air temperature at a location of a crop), humidity, radiation, barometric pressure, rainfall, wind direction, and / or wind speed. Further, the environmental monitoring device may be operated by and / or receive information from a third-party entity that generates environmental data for one or more locations including a location of a crop that is being monitored. The environmental monitoring device may include any of the features and / or components of the computing system. For example, an environmental monitoring device may include one or more processors, a memory, one or more input devices, and / or one or more output devices. Further, the environmental monitoring device may have different or similar configurations and / or architectures to that of computing system.
[0035] As used herein, a “crop growth event” or the like refers stages of plant growth from planting to harvesting which can often be determined by analyzing changes in the plant phenotype during growth. Examples of crop growth events of the methods described herein include, but are not limited to, emergence, silking, shedding, flowering, and physiological maturity. In certain embodiments of the methods described herein, the crop growth event is flowering time. In certain embodiments, the flowering time of the plant is predicted using a prediction of the estimated leaf appearance rate, total leaf number, minimum ear biomass, or any combination thereof. In certain embodiments, the crop growth event for the female parent is silking and the crop growth event for the male parent is shedding.
[0036] The optimal planting date is not particularly limited and may be selected based on a prediction of a desired outcome such as, for example, maximum yield, lowest cost of inputs, predicted harvest date, or combinations of multiple desired incomes. In certain embodiments of the methods described herein, the optimal planting date is the date in which the planting of the first parent and planting of the second parent are predicted to give the maximum hybrid seed yield. In certain embodiments, to maximize hybrid seed yield the female parent and male parent optimal planting dates comprise a planting date in which the female parent is predicted to silk and the male parent is predicted to shed pollen within at least 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 days of each other. In certain embodiments, the optimal planting date of the female parent and the male parent occur on the same day. In certain embodiments, the optimal planting date of the female parent and the male parent occur on different days.
[0037] As used herein a “management decision”, “field management decision”, “crop management decision” or the like refers to an intervention in the seed production field to alter thegrowth of one or more plants in the field. Non-limiting examples of management decisions of the methods described herein include herbicide application, irrigation, defoliation application, pesticide application, detasseling, and filed monitoring. In certain embodiments, the management decision comprises one or more of herbicide application, irrigation, defoliation application, detasseling, or any combination thereof.
[0038] Often hybrid seed production is conducted using non-stress field conditions such that in certain embodiments of the methods described herein the prediction models are trained to predict growth under non-stress conditions. However, in certain instances hybrid seed production may occur in a field with abiotic and / or biotic stress conditions, such as, for example, nitrogen stress, water stress, insect stress, or weed stress. Thus, in certain embodiments of the methods described herein the prediction models are trained to predict growth under one or more abiotic and / or biotic stress conditions. In certain embodiments of the methods described herein the prediction models are trained to predict growth under non-stress conditions and one or more abiotic and / or biotic stress conditions. In embodiments in which hybrid seed production occurs in fields with stress conditions, the management decisions may further comprise an intervention to reduce the biotic and / or abiotic stress such as, for example, nitrogen application or water application.
[0039] Further provided is a method of allocating seed production fields for hybrid seed production comprising inputting genotypic information from a first parent and a second parent for each of a plurality of hybrid products and location-specific environmental data for a plurality of hybrid seed production fields into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields, determining the optimal hybrid seed production field for at least one of the plurality of hybrid products based on the simulated plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields. In certain embodiments, before determining the optimal hybrid seed production field for at least one of the plurality of hybrid products the method further comprises inputting the simulated plant growth information into at least one additional learned prediction model. In certain embodiments, the at least one additional learned prediction model predicts one or more of transportation costs such as, for example, cost of transporting harvested seed from ahybrid seed production field and / or cost of supplying inbred parent seed for hybrid production to a hybrid seed production field, predicted hybrid seed harvest date, and matching plant conditioning capacities.
[0040] In certain embodiments, the at least one additional learned prediction model predicts transportation costs using a freight forecasting model. In certain embodiments, the freight forecasting model optimizes logistics and transportation planning by accurately forecasting freight needs and costs. Using historical data regarding shipment volumes, routes, and seasonal trends as inputs, the model forecasts future freight costs.
[0041] The plant of the methods described herein is not particularly limited and may be any plant for which seed production (e.g., commercial seed production) is desired. Examples of plant species of interest include, but are not limited to, maize (Zea mays), wheat (Triticum aestivum), sunflower (Helianthus annuus), soybean (Glycine max), Brassica sp. (e.g., B. napns, B. rapa, B. juncea), particularly those Brassica species useful as sources of seed oil, sorghum (Sorghum bicolor, Sorghum vulgare), potato (Solanum tuberosum), pea (Lathyrus spp), and cotton (Gossypium barbadense, Gossypium hirsutum). In certain embodiments, the plant is a hybrid plant, such as for example, hybrid maize, hybrid wheat, hybrid Brassica sp., or hybrid sorghum.
[0042] In certain embodiments, the female parent of the hybrid plant, the male parent of the hybrid plant, or both the male parent and female parent comprise a transgene. In certain embodiments, the female parent of the hybrid plant, the male parent of the hybrid plant, or both the male parent and female parent comprise a site-specific genome edit, such as those introduced in a genome edited breeding program, which may include 10’s to 100’s of site-specific genomic modifications. In certain embodiments, the female parent of the hybrid plant, the male parent of the hybrid plant, or both the male parent and female parent comprise a transgene and a genome edit. Similarly, when the plant for increased seed production of the methods described herein is a variety plant, the plant may comprise a transgene and / or a genome edit.
[0043] The method for training the learned prediction model for use in the methods described herein is not particularly limited and may be trained using any method known in the art to train a model to learn a relationship of genotype by environment interactions and / or genotype by environment by management interactions to predict a phenotype or crop growth. In certain embodiments, the learned prediction model is trained by growing a population of training individuals in a plurality of environments, phenotyping the population of training individuals inthe plurality of environments to generate a phenotype by and across environment(s) training data set(s); and associating the phenotype by and across environment s) training data set(s) with a genotype training data set comprising genetic information across the genome of each training individual, and subsequently, by using a biological model, estimating effects of genotypic markers and linking the estimation of effects of genotypic markers with the biological model to generate an association training data set that can be then used for on demand predictions of component traits and simulations of complex traits.
[0044] The number of training individuals is not particularly limited and should be reflective of the genotypic diversity of the population of interest (e.g., female maize inbred population or male maize inbred population). In certain embodiments, the training individuals comprise plants (e.g., inbred plants) developed from a traditional breeding program. As would be understood by a person of ordinary skill in the art, learned prediction models developed solely with plants from a traditional breeding program may have decreased accuracy when the female parent and / or the male parent of the hybrid plant comprises a genome edit (e.g., developed from a genome edited breeding program) as compared to the prediction for male or female parents developed by traditional breeding. In other words, certain genome edits may not be found in traditional breeding material making the genotypic effect difficult to detect. Accordingly, in certain embodiments the training individuals comprise plants (e.g., inbred plants) developed from a genome edited breeding program or plants (e.g., inbred plants) developed from both a traditional breeding program and a genome edited breeding program.
[0045] Traditional breeding refers to methods to develop new plants with desirable traits by choosing plants with favorable characteristics and crossing them to produce offspring with improved traits. Plants (e.g., inbreds or varieties) developed by traditional breeding for the methods provided herein include non-modified plants and plants comprising transgenes, such as those introduced by trait introgression or transformation of the plant. Genome edited breeding refers to methods to develop new plants with desirable traits that uses genome editing technologies, such as, for example CRISPR / Cas, TALENs, or ZFNs, to precisely modify the DNA of an organism to produce plants with improved traits. Plants (e.g., inbreds or varieties) developed by genome edited breeding for the methods provided herein comprise at least one genome edit. Plants containing at least one genome edit are created through trait integration or through direct editing of the plant material.
[0046] The methods described herein for increasing yield in a hybrid seed production field can be adapted to increase yield in a varietal (e.g., soybean) seed production field using one or more of the learned prediction models described herein. Accordingly, provided is a method for increasing yield in a varietal (e.g., soybean) seed production field comprising inputting genotypic information varietal plant for which commercial seed production is desired and location-specific environmental data for a seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for at least one crop growth event of the varietal plant in the seed production field, generating by the learned prediction model a prediction of the phenotype for the at least one crop growth event in the seed production field for the varietal plant, and determining, based on the prediction of the phenotype for the at least one crop growth event for the varietal plant, an optimal planting date for the varietal plant, a management decision, or any combination thereof. In certain embodiments, the method further comprises planting the varietal plant in the seed production field using the determined optimal planting date, thereby increasing seed yield as compared to a control (e.g., yield from a field in which the varietal plant is planted using a GDU average across all locations). In certain embodiments, the method further comprises (i) updating the locationspecific environmental data in the learned prediction model at least one time during the growing season, (ii) generating by the learned prediction model comprising the updated location-specific environmental data a prediction of the phenotype for at least one crop growth event in the seed production field for the varietal plant, and (iii) determining, based on the prediction of the phenotype for the at least one crop growth event for the varietal plant a management decision.
[0047] Additionally provided is a method for increasing yield in a seed production field comprising generating a dynamic crop growth model (CGM) (e.g., dynamic WGP-CGM) for a varietal plant, the dynamic CGM generated by inputting genotypic information from the varietal plant and location-specific environmental data for a seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the varietal plant in the seed production field, determining an optimal planting date for the varietal plant in the seed production field based on the dynamic CGM generated for the varietal plant. In certain embodiments, the method further comprises planting the varietalplant on the determined optimal planting date, updating the dynamic CGM for the varietal plant, at least one time during a growing season by inputting observed, forecasted, or observed and forecasted location-specific environmental data into the learned prediction model, and determining at least one field or crop management decision based on the updated dynamic CGM for the varietal plant.
[0048] Further provided is a method of allocating seed production fields for varietal seed production comprising inputting genotypic information from a varietal plant for each of a plurality of varietal plant products and location-specific environmental data for a plurality of seed production fields into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the varietal plant for each of the plurality of varietal products in the plurality of seed production fields, determining the optimal seed production field for at least one of the plurality of varietal products based on the simulated plant growth for the varietal plant for each of the plurality of varietal plant products in the plurality of seed production fields. In certain embodiments, before determining the optimal seed production field for at least one of the plurality of varietal plant products the method further comprises inputting the simulated plant growth information into at least one additional learned prediction model. In certain embodiments, the at least one additional learned prediction model predicts one or more of transportation costs such as, for example, cost of transporting harvested seed from a hybrid seed production field and / or cost of supplying inbred parent seed for hybrid production to a hybrid seed production field, predicted hybrid seed harvest date, and matching plant conditioning capacities.
[0049] Also provided are computer readable mediums having stored thereon instructions that, when executed by a processor (or computing device), cause the processor to perform the steps of the methods described herein to provide an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, a production cost or any combination thereof. Further provided are computer readable mediums having stored thereon instructions that, when executed by a processor (or computing device), cause the processor to perform the steps of the methods described herein to determine the optimal hybrid seed production field for at least one of a plurality of hybrid products.
[0050] Further disclosed herein are systems (e.g., computer systems) for use in seed production, such as hybrid seed production that include (a) one or more servers, each of the one or more server storing plant data, and (b) a computing device communicatively coupled to the one or more servers, the computing device including: (1) a memory, and (2) one or more processors configured to perform operations to: (a) train a learned prediction model, (b) generating by a learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for a female parent from the genotypic information of the female parent, (c) generating by a learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for a male parent from the genotypic information of the male parent, and (d) determining, based on the prediction of the phenotype for the at least one crop growth event for the female parent and the male parent in the hybrid seed production field, an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, or any combination thereof. The one or more processors may also be configured to perform operations to train the learned prediction model by associating a phenotype by environment training data set comprising phenotypic data from a plurality of candidate plant genotypes grown in a plurality of environments with a genotype training data set comprising genetic information across the genome of each candidate plant genotype, using a biological model, estimating effects of genotypic markers and linking the estimation of effects of genotypic markers with the biological model to generate an association training data set.
[0051] The type of system (e.g., computer system) is not particularly limited and may be any system comprising a computing device and one or more servers such as the system provided in Fig. 5. Referring to Fig. 5, a block diagram of a computer system 100 to provide an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, a production cost, or any combination thereof using at least one learned prediction model is shown. To do so, the system 100 may include a computing device 110 and a server 130 that is associated with a computer system. The system 100 may further include one or more servers 140 that are associated with other computer systems such that the computing device 110 may communicate with different computer systems running different platforms. However, it should be appreciated that, in some embodiments, a single server (e.g., a server 130) may run multiple platforms. The computing device 110 is communicatively coupled to the one or moreservers 130, 140 via a network 150 (e.g., a local area network (LAN), a wide area network (WAN), a personal area network (PAN), the Internet, etc.).
[0052] In general, the computing device 110 may include any existing or future devices capable of training a machine learning model. For example, the computing device may be, but not limited to, a computer, a notebook, a laptop, a mobile device, a smartphone, a tablet, wearable, smart glasses, or any other suitable computing device that is capable of communicating with the server 130.
[0053] The computing device 110 includes a processor 112, a memory 114, an input / output (I / O) controller 116 (e.g., a network transceiver), a memory unit 118, and a database 120, all of which may be interconnected via one or more address / data bus. It should be appreciated that although only one processor 112 is shown, the computing device 110 may include multiple processors. Although the I / O controller 116 is shown as a single block, it should be appreciated that the I / O controller 116 may include a number of different types of VO components (e g., a display, a user interface (e.g., a display screen, a touchscreen, a keyboard), a speaker, and a microphone).
[0054] The processor 112 as disclosed herein may be any electronic device that is capable of processing data, for example a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a system on a chip (SoC), or any other suitable type of processor. It should be appreciated that the various operations of example methods described herein (i.e., performed by the computing device 110) may be performed by one or more processors 112. The memory 114 may be a random-access memory (RAM), read-only memory (ROM), a flash memory, or any other suitable type of memory that enables storage of data such as instruction codes that the processor 112 needs to access in order to implement any method as disclosed herein. It should be appreciated that, in some embodiments, the computing device 110 may be a computing device or a plurality of computing devices with distributed processing.
[0055] As used herein, the term “database” may refer to a single database or other structured data storage, or to a collection of two or more different databases or structured data storage components. In the illustrative embodiment, the database 120 is part of the computing device 110. In some embodiments, the computing device 110 may access the database 120 via a network such as network 150. The database 120 may store data (e.g., input, output, intermediary data) used for calculating sampling rates. For example, the data may include genotypic data, such as single nucleotide polymorphisms (SNPs), genetic markers, haplotype, sequence information,phenotypic data, environment data, production costs, pedigree information, co-ancestry information, or combinations thereof that are obtained from one or more servers 130, 140.
[0056] The computing device 110 may further include a number of software applications stored in a memory unit 118, which may be called a program memory. The various software applications on the computing device 110 may include specific programs, routines, or scripts for performing processing functions associated with the methods described herein. Additionally, or alternatively, the various software applications on the computing device 110 may include general-purpose software applications for data processing, database management, data analysis, network communication, web server operation, or other functions described herein or typically performed by a server. The various software applications may be executed on the same computer processor or on different computer processors. Additionally, or alternatively, the software applications may interact with various hardware modules that may be installed within or connected to the computing device 110. Such modules may implement part of or all of the various exemplary method functions discussed herein or other related embodiments.
[0057] Although only one computing device 110 is shown in Fig. 5, the server 130, 140 is capable of communicating with multiple computing devices similar to the computing device 110. Although not shown in Fig. 1, similar to the computing device 110, the server 130, 140 also includes a processor (e.g., a microprocessor, a microcontroller), a memory, and an input / output (I / O) controller (e.g., a network transceiver). The server 130, 140 may be a single server or a plurality of servers with distributed processing. The server 130, 140 may receive data from and / or transmit data to the computing device 110.
[0058] In certain embodiments, the computing device 110 may generate predictions of plant genotype performance by using at least one learned prediction model to generate a predicted phenotype for the at least one crop growth event for the male parent and female parent in a hybrid seed production field. In certain embodiments, the computing device 110 may further generate an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, a production cost, or any combination thereof.
[0059] The following are examples of specific embodiments of some aspects of the invention. The examples are offered for illustrative purposes only and are not intended to limit the scope of the invention in any way.EXAMPLE 1
[0060] This example demonstrates the development of a learned prediction model to predict a phenotype for a crop growth event.
[0061] To predict the crop growth event flowering time, the crop growth model named the Maize Flowering Synchrony model (hereafter called MFS) was developed. MFS is parametrized by three unobserved physiological parameters (z.e., coblf, tin, and ebRl) that determine the phenotypes days-to-silk and days-to-shed as a function of thermal time accumulation (Figure 2). MFS is based on an exponential relationship between leaf number (LN) and thermal time (GDD) where LN is equal to:
[0062] LN = 2.5 *eGDD*coblf
[0063] with 2.5 representing the number of leaves at emergence, and coblf equal to 0.00225 but considered as a genotype-specific parameter. After planting, 30.6 GDD (C) are required for maize to germinate. Air temperature data was retrieved with the R package nasapower.
[0064] After a specific thermal time accumulation ttv, the onset of ear growth is initiated at a developmental stage “Vn”. It is assumed that the ear leaf coincides with the position of the largest leaf which is related to the parameter total leaf number (tin) as follows:
[0065] Vn = tln*0.61
[0066] As an example, if a specific inbred has a total leaf number equal to 23, the ear growth will be initiated at the developmental stage “VI 5”.
[0067] To model female flowering (i.e., silking), an exponential growth pattern is used to calculate accumulated ear biomass. The ear grows exponentially with an ear growth rate (egr equal to egr =tthirepresenting the thermal time between the development tthi stage “Vn” to the beginning of grain filling as a function of tin. As a starting point of ear growth development, the ear biomass is considered equal to 0.01g and grows up to a maximum of 5g during the silking period.
[0068] Silking is a change of state at the individual plant level and when the ear biomass exceeds the value of ebRl, the reproductive stage R1 i.e., silking) is considered reached. The ear biomass at a specific thermal time tt is equal to 0.01 * ee9r*tt:
[0069] The MFS crop growth model was incorporated into a CGM-WGP sampling algorithm presented, as described in Fig. 1. It is important to note that the three physiological parameters tin, cobl and ebRl are genotype specific and are modeled in CGM-WGP as a linear function of observed biallelic single nucleotide polymorphism (SNP) markers representing the inbredgenotypes. Those physiological parameters are then connected to the observed flowering phenotypes by the MFS which serves as a link function.
[0070] In the context of the CGA / -WGP infrastructure, the MFS crop growth model can be used as a link function between the physiological parameters and the observed phenotypic data as represented by the general equation below (1):Pij - NCCGM^ Ej), ^ (1)is the simulated phenotype (e.g., days-to-silk and days-to-shed) of the genotype i in the field j. Ttiis the genotype-specific unobserved physiological trait t for the genotype z; Ej represents the environmental features for the field j andis the residual variance for the simulated phenotype in the field j. The prior on the genotype-specific physiological trait Tti(z.c., coefficient of leaf appearance coblf, total leaf number tin and ear biomass at silking stage R1 ebRE) is:
[0072] Where f>Otis the intercept of the genotype-specific trait Z; Ztikis the SNP score of the genotype z at marker k of trait t (same markers were used for all traits); utkis the effect of the marker k for trait t (same markers were used for all traits);is prior variance for the trait t. The prior distribution for / ?Ofis a Normal distribution with meanand a residual variance (TQPriors for jj.tkare Normal distributions with mean of 0 and marker effect variance of OMtkwhich is equivalent to those used in ‘BayesA’ WGP model (Meuwissen et al., 2001). The variances a and <7^tfeare associated with a scaled inverse chi-square prior distribution with 4.001 degrees of freedom and scaling factors St. Stfollows a gamma distributions with shape and rate parameters calculated according to Habier et al. (2011). This hierarchical model can be visualized in a tree form.
[0073] Metropolis-Hastings within Gibbs algorithm was implemented to sample from the posterior distribution of all the parameters (Gelman, 2004). Sampling chains were run for 50,000 iterations with the first 30,000 discarded as ‘burn-in’ and samples from every 20thpost burn-in iteration stored, thus resulting in 1000 stored samples. The algorithm was implemented as a C++ routine.
[0074] The posterior predictive distribution for a genotype z in field j was obtained by entering the physiological parameters Ttiinto the MFS model together with the inputs of the field j,resulting in one set of phenotypic values per posterior sample. The mean of this distribution is used as the predicted value. Note that the predicted values of the set of physiological parameters are used as a joint set in the MFS model to describe the relationship between thermal time and inbred development. Following the WGP step, the set of new physiological parameters sampled from the posterior distribution, if accepted, are passed as a joint set to the MFS model together with the management and weather information (Fig. 1) to simulate the genotype specific flowering phenotypes. However, j ointly sampling the set of physiological parameters as described above can create a non-identifiability problem, meaning that different combinations for the three physiological parameters can lead to the same phenotype value.
[0075] To assess the accuracy and prediction quality of the MFS model data was collected from three maize fields in the US Corn Belt (Iowa, IA) from the years 2018, 2020, and 2021. Each field is composed of multiple plots, such that a given inbred was grown multiple times and thus as several observed phenotypes were produced. When 50% of the plants in each plot were shedding (for males) and silking (for females), the corresponding dates were recorded and converted in number of days from the planting date to the flowering date. Regarding the summaries of weather data, 2018 had high GDD while 2020 had the lowest GDD accumulation. A physiological parameter validation set, where the total leaf numbers were measured for 26 genotypes grown in 2021 in fields from both Iowa and Nebraska was also used.
[0076] The Pearson correlation coefficient and the mean absolute error (MAE) between the simulated and the measured phenotype (days-to-silk) were calculated to assess the goodness of fit of MFS-WGP. As shown in Table 1, MFS-WGP is able to predict flowering in a different year, and it has slightly higher predictive ability when the estimation set comprises of two years of data and when the year of the estimation set is closer in time to the year of the validation set.Table 1: Predictive ability of days-to-silk of female inbredsEXAMPLE 2
[0077] This example demonstrates improved silking and shedding synchrony using the learned prediction model to predict the phenotype of a crop growth event.
[0078] Prior methods for determining planting dates in hybrid production fields use the best linear unbiased prediction (BLUP) methodology to estimate the total growing degree days from planting to female silking or male shedding, providing one value per genotype to be used across all locations therefore limiting the capacity of site-specific predictions of flowering time. The learned prediction model allows for genotype by environment interactions or genotype by environment by management interactions to predict one or more phenotypes for plants in a specific seed production field.
[0079] To determine if the learned prediction models improve female silking and male shedding synchrony a comparison of BLUP methodology versus the prediction model was conducted across the US seed production fields in 2023. As shown in Table 2, the prediction model is able to predict female silking for a large number of inbreds and a large number of seed production fields with less error (difference in days between observed silking and predicted silking) as compared to prior methods based on best linear unbiased predictions of synchrony of female silking (-0.27 days vs -2.13 days, respectively). Similarly, as shown in Table 3, the prediction model is able to predict male shedding (nick) for a large number of inbreds and a large number of seed production fields with less error (difference in days between observed nick and predicted nick) as compared to prior methods based on best linear unbiased predictions of synchrony of male shedding (-0.51 days vs -2.36 days, respectively).Table 2: Comparison of Silking Accuracy Using BLUP Methodology or the Prediction Model* Mean is the average of the difference between the actual observed silking day of year and the predicted silking day of year across a large number of seed production fields in 2022 and 2023 for the “Prediction model” and the BLUP valuesTable 3: Comparison of Shedding Accuracy Using BLUP Methodology or the Prediction Model* Mean is the average of the difference between the actual 50% observed shedding day of year and the predicted 50% shedding day of year across a large number of seed production fields in 2022 and 2023 for the “Prediction model” and the BLUP valuesEXAMPLE 3
[0080] This example demonstrates improving hybrid yield using the learned prediction model to predict the phenotype of a crop growth event.
[0081] A female inbred, called Feml, was selected from a female inbred dataset. Data for the female dataset was collected from three maize fields, as described above, and comprised around 2000 female inbreds. When 50% of the plants in each plot were silking, the corresponding dates were recorded and converted in number of days from the planting date to the flowering date.
[0082] Two male inbreds, called Mall and Mal2, were selected from a male inbred dataset. The male inbred data set comprises around 2000 male inbreds with data collected from three maize fields as described above. When 50% of the plants in each experiment were shedding, the corresponding dates were recorded and converted in number of days from the planting date to the flowering date.
[0083] Feml, Mall, and Mal2 were all grown in three fields and each of their phenotypes was measured. The MFS-WGP was fitted separately for the female dataset but excluding the inbred parent Feml and for the male inbred dataset excluding the inbred patents Mall and Mal2.
[0084] Two maize hybrids, called Hybl and Hyb2, were created by crossing the inbred parents FemlxMall and FemlxMal2, respectively.
[0085] The difference between days-to-silk for the female Feml and days-to-shed for the two males (Mall and Mal2) (also called anthesis silking interval), when measured or simulated with MFS-WGP, is displayed in Figs. 2A and 2B for three different fields. When producing commercial seed for the hybrid Hybl, because its two parents have similar relative maturities and flowering dynamics, both the observed and simulated difference between days-to-silk for female Feml and days-to-shed for male Mall are synchronous (Fig. 2A). Because of the synchrony in flowering between the two parents, no differential planting is needed for the two parents to ensure optimized hybrid seed production. However, when producing hybrid Hyb2,chosen here as a contrasting example, the two parents have an average difference between days- to-silk for female Feml and days-to-shed for male Mal2 of approximately five days (Fig. 2B). Thus, differential planting dates must be considered to ensure reliable hybrid seed production. Since the male inbred is on average five days earlier than the female inbred. To optimize the chance of effective release of pollen at the period of exposure of the stigmas of the female, the male should be shedding two days prior and two days after the female starts silking. Thus, in this example, the male should be planted at two different dates: three and seven days after the female planting date. For both hybrids, the differences between the simulated, when removed from the data set, male and female flowering phenotypes are in accordance with the observed values. There is a shrinkage effect towards the simulated mean difference between the two phenotypes. As a reminder, those simulated days-to-silk and days-to-shed are direct outputs from the MFS crop growth model that simulates the growth and development as a function of thermal time accumulated (Fig. 3). Therefore, in addition to the flowering dates, the vegetative stages can be simulated which is beneficial for the management of the field.
[0086] To understand how differential in planting dates could vary in hybrid seed production, a more general representation of the difference between days-to-silk and days-to-shed for a possible complete set of 2500 hybrids is represented in Fig. 4. Those hybrids represent all crosses possible between 50 females and 50 males selected from the datasets described above. The difference between days-to-silk and days-to-shed varies from -11 to 11 days. For any given male inbred (any y-axis tick on Fig. 4), its planting date in commercial seed production fields would be impacted by the choice of the female inbred. The difference between days-to-silk and days-to-shed is comparable over the three locations. The hybrids Hybl and Hyb2 (Figs. 2A-2B), that represent two contrasted examples, are highlighted in Fig. 4.EXAMPLE 4
[0087] This example demonstrates improved hybrid yield using the learned prediction model to make an in-season management decision.
[0088] The inbred parent plants are planted based on the predicted flowering time as described in Example 3. After planting the dynamic crop growth model (CGM) is updated inbred specific genetic coefficients (component trait predictions) and forward-looking weather forecast (5 - 10 days as a function of the confidence in the weather forecast) during the growing season in orderto improve the timing of in season crop management decisions such as detasseling, herbicide or pesticide application, pollinating the female parent, timing of defoliation, and harvest.
[0089] For predicting the timing of detasseling, the dynamic CGM is updated with inbred specific genetic coefficients (component trait predictions) and forward looking 5-day weather forecast and the dynamic CGM is used to simulate the occurrence of the physiological stage that warrants the initiation of the detasseling management step.
[0090] For predicting herbicide or pesticide application, the dynamic CGM is updated with inbred specific genetic coefficients (component trait predictions) and forward looking 10-day weather forecast and the dynamic CGM is used to simulate the occurrence of the physiological stage that warrants the initiation of optimized application of herbicide and respectively pesticide crop management steps.
[0091] For predicting the timing of pollination of the female parent, the dynamic CGM is updated with inbred specific genetic coefficients (component trait predictions) and forward looking 10-day weather forecast and the dynamic CGM is used to simulate the optimal timing of adding (e.g., spraying) pollen on the field that would maximize yield.
[0092] For predicting timing of defoliation, the dynamic CGM is updated with inbred specific genetic coefficients (component trait predictions) and forward looking 10-day weather forecast and the dynamic CGM is used to simulate the occurrence of the physiological stage that warrants the initiation of the defoliation management step.
[0093] For predicting harvest, the dynamic CGM is updated with inbred specific genetic coefficients (component trait predictions) and forward looking 10-day weather forecast and the dynamic CGM is used to simulate the occurrence of the crop development stage that warrants the initiation of the crop harvest management step.
[0094] All publications and patent applications in this specification are indicative of the level of ordinary skill in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated by reference.
[0095] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Unless mentioned otherwise, the techniques employed or contemplated herein arestandard methodologies well known to one of ordinary skill in the art. The materials, methods and examples are illustrative only and not limiting.
[0096] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0097] Units, prefixes and symbols may be denoted in their SI accepted form. Unless otherwise indicated, nucleic acids are written left to right in 5’ to 3’ orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively. Numeric ranges are inclusive of the numbers defining the range. Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
Claims
We claim:
1. A method for increasing yield in a hybrid seed production field, the method comprising: a. inputting genotypic information from a female parent of a hybrid plant and locationspecific environmental data for a hybrid seed production field into a first learned prediction model, the first learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for at least one crop growth event of the female parent in the hybrid seed production field; b. inputting genotypic information from a male parent of the hybrid plant and the location-specific environmental data for the hybrid seed production field into a second learned prediction model, the second learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to predict a phenotype for the at least one crop growth event for the male parent in the hybrid seed production field; c. generating by the first learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the female parent; d. generating by the second learned prediction model a prediction of the phenotype for the at least one crop growth event in the hybrid seed production field for the male parent; and e. determining, based on the prediction of the phenotype for the at least one crop growth event for the female parent and the male parent in the hybrid seed production field, an optimal planting date for the female parent, an optimal planting date for the male parent, a management decision, or any combination thereof.
2. The method of claim 1, wherein the first learned prediction model and the second learned prediction model are the same learned prediction model.
3. The method of claim 1 or 2, wherein the first learned prediction model, the second learned prediction model, or both comprise a whole genome prediction crop growth model (WGP- CGM).
4. The method of any one of claims 1-3, wherein the at least one crop growth event is flowering time.
5. The method of claim 4, wherein the prediction for flowering time phenotype is determined using the estimated leaf appearance rate, total leaf number, minimum ear biomass, or any combination thereof.
6. The method of any one of claims 1-5, wherein the method comprises determining the optimal planting date for the female parent and the optimal planting date for the male parent.
7. The method of any one of claims 1-6, wherein the optimal planting date for the male parent and the female parent comprises a planting date in which the female parent is predicted to silk and the male parent is predicted to shed pollen within 4 days of each other.
8. The method of any one of claims 1-7, wherein the method further comprises planting the female parent and male parent in the hybrid seed production field using the determined optimal planting dates, thereby increasing hybrid seed yield as compared to a field in which the first parent and second parent are planted using a GDU average.
9. The method of any one of claims 1-8, wherein the location-specific environmental data comprises historical weather conditions, forecast weather conditions, soil conditions, or any combination thereof.
10. The method of any one of claims 1-9, wherein the method further comprises (i) updating the location-specific environmental data in the first learned prediction model, the second learned prediction model, or both at least one time during the growing season, (ii) generating by the first learned prediction model, the second learned prediction model, or both comprising the updated location-specific environmental data a prediction of the phenotype for at least one crop growth event in the hybrid seed production field for the female parent, the male parent, or both, and (iii) determining, based on the prediction of the phenotype for the at least one crop growth event for the female parent, the male parent, or both a management decision.11 . The method of claim 10, wherein the management decision comprises herbicide application, irrigation, defol application, pesticide application, detasseling, or any combination thereof.
12. The method of claim 11, wherein the method further comprises applying management decision to the female parent, the male parent, or both in the hybrid seed production field.
13. The method of any one of claims 1-12, wherein the first learned prediction model, the second learned prediction model, or both the first and second learned prediction model is trained by: a. growing a population of training individuals in a plurality of environments; b. phenotyping the population of training individuals in the plurality of environments to generate a phenotype by environment training data set; and c. associating the phenotype by environment training data set with a genotype training data set comprising genetic information across the genome of each training individual, using a biological model, estimating effects of genotypic markers and linking the estimation of effects of genotypic markers with the biological model to generate an association training data set.
14. The method of claim 13, wherein the training individuals are developed from a traditional breeding program, a genome edited breeding program, or a combination thereof.
15. A method for increasing yield in a hybrid seed production field, the method comprising: a. generating a dynamic crop growth model (CGM) for a first parent of a hybrid plant, the dynamic CGM generated by inputting genotypic information from the first parent of the hybrid plant and location-specific environmental data for a hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent in the hybrid seed production field; b. generating a dynamic CGM for a second parent of the hybrid plant, the dynamic CGM generated by inputting genotypic information from the second parent of the hybrid plant and location-specific environmental data for the hybrid seed production field into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment bymanagement interactions to simulate plant growth for the second parent in the hybrid seed production field; c. determining an optimal planting date for the first parent and an optimal planting date for the second parent in the hybrid seed production field based on the dynamic CGM generated for the first parent and the second parent of the hybrid plant; d. planting the first parent on the determined optimal planting date and planting the second parent on the determined optimal planting date; e. updating the dynamic CGM for the first parent, the second parent or both the first and second parent at least one time by inputting observed location-specific environmental data into the learned prediction model; and f. determining at least one field or crop management decision based on the updated dynamic CGM for the first parent, the updated CGM for the second parent or the updated CGM for both the first and second parent.
16. The method of claim 15, wherein the dynamic CGM predicts a flowering time for the first parent and the second parent, and the optimal planting date for the first parent and the optimal planting date for the second parent is determined from the prediction of flowering time.
17. The method of claim 16, wherein the optimal planting date for the first parent and the optimal planting date for the second parent comprises a date in which the first parent and second parent are predicted to flower and silk within 4 days of each other.
18. The method of claim 16 or 17, wherein the prediction for flowering time phenotype is determined from the dynamic crop growth model by using the estimated leaf appearance rate, total leaf number, minimum ear biomass, or any combination thereof.
19. The method of any one of claims 15-18, wherein the location-specific environmental data comprises historical weather conditions, forecast weather conditions, soil conditions, or any combination thereof.
20. The method of any one of claims 15-19, wherein the observed location-specific environmental data comprises historical weather conditions, forecast weather conditions, soil conditions, or any combination thereof.21 . The method of any one of claims 15-20, wherein the at least one field or crop management decision herbicide application, irrigation, defol application, pesticide application, detasseling, or any combination thereof.
22. The method of any one of claims 15-21, wherein the learned prediction model comprises a whole genome prediction crop growth model (WGP-CGM).
23. The method of any one of claims 15-22, wherein the learned prediction model is trained by: a. growing a population of training individuals in a plurality of environments; b. phenotyping the population of training individuals in the plurality of environments to generate a phenotype by environment training data set; and c. associating the phenotype by environment training data set with a genotype training data set comprising genetic information across the genome of each training individual, using a biological model, estimating effects of genotypic markers and linking the estimation of effects of genotypic markers with the biological model to generate an association training data set.
24. The method of claim 23, wherein the training individuals are developed from a traditional breeding program, a genome edited breeding program, or a combination thereof.
25. A method of allocating fields for hybrid seed production, the method comprising: a. inputting genotypic information from a first parent and a second parent for each of a plurality of hybrid products and location-specific environmental data for a plurality of hybrid seed production fields into a learned prediction model, the learned prediction model trained to learn a relationship of genotype by environment interactions or genotype by environment by management interactions to simulate plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields; and b. determining the optimal hybrid seed production field for at least one of the plurality of hybrid products based on the simulated plant growth for the first parent and the second parent for each of the plurality of hybrid products in the plurality of hybrid seed production fields.
26. The method of claim 25, wherein after (a) and before (b) the method further comprises inputting the simulated plant growth for the first parent and the second parent for one or more of the plurality of hybrid products in the plurality of hybrid seed production fields into at least one additional learned prediction model, the at least one additional learned prediction model trained to learn a relationship between hybrid production performance in an environment and production costs and wherein the optimal hybrid seed production field for the at least one of the plurality of hybrid products is based on both the simulated plant growth and the output of the at least one additional learned prediction model.
27. The method of claim 25 or 26, wherein the production costs comprise transportation costs.
Citation Information
Patent Citations
Machine learning in agricultural planting, growing, and harvesting contexts
US20190050948A1
Hybrid seed selection and seed portfolio optimization by field
US20190139158A1
Synchronized breeding and agronomic methods to improve crop plants
US20230030326A1
Methods and systems for crop land evaluation and crop growth management
US20240095621A1