Accelerated breeding and improved training data

By employing machine learning models and dividing regeneration material for both breeding and data generation, the breeding process is accelerated, addressing inefficiencies in traditional methods and enhancing prediction accuracy for faster cultivar development.

WO2026102419A1PCT designated stage Publication Date: 2026-05-15MONSANTO TECHNOLOGY LLC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MONSANTO TECHNOLOGY LLC
Filing Date
2025-11-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional plant breeding methods are time-consuming and inefficient, with prediction accuracy decreasing over successive breeding cycles due to the need for data from closely related lines, and there is a lack of efficient methods to accelerate the development of new cultivars.

Method used

Utilize machine learning models to predict and select plants or genomes for continued breeding, divide regeneration material into portions for both breeding and data generation, and incorporate environmental data to enhance prediction accuracy, allowing for faster genetic gain and yield improvement.

Benefits of technology

Accelerates the breeding process by increasing prediction accuracy and reducing the time required for generating new cultivars, enabling faster genetic gain and yield improvement through advanced data collection and model refinement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000035_0001
    Figure IMGF000035_0001
  • Figure IMGF000041_0001
    Figure IMGF000041_0001
  • Figure IMGF000041_0002
    Figure IMGF000041_0002
Patent Text Reader

Abstract

Described herein are materials for, and methods of, breeding that allow for efficient increases in crop improvement. The methods include one or more machine learning (ML) models to predict and / or select which plants or genomes from at least one pool of plants are useful for continuing forward in the breeding program. Another approach that can be utilized is to assess in silico the desirability of simulated progeny from plants in the pool to decide which plants would produce desired progeny.
Need to check novelty before this filing date? Find Prior Art

Description

PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WOAccelerated Breeding and Improved Training DataCross-Reference to Related Applications

[0001] This application claims the priority benefit of United States Provisional Application No. 63 / 718,436, filed November 8, 2024, which is incorporated herein by reference in its entirety.Background

[0002] Breeders are continually developing new cultivars through various plant genetic improvement programs. These programs commonly rely on both forward breeding methods and backcross-based trait introgression methods. These methods can be time-consuming and inefficient. As breeders look to accelerate crop variety development, it is critical to develop improved methods for developing new cultivars that increase efficiency and facilitate a faster generation of new cultivars.Summary

[0003] Both the creation and testing of new breeding lines contribute to the cost of a plant breeding program. The duration of their growth cycles, and dependence upon environmental conditions to reach maturity, seed multiplication rates and many other factors lead to long lag times between the sexual crosses performed to initiate the creation of new breeding lines and collecting the data that allows new lines to be assessed. A variety of approaches are used to reduce the time and resources needed to make and test lines. For example, plants may be grown in controlled environments that allow multiple sexual generations in a single year. Modem plant breeding can incorporate in silico genomic predictions, however, there is still no escape from the need for actual plants to be used in plant breeding programs. Disclosed herein are methods to significantly increase the speed and accuracy of accelerated breeding programs to enable faster increases in genetic gain and / or increased yield.

[0004] A general limitation in the use of performance prediction methods is that prediction accuracy is lower for models trained with data that is collected from lines that are less related to individuals whose performance is being predicted compared to data from lines that are more related. In practice, the accuracy of prediction is reduced over successive breeding cycles unless data from more related lines is continually collected and incorporated. Thus, it is necessary to update models with training sets from closely related individuals.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0005] Described herein are accelerated plant breeding methods that include using, among other things, one or more machine learning (ML) models to predict and / or select which plants or genomes from at least one pool of plants are useful for continuing forward in the breeding program. Another approach that can be utilized is to assess in silico the desirability of simulated progeny from plants in the pool to decide which plants would produce desired progeny.

[0006] In addition to selecting lines that are suitable to use for crossing and propagation of the breeding pools, lines may be selected to produce progeny for field testing or as candidate lines to create new commercial varieties. In some embodiments, at least two portions of regeneration material may be desired from a plant selected from the pool. In some embodiments, one of the portions of regeneration material from a single plant can be continued forward in a breeding cycle in the pool. In some embodiments, a second portion of the regeneration material can be used to develop a commercial line, placed into field trials to generate data for training the ML model or combinations thereof. The high genomic similarity of the regeneration material in both portions allows for higher accuracy for future predictions relating to future cycles of the accelerated breeding program. The accelerated breeding program can contain one or more pools of plants. A pool of plants can be described by the overall goal of the particular pool. For example, a pool of plants can be developed with the intention to create new varieties that are well adapted to a specific geographic region. Individual plants within a pool will have some genetic features in common but the pool will comprise combinations of various genetic features including can contain alleles that contribute to quantitative traits, alleles that have qualitative effects (also called native traits), transgenes, gene edits, and combinations thereof. Mathematical approaches such as and also be selected for genomic economic estimated breeding value (GEBV) can be used to make predictions about individuals within the pool based on their genetic composition. Such predictions can be used to prioritize crosses among particular individuals selected within the pool to generate the next generation and increase the average performance of the pool over successive rounds of crossing and selection (see Peixoto et al, 2024, Utilizing genomic prediction to boost hybrid performance in a sweet com breeding program, Front in Plant Sci, Vol 15 (https: / / doi.org / 10.3389 / fols.2024 1293307); Peixoto et al, 2024, Use of simulation to optimize a sweet corn breeding program: implementing genomic selection and double haploid technology, G3 Genes|Genomes|Genetics, Vol 14 (https: / / d0i.0rg / l 0, 1093 / g3 j oumal / jkae 128); Meuwissen etPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO al, 2001, Prediction of Total Genetic Value Using Genome-Wide Dense Marker Maps, Genetics, Vol 157, 4:1819-1829 (ht,tps: / / doi prg / 10, 1093 / genetics / l 57,4 J 819); and Hickey et al, 2014, Evaluation of Genomic Selection Training Population Designs and Genotyping Strategies in Plant Breeding Programs Using Simulation, Crop Sci, Vol 54 (https: / / doi.org / 10.2135 / cropsci2013.03.0195).

[0007] One of ordinary skill in the art will appreciate that environmental data including, for example, soil type, precipitation, wind strength, sun intensity, fertilizer utilization, cover crop status, crop rotations and combinations thereof can be collected and incorporated into mathematical models used to select individuals with genomes of interest to use in in the accelerated breeding program.

[0008] In some embodiments an accelerated breeding program uses pools of genomes that are at least minimally heterogeneous, and contain individuals that contain some level of heterozygosity as further described herein. Additional genetic variation is generated in each generation of the pool through genetic recombination at meiosis that creates individuals with new combinations of alleles that exist within the pool. Periodically, the accelerated breeding program benefit by incorporating individuals from outside the pool to provide additional genetic variation including new alleles, new combinations of alleles, edits, transgenes and combinations of these. Decisions to introduce new lines into a pool may be made to provide a distinct new trait, such as a disease resistance gene, or may be done based on statistical approaches, for example, to maintain the right average co-ancestry value, as calculated using for example the VanRaden method 1, or comparable methodology.

[0009] In some embodiments, the methods include genotyping plants; 1) in the accelerated breeding program, 2) plants resulting from crosses using regeneration material, and 3) combinations of 1 and 2. Genotyping can be accomplished using any method known in the art. In some embodiments whole genome sequencing can be used on one more individuals within the pool or related to the pool and in instances where 100% of the genomic sequence is not experimentally identified for any given individual within the pool, the missing sequences can be imputed based upon data sets used in the accelerated breeding program. As described herein an accelerated breeding program can contain one or more pools of plants. A pool of plants can be characterized by the number of generations, or crosses that are made per year, the diversity of the genomes maintained in the pool, the number of plants maintained in the pool (i.e., more than 100,PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO200, 300, 500, 1000, 1500, or 2000 distinct genomes), the desired outcome of the pool (i.e., plus disease resistance, yield alone, insect resistance, environmental resilience, etc.) and combinations thereof. At any point in the accelerated breeding program one or more genomes can be selected for extraction from the pool (see, FIGS 1 and 2). The selected genome can be used to create at least two portions of regeneration material. One portion will continue in the pool and the other portion can be used for one or more additional purposes. Such purposes include for example, generating data that can be used to educate the ML algorithm(s) used in the accelerated breeding program and developing commercial products through various breeding schemes. The portions can be created, for example, by dividing pollen from a plant having a selected genome into two or more portions and using the portions for the purposes described, taking cuttings from a selected genome and generating clones and / or regenerating plants from tissues. As mentioned, among other things, benefits of the creation of at least two portions of regeneration material can include increasing the speed of advancement of the pool goals and returning data for educating and improving the ML model(s) in less than 4, 6, 8, 12, 18, or 24 months.

[0010] One of ordinary skill in the art of plant breeding will appreciate the ability to collect phenotypic and trait data in real time during plant development. During development as a seed germinates into a plantlet, the plant grows to maturity, and finally the plant reaches senescence different sets of phenotypic and trait data become available to a plant breeder. In an accelerated breeding program, the plant breeder efforts to collect and return data expediently to operate ML models with the most recent data possible. For some traits, such as germination, this can be in less than 2, 3, 4, 5, 6, 7, 10, or 14 days from the planting date. For traits and phenotypes that do not express or become available until maturity and / or senescence, or post harvesting data collection and return to ML models can be in less than 2, 3, 4, 5, or 6 months from the planting date depending on the type of plant or crop.

[0011] In a typical breeding program, one or more plant growing seasons may be needed to prepare materials for field testing to collect phenotypic and trait data. Methods for preparing materials for field testing can vary by crop, typically occurs in a linear step-by-step progression, and includes, but is not limited to double haploid production, seed increase, and / or hybrid makeup (e.g., crossing parent plants of different heterotic groups) with the field test consisting of homozygous lines (e.g. inbreds) such as double haploids, or lines below a threshold of heterozygosity, such as <10% (e.g. lines selfed to at least F5 generation are approximately 6.25%PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO heterogeneous). Field testing typically occurs at least 18 or 24 months after selecting individuals to begin preparations (e.g. DH production, selfing to at least F5 generation, seed increase, hybrid makeup, etc.) for the field test. In an accelerated breeding program, methods of preparing materials for field testing can occur in linear or non-linear progression (i.e., in parallel or asynchronous), and comprise lines with low heterozygosity (<10%) or high heterozygosity, such as 12.5% (F4 generation), 25% (F3 generation), 50% (F2 generation), or 100% (Fl generation) (numbers provided are an average estimate stemming from inbred parents that are from distant heterotic groups). In an accelerated breeding program with continued cycling pools of genomes, parental crosses are continually made between individuals with some level of heterozygosity (i.e., outcrossing within the pool) to promote genetic variability. Individual plants extracted directly from accelerated breeding cycling pools can then be used to create outbreds. Field testing in an accelerated breeding program, with linear or non-linear preparations and utilization of low or high heterozygosity lines, can occur in less than 5, 6, 7, 8, 9, 10, 11, 12, 15, 18 or up to 24 months after selecting individuals to be entered into the field test.

[0012] In some embodiments the regeneration material is pollen and in yet other embodiments the pollen can be primed prior to it being used to continue in the pool. Priming refers to processes that select for pollen that can survive exposure to various conditions, such as heat, cold, mechanical manipulation, chemical exposure and combinations thereof. The pollen can be primed prior to being used to create plants for generating data for use in the ML model. In yet additional embodiments the pollen can be primed and used for both continuing the pool and for generating data. The pollen can be primed at the bicellular or tricellular state and any treatment known in the art can be used in the priming process. For example, the pollen can be primed using heat treatment, cold treatment, pressure treatment, chemical treatment, abrasion, and combinations thereof.

[0013] In some instances, the regeneration material is pollen, and the pollen can be used for continued breeding in the accelerated breeding program, and a portion of the pollen can be used to pollinate a paternal haploid inducer line. In some embodiments the pollen is stored and one of skill in the art will appreciate that storage conditions and processing can provide a form of priming. In yet, other embodiments the paternal haploid inducer line can express gene editing polypeptides and the haploid plant, or diploid plant resulting from the cross can contain an edited genome.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0014] In yet other embodiments the regeneration material can be tissue, and the portion of regeneration material used for data collection and / or commercial line development can be the result of vegetative propagation.

[0015] In yet other embodiments the regeneration material can be an explant, embryo, or seed, and the portion of regeneration material used for data collection and / or commercial line development and / or used for continued breeding in the accelerated breeding program can be a combination of pollen, tissue, explant, embryo, and seed.

[0016] In yet other embodiments that regeneration material collected from the plant selected from the pool for extraction from the accelerated breeding program can be a combination of pollen, and tissue.

[0017] One of skill in the art will appreciate that pollen as described herein can be stored prior to its use, thus allowing it to be used flexibly relative to the fertility timing of the female plants and the location of such plants. Also described herein is a tracking system developed to integrate the use of stored pollen into an accelerated breeding program.

[0018] Many of the methods described herein can be used generally for plants, however, some of the methods described are particularly useful for certain plant varieties. Crop plants include, for example, corn (maize) (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), wheat (Triticum aestivum), soybean (Glycine max), tobacco (Nicotiana tabacum), cotton (Gossypium barbadense, Gossypium hirsutum), oats, barley, and vegetables.

[0019] Another embodiment described herein includes using the regeneration material to create double haploid progeny that may be developed into commercial lines, used to create field data to update training sets or both, progeny may be genotyped and preferred progeny selected. The use of paternal haploid inducer (PHI) female plants to produce haploid progeny allows a suitable number haploid progeny to be produced from selected individuals in the accelerated breeding program. In some embodiments, the pollen from a single selected individuals is collected and applied to many female PHI plants to produce a population of haploid progeny that may be genotyped and a subset of progeny selected based on their predicted value to use for commercial lines, field testing or both. In other embodiments, pollen from 2 or more selected individuals in the accelerated breeding program is combined and applied to 2 or more PHI females to produce a population of haploid progeny that may be genotyped and selected.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0020] Another embodiment described herein includes methods of making paternal haploid plants using stored pollen. Such methods can include harvesting the pollen, storing the pollen for at least 30 minutes, at least an hour, at least a day, at least a week, at least a month, at least six months, at least a year or at least 5 years and applying the pollen to a paternal haploid inducer plant(s) and collecting seeds from the resulting plants. These methods can also include inducing genomic doubling of the haploids to produce double haploid plants. When double haploid plants are made according to these methods, the nuclear genomes of the double haploid plants are from the pollen and the cytoplasmic genomes are from the female inducer plant. As with all of the methods described herein, the pollen can be primed prior to application, including before use with a paternal haploid inducing line.

[0021] Plant breeding relies upon the cross pollination between parent plants, which historically required planting both parents in the same geographic location and reaching reproductive maturity at the same time. The advent of pollen storage processes eliminates the spatial and temporal constraints of pollination. A batch of pollen that has been collected and stored can be used to pollinate multiple plants and produce numerous progeny. However, maintaining pollen identity and tracking from collection through operations, containers, and equipment required for storage and subsequent pollination on the desired female recipient in a breeding program is challenging. One aspect of the disclosure provides methods of tracking, inventorying, and auditing pollen samples within a global breeding program enabling compliant use and distribution of viable pollen across operations and use cases within a breeding program. The methods can include collecting pollen at a first location, associating the pollen sample with its donor plant collected from, maintaining the association through subsequent processing steps, storing the pollen for at least 1 day using any method known in the art, applying the pollen to a second plant and collecting seed from the second plant. These methods can include freezing the pollen, thawing the pollen, priming the pollen, and combinations thereof.

[0022] The pollen tracking system can additionally include calculating the viability of the pollen using predictive modeling based upon the environmental conditions the pollen is exposed to throughout the collection, storage and distribution process. One of skill in the art will appreciate that pollen viability is also a trait that can be modeled and particular plants that are selected for extraction from an accelerated breeding program can be selected based, in part, upon the predicted stability of the pollen. See, Frova, C., Sari-Gorla, M. Quantitative trait loci (QTLs) for pollenPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO thermotolerance detected in maize. Molec. Gen. Genet. 245, 424-430 (1994). In some embodiments, the pollen tracking system can provide prescriptive use information depending upon the predicted viability of the pollen. For example, a particular pollen sample with a particular genome and stored under specific conditions can be predicted to be 50%, 60%, 70%, 80%, or 90% viable, and usage can then be modified based upon such predictions. For example, application rates can be increased or decreased depending upon such information.

[0023] Described herein are also methods of using vegetative propagation (VP) in combination with accelerated breeding methods. These methods include using a ML algorithm to identify genomes that can be selected for continuation in the breeding program. The selected genomes can be multiplied by taking tissue samples or cuttings and propagating the tissue samples to maturity. Such methods are particularly useful in breeding programs that are used with plants that do not produce high numbers of seeds. The resulting plants grown from the tissue sample or cuttings (VP plants), the originating plant (sometimes hereinafter referred to as the mother plant), and progeny thereof can be used to create data that can be used to educate the algorithm being used in the accelerated breeding program. Advantages of these methods include being able to create data from both the mother plant and the VP plant without waiting for an entire growth cycle to collect seeds and multiply the mother plants.

[0024] In summary, the invention is directed at the following numbered embodiments:

[0025] Embodiment 1 : A method of accelerated plant breeding comprising:Selecting at least one first plant from a cycling population of plants using a machine learning model;Making at least two portions of regeneration material from the at least one first plant;Using a first portion of the at least two portions of the regeneration material to cross with at least one second plant in a cycle of an accelerated breeding program;Using a second portion of the at least two portions of the regeneration material to produce plants and generate data; andUsing the data generated to educate the machine learning model, wherein using a first portion of the at least two portions of regeneration material to cross with the at least one secondPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO plant and using the second portion of the regeneration material to generate data to educate the machine learning model provides for increased prediction accuracy.

[0026] Embodiment 2: A method of increasing the prediction accuracy of a genomic prediction model comprising: selecting a genome from a cycling population of plants for use in generating data to train a machine learning model;Crossing a plant comprising the desired genome with a substantially homozygous plant;Collecting data from the progeny of the cross; andUsing the data to refine the machine learning model, wherein the time between selecting and using the data is less than 1.5 years.

[0027] Embodiment 3: A method of using mixed germplasm source data for updating machine learning models for genomic selection, comprising:Maintaining at least one cycling population of genomes; selecting from the at least one cycling population of genomes a desired population; dividing the desired population into a first sub-population of genomes for data generation and a second sub-population for generating another cycle in the at least one cycling population of genomes; directly crossing the first sub-population of genomes for data generation with one or more plants having known associated data; collecting data from the cross to produce outbred data; and retraining the machine learning models for genomic selection using the outbred data and a data set from an earlier cycle in the at least one cycling population of genomes.

[0028] Embodiment 4: The method according to any one of embodiments 1-3, wherein the cycling population of plants comprises a co-ancestry value of less than 1 as calculated by the VanRaden method 1.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0029] Embodiment 5: The method according to any one of embodiments 1-4, further comprising tracking the second portion of regeneration material and tracking the environmental conditions the regeneration material is exposed to.

[0030] Embodiment 6: The method according to any one of embodiments 1-5, further comprising genotyping and using the genotype data in the selecting step.

[0031] Embodiment 7: The method according to any one of embodiments 1-6, wherein the data generated is used to educate the machine learning model within 1 year of the selection of the at least one first plant for extraction, selecting step from a cycling population and the like.

[0032] Embodiment 8: The method according to any one of embodiments 1-7, wherein the data generated is used to educate the machine learning model within 6, 9 of 12 months of the selection of the at least one first plant for extraction, desired population and the like .

[0033] Embodiment 9: The method according to any one of embodiments 1-8, wherein the selected genome or regenerative material is extracted from the cycling population as pollen

[0034] Embodiment 10: The method according to any one of embodiments 1-9, wherein at least one of the at least two portions of regeneration material that is pollen is primed.

[0035] Embodiment 11 : The method according to any one of embodiments 1-10, wherein the pollen is primed during the bicellular or tricellular state.

[0036] Embodiment 12: The method according to any one of embodiments 1-11, additionally comprising crossing the pollen that is the second portion of regeneration material, the desired genome extracted from the cycling population or the desired genome with a paternal haploid inducer plant.

[0037] Embodiment 13: The method according to any one of embodiments 1-12, wherein the paternal haploid inducer plant comprises genome editing components.

[0038] Embodiment 14: The method according to any one of embodiments 1-13, further comprising using at least one of the at least two portions of pollen for outbred testing.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0039] Embodiment 15: The method according to any one of embodiments 1-14, wherein using a second portion of the regeneration material to generate data, further comprises generating a double haploid and collecting data.

[0040] Embodiment 16: The method according to any one of embodiments 1-15, wherein the regeneration material is tissue and the second portion of regeneration material is created using vegetative propagation.

[0041] Embodiment 17: The method according to any one of embodiments 1-16, wherein the pool of plants are selected from corn (maize) (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), wheat (Triticum aestivum), soybean (Glycine max), tobacco (Nicotiana tabacum), cotton (Gossypium barbadense, Gossypium hirsutum), oats, barley, and vegetables.

[0042] Embodiment 18: The method according to any one of embodiments 1-17, wherein the pool of plants, cycling population and the like are corn.

[0043] Embodiment 19: The method according to any one of embodiments 1-18, wherein the data is selected from data used to compute GEBV.

[0044] Embodiment 20: The method according to any one of embodiments 1-19, wherein the data comprises environmental data in combination with phenotypic data.

[0045] Embodiment 21 : The method according to any one of embodiments 1-20, wherein the priming is selected from heat treatment, cold treatment, pressure treatment, chemical treatment, abrasion, desiccation treatment, and combinations thereof.

[0046] Embodiment 22: The method according to any one of embodiments 1-21, wherein the cycling population comprises a co-ancestry value of less than 1 as calculated by the VanRaden method 1.

[0047] Embodiment 23: The method according to any one of embodiments 1-22, wherein the cycling population advances less than 5 generations before the data is used to refine the machine learning model.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0048] Embodiment 24: The method according to any one of embodiments 1-23, wherein the selected genome is in a form selected from the group consisting of pollen, a tissue cutting, or a seed.

[0049] Embodiment 25: The method according to any one of embodiments 1-24, wherein the selected genome is a seed and the method further comprises growing the seed into a plant.

[0050] Embodiment 26: The method according to any one of embodiments 1-25, wherein the data set from an earlier cycle comprises data from a substantially homozygotic individual.

[0051] Embodiment 27: The method according to any one of embodiments 1-26, wherein a crossing partner that crosses with a selected genome from the cycling population is a substantially homozygotic individual and is a maternal double haploid, paternal double haploid, a selfed plant, and combinations thereof.Detailed DescriptionFigure Description

[0052] FIG. 1 shows a diagram of an accelerated breeding program that is using machine learning (ML) indicated as part of the process occurring at the open arrows which allows for selection of genomes to advance. The selection of genomes for continued breeding and genomes to discard can happen multiple times during the year as indicated. At any point in the cycle desirable genomes can be identified for further testing or commercialization. Regeneration material can be retrieved from the accelerated breeding program for example by cuttings from a donor plant, retrieval of immature embryos from a donor plant (e.g., from a seed), removal of a portion of the pollen, or combinations thereof. The regeneration material can then be used to; increase seed count for further rounds of breeding, data collection, or line creation for desired target products.

[0053] FIG. 2 shows a diagram comparing the accelerated breeding program described herein (B) compared to traditional breeding programs (A). The traditional workflow includes selecting a seed or plant that is genomically similar (i.e., sibling, half-sibling) to the sample that is extracted from the population in the breeding cycling. The extracted genome or plant is allowed to mature into a double haploid through the use of maternal haploid induction system (MHIS) or can be grown and selfed. The now inbred line can be used in field testing and / or for seed increase. The dotted linePCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO indicates when data would be available for educating the ML model. The second workflow indicated by B shows how a portion of regeneration material is removed from the population and through VP, paternal haploid induction (PHIS), or selfing the portion can be used to collect data for educating the ML model. The relatedness of the portion of regeneration material in the population and the portion that is extracted from the population is greater than depicted in A, and therefore the processes shown in B lead to increased accuracy of predictions when the data is collected and returned to education the ML algorithm.

[0054] FIGS. 3A and 3B show the drift in accurate relationships among genotypes as generations, or distance, increase without model retraining (FIG 3A) and genetic relationship among highly related individuals (e.g., full siblings) compared to less related individuals (e.g., second cousins) (FIG 3B). One of ordinary skill in the art will appreciate that both the number of generations (i.e., amount of time) between model retraining (such as depicted in FIG 3A) and relatedness of individuals in a training set (FIG 3B) are key factors for accurately predicting relationships among genotypes and traits. See Bernardo 2025, Why does genomewide prediction become ineffective after several cycles of recurrent slection?, Crop Sci Vol 65 (https: / / doi.org / l 0, 1002 / csc2.70164); and Visscher et al, 2006, Assumption-Free Estimation of Heritability from Genome-Wide Identity - by-Descent Sharing between Full Siblings, PLOS Genetics, Vol 2 (http s : / / doi . org / 10.1371 / j ournal . pgen .0020041).

[0055] FIGS. 4A and 4B show the genomes selected for extraction from a cycling breeding population (dark leaves) which are referred to also using a male sign, but one of skill in the art will appreciate that pollen donor and pollen recipient roles can be reversed. The genomes selected for outbred testing are subsequently crossed with known substantially homozygotic lines, shown here as double haploids (DH) lines (FIGS. 4A and 4B). Known plants as used herein refers to plants that have ancestors that have been genotyped and tested in one or more of the following: fields, controlled environments, or combinations thereof, and as such have a confidence associated with their predicted breeding value, thus controlling the noise their genomes add when used to test the portion of regenerative material from the cycling population. The outbred crosses result in plants that can be used for data collection and education of the machine learning (ML) models within 1.5 years from the extraction from the cycling population.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0056] FIG. 5 shows an example of the use of pollen from selected lines in an accelerated breeding program (also referred to as a continuous cycle program) to make haploids for line development and testing. In this particular diagram kernels are seed chipped and their genomes are analyzed and certain seeds are planted [1], A portion of the pollen can be optionally stored, primed by various treatments, transported and / or directly used to pollinate maternal plants that are haploid inducers (Manonmani et al, 2025, Chapter 2: Harnessing Pollination in Crop Plants and Pollen Selection Application in Plant Breeding, in Next Generation Plant Breeding Professional Prints, NY USA(https: / / www.researchgate.net / publication / 391109622_Harnessing_Pollination_in_Crop_Plants_ and_Pollen_Selection_application_in_Plant_Breeding)). Other portions of the pollen can be used to pollinate plants in the next round of the accelerated breeding program. The resulting haploids can be genotyped and / or doubled using any method known in the art [2], The resulting plants can be used to generate commercial products, as breeding stock in the accelerated breeding program, used to generate data for further refinement of the models used for selection or combinations thereof.

[0057] FIGS 6A and 6B show schematically the impact of the use of vegetative propagation on an accelerated breeding program. FIG. 6A illustrate the necessity of not being able to collect training data and yet still keep the same genetics in the cycling population.

[0058] FIG. 7 shows the impact of various VP methods on relative maturity as compared with the plant of origin (mother). Time taken by soybean cuttings to reach different development stages. Cuttings derived from V2 and V3 donor plants were subjected to different rooting treatments: no hormone treatment (UT), Hormex3 at 0.3% IBA (indole-butyric acid)) concentration in powder and Clonex gel (IBA 0.3%), and recovered with two different growth media: Rapid Rooter (RR) plug and vermiculite (Vermi). For cuttings recovered in RR, once recovered, retransplanted into bigger pots with BM7 growing mix. For cuttings inserted in Vermiculite, no retransplanting. Development stages: R1 - beginning bloom, R2 - full bloom, R3 - beginning pod, R4 - full pod, R5 - beginning seed, R6 - full seed. Trendline is forRl stage. The days after donor seeds planting were recorded (n= 4). Data from a single germplasm line.

[0059] FIGS. 8A and 8B show the impact of using regeneration material that is divided into two or more portions and continued in a subsequent generation of the pool, as well as being used inPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO testing. FIG. 8A depicts the individuals identified as 2s that will continue in the cycling population and the related individuals identified as a 1 that will be used for data generation. FIG. 8B shows the impact of being able to use the entire population for both.

[0060] FIGS. 9A through 9D show the amount of heterozygosity as measured from genotypes of the genomes in the accelerated breeding program from two exemplary cycling pools at specific generations of cycling and extraction. The overall goal of the first exemplary cycling pool is genetic gain improvement, hence it is an elite cycling pool and it is defined by a set maturity range for the plants cycling within the pool. Heterozygosity levels in the elite cycling pool generation 1 with both heterotic groups combined primarily ranged between 0.07 to 0.33 with a low count of individuals in the pool <0.07 heterozygosity (FIG. 9A). Heterozygosity levels in one set of extracted plants from the elite cycling pool that were matured into double haploids (DH) via MHIS (prior to cycling generation 1), were essentially 0, as expected (FIG. 9D). Heterozygosity levels in another set of extracted plants from the elite cycling pool generation 1, that were selfed a single generation, ranged between 0.10 to 0.30 (FIG. 9C). The overall goal of the second exemplary cycling pool is genetic diversity, hence it is a diversity cycling pool and it is defined by the same maturity range for the plants cycling within the pool as the exemplary elite cycling pool. Heterozygosity levels in one set of extracted plants from the diversity cycling pool ranged between 0.10 to 0.43 (FIG. 9B).

[0061] FIG. 10 shows the genomic relationships by heterotic group of genomes in the accelerated breeding program from two exemplary cycling pools at specific generations of cycling and extraction (see FIGS. 9A-9D). The genomic relationships were calculated by the VanRaden method 1. For each heterotic group, the genomic relationships were calculated within the elite cycling pool generation 1 (A) and ranged from 0.75 to 1.00 (Heterotic Group 1) and 0.52 to 0.63 (Heterotic Group 2). The genomic relationships between the elite cycling pool generation 1 and extracted plants that were selfed from the elite cycling pool generation 1 (B) ranged from 0.37 to 0.92 (Heterotic Group 1) and 0.46 to 0.59 (Heterotic Group 2). The genomic relationships between the elite cycling pool generation 1 and extracted plants that were matured into DH via MHIS from the elite cycling pool prior to generation 1 (C) ranged from 0.27 to 0.85 (Heterotic Group 1) and 0.43 to 0.59 (Heterotic Group 2). The genomic relationships between the elite cycling poolPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO generation 1 and extracted plants from the diversity cycling pool generation 1 (D) ranged from 0.33 to 0.64 (Heterotic Group 1) and 0.34 to 0.50 (Heterotic Group 2).Overview

[0062] Marker assisted breeding (MAS) has been used for many years to support traditional breeding. Initially MAS focused on using markers that specifically identified regions of genomes that included favorable or desirable traits. Examples of traits that have been tracked in breeding programs include naturally occurring traits (i.e., native traits), for example, disease resistance, as well as traits that were inserted using genetic engineering (i.e., transgenes), for example the Cry 1 Ac gene that was originally derived from Bacillus thuringiensis. Expanded marker libraries, for instance genome wide marker libraries, can be used to support accelerated breeding programs. Accelerated breeding programs can also be supported using genomic sequencing, large marker libraries, and combinations thereof. One of skill in the art will appreciate that trait focused MAS, does not necessarily capture all of the allelic variation information that can be tracked in an accelerated breeding program, especially not in traits that are affected by multiple genes (e.g. quantitative traits). Therefore, when MAS is used, it may lead to slower product development, and possibly inferior product development. One of ordinary skill in the art will appreciate that using genomic sequencing techniques or larger marker libraries can be helpful for tracking many more allelic variations that contribute to the phenotype.

[0063] Genomic selection allows for data to be captured across the whole genome which supports a predictive approach rather than a design approach normally associated with MAS. For example, using the genotypic information of inbred lines a prediction of the genomes that will result from a cross can be made and depending upon the calculated probability of achieving a desired genome the parental cross can be made, or not. Genomes that do not give rise to an acceptable probability of increasing the desired characteristic can be discarded (See. FIG 1). One of ordinary skill in the art will appreciate that using whole genome selection allows for improved prediction accuracy.

[0064] One of ordinary skill in the art will also appreciate that the science of plant breeding has shifted to rely heavily upon machine learning (ML), including for example artificial intelligence (Al), statistical analysis (SA), as well as bioinformatics (BI). The success of any program relying on ML, BI, and Al, depends upon the algorithms used and the quality of the data used to train those algorithms. If a breeding program relies upon low quality data to train its models the resultPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO will be sub-optimal genome production. Therefore, experiments need to be designed to capture the right data of high quality.

[0065] A typical accelerated breeding program includes steps such as tissue collection, genotyping, imputation, cross allocation, germination, and seed production. Imputation refers to a statistical technique to improve genotype coverage by inferring un-genotyped markers of an individual based on directly genotyped markers and linkage disequilibrium patterns observed in an appropriate reference population. See Schurz et al, 2019, Evaluating the Accuracy of Imputation Methods in a Five-Way Admixed Population, Front in Genetics, Vol 10 (https: / / doi.org / 10.3389 / fgene.2019.00034).

[0066] As used herein accelerated breeding refers to any plant breeding program that targets population improvement by relying upon predictive statistical analyses and associated training data collection to identify desirable genomes either in silico (when the program is used to predict a desirable genome that is not already a known genome), in actual existing genomes, and combinations thereof and then use those genomes to make successive crosses (either in silico or physical), while generating new phenotypic, metabolic, proteomic, allelic, genomic, expression, environmental, economic data or combinations thereof that can be used to revise training sets used in the statistical algorithms. It is further understood that the use of the word accelerated refers to the speed of the number of generations of genomes that can be identified in a specific time frame. The term accelerated is used to refer to the comparison to other breeding programs that operate more traditionally with selection of plants based upon their field performance and the crossing of those that were selected and also can refer to other breeding programs that partially operate in data driven cycles, but yet do not achieve sufficient progress. The speed of an accelerated breeding program can be referred to in cycles per year, Typically, an accelerated breeding program has two or more, three or more, four or more, five or more, six or more, seven or more, breeding cycles per year. The speed of an accelerated breeding program can also be referred to in generation cycles of any given plant species, where generation cycle refers to the amount of time required from germination to post-reproduction collection of suitable material to grow the next generation. Postreproduction, or post-pollination, collection can include VP, or seed harvest, or combinations thereof.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0067] In one example described herein an accelerated breeding program uses more than 1, more than 2, more than 3, more than 4, more than 10, more than 15, more than 20, pools of genomes. Accelerated breeding programs can operate using pools of genomes that can be created and subjected to cycles of improvement such as 1, 2, 3, 4, 5, 10, 20, 30 or more cycles / year and at any point in the cycle one or more genomes can be selected and directed towards commercial development, data generation for accuracy purposes, added to a different pool of genomes, or combinations thereof. Some pools will be continuously subjected to cycles of improvement, for example pools that are selected for based upon genetic gain and diversity. A pool can include more than 10 distinct genomes, more than 30, 40, 50, 70, 100, 150, 200, 500, or more than a 1000 distinct genomes. A pool is selected based on specific selection criteria captured by a selection index. A selection index weighs multiple traits of economic importance into a single value that allows one to select for multiple traits at the same time, placing more emphasis on those of greater importance, and avoiding unfavorable correlated responses to selection. Typically in commercial grain production crops, yield carries by far the largest weight since this is crucial to profitability. Multiple selection indices exist, each aimed at a specific market or biological type.

[0068] In a specific example, a genome pool is maintained comprising genomes that include recombinant sequences that impart resistance to one or more insects and / or tolerance to one or more herbicides. That first genomic pool is cycled to improve genetic merit while maintaining the heterologous sequences (see description of heterologous as it relates to pools provided herein) and their associated functionality. This first pool can continuously cycle incrementally improving cycle after cycle, generation after generation.

[0069] A second pool is created with a new trait of interest inserted, such as a trait that provides resistance to a disease, a class of insect, an environmental condition or combinations thereof. That trait can be derived from heterologous nucleic acid sequences, native nucleic acid sequences (from the same species that the genomic pool is derived from), through gene editing of endogenous sequences, interfering nucleic acid sequences (RNAi, antisense and the like) or combinations thereof. This second pool can cycle for 1, 2, 3, or more generations using genome wide selection (GWS) to select crosses for genetic improvement using GEBV.

[0070] At any point one or more genomes can be identified for the creation of inbred lines. This process can be referred to as line derivation and one of ordinary skill in the art will appreciate thatPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO inbred lines can be created through various backcrossing strategies, successive self-pollination (i.e., selfing) or through double haploid production.

[0071] Genomes identified for crosses between the first pool and the second pool can be made, and the selected genomes can be crossed directly or initially subjected to line derivation prior to crossing.

[0072] Various aspects of accelerated breeding programs are known in the art. Many aspects are generally described by Anilkumar et al, 2022, Advances in integrated genomic selection for rapid genetic gain in crop improvement: a review, Planta, Vol 256 (https: / / doi.org / 10.1007 / s00425-022- 03996-y), which is herein incorporated by reference in its entirety.Data collection and timing

[0073] In modem plant breeding, the ability to collect phenotypic and trait data in real time as a plant grows is a fundamental skill for breeders. This approach enables breeders to make informed selections, adjust breeding strategies, and feed accurate data into predictive models, such as those utilizing ML. The real-time collection of phenotypic and trait data is not only essential for immediate decision-making but also for training and refining statistical and ML models in an accelerated breeding program with continued cycling pools of genomes.

[0074] Breeders utilize a variety of tools to streamline and standardize data collection, such as handheld devices, field tablets, digital imaging systems, and integrated data management platforms. These tools enable breeders to capture data in the field, sync it to centralized databases, and link phenotypic observations with genotypic and environmental information. Data can be pre- processed, normalized, and used directly or as input for further analysis and model training. The following describes the understanding and practical methods employed by one of ordinary skill in the art for the collection of plant phenotypes and traits at various growth stages, from germination through harvest and post-harvest analysis.

[0075] The process begins at the earliest stage of seed germination. As soon as seeds are sown, breeders monitor for signs of germination, which typically involves the emergence of the radicle or shoot from the seed coat to emerge from soil in pots or fields. Germination data may include the time to germination, percentage of seeds germinating, uniformity of emergence, and vigor of the seedlings. This data is collected soon after the seeds begin to grow, often within days ofPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO planting, and can be recorded manually or with automated imaging systems. Early germination data is crucial for assessing seed quality, viability, and initial plant establishment. It is well known amongst plant breeders that seed storage conditions, such as temperature and humidity, and timing of storage are also factors that can affect germination.

[0076] As the plants continue to grow, breeders collect a wide array of phenotypic traits throughout the season. These traits can include, and are not limited to plant height and growth rate, leaf traits (e.g., area, shape, and color), developmental stages (e.g., flowering time, pod initiation), disease and pest resistance, often observed by scoring visible symptoms or using sensors, environmental response traits, such as drought or heat tolerance, biomass accumulation and canopy architecture, root traits (sometimes assessed through destructive sampling or imaging). Data collection methods range from manual scoring, digital photography, and sensor-based measurements (e.g., drones, UAVs, soil and climate sensors) to automated platforms that record quantitative metrics in real time. This continuous data flow allows breeders to track the progression of individual plants and populations, enabling early identification of superior genotypes or those requiring further attention.

[0077] At the end of the growing cycle, breeders focus on yield, seed moisture, fruit quality, and post-harvest traits, which are critical for evaluating the commercial potential and overall success of breeding candidates. Key traits collected at this stage include, but are not limited to seed or fruit yield per plant or plot, harvest index (ratio of yield to total biomass), quality traits such as seed size, shape, moisture content, oil content, protein content, and nutritional composition, postharvest disease resistance or storability, processing traits (e.g., ease of threshing, milling quality). These data are typically gathered using harvesters equipped with yield monitors, laboratory analyses for quality traits, and post-harvest storage trials. Yield data is often the primary driver for selection, but increasingly, breeders incorporate a suite of post-harvest traits to ensure that new cultivars meet the needs of farmers, processors, and consumers.

[0078] The association of the collected data with the genetic information about the plants, the location of the plants, predicted viability of regenerable material from the plants, and the availability of appropriate breeding pairings, including nicking times, for the plants needs to be managed to allow for continuous cycling of the populations. Databases and tracking mechanismsPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO(bar codes, RFID tags, and the like) that are well known in the art can be used to assist in the management of this data as well.Data usage and model training

[0079] Accelerated breeding programs rely on a variety of different data types, environmental, phenotypic, genomic, metabolic, expression and combinations thereof. The quality of the data will have a significant impact on the quality of predictions made by the ML algorithms. Generally, the statistical model that is developed to produce a prediction or estimation is trained on, or refined, using a training pool of data which is a real population of plants for which genotypic and phenotypic data is available. That data is then used to develop the model and the model is applied to a population, either an in silco (theoretical) population, actual population of existing genomes or combinations thereof. One of skill in the art will appreciate that validation of the model can be achieved using a validation population, for which both the genotypic and phenotypic data is known. The model can be applied to that population to determine the accuracy of the GEBVs as compared to the population’ s observed phenotype value. Models with high validation accuracy are applied in production to predict the genetic merit of individuals that may or may not have been included in the training pool. Theoretical accuracies for predictions can be estimated based on the model, genetic architecture of the trait, relationships between individuals, the amount and type of information available for prediction, etc.

[0080] Any model or combination of models can be used during the accelerated breeding program. For example, normal distribution models such as, random regression best linear unbiased predictor (RR-BLUP), best linear unbiased prediction (BLUP) and Genomic-BLUP (GBLUP), models assuming markers with higher probability of large effects Bayesian (BayesA), models assuming markers with zero effects (BayesB), non-parametric methods reproducing kernel Hilbert space (RKHS and Random forest) and combinations thereof can be used in the accelerated breeding program. See, Heslot et al, 2022, Genomic Selection in Plant Breeding: A Comparison of Models, Crop Sci, Vol 52 (https: / / doi.org / 10.2135 / cropsci2011.06.0297), which is herein incorporated by reference in its entirety.

[0081] All models require the use of genomic and / or pedigree information to exploit the association between the trait expression and gene content. A relationship matrix is created to describe the similarity and relatedness between all individuals included in the model, whichPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO increases the accuracy of prediction for one individual based on the performance of all included traits / environments for all other related individuals

[0082] Accelerated breeding programs can also aim to simultaneously predict genetic merit for multiple traits and / or environments in the same model, which can be from 2 to more than 30. Each environment and / or trait are treated as separate variates within the same model, hence the term multi-variate instead of purely multi-trait or multi-environment. The same trait in different populations can be considered as a different trait and included in the model as such.

[0083] The implementation of a combined multi-trait multi-environment model in an accelerated com breeding program presents significant scientific and technical challenges, particularly due to the complexity of modeling genetic correlations among traits and substantial genotype-by- environment interactions. While multi-trait multi-environment approaches have demonstrated potential to increase prediction accuracy and genetic gain, their practical deployment is hindered by computational demands and the need for large, high-quality datasets (Mora-Poblete et al, 2023, Multi -trait and multi-environment genomic prediction for flowering traits in maize: a deep learning approach, Front in Plant Sci, Vol 14 (https: / / doi.org / 10.3389 / fpls.2023.1153040)). Recent studies show that, although advanced statistical models can improve prediction accuracy, their application is often limited by data availability and computational resources, which are not always accessible in practical breeding programs (Lozano et al, 2023, Regularized multi -trait multi -locus linear mixed models for genome-wide association studies and genomic selection in crops, BMC Bioinformatics, Vol 24 (https: / / doi.org / 10.1186 / sl2859-023-05519-2)). Furthermore, even the latest multi-trait multi-environment methods struggle with scalability and transferability across diverse populations, and often fail to fully exploit genetic correlations when training data are limited, or environments are highly heterogeneous (Cuevas et al, 2017, Bayesian Genomic Prediction with Genotype x Environment Interaction Kernel Models, G3 Genes|Genomes|Genetics, Vol 7 (https: / / doi.org / 10.1534 / g3.116.035584)).

[0084] Environmental data can be integrated with phenotypic and genetic data using multi-trait multi-environment models such as described in Gill et al, 2021, Multi-Trait Multi-Environment Genomic Prediction of Agronomic Traits in Advanced Breeding Lines of Winter Wheat, Front in Plant Sci, Vol 12 (https: / / doi.org / 10.3389 / fpls.2021.709545). The multi-variate genetic evaluation model is designed to include fixed and random effects. Fixed effects may include the average ofPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO each trait, and year-location. Random effects may include GCA for each parental line, SCA for respective hybrid progeny of any two parents, the permanent-environment effect, and the residual. The types of models that may be used include Bayesian approaches, random regression models, mixed linear models, ridge regression models, threshold models, survival models, ML, or other similar model types.

[0085] The methodology used in accelerated breeding programs can include elements of genetic evaluation, genomic prediction, optimal contribution selection (OCS), imputation, prescriptive testing, and end-to-end connectivity. This overall framework of elements (i.e., quantitative genetics framework) involves the use of combinations of various classes of actual data collected through genomic sequencing, phenotyping, geopositional information, environmental data, pedigree data, and combinations thereof, as well as data derived from predictions about such actual data classes.

[0086] One of ordinary skill in the art will appreciate that models used in the accelerated breeding program can be trained using a variety of data, for example genome sequence information including single nucleotide polymorphisms (SNP), quantitative trait loci (QTL), RNA sequence information (RNA-seq), short read genomic sequencing, marker data, long read genome sequence information, methylation status, gene expression values, insertions / deletions (indels), and combinations thereof.

[0087] Environmental data, including spatial data at both a micro and macro level can also be used to refine the accelerated breeding program’s genomic selection models. One of ordinary skill in the art will appreciate that use of multistage analysis is also possible. See for example, Bemardeli et al, 2021, Modeling spatial trends and enhancing genetic selection: An approach to soybean seed composition breeding, Crop Sci, Vol 61 (https: / / doi.org / 10.1002 / csc2.20364), which is herein incorporated by reference in its entirety.

[0088] One of ordinary skill in the art can appreciate that the tension between needing genetic diversity and rounds of selection driving towards, for instance yield gain need to be managed in an accelerated breeding program via approaches such as OCS. Multiple software packages are available to apply OCS, including MateSel (Kinghorn et al, Instructions for MateSel accessed 8 November 2025 at https: / / matesel.com / content / documentation / MateSelInstructions.pdf).PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0089] The step of imputation is performed using probabilistic models and a reference-based or pedigree-based method to infer un-observed genetic markers of an individual based on the observed markers of the individual and linkage disequilibrium patterns in a genotyped reference population or relatives.

[0090] Exemplary accelerated breeding programs include for example those described in, Kishore et al, 2011, Methods for increasing genetic gain in a breeding population, WO2012075125A1, which is herein incorporated by reference in its entirety.Genomic sampling

[0091] Any tissue that is capable of providing a nucleic acid sequence sample (i.e., RNA, DNA, combinations thereof) can be used to collect the data desired for the accelerated breeding program. One of ordinary skill in the art will appreciate that seed chipping can be performed such that the seed retains viability but yet is used to provide a genetic sample, moreover, if desired in the case of a paternal haploid system the seed can be used to provide both the maternal genetic sample (i.e., maternal tissue such as pericarp) and a sample of genetics from the pollen (paternal genome). In some cases it will not be desirable to wait until seeds are set and any other tissue can be used to extract the genomic sample during any stage in development. One of skill in the art will appreciate that such sampling will be done based upon the design of the accelerated breeding program and the variety of plant that is being bred. One of ordinary skill in the art will appreciate that there are a variety of methods of collecting genomic information. These methods include using marker platforms such as high-density marker platforms (e.g., Infinium™ chips, TaqMan and the like), various next generation sequencing techniques, combinations of single nucleotide polymorphism sequencing and sequencing, and any other method known in the art that provides adequate genomic data for the various models used to make selections during the breeding cycles.

[0092] One of ordinary skill in the art will appreciate that algorithms can be used to estimate the optimal % of testing needed to find the desired genomes. The number of seeds per parental group to genotype and be available for selection is determined by an optimization algorithm taking multiple factors into account, including but not limited to, operational limitations, financial resources allocated, population size targets, Mendelian sampling variance and genetic merit. This could result in 100% of the seeds per parental group to be genotyped, or alternatively, a subset of the parental group could be genotyped. Different genotyping platforms can be used for thePCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO population, depending on the use thereof, including, but not limited to, trait presence detection, parentage verification, imputation accuracy, prediction accuracy, and genetic contribution to the population. In some of the methods described herein, the plants in a pool are genotyped using a low density sequencing technique and then parents that are selected for crossing can be subjected to high density sequencing.Regenerative material

[0093] As used herein regenerative material refers to plant material that can be used to regenerate a plant comprising a genome (e.g., 2n or diploid) or half of a genome (e.g., In or haploid). Regenerative material can be used in the context of an accelerated breeding program to increase the availability of a genome or half of a genome. In some instances the haploid regenerated material can be induced to produce a double haploid (i.e., an inbred or homozygous diploid). In the context of an accelerated breeding program, the use of regenerative material can be useful to increase the availability of a desired genome for; field testing, particularly when replicate data is desired (increase in N), simultaneous use in multiple genomic pools, simultaneous use in an outbreeding scheme outside of a cycling genomic pool and yet still continuing in the cycling of the pool.

[0094] The vegetative propagation (VP) methods and tiller induction methods described herein can also be used in combination with any breeding program to expand the amount of genetically identical plants and associated seeds available for use in any workflow that is used in a breeding program. One of ordinary skill in the art will appreciate that during haploid induction or in certain stages of breeding monocots or dicots the ability of a given plant to produce enough desirable seed for the next stage in the breeding program is not sufficient. For example, in doubled haploid production it may be that only one or two seeds are produced on a haploid plant (DHO). Using vegetative propagation or tiller induction in combination with the one desired seed haploid plant will provide enough genetic material to advance the program.

[0095] In one example, a pool of genomes is being cycled to increase GEBV and one or more genomes are identified as desirable for continuation in the pool and also for further use outside of the pool. In an instance where the genome is a dicot, such as a soybean genome, the desirable plants are used as donor plants and one or more cuttings are taken and one or plants are generated from the cuttings. The newly generated plants can then be used in testing to provide new data thatPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO can be directly used to update the model used in the cycling of the pool and the identical clones can continue in the pool of genomes being continuously improved. The developmental timeline of the newly generated plants can also be exploited to potentially feed the genome back into the pool at a later cycle while minimizing any delay in the breeding program. Furthermore, the availability of the newly generated plants allows them to be transported and used in various geographies and in other breeding programs without the need for seed increase rounds or waiting through a whole growth cycle.

[0096] Regenerable material also includes taking tissue samples, including immature embryos or mature seeds, and forming calli and regenerating plants. One of ordinary skill in the art will appreciate that technique used for the formation of calli and regeneration of plants will largely depend upon the type of plant. See, Bridgen et al, 2018, Plant Tissue Culture Techniques for Breeding, Handbook of Plant Breeding, Vol 11, 127-144, Springer, Cham(https: / / doi.org / 10.1007 / 978-3-319-90698-0_6).

[0097] Other sources of regenerable material include pollen. Similar to the example described above, desirable genomes from a population of plants can be selected and the pollen from those plants can be collected. The pollen can be divided into 2, 3, 4, 5, or more portions and each portion can be used for distinct purposes. In a particular example, portions of the pollen can be used to pollinate 1, more than 1, more than 2, more than 3, more than 4, more than 10, more than 15, more than 20 plants in the genomic pool from which it came (pool of origination) and / or 1, more than 1, more than 2, more than 3, more than 4, more than 10, more than 15, more than 20 plants in any other genomic pool(s). In a particular example, a portion of the pollen can be used to continue the cycle in the genomic pool from which it came (pool of origination) and another portion can be used to fertilize a maternal plant that induces paternal haploid production. The paternal haploids can be converted to diploids using methods known in the art and inbred lines comprising genomes that are homozygous for the original desired genome from the continuous cycle are efficiently created. Moreover, the double haploid plants can be used to generate data that is used to refine the continuous cycle of the pool of origination. Similarly, a portion of pollen can be used to directly pollinate a plant having known GEBV and data from the resulting outbred can be used to further refine the ML models used in the accelerated breeding programPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO

[0098] Regenerable material is additionally useful for creating experiments where it is desired to control the variability of the genomes used in the experiments. One of skill in the art will appreciate that through the process of meiosis the pollen population in a plant and maternal reproductive tissue is not genetically identical, however, use of a single source for some experimentation can be desirable. Particularly if data from that experiment is used to educate ML algorithms.

[0099] The use of pollen and PHIS can also be particularly useful in genomic editing. The haploid induction line can additionally be transformed with a genome editing system. For example, nucleic acid sequences encoding gene editing amino acid sequences such as the group consisting of a polynucleotide-guided endonuclease, CRISPR-Cas endonucleases, base editing deaminases, zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), engineered site-specific meganuclease, or Argonaute. In instances where the induction line comprises a genome editing system the result can be edited genomes that can quickly be induced to form edited double haploids that do not contain the recombinant nucleic acid sequences encoding the genome editing system. See for example, Armstrong et al, 2019, Methods and compositions for genome editing via haploid induction, US11401524B2.

[0100] The use of regenerable material, particularly through VP, also includes the ability to establish differential treatment experiments to determine the performance of selected genomes under different treatments. For example, plants from regenerable material can be grown under various environmental (e.g. field or controlled) conditions, subjected to gene editing schemes, grown in different geographies and combinations thereof. The results of such experimentation can be further used to refine the accelerated breeding program. Such experimentation could be accomplished through traditional breeding methods, such as selfing, and recurrent selection to obtain inbred lines and then using that material in experimental designs, but those processes do not operate at the speeds necessary to continually improve the breeding program.

[0101] One of ordinary skill in the art will appreciate that the use of regenerable material as described herein provides the advantage of maintaining a high genomic similarity between the genomes that are continued in the genomic pools and the genomes that are used to gather data that is useful in subsequent rounds of selection. Thus, leading to increases in the accuracy of model predictions.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WODefinitions

[0102] As used herein, a “genome editing system” refers to any system (e.g., a meganuclease, a ZFN, a TALEN, a CRISPR / Cas9 system, a CRISPR / Cpfl system, a recombinase, a transposase), or a combination of genome editing systems known in the art, that is used in a method to introduce one or more insertions, deletions, substitutions, or inversions to a locus in a cell to generate a dominant negative allele or a dominant positive allele.

[0103] As used herein, “educate” in the context of education of a model can be a cyclical and adaptive process. This process begins with one of more training and / or validation data sets which includes one or more input data sets and may result in iterative adjustments to secure a desired model performance. The educated model may be retrained at one or more regular or irregular intervals as additional data is available. A user is said to educate a model by interacting with the model during model creation, optimization, validation, and / or training / retraining of said model.

[0104] As used herein, “gathering data” refers to any method or systematic collection of data that can be performed by various tools and can encompass various data types including but not limited to satellite imagery, drones (UAV) imagery, aerial photography, sensors, machine and equipment data (such as tractors, harvesters, irrigation systems, etc), sampling data (soil, water, plant samples and measurements, etc), application data (fertilizer, pesticide, etc), field surveys (disease and pest data, soil data, etc), digital platforms and apps (such as Climate FieldView™, etc), grower networks (including many farmers, agencies such as USDA and FAO, etc). The data can be used in the form as collected / obtained and can also be pre-processed (filtered, normalized, standardized, etc) and be the output of other calculations, models, or estimations.

[0105] As used herein, an “outbred” refers to any highly heterozygous individual, such as at least 10%, 12.5% (e.g., F4 generation), 25% (e.g., F3 generation), 50% (e.g., F2 generation), or 100% (e.g, Fl generation) heterozygosity. The generations and heterozygosity estimates provided are based upon two highly inbred lines from distinct heterotic groups being used to create the Fl generation. In the accelerated breeding programs described herein the heterozygosity of the generations will be corresponding less than these numbers (see FIGS. 9A-9D) and depend upon the parental makeup. One of ordinary skill in the art of plant breeding will appreciate continued crossing of heterozygous individuals, maintained by crossing unrelated individuals or less relatedPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO individuals in the population, can both introduce new genetic material and increase genetic variability in a pool of genomes.

[0106] As used herein, “pollen priming” is used to described treatments to the pollen prior to pollination which induces a selection event upon the treated pollen. Examples of a treatment include increased temperatures, decreased temperatures, increased or decreased pollen moisture content, increased or decreased pressure, chemical exposure, abrasion upon the pollen or combinations thereof. These examples relate to environmental conditions such as cold, freezing, drought, and heat stress which are all variable from location to location and season to season making it extremely difficult to select for these abiotic stress tolerances under a typical breeding pipeline. The priming method is to tap into the large pool of genetic diversity and to select for phenotypes under selected treatment conditions that are otherwise difficult to capture and accurately assess in field trials within a modern breeding program due to high rates of variation in the field conditions, making selection of said phenotype a lengthy and inefficient process.

[0107] As used herein, “pool or population” refers to a grouping of genomes (actual and predicted) that are used to make predictions using machine learning (ML) algorithms. A given pool can be developed to target one or more desired characteristics, through cycles of predictions, actual crosses, in silico crosses, and combinations thereof. The characteristics targeted can include yield, optimization for various environments, disease resistance, pest resistance, and combinations thereof. A heterogeneous pool of plants (as defined collectively as a population by methods known in the art, such as the VanRaden method) as used herein refers to a pool of plants that upon selection by the ML algorithm are used to create crosses and further optimize for the desired characteristic(s). Heterogeneous used in this context refers to the in-breeding coefficient as calculated by the VanRaden method 1 from VanRaden 2008, Efficient Methods to Compute Genomic Predictions, J Dairy Sci, Vol 91, 11:4414-4423 (https: / / doi.org / 10.3168 / jds.2007-0980), herein incorporated by reference as reflected as an average of the relationships in the population. Using the VanRaden method 1, the pools described herein will be sufficiently genetically diverse to allow for improvements in desired characteristics within the population (see FIG. 10) depending upon recombination. If the members of the population are too similar changes / improvements may not be observable. For example, a pool of plants will have an average co-ancestry value of lessPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO than 2.0, less than 1.5, less than 1.0, less than 0.7, less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2 or less than -0.5.

[0108] In additional examples, a population that is being used to select crosses can also be maintained not only based upon its population diversity (i.e., the VanRaden method 1 described herein), but also by the average heterozygosity of the individuals maintained in the population. For example, the average individual heterozygosity of a cycling population can be from about 10%- 90%, 10%-80%, 10%-70%, 10%-60%, 10%-50%, 20%-90%, 20%-80%, 30%-90%, 30%-80%, 30%-70%, 40%-60%, and ranges within these ranges. The relative heterozygosity can be specifically controlled to optimize precision of the predictive models used in relation to the training data used, for example, outbred testing or inbred testing.

[0109] A population that is maintained over time and that parents from the population are selected and intercrossed, generation after generation, is referred to as a cycling population, cycling pool, or continuous cycling. One of ordinary skill in the art will appreciate that periodically to maintain genomic diversity and / or to introduce one or more traits of interest, for example gene edited traits, transgenic traits, or native traits, new genomes will be added to the population.Examples

[0110] The present disclosure is further illustrated in the following Examples. It should be understood that these Examples, while indicating embodiments of the invention, are given by way of illustration only. Thus, various modifications to the crop model, the relationships to simulate / model the limited transpiration trait, methods of analyses, and applying such methods for crop improvement are disclosed.Example 1. Increasing Speed of Data Return to Increase Precision of ML Models[OHl] One of ordinary skill in the art will appreciate that 2 key tactics to improve the prediction accuracy of ML models used in breeding include, but are not limited to: 1) increasing the speed at which data is returned to the model for re-training (also referred to as educating or updating); and 2) using genomes with high similarity to the cycling population to collect the data (Bernardo 2025, Why does genomewide prediction become ineffective after several cycles of recurrent selection?, Crop Sci Vol 65 (https: / / doi.org / 10.1002 / csc2.70164)). While these two tactics involve differentPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO practical methods and steps they are both aimed at the goal of returning data from genomes that are not too genetically distant from the then current genomes in the cycling population.

[0112] To increase the speed of data return direct use of the genomes extracted from the cycling population is performed. The extracted genomes are crossed and the collection of useable data to the ML model is achieved in less than 3 years, 2 years, or less than 1 year from extraction from the cycling population. The direct use of genomes from the cycling population for data collection is herein referred to as outbred testing. Genomes selected for extraction from the model can be extracted by removing seeds (siblings or half siblings) from the cycling population. The seeds are grown and crossed with one or more substantially homozygous lines (i.e., inbred or double haploid) (See FIGS. 4A and 4B).

[0113] The data from the outbred testing can be combined with data collected from double haploid testing as well as all of the other data types mentioned herein.

[0114] Similarly, the speed of data return can be increased by using pollen from selected genomes in the cycling population to cross with one or more substantially homozygous lines (i.e., inbred or double haploid). This usage, without creating a more homozygous generation through haploid process or selfing is also herein referred to as outbred testing. Data from the progeny generation is then used to update the ML model.

[0115] The progeny resulting from the crosses and used for data collection will not be fully heterozygotic due to the heterozygosity of the extracted genome (seed) and the allelic population found in the pollen. However, the speed of being able to return field data to the cycling population results in the cycling populations only having advanced less than 5 cycles, or less than 4 cycles, or less than 3 cycles before the ML model is updated.

[0116] Siblings of cycling pool selected parents are theoretically 50% related since parents share 50% of their genetic material with offspring, but the true relatedness can vary widely. Statistically, 95% of full siblings have been shown to have a genetic relatedness from approximately 42% to 57% (Visscher et al, 2006, Assumption-Free Estimation of Heritability from Genome-Wide Identity-by-Descent Sharing between Full Siblings, PLOS Genetics, Vol 2 (https: / / doi.org / 10.1371 / journal.pgen.0020041)). Since individuals are not fully homozygous, the allelic combination (AA, AB, BA, BB) similarity may be considerably lower. Thus, by chance, wePCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO might select individuals for outbred testing that are rather lowly related to their siblings advancing to the next cycle in the population (see FIG 3B).

[0117] When pollen haploid inducers are used, pollen is taken directly from cycling pool individuals to mate to an inducer. There is recombination step before induction but 100% of the genetic material of the DH comes from an individual used as a cycling pool parent, thus no alleles are foreign, although allele combinations (AA, AB, BA, BB) may be different. If the cycling pool individual is 75% homozygous, the DH will be at least 75% genetically identical in terms of allelic combinations to the advanced cycling pool individual, even after the recombination event. The resulting DH will also be fully homozygous, which will capture more across-plot and within-plot heritability (Bernardo 2020, Breeding for quantitative traits in plants, 3rdedition, Stemma Press (https: / / bernardo-group.org / wp-content / uploads / 2019 / l l / BQTP3e_Sample_Pages.pdf)). Using PHIS ensures that, on average, individuals tested in the field to increase accuracy will be more related to the individuals in cycling pools, as well as the individuals that will eventually make it to the market. Additionally, the DH individuals tested can directly go into the standard breeding workflow and result in parents of commercial products.Example 2. Outbred Testing Usage in Accelerated Breeding

[0118] In a retrospective study, outbred testing was used from two exemplary cycling pools to quickly return training data to cycling population ML models in an accelerated corn breeding program. In an accelerated breeding program it is typical to maintain two distinct heterotic groups in cycling populations. In some instances, these can be referred to as female and male heterotic groups. The heterozygosity levels of these groups collectively from one exemplary cycling pool is shown in FIG. 9A.

[0119] Genetic diversity of the cycling pools is maintained by additional cycling pools that have a relatively higher heterozygosity as shown in FIG. 9B. Periodically, individuals from the diversity pool are incorporated in the crossing of the accelerated breeding pool.

[0120] In this example, individuals were extracted from each of the two distinct heterotic groups in the accelerated breeding program (FIG. 9A) and selfed (i.e., female heterotic group 1, male heterotic group 2). The distribution of resulting heterozygosity is shown in FIG. 9C. ForPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO comparison the heterozygosity of DH lines created from each of the female and male cycling pools are shown in FIG. 9D. Because the creation of DH lines via MHIS is a lengthier process, there are 6 or more cycling generations of time between the DH lines and the selfed individuals.

[0121] FIGS. 10A and 10B shows a comparison of the genomic relationship among the various resulting progeny. Within the two cycling pools, for heterotic groups 1 (female, see FIG. 10 column 1) and heterotic group 2 (male, see FIG. 10B column 1), the genomic relationship among individuals as measured by the VanRaden method 1 of ranges between approximately 1-0.75 (column 1, FIG. 10A) and 0.65-0.51 (column 1, FIG. 10B), respectively. The genomic relationship between the extracted selfed individuals combined with the population from which it was chosen from their respective pools they originate is shown in column 2 in FIGS. 10A and 10B, respectively. Similarly, the genomic relationship of the extracted double haploid lines tested in the same year and combined with the respective cycling population from which from their respective pools are shown in column 3 in FIGS. 10A and 10B, respectively.

[0122] The genomic relationship of the two diversity pools (Diversity Pool 1 and Diversity Pool 2) and the respective cycling populations_is shown in column 4 of FIGS. 10A and 10B.

[0123] The genetic diversity of the two representative pools (Diversity Pool 1 and Diversity Pool 2) in combination with the respective cycling pools is shown in column 4 of FIGS. 10A and 10B.

[0124] When data was returned, the elite cycling pool had been progressed four generations (i.e., now the elite cycling pool generation 5). To evaluate the impact of outbred testing in the genetic evaluation of elite cycling pool generation 5 predictions of selection candidates was done using the GBLUP model with field testing data from only the DH hybrids and compared to using the field testing data of the DH hybrids combined with the Outbred elite hybrids (Table 1). Across seven traits, including yield, plant height, etc. the accuracy advantage of educating the GBLUP model with the combined data from DH hybrids and Outbred elite hybrids compared to only the DH hybrids was found to be an average 3.1% accuracy increase for Heterotic Group 1 and an average 1.2% accuracy increase for Heterotic Group 2.PCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WOTable 1. Prediction accuracy advantages for seven traits from heterotic groups 1 and 2 comparing use of field testing data from a combination of DH hybrids and Outbred elite hybrids to only DH rybrids.

[0125] The results show that outbred genomes can be used directly to generate data for updated training data. This example, combines data from double haploids and outbreds to achieve superior predictions. One of skill in the art will appreciate that the data from the outbreds can be obtained faster than the double haploid data and therefore maintain higher genetic relatedness to the cycling pool. Moreover, the data from outbred testing and double haploid testing can be mixed such that the extraction from the cycling population (pool) were taken from different generations.Example 3. Paternal Haploid Usage in Accelerated Breeding

[0126] An accelerated com breeding program utilizing a paternal haploid inducer system (PHIS) can enable training data to be returned lyr faster compared to a maternal haploid induction system (MHIS) that is typical in a modern com breeding program (See FIG. 2). Multiple genetic mechanisms for both maternal and paternal haploid induction systems have been described (Song et al, 2024, Haploid induction: an overview of parental factor manipulation during seed formation, Front in Plant Sci, Vol 15, (https: / / doi.Org / 10.1016 / j.tplants.2021.02.010)) and as depicted in Widiez 2021, Haploid Embryos: Being Like Mommy or Like Daddy?, Trends in Plant Sci, Vol 26, 5:425-427 (https: / / doi.Org / 10.1016 / j.tplants.2021.02.010). Modem corn breeders have usedPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO maternal haploid induction in general due to higher efficiency doubling rates in generating DH lines but at the expense of added time in the overall breeding program progress to reach the fieldtesting stage of DH lines.

[0127] The PHIS process utilizes the pollen of a target individual to generate haploid offspring by crossing an inducer line, which in subsequent generation can be used to generate a dihaploid line by doubling. Unlike the MHIS process, the induction process is started using pollen from the cycling population individuals, which means the haploid pollen source from a given origin can be used for line derivation as well as continuous mating within the cycling pool process. Furthermore, the pollen used to derive lines can be stored and used in future mating schemes concurrently when the derived line is evaluated phenotypically in the field. This process can increase the accuracy of prediction in multiple ways, for example: 1) the pollen from which the line is derived and tested is itself used as a parent in cycling pool crosses, and therefore contributes to increased accuracy due to higher degree of relatedness between individuals (possibly within 3-6 generations) as opposed to 6-8 generations of lines from MHIS process (see FIG. 2); 2) when the stored pollen (from which the line is derived and subsequently phenotyped) is used as a future parent in the cycling pool mating scheme, it can close the gap in relatedness between the generations from which the phenotypic (training) data is collected and continuous cycling generation, thus increasing the overall accuracy of prediction by virtue of increased relatedness and also slows down the decay of long term prediction accuracy (Bernardo 2025, Why does genomewide prediction become ineffective after several cycles of recurrent slection?, Crop Sci Vol 65 (https Z / doi . org / 10.1002 / csc2 , 70164); Habier et al, 2010, The impact of genetic relationship information on genomic breeding values in German Holstein cattle, Genetics Selection Evolution, Vol 42 (https: / / doi.org / 10.1186 / 1297-9686-42-5)); and 3) The doubled individuals that are tested are highly homozygous, requiring fewer individuals to achieve the same level of accuracy as in outbred testing and have potential for commercialization one year earlier than the current standard process.Example 4. Creation of Male Haploid Induction line

[0128] Paternal haploid induction lines can be created using any method known in the art. For example, modifications can be made to the expression level and / or amino acid sequence of thePCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WOCenH3 gene in the donor plant resulting in retarded CenH3 activity. In certain examples, the CenH3 gene can be controlled using an inducible promoter such that the haploid induction capability is controllable by the inducible conditions. In yet other embodiments, the CenH3 protein may be degraded or impaired in a gamete, for example by gamete specific expression of an engineered E3 ligase that targets the CenH3 protein for degradation. In yet other embodiments the CenH3 promoter can be altered to become inactive or less active. Another method of creating a haploid inducer line is through frame shift, amino acid substitution or deletion of the CenH3 amino acid sequence.

[0129] Moreover, after an initial haploid induction line is created the mechanism of induction can be transferred to other genomes that may be useful in the breeding program.

[0130] One of skill in the art will appreciate that PHIS allows for the creation of a lines from pollen extracted from the cycling pool population. The selected pollen is used to fertilize a corn plant that is capable of producing paternal haploids. The paternal haploid induction line can be made using any method known in the art, for example, methods utilizing alteration of the centromere histone H3 activity / expression or mutations to the igl gene. See, Wang et al, 2019, Centromere histone H3- and phospholipase-mediated haploid induction in plants, Plant Methods, Vol 15 (https: / / doi.org / 10.1186 / sl3007-019-0429-5). Alternatively, the CenH3 protein can be targeted for degradation using any method known in the art. Such a system was successfully designed and described in Demidov et al, 2022, Haploid induction by nanobody -targeted ubiquitin- proteasome-based degradation of EYFP-tagged CENH3 in Arabidopsis thaliana, I Exp Botany, Vol 73 (https: / / doi.org / 10.1093 / jxb / erac359).

[0131] Y et another method of impairing CenH3 function in a gamete was described in Maheswari et al, 2014, Generation of haploid plants, WO2014110274A2.Example 5: Stored and Used Pollen for Paternal Haploid Production

[0132] This example describes the creation of paternal dihaploid plants from pollen that has been stored and used to pollinate a PHIS line.

[0133] Pollen was collected by placing tassel bags over the tassels when pollen was beginning to shed from the tassel. The next morning, pollen was harvested by bending the tassel down and vigorously shaking the tassels and bags to release pollen. The pollen was poured from the tasselPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO bags through a filter to remove anthers and other debris. A subset of the collected pollen was destructively sampled to determine the moisture content of the pollen via a halogen moisture analyzer. The remaining pollen was then placed into a container with a fine filter on the bottom that the pollen was too large to pass through and was subjected to a drying gas until a final moisture content of approximately 12.4% was obtained. The dried pollen was then put in a cryovial and placed in liquid nitrogen. The cryovial with pollen was stored in a cry opreservation freezer until it was time to apply. The pollen was thus primed using desiccation and temperature exposure.

[0134] The container(s) of pollen that is stored can then be divided into smaller portions and diverted to different geographies, used in different breeding schemes for the purpose of commercial line production, data collection or continuation in an accelerated breeding program. One of skill in the art will appreciate that by portioning the pollen it can be used for some or all of these purposes simultaneously. The pollen can also be stored and used in a time shifted pattern to pollinate plants in later cycles of an accelerated breeding program. Finally, the portions of the pollen can be primed prior to being used to pollinate plants.

[0135] To create a population of haploids from a hybrid parent, stored pollen was applied to PHIS plants grown in a greenhouse. The PHIS plants had previously been genetically engineered to alter the CenH3 gene such that the maternal chromosome formation was disrupted and a percentage of the surviving seeds contained only the genetics from the pollen.

[0136] The maternal plant was prepared for pollination by placing a wax paper shoot bag on top of the ear when the ear began to develop. This prevented unwanted pollination. The ears were monitored and when silk began to protrude (or just before), the tips of the ears were cut. The next day, many silk were protruding from the cut ear tips. The stored pollen was taken to the greenhouse under cryopreservation conditions. A shoot bag was removed from an ear and pollen was sprinkled onto the silk. The ear bag was placed back on the ear and the plant was allowed to develop to maturity, harvested and shelled.

[0137] Seeds were directly planted in flats containing soil. They were grown without treatment until they were small seedlings. Leaf samples were collected from each plant and put into 96- deepwell blocks. The blocks of samples were processed to extract DNA and the DNA was used to perform typing assays. Specifically, SNP assays were performed for 10 markers distributed throughout the corn genome. 5 markers were selected because the two parents of the hybrid havePCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO different SNPs and 5 markers were selected because the two parents have the same SNP and that SNP is different from the putative PHIS plant. Each of the haploid plants that were identified comprised distinct combinations of genetics from the parental lines (A and B).

[0138] Haploids can be converted to diploids by growing the plants and allowing for self- pollination. Alternatively, diploids can be induced for example using colchicine after they have germinated. This treatment may be performed to recover doubled haploid plants. Other treatments are known to those skilled in the art to allow recovery of doubled haploidsExample 6: Vegetative propagation to enable accelerated breeding program

[0139] One of ordinary skill in the art will appreciate that there are many methods of VP plants. For example, row crops can be VP hydroponically (see, Xu et al, 2013, Vegetative propagation of soybean plants in a hydroponic environment, US20140366441A1, herein incorporated by reference in its entirety), through cuttings (see, Rhanor et al, 2023, Method to produce seeds rapidly through asexual propagation of cuttings in legumes, WO2023192474, which is herein incorporated by reference in its entirety), and through induction of secondary branching (see, Ovadya et al, 2016, Multi-ear system to enhance monocot plant yield, US10772274B2, which is herein incorporated by reference in its entirety).

[0140] Typical VP practice was followed with some variations for soybean VP experiments.

[0141] A shade environment with -100 umol of light was prepared for cutting recovery for up to 9 days. This can be achieved by a temporary tent with shade cloth or plastic sheets. Prior to cutting, prep cutting media accordingly: for Rapid Rooting plugs (RR plugs, supplier info), presoak the cube for >5 mins. For vermiculite or Berger BM7 or equivalent growing media, direct watering in the pots to saturate the media. Germplasms representing diverse maturity groups was used for experiments.

[0142] Prior to each cutting, sharp scissors were cleaned and disinfected with 70% Ethanol spray. Each cutting was trimmed to retain one fully-expanded trifoliate leaf and intact shoot tip. Cuttings were kept in a beaker with clean tap water prior to treatments. Cuttings can be optionally treated with rooting powder, Clonex gel or IB A liquid solution prior to being planted. For Hormex rooting powder or Clonex gel treatment, 1-2 inch basal portion of each cutting was directly dipped into the power or gel. For IB A liquid solution treatment, basal portion of each cutting was soaked inPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WOIB A solution for 15 mins. Inserted the cuttings directly in presoaked RR plugs or pots with vermiculite or BM7, deep enough to cover the bottom node with growing media. Placed RR plugs with cuttings in flat trays filled with enough water to maintain plugs moist. Placed a plastic dome on each cutting tray. Placed the recovery tray under low light shade. The trays were monitored with RH / temp sensors and observed daily. Adjusted venting valve to achieve cutting recovery RH% of about 65%-85%. Managed the temp inside dome to <35C, with misting as needed. Recovery pots / RR plugs were checked and watered as needed daily.

[0143] When new growth and roots were fully established on cuttings (about 7-9 days), cuttings with RR were directly transplanted into 10” pots with BM7 or equivalent potting media. No retransplanting for the cuttings established directly on vermiculite or BM7 growing media pots. Established cutting plants were then grown in regular GH / GC condition with light intensity of 400- 600 pmol. Fertilized the plant with Jack’s professional 20-10-20 soluble fertilizer with every watering with EC (electrical conductivity) of -ImS / cm. Regular IPM practice was followed to manage pest.

[0144] Cuttings can be grown directly in the same pot as donor plant. Donor plant of cuttings (mother plant) were grown individually in 10 inch pots. Cuttings were prepped and treated same as above. Briefly, individual cutting from VI -R1 mother plant was treated with or without rooting hormone and directly inserted into the same donor pot. Or inserted the cutting in presoaked RR plug first and placed the cutting RR plug directly in the same donor pot. Care was taken to ensure the cutting was spaced away from the edge of the pot and the mother plant. Pots were watered immediately and then recovered under shade with low light (-100 umol) with daily monitoring for watering for about 9 days. Pots were then moved out of shade for regular full growth care as covered above.

[0145] VP cuttings caused a few days delay of developmental progression but have similar pod and seeds set potential as the donor mother plants (FIGs 7A-7C).Example 7. Vegetative Propagation Enhancement of Accelerated Breeding

[0146] In an accelerated breeding program simulations have shown genetic gain can be improved by approximately 21% by increasing the number of candidates (Table 1) evaluated in field testing. Plant breeders routinely face operational challenges achieving the desired number of candidatesPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO for field testing due to seed quantity limitations and / or timing constraints. VP helps sustain expected genetic gain when crossing seed targets would not be met in the absence of VP, and it helps exceed baseline genetic gain by enabling a larger number of candidates per pedigree to be available for field testing. In one season of selecting candidates for field testing utilizing VP the increased quantity of seed produced with VP compared to expected seed production without use of VP ranged from 92% to 183% more seed produced across three major candidate pools and 498 total pedigrees (Table 2).Table 1. Simulation of doubling the number of candidates on %genetic gain.Example 8. Identification of germplasm having a desired genotype

[0147] Plant samples, such as plant seeds, plant parts or combinations thereof, suitable for genomic analyses are collected from a range of plants, including one plant up to a population of plants. Techniques known in the art for collecting such samples include leaf punches and seed chipping.

[0148] Desired genotypes are identified for genomic analyses using techniques known in the art such as DNA sequencing (i.e., long reads, short reads, amplicons, and the like), genome wide sequencing, single nucleotide polymorphisms (SNPs), commercially available markers (i.e., TAQMAN® Thermo Fisher Scientific), phenotypic markers, and combinations thereof.

[0149] Genomic information of a given genome from a plant is inferred or imputed from multiple relatives (e.g., known genetic relationships within a family or pedigree), reference genomes, orPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO combinations thereof. Continuous imputation throughout a pedigree and over the course of plant generation cycles exponentially increases the amount and density of genomic data. Statistical models known in the art to predict (i.e., inferring, imputing) missing or a higher density of genotypes from pedigree information, reference genomes, or combinations thereof include Hidden Markov Models (HMM), Beagle (Browning et al, 2018, A One-Penny Imputed Genome from Next-Generation Reference Panels, Am J Human Genetics, Vol 103 (https: / / doi.Org / 10.1016 / j.ajhg.2018.07.015)), Minimac (Fuchsberger et al, 2015, minimac2: faster genotype imputation, Bioinformatics, Vol 31 (https: / / doi.org / 10.1093 / bioinformatics / btu704)), Random Forests and Machine Learning (ML), Bayesian models, K-nearest Neighbors (KNN), or combinations thereof.

[0150] A brief overview of why genomic information enhances selective breeding by identifying desirable traits more accurately is provided in Dirk-Jan de Koning 2016, Meuwissen et al. on Genomic Selection, Genetics, Vol 203 (https: / / doi.org / 10.1534 / genetics.116.189795), which is herein incorporated by reference in its entirety. One of ordinary skill in the art will appreciate that any method of genomic selection can be used in combination with the use of ML models that are trained using data as described herein. An overview of various exemplary genomic selection techniques that can be used in an accelerated breeding program is provided in Anilkumar et al, 2022, Advances in integrated genomic selection for rapid genetic gain in crop improvement: a review, Planta, Vol 256 (https: / / doi.org / 10.1007 / s00425-022-03996-y), which is herein incorporated by reference in its entirety.

[0151] In some embodiments an accelerated breeding program an include training a statistical model for associations between genotypes and traits of interests (i.e., phenotypes) using a set of individuals that are related to selection candidates. The trained model is then used to predict GEBVs of selection candidates using genotypes, imputed genomic information, or combinations thereof with final candidate selection based on their GEBVs. Selected candidates have the most likely superior genomes and are used as parents in crossing schemes (i.e., one to one, one to many, many to many) to generate subsequent progeny that have superior qualities. This process can be continuously cycled generation to generation of plants. The model used for selection and the desired characteristics that direct the population of plants can change, such that for example, a pool of plants is being selected generation after generation for yield, disease resistance, droughtPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO tolerance, maturity timing, seed composition (i.e., oil profile), seed size, flower or fruit size, and combinations thereof

[0152] Genomes were identified in a breeding program by genotyping seeds. More specifically a heterogeneous population of over 14,000 com seeds, having heterozygotic genomes, were sampled via seed chipping. The samples were genotyped with a -50,000 SNP marker Infinium®Illumina panel and genomic information enhanced via imputation using HMMs based on the population and up to 100 or more available corn whole genome reference sequences. GEBVs of each plant were predicted using a global genetic evaluation model (GGEM) with MLM (Mixed Linear Model) equations, then simulated annealing and genetic algorithms informed selection of -7,000 crosses between plants to generate new progeny. Approximately 22,000 com plants were grown for crossing with an average of 150 seeds per plant harvested, creating 3.3M new genomes. The cycle was repeated by again sampling progeny corn seeds via seed chipping. In this cycle, after calculating GEBVs an optional selection step was included to designate crosses to a corn MHIS in addition to crosses between plants of the population to generate the next generation of heterogeneous progeny. Plants crossed to the MHIS enabled production of DHs suitable for field testing to generate phenotypic data to continue training the GGEM and continue for advanced product development. Subsequent generation cycles continued repeating this process.Example 9. In silico predicted germplasm pools

[0153] At any step in a given breeding cycle the genomic information collected can be used to predict the genomic content of a future generation of plants without actually producing plants. Depending upon the predicted genomic outcomes one or more target genomic outcomes can be identified and a prediction can be made as to the probability of achieving that target genomic outcome can be calculated. One of ordinary skill in the art will appreciate that if the probability is too low the number of actual plants that would need to be grown and tested is impractical. However, the target genomic outcome once identified can be used to identify alternative breeding strategies to arrive at the target genomic outcome in an efficient manner.

[0154] An exemplary method includes those taught in Chavali et al, 2018, Methods and Systems for Identifying Progenies for use in Plant Breeding, US20190180845A1, which is herein incorporated by reference in its entirety. In the ‘845 reference, a method is depicted in which data is structured and then predictions are made and progenies are selected for continuation in testingPCT / US25 / 54866 10 November 2025 (10.11.2025)BCS246119 WO or validation. Using the teachings provided herein the same genomes could continue in multiple workflows. For example, a single source of regenerable material can be used to collect genomic data, as well as to produce data for use in modeling future generations.

Claims

What is claimed is:

1. A method of accelerated plant breeding comprising:Selecting at least one first plant from a cycling population of plants using a machine learning model;Making at least two portions of regeneration material from the at least one first plant;Using a first portion of the at least two portions of the regeneration material to cross with at least one second plant in a cycle of an accelerated breeding program;Using a second portion of the at least two portions of the regeneration material to produce plants and generate data; andUsing the data generated to educate the machine learning model, wherein using a first portion of the at least two portions of regeneration material to cross with the at least one second plant and using the second portion of the regeneration material to generate data to educate the machine learning model provides for increased prediction accuracy.

2. The method according to claim 1, wherein the cycling population of plants comprises a coancestry value of less than 1 as calculated by the VanRaden method 1.

3. The method according to claim 1, further comprising tracking the second portion of regeneration material and tracking the environmental conditions the regeneration material is exposed to.

4. The method according to claim 1, further comprising genotyping and using the genotype data in the selecting step.

5. The method according to claim 1, wherein the data generated is used to educate the machine learning model within 1 year of the selection of the at least one first plant for extraction.

6. The method according to claim 1, wherein the data generated is used to educate the machine learning model within 6 months of the selection of the at least one first plant for extraction.

7. The method according to claim 1, wherein the at least two portions of regeneration material are pollen.

8. The method according to claim 7, wherein at least one of the at least two portions of regeneration material that is pollen is primed.

9. The method according to claim 7, wherein the pollen is primed during the bicellular or tricellular state.

10. The method according to claim 7, additionally comprising crossing the pollen that is the second portion of regeneration material with a paternal haploid inducer plant.

11. The method according to claim 10, wherein the paternal haploid inducer plant comprises genome editing components.

12. The method according to claim 7, further comprising using at least one of the at least two portions of pollen for outbred testing.

13. The method according to claim 10, wherein using a second portion of the regeneration material to generate data, further comprises generating a double haploid and collecting data.

14. The method according to claim 1, wherein the regeneration material is tissue and the second portion of regeneration material is created using vegetative propagation.

15. The method according to claim 1, wherein the pool of plants are selected from com (maize) (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), wheat (Triticum aestivum), soybean (Glycine max), tobacco (Nicotiana tabacum), cotton (Gossypium barbadense, Gossypium hirsutum), oats, barley, and vegetables.

16. The method according to claim 1, wherein the pool of plants are corn.

17. The method according to claim 1, wherein the data is selected from data used to compute GEBV.

18. The method according to claim 17, wherein the data comprises environmental data in combination with phenotypic data.

19. The method according to claim 8, wherein the priming is selected from heat treatment, cold treatment, pressure treatment, chemical treatment, abrasion, desiccation treatment, and combinations thereof.

20. A method of increasing the prediction accuracy of a genomic prediction model comprising: selecting a genome from a cycling population of plants for use in generating data to train a machine learning model;Crossing a plant comprising the desired genome with a substantially homozygous plant;Collecting data from the progeny of the cross; andUsing the data to refine the machine learning model, wherein the time between selecting and using the data is less than 1.5 years.

21. The method according to claim 20, wherein the cycling population comprises a co-ancestry value of less than 1 as calculated by the VanRaden method 1.

22. The method according to claim 20, wherein the cycling population advances less than 5 generations before the data is used to refine the machine learning model.

23. The method according to claim 20, wherein the selected genome is in a form selected from the group consisting of pollen, a tissue cutting, or a seed.

24. The method according to claim 23, wherein the selected genome is a seed and the method further comprises growing the seed into a plant.

25. A method of using mixed germplasm source data for updating machine learning models for genomic selection, comprising:Maintaining at least one cycling population of genomes; selecting from the at least one cycling population of genomes a desired population; dividing the desired population into a first sub-population of genomes for data generation and a second sub-population for generating another cycle in the at least one cycling population of genomes; directly crossing the first sub-population of genomes for data generation with one or more plants having known associated data; collecting data from the cross to produce outbred data; and retraining the machine learning models for genomic selection using the outbred data and a data set from an earlier cycle in the at least one cycling population of genomes.

26. The method according to claim 25, wherein the data set from an earlier cycle comprises data from a substantially homozygotic individual.

27. The method according to claim 26, wherein the substantially homozygotic individual is a maternal double haploid, paternal double haploid, a selfed plant, and combinations thereof.