Methods and systems for use in plant breeding
A method using a performance objective model and mathematical solver optimizes plant genotype selection to enhance genetic diversity and meet multiple breeding objectives, addressing inefficiencies in traditional breeding decision-making.
Patent Information
- Application Number
- PCT/US2025/036782
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-15
AI Technical Summary
Breeders face a significant burden in sorting through a semi-infinite number of breeding combinations while ensuring genetic diversity and adherence to multiple breeding objectives, leading to inefficiencies in decision-making and potential genetic bottlenecks.
A method utilizing a performance objective model and a mathematical solver to calculate advancement values for plant genotypes, optimizing primary and secondary phenotypic data to maximize desired traits while adjusting for constraints, thereby selecting the most suitable genotypes for breeding decisions.
This approach allows for efficient selection of genotypes that enhance genetic diversity and meet multiple breeding objectives, reducing the time and resource burden on breeders and preventing genetic depletion.
Smart Images

Figure US2025036782_15012026_PF_FP_ABST
Abstract
Description
Attorney Reference No.: 108485-WO-SEC-1 METHODS AND SYSTEMS FOR USE IN PLANT BREEDING CROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to US provisional application No.63 / 668,879, filed July 9, 2024, which is incorporated by reference herein in its entirety. FIELD
[0002] This disclosure relates to methods and systems for genotype selection and advancement for plant breeding. BACKGROUND
[0003] During each round of selection related to breeding decisions, breeders are faced with a semi-infinite number of combinations of potential breeding crosses, populations to be created, and progeny to be selected / advanced in a manner that fits with their breeding objectives and ensures retention and / or improvement of genetic diversity for continued long-term genetic gain.
[0004] Each breeding or evaluation zone can have multiple breeding objectives and / or product concepts that need to be considered, sorted, filtered, classified / categorized, and evaluated. Historically, breeders invested significant time in sorting, filtering, and evaluating these different combinations according to their breeding objectives, business rules, and diversity targets as well as other logistical criteria such as crossing block capacity, prediction accuracy, and field testing allocation to make informed breeding decisions.
[0005] There remains an unmet need for algorithmic-based breeding decisions that can reduce this burden for enhanced breeding decisions. SUMMARY
[0006] In a first aspect, the disclosure provides a method of advancing plant genotypes to create a plant having one or more target phenotypes, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) aAttorney Reference No.: 108485-WO-SEC-1 secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) performing one or more phenotypic analyses on a selected genotype.
[0007] In a second aspect, the disclosure provides a method of selecting plant genotypes for an improved population, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes;Attorney Reference No.: 108485-WO-SEC-1 (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) pollinating a plant of the selected genotype to generate a population. Pollinating can refer to pollinating a flower or silk of a plant with donor pollen or crossing.
[0008] In a third aspect, the disclosure provides a method of selecting plant genotypes for an improved cross, the method comprising: (a) providing a performance objective model m with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) crossing a plant of the selected genotype to create an inbred or a hybrid plant.
[0009] In a fourth aspect, the disclosure provides a method of advancing plant genotypes to create a plant having an improved phenotype, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype,Attorney Reference No.: 108485-WO-SEC-1 wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) performing one or more phenotypic analyses on a selected genotype.
[0010] In the first, second, third, and / or fourth above aspects, the algorithm can be further provided a diversity target value to be satisfied by the target number of selected genotypes, wherein the objective function created by the algorithm is further defined by the diversity target value, and wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value.
[0011] In the first, second, third, and / or fourth above aspects, the algorithm can be further provided one or more germplasm use constraints, wherein the objective function created by the algorithm is further defined by the one or more germplasm constraints, and wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and the one or more germplasm use constraints.
[0012] In the first, second, third, and / or fourth above aspects, the plurality of candidate genotypes comprises: (a) a realized genotype from an inbred or hybrid plant; (b) a hypothetical genotype from an inbred or hybrid plant; (c) a genotype resulting from a realized breeding cross; (d) a genotype resulting from a hypothetical breeding cross; or (e) a combination of a. – d.Attorney Reference No.: 108485-WO-SEC-1
[0013] In the first, second, third, and / or fourth above aspects, the primary phenotypic data and / or the secondary phenotypic data for each candidate genotype can comprise hypothetical phenotypic data.
[0014] In the first, second, third, and / or fourth above aspects, the candidate genotype comprising the hypothetical phenotypic data can be a realized genotype.
[0015] In the first, second, third, and / or fourth above aspects, the candidate genotype comprising the hypothetical phenotypic data can be a hypothetical genotype.
[0016] In the first, second, third, and / or fourth above aspects, the algorithm can be a genetic algorithm.
[0017] In the first, second, third, and / or fourth above aspects, the algorithm can be a heuristic or metaheuristic search algorithm.
[0018] In the first, second, third, and / or fourth above aspects, adjusting the advancement value by the constraint imposed on the secondary performance objective can comprise imposing a penalty value on the advancement value of each candidate genotype. For example, the constraint imposed on the secondary performance objective can be a range of numerical values. In another example, the constraint imposed on the secondary performance objective can be a minimum or maximum numerical value representing a minimum or maximum threshold to be satisfied by the selected genotype.
[0019] In the first, second, third, and / or fourth above aspects, the primary performance objective can be a trait.
[0020] In the first, second third, and / or fourth above aspects, the secondary performance objective can be a trait.
[0021] In the first, second, third, and / or fourth above aspects, the diversity target value can be a minimum numerical value representing a minimum threshold to be satisfied by the selected genotype.
[0022] In the first above aspect, a method of advancing plant genotypes to create a plant having one or more target phenotypes can further comprise growing a plant of the selected genotype or growing multiple plants of multiple selected genotypes.
[0023] In the second above aspect, a method of selecting plant genotypes for an improved population can further comprise growing the population of plants having the selected genotype.Attorney Reference No.: 108485-WO-SEC-1
[0024] In the third above aspect, a method of selecting plant genotypes for an improved cross can further comprise growing the inbred or hybrid plant.
[0025] In the fourth above aspect, a method of advancing plant genotypes to create a plant having an improved phenotype can further comprise growing a plant of the selected genotype or growing multiple plants of multiple selected genotypes.
[0026] In the first, second, third, and / or fourth above aspects, the plurality of candidate genotypes can comprise candidate genotypes having one or more genomic modifications introduced via a site-specific genome-editing agent. A site-specific genome-editing agent can be a zinc finger nuclease, TALEN, homing endonuclease, Cas polypeptide, TnpB nuclease, or a hydrolytic endonucleolytic ribozyme. A genomic modification can comprise an insertion, a deletion, a single nucleotide polymorphism, an inversion, or a translocation.
[0027] In a fifth aspect, the disclosure provides a computing device comprising a processor configured to perform the method of the first, second, third, and / or fourth above aspects.
[0028] In a sixth aspect, the disclosure provides a computer-readable medium comprising instructions, which when executed by a computing device, cause the computing device to carry out the method of the first, second, third, and / or fourth above aspects.
[0029] In a seventh aspect, the disclosure provides a system for advancing plant genotypes to create a plant having one or more target phenotypes, the system comprising: (i) one or more servers for storing data, the data comprising: (a) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (b) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (c) primary phenotypic data for each candidate genotype related to the primary performance objective; (d) a secondary performance objective to be satisfied by the selected genotype; (e) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (f) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (ii) a computing device communicatively coupled to the one or more servers, the computing device including a memory and one or more processors configured to carry out: (a) providing a performance objective model with the plurality of candidate genotypes, the primary performance objective, the primary phenotypic data for each candidate genotype, the secondary performance objective, the secondary phenotypic data for each candidate genotype, and the target number of selected genotypes to be returned; (b) calculating anAttorney Reference No.: 108485-WO-SEC-1 advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by a performance objective model and defined by the advancement values calculated by the performance objective model and the target number of selected genotypes to be returned from the plurality of candidate genotypes; and (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective.
[0030] In the seventh above aspect, the data can further comprise a diversity target value to be satisfied by the target number of selected genotypes, the objective function created by the performance objective model is further defined by the diversity target value, and the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value.
[0031] In the seventh above aspect, the data can further comprise one or more germplasm use constraints, the objective function created by the performance objective model is further defined by the one or more germplasm constraints, and the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and the one or more germplasm use constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG.1A is a flowchart illustrating an example of an algorithm-based method for selecting one or more candidate plant genotypes.
[0033] FIG. 1B is a flowchart illustrating another example of an algorithm-based method for selecting one or more plant genotypes
[0034] FIG.1C is a further example of an algorithm-based method for selecting one or more plant genotypes.
[0035] FIG. 2 is a density plot representing the distribution of all (~36,000) candidates for selection relative to a set of reference checks for the primary trait according to Example 3.Attorney Reference No.: 108485-WO-SEC-1
[0036] FIG.3 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the primary trait without any secondary constraints according to Example 3.
[0037] FIG. 4 is a density plot representing the distribution of all (~36,000) candidates for selection relative to a set of reference checks for the secondary trait (Sec_Trait_4) with illustrated boundaries of the accepted range for the secondary trait according to Example 4.
[0038] FIG.5 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the secondary trait constrained by the accepted range according to Example 4.
[0039] FIG.6 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the primary trait with constraints for the secondary trait (Sec_Trait_4) according to Example 4.
[0040] FIG.7 is a graph representing the percentage gain of the mathematical solver selections vs the diversity value when maximizing for the primary trait without any secondary constraints according to Example 4.
[0041] FIG.8 is a graph representing the percentage gain of the mathematical solver selections vs the diversity value when maximizing for the primary trait with constraints for the secondary trait (Sec_Trait_4) according to Example 4.
[0042] FIG.9 is a graph representing the relationship between the primary trait and the secondary trait (Sec_Trait_4), and the mathematical solver selections for maximizing the primary trait without any secondary trait constraints according to Example 4.
[0043] FIG.10 is a graph representing the relationship between the primary trait and the secondary trait (Sec_Trait_4). The algorithm mathematical solver (black dots) for maximizing the primary trait with constraints for the secondary trait (Sec_Trait_4) are shown as well as selection candidates dropped (grey dots) by the mathematical solver due to the secondary trait constraints according to Example 4.
[0044] FIG.11 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the primary trait when constrained by a minimum diversity value according to Example 5.Attorney Reference No.: 108485-WO-SEC-1
[0045] FIG.12 is a graph representing the percentage gain of the mathematical solver selections vs the diversity value when maximizing for the primary trait when constrained by a minimum diversity value according to Example 5.
[0046] FIG.13 is a graph representing the relationship between the primary trait and the secondary trait (Sec_Trait_4). The mathematical solver selections (black dots) for maximizing the primary trait when constrained by a minimum diversity value (no constraints on Sec_Trait_4) are shown as well as selection candidates dropped (grey dots) by the mathematical solver due to the diversity constraint according to Example 5.
[0047] FIG.14 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the secondary trait (Sec-Trait-4) when constrained by a minimum diversity value according to Example 6.
[0048] FIG.15 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the primary trait when constrained by a secondary trait (Sec_Trait_4) and a minimum diversity value according to Example 6.
[0049] FIG.16 is a graph representing the percentage gain of the mathematical solver selections vs the diversity value when maximizing for the primary trait with a secondary trait (Sec_Trait_4) and diversity constraint according to Example 6.
[0050] FIG.17 is a graph representing the relationship between the primary trait and the secondary trait (Sec_Trait_4). The mathematical solver selections (black dots) for maximizing the primary trait when constrained by a secondary trait (Sec_Trait_4) and a minimum diversity value according to Example 6.
[0051] FIG.18 is a density plot representing the distribution of the mathematical solver selections relative to a set of reference checks for the primary trait when constrained by multiple secondary traits (Sec_Trait_1 – Sec_Trait_21) and a minimum diversity value according to Example 7.
[0052] FIG.19– FIG.39 are density plots representing the distribution of the mathematical solver selections relative to a set of reference checks for each secondary trait (Sec_Trait_1 – Sec_Trait_21) when constrained by a minimum diversity value according to Example 7.
[0053] FIG.40 is a graph representing the percentage gain of the mathematical solver selections vs the diversity value when maximizing for the primary trait with multiple secondary trait constraints (Sec_Trait_1 – Sec_Trait_21) and a minimum diversity value according to Example 7.Attorney Reference No.: 108485-WO-SEC-1 DETAILED DESCRIPTION
[0054] All patents, publications and patent applications mentioned in the specification are indicative of the level of those skilled in the art to which this disclosure pertains. All patents, publications and patent applications are herein incorporated by reference in the entirety to the same extent as if each individual patent, publication or patent application was specifically and individually indicated to be incorporated by reference in its entirety.
[0055] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosed methods belong. In this specification and in the claims which follow, reference will be made to a number of terms which shall be defined as set forth below unless otherwise specified.
[0056] As used herein the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells and reference to “the protein” includes reference to one or more proteins and equivalents thereof known to those skilled in the art, and so forth.
[0057] As used herein, "crossed", "cross", or "crossing" in the context of this disclosure refers to the fusion of gametes via pollination to produce progeny (i.e., cells, seeds, or plants), and encompasses both realized and hypothetical sexual crosses (the pollination of one plant by another). In the present disclosure, a “realized cross” refers to an in-situ cross between two parent genotypes and a “hypothetical cross” refers to an in-silico, simulated cross between two parent genotypes. Breeding cross can occur within (inbreds) and across (hybrids) heterotic pools.
[0058] As used herein, "dicotyledonous" or "dicot" refers to the subclass of angiosperm plants also knows as "dicotyledoneae", whose seeds typically comprise two embryonic leaves, or cotyledons, and includes references to whole plants, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and progeny of the same. "Monocotyledonous" or "monocot" refers to the subclass of angiosperm plants also known as "monocotyledoneae", whose seeds typically comprise only one embryonic leaf, or cotyledon, and includes references to whole plants, plant elements, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, and progeny of the same.
[0059] As used herein, “germplasm” refers to genetic material of or from an individual (e.g., a plant), a group of individuals (e.g., a plant line, variety, or family), or a clone derived from a line, variety, species, or culture, or more generally, all individuals within a species or for several species (e.g., maize germplasm collection or Andean germplasm collection). The germplasm can be partAttorney Reference No.: 108485-WO-SEC-1 of an organism or cell, or can be separate from the organism or cell. In general, germplasm provides genetic material with a specific molecular makeup that provides a physical foundation for some or all of the hereditary qualities of an organism or cell culture. Germplasm includes cells, seed, or tissues from which new plants may be grown, or plant parts, such as leaves, stems, pollen, or cells, that can be cultured into a whole plant.
[0060] As used herein, “hybrid” refers to the progeny obtained between the crossing of at least two genetically dissimilar parents.
[0061] As used herein, “inbred” refers to a line that has been bred for genetic homogeneity.
[0062] As used herein, “line” or “strain” refers to a group of individuals of identical parentage that are generally inbred to some degree and that are generally homozygous and homogeneous at most loci (isogenic or near isogenic). An “elite line” refers to any line that has resulted from breeding and selection for superior agronomic performance.
[0063] As used herein, a “population” refers to a group of interbreeding plants. More specifically, a population refers to a set of segregating genotypes resulting from the crossing of two inbred lines (breeding within inbred lines). These segregating genotypes are subsequently propagated via selfing or doubled haploid induction methods or any other suitable technique known to one skilled in the art.
[0064] As used herein, “progeny” refers to any subsequent generation of a plant. More specifically, progeny refers to a specific genotype (realized or hypothetical) within a population that results from, or is derived from, the crossing of two inbred lines and subsequent propagation via selfing or doubled haploid induction methods.
[0065] As used herein, “trait” refers to a physiological, morphological, biochemical, or physical characteristic of a plant or particular plant material or cell. In some instances, this characteristic is visible to the human eye, such as seed or plant size, or can be measured by biochemical techniques, such as detecting the protein, starch, or oil content of seed or leaves, or by observation of a metabolic or physiological process (e.g., by measuring uptake of carbon dioxide) or by the observation of the expression level of a gene or genes (e.g., by employing Northern analysis, RT- PCR, microarray gene expression assays, or reporter gene expression systems) or by agricultural observations such as stress tolerance, yield, or pathogen tolerance.
[0066] The practice of plant breeding relies on the creation of variability through development of a generally large set of new crosses, populations, varieties, progeny, and / or hybrid productsAttorney Reference No.: 108485-WO-SEC-1 followed by the evaluation of these new crosses, populations, varieties, inbreds, and / or hybrid products to narrow down to a desired selected set that meets breeding objectives and product targets. This evaluation can be performed as an in-silico evaluation (related to the use of predictions for projecting the potential performance of a cross, population, variety, inbred, and / or hybrid), an in-situ evaluation (where realized observations are taken on the set of interest), or a combination thereof. The decisions breeders make following these evaluations determine the resulting next set of genetic material for future genetic progress. This cyclical process is commonly referred to as recurrent selection in the case of varietal development and reciprocal recurrent selection in the case of hybrid development where genetic progress is conducted in female and male inbred pools.
[0067] Generally, the “best” performers are related to one another and share similar genetic backgrounds and therefore simply sorting and filtering out according to the desired performance attributes and breeding objectives is not sufficient as it can, over time, decrease diversity across generations (e.g., decreasing diversity in parents and grandparents used in crosses and for population development) and / or overuse parent and / or grandparent germplasm. This depletion of genetic variability results in genetic bottlenecks and a long-term decrease in genetic gain. Hence, aside from evaluating the large number of possible combinations of crosses, breeders consider the resulting need to take into consideration the impact that their decisions will have on the long-term diversity of their germplasm.
[0068] Further, breeders make these decisions in a timely manner to abide by growing season availability, including performing their desired crosses, generating sufficient seed from the selected set for further evaluation, and meeting operational timelines for nurseries. Relying on sorting and filtering practices is therefore not optimal when faced with a large number of selection possibilities and multiple breeding objectives that must be satisfied with limited time and resources.
[0069] For example, in a large pool of selection candidates, “S”, breeders narrow down their selections to a relatively small elite subset “E”. Assuming all selection candidates have favorable performance evaluation attributes, there could be several favorable combinations of “E” within the selection candidates “S”. The number of possible combinations (C) can be identified as:Attorney Reference No.: 108485-WO-SEC-1
[0070]
[0071] of 1000 favorable selection candidates, and the breeder intends to select the best 100 elite set E, there would be
[0072] combinations to sort through.
[0073] The methods and systems described herein allow for the evaluation and selection of crosses, populations, varieties, inbreds, and / or hybrids while allowing for dynamic optimization across multiple dimensions and objectives that can be leveraged when the optimization targets change and / or evolve in complexity as the scope of optimization targets increases.
[0074] Disclosed herein are methods and systems for advancing plant genotypes using a mathematical indexing approach that allows a breeder or an automated system to input a dynamic set of selection criteria (i.e., selection criteria can be changed run-to-run), wherein a heuristic or metaheuristic search algorithm collectively receives the selection criteria in a single step. The methods described herein are superior to tandem or step-wise type selection as they allow all candidates to be evaluated for their potential contribution to the overall performance objective in comparison to traditional approaches of using cutoff values for each secondary performance objective which removes any candidates that violate secondary performance objective threshold values, thus reducing the genetic search space to a subset of candidates that are favorable for all secondary trait thresholds that then would be subjected to additional diversity and germplasm use constraints. The methods disclosed herein consider or weight each secondary performance objective, which allows all candidates of selection to be considered while also integrating additional constraints related to diversity and germplasm use.
[0075] More specifically, the algorithm receives candidate plant genotypes and a breeder-selected (or automated system-selected) primary attribute of interest to be maximized (i.e., the primary attribute of interest is maximized with respect to the candidate plant genotypes used as input and maximized with respect to the primary attribute of interest), referred to herein as a “primary performance objective”, and phenotypic data for each candidate genotype for the primary attributeAttorney Reference No.: 108485-WO-SEC-1 of interest. The algorithm further receives one or more secondary attributes of interest, referred to herein as a “secondary performance objective”, phenotypic data for each candidate genotype related to the secondary attribute of interest, and a target number of candidate genotypes to be selected. The algorithm can also receive additional, optional constraints including a diversity target, with accompanying genotypic data, and / or a germplasm use constraint. The algorithm calculates an advancement value for each candidate genotype, the advancement value intended to maximize the primary performance objective while adjusting for constraints (numerical penalty values) imposed by the secondary performance objective. Once advancement values are calculated, the algorithm creates an objective function to be provided to a mathematical solver, the objective function defined by or includes the advancement values for each candidate genotype, the target number of selected genotypes to be returned from the candidate genotypes, and any optional constraints (e.g., diversity target and germplasm use). The mathematical solver then solves the objective function and collects the mathematical solution output, which is the target number of selected genotypes that maximize the primary performance objective and satisfy any optional constraints provided to the algorithm. “Genotypic data” as used herein can include genome sequence information such as SNP, QTL, RNA-seq, short read genomic sequencing, marker data, long read genome sequence information, methylation status, gene expression values, indels, haplotypes, and combinations thereof. In some aspects, the genotypic data includes a collection of genotypic markers, such as genome-wide markers, or single nucleotide polymorphisms (SNPs). In some examples, the genotypic data is imputed.
[0076] For example, as shown in FIG.1A, an algorithm receives candidate data including a unique identifier for each candidate genotype, a primary performance objective, one or more secondary performance objectives, phenotypic data for each candidate genotype for the primary and secondary performance objectives, and the target number of selected genotypes (e.g., 500 candidates) to be returned from the candidate genotypes. The algorithm can further receive optional constraints related to diversity and germplasm use (optional constraints depicted as dashed lines in FIG. 1A). The algorithm reads in the primary and secondary performance objectives, applies data processing steps, and generates an advancement value for each candidate genotype. More specifically, using the primary phenotypic data for each candidate genotype, the algorithm calculates its advancement value by first normalizing (i.e., centering around 0) the primary performance objective, then normalizing, and in some cases, penalizing each secondaryAttorney Reference No.: 108485-WO-SEC-1 performance objective with penalty values increasing the further the candidate genotype deviates from the secondary performance objectives, and finally adjusting the normalized primary performance objective values by the secondary performance objective penalties. Once the algorithm generates advancement values for each candidate genotype, it translates or parameterizes the advancement values and target number of selected genotypes into an objective function (i.e., mathematical equation), which is then provided to a mathematical solver. The mathematical solver then solves the objective function. In this case, with no optional constraints on diversity and germplasm use, the objective function returns the target number of selected genotypes with advancement values maximizing the primary performance objective.
[0077] Turning to FIG. 1B, an algorithm receiving candidate data additional constraints is illustrated. The algorithm receives candidate data including a unique identifier for each candidate genotype, a primary performance objective, one or more secondary performance objectives, phenotypic data for each candidate genotype for the primary and secondary performance objectives, and the target number of selected genotypes (e.g., 500 candidates) to be returned from the candidate genotypes. The algorithm further receives constraints related to diversity and germplasm use (e.g., genetic background constraint and phenotype confidence target). The algorithm reads in the primary and secondary performance objectives, applies data processing steps, and generates an advancement value for each candidate genotype. More specifically, using the primary phenotypic data for each candidate genotype, the algorithm calculates its advancement value by first normalizing the primary performance objective, then normalizing, and in some cases, penalizing each secondary performance objective, with penalty values increasing the further the candidate genotype deviates from the secondary performance objectives, and finally adjusting the normalized primary performance objective values by the secondary performance objective penalties. Once the algorithm generates advancement values for each candidate genotype, it translates the advancement values, target number of selected genotypes, diversity reference data for diversity, and germplasm use constraints into an objective function, which is then provided to a mathematical solver. The mathematical solver then solves the objective function, returning the target number of selected genotypes with advancement values maximizing the primary performance objective and meeting the diversity and germplasm use constraints.
[0078] FIG.1C illustrates an example of the methods described herein in which a diversity target value (e.g., a numerical value from 0-1) functions as the secondary constraint on the primaryAttorney Reference No.: 108485-WO-SEC-1 performance objective. The algorithm receives candidate data including a unique identifier for each candidate genotype, a primary performance objective, phenotypic data for each candidate genotype for the primary performance objective, and the target number of selected genotypes (e.g., 500 candidates) to be returned from the candidate genotypes. The algorithm further receives the reference data for diversity and optionally any germplasm use constraints. The algorithm reads in the primary performance objective, applies data processing steps (i.e., normalizing the primary performance objective), and generates an advancement value for each candidate genotype. Once the algorithm generates advancement values for each candidate genotype, it translates and parameterizes the advancement values, target number of selected genotypes, reference data for diversity, and optional germplasm use constraints into an objective function, which is then provided to a mathematical solver. The mathematical solver then solves the objective function, returning the target number of selected genotypes with advancement values maximizing the primary performance objective and meeting the diversity and optional germplasm use constraints.
[0079] The methods described herein can select candidate genotypes from a realized genotype from an inbred or hybrid plant, a hypothetical genotype from an inbred or hybrid plant, resulting from a realized breeding cross (e.g., breeding pool of females or males), and / or resulting from a hypothetical breeding cross. As used herein a “realized genotype” or a “realized candidate genotype” refers to an existing or known genotype with available or observed data, such as a genotype from existing plant progeny. As used herein, a “hypothetical genotype” or a “hypothetical candidate genotype” refers to a simulated genotype within a hypothetical population, breeding cross, or hybrid cross. The hypothetical genotype can be progeny that have yet to be materialized.
[0080] A hypothetical genotype can be projected (i.e., predicted or estimated) from an observed DNA fingerprint or a predicted DNA fingerprint. As used herein, “DNA fingerprint” refers to a unique genetic profile or signature of a plant cultivar. DNA fingerprinting involves the generation of a set of distinct DNA fragments from a single DNA sample. The generated DNA fragments are then used as a source of genotypic information. A variety of techniques can be used to generate DNA fingerprinting patterns. The choice of the technique depends on the organism being studied and on the question being addressed. All DNA fingerprinting techniques study patterns associated with genetic markers; however, individual techniques differ in terms of the number and type of genetic markers examined. For example, some approaches allow the examination of a marker at aAttorney Reference No.: 108485-WO-SEC-1 single locus (called single-locus markers), whereas others allow the simultaneous investigation of multiple loci (called multi-locus markers). Some approaches focus on co-dominant markers, which provide information about both alleles present at a given locus. In contrast, other techniques are concerned with dominant markers, which only report the presence or absence of a given allele and cannot provide information about whether an individual is homozygous for that allele.
[0081] Phenotypes of the candidate genotypes (i.e., primary and secondary phenotypic data) can also be hypothetical or observed / realized. As used herein, a “hypothetical phenotype” or “hypothetical phenotypic data” refers to a simulated or projected (i.e., predicted) trait or collection of traits in an organism, such as a plant. A hypothetical phenotype can be based on a realized or hypothetical genotype. When based on a hypothetical genotype, the phenotype is predicted from the relevant DNA fingerprint. As used herein, an “observed phenotype” or a “realized phenotype” refers to an observable or measurable trait or collection of traits in an organism, such as a plant, resulting from the interaction of its genotype and environment.
[0082] Observed phenotypes for candidate genotypes can be from historical performance data such as, but not limited to, performance during a previous growing season of the respective genotype.
[0083] Hypothetical phenotypes for the candidate genotypes can be predicted based on historical performance data of a related genotype such as, but not limited to, performance during a previous growing season of the related genotype.
[0084] If a hypothetical genotype is selected for advancement by the methods and systems described herein, a plant having the selected genotype can be created via a pollination event (pollinating a flower or silk of a plant with donor pollen) or crossing.
[0085] In a first aspect, the disclosure provides methods of advancing plant genotypes to create a plant having one or more target phenotypes, the methods comprising: (a) providing an algorithm with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (iv) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating anAttorney Reference No.: 108485-WO-SEC-1 advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the algorithm and defined by: (i) the advancement values calculated by the algorithm; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; and (e) advancing one or more selected genotypes from the target number of selected genotypes. Following genotype selection, if the selected genotype is a realized genotype, the breeder can perform one or more additional phenotypic analyses on the selected genotype such as, but not limited to, growing the selected candidate genotype plant, conducting field tests on the selected candidate genotype plant, crossing the selected candidate genotype plant, making a new population, advancing the selected candidate genotype in the breeding program, or recycle the genotype back into a breeding cycle. Following genotype selection, if the selected genotype is a hypothetical genotype, the breeder can decide to create the hypothetical cross and generate new realized genotypes that can then be subjected to additional algorithmic selections and / or carried forward for additional field testing evaluation, used as a parent of new crosses, advanced into a product such as for varietal crops or used as a parent of new hybrid products for hybrid crops.
[0086] In an example of this first aspect, the methods can further include providing the algorithm with one or more germplasm use constraints and / or a diversity target value to be satisfied by the target number of selected genotypes. The germplasm use constraints and / or diversity target value further define the objective function and the mathematical solution output of the mathematical solver is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and / or the one or more germplasm use constraints.
[0087] In a second aspect, the disclosure provides methods of selecting plant genotypes for an improved population, the methods comprising: (a) providing an algorithm with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidateAttorney Reference No.: 108485-WO-SEC-1 genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (iv) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the algorithm and defined by: (i) the advancement values calculated by the algorithm; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; and (e) advancing one or more selected genotypes from the target number of selected genotypes. Following genotype selection, if the selected genotype is a realized genotype, the breeder can pollinate a plant of the selected genotype to generate a population. Following genotype selection, if the selected genotype is a realized genotype, the breeder can perform one or more additional phenotypic analyses on the selected genotype such as, but not limited to, growing the selected candidate genotype plant, conducting field tests on the selected candidate genotype plant, crossing the selected candidate genotype plant, making a new population, advancing the selected candidate genotype in the breeding program, or recycle the genotype back into a breeding cycle. Following genotype selection, if the selected genotype is a hypothetical genotype, the breeder can the breeder can decide to create the hypothetical cross and generate new realized genotypes that can then be subjected to additional algorithmic selections and / or carried forward for additional field testing evaluation, used as a parent of new crosses, advanced into a product such as for varietal crops or used as a parent of new hybrid products for hybrid crops. Accordingly, provided herein are methods for creating an improved population or cross. As used herein, an “improved population” will be understood to mean a population having a higher primary performance objective (e.g., primary trait) while having improved or matching secondary performance objectives (e.g., secondary traits).Attorney Reference No.: 108485-WO-SEC-1
[0088] In an example of this second aspect, the methods can further include providing the algorithm with one or more germplasm use constraints and / or a diversity target value to be satisfied by the target number of selected genotypes. The germplasm use constraints and / or diversity target value further define the objective function and the mathematical solution output of the mathematical solver is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and / or the one or more germplasm use constraints.
[0089] In a third aspect, the disclosure provides methods of selecting plant genotypes for an improved cross, the methods comprising: (a) providing an algorithm with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii) primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (iv) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a mathematical solver, wherein the objective function is created by the algorithm and defined by: (i) the advancement values calculated by the algorithm; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the mathematical solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; and (e) advancing one or more selected genotypes from the target number of selected genotypes. Following genotype selection, if the selected genotype is a realized genotype, the breeder can perform one or more additional phenotypic analyses on the selected genotype such as, but not limited to, growing the selected candidate genotype plant, conducting field tests on the selected candidate genotype plant, crossing the selected candidate genotype plant, making a new population, advancing the selected candidate genotype in the breeding program, or recycle the genotype back into a breeding cycle.Attorney Reference No.: 108485-WO-SEC-1 Following genotype selection, if the selected genotype is a hypothetical genotype, the breeder can the breeder can decide to create the hypothetical cross and generate new realized genotypes that can then be subjected to additional algorithmic selections and / or carried forward for additional field testing evaluation, used as a parent of new crosses, advanced into a product such as for varietal crops or used as a parent of new hybrid products for hybrid crops. Accordingly, provided herein are methods for creating an improved population or cross. As used herein, an “improved population” will be understood to mean a population having a higher primary performance objective (e.g., trait) while having improved or matching secondary performance objectives (e.g., traits).
[0090] As used herein, an “improved cross” will be understood to mean a cross resulting in progeny having a higher primary performance objective (e.g., primary trait) while having improved or matching secondary performance objectives (secondary traits).
[0091] In an example of this third aspect, the methods can further include providing the algorithm with one or more germplasm use constraints and / or a diversity target value to be satisfied by the target number of selected genotypes. The germplasm use constraints and / or diversity target value further define the objective function and the mathematical solution output of the mathematical solver is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and / or the one or more germplasm use constraints.
[0092] In a fourth aspect, the disclosure provides systems for executing the methods described herein. The systems can include (i) one or more servers for storing data (e.g., unique identifiers for each candidate genotype, a primary performance objective, primary phenotypic data, secondary performance objectives, secondary phenotypic data, germplasm use constraints, and / or reference data for diversity), and (ii) a computing device communicatively coupled to the one or more servers, the computing device including a memory and one or more processors to perform operations including running an algorithm to generate an advancement value and objective function and executing a mathematical solver to return a target number selected candidates from the candidate genotypes.
[0093] The computing device of the system can be, for example, a computer, a notebook, a laptop, a mobile device, a smartphone, a tablet, wearable, smart glasses, or any other suitable computing device that is capable of communicating with a server. The computing device can include aAttorney Reference No.: 108485-WO-SEC-1 processor, a memory, an input / output (I / O) controller (e.g., a network transceiver), a memory unit, and a database, all of which may be interconnected via one or more address / data bus.
[0094] The system processor can be any electronic device that is capable of processing data, for example a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a system on a chip (SoC), or any other suitable type of processor. It should be appreciated that the various operations of example methods described herein (i.e., performed by the computing device) can be performed by one or more processors. The memory can be a random- access memory (RAM), read-only memory (ROM), a flash memory, or any other suitable type of memory that enables storage of data such as instruction codes that the processor needs to access in order to implement any method as disclosed herein. The the computing device can be a computing device or a plurality of computing devices with distributed processing.
[0095] As used herein, the term “database” refers to a single database or other structured data storage, or to a collection of two or more different databases or structured data storage components. The database can part of the computing device. Alternatively, the computing device can access the database via a network. The database can store data (e.g., input, output, intermediary data) used for selecting candidate genotypes for advancement. For example, the data may include genotypic data, such as single nucleotide polymorphisms (SNPs), genetic markers, haplotype, sequence information, phenotypic data, mean locus effects (MLE) data, breeder’s field notes, environmental data, predicted genetic values, pedigree information, co-ancestry information, or combinations thereof that are obtained from one or more servers.
[0096] The computing device can further include a number of software applications stored in a memory unit, which may be called a program memory. The various software applications on the computing device can include specific programs, routines, or scripts for performing processing functions associated with the methods described herein. Additionally or alternatively, the various software applications on the computing device can include general-purpose software applications for data processing, database management, data analysis, network communication, web server operation, or other functions described herein or typically performed by a server. The various software applications can be executed on the same computer processor or on different computer processors. Additionally or alternatively, the software applications can interact with various hardware modules that installed within or connected to the computing device. Such modules canAttorney Reference No.: 108485-WO-SEC-1 implement part of or all of the various exemplary method functions discussed herein or other related embodiments.
[0097] The server can be a single server or a plurality of servers with distributed processing, which receive data from and / or transmit data to the computing device. The network can be any suitable type of computer network that functionally couples at least one computing device with the server. The network can include a proprietary network, a secure public internet, a virtual private network and / or one or more other types of networks, such as dedicated access lines, plain ordinary telephone lines, satellite links, cellular data networks, or combinations thereof.
[0098] In the first, second, third, and fourth aspects, a candidate genotype of the plurality of candidate genotypes can comprise one or more genomic modifications introduced via a site- specific genome-editing agent, the genomic modification(s) being an insertion, a deletion, a single nucleotide polymorphism, an inversion, and / or a translocation. Site-specific genome-editing agents include a zinc finger nuclease, a TAL effector nuclease (TALEN), a homing endonuclease, a Cas polypeptide, a TnpB nuclease, and a hydrolytic endonucleolytic ribozyme. Cas Polypeptides
[0099] Cas polypeptides and effector proteins can be used for targeted genome-editing (via simplex and / or multiplex double-strand breaks and / or single-strand breaks) and targeted genome regulation (via tethering of epigenetic effector domains to either the Cas polypeptide or guide polynucleotide). A Cas polypeptide can also be engineered to function as a polynucleotide-guided recombinase, and via polynucleotide tethers can serve as a scaffold for the assembly of multiprotein and nucleic acid complexes (Mali et al., 2013, Nature Methods Vol.10: 957-963).
[0100] A “Cas polypeptide” or “Cas effector protein” will be understood to mean a polynucleotide-guided CRISPR-associated protein that, when in complex with a suitable guide polynucleotide, can recognize and bind to a target DNA polynucleotide. A Cas polypeptide or Cas effector protein can have, but is not required to have, the additional functionality of: unwinding, unwinding and cutting, unwinding and nicking, cutting, or nicking a DNA polynucleotide.
[0101] A Cas polypeptide includes, but is not limited to, Cas9, Cas12f (Cas-alpha, Cas14), Cas12l (Cas-beta), Cas12a (Cpf1), Cas12b (a C2c1 protein), Cas13 (a C2c2 protein), Cas12c (a C2c3 protein), Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas3, Cas3-HD, Cas 5, Cas6, Cas7, Cas8, Cas10, or combinations or complexes of these.Attorney Reference No.: 108485-WO-SEC-1
[0102] In another example of the methods and systems disclosed herein, a genome-editing effector enzyme is a Cas endonuclease comprising one or more domains enabling it to function as a double- strand break-inducing agent. A Cas endonuclease is a type of Cas polypeptide or Cas effector protein comprising one or more nuclease domains (for example, one or more RuvC nuclease domains) that, when in complex with a suitable guide polynucleotide, recognizes, binds, and induces a double-strand break in a target DNA polynucleotide.
[0103] In a further example of the methods and systems disclosed herein, a genome-editing effector polypeptide is a Cas polypeptide having no substantial nuclease activity and is referred to as a catalytically “inactivated Cas” or a “deactivated Cas” (“dCas”). A dCas can be a Cas endonuclease comprising one or more modifications or mutations that abolish or reduce its ability to cut a double-stranded polynucleotide. The modified form of the Cas polypeptide can include an amino acid change (e.g., deletion, insertion, or substitution) that reduces or eliminates the naturally-occurring nuclease activity of a Cas endonuclease. A deactivated Cas typically lacks any functional nuclease domains. In this way, a deactivated Cas does not induce a single-strand break or a double-strand break, but can still bind to a target DNA polynucleotide. For example, a deactivated Cas can comprise (i) a mutant, dysfunctional RuvC domain(s) based on dimer.
[0104] In yet another example of the method and systems disclosed herein, a genome-editing effector polypeptide is a Cas polypeptide having nickase activity (i.e., induces a single-strand break), and is referred to herein as a “Cas nickase” (“nCas”) or a “Cas polypeptide having nickase activity”. A nCas can be a Cas endonuclease comprising one or more modifications or mutations such that DNA nicking functionality is retained. A Cas nickase typically comprises one functional endonuclease domain that allows the Cas to break only one strand (i.e., make a nick) at a target DNA sequence. For example, a Cas nickase can comprise (i) a mutant, dysfunctional RuvC domain and (ii) a functional HNH domain (e.g., wild-type HNH domain). As another example, a Cas nickase can comprise (i) a functional RuvC domain (e.g., wild-type RuvC domain) and (ii) a mutant, dysfunctional HNH domain.
[0105] Non-limiting examples of Cas nickases suitable for use herein are disclosed in US20140189896 published on 03 July 2014.
[0106] Cas polypeptides of the methods and systems disclosed herein, include functional fragments. As used herein, a “functional fragment”, “fragment that is functionally equivalent”, and “functionally equivalent fragment” of a Cas polypeptide are used interchangeably herein, and referAttorney Reference No.: 108485-WO-SEC-1 to a portion or subsequence of the Cas polypeptide in which the ability to recognize, bind to, and optionally unwind, nick, or cut a target polynucleotide sequence is retained.
[0107] Cas polypeptides of the methods and systems herein, include functional variants. As used herein, a “functional variant”, “variant that is functionally equivalent”, and “functionally equivalent variant” of a Cas polypeptide are used interchangeably herein, and refer to a variant of the Cas polypeptide in which the ability to recognize, bind to, and optionally unwind, nick, or cut all or part of a target polynucleotide sequence is retained.
[0108] Cas polypeptides of the methods and systems disclosed herein can be a multifunctional Cas polypeptide. As used herein, a “multifunctional Cas polypeptide” includes reference to a single polypeptide that has Cas polypeptide functionality and at least one other functionality, such as but not limited to, the functionality to form a cascade (comprises a domain that can form a cascade with other proteins). A multifunctional Cas polypeptide can comprise an additional protein domain relative to (either internally, upstream (5’), downstream (3’), or any combination thereof) those domains typical of a Cas polypeptide.
[0109] A Cas polypeptide can be isolated from a native source, or from a recombinant source where the genetically modified host cell is modified to express the nucleic acid sequence encoding the protein. Alternatively, the Cas polypeptide can be produced using cell free protein expression systems, or be synthetically produced. Cas polypeptides can be isolated and introduced into a heterologous cell, or can be modified from its native form to exhibit a different type or magnitude of activity than what it would exhibit in its native source. Such modifications include, but are not limited to, fragments, variants, substitutions, deletions, and insertions.
[0110] Fragments and variants of Cas polypeptides can be obtained via methods such as site- directed mutagenesis and synthetic construction. Methods for measuring endonuclease activity are well known in the art such as, but not limiting to, WO2013166113 published 07 November 2013, WO2016186953 published 24 November 2016, and WO2016186946 published 24 November 2016.
[0111] A Cas polypeptide can be part of a fusion protein comprising one or more heterologous protein domains (e.g., 1, 2, 3, or more domains in addition to the Cas protein). Suitable fusion partners include, but are not limited to, a polypeptide that provides an activity that indirectly increases transcription by acting directly on the target DNA or on a polypeptide (e.g., a histone or other DNA-binding protein) associated with the target DNA. Additional suitable fusion partnersAttorney Reference No.: 108485-WO-SEC-1 include, but are not limited to, a polypeptide that provides for methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity. Further suitable fusion partners include, but are not limited to, a polypeptide that directly provides for increased transcription of the target DNA (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription regulator, etc.). In another example, a Cas polypeptide can be a fusion protein further comprising a nuclease domain, a transcriptional activator domain, a transcriptional repressor domain, an epigenetic modification domain, a cleavage domain, a nuclear localization signal, a cell-penetrating domain, a translocation domain, a marker, or a transgene that is heterologous to the target DNA polynucleotide sequence or to the cell from which said target DNA polynucleotide sequence is obtained or derived. In another example, a Cas polypeptide fusion protein comprises Clo51 or Fok1.
[0112] A Cas polypeptide as described herein can be expressed and purified by methods known in the art, for example, as described in WO / 2016 / 186953.
[0113] A Cas polypeptide as described herein can comprise a heterologous nuclear localization sequence (NLS) of sufficient strength to drive accumulation of the Cas polypeptide in a detectable amount in the nucleus of a yeast cell herein, for example.
[0114] A Cas polypeptide can be fused to a heterologous sequence to induce or modify its activity (e.g., binding, cutting, or nicking). Cas9 Polypeptides
[0115] In the methods and systems disclosed herein, a genome-editing effector polypeptide can be a Cas9 polypeptide, such as a Cas9 endonuclease, and a guided Cas polypeptide system comprises a Cas9 polypeptide and one or more guide polynucleotides that introduce one or more site-specific modifications in a target DNA polynucleotide sequence. The guided Cas polypeptide system can further comprise a donor DNA or a polynucleotide modification template. Some exemplary Cas9 endonucleases are described in, for example, WO2019165168.
[0116] Cas9 (formerly referred to as Cas5, Csn1, or Csx12) is a Cas polypeptide that forms a complex with a crNucleotide and a tracrNucleotide, or with a single guide polynucleotide, forAttorney Reference No.: 108485-WO-SEC-1 specifically recognizing and cutting or nicking a target DNA polynucleotide. The canonical Cas9 recognizes a 3’ GC-rich PAM sequence on the target double-stranded DNA, typically comprising an NGG motif.
[0117] A Cas9 polypeptide can comprise a RuvC nuclease with an HNH (H-N-H) nuclease adjacent to the RuvC-II domain. The RuvC nuclease and HNH nuclease each can cut a single DNA strand at a target sequence (the concerted action of both domains leads to DNA double-strand cleavage, whereas activity of one domain leads to a nick). In general, the RuvC domain comprises subdomains I, II and III, where domain I is located near the N-terminus of Cas9 and subdomains II and III are located in the middle of the protein, flanking the HNH domain (Hsu et al., 2013, Cell 157:1262-1278). Cas9 polypeptides are typically derived from a type II CRISPR system, which includes a DNA cleavage system utilizing a Cas9 endonuclease in complex with at least one guide polynucleotide component. For example, a Cas9 can be in complex with a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). In another example, a Cas9 can be in complex with a single guide RNA (Makarova et al.2015, Nature Reviews Microbiology Vol.13:1-15).
[0118] The type II CRISPR / Cas system from bacteria employs a crRNA and tracrRNA to guide the Cas endonuclease to its DNA target. The crRNA (CRISPR RNA) contains the region complementary to one strand of the double-strand DNA target and base pairs with the tracrRNA forming a RNA duplex that directs the Cas endonuclease to cut the DNA target. Cas-alpha Polypeptides
[0119] In the methods and systems disclosed herein, a genome-editing effector polypeptide can be a Cas-alpha (e.g., Cas12f) polypeptide, and a guided Cas polypeptide system comprises a Cas- alpha polypeptide and one or more guide polynucleotides that introduce one or more site-specific modifications in a target DNA polynucleotide sequence. The guided Cas polypeptide system can further comprise a donor DNA or a polynucleotide modification template. Some exemplary Cas- alpha endonucleases are described in, for example, US10934536 and WO2022082179.
[0120] A Cas-alpha endonuclease is a functional polynucleotide-guided, PAM-dependent dsDNA cleavage protein of fewer than 800 amino acids comprising: a C-terminal RuvC catalytic domain split into three subdomains, a bridge-helix, and one or more zinc finger motifs, and further comprising an N-terminal Rec subunit with a helical bundle, WED wedge-like (or “Oligonucleotide Binding Domain”, OBD) domain, and, optionally, a zinc finger motif.Attorney Reference No.: 108485-WO-SEC-1
[0121] Cas-alpha polypeptides can comprise one or more zinc finger coordination motifs that may form a zinc binding domain. Zinc finger-like motifs can aid in target and non-target strand separation and loading of the guide polynucleotide into the target DNA polynucleotide. Cas-alpha polypeptides comprising one or more zinc finger-like motifs can provide additional stability to a ribonucleoprotein complex on a target DNA polynucleotide. Cas-alpha endonucleases comprise C4 or C3H zinc binding domains.
[0122] A Cas-alpha polypeptide can function as a double- or single-strand break-inducing agent. A catalytically inactive Cas-alpha endonuclease (dCas-alpha) can be used to target or recruit to a target DNA polynucleotide sequence without inducing cleavage or nicking. A catalytically inactive Cas-alpha polypeptide can be combined with a base editing molecule, such as a cytidine deaminase or an adenine deaminase.
[0123] Guide polynucleotide-Cas polypeptide complex
[0124] As used herein, a “guide polynucleotide-Cas polypeptide complex” or a “guide polynucleotide / Cas polypeptide complex” refers to at least one guide polynucleotide and at least one Cas polypeptide that form a complex capable of directing the Cas polypeptide to a target DNA polynucleotide and enabling the Cas polypeptide to recognize, optionally bind to, and optionally cut or nick a target DNA polynucleotide (i.e., introduce a double-strand break or single-strand break, respectively). A guide polynucleotide-Cas polypeptide complex can comprise a Cas polypeptide and suitable guide polynucleotide of any of the known CRISPR system (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al.2015, Nature Reviews Microbiology Vol. 13:1-15; Zetsche et al., 2015, Cell 163, 1-13; Shmakov et al., 2015, Molecular Cell 60, 1-13). The terms “guide polynucleotide-Cas endonuclease complex” and “guide polynucleotide / Cas endonuclease complex” refer to a guide polynucleotide-Cas polypeptide complex having a Cas endonuclease that comprises one or more domains enabling it to function as a double-strand break- inducing agent.
[0125] The terms “guide RNA-Cas polypeptide complex”, “gRNA-Cas polypeptide complex”, “guide RNA / Cas polypeptide complex”, and “gRNA / Cas polypeptide complex” are used interchangeably herein and refer to at least one guide RNA and at least one Cas polypeptide that form a complex capable of directing the Cas polypeptide to a target DNA polynucleotide and enabling the Cas polypeptide to recognize, optionally bind to, and optionally cut or nick a target DNA polynucleotide (i.e., introduce a double-strand break or single-strand break, respectively).Attorney Reference No.: 108485-WO-SEC-1 The terms “guide RNA-Cas endonuclease complex” and “guide RNA / Cas endonuclease complex” refer to a guide RNA-Cas polypeptide complex having a Cas endonuclease that comprises one or more domains enabling it to function as a double-strand break-inducing agent.
[0126] In an example of the methods and systems disclosed herein, the Cas polypeptide of a guide polynucleotide-Cas polypeptide complex is a Cas endonuclease, and genome-editing of a target DNA polynucleotide via the guide polynucleotide-Cas endonuclease complex comprises non- homologous end joining or homology-directed repair following a Cas endonuclease-mediated double-strand break. Alterations or modifications of a target DNA sequence include, for example: (i) substitution of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) – (iii).
[0127] Uses for the guided Cas polypeptide systems described herein include, but are not limited to, modifying or replacing nucleotide sequences of interest (such as a regulatory elements), insertion of polynucleotides of interest, gene knock-out, gene-knock in, modification of splicing sites and / or introducing alternate splicing sites, modifications of nucleotide sequences encoding a protein of interest, amino acid and / or protein fusions, and gene silencing by expressing an inverted repeat into a gene of interest.
[0128] In another example of the methods and systems disclosed herein, the Cas polypeptide of a guide polynucleotide-Cas polypeptide complex is a dCas. A guide polynucleotide-dCas polypeptide complex can be used, for example, for base editing.
[0129] In yet another example of the methods and systems disclosed herein, the Cas polypeptide of a guide polynucleotide-Cas polypeptide complex is a nCas. A guide polynucleotide-nCas polypeptide complex can be used, for example, for base editing or prime editing.
[0130] A guide polynucleotide-Cas polypeptide complex of the methods and systems disclosed herein can be a ribonucleoprotein (RNP) complex, wherein the Cas polypeptide is provided as a protein and the guide polynucleotide is provided as a ribonucleotide.
[0131] In a guide polynucleotide-Cas polypeptide complex of the methods and systems disclosed herein, the Cas polypeptide can further comprise one copy or multiple copies of a subunit of an additional (secondary) Cas polypeptide. For example, the Cas polypeptide can be covalently or non-covalently linked, or assembled to, one or more copies of a secondary Cas polypeptide subunit, such as a Cas1 subunit, a Cas2 subunit, a Cas4 subunit, or a combination thereof, effectively forming a cleavage-ready cascade.Attorney Reference No.: 108485-WO-SEC-1
[0132] Target DNA polynucleotide(s)
[0133] As used herein, “target DNA polynucleotide”, “DNA target site”, “target DNA sequence”, “target site”, “target sequence”, “genomic target sequence”, “genomic target site”, and “target polynucleotide” are used interchangeably and refer to a polynucleotide sequence, such as but not limited to, a nucleotide sequence on a chromosome, episome, a locus, or any other DNA molecule in the genome (including chromosomal DNA, chloroplastic DNA, mitochondrial DNA, or plasmid DNA) of a cell, at which a guide polynucleotide-Cas polypeptide complex can recognize, bind to, and optionally nick or cut. The target DNA polynucleotide can be an endogenous site in the genome of a cell, or alternatively, the target DNA polynucleotide can be heterologous to the cell, and thereby not naturally occurring in the genome of the cell, or alternatively, the target DNA polynucleotide can be found in a heterologous genomic location compared to where it occurs in nature. As used herein, terms “endogenous target sequence” and “native target sequence” are used interchangeably to refer to a target sequence that is endogenous or native to the genome of a cell and is at the endogenous or native position of that target sequence in the genome of the cell. An “artificial target site” or an “artificial target sequence” are used interchangeably herein and refer to a target sequence that has been introduced into the genome of a cell. Such an artificial target sequence can be identical in sequence to an endogenous or native target sequence in the genome of a cell, but be located in a different position (i.e., a non-endogenous or non-native position) in the genome of a cell.
[0134] Methods for “modifying a target site” and “altering a target site” are used interchangeably herein and refer to methods for producing an altered target site.
[0135] As used herein, a “modified target DNA polynucleotide”, “modified DNA target site”, “modified target DNA sequence”, “modified target site”, “modified target sequence”, “modified genomic target sequence”, “modified genomic target site”, “modified target polynucleotide”, “altered target DNA polynucleotide”, “altered DNA target site”, “altered target DNA sequence”, “altered target site”, “altered target sequence”, “altered genomic target sequence”, “altered genomic target site”, and “altered target polynucleotide” are used interchangeably and refer to a target sequence as that comprises at least one alteration or modification when compared to a non- altered target sequence. Such alterations or modifications include, for example: (i) substitution of at least one nucleotide, (ii) a deletion of at least one nucleotide, (iii) an insertion of at least one nucleotide, or (iv) any combination of (i) – (iii).Attorney Reference No.: 108485-WO-SEC-1
[0136] As used herein, a “mutated gene” is a gene that has been altered through human intervention, and has a sequence that differs from the sequence of the corresponding non-mutated gene by at least one nucleotide addition, deletion, or substitution. The methods and systems described herein can be used to derive mutated genes. In a specific example, a mutated plant can comprise a mutated gene.
[0137] The terms “decreased”, “fewer”, “reduced”, “slower” and “increased”, “faster”, “enhanced”, “greater” as used herein refer to a decrease or increase in a characteristic of a modified cell or organism compared to an unmodified cell or organism (for example, a modified plant element or resulting plant compared to an unmodified plant element or resulting plant). For example, a decrease in a characteristic can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400%) or more lower than the untreated control and an increase can be at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, between 5% and 10%, at least 10%, between 10% and 20%, at least 15%, at least 20%, between 20% and 30%, at least 25%, at least 30%, between 30% and 40%, at least 35%, at least 40%, between 40% and 50%, at least 45%, at least 50%, between 50% and 60%, at least about 60%, between 60% and 70%, between 70% and 80%, at least 75%, at least about 80%, between 80% and 90%, at least about 90%, between 90% and 100%, at least 100%, between 100% and 200%, at least 200%, at least about 300%, at least about 400% or more higher than an untreated or unmodified control.
[0138] NHEJ and HDR
[0139] The guided Cas polypeptide system components described herein can be used can be used for genome-editing via inducing double-strand breaks at a DNA target site. For example, a genome editing system can comprise a Cas endonuclease, one or more guide polynucleotides, and optionally a donor DNA or polynucleotide modification template, and editing a target DNA polynucleotide comprises non-homologous end joining (NHEJ) or homology-directed repair (HDR) following a Cas endonuclease-mediated double-strand break. Once a double-strand break is induced in the DNA, the cell's DNA repair mechanism is activated to repair the break. The mostAttorney Reference No.: 108485-WO-SEC-1 common repair mechanism to bring the broken ends together is the nonhomologous end-joining pathway (Bleuyard et al., (2006) DNA Repair 5:1-12). The structural integrity of chromosomes is typically preserved by the repair, but deletions, insertions, or other rearrangements are possible (Siebert and Puchta, (2002) Plant Cell 14:1121-31; Pacher et al., (2007) Genetics 175:21-9). Alternatively, the double-strand break can be repaired by homologous recombination between homologous DNA sequences. Once the sequence around the double-strand break is altered, for example, by exonuclease activities involved in the maturation of double-strand breaks, gene conversion pathways can restore the original structure if a homologous sequence is available, such as a homologous chromosome in non-dividing somatic cells, or a sister chromatid after DNA replication (Molinier et al., (2004) Plant Cell 16:342-52). Ectopic and / or epigenic DNA sequences can also serve as a DNA repair template for homologous recombination (Puchta, (1999) Genetics 152:1173-81).
[0140] As used herein, “homologous recombination” (HR) includes the exchange of DNA fragments between two DNA molecules at the sites of homology. The frequency of homologous recombination is influenced by a number of factors. Different organisms vary with respect to the amount of homologous recombination and the relative proportion of homologous to non- homologous recombination. Generally, the length of a region of homology affects the frequency of homologous recombination events: the longer the region of homology, the greater the frequency. The length of a homology region needed to observe homologous recombination is also species- variable. In many cases, at least 5 kb of homology has been utilized, but homologous recombination has been observed with as little as 25-50 bp of homology.
[0141] In an example of the methods and systems disclosed herein, a genome-editing system comprises a Cas endonuclease, one or more guide polynucleotides, and a donor DNA. As used herein, “donor DNA” is a DNA construct that comprises a polynucleotide of interest to be inserted into the DNA target site of a Cas endonuclease. The donor DNA further comprises a first and a second region of homology flanking the polynucleotide of interest. The first and second regions of homology of the donor DNA share homology to a first and a second genomic region, respectively, present in or flanking the DNA target site of a cell or organism genome. Once a double-strand break is introduced in the DNA target site by the Cas endonuclease, the first and second regions of homology of the donor DNA can undergo homologous recombination with their corresponding genomic regions of homology resulting in exchange of DNA between the donorAttorney Reference No.: 108485-WO-SEC-1 and the target genome. As such, the provided methods result in the integration of the polynucleotide of interest of the donor DNA into the double-strand break in the DNA target site in the host genome, thereby altering the original target site and producing a modified genomic target site.
[0142] In another example of the methods and systems disclosed herein, a genome-editing system comprises a Cas endonuclease, one or more guide polynucleotides, and a polynucleotide modification template that provides the basis for template-directed repair of a double-strand break. As used herein, a “polynucleotide modification template” includes a polynucleotide that comprises at least one nucleotide modification (i.e., substitution, addition or deletion) when compared to the nucleotide sequence to be edited (i.e., DNA target site). Optionally, the polynucleotide modification template can further comprise homologous nucleotide sequences flanking the nucleotide modification, wherein the flanking homologous nucleotide sequences provide sufficient homology to the desired nucleotide sequence to be edited.
[0143] Base Editing
[0144] The guided Cas polypeptide system components described herein can be used for base editing. A base editing system comprises a base editing agent and a plurality (i.e., more than one) guide polynucleotides, and modifying a target DNA polynucleotide comprises introducing a plurality of nucleobase edits in the target polynucleotide sequence resulting in a variant nucleotide sequence. As used herein, any molecule or complex that effects a change in a nucleobase is a “base editing agent”.
[0145] One or more bases of a target DNA polynucleotide can be chemically altered to change the base from one type to another, for example, from a Cytosine to a Thymine, or an Adenine to a Guanine. A plurality of nucleobases, for example, 2 or more, 5 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more 90 or more, 100 or more, or even greater than 100, 200 or more, up to thousands of bases can be modified or altered, to produce a cell or an organism, such as a plant, with a plurality of modified bases.
[0146] Any base editing complex, such as a base editing agent (e.g., a deaminase or a protein having deaminase activity) fused to, linked to, or otherwise operably associated with (directly or indirectly) a guided Cas polypeptide can be used to recognize and bind to a desired locus in the genome of an organism and chemically modify one or more bases of a target DNA polynucleotide without creating a double-strand break.Attorney Reference No.: 108485-WO-SEC-1
[0147] A “deaminase” is an enzyme that catalyzes a deamination reaction. For example, deamination of adenine with an adenine deaminase results in the formation of inosine. Inosine selectively base pairs with cytosine instead of thymine. This results in a post-replicative transition mutation, such that the original A – T base pair transforms into a G – C base pair. In another example, cytosine deamination results in the formation of uracil, which can be repaired by cellular repair mechanisms back to a C – T base pair or to a T – A, G – C, or A – T base pair. This heterogeneity in repair can be suppressed by the introduction of a uracil glycosylase inhibitor, such that DNA repair or replication transforms the original C – T base pair into a T – A base pair (Burnett et al. (2022) Frontiers in Genome Editing. 4, 923718). In the case of both adenine and cytosine deaminases, the introduction of a nick promotes the respective base pair change (Burnett et al., 2022).
[0148] Site-specific base conversions can be achieved to engineer one or more nucleotide changes (i.e., single nucleotide polymorphisms) to create one or more edits in the genome. These include for example, a site-specific base edit mediated by an C•G to T•A or an A•T to G•C base editing deaminase (Gaudelli et al., Programmable base editing of A•T to G•C in genomic DNA without DNA cleavage." Nature (2017); Nishida et al. “Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems.” Science 353 (6305) (2016); Komor et al. “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage.” Nature 533 (7603) (2016):420-4.) A catalytically “dead” or inactive Cas endonuclease (“dCas”), fused to a cytidine deaminase or an adenine deaminase becomes a specific base editor that can alter DNA bases without inducing a double-strand DNA break. Base editors convert C->T (or G- >A on the opposite strand) or an adenine base editor that would convert adenine to inosine, resulting in an A->G change within an editing window specified by the guide polynucleotide.
[0149] A base editing deaminase, such as a cytidine deaminase or an adenine deaminase, can be fused to, linked to, or otherwise operably associated with (directly or indirectly) a guided Cas polypeptide, such as an RNA-guided dCas or a partially active (i.e., having nickase activity) Cas nickase (“nCas”) so that it does not cut a target site to which it is guided. The dCas forms a functional complex with a guide polynucleotide that shares homology with a polynucleotide sequence at the target site, and is further complexed with the deaminase molecule. The guided dCas or nCas recognizes and binds to a double-stranded target sequence, opening the double-strand to expose individual bases. In the case of a cytidine deaminase, the deaminase deaminates theAttorney Reference No.: 108485-WO-SEC-1 cytosine base and creates a uracil. Uracil glycosylase inhibitor (UGI) is provided to prevent the conversion of U back to C. DNA replication or repair mechanisms then convert the uracil to a thymine (U to T), and subsequent repair of the opposing base (formerly G in the original G-C pair) to an adenine, creating a T-A pair (Komor et al. Nature Volume 533, Pages 420-424, 19 May 2016).
[0150] Prime Editing
[0151] The guided Cas polypeptide system components described herein can be used for prime editing. A prime editing system comprises a prime editing agent and one or more guide polynucleotides, and modifying a target DNA polynucleotide comprises introducing one or more targeted insertions, deletions, or base-to-base conversions (also know as nucleobase swaps) without generating a double-strand break.
[0152] A prime editing agent can be, for example, a Cas polypeptide fused to a reverse transcriptase (RT), wherein the Cas polypeptide is modified to nick DNA rather than generating double-strand break. As used herein, “nick” refers to a single-strand break in a double-stranded DNA molecule. This nCas polypeptide-reverse transcriptase fusion can also be referred to as a “prime editor” or “PE”. A guide polynucleotide of a prime editing system can be a prime editing guide polynucleotide (pegRNA), which is larger than standard sgRNAs commonly used for CRISPR genome editing (e.g., >100 nucleobases). The pegRNA comprises a primer binding sequence (PBS) and a RT template containing the desired or target RNA sequence at its 3’ end.
[0153] During prime editing, the PE:pegRNA complex binds to a target DNA polynucleotide and the modified Cas polypeptide nicks one target DNA strand resulting in a flap. The PBS on the pegRNA binds to the DNA flap, and the target RNA sequence of the RT template is reverse transcribed using the reverse transcriptase. The edited strand is incorporated into the target DNA polynucleotide at the end of the nicked flap, and the target DNA strand is repaired with the new reverse transcribed DNA.
[0154] Performance Objective Model
[0155] The methods and systems described herein utilize heuristic or metaheuristic search algorithms (also described herein as a “performance objective model”), which solve for (e.g., maximize) the primary performance objective while satisfying the constraints of the one or more secondary performance objectives for each candidate genotype. This algorithm-based enhanced breeding decision enables the breeder or automated system to optimize selection decisions towardAttorney Reference No.: 108485-WO-SEC-1 the primary performance objective (e.g., a primary trait) while controlling for additional secondary performance objectives (e.g., secondary traits), an optional diversity target, and optional germplasm use constraints.
[0156] As shown in FIGS. 1A-1C, the algorithm input includes a unique identifier for each candidate genotype, a primary performance objective to be maximized, one or more secondary performance objectives with constraints, primary phenotypic data, secondary phenotypic data, a target number of candidate genotypes to be selected or returned by the mathematical solver, an optional diversity target value and accompanying genotypic data for each genotype candidate, and optional germplasm use constraints.
[0157] The algorithm reads in the unique candidate genotype identifiers and primary and secondary performance objectives, and accompanying phenotypic data for each candidate genotype, applies data processing steps for normalization and penalization of the primary performance objective based on the constraints of the secondary performance objective(s), and generates an advancement value for each candidate genotype. It will be understood that the advancement value is based solely on the primary and secondary performance objectives.
[0158] The algorithm then creates an objective function with multiple parameters including the advancement value for each candidate genotype, the target number of candidate genotypes to be selected or returned by the mathematical solver, the optional diversity target value and accompanying genotypic data for each genotype candidate, and the optional germplasm use constraints, and provides the objective function to a mathematical solver, which solves the objective function and returns the target number of candidate genotypes meeting the selection criteria. As used herein, “creating an objective function” or “defining an objective function”, refers to translating the advancement value of each candidate genotype, the target number of candidates genotypes to be returned by the mathematical solver, the optional diversity target value and accompanying genotypic data for each genotype candidate, and the optional germplasm use constraints into a mathematical equation that is subsequently provided to a mathematical solver. As used herein, “solving the objective function” will be understood to mean identifying the target number of candidate genotypes that maximize the primary performance objective and meet the optional diversity target value and optional germplasm use constraints. The resulting target number of selection candidates can then be evaluated simultaneously for enhanced breeding decisions such as crosses to make, populations to create, and progeny (e.g., hybrids) to advance.Attorney Reference No.: 108485-WO-SEC-1
[0159] In an example of the methods and systems described herein, the method can be further fine- tuned such as by eliminating selection candidates that do not meet secondary performance objective constraints prior to running the algorithm, decreasing the penalties on secondary traits where a candidate genotype readily meets the secondary trait constraints, and / or increasing the penalties on secondary traits where a candidate genotype does not readily meet the secondary trait constraints. Further, the breeder / user or system can alter the objective function such as by lessening the diversity target and / or adjusting the germplasm use constraints. In a specific example, after an initial run and algorithm-based selection, the breeder or system can adjust the penalty score of one or more secondary traits based on the breeder or system ranking of secondary traits. In some examples, such as when there are multiple secondary performance objectives, some secondary traits are more important and penalty values of lesser important traits are adjusted accordingly. In other examples where many of the candidate genotypes readily satisfy one or more secondary performance objectives (i.e., meet the threshold for the given secondary performance objective), the readily satisfied secondary performance objectives have a lesser penalty value leading the algorithm to emphasize or weight harder-to-achieve secondary performance objectives.
[0160] Using the algorithm-based methods and systems described herein, breeders or systems can assess multiple selection strategies and consolidate selections across multiple breeding objectives. In some examples, the algorithm can be run with multiple performance objectives simultaneously. In other examples, the algorithm is run sequentially where the output of selected candidate genotypes meeting the selection criteria is used as input for a sequential run to select a subset from the candidate genotypes those meeting one or more additional performance objectives. Alternatively or in addition to, in some examples, the algorithm is run on candidate genotypes for different performance objectives in parallel. The candidate genotypes meeting the selection criteria from these different runs may be pooled together and the target number of candidates meeting selection criteria selected. Further, these methods allow for rapid iteration across a near-infinite set of potential combinations that cannot realistically be conducted in an efficient timeframe. Other advantages of the disclosed methods include providing consistent methodology for applying selections across years as well as across several stages of the breeding pipeline, the potential to expand selections to cover even more dimensions (ex: “real-time / in-season” data), and the robustness of maintaining selections according to predefined performance objectives.Attorney Reference No.: 108485-WO-SEC-1
[0161] The methods and systems described herein can utilize several metaheuristic search methods such as evolutionary algorithms (e.g., genetic algorithm, genetic programming, evolutionary programming, evolution strategy, differential evolution, a coevolutionary algorithm, neuroevolution) or Tabu search, learning classifier systems, probabilistic techniques such as simulated annealing, and / or heuristic search algorithms such as Beam search or Greedy to define the optimal set of combinations to advance.
[0162] In a particular example of the methods and systems described herein, the algorithm used to solve the mathematical representation of the combinatorial problem, is a genetic algorithm with dynamic solution updates to the objective function(s) (i.e., the primary performance objective or objectives) and secondary parameters (i.e., the secondary performance objective or objectives).
[0163] The objective function for the methods and systems described herein can utilize a mathematical formula as described below:
[0164] Assuming a set ^^ of several candidate genotypes ^^ for potential selection, the intent is to maximize the response to selection on a primary performance objective primary trait (^^^^) (e.g., a primary trait) while adjusting for penalties associated with the secondary performance objective(s) ^^ (e.g., secondary traits) and meeting diversity target value ^^. The objective function can be defined as:^^′: a vector of binary values [0,1] such that ^^′^^ is equal to the target number of candidate selections to be returned ^^^^: a vector of phenotypic data for candidate genotypes related to the primary performance objective ^^^^: the average value of the primary performance objective ^^^^: the standard deviation of the primary performance objective ^^^^: a vector of phenotypic data for candidate genotypes related to the secondary performance objective ^^Attorney Reference No.: 108485-WO-SEC-1 ^^^^: the average value of secondary performance objective ^^ ^^^^: the standard deviation of secondary trait ^^ ^^^^: a vector of binary values [0,1] such as ^^^^: penalty value associated with secondary trait ^^ ^^: a binary variable defined according to the diversity target (t ) that can be used for rewarding / penalizing diversity gain / loss. diversity ismetric that indicates the level of relatedness among the selected set of candidates. Several diversity metrics can be used such as average co-ancestry, average genomic relationship, average kinship, effective population size, rate of homozygosity etc... ^^: the penalty associated with diversity target value when diversity target is not met. This value becomes the reward for excess diversity when diversity target is met and the strategy requires rewarding excess diversity Kj : a binary variable [0,1] such as with j indicating the type of germplasm use constrain used (ex: parents, grandparents and / or any combination of pedigree / germplasm use restrictions to be imposed) gj : the penalty associated with germplasm use limit when germplasm use limit is exceeded for its corresponding germplasm use constraint
[0165] To solve ^^′ (the vector of binary values [0,1]) that maximizes the objective function, a Binary Genetic Algorithm was used.Attorney Reference No.: 108485-WO-SEC-1
[0166] In another example, the objective function can be extended to include additional constraints related to accumulation of desired native traits, genome edits, or any other categorical attributes at desired frequencies through the incorporation of additional related penalty terms similar to the one used for the germplasm use constraint. For example, a constraint related to the accumulation ofdesired genome edit frequencies can be added by including the penalty term (−^^^^ ∗ ^^^^) to theobjective function: ′(^^1−^^1)′(^^^^^^((^^^^−^^^^))∗^^ ∗(^^^^′^^−^^^^ ^^^^^^)^^^^^^^^1∑^^ ^^ ^^=2^^′^^− ^^ ∗ ^^ − ^^^^ ∗ ^^^^ − ^^^^ ∗ ^^^^)where edit frequency and ^^^^is a vectorin the case of hard constraints on edit frequencies.
[0167] Alternatively, ^^^^could be a vector of ratios of the desired genome edit frequency to the realized genome edit frequency in such a way that:
[0169] As used herein, a “primary performance objective” refers to the highest priority phenotypic attribute or characteristic of interest to be achieved by a candidate plant genotype as defined in the algorithm of the methods and systems described herein. The primary performance objective can be a target trait, referred to herein as the “primary target trait”, “primary target trait of interest”, or “primary trait”. A target trait of interest includes traits of agronomic or economic importance or interest, such as but not limited to, those disclosed in Table 1.
[0170] In the methods and systems described herein, the primary performance objective of candidate genotypes can be maximized or minimized. Preferentially, the primary performance objective is maximized. As shown in FIGS.1A-1C, a primary performance objective is selected by a user or automated system, and is provided to the algorithm along with phenotypic data for each candidate genotype related to the primary performance objective (also known as “primary phenotypic data”). The algorithm subsequently utilizes the primary phenotypic data of each candidate genotype to calculate its advancement value, reflective of maximizing the primaryAttorney Reference No.: 108485-WO-SEC-1 performance objective. For example, if plant yield is selected as the primary performance objective, the algorithm would generate an advancement value for each candidate genotype based on its primary phenotypic data related to plant yield. Table 1: Traits adaptability moisture content root architecture (e.g., root lodging resistance) n )
[0171] Secondary Performance Objective(s)
[0172] As used herein, a “secondary performance objective” refers to one or more secondary phenotypic attributes or characteristics of interest to be achieved by a candidate plant genotype as defined in the algorithm of the methods and systems described herein. Each secondary performance objective can be a target trait described herein as the “secondary target trait”, “secondary target trait of interest”, or “secondary trait”. Secondary target traits include traits of agronomic or economic importance, such as but not limited to, those disclosed in Table 1.
[0173] As shown in FIGS.1A-1C, one or more secondary performance objectives are selected by a user or an automated system, and are provided to the algorithm along with phenotypic data for each candidate genotype related to the secondary performance objective(s) (also known as “secondary phenotypic data”). The algorithm subsequently utilizes the secondary phenotypic data of each candidate genotype to calculate its advancement value, specifically, adjusting theAttorney Reference No.: 108485-WO-SEC-1 advancement value for each candidate genotype by a constraint imposed on the secondary performance objective. Each of the secondary performance objectives serve as a constraint for the primary performance objective of the algorithm. As used herein, “constraint” refers to a desired limitation that is set on the secondary performance objectives by a user (e.g., breeder) or an automated system. The constraint of each secondary performance objective is a user-determined (or automated-system determined) numerical value, threshold, or range to be optimized by a candidate genotype. In the algorithm, each secondary performance objective constraint can impose a numerical penalty value on the advancement value of each candidate genotype, whereby candidate genotypes that maximize the primary performance objective while also satisfying or meeting the secondary performance objective constraints (e.g., being equal to or greater than a minimum threshold value for a secondary performance objective, being less than or equal to a maximum threshold value for a secondary performance objective, or being within an acceptable numerical range for a secondary performance objective). For example, a user may define the primary performance objective as plant yield and define secondary performance objectives as plant height, herbicide tolerance, and seed oil composition. The algorithm would calculate an advancement value for each candidate genotype (i.e., reflecting the degree to which the phenotypic data related to plant yield for each candidate genotype maximizes plant yield), and adjust (i.e., penalize) the advancement value based on the ability of the secondary performance objective phenotypic data of each candidate genotype to satisfy or meet the target constraints for plant height, herbicide tolerance, and seed oil composition. In this way, the algorithm indicates or returns candidate genotypes that maximize a primary performance objective, such as plant yield, while also providing acceptable phenotypic outcomes for secondary performance objectives, such as plant height, herbicide tolerance, and seed oil composition.
[0174] Optionally, the constraints imposed by the secondary performance objectives can be weighted in the algorithm. “Weighted”, in regards to a secondary performance objective, means that the secondary performance objectives can have varying degrees of importance within the algorithm, as ranked or indicated by a user or automated system. As such, the penalty imposed on one secondary performance objective can be more or less stringent than the penalty imposed on another secondary performance objective. For example, an acceptable range or value for seed oil composition may allow for a margin of error (i.e., a permissible or tolerable degree of deviation from the acceptable range) while the acceptable range or value for plant height does not. In thisAttorney Reference No.: 108485-WO-SEC-1 way, the penalty for plant height as a secondary performance objective is weighted in comparison to the penalty for seed oil composition as a secondary performance objective.
[0175] Diversity Target Value
[0176] The methods and systems described herein can further evaluate candidate genotypes using an optional diversity target value, which is defined by a user or automated system. As used herein, “diversity” refers to a metric to calculate or evaluate the degree of relatedness amongst the selection candidates. In a particular example of the methods and systems described herein, the diversity input (also known as reference data for diversity) for the algorithm is based on co- ancestry. Other suitable diversity metrics include effective population size, genic variance, pedigree- and / or genomic-based similarity, co-ancestry, or relationship, fixation rate (e.g., genome-wide or genomic region-specific).
[0177] In an example, the diversity value is a numerical value from 0-1 to be met or achieved by a selected genotype candidate. The diversity target value can be calculated according to the average relationship or diversity among parents, such as in the case of hypothetical candidate genotypes. Alternatively, the diversity target value can be calculated for realized candidate genotypes.
[0178] As shown in FIGS.1A-1C, reference data for diversity, which incudes a diversity target value and genotypic data for each genotype candidate, can be provided to the algorithm. The algorithm utilizes the diversity target value and genotypic data, in conjunction with the advancement values for each genotype candidate, to create or define an objective function to be provided to a mathematical solver.
[0179] Germplasm Use Constraints
[0180] The methods and systems described herein can further evaluate candidate genotypes using one or more optional germplasm use constraints.
[0181] A first exemplary germplasm use constraint, referred to herein as a “genetic background constraint” are breeder or system selected limits on use of a genetic background (e.g., plant parent, grandparent, etc.). For example, parent inbreds can be ranked with breeder-defined limits on the number of times a ranked inbred can be selected from the candidate genotypes (e.g., do not select “rank inbred 1” more than 10 times). Such rankings can be, but are not required to be, inbred performance from one growing season to the next. Using a genetic background constraint in the methods and systems described herein can promote genetic diversity by minimizing or reducing use of existing genetic backgrounds or closely related genetic backgrounds.Attorney Reference No.: 108485-WO-SEC-1
[0182] A second exemplary germplasm use constraint allows a breeder to minimize or reduce the uncertainty related to the primary phenotypic data and / or the secondary phenotypic data of each candidate genotype, either of which can be hypothetical or realized. This uncertainty, referred to herein as “risk mitigation factor” or “risk mitigation level”, can be a numerical measure associated with the primary and / or secondary phenotypic data such as accuracy or error measurement, a categorical score (ex: 1, 2, 3, …, 9) or a classification score (low, medium, high) that indicates the degree of confidence the breeder and / or system should take into consideration with regards to the primary and / or secondary performance objectives based on hypothetical or observed phenotypic data of the candidate genotypes.
[0183] This “phenotype confidence target” refers to the achievement of a desired or target confidence level according to the risk mitigation factors described above. A phenotype confidence target can be numerical risk mitigation levels below a desirable value or within a desirable range, a certain distribution of risk categorical scores or combination thereof (ex: 50% score 1, 30% score 2, 20% score 9 or 30% score 1 by score 1, 20% score 1 by score 2, 15% score 2 by score 3, 10% score 2 by score 4, 5% score 2 by score 9, 20% score 9 by score 9), a certain distribution of risk classification scores (50% low, 30% medium, 20% high), or any combination thereof. Risk mitigation factors and phenotype confidence targets can be based on statistical or subjective confidence in a candidate genotype. For example, for a realized candidate genotype, confidence can be based on observed performance of the genotype over a one or more growing seasons. In another example, a phenotype confidence target can encompass confidence in projected performance of a hypothetical candidate genotype, favoring hypothetical candidate genotypes having positive projected performance attributes for advancement.
[0184] As shown in FIGS.1A-1C, germplasm use constraints (e.g., genetic background constraint and phenotype confidence target) can be provided to the algorithm. The algorithm utilizes the germplasm use constraints, in conjunction with the advancement values for each genotype candidate and optional reference data for diversity, to create or define an objective function to be provided to a mathematical solver. EXAMPLES Example 1: Parameters and Constraints for Genotype Selection
[0185] The algorithm used in Examples 3-7 utilized the following:Attorney Reference No.: 108485-WO-SEC-1
[0186] Assuming a set ^^ of several candidate genotypes ^^ for potential selection, the intent is to maximize the response to selection on a primary trait (^^^^) while adjusting for penalties associated with the secondary trait(s) ^^ and meeting diversity target ^^. The objective function can be defined as:^^′: a vector of binary values [0,1] such that ^^′^^ is equal to the number of candidate selections to be returned (target number) ^^^^: a vector of phenotypic data for candidate genotypes related to the primary trait ^^ ^^^^: the average value of the primary trait ^^ ^^^^: the standard deviation of the primary trait ^^ ^^^^: a vector of phenotypic data for candidate genotypes related to the secondary trait ^^ ^^^^: the average value of secondary trait ^^ ^^^^: the standard deviation of secondary trait ^^ ^^^^: a vector of binary values [0,1] such as ^^^^: the penalty value associated with secondary trait ^^ ^^: a binary variable [0,1] that indicates no reward is provided for excessive diversity and a hard penalty is imposed when diversity target (t ) is not metwhere ^^ is a vector of parent contributions and ^^ the genomic based coancestry matrix representing the genomic relationship between the parents of the candidates under selectionAttorney Reference No.: 108485-WO-SEC-1 ^^: the penalty associated with diversity when diversity target value (t ) is not met K1 : a binary variable [0,1] such asK2 : a binary variable [0,1] such asassociated with any grandparent use exceeding grandparent use limit
[0187] To solve ^^′ (the vector of binary values [0,1]) that maximizes the objective function, a Binary Genetic Algorithm was used. Example 2: Candidate Genotypes for Algorithm-based Genotype Selection
[0188] Examples 3-7 investigated the algorithm-based genotype selection methods described herein to select for different breeding objectives. The data utilized in these Examples comprised approximately 36,000 candidate genotypes for potential selection with the intent to narrow the search space to 500 selections (i.e., the target number of candidate genotypes to be selected). The 36,000 candidate genotypes were generated from a list of 334 parents (CTVA_1 to CTVA_334) crossed at an average rate of 210 bi-parental crosses per parent.
[0189] The selection algorithm had input parameters related to a primary performance objective to be maximized (referred to in the Examples as the primary trait), the target number of candidate genotypes to be selected, primary phenotypic data for each candidate genotype, continuous and / or categorical secondary performance objectives (secondary traits), secondary phenotypic data for each candidate genotype, a diversity target value, and germplasm use constraints related to parent use and grandparent use. The primary trait to be maximized in Examples 3-7 was predicted yield and moisture index value. Example 3: Selection Method without Constraints for a Secondary TraitAttorney Reference No.: 108485-WO-SEC-1
[0190] In this Example, the algorithm and data set of Examples 1 and 2 were used to select the top 500 candidate genotypes for maximization of grain yield (the primary trait) without any secondary trait constraints. The methods described herein utilized the algorithm and calculated the advancement value for each candidate genotype based on the primary trait without any secondary trait constraints, created the objective function with the advancement values and a 500 candidate target number, and provided the objective function to the mathematical solver.
[0191] FIG.2 illustrates the distribution of all (~36,000) candidate genotypes relative to a set of reference checks for the primary trait (the reference checks not being part of the candidates for selection). FIG. 3 illustrates the distribution of the mathematical solver algorithm output (500 selections) relative to a set of reference checks for the primary trait (the checks not being part of the selected candidate genotypes).
[0192] These data demonstrate that without any secondary trait constraints, the mathematical solver selected the top 500 candidate genotypes based solely on maximization of the primary trait. This is equivalent to sorting the candidates according to the phenotypic data related to the primary trait and selecting the top 500 individuals based on the primary trait. Example 4: Selection Method with Constraints for a Secondary Trait
[0193] In this Example, the algorithm and data set of Examples 1 and 2 were used to select the top 500 candidate genotypes for maximization of the primary trait with a single secondary trait (Sec_Trait_4), which was predicted moisture value. The secondary trait values were constrained to between 16.3 and 19.3. Candidate genotypes with a secondary trait value outside of the accepted range (i.e., 16.3-19.3) incurred a penalty, with the penalty value increasing as the secondary trait value deviated further from the accepted range. Specifically, the penalty was 2 points on the normalized advancement value of the target trait for each standard deviation of the secondary trait value from the accepted range. Secondary trait values falling within the accepted range incurred a “0” penalty value. The algorithm calculated the advancement value for each candidate genotype based on the primary trait and the secondary trait, created the objective function with the advancement values and a 500 candidate target number, and provided the objective function to the mathematical solver.
[0194] FIG.4 illustrates the distribution of all candidate genotypes for selection for the secondary trait relative to a set of reference checks and the boundaries of the accepted range for the secondaryAttorney Reference No.: 108485-WO-SEC-1 trait. FIG.5 illustrates the distribution of the mathematical solver output (500 selections) relative to a set of reference checks for the secondary trait. These data demonstrate that a majority of the selections recommended by the algorithm fall within the accepted range (i.e., 16.3-19.3) for the secondary trait.
[0195] FIG. 6 illustrates the distribution of the mathematical solver output (500 candidates) relative to a set of reference checks for the primary trait while accounting for the secondary trait constraints. In comparison to FIG.3 (distribution of the algorithmic recommended selections for the primary trait without any secondary trait constraints), the recommended selections for the primary trait when constrained on the secondary trait (FIG. 6) were in the higher tail of the distribution. With the secondary trait constraints, only 181 individuals were among the top 500 individuals for the primary trait identified with no secondary trait constraints.
[0196] FIG.7 illustrates the percentage gain (~2.47%) of the mathematical solver recommended selections vs the diversity value over the average of the entire set of candidate genotypes when maximizing the primary trait only (i.e., no secondary trait constraints; Example 3). FIG. 8 illustrates the percentage gain (~2.11%) of the mathematical solver recommended selections vs the diversity value when maximizing the primary trait with the secondary trait constraints.
[0197] FIG. 9 illustrates the relationship between the primary trait and the secondary trait (Sec_Trait_4), and the distribution of mathematical solver recommended selections (black dots), maximizing for the primary trait without any constraints on the secondary trait. By contrast, as shown in FIG.10, when the secondary trait is constrained to the accepted range (i.e., 16.3-19.3), the algorithm maintained most of the selections within the secondary trait constraints, while dropping other individuals (greyed out dots).
[0198] Table 2 provides an overview of the inbred use (number of mathematical solver selections per parent) to maximize the primary trait with and without constraints on the secondary trait. When maximizing the primary trait without secondary trait constraints, the algorithm selections included 13 inbred parents with an average of 39 crosses per parent, which indicated a high degree of half- siblings being selected. When the primary trait is constrained by the secondary trait, the algorithm selections were still driven by 13 inbred parents with an average of 38 crosses per parent. The selection contribution of each parent changed. Parents CTVA_133, CTVA_69, CTVA_134 and CTVA_170 were removed, while parents CTVA_280, CTVA_140, CTVA_203, and CTVA_113 were added.Attorney Reference No.: 108485-WO-SEC-1
[0199] The candidate genotypes were grouped by ranks order of 500 candidates (from best to least performers according to the primary trait). The algorithm recommended the top 500 candidates for the primary trait when no constraints were used whereas the recommended candidate genotypes were spread across the different group ranks when subjected to constraints for secondary trait Sec_Trait_4. The distribution of selected candidate genotypes is shown in Table 3 (group orders where no candidate genotypes were recommended by the algorithm were dropped for ease of representation, for example group orders 6001-6500, 6501-7000, 7001-7500, 7501-8000). With the secondary trait constraints, only 181 candidate genotypes were among the top 500 candidates for the primary trait identified with no secondary trait constraints. Table 2: Number of Selections Per Parent Number of Selections Parent Without constraints on With constraints onTable 3: Rank Groups of Algorithm Selections Primary Optimize for Primary Trait without Optimize for Primary Trait withAttorney Reference No.: 108485-WO-SEC-1 501-1000 0 1 1501-2000 0 3Attorney Reference No.: 108485-WO-SEC-1 23001- 0 19 23500Example 5: Selection Method to Maximize a Primary Trait with a Diversity Target Value
[0200] In this Example, the algorithm and data set of Examples 1 and 2 were used to select the top 500 candidates to maximize the primary trait with a diversity target value. Specifically, the diversity target value was set to 0.33, which was selected based on coancestry (with diversity calculated as 1-coancestry). No secondary trait constraints were used in this Example.Attorney Reference No.: 108485-WO-SEC-1
[0201] FIG. 11 illustrates the distribution of the mathematical solver output (500 candidates) relative to a set of reference checks for the primary trait while accounting for the diversity constraint (≥0.33). In comparison to FIG. 3 (distribution of the algorithmic recommended selections for the primary trait without any secondary trait or diversity constraints), the recommended selections for the primary trait with a diversity constraint (FIG. 11) were not selected from the top end tail of the primary trait.
[0202] FIG.12 illustrates the percentage gain (~0.24%) of the mathematical solver recommended selections vs the diversity value when maximizing the primary trait with the diversity constraint. Although several feasible selections (grey dots) were identified that increased diversity beyond the 0.33 diversity threshold, a “best” selection was identified as a combination of selections with the highest rate of gain that met the diversity constraint.
[0203] FIG. 13 illustrates the relationship between the primary trait and the secondary trait (Sec_Trait_4), and the distribution of mathematical solver recommended selections (black dots), maximizing for the primary trait when only constrained by a minimum diversity target value (without any constraints on the secondary trait).
[0204] Table 4 provides an overview of the inbred use (number of mathematical solver selections per parent) to maximize the primary trait with a diversity constraint. The mathematical solver selections were spread across 81 inbred parents with an average of 6 crosses per parent, which indicated the degree of half-siblings significantly decreased as compared to Example 4. Table 4: Number of Selections Per Parent with Diversity Target Value Parent Selections Parent SelectionsAttorney Reference No.: 108485-WO-SEC-1 CTVA_201 8 CTVA_133 5 CTVA_222 8 CTVA_138 5Example 6: Selection Method to Maximize a Primary Trait with Constraints for a Secondary Trait and a Diversity Target Value
[0205] In this Example, the algorithm and data set of Examples 1 and 2 were used to select the top 500 candidate genotypes maximize the primary trait with a single secondary trait (Sec_Trait_4) constrained to between 16.3 and 19.3 and a minimum diversity target value of 0.33..
[0206] FIG.14 illustrates the distribution of the mathematical solver output (500 candidates) for the secondary trait relative to a set of reference checks and subject to the diversity constraint. FIG. 15 illustrates the distribution of the mathematical solver output (500 candidates) relative to a setAttorney Reference No.: 108485-WO-SEC-1 of reference checks for the primary trait while accounting for the secondary trait constraints (16.3 and 19.3) and a minimum diversity value of 0.33. These data demonstrate an improvement in the mathematical solver selected set in contrast to the reference checks when compared to FIG.2.
[0207] FIG.16 illustrates the percentage gain of the mathematical solver recommended selections vs the diversity value over the average of the entire set of selected candidate genotypes when maximizing the primary trait with secondary trait and diversity constraints. A “best” selection was identified as a combination of candidate genotype selections with the highest rate of gain that met the diversity constraint with a minimum accumulation of penalties on the secondary trait. FIG.17 illustrates the relationship between the primary trait and the secondary trait (Sec_Trait_4), and the distribution of mathematical solver recommended selections (black dots), when the target trait is constrained by the secondary trait accepted range and diversity target value.
[0208] Table 5 provides an overview of the inbred use (number of mathematical solver selections per parent) to maximize the primary trait with the secondary trait and diversity constraints. The mathematical solver selections were spread across 76 inbred parents with an average of 6 crosses per parent. Table 5: Number of Selections Per Parent with Secondary Trait and Diversity Constraints Parent Selections Parent SelectionsAttorney Reference No.: 108485-WO-SEC-1 CTVA_321 8 CTVA_95 5 CTVA_18 7 CTVA_99 5Example 7: Selection Method to Maximize a Primary Trait with Multiple Secondary Trait Constraints, a Target Diversity Value, and Germplasm Use Constraints
[0209] In this Example, the algorithm and data set of Examples 1 and 2 were used to select the top 500 candidate genotypes to maximize the primary trait with the following constraints: 21 secondary traits having the constraints and penalties as defined in Table 6, a minimum diversity target value of 0.33, and a limit on germplasm use (i.e., a genetic background constraint) (parent use < 10 crosses and grandparent use < 40 crosses). Table 6 also includes descriptions for secondary traits Sec_Trait_1 – Sec_Trait_21.
[0210] FIG. 18 – FIG. 39 illustrates the distribution of all candidate genotypes for selection of the primary trait (FIG. 18) and all secondary traits (FIGS. 19-39) relative to a set of reference checks. FIG.40 illustrates the percentage gain of the mathematical solver recommended selections vs the minimum diversity target value over the average of the entire set of selection candidatesAttorney Reference No.: 108485-WO-SEC-1 when maximizing the primary trait with secondary traits, a minimum diversity target value, and a limit on germplasm use. Table 7 illustrates the parent and grandparent use limits in the mathematical solver selected candidates, which indicates that the germplasm use constraints were satisfied.
[0211] In this Example, the method was able to make selections that met the constraint criteria for all the secondary traits, except Sec_Trait_4. Nevertheless, the method selections for Sec_Trait_4 were an improved set over the reference checks and close to the constraint zone of 16.3 to 19.3. Table 6: Secondary Trait Threshold Boundaries and Penalties Secondary Trait Descr Lower Upper Lower Upper Trait iption Boundary Boundary Penalty PenaltyAttorney Reference No.: 108485-WO-SEC-1 Sec_Trait_15 Predicted overall standability value 2 90 0.5p Minimum Mean Maximum Selection Parent Limits 1 167 6
Claims
Attorney Reference No.: 108485-WO-SEC-1 CLAIMS 1. A method of advancing plant genotypes to create a plant having one or more target phenotypes, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii)primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; andAttorney Reference No.: 108485-WO-SEC-1 (f) performing one or more phenotypic analyses on a selected genotype.
2. A method of advancing plant genotypes to create a plant having an improved phenotype, the method comprising: (g) providing a performance objective model with: (vii) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (viii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (ix) primary phenotypic data for each candidate genotype related to the primary performance objective; (x) a secondary performance objective to be satisfied by the selected genotype; (xi) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (xii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (h) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (i) creating an objective function for a solver, wherein the objective function is created by the performance objective model and defined by: (iii)the advancement values calculated by the performance objective model; and (iv) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (j) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (k) advancing one or more selected genotypes from the target number of selected genotypes; andAttorney Reference No.: 108485-WO-SEC-1 (l) performing one or more phenotypic analyses on a selected genotype.
3. A method of selecting plant genotypes for an improved population, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii)primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) pollinating a plant of the selected genotype to generate a population.Attorney Reference No.: 108485-WO-SEC-1 4. A method of selecting plant genotypes for an improved cross, the method comprising: (a) providing a performance objective model with: (i) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (ii) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes; (iii)primary phenotypic data for each candidate genotype related to the primary performance objective; (iv) a secondary performance objective to be satisfied by the selected genotype; (v) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (vi) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by the performance objective model and defined by: (i) the advancement values calculated by the performance objective model; and (ii) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective; (e) advancing one or more selected genotypes from the target number of selected genotypes; and (f) crossing a plant of the selected genotype to create an inbred or a hybrid plant.
5. The method of any one of claims 1-4, wherein the performance objective model is further provided:Attorney Reference No.: 108485-WO-SEC-1 (vii) a diversity target value to be satisfied by the target number of selected genotypes, wherein the objective function created by the performance objective model is further defined by the diversity target value, and wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value.
6. The method of claim 5, wherein the performance objective model is further provided: (viii) one or more germplasm use constraints, wherein the objective function created by the performance objective model is further defined by the one or more germplasm constraints, and wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and the one or more germplasm use constraints.
7. The method of any one of claims 1-6, wherein the plurality of candidate genotypes comprises: (a) a realized genotype from an inbred or hybrid plant; (b) a hypothetical genotype from an inbred or hybrid plant; (c) a genotype resulting from a realized breeding cross; (d) a genotype resulting from a hypothetical breeding cross; (e) a combination of a. – d.
8. The method of any one of claims 1-7, wherein primary phenotypic data and / or the secondary phenotypic data for each candidate genotype comprises hypothetical phenotypic data.
9. The method of claim 8, wherein the candidate genotype comprising the hypothetical phenotypic data is a realized genotype.Attorney Reference No.: 108485-WO-SEC-1 10. The method of claim 8, wherein the candidate genotype comprising the hypothetical phenotypic data is a hypothetical genotype.
11. The method of any one of claims 1-10, wherein the performance objective model is a genetic algorithm.
12. The method of claim 11, wherein the performance objective model is a heuristic search algorithm or a metaheuristic search algorithm.
13. The method of any one of claims 1-12, wherein adjusting the advancement value by the constraint imposed on the secondary performance objective comprises imposing a penalty value on the advancement value of each candidate genotype.
14. The method of any one of claims 1-13, wherein the constraint imposed on the secondary performance objective is a range of numerical values.
15. The method of any one of claims 1-13, wherein the constraint imposed on the secondary performance objective is a minimum or maximum numerical value representing a minimum or maximum threshold to be satisfied by the selected genotype.
16. The method of any one of claims 1-15, wherein the primary performance objective is a trait.
17. The method of any one of claims 1-16, wherein the secondary performance objective is a trait.
18. The method of any one of claims 5-17, wherein the diversity target value is a minimum numerical value representing a minimum threshold to be satisfied by the selected genotype.
19. The method of any one of claims 1, 2, or 5-18, further comprising growing a plant of the selected genotype.Attorney Reference No.: 108485-WO-SEC-1 20. The method of any one of claims 3 or 5-18, further comprising growing the population of plants having the selected genotype.
21. The method of any one of claims 4-18, further comprising growing the inbred or hybrid plant.
22. The method of any one of claims 1-21, wherein the plurality of candidate genotypes comprises candidate genotypes having one or more genomic modifications introduced via a site- specific genome-editing agent.
23. The method of claim 22, wherein the site-specific genome-editing agent is a zinc finger nuclease, TALEN, homing endonuclease, Cas polypeptide, TnpB nuclease, or a hydrolytic endonucleolytic ribozyme.
24. The method of claim 22, wherein the one or more genomic modifications comprises an insertion, a deletion, a single nucleotide polymorphism, an inversion, or a translocation.
25. A computing device comprising a processor configured to perform the method of any one of claims 1-24.
26. A computer-readable medium comprising instructions, which when executed by a computing device, cause the computing device to carry out the method of any one of claims 1-24.
27. A system for advancing plant genotypes to create a plant having one or more target phenotypes, the system comprising: (i) one or more servers for storing data, the data comprising: (a) a plurality of candidate genotypes, each candidate genotype having a unique identifier; (b) a primary performance objective to be maximized by a selected genotype, wherein the selected genotype is chosen from the plurality of candidate genotypes;Attorney Reference No.: 108485-WO-SEC-1 (c) primary phenotypic data for each candidate genotype related to the primary performance objective; (d) a secondary performance objective to be satisfied by the selected genotype; (e) secondary phenotypic data for each candidate genotype related to the secondary performance objective; and (f) a target number of selected genotypes to be returned from the plurality of candidate genotypes; (ii) a computing device communicatively coupled to the one or more servers, the computing device including a memory and one or more processors configured to carry out: (a) providing a performance objective model with the plurality of candidate genotypes, the primary performance objective, the primary phenotypic data for each candidate genotype, the secondary performance objective, the secondary phenotypic data for each candidate genotype, and the target number of selected genotypes to be returned; (b) calculating an advancement value for each candidate genotype based on the primary phenotypic data and the secondary phenotypic data of each candidate genotype, wherein the advancement value for each candidate genotype maximizes the primary performance objective and is adjusted by a constraint imposed on the secondary performance objective; (c) creating an objective function for a solver, wherein the objective function is created by a performance objective model and defined by the advancement values calculated by the performance objective model and the target number of selected genotypes to be returned from the plurality of candidate genotypes; and (d) executing the solver and collecting the mathematical solution output, wherein the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective.
28. The system of claim 27, wherein the data further comprises a diversity target value to be satisfied by the target number of selected genotypes, the objective function created by the performance objective model is further defined by the diversity target value, and the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value.Attorney Reference No.: 108485-WO-SEC-1 29. The system of claim 27, wherein the data further comprises one or more germplasm use constraints, the objective function created by the performance objective model is further defined by the one or more germplasm constraints, and the mathematical solution output is the target number of selected genotypes that maximize the primary performance objective and satisfy the diversity target value and the one or more germplasm use constraints.