Trait introgression process optimisation
The MILP-based optimization of breeding process variables addresses the complexity in trait introgression by optimizing screening strategies across generations, enhancing efficiency and confidence in achieving desired traits within resource limitations.
Patent Information
- Application Number
- PCT/AU2025/050343
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-10
- Filing Date
- 2025-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
The trait introgression process in plant breeding is hindered by the complexity of numerous variables and combinations, leading to prohibitive computational requirements and inefficiencies in optimizing the breeding process for introducing desired traits into elite varieties.
A method utilizing a cost model optimized through mixed integer linear programming (MILP) to calculate breeding process variables across multiple generations, incorporating decision variables for screening methods and constraints to ensure feasibility and efficiency, thereby optimizing the trait introgression process.
This approach reduces computational burden and improves the breeding process by determining optimal screening strategies across generations, ensuring cost-effectiveness and high confidence in achieving desired traits within limited resources and time constraints.
Smart Images

Figure AU2025050343_16102025_PF_FP_ABST
Abstract
Description
"Trait introgression process optimisation" Technical Field
[0001] This disclosure relates to computer implemented methods that optimise a trait introgression process. Background
[0002] The rapid change in climate, demographics and food demand patterns are putting plant breeders under immense pressure to develop sustainable crop varieties that can show resilience while facing natural and economic uncertainties. To develop sustainable plant varieties, breeders genetically modify the existing variety known as Elite Variety (EV) by introducing missing but desired traits from other varieties. The result is the Desired Variety (DV), a plant variety containing the desired traits while retaining the traits of the elite variety.
[0003] However, a multitude of variables affect the trait introgression process, which makes breeding very difficult. Further, a computer implementation of the process design has been hampered by the sheer number of combinations and variables that quickly lead to prohibitive requirements on computational resources. Summary
[0004] A method for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism comprises: accessing a cost model the cost model including breeding process variables, the cost model representing a cost of the breeding process over multiple generations of the organism based on the breeding process variables; optimising the cost model by calculating breeding process variable values that represent a reduced cost of the breeding process,wherein optimising the cost model comprises calculating breeding process variable values for breeding process variables for multiple generations of the organism; wherein the breeding process variables comprise decision variables representing a selected method for screening individuals for the one or more desired traits at each of the multiple generations of the organism.
[0005] It is an advantage that the method includes calculating the breeding process variable values for multiple generations, while, at the same time, the variables comprise decision variables for the screening mothed at each generation. This way, the method can optimise the screening methods for all generations in a single cost model. This results in an improved breeding process and makes the method computationally feasible.
[0006] In some embodiments, the cost model comprises a lower bound on a number of individuals to screen for each generation.
[0007] In some embodiments, the lower bound on the number of individuals to screen for each generation is based on the selected method for screening individuals at that generation.
[0008] In some embodiments, the number of individuals to screen for each generation is based on a number of retained individuals in that generation and a confidence in individuals of that generation including the one or more desired traits.
[0009] In some embodiments, in case the selected method for screening individuals is phenotype screening, the lower bound is a function of: a number of retained individuals, and a confidence in a final generation of the breeding process including the desired trait.
[0010] In some embodiments, in case the selected method for screening individuals is marker-assisted screening, the lower bound is based on a transition probability betweengenotype groups, wherein at least some of the genotype groups include multiple genotypes.
[0011] In some embodiments, the genotype groups are defined based on whether the one or more traits are present in homozygous or heterozygous form.
[0012] In some embodiments, the method comprises introgressing multiple traits and the genotype groups are defined based on the number of traits present in an individual.
[0013] In some embodiments, the genotype groups cover all possible genotypes in relation to the one or more multiple traits.
[0014] In some embodiments, the method comprises determining the lower bound that guarantees at least one individual of the retained individuals contains the multiple desired traits at a one of one or more confidence values.
[0015] In some embodiments, the transition probability is based on a distance between genetic markers associated with the one or more desired traits.
[0016] In some embodiments, the cost model comprises: an integer variable indicative of a number of individuals created at each generation; and a binary decision variable indicative of whether the one or more traits are present in retained individuals at a confidence value.
[0017] In some embodiments, optimising the cost model is constrained by the integer variable being greater than a sum of the binary decision variable weighted by a lower bound on the integer variable for multiple values of retained individuals.
[0018] In some embodiments, optimising the cost model comprises minimising expenditure.
[0019] In some embodiments, optimising the cost model comprises minimising a number of generations of the breeding process.
[0020] In some embodiments, optimising the cost model comprises maximising a confidence in a final generation of the breeding process including the desired trait.
[0021] In some embodiments, the cost model is a multi-objective model and optimising the cost model comprises simultaneously: minimising expenditure, minimising a number of generations of the breeding process, and maximising a confidence in a final generation of the breeding process including the desired trait.
[0022] In some embodiments, the cost model comprises a mixed integer programming problem comprising integer variables that indicate the number of individuals and binary variables that are the decision variables.
[0023] In some embodiments, optimising the cost function is constrained by a maximum cost value.
[0024] In some embodiments, the cost model represents multiple scenarios and each of the breeding process variables is associated with one of the multiple scenarios.
[0025] In some embodiments, the cost model comprises constraints to ensure that decisions made at any stage in the breeding process are applicable and valid across all possible scenarios considered in the cost model.
[0026] A method for introgressing a desired trait into an elite variant of an organism comprises: creating a model for a progeny genotype as one of multiple genotype classes, wherein each of the multiple genotype classes is indicative of a number of desired traits in each individual;calculating transition probabilities between genotype classes when crossing the progeny with the elite variant; and using the model to calculate breeding process variable values.
[0027] A computer-system comprises one or more processors configured to perform the above method for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism. Brief Description of Drawings
[0028] Figure 1 illustrates a breeding process with phenotype screening.
[0029] Figure 2 illustrates the breeding process with marker-assisted screening.
[0030] Figure 3 illustrates a method for introgressing one or more desired traits into an organism using a breeding process.
[0031] Figure 4 illustrates a method for introgressing a desired trait into an elite variant of an organism.
[0032] Figure 5 illustrates the position of flanking markers and trait marker on a chromosome of length L and their distance from each other.
[0033] Figure 6: Each sub-figure shows the optimal breeding profile suggested by the MILP model corresponding to a given MACL in the range[50,60,70,80,90,95], to be achieved by the second generation subject to the budget (^^-axis). The budget variesover the range [90^^, ⋯ ,150^^]. The blue bars show the confidence profiles produced bythe optimiser when using phenotyping as screening technique. The lighter the shade of blue, the earlier the generation. The number inside each bar segment indicates the confidence (as a percentage) in achieving the desired plant in respective generations.
[0034] Figure 7: Each sub-figure shows the optimal breeding profile suggested by the MILP model corresponding to a given MACL in the range[50,60,70,80,90,95], to beachieved subject to the budget (^^-axis). The budget varies over the range[90^^, ⋯ ,150^^].
[0035] Figure 8: Each sub-figure shows the optimal breeding profile suggested by the MILP model corresponding to a given MACL in the range [50,60,70,80,90,95] to be achieved by the second generation subject to the budget (^^-axis). The budget variesover the range [90^^, ⋯ ,150^^]. The yellow and blue bars show confidence profilesproduced by optimiser using screening technique MAS and phenotyping, respectively.
[0036] Figure 9: Each sub-figure shows the optimal breeding profile suggested by the MILP model corresponding to a given MACL in the range [50,60,70,80,90,95] for a 3-trait introgression case. The MILP runs respects the condition to achieve MACL by the second generation subject to the budget (^^-axis). The budget varies over the range[1.7^^, 1.8^^ ⋯ ,3.2^^]. The yellow and blue bars show confidence profiles produced byoptimiser using screening technique MAS and phenotyping, respectively.
[0037] Figure 10 illustrates a computer system that is configured to perform the methods disclosed herein.
[0038] Figure 11 illustrates another example of a computer system 1100 for implementing the disclose methods for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism. Description of Embodiments Genetics of breeding
[0039] This disclosure relates to trait introgression which is related to traits, such as resistance, yield, etc., into organisms, such as plants including wheat, for example. Traits are typically encoded by genes or variants of genes at specific loci in the genome, such as loci on a chromosome. An individual can carry one or multiple copies of the variant (i.e. allele). Diploid organisms, such as most animals and humans, havetwo copies of each chromosome and therefore carry zero, one or two copies of the allele. Individuals with two identical copies are referred to as homozygous and those with two different copies are referred to as heterozygous. For polyploid organisms, like durum wheat for example, homozygous individuals have multiple identical copies whereas heterozygous individuals have different copies.
[0040] When reference is made to a trait herein, this refers to a trait that is expressed in one or more genes or loci. For simplicity, the term trait is used interchangeably with the sequence at that loci. Typically, an elite variant is used that does not have the trait and the aim is to introgress the trait from a donor parent (also referred to as donor variant). However, the donor parent also has many undesirable traits that should be removed by back crossing with the elite variant. The trait can be encoded by a single nucleotide polymorphisms (SNP), structural variant, or other variant or difference compared to the elite variant.
[0041] When the elite variant is crossed with the donor parent, the offspring may inherit zero, one or more copies of the trait from the donor parent. If the donor parent is homozygous in the trait, the offspring will have at least one copy but often the donor parent is heterozygous, which means some offspring may not have any copy of the trait. Further, there can be chromosomal crossover, which leads to a further source of variations for the inheritance of alleles. The aim is to find the offspring that has inherited the trait and select only those individuals for further breeding.
[0042] To find the offspring with the trait, it is possible to observe the offspring visually. For example, the offspring may be subjected to a herbicide (such as glyphosate) and it can be observed whether the offspring is resistant. Or the height or yield of the offspring can be observed to select individuals with specific characteristics from the donor parent.
[0043] Another way is to genotype the offspring. The aim is to find the offspring with the desired genotype (at least one copy of the trait). There are different methods for genotyping, which have relative advantages and disadvantages. For example, fullgenotype sequencing would provide the largest amount of information but the costs and timing are prohibitive for most applications. Polymerase chain reaction testing for the traits directly could also be a possibility but is also impractical in many cases. A preferable technique is marker-based genotyping. This method does not sequence or detect the actual genetic information that encodes the trait but detects a marker that is strongly linked with the trait. The marker is a region in the genome that is more readily detectable than the actual trait itself. It is desirable that the marker is closely located on the genome to the trait so that it is unlikely that crossovers separate the marker from the trait loci. In some examples, there are two markers, referred to as flanking markers on either side of the loci of the trait. This provides for additional reliability during genotyping.
[0044] Other examples of genotyping, such as assay testing, gel based methods, and others are equally possible. Even further, genotyping or phenotyping can be performed using RNA expression analysis or protein sequencing.
[0045] The disclosed method is particularly useful in cases where multiple traits are to be introgressed. Each of those traits has its own loci on the genome associated with respective markers. So the genotype can be defined with reference to those traits. For example, for three traits on three respective loci. If those loci encode the traits, the individuals are labelled as ABC. Else, if the same loci do not encode the traits, the individuals are labelled as abc. More specifically, for a diploid organism with two copies, the individuals may be aa for the first loci, which means they have no copy of the trait. That is, the elite variant has the genotype aabbcc and the donor parent has the genotype AaBbCc assuming the donor parent is heterozygous. As a result, the offspring, inheriting any of the two copies at each location, can have any of the mutations, such as aabbcc, Aabbcc, aaBbcc, aabbCc, AaBbcc, AabbCc, and so on. Other example organisms, such as durum wheat may be tetraploid, meaning the combinations would be significantly more including aaaabbbbcccc, aaaAbbbbcccc, and so on.
[0046] As can be seen, the number of combinations quickly increases which makes even enumeration cumbersome. Even more so, the selection based on this large number of genotypes is extremely difficult and finding the optimal screening strategy for each generation is near impossible.
[0047] It is further noted that the first-generation offspring will have a large number of undesirable traits from the donor parent. Therefore, selected individuals of the first- generation offspring are crossed with the elite variant again. This is repeated to gradually reduce the amount of genetic information remaining. At each generation, a number of individuals is generated, and a number of individuals is selected for further breeding. Both numbers are difficult to optimise.
[0048] More particularly, at each generation, there is a question of which genotypes should be selected in order to achieve trait introgression and how many individuals are generated and how many selected for each genotype. A breeder aims to get the desired plants with high confidence. This is referred to as the Trait Introgression Process (TIP). Phenotype vs genotype screening
[0049] Figure 1 illustrates a breeding process with phenotype screening. An elite variant 101 is crossed with a donor parent 102 resulting in a generated F1 individuals 103. These are back-crossed with the elite variant 104 resulting in a first back-crossed generation 105. In this example, the trait to be bred into the elite variant is not visible but the similarity of the plants is shown by their height, noting that donor parent plant 102 is higher than elite variant plant 101. It is noted that the elite variant 101 has a number of advantageous traits over donor parent102, which should be preserved (such as yield, resistance, for example).
[0050] Figure 1 shows twelve individuals of the BC1 generation 105, which have different height. Those plants that have a height that is most similar to elite variant 101 are selected. This is not always a single trait but could be an overall impression or acombination of traits that make the selected plants most similar to the elite variant 101. The selected plants are indicated by a dashed circle in Figure 1.
[0051] Figure 2 illustrates the breeding process with marker-assisted screening. Again, elite variant 201 is crossed with donor parent 202 and the offspring 203 is back- crossed with elite variant 204. Twelve plats of the back-cross generation 205 are genotyped, such as by gel electrophoresis. Figure 2 shows the genotype of the twelve plants where black regions indicate a gene sequence that originates from donor parent 202 and white regions indicate a gene sequence that originates from the elite variant 201. The marker-assisted selection now selects the individual with the least amount of genetic material from the donor parent 202 in order to preserve the elite variant 201 as much as possible. It is noted that Figures 1 and 2 are simplifications for illustrative purposes and to introduce some of the concepts used in this disclosure. They do not accurately reflect the methods disclosed below relating to groups of genotypes etc.
[0052] Several decisions impact the performance of the TIP in terms of the total cost, time, required resources, and confidence in getting the desired plant at the end of a plant breeding program. These include the population size at each generation, the plant screening method (e.g., marker-assisted vs. phenotyping vs. hybrid), which traits to introgress, and how many traits to introgress. Minimising the time and cost of the introgression process while ensuring at least one plant with the desired trait is produced with acceptable confidence under limited resources is a complex decision problem. These decisions may be based on the breeder’s experience, combined with spreadsheets that use gross simplifications. Optimisation methodology
[0053] This disclosure provides an exact methodology that can improve the performance of the Trait Introgression Process (TIP) by reducing cost and time under a limited budget and resources. While some examples herein relate to plants, such as wheat, other examples relate to other plants, such as fruit, nuts, vegetables or legumes, or other organisms, such as animals including cattle or other livestock.
[0054] The proposed method can assist the breeders in selecting the correct number of plants to screen using the correct screening technique at each generation to maximise confidence in achieving the plants with desired traits while minimising the cost and time. For phenotyping, the model uses Mendel’s law to decide the number of plants to screen as a function of retained plants and the desired percent confidence, whereas for marker assisted selection (MAS), the number of plants to screen is a function of the recombination probabilities defined by the given marker data and transition probability from one genotype to another.
[0055] Instead of optimising the TIP for one generation in isolation, the current disclosure provides a solution that optimises the TIP for all generations in a single model. As a result, the calculated TIP variable values are far superior than alternatives that rely on isolated generations optimisation.
[0056] Several experiments were conducted to understand the impact of the limited budget, acceptable confidence, and different screening techniques at different generations on TIP performance. The results show that restriction of achieving acceptance confidence in early generations requires more budget and may result in failure of TIP, i.e., that obtaining a desired plant with acceptable confidence under the given budget is not feasible. However, relaxing this restriction ensures a feasible solution for the considered instances of TIP.
[0057] The results also demonstrate that, though MAS is expensive, it helps get a feasible solution for TIP even under the restriction of obtaining a desired plant with acceptable confidence in an early generation. The results also demonstrate that MAS is preferred over phenotyping under the given budget if doing so improves overall objective value and gives more confidence in each generation compared to cost.
[0058] The table below summarises some of the factors and decisions that may impact the success of the TIP:Method for introgressing traits
[0059] Figure 3 illustrates a method 300 for introgressing one or more desired traits into an organism using a breeding process. The method operates over multiple generations of the organism, which means that the multiple operations are being considered as one optimisation problem. This is in contrast to other methods that consider each generation individually.
[0060] The method 300 commences by a computer system accessing a cost model. The cost model may be stored on computer memory, such as on volatile or non-volatile memory. The cost model may also be accessed over a data communication link, such as via the internet or other computer network. The cost model may be expressed as equations or as a compiled or uncompiled encoding, such as source code or definition in a solver-specific syntax or file type.
[0061] The cost model includes breeding process variables, such as process variables discussed herein. This may include any combination of one or more of number of individuals created at each generation, number of individuals retained at each generation and method of screening.
[0062] The cost model represents a cost of the breeding process over the multiple generations of the organism based on the breeding process variables. This means, thecost model, when evaluated, calculates the cost of the breeding process. The breeding process variables are variables of the cost model and therefore influence the calculated cost. When the values of the breeding process variables are changed, the cost of the breeding process changes. As a result, the computer system optimises the cost model by calculating breeding process variable values that represent a reduced cost of the breeding process. The cost model may also comprise constraints which need to be met for any valid solution of the cost model. In other words, a set of variable values with a reduced cost that would lead to a violation of one of the constraints, would not be considered a possible solution.
[0063] In summary, method 300 provides the following advantages: 1. It finds the right screening technique for every generation rather than assuming a single screening technique for the entire TIP. For the MILP model, the cost of using resources can vary across generations based on whether, in a back-cross generation, the selection is made using mass or phenotyping. 2. It includes the decision on the maximum size of the retained and screened populations at each generation in MILP, rather than by using simulations. 3. It can optimise against more than one given type of doner parent in one go, whereas other methods can obtain only one type of donor parent at a time. 4. The disclosed method has the flexibility to budget for each generation and for the total TIP.
[0064] The computer system may perform an optimisation algorithm, which may be an algorithm specialised in integer linear programming, or mixed integer linear programming (MILP). Other optimisation approaches are equally possible. Options include iterative methods, such as Newton’s method, quadratic programming, gradient descent, simplex algorithms and others.
[0065] In effect, optimising the cost model comprises calculating breeding process variable values for breeding process variables for multiple generations of the organism so that the calculated values result in an optimal result for the cost as calculated by evaluating the cost model while satisfying all of the constraints in the cost model. In one example, the proposed MILP (optimisation model) finds the global optimal solution. In some cases, however, it terminates early are if the user applies one or more termination criteria for the early termination of optimisation process. For example, if the user defined time limit reached or the gap between the current solution in hand and the best expected solution (referred to as “MIP gap”) is less than the user defined threshold.
[0066] In some cases, “optimising” does not necessarily mean that a theoretically optimal global minimum is reached. Instead, optimising may mean that the values are iteratively varied to reduce the total cost until the process is terminated upon meeting a termination condition. That condition can be a maximum number of iterations, minimum change between iterations, etc.
[0067] It is noted that the breeding process variables comprise decision variables. These variables are referred as Boolean if they are binary in nature, continuous if they can take any real value over a defined domain, integer if they can take any integer value over the defined domain. There are number of decision variables of different types such as screening technique selection decision is of Binary type whereas the number of individuals to retain and screen at each generation are Integer type.. The decision variables may be discrete, which is why an integer linear problem is an appropriate choice to encode the problem to be solved.
[0068] As set out above, method 300 may comprise the optimisation of a MILP cost function with corresponding constraints. It is typically challenging to formulate such a MILP model because it is often difficult to find a suitable way to represent the multitude of inputs as appropriate variables in the MILP problem. For example, the breeding process is characterised by a number of breeding process variables. There is a number of different ways of encoding those variables in the MILP problem. Further,some formulations of the cost model lead to a prohibitive computational expense, such as processing time or memory requirements, while others require fewer processing steps, less memory or are parallelisable due to the structure of the formulated problem. Lower bound on the number of individuals
[0069] One example variable is the number of individuals to screen in each generation. This variable can be included into the cost model since each individual contributes to the cost of screening – more individuals to screen means higher cost and vice versa. However, it has been found that including this variable into the cost function leads the optimiser to reduce the number of individuals to a minimum, which is not desired for the overall cost optimisation because reducing the number of individuals at each step also reduces the likelihood that at least one individual of the final generation contains the desired traits.
[0070] To address this issue, the cost model comprises a lower bound on the number of individuals to screen for each generation. This refers to the number of individuals that are created in each generation and that are then processed through screening, such as phenotype or marker assisted screening. By setting a lower bound, the optimisation method does not undercut that bound which means there will be a sufficient number of individuals in each generation. This constraint significantly assists in reaching a solution faster, where the solution is feasible in the sense that there is at least one individual in the last generation that has the desire trait with at least a given likelihood. In other words, the solution space is limited for the optimiser, which means the optimiser finds the optimal solution in the remaining space faster.
[0071] The lower bound on the number of individuals to screen for each generation is based on the selected method for screening individuals at that generation. That is, the lower bound depends on the screening method. Further, the number of individuals to screen for each generation is based on a number of retained individuals in that generation and a confidence in individuals of that generation including the one or more desired traits.
[0072] As mentioned above, the lower bound depends on the screening method. In case the selected method for screening individuals is phenotype screening, the lower bound is a function of a number of retained individuals, and a confidence in a final generation of the breeding process including the desired trait.
[0073] On the other hand, in case the selected method for screening individuals is marker-assisted screening, the lower bound is based on a transition probability between genotype groups, wherein at least some of the genotype groups include multiple genotypes. The genotype groups are described in more detail further below, but briefly, they are defined based on whether the one or more traits are present in homozygous or heterozygous form. The genotype groups become particularly useful when introgressing multiple traits and the genotype groups are defined based on the number of traits present in an individual. The number of potential genotypes increases quickly with the number of traits to introgress when all combinations of homozygous and heterozygous traits are considered. Therefore, the definition of genotype groups reduces this burden. This is particularly important in the current method because transitions between genotypes are used. If there is a large number of genotypes, there is an even larger number of possible transitions between them. On the other hand, if the genotypes are grouped into groups, then only the transitions between the groups need to be considered, which reduces the overall computational complexity.
[0074] In that sense, the genotype groups cover all possible genotypes in relation to the one or more multiple traits so that backcrossing of any genotype will always be in one of the genotype groups (which may be a different group to its parents).
[0075] Returning to the concept of the lower bound, it has been found that good results are achievable by determining the lower bound that guarantees at least one individual of the retained individuals contains the multiple desired traits at a one of one or more confidence values. This means that there may be a single confidence value, which is a constant in the cost model or, in other examples, there may be multiple different confidence values and one of them can be selected by the optimiser.Transition probabilities
[0076] As alluded to before, the method uses transition probabilities between genotype groups. Those transition probabilities is based on a distance between genetic markers associated with the one or more desired traits. This reflects the fact that loci that are closely located on the genome are likely inherited together whereas more distant loci are more likely to be separated and only one of them being further inherited. This is particularly the case due to crossovers as a source of genetic mutation. Integer variables and constraints
[0077] Continuing with the formulation of the cost model, it is noted that the cost model comprises, for example, an integer variable indicative of the number of individuals created at each generation. This means for each generation, there is a separate integer variable in the cost model (i.e. the cost function or the constraints or both). In the cost model, there is also a binary decision variable indicative of whether the one or more traits are present in retained individuals at a confidence value.
[0078] Optimising the cost model is then constrained by the integer variable being greater than a sum of the binary decision variable weighted by a lower bound on the integer variable for multiple values of retained individuals. In other words, the binary decision variable indicates yes / no on whether the traits are present at the confidence value. This variable, as a 1 or 0, is then multiplied with the lower bound on the number of retained individuals to obtain a cost contribution of that variable to the overall cost. Optimisation
[0079] While reference is made to ‘cost’ above, this could refer to a number of different aspects, such as expenditure (i.e. monetary cost) or time. In one example, optimising the cost model comprises minimising expenditure, which mean the solution is optimised to achieve the minimum monetary cost. This can include the cost fordifferent screening methods times the number of screened individuals, etc. Other costs can also be included, such as environmental cost, area cost, labour cost, etc.
[0080] In a further example, optimising the cost model comprises minimising a number of generations of the breeding process. This reduces the time that it takes to introgress the desired traits into the individuals and therefore may also reduce some of the monetary cost factors.
[0081] In yet a further example, optimising the cost model comprises maximising a confidence in a final generation of the breeding process including the desired trait. This means that it is less likely that no individual holds the desired traits. There may still be a non-zero risk that this occurs but the disclosed method can reduce that risk.
[0082] In light of the above examples, it is desirable to combine the multiple objective into a multi-objective model with any possible combination of objections. This could mean, for example, that optimising the cost model comprises simultaneously minimising expenditure, minimising a number of generations of the breeding process, and maximising a confidence in a final generation of the breeding process including the desired trait. It is a particular strength of the disclosed method that these three objectives can be optimised, which is not possible with other methods. This means, with other methods, it is necessary to guess certain variable values or keep them constant. At the same time, there is no guarantee that those variable values were either conservative enough to achieve the goal or too conservative so thot resources are wasted. With the proposed methods, the multiple objectives can be optimised at the same time while guaranteeing that the formulated constraints are satisfied.
[0083] Providing further detail on the cost model formulation, in one example, the cost model comprises a mixed integer programming problem comprising integer variables that indicate the number of individuals and binary variables that are the decision variables. In this formulation, high quality solutions have been found that include breeding process variables that lead to the desired outcome at a givenconfidence where confidence is not optimised. In other cases, the method may also maximise confidence.
[0084] A further improvement of the optimisation may be in constraining the cost function by a maximum cost value. As a result, a solution is only provided if it is within the cost budget and otherwise, the method informs the user that no solution can be found for the given budget. This is useful in practice where breeders have a limited budget and want to know whether they can achieve their goals within that budget. Method for introgressing traits using genotype classes
[0085] Figure 4 illustrates a method 400 for introgressing a desired trait into an elite variant of an organism. The method 400 commences by creating 401 a model for a progeny genotype as one of multiple genotype classes as described above. Further details on example mathematical definitions are provided below. Each of the multiple genotype classes is indicative of a number of desired traits in each individual, so the individuals are grouped into genotype classes depending on how many or whether they include any of the desired traits. The method then calculates 402 transition probabilities between genotype classes when crossing the progeny with the elite variant. Details are also provided below, which include details on how to calculate those transition probabilities. Finally, the method 400 uses 403 the model together with the transition probabilities to calculate breeding process variable values as described above. It is noted, however, that other formulations of the cost model can equally be used with the transition probabilities. For example, the transition probabilities can be used as weights in weighted sums of the cost function or can be used as constraints to measure the likelihood of traits being inherited.
[0086] As illustrated in the above table, a number of parameters play a role in optimising the introgression process with the objective of minimising cost and time. The duration of TIP is represented by the number of generations required to introgress all the desired traits into the elite population. In an ideal situation, at the end of the TIP, there should be at least one plant that contains the desired traits with the pre-definedconfidence level. However, with limited resources and budget, sometimes it is not feasible to obtain even a single plant with the pre-defined confidence level. Hence, under limited budget and resources, there are three conflicting objectives: minimising the cost, minimising the number of generations, and maximising the confidence level in getting at least one plant with the desired traits.
[0087] This disclosure provides a cost model a (MILP or other) to introgress the desired traits into the elite population with the objective of minimising the time and cost while maximising the confidence limit. With the limited resources and budget availability, the model respects the constraints proposed by the breeders, for example, the availability of a particular resource in a generation. The details of all the model constraints are explained in detail below. A number of input parameters and sets are required to build the model. The full notation can be found below.
[0088] The cost model minimises the total cost and time while maximising the confidence in getting desired plants at the end of the process. The model supports breeders in taking the decisions of how many plants to take forward (retained plants), how many to screen, and what should be the confidence in getting the dv at each generation. The model ensures minimum acceptable confidence defined by the breeder in getting the desired individual at the end of TIP.
[0089] The number of plants to be screened and retained at each generation heavily depends on the screening strategy used at that generation. Instead of selecting the plants at an individual level, the method uses selection strategies at the group level. That is, at each generation, the model selects the type of plants to keep for the following generation, as to minimise the cost and number of generations while maximising the confidence achieved at the final generation. The details of classification into the exhaustive and mutually exclusive types are given below. Once the genotypes are categorised into different types, the method calculates the transition probabilities of conversion from one genotype to another by using our general version of the probability transition matrix.
[0090] The proposed model chooses one screening strategy from the set of available strategies for each generation and, based on this choice, will find the lower bound on the number of plants to screen for a generation, corresponding to the decisions number of retained plants and confidence in getting the DV. Specifically, for the phenotype screening technique, Mendel’s law is used for calculating the lower bound on the number of plants to screen for a generation as a function of the number of retained plants and confidence in getting the DV for that generation. If, on the other hand, the model uses MAS as the screening strategy, then the lower bounds are calculated using the probability of genotype in a given generation and the transition probability, which relies on the locations of the markers and their distances from the desired traits. Furthermore, the proposed model has the flexibility of introducing a limit on the budget available for breeding specific generations, and on the total budget for the whole introgression process, whereas previous work did not consider these constraints. Nomenclature
[0091] The following nomenclature is used in this disclosure to describe examples but without limiting the scope of this disclosure to this particular nomenclature.ℛ^^ the range for number of retained population denoted by {^^^^1, ^^^^1 +1, ⋯ , ^^^^^^} where ^^^^1, ^^^^^^ will corresponds to the lower and upper boundsrespectively on the number of retained population as advised by the breeder ℋ the set denoting the different genotypes. Each genotype is different in the number of traits in homozygous state out of all traits under consideration ^^ the set of possible acceptable confidence level in getting at least one plant containing desired traits at the end of tip. The process is considered as success if it stops with at least one plant containing desired traits with confidence in this set. One of the objective would be to maximise the achieved confidence levelℐ the index set of traits {1,2, ⋯ , ^^} to introgress^^ the index set of all markers^^ the set of generations used in introgression process to achieve desire quantity.^^ = {0,1,2, ⋯ , ^^^^^^^^}here ^^^^^^^^ the given upper bound on the number of generations.In theory, it can be 7-8 ^^ the set of available screening strategies as dictated by the decision maker the set of sequences of length ^^ ∈ {1 ⋯ |^^|} where each sequence denotes thepossible location of the selected ^^ markers in heterozygous state ℒ list of all possible genotypes where each genotype is represented by a binarysequence (^^^^^^1, ^^^^1, ^^^^^^1 = ^^^^^^2, ^^^^2, ^^^^^^2 = ^^^^^^3, ⋯ , ^^^^^^^^−2 =corresponds to the trait ^^ ∈ ℐ, ^^^^^^^^ , ^^^^^^^^ denote its left and right flanking marker, and^^^^^^denotes the marker of the trait^^^^^^ℎ number of individuals selected for screening at back-cross generation ^^ ∈ ^^when the strategy ^^ ∈ ^^ is used for genotype ℎ ∈ ℋ^^^^^^ℎ^^^^ Boolean variable taking value one only if, after screening, ^^ ∈ ℛ^^ individualsare retained from generation ^^ ∈ ^^ for next generation under strategy ^^ ∈ ^^ forgenotype ℎ ∈ ℋ to achieve confidence level ^^ ∈ ^^^^^^^^ℎ indicator variable that takes the value of one if screening technique ^^ ∈ ^^ isused in generation ^^ ∈ ^^ for genotype ℎ ∈ ℋ^^^^^^^^^^^^ℎ the cost of breeding all the plants of type ℎ ∈ ℋ that use screeningtechnique ^^ ∈ ^^ in generation ^^ ∈ ^^^^^^^^^^e maximum number of population that can be obtained in each generation ^^^^^^^^e minimum number of population that should be obtained in each generation ^^^^^^^^^ℎ^^^^^^^wer bound on the value of ^^^^^^ℎon the minimum number of plants to screen toensure that at least one plant out of ^^ ∈ {1, ⋯ , |ℛ^^|} individuals in generation ^^ ∈ ^^using strategy ^^ ∈ ^^ contains genotype ℎ ∈ ℋ with confidence ^^ ∈ ^^^^ total budget ^^^^^^^^^^^^^^^^ℎ aximum amount of resources available for strategy ^^ ∈ ^^ in generation ^^ ∈ ^^and genotype ℎ ∈ ℋ]^^, ^^ markerℎ genotype^^ generation ^^ trait ^^ number of individuals ^^ confidence level ^^ left ^^ right^^ℎ0 probability of that each plant in the initial population contains genotype ℎ ∈ ℋ^^ℎ^^minimum probability with which the selected individuals contain a particulargenotype ℎ ∈ ℋ at the end of generation ^^ ∈ ^^^^^^^^^^ right flank marker of trait ^^ ∈ ℐ^^^^^^^^ left flank marker of trait ^^ ∈ ℐ^^^^^^ℎ^^^^target trait where ^^ ∈ ℐ^^^^indicator variable which takes the value 1 if the marker at the ^^^^ℎlocation of the sequence is in a heterozygous state, 0 if homozygous ^^ total number of flanking markers^^ list of all genome sequence^^^^1, ^^^^^^1 = ^^^^^^2, ^^^^2, ^^^^^^2 =^^^^^^^^ , ^^^^^^ , ^^^^^^^^ ^^^^{0,1}, for each trait ^^ ∈ ℐ^^^^^^ℎ probability of occurrence of genotype ℎ ∈ ℋ at back-crossed generation ^^ ∈ ^^using screening technique ^^ ∈ ^^ℎ^^^^^^tal number of types ^^^^′ℎℎ transition probability from type ℎ to type ℎ′, ℎ, ℎ′ ∈ ℋℓ ℓ ∈ ^^ represents a genome sequence^^^^^^ probability of crossover on the left of marker ^^ ∈ ^^^^^^^^ probability of crossover on the right of marker ^^ ∈ ^^Acronyms MAS Marker assisted selectiong GM Genetically modified DV Desired VarietyEV Elite Variety TIP Trait Introgression Process MAS Marker-assisted Selection T-BCM Back Cross-breeding Method RP Recurrent Parent DP Donor Parent MABC Marker-assisted Back Crossing MASS Marker-assisted Screening strategy BC back-crosses MILP Mixed-integer Linear Program RTIP Restricted Trait Introgression Process RTIP-BTRC Restricted Trait Introgression Process Under Limited Budget, Time or Resource Constraints MACL Minimum Acceptable Confidence Glossary Recurrent Parent In recurrent back-crossing, the parent that is crossed with the first and the subsequent generations. Also known as back-cross parent Selfing is the most extreme form of inbreeding. One parent donates both gametes to the offspring Heterozygous refers to having inherited different versions (alleles) of a genomic marker from each biological parent. Thus, an individual who is heterozygous for a genomic marker has two different versions of that marker Homozygous an individual who is homozygous for a marker has identical versions of that marker, by contrast to heterozygous Assumptions
[0092] Some or all of the following assumptions may apply:1. The upper and lower bounds on the population size to be retained at each stage of breeding cycle are given. 2. The set of available techniques for screening the plants (e.g., MAS, phenotyping, hybrid) is given. 3. The duration and cost matrix for each task (tasks can be planting, harvesting, watering, screening quality etc.) are given. The matrix will be a function of the screened technique and the resource used. In the absence of such matrix the cost is taken as 1 per plant for all resources and all screening techniques. 4. The available resources to do all tasks and the upper bound on their usage is given. In the absence of such information, the model will assume that all resources are available and without any bound on their usage. 5. The initial population of RP and DP is given. 6. Each plant in the initial population contains genotype ℎ with given probability of ^^ℎ0. 7. The minimum probability ^^ℎ^^with which the selected individuals contains a particular genotype ℎ at the back-cross generation ^^ is given. 8. The distance between the flanking markers (both left and right) and donor markers are given, if MAS is available as the screening technique. 9. The number of generations to perform foreground selection is given. Currently, this is assumed to be two and remaining generations work with background selection. 10. It may be assumed that right flanking marker ^^^^^^^^of a trait ^^ is the same as theleft flanking marker of next trait ^^ + 1, i.e., ^^^^^^^^ =for all ^^ in ℐ, the set of alltraits to introgress.
[0093] It is noted that some of these assumptions may be removed by including corresponding variables into the model or by other means. Probability matrix and type conversions
[0094] Given the marker data for desired traits to introgress and the corresponding flanking markers, the genotype of a plant can be defined as the binary sequence representing the state (homozygous or heterozygous) of each marker arranged from leftto right. Mathematically, for any trait ^^ ∈ ℐ, if ^^^^^^^^ , ^^^^^^^^ denote its left and rightflanking marker, and ^^^^^^denotes the marker of the trait, then the genotype is given bythe binary sequence^^^^^^^^−1, ^^^^^^−1, ^^^^^^^^−1 = ^^^^^^^^, ^^^^^^, ^^^^^^^^). Any termof this sequence takesvalue one or zero based on the rule ^^^^= {1 if the marker at ith location of the sequence is in heterozygous state,0 otherwise, i. e. , homozygous for the recurrent parent allele
[0095] For example, in the case of a single desired trait, all possible genotypes can berepresented by the binary tuple^^^^1, ^^^^^^1) where ^^^^^^1, ^^^^1, ^^^^^^1 ∈ {0,1}guided by Rule (3.1). The TIP aims to incorporate all desired traits in the elite plant without touching its other traits. This means that, at the end of any back-crossing step, the most desired genotype, referred to as a success, is the one that contains the desired trait in the heterozygous state and all flanking markers at the homozygous state.
[0096] By contrast, the least desirable genotype, referred to as a failure, is the one that has all desired traits in the homozygous state, as this would mean that the process fails to introgress any of the desired traits. For example, in the single trait case, success would mean getting a single genotype {(0,1,0)}, whereas the set of failed genotypes is{(0,0,0), (1,0,0), (1,0,1), (0,0,1)}. The other possible genotypes for a single-traitcase have one or both flanking markers around the desired trait at homozygous statewhile the trait itself is at heterozygous state, i.e., the set {(1,1,0), (0,1,1), (1,1,1)}.These genotypes are neither considered a success, as they are not homozygous at both flanking markers, nor a failure, as they do contain the desired trait. 1. Type 1 (success): Genotypes of this type belong to the set where all desired traits are in the heterozygous state and all flanking markers are at the homozygous state, thatis, ^^^^^^ = 1 and ^^^^^^^^ , ^^^^^^^^ = 0 ∀^^ ∈ ℐ. More precisely, Type 1 genotypes belong tothe list L1 and the rest are a list of all possible genotypes whose elements come fromthe fact that ∀^^ ∈ ℐ, ^^^^^^^^ , ^^^^^^ , ^^^^^^^^ ∈ {0,1}. The index ^^ is right, as the length ofeach sequence in L1 is equal to the number of traits, ^^1 = [(^^^^^^1, ^^^^1, ^^^^^^1, ⋯ , ^^^^^^^^−1, ^^^^^^, ^^^^^^^^) ∈ ^^ ,2. Type 2: This is the set of genotypes where all desired traits and at least one flankingmarker are in the heterozygous state. That is, ^^^^^^ = 1 ∀^^ ∈ ℐ and there exists at leastone ^^ for which either one or both of ^^^^^^^^and ^^^^^^^^equal one. Type 2 genotypes are further categorised into sub-types based on the number of flanking markers in aheterozygous state and their locations, denoted by ^^ and the sequence ℓ ∈ ℒ^^,respectively. Note that ℒ^^denotes the set of sequences of length ^^, where each sequence denotes the possible location of the selected ^^ markers in heterozygous state.Thus for each ^^,where ^^ is the total number of flankingEach sub-type is a function of ^^, ℓ and is denoted by t2(^^, ℓ). The total number of suchsub-types is ∑^^∈ 1,2⋯,^^}Note that ^^ = |ℐ + 2 under the assumption ^^^^^^^^ =^^^^^^^^+1 for all ^^ ∈ ℐ. More precisely, t2(^^, ℓ) genotypes belong to the list^^(^^) = [ℓ = (^^^^^^1, ^^^^1, ^^^^^^1, ⋯ , ^^^^^^^^−1, ^^^^^^, ^^^^^^^^) ∈ ^^,3. Type 3 : The set of genotypes with at least one desired trait in the heterozygousstate and one in the homozygous state. That is, there exists ^^, ^^ ∈ ℐ with ^^ ≠ ^^ such that^^^^^^ = 1 and ^^^^^^ = 0. More precisely, Type 3 genotypes belong to the list4. Type 4 : The set of genotypes for which all desired traits are in the homozygousstate. That is, ^^^^^^ = 0 for all ^^ ∈ ℐ. More precisely, Type 4 genotypes belong to the list
[0097] The total number of types as defined above is^^() . These are ^^ exhaustive, i.e, all possible genotypes in any bc generation will belong to a type. Some plants can change their genotype, and hence their type, while going from one BC generation to another. However, not all type conversions are possible. For example, if marker ^^ is in heterozygous state in one type, and homozygous in another, then the probability of conversion from the former type to the latter is zero, as ^^ cannot change its state from homozygous to heterozygous in any bc generation. Therefore, the probability of conversion of Type 4 to another type is zero. The table below gives the complete conversion matrix from one type to another. Each cell of this table denotes whether if it is possible to convert a row type into a column type.
[0098] Using the above types, the probability, ^^^^^^ℎof occurrence of a genotype ℎ at bc generation ^^ using the MAS screening technique ^^ can be determined aswhere ℎ^^^^^^ = 3 (^^) notes the total number of types, and ^^^^ℎdenotes the ^^1ℎtransition probability from Typeto Type ℎ. Note that for ^^ = 1 (the F1 generation),all flanking markers and desired trait markers are in the heterozygous state. This meansthat all plants in the F1 generation belong to t2(^^, ℓ) where ℓ = {1, ⋯ , ^^} , and ^^denotes the total number of markers. Therefore, ^^1^^ℎ = 0 for all genotypes, exceptwhen ℎ belongs to t2(^^, ℓ).
[0099] From Equation (3.2), we can observe that the calculation ofdepends on the transition probabilities. The transition probabilities for the six-trait case are shown them in the below Table 3, which shows pairs of types of genotype in BC generation with transition probability for conversion. Each cell (row, column) of this table denotes the transition probability of a genotype from row type to the column type. ^^ in the table denotes the number of desired traits to introgress. The probability of crossover on the left and right of marker ^^ are denoted by ^^^^^^ , ^^^^^^respectively.
[0100] The cell (^^, ^^) of this table gives the transition probability of obtaining type incolumn ^^ from type in row ^^. The logic used to calculate the transition probabilities is as follows: 1. Transition from Type 1 to: (a) Type 1. The Type 1 genotype can produce a genotype of its own type only if there is no crossover between the first flanking marker and the last flanking marker in the bc generation. Otherwise, either one of the desired traits will convert to the homozygous state, or one of the flanking markers will convert to heterozygous state. Let ^^^^^^ , ^^^^^^denote the probabilities of crossover on the left and right of any marker, respectively. The transition probability of conversion from Type 1 to itself is the product of no-crossover on the left and right of any marker in between left-most andright-most flanking markers, and is given(b) Sub-Type 2. Genotypes of Type 1 cannot transform into any type where the genotype should have at least one flanking marker in heterozygous state. Therefore, the transition probability of converting from Type 1 to any of the Sub-Types 2 is zero. (c) Type 3. The transition probability of Type 1 to Type 3 can be calculated as 1 minus the transition probabilities to other types. (d) Type 4. Type 4 contains all the desired traits in homozygous state with no restriction on the state of the flanking markers. At each bc generation, either the desired trait remains in heterozygous state, or gets converted to the homozygous state. Thus, 1 ^^ the transition probability of any genotype of Type 1 to Type 4 is ( 2) , where ^^ is thenumber of traits considered. In our case, ^^ = 6.2. Transition from Sub-Type 2, t2(^^, ℓ) to:(a) Type 1. For genotypes of Type 1, all flanking markers are in homozygous state, whereas Sub-Type 2 contains genotypes with at least one flanking marker in theheterozygous state. Therefore, the transition probability from any Sub-Type 2 ^^2(^^, ℓ)to Type 1 is zero. (b) Sub-Type 2 ^^2(^^1, ℓ1). For ^^1 > ^^, the transition from ^^2(^^, ℓ) tois not possible because the number of heterozygous flanking markers inis higher. Similarly, the conversion from one Sub-Type 2 to another is notpossible when^ ^^ butThe transition from one Sub-Type 2 to another Sub-Type 2 is possible only when ^^1 ^For such cases, the transitionprobability from ^^2(^^, ℓ)consists of three terms, where calculation ofeach term ensures that the desired traits remain in heterozygous state: i. Term 1: Crossover on both sides of each flanking marker that are inheterozygous state, and should convert into homozygous state. Since⊆ ℓ, thereforewill contain all such markers. The desired probability is ∏ ^^∈ℓ\ℓ1 ^^^^^^^^ ^^ ^^. ii. Term 2: No crossover on any side of the flanking marker that is in heterozygous state, and should remain in heterozygous state. Note thatwill containall such markers, therefore the desired probability is ^^^^− ^^^^)(1 − ^^^^) . iii. Term 3: Either crossover on both sides or neither side of any of the flanking markers that is in homozygous state, and should remain in homozygous state. Note thatall such markers will belong to the list {1 ⋯ ^^}\ℓ, therefore the desired probability is(c) Type 3 and Type 4: Same logic as Type 1 to Type 3, and Type 1 to Type 4 conversion. 3. Transition from Type 3 to:(a) Type 1 and any Sub-Type 2. Genotypes of Type 3 cannot transform into any type where the genotype should have at least one more desired trait in heterozygous state. Therefore, the transition probability from Type 3 to Type 1 and any of the Sub- Types 2 is zero. (b) Type 3 and Type 4. Follows the same logic as Type 1 to Type 3, and Type 1 to Type 4 conversion. 4. Transition from Type 4. Type 4 contains all the desired traits in homozygous state. Therefore, the transition probability from Type 4 to any other type is zero. Type 4 can only be converted to itself in any bc generation. The MILP model
[0101] Let ^^^^^^ℎbe the number of individuals selected for screening at back-cross generation ^^ when the strategy ^^ is used for genotype ℎ; ^^^^^^ℎ^^^^a Boolean variable that takes the value one only if, after screening, ^^ individuals are retained from generation ^^ for the next generation under screening strategy ^^ for genotype ℎ to achieve confidence level ^^, and zero otherwise; and ^^^^^^ℎan indicator variable that takes the value of one if screening technique ^^ is used in generation ^^ for genotype ℎ.
[0102] The objective is to minimise the amount of resources used by minimising the cost, while achieving maximum confidence level in the minimal number of generations,where ^^^^^^^^^^^^ℎis the cost of breeding all the plants of type ℎ that use screening technique ^^ in generation ^^. Subject to:1. At most one screening technique is used. If generation ^^ is not the last generationof the TIP, we need to select the population that advances to next generation ^^ + 1 ofthe back-cross cycle. For selecting the plants, we need to screen them using one of the available screening strategies ^^ for each heterozygous trait ℎ. The corresponding mathematical constraint is:2. If the current generation is progressed, the plants should be screened. The process stops in generation ^^ if the success criterion is met at ^^. In that case, no more selection is needed for any generation after ^^. The following mathematical constraint is added in the model to cover this condition.3. The population is retained in generation ^^ only if a screening technique is chosen for ^^. If the population is retained in generation ^^, then a screening strategy must be activated for ^^.4. The minimum and maximum threshold of screened population at each generation should be respected,^^^^^^ℎ ^ ^^max ^^^^^^ℎ ∀^^ ∈ ^^, ∀^^ ∈ ^^, ∀ℎ ∈ ℋ , (3.8)where ^^^^^^^^d ^^^^^^^^e the minimum and maximum population thresholds, respectively. 5. Minimum number of required individuals to screen. For each combination of(^^, ^^, ℎ, ^^, ^^), we calculate the lower bound ^^^^^^^^^ℎ^^^^^^^the value of ^^^^^^ℎoutside theoptimisation model. Such a lower bound indicates the minimum number of plants to screen to ensure that at least one plant out of ^^^^^^ℎ^^^^retained plants in generation ^^ contains genotype ℎ with confidence ^^. To calculate these lower bounds, we use the fact that if ^^ denotes the probability of occurrence of a particular genotype in a plant, then the probability ^^ of obtaining at least one plant out of ^^ plants containing thatgenotype can be obtained by solving the equation ^^ = 1 − (1 − ^^)^^. Thus, for eachcombination (^^, ^^, ℎ, ^^, ^^), the lower bound ^^ =individuals to screen forensuring the existence of at least ^^^^^^ ^^ plants of genotype ℎ with probability ^^ =in generation ^^ using screening technique ^^ can be calculated by solving the inequalitywhere ^^ ∈ ^^ is the given confidence level. For phenotyping, the value of ^^^^^^ℎ iscalculated using Mendel’s law, whereas for MAS, the value comes from Equation (3.2), which depends on the transition probabilities defined in Subsection 3.2. The following constraint are in the model to respect this lower bound on the number of screen populations ^^^^^^ℎ:6. The budget-constraint should be respected.7. The capacity restriction in terms of resource utilisation should be respected, ^^^^^^ℎ ^ ^^^^^^max^^^^ℎ ∀^^ ∈ ^^, ∀^^ ∈ ^^, ∀ℎ ∈ ℋ , (3.12)where ^^^^^^^^^^^^^ℎ^^^the maximum amount of resources available for using strategy ^^ in generation ^^ and genotype ℎ.8. Confidence level attained between consecutive generations. The attained confidence should increase monotonically as generations pass.∀^^ ∈ ^^, ∀^^ ∈ ^^, ∀^^ ∈ ^^, ∀ℎ ∈ ℋ . (3.13)9. The minimum required confidence limit should be achieved for each genotype.Experiments and results
[0103] To understand 1) the impact of limited budget and resources on the total length of the breeding process, 2) the confidence that can be placed on achieving the introgression of the desired traits, and 3), the advantage of choosing different screening techniques at different generations over single technique across generations, we have conducted experiments against the simplest case where the cost of each resource used under a screening technique is the same across generations. We ran the model for two scenarios: 1. Scenario 1 (S1): Without screening technique decisions. Under this scenario, the only available technique is phenotyping. 2. Scenario 2 (S2): With screening technique decisions where it is possible to choose from range of available screening techniques such as phenotyping, MAS, and hybrid technique.
[0104] In Scenario (S1), the cost can be taken as unit cost, which can be justified by the fact that, in each generation, every plant requires the same combination of resources in the same quantity. Unit costs will make the total cost of a generation equivalent to the number of plants screened in that generation. In Scenario (S2), the cost forphenotyping is taken as unit cost, while the cost for MAS is taken as a multiple of the unit cost, set for the purpose of the experiment. The table below gives the configurable parameters for our experiments and their respective ranges.
[0105] The table below shows scenarios and experiments of introgression model.Results for scenario S1
[0106] We run two types of experiments for S1 to access the impact of imposing a hard restriction on achieving the MACL at the second generation (denoted as Experiment 1, or E1 for short) against the case where the only restriction is to achieve MACL at the end of TIP (E2). In E2, the model decides at which generation the MACL can be achieved with a given budget, while striving to achieve as high a confidence as possible in earlier generations. Our experiments show that: 1. E1 requires more budget for success of TIP. Experiment E1 works with the hard restriction of achieving the desired plant with confidence of at least MACL. This restriction impacts the required budget for success of TIP, which is evident from the results where a budget of at least 105^^ is needed for the success of TIP (Figure 6). By contrast, E2 shows that TIP succeeds with a lower budget of 90^^ thanks to this experiment’s flexibility to decide on the generation in which the MACL can be achieved, subject to the available budget. 2. E2 and E1 have the same solution wherever E1 is feasible. Experiments also suggest that the optimal confidence profile (i.e., the confidence level in getting the desired plant at the end of TIP and intermediate generations) produced by E1 and E2 are the same wherever it is possible to achieve MACL in the second generation. This is because E2 does not impose such restrictions but aims to achieve the MACL as early as possible. For example, for budget 120^^, E1 and E2 have the same profile for MACL values of 50,60 and 70 (sub-figures a to c in Figure 6 and in Figure 7). By comparison,E1 fails for ^^^^^^^^ > 70 and E2 requires at least three generations (sub-figure d and ein Figure 7). 3. The confidence profile is sensitive to MACL and available budget in E2: Experiment E2 shows significant variability in the optimal confidence profile with respect to the value of MACL. For example, corresponding to an available budget of120^^, the optimal confidence profile for ^^^^^^^^ = 50 is (75,76,80,89) (sub-figure a inFigure 7), for ^^^^^^^^ = 90 is (75,76,79,93) (sub-figure e in Figure 7), and for^^^^^^^^ = 95 is (75,76,78,95) (sub-figure f in Figure 7).
[0107] The variation of the optimal confidence profiles in E2 is more noticeable for lower budgets. The reason for such variations in experiment E2 is that the MILP strives to maximise the confidence level achievable in earlier generations as much as possible while guaranteeing a plant with a confidence level of at least MACL. For example, corresponding to the available budget of 95^^, the optimal confidence profile for^^^^^^^^ = 50 is (31,32,42,66) (sub-figure a in Figure 7), for ^^^^^^^^ = 70 is(31,32,33,54,77) (sub-figure c in Figure 7), and for ^^^^^^^^ = 90 is(30,31,51,84,97,98)(sub-figure e in Figure7). Therefore, for the higher value of MACL, the confidence level achieved in earlier generations is not as good as the confidence limit achievable with the lower values of MACL, increasing the total number of generations. For example, corresponding to an available budget of 105^^, theoptimal confidence level achieved for ^^^^^^^^ = 50 is (53,54) with a total of twogenerations (sub-figure a in Figure7), whereas the confidence profile for ^^^^^^^^ = 60 is(50,51,81)with a total of three generations (sub-figure b in Figure7) and theconfidence profile for ^^^^^^^^ = 90 is (49,50,83,97,98) with five generations in total(sub-figure e in Figure 7). Results for scenario S2
[0108] To understand how being able to choose from among different screening strategies can help experiment E1 where it fails, we have conducted a third set of experiments (E3) where, at each generation, the model can decide which screening technique to use. For such experiments we consider the same set of restrictions and parameters as for E1 (Table above), together with the following additional assumptions: 1. The flanking markers on both sides of each of the six traits to introgress are given with probability of recombination 0.5. The value is taken for experimental purposes only and does not impose any limitation on the model.2. The left flanking marker coincides with the right flanking marker of the trait on its immediate left. This assumption is for experimental purposes only and does not impose any limitation on the model. 3. The cost of resources used in MAS is twice the cost of resources used in phenotyping.
[0109] Intuitively, despite MAS having a higher cost, its use may reduce the overall cost due to its capacity of identifying the right type of plant with the help of a given marker. In fact, based on the location of flanking markers and recombination probabilities, using MAS may result in an increased success rate, even if constrained by a low budget.
[0110] Our results match these expectations and demonstrate an increase in the success rate of TIP (sub-figures a to f in Figure 8). Our results also show that, under the hard constraint of achieving MACL in the second generation, the model produces a feasible confidence profile by selecting the MAS screening strategy for a budget rangeof 90^^ − 140^^ where E1 had previously failed (see yellow bars in sub-figures a to f inFigure 8 ). However, with the given set of parameters, MAS does not prove to be more efficient wherever phenotyping can produce a feasible profile.
[0111] We believe that the restriction of achieving MACL in the second generation and the fact that phenotyping has cheaper resources are primarily responsible for this behaviour. To verify our belief, we did another experiment with only three traits and abudget range of 1.7^^ − 3.2^^, where we dropped the hard constraint that requiresachieving MACL in the second generation. The results support our belief, and the model chooses MAS over phenotyping if doing so increases the overall confidenceprofile (sub-figures a to f in Figure 9). For example, at ^^^^^^^^ = 50, the modelsuggested using MAS for a budget range of 1.7^^ − 1.9^^ and switches to phenotypingat 2^^, while the breeding profile obtained for 1.9^^ using MAS is feasible for 2^^. The reason for this switch was that it is impossible to achieve a better breeding profile of (51,52,60) with MAS with a budget of 2^^, as it is twice as expensive as phenotyping.However, if budget is increased to 2.3^^, MAS is selected again and results in a better breeding profile and improved overall objective value, even if at a slightly higher cost.
[0112] The model does switch between phenotyping and MAS for all considered values of MACL to achieve a higher confidence level at earlier generations and produce a feasible breeding profile. Switching screening techniques may increase the cost; however, the model will suggest switching between phenotyping and MAS only if it improves the overall objective. We observe that with increase in the value of MACL, the chances of getting a feasible profile using phenotyping decreases. For example, for a budget of 2^^, phenotyping produced an optimal breeding profile of(51,52,60)for^^^^^^^^ = 50, which is not feasible for ^^^^^^^^ > 50 as the confidence achieved in thesecond generation is 52 < 60. Similarly, for a budget of 2.1^^, phenotyping producedan optimal breeding profile of (61,62,68) for ^^^^^^^^ < 70, which is not feasible for^^^^^^^^ >= 70 as the confidence achieved in second generation is 62 < 70. Similarreasoning holds for a budget of 2.2^^ and higher. Stochastic impact of resource availability, market and climate change:
[0113] Given the uncertainty in how climate change will unfold and how the market will respond, the costs and available resources may not be predictable in advance. However, we can assess the probabilities of these metrics over the expected timeframe for the breeding pipeline, which typically spans 7 to 8 years, by taking climate uncertainty into account. For example, when introgressing drought tolerance—a trait highly sensitive to environmental conditions—unexpected climate events like prolonged dry spells or irregular rainfall can drastically alter the effectiveness of field trials and require different types and numbers of resources for irrigation or protected cultivation. Note that not only type and number but the cost may also increase due to the change in demand-supply dynamics for required resources; for example, the change in climate may make some land area or particular resource unfit to use for further testing down the breeding line. Adding to this, market response uncertainty—such as shifts in consumer demand or policy incentives for climate-resilient varieties—can influence the long-term value of a trait, impacting prioritization decisions within thebreeding pipeline. Thus, considering the deterministic inputs such as cost and resource metrics may give a result that may not be viable to implement at the end or is suboptimal.
[0114] To manage uncertainties in breeding decisions, one approach is to create representative scenarios based on varying factors like cost and resource availability. Breeders may solve a deterministic version of the problem for each scenario, resulting in tailored solutions. However, this method can lead to suboptimal decisions since it doesn't account for the likelihood of each scenario.
[0115] A more effective method is stochastic optimization, which considers all scenarios and their probabilities simultaneously. This allows for a single robust solution that performs well across different conditions while mitigating risks. By balancing potential gains and hedging against adverse outcomes, stochastic optimization helps breeding programs make proactive, data-driven decisions. This is especially important in resource-constrained environments, where errors due to overlooked variability can significantly impact genetic gains.
[0116] Provided herein is a stochastic optimisation method to address the above- mentioned challenges and solution. The stochastic version is based on: • Scenarios representing the uncertainties, together with the probability of their occurrence • The cost and resource metric for each scenario • All decision variables and constraints used in the deterministic version will now have another index for scenario ^^ ∈ ^^ where ^^ represents a scenario from thescenario set ^^. • Non-anticipativity constraints that ensure that the decisions are taken considering all scenarios together. Non-anticipativity constraints are used in this stochastic optimization for breeding programs because they ensure consistency across all scenarios. Essentially, these constraints prevent the decision variables from being anticipated in advance for specific scenarios. In other words, decisions made at any stage in the breeding process must be applicable and valid across all possible scenarios considered in the model. For instance, if a breeder decides to allocate acertain resource or choose a specific screening method at an early generation, non-anticipativity constraints ensure that this decision does not depend on the knowledge of which particular scenario will unfold. This approach maintains the robustness and applicability of the solution, as it doesn't allow decisions to be tailored to, or 'anticipate', specific future scenarios. In the context of breeding programs, these constraints are beneficial because they help create a unified strategy that is feasible and effective regardless of the variability in climate, market responses, or resource availability. Therefore, non-anticipativity constraints help mitigate risks by ensuring that the breeding decisions are made holistically, considering all scenarios and their probabilities, rather than making scenario-specific adjustments that could lead to suboptimal or infeasible results in reality. Conclusions
[0117] Plant breeding programs are used to develop new crop varieties that are more resilient to changing environmental conditions. These new varieties exhibit desirable traits, such as pest and drought resistance and increased yields. Traits from donor parents are introduced to elite varieties through the trait introgression process, which produces the desired plant variety with high confidence after a number of generations. This process is increasingly assisted by novel molecular breeding technologies, which can reduce the time it takes to develop new plant varieties when compared to traditional phenotypic selection. On one hand, the many decisions that impact the TIP performance in terms of the total cost, time, required resources, and confidence in getting the desired plant at the end of TIP are often taken by using spreadsheets and expert knowledge; this approach leaves a lot of room for improvement. On the other hand, much of the existing research reported in the literature does not consider the TIP problem in all of its complexity, so the solution that combines expert knowledge and accurate modelling is needed.
[0118] The present disclosure provides a MILP that simultaneously considers the three objectives of minimising cost, minimising the number of generations needed to complete the TIP, and maximising the confidence in achieving the DV at the end of theprocess, while also informing the decision makers about the quality of the solution. Moreover, the proposed MILP is flexible in the sense that it allows for the use of different plant screening strategies at different generations in the same breeding program, a feature that has not been reported previously. We have also presented preliminary results that show that, by setting the Minimum Acceptable Confidence Level, the suggested number of generations and the number of plants could decrease, or that the confidence level needed to obtain a plant with the desired traits could increase. In other examples, the cost model may include other decisions related to the breeding process.
[0119] Figure 10 illustrates a computer system 1000 that is configured to perform the methods disclosed herein. Computer system 1000 comprises processor 1001 and non- volatile memory 1002 (simply referred to as memory 1002 herein). Memory 1002 is non-transitory and stores computer readable code that, when executed by processor 1001, causes processor 1001 to perform the methods disclosed herein, including method 300 and method 400. Processor 1001 may also store data on memory 1002, such as variable values and the cost model as well as screening results and other information relating to the breeding process. Computer system 1000 further comprises random access memory 1003, which is volatile and can be used to store data and program code temporarily for faster access.
[0120] As mentioned above, processor 1001 may accessing the cost model from memory 1002 and may also access the cost model from other sources, including online cloud storage, external drives, network attached storage and the like. The cost model includes breeding process variables and model represents a cost of the breeding process over multiple generations of the organism based on the breeding process variables. Processor 1001 optimises the cost model by calculating breeding process variable values that represent a reduced cost of the breeding process.
[0121] As already described herein, optimising the cost model comprises calculating breeding process variable values for breeding process variables for multiple generations of the organism and the breeding process variables comprise decision variablesrepresenting a selected method for screening individuals for the one or more desired traits at each of the multiple generations of the organism.
[0122] Processor 1001 may then store the calculated variable values on memory 1002, display them on a display device (not shown) or create a document in electronic or paper form, to convey the breeding process variables, including the values of the decision variables to a user, such as a breeder.
[0123] In yet a further example, the processor 1001 may generate a user interface, which may be displayed locally or remotely as a web service. The user interface comprises input fields to enable the user to enter variables of the breeding process, such as maximum budget. The user interface may also comprise output fields to display the calculated breeding process variables to the user. In yet another example, the user interface provides input fields for the user to input results of a screening method for one generation, such as number of retained individuals, and processor 1001 can re-evaluate the cost model to calculate an updated value of the breeding process variables based on the entered screening results.
[0124] Further, processor 1001 may control an automated line of breeding individuals, which is particularly applicable to plants. That is, the processor 1001 controls an automated planting mechanism to generate the specified number of individuals and may then provide those individuals after a pre-determine period of growing time, to the screening method indicated by the values of the breeding process variables.
[0125] It is noted that the methods disclosed herein may equally be implemented by multiple processors or on a distributed computing platform.
[0126] Further, as stated above, the optimizer handles a multi-objective problem and may solve the problem as a weighted sum problem. The weighten sum problem comprises different weights for the different objectives. These different weights may be based on the user input indicative of a decision maker's belief of how important eachobjective is. The optimization model provides an optimized breeding scheme based on these weights, considering the other constraints. Therefore, the user can input a number of different sets of weight for different scenarios or different aims.
[0127] The optimizer can be run in parallel for different sets of weights to get a robust breeding scheme for each set and evaluate the trade-offs between the objectives. Once the optimiser has determined the breeding parameters for the sets of weights, the breeder can decide which breeding profile should be implemented in practice, looking at the trade-offs between objective values. For example, what happens if cost is prioritized vs. time prioritized vs. confidence level prioritized? This will guide the breeder in deciding the most suitable balance between the three objectives that are feasible for him in practice.
[0128] The above process can be supperted by a user interface. The user interface may comprises input fields for user input indicative of the weights of each of the objectives. This may be a numerical entry or slider or other graphical user element. There may be a graphical user element that provides no absolute numerical values but a relative weighting. For example, a user interface may comprises a visual indication of the different weights as a two-dimensional area, such as a radar chart. The user can then set a point in the area to indicate the relative weighting of the objectives. That is, the distance of the user-set point to each objective determines the weighting of that objective relative to the other objectives. In one example, the objective weights are normalised to add up to 1. The user interface may comprise multiple areas so that the user can set multiple sets of relative weights for parallel processing.
[0129] The user interface may also visualize the different results. That is, the user interface may display the results for the output variables for the different sets of objectives. The output values may be represented by a bar chart, which may include a graphical indication (e.g., length of each bar) of the output variable values, which may include number of generations, achieved confidence and overall cost. The output values or bar charts may be updated in real-time as the user changes the weights, such as bymoving the user-set point in the two-dimensional area. The user can then decide what best suits the need.
[0130] The breeder can plan resource allocation based on informed decisions and objective values; for example, if the optimization suggests “X” number of plants to select in a generation “g” in the future, then the breeder can plan in advance for the resources to ensure that “X” plants can be handled in generation “g” at the appropriate point in time.
[0131] Based on the informed decision on the number of individuals to select / retain the screening technique to use, and how much benefit it gives in terms of confidence achieved vs cost vs time, the breeder can decide whether it is beneficial to invest in the infrastructure / tools (for example, investing in genomic screening setup) and build in- house capability or outsourcing certain tasks to specialized service providers or outsourcing certain tasks to specialized service providers.
[0132] The recommended number of selected / retained individuals of different genotypes guides the breeders in selecting the most promising plants to evaluate and advance in the breeding process.
[0133] If there is a scope for flexibility in infrastructure investment decisions, the breeder can run the model for different values of parameters, such as the minimum and maximum number of populations that can be handled (bounds on x variables). This way, the model result can guide them to determine whether the minimum investment in infrastructure is sufficient to achieve the desired / acceptable results.
[0134] Figure 11 illustrates another example of a computer system 1100 for implementing the disclose methods for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism. Computer system 1100 comprises a cost model module 1101 to access and manage a cost model. This involves retrieving and updating breeding process variables in the cost model, which represents a cost of the breeding process over multiple generations of theorganism based on the breeding process variables. Computer system 1100 further comprises an optimisation module 1102 to optimise the cost model by calculating breeding process variable values that represent a reduced cost of the breeding process. The optimisation module 1102 calculates the breeding process variable values for the breeding process variables for multiple generations of the organism. This includes calculating the values of decision variables representing a selected method for screening individuals for the one or more desired traits at each of the multiple generations of the organism.
[0135] The different modules in Figure 11 may be implemented as software modules or hardware models or in other ways. They may be implemented in a single computer or single processor or on different virtual machines, processors, computer systems servers, images and the like.
[0136] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.
Claims
CLAIMS:
1. A method for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism, the method comprising: accessing a cost model the cost model including breeding process variables, the cost model representing a cost of the breeding process over multiple generations of the organism based on the breeding process variables; optimising the cost model by calculating breeding process variable values that represent a reduced cost of the breeding process, wherein optimising the cost model comprises calculating breeding process variable values for breeding process variables for multiple generations of the organism; wherein the breeding process variables comprise decision variables representing a selected method for screening individuals for the one or more desired traits at each of the multiple generations of the organism.
2. The method of claim 1, wherein the cost model comprises a lower bound on a number of individuals to screen for each generation.
3. The method of claim 2, wherein the lower bound on the number of individuals to screen for each generation is based on the selected method for screening individuals at that generation.
4. The method of claim 3, wherein the number of individuals to screen for each generation is based on a number of retained individuals in that generation and a confidence in individuals of that generation including the one or more desired traits.
5. The method of any one of claims 2 to 4, wherein, in case the selected method for screening individuals is phenotype screening, the lower bound is a function of: a number of retained individuals, and a confidence in a final generation of the breeding process including the desired trait.
6. The method of claim 2, wherein, in case the selected method for screening individuals is marker-assisted screening, the lower bound is based on a transition probability between genotype groups, wherein at least some of the genotype groups include multiple genotypes.
7. The method of claim 6, wherein the genotype groups are defined based on whether the one or more traits are present in homozygous or heterozygous form.
8. The method of claim 6 or 7, wherein the method comprises introgressing multiple traits and the genotype groups are defined based on the number of traits present in an individual.
9. The method of claim 8, wherein the genotype groups cover all possible genotypes in relation to the one or more multiple traits.
10. The method of any one of claims 6 to 9, wherein the method comprises determining the lower bound that guarantees at least one individual of the retained individuals contains the multiple desired traits at a one of one or more confidence values.
11. The method of any one of claims 6 to 10, wherein the transition probability is based on a distance between genetic markers associated with the one or more desired traits.
12. The method of any one of the preceding claims, wherein the cost model comprises: an integer variable indicative of a number of individuals created at each generation; and a binary decision variable indicative of whether the one or more traits are present in retained individuals at a confidence value.
13. The method of claim 12, wherein optimising the cost model is constrained by the integer variable being greater than a sum of the binary decision variable weighted by a lower bound on the integer variable for multiple values of retained individuals.
14. The method of any one of the preceding claims, wherein optimising the cost model comprises minimising expenditure.
15. The method of any one of the preceding claims, wherein optimising the cost model comprises minimising a number of generations of the breeding process.
16. The method of any one of the preceding claims, wherein optimising the cost model comprises maximising a confidence in a final generation of the breeding process including the desired trait.
17. The method of any one of the preceding claims, wherein the cost model is a multi-objective model and optimising the cost model comprises simultaneously: minimising expenditure, minimising a number of generations of the breeding process, and maximising a confidence in a final generation of the breeding process including the desired trait.
18. The method of any one of the preceding claims, wherein the cost model comprises a mixed integer programming problem comprising integer variables that indicate the number of individuals and binary variables that are the decision variables.
19. The method of any one of the preceding claims, wherein optimising the cost function is constrained by a maximum cost value.
20. The method of any one of the preceding claims, wherein the cost model represents multiple scenarios and each of the breeding process variables is associated with one of the multiple scenarios.
21. The method of claim 20, wherein the cost model comprises constraints to ensure that decisions made at any stage in the breeding process are applicable and valid across all possible scenarios considered in the cost model.
22. A method for introgressing a desired trait into an elite variant of an organism, the method comprising: creating a model for a progeny genotype as one of multiple genotype classes, wherein each of the multiple genotype classes is indicative of a number of desired traits in each individual; calculating transition probabilities between genotype classes when crossing the progeny with the elite variant; and using the model to calculate breeding process variable values.
23. A computer-system comprising one or more processors configured to perform the method for introgressing one or more desired traits into an organism using a breeding process over multiple generations of the organism according to any one of claims 1 to 20.