Method for decision support in selecting crossings for population

The method optimizes plant crossing selection through multi-objective optimization for genetic gain and resource efficiency, addressing the challenges of traditional breeding methods by reducing generations and crossings.

WO2025262645A1PCT designated stage Publication Date: 2025-12-26BASF AGRICULTURAL SOLUTIONS US LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056277
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-06-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Selecting crossings for plant population development is challenging due to the difficulty in objectively deciding which parents to choose and the time-consuming nature of traditional breeding methods, especially with climate-induced biotic and abiotic stress factors, requiring improved decision support.

Method used

A computer-implemented method for decision support in selecting crossings using multi-objective optimization to maximize short-term and long-term genetic gain while minimizing resource requirements, involving pareto front evaluation and user or automated selection of solutions.

Benefits of technology

Accelerates plant population development by optimizing resource allocation and reducing the number of generations and crossings needed to achieve desired genetic gains, with semi-automated or fully automated selection processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025056277_26122025_PF_FP_ABST
    Figure IB2025056277_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides a computer-implemented method for decision support in selecting crossings for population development, the method comprising: receiving a set of candidate parents, a set of traits, and one or more constraints comprising a resource constraint representative of a maximum number of crossings and / or a maximum progeny population size; carrying out a multi-objective optimization; evaluating a pareto front of the set of solutions and selecting a subset of one or more solutions within the set of solutions based on a result of evaluating the pareto front; selecting a solution of the subset of solutions; and providing the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for decision support in selecting crossings for population

[0002] FIELD OF THE INVENTION

[0003] The invention provides a method for decision support in selecting crossings for population development, a system, a computer program product, and a computer-readable medium.

[0004] BACKGROUND

[0005] Plant breeding is the science of changing the traits of plants in order to produce desired characteristics which is used for cultivar development, crop improvement, and seed improvement. Plant breeding can be carried out by using different techniques including simply selecting plants with desirable characteristics for propagation, methods of using knowledge of genetics and chromosomes as well as complex molecular techniques. The modern plant breeding techniques include, for instance, marker assisted selection, genetic modification, genomic selection, and reverse breeding and doubled haploids. The goals of plant breeding are to produce crop varieties that boast unique and superior traits for a variety of applications. The interested traits include increased nutrition, flavor, and yield of crops, increased tolerance to environmental pressures (e.g. salinity, drought), increased tolerance to herbicides, and resistance to fungi, bacteria, and insect. Traditionally, breeding involves the creation of multi-generation genetically diverse populations on which human selection is practiced to generate adapted plants with new combinations of specific desirable traits. The selection process is driven by biological assessment in relevant target environments and available knowledge of genes and genomes. Progress is assessed based on gain under selection, which is a function of genetic variation, selection intensity, and time.

[0006] At present selecting crossings of parents, e.g. for plant population development, is very challenging. Generally, plant population development is done with the aim of increasing certain traits. However, it is very difficult to objectively decide which parents to select for crossing and how to pursue population development. In addition, the improved plant varieties should face the various biotic and abiotic stress factors that are increasing with the current climate changes. Moreover, the breeding and selecting can be very time-consuming. Even with the very latest in biotech-assisted conventional breeding, incorporation of a trait may take an average of seven generations for clonally propagated crops, nine for self-fertilizing, and seventeen for cross-pollinating.

[0007] Approaches for selecting crossings may be done by a person, e.g. a breeder, in some cases supported by statistical approaches. However, overall the selection process requires improvement. Particularly, there is a need for providing automatic decision support in selecting crossings for population development.

[0008] Thus, it is an object of the present invention to provide a method that allows for overcoming at least some of the above challenges, particularly providing a method and system for decision support in selecting crossings for population development.

[0009] SUMMARY

[0010] The object is achieved by the present invention. The invention provides methods, systems, computer program products, and computer-readable mediums according to the independent claims. Preferred embodiments are laid down in the dependent claims.

[0011] The present disclosure provides a computer-implemented method for decision support in selecting crossings for population development, the method comprising: receiving a set of candidate parents, a set of traits, and one or more constraints comprising a resource constraint representative of a maximum number of crossings and / or a maximum progeny population size; carrying out a multi-objective optimization, wherein carrying out the optimization comprises, constrained by the received one or more constraints, optimizing at least a first optimization objective, a second optimization objective, and a third optimization objective, and outputting a set of solutions, wherein each solution is a list of crossings of parents (each crossing being a crossing of two parents) of the set of candidate parents, wherein the first optimization objective is a short-term genetic gain objective associated with the set of traits and the optimization comprises maximizing the short-term genetic gain, wherein the second optimization objective is a long-term genetic gain objective and the optimization comprises maximizing the long-term genetic gain, and wherein the third optimization objective is a resource requirements objective associated with a predicted required number of progenies and the optimization comprises minimizing the resource requirements; evaluating a pareto front of the set of solutions and selecting a subset of one or more solutions within the set of solutions based on a result of evaluating the pareto front. outputting, to a user, a representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the subset of one or more solutions, and receiving a user input selecting a solution of the subset of solutions, or automatically ranking the solutions of the subset of solutions based on an objective function used in the multi-objective optimization and automatically selecting the highest-ranked solution; providing the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.

[0012] Thus, a method is provided that provides a list of crossings of parents, i.e., the selected solution, and, for each of the crossings a predicted required number of progenies. The method as disclosed herein allows for semi-automated or fully automated selection of the solution.

[0013] Based on the solution, breeding may be carried out to obtain respective progenies of the respective crossing of parents. This may be considered to be population development, the population development being based on the list of crossings.

[0014] This relates particularly to plant population development.

[0015] With conventional methods, it would take more generations and / or more crossings and associated resources to obtain similar results. Therefore, the present method allows for better resource allocation and accelerates population development. To some degree, the present method alleviates trial and error of existing methods. In addition, the method of the present disclosure allows for low demand in processing and computing resources.

[0016] As such, the above-mentioned challenges and other challenges are addressed by the method disclosed herein.

[0017] In other words, the present method comprises the above multi-objective optimization and subsequently carrying out a sequential selection process, wherein, in a first selection step, a subset of one or more solutions within the set of solutions may be is selected based on evaluation of the pareto front, e.g., by finding solutions at or near the pareto front and selecting them as subset. The solution(s) may be associated (and, in case of an intended subsequent user selection, presented together) with respective compromises of at least one of the optimization objectives. The first selection step may, for example, be fully automated.

[0018] In a second selection step, a solution may be selected among the subset of solutions. This may be done in different ways.

[0019] For example, the method may comprise automatically determining a ranking of the solutions based on an objective function used in the multi-objective optimization and automatically selecting the highest- ranked solution. This scenario may allow for full automation of the selection of a solution. Alternatively, receiving a user input may be involved in the second selection step. A representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the subset of one or more solutions may be outputted to a user. A user input selecting a solution of the subset of solutions may then be received. Thus, in the user-selection scenario, display at a graphical user interface may allow for the user to review the compromises. In both scenarios, after the subset selection, further evaluation and selection of the solution is based on quantitative analysis and potentially reporting, of the impact of such compromises on the objectives.

[0020] The second selection step may involve receiving user input, and, the method may comprise, after the first selection step, outputting, to a user, a representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the subset of one or more solutions, and the second selection step may comprise receiving a user input selecting the at least one solution of the subsets of solutions.

[0021] Optionally, as will be explained below, potential solutions of the subset of solutions may be checked as to whether the respective associated compromises are acceptable. Predetermined criteria may be applied to that end. Otherwise, the optimization and selection of a subset may be executed again with (minor) adjustments, e.g. adjusted constraints and / or objective in the optimization and / or adjusted evaluation of a pareto front, e.g. search parameters, and / or adjusted reviewed objectives in the selection in order to identify an optimal solution for the modified objectives.

[0022] It is to be understood, while above three optimization objectives are mentioned as an example, more than three optimization objectives may be used.

[0023] The optimization of the optimization objectives may be independent or dependent.

[0024] It is to be understood that, herein, “population”, “parent”, “progeny” and like terms refer to plants, particularly plant breeding. Thus, the methods herein are to be understood as not entailing animals or humans. The plant population shall mean the number of plants in an area of land. It is calculated by dividing the number of seeds planted by the number of acres. Plant populations exhibit genetic variability due to random mutation, gene flow, and natural selection. Progeny may refer to a genetic descendant or offspring, and to any generation plant. A progeny plant may be from any filial generation, e.g., F1.

[0025] In the present disclosure, the term “user” is to be understood broadly and can be any human providing input via a user interface. In general, the method may be a distributed method and user interaction or user input may involve different users, different times, different input devices, and the like. The term “user” may entail a software scientist and / or breeding scientist and / or a breeder. The method, as such, works irrespective of the user identity and role.

[0026] The resources, herein, may refer to resources associated with crossing parents to obtain the progeny, upkeep of the progeny, and study of the progeny. Resource constraints may be associated with a maximum number of potential parents and, accordingly, maximum number of possible crossings, and / or an overall maximum number of progenies. The number of progenies is referred to as the progeny population size, which can for example be a per-crossing and / or a total progeny population size. Thus, a resource constraint may be associated with a maximum progeny population size. In any case, it may be desirable to reduce the number of progenies to reduce resource requirements. This may be accomplished, in the present disclosure, by a corresponding optimization objective, the resource requirements objective, which may aim at minimizing the predicted required number of progenies.

[0027] In particular, a resource according to the present invention shall mean any factor which is necessary to accomplish a breeding goal or carry out a breeding activity. In short, they are the components to achieve breeding goal including tangible and intangible assets, but is not limited to those. The resource can be plants, plant materials, and plant germplasm used for breeding.

[0028] Providing a predicted required number of progenies for a crossing may comprise retrieving and / or determining said predicted required number of progenies.

[0029] The predicted required number of progenies may be the result of a genomics-based method. The predicted required number of progenies may be the result of statistical methods, e.g. be probabilitybased.

[0030] The predicted required number of progenies for a crossing may be a predicted number of progenies obtained by said crossing is needed to have a certain confidence that at least a certain number of the progenies obtained by said crossing meet predetermined criteria, e.g., criteria associated with certain goals, such as a short term genetic gain target. This may be determined using genomics-based methods, e.g. including statistical methods.

[0031] It will be understood from the detailed discussion above and / or below, that the method may comprise predicting the required number of progenies. This prediction may be genomics-based and may optionally involve, particularly usually involves, statistical methods.

[0032] The short-term genetic gain may be an overall improvement in aggregated trait values over the course of one or more breeding generations, as an example. By ways of the short-term genetic gain objective, the method aims at maximizing the short-term genetic gain.

[0033] As will be discussed in more detail below, a long-term genetic gain may be indicative of a projected genetic diversity over multiple generations. By ways of the long-term genetic gain objective, the method aims at maximizing the long-term genetic gain.

[0034] Genetic gain or genetic gain from selection may be seen as the improvement in average genetic value in a population or the improvement in average phenotypic value due to selection within a population over cycles of breeding (Hazel and Lush, 1942).

[0035] The multi-objective optimization may be based on an objective function. A feasible solution that minimizes or maximizes (depending on the formulation of the optimization problem) the objective function may be considered an optimal solution. Usually there are several local minima / maxima, such that several solutions will be yielded, i.e., a set of solutions.

[0036] Evaluating the pareto front may involve determining the pareto front and optionally finding solutions at the pareto front and / or finding solutions within a certain distance from the pareto front. Thus, solutions at or near the pareto front may be selected. The term “near the pareto front” is to be understood as meaning “within a predetermined distance from the pareto front”.

[0037] Ranking the solutions based on an objective function may, for example, rank solutions according to their respective scores on the scale of the objective function.

[0038] A crossing may be indicated, in the list, by the pair of parents to be crossed. As such, a list of crossings (i.e. a solution) may comprise a list of pairs of parents to be crossed. The invention also provides a method comprising the methods for decision support of the present disclosure, and further comprising using the list of crossings and selecting and crossing the parents in the manner indicated in the list for generating a progeny population and / or providing resources for population development based on a / the list of crossings.

[0039] That is, a method may be provided that employs the selected solutions, specifically the provided selected list of crossings and the predicted required number of progenies for each of the crossings of said list of crossings.

[0040] Based thereon, breeding may be carried out to obtain respective progenies of the respective crossing of parents. This may be considered to be (breeding) population development based on the list of crossings. The respective resources for the population development may be provided, such as spatial resources, material resources, and time-related resources, the resources allowing for crossing the parents and then growing and studying the progenies.

[0041] With conventional methods, it would take more generations and / or more crossings and associated resources to obtain similar results. Therefore, the present method allows better resources allocation and allows for more efficient (plant) population development.

[0042] The invention also provides a system comprising a computing system configured to carry out the method of the present disclosure.

[0043] The invention also provides a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the methods of the present disclosure.

[0044] The invention also provides a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the methods of the present disclosure.

[0045] The features and advantages outlined above in the context of the method similarly apply to the system, the computer program product, and the computer-readable medium described herein.

[0046] According to the present disclosure, evaluating the pareto front comprises using an algorithm, such as a search algorithm, to select a subset of one or more solutions within the set of solutions.

[0047] In an example, a search algorithm may find and select a solution that scores best across all objective functions. In an example, if the constraints are too strict, a user may restart the search algorithm procedure with more relaxed constraints (or with less constraints).

[0048] The algorithm may be an algorithm that finds the pareto front and distance of the solutions from the pareto front. It may yield solutions at or within a predetermined distance from the pareto front. It is noted that the solutions will be associated with different trade-offs or compromises. The algorithm will narrow down the potential solution list based on the solution’s relation to, e.g. distance from, the pareto front. Such a selection can be considered to be an objective selection that does not require to know and take into account the actual content of each solution or meaning / consequence of the tradeoff. This allows for ease of transfer to different scenarios and does not require availability / knowledge of all underlying information and principles.

[0049] The method according to the present disclosure may comprise, prior to carrying out the optimization, pre-selecting a subset of candidate parents among the set of candidate parents based on genomic information, wherein the optimization is carried out only for crossings of the subset of candidate parents.

[0050] As an example, the genomic information may comprise genomic breeding values, trait profile (e.g. native and transgenic traits), marker frequency, marker-based additive and / or dominance effects, and quantitative genetic loci (QTL) associated with traits of agronomic importance (disease resistance, flavor, shape, color, maturity, and others).

[0051] In other words, commonly used genomic breeding values or values derived therefrom can be employed to identify, in advance, the most suitable input data into the optimization, which reduces computational resources of the optimization and subsequent steps, and overall is likely to yield better results. It can be seen as a filter for input data that filters for specific genomic properties of the potential parents.

[0052] The method according to the present disclosure may particularly comprise simulating, making use of genomic information, mean and variance of a progeny for every possible crossing. The optimization may be carried out only for a subset of crossings for which feasibility for a set of genetic gain targets has been confirmed based on the simulated mean and variance of a progeny.

[0053] For example, only parents associated with said subset of crossings may be input into the optimization.

[0054] The mean and variance of a progeny may relate to any trait that shows a quantitative distribution of the observed genomic breeding values. Such traits can be assumed controlled by many genes with additive and / or dominance effects. These effects are specific to the considered breeding program and are derived using statistical methods known as genomic selection methods.

[0055] Mean and variance of a progeny may be predicted using simulation or prediction equations, either using pre-calculated marker effects that were previously calculated using common methods.

[0056] According to the present disclosure, the short-term genetic gain and / or the long-term genetic gain may be determined making use of (the) genomic information.

[0057] The method of the present disclosure may comprise determining the expected short-term genetic gain and / or long-term genetic gain. Known simulation methods, which are generally statistical methods, may be used for determining the genetic gains. Accordingly, the determining will generally be carried out automatically.

[0058] According to the present disclosure, the short-term genetic gain may be determined based on, e.g. stochastically, predicted superior progeny values for each of the crossings, in particular wherein the superior progeny value for the respective crossing is an aggregated value aggregated over the set of traits with optional weighting of the traits.

[0059] The superior progeny value may be considered a short-term genetic gain potential score of a given crossing combination. It may be calculated using known methods. Further details are also provided below.

[0060] The superior progeny value may be a value of a particular cross thus depends on the expected performance of its best progeny. Superior progeny value may be a linear combination of the mean of the cross’s progeny and their standard deviation. A description and method for calculating the superior progeny value is, for instance, described in Shengqiang Zhong and Jean-Luc Jannink, Genetics 177, 567-576 (2007).

[0061] According to the present disclosure, the long-term genetic gain may be representative of a genetic diversity, wherein the genetic diversity is expressed by a minor allele frequency score and / or an inbreeding coefficient. Genetic diversity or long-term genetic gain may express an overall expected level of genotypic variance among individuals in a population.

[0062] The minor allele frequency score may be a score (e.g. a value) that quantifies a minor allele frequency. Specifically, the minor allele frequency score may identify the number of rare alleles in a population basis which a specific individual and its potential crossing combinations carry. Individuals with higher scores are desirable since they allow carrying forward rare alleles that can be valuable for future generations in the breeding program.

[0063] The inbreeding coefficient is a measure that quantifies how closely related individuals are in a population due to common ancestry. The higher this coefficient, the lesser is the genetic variability in a population. The inbreeding coefficient is well known to a person skilled in the art. The description of a method for calculating the inbreeding coefficient is, for instance, described in Deniz Akdemir and Julio I. Sanchez, Efficient Breeding by Genomic Mating, Frontiers in Genetics, Vol. 7, 210, (2016).

[0064] The above can be used as predictors for long-term genetic gain and allow to quantitatively take it into account in an automatic determination of solutions, particularly in an optimization.

[0065] According to the present disclosure, the inbreeding coefficient of a solution may depend on at least one of a number of times a parent is used for crossing in the solution, a number of unique parents used for crossing in the solution, and the similarity of the parents of the respective crossing.

[0066] According to the present disclosure, the long-term genetic gain objective may be based on only one of the inbreeding coefficient and the minor allele frequency score. This may already yield good results while reducing complexity and resource requirements of the determination of solutions.

[0067] According to the present disclosure, the long-term genetic gain objective may be based on both, the inbreeding coefficient and the minor allele frequency score. This may improve the overall results, i.e., improve the selected solution, particularly in terms of the long-term genetic gain. Optionally the method may comprise normalization of the inbreeding coefficient and the minor allele frequency score and combining the normalized inbreeding coefficient and minor allele frequency score into a single optimization objective - which may be called diversity measurement. This allows for making it easier and more efficient for the optimization to find a result.

[0068] The subset of one or more solutions may be a first subset of one or more solutions and the method according the present disclosure may comprise determining prior to selecting a solution of the subset of solutions, particularly in response to a user input and / or by automatic determination based on the optimization objective, whether any of the first subset of solutions meets predetermined criteria and, in case none of the first subset of solutions meets the predetermined criteria: a) modifying the set of candidate parents, the set of traits, and / or at least one of the one or more constraints; b) carrying out the optimization with the modified set of candidate parents, the modified set of traits, and / or the modified constraint(s), the optimization comprising outputting a second set of solutions; c) evaluating, particularly using a search algorithm, a pareto front of the second set of solutions and selecting a second subset of one or more solutions within the second set of solutions based on a result of evaluating the pareto front; d1) outputting, to a user, a representation of the potential crossings, their associated shortterm genetic gain and long-term genetic gain, and a representation of the second subset of one or more solutions, and receiving a user input selecting a solution of the second subset of solutions or of the first and second subset of solutions, or d2) automatically ranking the solutions of the second subset of solutions based on an objective function used in the multi-objective optimization and automatically selecting the highest-ranked solution among the second subset of solutions or among the first and second subset of solutions; providing the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.

[0069] Optionally, the method may only go to steps d1) or d2) in case it is determined that the second subset of solutions meets the predetermined criteria, and otherwise the method may return to step a).

[0070] That is, an iterative process may be provided that allows for repeating optimization with adjusted settings in case the subset of solutions has no solution that is considered acceptable based on predetermined criteria. Solutions may not meet criteria due to the constraints or combination of constraints. For example, some constraints may be loosened or tightened selectively. Moreover, the selection of candidate parents may have been such that no sufficiently good results could be yielded, e.g. due to the parent’s trait profile. In this case using a different set of candidate parents may improve the results. Also the set of traits, which are in some cases conflicting, may be a reason for not obtaining sufficiently good solutions. Thus, the set of traits may be changed.

[0071] The meeting of criteria may be determined automatically. Alternatively or in addition, a user may evaluate whether criteria are met. User input may be received, which may trigger automatic modifying or provide input for the modifying of step a). The modifying may be done based on received user input and / or based on automatic adherence to rules for modifying. Said rules may be based on theoretical considerations, or be empirically or semi-empirically determined rules. Also, the modification can be obtained using artificial intelligence algorithms, e.g. artificial intelligence algorithms that can extensively evaluate the larger pool of candidates available in a breeding program.

[0072] Optionally, after the first repetition of the optimization, the method may proceed with the (final) selection of a solution, irrespective of whether criteria are met. Alternatively, repetitions may continue until the criteria are met, optionally unless a predetermined number of repetitions has been reached before the criteria are met.

[0073] As an example, solutions with maximum acceptable compromises (in different aspects) may be used as benchmarks for determining whether criteria are met. For example, maximum (acceptable) liability and minimum (acceptable) benefit can be expressed in the form of quantitative criteria that can be applied to the solutions. If no solution meets those criteria, they should not be pursued, which can be ensured by the above iterative process.

[0074] According to the present disclosure, the constraints may comprise constraints related to the inclusion and / or exclusion of specific crosses and / or constraints related to achieving a maximum value for one or more predetermined traits, and / or constraints related to achieving a constant value for one or more predetermined traits, and / or constraints related to achieving a minimum value for one or more predetermined traits, and / or constraints related to a predetermined proportion of crosses with a desirable characteristic that should make up the final crossing list.

[0075] The one or more constraints may comprise at least one of a minimum short-term genetic gain, a minimum solution size, a maximum solution size, an inclusion constraint associated with one or more crossings to be included in each of the solutions, a constraint associated with one or more crossings to be included or excluded from each of the solutions, a secondary traits constraint associated with one or more secondary traits.

[0076] Constraints related to the inclusion and / or exclusion of specific crosses may be seen as constraints related to must-have / may-not-have specific crosses. The desirable characteristics may comprise target crossing outcomes for categorical traits (such as disease resistance, etc.).

[0077] According to the present disclosure, the subset of solutions may comprise solutions at or within a predetermined distance from the pareto (dominant) front of solutions of the multi-objective optimization. That is, the pareto front of solutions may be determined and a predetermined distance, e.g., fixedly set or selected by a user may be compared to an actual distance of each of the solutions from the pareto front. Only if the solution is at the pareto front or if the actual distance is equal to or smaller than the predetermined distance, the solution may be selected to be part of the subset of solutions.

[0078] This allows for an objective selection that yields good results. If allowing for a predetermined distance from the pareto front, it is at the same time possible to reduce the likelihood of a very small or zero subset of solutions, which could result from, e.g., only selecting solutions exactly at the pareto front.

[0079] According to the present disclosure, the representation of the potential crossings and their associated short-term genetic gain and long-term genetic gain may comprise representing each crossing within a two-dimensional coordinate system, the coordinate system having a first axis representative of the short-term genetic gain and a second axis representative of a long-term genetic gain.

[0080] As such, information required for making a selection of a solution with all relevant information, particularly taking into account trade-offs, can be made, which improves decision accuracy and outcome.

[0081] Variation of the predetermined distance may allow for tuning the subset size in a quantitative manner related to quality of the solutions.

[0082] According to the present disclosure, the representation of the subset of solutions comprises, for each solution of the subset, an indication of a group of crossings associated with the solutions, particularly the group being represented in the coordinate system such as by heatmaps or shapes (in other words: symbols, indication, representations, for example boxes) enclosing the crossings of the group of crossings or other highlights highlighting crossings of the group of crossings.

[0083] The shapes or other highlights allow for providing all information required to evaluate the tradeoffs made by selecting one of the solutions. That is, a user will be provided with the information, which crossings belong to which solution while still being able to review the quality (e.g. tradeoffs) of each crossing, and thereby being able to make the most informed decision and selection among the solutions.

[0084] The shapes enclosing the crossings may be of any shape. For example, shapes and / or colors of the shapes enclosing the crossings may be employed to distinguish between solutions. Other representations / highlights are also conceivable, but the shapes may allow for little interference with the remaining areas, i.e., reduced information loss.

[0085] According to the present disclosure, optionally, the representation of the potential crossings, and e.g. their associated diversity, may further comprise representing each crossing as a geometric shape, such as a circle, within the coordinate system, the size of the geometric shape being correlated with a first metric of the long-term genetic gain and the second axis value is correlated with a second metric of the long-term genetic gain, in particular wherein the first metric is associated with one of an inbreeding coefficient and a minor allele score, and the second metric is associated with the other one of the inbreeding coefficient and the minor allele score.

[0086] Such representation allows for providing all information (which is multi-dimensional) that is required to select solutions based on trade-offs in one representation, which improves selection accuracy and reduces trial and error compared to other representations that fail to provide all information concurrently. TERMINOLOGY

[0087] "As used herein, terms in the singular and the singular forms like "a", "an" and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "plant", "the plant" or "a plant" also includes a plurality of plants; also, depending on the context, use of the term "plant" can also include genetically similar or identical progeny of that plant or plants derived therefrom by crossing. Also as used herein, the word "comprising" or variations such as "comprises" or "comprising" will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.

[0088] As used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or"). The term “comprising” also encompasses the term “consisting of’.

[0089] Below, some terminology will be provided. The terminology is provided for illustrative purposes, and particularly for understanding the subsequent detailed examples. The terminology is not to be understood as limiting to the claim scope.

[0090] Related to combinatorial optimization:

[0091] • Solution / alternative: A specific and unique outcome to a decision making problem

[0092] • Decision variables: A set of unknowns which, when assigned specific known values from a finite set, combinedly depict the constitution of a particular solution

[0093] • Objective function: A metric used to measure the desirability of a solution with respect to a certain criterion

[0094] • Hard constraint: A condition used to assess whether or not a solution is realistically and practically implementable (i.e. a condition used to assess the feasibility of a given solution)

[0095] • Soft constraint: A condition on the ideal constitution of solutions, the violation thereof incurs certain penalty costs in the / one-of-the objective function(s)

[0096] • Optimization: The process of finding the most desirable / preferred feasible solution in a decision making problem

[0097] • Combinatorial optimization problem: A decision making problem with a finite amount of solutions in which the desirability and feasibility of each solution can be measured and assessed

[0098] • Optimization model: The abstract blueprints of an optimization problem, used to map decision-making goals to computational programming code.

[0099] • Move: A transformation used to slightly alter the inherent composition of a given solution

[0100] • Neighbor solution: A solution that can be generated from any other given solution after performing a single move

[0101] • Solution neighborhood: The set of all potential neighbor solutions to a given solution

[0102] • Search algorithm: A finite sequence of moves used to find one or more feasible desirable solutions in an optimization problem

[0103] • Initial solution: The starting point of a search algorithm.

[0104] • Domain space: The set of all solutions • Feasible domain space: The set of all feasible solutions

[0105] • Objective space: The set of all objective function(s) measurements forming a one-to-one or one-to-many relationship with all solutions in the domain space

[0106] • Search space: The set of all potential solutions in the radar of a search algorithm

[0107] • Constraint relaxation: The act of reducing the impact of a hard constraint on the feasible domain space as a means to improve the navigation capabilities of the search algorithm throughout the search space.

[0108] • Dominance: A condition used to assess whether a given solution scores better than another given solution across all objective functions in a multi-objective optimization problem

[0109] • Archive: A dynamic set keeping track of all feasible non-dominated solutions found during the course of a search algorithm

[0110] • Non-dominated front: The final archive obtained, upon termination of the search algorithm

[0111] • Pareto front: The best possible non-dominated front that intrinsically exists (but that has no certainty of being uncovered).

[0112] Related to plant breeding:

[0113] • Scenario: A plant breeding program categorized by type of crop, breeding season, and geographical location on the globe

[0114] • Breeding population: A population of plant individuals of a given crop in the experimental phase of a scenario

[0115] • Crossing combination: A specific pair of plant individuals from the same cultivar to be used as potential parents breeding material

[0116] • Progeny: A child individual(s) obtained from the breeding process of a crossing combination

[0117] • Progeny population size: The total number of progenies bred from a specific crossing combination.

[0118] • Genotype: The genetic code of a plant individual

[0119] • Phenotype: The observable physical characteristics of a plant individual

[0120] • Trait: A phenotypic characteristic of interest to the breeders, farmers, decision makers and / or market

[0121] • Short-term genetic gain: The overall improvement in aggregated trait values over the course of generations

[0122] • Elite progeny: A progeny individual that meets or exceeds the targeted short-term genetic gain benchmark

[0123] • Superior progeny value: The estimated short-term genetic gain potential score of a given crossing combination

[0124] • Genetic diversity: The overall level of genotypic variance among the individuals in a cultivar. This will sometimes also be referred to as the long-term genetic gain.

[0125] • Primary trait: A trait whose attributes feature in the short-term genetic gain objective function in the optimization model

[0126] • Secondary trait: A trait whose attributes feature in a hard constraint in the optimization model

[0127] • Inbreeding: The process of breeding plant individuals that are closely genetically related. • Cultivar: In general, a cultivar may be a kind of cultivated plant that people have selected for desired traits and which retains those traits when propagated. It may be an organism and especially one of an agricultural or horticultural variety or strain originating and persistent under cultivation. Cultivars are usually produced through artificial selection and are different forms of the same species. In the present disclosure, cultivar can have equal meaning of breeding population and selected crossings.

[0128] Further features, examples, and advantages will become apparent from the detailed description making reference to the accompanying drawings.

[0129] BRIEF DESCRIPTION OF THE DRAWINGS

[0130] In the accompanying drawings,

[0131] Fig. 1 illustrates an exemplary method according to the present disclosure;

[0132] Fig. 2 illustrates another exemplary method according to the present disclosure;

[0133] Fig. 3 illustrates another exemplary method according to the present disclosure;

[0134] Fig. 4 illustrates a system according to the present disclosure;

[0135] Fig. 5 illustrates another exemplary method according to the present disclosure;

[0136] Figs. 6a and 6b illustrate statistical considerations on population size and genetic gain;

[0137] Figs. 7a to 7c illustrate exemplary visualizations of example solutions;

[0138] Fig. 8 illustrates simulation results of an evolution of a breeding population relative to their overall phenotype performance;

[0139] Figs. 9 and 10 illustrate exemplary secondary traits constraints input data.

[0140] DETAILED DESCRIPTION

[0141] Fig. 1 illustrates an exemplary computer-implemented method 100 for decision support in selecting crossings for population development according to the present disclosure.

[0142] In step S11 , the method comprises receiving a set of candidate parents, a set of traits, and one or more constraints comprising a resource constraint representative of a maximum number of crossings and / or a maximum progeny population size.

[0143] In other words, an input may be provided that specifies which parents are candidates for potential crossings. Furthermore, a desired set of traits is provided as input. Constraints for the optimization may be provided as input, which comprise the resource constraints. Resource constraints will generally comprise a maximum number of crossings and / or maximum progeny population size, i.e., a maximum number of progenies.

[0144] In step S12, the method comprises carrying out a multi-objective optimization.

[0145] Carrying out the optimization comprises, constrained by the received one or more constraints, optimizing, in step S12a, at least a first optimization objective, a second optimization objective, and a third optimization objective, and outputting, in step S12b, a set of solutions, wherein each solution is a list of crossings (each crossing being a crossing of two parents) of the set of candidate parents. Optionally, the optimization objectives may be optimized independently. The use of three optimization objectives is merely illustrative. More optimization objectives may be employed. The first optimization objective is a short-term genetic gain objective associated with the set of traits and the optimization comprises maximizing the short-term genetic gain. The second optimization objective is a long-term genetic gain objective and the optimization comprises maximizing the longterm genetic gain. The third optimization objective is a resource requirements objective associated with a predicted required number of progenies and the optimization comprises minimizing the resource requirements.

[0146] In other words, the optimization will be carried out and an objective function may be used to that end. Multiple solutions may be output by the optimization. Each solution is a list of crossings. The list of crossings may comprise, for each list entry, the two parents to be crossed. As will be seen below, the overall output of the method will yield, for each pair of parents / each crossing, how many progenies would probably be required, e.g., according to statistical predictions.

[0147] In step S13, the method comprises evaluating a pareto front of the set of solutions and selecting a subset of one or more solutions within the set of solutions based on a result of evaluating the pareto front.

[0148] As an example, the subset of solutions may comprise solutions at or within a predetermined distance from the pareto front of solutions of the multi-objective optimization.

[0149] Evaluating the pareto front may comprise using an algorithm, such as a search algorithm, to select a subset of one or more solutions within the set of solutions.

[0150] In step S14 a solution of the subset of solutions is selected, particularly automatically or by receiving a user input.

[0151] Specifically, the method may comprise, in step S14a, outputting, to a user, a representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the subset of one or more solutions, and receiving, in step S14b, a user input selecting a solution of the subset of solutions. Figs. 7b and 7c illustrate how solutions might be represented to allow an informed user selection.

[0152] Alternatively, the method may comprise, in step S14c, automatically ranking the solutions of the subset of solutions based on an objective function used in the multi-objective optimization and, in step S14d, automatically selecting the highest-ranked solution.

[0153] The method may comprise, in step S15, providing the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings. The list may, among other uses, also be output to a user via a display device. The list may then be used for population development as described below.

[0154] In an optional step S10, prior to step S11 , i.e., prior to carrying out the optimization, the method may comprise pre-selecting a subset of candidate parents among the set of candidate parents based on genomic information, wherein the optimization is carried out only for crossings of the subset of candidate parents. In particular, the method may comprise simulating, making use of genomic information, mean and variance of a progeny for every possible crossing. The optimization may be carried out only for a subset of crossings for which feasibility for a set of genetic gain targets has been confirmed based on the simulated mean and variance.

[0155] In other words, prior to carrying out optimization, the potential input into the optimization may be reduced by pre-selecting candidate parents using genomic information. As will be explained below, e.g. in the context of Figs 6a and 6b, this may be done using probabilities / statistics.

[0156] Fig. 2 shows another computer-implemented method 200 for decision support in selecting crossings for population development according to the present disclosure. Said method may be considered as having an iterative nature. The steps S11 to S13 (and optionally S10) may the same as in Fig. 1 and reference is made to the above description of said steps. This also applies to steps S14 and S15. The subset of one or more solutions is a first subset of one or more solutions.

[0157] In the present example, prior to selecting a solution of the subset of solutions (such as explained in the context of Fig. 1 in step S14) it is determined, in step S16, whether any of the first subset of solutions meets predetermined criteria. If this is the case, the method proceeds with step S14 as explained in the context of Fig. 1 . In case none of the first subset of solutions meets the predetermined criteria, the method proceeds with step S17.

[0158] In step S17 (step a) in the claims), the method comprises modifying the set of candidate parents, the set of traits, and / or at least one of the one or more constraints.

[0159] In subsequent step S18 (step b) in the claims), the method comprises carrying out the optimization with the modified set of candidate parents, the modified set of traits, and / or the modified constraint(s), the optimization comprising outputting a second set of solutions. This step may be carried out in the manner outlined in the context of Fig. 1 for step S12.

[0160] In subsequent step S19 (step c) in the claims), the method comprises evaluating, particularly using a search algorithm, a pareto front of the second set of solutions and selecting a second subset of one or more solutions within the second set of solutions based on a result of evaluating the pareto front. This step may be carried out in the manner outlined in the context of Fig. 1 for step S13.

[0161] In subsequent step S20a (step d1) in the claims), the method comprises outputting, to a user, a representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the second subset of one or more solutions, and receiving, in step S20b, a user input selecting a solution of the second subset of solutions or of the first and second subset of solutions. This step may be carried out in the manner outlined in the context of Fig. 1 for steps S14a and S14b.

[0162] Alternatively, in step S20c (step d2) in the claims), the method comprises automatically ranking the solutions of the second subset of solutions based on an objective function used in the multi-objective optimization and automatically selecting, in step S20d, the highest-ranked solution among the second subset of solutions or among the first and second subset of solutions. This step may be carried out in the manner outlined in the context of Fig. 1 for steps S14c and S14d.

[0163] The method may then proceed to step S15 of providing the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.

[0164] Optionally, the method may only proceed to step S20a or S20c only in case it is determined that the second subset of solutions meets the predetermined criteria in the manner described above in the context of step S16 and otherwise returning to step S17 (step a) in the claims). This may be repeated until the predetermined criteria are met, as indicated by the dashed arrows in Fig. 2.

[0165] Determining whether the solutions meet predetermined criteria may be in response to a user input and / or by automatic determination based on the optimization objective.

[0166] The method as shown in Fig. 2 and outlined above allows for further reducing trial and error in the selection of the solution and, thus, saves resources. Moreover, it is even less error prone than the method as described in Fig. 1.

[0167] Fig. 3 illustrates a method that comprises a computer-implemented method 300 for decision support in selecting crossings for population development according to the present disclosure, which method may, for example, be or comprise the method 100 or 200 outlined above, and further comprises using, in step S31 , the list of crossings and selecting, in step S32, and crossing, in step S33, the parents in the manner indicated in the list for generating a progeny population and / or providing, in step S34, resources for population development based on the list of crossings.

[0168] A system 1 according to the present disclosure is schematically illustrated in Fig. 4. Such a system may be employed in any of the methods according to the present disclosure, particularly one of the methods described in the context of Figs. 1 to 3. The system comprises a computing system 1a configured to carry out the method of the present disclosure, particularly all automatic steps thereof. The computing system may be a distributed computing system or a non-distributed computing system, for example. The system may further comprise a display device 1 b and a user input device 1c, which may be incorporated in the display device or separate from the display device. When the method comprises outputting a representation of potential crossings to a user, this may be done via the display device. When the method comprises receiving a user input, such as for selecting a solution, this may be done via the user input device. The list of crossings provided by the method may, for example, be displayed on the display device. Alternatively or in addition, it may be stored in an optional storage device 1d or transmitted to another computing system. It is noted that one or more of elements 1 b to 1d may be part of the computing system 1a.

[0169] Further features, explanations and advantages associated with the method and system of the present disclosure will be outlined below.

[0170] An overview of an exemplary method according to the present disclosure is illustrated in Fig. 5.

[0171] Input parameters (parents list, trait parameters, target progeny size, budget, constraints) may be provided. This may optionally involve user input, such as by a breeder.

[0172] Then, marker effects, accuracy, pedigree, phenotypic traits are retrieved. This may involve user input, such as by a biometrician.

[0173] Next, mean and variance for every possible cross can be simulated. Published formulae maybe used as alternative. This may optionally involve user input, such as by a biometrician.

[0174] Next, the feasibility of each cross for a set of genetic gain targets may be determined. This may optionally involve user input such as by a biometrician.

[0175] Then an optimization algorithm to evaluate potential set of solutions is carried out.

[0176] The steps of determining feasibility and carrying out an optimization algorithm may be carried out repeatedly until acceptable results are obtained. This might be determined automatically and / or involve input of a user, such as a biometrician.

[0177] An evaluation algorithm outputs and reports short- / long-term results, e.g. to a user such as a biometrician.

[0178] Optimal solutions in the pareto frontier are determined automatically, e.g. by an algorithm, and reported. This yields the subset of solutions.

[0179] A best solution may be selected, for example based on highest objective function, and may be reported to requestor, which may, for example, be a breeder or biometrician.

[0180] A final crossing list with population sizes corresponding to the solution may be output and used in the field, e.g. by a breeder.

[0181] As will be understood from the above, the choice of parents in breeding is crucial decision in a breeding program. If one is able to predict superior crosses, this allows gains in efficiency. These are concepts associated with breeding. The present disclosure allows to also take into account genomics, such as marker information permits to predict genetic variability within a population. While this allows a wealth of information there is still a huge space of possibilities for crossings and there may be many, in part conflicting parameters that need to be considered, which, in the present case, is tackled by optimization and subsequent selection steps. As such, the present method brings together breeding, genomics, and optimization to address the above-identified challenges.

[0182] In the following, some considerations on the optimization algorithm will be presented.

[0183] Figs. 6a to 6b illustrate statistical considerations on population size and genetic gain, which may, among others, be employed for preselection of parents in step S10 above. The statistical probability to generate a superior progeny from a cross can be determined by calculating the area under the curve described by the predicted mean and variance of such a cross. This also indicates the feasibility of a potential cross to contribute to the targeted short term genetic gain. In specific, Figure 6a shows the typical differences among different crosses for the mean and variance predictions. Figure 6b shows that the probability of a progeny population to meet the short term genetic gain target can be translated into the size of such population in order to achieve this target with high confidence.

[0184] It is an aim of the present disclosure to obtain crossing lists, supported by the selection support tool, that meet: Multi-trait goals (short-term); Diversity (long-term); Resource Requirements (e.g. available resources); a biologically feasible genetic gain objective.

[0185] Known approaches can provide an expected mean of a progeny from a given cross. In addition, genomics can provide the expected distribution (variance) from a given cross (Figure 6a). Further details are provided below, for example in section 1.1.

[0186] Determining the number of progenies needed to have a certain confidence that at least a certain number of progenies meet a predetermined threshold is based on probabilities (Figure 6b). Further details are provided below, for example in section 1.2.

[0187] It is noted that the list of potential bi-parental crosses is extremely high and increases significantly when increasing the number of potential crossing partners even by small numbers. Thus, it is important to make a selection, but it is also difficult to obtain the best list a priori, particularly when trying to meet many conflicting objectives, such as yield, strength, diversity, resistance to environmental factors, resources (e.g. required for the actual progenies).

[0188] As an example, in the optimization stage, yield may be maximized, yield and diversity may be maximized (as such), and yield and diversity may be maximized fora given number of progenies that can be used.

[0189] Long-term genetic gain, e.g. involves increased diversity in the crossing block, e.g., minimizing an inbreeding metric or maximizing a rare allele score.

[0190] An example with numbers is provided below and illustrated in Figs. 7a to 7c.

[0191] In this example, 65 parents are used, which means there are 2080 potential crosses.

[0192] 6 Traits are considered: Yield, Lint, Length, Strength, Mic, SFI. All traits are considered to have equal importance.

[0193] The aim is to maximize Yield, Lint, Length, Strength and to minimize SFI, and to keep MIC constant.

[0194] Some constraints / aims are selected as input, e.g. constraints concerning resources, e.g., the constraint of at most 15 entries for each population and at most 60 crosses in total. A genetic gain target is provided, e.g. a genetic gain target - 10% and 20% above the current NL-3 mean. A target is provided to obtain at least 3 progenies for each cross.

[0195] Fig. 7a illustrates the genetic relationship among candidate parents. This provides insights on the potential use of all parents or a subset of them for the objectives above. In general, very similar individuals (circles located close to each other in the graphic space) are redundant and only one individual should be used to represent that neighborhood (space in the graph). The whole space in the graph represents the whole available genetic diversity provided by the candidate parents.

[0196] Results illustrating short-term and long-term gain of potential crossings are illustrated in Fig. 7b. In this example, the following numbers are yielded and illustrated as explained below:

[0197] ■ 2080 crosses

[0198] ■ 1291 - 10% genetic gain

[0199] ■ 871 - 20% genetic gain

[0200] ■ Short-term genetic gain

[0201] ■ SPV -> X-axis (maximize)

[0202] ■ Long-term genetic gain

[0203] ■ IC -> Y-axis (minimize)

[0204] ■ RAF -> bubble size (maximize)

[0205] Fig. 7c illustrates selection of two different lists of potential crosses (subset of solutions, each list being a solution).

[0206] Each crossing list will have trade-offs, while meeting general objectives. A person or an algorithm may select a final list, i.e., the solution or list to be output. If none of the solutions are considered adequate, the method may return to setting the constraints / aims that are selected and input (outlined above) to obtain a different set of solutions.

[0207] A potential solution, i.e. crossing list alongside the predicted number or progenies for each crossing (i.e., a potential output of the method of the present disclosure), is depicted in the table below, e.g., as obtained from the above example. It shows the solution ID, the parents, and the population size for each crossing, all output from the method of the present disclosure.

[0208]

[0209] Further examples and embodiments are outlined below. Initial considerations

[0210] For thousands of years, farmers have adopted a methodological process toward the selection of agricultural plants with particular desirable characteristics (e.g. nutrition constituents, yield, shape, tolerance to environmental shifts, resistance to disease and herbicides, or increased storage / shelf life), also known as traits. Ideally, such plant specimens, when employed as progenitors for subsequent generations, result in the accumulation of desirable traits over time.

[0211] Today, plant populations are typically developed for the purpose of direct market release or as candidate parents in upcoming generations. Prior to the start of an agricultural phase, farmers are then typically required to order a selection of seeds to be planted on their fields. These seeds are developed and improved by plant breeding experts through the creation, management and evolution of various plant breeding programs known as cultivars.

[0212] Plant breeding is a rather complicated matter. On the one side, it would take a very long time for plant populations to improve significantly across multiple traits by means of mere trial-and-error approaches while, on the other side, observable trait attributes are not only induced by the genomic configuration of the plant itself, but also by the environment in which the plant is grown, and the interaction of the genotype with the environment. Significant crop improvement goals generally may be difficult to achieve without the combined expertise of research scientists and universities, large investment funds, information sharing, big data analysis, stochastic modelling and computerized optimization.

[0213] It is a goal to allow for deciding how the next population of a given crop should be developed from an (available) initial population. Uses of an advanced semi-automated software engine herein referred to as the Genomic Assisted Population Development (GAPD) is proposed. A purpose of this software is to assist breeders, farmers, and other various human decision makers, in making better decisions in the context of a plant breeding program.

[0214] The underlying process of the GAPD according to the present disclosure can be broken down into three broad steps: Firstly, raw input data are fed to the blueprints of a stochastic multi-objective combinatorial optimization model. Secondly, a multi-objective solution search metaheuristic, capable of extracting high-quality compromise solutions residing within the feasible domain space underlined by this mathematical model, is launched to explore the decision space delimited by this model. Lastly, the decision maker(s) may pick the most preferred compromise solution returned by this metaheuristic or an automated selection of a solution may be carried out, following which the corresponding solution may be communicated to the breeders and implemented.

[0215] 1. Decision making goals

[0216] Decision making problems gravitate around the attainment of pre-established goals. While there always exists room for uncertainty and potential errors in decision making, there is perhaps no bigger source of inefficiency than the inability of decision makers to identify the fundamental goals of the underlying problem. Not only does that lead to inaccurate / inefficient output results, it can also create long setbacks for all those involved in the decision making problem. In this section the aim is to provide a description of the fundamental goals featuring in the GAPD.

[0217] 1.1 On progeny population sizes and limited resources

[0218] Although progeny individuals inherit a great deal of their genotypic and phenotypic characteristics from their respective parents, due to factors such as random mutations and environmental conditions, the a priori knowledge of progeny phenotype attributes are still far from deterministic. Despite great advances in the fields of genomic and prototypic prediction, such attributes can therefore only be defined stochastically.

[0219] While progeny candidates share characteristics from both parents, the underlying target genotype can only be achieved if it actually manifests itself among the offspring population. Many studies hence suggest that sufficient amount of crosses should be conducted in order so that at least one progeny meets or exceeds these genetic gain targets with a certain probability. Practically speaking, plant breeders want to be able to increase the chances that a promising crossing combination will, a posteriori, produce at least a certain number of elite progeny individuals. Such elite individuals would then qualify as candidates to be released to the market and / or to become members of the subsequent populations. The number of progenies to be produced from a chosen crossing combination is referred to as the progeny population size.

[0220] One should note that there can certainly be other key reasons to produce multiple progenies from the same crossing combination, such as the death or malformation of progenies before they reach their full phenotype transformation state, or the opportunity to select other (relatively less important) trait characteristics not taken into account in the optimization model.

[0221] The calculation of crossing combinations population sizes is not particularly trivial. Out of the provided multi-dimensional Gaussian distribution of any given crossing combination (see also Figure 6a), the probability that a progeny meets the required genetic gain target is calculated by measuring the geometric hypervolume that falls beyond the multi-dimensional linear threshold induced through the specified target genetic gain. Theoretically speaking, this hypervolume is ideally evaluated with the use of definite integrals.

[0222] This multi-dimensional genetic gain threshold is constructed on a so-called global frame of reference. Upon these frames, the crossing combination probability distributions are fitted independently from one another by employing so-called local frames of reference. With the calculation of every crossing combination population size, the corresponding local frame of reference is generated and superposed onto the global one, from where the calculation of the hypervolume of interest ensue.

[0223] 1.2 On superior progeny values and short-term genetic gain

[0224] As mentioned earlier, plant breeders measure the increase in short-term genetic gain, from one breeding generation to the next, by observing and recording the average weighted phenotypic attributes differential across all individuals and all considered traits making up these two corresponding sequential populations of individuals.

[0225] In order to measure this gain, we employ the aggregated so-called superior progeny values of all crossing combinations featuring in a solution. First we assume that the probabilistic crossing combination trait phenotypic values are each normally distributed with means and variances provided as part of the GAPD input data, and that we define an arbitrary percentage target threshold such that any progeny displaying phenotypic trait values produced beyond it is considered to be ’’superior” (i.e. in the breeders / decision makers’ eyes) with respect to those specific traits. The crossing combination trait value delimiters on the axes of the respective normal distributions which define the area under the curve delimited by such percentage target threshold are then defined as the superior progeny values of these distributions (see also Figure 6b).

[0226] The aggregated superior progeny value of any given crossing combination is thereby obtained by multiplying each of its normalized superior progeny values by the preferential weight we assign to the corresponding trait, and then by summing up all those terms. In other words, we multiply the vector of the crossing combination’s normalized superior progeny values by the transpose of the vector of preconfigured trait preferential weights. The short-term genetic gain objective of any given solution in the domain space of the GAPD problem is then evaluated by summing up the resulting corresponding crossing combinations’ superior progeny values featured within this solution.

[0227] 1.3 On cultivar diversity and long-term genetic gain

[0228] According little importance to the population diversity factor is the right strategy to adopt only if the time span of a breeding program is that of a few years. But for most breeding programs, while shortterm genetic gain is crucial toward the improvement of a cultivar, focusing the bulk of efforts on selecting only crossing combinations that maximize the superior progeny metric value of a solution does not address the genetic diversity element that this solution brings to the breeding population.

[0229] Practically speaking, a breeding program adopting a strategy that solely focuses on the genetic performance of a subgroup of elite individuals (that share similar capacities for producing progenies with promising phenotypic improvement, and thereby usually share similar genetic material) maximizes the exploitation of specific sub-regions of the genome, while paying little to no attention to the preservation of the genomic configuration in the remaining regions. Ignoring the genetic diversity in a cultivar can be undesirable for four main reasons:

[0230] 1 . While big jumps in improvement can be observed across the desirable traits over the course of a few generations, the exploitation of only specific parts of the genome will inevitably converge after a relatively few numbers of generations, following which no further phenotypic performance improvement is possible. This is mainly due to the increasingly limited number of combinations available to configure the DNA of new population progenies.

[0231] 2. A restricted genetic pool increases the rate of inbreeding, which increases the redistribution of alleles from heterozygous to homozygous states. This can lead to the increased expression of deleterious recessive genes and, consequently, to reduced genetic gain performance altogether.

[0232] 3. While the measure of what “good performance” represents is unlikely to change too much in the short-term, that is however not a certainty in the long-term. In particular, the trait phenotype characteristics deemed worthy at present, in addition to the respective levels of desirability that we allocate to each of those traits, may very well shift in time as the market evolves and farming strategies adapt thereof.

[0233] 4. Lastly, a diverse cultivar provides more flexibility to recalibrate breeding strategies in the advent of new or changing environments. More specifically, because phenotypic output is not only dependent on genomic configuration, but also on the environment in which the crop is grown, individuals displaying the highest phenotypic performance with respect to the environments in which they are currently harvested might no longer strive under the influences of new, potentially very different farming regions, or as a result of significant shifts in environmental conditions due to climate change in existing farming regions.

[0234] 1.4 On the implementation of secondary traits

[0235] However, due to good leaps in the quality and efficiency of the GAPD search algorithm, it is important to also give importance to the presence of secondary traits in the chosen solutions. Because the shortterm genetic gain objective function is constructed on a fusion of multiple primary traits, it would not be wise to dilute the importance of those traits by further incorporating secondary traits inside that function, or to over-complicate the function with so many different elements.

[0236] Instead, secondary traits are incorporated as hard constraint in the optimization model formulation, with the option of having the full flexibility to relax these constraints whenever their strictness level too negatively impacts the performance of the algorithm.

[0237] 1.5 On solution sizes

[0238] In the GAPD, the solution size simply refers to the number of crossing combinations featuring in a given solution. Solution sizes being flexible in the GAPD, their a priori parametric configuration plays a major role in the optimization procedure. In line with significantly increasing the size of the domain space (in contrast to using a fixed solution size), having multiple solution sizes allows us to uncover more interesting, broader, and better spread non-dominated fronts. We also rightfully observe that larger solutions sizes are heavily correlated to higher solution diversity and operating costs, but inversely correlated to short-term genetic gain.

[0239] 1.6 On the simultaneous optimization of multiple goals

[0240] In optimization modelling, one of the biggest area for potential failure evolves around the misinterpretation of the quality assessment of solutions. This is because solution quality criteria do not necessarily adhere to what seems the most logical to the modeler, but can also rely heavily on confusing subjective preferences conjured by decision makers. A good balance between these two paradigms is therefore wise. The optimization process in the GAPD problem focuses on running multiple so-called short-term genetic gain target trials (simply referred to in these notes as trials), each made up of a unique domain space and feasible domain space (defined by both the chosen trial-dependent subset sizes and the hard constraints induced through the trial-dependent crossing combination population sizes). In each trial, we seek to find solutions in which the SPV, minor allele frequency and inbreeding coefficient metrics are simultaneously optimized. More specifically, each trial returns a non-dominated front in biobjective space, where the SPV objective is optimized on the one side, and the combined minor allele frequency and inbreeding coefficient sub-objectives are optimized on the other side. Each nondominated front is hence made up of several solutions providing the decision makers with various trade-off alternatives between short-term and long-term genetic gain.

[0241] 2 Mathematical modelling

[0242] 2.1 Subset size upper bound

[0243] As mentioned earlier, there exists a total of ^N(N - 1) possible crossing combinations in any GAPD problem, where N is the number of parent individuals. We have also established that (1 ) each crossing combination ) is assigned a meticulously derived population size for each trial ~g, and

[0244] (2) the total number of progenies produced in any solution of a GAPD problem (i.e. the sum of the selected crossing combinations’ population sizes) are bounded above by the cost / budget constraint value C. We are therefore able to state that there exists a maximum number of crossing combinations

[0245] 2 < Kg - 1) beyond which there exists no feasible solutions (i.e. such that all possible solutions of size K + 1 are infeasible).

[0246] It is useful to establish this bound for use in the strategic configuration of the GAPD optimization algorithms later on. This can simply be achieved by first sorting the population sizes in ascending order, then by adding the population sizes one at a time from the top until the cost constraint is violated.

[0247] Mathematically speaking, if we denote to be the ordered rank r of the population size of Crossing Combination ) in trial ~g, then Kg is the value for which the inequality hold true.

[0248] 2.2 Input data and parametric configuration

[0249] Consider a set of candidate parent individuals !P = {Pt, PN} forming the (initial) cultivar population of a GAPD problem, with candidate crossing combinations C = {c12, c13, and define the set of secondary traits S = {st, . . . , ss}. Let Usrepresent the combined set of unique (categorical) secondary traits values assigned to all the crossing combinations in C, where S G S. Moreover, define the Inclusion Set I to force certain crossing combinations into any solution, and inversely define the Exclusion Set s to prevent certain crossing combinations from featuring in any solution.

[0250] Next, let the parameter fli7G N represent the progeny population size of Crossing Combination ci7and let vi7represent the superior progeny value of ci7, where i,j G {1, . . . , N}. Due to limited amount of resources, we also define the parameter B > 0 as the maximum total number of progenies (i.e. operating costs) produced across all selected crossing combinations.

[0251] Thereafter, define the so-called G-Matrix G as a symmetrical N by N matrix whose entries depict the level of genetic similarity between parent Individuals i,j G {1, ... , / V], where a higher value indicates more similarity. Although it might appear counter-intuitive at first, it is good to note that, for practical reasons, the diagonal entries in this matrix can also slightly vary from one another (even though those are typically the highest).

[0252] More advanced perhaps, we also define the user parameter a 0 to factor in the penalty given to parent individuals that feature multiple times in a given solution, and the user parameter / 3 1 to favour solutions that contain a higher numberof unique / distinct parents (i.e. the total number of parents used at least once in the solution). In other words solutions containing a lot of unique parents, and few parents that are overused, will tend to be favored in terms of their genetic diversity.

[0253] Moving on, define the set of all secondary traits constraint objects A = {ht, . . . , h|AJ, and define the linking parameter if Secondary Trait Constraint h is associated to Secondary Trait s, otherwise, for all h G A and all s e S. This parametric configuration of course requires both that |A| > |S|, and that for all h G A.

[0254] Hereafter, let 6hG USrepresent the value assessed in Constraint h G A for the corresponding Secondary Trait S G S, and let $iJSrepresent the value assigned to Crossing Combination ci7for Secondary Trait S G S. We then construct the parameter Finally let G (0, 1] depict the minimum and maximum proportion (respectively) of crossing combinations in a given solution required to match the value 9hin Secondary T rait Constraint h G A, where of course

[0255] 2.3 Decision variables

[0256] In the GAPD problem, the domain space is delimited by the set of decision variables

[0257] 1, if parent Individual i is selected to breed with parent Individual j, 0, otherwise identical to x.f, and

[0258] 1, if parent Individual i is selected at least once in the solution, 0, otherwise

[0259] In addition, for ease of representation in the model formulation later on, we derive the solution size as the total number of unique parents featuring in the solution as and the overall number of times that parent Individual i G J3employed in the solution as

[0260] 2.4 Objective functions

[0261] 2.4.1 Short-term genetic gain (Max) 2.4.2 Option 1 : Inbreeding coefficient (Min)

[0262] In the optimization model you will work with, the diversity component will solely be evaluated by a so- called inbreeding coefficient metric (without much loss of generality with regard to the scope of your thesis). There exist additional metrics that can be combined in the objective function (such as the minor alleles frequency metric) to complement the inbreeding coefficient metric, but we will not consider those at this point.

[0263] While there exists many mathematical variations in the literature behind the inbreeding coefficient metric, the described herein standalone version aims to meet the following two conditions / axioms in all conceivable scenarios:

[0264] I) If (1) all non-diagonal entries in the G-Matrix have the same values, (2) all diagonal entries in the G-Matrix have the same values, (3) Solution B contains more unique parent individuals than Solution A, regardless of the number of crossing combinations in either solution (i.e. solution sizes), and (4) or is set to 0 if Solution B has a higher size than Solution A, then Solution B must have a smaller / better inbreeding coefficient score than Solution A. The only exception to this condition occurs when / 3 = 1 , in which case do both solutions have the same score.

[0265] II) If Solution B contains more crossing combinations than Solution A, but both solutions share the same set of unique parents, and a is set to 0, then the inbreeding coefficient score of Solution B must be identical to that of Solution A.

[0266] We henceforth calculate the inbreeding coefficient of solution x with the function

[0267] 2.4.3 Option 2: Inbreeding coefficient Hybrid (min)

[0268] N-l N z z XQ (0.5Gjj + 0.25 G a + 0.25Gyy) t=i j=i+i

[0269] 2.4.4 Minor alleles frequencies metric (Max)

[0270] Consider again an initial GAPD population of N individuals {P1;.... Pw}, and recall that 8imG {0, 1, 2} represents the known diploid value of Individual i G {1, ... , / V} at Marker m G {1, ... , / V}. Although these diploid values are actually part of the raw input data, it is nevertheless useful to mention how they were numerically derived from their original allele values in case only such values are provided in certain GAPD scenarios.

[0271] Basically, every marker position of any given individual carries a combination of two alleles, say Ai and A2, such that there exists a total of three configurations, namely (Ai,Ai), (AI, A2), (or (A2,Ai)), and (A2, A2). For any given marker, the first step then consists in counting how many A1 and A2 alleles are there respectively across the population of initial individuals. We then observe which count is the smaller (if not a tie). This identifies the less frequently occurring allele; i.e. the minor allele.

[0272] As such, each individual entry is assigned a numerical score for that marker of either one of {0, 1, 2}, based on if and how many minor alleles this individual possesses for that particular marker. For example, suppose that the minor allele at Marker m is A2. Then, the minor allele score of Individual i at that marker is assigned a value of 8im= 0 if mm= (AltAr~), 8im= 1, if mm= (AltA2) or mm= (A2, AI), and 8im= 2 if mm= (A2,A2). The interested reader is referred to the R package BCS. Generics for a full derivation of this procedure.

[0273] The so-called minor allele frequency of Marker m is henceforth calculated as where it should be noted that < / >monly takes on extreme values 0 or 2 if (respectively): (1) the entire population carries the same allele for that specific marker (while the other allele does not feature at all), consequently also pointing out its diversity saturation; and (2) the allele types are exactly evenly spread out across all individuals. As seen above, the assigned value to either of these cases is, by default, set to 1 so that, by definition of the logarithmic function, its contribution toward the minor allele frequency long-term genetic gain score will (appropriately) be zero.

[0274] We are thereafter interested in evaluating the contribution of minor / rare alleles that each individual packs within itself. Unlike more traditional minor allele contribution metrics, which consider additive linear preference over all possible allele rarity levels (e.g. De Beukelaer 2017).

[0275] 2.4.5 Operating costs (Min)

[0276] N-l N z i=l: j s=i+ixijUij

[0277] 2.5 Hard constraints

[0278] 2.5.1 Inclusion set

[0279] 2.5.2 Exclusion set

[0280] 2.5.3 Operating costs

[0281] 2.5.4 Secondary traits

[0282] 2.6 Diversity sub-objective metrics bounds

[0283] While the diversity objective function may very well be assessed through the original numerical scales whenever either one of the minor alleles frequency or inbreeding coefficient metric is adopted in the model formulation, this is certainly not the case in instances where both these metrics are simultaneously adopted. Clearly this is because the objective space ranges of these metrics (mapped from the search space of a given GAPD scenario) are made up of very different sets of values. When both metrics are employed simultaneously, the diversity score of any given solution can therefore not be assessed by merely adding the two metric scores multiplied by their respective user-defined preferential weights. This would result in one of the sub-objective being significantly favored over the other. And this independently of the assigned preferential weights, which should only be configured in respect of the preference that we give to each sub-objective, never as a means to “fairly” compensate for these variations in sub-objective ranges.

[0284] The reason we normalize a metric is hence simply to map the entire spectrum of values spanning a sub-objective range in IR onto the interval [0, 1] or (0, 1), thereby allowing a much “fairer” objective assessment of the two (or more) combined sub-objectives. In order to do this, we first need to find adequate numerical lower and upper bounds for each diversity metric and for each trial. For the minor alleles frequency sub-objective in trial g, we (ideally) need to find the lower and upper bound values (where Sg is the search space of trial g), and

[0285] Similarly, we define L[<5^] and I / ^] as the lower and upper bound values of the inbreeding coefficient sub-objective in trial g.

[0286] 2.7 Numerical examples

[0287] 2.7.1 Secondary traits constraints

[0288] The implementation of secondary traits constraints is perhaps also a bit tricky. Since the topic will strongly evolve around those, It may therefore be useful to clarify the abstract notation in the previous section through the use of a simple hypothetical example.

[0289] Consider again a cultivar set J> = {Pi, P2, P3, P4} with corresponding crossing combinations set combinations C = {c12, c13,c14, c23, c24. c34}, and assume that two secondary traits S = {s1,s2} and secondary traits constraints A = {h1,h2,h3} are in use.

[0290] Given the secondary traits constraints input data provided in Figs. 9 and 10, we can then extract the following parameters:

[0291] Tl1= {" A" B" C" D"} and U2= {"(0,1", "(0,2)", "(2,1)"}

[0292] (respectively). 2.8 Optimization model formulation

[0293] Based on everything discussed up to this point, the aim in this GAPD optimization problem is therefore to maximise maximise minimise such that xa= 0 i G J>, (9)

[0294] Yi 6 {0,1} i e P. (10) 3 Post-optimization solution selection

[0295] As mentioned earlier, when dealing with multi-objective optimization models, the output of the search algorithm consists in not one, but rather multiple high-quality solutions (this set is referred to in the literature as the non-dominated front). Each of these solutions embody different trade-off compromises across the objective functions. Therefore, no solution is objectively better than any others within that output set. Out of this set, the decision maker is then required to select the solution with the objective function score compromises that best matches his / her preferences.

[0296] Because the GAPD framework can be completely automated, however, there can be no human-in- the-loop intervention at any point in the entire framework (i.e. anywhere between entering the raw data objects at the very beginning, to analyzing the returned solution at the very end). This means that we had to develop an approach that allows the framework to pick a single solution from the non-dominated front returned by the search algorithm.

[0297] The approach developed by the inventors and described herein is referred to as post-optimization weighted multi-objective scalarization. It combines a method of normalizing the objective function values and multiplying them by numerical weights that are representative of the decision maker’s preferences. This approach is not only efficient at selecting the best solution from a GAPD nondominated front, but also relatively easily implementable into the framework.

[0298] 3.1 Abstract description

[0299] Suppose that the non-dominated front is given by the set of solution N = {X1,X2, ...,XN], and let the preferential weights of the decision maker be given by the values {w1,w2,w3}, where each of these weights correspond to the GAPD objective functions depicted in the optimization model in Section 2.8 (namely short-term genetic gain, genetic diversity and operational costs).

[0300] Next, let [Oi] and [OiJ represent the upper and lower bound values of the first objective function of all solutions in N. Similarly, let [O2] [O2J, [O3I and [O3J respectively represent the upper and lower bound values of the second and third objective functions of all solutions in N. Furthermore, let

[0301] °2(xi') and O3(xt) represent the first, second, and third objective scores (respectively) of solution xt G N.

[0302] The so-called scalarized objective value of solution x; is then calculated as where it must be noted that the first two terms are added due to the corresponding objective functions being maximized, while the third term is subtracted as a result of the corresponding objective function being minimised.

[0303] Each solution N is then ranked according to their scalarized objective values. The GAPD framework will therefore pick the solution which possess the highest rank.

[0304] 3.2 Numerical example

[0305] To illustrate the principles described in the section above, consider a (hypothetical) non-dominated front output of a GAPD optimization process consisting of 5 solutions, along with their respective objective function values, laid out in Table 1.

[0306] Table 1: Hypothetical non-dominated front of a GAPD optimization problem consisting of five solutions, along with their respective objective function values. Next, suppose that the decision maker’s preferences with regard to all 3 objective functions is reflected by means of the weights w1= 2,wz= 1.6, w3= 1.1 (i.e. in this case the decision maker cares the most about short-term genetic gain, and second most about genetic diversity). The scalarized objective value of, for example, Solution xi, is then calculated as

[0307] Similarly, we can derive the scalarized objective values of the other 4 solutions in A / , as shown in Table 2. The so-called preferential ranks of these solutions are also provided in this table. Looking at the results, it follows that Solution X3 is the best. This solution will therefore be reported back to the decision maker as part of the GAPD framework output.

[0308] Table 2: Scalarized objective values and preferential ranks of the non-dominated front solutions. 4 Alternative GAPD multi-objective optimization methods

[0309] In the GAPD multi-objective optimization model, we employ an approach that aims to find the Pareto front out of the search space of any given problem instance using a metaheuristic search engine. As discussed previously, we refer to the non-dominated front as an approximation of the Pareto front. The more complex the problem instance, the less likely it is that the non-dominated front will converge exactly to the Pareto front. Nevertheless, we have observed that the optimization approach still produces non-dominated fronts of excellent quality, even when dealing with GAPD problem instances of high-complexity. However, other approaches to solve the optimization problems are also conceivable.

[0310] This approach is just one of many that we could employ to solve GAPD multi-objective optimization problems. Examples of other potential techniques include:

[0311] Hypervolume Chebyshev scalarization (e.g. https: / / arxiv.org / pdf / 2006.04655.pdf)

[0312] Mathematical programming (e.g. https: / / link.springer.com / journal / 10107)

[0313] Reactive search Optimization (e.g. https: / / core.ac.uk / download / pdf / 11829614.pdf)

[0314] Directed search domain (e.g. https: / / www.tandfonline.com / doi / abs / 10.1080 / 0305215X.2010.497185)

[0315] Branch and bound (e.g. https: / / www.sciencedirect.com / science / article / abs / pii / S037722171730067X)

[0316] Indirect / Self organisation Optimization (e.g. https: / / www.iosotech.com / pab7.htm)

[0317] Ziont-Wallenius method (e.g. https: / / www.sciencedirect.com / science / article / pii / S0307904X11000849) While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered exemplary and not restrictive.

[0318] The invention is not limited to the disclosed embodiments. In view of the foregoing description and drawings it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention, as defined by the claims.

Claims

CLAIMS1. A computer-implemented method (100, 200) for decision support in selecting crossings for population development, the method comprising: receiving (S11) a set of candidate parents, a set of traits, and one or more constraints comprising a resource constraint representative of a maximum number of crossings and / or a maximum progeny population size; carrying out a multi-objective optimization (S12), wherein carrying out the optimization comprises, constrained by the received one or more constraints, optimizing (S12a) at least a first optimization objective, a second optimization objective, and a third optimization objective, and outputting (S 12b) a set of solutions, wherein each solution is a list of crossings of parents of the set of candidate parents, wherein the first optimization objective is a short-term genetic gain objective associated with the set of traits and the optimization comprises maximizing the short-term genetic gain, wherein the second optimization objective is a long-term genetic gain objective and the optimization comprises maximizing the long-term genetic gain, and wherein the third optimization objective is a resource requirements objective associated with a predicted required number of progenies and the optimization comprises minimizing the resource requirements; evaluating (S13) a pareto front of the set of solutions and selecting a subset of one or more solutions within the set of solutions based on a result of evaluating the pareto front; outputting (S14a), to a user, a representation of the potential crossings, their associated short-term genetic gain and long-term genetic gain, and a representation of the subset of one or more solutions, and receiving (S14b) a user input selecting a solution of the subset of solutions, or automatically ranking (S14c) the solutions of the subset of solutions based on an objective function used in the multi-objective optimization and automatically selecting (S14d) the highest-ranked solution; providing (S15) the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.

2. The method of claim 1 , wherein evaluating the pareto front comprises using an algorithm, such as a search algorithm, to select a subset of one or more solutions within the set of solutions.

3. The method of claim 1 or 2, comprising, prior to carrying out the optimization, pre-selecting (S10) a subset of candidate parents among the set of candidate parents based on genomic information, wherein the optimization is carried out only for crossings of the subset of candidate parents, the method particularly comprising simulating, making use of genomic information, mean and variance of a progeny for every possible crossing, wherein the optimization is carried out only for a subset of crossings for which feasibility for a set of genetic gain targets has been confirmed based on the simulated mean and variance.

4. The method of any of the preceding claims, wherein the short-term genetic gain and / or the long-term genetic gain is determined making use of genomic information.

5. The method of any of the preceding claims, wherein the short-term genetic gain is determined based on, e.g. stochastically, predicted superior progeny values for each of the crossings, in particularwherein the superior progeny value for the respective crossing is an aggregated value aggregated over the set of traits with optional weighting of the traits.

6. The method of any of the preceding claims, wherein the long-term genetic gain is representative of a genetic diversity, wherein the genetic diversity is expressed by a minor allele frequency score and / or an inbreeding coefficient.

7. The method of claim 6, wherein the inbreeding coefficient of a solution depends on at least one of a number of times a parent is used for crossing in the solution, a number of unique parents used for crossing in the solution, and the similarity of the parents of the respective crossing.

8. The method of claim 6 or 7, wherein the long-term genetic gain objective is based on both, the inbreeding coefficient and the minor allele frequency score, optionally the method comprising normalization of the inbreeding coefficient and the minor allele frequency score and combining the normalized inbreeding coefficient and minor allele frequency score into a single optimization objective, or wherein the long-term genetic gain objective is based on only one of the inbreeding coefficient and the minor allele frequency score.

9. The method of any of the preceding claims, wherein the subset of one or more solutions is a first subset of one or more solutions and wherein the method comprises, determining (S16) prior to selecting a solution of the subset of solutions, particularly in response to a user input and / or by automatic determination based on the optimization objective, whether any of the first subset of solutions meets predetermined criteria and, in case none of the first subset of solutions meets > predetermined criteria: a) modifying (S17) the set of candidate parents, the set of traits, and / or at least one of the one or more constraints; b) carrying (S18) out the optimization with the modified set of candidate parents, the modified set of traits, and / or the modified constraint(s), the optimization comprising outputting a second set of solutions; c) evaluating (S19), particularly using a search algorithm, a pareto front of the second set of solutions and selecting a second subset of one or more solutions within the second set of solutions based on a result of evaluating the pareto front; d1) outputting (S20a), to a user, a representation of the potential crossings, their associated shortterm genetic gain and long-term genetic gain, and a representation of the second subset of one or more solutions, and receiving (S20b) a user input selecting a solution of the second subset of solutions or of the first and second subset of solutions, or d2) automatically ranking (S20c) the solutions of the second subset of solutions based on an objective function used in the multi-objective optimization and automatically selecting (S20d) the highest-ranked solution among the second subset of solutions or among the first and second subset of solutions; providing (S15) the list of crossings of the selected solution and a predicted required number of progenies for each of the crossings comprised in the list of crossings.

10. The method of any of the preceding claims, the constraints comprising constraints related to the inclusion and / or exclusion of specific crosses and / or constraints related to achieving a maximum value for one or more predetermined traits, and / or constraints related to achieving a constant value for one or more predetermined traits, and / or constraints related to achieving a minimum value for one or more predetermined traits, and / or constraints related to a predetermined proportion of crosses with a desirable characteristic that should make up the final crossing list.

11. The method of any of the preceding claims, wherein the subset of solutions comprises solutions at or within a predetermined distance from the pareto front of solutions of the multi-objective optimization.

12. The method of any of the preceding claims, wherein the representation of the potential crossings and their associated short-term genetic gain and long-term genetic gain comprises representing each crossing within a two-dimensional coordinate system, the coordinate system having a first axis representative of the short-term genetic gain and a second axis representative of a long-term genetic gain, and / or wherein the representation of the subset of solutions comprises, for each solution of the subset, an indication of a group of crossings associated with the solutions, particularly the group being represented in the coordinate system such as by heatmaps or shapes enclosing the crossings of the group of crossings or other highlights highlighting crossings of the group of crossings, wherein, optionally, the representation of the potential crossings, and optionally their associated diversity, further comprises representing each crossing as a geometric shape, such as a circle, within the coordinate system, the size of the geometric shape being correlated with a first metric of the longterm genetic gain and the second axis value is correlated with a second metric of the long-term genetic gain, in particular wherein the first metric is associated with one of an inbreeding coefficient and a minor allele score, and the second metric is associated with the other one of the inbreeding coefficient and the minor allele score.

13. A method (300) comprising the method (100, 200) of any of the preceding claims, further comprising using (S31) the list of crossings and selecting (S32) and crossing (S33) the parents in the manner indicated in the list for generating a progeny population and / or providing resources for population development based on a / the list of crossings.

14. A system (1) comprising a computing system (1a) configured to carry out the method of any of claims 1 to 12.

15. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 12.

16. A computer-readable medium comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 12.

Citation Information

Patent Citations

  • Methods for increasing genetic gain in a breeding population

    US20120151625A1

  • Improved computer implemented method for breeding scheme testing

    US20190172548A1

  • System and method for genomic prediction

    US20230290438A1