METHODS AND SYSTEMS FOR USE IN THE IMPLEMENTATION OF RESOURCES IN PLANT BREEDING
Patent Information
- Application Number
- MX2021011729
- Authority / Receiving Office
- MX · MX
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-28
- Filing Date
- 2021-09-24
- Publication Date
- 2026-06-12
- Estimated Expiration
- 2040-03-27
AI Technical Summary
Existing resource allocation methods in plant breeding do not account for phenotypic and genotypic variations among plant origins, leading to inefficient distribution and potential underutilization of resources.
A method and system that utilize an algorithm to allocate resources based on trait performance probabilities, risk, and diversity, optimizing resource distribution by focusing on origins with higher probabilities of producing high-performing and genetically appropriate progeny.
Enhances the efficiency of resource utilization in plant breeding by directing resources to origins with higher potential for desired traits and genetic diversity, improving the overall breeding process.
Abstract
Description
METHODS AND SYSTEMS FOR USE IN THE IMPLEMENTATION OF RESOURCES IN PLANT BREEDING iviA / a / zuz ι / uii rza CROSS REFERENCE TO RELATED APPLICATION This application claims the benefit and priority of U.S. Provisional Application No. 62 / 825.513, filed on March 28, 2019. The entire disclosure of the foregoing application is incorporated herein by reference. FIELD OF INVENTION This disclosure generally refers to methods and systems for use in the implementation of resources in plant breeding and, in particular, to methods and systems for use in the allocation of resources in plant breeding contexts where the allocation is based on the performance and / or genetic distributions of the origins. BACKGROUND This section provides background information related to this disclosure, which is not necessarily the state of the art. In plant development, modifications are often made to plants through either selective breeding or genetic manipulation. Depending on the specific selection or manipulation, the resulting plant material is introduced into a breeding line, where the plants are then created, cultivated, and evaluated. When the plants' performance reaches an expected threshold, exceeds it, or reaches its maximum performance—for example, for a given phenotype—or when the frequencies of genotypes are at or above a certain threshold, the plants may be considered target plants for further development and / or commercial implementation. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described in this document are for illustrative purposes only, show selected modalities, are not all possible implementations, and are not intended to limit the scope of this disclosure. FIG. 1 is an exemplary system of the present disclosure suitable for allocating resources within plant breeding lines, based, at least in part, on phenotypic and / or genotypic information; FIG. 2 is an exemplary graph of trait performance probability distributions for multiple pairs of origins, and which form a basis for resource allocation in the system of FIG. 1; FIG. 3 is a block diagram of an exemplary computer device that can be used in the system of FIG. 1; and FIG. 4 is an exemplary method, suitable for use with the system of FIG. 1, for the allocation of resources within plant breeding lines on the basis, at least in part, of phenotypic and / or genotypic information. The corresponding reference numbers indicate corresponding parts in the various views of the drawings. DETAILED DESCRIPTION Exemplary modalities will be described in more detail below with reference to the accompanying drawings. The descriptions and specific examples included herein are for illustrative purposes only and are not intended to limit the scope of this disclosure. Various breeding techniques are commonly used in the agricultural industry to produce desired plants. Each technique, and each associated process, requires resources to create, cultivate, or evaluate plant material. Some of these resources, included in a plant breeding program, include, but are not limited to, land such as field rows and plots, greenhouse space, genotyping laboratory units, and duplicated haploid units (DHUs). For example, when a number of origins are selected for duplicated haploid (DH) production, the capacity for that production is required, which is determined by available fields, laboratories, manpower, funding, and other resources.or other resources, to carry out this process, which can be divided into individual units, in this case, DHUs, and then distributed evenly among the selected origins. For example, if 200 origins are selected and 1000 DHUs are available, and the DHU resources are divided among them, each origin is allocated 5 DHUs. However, this even distribution does not account for any variation in potential value or potential genetic / phenotypic variation within the different origins. Uniquely, the methods and systems described herein allocate resources within a breeding line based on one or more phenotypic and / or genotypic characteristics of the origins. Specifically, a decision engine employs an algorithm that considers the performance probabilities of traits for the origins (e.g., expressed as a binomial distribution, etc.), as well as the risk and / or genotypic components and / or diversity associated with the selected origin cohort. The variation in the potential value of specific origins can also be predicted by simulating genetic / phenotypic variation.Through this algorithm, the resources available for the breeding process are allocated among the origins, with more resources dedicated to origins with a higher probability of producing progeny that perform above one or more thresholds, and / or with a higher probability of producing progeny that express certain genetic components at rates considered appropriate and / or desired for the breeding line. In this way, the breeding line is improved (as a practical application of the methods and systems of the present invention) by allocating resources more efficiently in order to produce high-performing and / or more genetically appropriate progeny. That said, progenies are generally organisms descended from crosses between one or more parental organisms of the same species—that is, origins. Progenies can refer, for example, to the universe of all possible progenies of a particular breeding program, a subset of all possible progenies specific to one or more origins, all the offspring of an origin in a given generation, certain offspring of an origin, and so on. Furthermore, as used herein, the term origin refers to the set of progeny parents and is therefore interpreted as singular or plural, as appropriate. Phenotypic data, trait distribution, ancestry, genetic sequence, commercial success, and other information about the progenies are either known or can be simulated and can be stored in the memory described herein. Phenotypic data as used in this document include, without limitation, information regarding the phenotype of a given progeny (e.g., a plant, etc.) or a population of progenies (e.g., a group of plants, etc.). Phenotypic data may include the size and / or vigor of the progeny (e.g., plant height, stem circumference, stem strength, etc.), yield, time to maturity, resistance to biotic stress (e.g., resistance to diseases or pests, etc.), resistance to abiotic stress (e.g., drought or salinity, etc.), growing climate, or any additional phenotypes and / or combinations thereof. It should be appreciated that the methods and systems in this document generally involve phenotypic data associated with one or more origins, progenies, etc., and related phenotypic variations. That said, it should be noted that genotypic data may be used in place of, or in connection with, or in combination with the phenotypic data described herein (for example, to further supplement the phenotypic data, and / or to further inform the models, algorithms, and / or predictions in this document, etc.), in one or more exemplary implementations, to assist in the selection of progeny groups and / or in the identification of progeny sets consistent with the description in this document.This can take the form of using an algorithm, for example, to predict phenotypic values and / or variances for a given cross, based on known or simulated genotypic data associated with that cross. Figure 1 illustrates an exemplary system 100 for allocating resources within plant breeding lines, based, at least in part, on known or simulated phenotypic and / or genotypic information, and in which one or more aspects of this disclosure may be implemented. Although, in the described modality, the parts of system 100 are presented in one arrangement, other modalities may include the same or different parts arranged differently according to, for example, the resources available for allocation to progenies, the number of origins, particular types of origins, particular types of progenies, genotypes of interest, and / or phenotypes of interest, etc. As shown in FIG. 1, the 100 system generally includes a breeding line 102, which is provided to advance origins, progenies, etc., through testing and selection, to further development and / or commercial use. The breeding line 102 generally defines a pyramidal progression, whereby a large number of potential origins are entered and then successively reduced (e.g., by selection reduction, etc.) to a preferred or desired number of origins, progenies, or plants.While breeding line 102 is configured to allocate resources within it, as provided in this document, breeding line 102 may be configured to employ one or more of other techniques that may include a wide range of methods known in the art to create, select, or advance origins or progenies within breeding line 102, often in accordance with the particular plant and / or organism for which breeding line 102 is provided. In certain breeder line configurations (e.g., large industrial breeder lines, etc.), testing, selection, and / or advancement decisions may be directed at hundreds, thousands, or more of origins, progenies, etc., in multiple phases and at various locations over several years, to arrive at a reduced set of origins, progenies, etc., which are then selected for the development of commercial products. In short, the illustrated breeder line 102 is configured, through the evaluations, selections, etc., included within it, to reduce a large number of origins, progenies, etc., to a relatively small number of commercially viable, high-performing products. In this exemplary form, breeding line 102 can be described with reference to, and is generally directed towards, maize and its traits and / or characteristics. However, it should be noted that the systems and methods described herein are not limited to maize and can be employed in a breeding line or program related to other plants, for example, to improve fruits, vegetables, grasses, trees, or ornamental crops, including, without limitation, maize (Zea mays), soybean (Glycine max), cotton (Gossypium hirsutum), peanut (Arachis hypogaea), barley (Hordeum vulgare), oats (Avena sativa), bermudagrass (Dactylis glomerata), and rice (Oryza sativa, including indica and japonica varieties). sorghum {Sorghum bicolor)·, sugar cane (genus Saccharum)·, tall fescue {Festuca arundinacea)·, grass species (e.g., species: Agrostis stolonifera, Poa pratensis, Stenotaphrum secundatum, etc.); Wheat (Triticum aestivum) and alfalfa (Medicago sativa), members of the genus Brassica, which include broccoli, cabbage, cauliflower, cantaloupe and rapeseed, carrot, Chinese cabbage, cucumber, dried beans, eggplant, fennel, green beans, squash, leek, lettuce, melon, okra, onion, peas, pepper, pumpkin, radish, spinach, zucchini, sweetcorn, tomato, watermelon, honeydew melon, cantaloupe and other melons, banana, castor bean, coconut, coffee, cucumber, poplar, southern pine, radiata pine, Douglas fir, eucalyptus, apple and other tree species, orange, grapefruit, lemon, lime and other citrus fruits, clover, flaxseed, olive, palm, Capsicum, Piper and Pimenta peppers, sugar beet, sunflower, sweetgum, tea, Tobacco and other fruits, vegetables, tubers, and root crops. The methods and systems of the present invention can also be used with non-cultivated species, especially those used as model methods and / or systems, such as Arabidopsis.Furthermore, the methods and systems described in this document can be employed beyond plants, for example, for use in animal breeding programs or other breeding programs that are not for plants and / or crops. As shown in FIG. 1, the breeding line 102 includes an origin start phase 104 and a cultivation and testing phase 106, which together identify and / or select one or more origins or progenies to advance to a validation phase 108. In the validation phase 108, the progenies are then introduced into pre-commercial trials as progenies, lines, or hybrids, for example, according to the particular type of progeny, or other suitable processes (e.g., a characterization and / or commercial development phase, etc.) with the ultimate goal of planting and / or commercializing the progenies. It should be noted that the breeding line 102 can include a variety of conventional processes known to those skilled in the art in the three different phases 104, 106, and 108 illustrated in FIG. 1. In the origin initiation phase 104, a set of potential origins is narrowed down to a selected set of origins, for example, based on origin selection systems and / or based (at least in part) on the methods and systems disclosed in Applicant's jointly owned U.S. Patent Application 15 / 618.023, entitled “Methods for Identifying Crosses for Use in Plant Breeding,” the full description of which is incorporated herein by reference. It should be appreciated that other selection techniques may be employed to select origins in the origin initiation phase 104, which may be supported by a variety of origin-related data and / or origin predictions, etc. Once the origins have been selected, they proceed to the cultivation and evaluation phase 106, in which the progenies are planted or otherwise introduced into one or more cultivation spaces, such as greenhouses, shade houses, nurseries, growing plots, fields (or test fields), etc. As should be understood, the cultivation and testing phase 106 includes a number of resources for cultivating and evaluating the progenies of the selected origins. These resources may include, for example, double haploid units (DHUs), which are the resources required to cultivate and evaluate the progeny of the origins. It should be noted that other resources may be included in the cultivation and evaluation phase 106, which are subject to the techniques explained herein.Here, the resources within the cultivation and testing phase 106 are, in general, allocated by an allocation engine 110, to the source pairs identified in the selected sources, as described below. Once the progenies grow into the culture and evaluation phase 106, each is evaluated (again as part of the culture and evaluation phase 106 in this example) to derive and / or collect phenotypic and / or genotypic data for the progeny, and the phenotypic and / or genotypic data are stored in one or more data structures. Common examples of phenotypes that can be evaluated by such tests include, but are not limited to, disease resistance, abiotic stress resistance, yield, seed and / or flower color, moisture content, size, shape, surface area, volume, mass, and / or quantity of chemical substances in at least one seed tissue, for example, anthocyanins, proteins, lipids, carbohydrates, etc., in the embryo, endosperm, or other seed tissues. For example, when a progeny (e.g., grown from a seed, etc.) has been selected or otherwise modified to produce a particular chemical substance (e.g.(a pharmaceutical product, a toxin, a fragrance, etc.), the progeny can be analyzed to quantify the desired chemical substance. When the progeny is deemed successful, based on phenotypic and / or genotypic data and a variety of thresholds and / or bases, it advances to validation phase 108. In this phase, the progeny undergoes pre-commercial testing or other appropriate processes (e.g., a characterization phase and / or commercial development) with the objective of planting and / or commercializing the progeny. That is, the progeny group may then be subjected to one or more additional or subsequent tests and / or selection methods, trait integration operations, hybridization with other inbred lines, and / or clustering techniques to prepare the progeny, or the plant material derived from them, for further testing and / or commercial activities. Referring again to resource allocation, and with continued reference to FIG. 1, allocation engine 110 includes (and / or is associated with) at least one computing device, which may be a standalone computing service, or it may be a computing device integrated with one or more other computing devices. Allocation engine 110 is then configured, by means of computer-executable instructions and / or one or more algorithms provided herein (or variants thereof), to perform the operations described herein, for example, as part of resource allocation on playback line 102. In addition, system 100 also includes an origins data structure 112 coupled to the allocation engine 110. In this example, the origins data structure 112 includes data related to origins, as well as related ancestors and / or origins, progenies, etc. The data can include various types of data for progenies, origins, etc., related, for example, to the origin of the plant material, the evaluation of the plant material, etc. An example of the data type included in data structure 112 is genetic marker data for origins, going back two, three, five, six, ten, or more years, etc. More generally, data structure 112 can include data consistent with a current crop / evaluation cycle and can include data related to previous crop / evaluation cycles.For example, data structure 112 may include data indicative of various different characteristics and / or traits of the plants for the current year and / or the last, one, two, five, ten, fifteen or more or less years of the plants through the cultivation and evaluation phase 106, or other cultivation spaces included within or outside the breeding line 102, and also present data from the cultivation and evaluation phase 106. In general, the origins data structure 112 includes phenotypic data, which have been measured, simulated, or both, for origins, with which phenotypic variances can be generated for each origin. An example of such variation is illustrated in Figure 2. Curve 202 represents the known or simulated phenotypic variance of a first pair of origins, and curve 204 represents the known or simulated phenotypic variance of a second, different pair of origins. In this example, the first pair of origins includes low biparental genetic similarity between the included parents, so the combination will generally produce a diverse set of progeny based on the number of loci where recombination could occur. Conversely, the second pair of origins includes relatively high biparental similarity between its parents, so the combination will generally produce a less diverse set of progeny (compared to the first origin) based on a reduced number of loci where recombination could occur. As shown in Figure 2, the highest probability of producing higher-performing progeny (as predicted by the simulation, for example) is associated with the pair of origins for curve 202, since the curve includes a larger area under the curve at the far right, which extends beyond the performance threshold 206. In this example, a higher value on the x-axis indicates higher-performing progeny, and the fact that curve 202 has more area under the curve beyond the threshold 206 indicates that it has a higher probability of producing progeny in that performance region. These variances can be predicted by simulation before allocating resources to generate a breeding population. Based on the predicted progeny performance, breeding resources can be allocated in an optimized manner to increase the probability of producing the highest-performing progeny within the line. In this exemplary mode, allocation engine 110 is configured to rely on known or simulated phenotypic variances for a given set of source pairs to allocate available breeding resources among those pairs. Specifically, allocation engine 110 is configured to employ the algorithm provided below as Equation (1) to minimize or reduce an output (through different permutations of resource allocations). minimize + A2[P(0¿> r?)(l — IP(e¿ > η))^] + λ3\\Τ1Ηχ - ξ||χ(1) The equation above is uniquely constructed to indicate resource allocation. It includes three main terms, which respectively include performance -7^(^ > η)χι, risk A2[P(0¿ > 7?)(1 - P(0¿ > η^υ^], and diversity λ3\\TIHχ - ξ||ι, where equation (1) is expected to be minimized or relatively minimized for a given set of origins. Each of the terms includes a weighting variable, λι, λ2, and λ8, which is determined based on the preference of a decision marker, extraction of historical successes, machine learning methodologies, random chance, and / or any other appropriate method. Once the set of origins is acquired through the equation above, resource allocation can be determined among the origins based on the known or simulated progeny performance of each individual breeding population.In this regard, it is expected to be adjustable by the variance of the given populations and the knowledge of the breeder, to the parental performance, to ensure the generation of progeny of desired and / or improved performance. Apart from the weights, the first term of Equation (1) describes a probability that the reproductive performance for the nth origin, η, will be greater than a target threshold, η. This is a probability distribution of trait performance and / or the probability of expressing certain genetic components. For example, the term might represent the probability that progeny from origin η will demonstrate performance greater than the desired performance threshold, qYLD, or the probability that progeny from origin η will demonstrate upright stem posture greater than the desired upright stem posture threshold, qSTLK. This can even be applied to more seemingly binary traits, such as the presence or absence of a specific haplotype, in which case the probability distribution might take a binomial form, and the threshold, η, might assume the more trivial role of indicating the binary outcomes. The probability distributions of trait values for two given (origin) populations are represented, for example, in FIG. 2, as the two curves 202, 204 for the different origins, i.e., the first origin, or origin l, which is referenced by curve 202, and the second origin, or origin_2, which is referenced by curve 204. The value acquired through known or simulated phenotypic data is a potential distribution for the progeny resulting from the specific origin, shown along with a corresponding probability of that value being exhibited by any given progeny resulting from the origin, or, generally, a binomial distribution. For example, the value along the x-axis could be any trait of interest, such as yield values, etc., and the values along the y-axis would be the probability density of the given trait value on the x-axis. Continuing with the reference to FIG.2. When the threshold or η is set to a value of 114, for example (as indicated by reference point line 206), a probability of the source having a value above the threshold is determined based on the illustrated curve. This is generally understood as the area under the curve to the right of the threshold at the value of 114. The probability is then multiplied by xi, which represents the resources (e.g., number of DHUs, etc.) to be allocated to the i-th source. The second term of Equation (1) includes the risk associated with allocating resources to the nth origin. In particular, the risk is again based on the probability that the reproductive value for the nth origin, η, will be greater than the target threshold, η. However, the risk probability is included as the variance of the reproductive value (i.e., P x (1 - P)), as represented in the curve in FIG. 2, for example. This is multiplied, again, by x„, which are the resources (e.g., number of DHUs, etc.) to be allocated to the nth origin, which is further multiplied by U„, which is the confidence level in the genotypic and phenotypic learning for the nth origin. This confidence level, U„, can be better understood as how much confidence can be given to the known or simulated genotypic and phenotypic values and distributions attributed to the nth origin (as shown in FIG. 2).2, for example) depending on how much data has been collected on the [h] origin, how well the genetic background of the [h] origin is represented in any relevant training set, and the underlying confidence / error intervals for any analyses, predictive models, etc., involved in this process. The confidence level, Ui, provides a basis for quantifying the risk associated with allocating resources to the particular origin. The risk can also take into account traits indicative of risk, such as stability, disease resistance, etc. The third term of Equation (1) incorporates a diversity of origins included with the allocation of resources to the i-th origin. Specifically, a transition probability matrix from the progeny heterotic groups to the origin heterotic groups, T, is multiplied by an incidence matrix to map the origin heterotic groups to the origins, IH, and the selected origins, x. This is then reduced by a portfolio of breeding objectives, ξ. In effect, then, the third term represents the deviation of the selected origin from a portfolio of objectives. In the exemplary mode, Equation (1) is used by the allocation engine 110, and is limited by several conditions. First, x is a positive integer, as indicated in Equation (2) below, and ey, as used in the following equations, is an indicator of x, as indicated in Equation (3). xe Z+(2) ye {0,1} (3) The sum of x, which is the amount of resources allocated to each i-th source, must equal an, which is the total number of resource units, e.g., DHUs, field plots, pots in a greenhouse, laboratory resources, etc., to be allocated by Equation (4). In other words, when 1000 DHUs are provided for allocation in Equation (1), each DHU must be allocated to a source. Likewise, Equation (5) dictates that the sum of y must equal the total number of selected sources, m. That is, a group of sources is identified in Equation (1) for which resources will be allocated, and Equation (1) must allocate at least one resource to each source, so that each source is represented by y. 1Γχ = n (4) iviA / a / zuz ι / ui lTy = m(5) In addition to the above, Equation (6) imposes an upper limit, Uupper, and a lower limit, Ulower, on the number of resources allocated to an ith source, and Equation (7) imposes a limit on x and y, in relation to the upper limit. ^lower — X — ^upper(θ)X / Uupper — y —x(7) Gender constraints are also imposed through Equations (8) and (9), as provided below. Specifically, a male incidence vector of origins, M, summed for allocated resources, y, must be greater than or equal to a number of chosen origins, m, multiplied by a male gender threshold, crm, set by the breeder, or otherwise. The threshold is set as a percentage, such as, for example, 40%, 60%, or a percentage in between, or another percentage, based on a breeding line status 102 and / or a future target. Likewise, a female incidence vector of origins, F, summed for allocated resources, y, must be greater than or equal to a number of chosen origins, m, multiplied by a female gender threshold, at, set by the breeder, or otherwise. MTy > maM(8) FTy > maF(9) Finally, in this exemplary modality, Equation (10) imposes a limit on the number of occurrences of the parents, where a parental incidence vector of origins, lp, which is summed for the allocated resources, y, must be less than or equal to a number of chosen origins, m, multiplied by a parental threshold, ap, as established by the breeder or otherwise. The parental threshold, ap, is set as a percentage, such as, for example, 5% or another percentage, depending on the status of breeding line 102 or the decision-making preference, to ensure that there is a desired and / or healthy amount of diversity in the breeding process for future genetic gains. lPy < map(10) Although described above in the context of the equations, the variables and / or terms included in Equations (1)–(10) are provided in Table 1, along with a definition of the variables and / or terms. It should be appreciated that the terms and variables are not strictly limited to the following definitions, but include any and all readily observable variations, as those skilled in the art will understand. Table 1 Term Description n number of resource units m number of selected origins n target threshold for reproductive values 0¡ reproductive value for the / th origin P¡ probability of reproductive value greater than the threshold for / th origin U¡ genetic learning confidence level for / th origin ξ portfolio of reproductive objectives Ir incidence matrix mapping parents to origins Ih incidence matrix mapping heterotic groups of origins to origins T transition probability matrix from heterotic groups of progeny to heterotic groups of origins M male incidence vector F female incidence vector x¡ amount of resource allocated to / th origin y¡ binary decision, 1 if / th origin is allocated with positive resource, otherwise, 0 The allocation engine 110 is configured to then solve the above equations, which effectively allocates resources, such as DHU, among the sources based on performance, risk, and diversity. Once the allocation engine 110 determines the allocation, it is further configured to broadcast or transmit the allocation, by source, to one or more players. In response, the players, on line 102, then utilize the resource from the sources, as defined by the allocation provided by the allocation engine 110, thereby populating playback line 102. Furthermore, it should be noted that the 110 allocation engine can be configured to provide (for example, generate and display on a player's computing device, etc.) and / or respond to a user interface through which the player (generally speaking, a user) can provide one or more inputs, which the 110 allocation engine then uses to perform resource allocations between sources. User interfaces for receiving inputs can be provided either directly on a computing device (for example, a 300 computing device as described below, etc.) associated with the player, where the 110 allocation engine is employed, or through one or more network applications through which a remote user (again, potentially a player) can interact with the 110 allocation engine (for example, an application programming interface (API), etc.). Figure 3 illustrates an exemplary computing device 300 that can be used in system 100, for example, in connection with various stages of the breeding line 102, or in connection with the allocation engine 110 and / or the progeny data structure 112, etc. For example, in different parts of the breeding line 102, breeders or other users interact with the computing devices, in accordance with computing device 300, to enter data and / or access data in the progeny data structure 112 to support breeding decisions and / or tests completed or performed by such breeders or other users. In connection with this, the allocation engine 110 of system 100 includes and / or is implemented in at least one computing device compatible with computing device 300.In this regard, the computing device 300 can be uniquely or specifically configured, by means of executable instructions, to implement the various algorithms and other operations described in this document with respect to the allocation engine 110. It should be appreciated that the system 100, as described in this document, can include a variety of different computing devices, either compatible with computing device 300 or different from computing device 300. The exemplary computing device 300 may include, for example, one or more servers, workstations, personal computers, laptops, tablets, smartphones, other suitable computing devices, combinations thereof, etc. Furthermore, the computing device 300 may include a single computing device, or it may include multiple computing devices located in close proximity or distributed across a geographic region, and linked together via one or more networks. Such networks may include, without limitation, the Internet, an intranet, a private or public local area network (LAN), a wide area network (WAN), a mobile network, telecommunications networks, combinations thereof, or other suitable networks, etc.In one example, the progeny data structure 112 of system 100 includes at least one server computing device, while the allocation engine 110 includes at least one separate computing device, which is coupled to the progeny data structure 112, either directly and / or by one or more LANs, etc. That said, the illustrated computing device 300 includes a processor 302 and a memory 304 that is coupled to (and communicating with) the processor 302. The processor 302 may include, without limitation, one or more processing units (e.g., in a multi-core configuration, etc.), including a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuit or processor capable of performing the functions described herein. The foregoing list is provided by way of example only and is not intended to limit in any way the definition and / or meaning of "processor." Memory 304, as described herein, is one or more devices that allow the storage and retrieval of information, such as executable instructions and / or other data. Memory 304 may include one or more computer-readable storage media, such as, but not limited to, dynamic random-access memory (DRAM), static random-access memory (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), solid-state drives, flash drives, CD-ROMs, USB flash drives, tapes, hard disks, and / or any other type of volatile or non-volatile computer-readable physical or tangible medium.Memory 304 can be configured to store, without limitation, progeny data structure 112, phenotypic data, test data, source data (e.g., trait performance distributions, etc.), weights, thresholds, and / or other types of data (and / or data structures) suitable for use as described herein, etc. In various embodiments, computer-executable instructions can be stored in memory 304 for execution by processor 302 to cause processor 302 to perform one or more of the functions described herein, such that memory 304 is a physical, tangible, non-transient, computer-readable storage medium. Such instructions often improve the efficiencies and / or performance of processor 202 performing one or more of the various operations of this invention.It should be noted that memory 304 may include a variety of different memories, each implemented in one or more of the functions or processes described in this document. In the exemplary embodiment, the computing device 300 also includes an output device 306 that is coupled to (and in communication with) the processor 302. The output device 306 outputs, or presents, to a user of the computing device 300 (e.g., a player, etc.), for example, by displaying and / or otherwise outputting information such as, without limitation, selected offspring, offspring as commercial products, and / or any other type of data as desired. It should be further appreciated that, in some embodiments, the output device 306 may comprise a display device so that various interfaces (e.g., network or other applications, etc.) can be displayed on the computing device 300, and in particular on the display device, to show such information and data, etc.In some examples, the computing device 300 can display interfaces on a display device of another computing device, which includes, for example, a server hosting a website with multiple web pages, or interacting with a web application used on the other computing device, etc. The output device 306 can include, without limitation, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED display (OLED), an electronic ink display, combinations thereof, etc. In some configurations, the output device 306 can include multiple units. The computing device 300 further includes an input device 308 that receives user input. The input device 308 is coupled to (and in communication with) the processor 302 and may include, for example, a keyboard, a pointing device, a mouse, a stylus, a touch-sensitive panel (e.g., a touchpad or touchscreen, etc.), another computing device, and / or an audio input device. In addition, in some exemplary embodiments, a touchscreen, such as one included in a tablet or similar device, may function as both an output device 306 and an input device 308. In at least one exemplary embodiment, the output device 306 and the input device 308 may be omitted. In addition, the illustrated computing device 300 includes a network interface 310 coupled to (and communicating with) the processor 302 (and, in some configurations, also with the memory 304). The network interface 310 may include, without limitation, a wired network adapter, a wireless network adapter, a telecommunications adapter, or other devices capable of communicating with one or more different networks. In at least one configuration, the network interface 310 is used to receive input to the computing device 300. For example, the network interface 310 may be coupled to (and communicating with) data collection devices in the field in order to collect data for use as described herein. In some exemplary configurations, the computing device 300 may include the processor 302 and one or more network interfaces incorporated in or with the processor 302. Figure 4 illustrates an exemplary method 400 for selecting progenies in a progeny identification process. Exemplary method 400 is described here in relation to system 100 and can be implemented, in whole or in part, in allocation engine 110 of system 100. Furthermore, for illustrative purposes, exemplary method 400 is also described with reference to the distributions in Figure 2 and computing device 300 in Figure 3. However, it should be appreciated that method 400, or other methods described herein, are not limited to system 100, the distributions in Figure 2, or computing device 300. Likewise, conversely, the systems, data structures, and computing devices described herein are not limited to exemplary method 400. To begin, a breeder (or other user) initially identifies a plant type (e.g., corn, soybean, etc.) and one or more desired phenotypes, potentially consistent with one or more desired characteristics and / or traits to be advanced in the identified plant, or a desired performance in a commercial plant product. In turn, based on the foregoing and / or one or more other criteria, the breeder or user, alone or through various processes, selects multiple origins as a starting point. The origin may be selected by any suitable means, in view of the foregoing, including, again, through the methods described in Applicant's U.S. Joint Ownership Application No. 15 / 618.023, which is incorporated herein by reference in its entirety. In this exemplary modality, 200 sources are selected, which can be denoted m, and the available resources include 1000 DHU, which can be denoted n. By way of explanation, these numbers can provide 1.323 x 10215 different possible ways of distributing 1000 DHU among the 200 sources (where each source is included in at least one DHU and is also allowed to include up to a maximum number of the remaining resources). For the multiple selected origins, data structure 112 includes various data representative of the origins. Among the data, data structure 112 includes a trait performance distribution, which generally provides a probability that the origin will exhibit a specific trait value. The probability is usually determined based on tests and / or prediction models, for example, trained on historical data, including past gene products and the distribution of the specific trait of interest. As shown in FIG. 2, for example, the trait performance distribution is illustrated as a binomial distribution for the two origins, in curves 202 and 204, which is indicative of a probability that the respective origins will perform at the indicated value. Thus, for example, the first origin, or origin _1 (identified in curve 202), has a probability of 0.08 has a performance of 104, while the second source, or source _2 (identified in curve 204), has a probability of 0.03 of having a performance of 107. As can be seen in FIG. 2, the probability of performance above the exemplary target threshold 206 (which has a performance value of 114) is higher for the second source (or source _2, identified in curve 204) than for the first source (or source _1, identified in curve 202). It should be appreciated that a distribution and / or other probability expression of the type described here is included in data structure 112 for each of the multiple selected sources. Furthermore, data structure 112 also includes a confidence level for genetic learning, referred to earlier as U1. This confidence level can be based on the frequency with which genetic material similar to a given origin is present within previously tested sets in breeding line 102 and / or historical datasets used to train one or more suitable predictive models employed within the overall breeding process and / or the resource allocation process described herein. The confidence level further explains the robustness of the one or more predictive models employed, which can be based, for example, on how well the origin is known, and / or the confidence in the origin provided by the distribution.Simply put, this frequency can be used in comparison to the average frequency of genetic families within the training sets to create an estimate of how much confidence there is in the model. For example, if a particular genetic family is represented 1.5 times more often within the training set than the average family, 1.5 could be used as the confidence level for this particular line. Likewise, another family might be represented 0.75 times, and a cross between these two lines could be characterized by Ui = 1.5 x 0.75 = 1.125, where the confidence level for the origin is a simple multiplication of the confidence levels for the parents. It is important to note that genetic confidence can also be derived in much more sophisticated ways. For example, the confidence for each parent of the cross could be derived as a result of a Bayesian analysis of the entire germplasm pool.The posterior origin confidence level could be derived itself using a more sophisticated convolution of the parental confidences or, even more directly, it could be derived from the confidence results of any machine learning algorithm and / or simulation engines that have been used to assess this variance of the expected reproductive value of the origin. Furthermore, data structure 112 includes a portfolio of objectives for breeding objective sets, for example, by the breeder at the beginning of the start-up phase 104 (or later), which is ξ. The portfolio of objectives may include any of a number of objectives and distributions that define what a target, desired, or ideal germplasm pooling might look like in breeding line 102. Some of these objectives may include gender distributions (heterotic pooling) across breeding line 102, the distribution of different germplasm pools within breeding line 102, and the desired distribution of parents at different stages of the breeding life cycle (for example, to balance the use of older, proven parents with younger, less proven parents with newer genetics; etc.).For an example profile, one operator might decide that a line should have at least 45% male and 45% female lines, with the remainder selected based on performance. Meanwhile, another operator might decide that the origins in the line should have a perfect 50 / 50 split between heterotic male and female groups. In yet another example, a target profile could be based on the maturity distribution of the origins within a specific breeding line. For instance, if a line were responsible for a six-day crop maturity period, a potential target maturity profile for the material to be added to the line might specify that 25% of all origins should fall within the first two days of that period, 50% should fall within the middle two days, and 25% of the origins should fall within the last two days of the period.This target profile would help ensure that most lines produced by origins with this intermediate parental maturity (average of the individual maturity of the two parents) fall within the six-day window for the line. Despite these specific examples, it should be noted that the target profile can include any profile considered desirable by the producer and / or a person involved in resource allocation among the origins included in the allocation. Objectives can be set in several ways. Most simply, objectives can be set through human input to align breeding line 102 with certain business goals or constraints. These objectives can be communicated to data scientists and then manually transferred to the allocation engine 110, or they can be stored in a database or API using a web user interface or other tool. With the development of more advanced analytics and simulations, objectives could be set algorithmically based on a defined plan, roadmap, or strategy to have a desired and / or higher probability of improving, leveraging, and / or maximizing breeding line 102 and / or the associated business performance with the allocated resources, and potentially align closely with future market needs for a given plant, etc.The targets could be stored in a database or API for later retrieval by the 110 allocation engine, as desired and / or required for performance as described in this document. As shown in FIG. 4, in method 400 (402), the allocation engine 110 accesses the data included in the data structure 112 for the multiple selected origins. The data includes, for example, a probability distribution for trait performance for each of the selected origins. Other data, for each origin, may include gender data, parental and / or heterotic data, etc. Then, allocation engine 110 determines, in step 404, a resource allocation of the available resources (i.e., the 1000 DHU in this example) for the multiple selected sources. Specifically, in this example, allocation engine 110 uses the allocation algorithm in Equation (1) (reproduced below). It should be noted that, in other versions of the method, different algorithms (derived from Equation (1) or not) can be used to allocate the available resources among a set of sources. ΣΝ >η)Χί + λ2[^(θι > 77)(1- IP(0¿> η^υ^] + λ3\\ΤΙΗχ - ξΙ^ ί=1 As explained above, the algorithm in Equation (1) includes three terms, which are generally related to performance, risk, and diversity. It is important to note that the resource allocation process described in this document can be applied not only to high-level decisions such as how to distribute DHUs or how to allocate test plots, but also to auxiliary and secondary decisions. For example, even once this process has been used to allocate DHUs, as discussed above, among a set of origins based on the expectation of how the performance distributions of different origins (e.g., yield, etc.) from known or simulated phenotypic data indicate the probability that their progeny will reach or exceed a certain performance level, it can also be applied to subprocesses within the double haploid (DH) process. For example, when a subprocess within a DH process produces more seeds from DH lines, it should be noted that after production, for example, there can only be a finite number of greenhouse spaces in which the DH process can be carried out normally. The reproductive value (on the line in Figure 2) that would be relevant in the process is the probability distribution of the number of grains that a given inbreeding would produce per plant. Based on the probability that a given line will produce more than a set limit, for example, 180 grains, per plant, the limited greenhouse spaces can be allocated to different lines to improve and / or maximize the number of grains produced, while ensuring that each line has a required and / or minimum number of grains at the end of the process. Due to the complexity involved in resource allocation, the algorithms and computing technologies described in this document are based on their commercial applications. However, for illustrative purposes, a simplified example is presented herein. In this regard, it is instructive to consider a case in which three greenhouse sites must be divided between two DH lines in order to produce more seeds, as described above. The relevant values for the problem are as follows: Table 2 Term Value n 3 greenhouse units (one plant per unit) η 180 grains per plant Pi 0.3 P2 0.9 Ui 0.5 U2 1.25 ξ Each line must have at least one resource Ai 0.3 a2 0.3 A3 0.4 In general, here, the third term (diversity) would impose a target distribution across the origins, which in this example would likely be a desired number of grains for each origin, determined through another process or analysis. To keep this example simple for illustrative purposes, this term will be simplified by stating that each line must have at least one resource allocated. With this objective, the third term would be +1 * λ³ for solutions where one or the other line has no resources allocated, and +0 when both lines receive at least one resource. Given the other values defined above, this would prevent solutions with a non-zero third term from producing the minimized solution, so this example can focus only on the two possible solutions where both lines receive resources. Expanding Equation (1) for a total of two lines (N = 2) yields: minimize [[-A^^i + λ2(Ρ1(1 - P^U^ + λ3* 0] +[-λ1Ρ2χ2+ λ2(Ρ2(1 - P2)U2x2~) + λ3* 0]] If the values from Table 2 are entered into this expanded equation for each of the two possible ways of distributing the resources, results will be obtained for each potential solution. Minimizing the result, in this case, will mean selecting the resource allocation that yields the smallest number in this equation. Solution 1. Line 1 gets two resources, and line 2 gets one resource. [—0.3 * 0.3 * 2 + 0.3 * 0.3 * 0.7 * 0.5 * 2 + 0.4 * 0] + [—0.3 * 0.9 * 1 + 0.3 * 0.9 * 0.1 * 1.25 * 1 + 0.4 * 0] = -0.353 Solution 2. Line 1 gets one resource, and line 2 gets two resources. [—0.3 * 0.3 * 1 + 0.3 * 0.3 * 0.7 * 0.5 * 1 + 0.4 * 0] + [—0.3 * 0.9 * 2 + 0.3 * 0.9 * 0.1 * 1.25 * 2 + 0.4 * 0] = -0.531 As can be seen above, Solution 2, in which line 1 receives one resource and line 2 receives two resources, yields the minimum solution to Equation (1). This indicates that this solution achieves the highest probability of producing the most seeds while ensuring that each line receives at least one resource. Furthermore, it can be seen that, in this particular situation, although the uncertainty surrounding the reliability of line 2 was much greater than that of line 1, the large difference in their probability of success offset the uncertainty. While the nature of this example is simplified for illustrative purposes in this document, it remains exemplary of both the impact of the methodology and its versatility (and practical applicability) in terms of the different types of plant breeding allocations that will be undertaken. With reference to Figure 4, allocation engine 110 then allocates the DHU accordingly for the multiple selected origins in a manner consistent with the determined resource allocation. Specifically, in the previous example, with respect to Table 2, the three greenhouse units are allocated as follows: one to Line 1 and two to Line 2, so that the physical material consistent with the lines is either physically removed or planted in the specific greenhouse units. In practice, for example, when the lines are both maize plants, a plant with an 'inducer' genotype (i.e., a plant that has a relatively high probability of producing haploid progeny when crossed with a diploid maize plant) is used to pollinate the silks of one progeny plant from Line 1 and two progenies from Line 2 (where each greenhouse unit is allocated one plant).The resulting haploid progeny are exposed to a mitotic inhibitor (e.g., colchicine, etc.) to disrupt normal cell division and induce chromosome duplication in the nucleus. Therefore, the resulting plants have two identical chromosomes with elite genetics. A person skilled in the technique will understand that the DHU could also be allocated to create haploid plants in vivo through parthenogenesis (apomixis) or pseudogamy; or in vitro, through gynogenesis and / or androgenesis. For example, in the case of the reproduction of Brassica napus and Brassica juncea, haploid plants can be created using microspore culture, other culture, and ovum / ovule culture to generate subsequent duplicate haploid plants. It should further be understood that the allocation or granting of resources, in accordance with the allocation determined in Method 400, may be carried out in other ways, depending, for example, on the types of resources to be allocated / granted and the plants to be grown. Furthermore, resource allocation can be performed by the allocation engine 110, by users associated with the allocation determined in method 400 (e.g., players, etc.), or by a combination thereof. For example, the allocation engine 110 might issue a report as part of the allocation in method 400, indicating the determined allocation (e.g., where the report accounts for the resources available for allocation and the sources granted to those allocations, etc.), after which one or more users associated with the playback line 102 might physically enforce the determined allocation across multiple resources.In this example, the physical resources on playback line 102 are altered and / or implemented by allocating resources consistently with the determined allocation, thereby transforming the resources from generic to specific (i.e., each resource is implemented with the specific source designated in the allocation). It should be noted that the involvement of allocation engine 110 and / or one or more users, or combinations thereof, may differ depending on the specific type and number of resources being allocated, the specific playback line 102, the sources selected and allocated as described in this document, and so on. In light of the above, the unique systems and method described herein provide intelligent resource allocation for breeding lines. In particular, resources (and their use) can generally be time-consuming, costly, or even limited for specific breeding lines (e.g., depending on the type of plants being bred in the given lines, etc.). This document employs one or more algorithms that consider the performance probabilities of traits for the origins (e.g., expressed as a binomial distribution, etc.), as well as the risk and / or genotypic components and / or diversity associated with the selected origins. The described algorithms allocate resources (which include growing space (e.g., field plots, etc.), field equipment, laboratory space, laboratory equipment, personnel, etc.) accordingly.Breeding lines (or a combination or subset thereof) are allocated with a higher probability of producing progeny that perform above one or more thresholds, and / or with a higher probability of producing progeny that express certain genetic components at rates considered appropriate and / or desired for the breeding lines. Therefore, breeding lines, which are based on data related to origins not previously relied upon for resource allocation (and, by extension, the process that implements the data) (i.e., using particular information and techniques), allow for the improvement described herein (i.e., the improvement of existing technologies and processes for resource allocation to promote the identified origins with greater potential to more resources) over the conventional even distribution of resources among the identified origins. That said, it should be appreciated that the functions described in this document, in some forms, can be described in computer-executable instructions stored on a computer-readable medium and executable by one or more processors. A computer-readable medium is a non-transient, computer-readable medium. By way of example, and without limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Combinations of the foregoing should also be included within the scope of computer-readable media. It should also be appreciated that one or more aspects of this disclosure transform a general purpose computing device into a special purpose computing device when configured to perform the functions, methods, and / or processes described herein. As will be further appreciated on the basis of the descriptive memorandum above, the above-described modalities of disclosure can be implemented using computer programming or engineering techniques, including software, firmware, computer hardware, or any combination or subset thereof, wherein the technical effect can be achieved by performing at least one of the following operations: (a) for multiple sources, by accessing a data structure that includes representative data from the multiple sources, wherein the data includes, for each of the multiple sources, a trait performance expression and / or genotypic components;(b) determining, by means of at least one computer device, a resource allocation, which allocates n resources among the multiple origins, based on a probability associated with the performance expressions of traits and / or genotypic components for the origins, where n is an integer; and (c) allocating the n resources in a breeding line for the multiple origins, based on the determined resource allocation, whereby the origins impose themselves on the resources in a manner consistent with the resource allocation; and / or (d) where: (i) the determination of the resource allocation includes determining the resource allocation on the basis of a comparison of:; value Y* > r / X + Á2[P(0¿> ^)(1 - P(0É> η)^] +λ3\\T1Hχ - ξΙΚ '¿ = 1 for multiple allocations of potential resources; (i) at least one of the n resources is allocated in the resource allocation to each of the multiple sources; and wherein each of the n resources is allocated in the resource allocation to one of the multiple sources; (ii) the determination of the resource allocation for a hybrid crop in which the heterotic male and female groups are kept separate includes determining the resource allocation, subject to: MTy > maM, FTy > maF, and aM+ aF< 1; (iv) the determination of resource allocation includes determining the allocation of resources based on a predefined portfolio of objectives, whereby a relative value for each potential resource allocation is reduced based on a deviation of the resource allocation from the predefined portfolio of objectives; and / or (v) the determination of resource allocation includes determining the allocation of resources based on confidence in the expression of performance traits and / or genotypic components for each of the multiple origins. Examples and embodiments are provided to ensure that this disclosure is complete and fully conveys its scope to those skilled in the art. Numerous specific details, such as examples of specific components, devices, and methods, are set out to provide a comprehensive understanding of the embodiments described herein. It will be evident to those skilled in the art that specific details are not required, that the exemplary embodiments can be implemented in many different ways, and that none of these should be construed as limiting the scope of the disclosure. In some exemplary embodiments, well-known processes, well-known device structures, and well-known technologies are not described in detail.Furthermore, the advantages and improvements that can be achieved with one or more exemplary modalities described herein may provide all or none of the advantages and improvements mentioned above, and still fall within the scope of this disclosure. The specific values described herein are illustrative in nature and do not limit the scope of this disclosure. The disclosure herein of particular values and ranges of values for given parameters is not exclusive of other values and ranges of values that may be useful in one or more of the examples described herein. Furthermore, it is intended that any two particular values for a specific parameter stated herein may define the endpoints of a range of values that may also be suitable for the given parameter (i.e., the disclosure of a first and second value for a given parameter may be interpreted to describe that any value between the first and second values could also be used for the given parameter).For example, if parameter X is exemplified here to have the value A and is also exemplified to have the value Z, it is contemplated that parameter X may have a range of values from approximately A to approximately Z. Similarly, it is contemplated that disclosing two or more ranges of values for a parameter (where such ranges are either nested, overlapping, or distinct) subsumes all possible combinations of ranges for the value that could be claimed using endpoints of the disclosed ranges. For example, if parameter X is exemplified here to have values in the range 1–10, or 2–9, or 3–8, it is also contemplated that parameter X may have other ranges of values, including 1–9, 1–8, 1–3, 1–2, 2–10, 2–8, 2–3, 3–10, and 3–9. The terminology used in this document is intended to describe particular exemplary modalities only and is not intended to be exhaustive. As used herein, the singular forms "a," "an," and "the" may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises," "comprising," "includes," and "has" are inclusive and thus specify the presence of stated features, whole numbers, stages, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, whole numbers, stages, operations, elements, components, and / or groups thereof. The stages, processes, and operations of the method described herein should not necessarily be interpreted as requiring their execution in the particular order discussed or illustrated, unless specifically identified as an order of execution.It should also be understood that additional or alternative stages may be used. When a feature is referred to as being on, attached to, connected to, coupled to, associated with, in communication with, or included with another feature or layer, it may be directly on, attached, connected, or coupled, or associated with, or in communication with, or included with the other feature, or intermediate features may be present. As used herein, the terms "and / or at least one of" include any and all combinations of one or more of the listed features being associated. None of the elements listed in the claims are purported to be a means-plus-function element within the meaning of 35 USC §112(f), unless an element is expressly recited using the phrase "means to," or in the case of a method claim, using the phrases "operation to" or "step to." Although the terms first, second, third, etc., may be used in this document to describe various features, these features should not be limited by these terms. These terms may only be used to distinguish one feature from another. Terms such as first, second, and other numerical terms, when used in this document, do not imply a sequence or order unless clearly indicated in the context. Therefore, a first feature discussed in this document could be called a second feature without departing from the exemplary modalities. The preceding description of the modalities has been provided for illustrative and descriptive purposes. It is not intended to be exhaustive or to limit disclosure. The individual elements or features of a particular modality are generally not limited to that particular modality, but, where applicable, are interchangeable and may be used in a selected modality, even if not specifically shown or described. The same may also be varied in many ways. Such variations should not be considered a departure from disclosure, and all such modifications are intended to be included within the scope of disclosure. «7 / 1 I Π / I 7Π7 / Β / Y» NOVELTY OF THE INVENTION Having described the present invention as above, it is considered novel and, therefore, the contents contained in the following are claimed as property:
Claims
1. A computer-implemented method for allocating resources in a breeding line to multiple origins, the method being characterized in that it comprises: for multiple origins, accessing a data structure that includes representative data from the multiple origins, wherein the data includes, for each of the multiple origins, a performance expression of traits and / or genotypic components; to determine, by means of at least one computer device, a resource allocation, which allocates n resources among the multiple origins, on the basis of a probability associated with the performance expressions of traits and / or genotypic components for the origins, as defined by: ΣN > η)χι +Á2[P(0¿ > η)(1- P(0¿ > η^Ι^ϧί] + λ3\\TIHχ - ξ||χ ί=ι where n is an integer number of available resources; η is a target threshold for the reproductive value; Θ, is a variable for a reproductive value, or a vector thereof, for the specific origin;P¡ is the probability of finding a reproductive value, or a vector thereof, greater than some threshold for the specific origin; U¡ is a confidence level of genetic learning for a specific origin; ξ is a portfolio of breeding objectives; yx¡ is an integer decision variable for the resources allocated to the specific origin; and physically allocate the n resources in a breeding line for the multiple origins, based on the determined resource allocation.
2. The method of claim 1, characterized in that at least one of the n resources is allocated in the resource allocation to each of the multiple sources; and wherein each of the n resources is allocated in the resource allocation to one of the multiple sources.
3. The method of claim 1, characterized in that determining the allocation of resources for a hybrid crop in which the male and female heterotic groups are kept separate includes determining the allocation of resources, subject further to: MTy > maM, FTy > maF, and aM + aF < 1; wherein M is the male incidence vector; aM is the minimum fraction of m origins that are designated to be devoted to male crosses; F is the female incidence vector; and aF is the minimum fraction of m origins that are designated to be devoted to female crosses; whereby the n resources can be appropriately allocated to each heterotic group without exceeding the maximum of m origins.
4. The method of claim 1, characterized in that at least one of the n resources is allocated in the resource allocation to each of the multiple sources; and wherein each of the n resources is allocated in the resource allocation to one of the multiple sources.
5. The method of claim 4, characterized in that determining the allocation of resources for a hybrid crop in which the male and female heterotic groups are kept separate includes determining the allocation of resources, subject further to: MTy > maM, FTy > maF, and aM + aF < 1; wherein M is the male incidence vector; aM is the minimum fraction of m origins that are designated to be devoted to male crosses; F is the female incidence vector; and aF is the minimum fraction of m origins that are designated to be devoted to female crosses; whereby the n resources can be appropriately allocated to each heterotic group without exceeding the maximum of m origins.
6. The method of claim 1, characterized in that the determination of the resource allocation includes determining the resource allocation on the basis of a predefined portfolio of objectives, whereby a relative value for each potential resource allocation is reduced based on a deviation of the resource allocation from the predefined portfolio of objectives.
7. The method of claim 1, characterized in that the determination of resource allocation further includes determining resource allocation based on confidence in the expression of trait performance and / or genotypic components for each of the multiple origins.
8. The method of claim 1, characterized in that the physical allocation of the n resources in the breeding line includes planting at least one plant product based on at least one of the multiple origins and at least one progeny of the multiple origins, in a cultivation space consistent with the determined resource allocation.
9. A system for allocating resources in a breeding line, the system being characterized in that it comprises: a data structure that includes representative data from multiple selected origins, wherein the data includes a performance expression of traits and / or genotypic components for each of the multiple selected origins; and a computer device coupled in communication with the data structure and configured to: access data in the data structure for each of the multiple selected origins; and determine a resource allocation, which allocates n resources among the multiple selected origins, on the basis of a probability associated with the performance expression of traits and / or genotypic components for the origins, wherein n is an integer.
10. The system of claim 9, characterized in that at least one of the n resources is allocated, in the resource allocation, to each of the multiple sources; and wherein each of the n resources is allocated, in the resource allocation, to one of the multiple sources.
11. The system of claim 9, characterized in that the computer device is configured to determine the allocation of resources on the basis of a reduction and / or minimization of the value for each potential allocation, wherein the value for each potential allocation is defined as: value ΣΓ=ι Wi > + λ2[Ρ(A > η)(1 - m > η))υΛ]+λ3\\T1Hχ - ξ||ι; wherein n is an integer number of available resources; η is a target threshold for the reproductive value; Θ, is a variable for a reproductive value, or a vector thereof, for the specific origin; P¡ is the probability of finding a reproductive value, or a vector thereof, greater than some threshold for the specific origin; U¡ is a level of confidence in genetic learning for a specific origin; ξ is a portfolio of reproductive objectives; yx¡ is an integer decision variable for the resources allocated to the specific origin.
12. The system of claim 9, characterized in that the computing device is configured to determine resource allocation, further consistent with: l'x = n; lry = m; MTy > maM\ FTy > maF\ lPy < map; ^inferior — X — ^-superior > X / ^superior — y — X i X e Z+; y ye {oi}; iviA / a / zuz ι / uii rza where x is the integer resource allocation variable indicating the resources allocated to each source; n is the total number of resources available for allocation; y is the binary selection variable indicating which sources have been selected; m is the target number of sources in the selected set; lp is the parental incidence vector for the sources; ap is the threshold set for the maximum parental usage rate within the selected set; uinferior is the lower limit of the number of resources that can be allocated to the / th source; and usuperior is the upper limit of the number of resources that can be allocated to the / th source.
13. The system of claim 9, characterized in that it further comprises a breeding line; and wherein the breeding line includes the n resources allocated to one or more cultivation spaces according to said determined resource allocation.
14. The system of claim 13, characterized in that the computer device is configured to determine resource allocation based on a reduction and / or minimization of the value for each potential allocation, wherein the value for each potential allocation is defined as: value ΣΓ=ι > η)χι + Á2[P(0¿ > / / )(1 - P(0¿ > η^υ^] + λ31|T1Hχ - ξ||ι; wherein η is an integer number of available resources; η is a target threshold for the reproductive value; Θ, is a variable for a reproductive value, or a vector thereof, for the specific origin; P, is the probability of finding a reproductive value, or a vector thereof, greater than some threshold for the specific origin; U¡ is a level of confidence in genetic learning for a specific origin; ξ is a portfolio of reproductive objectives; yx¡ is an integer decision variable for the resources allocated to the origin specific.
15. The system of claim 14, characterized in that at least one of the n resources is allocated, in the resource allocation, to each of the multiple sources; and wherein each of the n resources is allocated, in the resource allocation, to one of the multiple sources.
16. The system of claim 15, characterized in that the computing device is configured to determine resource allocation, further consistent with: 1T x = n; lry = m; MTy > maM\ FTy > maF', IPy < map; ^inferior — % — ^superior i / ^superior — y — X e Z+; iviA / a / zuz ι / uii rza y ye {OL}; where x is the integer resource allocation variable indicating the resources allocated to each source; n is the total number of resources available for allocation; y is the binary selection variable indicating which sources have been selected; m is the target number of sources in the selected set; lp is the parental incidence vector for the sources; ap is the threshold set for the maximum parental usage rate within the selected set; uinferior is the lower limit of the number of resources that can be allocated to the / th source; and usuperior is the upper limit of the number of resources that can be allocated to the / th source.
17. A non-transient, computer-readable storage medium characterized in that it comprises a method for allocating resources in a breeding line, which, by means of at least one processor, causes the at least one processor to perform the steps of: for multiple origins, accessing a data structure that includes representative data of the multiple origins, wherein the data includes, for each of the multiple origins, a performance expression of traits and / or genotypic components; determining a resource allocation, which allocates n resources among the multiple origins, on the basis of a probability associated with the performance expressions of traits and / or genotypic components for the origins, where n is an integer; and allocating the n resources in a breeding line for the multiple origins, according to the determined resource allocation.
18. The computer-readable non-transient storage medium of claim 17, characterized in that the step of the method of determining the allocation of resources further causes the at least one processor to determine the allocation of resources on the basis of a comparison of: ΣN -Wh > + λ2[P(A > / ))(1 - P(0í > +M77Hx - ξIIi ί=ι for the multiple allocations of potential resources; wherein n is a number of available resources; η is a target threshold for the reproductive value; 9¡ is a variable for a reproductive value, or a vector thereof, for the specific origin; P¡ is the probability of finding a reproductive value, or a vector thereof, greater than some threshold for the specific origin; U¡ is a level of confidence of genetic learning for a specific origin; ξ is a portfolio of reproductive objectives; yx¡ is an integer decision variable for the resources allocated to the specific origin. iviA / a / zuz ι / uii rza 19. The computer-readable non-transient storage medium of claim 18, characterized in that at least one of the n resources is allocated in the resource allocation to each of the multiple sources; and wherein each of the n resources is allocated in the resource allocation to one of the multiple sources.
20. The computer-readable, non-transient storage medium of claim 19, characterized in that the step of the method for determining resource allocation further causes at least one processor to determine the resource allocation for a hybrid crop in which the male and female heterotic groups are kept separate, subject to: MTy > maM, FTy > maF, and aM + aF < 1; wherein M is the male incidence vector; aM is the minimum fraction of m origins that are designated to be dedicated to male crosses; F is the female incidence vector; and aF is the minimum fraction of m origins that are designated to be dedicated to female crosses; whereby the n resources can be appropriately allocated to each heterotic group without exceeding the maximum of m origins.