Methods and systems for use in trait development in agricultural crops
By combining multimodal deep learning models with genotype, weather, and soil data, the problems of resource waste and inefficiency in agricultural crop breeding have been solved, and efficient trait prediction and optimization have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MONSANTO TECHNOLOGY LLC
- Filing Date
- 2024-10-14
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies in agricultural crop breeding face the challenge of relying heavily on testing for decision-making, leading to resource waste and inefficiency, and making it difficult to efficiently predict and optimize plant trait performance.
A deep learning model using a multimodal architecture combined with genotype, weather, and soil data is used to predict and optimize the phenotypic performance of agricultural crops through training and validation, and to utilize computing devices for gene sequence improvement and variety selection.
It enables efficient and accurate prediction and optimization of agricultural crop traits, reduces resource waste, and improves breeding efficiency and phenotypic gains of varieties.
Smart Images

Figure CN122121732A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims the benefit and priority of Greek Patent Application No. 20230100865, filed on October 19, 2023. The entire disclosure of the aforementioned application is incorporated herein by reference. Technical Field
[0002] This disclosure relates in general to methods and systems for use in the development of traits in agricultural crops. Background Technology
[0003] This section provides background information relating to this disclosure, which is not necessarily prior art.
[0004] Plant improvement is often known to be achieved through selective breeding or genetic manipulation. Based on specific improvements, the resulting plants can exhibit one (or more) desired traits. These traits can be tested across a variety of different environments, and when they are confirmed, the improved plants can be advanced for further plant development and / or commercialization, thereby propagating the plants and selling them to growers. Summary of the Invention
[0005] This section provides a general overview of this disclosure and is not a full disclosure of its entire scope or all its features.
[0006] The exemplary embodiments of this disclosure generally relate to a system for use in interpreting traits of interest in agricultural crops. In one exemplary embodiment, such a system typically includes a computing device comprising a memory and at least one processor, wherein the memory includes executable instructions, a trained predictive architecture, and a repository. The repository includes genotypic data of a plurality of known varieties, as well as weather and soil data associated with the growth of these known varieties. The at least one processor is configured via the executable instructions and the trained predictive architecture to: identify a plurality of proposed varieties of the crop, wherein each of the plurality of proposed varieties includes distinct gene sequences compared to other proposed varieties among the plurality of proposed varieties and the known varieties; use the trained model to predict the trait of interest for each of the plurality of proposed varieties based on the data included in the repository; select a proposed variety among the plurality of proposed varieties according to a phenotypic gain-based acquisition function; and guide seeds representing the selected proposed variety among the plurality of proposed varieties to an experimental stage to evaluate the trait of interest of the selected proposed variety among the plurality of proposed varieties.
[0007] Example embodiments of this disclosure also relate generally to methods for use in interpreting traits of interest in agricultural crops. In one exemplary embodiment, the method typically includes: identifying a plurality of proposed varieties of a crop, wherein each of the plurality of proposed varieties includes distinct gene sequences compared to other proposed varieties and known varieties among the plurality of proposed varieties; predicting the trait of interest for each of the plurality of proposed varieties using a trained model based on data included in a repository; selecting a proposed variety among the plurality of proposed varieties according to a phenotypic gain-based acquisition function; and directing seeds representing the selected proposed variety among the plurality of proposed varieties to an experimental phase to evaluate the trait of interest in the selected proposed variety among the plurality of proposed varieties.
[0008] As will be apparent from the description provided herein, further applicability will be readily apparent. The descriptions and specific examples within this invention are intended for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description
[0009] The accompanying drawings described herein are for illustrative purposes only, and not for all possible implementations, and are not intended to limit the scope of this disclosure.
[0010] Figure 1 This disclosure provides an example system applicable to trait development in agricultural crops; Figure 2 It is possible Figure 1 The example architecture used in the system includes multiple different data modalities in its respective modal layers for use in interpreting traits in agricultural crops; Figure 3 It is possible Figure 1 A block diagram of an example computing device used in the system; and Figure 4 It is used in the development and / or creation of new varieties of agricultural crops, and is suitable for use with Figure 1 This is an example method used in systems that are similar to this one.
[0011] Throughout the various views in the accompanying drawings, corresponding reference numerals indicate the corresponding parts. Detailed Implementation
[0012] Example embodiments will now be described more fully with reference to the accompanying drawings. The descriptions and specific examples included herein are intended for illustrative purposes only and are not intended to limit the scope of this disclosure.
[0013] In conjunction with agricultural crop development, a variety of different techniques can be used to promote desired traits in crops. These techniques often rely on specific decisions that essentially predict the expected performance of the improved plant in at least one specific trait (e.g., yield, disease resistance, plant physiological traits, etc.). Thus, given the thousands of potential plants, sources, etc., the number of decisions made in each experiment to consider the expected trait performance can reach billions, trillions, or even infinity. These decisions also become dependent on testing some of them, which then leads to other decisions, resulting in a often vast amount of physical and biological resources allocated to breeding decisions.
[0014] Uniquely, the methods and systems in this paper provide an interpretation of agricultural crop traits based on a multimodal architecture, through which a wide range of decisions can be considered.
[0015] Figure 1 An example system 100 is shown for integrating traits in agricultural crops based at least in part on the balance of crop loss and harvest characteristics. Although the parts of system 100 are presented in one arrangement in the described embodiment, other embodiments may include the same or different parts arranged in other ways, depending on (e.g.) available data (e.g., data from different modalities, etc.), crop type, trait of interest, applicable rules and regulations, etc.
[0016] like Figure 1 As shown, system 100 typically includes a crop development cycle 102, which is provided to identify, select, and / or create new agricultural crops, and to select crops to be advanced for further experimentation and characterization, etc. (e.g., and ultimately to be created). Generally, as shown, crop development cycle 102 includes a prediction phase 104, a selection phase 106, and an experimental phase 108, etc. In this example embodiment, crop development cycle 102 is configured through these phases (or periods) to: collect data on crop performance in field 110; store the data in a repository 112; predict the varietal performance of the crop (e.g., when it involves at least one phenotypic trait, etc.) by computing device 114 (in prediction phase 104) (e.g., based on its genetic improvement, the selected environment, etc.); and further select among new varieties within crop development cycle 102 (in selection phase 106) by computing device 112 through physical testing in field 110 (in experimental phase 108).
[0017] Generally, crops are associated with different genetic compositions, and thus changes or improvements to the crop's gene sequence define the difference between one variety and another. Modifications to the crop's gene sequence define additional varieties in order to improve crop performance, for example, based on yield, environment, disease resistance, etc. Crop development cycle 102 is configured to be employed in intelligent modification of gene sequences to improve the performance of one or more crops.
[0018] Crop development cycle 102 can be for various types of crops, which can be any suitable type of plant, etc. Crops can be of the same type (e.g., corn or maize) or can be different types or varieties, etc. Generally speaking, crops (or the plants of crops) can include, but are not limited to, soybeans ( Glycine max ),cotton( Gossypium hirsutum ),peanut( Arachis hypogaea ),barley( Hordeum vulgare );oat( Avena sativa Orchard grass ( Dactylis glomerata ); rice Oryza sativa (including indica and japonica rice varieties); sorghum ( Sorghum bicolor );sugar cane( Saccharum sp ); tall fescue ( Festuca arundinacea ); Lawngrass species (e.g., species: Agrostis stolonifera, Poa pratensis, Stenotaphrum secundatum etc.); wheat ( Triticum aestivum ) and alfalfa Medicago sativa ), Brassica ( Brassica Members of the genus *Broccolis* include broccoli, cabbage, cauliflower, rapeseed and canola, carrots, bok choy, cucumbers, dried beans, eggplant, fennel, green beans, gourds, leeks, lettuce, melons, okra, onions, peas, peppers, squash, radishes, spinach, zucchini, sweet corn, tomatoes, watermelons, honeydew melons, cantaloupes and other melons, bananas, castor beans, coconuts, coffee, cucumbers, etc. poplar Southern pine, radiata pine, Douglas fir eucalyptus Apples and other tree species, oranges, grapefruits, lemons, limes and other citrus fruits, clover, flaxseed, olives, palm trees, chili , black pepper and sweet Peppers, beets, sunflowers, sweetgum, tea, tobacco, and other fruits, vegetables, tubers, and root crops.
[0019] The repository 112 in system 100 is populated with historical data on various crops and their different varieties. Generally, repository 112 includes genotype data, weather data, soil data, management data, and phenotypic data.
[0020] In conjunction with this, historical data is compiled, at least in part, based on the varieties of crops planted, grown, evaluated, and harvested from field 110. Field 110 is suitable for growing one or more different types of crops. Field 110 may include dozens, hundreds, thousands, or more or fewer plots of land, individually or collectively covering dozens, hundreds, thousands, or more or fewer acres (e.g., they may have any suitable size, etc.). Thus, an individual plot of land in field 110 may be less than one acre in some examples (e.g., it may be referred to as a plot, etc.), or more than one acre in other examples, or even more than several acres in yet another example. Field 110 may be located outdoors and exposed to natural conditions (e.g., weather, etc.), or it may be located indoors and controlled by planned conditions. Thus, field 110 may include growing space for any different types or sizes of plants (and / or crops).
[0021] Additionally, historical data can include data from one year or several years (e.g., Y1, Y2, Y3... YN, where N is an integer, etc.). In one example, historical data includes: three years or three growing seasons of data for a specific corn variety, five years of data for another corn variety, and only one year of data for yet another corn variety. Historical data can be organized by variety or crop, crop type, year, etc.
[0022] Specifically, in repository 112, genotype data may include a specific representation of features or markers of one or more DNA sequences of a crop variety. Markers may include any suitable number of base pairs from one or more sequences, and, depending on the specific variety, the one or more sequences may be represented as a specific vector or matrix. Thus, genotype data is a vector expressing the sequence as values, where each value represents a marker in the nucleotide sequence. In this example, the vector contains thousands of markers from male contributors and thousands of markers from female contributors of the variety, where positions in the vector indicate specific markers, and values in the vector indicate what the plant (or crop) represents of that particular sequence. Alternatively, genotype input may include numerical or categorical values representing genomic features (genomic sequences, genes, genetic alleles, cis-regulatory elements, RNA-coding sequences, etc.), and associated genomic metadata such as gene ontology (GO terminology), annotations, expression values, etc. The genotype data for the variety is generally sufficient to reproduce, to a certain extent, the sequence of plant 106 represented by the vector.
[0023] As indicated above, plant (or crop) varieties (during experimental phase 108) are planted and grown in field 110. These varieties are also measured or evaluated (or tested) against one or more types of phenotypic data for the crop varieties in field 110 during the growing season, at harvest, or at other times. Phenotypic data may be specific to plant / crop traits, where traits of interest may include, but are not limited to: size and / or robustness (e.g., plant height, ear height, lodging resistance, sustainability, stem circumference, stem strength, etc.), yield, maturity, resistance to stress (e.g., disease resistance or insect resistance, etc.), resistance to abiotic stresses (e.g., drought resistance or salt tolerance, etc.), growing climate, or any other suitable phenotypic data, and / or combinations thereof. Traits of interest may additionally or alternatively include, but are not limited to: yield, thousand-grain weight, marketable seed units, average seed pixel area, tassel size and skeletonization, gray spot, anthracnose, Goss's wilt, diploidia, Fusarium wilt, Fusarium head blight, northern leaf blight, brown rust / southern rust, scorch spot, greensnap, moisture content, plant height, lodging, chloride ions, southern stem canker, white mold, sudden death syndrome, soybean cyst nematode, root-knot nematode, Phytophthora blight, iron deficiency chlorosis, frog-eye leaf spot, brown stem rot, maturity, and / or combinations thereof.
[0024] Phenotypic data is associated with specific variety identification data (e.g., crop plant identifiers, etc.), for example, to identify specific phenotypic data to specific varieties and / or specific fields 110.
[0025] As the crop grows in field 110, field 110 experiences various types of weather, from precipitation to sunshine, wind, and other conditions. In this example embodiment, weather data is measured and / or collected to indicate the conditions of the field at or before the crop grows in field 110. Weather data may be collected from field 104 (e.g., via sensors in field 110) or may be collected from a third-party source. The weather data is stored in repository 112. The weather data is associated with identification data of one or more fields in field 110 (e.g., via field identifiers, etc.), for example, to identify specific weather data over time to one or more specific fields 110. For example, weather data may include, but is not limited to: atmospheric pressure (e.g., average, minimum, maximum, etc.), wind (e.g., maximum gust, minimum wind speed, maximum wind speed, average wind speed, wind direction, etc.), precipitation (e.g., precipitation rate, maximum precipitation rate, average precipitation rate, precipitation amount, etc.), solar radiation (e.g., total radiation, maximum net radiation, etc.), cloud cover (e.g., average cloud cover, etc.), snow cover (e.g., coverage, depth, density, etc.), soil temperature at different levels (e.g., levels 1 to 4, etc.) (e.g., average, minimum, maximum, etc.), soil moisture content at different levels, temperature (e.g., minimum, maximum, etc. for each time interval (e.g., day, hour, etc.), dew point temperature (e.g., average, minimum, maximum, etc.), relative humidity (e.g., average, minimum, maximum, etc.), etc.
[0026] Weather data can be expressed in time series over regular or irregular time intervals (typically from the planting date until the observation of the trait of interest (e.g., yield, harvest, etc.)).
[0027] Similarly, the soil in field 110 may experience one or more soil conditions before, during, or after crop growth in field 110. As described above, soil data (e.g., via sensors in field 110, etc.) are measured and / or collected and stored in repository 112. The soil data are associated with identification data for the corresponding field 110 (e.g., via field identifiers, etc.), for example, to identify specific soil data over time to one or more specific fields. For example, soil data may include, but is not limited to, measurements related to organic matter (OM), cation exchange capacity (CEC), pH (e.g., acidity, etc.), sand content, clay content, silt content, available soil moisture content (volume fraction up to the wilting point), and bulk density, which may be captured at one or more discrete times or time intervals before, during, or after the growth of one or more plants / crops in field 110.
[0028] In addition, farmland 110 typically undergoes one or more management practices, which are reflected in management data. Management data may include, but is not limited to: irrigation practices, no-till practices, cover crops, treatment applications (e.g., chemical spraying), planting density, etc. Management data may represent agricultural farmland over a year or several years (e.g., farmland identical or similar to farmland for which weather and soil data are used, or a subset thereof, etc.). As described above, management data is collected and stored in repository 112.
[0029] It should be understood that the aforementioned historical data may be compiled for hundreds or thousands of plants over multiple growing seasons (e.g., three, five, ten, twenty, or more growing seasons). Therefore, repository 112 includes a large amount of data that identifies the specific genetic makeup of these varieties to the field 110 where they are tested, as well as weather and soil data, and also identifies them to plant performance. In this example embodiment, repository 112 includes tens of millions of data points from hundreds of thousands or millions of plants.
[0030] like Figure 1 As shown, in this example implementation, prediction phase 104 includes a multimodal architecture 116. The multimodal architecture 116 is included in computing device 114 and trained to predict at least one phenotypic trait of multiple crops based on multiple modalities of input data. The multimodal architecture 116 is initially trained based on data (e.g., historical data, etc.) included in repository 112.
[0031] Specifically, the training data includes genotypic data of historical crop varieties, weather data, soil data, management data, and phenotypic data. Genotypic data, weather data, and soil data each define one modality of the multimodal architecture 116, and together they define the input to architecture 116. Phenotypic data includes traits of interest, where predictions of the traits of interest are the expected output of architecture 116. Figure 2 The multimodal architecture 116 is shown in more detail below.
[0032] As shown in the figure, architecture 116 includes a genetic data modality 202, a weather modality 204, and a soil modality 206 in the modality layer. Genetic modality 202 may include, but is not limited to, a bidirectional long short-term memory (LSTM) model configured to generate latent genetic feature data; weather modality 204 includes an LSTM model configured to generate latent weather feature data; and soil modality 206 includes a one-dimensional convolutional neural network (CNN) model configured to generate latent soil feature data. As shown, latent weather feature data and latent soil feature data are combined into latent environmental feature data. Then, in this example, latent genetic feature data, latent environmental feature data, and management measure data 208 are combined in feature fusion between the modality layer and the aggregation layer. The aggregation layer includes a feedforward neural network with residual connections configured to generate the trait output of interest by combining latent features from the modality layer.
[0033] In other example implementations, the multimodal architecture 116 may include other suitable techniques, such as, for example, Bayesian neural networks and conformal prediction, deep neural network (DNN) models, recurrent neural network (RNN) models, generative pre-trained (GPT) models, multilayer perceptron (MLP) models, etc., or may include one or more different models per modality or across multiple modalities. U.S. Provisional Application 63 / 465,239, filed May 9, 2023, includes additional variations of the architecture that may be used herein, the entire disclosure of which is incorporated herein by reference.
[0034] It should be understood that, in this example implementation, the multimodal aspect of architecture 116 is innovative, important, and / or critical. By allowing the processing of multiple different modalities of data through different models in the modality layer (as described above), the different models in the modality layer are configurable, customizable, and / or selectable for one or more specific features and / or profiles of the data for that specific modality before aggregation in the aggregation layer. In this way, in this example implementation, different data types (e.g., ...) are evaluated in separate models. , Genetic data, weather data, and soil data, etc., as described in this article, are combined such that the influence of one type of data is not limited or otherwise diminished by other data (e.g., because a specific model was used for specific different data, etc.).
[0035] In summary, such as Figure 2 As shown, architecture 116 also includes a loss function 210, which is used to train architecture 116.
[0036] In this example implementation, the loss function 210 is presented as follows: In the above, D is the entire dataset; G is a vector of genomic representation (or genotype) variables (e.g., SNPs (single nucleotide polymorphisms)) for any arbitrary variety; and Y is the phenotypic observation of interest (e.g., yield) for the corresponding variety of G. Furthermore, Known as likelihood weights, they represent the importance of high-amplitude, low-probability phenotypes (or events), where... It is the probability density of the occurrence of a specific phenotype (given G), and is based on the labels of the training data (e.g., traits of interest, etc.). It is the probability density of the occurrence of a specific genotype G (given G).
[0037] Based on the above, computing device 114 is configured to train the multimodal architecture 116, either entirely or partially, based on training data, such that the multimodal architecture 116 predicts traits of interest and minimizes the loss function. In this example implementation, computing device 114 is configured to input data from the training dataset into each modality of the modalities, i.e. , Each modality model is trained on specific data to target known traits of interest from the training data, thus, in this embodiment, architecture 116 is trained as a fully connected architecture. That is, it should be understood that in other embodiments, or in conjunction with updates based on additional data, the different modalities of architecture 116 can be trained separately and then included in architecture 116. For example, each modality model in the modality layer can be trained separately based on the specific modality and the trait of interest from the data in the training dataset, and then, once trained, the modality models can be combined with an aggregation layer to train the models included in the aggregation layer.
[0038] Consistent with the above, modal layers and aggregation layers can be combined to provide the realization of hidden information features in intermediate, individual modal data.
[0039] Once the multimodal architecture 116 is trained, the computing device 114 is configured to use a reserved portion of the training data (i.e., , A validation dataset is used to similarly partially or entirely validate architecture 116. Based on validation that architecture 116 provides sufficient performance (e.g., based on one or more thresholds, etc.), architecture 116 is stored in repository 112 for use in predicting traits of interest based on data consistent with the input modality of the data. Sufficient performance can be defined based on one or more accuracy levels, biases, etc. (e.g., compared to thresholds, etc.).
[0040] Continue to refer to Figure 1Utilizing a trained (and validated) multimodal architecture 116, computing device 114 is configured to identify new gene sequences for a proposed variety of crop, typically by modifying gene sequences of crop varieties already included in repository 112 (e.g., to achieve one or more desired traits of interest). In conjunction with this, in this example, gene sequence modification can be performed in two ways: through parental hybridization or through genome editing. Parental hybridization produces large segments of inherited and discarded gene sequences, which the model uses to infer and predict associated phenotypes. As an example, in hybrid crops (such as corn / maize), two inbred parent lines will produce predicted deterministic hybrid offspring. Furthermore, any number of generations can be evaluated in newly created inbred lines, which can then be combined to obtain hybrid representations for evaluation. Genome editing provides direct alterations to specific regions of the prediction.
[0041] In this way, computing device 114 is configured to define hundreds of thousands, or millions, or more proposed new varieties of agricultural crops.
[0042] In this example implementation, computing device 114 is configured to input the proposed new variety into a training architecture 116, whereby (configured by architecture 116) computing device 114 is configured to predict traits of interest (e.g., yield, ear height, etc.) of the proposed new variety.
[0043] In this approach, the trait of interest is predicted based on the aforementioned weather, soil, and / or management data, where the data indicates the region of interest for the design of evaluating the new gene sequence. For example, a product concept can be defined as the target of the newly proposed variety. The product concept may include a specific region / location, which can further define the specific weather data, soil data, and management data (e.g., standard agronomic practices for that region / location) to be used as inputs to architecture 116 for predicting the trait of interest. The product concept may further define and / or include additional traits or phenotypic expressions of interest (e.g., ear height, yield, etc.).
[0044] Subsequently, computing device 114 is configured to select a proposed variety from a proposed variety based on an acquisition function, which can select a proposed variety from a proposed variety that provides an exploration of its genetic composition and / or utilizes certain phenotypic data (e.g., enhanced traits of interest, etc.).
[0045] In this example, computing device 114 is configured to generate phenotypic gains for new varieties or genome representations. The potential value of this aspect was evaluated based on the following: in This is the increase in the maximum observed phenotype value compared to the initial maximum observed phenotype before testing future genome representation, and In the genomic domain of interest The total variance or uncertainty of the integrated phenotype. Improving phenotypic gain requires specific “collection” or field testing of the identified new genome representation. For the explicit purpose of optimizing or improving phenotypic gain at the fastest possible rate, the collection function used to create and test new genome representations can be defined as any of the following: or in It is by Figure 2 An ensemble (i.e., multiple implementations) of models predicts the average phenotypic value for a given genotype G across various environments and management measures (e.g., defined by product concepts, etc.). It is used to weigh increases The selected scalar value is chosen to represent the importance of exploring the gene space D. This underscores the importance of the previously described high-amplitude, low-probability phenotype, in addition to Defined by the ensemble of trained models, and This is the prediction variance for a given phenotype G, and identifies undersampled regions in the gene-phenotype space. Acquisition functions can be defined for one or more distinct phenotypes, and then aggregated for multi-object acquisition (e.g., defined by one or more product concepts). The chosen new genome representation to be created is used to optimize the acquisition values. The selected new genome representations are then field-tested to observe phenotypes, which in turn updates the phenotype gain equation, thereby both increasing the numerator and decreasing the denominator.
[0046] Based on the above, phenotypic gain can be understood as a metric that is being optimized or improved, and the acquisition function is used to achieve this optimization or improvement of phenotypic gain.
[0047] Next, in system 100, selected varieties from the proposed varieties are created as seeds to be planted, and then planted in field 110 in experimental phase 108. For example, seeds and / or plants can be modified via conventional techniques (selective breeding, gene manipulation, etc.) to produce plants with the sequence of the proposed variety. In conjunction with this, as an example, genetic modification of plants can include point mutations, insertion / deletion mutations, substitution mutations, frameshift mutations, etc., in the plant genome that produce the desired sequence. Then, the "offspring" seeds of the genetically modified "parental" plant can include the genetic modifications made in the parental plant and can express the desired phenotypic trait. Conventional sequencing techniques (such as polymerase chain reaction (PCR)) can be used to determine whether the offspring seeds or the offspring plants produced from said offspring seeds contain the modified sequence. Qualitative observation can also be used to determine whether the offspring plants produced from said offspring seeds express the desired phenotypic trait.
[0048] Figure 3 An example computing device 300 that can be used in system 100 is shown. In conjunction with this, computing device 114 and / or repository 112 may include at least one computing device consistent with and / or implemented therein. In conjunction with this, computing device 300 may be uniquely or specifically configured via executable instructions to implement the various algorithms and other operations described herein with respect to multimodal architecture 116. It should be understood that, as described herein, system 100 may include a variety of different computing devices that are consistent with or different from computing device 300. That is, repository 112 and computing device 114 in system 100 may include computing device 300 and / or may be consistent with it.
[0049] Example computing device 300 may include, for example, one or more servers, workstations, personal computers, laptop computers, tablet computers, smartphones, other suitable computing devices, combinations thereof, etc. Furthermore, computing device 300 may include a single computing device, or it may include multiple computing devices located in close proximity or distributed across a geographical area and coupled to each other via one or more networks. Such networks may include, but are not limited to: the Internet, intranets, private or public local area networks (LANs), wide area networks (WANs), mobile networks, telecommunications networks, combinations thereof, or one or more other suitable networks, etc. In one example, the repository 112 of system 100 includes at least one server computing device that is directly and / or coupled to repository 112 via one or more LANs, etc.
[0050] In this context, the computing device 300 shown includes a processor 302 and a memory 304 coupled to (and in communication with) the processor 302. The processor 302 may include, but is not limited to, one or more processing units (e.g., in a multi-core configuration), including a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuitry or processor capable of implementing the functions described herein. The foregoing enumeration is merely exemplary and is therefore not intended to limit the definition and / or meaning of a processor in any way.
[0051] As described herein, memory 304 is one or more devices that enable the storage and retrieval of information such as executable instructions and / or other data. Memory 304 may include one or more computer-readable storage media, such as, but not limited to: dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), solid-state devices, flash drives, CD-ROMs, thumb drives, magnetic tape, hard disks, and / or any other type of volatile or non-volatile physical or tangible computer-readable media. Memory 304 may be configured to store, but not limited to, weather data, soil data, genotype data, latent trait data, phenotypic data, models (e.g., trained models, untrained models, etc.), and / or other types of data (and / or data structures) suitable for use as described herein. In various embodiments, computer-executable instructions may be stored in memory 304 for execution by processor 302 to cause processor 302 to perform one or more functions described herein (e.g., one or more operations included in method 400, etc.), such that memory 304 is a physical, tangible, and non-transitory computer-readable storage medium. Such instructions generally improve the efficiency and / or performance of processor 302 performing one or more of the various operations described herein. It should be understood that memory 304 may include a variety of different memories, each implemented within one or more functions or processes described herein.
[0052] In an example implementation, computing device 300 also includes an output device 306 coupled to (and in communication with) processor 302. Output device 306 outputs or presents information to users of computing device 300 (e.g., breeders, etc.), such as, but not limited to, selected progeny, progeny as commercial products, traits of interest, plant performance metrics, and / or any other type of data desired, for example, by displaying and / or otherwise outputting. It should also be understood that in some implementations, output device 306 may include a display device such that various interfaces (e.g., applications (web-based or otherwise)) can be displayed at computing device 300, and specifically at a display device, to display such information and data, etc. Furthermore, in some examples, computing device 300 may display an interface at a display device of another computing device, such as a server hosting a website with multiple web pages, or interacting with a web application employed at the other computing device. Output device 306 may include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, organic LED (OLED) displays, “e-ink” displays, combinations thereof, etc. In some implementations, output device 306 may include multiple units.
[0053] The computing device 300 further includes an input device 308 for receiving input from a user. The input device 308 is coupled to (and communicates with) the processor 302 and may include, for example, a keyboard, pointing device, mouse, stylus, touch-sensitive panel (e.g., touchpad or touchscreen), another computing device, and / or an audio input device. Furthermore, in some example embodiments, a touchscreen, such as that included in a tablet computer or similar device, may serve as both an output device 306 and an input device 308. In at least one example embodiment, both the output device 306 and the input device 308 may be omitted.
[0054] Furthermore, the computing device 300 shown includes a network interface 310 coupled to and in communication with the processor 302 (and in some embodiments, also coupled to the memory 304). The network interface 310 may include, but is not limited to, a wired network adapter, a wireless network adapter, a telecommunications adapter, or other devices capable of communicating with one or more different networks. In at least one embodiment, the network interface 310 is used to receive input from the computing device 300. For example, the network interface 310 may be coupled to (and in communication with) a field data collection device (e.g., one or more sensors among sensors 114, 116, etc.) to collect data for use as described herein. In some example embodiments, the computing device 300 may include a processor 302 and one or more network interfaces incorporated into or associated with the processor 302.
[0055] Figure 4 An example method 400 for use in promoting the development of traits in agricultural crops is shown. Example method 400 is described herein in conjunction with system 100 and can be implemented, wholly or partially, in the computing device 114 of system 100. Furthermore, for illustrative purposes, reference is also made to… Figure 3 The example method 400 is described using computing device 300. However, it should be understood that method 400 or other methods described herein are not limited to system 100 or computing device 300. And conversely, the systems, data structures / repositories, and computing devices described herein are not limited to example method 400.
[0056] First, plant technicians (or other users) (e.g., breeders, project managers, etc.) identify the plant type (e.g., corn, soybean, etc.) and the traits of interest (e.g., yield, height, etc.). Plant technicians can also define specific environments (e.g., in deployment scenarios, product concepts, etc.) based on the specific goals or varieties of the plant (and / or its corresponding crop) to be developed.
[0057] In this example, the plant technician has decided to develop one or more varieties of maize, and the trait of interest is yield (e.g., where development is related to increasing the yield of maize varieties, etc.).
[0058] Based on this, at 402, in response to input from a plant technician (or other user), computing device 114 accesses data from repository 112. The accessed data includes genotypic data for hundreds, thousands, or more varieties of maize over multiple growing seasons, weather data, soil data, management data, and data on traits of interest. Again, in this example, the trait of interest is yield. Therefore, the accessed data includes, for example, genotypic data indicating specific gene sequences of a variety. Genotypic data may include numerical vectors, where each value indicates a specific sequence at a particular marker in the gene sequence. Markers may be limited to: those markers known to affect yield in a specific crop type, and / or additional markers known to be related to other traits of interest or whose effect on one or more traits of interest is unclear.
[0059] As indicated above, weather data includes time series of weather data indicating one or more specific weather conditions in a particular field (e.g., one or more fields in field 110, other fields, etc.) under which a particular variety produces the trait of interest. Similarly, soil data indicates soil conditions in fields (e.g., one or more fields in field 110, other fields, etc.) under which varieties grow to provide the trait of interest. In addition to genotypic data, weather data, and soil data, management data, if any, may be accessed, indicating management practices (if any) applied to specific fields (e.g., one or more fields in field 110, other fields, etc.) to produce the trait of interest.
[0060] It should be understood that genotype data, weather data, and / or soil data can be processed to allow for and / or enhance compatibility with multimodal architectures.116 In conjunction with this, data can be filtered to include only certain genotype data, or certain weather data (e.g., average temperature, etc.) and / or certain soil data, etc.
[0061] At 404, computing device 114 trains architecture 116. In conjunction with this, in this example, each model in the modal layer is trained individually (e.g., trained individually first and then combined via feature fusion, trained together, etc.). Figure 2 The model for gene modality 202, or the bidirectional LSTM model, is trained based on gene data and traits of interest. Therefore, the model for gene modality 202 is trained to generate latent trait genotype data from gene data, thereby associating latent trait genotype data (other than weather or soil) with specific yields. Similarly, the LSTM model for weather modality 204 is trained based on weather data and traits of interest to generate latent trait weather data from weather data, thereby associating latent trait weather data with specific yields. Furthermore, the one-dimensional CNN model for soil modality 206 is trained based on soil data and traits of interest to generate latent trait soil data from soil data, thereby associating latent trait soil data (other than gene composition or weather) with specific yields.
[0062] It should be understood that, as part of the training process, each trained model from the modal layer can be validated to ensure sufficient performance as defined, for example, by percentage, bias, etc. (e.g., relative to a threshold, etc.).
[0063] In addition, as Figure 4As part of the training in method 400, for example, the trained model is included in architecture 116, through which latent feature data of weather and soil modalities are combined, and then the training dataset is provided to architecture 116 by computing device 114 to train the neural network model of the aggregation layer by combining the trained models of the modal layers. The overall training of architecture 116 relies on the loss function described above to refine the hyperparameters and / or weights of the models included in architecture 116.
[0064] After the neural network model is trained, at 406, computing device 114 validates the overall architecture 116 based on a reserved portion of the training data. At 408, the trained architecture 116 is stored by computing device 114 in memory (such as, for example, repository 112). In this way, architecture 116 is trained to predict the trait of interest (yield in this example) based on inputs of genotype data, weather data, and soil data.
[0065] Next, at 410, computing device 114 typically identifies the gene sequence of the proposed crop variety by modifying the gene sequences of crop varieties already included in the repository 112. Computing device 114 inputs the proposed variety and its specifically associated genotype data, along with appropriate weather and soil data (and management data), into a training architecture 116, whereby at 412, computing device 114 uses the trained architecture to predict the traits of interest for the proposed variety.
[0066] At 414, computing device 114 selects the desired variety from the desired varieties based on the acquisition function described above, etc.
[0067] Next, at 416, computing device 114 guides the proposed variety of the proposed varieties to experimental stage 108, whereby the proposed variety of the proposed varieties is created as a physical plant, which is grown in field 110 and tested to verify and / or evaluate the performance of the traits of interest.
[0068] Method 400 can be repeated from time to time, with or without retraining, to continue exploring new varieties of crops.
[0069] In light of the above, the unique systems and methods described in this paper provide high-level interpretations of traits of interest, which can be used to define specific genetic or environmental characterization maps of plants.
[0070] In particular, the embodiments described herein improve the underlying techniques for interpreting traits of interest by including a modal layer and a aggregation layer, thereby enabling the modeling of a wide range of different types of data (e.g., ,The data is processed specifically using marker-based genetic data, time-based weather data, and depth-based soil data. In this way, accurate and precise prediction of specific traits of interest (e.g., ear height in maize) depends on different data, but unlike conventional single-model or linear models, it does not discard or weaken the influence of one data type on another, where data is combined before being introduced into or input into the model. This helps improve the underlying technology of the architecture presented in this paper and the resulting predictions.
[0071] In this context, it should be understood that, in some embodiments, the functionality described herein can be described using computer-executable instructions stored on a computer-readable medium and executable by one or more processors. A computer-readable medium is a non-transitory computer-readable medium. For example, and not as a limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible by a computer. Combinations of the above should also be included within the scope of computer-readable media.
[0072] It should also be understood that one or more aspects of this disclosure, when configured to perform the functions, methods and / or processes described herein, transform a general-purpose computing device into a special-purpose computing device.
[0073] Based on the foregoing description, it will be further understood that the embodiments described above in this disclosure can be implemented using computer programming or engineering techniques (including computer software, firmware, hardware, or any combination or subset thereof), wherein the technical effects can be achieved by performing at least one of the operations stated in the claims, such as: (a) identifying a plurality of proposed varieties of a crop, wherein each of the plurality of proposed varieties includes a different gene sequence compared to other proposed varieties and known varieties; (b) using a trained model to predict the trait of interest for each of the plurality of proposed varieties based on data included in a repository; (c) selecting a proposed variety from the plurality of proposed varieties according to a phenotypic gain-based acquisition function; and / or (d) directing seeds representing the selected proposed variety from the plurality of proposed varieties to an experimental stage to evaluate the trait of interest for the selected proposed variety from the plurality of proposed varieties.
[0074] Examples and embodiments are provided so that this disclosure will be complete and will fully communicate the scope to those skilled in the art. Numerous specific details, such as examples of specific components, devices, and methods, are set forth to provide a complete understanding of embodiments of this disclosure. It will be apparent to those skilled in the art that specific details are not required, exemplary embodiments may be embodied in many different forms, and neither should be construed as limiting the scope of this disclosure. In some example embodiments, well-known processes, well-known device structures, and well-known techniques are not described in detail. Furthermore, advantages and improvements that may be achieved using one or more example embodiments disclosed herein may provide all or none of the advantages and improvements mentioned above, and still fall within the scope of this disclosure.
[0075] The specific values disclosed herein are merely examples and do not limit the scope of this disclosure. The disclosure of specific values and ranges of values for a given parameter herein does not exclude other values and ranges of values that may be used in one or more instances disclosed herein. Furthermore, it is contemplated that any two specific values of a particular parameter described herein may define endpoints of a range of values that may also be applicable to the given parameter (i.e., the disclosure of a first and a second value of a given parameter may be interpreted as disclosing that any value between the first and the second value may also be applicable to the given parameter). For example, if parameter X is exemplified herein as having a value A and also exemplified as having a value Z, then parameter X is contemplated to have a range of values from about A to about Z. Similarly, it is contemplated that the disclosure of two or more ranges of values for a parameter (whether such ranges are nested, overlapping, or distinct) covers all possible combinations of ranges of values that may be claimed using endpoints of the disclosed range. For example, if the parameter X is exemplified in this paper as having a value in the range of 1-10, or 2-9, or 3-8, it is also conceivable that the parameter X could have other value ranges, including 1-9, 1-8, 1-3, 1-2, 2-10, 2-8, 2-3, 3-10, and 3-9.
[0076] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context explicitly indicates otherwise. The terms “comprises,” “comprising,” “including,” and “having” are inclusive and therefore indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein should not be construed as necessarily requiring them to be performed in the particular order discussed or shown, unless specifically indicated otherwise. It should also be understood that additional or alternative steps may be employed.
[0077] When a feature is described as being “on,” “joined to,” “connected to,” “coupled to,” “associated with,” “communicating with,” or “included” in another element or layer, it may be directly located on, joined to, connected to, or coupled to, associated with, communicate with, or included in another feature, or there may be intervening features present. As used herein, the terms “and / or” and “at least one of” include any and all combinations of one or more of the associated listed items.
[0078] The elements stated in the claims are not intended to be means plus functional elements as defined in 35 USC §112(f), unless the element is explicitly stated using the phrase “means for…” or in the case of a method claim using the phrase “operation for…” or “steps for…”.
[0079] Although the terms first, second, third, etc., may be used herein to describe various features, these features should not be limited by these terms. These terms may only be used to distinguish one feature from another. Unless the context clearly indicates otherwise, terms such as “first,” “second,” and other numerical terms, as used herein, do not imply order or sequence. Therefore, a first feature described herein may be referred to as a second feature without departing from the teachings of the exemplary embodiments described.
[0080] The above description of embodiments has been provided for illustrative and descriptive purposes. This description is not intended to be exhaustive or limiting of this disclosure. Individual elements or features of a particular embodiment are generally not limited to that particular embodiment, but are interchangeable where appropriate and can be used in the chosen embodiment, even if not specifically shown or described. The same can be varied in many ways. Such variations should not be considered as departing from this disclosure, and all such modifications are intended to be included within the scope of this disclosure.
Claims
1. A system for use in interpreting traits of interest in agricultural crops, the system comprising: A computing device, the computing device including a memory and at least one processor; The memory includes executable instructions, a trained predictive architecture, and a repository, which includes genotype data for multiple known varieties, as well as weather and soil data associated with the growth of the known varieties. The at least one processor is configured via the executable instructions and the trained prediction architecture as follows: Identifying multiple proposed varieties of a crop, wherein each of the multiple proposed varieties includes a different gene sequence compared to other proposed varieties and known varieties; The trained model is used to predict the trait of interest for each of the plurality of proposed varieties based on the data included in the repository; The proposed variety is selected from the plurality of proposed varieties based on the acquisition function based on phenotypic gain; as well as Seeds representing a selected variety from the plurality of proposed varieties are guided to the experimental stage to evaluate the trait of interest of the selected variety from the plurality of proposed varieties.
2. The system of claim 1, wherein the at least one processor is configured, via the executable instructions and the trained prediction architecture, to input genotype data representing the proposed variety, together with weather and soil data associated with the specific region for which the proposed variety is designated, in order to predict the trait of interest.
3. The system of claim 1 or claim 2, further comprising a plurality of fields in the experimental phase, wherein the seeds are planted in the plurality of fields.
4. The system of claim 1, wherein the at least one processor is configured via the executable instructions to train the prediction architecture based on a loss function of the associated likelihood of indicative phenotypic observations and high-amplitude, low-probability phenotypic values.
5. The system of claim 1, wherein the prediction architecture comprises a multimodal architecture, the multimodal architecture comprising a first modality dedicated to genotype data, a second modality dedicated to weather data, and a third modality dedicated to soil data.
6. The system of claim 5, wherein each of the modalities comprises at least one of the following: a deep neural network (DNN) model, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative pre-trained (GPT) model, a transformer or language model, a multilayer perceptron (MLP) model, and a long short-term memory (LSTM) model.
7. The system of claim 5, wherein the multimodal architecture further comprises an aggregation layer configured to combine latent features from each of the modalities.
8. The system of claim 1, wherein the aggregation layer comprises a neural network model.
9. A computer-implemented method for use in interpreting traits of interest in agricultural crops, the method comprising: A computing device identifies multiple proposed varieties of a crop, each of which includes a different gene sequence compared to other proposed varieties of the crop and compared to known varieties of the crop. The computing device uses a trained prediction architecture to predict the trait of interest for each of the plurality of proposed varieties; The proposed variety is selected from the plurality of proposed varieties based on the acquisition function based on phenotypic gain; as well as Seeds representing a selected variety from the plurality of proposed varieties are guided to the experimental stage to evaluate the trait of interest of the selected variety from the plurality of proposed varieties.
10. The computer-implemented method of claim 9, wherein predicting the trait of interest comprises: The trait of interest is predicted based on genotype data representing the proposed variety, along with weather and soil data associated with the specific region where the proposed variety is designated for use.
11. The computer-implemented method of claim 9 or claim 10, further comprising: Seeds representing a selected variety from the plurality of proposed varieties were planted in multiple fields during the experimental phase.
12. The computer-implemented method according to any one of claims 9 to 11, further comprising: The prediction architecture is trained using a loss function based on the correlation likelihood between indicative phenotypic observations and high-amplitude, low-probability phenotypic values.
13. The computer-implemented method of any one of claims 9 to 12, wherein the prediction architecture comprises a multimodal architecture, the multimodal architecture comprising a first modality dedicated to genotype data, a second modality dedicated to weather data, and a third modality dedicated to soil data.
14. The computer-implemented method of claim 13, wherein each of the modalities comprises at least one of the following: a deep neural network (DNN) model, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative pre-trained (GPT) model, a transformer or language model, a multilayer perceptron (MLP) model, and a long short-term memory (LSTM) model.
15. The computer-implemented method of claim 13 or claim 14, wherein the multimodal architecture further includes an aggregation layer; and wherein the method further includes combining potential features from each of the modalities through the aggregation layer.
16. A non-transitory computer-readable storage medium comprising executable instructions that, when executed by at least one processor to interpret traits of interest in an agricultural crop, cause the at least one processor to: Identify multiple proposed varieties of a crop, wherein each of the multiple proposed varieties includes a different gene sequence compared with other proposed varieties of the crop and compared with known varieties of the crop; The trained prediction architecture is used to predict the trait of interest for each of the plurality of proposed varieties; The proposed variety is selected from the plurality of proposed varieties based on the acquisition function based on phenotypic gain; as well as Seeds representing a selected variety from the plurality of proposed varieties are guided to the experimental stage to evaluate the trait of interest of the selected variety from the plurality of proposed varieties.
17. The non-transitory computer-readable storage medium of claim 16, wherein the executable instructions, when executed by the at least one processor, cause the at least one processor to: use the trained prediction architecture to predict the trait of interest based on genotype data representing the proposed variety, together with weather and soil data associated with the specific region for which the proposed variety is designated.
18. The non-transitory computer-readable storage medium of claim 16, wherein the executable instructions, when executed by the at least one processor, further cause the at least one processor to: train the predictive architecture based on a loss function of the associated likelihood of indicative phenotypic observations and high-amplitude low-probability phenotypic values.
19. The non-transitory computer-readable storage medium of claim 16, wherein the prediction architecture comprises a multimodal architecture, the multimodal architecture comprising a first modality dedicated to genotype data, a second modality dedicated to weather data, and a third modality dedicated to soil data, and Each of the modalities mentioned therein includes at least one of the following: a deep neural network (DNN) model, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative pre-trained (GPT) model, a transformer or language model, a multilayer perceptron (MLP) model, and a long short-term memory (LSTM) model.
20. The non-transitory computer-readable storage medium of claim 19, wherein the multimodal architecture further comprises an aggregation layer; and wherein the executable instructions, when executed by the at least one processor, use the aggregation layer to cause the at least one processor to combine potential features from each of the modalities.