Creating plants with designed genomes

The BINN addresses the challenges of conventional models by integrating biological domain knowledge to enhance the prediction and design of plant genomes for desired phenotypic traits, achieving accurate and efficient genetic manipulation.

WO2026025076A1PCT designated stage Publication Date: 2026-01-29MONSANTO TECHNOLOGY LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039336
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-07-25
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional computational techniques struggle to effectively model genotype-to-phenotype relationships in plants due to high-dimensional datasets, data variability, and nonlinear interactions, leading to degraded analysis and limited performance in predicting specific phenotypic traits.

Method used

A biology-informed neural network (BINN) is employed, incorporating domain-specific biological knowledge to define connection rules between genotypic and phenotypic data, reducing overfitting and enhancing generalization by limiting tunable parameters and capturing nonlinear interactions.

Benefits of technology

The BINN architecture improves the prediction and design of plant genomes for specific traits by accurately modeling genotype-phenotype relationships, enabling efficient and precise genetic manipulation for improved agronomic traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039336_29012026_PF_FP_ABST
    Figure US2025039336_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Methods arc provided for making a plant genome having a desired nucleic acid sequence. One example method includes identifying genomic segments of a plant that include at least one desired nucleic acid sequence, using a trained neural network, which is biology-informed through first pre-training connections between a genetic data layer and at least one intermediate layer and second pre-training connection between the at least one intermediate layer and a phenotypic layer. The method also includes recombining the at least one at least one desired nucleic acid sequence with a second nucleic acid sequence to form a target plant genome.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 5089-000200-WO-POA CREATING PLANTS WITH DESIGNED GENOMES CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of, and priority to, Greek Patent Application No. 20240100525, filed on July 26, 2024. The entire disclosure of the above- referenced application is incorporated herein by reference. FIELD

[0002] The present disclosure generally relates to methods and systems for use in making plants with genomes that include desired nucleic acid segments that are identified using biology-informed neural networks (e.g., graph neural networks, etc.). BACKGROUND

[0003] This section provides background information related to the present disclosure which is not necessarily prior art.

[0004] Global population growth creates a need to more efficiently grow crop plants. Crop plants that provide a maximum yield using the smallest amount of resources are desirable to achieve the efficient usage of farm inputs. Additionally, extreme weather patterns create the need for plants that maximize yield in extreme climates, or climates that are sporadically extreme. Traditional plant breeding techniques can be used to achieve improvements, however, such techniques can take years to provide minor incremental improvements. SUMMARY

[0005] This section provides a general summary of the disclosure and is not a comprehensive disclosure of its full scope or all of its features.

[0006] Disclosed herein are efficient methods of making plant genomes and plants that exhibit improved agronomic traits. Such methods include the identification of nucleic acid sequences that when combined into a genome create a plant with one or more desired phenotypic traits. The desired nucleic acid sequences can include specific genes, clusters of genes, non- coding regions, clusters of non-coding regions, and combinations thereof. The identification ofAttorney Docket No. 5089-000200-WO-POA the desired nucleic acid sequences is accomplished using a biologically informed neural network (BINN), wherein the biological data used to train the model has at least one known connection (layer) (i.e., is informed) to genetic sequences used in the model. Thus, allowing for efficient use of data and computational resources and the creation of genomes and plants at an accelerated pace.

[0007] The methods of making a genome described herein include a step of identifying a genomic segment in a plant that includes a desired nucleic acid sequence that has been identified using a BINN. The plant that has the identified genomic segment can be completely distinct from any of the plants, or plant genomes, that were used to train, or were used during the analysis by the BINN algorithm. One of ordinary skill in the art will appreciate that the genomic segment can be found through any method known in the art, including through sequencing and analysis of plants that are distantly related to those used in the BINN, for example, non-domesticated related crop plants. Accordingly, the methods described herein include recombining the identified genomic segment with other nucleic acid sequences to form a plant genome. It is also appreciated that the genomic segment that is identified is found within the genomes of a breeding program, and in some instances the breeding program includes the population of genomes that were analyzed using the BINN.

[0008] As mentioned, the methods described herein can include a step of recombining the at least one desired nucleic acid sequence with another sequence. The recombining step can include the use of multiple technologies, for example, molecular biology techniques used for genomic manipulation, such as gene editing, homologous recombination and the like. The recombining step can also include the use of standard breeding practices, including marker assisted breeding and the like (Khandagale and Nadaf, Plant Biotechnology Reports, 10, (2016)). The breeding practices can be aided by precision breeding algorithms, such as described by Miller et al., Frontiers in Plant Science, May 2023, and Alpha sim-R (Gaynor, et al., 2021), and the like.

[0009] Data sets used to train the algorithm will include data that is connected to nucleic acid sequences that are connected directly to, or indirectly through multiple biologically informed nodes to, the phenotype (P). Indirect connections for example are connections such as nucleic acid sequences associated with metabolite formation and metabolite data connected to a desired P. The P can be any measurable observation of a plant. The measurable observation canAttorney Docket No. 5089-000200-WO-POA be for example an economically important agronomic trait such as yield, drought tolerance, stature, disease resistance, pest resistance, and the like. The P can also be a trait that is fundamental to developing plant biotechnology. For example, P can be flower morphology, transformability (the ability to be genetically manipulated), sterility, haploid induction, haploid doubling, self-incompatibility, and the like. The actual data can be RNA expression data, metabolomics data, protein conformational data, enzymatic activity data and combinations thereof. Moreover, such data can include the element of time to allow the BINN to identify desired genomic sequences that are active or inactive during various developmental stages of plants. Data sets can also include data derived from disparate tissues such as roots and leaves.

[0010] One of ordinary skill in the art will appreciate that when molecular biology techniques are used during the recombination step, transformation of plant cells may be necessary. Any method of transforming cells can be used, for example, electroporation, biolistics, agrobacterium mediated, viral mediated and combinations thereof. Transformation can be transient, meaning that a DNA sequence, such as a plasmid, is temporarily introduced into a plant cell to allow for the delivery of proteins through expression and nucleic acid sequences. Transient expression allows for editing components and guide RNA molecules to be introduced into the cell to edit the genome. Transformation can also be used to introduce permanent genomic changes into the cell genome. The nucleic acid sequences that are introduced into the genome can be; the same as the nucleic acid sequences found elsewhere within the same genome, nucleic acid sequences found in a different cell from a distinct plant of the same genus, nucleic as sequences that are completely not from the same genus and combinations thereof. The nucleic acid sequences introduced can include sequences that encode proteins, for example enzymes, structural proteins, storage proteins and the like. The coding sequences can be found in engineered cassettes that express open reading frames, whole genomic sequences that are introduced (through transformation and / or breeding) into the genome and combinations thereof.

[0011] Plants made to incorporate the desired nucleic acid sequences can include additional features, such as the addition of, or rearrangement of, regulatory elements such as promoters, leaders (also known as 5’ UTRs), enhancers, introns, and transcription termination regions (or 3´ UTRs) which play an integral part in the overall expression of genes in living cells. The term “regulatory element” (or “expression element”) as used herein, refers to a DNA polynucleotide having gene regulatory activity. The term “gene regulatory activity,” as usedAttorney Docket No. 5089-000200-WO-POA herein, refers to the ability to affect the expression of an operably linked transcribable DNA polynucleotide, for instance by affecting the transcription and / or translation of the operably linked transcribable DNA polynucleotide. Regulatory elements, such as promoters, leaders, enhancers, introns and 3´ UTRs that function in plants are useful for modifying plant phenotypes through genetic engineering. These additional features can be used to alter one or more of the desired nucleic acid sequences identified by the BINN.

[0012] Sensitivity analysis can be used to identify additional data for use as intermediate data. Sensitivity analysis can be done on the nodes in the model itself. For example, an existing BINN can have data extracted from it to determine the magnitude of the change that is caused by the loss of the data. If the BINN changes dramatically, there is more sensitivity associated with the removed data. Stated another way, sensitivity analysis can be performed by systematically perturbing or ablating individual pathway nodes in a trained BINN, be zeroing out or masking its intermediate activations, and then measuring the change in the model’s predictions relative to the unperturbed baseline. The larger the shift in output, the more sensitive and thus the more important that pathway is for the final prediction. This sensitivity score allows one to rank pathways by their impact and pinpoint (identify) what intermediate data would most benefit from additional measurement or refinement.

[0013] One of ordinary skill in the art will appreciate that nodes can be associated with genes, metabolites, expression levels, protein structures, and the like. In instances where a node is identified as sensitive, for example removal of the node causes a shift that exceeds a certain threshold, then that node indicates biological relevance for the P. Such a threshold can be chosen relative to the noise or variability of the model’s predictions, a percent change, e.g., 5%, or a top-k ranking, i.e., rank all nodes and focus on the 10% most sensitive. If a particular class of data, for example data associated with a gene, that data is a target for manipulation to impact the desired phenotype.

[0014] Conversely, if data is removed and no significant change occurs, however, such data is believed to be important for the biological phenotype in question, it could indicate that the data quality, quantity or a combination thereof needs to be altered. For example, if biological data is known that shows that a P is lost or significantly changed when a plant cell that includes a knock-out (removal of the gene) of a gene is tested, but yet in the BINN model theAttorney Docket No. 5089-000200-WO-POA node associated with that gene is not sensitive the data underlying the BINN can be analyzed and altered, supplemented or combinations thereof.

[0015] Further areas of applicability will become apparent from the description provided herein. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure. DRAWINGS

[0016] The drawings described herein are for illustrative purposes only of selected embodiments, are not all possible implementations, and are not intended to limit the scope of the present disclosure.

[0017] FIG. 1 illustrates an example system of the present disclosure suitable for indicating phenotypic data for a plant, based on specific genotypic data for the plant, through use of a biology-informed network;

[0018] FIGS. 2A-C illustrate multiple different embodiments of networks, which may be implemented in the system of FIG. 1, where each relies on different data to define intermediate layer(s) of the network;

[0019] FIG. 3 is a block diagram of an example computing device that may be used in the system of FIG. 1;

[0020] FIG. 4 illustrates an example method for indicating phenotypic traits based on genotypic data, which is suitable for use with the system of FIG. 1; and

[0021] FIG. 5 is a graphical representation of mean squared error (MSE) for different models applied for indicating phenotypic data for plants (including, for example, a penalized linear regression (RR), a fully-connected network (FCN), a biology-informed neural network (BINN) trained with standard MSE (BINN MSE), and a custom BINN with soft-constraint loss (BINN soft)); and

[0022] FIGS. 6A-6H include graphical representations illustrating relative performance (based on Spearman correlations) of an example BINN model as described herein an other models, in connection with flower-time phenotypes (Anthesis.sp.NE in FIGS. 6A and 6B, Anthesis.sp.MI in FIGS. 6C and 6D, Silking.sp.NE in FIGS. 6E and 6F, and Silking.sp.MI in FIGS. 6G and 6H), across multiple different genetic populations (including an aggregate of all populations (identified as “all” in the figures), a stiff stalk (SS) heterotic population, a non-stiffAttorney Docket No. 5089-000200-WO-POA stalk (NSS) heterotic population, an iodent (IDT) heterotic population, a popcorn population, a sweet corn population, a tropical population, and an others population (e.g., a mix of uncategorized lines, etc.)).

[0023] Corresponding reference numerals indicate corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION

[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. The description and specific examples included herein are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure.

[0025] Various computer-based approaches for linking genotypic data to observable traits or phenotypes of plants have been explored by means of a range of analytical techniques. Yet, the conventional computational techniques encounter challenges when addressing complexities present in biological systems of the plants, which presents, for example, as high- dimensional datasets, data variability, etc. Conventional computational techniques fail to overcome these challenges, and provide degraded analysis in identifying specific phenotypes for specific genotypic data of the plants.

[0026] In connection with the above, effectively modeling genotype-to-phenotype relationships is substantially hindered through the deficiencies in addressing the above complexities. Conventional modeling, such as Genomic Best Linear Unbiased Prediction (GBLUP) techniques, depend on linear models that assume additive genetic effects and do not adequately address complex nonlinear interactions in the biological systems of plants. Although deep learning models are an option to address nonlinearities, performance is limited, again, by the high dimensionality of genotypic data, sparse datasets, and unpredictable noise present in biological systems. Additionally, widely used machine learning architectures, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), are not well- suited for biological applications because such models are limited in capturing long-range genetic interactions and also fail to incorporate domain-specific biological insights.

[0027] Uniquely, however, the systems and methods herein define a novel approach to address the above challenges by imposition of biological domain knowledge into the neural network, through specific bio-informed connection rules for one or more intermediate dataAttorney Docket No. 5089-000200-WO-POA layer(s), whereby the neural network is built consistent with the connection rules, and then trained, to provide a specific biology-informed neural network (BINN).

[0028] In particular, biological domain knowledge is ingested and used to initially define one or more intermediate layers (e.g., gene ontology, gene expression, metabolic pathways, etc.). The domain knowledge is leveraged to define connection rules for nodes of the neural network between the genotypic data for a plant and an intermediate layer, and then also between the intermediate layer(s) and the phenotypes of the plant (and between the same or additional intermediate layers, as needed), etc. Each node is a sub-network of any suitable type (e.g., a neural network, etc.) with a single output, where the biological domain knowledge defines the connections between certain ones of the nodes based on dependencies. The architecture of the network is then built consistent with the connection rules and trained to predict, select, or design genomes for one or more specific phenotypic traits of the plant. In connection therewith, by explicitly defining the connections through rules, the architecture reduces the number of tunable parameters within the neural network, which limits overfitting and / or enhances the model’s ability to generalize from sparse and noisy training datasets, to provide enhanced performance.

[0029] In this manner, the systems and methods herein define a technical solution to the technical problems associated with prediction, selection, and design of genomes for specific phenotypic traits. As such, what is described herein provides a clear deviation over conventional techniques, whereby the present disclosure uses limited rules in a process specifically designed to achieve an improved technological result in known techniques in plant technology.

[0030] FIG. 1 illustrates an example system 100, in which one or more aspects of the present disclosure may be implemented. Although, in the described embodiment, parts of the system 100 are presented in one arrangement, other embodiments may include the same or different parts arranged otherwise depending, for example, on availability of data, particular crops / plants, types of relational data, privacy rules and regulations, etc.

[0031] As shown in FIG. 1, the system 100 generally includes a platform computing device 102, which is provided for embedding biological domain knowledge into one or more models (e.g., a neural network, etc.) to define a biology-informed model architecture.

[0032] In this example embodiment, the platform computing device 102 is coupled in communication (as indicated by the arrowed lines in FIG. 1) with multiple databases, including:Attorney Docket No. 5089-000200-WO-POA a database 104, which includes various phenotypic datasets for various plants and / or years; a database 106, which includes genetic information for plants of interest; and a database 108, which includes domain knowledge associated with the specific plants of interest, annotations, etc.

[0033] It should be appreciated that the databases 104-108 may be included, in whole or in part, in the platform computing device 102, or may be sperate and / or standalone database computing devices, etc.

[0034] In general, the phenotypic database 104 includes, without limitations, data representative of specific phenotypes, such as measurements or observations of specific traits from various growing spaces (e.g., field trials, commercial production, etc.). The data may include any suitable information indicative of the genetics of specific plants and the expression of traits in the growing spaces, e.g., grain yield, disease resistance, plant-height, and ear-height, etc. It should be understood that other phenotypic data may be included in the database 104, especially, phenotypic data for a phenotype of a plant to be predicted, etc. Further, the genetic information database 106 may include, without limitation, genotypic data, including single nucleotide polymorphisms (SNPs) and other genetic markers relevant to the specific plants for which the architecture 110 is to be trained.

[0035] Additionally, the domain knowledge repository database 108 includes biological information which is interpretable as one or more causal links between genetic data and phenotypic traits, including, without limitations, gene expression level data, genome assemblies, known metabolic pathways, functional annotations, gene ontology resources, including, for example, GO terms, etc., and other biological insights that inform the modeling process. The biological insights indicate, among other things, a qualitative description of the plants and processes that regulate plants. As such, the biometric domain knowledge may instruct, for example, what genes (i.e., input data) are expressed under different conditions, what functions the genes perform for the organisms and with what other genes with which the genes interact, as well as what large-scale biological processes the genes influence.

[0036] It should be appreciated that the databases 104-108 may also include any suitable data such as (or related to) intermediate data for use in compiling connection rules for intermediate layer(s).Attorney Docket No. 5089-000200-WO-POA

[0037] That is, the intermediate layer(s) may relate specifically to, for example, transcription, translation, protein folding, enzyme kinetics, metabolism, etc. The data may be indicative of gene expression, RNA stability, codon usage, functional domain, gene / protein family, impact of amino acid change, protein shape, protein complexity, protein-protein interactions, catalysis rate, enzyme specificity, gene pathways or circuits, microscopy, hyperspecttrometry, report assays, imagery, or other suitable data, etc. As such, in this example embodiment, the biological domain knowledge is represented as specific details indicative of intermediate data included in the databases 104-108, etc.

[0038] While the data described above is illustrated as being included in three different, separate databases (e.g., databases 104-108, etc.), it should he appreciated that the data may be included in any suitable number of databases. For example, the genetic information database 106 may be integrated with the domain knowledge repository database 108 in one or more embodiments.

[0039] Regardless of combinations and / or locations, the platform computing device 102 is configured to access the data from the databases 104-108. In connection therewith (e.g., upon accessing the data, etc.), it should be understood that the platform computing device 102 may be configured to perform one or more types of data preprocessing, such as normalization, filtering, and quality control, to ensure that the data is suitable for further analysis as explained herein.

[0040] In this example embodiment, the platform computing device 102 is configured to build and train a neural network to predict, select, and / design specific genotypic data for a specific phenotype.

[0041] With reference again to FIG. 1, the system 100 includes a general architecture 110, which is representative of the neural network to be built and trained (e.g., the BINN, etc.). As shown, the architecture 110 includes genotypic data as an input layer (illustrated at the top of the architecture 110), where A, T, C, G represent specific nucleotide bases at specific locations in the genome of a subject plant, and phenotypic data as an output layer (illustrated at the bottom of the architecture 110), where P is representative of the phenotypic trait(s) of the plant with the specific genes. The general architecture 110 includes a single intermediate layer, which is specific to an intermediate data type. Data specific to the intermediate layer is accessed from the databases 104-108 and leveraged to compile one or more specific connection rules between theAttorney Docket No. 5089-000200-WO-POA input layer and the intermediate layer, and also between the intermediate layer and the output layer. The connection rules define specific dependencies of the nodes to be connected. It should be appreciated that each node (i.e., each illustrated circle in the intermediate layer) is a sub- network, which includes a suitable dedicated neural network (e.g., convolutional neural network (CNN), fully connected neural network (FCNN), etc.), which is representative of the logic to be trained to provide an output to the next layer. The single intermediate layer architecture 110 as shown in FIG. 1 may be suited where the type of intermediate data is gene expression or gene ontology, etc., as explained below.

[0042] In addition, as shown in FIG. 1, the intermediate layer in the architecture 110 may be implemented as a different number of layers, which may include the same or a different configuration. FIG. 1 illustrates three specific variations of the architecture 110, referenced 110a, 110b, and 110c. The architecture 110a, for example, includes multiple intermediate layers representative of the same or different types of intermediate data, where the number of layers is 2, 3, 5, or more or less, etc. The architecture 110b, for example, is a staggered architecture, where the unique data from a specific type of intermediate data is included at different layers. And, the architecture 110c, for example, is a parallel architecture, where different types of intermediate data are used, but the layers do not explicitly interact with one another (e.g.,. provide inputs to or receive outputs from, etc.).

[0043] It should be appreciated that, from the illustration of the architecture 110, various different combinations of intermediate layers may be employed in still other embodiments, where the intermediate layer(s) may be stacked or not, and interact, or not, between the input layer (genetic data) and the output layer (phenotypic trait(s)). The number and relative position of the intermediate layers may be based on the type of intermediate data to be employed, for example.

[0044] As it relates to the above, the intermediate layer(s) leverage the domain knowledge understanding that intermediate quantities H0,i are impacted by specific parts of the genome, and that the intermediate quantities affect the phenotypic pathways Hm,i. This means that the biological system of the plant can be represented as a connection graph with Hm,i being the nodes and operators Fm,i carrying information regarding edges of the graph. This provides the following expression:Attorney Docket No. 5089-000200-WO-POA Fm,i→ V (Hm,i) = {Hm−1,cm,i,1, ..., Hm−1,cm,i,Km,i},where V (Hm,i) corresponds to the edges connecting Hm,ito intermediate quantities Hm−1,j, of layer m- affect quantity Hm,i. Under the above approximation, the biological system of the plant is represented through the expression: H0,i← {Gk}k∈C0,i, ∀H0,i∈ H0,Hn,i← V (Hn,i), ∀Hn,i∈ Hn, and 1 ≤ n ≤ N, P ← {HN} where {Gk}k∈C0,iare the genes associated with H0,i. The representation is then used to construct a computational neural network architecture to mimic qualitative associations between intermediate quantities in the domain knowledge for the plant. The intermediate nodes are considered as latent-space sub-networks of the neural network of the architecture 110, which are connected by suitable connection rules compiled from the domain knowledge, i.e., edges. The sub-networks specific to the intermediate variables or connections thereto are then trained, via tunable parameters thereof.

[0046] In this way, the architecture 110, for example, is embedded with domain knowledge, and thus, is biologically-informed. By doing so, the biology-informed architecture includes significantly fewer parameters to be trained (as compared to, for example, a genotype-phenotype fully-connected neural network), since connections are dictated by the approximations of operators Fm,i. Therefore, potential for overfitting, as seen in conventional models, due to overparameterization compared to the number of available data, is limited. That is, the architecture relies on inductive bias provided by embedding biological priors to steer learning toward meaningful genotype-phenotype relationships and enhances both accuracy and generalization. Additionally, the BINN architecture (e.g., architecture 110, etc.) employs sub-networks at the nodes (e.g., fully connected neural network (FCNN), etc.), which boosts expressivity (trading off direct interpretability) and enables the neural network to learn potential nonlinear interactions that the one-dimensional connections cannot capture. This is valuable when allelic combinations within a locusAttorney Docket No. 5089-000200-WO-POA contribute non-additively to the phenotype. Furthermore, a final integrator network which fuses the sub-network outputs, as illustrated herein, can capture global, higher-order cross- gene interactions.

[0047] Consequently, the computing device 102 is configured to, based on one or more inputs from the user, and the associated domain knowledge data for the intermediate layer(s), to define connection rules based on the intermediate data, between the input layer and the output layer. The connection rules predefine dependencies between input, output, and intermediate layers in the network whereby the architecture 110 is biology-informed.

[0048] The computing device 102 is configured to then build the architecture 110 consistent with the connection rules, whereby nodes are connected to certain nodes, etc. The computing device 102 is configured to train the architecture 110 based on training data from the databases 104-108.

[0049] In one specific example, as shown in FIG. 2A, an example architecture 200 is contemplated to rely on gene expression as the type of intermediate layer. That is, a number of gene expressions are defined as nodes of the architecture 200, each designated as one of e1-e5. Based on the domain knowledge from the databases 104-108, which may include hundreds, thousands, ten of thousands of experiments, or more, etc. for gene expressions, rules are compiled to define connections between, for example, the gene expression e1 (intermediate layer) and the phenotypic data P, and also between the gene expressions e2 and e4 and the phenotypic data P, as shown. There are, however, no connections between the gene expressions e3 and e5 and the phenotypic data P, in the rules. The rules are then compiled to define connections between the genes (as labeled g1-g…) and gene expressions (e1-e5, in the intermediate layer). As shown, for example, the genes g2, g3, and g10 are understood from the gene expression literature / experiments to impact the gene expression e2. Similar links are shown for other ones of the genes.

[0050] It should be understood that the process of compiling connection rules may be automated, in whole or in part, whereby user interactions and / or input is relied on by the platform computing device 102 to compile the connection rules.

[0051] Once the connection rules are compiled, the platform computing device 102 is configured to build the architecture 200, for example, consistent with connection rules and then to train the architecture 200 with the accessed training data (e.g., trial data indicative of theAttorney Docket No. 5089-000200-WO-POA specific phenotypic data for the genotypic data, etc.). In this specific embodiment, a loss function (as illustrated in FIG. 2A) is also defined for a mean-squared error (MSE). Hard and soft loss functions (as illustrated in FIG. 2A) are also defined, which rely on the MSE and also a weighted term λ to impose performance criteria related to known intermediate values from the training data (e.g., known gene expression links, etc.). In this example, the “hard” loss function imposes an MSE term on the intermediate values, and the “soft” loss function relies on a correlation constraint on the intermediate values. In this way, the architecture 200 as shown in FIG. 2A is built being biology-informed (through use of the gene expression intermediate layer) and trained to output the phenotypic data for a given input genotypic data for a plant.

[0052] FIG. 2B illustrates a different embodiment, in which the connection rules are compiled based on metabolic pathway data from the domain knowledge repository database 108. That is, the metabolic pathways are employed as the intermediate data in a staggered architecture configuration, which is shown as the architecture 202, to predict time to bud outgrowth, in this example. As shown, the domain knowledge, for example, in the form of literature, indicates specific mathematical expressions for Cytokinins (CK), Auxins (A), Sucrose (S), and Strigolactones (SL). The mathematical expressions are not retained, but the dependencies are used to compile the connection rules. That is, the top mathematical expression in FIG. 2B indicates that Cytokinins (CK) is dependent on Auxins (A) and Sucrose (S), and also Cytokinins (CK). As such, connection rules include a sub-network from the genotypic data to Sucrose (S) and a sub-network from the genotypic data to Auxins (A), where the outputs of the sub-networks along with the genotypic data for Cytokinins (CK) feed into a sub-network for Cytokinins (CK). Similar connection rules are compiled for Strigolactones (SL), and the output of Cytokinins (CK) and Strigolactones (SL) as well as Sucrose (S) feeding into the phenotypic sub-network to define the phenotypic data P.

[0053] Based on the connection rules, the platform computing device 102 is configured to build the architecture 202 shown in FIG. 2B, which is thereby biology-informed, and to then train the architecture 202 based on the accessed training data (e.g., trial data indicative of the specific phenotypic data for the genotypic data, etc.) to output the phenotypic data P for a given input genotypic data (e.g., nucleic acid combination, etc. as indicated in the input portion of the architecture 202) for a plant. In connection therewith, the platform computing device 102 is configured to use any suitable loss function to train the architecture 204.Attorney Docket No. 5089-000200-WO-POA

[0054] FIG. 2C illustrates a still different embodiment, in which the connection rules are compiled based on gene ontology (GO) to provide specific dependencies between genotypic data (SNPs) and phenotypic data (P) (e.g., yield, ear height, plant height, etc.). The architecture 204 includes multiple layers of GO, which, as shown, include connections from genes to SNPs and from the SNPs to the GO, and to the phenotype. The connection rules are compiled to represent each dependency indicated in the domain knowledge and / or literature for gene ontology, the specific plant and specific phenotype.

[0055] As above, once the connection rules are compiled, the platform computing device 102 is configured to build the architecture 204 consistent with the connection rules and to then train the architecture 202 based on the accessed training data (e.g., trial data indicative of the specific phenotypic data for the genotypic data, etc.) to output the phenotypic data for a given input genotypic data for a plant. In connection therewith, the platform computing device 102 is configured to use any suitable loss function to train the architecture 204.

[0056] Referring again to FIG. 1, the platform computing device 102 is configured to employ the trained BINN architecture 110, for example, to predict, select or design genomes for one or more plants. That is, the platform computing device 102 is configured to select one or more of the genomes based on the predicted phenotypic traits, from the trained BINN architecture 110, for the genomes. The platform computing device 102 is further configured to output the selected one or more genomes to a breeding pipeline process 112, whereby the selected one or more genomes are constructed, planted in the ground, grown, and tested, etc.

[0057] In particular, for example, the trained BINN architecture 110 herein may be used in conjunction with a simulation of the genetic progeny of theoretical crosses to determine the composition of genomic elements that will perform optimally for traits of interest. Following the prediction of the most likely genomes that will result from a given cross of plant parents, the genomic representation of those genomes may be input into the trained BINN architecture to generate a prediction of performance for a given phenotypic trait. Upon determination of the optimal combination of genomic elements for each phenotypic trait of interest following the predictions, a crossing schema may be established, and during the commencement of the crossing schema, genotyping may be conducted. These genotypes may be fed into the trained BINN architecture 110, and once a predictedAttorney Docket No. 5089-000200-WO-POA performance threshold is met, the individual genotype may be considered as eligible as a product or for inclusion into further plant breeding pipeline operations. Furthermore, the trained BINN architecture 110 may be used to interrogate the performance of novel genotypes that cannot be achieved via crossing such as gene editing or transformation. Upon prediction of the specific novel genotype by the trained BINN architecture 110, the genome may be created via genome editing, transformation, or some combination thereof that could include crossing.

[0058] In connection with the above, especially as it relates to designing genomes that will perform optimally for a given trait or combination of traits, the analysis continues beyond the prediction of the phenotype, via the trained BINN architecture 110. In particular, in various embodiments, the design process takes into consideration one or more environmental impacts that the genomes are likely to encounter as the expression of any given phenotype is a function of its genotype by the discrete environment within which it is placed. Thus, after a set of theoretical genomes have been passed through the trained BINN architecture 110 and predictions are made, a subset of the highest predicted theoretical genomes is selected for further assessment by an environmental model. The environmental model accepts the above genomes as input but makes predictions in one or more environments, which may include: static features such as, for example, soil type, topography, latitude, longitude, altitude, etc.; time-series features such as rainfall, temperature, humidity, wind gusts, solar radiation, etc.; and management practices; etc. An example environmental model is described in Applicant’s US Pat. Appl. Publ. No. 2024 / 0378354 (published on November 14, 2024).

[0059] The above further analysis, via the environmental model, in combination with the BINN architecture 110, allows for the identification of potential performance and / or potential performance risks with regard to the genomes (e.g., theoretical genotypes, etc.), in general, and further in interacting with its potential environment. Following these predictions, a further subset of genomes is selected for creation via recombination, editing, transformation, or some combination thereof, as explained more fully herein.

[0060] FIG. 3 illustrates an example computing device 300 that may be used in the system 100, etc. In connection therewith, the computing device 102 of the system 100 includes one or more computing devices at least partially consistent with computing device 300. The computing device 300 may be configured, by executable instructions, to implement the variousAttorney Docket No. 5089-000200-WO-POA algorithms and other operations described herein with regard to the computing device 102. It should be appreciated that the system 100, as described herein, may include a variety of different computing devices, either consistent with computing device 300 or different from computing device 300.

[0061] The example computing device 300 may include, for example, one or more servers, workstations, personal computers, laptops, tablets, smartphones, other suitable computing devices, virtual workspaces, combinations thereof, etc. In addition, the computing device 300 may include a single computing device, or it may include multiple computing devices located in close proximity or distributed over a geographic region, and coupled to one another via one or more networks. Such networks may include, without limitations, the Internet, an intranet, a private or public local area network (LAN), wide area network (WAN), mobile network, telecommunication networks, combinations thereof, or other suitable network(s), etc. In one example, the computing device 102 and / or the databases 104, 106, 108, and / or the architectures 110, 200, 202, 204 may include and / or may be implemented in at least one computing device consistent with the computing device 300. In addition, the databases 104, 106, 108, and / or the architecture 110 of the system 100 may each include (and / or be implemented in) at least one server computing device, while the computing device 102 includes at least one separate computing device, which is coupled to the databases 104, 106, 108, and / or the architecture 110, directly and / or by one or more LANs, etc.

[0062] With that said, the illustrated computing device 300 includes a processor 302 and a memory 304 that is coupled to (and in communication with) the processor 302. The processor 302 may include, without limitation, one or more processing units (e.g., in a multi-core configuration, etc.), including a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuit or processor capable of the functions described herein. The above listing is example only, and thus is not intended to limit in any way the definition and / or meaning of processor.

[0063] The memory 304, as described herein, is one or more devices that enable information, such as executable instructions and / or other data, to be stored and retrieved. The memory 304 may include one or more computer-readable storage media, such as, without limitation, dynamic random access memory (DRAM), static random access memory (SRAM),Attorney Docket No. 5089-000200-WO-POA read only memory (ROM), erasable programmable read only memory (EPROM), solid state devices, flash drives, CD-ROMs, thumb drives, tapes, hard disks, and / or any other type of volatile or nonvolatile physical or tangible computer-readable media. The memory 304 may be configured to store, without limitation, phenotypic data, genotypic data (e.g., SNPs, etc.), predicted values, architectures / models, and / or other types of data (and / or data structures) suitable for use as described herein, etc. In various embodiments, computer-executable instructions may be stored in the memory 304 for execution by the processor 302 to cause the processor 302 to perform one or more of the functions described herein, such that the memory 304 is a physical, tangible, and non-transitory computer-readable storage media. It should be appreciated that the memory 304 may include a variety of different memories, each implemented in one or more of the functions or processes described herein.

[0064] Furthermore, in various embodiments, computer-executable instructions may be stored in the memory 304 for execution by the processor 302 to cause the processor 302 to perform one or more of the functions described herein (e.g., one or more of the operations of method 300, etc.), such that the memory 304 is a physical, tangible, and non-transitory computer readable storage media. Such instructions often improve the efficiencies and / or performance of the processor 302 that is performing one or more of the various operations herein (e.g., the performance of the computing device 300, etc.), whereby in connection with such performance the computing device 300 may be transformed into a special purpose computing device. It should be appreciated that the memory 304 may include a variety of different memories, each implemented in one or more of the functions or processes described herein.

[0065] In the example embodiment, the computing device 300 also includes a presentation unit 306 that is coupled to (and is in communication with) the processor 302. The presentation unit 306 outputs, or presents, to a user of the computing device 300 (e.g., a breeder, etc.) by, for example, displaying and / or otherwise outputting information such as, but not limited to, predicted phenotypic data, etc. It should be further appreciated that, in some embodiments, the presentation unit 306 may comprise a display device such that various interfaces (e.g., applications (network-based or otherwise), etc.) may be displayed at computing device 300, and in particular at the display device, to display such information and data, etc. And in some examples, the computing device 300 may cause the interfaces to be displayed at a display device of another computing device, including, for example, a server hosting a website having multipleAttorney Docket No. 5089-000200-WO-POA webpages, or interacting with a web application employed at the other computing device, etc. Presentation unit 306 may include, without limitation, a liquid crystal display (LCD), a light- emitting diode (LED) display, an organic LED (OLED) display, an “electronic ink” display, combinations thereof, etc. In some embodiments, presentation unit 306 may include multiple units.

[0066] The computing device 300 further includes an input device 308 that receives input from the user. The input device 308 is coupled to (and is in communication with) the processor 302 and may include, for example, a keyboard, a pointing device, a mouse, a touch sensitive panel (e.g., a touch pad or a touch screen, etc.), another computing device, and / or an audio input device. Further, in some example embodiments, a touch screen, such as that included in a tablet or similar device, may perform as both presentation unit 306 and input device 308. In at least one example embodiment, the presentation unit 306 and input device 308 may be omitted.

[0067] In addition, the illustrated computing device 300 includes a network interface 310 coupled to (and in communication with) the processor 302 (and, in some embodiments, to the memory 304 as well). The network interface 310 may include, without limitation, a wired network adapter, a wireless network adapter, a telecommunications adapter, or other device capable of communicating to one or more different networks. In at least one embodiment, the network interface 310 is employed to receive inputs to the computing device 300. For example, the network interface 310 may be coupled to (and in communication with) in-field data collection devices, in order to collect data for use as described herein. In some example embodiments, the computing device 300 may include the processor 302 and one or more network interfaces incorporated into or with the processor 302.

[0068] FIG. 4 illustrates an example method 400 of determining phenotypic data based on genetic data. The example method 400 is described herein in connection with the system 100, and may be implemented, in whole or in part, in the platform computing device 102 of the system 100. Further, for purposes of illustration, the example method 400 is also described with reference to the computing device 300 of FIG. 3. However, it should be appreciated that the method 400, or other methods described herein, are not limited to the system 100 or the computing device 300. And, conversely, the systems, data structures, and the computing devices described herein are not limited to the example method 400.Attorney Docket No. 5089-000200-WO-POA

[0069] As the outset, it should be appreciated that the databases 104-108 include various data related one or more plants, or other type of organisms, etc.

[0070] At 402, the platform computing device 102 receives a request to predict a specific phenotypic trait for one or more genomes for a plant. The genomes may be genomes physically existing as a plant, or may include synthetic genomes yet to be realized in a physical plant. As such, the request may be related to prediction of a performance of the genome to support a decision to select, or not, a genome for planting, breeding, etc., or the request may be related to evaluating a genome in connection with plant design. The particular requests may be received from a breeder, a designer or another user associated with the breeding pipeline process 112.

[0071] The phenotypic traits may include, without limitation, yield, time to budding, flowering date, plant height, disease resistance, etc.

[0072] At 404, the platform computing device 102 accesses the relevant data from the trial database 104, the genetic information database 106, and the domain knowledge repository database 108.

[0073] The relevant data is generally defined based on the request. For example, where the request is related to corn, the relevant data is limited to data representative of corn. The data generally includes, as explained above, data representative of specific phenotypes, such as measurements of observable traits from field trials, genetic data, and domain knowledge such as the gene ontology resources, metabolic pathways, gene expression, etc.

[0074] Optionally, as indicated by the dotted line in FIG. 4, at 406, the platform computing device 102 preprocesses and integrates the accessed data. Preprocessing may include, for example, normalization, filtering, and quality control, etc., as explained above, to provide (e.g., transform, etc.) the data into a suitable form for further analysis as explained herein.

[0075] At 408, the platform computing device 102 compiles connection rules for the BINN architecture 110, based on the domain knowledge representative of the intermediate data.

[0076] It should be appreciated that the platform computing device 102, or associated users, may rely on content of the intermediate knowledge initially to define the configuration of the BINN architecture 110. For example, where only a single set of intermediate data is selected, and the intermediate data indicates only one intermediate layer (e.g., gene expression as in FIG.Attorney Docket No. 5089-000200-WO-POA 2A, which indicates one intermediate layer, as compared to metabolic pathways as in FIG. 2B where interdependencies suggest multiple staggered layers; etc.).

[0077] It should be appreciated that the particular architecture selected is based on the characteristics of the intermediate data. That is, for example, the architecture or model configuration is selected according to, without limiting, a number of available omics layers, the natural hierarchy and potential inter-dependencies. For instance, when a single omics intermediate data type dominates predictive power, a single-layer model configuration is more appropriate. When omics measurements form a clear sequence (e.g., genotype->expression- >metabolites->phenotype, etc.), a stacked-layer BINN is appropriate; by contrast, when a single omics modality is available but with rich intra-layer dependencies in the intermediate data, it may be preferred to allow pathway subnetworks to exchange information via sparse cross- connections creating a staggered-layer architecture for the BINN. As such, the specific architecture may be selected, employed, for various reason related to the intermediate data, and also the input data and phenotypic data to be output, etc.

[0078] Next, the platform computing device 102, alone or in combination with user input, leverages the domain knowledge, such as, for example, literature, experimentations, etc., to compile connection rules for the nodes of the BINN architecture 110. Compiling the connection rules is accomplished in multiple iterations, which may be in series or in parallel. For example, for gene expression, the platform computing device 102, alone or in combination with user input, compiles connection rules from ones of the genes in FIG. 2A to ones of the gene expressions. As shown, for example, it is determined that three specific nucleotide base locations impact gene expression e4, whereby the connection rules define a connection from each of the nucleotide base locations to the gene expression node for e4. At the same time, or prior or after, the platform computing device 102, alone or in combination with user input, compiles connection rules from the gene express intermediate layer to the phenotypic trait, in a similar manner. The domain knowledge is thus leveraged to compile connection rules, which are then used to build the specific biology-informed architecture.

[0079] It should be appreciated that the platform computing device 102, alone or in combination with user input, may continue in the above manner to compile connection rules between the input layer, each intermediate layer, and each output layer.Attorney Docket No. 5089-000200-WO-POA

[0080] Thereafter, the platform computing device 102 builds, at 410, the BINN architecture 110 based on the compiled connection rules. The connection rules are implemented in the neural network, for example, whereby the connection rules embed the domain knowledge upon which the rules are based on the neural network.

[0081] At 412, the platform computing device 102 trains the BINN architecture 110 based on a training dataset from the phenotypic database 104 and genetic information database 106, and then validates, at 414, the BINN architecture 110 based on the reserved validation dataset (from the training dataset). The platform computing device 102 leverages one or more loss functions to train the BINN architecture 110. One example loss function is the mean squared error (MSE), to measure the average squared difference between predicted values and actual values. The MSE may be used in combination with one or more weighting functions in other examples.

[0082] Subsequently, the platform computing device 102 compares the performance of the BINN architecture 110 to one or more thresholds, and when the performance satisfies the one or more thresholds, the platform computing device 102 stores the trained BINN architecture 110 for use in predicting the phenotypic trait. The threshold may be any suitable threshold where the training has converged, whereby the improvement of the BINN architecture performance is not improved over, for example, 1%, through a fixed number of additional datapoints from the training dataset.

[0083] Next, as shown in FIG. 4, the platform computing device 102 predicts, at 416, using the trained BINN architecture 110, the phenotypic traits of one or more genomes of plants. That is, the genetic data for the genomes is input to the trained BINN architecture 110, and the expression of phenotypic trait(s) is output from the trained BINN architecture 110. The platform computing device 102 then selects, at 418, ones of the genomes based on the predicted phenotypic trait(s) (from 416) (e.g., selects a top 10%, top 15%, etc.), and directs, at 420, the selected ones of the genomes to the breeding pipeline process 112, as explained above.

[0084] FIG. 5 illustrates the relative performance (at 500) of the BINN architecture herein and multiple other models, including a penalized linear regression (RR) 502 and a three- layer fully-connected network (FCN) 504. For the BINN variants, one was trained with standard MSE (BINN-MSE) 506 and one with custom biologically-informed soft-constraint loss (as explained above, BINN soft) 508. In connection therewith, FIG. 5 illustrates test‐set MSEsAttorney Docket No. 5089-000200-WO-POA across nine training‐set sizes logarithmically spaced from 500 to 20,000 samples. Both BINN variants (506, 508) outperform the RR 502 at every sample size, and once the training set exceeds roughly lines, the errors converge to the level of the unconstrained FCN 504. Notably, the BINN soft 508 achieves the lowest MSE in the low‐data regime, demonstrating that incorporating a Pearson correlation term, calculated for only a small percentage of the lines, provides smoother, more informative gradient signals. The results indicate that BINN improves robustness during training, yielding superior accuracy when phenotypic data are scarce and intermediate traits are not known or incomplete.

[0085] That said, the BINN architectures described herein learn with less input data (i.e., sparse datasets), and outperforms in high-dimensionality large noise scenarios (as compared to linear regression and fully connected neural networks (FCNs), as illustrated in FIG. 5). It should also be appreciated, from a technical perspective, that BINN architectures are computationally cheaper than FCNs.

[0086] Further, the platform computing device 102, consistent with the above, is configured to use the BINN with standard MSE. For validation, the platform computing device 102 is configured to compare the BINN architecture to a traditional ridge regression G2P model and a ridge regression E2P model. E2P models perform best as expressions correlate well with these two phenotypes. It should be appreciated that it is not practical to design genotypes with expression data, as control lies with the genotypes. As such, only conventional models with G as the only input are useful for validating the BINN architecture, whereby the validation is based on G2P and G2E2P models (where E2P serves as a loose upper bound for the G2E2P model).

[0087] In connection therewith, FIGS. 6A-6H illustrate operation of multiple different models, including a genotype-to-expression-to-phenotype (G2E2P) BINN model, using a standard GBLUP genomic‐prediction model as the G2P baseline. In connection therewith, E2P (Elastic Net on expression), G2P (GBLUP on genotype), and G2E2P (BINN with expression‐ informed sparsity) were assessed, against the standard GBLUP, in a sparse‐data regime via five random splits, each using 20% of the lines for training / validation and the remaining 80% for testing. The figures illustrate Spearman correlation on the test‐set across four flowering‐time phenotypes (Anthesis.sp.NE in FIGS. 6A and 6B, Anthesis.sp.MI in FIGS. 6C and 6D, Silking.sp.NE in FIGS. 6E and 6F, and Silking.sp.MI in FIGS. 6G and 6H), for multiple differentAttorney Docket No. 5089-000200-WO-POA genetic populations (including an aggregate of all populations (identified as “all” in the figures), a stiff stalk (SS) heterotic population, a non-stiff stalk (NSS) heterotic population, an iodent (IDT) heterotic population, a popcorn population, a sweet corn population, a tropical population, and an others population (e.g., a mix of uncategorized lines, etc.)). As shown, the E2P models exceed the performance of the baseline G2P in all test cases. Using a paired t-test approach, it was determined that BINN outperforms GBLUP with statistical significance at a rate of about 72%. Moreover, as shown in the comparisons, the BINN provided even better performance (e.g., as compared to the GBLUP, etc.) within each of the particular genetic populations. For instance, IDT has improvements in Anthesis and Silking in both Nebraska (NE) and Michigan (MI) environments ranging from 40.6% to 49.8% (rather than the 2-3.9% for all). This therefore illustrates the improved performance of the BINN architecture relative to other models in various genetic contexts (and, more particularly, the relative improvement within the different genetic contexts). EMBODIMENTS

[0088] Additional example embodiments of the present disclosure are set forth below (without limitation).

[0089] Example embodiment 1. A method of making a plant genome having a desired nucleic acid sequence, comprising: identifying genomic segments of a plant that include at least one desired nucleic acid sequence, using a trained neural network, which is biology- informed through first pre-training connections between a genetic data layer and at least one intermediate layer and second pre-training connection between the at least one intermediate layer and a phenotypic layer; and recombining the at least one at least one desired nucleic acid sequence with a second nucleic acid sequence to form a target plant genome.

[0090] Example embodiment 2. The method according to example embodiment 1, wherein the at least one intermediate layer comprises more than one node.

[0091] Example embodiment 3. The method according to example embodiment 2 or example embodiment 3, further comprising conducting a sensitivity analysis of the more than one node.Attorney Docket No. 5089-000200-WO-POA

[0092] Example embodiment 4. The method according to example embodiment 3, wherein a highly sensitive node is identified and the at least one desired nucleic acid sequence is associated with the node.

[0093] Example embodiment 5. The method according to any one of example embodiments 1-4, wherein the phenotypic layer is indicative of at least one measurable observation of the plant.

[0094] Example embodiment 6. The method according to example embodiment 5, wherein the phenotypic layer is representative of: metabolite data, agronomic traits, expression data, proteomic data, and / or combinations thereof.

[0095] Example embodiment 7. The method according to any one of example embodiments 1-6, wherein recombining is selected from gene editing, transformation, crossing, and combinations thereof.

[0096] Example embodiment 8. The method according to any one of example embodiments 1-7, further comprising: building, by the platform computing device, the architecture of the neural network consistent with the biology-informed connection rules, whereby the neural network is embedded with first connections between an input genetic layer and the at least one intermediate layer and second connections between the at least one intermediate layer and an output phenotypic layer; and training, by the platform computing device, the neural network with the embedded first and second connections, based on at least a portion of the phenotypic data and the genetic data.

[0097] Example embodiment 9. The method according to any one of example embodiments 1-8, wherein the identifying genomic segments comprises analyzing single nucleotide polymorphisms, methylation patterns, marker library results, nucleic acid sequences and combinations thereof.

[0098] Example embodiment 10. The method according to any one of example embodiments 1-9, further comprising: accessing, by a computing device, data comprising genomic information associated with the plant; and causing, by the computing device, an algorithm to prescribe a recombining strategy.

[0099] Example embodiment 11. The method according to example embodiment 10, wherein the algorithm additionally accesses the breeding history of the at least one plant to prescribe a recombining strategy.Attorney Docket No. 5089-000200-WO-POA

[0100] Example embodiment 12. The method according to any one of example embodiments 1-11, wherein the recombining of the at least one desired nucleic acid sequence with a second nucleic acid sequence occurs in a plant that comprises a genome that is distinct from the genetic information used to identify the desired nucleic acid sequence.

[0101] Example embodiment 13. The method according to any one of example embodiments 1-12, wherein the at least one desired nucleic acid sequence comprises a nucleic acid sequence that encodes one or more proteins that alter the phenotype .

[0102] Example embodiment 14. The method according to any one of example embodiments 1-13, wherein the at least one desired nucleic acid sequence comprises a nucleic acid sequence that controls gene expression.

[0103] Example embodiment 15. The method according to any one of example embodiments 1-14, wherein at least one desired nucleic acid sequence comprises a nucleic acid sequences that contributes to a desired phenotype selected from yield, plant height, maturity, planting density, fiber quality, oil content, flavor profiles, nutrient content, protein content, oil profile, nitrogen fixation, silking, ear height, disease resistance, climate tolerance, abiotic stress, photosynthetic efficiency, transformability, haploid formation, haploid doubling, pollen stability, chromosome recombination rate, apomixis / parthenogenesis, insect resistance, or combinations thereof.

[0104] Example embodiment 16. The method according to any one of example embodiments 1-15, wherein the new plant comprises a genome that includes at least two types of engineered nucleic acid sequences selected from intragenic sequences, transgenic sequences, cisgenic sequences and combinations thereof.

[0105] Example embodiment 17. The method according to any one of example embodiments 1-16, wherein recombining comprises at least two generations of crosses.

[0106] Example embodiment 18. The method according to any one of example embodiments 1-17, wherein the new plant is selected from corn, soybean, cotton, wheat, rice, canola, oilseed rape, sugar beet, sorghum, millet, alfalfa, vegetable crops, forest trees, and fruit crops.

[0107] Example embodiment 19. The method according to any one of example embodiments 7-18, wherein gene editing comprises contacting a Cas related protein with a gRNA.Attorney Docket No. 5089-000200-WO-POA

[0108] Example embodiment 20. The method according to any one of example embodiments 1-19, wherein a single recombination is made.

[0109] Example embodiment 21. A computer-implemented method for use in genomic prediction specific to plants, the method comprising: accessing, by a platform computing device, data specific to a plant, the data including genetic information for the plant, phenotypic data for the plant, and domain knowledge data representative of a causal link between the genetic data and the phenotypic data; for a specific phenotypic trait of the plant, based on the domain knowledge data, compiling biology-informed connection rules for at least one intermediate layer of an architecture of a neural network, whereby the biology-informed connection rules impose the causal link into the at least one intermediate layer; building, by the platform computing device, the architecture of the neural network consistent with the biology- informed connection rules, wherein the neural network includes first connections between an input genetic layer and the at least one intermediate layer and second connections between the at least one intermediate layer and an output phenotypic layer; training, by the platform computing device, the neural network with the first and second connections, based on at least a portion of the phenotypic data and the genetic data; predicting, by the platform computing device, using the trained neural network, the phenotypic trait of a genome of the plant; and based on the predicted phenotypic trait of the genome satisfying a threshold, constructing and growing, in a growing space, the plant including the genome.

[0110] Example embodiment 22. The computer-implemented method of example embodiment 21, wherein the architecture includes a single intermediate layer; and wherein the architecture includes multiple sub-networks, each at a node of the single intermediate layer.

[0111] Example embodiment 23. The computer-implemented method of example embodiment 21 or example embodiment 22, wherein the architecture includes a stagger architecture, in which the at least one intermediate layer includes multiple intermediate layers; and wherein the first connections include at least one connection between the first input layer and more than one of the multiple intermediate layers.

[0112] Example embodiment 24. The computer-implemented method of any one of example embodiments 21-23, wherein the domain knowledge data includes gene expression data or gene ontology data.Attorney Docket No. 5089-000200-WO-POA

[0113] Example embodiment 25. The computer-implemented method of any one or example embodiments 21-24, wherein compiling biology-informed connection rules for at least one intermediate layer includes: compiling first biology-informed connection rules defining the first connections between the input genetic layer and the at least one intermediate layer; and compiling second biology-informed connection rules defining the second connections between the at least one intermediate player and the output phenotypic layer.

[0114] Example embodiment 26. The computer-implemented method of any one or example embodiments 21-25, wherein training the built architecture of the neural network is further based on at least one loss function.

[0115] Example embodiment 27. The computer-implemented method of example embodiment 26, wherein the at least one loss function includes a means squared error (MSE) loss function.

[0116] Example embodiment 28. The computer-implemented method of example embodiment 27, wherein the at least one loss function further includes a correlation constraint on the domain knowledge data.

[0117] In view of the above, the systems and methods herein provide a BINN architecture that enables efficient data handling and integration, supporting advanced biological analyses and predictive modeling that are robust to data limitations and biological complexity.

[0118] With that said, it should be appreciated that the functions described herein, in some embodiments, may be described in computer executable instructions stored on a computer readable media, and executable by one or more processors. The computer readable media is a non-transitory computer readable media. By way of example, and not limitation, such computer readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Combinations of the above should also be included within the scope of computer-readable media.

[0119] It should also be appreciated that one or more aspects of the present disclosure transform a general-purpose computing device into a special-purpose computing device when configured to perform the functions, methods, and / or processes described herein.Attorney Docket No. 5089-000200-WO-POA

[0120] As will be appreciated based on the foregoing specification, the above- described embodiments of the disclosure may be implemented using computer programming or engineering techniques, including computer software, firmware, hardware or any combination or subset thereof, wherein the technical effect may be achieved by performing at least one of the following operations: (a) identifying genomic segments of a plant that include at least one desired nucleic acid sequence, using on a trained neural network, which is biology-informed through first pretraining connections between a genetic data layer and at least one intermediate layer and second pretraining connection between the at least one intermediate layer and a phenotypic layer; (b) recombining the at least one at least one desired nucleic acid sequence with a second nucleic acid sequence to form a target plant genome; (c) accessing data specific to a plant, the data including genetic information for the plant, phenotypic data for the plant, and domain knowledge indicative of a causal link between the genetic data and the phenotypic data; (d) for a specific phenotypic trait of the plant, based on the domain knowledge, compiling biology-informed connection rules for at least one intermediate player of an architecture of a neural network; (e) building the architecture of the neural network consistent with the biology- informed connection rules, whereby the neural network is embedded with first connections between an input genetic layer and the at least one intermediate layer and second connections between the at least one intermediate layer and an output phenotypic layer; (f) training the neural network with the embedded first and second connections, based on at least a portion of the phenotypic data and the genetic data; (g) predicting using the trained neural network, the phenotypic trait of a genome of the plant; and / or (h) based on the predicted phenotypic trait of the genome satisfying a threshold, constructing and growing, in a growing space, the plant including the genome.

[0121] Examples and embodiments are provided so that this disclosure will be thorough, and will fully convey the scope to those who are skilled in the art. Numerous specific details are set forth such as examples of specific components, devices, and methods, to provide a thorough understanding of embodiments of the present disclosure. It will be apparent to those skilled in the art that specific details need not be employed, that example embodiments may be embodied in many different forms and that neither should be construed to limit the scope of the disclosure. In some example embodiments, well-known processes, well-known device structures, and well-known technologies are not described in detail. In addition, advantages andAttorney Docket No. 5089-000200-WO-POA improvements that may be achieved with one or more example embodiments disclosed herein may provide all or none of the above mentioned advantages and improvements and still fall within the scope of the present disclosure.

[0122] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms "a,” “an,” and “the” may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms “comprises,” “comprising,” “including,” and “having,” are inclusive and therefore specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order discussed or illustrated, unless specifically identified as an order of performance. It is also to be understood that additional or alternative steps may be employed.

[0123] When a feature is referred to as being “on,” “engaged to,” “connected to,” “coupled to,” “associated with,” “in communication with,” or “included with” another element or layer, it may be directly on, engaged, connected or coupled to, or associated or in communication or included with the other feature, or intervening features may be present. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0124] Although the terms first, second, third, etc. may be used herein to describe various features, these features should not be limited by these terms. These terms may be only used to distinguish one feature from another. Terms such as “first,” “second,” and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first feature discussed herein could be termed a second feature without departing from the teachings of the example embodiments.

[0125] The foregoing description of the embodiments has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure. Individual elements or features of a particular embodiment are generally not limited to that particular embodiment, but, where applicable, are interchangeable and can be used in a selected embodiment, even if not specifically shown or described. The same may also be varied in manyAttorney Docket No. 5089-000200-WO-POA ways. Such variations are not to be regarded as a departure from the disclosure, and all such modifications are intended to be included within the scope of the disclosure.

Claims

Attorney Docket No. 5089-000200-WO-POA CLAIMS What is claimed is:

1. A method of making a plant genome having a desired nucleic acid sequence, comprising: identifying genomic segments of a plant that include at least one desired nucleic acid sequence, using a trained neural network, which is biology-informed through first pre-training connections between a genetic data layer and at least one intermediate layer and second pre- training connection between the at least one intermediate layer and a phenotypic layer; and recombining the at least one at least one desired nucleic acid sequence with a second nucleic acid sequence to form a target plant genome.

2. The method according to claim 1, wherein the at least one intermediate layer comprises more than one node.

3. The method according to claim 2, further comprising conducting a sensitivity analysis of the more than one node.

4. The method according to claim 3, wherein a highly sensitive node is identified and the at least one desired nucleic acid sequence is associated with the node.

5. The method according to claim 1, wherein the phenotypic layer is indicative of at least one measurable observation of the plant.

6. The method according to claim 5, wherein the phenotypic layer is representative of: metabolite data, agronomic traits, expression data, proteomic data, and / or combinations thereof.

7. The method according to claim 1, wherein recombining is selected from gene editing, transformation, crossing, and combinations thereof.Attorney Docket No. 5089-000200-WO-POA 8. The method according to claim 1, further comprising: building, by the platform computing device, the architecture of the neural network consistent with the biology-informed connection rules, whereby the neural network is embedded with first connections between an input genetic layer and the at least one intermediate layer and second connections between the at least one intermediate layer and an output phenotypic layer; and training, by the platform computing device, the neural network with the embedded first and second connections, based on at least a portion of the phenotypic data and the genetic data.

9. The method according to claim 1, wherein the identifying genomic segments comprises analyzing single nucleotide polymorphisms, methylation patterns, marker library results, nucleic acid sequences and combinations thereof.

10. The method according to claim 1, further comprising: accessing, by a computing device, data comprising genomic information associated with the plant; and causing, by the computing device, an algorithm to prescribe a recombining strategy.

11. The method according to claim 10, wherein the algorithm additionally accesses the breeding history of the at least one plant to prescribe a recombining strategy.

12. The method according to claim 1, wherein the recombining of the at least one desired nucleic acid sequence with a second nucleic acid sequence occurs in a plant that comprises a genome that is distinct from the genetic information used to identify the desired nucleic acid sequence.

13. The method according to claim 1, wherein the at least one desired nucleic acid sequence comprises a nucleic acid sequence that encodes one or more proteins that alter the phenotype .Attorney Docket No. 5089-000200-WO-POA 14. The method according to claim 1, wherein the at least one desired nucleic acid sequence comprises a nucleic acid sequence that controls gene expression.

15. The method according to claim 1, wherein at least one desired nucleic acid sequence comprises a nucleic acid sequences that contributes to a desired phenotype selected from yield, plant height, maturity, planting density, fiber quality, oil content, flavor profiles, nutrient content, protein content, oil profile, nitrogen fixation, silking, ear height, disease resistance, climate tolerance, abiotic stress, photosynthetic efficiency, transformability, haploid formation, haploid doubling, pollen stability, chromosome recombination rate, apomixis / parthenogenesis, insect resistance, or combinations thereof.

16. The method according to claim 1, wherein the new plant comprises a genome that includes at least two types of engineered nucleic acid sequences selected from intragenic sequences, transgenic sequences, cisgenic sequences and combinations thereof.

17. The method according to claim 1, wherein recombining comprises at least two generations of crosses.

18. The method according to claim 1, wherein the new plant is selected from corn, soybean, cotton, wheat, rice, canola, oilseed rape, sugar beet, sorghum, millet, alfalfa, vegetable crops, forest trees, and fruit crops.

19. The method according to claim 7, wherein gene editing comprises contacting a Cas related protein with a gRNA.

20. The method according to claim 1, wherein a single recombination is made.Attorney Docket No. 5089-000200-WO-POA 21. A computer-implemented method for use in genomic prediction specific to plants, the method comprising: accessing, by a platform computing device, data specific to a plant, the data including genetic information for the plant, phenotypic data for the plant, and domain knowledge data representative of a causal link between the genetic data and the phenotypic data; for a specific phenotypic trait of the plant, based on the domain knowledge data, compiling biology-informed connection rules for at least one intermediate layer of an architecture of a neural network, whereby the biology-informed connection rules impose the causal link into the at least one intermediate layer; building, by the platform computing device, the architecture of the neural network consistent with the biology-informed connection rules, wherein the neural network includes first connections between an input genetic layer and the at least one intermediate layer and second connections between the at least one intermediate layer and an output phenotypic layer; training, by the platform computing device, the neural network with the first and second connections, based on at least a portion of the phenotypic data and the genetic data; predicting, by the platform computing device, using the trained neural network, the phenotypic trait of a genome of the plant; and based on the predicted phenotypic trait of the genome satisfying a threshold, constructing and growing, in a growing space, the plant including the genome.

22. The computer-implemented method of claim 21, wherein the architecture includes a single intermediate layer; and wherein the architecture includes multiple sub-networks, each at a node of the single intermediate layer.

23. The computer-implemented method of claim 21, wherein the architecture includes a stagger architecture, in which the at least one intermediate layer includes multiple intermediate layers; and wherein the first connections include at least one connection between the first input layer and more than one of the multiple intermediate layers.Attorney Docket No. 5089-000200-WO-POA 24. The computer-implemented method of claim 21, wherein the domain knowledge data includes gene expression data or gene ontology data.

25. The computer-implemented method of claim 21, wherein compiling biology- informed connection rules for at least one intermediate layer includes: compiling first biology-informed connection rules defining the first connections between the input genetic layer and the at least one intermediate layer; and compiling second biology-informed connection rules defining the second connections between the at least one intermediate player and the output phenotypic layer.

26. The computer-implemented method of claim 21, wherein training the built architecture of the neural network is further based on at least one loss function.

27. The computer-implemented method of claim 26, wherein the at least one loss function includes a means squared error (MSE) loss function.

28. The computer-implemented method of claim 27, wherein the at least one loss function further includes a correlation constraint on the domain knowledge data.

Citation Information

Patent Citations

  • Plant breeding method

    US20050144664A1

  • Machine learning driven gene discovery and gene editing in plants

    US20220301658A1

  • Reduction of computation complexity of neural network sensitivity analysis

    US9483727B2

  • Processes for modulating plant gene expression

    WO2023227632A1