Means and method for identifying genetic alterations contributing to a complex trait
A GNN-based method with multi-omics embeddings identifies and optimizes genome edits for complex plant traits, enhancing breeding efficiency and trait prediction in plants.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VLAAMS INTERUNIVERSITAIR INST VOOR BIOTECHNOLOGIE VZW
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods struggle to identify specific combinations of gene edits contributing to complex traits in plants, such as yield and stress tolerance, due to the complex interactions between multiple genes and environmental factors, making it difficult to predict and enhance these traits effectively.
A computer-implemented method using a Graph Neural Network (GNN) with multi-omics embeddings to analyze genotyping and phenotyping data, identifying precise combinations of genome edits that enhance complex traits by predicting their contributions to plant features, and guiding iterative breeding strategies.
Enhances the prediction of complex traits by accurately identifying contributing genome edits, optimizing breeding processes, and dynamically refining predictions through continuous feedback loops, leading to improved plant varieties with desired characteristics.
Smart Images

Figure IMGF000025_0001 
Figure 00000024_0000 
Figure 00000025_0000
Abstract
Description
[0001] HiNel / AI-traitselect / 868
[0002] MEANS AND METHOD FOR IDENTIFYING GENETIC ALTERATIONS CONTRIBUTING TO A COMPLEX TRAIT
[0003] Field of the invention
[0004] The present methods and systems generally relate to the agronomical field and relate to subfields of computational biology and bioinformatics. More, specifically the invention provides a computer-implemented algorithm which can predict the genetic alterations which are associated with a complex trait in plants.
[0005] Introduction to the invention
[0006] Many traits that are important for fitness and agricultural value of plants are complex quantitative traits. These traits are called complex because they are controlled by many genes and by environmental factors. Although the genetics of quantitative traits has been studied for over 100 years, very few of the polymorphisms that cause variation in these traits were known until recently. Gene editing is the preferred tool for the creation of novel alleles for plant genetic research and breeding. Subsequent phenotypic analysis is used to associate genes to (molecular) functions and traits such as nutritional quality, yield, and resistance to (a)biotic stress. Because drought, heat, salt stress and yield are likely to become more important due to climate change, plant responses to these abiotic stresses are intensely studied to improve agricultural production and contribute to the sustainability of the global food system. Many of these plant traits are controlled by a complex interaction between many different genes and require the identification of the right combinations of alleles to have pronounced and desired effects. Having the right combinations of alleles is also valuable in overcoming genetic redundancy, which is common in plants. Hence, the ability to stack multiple gene edits in the same plant is crucial for the engineering of complex traits. Multiplexing, i.e. the simultaneous targeting of many genes, is one of the major advantages of the CRISPR / Cas9 gene editing system. The Cas9 protein can be delivered separately or can be combined with a single T-DNA containing multiple guide RNAs (gRNAs), only differing in the spacer sequence and targeting a variety of genes, including multiple members of gene families. The result is often a population of plants with multiple gene edits and it is difficult to define the specific contribution of two or more genes to a complex genetic trait. The present invention satisfies this need and provides a computer-implemented method which can identify specific gene edited combinations which are causal for a specificHiNel / AI-traitselect / 868
[0007] complex trait of interest. Such specific combinatorial edits can then be conveniently introduced into different relevant germplasm of the same crop or in germplasm of other crops for evaluating the complex trait in greenhouse conditions and subsequently in the field.
[0008] Summary of the invention
[0009] The present invention provides a computer-implemented method for enhancing complex traits in plants, comprising:
[0010] • Providing one or more populations of plants or seeds edited using multiplex CRISPR technology targeting a predefined set of genetic elements associated with one or more desired complex traits;
[0011] • Performing genotyping and phenotyping to identify specific gene edits in each plant or seed and to collect plant features associated with the complex traits of interest, using automated methods, but not limited to them;
[0012] • Utilizing an Al model, which may include but is not limited to a Graph Neural Network (GNN), using multi-omics embeddings to estimate the contributions of individual gene edits to the complex traits and to identify precise combinations of gene edits that maximize the trait's beneficial outcomes based on extracted plant features.
[0013] In yet another aspect the invention provides a computer-implemented method for identifying at least two genome edits contributing to a selected complex trait in plants comprising:
[0014] i. selecting genetic elements predicted to be associated with a selected complex trait by applying an artificial intelligence (Al) model, trained on biological data from plants, plant cells, plant organs or plant seeds, wherein genetic elements in said Al model are associated with features of complex traits, and wherein said Al model comprises embeddings,
[0015] ii. genotyping a population of genome edited plants, plant cells, plant organs or plant seeds wherein the genome edits are present in the genetic elements selected in step i) and phenotyping each of said genotyped plants, plant cells, plant organs or plant seeds, wherein phenotyping comprises measuring said selected complex trait, and
[0016] iii. calculating the contribution of genome edits in the genotyped and phenotyped population of plants, plant cells, plant organs or plant seeds by applying the Al modelHiNel / AI-traitselect / 868
[0017] from step i) to the selected complex trait in plants, plant cells, plant organs or plant seeds, and
[0018] iv. identifying combinations of at least two genome edits that contribute to a selected complex trait in plants.
[0019] In a particular aspect the Al model guides the crossing of gene-edited plant populations to generate new combinations and also to iteratively narrow the list of genetic targets.
[0020] In another aspect wherein the populations are dynamic, utilizing the Al model to guide crosses wherein the endonuclease (e.g. Cas endonuclease) is still active due to the presence of gRNAs, enabling the generation of new populations of edited plants within a reduced and optimized target gene space and selecting plants or seeds that display genotypes significantly enhancing the desired complex trait as determined by refined predictions of the Al model.
[0021] In a specific aspect the Al methods utilize embeddings based on protein sequence, gene expression, functional annotations, epigenetic data and other data such as gene interaction networks, and evolutionary conservation, combined with plant features extracted through automated phenotyping or other data collection methods, to enhance the predictive analysis of gene-trait interactions.
[0022] In yet another aspect the Al model's predictions are iteratively refined by re-training on data from successive breeding cycles, including updates from phenotyping and genotyping data, and guiding future crosses based on these predictions.
[0023] In yet another aspect the genotyping and phenotyping processes are integrated into a continuous feedback loop with the Al model, enabling dynamic updates to the breeding strategy based on data from both plant features and genetic edits.
[0024] In yet another aspect dynamic plant populations are maintained, allowing the Al model to guide crosses to generate successive populations of optimized plants.
[0025] Scheme of the computer-implemented invention Genotyping and phenotyping a population of genome edited plants, plant cells, plant organs or plant seeds (Step 1 and Step 2) wherein said genome edits are in a predefined set of genetic elements and wherein said phenotyping comprises measuring the complex trait in each of the genotyped plants, plant cells, plant organs or plant seeds, applying an artificial intelligence (Al) model (Step 3), trained onHiNel / AI-traitselect / 868
[0026] biological data related to genome edits associated with features of a complex trait, wherein said Al model comprises embeddings for the data analysis in gene-level embeddings (depicted by circles in the figure), plant feature-level embeddings (depicted by squares in the figure) and complex trait-level embeddings (depicted by stars in the figure), and calculating the contribution of single or multiple genome edits to a complex trait in plants and, identifying combinations of at least two genome edits that contribute to a complex trait in plants (Step 5). The retrained artificial intelligence (Al) model (see Detailed View of Step 3) is employed to predict specific genetic elements that are expected to influence certain plant features and traits. These predictions may then guide the generation of novel plant populations comprising genomic edits introduced at the predicted genetic elements (Step 4).
[0027] Figure 2: Training results on 25 different random subsets of data. The models with embeddings performed generally better with a mean Pearson correlation of 0.62 vs 0.59 for models without embeddings on independent test sets (this difference is significant according to a Wilcoxon signed rank test). Training data contained 3050 plants, the validation set 542 plants and the test set 545 plants.
[0028] Figure 3: UMAP plot of the different plants in the dataset based on their genotypes. The blue dots are the training and validation dataset, while the orange dots represent the independent test dataset. The figure shows that the test dataset provides a good coverage of the dataset. Training data contained 3050 plants, the validation set 542 plants and the test set 545 plants.
[0029] Figure 4: SHAP values showing the impact of mutations in different genes. On the y-axisall genes are shown (most genes are redacted) ordered by importance, while the x-axis shows the impact on the model predictions for PLA. Each dot represents a plant from the dataset. The color of the dot represents the (out-of-frame) mutation frequency of that gene in the specific plant (blue means that the gene is not mutated, purple represents a heterozygous mutation and red means a homozygous mutation). Positive SHAP values indicate that the mutation of that gene pushes the PLA to higher values. A red dot with a high SHAP value for a specific gene means that the gene has a highly positive effect on the PLA when homozygously mutated. The SHAP values of a specific gene can be dependent on other genes.HiNel / AI-traitselect / 868
[0030] Detailed description of the invention
[0031] The present disclosure relates generally to a machine learning engine for the identification of multiple genomic edits associated with complex traits in plants. The present disclosure also relates to a system (or apparatus) implementing the artificial intelligence (Al) platform.
[0032] Example embodiments will be described more fully hereinafter. It should be understood that such systems, computer readable media, and methods may be embodied in many different forms and should not be construed as limited to the example embodiments set forth herein. Rather, these example embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the claims to those of ordinary skill in the art.
[0033] The term "machine learning" as used herein generally refers to a type of artificial intelligence (Al) that provides computers with the ability to learn without being explicitly programmed. Machine learning is a branch of Al focusing on systems that can learn from data, identify patterns, and make decisions with minimal human intervention.
[0034] As used herein, the term "complex trait" refers to a characteristic that is influenced by multiple genes and often by environmental factors as well. Unlike simple traits, which are controlled by a single gene, complex traits involve interactions between various genetic and environmental elements, making them more challenging to study and predict. Examples of complex traits in plants include yield, drought tolerance, abiotic stress tolerance, short stature and disease resistance. These traits are important for plant breeding and agricultural practices as they can significantly impact crop productivity and sustainability.
[0035] As used herein "mutation" refers to a change in the nucleotide sequence of a genetic element for example such a nucleotide sequence change leads to another amino acid sequence of a native protein, mutations can also occur in a non-coding region such as the promoter or terminator region or can also occur in the transcription factor binding site.
[0036] Terms such as "first", "second", and "within" are used merely to distinguish one component (or part of a component or state of a component) from another. Such terms are not meant to denote a preference or a particular orientation and are not meant to limit embodiments of the disclosure. In the following detailed description of the example embodiments, numerousHiNel / AI-traitselect / 868
[0037] specific details are set forth in orderto provide a more thorough understanding of the invention. However, it will be apparent to one of ordinary skill in the art that embodiments of the disclosure may be practiced without these specific details. In other instances, well-known features have not been described in detail to avoid unnecessarily complicating the description.
[0038] A user may be any person or entity that interacts with the database, the Al platform, or both. Examplesofa user may include, but are not limited to, a principal investigator, a scientist, a postdoctoral candidate, a graduate student, or an agricultural company, for example. There can be one or multiple users.
[0039] The present invention provides in a first embodiment a computer-implemented method for identifying at least two genome edits contributing to a selected complex trait in plants comprising:
[0040] i. selecting genetic elements predicted to be associated with a specific complex trait by applying an artificial intelligence (Al) model, trained on biological data from plants, plant cells, plant organs or plant seeds, wherein genetic elements in said Al model are associated with features of complex traits, and wherein said Al model comprises embeddings,
[0041] ii. genotyping a heterogenous population of different genome edited plants, plant cells, plant organs or plant seeds wherein the genome edits are present in the genetic elements selected in step i), and phenotyping each of said genotyped plants, plant cells, plant organs or plant seeds, wherein phenotyping comprises measuring said selected complex trait, and
[0042] iii. calculating the contribution of genome edits in the genotyped and phenotyped population of plants, plant cells, plant organs or plant seeds by applying the Al model from step i), to the selected complex trait in plants, plant cells, plant organs or plant seeds, and
[0043] iv. identifying combinations of at least two genome edits that contribute to a selected complex trait in plants.
[0044] In the invention a 'selected complex trait" can be used interchangeable with a 'specific complex trait', or a 'particular complex trait" or a "desired complex trait" or a "complex trait of interest". In the invention 'a heterogeneous population of genome edited plants' is a population of different genome edited plants wherein plants with one genome edit in only one geneticHiNel / AI-traitselect / 868
[0045] element occur, plants with two genome edits in two different genetic elements occur, plants with three, four, five or more different genome edits occur in three, four, five or more different genome edits are present in one plant.
[0046] The wording "wherein genetic elements in said Al model are associated with features of complex traits" means that a multiple of genetic elements are associated with a multiple of complex traits. Thus, the multiple of complex traits can be considered as a genus of different complex traits wherein two or more genetic elements are associated with specific complex traits present in said multiple of complex traits. A genetic element can be associated with one or more complex genetic traits but one specific complex trait is associated with a specific set of two or more genetic elements.
[0047] The wording "wherein phenotyping comprises measuring said selected complex trait" means thatthe expression ofthe selected complex trait is measured (or quantified) in each ofthe plants in the population of genotyped plants.
[0048] Thus the invention starts from the identification of genetic elements which are predicted to be associated with a particular complex trait (particular is herein interchangeably used with selected or chosen). The identification of genetic elements associated with a high likelihood of being associated with a particular complex trait is done with the artificial intelligence model of the invention. A high likelihood of being associated with a complex trait can be a likelihood of at least 30%, at least 40%, at least 50%, at least 60% or higher. The Al model of the invention is built with genetic elements associated with a multiple of complex traits (as opposed to the association with a selected complex trait). In other words, the Al model incorporates a multiple of genetic elements associated with a multiple of complex traits. The Al model incorporates embeddings which leads to a better prediction of the association between particular genetic elements and a specific complex trait.
[0049] In a next step a heterogenous population of different genome edited plants, plant cells, plant organs or plant seeds which have been subjected to gene editing at a predefined set of genetic loci will be genotyped and phenotyped. Such a population of genome edited plants is typically a population of 100, 200, 300, 400, 500 or more plants. Such predefined set of genetic loci or genetic elements which is an equivalent term are user defined. Generally such single or multiple genetic elements have been described in literature, or have been identified in proprietary research, as being associated with complex genetic traits or such single genetic elements haveHiNel / AI-traitselect / 868
[0050] been identified in GWAS studies or QTL analysis but such genetic elements can also be identified with the artificial intelligence model of the present invention. Genome editing is carried out with gene editing tools such as CRISPR-CAS technology, or Zinc Finger Nucleases (ZFNs), or Base Editors which are modified versions of CRISPR that can convert one DNA base into another without making double-strand breaks, or Prime Editors which can make a wide variety of genetic changes, including insertions, deletions, and all 12 possible base-to-base conversions, without requiring double-strand breaks or by using meganucleases which enzymes recognize and cut large DNA sequences, making them useful for specific and targeted gene editing. Genome editing can be carried out in the coding sequences of genes, regulatory sequences of genes (promoter, enhancer, introns, 3' sequences of genes) as well as in other genetic elements such as circular RNAs, microRNAs or long non-coding RNAs and the like. Depending on the predicted impact of each genetic element, editing at predefined genetic loci can aim to either knock out a genetic element in a heterozygous or homozygous state or when a genetic element is a gene, it can regulate its expression by upregulating or downregulating the gene or affect protein functionality. The modulation of gene expression can be tailored to occur in a tissue-specific manner, targeting particular tissues such as roots, leaves, or stems. Additionally, the upregulation or downregulation of gene activity can be restricted to a specific tissue or time in the development of the plant, ensuring that the gene edits influence the desired traits precisely when and where they are needed.
[0051] In a next step genotyping and phenotyping is performed, which can occur simultaneously or separately. Genotyping identifies the specific edits of genetic elements generated in each plant, plant cell, plant organ or plant seed within the population, using technologies such as nextgeneration sequencing (NGS) or more typically with PCR-based methods to characterize the edited loci. Phenotyping involves measuring the plant's traits of interest, and plant features associated with a complex trait. Plant features can be determined using automated systems (such as high-throughput imaging, sensors, or drones), but the method is not limited to automation; manual measurements and other data collection techniques may also be employed. The plant features collected may include plant height, leaf area, root structure, flowering time, and other morphological, physiological, and biochemical characteristics, which represent the phenotypic expression of complex traits. Generally a complex trait comprises several plant features which are associated with this complex trait. These genotypic andHiNel / AI-traitselect / 868
[0052] phenotypic data are then analyzed together, allowing for a comprehensive assessment of how specific edits of genetic elements influence plant features and together complex traits.
[0053] The computer-implemented method of the invention identifies at least two, at least three, at least four, or more genome edits which contribute to a complex trait. A complex trait is by definition associated with two or more genome edits (or genomic alterations or genomic mutations or gene mutations) as opposed to a single Mendelian trait which is associated with one genome edit.
[0054] In another embodiment the invention provides a computer-implemented method for identifying at least two genome edits contributing to a complex trait in plants comprising:
[0055] 1. genotyping and phenotyping a population of genome edited plants, plant cells, plant organs or plant seeds wherein said genome edits are in a predefined set of genetic elements and wherein said phenotyping comprises measuring the complex trait in each of the genotyped plants, plant cells, plant organs or plant seeds,
[0056] 2. applying an artificial intelligence (Al) model, trained on biological data related to genome edits associated with features of a complex trait, wherein said Al model comprises embeddings, and
[0057] 3. calculating the contribution of single or multiple genome edits to a complex trait in plants and,
[0058] 4. identifying combinations of at least two genome edits that contribute to a complex trait in plants.
[0059] The Al model employed in the method is based on an artificial intelligence model such as for example a Graph Neural Network (GNN) but is not limited to this structure meaning that a GNN can also be combined with another Al model. The Al model is designed to predict relationships between edits of genetic elements, plant features associated with a complex trait, and complex traits, using multiple levels of embeddings to enhance the predictive power of the Al model. Thus the model comprises genetic elements, plant features, and complex traits as nodes. Genetic elements are connected to one or more plant features, which features represent measurable characteristics such as plant height, root depth, and leaf area. These features, in turn, constitute the expression of complex traits such as yield, drought tolerance, or disease resistance. Additionally, genetic elements are connected with each other through gene interaction networks, representing their interactions and regulatory relationships. ThisHiNel / AI-traitselect / 868
[0060] interconnected, hierarchical structure enables the Al model to predict how network engineering through genome editing influences specific plant features and, subsequently, the overall expression of complex traits.
[0061] A graph neural network (GNN) is a machine learning technique that is used to perform various operations on graphical data. Thus a GNN processes data that can be represented by graphs. The key design element of GNNs is the use of pairwise message passing, such that graph nodes iteratively update their representations by exchanging information with their neighbors.
[0062] A graph is the most basic and essential part of the graph neural network. A graph is a data structure that is made of two components that are vertices (or nodes) and edges. Essentially there are 3 types of graph neural networks: a recurrent GNN, a type in which the connection between the various nodes generates a cyclic allowance of output from the other nodes, a spatial GNN and a spectral GNN.
[0063] The various functions performed by the GNN are: i) node classification which is the process of training a model to predict labels of nodes, based on the features of each node and the features and connections of its neighboring nodes, ii) link prediction, which is defining the relationship between the two nodes in a graph, and this checks whether the two nodes are connected or not and iii) graph classification, which classifies the graph on the basis of nodes and connections present on the different graphs.
[0064] In the instant invention, to improve the accuracy and depth of the Al model's predictions, thus the GNN, embeddings are applied. Multi-omics embeddings represent the properties of each genetic element, such as genes, in the model. These embeddings are created from various biological data sources:
[0065] • Protein sequence and structure embeddings: Capturing the evolutionary relationships and functional roles of the encoded proteins.
[0066] • Gene expression embeddings: derived from transcriptomics data (e.g., RNA-seq), these embeddings reflect how genes are expressed across different conditions and tissues. • Functional annotations: Sourced from databases like Gene Ontology (GO) or InterPro, these embeddings describe the molecular functions, biological processes, and cellular locations of each gene.HiNel / AI-traitselect / 868
[0067] • Epigenetic and regulatory embeddings: These embeddings may include data from chromatin accessibility or transcription factor binding sites, helping to predict how gene expression is regulated.
[0068] In the graph model, plant features act as intermediate nodes between genetic elements and traits. This structure allows the Al model to first predict how edits in genetic elements impact individual plant features (such as root depth or leaf thickness), and then use the feature-level information to predict the complex trait outcomes.
[0069] The computer-implemented method of the invention is schematically depicted in Figure 1 and can be depicted with the scheme and numbers in Figure 1 as follows:
[0070] The present invention a computer-implemented method for identifying at least two genome edits contributing to a selected complex trait in plants comprising:
[0071] i. selecting genetic elements predicted to be associated with a specific complex trait by applying an artificial intelligence (Al) model (Figure 1, Step 4), trained on biological data from plants, plant cells, plant organs or plant seeds, wherein genetic elements in said Al model are associated with features of complex traits, and wherein said Al model comprises embeddings for the data analysis in gene-level embeddings, and / or plant feature-level embeddings and / or complex trait-level embeddings,
[0072] ii. genotyping a heterogenous population of genome edited plants, plant cells, plant organs or plant seeds (Figure 1, Step 1 and 2) wherein the genome edits are present in the genetic elements selected in Step 4 (Figure 1), and phenotyping each of said genotyped plants, plant cells, plant organs or plant seeds (Figure 1, Step 1 and 2), wherein phenotyping comprises measuring said selected complex trait, and
[0073] iii. calculating the contribution of genome edits in the genotyped and phenotyped population of plants, plant cells, plant organs or plant seeds by applying the Al model (Figure 1, Step 3) to the data in Step 2 (Figure 3) selected complex trait in plants, plant cells, plant organs or plant seeds, and
[0074] iv. identifying combinations of at least two genome edits that contribute to a selected complex trait in plants (Figure 1, Step 5).
[0075] The information obtained by the Al model will be invaluable to guide the breeding process. Based on the Al model's predictions, optimal combinations of edits of genetic elements are suggested to improve the desired complex traits. Such predictions done by the Al model areHiNel / AI-traitselect / 868
[0076] validated in the next breeding cycle, where new populations of edited plants in genetic elements are generated, phenotyped, and again analyzed. The Al model is then retrained with the new data, refining its predictions further.
[0077] A specific aspect of the invention is the use of dynamic Populations and Al-Guided Crosses The predictions from the Al model determines which crosses should be made to generate new populations of edited plants. By doing so, the Al model dynamically optimizes the target gene space, speeds up the breeding process with minimal plant growth space, ensuring that each new population progressively moves closer to the desired genetic and phenotypic outcomes.
[0078] Feedback Loop
[0079] In a specific embodiment the Al model continuously updates its predictions based on the latest phenotyping and genotyping data. As new data is collected, the embeddings are refined and connections can be added or removed, improving the model's ability to capture the complex relationships between edits of genetic elements and phenotypic outcomes (Figure 1, detailed view of step 3). For example:
[0080] • Embeddings may be updated with new transcriptomics or regulatory data, refining predictions or expanding to other organs.
[0081] • Connections between genetic elements are added to the graph model based on validated interactions found in breeding cycles (the creation, genotyping, phenotyping and breeding in a population of genome edited plants, plant cells, plant organs or plant seeds) or transcriptomic data.
[0082] • Genetic element-to-plant feature connections are added to the graph model based on validated effects of edits in genetic elements on plant features as found in breeding cycles.
[0083] Iterative Predictions and Validation
[0084] With each breeding cycle, the model suggests the next optimal set of edits in genetic elements to be created based on the updated embeddings and connections. The newly generated plants (with edits in these newly chosen genetic elements) are phenotyped and genotyped, and theHiNel / AI-traitselect / 868
[0085] results are used to validate the Al model's predictions. This iterative approach ensures that the model becomes more accurate over time, as it continuously learns from the feedback provided by real-world data.
[0086] Dynamic Breeding Strategy
[0087] The Al model's ability to incorporate new data on plant features and traits allows for a dynamic breeding strategy. Rather than following a static list of genetic targets, the model dynamically adjusts the gene-editing strategy as new insights emerge. The continuous refinement of the multi-level embeddings and the genetic element-plant feature-trait graph allows for more precise selection of editing targets, ultimately leading to the development of superior plant varieties with enhanced traits.
[0088] At the conclusion of each breeding cycle, the Al model, utilizing embeddings together with genetic element-to-genetic element and genetic element-to-feature connections provides a refined prediction of which plants are most likely to exhibit the desired complex traits. These embeddings and connections allow for a more informed and accurate selection process by linking the effects of edits to specific plant features and how these features collectively contribute to trait expression. This embedding-driven approach ensures that the selected plants demonstrate the highest potential for trait enhancement, resulting in superior crop varieties optimized for large-scale cultivation or further genetic improvement.
[0089] In yet another embodiment the invention provides a computer-readable storage medium which stores computer-executable instruction that, when executed by at least one processor, causes the processor to perform the herein described methods.
[0090] In yet another embodiment the invention provides an apparatus comprising control circuitry configured to perform the herein described methods.
[0091] Systems of the disclosure can include an intranet-based computer system that is capable of communicating with various software. A computer system includes any type of computing device or communication device. Examples of such a system can include, but are not limited to, super computers, a processor array, distributed parallel system, a desktop computer with LAN, WAN, Internet or intranet access, a laptop computer with LAN, WAN, Internet or intranet access,HiNel / AI-traitselect / 868
[0092] a smart phone, a server, a server farm, an android device (or equivalent), a tablet, smartphones, and a personal digital assistant (PDA). Further, as discussed above, such a system can have corresponding software (e.g., user software, sensor device software). The software of one system can be a part of, or operate separately but in conjunction with, the software of another system.
[0093] Embodiments of the disclosure include a storage repository. The storage repository can be a persistent storage device (or set of devices) that stores software and data. Examples of a storage repository can include, but are not limited to, a hard drive, flash memory, some other form of solid-state data storage, or any suitable combination thereof. The storage repository can be located on multiple physical machines, each storing all or a portion of the database, Al platform, protocols, algorithms, or other stored data according to some example embodiments. Each storage unit or device can be physically located in the same or in a different geographic location. In embodiments, the storage repository may be stored locally, or on cloud-based serveries such as Amazon Web Services.
[0094] In one or more example embodiments, the storage repository stores one or more databases, Al Platforms, protocols, algorithms, and stored data. The protocols can include any of a number of communication protocols that are used to send, receive, or send and receive data between the processor, datastore, memory and the user. A protocol can be used for wired and / or wireless communication. Examples of a protocols can include, but are not limited to, Modbus, profibus, Ethernet, and fiberoptic.
[0095] Systems of the disclosure can include a hardware processor. The processor of the computer executes software, algorithms, and firmware in accordance with one or more example embodiments. The processor can be a central processing unit, a multi-core processing chip, SoC, a multi-chip module including multiple multi-core processing chips, or other hardware processor in one or more example embodiments. The processor is known by other names, including but not limited to a computer processor, a microprocessor, and a multi-core processor. The processor can also be an array of processors.
[0096] In one or more example embodiments, the processor executes software instructions stored in memory. Such software instructions can include generating machine learning models, executingHiNel / AI-traitselect / 868
[0097] machine learning models, performing analysis on data received from the database, and so forth. The memory includes one or more cache memories, main memory, or any other suitable type of memory. The memory can include volatile or non-volatile memory.
[0098] The processing system can be in communication with a computerized data storage system which can be stored in the storage repository. The data storage system can include a non-relational or relational data store, such as a MySQL or other relational database. Other physical and logical database types could be used. The data store may be a database server, such as Microsoft SQL Server, Oracle, IBM DB2, SQLITE, or any other database software, relational or otherwise. The data store may store the information identifying syntactical tags and any information required to operate on syntactical tags. In some embodiments, the processing system may use object-oriented programming and may store data in objects. In these embodiments, the processing system may use an object-relational mapper (ORM) to store the data objects in a relational database. The systems and methods described herein can be implemented using any number of physical data models. In one example embodiment, an RDBMS can be used. In those embodiments, tables in the RDBMS can include columns that represent coordinates. The tables can have pre-defined relationships between them. The tables can also have adjuncts associated with the coordinates.
[0099] In embodiments, the systems of the disclosure can include one or more I / O (input / output) devices allow a user to enter commands and information into the system, and also allow information to be presented to the user or other components or devices. Examples of input devices include, but are not limited to, a keyboard, a cursor control device (such as a mouse), a microphone, a touchscreen, and a scanner. Examples of output devices include, but are not limited to, a display device (e.g., a display, a monitor, or projector), speakers, outputs to a lighting network (such as a DMX card), a printer, and a network card. For example, the input devices can be used to enter data on native proteins and mutation sequences and assays. The input devices can also enter wanted functional data for a protein.
[0100] Various techniques are described herein in the general context of software.
[0101] Generally, software includes routines, programs, objects, components, data structures, and so forth that perform particular tasks or implement particular abstract data types. AnHiNel / AI-traitselect / 868
[0102] implementation of these modules and techniques can be stored on or transmitted across some form of computer readable media. Computer readable media is any available non-transitory medium or non-transitory media that is accessible by a computing device. By way of example, and not limitation, computer readable media includes computer storage media.
[0103] In embodiments, the Al Platform comprises a machine learning method, such as a neural network. In some embodiments, the Al platform includes neural networks, genetic algorithms, decision trees, fuzzy logic, symbolic rules, gradient boosting, support vector machines, and other machine learning based systems. Pluralities and / or combinations of the above may also be used. In embodiments, the Al Platform can use ML frameworks such as, Keras, Caffe, Pytorch, JAX, TensorFlow, the Microsoft Cognitive Toolkit, MXNet, Chainer, and Theano, with a Python implementation as the predominant data science language. In embodiments, the Al platform will allow for agnostic integration with other algorithms (such as gradient boosting, SVM, Gaussian processes) and their respective frameworks (XGBoost, SciKit Learn, GPy etc.) by separating data preparation from model creation and by using a NumPy data format common to all of these frameworks. In some embodiments, data preparation tools can be released as a Python package.
[0104] Embodiments of the disclosure use protein feature encodings to add physical or biological knowledge to amino acid sequences to create representations amenable to machine learning. As the choice of encoding varies based on the size and diversity of the input, as well as the task, several encoding methods can be implemented, allowing users to test and select the encodings most relevant to their problem. The Al Platform can include the following encodings, for example: one-hot, autoencoders, amino acid property encoders, learned BLOSUM / MSA evolutionary encodings, sequence mutation representation relative to WT, secondary structure / solvent accessible surface area encodings, learned AA embeddings, POOL, Phoenix, and / or structural / graph / topological encodings.
[0105] The above-described embodiments of the present invention can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. It should be appreciated that any component orHiNel / AI-traitselect / 868
[0106] collection of components that perform the functions described above can be generically considered as one or more controllers that control the above-discussed functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.
[0107] One or more processors may be interconnected by one or more networks in any suitable form, including as a local area network or a wide area network, such as an enterprise network or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks, or fiber optic networks.
[0108] One or more algorithms for controlling methods or processes provided herein may be embodied as a readable storage medium (or multiple readable media) (e.g., a computer memory, one or more floppy discs, compact discs (CD), optical discs, digital video disks (DVD), magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement the various methods or processes described herein.
[0109] In some embodiments, a computer readable storage medium may retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. Such a computer readable storage medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the methods or processes described herein. As used herein, the term "computer-readable storage medium" encompasses only a computer-readable medium that can be considered to be a manufacture (e.g., article of manufacture) or a machine. Alternatively, or additionally, methods or processes described herein may be embodied as a computer readable medium other than a computer-readable storage medium, such as a propagating signal.
[0110] The terms "program" or "software" are used herein in a generic sense to refer to any type of code or set of executable instructions that can be employed to program a computer or otherHiNel / AI-traitselect / 868
[0111] processor to implement various aspects of the methods or processes described herein. Additionally, it should be appreciated that according to one aspect of this embodiment, one or more programs that when executed perform a method or process described herein need not reside on a single computer or processor but may be distributed in a modular fashion amongst several different computers or processors to implement various procedures or operations.
[0112] Executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0113] Also, data structures may be stored in computer-readable media in any suitable form. Nonlimiting examples of data storage include structured, unstructured, localized, distributed, shortterm and / or long term storage. Non-limiting examples of protocols that can be used for communicating data include proprietary and / or industry standard protocols (e.g., HTTP, HTML, XML, JSON, SQL, web services, text, spreadsheets, etc., or any combination thereof). For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that conveys relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including using pointers, tags, or other mechanisms that establish relationship between data elements.
[0114] While several embodiments of the present invention have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the present invention. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings of the present invention is / are used. Those skilled in the art will recognize or be able to ascertain using no moreHiNel / AI-traitselect / 868
[0115] than routine experimentation, many equivalents to the specific embodiments of the invention described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, the invention may be practiced otherwise than as specifically described and claimed. The present invention is directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present invention.
[0116] As used herein in the specification and in the claims, the phrase "at least one," in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase "at least one" refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently "at least one of A and / or B") can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0117] In the claims, as well as in the specification above, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of and "consisting essentially of shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.HiNel / AI-traitselect / 868
[0118] 1. improve the
[0119]
[0120] nee of the neural network
[0121] Fifty-nine (59) different genetic elements (represented by genes in this case) were selected for having a potential impact on the complex trait leaf growth (see Lorenzo et al (2023) The Plant Cell, 35(1), 218-238). Five plasmids each containing 12 different gRNAs (each gRNA targets a specific gene) were constructed and corn plants constitutively expressing the Cas9 endonuclease were transformed with these plasmids. These transformants were further crossed (to create different combinations of edited genes) and the subsequent generations were phenotyped (measuring leaf size, represented as pseudo leaf area or PLA, which is the leaf length multiplied by the leaf width) and genotyped (using amplicon sequencing to determine the edits). Within the project, multiple experiments were performed in which plant seedlings were grown under water deficit conditions (to simulate drought). Our dataset contains more than 4000 plants with 0 to 17 genes knocked out, that were phenotyped and genotyped.
[0122] In order to identify the set of genetic elements (here genes) that influence the complex trait of interest (leaf growth, measured as PLA), the genotypes and phenotypes were coupled and analyzed. To do this, we created a deep learning model with and without embeddings. To create embeddings, we used AgroNT, a DNA language model with about 1 billion parameters which has been trained on 48 different plant genomes (Mendoza-Revilla et al (2024) Communications Biology, 7(1), 835). This neural network has learned a general representation of nucleotide sequences and was able to perform accurate predictions of regulatory sequences, promoter and terminator strength, tissue-specific gene expression and prioritization of functional variants. We investigated whether including embeddings would add some useful information about functional redundancy or relatedness between genes into the model. For each gene targeted for mutation, an embedding was created with AgroNT based on the nucleotide sequence. Subsequently, we trained two deep learning models (a standard model without embeddings and a model containing embeddings) to predict the PLA based on the genotype of a plant. The two models were identical in their architecture, except for the addition of embeddings and some layers to process these embeddings. We found that the model which includes embeddings was superior, outperforming the model without embeddings in 84% of the trials (see Figure 2).HiNel / AI-traitselect / 868
[0123] The final neural network model (with embeddings) showed a Pearson correlation of 0.63 for the predictions of PLA compared to the actual measured values on an independent test set (see Figure 3).
[0124] 2.ldentification of genome edits which contribute to the complex leaf growth trait
[0125] After training, the model was interrogated using explainable Al techniques (Lundberg et al (2017) arXiv preprint arXiv:1705.07874). This resulted into the detection of several genes that had large effects on the complex trait of interest (leaf growth, represented by PLA). The SHAP values (Figure 4) highlight the importance of certain genes (a highly positive SHAP value for a edited gene indicates that the edit has a positive effect on the PLA). The effects of mutations of these genes can be conditional on the presence of mutations in other genes. Our neural network predicts edits in GRF10 to be strongly associated with PLA. To a lesser extent than for GRF10, the model also puts forward edits in TCP42 to have a positive effect on PLA. In Impens et al (2023) New Phytologist 239(4), 1521-1532) a mutation of GRF10 was found to increase PLA (although not significantly compared to the controls in the numbers tested), while a mutation in TCP42 did not affect PLA. This indicates that the positive effect of mutations in TCP42 are conditional on other genes. Indeed, the combined mutation of GRF10 and TCP42 increased PLA even more than mutations in GRF10 alone and was significant compared to the controls. Independent research confirmed the importance of BIN2 and BIN2.4 by expressing an RNAi molecule targeting the members of the BIN2 family (including BIN2 and BIN2.4). This silencing of the BIN2 family resulted in an increased leaf length, which is a component of PLA.
Claims
HiNel / AI-traitselect / 868Claims1. A computer-implemented method for identifying at least two genome edits contributing to a selected complex trait in plants comprising:i) selecting genetic elements predicted to be associated with a specific complex trait by applying an artificial intelligence (Al) model, trained on biological data from plants, plant cells, plant organs or plant seeds, wherein genetic elements in said Al model are associated with features of complex traits, and wherein said Al model comprises embeddings,ii) genotyping a heterogenous population of genome edited plants, plant cells, plant organs or plant seeds wherein the genome edits are present in the genetic elements selected in step i), and phenotyping each of said genotyped plants, plant cells, plant organs or plant seeds, wherein phenotyping comprises measuring said selected complex trait, andiii) calculating the contribution of genome edits in the genotyped and phenotyped population of plants, plant cells, plant organs or plant seeds by applying the Al model from step i) to the selected complex trait in plants, plant cells, plant organs or plant seeds, andiv) identifying combinations of at least two genome edits that contribute to a selected complex trait in plants.
2. A method according to claim 1 wherein the contribution of genome edits that contribute to a selected complex trait of step iv) are introduced into the Al model of step i).
3. A method according to claims 1 and 2 wherein said genome edited plants are generated by a multiplex CRISPR method.
4. A method according to claims 1, 2 or 3 wherein the artificial intelligence model comprises a graph neural network.
5. A method according to any one of claims 1 to 4 wherein said embeddings are at the level of genetic elements and / or a plant feature related to the complex trait and / or the complex trait level.
6. A computer-readable storage medium which stores computer-executable instruction that, when executed by at least one processor, cause the processor to perform a method of any one of claims 1 to 5.22HiNel / AI-traitselect / 8687. An apparatus comprising control circuitry configured to perform a method of any one ofclaims 1 to 5.