Explainable Machine Learning for Plant Gene Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional breeding methods for plants face limitations in controlling genetic diversity and efficiently identifying gene modifications that result in desired phenotypes due to the random nature of recombination and mutagenesis, leading to lengthy and unpredictable processes.
Innovation Solution
A method utilizing explainable machine learning to analyze gene expression profiles and identify candidate gene targets for editing, employing models like deep neural networks and Gaussian processes to predict phenotypes and generate ideal gene expression profiles, thereby enabling targeted genome edits for specific phenotypic changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional breeding methods using random recombination and mutagenesis are used, then genetic diversity is generated, but the process becomes lengthy and unpredictable
Solution Approach 1:
The patent applies preliminary action by using machine learning models to predict phenotypic outcomes before actual breeding experiments are conducted. The system analyzes gene expression profiles and predicts which genetic modifications will produce desired phenotypes, allowing breeders to plan and execute targeted experiments rather than relying on random mutagenesis and lengthy selection processes.
2Adaptability or versatility
If random mutagenesis is used to generate genetic variants, then genetic variation is increased, but identification of mutations of interest becomes long and labor-intensive
Solution Approach 1:
The patent replaces the mechanical and manual process of screening mutagenized plants with an automated machine learning system. The system uses nonlinear algorithms to analyze gene expression profiles and automatically identify which mutations are likely to produce desired phenotypes, substituting manual observation and selection with computational prediction.
3Measurement precision
If conventional genetic diversity analysis is used, then assessment of genetic divergence is achieved, but gene discovery and identification of gene modifications conducive to desired phenotype remain limited
Solution Approach 1:
The patent introduces machine learning models as an intermediary between genetic diversity analysis and gene discovery. The system takes gene expression profiles as input and uses nonlinear algorithms to predict phenotypic outcomes, thereby bridging the gap between measuring genetic divergence and actually discovering which genes are responsible for desired traits.
Data Source
AI summary
The present disclosure relates to leveraging explainable machine learning methods and feature importance mechanisms as a mechanism for gene discovery and furthermore leveraging the outputs of the gene discovery to recommend ideal gene expression profiles and the requisite genome edits that are conducive to a desired phenotype. Particularly, aspects of the present disclosure are directed to obtaining gene expression profiles for a set of genes measured in a tissue sample of a plant, inputting the gene expression profiles into a prediction model constructed for a task of predicting a phenotype as output data, generating, using the prediction model, the prediction of the phenotype for the plant, analyzing, by an explainable artificial intelligence system, decisions made by the prediction model to predict the phenotype, and identifying a set of candidate gene targets for the phenotype as having a largest contribution or influence on the prediction based on the analyzing.


