Machine Learning Model for Enhancer Activity Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods struggle to accurately model enhancer selectivity and tissue-specific gene expression due to the complexity of enhancer grammar and the repetitive nature of the genome, making it difficult to identify sequence determinants of enhancer-mediated gene expression.
Innovation Solution
A system and method using a training dataset of tandem n-mer repeat sequences to model enhancer activity, where each sequence is associated with measured activity in various in vivo states, allowing for the identification of binding components and their grammar, and optimizing models for specific tissue activities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If genome-scale experiments are used to study enhancer activity, then large amounts of data about transcription factor binding and expression are obtained, but the complexity of the repetitive genome and enhancer grammar makes it difficult to identify sequence determinants of tissue-specific gene expression
Solution Approach 1:
The patent segments the complex enhancer grammar problem into manageable components by using a hierarchical modeling approach. The model breaks down enhancer sequences into individual transcription factor binding sites and their grammatical relationships, allowing systematic analysis of sequence determinants without being overwhelmed by the full complexity of natural enhancers.
Solution Approach 2:
The patent changes the parameter of sequence complexity by using simplified synthetic enhancer sequences with controlled numbers and arrangements of transcription factor binding sites. This parameter change allows the model to learn grammatical rules from less complex sequences that can then be applied to predict activity of more complex natural enhancers.
2Loss of information
If the human genome is analyzed to find all combinations of transcription factors, then complete genomic information is available, but the genome is too short to encode all possible combinations, orientations and spacings of approximately 1,639 human transcription factors
Solution Approach 1:
The patent performs preliminary action by using in vitro transcription factor binding data and expression data to pre-characterize the binding properties of 1,639 human transcription factors before analyzing genomic sequences. This preliminary characterization allows the model to predict tissue-specific enhancer activity without requiring the genome to physically contain all possible combinations.
Solution Approach 2:
The patent creates computational copies of transcription factor binding sites in various combinations and arrangements through synthetic enhancer sequence design. These copied and recombined binding sites allow the model to learn from a virtual library of all possible combinations without requiring physical genomic space for all variations.
3Reliability
If models are trained on natural enhancer sequences, then real biological complexity is captured, but the repetitive nature of the genome and tissue-specific activity patterns make it difficult to learn generalizable rules for enhancer grammar
Solution Approach 1:
The patent applies local quality by training the model on tissue-specific enhancer sequences from different cell types, allowing the model to learn tissue-specific grammatical rules. Each training set has local quality tailored to specific tissue contexts, enabling the model to adapt and generalize across different biological conditions while maintaining biological accuracy.
Data Source
AI summary
A method of modeling enhancer activity obtains a training dataset in electronic form. The dataset comprises a plurality of training enhancers and, for each such enhancer, a corresponding measured amount of activity of the enhancer in each of one or more states. Each enhancer is a tandem repeat. A model comprising a plurality of parameters is trained by inputting each training enhancer into the model. Upon input, the model applies the parameters to the training enhancer to generate, as model output, a corresponding predicted activity of the training enhancer for a first state of the one or more states. The plurality of parameters is refined based differentials between corresponding measured and predicted amounts of activity in the first state for each of the training enhancer. A plurality of test enhancers is generated using the trained model and their activity tested thereby modeling enhancer activity.


