Machine Learning Model for Enhancer Activity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods struggle to accurately model enhancer selectivity and tissue-specific gene expression due to the complexity of enhancer grammar and the repetitive nature of the genome, making it difficult to identify sequence determinants of enhancer-mediated gene expression.

Innovation Solution

A system and method using a training dataset of tandem n-mer repeat sequences to model enhancer activity, where each sequence is associated with measured activity in various in vivo states, allowing for the identification of binding components and their grammar, and optimizing models for specific tissue activities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If genome-scale experiments are used to study enhancer activity, then large amounts of data about transcription factor binding and expression are obtained, but the complexity of the repetitive genome and enhancer grammar makes it difficult to identify sequence determinants of tissue-specific gene expression

Engineering Contradiction:
Improveamount of dataVSAvoidcomplexity of enhancer grammar
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the complex enhancer grammar problem into manageable components by using a hierarchical modeling approach. The model breaks down enhancer sequences into individual transcription factor binding sites and their grammatical relationships, allowing systematic analysis of sequence determinants without being overwhelmed by the full complexity of natural enhancers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of sequence complexity by using simplified synthetic enhancer sequences with controlled numbers and arrangements of transcription factor binding sites. This parameter change allows the model to learn grammatical rules from less complex sequences that can then be applied to predict activity of more complex natural enhancers.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If the human genome is analyzed to find all combinations of transcription factors, then complete genomic information is available, but the genome is too short to encode all possible combinations, orientations and spacings of approximately 1,639 human transcription factors

Engineering Contradiction:
Improvegenomic informationVSAvoidnumber of transcription factor combinations
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent performs preliminary action by using in vitro transcription factor binding data and expression data to pre-characterize the binding properties of 1,639 human transcription factors before analyzing genomic sequences. This preliminary characterization allows the model to predict tissue-specific enhancer activity without requiring the genome to physically contain all possible combinations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates computational copies of transcription factor binding sites in various combinations and arrangements through synthetic enhancer sequence design. These copied and recombined binding sites allow the model to learn from a virtual library of all possible combinations without requiring physical genomic space for all variations.

Inventive Principle:
Principle #26Copying

3Reliability

If models are trained on natural enhancer sequences, then real biological complexity is captured, but the repetitive nature of the genome and tissue-specific activity patterns make it difficult to learn generalizable rules for enhancer grammar

Engineering Contradiction:
Improvebiological accuracyVSAvoidgeneralizability of rules
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by training the model on tissue-specific enhancer sequences from different cell types, allowing the model to learn tissue-specific grammatical rules. Each training set has local quality tailored to specific tissue contexts, enabling the model to adapt and generalize across different biological conditions while maintaining biological accuracy.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250022532A1Regulating enhancer activity using machine-learning
Publication Date: 2025.01.16 SHAPE THERAPEUTICS INC
  • US20250022532A1 patent drawing
  • US20250022532A1 patent drawing
  • US20250022532A1 patent drawing

AI summary

A method of modeling enhancer activity obtains a training dataset in electronic form. The dataset comprises a plurality of training enhancers and, for each such enhancer, a corresponding measured amount of activity of the enhancer in each of one or more states. Each enhancer is a tandem repeat. A model comprising a plurality of parameters is trained by inputting each training enhancer into the model. Upon input, the model applies the parameters to the training enhancer to generate, as model output, a corresponding predicted activity of the training enhancer for a first state of the one or more states. The plurality of parameters is refined based differentials between corresponding measured and predicted amounts of activity in the first state for each of the training enhancer. A plurality of test enhancers is generated using the trained model and their activity tested thereby modeling enhancer activity.