ML Enhancer-Promoter Prediction via Engineered Epigenomic Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining functional enhancer-promoter interactions, such as the Activity by Contact (ABC) model, have limited accuracy in predicting functional enhancer-promoter pairs, which hinders the development of targeted therapies for diseases.
Innovation Solution
The implementation of machine learning models that analyze specific features from epigenomic datasets, including engineered features, to accurately predict functional enhancer-promoter pairs by distinguishing between functional and non-functional interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional methods such as the Activity by Contact (ABC) model are used to predict functional enhancer-promoter pairs, then the method is simple and easy to implement, but the prediction accuracy is limited and varying
Solution Approach 1:
The patent segments the feature set into two distinct categories: (1) features directly extracted from epigenomic datasets, and (2) engineered features derived from combinations of the first set. This segmentation allows the model to systematically process different types of information while improving prediction accuracy through the addition of engineered features that capture complex relationships between epigenomic markers.
Solution Approach 2:
The patent creates composite features by combining multiple epigenomic features (such as chromatin accessibility, histone modifications, and transcription factor binding) into integrated predictive models. These composite features represent complex biological interactions that individual features alone cannot capture, thereby improving prediction accuracy while maintaining a structured approach to model complexity.
2Measurement precision
If more features are extracted and engineered from epigenomic datasets, then the prediction accuracy improves, but the computational complexity and data processing requirements increase
Solution Approach 1:
The patent performs preliminary feature engineering by pre-defining the structure and relationships of epigenomic features before model training. Engineered features are systematically constructed from the first set of features, establishing a robust feature hierarchy that guides the machine learning process. This preliminary structuring reduces computational complexity during model execution while maintaining high prediction accuracy.
Solution Approach 2:
The patent transforms raw epigenomic data into multiple derived parameters through feature engineering, creating a hierarchical parameter structure where engineered features are functions of base features. This parameter transformation allows the model to capture non-linear relationships and interactions while managing computational complexity through a systematic parameter hierarchy that can be efficiently processed by machine learning algorithms.
Data Source
AI summary
Disclosed herein are methods for implementing machine learning models to analyze features from epigenomic datasets to determine whether enhancer-promoter pairs are functional or non-functional. Features can include a first set of features extracted from the epigenomic datasets. Furthermore, features can include a second set of features engineered from features of the first set. Machine learning models that incorporate features, including the first set of features and engineered second set of features, can predict, with improved metrics, whether enhancer-promoter pairs are functional or non-functional.


