Cell-Specific CRE Prediction Using MPRA-Trained Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately quantify the gene-regulatory potential of DNA sequences at nucleotide resolution, particularly in a cell or tissue-specific manner, due to the intractability of testing every element in the human genome using Massively Parallel Reporter Assays (MPRAs.
Innovation Solution
A computer-implemented method using a machine learning network trained on MPRA data sets to predict the activity of cis-regulatory elements (CREs) with cell-type, cell state, or environment specificity, employing neural networks to process nucleic acid sequences and generate predictions of CRE activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Massively Parallel Reporter Assays (MPRAs) are used to directly characterize cis-regulatory function, then measurement precision of CRE activity is improved, but device complexity and productivity deteriorate due to the intractability of testing every element in the human genome
Solution Approach 1:
The patent creates computational copies of MPRA functionality through machine learning models. Instead of physically testing every genomic element via MPRA, the system trains ML models on a subset of experimentally characterized sequences, then uses these models to predict CRE activity across the entire genome. This computational copying approach maintains measurement precision while achieving genome-wide coverage.
Solution Approach 2:
The patent performs preliminary MPRA experiments on a representative subset of genomic sequences to generate training data before scaling to genome-wide analysis. By pre-characterizing a curated set of CREs and using this data to train predictive models, the system establishes a foundation that enables subsequent high-throughput prediction without requiring exhaustive experimental testing of every genomic element.
2Productivity
If computational methods are used to predict CRE activity, then productivity is improved by enabling genome-wide analysis, but measurement precision deteriorates compared to direct MPRA measurement
Solution Approach 1:
The patent creates computational copies of MPRA functionality through machine learning models. Instead of physically testing every genomic element via MPRA, the system trains ML models on a subset of experimentally characterized sequences, then uses these models to predict CRE activity across the entire genome. This computational copying approach maintains measurement precision while achieving genome-wide coverage.
Solution Approach 2:
The patent replaces the mechanical MPRA experimental system with a computational machine learning system. The ML models substitute for the physical assay machinery, using learned patterns from training data to predict CRE activity. This substitution enables genome-wide analysis while maintaining accuracy by capturing the underlying biological relationships in computational form.
3Adaptability or versatility
If cell-type specific CREs are designed to enhance gene expression in targeted cells, then adaptability is improved, but device complexity increases due to the need for cell-specific optimization
Solution Approach 1:
The patent applies local quality by designing CREs with sequence features specifically optimized for target cell types. The system identifies and incorporates cell-type-specific regulatory motifs and sequence characteristics that enable selective activity in desired cell populations. This localized optimization of sequence properties achieves cell-type specificity without requiring entirely different CRE designs for each cell type.
Solution Approach 2:
The patent changes sequence parameters of CREs to achieve cell-type specificity. By adjusting nucleotide composition, motif density, and sequence features based on cell-type-specific patterns learned from training data, the system generates CREs with tailored activity profiles. These parameter modifications enable precise control of gene expression in target cells while maintaining a unified design framework.
Data Source
AI summary
Described in certain embodiments herein are computer implemented methods, systems, and computer program products that can be used to identify or engineered cell specific cis-regulatory elements (CREs). Also described herein are cell specific CREs and uses thereof.


