Human Regulome Database Integrating Non-Coding Genomic Variants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods struggle to identify and integrate functional information for non-protein coding regions in the human genome, making it difficult to interpret genome sequences and associate variants with disease phenotypes, especially since existing tools focus on protein-coding regions and lack comprehensive resources for low-throughput data from individual labs and consortia.
Innovation Solution
A Resource for the Human Regulome database is created to collect and integrate high-quality experimental results from intergenic and non-coding regions, annotating regulatory elements with controlled vocabularies, linking sequence variations to gene regulation and disease phenotypes, and using transcription factor binding as a biologically relevant biomarker.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If high-throughput methods are used to identify DNA elements, then broad coverage of the genome is achieved, but the direct mechanism of regulation between nucleotides and target genes cannot be identified
Solution Approach 1:
The patent combines high-throughput sequencing data with low-throughput experimental validation data into a unified database resource. This merging allows the system to maintain broad genome coverage while recovering detailed regulatory mechanism information that would be lost in high-throughput alone. The integrated resource links DNA elements to their target genes and regulatory mechanisms through multiple data types including ChIP-Seq, ATAC-Seq, and experimental validation records.
Solution Approach 2:
The patent introduces an intermediary database resource that acts as a bridge between high-throughput genomic data and functional validation data. This intermediary resource stores and organizes information about DNA elements, their regulatory mechanisms, and target gene associations, allowing researchers to access both broad coverage and detailed mechanistic information through a single integrated system.
2Productivity
If computational algorithms are used to predict regulatory regions, then hypotheses of functional nucleotides are generated, but biological significance must still be evaluated manually
Solution Approach 1:
The patent performs preliminary computational analysis and filtering of regulatory region hypotheses before manual evaluation. The system uses computational algorithms to predict regulatory regions, then pre-evaluates them using multiple data types and criteria, so that only the most promising hypotheses require manual biological significance evaluation. This preliminary action reduces the time burden on manual reviewers while maintaining high productivity in hypothesis generation.
3Measurement precision
If existing tools focus on protein-coding regions, then annotation of coding genes is improved, but interpretation of non-coding variants and their association with disease phenotypes is limited
Solution Approach 1:
The patent creates a universal database resource that handles both protein-coding and non-coding genomic regions with equal capability. The system is designed to annotate and analyze all types of genomic variants including coding genes, regulatory DNA elements, and non-coding RNAs. This multi-functional resource eliminates the limitation of tools that focus only on coding regions, allowing comprehensive interpretation of both coding and non-coding variants and their associations with disease phenotypes.
4Reliability
If low-throughput experimental data from individual labs is not integrated, then each lab maintains data quality control, but comprehensive resources for regulatory elements are insufficient
Solution Approach 1:
The patent merges data from multiple individual laboratories and consortia into a single integrated database resource. The system maintains data quality control through standardized validation criteria while combining datasets to achieve comprehensive coverage of regulatory elements. This merging allows the resource to preserve the reliability of individual lab data while achieving the quantity and comprehensiveness that no single lab could produce alone.
Data Source
AI summary
Measuring of the binding of a transcription factor (using, for example, chromatin immunoprecipitation) according to the present invention is provides an improved marker for a disease. These markers can be used in diagnostics for diseases where a transcription factor binding event plays a role. Additionally, they can be used to adjust disease risk profiles for healthy individuals as with typical genetic variants.


