Human Regulome Database Integrating Non-Coding Genomic Variants

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods struggle to identify and integrate functional information for non-protein coding regions in the human genome, making it difficult to interpret genome sequences and associate variants with disease phenotypes, especially since existing tools focus on protein-coding regions and lack comprehensive resources for low-throughput data from individual labs and consortia.

Innovation Solution

A Resource for the Human Regulome database is created to collect and integrate high-quality experimental results from intergenic and non-coding regions, annotating regulatory elements with controlled vocabularies, linking sequence variations to gene regulation and disease phenotypes, and using transcription factor binding as a biologically relevant biomarker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If high-throughput methods are used to identify DNA elements, then broad coverage of the genome is achieved, but the direct mechanism of regulation between nucleotides and target genes cannot be identified

Engineering Contradiction:
Improvegenome coverageVSAvoidregulatory mechanism information
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The patent combines high-throughput sequencing data with low-throughput experimental validation data into a unified database resource. This merging allows the system to maintain broad genome coverage while recovering detailed regulatory mechanism information that would be lost in high-throughput alone. The integrated resource links DNA elements to their target genes and regulatory mechanisms through multiple data types including ChIP-Seq, ATAC-Seq, and experimental validation records.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary database resource that acts as a bridge between high-throughput genomic data and functional validation data. This intermediary resource stores and organizes information about DNA elements, their regulatory mechanisms, and target gene associations, allowing researchers to access both broad coverage and detailed mechanistic information through a single integrated system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If computational algorithms are used to predict regulatory regions, then hypotheses of functional nucleotides are generated, but biological significance must still be evaluated manually

Engineering Contradiction:
Improvehypothesis generation rateVSAvoidmanual evaluation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary computational analysis and filtering of regulatory region hypotheses before manual evaluation. The system uses computational algorithms to predict regulatory regions, then pre-evaluates them using multiple data types and criteria, so that only the most promising hypotheses require manual biological significance evaluation. This preliminary action reduces the time burden on manual reviewers while maintaining high productivity in hypothesis generation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If existing tools focus on protein-coding regions, then annotation of coding genes is improved, but interpretation of non-coding variants and their association with disease phenotypes is limited

Engineering Contradiction:
Improvecoding gene annotation precisionVSAvoidnon-coding variant functional information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent creates a universal database resource that handles both protein-coding and non-coding genomic regions with equal capability. The system is designed to annotate and analyze all types of genomic variants including coding genes, regulatory DNA elements, and non-coding RNAs. This multi-functional resource eliminates the limitation of tools that focus only on coding regions, allowing comprehensive interpretation of both coding and non-coding variants and their associations with disease phenotypes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If low-throughput experimental data from individual labs is not integrated, then each lab maintains data quality control, but comprehensive resources for regulatory elements are insufficient

Engineering Contradiction:
Improvedata qualityVSAvoidregulatory element data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent merges data from multiple individual laboratories and consortia into a single integrated database resource. The system maintains data quality control through standardized validation criteria while combining datasets to achieve comprehensive coverage of regulatory elements. This merging allows the resource to preserve the reliability of individual lab data while achieving the quantity and comprehensiveness that no single lab could produce alone.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9946835B2Method and system for the use of biomarkers for regulatory dysfunction in disease
Publication Date: 2018.04.17 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US9946835B2 patent drawing
  • US9946835B2 patent drawing
  • US9946835B2 patent drawing

AI summary

Measuring of the binding of a transcription factor (using, for example, chromatin immunoprecipitation) according to the present invention is provides an improved marker for a disease. These markers can be used in diagnostics for diseases where a transcription factor binding event plays a role. Additionally, they can be used to adjust disease risk profiles for healthy individuals as with typical genetic variants.