Genomic Region Analysis via Information Content

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing genomic regions are hindered by the complexity and volume of epigenomic data, making it impractical to identify genomic regions of interest for research or therapeutic applications, as existing computer-implemented techniques are overwhelmed and unable to provide timely analysis.

Innovation Solution

The use of information theory techniques to determine the information content of genomic regions by evaluating chromatin states, which are defined by sets of chromatin characteristics, allowing for the identification of patterns and relationships in chromatin state occurrences across different cell types and organisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing computer-implemented techniques are used to analyze genomic regions, then comprehensive epigenomic data can be processed, but the analysis becomes overwhelmed and unable to provide timely results

Engineering Contradiction:
Improveanalysis accuracyVSAvoidanalysis speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts only the most relevant features from comprehensive epigenomic data using information theory-based metrics. Instead of analyzing all epigenomic data points, the system identifies and extracts key chromatin state features that carry the most information about genomic region functionality, thereby reducing computational complexity while maintaining analysis accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces information theory-based metrics as an intermediary layer between raw epigenomic data and genomic region identification. These metrics serve as a mediator that transforms complex multi-dimensional epigenomic data into simplified scoring systems that can be efficiently processed, bridging the gap between comprehensive data and timely analysis

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive epigenomic data is analyzed to identify genomic regions of interest, then accurate identification can be achieved, but the complexity and volume of data make it impractical

Engineering Contradiction:
Improveregion identification accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameters used to represent epigenomic data by applying information theory transformations. Instead of working with raw chromatin state annotations and epigenomic marks, the system transforms these into information content scores and statistical metrics that capture essential patterns while reducing data dimensionality and processing complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the complex epigenomic analysis task into distinct computational steps: calculating information content for individual genomic regions, aggregating scores across cell types, identifying statistical patterns, and ranking regions of interest. This segmentation allows each sub-task to be processed independently and efficiently, reducing overall system complexity

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If all genomic regions are analyzed in detail, then complete coverage is achieved, but the time required for analysis becomes prohibitive

Engineering Contradiction:
Improveanalysis completenessVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by focusing computational resources on calculating information theory metrics for only the most promising genomic regions rather than performing exhaustive detailed analysis on all regions. The system identifies and prioritizes regions with highest information content scores, analyzing them in detail while using summary statistics for the remainder, thereby achieving effective coverage with reduced time investment

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11195596B2Analyzing characteristics of genomic regions of a genome
Publication Date: 2021.12.07 MASSACHUSETTS INST OF TECH
  • US11195596B2 patent drawing
  • US11195596B2 patent drawing
  • US11195596B2 patent drawing

AI summary

Embodiments of techniques for analyzing one or more genomic regions of a genome of an organism. Data about a genomic region may be analyzed to determine an information content of the genomic region, which may indicate an amount of information provided by the genomic region. The data about the genomic region may be or include data identifying a chromatin state for the genomic region. A chromatin state may be one of a set of chromatin states that each define a different set of one or more chromatin characteristics. Chromatin characteristics may be structural and/or functional features of genomic regions. A chromatin state of a genomic region may be determined from, and describe, the genomic region such that when a genomic region has a set of one or more chromatin characteristics, a chromatin state associated with that combination of one or more chromatin characteristics is identified for the genomic region.