Transcriptionally Active Region Quantification via Genomic Binning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sequencing technologies face challenges in identifying transcriptionally active regions and quantifying sequence abundance, particularly for non-coding RNAs, due to inexact mapping and high computational intensity in high-throughput sequencing settings.

Innovation Solution

A sequence data analysis system using a reference sequence-based design that aligns sequencing output sequences to reference sequences, calculates coverage counts, and determines contiguous high-coverage regions to identify transcriptionally active regions and quantify their abundance, reducing computational complexity through coordinate comparisons rather than traditional sequence alignment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional sequence alignment methods are used to identify transcriptionally active regions, then mapping accuracy may be maintained, but computational intensity and processing time increase significantly

Engineering Contradiction:
Improvemapping accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the reference genome into discrete bins or intervals, and assigns sequencing reads to these bins based on positional information rather than performing full sequence alignment. This segmentation transforms the continuous sequence matching problem into discrete bin assignment, dramatically reducing computational complexity while maintaining the ability to identify transcriptionally active regions through coverage depth analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mapping approach where reads are mapped to genomic coordinates or bins without requiring precise sequence alignment. This intermediary representation (genomic position/ bin assignment) serves as a mediator between the raw sequencing data and the final expression quantification, enabling faster processing while preserving the essential information needed for identifying active transcription regions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If comprehensive sequence alignment is performed for all sequencing output, then accurate identification of transcriptionally active regions is achieved, but computational resources and storage requirements increase

Engineering Contradiction:
Improveidentification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary binning or interval assignment of sequencing reads based on their genomic coordinates before any detailed analysis. This preliminary action organizes the data into discrete bins that represent potential transcriptionally active regions, allowing subsequent analysis to focus only on these pre-identified regions rather than processing the entire genome sequence data, thus reducing computational complexity while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transitions from sequence-space analysis to genomic coordinate-space analysis by mapping reads to specific genomic positions or bins. This dimensional change from comparing sequence strings to analyzing coverage depth across genomic intervals simplifies the computational problem while preserving the ability to accurately identify transcriptionally active regions through coverage-based metrics.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of manufacture

If traditional RNA expression profiling methods are used, then established protocols are followed, but flexibility for small and non-coding RNA expression profiling is limited

Engineering Contradiction:
Improveprotocol standardizationVSAvoidmethod flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent employs a universal binning-based approach that can handle diverse RNA types (coding and non-coding) uniformly. By representing all sequencing reads in terms of their genomic bin assignments rather than requiring type-specific alignment protocols, the method achieves multi-functionality across different RNA species while maintaining a standardized analytical framework, thus enhancing versatility without sacrificing protocol standardization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8566039B2Method and system to characterize transcriptionally active regions and quantify sequence abundance for large scale sequencing data
Publication Date: 2013.10.22 GENOMIC HEALTH INC
  • US8566039B2 patent drawing
  • US8566039B2 patent drawing
  • US8566039B2 patent drawing

AI summary

This invention provides a quantitative method to determine transcriptionally active regions and quantify sequence abundance from large scale sequencing data. The invention also provides a system based on reference sequences to design and implement the method. The system processes large scale sequence data from high throughput sequencing, generates transcriptionally active region sequences as necessary, and quantifies the sequence abundance of the gene or transcriptionally active region. The method and system are useful for many analyses based on RNA expression profiling.