DNA Fragment Identification via Genomic Coverage Histograms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic sequencing methods, particularly next-generation sequencing (NGS), face challenges in accurately identifying long DNA fragments and structural variations, especially in regions with repetitive elements and long translocations, which are crucial for diagnostic and research purposes but are not efficiently addressed by existing commercial technologies.

Innovation Solution

The method involves grouping short reads to identify long DNA fragments by creating histograms of genomic coverage, allowing for the detection of structural variations by correlating common increases or decreases in these histograms, thereby determining the haploid genome and identifying structural rearrangements such as insertions, deletions, and translocations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short reads are used for sequencing, then sequencing cost and time are reduced, but the ability to identify long DNA fragments and structural variations is compromised

Engineering Contradiction:
Improvesequencing speedVSAvoidstructural variation detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments long DNA fragments into multiple short reads for sequencing, then uses computational methods to reassemble and analyze them. By dividing the long fragment into sequencable short segments and using the spatial distribution and pairing information of these segments, the system achieves both high-throughput sequencing and accurate structural variation detection

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a spatial dimension to short read analysis by examining the physical distance and orientation between paired reads. Instead of analyzing reads in isolation, the system uses the two-dimensional spatial relationship (distance and orientation) between mate-pair reads to infer long-range genomic structure and detect structural variations that single reads cannot reveal

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If conventional mate-pair sequencing is used, then some structural variations can be detected, but translocations beyond mate-pair distance range cannot be identified

Engineering Contradiction:
Improvestructural variation detection rangeVSAvoidsequencing methodology complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple types of information including read pairing relationships, spatial distance data, orientation information, and histogram analysis to create a comprehensive structural variation detection system. By combining these different data dimensions and analytical approaches, the system extends detection range beyond what any single method could achieve

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If short sequence data is used, then sequencing cost is reduced, but haploid genome identification becomes difficult

Engineering Contradiction:
Improvesequencing efficiencyVSAvoidhaplotype information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent uses feedback loops where initial assembly results inform subsequent refinement steps. The system iteratively improves haplotype phasing by using detected structural variations and read pairing information to resolve ambiguities, then uses this improved information to further refine the assembly, progressively recovering haplotype-specific information from short reads

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9514272B2Identification of DNA fragments and structural variations
Publication Date: 2016.12.06 COMPLETE GENOMICS INC
  • US9514272B2 patent drawing
  • US9514272B2 patent drawing
  • US9514272B2 patent drawing

AI summary

Various short reads can be grouped and identified as coming from a same long DNA fragment (e.g., by using wells with a relatively low-concentration of DNA). A histogram of the genomic coverage of a group of short reads can provide the edges of the corresponding long fragment (pulse). The knowledge of these pulses can provide an ability to determine the haploid genome and to identify structural variations.