DNA Fragment Identification via Genomic Coverage Histograms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomic sequencing methods, particularly next-generation sequencing (NGS), face challenges in accurately identifying long DNA fragments and structural variations, especially in regions with repetitive elements and long translocations, which are crucial for diagnostic and research purposes but are not efficiently addressed by existing commercial technologies.
Innovation Solution
The method involves grouping short reads to identify long DNA fragments by creating histograms of genomic coverage, allowing for the detection of structural variations by correlating common increases or decreases in these histograms, thereby determining the haploid genome and identifying structural rearrangements such as insertions, deletions, and translocations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short reads are used for sequencing, then sequencing cost and time are reduced, but the ability to identify long DNA fragments and structural variations is compromised
Solution Approach 1:
The patent segments long DNA fragments into multiple short reads for sequencing, then uses computational methods to reassemble and analyze them. By dividing the long fragment into sequencable short segments and using the spatial distribution and pairing information of these segments, the system achieves both high-throughput sequencing and accurate structural variation detection
Solution Approach 2:
The patent adds a spatial dimension to short read analysis by examining the physical distance and orientation between paired reads. Instead of analyzing reads in isolation, the system uses the two-dimensional spatial relationship (distance and orientation) between mate-pair reads to infer long-range genomic structure and detect structural variations that single reads cannot reveal
2Measurement precision
If conventional mate-pair sequencing is used, then some structural variations can be detected, but translocations beyond mate-pair distance range cannot be identified
Solution Approach 1:
The patent merges multiple types of information including read pairing relationships, spatial distance data, orientation information, and histogram analysis to create a comprehensive structural variation detection system. By combining these different data dimensions and analytical approaches, the system extends detection range beyond what any single method could achieve
3Productivity
If short sequence data is used, then sequencing cost is reduced, but haploid genome identification becomes difficult
Solution Approach 1:
The patent uses feedback loops where initial assembly results inform subsequent refinement steps. The system iteratively improves haplotype phasing by using detected structural variations and read pairing information to resolve ambiguities, then uses this improved information to further refine the assembly, progressively recovering haplotype-specific information from short reads
Data Source
AI summary
Various short reads can be grouped and identified as coming from a same long DNA fragment (e.g., by using wells with a relatively low-concentration of DNA). A histogram of the genomic coverage of a group of short reads can provide the edges of the corresponding long fragment (pulse). The knowledge of these pulses can provide an ability to determine the haploid genome and to identify structural variations.


