Hierarchical Genome Assembly Using SMRT Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid sequencing technologies face challenges in accurately assembling genome sequences due to errors, repeats, and the need for multiple libraries and sequencing methods, which complicates the identification of true variants and consensus calling, especially in regions with low coverage or high error rates.
Innovation Solution
A computer-implemented method for nucleic acid sequence identification using a hierarchical genome assembly process (HGAP) that processes long-insert SMRT sequencing data to generate pre-assembled reads, which are then assembled into contigs, utilizing a consensus algorithm to improve accuracy and handle errors, allowing for single-library, single-method genome assembly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple libraries and sequencing methods are used to improve assembly accuracy, then consensus sequence accuracy is improved, but device complexity and process complexity increase
Solution Approach 1:
The patent combines multiple sequencing methods (SMRT long-read sequencing and Illumina short-read sequencing) and multiple library types (standard insert libraries and mate-pair libraries) into a unified assembly pipeline. This merging allows the system to leverage the strengths of each method—long reads for spanning repeats and structural variants, short reads for high base-level accuracy—while managing complexity through integrated software tools that coordinate the diverse data sources.
Solution Approach 2:
The assembly process is segmented into distinct stages: initial assembly using SMRT long reads to establish scaffolds, followed by polishing with Illumina short reads to correct base-level errors. This segmentation allows each component to be optimized independently while contributing to the overall accuracy goal, resolving the contradiction by breaking down the complex task into manageable steps.
2Measurement precision
If multiple libraries and sequencing methods are used to improve assembly accuracy, then consensus sequence accuracy is improved, but ease of operation deteriorates
Solution Approach 1:
The patent introduces intermediate data structures and processing steps that mediate between the diverse input formats from different sequencing platforms. Tools like minimus2 and custom scripts serve as intermediaries that standardize data representation, handle format conversions, and coordinate the assembly of reads from multiple sources, thereby simplifying the operational complexity for users.
3Loss of information
If long-insert library is used to resolve repeats, then assembly completeness is improved, but manufacturing precision requirements increase
Solution Approach 1:
The patent employs parameter changes in library preparation, specifically using long-insert sizes (e.g., 10-20 kb or larger) to span repetitive regions that shorter inserts cannot cover. This parameter change in insert size allows the assembly to capture complete repeat structures and structural variants, improving assembly completeness while the associated precision requirements are managed through optimized library construction protocols.
Data Source
AI summary
The present invention is generally directed to a hierarchical genome assembly process for producing high-quality de novo genome assemblies. The method utilizes a single, long-insert, shotgun DNA library in conjunction with Single Molecule, Real-Time (SMRT®) DNA sequencing, and obviates the need for additional sample preparation and sequencing data sets required for previously described hybrid assembly strategies. Efficient de novo assembly from genomic DNA to a finished genome sequence is demonstrated for several microorganisms using as little as three SMRT® cells, and for bacterial artificial chromosomes (BACs) using sequencing data from just one SMRT® Cell. Part of this new assembly workflow is a new consensus algorithm which takes advantage of SMRT® sequencing primary quality values, to produce a highly accurate de novo genome sequence, exceeding 99.999% (QV 50) accuracy. The methods are typically performed on a computer and comprise an algorithm that constructs sequence alignment graphs from pairwise alignment of sequence reads to a common reference.


