Genome Assembly Using Score-Driven Layouts and Long-Range Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genome sequencing technologies face challenges in achieving accurate and efficient assembly of haplotypic sequences due to limitations in read length, contextual information, and error resilience, particularly with short-range non-contextual sequence reads, which hinder the generation of reliable whole-genome haplotypic sequences.
Innovation Solution
The method combines short-range non-contextual sequence reads with long-range mate-pairs and optical mapping techniques using a score-function-based approach to generate globally optimal layouts and consensus sequences, incorporating Bayesian methods and empirical Bayesian models to handle errors and improve assembly accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short-range non-contextual sequence reads are used for genome assembly, then sequencing speed and throughput are improved, but assembly accuracy and reliability deteriorate due to lack of long-range contextual information
Solution Approach 1:
The patent combines short-range sequence reads with long-range mate-pair information and optical mapping data into a unified assembly framework. This merging of multiple data sources with different range characteristics enables the system to achieve both high throughput from short reads and high accuracy from long-range constraints, directly resolving the contradiction between productivity and reliability.
Solution Approach 2:
The patent introduces long-range mate-pair sequences and optical mapping data as intermediary elements that bridge the gap between short-range reads. These intermediaries provide the missing contextual information that connects distant genomic regions, enabling accurate assembly without sacrificing the high throughput benefits of short-read sequencing.
2Ease of manufacture
If short-range sequence reads are used without contextual information, then sequencing cost is reduced, but the ability to resolve haplotypic ambiguity and structural variations deteriorates
Solution Approach 1:
The patent segments the sequencing strategy into two complementary components: inexpensive short-range reads for coverage and costly long-range mate-pair/optical mapping for context. This segmentation allows the system to obtain haplotypic information only where necessary, minimizing overall cost while preventing information loss in critical regions.
Solution Approach 2:
The patent applies different quality levels of sequencing data to different genomic regions based on their assembly difficulty. Regions with repetitive elements or structural variations receive enhanced long-range coverage, while unique regions rely on standard short-read coverage. This local quality differentiation maintains haplotypic information where needed without unnecessarily increasing overall sequencing cost.
3Loss of time
If conventional assembly algorithms are used without error correction, then processing time is reduced, but assembly reliability deteriorates due to error propagation
Solution Approach 1:
The patent performs error correction as a preliminary step before main assembly, using long-range mate-pair information and optical mapping data to identify and correct errors in short reads. By addressing errors beforehand, the system prevents error propagation during assembly without requiring extensive post-assembly validation, thus maintaining both speed and reliability.
Solution Approach 2:
The patent implements a feedback mechanism where long-range constraints from mate-pair and optical mapping data continuously guide the assembly process. Errors that violate these long-range constraints are identified and corrected, creating a self-correcting assembly system that maintains high reliability without excessive processing time.
Data Source
AI summary
Exemplary embodiments of the present disclosure relate generally to methods, computer-accessible medium and systems for assembling haplotype and/or genotype sequences of at least one genome, which can be based upon, e.g., consistent layouts of short sequence reads and long-range genome related data. For example, a processing arrangement can be configured to perform a procedure including, e.g., obtaining randomly located short sequence reads, using at least one score function in combination with constraints based on, e.g., the long range data, generating a layout of randomly located short sequence reads such that the layout is globally optimal with respect to the score function, obtained through searching coupled with score and constraint dependent pruning to determine the globally optimal layout substantially satisfying the constraints, generating a whole and/or a part of a genome wide haplotype sequence and/or genotype sequence, and converting a globally optimal layout into one or more consensus sequences.


