Genome Assembly via Mutagenic Diversity for Repetitive Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genome sequencing technologies face challenges in assembling complex genomes due to repetitive and low-information regions, which are exacerbated by high error rates in long-read platforms and the requirement for large DNA input, leading to incomplete and inaccurate genome assemblies.
Innovation Solution
Introducing random mutations into nucleic acid samples to increase sequence diversity, allowing for the assembly of previously difficult-to-assemble regions using a combination of mutagenic reactions, size selection, and amplification techniques, such as multiple strand displacement amplification with nucleotide analogs, to generate sequencing libraries that can be accurately assembled.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Length of moving object
If long-read sequencing platforms are used to increase read length for assembling repetitive regions, then the ability to span repetitive regions is improved, but the error rate increases significantly
Solution Approach 1:
The patent segments the genome assembly problem into two distinct phases: first, long reads are used to span repetitive regions and establish the overall structure; second, short accurate reads are used to polish and correct errors in the assembled sequence. This segmentation allows each sequencing technology to be used for its strength while mitigating their respective weaknesses.
Solution Approach 2:
The patent merges the outputs of long-read and short-read sequencing assemblies by using the long-read assembly as a scaffold and incorporating accurate short-read data to correct errors. This combination produces a final assembly that benefits from both the long span capability and high accuracy of the two approaches.
2Loss of information
If long-read sequencing is used to assemble complex genomes, then coverage of repetitive regions is improved, but the requirement for large amounts of input DNA increases
Solution Approach 1:
The patent applies partial action by using a small subset of long reads (only enough to span repetitive regions and establish structure) rather than requiring full coverage with long reads. The remaining gaps and errors are filled by short reads, reducing the total long-read input DNA requirement while still achieving complete coverage of repetitive regions.
3Reliability
If short-read sequencing platforms are used for high accuracy, then the error rate is reduced, but the ability to assemble repetitive regions is lost
Solution Approach 1:
The patent performs preliminary assembly using long reads to span repetitive regions and establish the overall genomic structure before applying short reads. This preliminary action creates a framework that guides the subsequent accurate polishing with short reads, ensuring that repetitive regions are correctly positioned before accuracy optimization.
Solution Approach 2:
The long-read assembly serves as an intermediary scaffold that bridges the gap between short reads and the final accurate assembly. It provides the structural framework that short reads alone cannot generate, while the short reads provide the accuracy that the long-read assembly lacks. The intermediary scaffold enables the combination of both approaches.
4Device complexity
If traditional assembly methods are used for complex genomes, then the process is simpler, but the assembly completeness and accuracy deteriorate
Solution Approach 1:
The patent segments the assembly process into distinct phases with different objectives: structure establishment using long reads, error identification, and accuracy polishing using short reads. This segmentation transforms a single complex accurate assembly problem into manageable stages, improving accuracy without overwhelming complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables the assembly of larger, more accurate contigs with reduced error rates and increased sequence coverage, overcoming the limitations of existing technologies by generating tunable mutation rates and improving the assembly of complex genomic regions.
Implementation Method 1
the multiple strand displacement amplification reaction uses Phi29 DNA polymerase
Implementation Method 2
the introducing mutations step of the method includes performing a multiple displacement amplification reaction using a nucleotide analog
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods of sequencing and assembling a nucleic acid sequence from a nucleic acid sample containing repetitive or low-information regions, which are typically difficult to sequence and/or assemble are provided. The methods of sequencing and assembling introduce mutations into the sample to increase sequence diversity between various repetitive regions present in the nucleic acid sample. This sequence diversity allows various segments to assemble independently of different, but similar sequences present in the nucleic acid sample.