Artificial Transposon Sequences for Reference-Free DNA Assembly
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid sequencing methods face challenges in assembling contiguous reads from repetitive sequences and require large DNA amounts, inefficient library preparation, and a reference genome for accurate representation, especially in sequencing long molecules like genomic DNA.
Innovation Solution
The use of artificial transposon sequences with unique barcodes and linkers for tagging and fragmenting DNA, allowing for the assembly of sequencing data without a reference genome by identifying paired barcode sequences to reconstruct the original proximity of DNA fragments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional nucleic acid sequencing methods are used, then sequencing can be performed, but assembly of contiguous reads from repetitive sequences is difficult and requires large DNA amounts
Solution Approach 1:
The transposon sequence is divided into multiple segments including first and second barcodes separated by a linker. This segmentation allows the barcode to be distributed across multiple sequencing reads, enabling accurate assembly of contiguous reads even from repetitive sequences by matching barcode segments across reads.
Solution Approach 2:
The barcode acts as an intermediary element that links different sequencing reads together. By inserting the barcoded transposon into the target nucleic acid and distributing barcode segments across multiple reads, the intermediary barcode enables accurate reconstruction of the original nucleic acid sequence without requiring large amounts of DNA.
2Adaptability or versatility
If reference genome-based assembly is used, then accurate representation can be achieved, but de novo assembly without reference genome is not possible
Solution Approach 1:
The barcode embedded in the transposon sequence enables the sequencing system to perform self-service assembly without external reference. The first and second barcode segments distributed across multiple reads automatically provide the information needed for de novo assembly, making the system versatile for both reference-based and reference-free approaches while maintaining accuracy.
3Productivity
If short reads are sequenced, then high throughput is achieved, but order information and contiguous assembly are lost
Solution Approach 1:
The barcode is preliminarily inserted into the target nucleic acid via transposon integration before sequencing. This preliminary action embeds order information directly into the short reads through the distributed barcode segments, allowing high-throughput sequencing to maintain contiguous assembly capability by matching barcode segments across reads during data processing.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Artificial transposon sequences having code tags and target nucleic acids containing such sequences. Methods for making artificial transposons and for using their properties to analyze target nucleic acids.