DNA Barcode Mapping for Short-Read Protein Variant Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for mapping protein-coding sequences in mutagenic libraries face challenges with high costs and low accuracy when using long-read sequencing, particularly in distinguishing closely related mutational variants, and are incompatible with short-read sequencing platforms.
Innovation Solution
Incorporating a first barcode (SynBC) within the protein-coding region and a second barcode (randomized) outside the coding region, allowing for the use of short-read sequencing to identify protein-coding regions by sequencing only these barcodes, thereby enhancing accuracy and resolving mutational variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If long-read sequencing is used to sequence the entire DNA region including barcodes and protein-coding regions, then the relationship between barcodes and mutated regions can be established, but the cost increases and accuracy decreases for distinguishing closely related variants
Solution Approach 1:
The DNA molecule is segmented into distinct functional regions: a barcode region containing unique identifying sequences (first and second barcodes) and a protein-coding region. This segmentation allows selective sequencing of only the barcode region using short-read sequencing platforms, avoiding the need to sequence the entire DNA molecule. The barcode region can be amplified and sequenced independently, reducing sequencing costs while maintaining the ability to identify and distinguish between different mutational variants through unique barcode sequences.
2Measurement precision
If long-read sequencing is used to sequence from barcode through protein-coding region, then mapping can be accomplished, but the error rate increases and cost increases
Solution Approach 1:
The barcode region is extracted as a separate, sequenceable element that can be independently amplified and sequenced using short-read sequencing. By taking out the barcode sequencing task from the full DNA sequencing process, the method eliminates the need to sequence through the entire protein-coding region, thereby reducing sequencing errors associated with long reads while maintaining accurate mapping through the unique barcode sequences.
3Stability of the object's composition
If synthetic DNA barcodes are placed outside the protein-coding region, then the open reading frame is maintained, but the relationship between barcode and mutated region must be established through experimental mapping
Solution Approach 1:
The method merges the barcode region with the protein-coding region in a fixed, known spatial relationship within the same DNA molecule. The first and second barcodes are positioned such that they are always associated with the protein-coding region containing the mutational variants. This merging eliminates the need for experimental mapping to establish relationships, as the spatial association is predetermined by the molecular construction, thereby reducing complexity while maintaining ORF integrity.
4Measurement precision
If fully-custom oligonucleotide synthesis is used to create single molecules with barcodes and protein-coding regions, then the relationship is established in synthesis, but the cost becomes prohibitively expensive for large-scale experiments
Solution Approach 1:
Instead of synthesizing entire DNA molecules with barcodes and protein-coding regions through fully-custom oligonucleotide synthesis, the method uses partial action by synthesizing and amplifying only the barcode region. The barcode region can be generated through PCR amplification from templates, avoiding the need for expensive de novo synthesis of complete DNA molecules. This partial approach maintains the ability to establish relationships between barcodes and protein-coding regions while dramatically reducing synthesis costs for large-scale experiments.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Provided are methods for resolving the relationship between unique user-designed and/or random synthetic DNA barcodes and a protein-coding mutated region of interest with enhanced accuracy that is amenable to short-read sequencing platforms. In addition, the methods introduce increased sequence divergence between mutational variants of a region of interest in order to enhance the resolvability of mutational variants within a mutagenic library when error-prone long-read sequencing platforms are used.