DNA Sequence Partitioning Algorithm for Error Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA synthesis technologies are limited by high error rates and are unable to efficiently synthesize arbitrarily long DNA sequences, particularly those longer than 2,000 bp, due to constraints on DNA length and the cost of designing and assembling complex DNA libraries.
Innovation Solution
The development of a novel partitioning algorithm and data structures that allow for the division of long DNA sequences into smaller, manageable segments that can be easily synthesized and assembled, utilizing an inventory database to reuse existing DNA sequences and optimize synthesis efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Length of stationary object
If de novo DNA synthesis is used to synthesize long sequences, then the length of synthesized DNA can be increased, but the error rate increases and becomes unacceptable
Solution Approach 1:
The patent divides long DNA sequences into multiple smaller subsequences (e.g., 200-2000 bp each) that can be synthesized with high accuracy using existing de novo synthesis technology. These subsequences are then assembled together through assembly operations to reconstruct the full long sequence, thereby avoiding the high error rates associated with attempting to synthesize the entire long sequence in a single de novo synthesis operation.
2Length of stationary object
If de novo synthesis is used for sequences longer than 2,000 bp, then longer sequences can be obtained, but synthesis speed decreases and costs increase
Solution Approach 1:
The patent segments long DNA sequences into smaller subsequences that fall within the optimal synthesis length range (200-2000 bp) where synthesis speed is high. By synthesizing these shorter subsequences rapidly and then assembling them, the overall process achieves both long final sequence length and maintains high synthesis speed, avoiding the slow synthesis rates associated with attempting to synthesize very long sequences in a single operation.
Solution Approach 2:
The patent performs preliminary identification and planning of subsequence boundaries and assembly strategies before the actual synthesis process. This includes analyzing the target long sequence to determine optimal partitioning points and selecting appropriate assembly methods in advance, which streamlines the subsequent synthesis and assembly operations and improves overall productivity.
3Length of stationary object
If complex DNA libraries are constructed from smaller DNA oligos, then long sequences can be assembled, but the cost becomes prohibitive
Solution Approach 1:
The patent applies different strategies to different regions of the target DNA sequence based on local characteristics. By analyzing specific features of different subsequences (such as GC content, secondary structure propensity, or homology to existing sequences), the system can select optimal synthesis and assembly approaches for each local region, improving overall cost-efficiency while achieving the goal of constructing long sequences or libraries.
Data Source
AI summary
A method, apparatus, and computer-readable medium for optimized partitioning for assembly of nucleic acid sequences, receiving nucleic acid sequences corresponding to target nucleic acids for assembly and synthesis parameters, querying an inventory database based on the nucleic acid sequences to determine first matching nucleic acid subsequences, the inventory database corresponding to nucleic acid subsequences available in an inventory, identifying second matching nucleic acid subsequences based on one or more overlaps between the nucleic acid sequences, generating an acyclic directed graph data structure corresponding to potential partitions of the nucleic acid sequences based on the first matching nucleic acid subsequences, the second matching nucleic acid subsequences, and one or more synthesis parameters, and determining an optimal partitioning of the nucleic acid sequences based on an optimal path through the acyclic directed graph data structure that minimizes a total weight of traversed edges and nodes within the acyclic directed graph data structure.


