Nucleic Acid Sample Pooling Under Maximum Overlap Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for sample multiplexing in next-generation sequencing (NGS) are computationally inefficient and costly, particularly when mixing DNA samples without barcodes, and automating large-scale mixing is non-trivial due to numerous pooling constraints.
Innovation Solution
A novel algorithm using a linear optimization solver to determine optimal pooling of nucleic acid samples by analyzing DNA similarities and applying constraints, such as maximum overlap thresholds, to ensure accurate mapping back to original sample libraries, without the need for barcodes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If barcode sequencing is used for sample multiplexing, then multiple samples can be identified and sequenced simultaneously, but additional DNA ligation steps and computational demultiplexing are required, increasing time and economic costs
Solution Approach 1:
The patent extracts and removes the barcode sequencing steps (ligation and computational demultiplexing) from the multiplexing process. Instead of using barcodes, the system directly sequences pooled DNA samples and uses bioinformatics to identify and separate sequences by their origin, eliminating the need for additional ligation steps and reducing process complexity while maintaining multiplexing capability
Solution Approach 2:
The patent introduces bioinformatics analysis as an intermediary between sample pooling and sequence identification. Rather than using physical barcodes to mediate sample identification, the system uses computational methods to analyze and attribute sequences to their source samples after pooling, replacing complex wet-lab steps with computational processing
2Device complexity
If DNA sample libraries are mixed together without barcodes, then time and economic costs are reduced, but the DNA of each sample library must be sufficiently different, making automation non-trivial for large numbers of samples
Solution Approach 1:
The patent performs preliminary bioinformatics analysis to calculate sequence similarity matrices and determine optimal pooling strategies before actual sequencing. By pre-computing which samples can be pooled together based on their sequence divergence, the system enables automated decision-making for large-scale multiplexing without requiring manual assessment of DNA differences
Solution Approach 2:
The patent changes the parameter of sample identification from relying on physical differences (barcode sequences) to relying on computational parameters (sequence similarity thresholds). By establishing quantitative criteria for sample distinguishability and using algorithmic optimization, the system makes automation feasible for large numbers of samples while maintaining process simplicity
3Productivity
If more samples are pooled together to reduce sequencing runs, then productivity increases, but the computational complexity of ensuring accurate mapping back to original sample libraries increases
Solution Approach 1:
The patent performs preliminary computation of sequence similarity matrices and optimal pooling strategy determination before sequencing. By pre-calculating which samples can be successfully pooled together based on their sequence divergence and establishing maximum overlap constraints, the system ensures that even large pools can be accurately mapped back to their original sources without losing sample identification accuracy
Solution Approach 2:
The patent implements a feedback mechanism where bioinformatics analysis of sequence data provides information about pooling effectiveness and sample identification accuracy. This feedback is used to refine and optimize pooling strategies, ensuring that as more samples are pooled together, the system can still maintain accurate mapping by adjusting pooling parameters based on observed sequence divergence patterns
Data Source
AI summary
A method, apparatus, and computer-readable medium for optimal pooling of nucleic acid samples for next generation sequencing, including receiving sample records corresponding to samples, each sample record comprising a sample identifier and a nucleic acid reference sequence of the sample, determining unique nucleic acid reference sequences in the sample records, computing, a nucleic acid overlaps between the unique nucleic acid reference sequences, and determining an optimal grouping of the plurality of samples into a plurality of sample pools based at least in part on the nucleic acid reference sequence of each sample record, the nucleic acid overlaps between the unique nucleic acid reference sequences, and one or more constraints, the one or more constraints including a maximum overlap constraint.


