Nucleic Acid Probe Design Using MSA Clustering to Reduce Capture Bias
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional probe designs for capturing nucleic acids are inefficient and biased, particularly when dealing with highly variable genomes like influenza, as they require numerous probes to cover genetic diversity and often miss less abundant variants due to target capture bias and sequencing limitations.
Innovation Solution
A method involving multiple sequence alignment (MSA) to design a minimal set of representative subsequences, which are used to synthesize probes that can efficiently capture a broad range of nucleic acid variants by shifting start positions and clustering aligned subsequences to reduce bias and increase coverage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventionally designed probes are used to capture nucleic acids, then the design is straightforward and simple, but target capture bias is introduced and genetic diversity coverage is limited
Solution Approach 1:
The probe set is segmented into multiple families, where each family targets a specific subset of genetic variants. This segmentation allows the system to cover broader genetic diversity while maintaining manageable design complexity for each individual probe family
Solution Approach 2:
Probes are designed with universal binding capabilities that allow a single probe to bind to multiple variant sequences. This multi-functionality reduces target capture bias by ensuring that probes can effectively bind to diverse genetic variants without requiring completely separate probe designs for each variant
2Reliability
If a collection of conventionally designed probes is used to reduce bias, then coverage of genetic diversity improves, but the number of probes required becomes enormous and the system becomes inefficient
Solution Approach 1:
Multiple probe families are merged into a single pooled probe set that can be used together in one experiment. This combining approach achieves comprehensive genetic diversity coverage without requiring separate experiments for each probe family, thereby reducing overall system complexity and the total number of probes needed
Solution Approach 2:
Probe sequences are designed to be copied across multiple families with systematic variations. This allows a minimal set of core sequences to be replicated and adapted across different probe families, reducing the total number of unique probe sequences needed while maintaining broad coverage capability
3Measurement precision
If Sanger sequencing is used to determine genetic sequences, then base calling accuracy is high for stable regions, but less abundant bases are missed and sequencing bias is introduced for highly variable regions
Solution Approach 1:
Probe hybridization is performed as a preliminary enrichment step before sequencing. This preliminary action concentrates target variants of interest, ensuring that even low-abundance variants are sufficiently enriched to be detected by subsequent sequencing methods without being lost in the background
4Reliability
If whole metagenome sequencing is used to sequence all nucleic acids, then unbiased genetic sequencing is achieved, but the cost is high and efficiency is reduced for targeted applications
Solution Approach 1:
Instead of sequencing all nucleic acids uniformly, the method applies targeted enrichment to specific genetic regions of interest using probe families designed for those regions. This local approach focuses sequencing resources on relevant targets, improving efficiency and reducing cost while maintaining unbiased coverage of the targeted regions through the multi-family probe design
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach minimizes target capture bias and maximizes genetic diversity coverage, enabling effective enrichment of highly variable genomes like influenza, even at low concentrations, with improved sensitivity and reduced probe number requirements.
Implementation Method 1
The disclosed approach utilizes available sequence data to capture genetic region(s) of interest. This conventional probe design depends on a single template that can be a reference sequence or majority-rule consensus sequence (which is typically computed from available sequence data for a target genetic region). This design approach has remained unchanged since the advent of Southern/Northern blotting, which relies on probe-target molecule binding to signal the presence of target nucleic acids.
Data Source
AI summary
The disclosure provides methods and systems for designing and synthesizing probes to capture a representative sample of genomic variants of a target genome from a sample. The methods include providing a multiple sequence alignment (MSA), designing a plurality of representative subsequences, and optionally synthesizing a nucleic acid probe. The designing step can comprise designating a plurality of intervals in the MSA, shifting start positions for each MSA subset, clustering the aligned subsequences within each adjusted subset, and determining a representative sequence for each reduced MSA subset. The disclosure also encompasses methods of isolating a plurality of nucleic acid variants of a targeted genomic subregion from a sample using the disclosed probe design, as well as the probe compositions themselves.


