CFTR PolyTG-PolyT Variant Detection via Simulated Sequencing Fingerprints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current next-generation sequencing (NGS) methods face challenges in accurately detecting variants in the CFTR gene's polyTG/polyT region due to its complex structure and instability, leading to difficulties in distinguishing between haplotypes and obtaining reliable alignments.
Innovation Solution
A computational strategy is developed using simulated variant reads and sequencing data to generate high-quality fingerprint profiles for the CFTR polyTG/polyT tract, involving the generation of simulated variants with all possible polyT and polyTG combinations, alignment using an NGS analysis pipeline, and construction of a lookup table for accurate variant detection in clinical samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If NGS methods are used to sequence the CFTR polyTG/polyT region, then high-throughput sequencing and cost-effectiveness are improved, but accurate variant detection and alignment reliability deteriorate due to the complex and unstable structure of the polyTG/polyT region
Solution Approach 1:
The patent applies preliminary action by generating simulated sequencing reads with known variants before actual clinical sequencing. The simulation process creates ground truth data representing all possible polyTG/polyT combinations and their expected sequencing patterns, which are then used to train and validate the analysis pipeline before it is applied to real patient samples. This preliminary simulation allows the system to learn and optimize its variant detection algorithms in advance.
Solution Approach 2:
The patent uses copying by creating simulated sequencing reads that replicate the complex sequencing patterns expected in the polyTG/polyT region. These simulated reads serve as virtual copies of the difficult-to-sequence real data, allowing the analysis pipeline to be trained on numerous replicate examples of variant patterns without requiring actual sequencing of every possible combination, thereby improving detection accuracy through virtual replication.
2Measurement precision
If Sanger sequencing is used to confirm NGS variants in the CFTR polyTG/polyT region, then measurement precision is improved, but loss of time and increased computational burden occur
Solution Approach 1:
The patent replaces the mechanical Sanger sequencing confirmation process with a computational approach. Instead of physically performing time-consuming Sanger sequencing to verify NGS variants, the system uses in silico simulation to generate expected sequencing patterns and compares actual NGS reads against these simulated ground truth data. This substitution eliminates the need for wet-lab confirmation while maintaining high accuracy through computational comparison of sequencing patterns.
Solution Approach 2:
The patent creates virtual copies of the sequencing data through simulation, generating simulated reads that represent all possible variants in the polyTG/polyT region. These simulated copies serve as the ground truth reference, allowing the system to confirm variants computationally by comparing actual sequencing data against the simulated expectations, thereby eliminating the need for time-consuming Sanger confirmation.
3Measurement precision
If targeted genomic sequencing is used to focus on specific regions, then sequencing depth and cost-effectiveness are improved, but device complexity and data analysis burden increase
Solution Approach 1:
The patent applies segmentation by dividing the complex analysis task into distinct modular components: (1) simulation of sequencing reads for all possible variants, (2) generation of expected patterns from simulated data, (3) comparison of actual reads against simulated expectations, and (4) variant calling. This segmentation allows each component to be optimized independently and simplifies the overall pipeline by breaking down the complex analysis into manageable, sequential steps that can be executed systematically.
Data Source
AI summary
The present disclosure relates to a computational strategy with a digital read-out for detecting variants in a targeted genomic region. Particularly, aspects are directed to generating simulated variants for the targeted genomic region, detecting zero or more variations in a first allele, and zero or more variations in a second allele by inputting the simulated sequencing files into a sequencing analysis pipeline for alignment to a reference genome, identifying the zero or more variations between the first and second allele's simulated sequencing read file and the reference genome, and outputting a variant table comprising the zero or more detected variants and their corresponding variant features, concatenating the variant features for the zero or more detected variants in the first allele and in the second allele to generate a sequencing fingerprint profile, and outputting, the sets of simulated variants and their corresponding sequencing fingerprint profiles into a truth fingerprint table.


