CFTR PolyTG-PolyT Variant Detection via Simulated Sequencing Fingerprints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current next-generation sequencing (NGS) methods face challenges in accurately detecting variants in the CFTR gene's polyTG/polyT region due to its complex structure and instability, leading to difficulties in distinguishing between haplotypes and obtaining reliable alignments.

Innovation Solution

A computational strategy is developed using simulated variant reads and sequencing data to generate high-quality fingerprint profiles for the CFTR polyTG/polyT tract, involving the generation of simulated variants with all possible polyT and polyTG combinations, alignment using an NGS analysis pipeline, and construction of a lookup table for accurate variant detection in clinical samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If NGS methods are used to sequence the CFTR polyTG/polyT region, then high-throughput sequencing and cost-effectiveness are improved, but accurate variant detection and alignment reliability deteriorate due to the complex and unstable structure of the polyTG/polyT region

Engineering Contradiction:
Improvesequencing throughputVSAvoidvariant detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by generating simulated sequencing reads with known variants before actual clinical sequencing. The simulation process creates ground truth data representing all possible polyTG/polyT combinations and their expected sequencing patterns, which are then used to train and validate the analysis pipeline before it is applied to real patient samples. This preliminary simulation allows the system to learn and optimize its variant detection algorithms in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating simulated sequencing reads that replicate the complex sequencing patterns expected in the polyTG/polyT region. These simulated reads serve as virtual copies of the difficult-to-sequence real data, allowing the analysis pipeline to be trained on numerous replicate examples of variant patterns without requiring actual sequencing of every possible combination, thereby improving detection accuracy through virtual replication.

Inventive Principle:
Principle #26Copying

2Measurement precision

If Sanger sequencing is used to confirm NGS variants in the CFTR polyTG/polyT region, then measurement precision is improved, but loss of time and increased computational burden occur

Engineering Contradiction:
Improvevariant confirmation accuracyVSAvoidconfirmation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical Sanger sequencing confirmation process with a computational approach. Instead of physically performing time-consuming Sanger sequencing to verify NGS variants, the system uses in silico simulation to generate expected sequencing patterns and compares actual NGS reads against these simulated ground truth data. This substitution eliminates the need for wet-lab confirmation while maintaining high accuracy through computational comparison of sequencing patterns.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates virtual copies of the sequencing data through simulation, generating simulated reads that represent all possible variants in the polyTG/polyT region. These simulated copies serve as the ground truth reference, allowing the system to confirm variants computationally by comparing actual sequencing data against the simulated expectations, thereby eliminating the need for time-consuming Sanger confirmation.

Inventive Principle:
Principle #26Copying

3Measurement precision

If targeted genomic sequencing is used to focus on specific regions, then sequencing depth and cost-effectiveness are improved, but device complexity and data analysis burden increase

Engineering Contradiction:
Improvesequencing depthVSAvoidanalysis pipeline complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex analysis task into distinct modular components: (1) simulation of sequencing reads for all possible variants, (2) generation of expected patterns from simulated data, (3) comparison of actual reads against simulated expectations, and (4) variant calling. This segmentation allows each component to be optimized independently and simplifies the overall pipeline by breaking down the complex analysis into manageable, sequential steps that can be executed systematically.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240412808A1Detection of cystic fibrosis transmembrane conductance regulator polytg/polyt variations by an NGS-based method
Publication Date: 2024.12.12 LABORATORY CORPORATION OF AMERICA HOLDINGS INC
  • US20240412808A1 patent drawing
  • US20240412808A1 patent drawing
  • US20240412808A1 patent drawing

AI summary

The present disclosure relates to a computational strategy with a digital read-out for detecting variants in a targeted genomic region. Particularly, aspects are directed to generating simulated variants for the targeted genomic region, detecting zero or more variations in a first allele, and zero or more variations in a second allele by inputting the simulated sequencing files into a sequencing analysis pipeline for alignment to a reference genome, identifying the zero or more variations between the first and second allele's simulated sequencing read file and the reference genome, and outputting a variant table comprising the zero or more detected variants and their corresponding variant features, concatenating the variant features for the zero or more detected variants in the first allele and in the second allele to generate a sequencing fingerprint profile, and outputting, the sets of simulated variants and their corresponding sequencing fingerprint profiles into a truth fingerprint table.