Sequencing Read Selection Using Phred Scores for Accurate Variant Calling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing (NGS) systems suffer from high error rates and inaccuracies in base sequence identification, leading to significant noise and incorrect base pairs in sequencing data, which can result in substantial computational burdens and reduced processing efficiency.
Innovation Solution
A method and system for measuring error rates and read quality on a read-by-read basis, filtering low-quality reads, and calculating sequencing metrics such as coverage depth, variant frequencies, and strand bias, using Phred scaled quality scores to select high-quality reads and determine clinically relevant variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If NGS systems sequence entire genomes in parallel to reduce processing requirements, then sequencing speed and throughput are improved, but error rates increase to 10-25% due to base calling inaccuracies and sample preparation errors
Solution Approach 1:
The patent applies preliminary action by filtering reads based on quality scores before performing variant calling and analysis. Low-quality reads are identified and excluded in advance using Phred quality score thresholds, preventing erroneous data from propagating through subsequent computational steps. This preliminary filtering resolves the contradiction by preparing the data in advance to maintain both high throughput and accuracy.
2Measurement precision
If all sequencing reads are processed to calculate sequencing metrics, then comprehensive analysis is achieved, but computational time and memory consumption increase significantly due to high error rates
Solution Approach 1:
The patent extracts and removes low-quality reads from the dataset before performing variant calling and sequencing metric calculations. By applying quality score thresholds and filtering out reads that do not meet the criteria, the system eliminates erroneous data that would otherwise consume computational resources and potentially skew results. This extraction approach maintains measurement precision while reducing the computational burden.
3Reliability
If high quality score thresholds are applied to filter reads, then sequencing data accuracy is improved, but the quantity of usable sequencing data decreases
Solution Approach 1:
The patent employs parameter changes by allowing flexible adjustment of quality score thresholds based on specific application requirements. Different threshold values can be set to balance accuracy and data volume trade-offs. The system can adaptively modify filtering parameters to optimize the balance between maintaining high accuracy and preserving sufficient data quantity for meaningful statistical analysis and variant calling.
Data Source
AI summary
The systems and methods discussed herein can calculate sequencing statistics such as coverage depth for sequencing data. The present solution can determine variant frequencies and identify clinically relevant variants. The present solution can read BAM and VCF input files and Phred scaled quality scores. The present solution can select relatively high quality reads based on the quality scores and can calculate reference and alternative allele counts for SNPs, insertions and deletions (INDELs), and structural variants.


