Sequencing Read Selection Using Phred Scores for Accurate Variant Calling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing (NGS) systems suffer from high error rates and inaccuracies in base sequence identification, leading to significant noise and incorrect base pairs in sequencing data, which can result in substantial computational burdens and reduced processing efficiency.

Innovation Solution

A method and system for measuring error rates and read quality on a read-by-read basis, filtering low-quality reads, and calculating sequencing metrics such as coverage depth, variant frequencies, and strand bias, using Phred scaled quality scores to select high-quality reads and determine clinically relevant variants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If NGS systems sequence entire genomes in parallel to reduce processing requirements, then sequencing speed and throughput are improved, but error rates increase to 10-25% due to base calling inaccuracies and sample preparation errors

Engineering Contradiction:
Improvesequencing speedVSAvoidbase sequence accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by filtering reads based on quality scores before performing variant calling and analysis. Low-quality reads are identified and excluded in advance using Phred quality score thresholds, preventing erroneous data from propagating through subsequent computational steps. This preliminary filtering resolves the contradiction by preparing the data in advance to maintain both high throughput and accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If all sequencing reads are processed to calculate sequencing metrics, then comprehensive analysis is achieved, but computational time and memory consumption increase significantly due to high error rates

Engineering Contradiction:
Improvesequencing metrics accuracyVSAvoidcomputational processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and removes low-quality reads from the dataset before performing variant calling and sequencing metric calculations. By applying quality score thresholds and filtering out reads that do not meet the criteria, the system eliminates erroneous data that would otherwise consume computational resources and potentially skew results. This extraction approach maintains measurement precision while reducing the computational burden.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If high quality score thresholds are applied to filter reads, then sequencing data accuracy is improved, but the quantity of usable sequencing data decreases

Engineering Contradiction:
Improvebase sequence accuracyVSAvoidusable sequencing data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent employs parameter changes by allowing flexible adjustment of quality score thresholds based on specific application requirements. Different threshold values can be set to balance accuracy and data volume trade-offs. The system can adaptively modify filtering parameters to optimize the balance between maintaining high accuracy and preserving sufficient data quantity for meaningful statistical analysis and variant calling.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260031185A1Genomic sequencing selection system
Publication Date: 2026.01.29 QUEST DIAGNOSTICS INVESTMENTS INC
  • US20260031185A1 patent drawing
  • US20260031185A1 patent drawing
  • US20260031185A1 patent drawing

AI summary

The systems and methods discussed herein can calculate sequencing statistics such as coverage depth for sequencing data. The present solution can determine variant frequencies and identify clinically relevant variants. The present solution can read BAM and VCF input files and Phred scaled quality scores. The present solution can select relatively high quality reads based on the quality scores and can calculate reference and alternative allele counts for SNPs, insertions and deletions (INDELs), and structural variants.