Variant-Specific Read Matching for Low-Frequency Variant Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for detecting variants in a data set often result in messy data sets with true and false positive variants, missed relevant instances, and false negatives due to aligning reads to a reference data set and characterizing mismatches, rather than systematically checking for catalogued variants.

Innovation Solution

A method that utilizes a variant-specific unique data set, where the reference data set does not include the variant, involves determining matching criteria for each read, and identifying the variant based on a quantity and quality of reads that satisfy these criteria, allowing for highly sensitive and specific detection of variants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods align reads to reference data set and characterize mismatches, then variant detection is performed, but messy data sets with true and false positive variants and false negatives are produced

Engineering Contradiction:
Improvevariant detection accuracyVSAvoiddata set quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

Instead of aligning reads to reference and characterizing mismatches as variants, the patent inverts the approach by directly querying reads against a catalog of known variants. This reversal eliminates the need to interpret messy mismatch data and directly identifies known variants with high precision, resolving the contradiction between measurement precision and data set quality.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent performs preliminary action by pre-cataloging known variants before the actual detection process. This pre-prepared variant catalog allows direct querying during detection, avoiding the need to re-discover variants from scratch during each alignment process, thereby improving both accuracy and data quality.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If conventional methods align reads to reference data set, then variant detection is performed, but relevant variant instances may be missed

Engineering Contradiction:
Improvevariant detection accuracyVSAvoidmissed variant instances
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent inverts the conventional approach by querying reads against a pre-built catalog of known variants rather than deriving variants from alignment mismatches. This ensures that all known variant instances are systematically checked for, preventing missed detections while maintaining high precision through direct matching.

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If variant-specific unique data set is used with multiple matching criteria, then sensitivity and specificity are improved, but computational complexity increases

Engineering Contradiction:
Improvevariant detection sensitivity and specificityVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the variant detection process into distinct components: a pre-built catalog of known variants, multiple independent matching criteria (first matching criterion for exact matches, second matching criterion for approximate matches), and a decision logic that combines results. This segmentation allows the complex multi-criteria system to be managed as modular, independent components, reducing overall system complexity while maintaining high sensitivity and specificity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12579117B2Variant identification by unique data set detection
Publication Date: 2026.03.17 COLOR HEALTH INC
  • US12579117B2 patent drawing
  • US12579117B2 patent drawing
  • US12579117B2 patent drawing

AI summary

Embodiments disclosed herein generally relate to detecting variants in a data set. A variant-specific unique data set, which includes a variant-inclusive portion that includes a particular variant and one or more other portions is accessed. The variant-specific unique data set corresponds to a particular region of a reference data set. A plurality of reads is received, with each read of the plurality of reads having been generated by processing a material collected from a subject. For each read of a subset of the plurality of reads, a first matching criterion and a second matching criterion can be determined to be satisfied, with the first matching criterion being more stringent than the second matching criterion. A data set associated with the subject can be determined to include the particular variant based on a quantity of reads in the subset. A result identified based on the particular variant can be output.