Control Nucleic Acid Sequences for Systematic Error Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing control nucleic acid sequences used in nucleic acid sequencing are oversensitive and fail to detect certain systematic errors, particularly in homopolymer regions, leading to reduced sequencing accuracy and compromised data quality.

Innovation Solution

Designing control nucleic acid sequences that identify loci with systematic errors using a variant caller, incorporating a representative set of loci involving errors in A, T, C, and G homopolymers of varying lengths, to improve error detection and sequencing accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If control nucleic acid sequences are designed to be highly sensitive to detect errors, then error detection capability is improved, but the sequences become oversensitive and fail to properly capture certain error modes

Engineering Contradiction:
Improveerror detection capabilityVSAvoiderror mode detection accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the parameters of control nucleic acid sequences by incorporating diverse homopolymer lengths (2-10 nucleotides) and varying GC content (30-70%), moving away from the conventional restricted design. This parameter diversification allows the sequences to detect multiple error modes without becoming oversensitive to any single error type, resolving the contradiction between sensitivity and reliability.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If control sequences focus on detecting errors in specific homopolymer lengths, then detection precision for those lengths is improved, but the ability to detect other error modes is reduced

Engineering Contradiction:
Improvedetection precision for specific homopolymer lengthsVSAvoiderror mode coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent designs control sequences that serve multiple functions by incorporating a comprehensive set of homopolymer lengths (2-10 nucleotides) and diverse GC content ranges (30-70%). Each control sequence can detect multiple error modes simultaneously, making the control system universal rather than specialized for a single error type, thus resolving the contradiction between precision and versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If conventional control sequences with restricted homopolymer lengths are used, then sequencing run time is reduced, but systematic errors in long homopolymers are not detected

Engineering Contradiction:
Improvesequencing run efficiencyVSAvoidsystematic error detection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent incorporates control sequences with long homopolymers (up to 10 nucleotides) and diverse GC content into the sequencing run beforehand. These pre-included control sequences systematically test for errors that conventional sequences would miss, allowing detection of systematic errors without significantly impacting overall sequencing productivity since they are integrated into the standard workflow.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12098424B2Control nucleic acid sequences for use in sequencing-by-synthesis and methods for designing the same
Publication Date: 2024.09.24 LIFE TECHNOLOGIES CORP
  • US12098424B2 patent drawing
  • US12098424B2 patent drawing
  • US12098424B2 patent drawing

AI summary

A method for designing test or control sequences may include identifying, using a variant caller, loci with systematic errors present in a plurality of sequencing runs included in a training set of sequencing runs obtained using sequencing-by-synthesis; and selecting a representative set of loci, including selecting from the identified loci an approximately equal number of loci involving errors in A, T, C, and G homopolymers and selecting from the identified loci an approximately equal number of loci involving homopolymers having a length of two, three, and four.