Control Nucleic Acid Sequences for Systematic Error Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing control nucleic acid sequences used in nucleic acid sequencing are oversensitive and fail to detect certain systematic errors, particularly in homopolymer regions, leading to reduced sequencing accuracy and compromised data quality.
Innovation Solution
Designing control nucleic acid sequences that identify loci with systematic errors using a variant caller, incorporating a representative set of loci involving errors in A, T, C, and G homopolymers of varying lengths, to improve error detection and sequencing accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If control nucleic acid sequences are designed to be highly sensitive to detect errors, then error detection capability is improved, but the sequences become oversensitive and fail to properly capture certain error modes
Solution Approach 1:
The patent changes the parameters of control nucleic acid sequences by incorporating diverse homopolymer lengths (2-10 nucleotides) and varying GC content (30-70%), moving away from the conventional restricted design. This parameter diversification allows the sequences to detect multiple error modes without becoming oversensitive to any single error type, resolving the contradiction between sensitivity and reliability.
2Measurement precision
If control sequences focus on detecting errors in specific homopolymer lengths, then detection precision for those lengths is improved, but the ability to detect other error modes is reduced
Solution Approach 1:
The patent designs control sequences that serve multiple functions by incorporating a comprehensive set of homopolymer lengths (2-10 nucleotides) and diverse GC content ranges (30-70%). Each control sequence can detect multiple error modes simultaneously, making the control system universal rather than specialized for a single error type, thus resolving the contradiction between precision and versatility.
3Productivity
If conventional control sequences with restricted homopolymer lengths are used, then sequencing run time is reduced, but systematic errors in long homopolymers are not detected
Solution Approach 1:
The patent incorporates control sequences with long homopolymers (up to 10 nucleotides) and diverse GC content into the sequencing run beforehand. These pre-included control sequences systematically test for errors that conventional sequences would miss, allowing detection of systematic errors without significantly impacting overall sequencing productivity since they are integrated into the standard workflow.
Data Source
AI summary
A method for designing test or control sequences may include identifying, using a variant caller, loci with systematic errors present in a plurality of sequencing runs included in a training set of sequencing runs obtained using sequencing-by-synthesis; and selecting a representative set of loci, including selecting from the identified loci an approximately equal number of loci involving errors in A, T, C, and G homopolymers and selecting from the identified loci an approximately equal number of loci involving homopolymers having a length of two, three, and four.


