Partially Double-Stranded Identifier Molecules for NGS Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing (NGS) methods face challenges in accurately detecting mutations in circulating tumor DNA and identifying variants in complex mixed metagenomic samples due to amplification artifacts and intrinsic errors, which compromise the fidelity of variant calling.

Innovation Solution

The use of partially double-stranded identifier molecules with unique identifier sequences and adapter molecules, which are ligated to DNA fragments to generate sequencing libraries, allowing for error correction and increased barcode fidelity through customizable barcode combinations and hamming distances of at least two, enabling sensitive and specific identification of variant alleles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional NGS methods are used for sequencing, then sequence data can be obtained, but amplification artifacts and intrinsic errors compromise the fidelity of variant calling

Engineering Contradiction:
Improvevariant calling fidelityVSAvoiddetection accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-defining a set of valid barcode combinations with specified Hamming distances before sequencing. This pre-established error-correcting code structure allows post-sequencing validation and correction of amplification artifacts and intrinsic errors, thereby improving variant calling fidelity without requiring changes to the core sequencing process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback through the validation process where observed barcode combinations are compared against the pre-defined set of valid combinations. When discrepancies are detected (such as single-base errors from amplification artifacts), the system uses the Hamming distance properties to identify and correct these errors, feeding back improved accuracy to the variant calling process

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the number of barcodes is increased to improve detection sensitivity, then more variants can be identified, but cross-talk between barcodes increases

Engineering Contradiction:
Improvedetection sensitivityVSAvoidbarcode cross-talk
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent changes the structural parameters of barcodes by enforcing minimum Hamming distances between valid barcode combinations. This parameter constraint ensures that even as the number of barcodes increases, each barcode remains sufficiently distinct from others, preventing cross-talk while maintaining high detection sensitivity through increased barcode diversity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies beforehand cushioning by pre-designing the barcode set with built-in error tolerance through Hamming distance constraints. This cushioning effect creates a buffer zone that prevents cross-talk even when sequencing errors or amplification artifacts occur, allowing the system to accommodate more barcodes without increasing cross-talk

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Measurement precision

If sequencing depth is increased to improve variant detection, then more accurate variant calling is achieved, but amplification artifacts and intrinsic errors are also amplified

Engineering Contradiction:
Improvevariant detection accuracyVSAvoidamplification artifacts
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent uses feedback by comparing observed barcode combinations against the pre-defined valid set. This allows the system to identify and correct amplification artifacts and intrinsic errors that become more prevalent with increased sequencing depth, maintaining variant detection accuracy without being overwhelmed by errors

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent converts the harmful effect of amplification artifacts and intrinsic errors into a beneficial validation process. By using the known structure of valid barcode combinations and Hamming distance properties, the system transforms errors into opportunities for error detection and correction, thereby improving overall variant calling accuracy

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances the sensitivity and specificity of variant detection by reducing cross-talk and increasing the number of available barcodes, allowing for optimized sequencing depth and error correction, thereby improving the accuracy of mutation identification in complex DNA samples.

Implementation Method 1

a first plurality of partially double-stranded identifier molecules and a second plurality of partially double-stranded adapter molecules are simultaneously ligated to complex mixtures of individual target DNA fragments

Methodology Applied
Scientific EffectLigation: Chemical Bonding

Data Source

PatentUS20230407370A1Geometric synthesis methods and compositions for double-stranded nucleic acid sequencing
Publication Date: 2023.12.21 CAMENA BIOSCI LTD
  • US20230407370A1 patent drawing
  • US20230407370A1 patent drawing
  • US20230407370A1 patent drawing

AI summary

The present disclosure provides compositions, kits and methods for sequencing double-stranded nucleic acids. The compositions, kits and methods comprise partially double-stranded identifier molecules and partially double-stranded adapter molecules. The compositions, kits and methods can be used to determine the abundance and/or identity of specific transcripts in a plurality of double-stranded nucleic acids, as well as identifying the frequency of mutations within certain transcripts in the plurality of double-stranded nucleic acids.