Semi-Random Barcode Primers for Accurate NGS Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next generation sequencing (NGS) technologies face high error rates and sequencing-dependent bias, complicating the identification of rare mutations and accurate transcript counting in DNA and RNA sequencing due to amplification noise and sequencing errors.

Innovation Solution

The use of semi-random barcodes to tag sequencing fragments before amplification, allowing for the identification and correction of errors in library preparation, amplification, and sequencing processes, thereby reducing bias and improving accuracy in mutation detection and transcriptome profiling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If random sequence barcodes are used to tag template molecules, then amplification bias can be corrected, but sequencing errors frequently cause misidentification of barcodes

Engineering Contradiction:
Improveamplification bias correctionVSAvoidbarcode identification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of the barcode sequence from completely random to semi-random with specific constraints. The barcode consists of repeating units of 3-6 nucleotides where each unit has limited variability (e.g., 2-4 possible sequences per position). This parameter change maintains the diversity needed for bias correction while reducing the total number of possible sequences, thereby minimizing sequencing errors that lead to misidentification.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The barcode is constructed as a composite structure with repeating modular units. Each unit consists of a core sequence with optional variations at specific positions. This composite design allows the barcode to function as both a unique identifier and an error-correcting code, where the repeating pattern provides redundancy that helps distinguish true barcodes from sequencing errors.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If sequence predefined barcodes are used to reduce misidentification, then barcode accuracy improves, but the cost for generating large varieties of barcodes becomes very high

Engineering Contradiction:
Improvebarcode identification accuracyVSAvoidbarcode generation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent changes the approach from using completely predefined unique sequences to semi-random sequences generated from limited alphabets. Instead of synthesizing millions of completely unique random sequences, the system uses repeating units with constrained variability, dramatically reducing the number of unique oligonucleotides that need to be synthesized while still providing sufficient diversity for accurate template tagging.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The barcode structure uses repeating units that are copied multiple times within each barcode sequence. For example, a 18-mer barcode might consist of 6 repetitions of a 3-mer unit or 3 repetitions of a 6-mer unit. This copying approach reduces the synthesis burden since the same or similar sequences are reused, while still providing unique combinations through different arrangements and variations of the repeating units.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If NGS technology is used for DNA sequencing, then mutation detection capability is enabled, but high error rates prevent confident identification of rare mutations

Engineering Contradiction:
Improvemutation detection capabilityVSAvoidrare mutation identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where barcodes are used to group sequencing reads by template molecule. Within each barcode group, the system can identify and correct errors by comparing multiple reads that originated from the same template. This feedback loop allows rare mutations to be distinguished from sequencing errors, as true mutations will appear consistently across multiple reads from the same template while errors will be random and inconsistent.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The barcode acts as an intermediary that links multiple sequencing reads to their original template molecule. By tagging each template with a unique barcode before amplification and sequencing, the system can trace all reads back to their source template. This intermediary structure enables error correction by allowing the system to identify which reads belong together and should be compared against each other for accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Quantity of substance

If library amplification is performed to increase sequencing depth, then detection sensitivity improves, but amplification noise complicates accurate transcript counting

Engineering Contradiction:
Improvesequencing depthVSAvoidtranscript counting accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The barcode provides a feedback mechanism that allows the system to track the history of each template molecule through amplification. By grouping reads by barcode, the system can identify which reads come from the same original template and apply correction algorithms that account for amplification bias. This feedback enables accurate transcript counting even after multiple amplification cycles, as the system can distinguish between biological variation and amplification artifacts.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The barcode serves as an intermediary tag that is attached to each template molecule before amplification and remains associated with all its progeny. This intermediary structure allows the system to trace amplification relationships and apply normalization factors based on the original template distribution, thereby correcting for amplification noise and enabling accurate transcript quantification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260022372A1Semi-random barcodes for nucleic acid analysis
Publication Date: 2026.01.22 QIAGEN SCIENCES LLC
  • US20260022372A1 patent drawing
  • US20260022372A1 patent drawing
  • US20260022372A1 patent drawing

AI summary

The present disclosure provides oligonucleotides that comprise semi-random barcode sequences. Such oligonucleotides may be incorporated into reverse transcription primers, PCR primers, or portions of sequencing adapters in preparing sequencing libraries. The resulting sequencing libraries can be used for accurate sequencing, including DNA or RNA counting and mutation detection. Methods and kits for preparing sequencing adapters and sequencing libraries are also provided.