Chimeric Artefact Detection Using Dual-End Identifier Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting chimeric artefact polynucleotides during amplification of mixed samples, particularly in long-read sequencing, are inadequate due to high error rates and the formation of chimeric molecules, which complicates the identification of true Unique Molecular Identifiers (UMIs).

Innovation Solution

A method involving the use of capture polynucleotides with identifier sequences and template switch oligonucleotides (TSOs) to generate libraries flanked by unique identifier sequences, allowing for the detection of chimeric artefacts by identifying underrepresented or mismatched pairings of these sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If long-read sequencing technologies are used to sequence native RNA, then sequencing resolution and read length are improved, but error rates increase making it difficult to identify true UMI sequences

Engineering Contradiction:
Improveread lengthVSAvoiderror rate
Core Design Contradiction:
Length of moving objectVSReliability

Solution Approach 1:

The UMI sequence is divided into multiple discrete nucleotide blocks (e.g., 3-5 blocks of 2-4 nucleotides each). Each block is independently sequenced and can be individually error-corrected, allowing the true UMI to be reconstructed even when individual blocks contain sequencing errors. This segmentation enables reliable UMI identification despite the high error rates of long-read sequencing technologies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method performs preliminary error correction by comparing each discrete nucleotide block against a pool of known UMI sequences before final UMI assignment. This preliminary action of blocking and comparing allows the system to anticipate and correct sequencing errors, distinguishing true UMIs from artefactual ones before downstream analysis.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If PCR-based library methods are used to overcome sample size limitations, then sensitivity is improved, but PCR amplification bias and chimeric molecule formation increase

Engineering Contradiction:
Improvesample sizeVSAvoidchimeric molecule formation
Core Design Contradiction:
Quantity of substanceVSObject-generated harmful factors

Solution Approach 1:

By segmenting the UMI into discrete nucleotide blocks that are flanking the cDNA molecule, the method creates unique molecular fingerprints that can track individual original RNA molecules through PCR amplification. This enables computational identification and removal of chimeric molecules that result from PCR artefacts, as they will have inconsistent UMI block pairings compared to true amplification products.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method uses the paired discrete UMI blocks as feedback signals to verify the authenticity of amplified molecules. During data analysis, the system checks whether the UMI block pairings are consistent with the original capture polynucleotide and TSO pairings, providing feedback that allows differentiation between true PCR amplification products and chimeric artefacts.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If discrete nucleotide blocks with at least two nucleotide substitutions are used in identifier sequences, then sequencing error detection is improved, but library complexity and synthesis requirements increase

Engineering Contradiction:
Improveerror detection accuracyVSAvoidlibrary complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method applies local quality control by ensuring that each discrete nucleotide block within the UMI has sufficient sequence diversity (at least two nucleotide substitutions from other blocks). This local differentiation within blocks provides error-detecting capability without requiring the entire library to be overly complex, as each block independently contributes to error correction.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent specifies particular parameters for the nucleotide blocks (2-4 nucleotides per block, 3-5 blocks per UMI, minimum two nucleotide substitutions between blocks) to optimize the balance between error detection capability and library complexity. These parameter changes create a UMI structure that is sufficiently complex to detect errors but manageable for synthesis and sequencing.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively identifies and removes chimeric artefacts, improving the accuracy of sequencing by reducing PCR amplification bias and enhancing the reliability of UMIs, especially in long-read sequencing technologies.

Implementation Method 1

capturing RNA molecules of the sample on a set of capture polynucleotides

Methodology Applied
Scientific EffectHybridization:

Implementation Method 2

performing reverse transcription of captured sample RNA molecules using a template switch reverse polymerase

Methodology Applied
Scientific EffectReverse transcription:

Implementation Method 3

annealing a set of template switch oligonucleotides (TSOs) to the 3' end non-templated nucleotides

Methodology Applied
Scientific EffectAnnealing: Annealing

Implementation Method 4

amplifying the cDNA using the capture polynucleotide PCR handle sequences and TSO PCR handle sequences

Methodology Applied
Scientific EffectPCR amplification:

Data Source

PatentUS20250223585A1Chimeric artefact detection method
Publication Date: 2025.07.10 OXFORD UNIVERSITY INNOVATION LTD
  • US20250223585A1 patent drawing
  • US20250223585A1 patent drawing
  • US20250223585A1 patent drawing

AI summary

The invention relates to methods for detecting chimeric artefact polynucleotides produced during amplification of a mixed sample of polynucleotide. The methods comprise adding identifier sequences to both ends of a sample polynucleotide. Also provided are arrays of annealed oligonucleotide strand pairs for providing a mixed pool of identifier sequences; kits and methods for producing a library of polynucleotides, or libraries of polynucleotides having identifier sequences at both ends of the polynucleotides; and arrays of template switch oligonucleotides.