Barcoded Random Primers for RNA Isoform Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for high-throughput nucleic acid sequencing using short sequence reads are inadequate for reconstructing RNA isoform models due to computational ambiguities and limitations in identifying correct isoform models, especially when dealing with complex splicing patterns and rare isoforms, leading to unidentifiable models.

Innovation Solution

The use of modified surfaces with oligonucleotides containing randomized or partially randomized regions and barcodes to generate barcoded random primers, which prime nucleic acid synthesis and allow for the grouping of sequences by barcode, enabling the assembly of correct or nearly correct sequences of template molecules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If standard RNA-seq library preparation methods are used to generate short sequence reads, then almost every sequence in the transcriptome including splice junctions can be included, but the short reads are insufficient to reconstruct or infer the structures of RNA molecules due to computational ambiguity

Engineering Contradiction:
Improvecoverage of transcriptome sequencesVSAvoidaccuracy of RNA isoform reconstruction
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent divides the RNA molecule structure inference problem into segments by using paired-end reads that capture both ends of fragmented RNA molecules. The library preparation method creates separate sequence reads from different regions (exons, introns, splice junctions) that can be computationally assembled to reconstruct the complete RNA isoform structure, resolving the ambiguity of short reads through structured segmentation of the transcriptome.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension to sequence analysis by incorporating paired-end read information and spatial relationships between sequence fragments. Instead of analyzing single short reads in one dimension, the method uses the spatial and structural relationships between multiple reads from the same RNA molecule to reconstruct isoform models, transforming the problem from linear sequence assembly to multi-dimensional structural inference.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If short sequence reads of ~125 nt are generated to include all sequences in the transcriptome, then comprehensive transcriptome coverage is achieved, but alternative features greater than 125 nt apart give rise to computational ambiguity

Engineering Contradiction:
Improvetranscriptome sequence coverageVSAvoidcomputational ambiguity in isoform identification
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies preliminary action by performing computational assembly and isoform model construction before final isoform identification. The method pre-processes the short reads by assembling them into longer contigs and constructing candidate isoform models, which then serve as the basis for unambiguous isoform identification. This preliminary assembly step resolves the computational ambiguity that would otherwise prevent accurate isoform detection from short reads.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If standard library preparation methods are used, then RNA sequences from all DNA regions including non-coding RNAs are included, but the correct isoform model is unidentifiable when linear equations relating isoform frequencies to exon and splice junction frequencies are under-constrained

Engineering Contradiction:
Improveinclusion of all RNA typesVSAvoididentifiability of correct isoform model
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces computational assembly algorithms and isoform model construction as intermediary steps between raw short read data and final isoform identification. These intermediaries process the under-constrained frequency equations by assembling reads into structured models that provide additional constraints, enabling the correct isoform model to be identified even when direct frequency analysis would be ambiguous.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10253363B2Materials and methods to analyze RNA isoforms in transcriptomes
Publication Date: 2019.04.09 VACCINE RES INST OF SAN DIEGO
  • US10253363B2 patent drawing
  • US10253363B2 patent drawing
  • US10253363B2 patent drawing

AI summary

The present invention includes, but is not limited to, methods, assays, and compositions for preparing libraries of nucleic acid molecules from biological samples, and for detecting and measuring the abundances of nucleic acid molecules, including RNA and DNA. The methods, assays, and compositions of the present invention provide in whole or in part for the detecting and measuring the abundances of nucleic acid molecules, and particularly but not limited to those nucleic acids that have structures that cannot be determined reliably by routine alignment to a reference sequence or concatenation of smaller sequences by matching homologous ends.