Barcoded Random Primers for RNA Isoform Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for high-throughput nucleic acid sequencing using short sequence reads are inadequate for reconstructing RNA isoform models due to computational ambiguities and limitations in identifying correct isoform models, especially when dealing with complex splicing patterns and rare isoforms, leading to unidentifiable models.
Innovation Solution
The use of modified surfaces with oligonucleotides containing randomized or partially randomized regions and barcodes to generate barcoded random primers, which prime nucleic acid synthesis and allow for the grouping of sequences by barcode, enabling the assembly of correct or nearly correct sequences of template molecules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If standard RNA-seq library preparation methods are used to generate short sequence reads, then almost every sequence in the transcriptome including splice junctions can be included, but the short reads are insufficient to reconstruct or infer the structures of RNA molecules due to computational ambiguity
Solution Approach 1:
The patent divides the RNA molecule structure inference problem into segments by using paired-end reads that capture both ends of fragmented RNA molecules. The library preparation method creates separate sequence reads from different regions (exons, introns, splice junctions) that can be computationally assembled to reconstruct the complete RNA isoform structure, resolving the ambiguity of short reads through structured segmentation of the transcriptome.
Solution Approach 2:
The patent adds a new dimension to sequence analysis by incorporating paired-end read information and spatial relationships between sequence fragments. Instead of analyzing single short reads in one dimension, the method uses the spatial and structural relationships between multiple reads from the same RNA molecule to reconstruct isoform models, transforming the problem from linear sequence assembly to multi-dimensional structural inference.
2Quantity of substance
If short sequence reads of ~125 nt are generated to include all sequences in the transcriptome, then comprehensive transcriptome coverage is achieved, but alternative features greater than 125 nt apart give rise to computational ambiguity
Solution Approach 1:
The patent applies preliminary action by performing computational assembly and isoform model construction before final isoform identification. The method pre-processes the short reads by assembling them into longer contigs and constructing candidate isoform models, which then serve as the basis for unambiguous isoform identification. This preliminary assembly step resolves the computational ambiguity that would otherwise prevent accurate isoform detection from short reads.
3Adaptability or versatility
If standard library preparation methods are used, then RNA sequences from all DNA regions including non-coding RNAs are included, but the correct isoform model is unidentifiable when linear equations relating isoform frequencies to exon and splice junction frequencies are under-constrained
Solution Approach 1:
The patent introduces computational assembly algorithms and isoform model construction as intermediary steps between raw short read data and final isoform identification. These intermediaries process the under-constrained frequency equations by assembling reads into structured models that provide additional constraints, enabling the correct isoform model to be identified even when direct frequency analysis would be ambiguous.
Data Source
AI summary
The present invention includes, but is not limited to, methods, assays, and compositions for preparing libraries of nucleic acid molecules from biological samples, and for detecting and measuring the abundances of nucleic acid molecules, including RNA and DNA. The methods, assays, and compositions of the present invention provide in whole or in part for the detecting and measuring the abundances of nucleic acid molecules, and particularly but not limited to those nucleic acids that have structures that cannot be determined reliably by routine alignment to a reference sequence or concatenation of smaller sequences by matching homologous ends.


