Modular RNA modification model training framework

The modular RNA construct addresses the limitations of conventional RNA synthesis by assembling longer sequences from shorter oligonucleotides with precise modifications, enhancing detection accuracy and reducing bias for nanopore sequencing.

WO2026090312A1PCT designated stage Publication Date: 2026-04-30NORTHEASTERN UNIV (US)
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NORTHEASTERN UNIV (US)
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional RNA synthesis methods are limited by short sequence lengths, low yields, and lack of site-specific modification control, particularly for long-read sequencing technologies like nanopore direct RNA sequencing, which require minimum read lengths and context control for accurate modification detection.

Method used

A modular RNA construct is assembled from shorter synthetic oligonucleotides using annealing and ligation, incorporating precise RNA base modifications within a controlled sequence context, suitable for long-read sequencing and benchmarking.

Benefits of technology

The modular RNA construct enables longer, customizable RNA sequences with site-specific modifications, overcoming length constraints and providing robust training data for machine learning models, reducing sequence bias, and enhancing detection accuracy in nanopore sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052101_30042026_PF_FP_ABST
    Figure US2025052101_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to nucleic acid engineering and, more specifically, to synthetic RNA constructs and assembly methods that enable precise, site-specific inclusion of one or more RNA base modifications in a controlled and extendable sequence context suitable for long-read sequencing and quantitative benchmarking. Disclosed are methods of making a modular RNA construct. The methods may comprise annealing a modification loop bait RNA oligonucleotide with a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide; and ligating the annealed oligonucleotides to form the modular RNA construct.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] MODULAR RNA MODIFICATION MODEL TRAINING FRAMEWORK RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application serial number 63 / 710,332, filed October 22, 2024.

[0003] BACKGROUND

[0004] More than 170 naturally occurring RNA base-level modifications have been documented and are increasingly implicated in cellular function and disease. Yet, robust study of modification signatures, benchmarking of detection strategies, and training of computational models remain constrained by technical limitations in RNA synthesis, sequence length, and context control.

[0005] Conventional synthetic RNA standards are typically short and become impractical above approximately 120 nucleotides due to low yields and decreased synthesis success, particularly when incorporating modified nucleotides. In vitro transcription can generate long, unmodified RNA or enzyme-introduced modifications, but generally lacks site-level control and often produces “all-or-nothing” modification states that do not reflect biological heterogeneity. These constraints are especially limiting for single-molecule, long-read platforms such as nanopore direct RNA sequencing, which require minimum read lengths and benefit from longer sequence context to improve modification resolution and signal interpretation. There is a need for new, advanced techniques for long-read sequencing.

[0006] SUMMARY OF THE INVENTION

[0007] The disclosure relates to nucleic acid engineering and, more specifically, to synthetic RNA constructs and assembly methods that enable precise, site-specific inclusion of one or more RNA base modifications in a controlled and extendable sequence context suitable for long-read sequencing and quantitative benchmarking.

[0008] Disclosed is a method of making a modular RNA construct. The method may comprise annealing a modification loop bait RNA oligonucleotide with a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide; and ligating the annealed oligonucleotides to form the modular RNA construct. In some embodiments, the modification loop bait comprises in 5' to 3' order: 1) a first annealing region, 2) a second annealing region, 3) a first unpaired region, 4) a third annealing region, and 5) a fourth annealing region. In some embodiments, the hairpin comprises in 5' to 3' order: 1) a fifth annealing region, 2) a first loop, 3) a sixth annealing region complementary to the fifth annealing region, and 4) a seventh annealing region complementary to the fourth annealing region. In some embodiments, the hairpin comprises a desthiobiotin moiety. In some embodiments, the method further comprises purifying the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions. In some embodiments, the hairpin comprises a 5' purification tag. In some embodiments, the method further comprises purifying the modular RNA construct by affinity capture of the 5' purification tag.

[0009] In some embodiments, the first loop comprises a repetitive GAGA motif. In some embodiments, the modification loop comprises in 5' to 3' order: 1) an eighth annealing region complementary to the third annealing region, 2) a second loop, and 3) a ninth annealing region complementary to the second annealing region. In some embodiments, the first unpaired region allows formation of the second loop. In some embodiments, the first unpaired region accommodates, houses, accepts, or is configured to receive the second loop. In some embodiments, the second loop comprises an RNA modification. In some embodiments, the RNA modification is N6-methyladenosine (m6A), pseudouridine (T), 5-methylcytidine (m5C), 2'-O-methylation, N1 -methyladenosine (mlA), N7-m ethylguanosine (m7G), inosine (I), or phosphorylation. In some embodiments, the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides. In some embodiments, the defined sequence motif or randomized nucleotides are each independently 3 to 15 nucleotides. In some embodiments, the method further comprises detecting the RNA modification.

[0010] In some embodiments, the poly A comprises in 5' to 3' order: 1) a tenth annealing region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A further comprises a second unpaired region. In some embodiments, the second unpaired region is a barcode. In some embodiments, the barcode distinguishes among different modular RNA constructs in a pooled sample. In some embodiments, the method further comprises preparing a sequencing library.

[0011] In some embodiments, each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 16-21 nucleotides in length. In some embodiments, each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof. In some embodiments, the modular RNA construct is at least 100 nucleotides in length. In some embodiments, the modular RNA construct is 100-200 nucleotides, 200-300 nucleotides, 300-400 nucleotides, 400-500 nucleotides, 500-600 nucleotides, 600-700 nucleotides, 700-800 nucleotides, 800-900 nucleotides, or 900-1000 nucleotides in length.

[0012] In some embodiments, the method further comprises sequencing the modular RNA construct using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the method further comprises assembling the oligonucleotides in a one-pot reaction. In some embodiments, the method further comprises assembling the oligonucleotides in a stepwise manner.

[0013] Disclosed herein is a kit comprising a modification loop bait RNA oligonucleotide, a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide.

[0014] In some embodiments, the modification loop bait comprises in 5' to 3' order: 1) a first annealing region, 2) a second annealing region, 3) a first unpaired region, 4) a third annealing region, and 5) a fourth annealing region. In some embodiments, the hairpin comprises in 5' to 3' order: 1) a fifth annealing region, 2) a first loop, 3) a sixth annealing region complementary to the fifth annealing region, and 4) a seventh annealing region complementary to the fourth annealing region. In some embodiments, the hairpin comprises a desthiobiotin moiety.

[0015] In some embodiments, the hairpin comprises a 5' purification tag. In some embodiments, the modification loop comprises in 5' to 3' order: 1) an eighth annealing region complementary to the third annealing region, 2) a second loop, and 3) a ninth annealing region complementary to the second annealing region. In some embodiments, the first unpaired region allows formation of the second loop. In some embodiments, the first unpaired region accommodates, houses, accepts, or is configured to receive the second loop.

[0016] In some embodiments, the second loop comprises an RNA modification. In some embodiments, the RNA modification is m6A, pseudouridine, or phosphorylation. In some embodiments, the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides. In some embodiments, the defined sequence motifs or randomized nucleotides are each independently 3 to 15 nucleotides. In some embodiments, the poly A comprises in 5' to 3' order: 1) a tenth annealing region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A further comprises a second unpaired region. In some embodiments, the second unpaired region is a barcode. In some embodiments, the barcode distinguishes among different modular RNA constructs in a pooled sample. In some embodiments, each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

[0017] In some embodiments, each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the method further comprises instructions for annealing, ligation, affinity purification, and sequencing library preparation.

[0018] BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Fig. 1 shows a complete modification cassette, including the modification loop. There are 4 component oligos from 5' to 3': Modification Bait, GAGA hairpin loop. Modification loop, and the Poly-A Splint.

[0020] Figs. 2A-2F show descriptions of the components. Fig. 2A shows a modification bait loop. Fig. 2B shows a GAGA hairpin loop. Fig. 2C shows a modification loop (10 random nucleotides pre and post modification pictured). Fig. 2D shows a poly A splint. Fig. 2E shows each of the 4 individual components and how they anneal together to form the final product with the ligation action of RNL2. Fig. 2F shows a fully ligated product with a total length of 231 nucleotides.

[0021] DETAILED DESCRIPTION

[0022] To date over 170 naturally occurring base level RNA modifications have been documented. Modifications have been shown to have biological impact and abnormal behavior in disease states. However, because of the fragility of RNA coupled with the wide scope of modifications, creating strategies for studying these modifications is challenging. No matter what modification identification strategy is being used, one of the best controls is the use of synthetic oligonucleotides with and without modifications at a given position. This allows for robust benchmarking of the precision, accuracy, recall, and quantifiability of these strategies. A current shortcoming of the field of solid-phase synthesis oligos is the short lengths achievable with RNA. Beyond 120 nucleotides, synthetic RNA oligos manufactured in one shot are severely decreased in yield, if not entirely impossible. Modifications also cause a drop in yield and a drop in the probability of successfully synthesizing the oligo.

[0023] With these limitations in mind we designed a synthetic oligonucleotide strategy that combines shorter synthetic oligos into a cassette that can be ligated together and sequenced. The present RNA modification oligonucleotide framework overcomes these limitations by assembling a longer RNA cassette from multiple shorter, synthetically accessible oligonucleotides. The framework delivers a fixed and predictable architecture that can incorporate one or more selected base modifications at defined positions and within tunable local sequence contexts, including randomized nucleotides to reduce sequence bias and better support model training. The constructs are readily ligated, purified using affinity tags, and prepared for downstream sequencing workflows, including nanopore direct RNA sequencing, thereby enabling modification benchmarking, quantification, and machine-learning-based signature discovery in a controlled, scalable, and cost-effective manner.

[0024] Short synthetic oligos: Shorter synthetic RNA oligonucleotides are currently the gold standard of positive controls in RNA modification sequencing. They can prove effective for platforms like Illumina sequencing and mass spectrometry, but are not suitable for long read technologies.

[0025] In-Vitro Transcription: Completely unmodified RNA can be produced through the process of in-vitro transcription. This can produce a powerful fully negative control for modification analysis and model training. Additionally through enzymatic means in-vitro transcribed RNA can be modified in accordance with the modifying enzymes added. While this can be helpful for non-specific pattern analysis, it lacks the site level specificity that model training requires. Additionally the all or nothing modification status of enzymatically modified in-vitro transcribed RNA does not reflect biology, which has much more nuanced RNA modification landscapes.

[0026] Our modular modification cassette design overcomes the shortcomings of short synthetic standards by combining four shorter RNA oligonucleotides into a larger construct (Fig. 1). The total construct size is larger than 200 nucleotides (e.g., 231 nucleotides) including a 15 nucleotide polyA tail. Oligos are annealed through designed oligonucleotide paired annealing regions and then stitched together using RNA ligase (e.g., RNL 2). The final product is modular, requiring only a single oligonucleotide of length 55 to be synthesized for each modification (or combination of modifications) of interest. This significantly reduces the financial burden of producing sufficient quantities of high quality modified oligos.

[0027] How is this different from current strategies? Our modular RNA modification oligonucleotide framework provides all context modification / canonical combinations at a selectable size. Where IVT can either provide fully canonical or fully modified, our strategy can give specific modification control in a variable context setting.

[0028] Additionally our strategy overcomes the length constraints of short oligonucleotide modification frameworks by joining multiple short oligos into a longer construct.

[0029] Why does length matter? Length is not always a key consideration for synthetic standards, but for some technologies such as nanopore sequencing, a minimum length threshold has to be achieved. Nanopore sequencing is of particular interest for RNA modification work because it is currently the only single molecule technology that can directly measure RNA modifications in their native sequence context.

[0030] Direct RNA Sequencing (DRS) is not without its shortcomings. Below 100 nucleotides the ability of the system to produce reliable sequence information is diminished, modifications add a layer of complexity that compounds the difficulties of short RNA sequencing. Thus providing additional length makes the identification of RNA modifications feasible.

[0031] This disclosure allows for the design, construction, and sequencing of customizable modification cassettes. This means that RNA modifications are able to be studied in a wider variety of sequence contexts. Additionally, this strategy is not limited to single modifications and can be extended to multiple modifications, which is an exceptionally challenging area to produce training data for Machine Learning Models.

[0032] This disclosure provides the following advantages: fully random context bound by fully contextualized sequence on both the 5' and 3' end allows for punctuated identification of signal for modification training. Additionally, the random context minimizes sequence dependent bias from entering the training dataset as may be seen with enzymatic approaches.

[0033] This disclosure provides customizable model development for mRNA vaccine production, quantification and quality control and provides therapeutic and diagnostic panels to identify rare diseases linked to non-traditional RNA modifications.

[0034] Definitions

[0035] For convenience, certain terms employed in the specification, examples, and appended claims are collected here. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0036] The term “about” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the values measured or determined, ie., the limitations of the measurement system. Where the terms “about” or “approximately” are used in the context of compositions containing amounts of ingredients or conditions such as temperature, these values include the stated value with a variation of 0-10% around the value (X ± 10%).

[0037] The terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are inclusive in a manner similar to the term “comprising.” The term “consisting” and the grammatical variations of consist encompass embodiments with only the listed elements and excluding any other elements. The phrases “consisting essentially of’ or “consists essentially of’ encompass embodiments containing the specified materials or steps and those including materials and steps that do not materially affect the basic and novel characteristic(s) of the embodiments.

[0038] Ranges are stated in shorthand to avoid having to set out at length and describe each and every value within the range. Therefore, when ranges are stated for a value, any appropriate value within the range can be selected, and these values include the upper value and the lower value of the range. For example, a range of two to thirty represents the terminal values of two and thirty, as well as the intermediate values between two to thirty, and all intermediate ranges encompassed within two to thirty, such as two to five, two to eight, two to ten, etc.

[0039] The term “preventing” is art-recognized, and when used in relation to a condition is well understood in the art, and includes administration of a composition which reduces the frequency of, or delays the onset of, symptoms of a medical condition in a subject relative to a subject which does not receive the composition. Thus, prevention of cancer includes, for example, reducing the incidence of cancer in a population of patients receiving a prophylactic treatment relative to an untreated control population, and / or delaying the onset of cancer in a treated population versus an untreated control population, e.g., by a statistically and / or clinically significant amount.

[0040] The term “ subject ' as used herein refers to a living mammal and may be interchangeably used with the term “patient”. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice and guinea pigs, and the like. The term does not denote a particular age or gender.

[0041] The term “therapeutically effective amount” of a compound with respect to the subject method of treatment refers to an amount of the compound(s) in a preparation which, when administered as part of a desired dosage regimen (to a mammal, preferably a human) alleviates a symptom, ameliorates a condition, or slows the onset of disease conditions according to clinically acceptable standards for the disorder or condition to be treated or the cosmetic purpose, e.g., at a reasonable benefit / risk ratio applicable to any medical treatment. A therapeutically effective amount herein may vary according to factors such as the disease state, age, sex, and weight of the patient, and the ability of the antibody to elicit a desired response in the individual.

[0042] As used herein, the term “treating or “treatment” includes reducing, arresting, or reversing the symptoms, clinical signs, or underlying pathology of a condition to stabilize or improve a subject’s condition or to reduce the likelihood that the subject’s condition will worsen as much as if the subject did not receive the treatment.

[0043] The term "complementary" and "complementarity" are interchangeable and refer to the ability of polynucleotides to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in antiparallel polynucleotide strands or regions. Complementary polynucleotide strands or regions can base pair in the Watson-Crick manner (e.g., A to T, A to U, C to G). 100% (or total) complementary refers to the situation in which each nucleotide unit of one polynucleotide strand or region can hydrogen bond with each nucleotide unit of a second polynucleotide strand or region. Less than perfect (or partial) complementarity refers to the situation in which some, but not all, nucleotide units of two strands or two regions can hydrogen bond with each other and can be expressed as a percentage.

[0044] The term “hybridization” or “annealing” is used to refer to the structure formed by 2 independent strands of RNA that form a double stranded structure via base pairings from one strand to the other. These base pairs are considered to be G-C, A-U, and G-U. (A - Adenine, C - Cytosine, G - Guanine, U - Uracil). As in the case of complementarity, hybridization can be total or partial. “Anneal” means to hybridize complementary nucleic acid sequences under appropriate buffer, temperature, and ionic conditions.

[0045] “Ligate” means to covalently join nucleic acid strands using a ligase.

[0046] The term “oligonucleotide” refers to RNA, DNA, or RNA: DNA oligonucleotides. The term “nucleotide” can refer to both, ribonucleotide or deoxyribonucleotide, unless otherwise explained.

[0047] “Short RNA” means a single-stranded RNA of length 10-60 nucleotides, including but not limited to miRNAs, piRNAs, siRNAs, and fragments thereof.

[0048] “miRNA” refers to mature microRNA molecules typically 18-30 nucleotides, including both 5p and 3p strands and sequence variants (isomiRs).

[0049] “Desthiobiotin” refers to a reversible affinity tag for streptavidin-based capture and elution used to enrich desired constructs.

[0050] “Ionic current signal” means the time-series current measurements generated as a nucleic acid translocates through a nanopore.

[0051] “Modified nucleotide” or “RNA modification” means a chemical modification at a nucleotide base or ribose / phosphate moiety, including but not limited to N6-methyladenosine (m6A), pseudouridine (T), and other base-level modifications that can be incorporated by solid-phase synthesis.

[0052] “Base modification” means a covalent chemical alteration to a nucleobase (e.g., m6A, pseudouridine), phosphorylation, or other chemical changes detectible via altered ionic current.

[0053] “Modification loop” means an RNA oligonucleotide segment that carries at least one predefined modification and is designed to anneal to complementary regions of a separate oligonucleotide to form part of the cassette.

[0054] “Modification loop bait” means an RNA oligonucleotide that provides the backbone of the cassette and contains multiple annealing segments that are complementary to adjacent cassette components.

[0055] “Random nucleotide” and “N” mean a position synthesized to contain any of the four canonical ribonucleotides at an approximately equal probability, thereby generating a pool of sequence variants that sample local sequence context.

[0056] A “hairpin” is a secondary structure making a stem-loop, formed when a single RNA or DNA strand folds back so that two complementary regions base-pair to make a double-stranded stem capped by an unpaired loop. A hairpin comprises a largely Watson-Crick base-paired stem and an apical loop of unpaired (or non-canonical) nucleotides. A “loop” refers to the unpaired segment of nucleotides at the apex of a stem-loop (hairpin) structure.

[0057] “Poly(A)” usually refers to the poly(A) tail: a stretch of adenine nucleotides. It may comprise 3-50 adenine nucleotides. “Barcoding segment” means an identifiable sequence segment used for sample multiplexing, tracking, or deconvolution of pooled constructs.

[0058] An “RNA ligation adapter” is a short, synthetic oligonucleotide that is covalently attached to an end of RNA molecules during next-generation sequencing (NGS) library preparation. By adding known sequences to otherwise unknown RNA ends, adapters provide the handles required for reverse transcription, PCR amplification, sample indexing, and attachment to the sequencing platform.

[0059] A “modular RNA construct” means an assembled RNA cassette formed from multiple synthetic RNA oligonucleotides. For example, it comprises a modification loop bait, a hairpin, a modification loop, and a poly-A splint, whose complementary annealing segments hybridize and are ligated to create a single contiguous RNA molecule. The construct may include a 5' purification tag and a poly(A) tail, and may incorporate one or more modified nucleotides at defined positions. It may be designed so that components can be interchanged (e.g., varying the modification loop) without redesigning the remaining components.

[0060] Methods

[0061] The technology is a modular RNA modification cassette system, a synthetic, build-to-order RNA construct assembled from multiple short pieces of RNA (called oligonucleotides, short strands of nucleic acids) to create a longer RNA molecule that contains precisely placed chemical RNA modifications (chemical changes to RNA bases such as methylation) within controlled and / or randomized surrounding sequence. Its purpose is to generate reliable, scalable standards and training materials to detect, quantify, and model RNA modifications across diverse sequence contexts, particularly for long-read sequencing platforms (notably nanopore direct RNA sequencing) and for machine learning model development and benchmarking.

[0062] The disclosure is a practical, modular way to build long, precisely modified RNA controls at scale. It keeps the synthetic chemistry tractable (shorter oligos), the assembly simple (anneal and ligate with bead-based cleanup), the sequencing compatibility high (poly-A tail, sufficient length), and the data utility strong (broad, unbiased sequence contexts), which together advances both measurement science and downstream clinical and manufacturing applications.

[0063] In practical terms, the disclosure overcomes the two biggest limitations in the field: (1) the difficulty of synthesizing long, modified RNAs in one piece (low yield beyond -120 nucleotides), and (2) the lack of flexible, context-rich controls with site-specific modifications. By modularly assembling shorter synthetic pieces, it yields a robust, longer RNA standard with user-defined modifications and surrounding sequence variation, suitable for research, diagnostics, and manufacturing quality control.

[0064] The system assembles four designed RNA oligonucleotides into a single cassette, including a short poly-A tail (a run of adenines used for compatibility with common library prep and capture steps). Each piece has predefined regions that anneal (bind via base-pairing) to partners and are then ligated (covalently joined) using an RNA ligase enzyme (e.g., T4 RNA Ligase 2). A bead-based purification step is enabled by a built-in affinity tag.

[0065] The four components and their functions are:

[0066] • Modification Loop Bait. This is the backbone strand that provides multiple

[0067] 10-nucleotide annealing sites to bring the other components together in the correct geometry. It also includes unpaired segments designed to create a local “bubble” that accommodates the modification-containing loop.

[0068] • GAGA Hairpin with Purification Tag. A hairpin is a self-folding RNA structure with a loop and stem. Here, the hairpin includes a desthiobiotin tag at the loop so the assembled cassette can be captured and cleaned up using streptavidin-coated beads (affinity purification), avoiding gel-based cleanups. This piece also includes a 5' phosphate group for ligation.

[0069] • Modification Loop. This is the centerpiece: it contains the target RNA modification (e.g., m6A, a methyl group at the N6 position of adenosine) flanked by either defined sequence motifs or randomized nucleotides (“N”) that introduce a controlled diversity of surrounding sequence context. “Random N” positions mean each site has a roughly 25% chance of being A, C, G, or U, enabling systematic sampling of many local sequence environments. The loop is synthesized by solid-phase synthesis (a standard chemical method for making custom oligos) so that modified bases can be placed at exact positions.

[0070] • Poly-A Splint. This strand adds a poly-A tail for compatibility with standard library preparation workflows (e.g., for RT-qPCR, Illumina, Nanopore Direct RNA, or PacBio). It also includes an unpaired segment that can serve as a barcode (a unique sequence tag) to distinguish different constructs in multiplexed experiments, and an annealing region to connect to the bait.

[0071] Assembly proceeds either in a single “one-pot” reaction or in controlled steps: the pieces are mixed to anneal at their complementary regions, then ligated enzymatically to form a continuous RNA. The desthiobiotin tag enables bead capture and wash steps after ligation, yielding a clean, ready -to-sequence construct. The total length, well above -100 nucleotides, meets the reliability threshold of nanopore direct RNA sequencing (the only widely used, single-molecule platform that can read native RNA and its modifications in situ), where shorter RNAs underperform and modification signals can be harder to interpret.

[0072] This modular design addresses longstanding pain points in RNA modification research and standardization.

[0073] • Specific yet flexible modification control. Unlike in vitro transcription (IVT) plus enzymes, which tends to generate either fully unmodified or globally modified transcripts without precise site control, this cassette places a particular modification at an exact site while offering tunable flanking sequence context. “Canonical nucleotides” here means unmodified A, C, G, U. The approach supports single or multiple modifications.

[0074] • Longer, sequencing-ready standards from short, high-yield building blocks. Direct chemical synthesis of long, modified RNAs suffers from sharply reduced yields beyond -120 nucleotides and with certain modifications. By stitching together shorter, high-quality oligos, the cassette achieves sequencing-relevant lengths without the yield penalties of one-piece synthesis.

[0075] • Reduced sequence bias for machine learning training. The use of randomized flanking positions (configurable, e.g., 6-20 “N”s around the modification) captures a broad diversity of sequence contexts, minimizing sequence-dependent bias and supplying better-balanced training data for ML classifiers that learn modification signatures from raw signal.

[0076] • Cost and workflow efficiency. Only the modification loop needs to be resynthesized to explore different modifications or motif sets; the other components are reusable. The desthiobiotin-streptavidin cleanup streamlines purification compared to gel-based methods.

[0077] • Compatibility with multiple platforms. The poly-A tail and overall design make the cassette plug-and-play for nanopore Direct RNA Sequencing (DRS) and adaptable to other workflows (RT-qPCR, Illumina, PacBio). Length is especially important for nanopore DRS signal stability and modification detection performance.

[0078] Representative applications include: • Training and benchmarking of modification-calling algorithms. The cassette provides known-truth, single-site modified standards across many sequence contexts to measure precision, recall, accuracy, and quantifiability, and to generate robust training corpora.

[0079] • Spike-in controls for quantification. Defined, barcoded standards can be added to biological samples to calibrate detection thresholds, estimate modification stoichiometry (the fraction of molecules modified), and monitor batch-to-batch assay performance.

[0080] • Therapeutic and diagnostic development. By enabling precise and diverse modification contexts, the system supports development and quality control of mRNA vaccines, assays for rare diseases associated with non-canonical RNA modifications, and broader epitranscriptomic studies (the layer of gene regulation mediated by RNA modifications).

[0081] The sequencing can be performed by any direct sequencing method that comprises a nanopore, for instance Oxford Nanopore technologies. The nanopore direct sequencing and the materials and protocols to perform it are known in the art. For instance, in US Patent Number 6,015,714. In some embodiments, the oligonucleotide adapter configured to perform nanopore direct sequencing is a double-stranded sequencing adapter DNA oligonucleotide with a helicase protein bound to one of the strands and having the complementary strand, a first DNA adapter oligonucleotide hybridization region. In some embodiments, the nanopore direct sequencing comprises a membrane, said membrane can be either solid-state or biological membranes.

[0082] Any known nanopore direct sequencing method or product can be used, for instance the one disclosed in US Patent Number 6,015,714 or US6,362,002.

[0083] The analysis or performing algorithm used can be any commercial one known by a skilled of many performing algorithms known in the art suitable for nanopore direct RNA sequencing. The first step is extracting the reads. This step can be done by commercial software, for instance MinKNOW or any software configured to analyze the sequencing results of the nanopore direct sequencing. Next step of the analysis is the base calling, which can be done by a skilled person using any of several known performing algorithms in the field, such as Guppy or Bonito. Last step of the analysis is mapping, which can be done by several known performing algorithms. For example, Minimap2 or BWA which is a versatile sequence alignment program that aligns nucleic acid sequences against a large reference database. In some embodiments, the performing algorithm is configured to capture (and sequence) more miRNA in a quantitative way.

[0084] Disclosed is a method of making a modular RNA construct. The method may comprise annealing a modification loop bait RNA oligonucleotide with a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide; and ligating the annealed oligonucleotides to form the modular RNA construct.

[0085] In some embodiments, the modification loop bait comprises in 5' to 3' order: 1) a first annealing region, 2) a second annealing region, 3) a first unpaired region, 4) a third annealing region, and 5) a fourth annealing region. In some embodiments, the hairpin comprises in 5' to 3' order: 1) a fifth annealing region, 2) a first loop, 3) a sixth annealing region complementary to the fifth annealing region, and 4) a seventh annealing region complementary to the fourth annealing region. In some embodiments, the hairpin comprises a desthiobiotin moiety. In some embodiments, the method further comprises purifying the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions. In some embodiments, the hairpin comprises a 5' purification tag. In some embodiments, the method further comprises purifying the modular RNA construct by affinity capture of the 5' purification tag.

[0086] In some embodiments, the first loop comprises a repetitive GAGA motif. In some embodiments, the modification loop comprises in 5' to 3' order: 1) an eighth annealing region complementary to the third annealing region, 2) a second loop, and 3) a ninth annealing region complementary to the second annealing region. In some embodiments, the first unpaired region allows formation of the second loop. In some embodiments, the first unpaired region accommodates, houses, accepts, or is configured to receive the second loop. In some embodiments, the second loop comprises an RNA modification. In some embodiments, the RNA modification is N6-methyladenosine (m6A), pseudouridine (T), 5-methylcytidine (m5C), 2'-O-methylation, N1 -methyladenosine (mlA), N7-m ethylguanosine (m7G), inosine (I), or phosphorylation. In some embodiments, the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides. In some embodiments, the defined sequence motif or randomized nucleotides are each independently 3 to 15 nucleotides. In some embodiments, the method further comprises detecting the RNA modification. In some embodiments, the poly A comprises in 5' to 3' order: 1) a tenth annealing region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A further comprises a second unpaired region. In some embodiments, the second unpaired region is a barcode. In some embodiments, the barcode distinguishes among different modular RNA constructs in a pooled sample. In some embodiments, the method further comprises preparing a sequencing library.

[0087] In some embodiments, each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 16-21 nucleotides in length. In some embodiments, each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof. In some embodiments, the modular RNA construct is at least 100 nucleotides in length. In some embodiments, the modular RNA construct is 100-200 nucleotides, 200-300 nucleotides, 300-400 nucleotides, 400-500 nucleotides, 500-600 nucleotides, 600-700 nucleotides, 700-800 nucleotides, 800-900 nucleotides, or 900-1000 nucleotides in length.

[0088] In some embodiments, the method further comprises sequencing the modular RNA construct using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the method further comprises assembling the oligonucleotides in a one-pot reaction. In some embodiments, the method further comprises assembling the oligonucleotides in a stepwise manner.

[0089] In some embodiments, only the modification loop varies among a panel of constructs, and the modification loop bait, hairpin, and poly-A splint are common to the panel. In some embodiments, the plurality of modular RNA constructs includes con-structs with randomized nucleotides flanking the modified position to provide agnostic meas-urements for signal learning.

[0090] Disclosed herein is a method for training a model to detect RNA modifications, comprising: generating a plurality of modular RNA constructs described herein that comprise at least one modified nucleotide in a range of sequence contexts; sequencing the plurality of modular RNA constructs to obtain signal data associated with modified and unmodified positions; and using the signal data to train model parameters to distinguish modified from unmodified nucleotides. Disclosed herein is a method for quantifying an RNA modification in a test sample, comprising: spiking the test sample with a modular RNA construct described herein having a known fraction of modified molecules at a predetermined site; sequencing the spiked test sample; and estimating a modification level in the test sample by calibrating against the known fraction in the modular RNA construct.

[0091] Modular RNA Construct

[0092] Disclosed herein is a modular RNA construct comprising: a modification loop bait RNA oligonucleotide having a plurality of annealing segments; a hairpin RNA oligonucleotide that is self-annealing and comprises a hairpin loop and optionally a 5' purification tag; a modification loop RNA oligonucleotide comprising at least one modified nucleotide at a predetermined position and comprising 5' and 3' annealing segments complementary to the modification loop bait; and a poly-A splint RNA oligonucleotide comprising an annealing segment complementary to the modification loop bait, an unannealed barcoding segment, and a poly(A) tail; wherein the annealing segments hybridize such that the modification loop bait anneals to each of the hairpin, the modification loop, and the poly-A splint, and wherein the oligonucleotides are ligated to form the modular RNA construct. In some embodiments, each annealing segment is 10 nucleotides in length. In some embodiments, the hairpin comprises an 8-nucleotide loop and a 5' desthiobiotin purification tag. In some embodiments, the modification loop comprises randomized nucleotides flanking the modified nucleotide. In some embodiments, the modification loop comprises from three to ten randomized nucleotides on each side of the modified nucleotide. In some embodiments, the modified nucleotide is selected from N6-methyladenosine and pseudouridine. In some embodiments, the poly-A splint comprises a poly(A) tail configured for poly(A)-dependent library preparation. In some embodiments, the modular RNA construct has a length suitable for nanopore direct RNA sequencing. In some embodiments, the modification loop comprises two or more modified nucleotides. In some embodiments, the hairpin, the modification loop, and the poly-A splint each comprise a 5' phosphate. In some embodiments, the modification loop bait comprises an unannealed spacer to permit the modification loop to form a bubble in the secondary structure. In some embodiments, the unannealed barcoding segment distinguishes among different modular RNA constructs in a pooled sample. In some embodiments, the poly(A) tail comprises at least 15 adenosine residues. Kits

[0093] Disclosed herein is a kit comprising a modification loop bait RNA oligonucleotide, a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide.

[0094] In some embodiments, the modification loop bait comprises in 5' to 3' order: 1) a first annealing region, 2) a second annealing region, 3) a first unpaired region, 4) a third annealing region, and 5) a fourth annealing region. In some embodiments, the hairpin comprises in 5' to 3' order: 1) a fifth annealing region, 2) a first loop, 3) a sixth annealing region complementary to the fifth annealing region, and 4) a seventh annealing region complementary to the fourth annealing region. In some embodiments, the hairpin comprises a desthiobiotin moiety.

[0095] In some embodiments, the hairpin comprises a 5' purification tag. In some embodiments, the modification loop comprises in 5' to 3' order: 1) an eighth annealing region complementary to the third annealing region, 2) a second loop, and 3) a ninth annealing region complementary to the second annealing region. In some embodiments, the first unpaired region allows formation of the second loop. In some embodiments, the first unpaired region accommodates, houses, accepts, or is configured to receive the second loop.

[0096] In some embodiments, the second loop comprises an RNA modification. In some embodiments, the RNA modification is m6A, pseudouridine, or phosphorylation. In some embodiments, the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides. In some embodiments, the defined sequence motifs or randomized nucleotides are each independently 3 to 15 nucleotides.

[0097] In some embodiments, the poly A comprises in 5' to 3' order: 1) a tenth annealing region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A further comprises a second unpaired region. In some embodiments, the second unpaired region is a barcode. In some embodiments, the barcode distinguishes among different modular RNA constructs in a pooled sample. In some embodiments, each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

[0098] In some embodiments, each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length. In some embodiments, the method further comprises instructions for annealing, ligation, affinity purification, and sequencing library preparation. In some embodiments, the kit comprises streptavidin-coated beads and buffers for reversible capture and elution. In some embodiments, the adapters and baits are provided with RNase-free HPLC purification and are supplied in concentrations suitable for formation of cassettes under ligation conditions specified in the instructions.

[0099] EXAMPLES

[0100] The invention now being generally described, it will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present invention, and are not intended to limit the invention.

[0101] Example 1: Exemplified Modular RNA cassette

[0102] Component Description:

[0103] The full construct consists of four synthetic oligos, two with fixed identity, one with a primarily fixed identity and a component for barcoding, and a highly variable modification loop. From 5’ to 3’ the components are the modification loop bait (Fig. 2A), the GAGA Hairpin (Fig. 2B), the modification loop (Fig. 2C), and the Poly-A splint (Fig. 2D). Each component anneals to at least one other component in the cassette complex (Table 1).

[0104] Table 1. Length and annealing partners for cassette oligos

[0105] Component Length 5' Annealing 3 ' Annealing Component Component Modification Loop 56 N / A GAGA Hairpin Bait

[0106] GAGA Hairpin 60 Self Modification Loop Bait

[0107] Modification Loop 55 Modification Loop Modification Loop Bait Bait

[0108] Poly-A Splint 60 Modification Loop N / A

[0109] Bait

[0110]

[0111] Modification Loop Bait: The modification loop bait’s primary purpose is to serve as the backbone of the cassette, annealing to all the other oligonucleotide components. Each of the annealing segments is 10 nucleotides in length, allowing for high sequence specificity and efficiency. The sequence of the modification bait loop is as follows:

[0112] 5 ’ CACAACCUACCCUCGGUGUAUAUGCAAGAAGAAGAACAUGGACUGCGAGC UAAGUG3’ (SEQ ID NO: 1)

[0113] CACAA (SEQ ID NO: 2) = Un-annealed overhang

[0114] CCUACCCUCG (SEQ ID NO: 3) = Annealing complement for the Poly-A Splint GUGUAUAUGC (SEQ ID NO: 4) = Annealing complement for the Modification Loop 3 ’ End

[0115] AAGAAGAAGAA (SEQ ID NO: 5) = Un-annealed spaces to allow modification Loop to anneal and create a bubble in the secondary structure.

[0116] CAUGGACUGC (SEQ ID NO: 6) = Annealing complement for the Modification Loop 5’ End

[0117] GAGCUAAGUG (SEQ ID NO: 7) = Annealing complement for the GAGA Hairpin.

[0118] GAGA Hairpin Loop: The GAGA hairpin is a self-annealing oligonucleotide with an 8 nucleotide loop. The GAGA hairpin shares a 5’ end with the modification loop bait, and a 3’ end with the modification loop. Additionally the GAGA Hairpin includes a purification tag (ex. desthiobiotin) at the 5’ end of the hairpin loop. The purification tag allows for a bead based cleanup rather than a gel based cleanup or size selection. A 5’ Phosphate is included in the oligonucleotide sequence for ligation.

[0119] The sequence is as follows:

[0120] 5’Phosphate- AUGGCAUUAAGAUCAGUUGCA(DesthioBiotin)GAGAGAGAUGCAACUGAUCUU AAUGCCAUCACUUAGCUC 3’ (SEQ ID NO: 8) AUGGCAUUAAGAUCAGUUGCA (SEQ ID NO: 9) = First Self annealing section (Desthio-Biotin)GAGAGAGA (SEQ ID NO: 10) = Hairpin Loop including DesthioBiotin

[0121] UGCAACUGAUCUUAAUGCCAU (SEQ ID NO: 11) = Complement of the first self-annealing portion

[0122] CACUUAGCUC (SEQ ID NO: 12) = Annealing complement for the Modification Loop Bait

[0123] Modification Loop: The modification loop is the crux of the construct because it includes the modification of interest. Any modification that can be reliably produced through solid phase synthesis is a potential target for this technique. Additionally it has become increasingly clear that the nucleotide sequence surrounding modifications is an important element both for the modification operation in the cell, and in measuring the signal of a modification. One option is to sequence oligonucleotides with the exact motif desired, or a set of motifs. This provides nuanced control and guarantees a higher yield of potentially useful motifs, but it suffers from requiring pre-existing knowledge of motif information.

[0124] Alternatively solid phase synthesis can be undertaken with random canonical nucleotides, labeled as N, that represent a 25% chance for each of the 4 nucleotides to be included at a position. Thus with ‘random’ nucleotides all sequence contexts can be captured, and sequenced given the known structure of the cassette.

[0125] In our design the number of random nucleotides can range from 6, 3 before and 3 after the modification, to 20, 10 before and 10 after the modification. In the extreme case 20 random nucleotides results in 1,099,511,627,776 possible combinations. With current throughput, this is a challenging number of options to get sufficient coverage of to learn in a machine learning context. But with advances in sequencing technologies throughput has been increasing rapidly, making this oligonucleotide strategy adaptable to the current state of sequencing technology, current and future.

[0126] An example sequence for an m6A modification loop with 6 random nucleotides flanking on both sides follows:

[0127] 5’ Phosphate - GCAGUCCAUGNNNNNN(m6A)NNNNNNGCAUAUACAC (SEQ ID NO: 13)

[0128] GCAGUCCAUG (SEQ ID NO: 14) = Annealing compliment to Mod Loop Bait sequence

[0129] NNNNNN (SEQ ID NO: 15) = 6 random nucleotides pre modification to provide increased sequence context for signal learning

[0130] (m6A) = Modification in the modification loop

[0131] NNNNNN (SEQ ID NO: 16) = 6 random nucleotides post modification to provide increased sequence context for signal learning

[0132] GCAUAUACAC (SEQ ID NO: 17) = Annealing compliment to Mod Loop Bait Sequence

[0133] Poly-A Splint: The Poly-A splint is designed to give additional length on the 3’ end of the cassette, allow the cassette to easily plug into standard sequencing protocols such as RT-qPCR, Illumina sequencing, Nanopore Direct RNA sequencing, and Pacbio sequencing. The Poly-A splint includes a section of annealed nucleotides that bind it to the Mod Loop Bait oligonucleotide, a poly-A tail for use in sequencing library preparation, and a section of un-annealed nucleotides that can be used as a barcoding system. The sequence is as follows: 5’Phosphate-CGAGGGUAGGAGGAAGGAGGUAAAUACAAAUAGGUAGGGAGAAGAAAAAAA AAAAAAAAA (SEQ ID NO: 18)

[0134] CGAGGGUAGG (SEQ ID NO: 19) = Annealing compliment to Mod Loop Bait sequence

[0135] AGGAAGGAGGUAAAUACAAAUAGGUAGGGAGAAG (SEQ ID NO: 20) = Unannealed barcodable segment of nucleotides

[0136] AAAAAAAAAAAAAAAA (SEQ ID NO: 21) = Poly-A tail for library preparation

[0137] Cassette Construction:

[0138] The cassette can be assembled in either a one pot system, or step wise to allow for increased control over variables. The purification tag means that at each step a bead cleanup can be used to remove unwanted product.

[0139] Applications:

[0140] The applications for this system may be a tool for training modification identification technologies and as a spike in control for modification quantification.

[0141] For modification training, the all context nature of the modification loop can provide agnostic measurements for signal learning. This prevents bias of previous research from guiding learning of modification signatures moving forward. Similarly the spike in controls can.

[0142] Protocol:

[0143] The protocol for the modular hairpin system has a couple additional steps to the previous version. First, the modular hairpin pool is created by having lOuM of the hairpin and equimolar proportions of each specific bait adding up to lOuM in annealing buffer (lOmM Tris-HCl pH 8.0, 50mM NaCl). This is denatured at 95°C for 3 minutes, cooled on ice for 10 minutes, then there is a Streptavidin Cl bead clean up (Thermo Fisher 65001), eluting in the same amount of volume as was there previously. This eluted product will be referred to as the pooled modular hairpin and we will assume a lOuM concentration of the created bait hairpin pool. The pooled modular hairpin is ligated at 37°C for Ihr with IX Quick Ligation Buffer (NEB B6058S) and 500 U / mL T4 RNA Ligase 2 (NEB M0239S). Following another identical Streptavidin Cl bead clean up (Thermo Fisher 65001), again eluting in the same initial volume of water, this can be stored in the -80°C freezer for up to 6 months. When prepared for the next step, in a 20uL reaction, have 5uM pool hairpin bait, 20uM universal poly A adapter, IX Quick Ligation Buffer (NEB B6058S), and 500 U / mL T4 RNA Ligase 2 (NEB M0239S) and 5uM of the modification loop. Allow to ligate at 37°C for 1 hour and follow with another Streptavidin Cl bead clean up (Thermo Fisher 65001). At this point the RNA is ready for a standard ONT DRS library preparation followed by sequencing.

[0144] INCORPORATION BY REFERENCE

[0145] Each of the patents, published patent applications, and non-patent references cited herein are hereby incorporated by reference in their entirety.

[0146] EQUIVALENTS

[0147] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Such equivalents are intended to be encompassed by the following claims.

Claims

We claim:

1. A method of making a modular RNA construct, comprising:annealing a modification loop bait RNA oligonucleotide with a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide; andligating the annealed oligonucleotides to form the modular RNA construct.

2. The method of claim 1, wherein the modification loop bait comprises in 5' to 3' order:1) a first annealing region,2) a second annealing region,3) a first unpaired region,4) a third annealing region, and5) a fourth annealing region.

3. The method of claim 2, wherein the hairpin comprises in 5' to 3' order:1) a fifth annealing region,2) a first loop,3) a sixth annealing region complementary to the fifth annealing region, and4) a seventh annealing region complementary to the fourth annealing region.

4. The method of claim 3, wherein the hairpin comprises a desthiobiotin moiety.

5. The method of claim 4, further comprising purifying the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions.

6. The method of claim 3, wherein the hairpin comprises a 5' purification tag.

7. The method of claim 6, further comprising purifying the modular RNA construct by affinity capture of the 5' purification tag.

8. The method of any one of claims 1-7, wherein the first loop comprises a repetitive GAGA motif.

9. The method of any one of claims 2-8, wherein the modification loop comprises in 5' to 3' order:1) an eighth annealing region complementary to the third annealing region,2) a second loop, and3) a ninth annealing region complementary to the second annealing region.

10. The method of claim 9, wherein the first unpaired region allows formation of the second loop.

11. The method of claim 9 or 10, wherein the second loop comprises an RNA modification.

12. The method of claim 11, wherein the RNA modification is N6-methyladenosine (m6A), pseudouridine (T), 5-methylcytidine (m5C), 2'-O-methylation, N1 -methyladenosine (ml A), N7-m ethylguanosine (m7G), inosine (I), or phosphorylation.

13. The method of claim 11 or 12, wherein the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides.

14. The method of claim 13, wherein the defined sequence motif or randomized nucleotides are each independently 3 to 15 nucleotides.

15. The method of any one of claims 11-14, further comprising detecting the RNA modification.

16. The method of any one of claims 1-15, wherein the poly A comprises in 5' to 3' order:1) a tenth annealing region, and2) a stretch of adenine nucleotides.

17. The method of claim 16, wherein the poly A further comprises a second unpaired region.

18. The method of claim 17, wherein the second unpaired region is a barcode.

19. The method of claim 18, wherein the barcode distinguishes among different modular RNA constructs in a pooled sample.

20. The method of claim 18 or 19, further comprising preparing a sequencing library.

21. The method of any one of claims 2-20, wherein each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 16-21 nucleotides in length.

22. The method of any one of claims 3-21, wherein each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

23. The method of any one of claims 2-22, wherein each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

24. The method of any one of claims 15-23, wherein the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

25. The method of any one of claims 1-24, wherein ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof.

26. The method of any one of claims 1-25, wherein the modular RNA construct is at least 100 nucleotides in length.

27. The method of any one of claims 1-26, wherein the modular RNA construct is 100-200 nucleotides, 200-300 nucleotides, 300-400 nucleotides, 400-500 nucleotides, 500-600 nucleotides, 600-700 nucleotides, 700-800 nucleotides, 800-900 nucleotides, or 900-1000 nucleotides in length.

28. The method of any one of claims 1-27, further comprising sequencing the modular RNA construct using a sequencing platform.

29. The method of claim 28, wherein the sequencing platform is a nanopore sequencing platform.

30. The method of any one of claims 1-29, further comprising assembling the oligonucleotides in a one-pot reaction.

31. The method of any one of claims 1-29, further comprising assembling the oligonucleotides in a stepwise manner.

32. A kit comprising a modification loop bait RNA oligonucleotide, a hairpin RNA oligonucleotide, a modification loop RNA oligonucleotide, and a poly-A splint RNA oligonucleotide.

33. The kit of claim 32, wherein the modification loop bait comprises in 5' to 3' order:1) a first annealing region,2) a second annealing region,3) a first unpaired region,4) a third annealing region, and5) a fourth annealing region.

34. The kit of claim 33, wherein the hairpin comprises in 5' to 3' order:1) a fifth annealing region,2) a first loop,3) a sixth annealing region complementary to the fifth annealing region, and 4) a seventh annealing region complementary to the fourth annealing region.

35. The kit of claim 34, wherein the hairpin comprises a desthiobiotin moiety.

36. The kit of claim 34, wherein the hairpin comprises a 5' purification tag.

37. The kit of any one of claims 33-36, wherein the modification loop comprises in 5' to 3' order:1) an eighth annealing region complementary to the third annealing region, 2) a second loop, and3) a ninth annealing region complementary to the second annealing region.

38. The kit of claim 37, wherein the first unpaired region allows formation of the second loop.

39. The kit of claim 37 or 38, wherein the second loop comprises an RNA modification.

40. The kit of claim 39, wherein the RNA modification is m6A, pseudouridine, or phosphorylation.

41. The kit of claim 39 or 40, wherein the RNA modification is flanked on each end by a defined sequence motif or randomized nucleotides.

42. The kit of claim 41, wherein the defined sequence motifs or randomized nucleotides are each independently 3 to 15 nucleotides.

43. The kit of any one of claims 32-42, wherein the poly A comprises in 5' to 3' order:1) a tenth annealing region, and2) a stretch of adenine nucleotides.

44. The kit of claim 43, wherein the poly A further comprises a second unpaired region.

45. The kit of claim 44, wherein the second unpaired region is a barcode.

46. The kit of claim 45, wherein the barcode distinguishes among different modular RNA constructs in a pooled sample.

47. The kit of any one of claims 33-46, wherein each annealing region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

48. The kit of any one of claims 34-47, wherein each loop is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

49. The kit of any one of claims 33-48, wherein each unpaired region is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

50. The kit of any one of claims 43-49, wherein the stretch of adenine nucleotides is 5-10 nucleotides in length, 10-15 nucleotides in length, or 15-20 nucleotides in length.

51. The kit of any one of claims 32-50, further comprising instructions for annealing, ligation, affinity purification, and sequencing library preparation.

Citation Information

Patent Citations

  • Methods to capture and / or remove highly abundant RNAS from a heterogenous RNA sample

    US20150218620A1

  • RNA molecules

    US20200263174A1

  • Methods of synthesizing RNA molecules

    US20220411841A1

  • Method of producing hairpin single-stranded RNA molecule

    US20240167033A1

  • Methods for processing and amplifying nucleic acids

    US9598727B2