Capture, detection, adaptation, and sequencing of microrna using nanopore technology
By forming a cassette with a hairpin and poly A oligonucleotide for short RNAs, nanopore sequencing achieves improved read lengths and modification detection, addressing the limitations of existing sequencing methods for microRNAs.
Patent Information
- Application Number
- PCT/US2025/052086
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-22
- Filing Date
- 2025-10-22
- Publication Date
- 2026-04-30
AI Technical Summary
Existing sequencing methods for short RNAs, particularly microRNAs, face challenges in achieving accurate read lengths and detecting native modifications due to signal truncation and indirect sequencing approaches, which hinder robust detection and analysis.
A method involving hybridization of a hairpin nucleotide sequence and poly A oligonucleotide to short RNAs, followed by ligation to form a cassette, which is then sequenced using nanopore technology, allowing for direct sequencing and modification detection.
Enables accurate, direct sequencing of short RNAs with improved signal normalization and modification detection, facilitating earlier and more precise disease detection and deeper understanding of gene regulation.
Smart Images

Figure US2025052086_30042026_PF_FP_ABST
Abstract
Description
[0001] CAPTURE, DETECTION, ADAPTATION, AND SEQUENCING OF MICRORNA USING NANOPORE TECHNOLOGY RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application serial number 63 / 710,330, filed October 22, 2024.
[0003] BACKGROUND
[0004] Short noncoding RNAs, such as miRNAs (approximately 18-30 nucleotides), are critical post-transcriptional regulators implicated in development, homeostasis, and disease, including cancer and neurodegenerative disorders. Despite their significance, direct sequencing of native miRNAs remains technically challenging. Sequencing-by-synthesis approaches infer sequence indirectly by copying a target, which obscures native modifications and introduces bias through reverse transcription and amplification. Direct RNA sequencing using nanopore platforms enables measurement of native molecules but is poorly suited to molecules shorter than approximately 60-70 nucleotides. During nanopore direct RNA sequencing, a motor protein regulates translocation but commonly detaches before the 5'-terminal 12-15 nucleotides are read, causing loss of signal for terminal bases. For miRNAs of ~22 nucleotides, this signal truncation renders basecalling unreliable and prevents robust detection of base modifications.
[0005] Existing enrichment strategies do not adequately solve technical barriers, such as increasing effective read length to enable accurate signal normalization and basecalling; capturing native short RNA molecules with high specificity and ligation efficiency; and integrating selective enrichment with standard nanopore workflows for high-throughput processing. There is a need for new, advanced techniques to sequence short RNAs effectively.
[0006] SUMMARY OF THE INVENTION
[0007] Disclosed is a method of preparing a sequencing library for a target short RNA, comprising: hybridizing a hairpin to at least a portion of the target short RNA and a poly A; ligating the poly A to the 3 ' end of the target short RNA; ligating the target short RNA to the 3' end of the hairpin to form a cassette comprising the poly A, the target short RNA, and the hairpin. Disclosed is a method of preparing a sequencing library for a target short RNA. The method may comprise: hybridizing a hairpin nucleotide sequence to a poly A oligonucleotide via a first hybridization region and hybridizing the hairpin nucleotide sequence to at least a portion of the target short RNA via a second hybridization region; ligating the poly A oligonucleotide to the 3' end of the target short RNA; and ligating the target short RNA to the 3' end of the hairpin nucleotide sequence, thereby forming a cassette comprising the poly A oligonucleotide, the target short RNA, and the hairpin nucleotide sequence.
[0008] In some embodiments, the method further comprises sequencing the cassette using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the hairpin nucleotide sequence comprises in 5' to 3' order: 1) the first hybridization region, 2) the second hybridization region complementary to at least a portion of the target short RNA, 3) a third hybridization region, 4) a loop, and 5) a fourth hybridization region complementary to the third hybridization region. In some embodiments, the hairpin nucleotide sequence is 30 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is RNA.
[0009] In some embodiments, the poly A oligonucleotide comprises in 5' to 3' order: 1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A oligonucleotide is 10 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the poly A is RNA. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof.
[0010] In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA. In some embodiments, the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter. In some embodiments, the target short RNA comprises a base modification; and the method further comprises detecting the base modification.
[0011] Disclosed herein is a method for preparing a sequencing library for a target short RNA. The method may comprises: hybridizing a Corresponding DNA or RNA Bait (CDB) to at least a portion of a target short RNA via a second hybridization region; ligating a 3' Nanopore Ligation Adapter (NLA) to the 3' end of the target short RNA; and ligating a 5' Enrichment Adapter (EA) to the 5' end of the target short RNA, thereby forming a cassette comprising the NLA, the target short RNA, and the EA.
[0012] In some embodiments, the method further comprises sequencing the cassette using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the CDB comprises in 5' to 3' order: 1) a first hybridization region, 2) the second hybridization region complementary to at least a portion of the target short RNA, and 3) a third hybridization region. In some embodiments, the CDB is 20 to 100 nucleotides in length. In some embodiments, the CDB is 20 to 40 nucleotides in length, 40 to 50 nucleotides in length, 50 to 60 nucleotides in length, 60 to 70 nucleotides in length, 70 to 80 nucleotides in length, 80 to 90 nucleotides in length, or 90 to 100 nucleotides in length. In some embodiments, the CDB comprises a randomized region capable of capturing a plurality of short RNA species. In some embodiments, the CDB is RNA or DNA.
[0013] In some embodiments, the NLA comprises in 5' to 3' order: 1) a fourth hybridization region complementary to the first hybridization region, and 2) a fifth hybridization complementary to an RNA ligation adapter. In some embodiments, the NLA is 10 to 50 nucleotides in length. In some embodiments, the NLA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the NLA is RNA or DNA. In some embodiments, the EA comprises in 5' to 3' order: 1) a sixth hybridization region, 2) a loop, 3) a seventh hybridization region complementary to the sixth hybridization region, and 4) an eighth hybridization region complementary to the third hybridization region.
[0014] In some embodiments, the EA is 10 to 50 nucleotides in length. In some embodiments, the EA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the EA is RNA or DNA. In some embodiments, the EA comprises a desthiobiotin moiety. In some embodiments, the method further comprises enriching the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions.
[0015] In some embodiments, the NLA and the EA together add at least 100 nucleotides of additional sequence to the cassette. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof. In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA. In some embodiments, the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter. In some embodiments, the target short RNA comprises a base modification and the method further comprises detecting the modification.
[0016] Disclosed herein is a kit comprising a hairpin nucleotide sequence and a poly A oligonucleotide. In some embodiments, the hairpin nucleotide sequence comprises in 5' to 3' order: 1) a first hybridization region, 2) a second hybridization region complementary to at least a portion of a target short RNA, 3) a third hybridization region, 4) a loop, and 5) a fourth hybridization region complementary to the third hybridization region. In some embodiments, the hairpin nucleotide sequence is 30 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length.
[0017] In some embodiments, the hairpin is RNA. In some embodiments, the poly A oligonucleotide comprises in 5' to 3' order: 1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A oligonucleotide is 10 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is RNA. In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA.
[0018] BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figs. 1A-1C show an exemplified system. Fig. 1A shows a brief overview of the system using miRNA 455-3p and how everything fits together to create a 106nt construct. Fig. IB shows a visualization of the system using a 6% TBE-UREA gel (Thermo Fisher EC68652BOX) that should denature the hairpin. Double ligation represents all the oligonucleotides being present and thus able to ligate. Single ligation only has the miRNA-455 3p hairpin and the synthetic miRNA 455-3p such that only one ligation can occur. Both double and single ligation bands are annotated for clarity. Fig. 1C shows that the double ligation product went through library preparation and was sequenced, confirming the full-length product was produced.
[0020] Figs. 2A-2C show miRNA sequencing cassette. Fig. 2A shows a miRNA sequencing cassette composed of Nanopore Ligation Adapter (NLA), target miRNA, Enrichment Adapter (EA) and Corresponding DNA Bait (CDB). These components are annealed and ligated together using a combination of T4 DNA ligase and RNL2. The CDB contains complementary overhangs of 10 nucleotides for both the NLA and the EA. Annealing can be performed in total RNA or from a short RNA isolation kit. The EA also has a Desthiobioton synthetically added to the hairpin structure allowing for enrichment with Cl beads. Fig. 2B shows that sequencing and alignment demonstrates full length products are being generated with high alignment identity. Fig. 2C shows the ionic current read out for an exemplary biological miR-125b-5p molecule sequenced from a postmortem pre-frontal cortex sample.
[0021] Figs. 3A-3H show ionic current model framework for tRNA classification and analysis that is adapted for miRNA. Fig. 3A shows diagram of the current tRNA adaptation framework including the 5' splint and the 3' splint used to extend the total length of the tRNA. The sequencing adapter is the same sequencing adapter used in the traditional ONT DRS sequencing protocol. Fig. 3B shows diagram of miRNA cassette adaptation strategy including the 5' splint and the 3' splint to extend the total length of the miRNA. The sequencing adapter is included. For the miRNA, the unligated bait used for the capture of and splinted ligation of miRNA to the cassette is included. This component of the cassette would not be visible in ionic current data since it is not ligated. Fig. 3C shows schematic of model architecture for joint classification and segmentation of the ionic current. The model is built on a transformer encoder backbone that processes variable-length input sequences.
[0022] Classification is performed using the learned features of a prepended [CLS] token, which represents the whole sequence, and passing it through a classification layer. Segmentation is formulated as a sequence-to-sequence task by applying a linear projection to the encoder outputs corresponding to each of the input tokens. Fig. 3D shows output of the seq2seq classification of ionic current chunks for segmentation of a sample tRNA. Each section of ionic current signal is classified as one of the four components: sequencing adapter, 3' adapter, target RNA molecule, and 5' adapter. This signal level classification is produced in conjunction with the molecular level classification. Fig. 3E shows confusion matrix for classification of 52 E. Coli tRNA classes with unique sequences with a median accuracy > 98%. Fig. 3F shows sliced out signal from the classification provided by the seq2seq model coupled with the corresponding basecalls. The signal is paired with sequence through the move table, an estimation of signal to sequence mapping produced by ONT's Dorado basecaller. This sequence, derived from signal slicing, is used for alignment. Fig. 3G shows diagram of the modified Wagner-Fischer algorithm to calculate an edit distance-based alignment. Using the classification of the RNA molecule to define the reference and the sliced signal to provide the target sequence region, the computational burden of edit distance calculations is mitigated and the optimal alignment for each read can be achieved. Fig. 3H shows an IGV visualization of alignment results from the tRNA classification and seq2seq defined signal slicing followed by the modified Wagner-Fischer alignment algorithm. For a biological E. Coli tRNA set a median alignment identity of 89%.
[0023] DETAILED DESCRIPTION
[0024] Micro RNAs (miRNAs) function as specific transcriptomic regulators in the cell and can have varied effects from complete gene silencing to activating, and controlling the rate of, translation. Recently miRNA has been getting more attention as it is examined for its diagnostic and therapeutic potential for various cancers, Alzheimer's Disease, etc. However, due to miRNA’ s small size (approximately 22nt), there is no robust method for directly sequencing miRNA. All the contemporary sequencing strategies suffer from the same drawbacks that are inherent with sequencing by synthesis (SBS), namely you are measuring the creation of a copy of your target molecule. As a result, despite some evidence that miRNA modifications are present, they have gone largely ignored due to the difficulty in studying them. To overcome these technique’s shortcomings, the use of direct RNA sequencing (DRS), with the Oxford Nanopore Technologies (ONT) sequencing platform, is utilized.
[0025] ONT’s sequencing platform works by measuring the ionic current as a nucleic acid passes through a biological nanopore, which in DRS happens in the 3’ to 5’ direction. The RNA passes through in a controlled, and readable, manner by being adapted with the RMX adapter, which has a protein slowing down the RNA molecule’s passage through the pore. Each base has a different ionic current signature as it passes through the nanopore. A basecaller takes this signal and identifies patterns in the ionic current data indicative of a nucleic acid sequence using a pre-trained Recurrent Neural Network (RNN) model. The ionic current data is variable between sequencing events, and significant normalization must occur before basecalling. To make the base calling more accurate, the basecaller normalizes the entire signal to get a better baseline from which to judge the spikes and dips in the ionic current. Therefore, increased signal length, up to a point, has a direct benefit in the ability of the basecaller to normalize the signal. This can have a marked impact on basecalling accuracy. Modification detection specifically benefits from signal normalization since each modification is almost like adding another base that must be identified by its ionic signature. To accomplish this, an additional analysis pipeline is often required. Once most of the molecule passes through the pore, the protein slowing down its progress detaches, which makes the terminal 12-15nt at the 5’ end uninterpretable by the basecaller. The present disclosure relates to nucleic acid analysis and sequencing technologies. More particularly, it concerns compositions and methods for capture, enrichment, adaptation, and direct sequencing of short RNA molecules, including microRNAs (miRNAs), using nanopore sequencing platforms. The disclosure further relates to computational models and signal processing frameworks for classification, segmentation, alignment, and modification detection from ionic current data produced by nanopore sequencing.
[0026] Directly sequencing short RNA strands using Oxford Nanopore Technology’s (ONT) is difficult as anywhere between 12-20nt of the terminal bases are lost. Disclosed herein is the targeting of specific miRNA, short single stranded RNAs that are approximately 22nt, using a bait with an overhang on either side. These overhangs allow us to attach two additional synthetic oligonucleotides such that the final molecule is elongated by 145nt.
[0027] This disclosure provides the following advantages:
[0028] 1) Allows for direct sequencing of native short RNA molecules.
[0029] 2) Allows for modification detection in short RNA molecules.
[0030] 3) Not achievable with any other existing strategy.
[0031] 4) A tool to discover new biology.
[0032] This disclosure provides non-invasive and early disease detection through identifying elevated levels of specific miRNAs. It will also provide further research on miRNA or other short RNAs.
[0033] This technology is a new way to read very small RNA molecules, e.g., microRNAs (miRNAs), directly and accurately using nanopore sequencing by first making them longer so the sequencer can see the whole thing. : It uses a short “bait” strand that grabs a specific miRNA and two synthetic “adapters” that attach to it. These pieces make the miRNA much longer so a nanopore sequencer can read it end-to-end. A device like the Oxford Nanopore device senses changes in electrical current as the RNA passes through a tiny pore. Each letter (and certain chemical modifications) creates a distinct signal.
[0034] Traditional methods often miss part of such short molecules or rely on copying them first, which can lose information. This method captures the entire miRNA sequence. It can also detect natural “decorations” on miRNAs that may affect how genes are controlled but are hard to study with older methods. Because specific miRNAs change in diseases like cancer or Alzheimer’s, directly measuring and characterizing them can support earlier and more precise detection.
[0035] A designed “bait” strand pairs with the target miRNA, pulling it out of a complex sample. Two adapters are ligated to the miRNA, one on each end, making the molecule long enough for stable, accurate reading. A built-in desthiobiotin tag lets the lab pull out correctly assembled molecules so most sequenced reads are useful. The elongated RNA runs through a nanopore that measures electrical current changes, creating a signal trace. Software and machine learning convert the signal into sequence, segment the parts (adapters vs. miRNA), and detect modifications.
[0036] This technology enables accurate identification and counting of specific miRNAs from real biological samples. It can detect base modifications on miRNAs directly from the signal. It can also enable panel-based profiling of many miRNAs at once with high sensitivity. This is a potential route to non-invasive, earlier disease detection and deeper understanding of gene regulation.
[0037] Disclosed is a method of preparing a sequencing library for a target short RNA, comprising: hybridizing a hairpin to at least a portion of the target short RNA and a poly A; ligating the poly A to the 3 ' end of the target short RNA; ligating the target short RNA to the 3' end of the hairpin to form a cassette comprising the poly A, the target short RNA, and the hairpin.
[0038] Definitions
[0039] For convenience, certain terms employed in the specification, examples, and appended claims are collected here.
[0040] As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0041] The term “about” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the values measured or determined, ie., the limitations of the measurement system. Where the terms “about” or “approximately” are used in the context of compositions containing amounts of ingredients or conditions such as temperature, these values include the stated value with a variation of 0-10% around the value (X ± 10%).
[0042] The terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are inclusive in a manner similar to the term “comprising.” The term “consisting” and the grammatical variations of consist encompass embodiments with only the listed elements and excluding any other elements. The phrases “consisting essentially of’ or “consists essentially of’ encompass embodiments containing the specified materials or steps and those including materials and steps that do not materially affect the basic and novel characteristic(s) of the embodiments. Ranges are stated in shorthand to avoid having to set out at length and describe each and every value within the range. Therefore, when ranges are stated for a value, any appropriate value within the range can be selected, and these values include the upper value and the lower value of the range. For example, a range of two to thirty represents the terminal values of two and thirty, as well as the intermediate values between two to thirty, and all intermediate ranges encompassed within two to thirty, such as two to five, two to eight, two to ten, etc.
[0043] The term “preventing” is art-recognized, and when used in relation to a condition is well understood in the art, and includes administration of a composition which reduces the frequency of, or delays the onset of, symptoms of a medical condition in a subject relative to a subject which does not receive the composition. Thus, prevention of cancer includes, for example, reducing the incidence of cancer in a population of patients receiving a prophylactic treatment relative to an untreated control population, and / or delaying the onset of cancer in a treated population versus an untreated control population, e.g., by a statistically and / or clinically significant amount.
[0044] The term “ subject ' as used herein refers to a living mammal and may be interchangeably used with the term “patient”. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice and guinea pigs, and the like. The term does not denote a particular age or gender.
[0045] The term “therapeutically effective amount” of a compound with respect to the subject method of treatment refers to an amount of the compound(s) in a preparation which, when administered as part of a desired dosage regimen (to a mammal, preferably a human) alleviates a symptom, ameliorates a condition, or slows the onset of disease conditions according to clinically acceptable standards for the disorder or condition to be treated or the cosmetic purpose, e.g., at a reasonable benefit / risk ratio applicable to any medical treatment. A therapeutically effective amount herein may vary according to factors such as the disease state, age, sex, and weight of the patient, and the ability of the antibody to elicit a desired response in the individual.
[0046] As used herein, the term “treating or “treatment” includes reducing, arresting, or reversing the symptoms, clinical signs, or underlying pathology of a condition to stabilize or improve a subject’s condition or to reduce the likelihood that the subject’s condition will worsen as much as if the subject did not receive the treatment.
[0047] The term "complementary" and "complementarity" are interchangeable and refer to the ability of polynucleotides to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in antiparallel polynucleotide strands or regions. Complementary polynucleotide strands or regions can base pair in the Watson-Crick manner (e.g., A to T, A to U, C to G). 100% (or total) complementary refers to the situation in which each nucleotide unit of one polynucleotide strand or region can hydrogen bond with each nucleotide unit of a second polynucleotide strand or region. Less than perfect (or partial) complementarity refers to the situation in which some, but not all, nucleotide units of two strands or two regions can hydrogen bond with each other and can be expressed as a percentage.
[0048] The term “hybridization” is used to refer to the structure formed by 2 independent strands of RNA that form a double stranded structure via base pairings from one strand to the other. These base pairs are considered to be G-C, A-U, and G-U. (A - Adenine, C - Cytosine, G - Guanine, U - Uracil). As in the case of complementarity, hybridization can be total or partial.
[0049] The term “oligonucleotide” refers to RNA, DNA, or RNA: DNA oligonucleotides. The term “nucleotide” can refer to both, ribonucleotide or deoxyribonucleotide, unless otherwise explained.
[0050] “Short RNA” means a single-stranded RNA of length 10-60 nucleotides, including but not limited to miRNAs, piRNAs, siRNAs, and fragments thereof.
[0051] “miRNA” refers to mature microRNA molecules typically 18-30 nucleotides, including both 5p and 3p strands and sequence variants (isomiRs).
[0052] “Adapter,” “Enrichment Adapter (EA),” and “Nanopore Ligation Adapter (NLA)” refer to synthetic nucleic acid constructs ligated to the 5' and / or 3' ends of a short RNA to increase overall length, facilitate sequencing, and enable enrichment.
[0053] “Bait” or “Corresponding DNA / RNA Bait (CDB)” refers to a synthetic oligonucleotide complementary to at least a portion of a target short RNA, configured to hybridize and direct ligation within the cassette.
[0054] “Desthiobiotin” refers to a reversible affinity tag for streptavidin-based capture and elution used to enrich desired constructs.
[0055] “Ionic current signal” means the time-series current measurements generated as a nucleic acid translocates through a nanopore. “Segmentation” means assigning portions of ionic current to specific components of the cassette (e.g., adapters, target short RNA).
[0056] “Classification” means predicting the class identity (e.g., miRNA species) from ionic current.
[0057] “Base modification” means a covalent chemical alteration to a nucleobase (e.g., m6A, pseudouridine), phosphorylation, or other chemical changes detectible via altered ionic current.
[0058] “Short RNA” means a single-stranded RNA of length 10-60 nucleotides, including but not limited to miRNAs, piRNAs, siRNAs, and fragments thereof.
[0059] “miRNA” refers to mature microRNA molecules typically 18-30 nucleotides in length, including both 5p and 3p strands and sequence variants (isomiRs).
[0060] A “hairpin” is a secondary structure making a stem-loop, formed when a single RNA or DNA strand folds back so that two complementary regions base-pair to make a double-stranded stem capped by an unpaired loop. A hairpin comprises a largely Watson-Crick base-paired stem and an apical loop of unpaired (or non-canonical) nucleotides. A “loop” refers to the unpaired segment of nucleotides at the apex of a stem-loop (hairpin) structure.
[0061] “Poly(A)” usually refers to the poly(A) tail: a stretch of adenine nucleotides. It may comprise 3-50 adenine nucleotides.
[0062] An “RNA ligation adapter” is a short, synthetic oligonucleotide that is covalently attached to an end of RNA molecules during next-generation sequencing (NGS) library preparation. By adding known sequences to otherwise unknown RNA ends, adapters provide the handles required for reverse transcription, PCR amplification, sample indexing, and attachment to the sequencing platform.
[0063] Methods
[0064] The sequencing can be performed by any direct sequencing method that comprises a nanopore, for instance Oxford Nanopore technologies. The nanopore direct sequencing and the materials and protocols to perform it are known in the art. For instance, in US Patent Number 6,015,714. In some embodiments, the oligonucleotide adapter configured to perform nanopore direct sequencing is a double-stranded sequencing adapter DNA oligonucleotide with a helicase protein bound to one of the strands and having the complementary strand, a first DNA adapter oligonucleotide hybridization region. In some embodiments, the nanopore direct sequencing comprises a membrane, said membrane can be either solid-state or biological membranes.
[0065] Any known nanopore direct sequencing method or product can be used, for instance the one disclosed in US Patent Number 6,015,714 or US6,362,002.
[0066] The analysis or performing algorithm used can be any commercial one known by a skilled of many performing algorithms known in the art suitable for nanopore direct RNA sequencing. The first step is extracting the reads. This step can be done by commercial software, for instance MinKNOW or any software configured to analyze the sequencing results of the nanopore direct sequencing. Next step of the analysis is the base calling, which can be done by a skilled person using any of several known performing algorithms in the field, such as Guppy or Bonito. Last step of the analysis is mapping, which can be done by several known performing algorithms. For example, Minimap2 or BWA which is a versatile sequence alignment program that aligns nucleic acid sequences against a large reference database. In some embodiments, the performing algorithm is configured to capture (and sequence) more miRNA in a quantitative way.
[0067] Disclosed herein are methods of preparing a sequencing library for a target short RNA, comprising: hybridizing a hairpin to at least a portion of the target short RNA and a poly A; ligating the poly A to the 3 ' end of the target short RNA; ligating the target short RNA to the 3' end of the hairpin to form a cassette comprising the poly A, the target short RNA, and the hairpin.
[0068] Disclosed herein are methods for preparing a sequencing library for a target short RNA, comprising hybridizing a Corresponding DNA or RNA Bait (CDB) to at least a portion of a target short RNA; ligating a 3' Nanopore Ligation Adapter (NLA) to the 3' end of the target short RNA; ligating a 5' Enrichment Adapter (EA) to the 5' end of the target short RNA to form a cassette comprising the NLA, the target short RNA, and the EA.
[0069] Disclosed is a method of preparing a sequencing library for a target short RNA, comprising: hybridizing a hairpin to at least a portion of the target short RNA and a poly A; ligating the poly A to the 3 ' end of the target short RNA; ligating the target short RNA to the 3' end of the hairpin to form a cassette comprising the poly A, the target short RNA, and the hairpin.
[0070] Disclosed is a method of preparing a sequencing library for a target short RNA. The method may comprise: hybridizing a hairpin nucleotide sequence to a poly A oligonucleotide via a first hybridization region and hybridizing the hairpin nucleotide sequence to at least a portion of the target short RNA via a second hybridization region; ligating the poly A oligonucleotide to the 3' end of the target short RNA; and ligating the target short RNA to the 3' end of the hairpin nucleotide sequence, thereby forming a cassette comprising the poly A oligonucleotide, the target short RNA, and the hairpin nucleotide sequence.
[0071] In some embodiments, the method further comprises sequencing the cassette using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the hairpin nucleotide sequence comprises in 5' to 3' order: 1) the first hybridization region, 2) the second hybridization region complementary to at least a portion of the target short RNA, 3) a third hybridization region, 4) a loop, and 5) a fourth hybridization region complementary to the third hybridization region. In some embodiments, the hairpin nucleotide sequence is 30 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is RNA.
[0072] In some embodiments, the poly A oligonucleotide comprises in 5' to 3' order: 1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A oligonucleotide is 10 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the poly A is RNA. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof.
[0073] In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA. In some embodiments, the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter. In some embodiments, the target short RNA comprises a base modification; and the method further comprises detecting the base modification.
[0074] Disclosed herein is a method for preparing a sequencing library for a target short RNA. The method may comprises: hybridizing a Corresponding DNA or RNA Bait (CDB) to at least a portion of a target short RNA via a second hybridization region; ligating a 3' Nanopore Ligation Adapter (NLA) to the 3' end of the target short RNA; and ligating a 5' Enrichment Adapter (EA) to the 5' end of the target short RNA, thereby forming a cassette comprising the NLA, the target short RNA, and the EA.
[0075] In some embodiments, the method further comprises sequencing the cassette using a sequencing platform. In some embodiments, the sequencing platform is a nanopore sequencing platform. In some embodiments, the CDB comprises in 5' to 3' order: 1) a first hybridization region, 2) the second hybridization region complementary to at least a portion of the target short RNA, and 3) a third hybridization region. In some embodiments, the CDB is 20 to 100 nucleotides in length. In some embodiments, the CDB is 20 to 40 nucleotides in length, 40 to 50 nucleotides in length, 50 to 60 nucleotides in length, 60 to 70 nucleotides in length, 70 to 80 nucleotides in length, 80 to 90 nucleotides in length, or 90 to 100 nucleotides in length. In some embodiments, the CDB comprises a randomized region capable of capturing a plurality of short RNA species. In some embodiments, the CDB is RNA or DNA.
[0076] In some embodiments, the NLA comprises in 5' to 3' order: 1) a fourth hybridization region complementary to the first hybridization region, and 2) a fifth hybridization complementary to an RNA ligation adapter. In some embodiments, the NLA is 10 to 50 nucleotides in length. In some embodiments, the NLA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the NLA is RNA or DNA. In some embodiments, the EA comprises in 5' to 3' order: 1) a sixth hybridization region, 2) a loop, 3) a seventh hybridization region complementary to the sixth hybridization region, and 4) an eighth hybridization region complementary to the third hybridization region.
[0077] In some embodiments, the EA is 10 to 50 nucleotides in length. In some embodiments, the EA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the EA is RNA or DNA. In some embodiments, the EA comprises a desthiobiotin moiety. In some embodiments, the method further comprises enriching the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions.
[0078] In some embodiments, the NLA and the EA together add at least 100 nucleotides of additional sequence to the cassette. In some embodiments, ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof. In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA. In some embodiments, the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter. In some embodiments, the target short RNA comprises a base modification and the method further comprises detecting the modification.
[0079] In some embodiments, the method further comprising basecalling the ionic current signal to produce basecalls corresponding to the target short RNA, wherein the cassette length enables improved signal normalization and basecalling at the 5' terminus of the target short RNA. In some embodiments, the method further comprising segmenting the ionic current signal into regions corresponding to the NLA, the EA, and the target short RNA using a trained machine learning model. In some embodiments, the method further comprising classifying the target short RNA species based on the segmented ionic current using the machine learning model. In some embodiments, the method further comprising aligning the sequence corresponding to the segmented target short RNA region to a reference sequence using an optimal edit-distance algorithm. In some embodiments, the method detects at least 2 attomoles of a target miRNA from 1-10 pg of total RNA input.
[0080] In some embodiments, the EA and / or NLA comprise RNA, DNA, or chimeric RNA:DNA, and ligation conditions are selected to achieve directional assembly with at least 50% ligation efficiency. In some embodiments, sequencing and analysis are performed to detect a base modification within the target short RNA with a confidence threshold based on a modification-aware signal model.
[0081] Disclosed herein is a method of quantifying a target miRNA in a biological sample, comprising: forming a cassette as in claim 1 with a bait complementary to the target miRNA; enriching the cassette; sequencing using nanopore direct RNA sequencing; classifying or aligning the target miRNA reads; and quantifying the abundance of the target miRNA based on read counts or classification outputs. In some embodiments, the quantification is performed across a panel of miRNAs to generate a profile of miRNA expression. In some embodiments, the method is used in detecting or monitoring a disease state characterized by altered miRNA expression.
[0082] Because the method sequences native short RNAs and produces elongated signal suitable for normalization, ionic current features associated with base modifications are more readily detected. In certain embodiments, sample preparation includes known modified and unmodified controls to calibrate modification calling. Analytical models detect shifts in current amplitude, dwell time, and event structure indicative of modifications such as methylations, pseudouridine, and terminal phosphorylations.
[0083] In embodiments tested, the cassette method provides full-length products with high alignment identity and classification accuracy. Limit of detection for certain miRNA species can reach the attomole range from microgram quantities of input total RNA, enabling detection of biologically relevant expression levels across tissues and disease states. The enrichment strategy ensures high sequencing throughput efficiency by increasing the fraction of reads containing target short RNAs rather than adapter-only species. Kits
[0084] Disclosed herein is a kit comprising a hairpin nucleotide sequence and a poly A oligonucleotide. In some embodiments, the hairpin nucleotide sequence comprises in 5' to 3' order: 1) a first hybridization region, 2) a second hybridization region complementary to at least a portion of a target short RNA, 3) a third hybridization region, 4) a loop, and 5) a fourth hybridization region complementary to the third hybridization region. In some embodiments, the hairpin nucleotide sequence is 30 to 200 nucleotides in length. In some embodiments, the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length.
[0085] In some embodiments, the hairpin is RNA. In some embodiments, the poly A oligonucleotide comprises in 5' to 3' order: 1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides. In some embodiments, the poly A oligonucleotide is 10 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length. In some embodiments, the poly A oligonucleotide is RNA. In some embodiments, the target short RNA is 10-60 nucleotides in length. In some embodiments, the target short RNA is a microRNA.
[0086] In some embodiments, the kit comprises streptavidin-coated beads and buffers for reversible capture and elution. In some embodiments, the adapters and baits are provided with RNase-free HPLC purification and are supplied in concentrations suitable for formation of cassettes under ligation conditions specified in the instructions.
[0087] Computer-Implemented Methods
[0088] Disclosed herein is a computer-implemented method for analyzing nanopore ionic current data from a cassette comprising adapters and a target short RNA, comprising: receiving ionic current data; processing the data with a transformer-based model to (i) segment the data into regions corresponding to adapters and the target short RNA, and (ii) classify the target short RNA species; and outputting at least one of a species classification, an alignment to a reference sequence, or a modification-aware analysis of the target short RNA. In some embodiments, the model is trained on in vitro transcribed RNA and fine-tuned using ionic current from biological samples. In some embodiments, the computer-implemented method further comprising generating an optimal alignment of the segmented target short RNA signal to a predicted reference sequence using a modified edit-distance dynamic programming algorithm. In some embodiments, the segmentation and classification are performed without reliance on basecalling, and the output comprises a species label and confidence score for the target short RNA.
[0089] EXAMPLES
[0090] The invention now being generally described, it will be more readily understood by reference to the following examples, which are included merely for purposes of illustration of certain aspects and embodiments of the present invention, and are not intended to limit the invention.
[0091] Example 1: Exemplified miRNA sequencing
[0092] With miRNA being approximately 22nt, losing even 12nt is usually over half of the molecule. Our strategy substantially elongates the 5’ end so that the entire miRNA is read by adapting the 5’ end of the miRNA. This has the added benefit of making the prepared library long enough for sequence normalization. Lastly, our method adds a section that is complementary to the RMX sequencing adapter so that it is easy to plug it into one of ONT’s DRS protocols. A universal poly A and a hairpin with a specific miRNA bait and a complementary section to the universal poly A oligonucleotide are added to miRNA-containing total RNA (Fig. 1A). The hairpin behaves as a bait and captures the specific miRNA it hybridizes to, and, along with the universal poly A adapter, is ligated together using RNA Ligase 2 (RNL2). This creates sufficient length (> 145 nucleotides) to utilize preexisting bioinformatic tools to extract information directly from the raw ionic current signal. For example, you can use the following hairpin bait, along with the universal poly A adapter, to capture miRNA 455-3p (5’ - GCAGUCCAUGGGCAUAUACAC - 3’) (SEQ ID NO: 1), a 21nt Alzheimer’s Disease associated miRNA:
[0093] miRNA 455-3p hairpin bait: 5’-GAGCUAAUAUGUGUAUAUGCCCAUGGACUGCCGUGACAUAGGAGAGAGACUA UGUCACG-3’ (SEQ ID NO: 2)
[0094] Universal poly A adapter: 5’- / 5Phos / AUAUUAGCUCAAAAAAAAAAAAAAA-3’ (SEQ ID NO: 3)
[0095] Here, we construct the miRNA 455-3p hairpin by sandwiching the miRNA 455-3p complement with 5’-GAGCUAAUAU-3’ (SEQ ID NO: 4) and 5’-CGUGACAUAGGAGAGAGACUAUGUCACG-3’ (SEQ ID NO: 5). Additionally, all synthetic oligonucleotides may be ordered with RNAse free HPLC purification. The system fits together like a cassette (Fig. 1A). Using a synthetic miRNA 455-3p, visualized on a 6% TBE-UREA gel (Thermo Fisher EC68652BOX) we can see both the full construct with the miRNA bait, synthetic miRNA 455-3p, and the universal poly A oligonucleotide (Fig. IB). Additionally, we sequenced the double ligation product on the MinlON and aligned it to our expected full-length construct, confirming our ability to sequence synthetic miRNA 455-3p (Fig. 1C). Using Integrated DNA Technologies (IDT) as the benchmark for technical limitations in RNA oligonucleotide ordering, the maximum length RNA oligonucleotide we could order is 60nt. As a result, with this design, the maximum length our target could be if it perfectly hybridizes to the bait is 22nt. If we intentionally let a bubble form in the center of the target, we could potentially capture a larger target.
[0096] Example 2: Exemplified longer miRNA sequencing
[0097] For a more modular system, we split the miRNA hairpin into 2 different synthetic oligonucleotides, a universal hairpin and a specific miRNA bait (Fig. 2A). Synthesizing longer RNA oligonucleotides is expensive and time consuming, especially with the necessary purification. Using non-ultramer length RNA oligonucleotides from IDT as a benchmark, the maximum length RNA oligonucleotide we could order is 60nt. Therefore, in this all RNA system, the largest size the target could be is 35nt. Similarly to the non-modular system, if we intentionally let a bubble form we could aim for longer targets. To demonstrate the no bubble modular bait design, for miRNA455-3p, you use the following oligonucleotides:
[0098] Universal hairpin: 5’- / 5Phos / AUGGCAUUAAGAUCAGUUGCA / ideSBioTEG / GAGAGAGAUGCAACUGAUC UUAAUGCCAUCACUUAGCUC-3’ (SEQ IDNO: 6)
[0099] miRNA 455-3p specific bait: 5’-CACAACCUACCCUCGGUGUAUAUGCCCAUGGACUGCGAGCUAAGUG - 3’ (SEQ ID NO: 7)
[0100] Universal poly A oligonucleotide: 5’- / 5Phos / CGAGGGUAGGAGGAAGGAGGUAAAUACAAAUAGGUAGGGAGAAGAAAA AAAAAAAAAAAA-3 ’ (SEQ ID NO: 8)
[0101] The bait is constructed by sandwiching -5’CACAACCUACCCUCG3’- (SEQ ID NO: 9) and -5’GAGCUAAGUG3’- (SEQ ID NO: 10) around a portion that complements the target. Additionally, all synthetic oligonucleotides may be ordered with RNAse free HPLC purification. Other than this design being modular, it also includes a purification tag, in this case a desthiobiotin modification, on the universal hairpin (Fig. 2A). This results in the larger majority of the sequenced product being full-length. Consequently, sequencing is more efficient as a high majority of the reads has the relevant target miRNAs rather than just the known universal polyA adapter. Iterating on this design, we have separated the 5’ adapter into a 5’ adapter and bait (Fig. 2A). This strategy leverages a panel of complementary ‘bait’ molecules to capture miRNA of interest - Corresponding DNA Bait (CDB). The bait and miRNA are annealed for sequencing in conjunction with both 5' (Enrichment Adapter (EA)) and 3' (Nanopore Ligation Adapter (NLA)) adapters to provide sufficient length to the miRNA for accurate base calling and alignment (Fig. 2A). The NLA has a total length of 60 nucleotides with the 25 NLA’ S most nucleotides being DNA to facilitate nanopore sequencing adaptation. The EA is 60 nucleotides bringing the total length added to each miRNA to 120 nucleotides. The synthetic nature of the adapters allows for enrichment of the 5’ adapter through the use of a desthiobiotin and Cl bead clean-up. Our data suggests that this technique has a limit of detection of 2amol for a single miRNA (for example, miR-125b-5p) from an input of 5 pg of total RNA. This limit of detection allows us to capture biologically relevant quantities of miRNA which are estimated to span multiple orders of magnitude and can be subject to disease and tissue specific patterns of expression.
[0102] Additionally, the cassette-based strategy allows for the selective sequencing of target miRNAs and minimizes the impact of off target RNAs on sequencing throughput. We have demonstrated that the addition of our custom adapters provides sufficient length for successful basecalling and alignment (Fig. 2B) and provides sufficient length for ionic current analysis, a hallmark of nanopore sequencing and a promising area of RNA research (Fig. 2C).
[0103] We have tested a single ligation, combining T4 DNA Ligase with RNL2 during the first ligation, having the bait be DNA instead of RNA (to order from IDT: 5’ / 5Phos / ATGGCATTAAGATCAGTTGCA / ideSBioTEG / GAGAGAGArUrGrCrArArCrUrGr ArUrCrUrUrArArUrGrCrCrArUrCrArCrUrUrArGrCrUrC 3’ (SEQ ID NO: 11)), ordering individually purified baits, using Integrated DNA Technologies (IDT) oligo pools for baits, doing SPRI bead-based clean ups instead of the 2nd and 3rd desthiobiotin enrichments, having an excess of baits in comparison to adapters, capturing and sequencing a synthetic modified miRNA, sequencing miRNA from a cell line, and testing differently extracted total RNA and size selected total RNA. We have also tested a scenario where there is no specific target and instead the bait has 22 random oligonucleotides. This lets us observe all 22nt RNAs in the cell, potentially creating a relatively easy, cheap, and efficient way to discover new miRNAs.
[0104] The protocol for the modular hairpin system has a couple additional steps to the previous version. First, the modular hairpin pool is created by having lOuM of the hairpin and equimolar proportions of each specific bait adding up to lOuM in annealing buffer (lOmM Tris-HCl pH 8.0, 50mM NaCl). This is denatured at 95°C for 3 minutes, cooled on ice for 10 minutes, then there is a Streptavidin Cl bead clean up (Thermo Fisher 65001), eluting in the same amount of volume as was there previously. This eluted product will be referred to as the pooled modular hairpin and we assume a lOuM concentration of the created bait hairpin pool. The pooled modular hairpin is ligated at 37°C for Ihr with IX Quick Ligation Buffer (NEB B6058S) and 500 U / mL T4 RNA Ligase 2 (NEB M0239S). Following another identical Streptavidin Cl bead clean up (Thermo Fisher 65001), again eluting in the same initial volume of water, this can be stored in the -80°C freezer for up to 6 months. When prepared for the next step, in a 20uL reaction, have 5uM pool hairpin bait, 20uM universal poly A adapter, IX Quick Ligation Buffer (NEB B6058S), and 500 U / mL T4 RNA Ligase 2 (NEB M0239S), filling the remaining volume with total RNA. Allow to ligate at 37°C for 1 hour and follow with another Streptavidin Cl bead clean up (Thermo Fisher 65001). At this point the RNA is ready for a standard ONT DRS library preparation followed by sequencing.
[0105] Example 3: Sequencing model
[0106] Our group has developed a deep learning framework for segmentation and classification of tRNA directly from Nanopore ionic current signal. This framework leverages the rich ionic current information produced from nanopore sequencing and utilizes the common sequencing adapters we have designed (Fig. 3A) to segment the signal into each adapter and the variable tRNA region. Our miRNA sequencing strategy uses a similar strategy with a common sequencing adapter with a variable miRNA region, meaning the underlying strategy is well suited for miRNA as well as tRNA. Our model is a transformerbased architecture capable of providing purely ionic current-based classification of tRNA and can be extended to miRNA (Fig. 3C). The model outputs both a classification and a sequence-to-sequence segmented ionic current output, allowing us to attribute specific regions of ionic current to specific components of our sequencing construct (Fig. 3D). The median classification accuracy on a test set composed of E. coli tRNA is 98.86% (Fig. 3E).
[0107] With the sequence-to-sequence output of the model we are able to segment the ionic current corresponding to each of our custom tRNA adapters (5' and 3') as well as isolate the ionic current that corresponds to the tRNA. Using the signal-to-sequence move table provided by ONT's basecaller, we can isolate the exact sequence of the tRNA (Fig. 3F) and align it to the reference sequence of the classification provided by our model (Fig. 3G). This strategy allows us to avoid the difficulties of aligning short noisy RNA sequences to multiple references with minimal differences, instead relying on our high accuracy classifications to determine which tRNA reference sequence to use. By using a modified Wagner-Fischer pairwise alignment algorithm we can guarantee an optimal alignment for each tRNA read to the predicted class of tRNA (Fig. 3H).
[0108] We have developed a method for training these models that utilizes in vitro transcribed (IVT) RNA as a baseline and polishes the model using nanopore signal from biological molecules. This has led to a highly flexible model training paradigm that does not rely on the costly purification of biological molecules. We apply the same strategy to miRNA, using the panels as a substitute for the tRNA populations we have previously made models for. For each miRNA panel, we generate a library of DNA templates with T7 primers to make a panel-specific IVT miRNA training set. While some model adjustments may be required for the difference in lengths between tRNA and miRNA, it also means that the training paradigm is more computationally tractable. With completed models capable of slicing out miRNA-specific sections of sequence, we are able to overcome one of the primary hurdles for alignment; short noisy RNA sequences being compared to references with minimal differences. Utilizing the same modified Needleman-Wunsch algorithm we used for tRNA sequencing, which obtained an alignment identity over 94% for a training dataset, we are able to construct robust alignments of miRNA. Beyond alignment, the classification capabilities of our models has a significant impact on our ability to quantify biological miRNA species, providing a direct classification without the intermediate steps of basecalling and alignment.
[0109] Altogether, this strategy makes for a good alternative to the current SBS RNA-seq and RT-qPCR methods being used to study miRNA. It functions by ligating a 5’ end adapter to short RNAs so that they can be adequately sequenced and analyzed. More broadly, this enables us to sequence multiple, specific, short RNAs that were not possible to directly observe prior to this system.
[0110] INCORPORATION BY REFERENCE
[0111] Each of the patents, published patent applications, and non-patent references cited herein are hereby incorporated by reference in their entirety. EQUIVALENTS
[0112] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
We claim:
1. A method of preparing a sequencing library for a target short RNA, comprising:hybridizing a hairpin nucleotide sequence to a poly A oligonucleotide via a first hybridization region and hybridizing the hairpin nucleotide sequence to at least a portion of the target short RNA via a second hybridization region;ligating the poly A oligonucleotide to the 3' end of the target short RNA; and ligating the target short RNA to the 3' end of the hairpin nucleotide sequence, thereby forming a cassette comprising the poly A oligonucleotide, the target short RNA, and the hairpin nucleotide sequence.
2. The method of claim 1, further comprising sequencing the cassette using a sequencing platform.
3. The method of claim 2, wherein the sequencing platform is a nanopore sequencing platform.
4. The method of any one of claims 1-3, wherein the hairpin nucleotide sequence comprises in 5' to 3' order:1) the first hybridization region,2) the second hybridization region complementary to at least a portion of the target short RNA,3) a third hybridization region,4) a loop, and5) a fourth hybridization region complementary to the third hybridization region.
5. The method of any one of claims 1-4, wherein the hairpin nucleotide sequence is 30 to 200 nucleotides in length.
6. The method of any one of claims 1-5, wherein the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length.
7. The method of any one of claims 1-6, wherein the hairpin nucleotide sequence is RNA.
8. The method of any one of claims 1-7, wherein the poly A oligonucleotide comprises in 5' to 3' order:1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides.
9. The method of any one of claims 1-8, wherein the poly A oligonucleotide is 10 to 50 nucleotides in length.
10. The method of any one of claims 1-9, wherein the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length.
11. The method of any one of claims 1-10, wherein the poly A is RNA.
12. The method of any one of claims 1-11, wherein ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof.
13. The method of any one of claims 1-12, wherein the target short RNA is 10-60 nucleotides in length.
14. The method of any one of claims 1-13, wherein the target short RNA is a microRNA.
15. The method of any one of claims 2-14, wherein the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter.
16. The method of any one of claims 1-14, wherein the target short RNA comprises a base modification; and the method further comprises detecting the base modification.
17. A method for preparing a sequencing library for a target short RNA, comprising: hybridizing a Corresponding DNA or RNA Bait (CDB) to at least a portion of a target short RNA via a second hybridization region;ligating a 3' Nanopore Ligation Adapter (NLA) to the 3' end of the target short RNA; andligating a 5' Enrichment Adapter (EA) to the 5' end of the target short RNA, thereby forming a cassette comprising the NLA, the target short RNA, and the EA.
18. The method of claim 17, further comprising sequencing the cassette using a sequencing platform.
19. The method of claim 18, wherein the sequencing platform is a nanopore sequencing platform.
20. The method of any one of claims 17-19, wherein the CDB comprises in 5' to 3' order:1) a first hybridization region,2) the second hybridization region complementary to at least a portion of the target short RNA, and3) a third hybridization region.
21. The method of any one of claims 17-20, wherein the CDB is 20 to 100 nucleotides in length.
22. The method of any one of claims 17-21, wherein the CDB is 20 to 40 nucleotides in length, 40 to 50 nucleotides in length, 50 to 60 nucleotides in length, 60 to 70 nucleotides in length, 70 to 80 nucleotides in length, 80 to 90 nucleotides in length, or 90 to 100 nucleotides in length.
23. The method of any one of claims 17-22, wherein the CDB comprises a randomized region capable of capturing a plurality of short RNA species.
24. The method of any one of claims 17-23, wherein the CDB is RNA or DNA.
25. The method of any one of claims 17-24, wherein the NLA comprises in 5' to 3' order:1) a fourth hybridization region complementary to the first hybridization region, and 2) a fifth hybridization complementary to an RNA ligation adapter.
26. The method of any one of claims 17-25, wherein the NLA is 10 to 50 nucleotides in length.
27. The method of any one of claims 17-26, wherein the NLA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length.
28. The method of any one of claims 17-27, wherein the NLA is RNA or DNA.
29. The method of any one of claims 17-28, wherein the EA comprises in 5' to 3' order:1) a sixth hybridization region,2) a loop,3) a seventh hybridization region complementary to the sixth hybridization region, and4) an eighth hybridization region complementary to the third hybridization region.
30. The method of any one of claims 17-29, wherein the EA is 10 to 50 nucleotides in length.
31. The method of any one of claims 17-30, wherein the EA is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length.
32. The method of any one of claims 17-31, wherein the EA is RNA or DNA.
33. The method of any one of claims 17-32, wherein the EA comprises a desthiobiotin moiety.
34. The method of claim 33, wherein the method further comprises enriching the cassette by capturing the cassette on streptavidin-coated beads; and eluting the cassette under reversible conditions.
35. The method of any one of claims 17-34, wherein the NLA and the EA together add at least 100 nucleotides of additional sequence to the cassette.
36. The method of any one of claims 17-35, wherein ligating comprises using T4 RNA Ligase 2, T4 DNA ligase, or a combination thereof.
37. The method of any one of claims 17-36, wherein the target short RNA is 10-60 nucleotides in length.
38. The method of any one of claims 17-37, wherein the target short RNA is a microRNA.
39. The method of any one of claims 18-38, wherein the sequencing platform is configured for direct RNA sequencing using a motor associated sequencing adapter.
40. The method of any one of claims 17-39, wherein the target short RNA comprises a base modification and the method further comprises detecting the modification.
41. A kit comprising a hairpin nucleotide sequence and a poly A oligonucleotide.
42. The kit of claim 41, wherein the hairpin nucleotide sequence comprises in 5' to 3' order:1) a first hybridization region,2) a second hybridization region complementary to at least a portion of a target short RNA,3) a third hybridization region,4) a loop, and5) a fourth hybridization region complementary to the third hybridization region.
43. The kit of claim 41 or 42, wherein the hairpin nucleotide sequence is 30 to 200 nucleotides in length.
44. The kit of any one of claims 41-43, wherein the hairpin nucleotide sequence is 30 to 50 nucleotides in length, 50 to 100 nucleotides in length, 100 to 150 nucleotides in length, or 150 to 200 nucleotides in length.
45. The kit of any one of claims 41-44, wherein the hairpin is RNA.
46. The kit of any one of claims 41-45, wherein the poly A oligonucleotide comprises in 5' to 3' order:1) a fifth hybridization region complementary to the first hybridization region, and 2) a stretch of adenine nucleotides.
47. The kit of any one of claims 41-46, wherein the poly A oligonucleotide is 10 to 50 nucleotides in length.
48. The kit of any one of claims 41-47, wherein the poly A oligonucleotide is 10 to 20 nucleotides in length, 20 to 30 nucleotides in length, 30 to 40 nucleotides in length, or 40 to 50 nucleotides in length.
49. The kit of any one of claims 41-48, wherein the poly A oligonucleotide is RNA.
50. The kit of any one of claims 42-49, wherein the target short RNA is 10-60 nucleotides in length.
51. The kit of any one of claims 42-50, wherein the target short RNA is a microRNA.