Methods for generating cDNA libraries from RNA
The method for generating double-stranded cDNA libraries addresses long processing times by using ribonucleotide tailing and adapter ligation with T4 RNA ligases, enabling rapid, automated, and cost-effective cDNA library generation for biomarker detection.
Patent Information
- Application Number
- JP2025514645
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-12
- Filing Date
- 2023-09-11
- Publication Date
- 2025-09-11
AI Technical Summary
Current biological sample processing times for detecting biomarkers, pathogens, or genetic traits are long and often require operator presence, necessitating the development of more feasible methods adaptable to partially or fully automated systems and faster turnaround.
A method for generating double-stranded cDNA libraries involving steps such as obtaining RNA molecules, generating first-strand cDNAs with ribonucleotide tails, ligating single-stranded adapters, and amplifying cDNAs using a truncated form of T4 RNA ligase 2 or T4 RNA ligase 1, which are adaptable to automated systems and reduce processing time.
The method enables rapid cDNA library generation in under 8 hours, reduces reaction steps and liquid handling, and is adaptable to automation, improving efficiency and reducing costs while maintaining high sequencing accuracy.
Smart Images

Figure 2025530278000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims priority to U.S. Provisional Application No. 63 / 375,372, filed September 12, 2022, which is incorporated herein by reference in its entirety. [Background technology]
[0002] The disclosure herein relates to the field of molecular biology, for example, to methods and compositions for generating double-stranded cDNA libraries from RNA samples.
[0003] Current biological sample processing times for detecting biomarkers, pathogens, or genetic traits in clinical samples are long and often require operator presence to perform the reaction. Therefore, there is a need to develop more feasible methods for nucleic acid analytical reactions that are adaptable to partially or fully automated systems and faster turnaround. Summary of the Invention
[0004] Provided herein are methods, compositions, and systems that improve upon existing methods for processing and preserving DNA or RNA of interest from biological samples for molecular analysis.
[0005] The present specification provides a method for preparing a cDNA library, comprising: (A) obtaining a plurality of RNA molecules from a sample; (B) generating first-strand cDNAs complementary to the RNA molecules from the plurality of RNA molecules using random primers and single-stranded first adapters; (C) incorporating a ribonucleotide tail at the 3' end of the first-strand cDNA using a mixture of ribonucleotide triphosphate bases and terminal deoxynucleotidyl transferase; (D) generating first-strand cDNAs with single-stranded adapters at both ends by adding a single-stranded second adapter to the 3' end of the ribonucleotide tail using a ligase that prefers 3' ribonucleotide ends; and (E) amplifying the first-strand cDNAs to prepare a cDNA library. In some embodiments, the first-strand cDNAs are synthesized using random primers ligated to the first adapters. In some embodiments, the first adapters are single-stranded. In some embodiments, the random primers are synthesized together with the first adapters.
[0006] In some embodiments, the RNA is obtained from a sample, and the sample is a biological sample.
[0007] In some aspects, the biological sample is a fresh biological sample, a frozen biological sample, or a forensic sample.
[0008] In some embodiments, obtaining the plurality of RNA molecules comprises extracting total RNA from the sample.
[0009] In some embodiments, the first adaptor is a universal adaptor.
[0010] In some aspects, the method further comprises, after generating first-strand cDNA, removing unused primers and degrading dNTPs with an exonuclease and a phosphatase, rendering them unavailable as substrates for terminal deoxynucleotidyl transferase (TdT) in a subsequent step. In some embodiments, the endonuclease is exonuclease 1, and the phosphatase is shrimp alkaline phosphatase (rSAP). rSAP also removes all phosphate groups from the 5' ends of RNA in the sample, preventing the formation of chimeric artifacts that can occur when T4 RNA ligase 1 is used for ligation. This is not an issue with T4 RNA ligase 2 (T4RL-2), a truncated form that prefers only double-stranded RNA (dsRNA or duplex RNA), and therefore cannot use 5'-phosphorylated oligos as substrates.
[0011] In some embodiments, the step of incorporating a ribonucleotide tail comprises incorporating fewer than 10 ribonucleotides at the 3' end of the first strand cDNA by terminal deoxynucleotidyl transferase (TdT), whose preferred substrates are the 3'-DNA terminus and deoxynucleotide triphosphates (dNTPs), such that ribonucleotide triphosphates (rNTPs) are not preferred substrates and therefore fewer than 10 are added to the 3'-end of the cDNA strand.
[0012] In some embodiments, the ribonucleotide tail is incorporated using terminal deoxynucleotidyl transferase (TdT).
[0013] In some embodiments, adding the single-stranded second adaptor comprises adding or ligating the adaptor to the 3'-terminal ribonucleotide incorporated in step (C) above.
[0014] In some embodiments, amplifying comprises performing a polymerase chain reaction (PCR) using primers that anneal to the first adapter and the second adapter to generate double-stranded cDNA for the cDNA library.
[0015] In some embodiments, the amplified double-stranded cDNA library obtained from step (E) is an amplified first cDNA library containing target sequences and unwanted non-target sequences, wherein the method further comprises removing from the amplified first cDNA library a subset of the amplified first cDNA library containing unwanted non-target sequences.
[0016] In some embodiments, removing a subset of the amplified first cDNA library that contains unwanted non-target sequences from the amplified first cDNA library is performed using a nucleic acid-guided endonuclease.
[0017] In some embodiments, the nucleic acid-guided endonuclease comprises a CAS9 enzyme.
[0018] In some embodiments, the target nucleotide sequence comprises a pathogen sequence, and removing a subset of the amplified first cDNA library comprises removing non-pathogen host genomic nucleic acid.
[0019] In some embodiments, removing a subset of the amplified first cDNA library comprises removing contaminating human genomic nucleic acid.
[0020] In some aspects, the contaminating human genomic nucleic acid is ribosomal nucleic acid.
[0021] In some aspects, the contaminating human genomic nucleic acid is a repetitive nucleic acid sequence.
[0022] In some embodiments, the target nucleotide sequence comprises fetal nucleic acid, and removing the subset of the amplified first cDNA library comprises removing non-target contaminating maternal nucleic acid.
[0023] In some embodiments, the target nucleotide comprises a genomic polymorphism, and removing a subset of the amplified first cDNA library comprises removing sequences that include wild-type sequences.
[0024] In one aspect, provided herein is a system for automated amplification of a target nucleotide sequence from a biological sample containing a non-target nucleotide sequence, the system comprising: (a) a control panel operably connected to a programmable machine having (b) a liquid handler unit having an automated liquid aspirating and dispensing side and an optical head side, the control panel configured to receive input commands to operate the machine. [Brief explanation of the drawings]
[0025] Some understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings.
[0026] [Figure 1] FIG. 1 illustrates a schematic workflow starting from nucleic acid extraction to isolation, depletion, and amplification to obtain enriched cDNA. [Figure 2] Figure 1 depicts a schematic of the workflow showing the stepwise manipulation of cDNA synthesis and single-stranded cDNA, resulting in the amplification and generation of an amplified double-stranded cDNA library (where rN is a mixture of adenosine ribonucleotides, uridine ribonucleotides, guanosine ribonucleotides, and a portion of cytidine ribonucleotides, the four possible bases of ribonucleotides, incorporated by terminal deoxynucleotidyl transferase TdT). A barcode is added during the first PCR, and primers are used in the second PCR next to the added barcode from the first PCR. [Figure 3]Figure 1 shows an exemplary schematic of the workflow, detailing ribonucleotide tailing by terminal deoxynucleotidyl transferase (TdT) and adapter ligation, where the 3'-terminal oligoribonucleotide (RNA) tail is the preferred substrate for the ligase. The Ampure XP SPRI kit (commercially available) is used for the bead-based cleanup step. Steps 1 through 5 are performed sequentially without any isolation of the desired product between steps, and conditions were set to ensure that the previous step does not interfere with the subsequent step. Additional possibilities include using T4 RNA ligase 1 with a 5' phosphate oligo adapter or T4 RNA ligase 2 (a double-stranded RNA enzyme) with a double-stranded RNA adapter. [Figure 4A] Figure 1 depicts exemplary structures of an adenylated deoxy-oligonucleotide adapter, or App-DNA adapter sequence (DNA with diphosphate bonds, shown as black bars) to be used with T4 RNA ligase 2 truncated (and all its mutants, i.e., K227Q), or alternatively, a 5'-phosphate DNA adapter sequence to be used in conjunction with adenosine triphosphate (ATP) and T4 RNA ligase 1. The third is an RNA duplex adapter (RNA shown as red bars) with all four ribonucleotides present in the overhang that pair with the ribonucleotide tail of the cDNA and can be ligated by T4 RNA ligase 2 even in the presence of ATP. For all three enzymes, T4 RNA ligase 2 truncated, T4 RNA ligase 1, and T4 RNA ligase 2, the 3'-RNA tail on the cDNA is the preferred substrate. [Figure 4B] Figure 4B shows additional possible adapter designs that may need to be used for T4 RNA ligase 1 (5' phosphate oligo...5' adenylation). Figure 4B also depicts a dual adapter that may be required for T4 RNA ligase 2. [Figure 5A] 1 depicts a DNA gel digest of the pre-depletion cDNA library showing yields, with cycle numbers indicated on the right. [Figure 5B] DNA gel analysis of pre-depletion cDNA libraries for 50 ng, 1 ng, and 0.25 ng of input RNA using either T4 RNA ligase 2 shortened with a 5'App-deoxyoligonucleotide adapter (columns 1-3 from left to right) or T4 RNA ligase 1 with adenosine triphosphate (ATP) and a 5'phosphorylated deoxyoligonucleotide adapter (columns 4-9 from left to right) in the addition reaction. [Figure 6] FIG. 1 depicts a DNA gel analysis of a pre-depleted cDNA library showing the yield in an exemplary preparation, Preparation A. No adapter dimers are detected. [Figure 7] Figure 1 shows comparator data (A, internal data) of the yield of cDNA samples using the disclosed method, along with comparators (B-C). Gel electrophoresis images of degraded DNA from a cDNA library sample obtained by the disclosed method (A) compared to that produced using a commercially available third-party protocol or kit (sample BC). [Figure 8] FIG. 1 shows comparative data of post-depletion ribosomal DNA sequences remaining in a sample generated using the presently disclosed method (Sample A, internal) and samples generated utilizing commercially available third-party protocols or kits (Comparator Samples B, C, and D). [Figure 9A] FIG. 1 shows the alignment % and overlap % in Sample A (internal) and Sample C (comparator) described above with high input RNA preparations (input concentrations are shown on the axes). [Figure 9B] FIG. 1 shows the alignment % and overlap % in Sample A (internal) and Sample B, Sample C, and Sample D (comparator) described above with low input RNA preparations (input concentrations are shown on the axes). [Figure 10A]Figure 1 shows an analysis of sequence reads from the samples A (internal) and C (comparator) described above with high-input RNA preparations (input concentrations are shown on the axes), categorized into % assigned sequences and % unassigned (multi-mapping), unassigned (no feature), and unassigned (ambiguous), with the removed samples from Preparation A and Preparation C. Higher assigned categories represent enriched DNA for target sequences and removal of unwanted sequences. Categories including unassigned (ambiguous) may, in some embodiments, represent junk or non-useful sequences that should be ignored. [Figure 10B] FIG. 1 shows analysis of sequence reads of depleted samples from Sample A, Sample B, Sample C, and Sample D, sorted as described above, with different preparations, low-input RNA preparations (input concentrations are shown on the axes). [Figure 11A] Figure 1 shows an analysis of sequence reads using the Star Alignment program for removed samples from Preparation A (internal) and Preparation C (comparator) categorized into % uniquely mapped and % mapped to multiple loci, as well as the unmapped category as indicated, from within Samples A and C described above, which have high-input RNA preparations (input concentrations are shown on the axes). Higher assigned categories represent enriched DNA for target sequences and removal of unwanted sequences. Categories including unassigned (ambiguous) may, in some embodiments, represent junk or unhelpful sequences that are ignored. [Figure 11B] FIG. 1 shows analysis of sequence reads of depleted samples from Sample A (internal), Sample B, Sample C, and Sample D, sorted as described above, with different preparations: low-input RNA preparations (input concentrations shown on the axes). [Figure 12A]Figure 12A shows the genes detected per million reads for each library preparation method at each total RNA input level. The graph depicts the efficiency measure of the library preparation method at each varying total RNA input level (A, internal), even allowing comparison with different library preparation methods (i.e., comparator product C, comparator product B, and comparator product D, which select for polyadenylated RNA as described above). Figure 12A shows that sample A (commercially synthesized, sequenced using the currently developed protocol with internally designed and CRISPR guides) with higher inputs (100 ng, 50 ng, and 10 ng) detects more reads per million than sample C, which is comparator data from a commercial source using the Stranded Total RNA-Seq Kit v3-Pico Input Mammalian Kit with the same input of total RNA. [Figure 12B] Figure 12 shows the genes detected per million reads for each library preparation method at each total RNA input level. The graph depicts the efficiency measure of the library preparation method at each varying total RNA input level (A, internal), even allowing comparison with different library preparation methods (i.e., Comparator Product C, Comparator Product B, and Comparator Product D, which select for polyadenylated RNA as described above). Figure 12B shows a comparison of the efficiency of different library preparation methods at lower total RNA input levels (1 ng, 100 pg, and 10 pg). At 1 ng and 100 pg total RNA input levels, the sequencing library preparation using internally designed CRISPR guides synthesized by a commercial entity (Sample A) detects more or slightly more genes per million reads than all other methods (B, C, D). At a 10 pg total RNA input level, the internal sequencing preparation (sample B) using CRISPR guides synthesized by a different commercial entity for ribonucleotide removal detects slightly fewer genes per million reads than sample C or sample D from sequencing library preparations. [Figure 13] 10 shows data showing ribosomal RNA alignment as a percentage of the fragment of mapped bases with the ribosomal DNA sequence for the preparation of Sample A described above and Sample N produced with a commercial protocol. [Figure 14] FIG. 10 shows alignment rates of reads of an exemplary sequenced library preparation across the range of input RNA for Sample A and Sample N after depletion. [Figure 15] FIG. 1 shows the % overlap of sequenced library preps for Sample A and Sample N after depletion. [Figure 16] FIG. 1 shows the average coverage of gene reads in sequenced library preps of sample A and sample N after depletion at the indicated input ranges. [Figure 17] FIG. 1 shows the number of genes with read counts greater than 10 in the indicated samples A and N. [Figure 18] FIG. 1 shows the number of genes in the sequenced library prep with read counts greater than 10 in the indicated sample A. [Figure 19] Figure 1 shows the level of complexity with variable amounts of RNA input using T4 RNA ligase 1. T4 RNA ligase 1 was tested with the aim of replacing T4 RNA ligase 2 with a truncated version. [Figure 20] Assay of 96 samples using control RNA, i.e., normalized input. The assay shows preliminary results of approximately 29% rRNA remaining after depletion. At the time the assay was completed, further optimization was underway. [Figure 21] FIG. 1 shows MDS plots of sequencing readouts from test runs using human liver RNA samples (left) and cell extract RNA (right) library preparations from normalized input RNA. [Figure 22A]FIG. 1 shows star alignment scores for a human liver RNA library preparation in a control experiment using the outlined protocol, including a ribosomal RNA removal step. The number of reads is shown, demonstrating the quality of the reads following the protocol including the ribosomal RNA removal step. [Figure 22B] FIG. 1 shows star alignment scores of sequenced cell extract RNA library preparations, showing the number of reads and demonstrating the quality of the reads following a protocol including a ribosomal RNA removal step. [Figure 23A] FIG. 1 shows star alignment scores for sequences of a human liver RNA library preparation in a control experiment using the outlined protocol, including the ribosomal RNA removal step, showing a high percentage of sequencing reads, demonstrating the quality of the reads following the protocol including the ribosomal RNA removal step. [Figure 23B] Figure 1 shows star alignment scores of sequenced cell extract RNA library preparations, showing a high percentage of sequencing reads and demonstrating the quality of reads following a protocol including a ribosomal RNA removal step. [Figure 24] FIG. 1 shows the duplicate read rate, which represents another indicator of the quality of the reads from the sequenced library preparations. DETAILED DESCRIPTION OF THE INVENTION
[0027] Disclosed herein are methods, systems, and compositions for rapidly and reliably generating double-stranded cDNA libraries from total RNA extracted from biological samples. In some embodiments, the methods can be completed in a shorter total time (e.g., 8 hours or less) compared to conventional methods. In some embodiments, the methods require fewer reaction steps, liquid transfer steps, and / or reaction vessels, thus improving process efficiency. For example, frequent reaction vessel changes are time-consuming and increase production costs compared to currently used conventional methods. In some embodiments, the methods involve fewer wash and elution steps. In some embodiments, the methods have fewer steps requiring continuous operator participation or supervision. In some embodiments, for example, advantages of the processes described herein include the ability for an operator to step away from the reaction without interrupting the reaction process. In some embodiments, the reactions are adaptable to full or partial automation. In some embodiments, the reactions are adaptable to bench-to-bedside protocols. In some embodiments, the methods described herein require fewer active steps to complete compared to currently used conventional methods.
[0028] In one aspect, it has been observed that the lack of preparedness experienced during the coronavirus disease (COVID) outbreak, particularly in detecting the highly contagious SARS-CoV-2 pathogen, focused the world's attention on the problem to be solved. It took approximately 60 days for the first reverse transcription-polymerase chain reaction (RT-PCR) test for SARS-CoV-2 infection (developed by the United States (US) Centers for Disease Control and Prevention (CDC)) to become available. It was then estimated that more than 6 million tests per day were needed, but it took over 270 days to conduct 800,000 of those tests per day. This severely limited the ability to test and isolate symptomatic individuals or those in close contact with confirmed positive cases, thereby allowing disease transmission to spread exponentially. Population-scale deployment of testing strategies is necessary on "day 0," i.e., when the first case is reported, and this requires further efforts at accelerating testing and production processes at all levels. Next-generation sequencing (NGS) offers day 0 capabilities that have the potential to enable viable, broad-spectrum, and large-scale testing strategies. NGS-based methods can identify a much broader range of pathogens, resistance mutations, and biothreat agents than other molecular diagnostic methods.
[0029] In one aspect, the sensitivity of SARS-CoV-2 detection by NGS can be comparable to that by RT-PCR when sequences that are likely irrelevant or do not contribute to pathogen detection or host response are removed from relevant samples, depending on the objectives of the task. Furthermore, we have shown that such a strategy can also be used for variant typing, co-infection detection, and assessment of individual human host responses in a single workflow using existing open-source analytical pipelines. In some aspects, the proposed NGS framework described herein can be considered pathogen-agnostic, potentially fundamentally changing how both large-scale pandemic response and clinical microbiology testing are pursued in the future.
[0030] Current standards for ligation-based library preparation techniques require numerous steps and numerous washes and recovery steps. Generally, these current standards require extraction, fragmentation, first-strand cDNA synthesis, second-strand cDNA synthesis, end repair, A-tailing, directional adapter ligation, CRISPR cleavage, and post-cleavage amplification. Multiple bead-based cleanups are also required. While these are part of established library construction methods that enable strand-specific sequencing and unique dual indexing, they can have many drawbacks. In some embodiments, conventional systems are two-day protocols. In some embodiments, conventional systems involve 78 liquid transfer steps. Second, due to the inefficiencies and "losses" associated with all cleanup steps, minimizing these steps led to the method described herein, which can use a minimum initial sample input of 5 ng total RNA. Third, conventional methods are not amenable to true walk-away automation, starting from the sample, including extraction, library construction, and removal. To solve this problem, the present method is designed to perform multiple stepwise reactions in a single vessel, with minimal intervention between wash and elution steps other than sample addition and temperature control, and the reactions are minimized throughout the workflow. In some embodiments, an ideal system may include a cartridge pre-loaded with all the reagents needed to go from "sample" to cleared library ready for clinical use with minimal manual intervention.
[0031] The methods described herein offer various technical advantages that improve the process of generating nucleic acid libraries for sequencing (e.g., next-generation sequencing), including, but not limited to, fewer processing steps than currently used conventional methods, additive-only processing steps (meaning minimal loss of desired cDNA material that can occur during cleanup steps), and / or consolidation.
[0032] In some embodiments, a push button cartridge may be used to automate the process.
[0033] In one embodiment, rapid and accurate diagnosis is achieved, and the sequencing capacity requirement must be reduced from 40 million sequenced molecules to approximately 10 million molecules. This can be achieved by removing additional uninformative molecules with a more comprehensive set of sgRNAs.
[0034] The following paragraphs provide a simplified overview and exemplary summary of the methods for preparing the libraries described herein.
[0035] In one embodiment, the method includes obtaining a biological sample. In one embodiment, the biological sample includes isolated cells, such as peripheral blood mononuclear cells (PBMCs), red blood cells, or cells from body fluids or tissues. In some embodiments, the biological sample includes excised tissue, tumor samples, peritoneal cells, bone marrow cells, or other isolated cells. In some embodiments, the biological sample includes frozen tissue or forensic tissue containing fragmented or damaged nucleic acid material. In some embodiments, the biological sample may include cell-free nucleic acid.
[0036] In some embodiments, the methods described herein may include the following steps: In some embodiments, cell lysis and DNase digestion are performed to extract RNA from the sample. This may be followed by RNA fragmentation using heat and Mg++ for approximately 3 minutes. The fragmentation time may vary depending on the desired library length. In some embodiments, random primers with universal adapter tails at the 5' end are used with reverse transcriptase to generate first-strand cDNA. In some embodiments, exonuclease I is used to clean up excess primers, and shrimp alkaline phosphatase (rSAP) is used to sequester excess dNTPs and remove 5' phosphates from the sample's oligoribonucleotides (RNAs) so that they cannot serve as substrates for T4 RNA ligase I when used in addition or ligation reactions. In some embodiments, a heating step at this point (approximately 95°C for 10 minutes) 1) incubates exonuclease I and rSAP and 2) denatures the cDNA / RNA hybrid duplex, ensuring that the cDNA is single-stranded, which is the preferred form for TdT and T4 RNA ligase 1 or T4 RNA ligase 2 truncated forms. All three enzymes prefer single-stranded substrates, with TdT preferring the 3' end of single-stranded DNA and both ligases preferring the 3' end of single-stranded RNA. T4 RNA ligase 2 acts only on double-stranded RNA and may require a double-stranded RNA adapter to pair with the ribonucleotide tail on the cDNA. In some embodiments, terminal transferase (TdT) is used to add approximately three ribonucleotides to the 3' end of the cDNA strand. A truncated T4 RNA ligase (T4 RNL2 truncated) can be used to add / ligate a 5'-adenylated deoxy-oligonucleotide (5'App-DNA oligo, see Figure 4A) adapter to the 3' ribonucleotide tail of a cDNA molecule. As a second possibility, instead of T4 RNL2 truncated, T4 RNA ligase 1 (with adenosine triphosphate, ATP) can also be used to ligate a 5'-terminal phosphate DNA adapter (5'-p-DNA oligo, see Figure 4B) to the 3' ribonucleotide tail of a cDNA strand.As a third possibility, T4 RNA ligase 2 can be used with 5'-end phosphate double-stranded RNA adapters (5'-p-dsRNA oligos or 5'-p-double-stranded RNA, see Figure 4C). Apply a 1.8x magnetic bead-based cleanup (e.g., SPRI) to remove excess adapters, buffer components, and isolate the cDNA product with the desired double-stranded adapters.
[0037] In some embodiments, a first PCR is used to generate double-stranded library molecules using full-length PCR primers containing sample barcodes complementary to the 3' and 5' adapters of the cDNA strand (see Figure 2). The adapter-bearing double-stranded DNA is then incubated with a CRISPR / CAS ribonucleoprotein complex containing guide sequences targeting unwanted molecules (a target list for each application). A 0.6x-0.8x SPRI cleanup is then performed. A second, post-cleavage PCR is applied with primers targeting the fixed portion of the full-length adapter following the sample barcode to enrich for the desired uncleaved library molecules. After the second PCR, a further SPRI cleanup is performed. The final library is then sized and quantified for dilution / sequencer loading. Some versions of DNA require enzymatic or physical fragmentation of the double-stranded DNA, followed by denaturation before the random priming step described above. The fragmentation step can be adjusted to generate longer or shorter NGS library molecules using this technique for long-read or short-read sequencing.
[0038] In the cDNA library method disclosed herein, in one embodiment, the cDNA primer is an adapter sequence with 8 random bases at the 3' end, and in contrast to conventional methods, it targets a position in the RNA itself, and the sequence may not contain a deterministic / fixed sequence that functions as an adapter. In some embodiments, the cDNA primer contains a fixed nucleotide sequence.
[0039] In the cDNA library methods disclosed herein, in one embodiment, tailing is performed using at least two different rNTPs, at least three different rNTPs, or a mixture of all four different rNTPs and TdT, in contrast to conventional methods that use only one nucleotide type.
[0040] In one embodiment, it is envisioned that random bases added to cDNA, in combination with random bases (three Ns) 5' of (App-adapters) or 5' phosphorylated adapters (p-adapters), together form unique molecular identifiers (UMIs) 6-8 bases in length that can be used for deduplication. This may not be possible with T4 RNA ligase 2 and dsRNA adapters, as random bases cannot be easily generated in the paired portion of the double-stranded RNA adapter.
[0041] In one embodiment, exonuclease I and shrimp alkaline phosphatase (rSAP) are added after cDNA synthesis to digest excess adapter-N8 primers. Since the cDNA / RNA heteroduplex is not a substrate for exonuclease I, it is protected from digestion. rSAP is added for two purposes: first, to remove dNTPs from the reverse transcription reaction to prevent them from being present in the TdT tailing reaction in the subsequent step. Otherwise, TdT may preferentially use dNTPs over rNTPs to add DNA bases to the 3' end of the cDNA. Second, to remove any 5'-phosphates from the total RNA sample, which can potentially prevent T4 RNA ligase I from forming RNA-RNA chimeras between two different portions of the RNA sample. In some embodiments, this is followed by heating to 95°C for 10 minutes (e.g., about 95°C for about 10 minutes) to a) inactivate Exo I, which requires at least 20 minutes at 80°C; b) denature the cDNA / RNA; and c) the presence of Mg degrades the RNA to about 150-200 bases, allowing it to be largely washed away in the final cleanup. In some embodiments, adding NaCO / NaHCO buffer may be avoided, as downstream reactions will not function in this buffer, necessitating an isolation step to remove it. In some embodiments, RNase H can also be used, but this adds additional time and enzymes to the cost and process of the preparation, while heating to 95°C for 10 minutes (e.g., about 95°C for about 10 minutes) accomplishes this task without interrupting the flow of the preparation.
[0042] In some embodiments, cobalt has been shown to have an inhibitory effect on downstream ligation reactions and is therefore omitted from the TdT tailing reaction. Previously reported methods use a double-stranded DNA adapter containing the dinucleotide dC to be added to the tailed riboG by T4 DNA ligase and allowed to react for 6 hours to overnight at 16°C. Because T4 DNA ligase is a double-strand-specific DNA ligation enzyme, this stipulation requires that the tailed base be complementary to the base of the double-stranded adapter. Random-base pairing is less successful than single-stranded ligation because stochastic variation in the base at each position prevents efficient pairing, so mixing all four bases may be less efficient. This also applies to ligation of dsRNA adapters with T4 RNA ligase 2.
[0043] The methods disclosed herein use a truncated form of T4 RNA ligase 2 (or the K227Q mutant, or various other variants of the enzyme available from any vendor, such as New England Biolabs), which is a single-strand-specific RNA ligase, and the 3' acceptor for addition or ligation must be the RNA end. Note that in some embodiments, the same concept may work with DNA instead of RNA. However, this process is so inefficient that no one would use it for this purpose. At the same time, the addition of 5' adenylation (App-oligo, shown in Figure 4A) can be either DNA or RNA. In some embodiments, DNA is used as the adapter sequence herein. Furthermore, in some embodiments, ligation can also be performed using T4 RNA ligase 1 in combination with ATP as a cofactor. T4 RNA ligase 1 uses a 5' donor phosphate-terminated DNA or RNA substrate (Figure 4B, shown with DNA) to ligate to a 3' acceptor RNA end (3' RNA is also the preferred substrate for this enzyme; 3' DNA is very inefficient). Finally, T4 RNA ligase 2 acts only on double-stranded RNA and requires a duplex RNA adapter to pair with ribonucleotide-tailed cDNA molecules.
[0044] In some embodiments, 5' adenylated oligonucleotides (5' App-DNA oligos) or 5' phosphorylated oligonucleotides (5' p-DNA oligos) or double-stranded RNA adapters are ligated to the tailed ribonucleotide bases on the cDNA strand (placed there by the TdT tailing reaction). This reaction requires the presence of 10-20% PEG 8000, which acts as a crowding agent, and occurs at 25-28°C for 30 minutes to 1 hour (perhaps 1 hour for low inputs).
[0045] One advantage of the above system is that all reactions can be performed continuously without isolation of reaction components until just before the first PCR reaction. In some embodiments, steps 1, 2, 3, 4, 5, or more are performed in a single container. The final pre-PCR reaction mixture (containing cDNA and all reactants, particularly PEG8000) may be cleaned up before being added to the PCR reaction. In some embodiments, PCR is performed in large volumes to dilute other components, particularly large amounts of PEG. In some embodiments, if T4 DNA ligase is used instead of PEG8000-shortened T4 RNA ligase 2, it may be possible to add it directly to a PCR reaction amplified using Taq polymerase.
[0046] In the methods disclosed herein, cDNA may be cleaned up (e.g., purified) from PEG 8000 and other reactants from the reaction series, but this cleaned-up cDNA material may also be added directly to a PCR reaction using a high-fidelity polymerase (e.g., Roche-KAPA or Watchmaker-Equinox, approximately 100X Taq fidelity) instead of the low-fidelity Taq polymerase used in this application, both of which are able to bypass ribonucleotide bases added to the cDNA molecule.
[0047] In some aspects, the methods disclosed herein can be adapted for ablation workflows.
[0048] In some aspects, the methods provide for reactions to be performed in less time than currently available methods, kits, and systems.
[0049] In some aspects, the methods provide reactions that require fewer liquid handling and transfer steps than currently used methods, kits, and systems.
[0050] In some embodiments, the method is adaptable to hands-free mechanized robotic liquid handling systems for the majority of the reaction.
[0051] In some aspects, the methods are designed to be more efficient and less costly than currently used methods, kits, and systems.
[0052] A representative time scale broken down description of each step in the method herein is shown in Figure 1. Some of the detailed processes are described below.
[0053] Nucleic acid extraction and RNA fragmentation:
[0054] The samples described herein are biological samples containing single-copy and multi-copy sequences. Biological samples can be any biological sample containing nucleic acids, including cellular tissues or intracellular material. In some cases, the sample is fragmented and differentially degraded. In some embodiments, samples or biological samples include blood, serum, plasma, nasal swabs or nasopharyngeal washes, saliva, urine, gastric fluid, spinal fluid, tears, stool, mucus, sweat, earwax, oil, glandular secretions, cerebrospinal fluid, tissue, semen, vaginal fluid, interstitial fluid, including interstitial fluid from tumor tissue, ocular fluid, spinal fluid, throat swabs, breath, hair, fingernails, skin, biopsies, placental fluid, amniotic fluid, umbilical cord blood, emphatic fluids, cavity fluid, sputum, pus, microbiota, meconium, breast milk, and / or other excretions. In some cases, blood samples contain circulating tumor cells or cell-free nucleic acids, such as tumor RNA, fetal RNA, or cell-free RNA. In some embodiments, RNA is extracted from a provided tissue, cell, or biological sample.
[0055] Provided herein are methods, compositions, and kits for selective enrichment of nucleic acid extraction and modification to further enable downstream functions, such as selective enrichment of pathogen nucleic acids, commensal nucleic acids, microbiome nucleic acids, high information regions, cancer alleles, or other nucleic acids of interest in a sample.
[0056] In some cases, the sample includes a combination of host samples from humans, cows, horses, sheep, pigs, monkeys, dogs, cats, gerbils, birds, mice, rats, etc., or any mammalian laboratory model for a disease, illness, or other phenomenon associated with rare nucleic acids. In some cases, the host nucleic acid is derived from a human. A host can be considered an organism that harbors a parasite, a pathogen, or a benign or relatively benign microorganism. In some embodiments, the sample may be contaminated with a second nucleic acid sample. Some examples of second nucleic acids, such as the nucleic acid of interest, may be derived from a pathogen, a microbiome, a tumor, fetal RNA in a maternal sample, an allele, or a mutant allele. In some cases, the second nucleic acid is derived from a non-host. In some cases, the second nucleic acid is derived from a prokaryote. In some cases, the second nucleic acid is derived from one or more selected from the group consisting of a eukaryote, a virus, a bacterium, a fungus, and a protozoan. In some embodiments, the second nucleic acid may be derived from a tumor cell. In some embodiments, the second nucleic acid may be fetal RNA in a maternal sample. In some embodiments, the second nucleic acid can be an allele or a mutant allele. The microbiome is also a source of the second nucleic acid consistent with the disclosure herein, as are other examples that will be apparent to those skilled in the art.
[0057] In some embodiments, one or more RNA molecules present in a sample are synthetically prepared, and the RNA may include 2'-modified nucleosides, such as 2'-O-modified ribose, 2'-O-methyl nucleosides, or 2'-O-methoxyethyl nucleosides.
[0058] Total RNA extraction and RNA fragmentation are performed according to known methods. In some embodiments, fragmentation is performed using a Hybrid A N8 primer (commercially available or synthetic). In some embodiments, the Hybrid A N8 primer is HPLC purified. In some embodiments, fragmentation yields fragments of 10 to 10,000 bases in length from the extracted total RNA. In some embodiments, the fragments are 10 bp to about 1,000 bp. Optionally, the second nucleic acid is capped with an adapter having a size ranging from about 10 bp to about 1,000 bp. In some embodiments, the fragments are about 25 bp to about 2,000 bp. In some embodiments, the fragments are about 25 bp to about 2,000 bp. In some embodiments, the fragments are about 50 bp to about 5,000 bp. In some embodiments, the fragments are about 100 bp to about 10,000 bp.
[0059] Preparation of single-stranded cDNA with 5' and 3' adapters and UMI. In some aspects, provided herein are methods for preparing cDNA that improve the efficiency of generating cDNA libraries. In one aspect, first-strand cDNA generated by reverse transcriptase is ligated to two single-stranded adapters on both ends, i.e., the 5' and 3' ends. In some embodiments, one of the two single-stranded adapters may be a universal adapter. In some embodiments, one of the two single-stranded adapters may contain a sequence from a strand of a double-stranded universal adapter. In some embodiments, at least one of the two single-stranded adapters contains a unique sequence. In some aspects, the adapter-ligated single-stranded cDNA contains a unique molecular identifier (UMI). In some embodiments, the described method is a rapid and easy method for adding / ligating single-stranded adapters to single-stranded cDNA molecules, where the adapters are distinct (e.g., the 5' adapter is distinct from the 3' adapter). In some aspects, the 3' adapter may have a unique sequence. In some embodiments, the three adapters may have one, two, three, four, or more unique nucleotides added to their ends before ligating to the single-stranded cDNA molecules. In some embodiments, the method does not include a double-stranded DNA ligation reaction. In some embodiments, the method does not include a T4 DNA ligation enzyme reaction. Thus, in such embodiments, the method does not include at least one or more time-consuming reaction steps, thereby reducing overall reaction time and improving efficiency. In some embodiments, excess unused primers are removed from the reaction by adding exonuclease I and shrimp alkaline phosphatase (rSAP) to remove dNTPs and eliminate chimeric artifacts that may result from the use of T4 RNA ligase I.
[0060] First-Strand cDNA Synthesis: In some embodiments, first-strand cDNA synthesis is performed using polyA+ RNA with an oligodT primer. In some embodiments, as disclosed herein, a set of oligonucleotide random primers is generated to generate cDNA that can include a 5' adapter and a 3' adapter, where a 5' single-stranded adapter is attached to the random primer and used for reverse transcription, e.g., the 5' adapter sequence terminates at the 3' end in the random primer sequence. Similarly, the above method can be adapted to generate oligodT primers with 5' adapter sequences. In some embodiments, fragmented total RNA, or alternatively purified polyA+ RNA, is subjected to a reverse transcription reaction using an appropriate reverse transcriptase and primers, where the primers include a 5' adapter sequence. In some embodiments, the 5' adapter is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more deoxyribonucleotides in length. In some embodiments, the 5' adapter includes a barcode sequence. In some embodiments, the barcode is a 2-5 nucleotide sequence. In some embodiments, the random primer is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or more deoxyribonucleotides in length. In some embodiments, a high-fidelity reverse transcriptase is used in the reaction. An example of a high-fidelity reverse transcriptase is MMLV RT. In some embodiments, the amount of starting material may be 1-100 pg of RNA, 1-1000 pg of RNA, 1 pg-10 mg of RNA, 1 pg-100 mg of RNA, 1 pg-500 mg of RNA, or 1 pg-1 mg of RNA. The RNA target may be 100 bp to 20 kb in length. The temperature at which the first-strand synthesis reaction is performed can vary according to the enzyme manufacturer's recommendations. Typically, a temperature range of 42-55°C is recommended for the highest fidelity PCR reaction.
[0061] Ribotailing of First-Strand cDNA: In some embodiments, the 3'-end of a cDNA strand is modified by the addition of between two and four ribonucleotides. In some embodiments, incorporating a short ribonucleotide tail includes incorporating two, three, four, or more ribonucleotides at the 3'-end of the first-strand cDNA. In some embodiments, incorporating a short ribonucleotide tail includes incorporating two, three, four, or more ribonucleotides at the 3'-end of the first-strand cDNA. In some embodiments, the ribonucleotides incorporated at the 3'-end are not repetitive oligonucleotides. In some embodiments, the ribonucleotides incorporated at the 3'-end are not all guanidine ribonucleotides. However, it may be appreciated that mixing all four bases is less efficient because pairing random bases with random bases reduces yield and efficiency. In some embodiments, terminal transferase (TdT) is used to incorporate two, three, four, or more ribonucleotides. In some embodiments, single-strand-specific RNA ligase is used to incorporate two, three, four or more ribonucleotides into the 3'-end of first-strand cDNA. In some embodiments, cobalt is omitted from the TdT tailing reaction. In some embodiments, the inclusion of cobalt can affect or minimize the efficiency of downstream ligation reactions.
[0062] Addition of an adapter: In some embodiments, addition of a single-stranded second adapter comprises adding the adapter to the incorporated 3'-terminal ribonucleotide. In some embodiments, the addition is performed by a single-stranded RNA ligase in which the 3' acceptor for addition is a ribonucleotide. In some embodiments, the ligase is T4 RNA ligase 2 truncated (or the KQ mutant from New England Biolabs, or any other mutant). In some embodiments, the terminal nucleotide is an adenylate base. In some embodiments, the adenylate (A) base is a deoxyribonucleotide or a ribonucleotide. In one embodiment, to improve the efficiency of the ligase reaction, the 5' oligonucleotide at the addition / ligation end is designed to be an adenine (A) base. In some embodiments, the adenylate base is a deoxyribonucleotide. In some embodiments, the adenylate base is a ribonucleotide. In some embodiments, the adapter comprises an adenylate base. In some embodiments, the adenylate base is incorporated at the 5' end of the adapter. In some embodiments, 5' adenylation is a modification of a DNA oligonucleotide adapter. In some embodiments, engineered preadenylation is used herein to generate 5'-preadenylated DNA / RNA oligonucleotides. In some embodiments, adenylated oligos with pyrophosphate linkages are substrates for T4 RNA ligase in the absence of ATP, significantly reducing undesired self-ligation and other by-products. Preferably, the adapter sequence is 10 to 30 nucleotides in length. The adapter sequence may be 10 to 25 nucleotides in length. The adapter sequence may be 10 to 20 nucleotides in length. The adapter sequence may be 15 to 30 nucleotides in length. The adapter sequence may be 20 to 30 nucleotides in length. The adapter sequence may be 15 to 25 nucleotides in length. The adapter sequence may be 15 to 20 nucleotides in length. The adapter sequence may be 20 to 25 nucleotides in length. The adapter sequence may be 22 to 25 nucleotides in length. The adapter sequence may be 15 nucleotides in length.The adapter sequence may be 16 nucleotides in length. The adapter sequence may be 17 nucleotides in length. The adapter sequence may be 18 nucleotides in length. The adapter sequence may be 19 nucleotides in length. The adapter sequence may be 20 nucleotides in length. The adapter sequence may be 21 nucleotides in length. The adapter sequence may be 22 nucleotides in length. The adapter sequence may be 23 nucleotides in length. The adapter sequence may be 24 nucleotides in length. The adapter sequence may be 25 nucleotides in length.
[0063] In some embodiments, the adapter is an App-B adapter. An exemplary App-B adapter may have a nucleotide sequence represented as 5'-App-NNN-adapter-3'. In some embodiments, adenylated App-RNA oligonucleotide RNA is added or ligated to ribotailed cDNA using the shortened T4 RNA ligase 2 described above. In some embodiments, adenylated App-DNA oligonucleotides are added to ribotailed cDNA strands using the shortened T4 RNA ligase 2 described above. In some embodiments, the adenylate bases used in the reaction are as shown in Figure 4. In some embodiments, the reaction mixture contains TdT reaction components, and the adenylated App-DNA oligonucleotide ligase reaction reagents are added for the appropriate incubation time and at the appropriate temperature for each reaction to occur, as determined by the manufacturer of each enzyme. In some embodiments, the reaction is performed at a temperature and conditions optimized for the reaction. In some embodiments, the reaction is performed in the presence of PEG8000, which acts as a crowding agent. In some embodiments, about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20% PEG 8000 is used. In some embodiments, about 10%-20% PEG is used. In some embodiments, about 15% PEG 8000 is used. In some embodiments, the reaction mixture can be viscous, so care is taken to effectively mix the reaction components.
[0064] In some embodiments, the above steps are followed by a clean-up reaction.
[0065] In some embodiments, a cleanup reaction is performed to remove spent or unnecessary reactants or by-products from the nucleic acid product obtained at the end of the reaction. In one embodiment, the nucleic acid product is DNA. In some embodiments, the nucleic acid product is RNA. In some embodiments, the DNA is PCR-amplified DNA. In some embodiments, a cleanup reaction may be performed after a PCR reaction, a reverse transcriptase reaction, an end-labeling reaction, genomic DNA or RNA extraction, subtractive hybridization, or any such application, to remove dNTPs, spent or excess enzymes, salts, and other reagents, and DNA fragments smaller than the optimal size of the nucleic acid product from the product to obtain a clean PCR product. In some embodiments, a cleanup reaction is performed to remove reagents, reactants, or by-products that are likely to interfere with subsequent downstream reactions, such as sequencing, restriction digestion, labeling, addition, cloning, in vitro transcription, blotting, or in situ hybridization. In some embodiments, the cleanup reaction is performed using a commercially available kit. In some embodiments, the cleanup reaction can be performed without using a commercially available kit. In some aspects, the cleanup reaction may involve binding of the desired nucleic acid product to a solid phase, e.g., a membrane, washing to remove unwanted spent or excess reagents, reactants, or by-products, and then elution of the desired nucleic acid.
[0066] In some embodiments, amplification involves performing a polymerase chain reaction (PCR) using primers that anneal to the first and second adapters to generate double-stranded cDNA. In some embodiments, amplification of the double-stranded cDNA is performed by PCR to generate a double-stranded cDNA library. In the methods described herein, the cDNA must be cleaned up from PEG8000 and other reactants from the series of reactions. However, this cleaned-up cDNA material is added directly to a PCR reaction using a high-fidelity polymerase (Roche-KAPA or Watchmaker-Equinox, approximately 100x Taq fidelity) instead of the low-fidelity Taq polymerase used in this application; both of these high-fidelity enzymes can bypass the ribose bases added to the cDNA molecules.
[0067] In some embodiments, amplification continues for 5, 8, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, or 60 cycles. As known to those skilled in the art, a cycle is a thermal cycle that cycles through the functions of primer annealing, polymerase reaction, and denaturation, followed by primer annealing during the amplification reaction such that the DNA is exponentially amplified with each cycle.
[0068] In some embodiments, the amount of DNA to initiate an amplification reaction is as little as 5 pg of DNA. In some embodiments, the amount of DNA to initiate an amplification reaction is as little as 10 pg. In some embodiments, the amount of DNA to initiate an amplification reaction is as little as 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, In some embodiments, the amount of DNA to initiate an amplification reaction is as little as 0.2 ng of DNA. In some embodiments, the amount of DNA to initiate an amplification reaction is as little as 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1 ng of DNA. In some embodiments, the amount of DNA to initiate the amplification reaction is as little as 2 ng of DNA, 5 ng of DNA, or 10 ng of DNA.In some embodiments, the amount of DNA to initiate the amplification reaction is 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 90, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600 , 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000ng.
[0069] The method for generating double-stranded DNA described herein may include the sequence of amplified total RNA, most of which may not contain the target sequence that is desired to be amplified for diagnostic purposes or for preparing a library for future use.In some embodiments, the obtained PCR-amplified DNA is subjected to the removal of unwanted sequences using selective removal of unwanted sequences.In some embodiments, the method involves selective sequence enrichment by removing unwanted sequences.
[0070] Removal Workflow Endonucleases for targeted cleavage of nucleic acids: The methods disclosed herein include targeting cleavage of a first nucleic acid using a site-specific, targetable, and / or engineered nuclease or nuclease system. In one embodiment, a target can be considered a specific sequence on a nucleic acid that needs to be cleaved. In some embodiments, a target can refer to the type or source of nucleic acid desired to be cleaved. A nuclease is considered targetable if it can be designed to act only on a specific target, e.g., cleave the target. Such nucleases can generate double-stranded breaks (DSBs) at desired locations in genomes, cDNA, or other nucleic acid molecules. In other examples, nucleases can generate single-stranded breaks. In some cases, two nucleases are used, each generating a single-stranded break. Many cleavage enzymes consistent with the disclosure herein share the property of producing molecules with accessible ends for single-stranded or double-stranded exonuclease activity.
[0071] As used herein, an endonuclease may be a restriction enzyme that is specific for at least one site on a first nucleic acid and does not cleave a second nucleic acid. The endonucleases described herein may also be specific for repetitive nucleic acid sequences, such as transposons or other repeats, centromere regions, or other repetitive sequences in a host genome. For example, some restriction endonucleases consistent with the disclosure herein are Alu-specific restriction enzymes. A restriction is Alu-specific, or even "specific" for other targets, if it cleaves the target and does not cleave other substrates, or cleaves other targets only rarely, thereby differentially removing its "specific" target. The presence of non-Alu or other non-target cleavage, such as because the cleavage site occurs rarely elsewhere in the host genome or transcriptome, or in pathogen or other rare nucleic acids present in the sample, does not make the endonuclease "non-specific," as long as differential removal of undesired nucleic acids is achieved.
[0072] The first nucleic acid may include a restriction enzyme Alu recognition site. The second nucleic acid does not include an Alu recognition site. In some embodiments, the first nucleic acid includes at least one sequence that maps to at least one nucleic acid recognition site selected from the group consisting of AluI, AsuHPI, BpulIOI, BssECI, BstDEI, BstMAI, Hinfl, and BstTUI. In some embodiments, the second nucleic acid does not include at least one recognition site selected from the group consisting of AluI, AsuHPI, BpulIOI, BssECI, BstDEI, BstMAI, Hinfl, and BstTUI recognition sites.
[0073] Endonucleases consistent with the present disclosure may include at least one selected from a Clustered Regulatory Interspaced Short Palindromic Repeat (CRISPR) / Cas system protein-gRNA complex, a zinc finger nuclease (ZFN), and a transcription activator such as an effector nuclease (TALEN). In some embodiments, the gRNA, i.e., the guide RNA, is complementary to at least one site on the first nucleic acid to generate a cleaved first nucleic acid capped at only one end. Other programmable, nucleic acid sequence-specific endonucleases are also consistent with the present disclosure. Such endonucleases may be targetable and further engineered to act on specific targets.
[0074] Zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), engineered homing endonucleases, and RNA or DNA guided endonucleases, for example, CRISPR / Cas such as Cas9 or CPF1, and / or Argonaute system, are particularly suitable for carrying out some of the methods of the present disclosure.In addition, or alternatively, RNA targeting systems, such as the CRISPR / Cas system comprising c2c2 nuclease, can also be used.
[0075] The methods disclosed herein may include cleaving a target nucleic acid using a CRISPR system, such as a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR system. The CRISPR / Cas system may be a multiprotein system or a single effector protein system. Multiprotein, or Class 1, CRISPR systems include Type I, Type III, and Type IV systems. Alternatively, Class 2 systems include a single effector molecule and include Type II, Type V, and Type VI.
[0076] The CRISPR system used in some methods disclosed herein may include a single or multiple effector proteins. The effector protein may include one or multiple nuclease domains. The effector protein may target DNA or RNA, which may be single-stranded or double-stranded. The effector protein may generate a double-stranded or single-stranded break. The effector protein may include a mutation in the nuclease domain, thereby generating a nickase protein. The effector protein may include a mutation in one or more nuclease domains, thereby generating a catalytically inactive nuclease that can bind to but not cleave the target sequence. The CRISPR system may include a single or multiple guide RNAs. The gRNA may include a crRNA. The gRNA may include a chimeric RNA having the sequences of a crRNA and a tracrRNA. The gRNA may include separate crRNAs and tracrRNAs. The target nucleic acid sequence may include a protospacer adjacent motif (PAM) or a protospacer flanking site (PFS). The PAM or PFS can be 3' or 5' of the target or protospacer site. Cleavage of the target sequence can generate a blunt end, a 3' overhang, or a 5' overhang. In some cases, the target nucleic acid does not contain a PAM or PFS.
[0077] The gRNA may comprise a spacer sequence. The spacer sequence may be complementary to the target sequence or the protospacer sequence. The spacer sequence may be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or 36 nucleotides in length. In some examples, the spacer sequence may be less than 10 nucleotides or more than 36 nucleotides in length.
[0078] The gRNA may comprise a repeat sequence. In some cases, the repeat sequence is a part of the double-stranded portion of the gRNA. The repeat sequence may be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides in length. In some examples, the spacer sequence may be less than 10 nucleotides or more than 50 nucleotides in length.
[0079] The gRNA may include one or more synthetic nucleotides, non-naturally occurring nucleotides, modified nucleotides, deoxyribonucleotides, or any combination thereof. Additionally or alternatively, the gRNA may include a hairpin, a linker region, a single-stranded region, a double-stranded region, or any combination thereof. Additionally or alternatively, the gRNA may include a signaling molecule or a reporter molecule.
[0080] CRISPR nuclease can be expressed endogenously or recombinantly.CRISPR nuclease can be encoded on chromosome, extrachromosome, or on plasmid, synthetic chromosome or artificial chromosome.CRISPR nuclease can be provided as polypeptide or mRNA that encodes polypeptide.In this example, polypeptide or mRNA can be delivered by standard mechanisms known in the art, such as by using cell-penetrating peptide, nanoparticle or virus particle.
[0081] The gRNA can be encoded by genetic DNA or episomal DNA. The gRNA can be provided or delivered simultaneously with or sequentially to the CRISPR nuclease. The guide RNA can be chemically synthesized, transcribed in vitro, or otherwise produced using standard RNA production techniques known in the art.
[0082] The CRISPR system may be a type II CRISPR system, such as a Cas9 system. The type II nuclease may optionally include a single effector protein containing RuvC and HNH nuclease domains. In some cases, a functional type II nuclease may include two or more polypeptides, each of which includes a nuclease domain or a fragment thereof. The target nucleic acid sequence may include a 3' protospacer adjacent motif (PAM). In some cases, the PAM may be 5' of the target nucleic acid. The guide RNA (gRNA) may include a single chimeric gRNA, which contains both a crRNA sequence and a tracrRNA sequence. Alternatively, the gRNA may include a pair of two RNAs, such as a crRNA and a tracrRNA. The type II nuclease may generate a double-stranded break, which optionally results in two blunt ends. In some cases, the type II CRISPR nuclease is engineered to be a nickase, so that the nuclease generates only a single-stranded break. In such cases, two different nucleic acid sequences can be targeted by gRNA, so that two single-strand breaks are generated by the nickase. In some cases, the two single-strand breaks effectively result in a double-strand break. In some cases where type II nickase is used to generate two single-strand breaks, the resulting free nucleic acid ends can be blunt, have a 3' overhang, or have a 5' overhang. In some cases, type II nuclease can be catalytically inactive, so that it binds to the target sequence but does not cut it. For example, type II nuclease can have mutations in both the RuvC and HNH domains, thereby making both nuclease domains non-functional. Type II CRISPR systems can be one of three subtypes, specifically type II-A, type II-B, and type II-C.
[0083] The CRISPR system can be a Type V CRISPR system, such as a Cpf1, C2c1, or C2c3 system. The Type V nuclease can optionally include a single effector protein containing a single RuvC nuclease domain. In other cases, a functional Type V nuclease contains a RuvC domain split between two or more polypeptides. In such cases, the target nucleic acid sequence can include a 5' PAM or a 3' PAM. The guide RNA (gRNA) can include a single gRNA or a single crRNA, such as in the case of Cpf1. In some cases, a tracrRNA is not required. In other examples, such as when C2c1 is used, the gRNA can include a single chimeric gRNA containing both a crRNA sequence and a tracrRNA sequence, or the gRNA can include a set of two RNAs, such as a crRNA and a tracrRNA. Type V CRISPR nucleases can generate double-stranded breaks, which in some cases generates a 5' overhang. In some cases, type V CRISPR nuclease is engineered to be a nickase, so that the nuclease only generates single-strand breaks. In such cases, two different nucleic acid sequences can be targeted by gRNA, so that two single-strand breaks are generated by the nickase. In some cases, the two single-strand breaks effectively produce a double-strand break. In some cases, when type V nickase is used to generate two single-strand breaks, the resulting free nucleic acid ends can be blunt, have a 3' overhang, or have a 5' overhang. In some cases, type V nuclease can be catalytically inactive, so that it binds to the target sequence but does not cut it. For example, type V nuclease can have a mutation in the RuvC domain, thereby making the nuclease domain non-functional.
[0084] The CRISPR system may be a Type VI CRISPR system, such as a C2c2 system. The Type VI nuclease may include a HEPN domain. In some examples, the Type VI nuclease may include two or more polypeptides, each of which includes a HEPN nuclease domain or a fragment thereof. In such cases, the target nucleic acid sequence may be RNA, such as a single-stranded RNA. When using a Type VI CRISPR system, the target nucleic acid may include a protospacer flanking site (PFS). The PFS may be the 3' or 5', or target or protospacer sequence. The guide RNA (gRNA) may include a single gRNA or a single crRNA. In some cases, a tracrRNA is not required. In other examples, the gRNA may include a single chimeric gRNA containing both a crRNA sequence and a tracrRNA sequence, or the gRNA may include a pair of two RNAs, such as a crRNA and a tracrRNA. In some examples, the Type VI nuclease may be catalytically inactive so that it binds to the target sequence but does not cleave it. For example, a type VI nuclease can have a mutation in the HEPN domain, thereby rendering the nuclease domain non-functional.
[0085] Non-limiting examples of suitable nucleases for use in the present disclosure, including nucleic acid-guided nucleases, include C2c1, C2c2, C2c3, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cpf1, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, C sc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx100, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, their orthologs, or modified versions thereof.
[0086] In some methods disclosed herein, the Argonaute (Ago) system may be used to cleave a target nucleic acid sequence. The Ago protein may be derived from a prokaryote, eukaryote, or archaea. The target nucleic acid may be RNA or DNA. The DNA target may be single-stranded or double-stranded. In some cases, the target nucleic acid does not require a specific target flanking sequence, such as a protospacer adjacent motif or a sequence equivalent to a protospacer flanking sequence. The Ago protein may generate a double-stranded break or a single-stranded break. In some cases, when an Ago protein forms a single-stranded break, two Ago proteins may be used in combination to generate a double-stranded break. In some cases, the Ago protein contains one, two, or more nuclease domains. In some cases, the Ago protein contains one, two, or more catalytic domains. One or more nuclease or catalytic domains may be mutated in the Ago protein, thereby generating a nickase protein capable of generating a single-stranded break. In other examples, mutations in the catalytic domain of one or more nuclease or Ago proteins generate catalytically inactive Ago proteins that can bind to, but may not cleave, the target nucleic acid.
[0087] Ago proteins can be targeted to target nucleic acid sequences by guide nucleic acids. In many examples, the guide nucleic acid is guide DNA (gDNA). The gDNA can have a 5' phosphorylated end. The gDNA can be single-stranded or double-stranded. Single-stranded gDNA can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In some examples, the gDNA can be less than 10 nucleotides in length. In some examples, the gDNA can be more than 50 nucleotides in length.
[0088] Argonaute-mediated cleavage can generate blunt ends, 5' overhangs, or 3' overhangs. In some cases, one or more nucleotides are removed from the target site during or after cleavage.
[0089] Argonaute protein can be expressed endogenously or recombinantly.Argonaute can be encoded on chromosome, extrachromosomally, or on plasmid, synthetic chromosome, or artificial chromosome.In addition, or alternatively, Argonaute protein can be provided as a polypeptide or mRNA that encodes a polypeptide.In such cases, polypeptide or mRNA can be delivered by standard mechanisms known in the art, such as by using peptides, nanoparticles, or viral particles.
[0090] The guide DNA can be provided by genetic DNA or episomal DNA. In some cases, gDNA is reverse transcribed from RNA or mRNA. In some cases, the guide DNA can be provided or delivered simultaneously with or sequentially to the Ago protein. The guide DNA can be chemically synthesized, constructed, or otherwise produced using standard DNA production techniques known in the art. The guide DNA can be cleaved, released, or otherwise derived from genetic DNA, episomal DNA molecules, isolated nucleic acid molecules, or any other source of nucleic acid molecules.
[0091] Nuclease fusion protein can be recombinantly expressed.Nuclease fusion protein can be encoded on chromosome, extrachromosomally, or on plasmid, synthetic chromosome, or artificial chromosome.Nuclease and chromatin remodeling enzyme can be separately engineered and then covalently linked.Nuclease fusion protein can be provided as a polypeptide or mRNA that encodes a polypeptide.In such an example, polypeptide or mRNA can be delivered by standard mechanisms known in the art, such as by using peptides, nanoparticles, or viral particles.
[0092] Guide nucleic acid (e.g., gRNA) can be complexed with a compatible nucleic acid-guided nuclease and hybridize with a target sequence, thereby directing the nuclease to the target sequence.The target nucleic acid-guided nuclease that can be complexed with a guide nucleic acid can be referred to as a nucleic acid-guided nuclease that is compatible with the guide nucleic acid.Similarly, guide nucleic acid that can be complexed with a nucleic acid-guided nuclease can be referred to as a guide nucleic acid, such as a gRNA, that is compatible with a nucleic acid-guided nuclease (e.g., a Cas enzyme molecule).
[0093] The guide nucleic acid can be DNA. The guide nucleic acid can be RNA. The guide nucleic acid can include both DNA and RNA. The guide nucleic acid can include non-natural nucleotide modifications. When the guide nucleic acid includes RNA, the RNA guide nucleic acid can be encoded by a DNA sequence in a polynucleotide molecule such as a plasmid, a linear construct, or an editing cassette, as disclosed herein.
[0094] The guide nucleic acid may include a guide sequence. A guide sequence is a polynucleotide sequence that is sufficiently complementary to a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a complexed nucleic acid-guided nuclease to the target sequence. The degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or greater. Optimal alignment can be determined using any suitable algorithm for aligning sequences. In some embodiments, a guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides in length. In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, or 20 nucleotides in length. Preferably, a guide sequence is 10 to 30 nucleotides in length. A guide sequence can be 10 to 25 nucleotides in length. A guide sequence can be 10 to 20 nucleotides in length. A guide sequence can be 15 to 30 nucleotides in length. A guide sequence can be 20 to 30 nucleotides in length. A guide sequence can be 15 to 25 nucleotides in length. The guide sequence can be 15-20 nucleotides in length. The guide sequence can be 20-25 nucleotides in length. The guide sequence can be 22-25 nucleotides in length. The guide sequence can be 15 nucleotides in length. The guide sequence can be 16 nucleotides in length. The guide sequence can be 17 nucleotides in length. The guide sequence can be 18 nucleotides in length. The guide sequence can be 19 nucleotides in length. The guide sequence can be 20 nucleotides in length. The guide sequence can be 21 nucleotides in length. The guide sequence can be 22 nucleotides in length. The guide sequence can be 23 nucleotides in length.The guide sequence may be 24 nucleotides in length. The guide sequence may be 25 nucleotides in length.
[0095] The guide nucleic acid may include a scaffold sequence. Generally, a "scaffold sequence" includes any sequence having a sufficient sequence to promote the formation of a targetable nuclease complex, where the targetable nuclease complex includes a nucleic acid-guided nuclease and a guide nucleic acid, and the guide nucleic acid includes a scaffold sequence and a guide sequence. A sufficient sequence within the scaffold sequence to promote the formation of a targetable nuclease complex may include a degree of complementarity along the length of two sequence regions within the scaffold sequence, such as one or two sequence regions involved in the formation of a secondary structure. In some cases, one or two sequence regions are contained or encoded on the same polynucleotide. In some cases, one or two sequence regions are contained or encoded on separate polynucleotides. Optimal alignment can be determined by any suitable alignment algorithm and can further account for secondary structures, such as self-complementarity within one or two sequence regions. In some embodiments, the degree of complementarity between one or two sequence regions along the length of the shorter of the two when optimally aligned is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more. In some embodiments, at least one of the two sequence regions is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, at least one of the two sequence regions is about 10-30 nucleotides in length. At least one of the two sequence regions may be 10-25 nucleotides in length. At least one of the two sequence regions may be 10-20 nucleotides in length. At least one of the two sequence regions may be 15 to 30 nucleotides in length. At least one of the two sequence regions may be 20 to 30 nucleotides in length. At least one of the two sequence regions may be 15 to 25 nucleotides in length. At least one of the two sequence regions may be 15 to 20 nucleotides in length.At least one of the two sequence regions may be 20-25 nucleotides in length. At least one of the two sequence regions may be 22-25 nucleotides in length. At least one of the two sequence regions may be 15 nucleotides in length. At least one of the two sequence regions may be 16 nucleotides in length. At least one of the two sequence regions may be 17 nucleotides in length. At least one of the two sequence regions may be 18 nucleotides in length. At least one of the two sequence regions may be 19 nucleotides in length. At least one of the two sequence regions may be 20 nucleotides in length. At least one of the two sequence regions may be 21 nucleotides in length. At least one of the two sequence regions may be 22 nucleotides in length. At least one of the two sequence regions may be 23 nucleotides in length. At least one of the two sequence regions may be 24 nucleotides in length. At least one of the two sequence regions may be 25 nucleotides in length.
[0096] The scaffold sequence of the subject guide nucleic acid may comprise a secondary structure. The secondary structure may comprise a pseudoknot region. In some cases, the compatibility of the guide nucleic acid with the nucleic acid-guided nuclease is at least partially determined by the sequence within or adjacent to the pseudoknot region of the guide RNA. In some cases, the binding kinetics of the guide nucleic acid to the nucleic acid-guided nuclease is partially determined by the secondary structure within the scaffold sequence. In some cases, the binding kinetics of the guide nucleic acid to the nucleic acid-guided nuclease is partially determined by the nucleic acid sequence within the scaffold sequence.
[0097] In aspects of the present disclosure, the term "guide nucleic acid" refers to a polynucleotide that includes 1) a guide sequence that can hybridize to a target sequence, and 2) a scaffold sequence that can interact with or complex with a nucleic acid-guided nuclease, as described herein.
[0098] A guide nucleic acid can be compatible with a nucleic acid-guided nuclease when the two elements can form a functional targetable nuclease complex that can cleave the target sequence. In many cases, a scaffold sequence that is compatible with a compatible guide nucleic acid can be found by scanning the sequences adjacent to the natural nucleic acid-guided nuclease locus. In other words, a natural nucleic acid-guided nuclease can be encoded on the genome in the vicinity of the corresponding, compatible guide nucleic acid or scaffold sequence.
[0099] Nucleic acid-guided nucleases can be adapted to guide nucleic acids not found in the nuclease's endogenous host. Such orthogonal guide nucleic acids can be determined through empirical testing. Orthogonal guide nucleic acids can be derived from different bacterial species, or can be synthetic nucleic acids, or can otherwise be engineered to be non-natural.
[0100] Orthogonal guide nucleic acids compatible with common nucleic acid-guided nucleases may include one or more common features. The common feature may include a sequence outside the pseudoknot region. The common feature may include the pseudoknot region. The common feature may include a primary sequence or a secondary structure.
[0101] A guide nucleic acid can be engineered to target a desired target sequence by modifying the guide sequence so that it is complementary to the target sequence, thereby allowing hybridization between the guide sequence and the target sequence. A guide sequence with an engineered guide nucleic acid can be referred to as an engineered guide nucleic acid. Engineered guide nucleic acids are often non-natural and do not occur in nature.
[0102] In some embodiments, guide RNA molecules directly interfere with sequencing, for example, by binding to target sequences and preventing the nucleic acid polymerization that would occur across the bound sequence.In some embodiments, guide RNA molecules work in tandem with RNA-DNA hybrid binding moieties, such as proteins.In some embodiments, guide RNA molecules direct the modification of the sequencing library elements that they may bind to, such as methylation, base deletion, or cleavage, so that in some embodiments, the sequencing library elements that they bind to are unsuitable for further sequencing reactions.In some embodiments, guide RNA molecules direct the endonucleolytic cleavage of the DNA molecules that they bind to, for example, by proteins with endonuclease activity, such as Cas9 proteins. Zinc finger nucleases (ZFNs), transcription activator-like effector nucleases, and Clustered Regulatory Interspaced Short Palindromic Repeat / Cas-based RNA-guided DMA nucleases (CRISPR / Cas9) are particularly suited to some embodiments of the present disclosure.
[0103] The guide RNA molecule comprises a sequence that base pairs with the target sequence (first nucleic acid) to be removed from sequencing. In some embodiments, the base pairing is complete, while in some embodiments, the base pairing is partial or comprises unpaired bases along with bases that pair with non-target sequences.
[0104] A guide RNA may contain a region or regions that form an RNA "hairpin" structure. Such a region or regions may partially or completely comprise a palindromic sequence, such that the 5' and 3' ends of the region can hybridize to each other to form a double-stranded "stem" structure, which in some embodiments is capped by a non-palindromic loop that tethers each of the single strands in the double-stranded loop to each other.
[0105] In some embodiments, the guide RNA comprises a stem loop, such as the tracrRNA stem loop. The stem loop, such as the tracrRNA stem loop, can be complexed or bound to a nucleic acid endonuclease, such as the Cas9 DNA endonuclease. Alternatively, the stem loop can be complexed with an endonuclease other than Cas9, or a nucleic acid-modifying enzyme other than an endonuclease, such as a base-removing enzyme, a methyltransferase, or an enzyme with other nucleic acid-modifying activity that interferes with one or more DNA polymerase enzymes.
[0106] The tracrRNA / CRISPR / endonuclease system was identified as an adaptive immune system in eubacterial and archaeal prokaryotes that allows cells to acquire resistance through repeated infection with viruses of known sequence. For example, Deltcheva E, Chylinski K, Sharma CM, Gonzales K, Chao Y, Pirzada ZA et al. (2011) “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III” Nature 471(7340):602-7.doi:10.1038 / nature09886.PMC 3070239.PMID 21455174, Terns MP, Terns RM (2011) “CRISPR-based adaptive immune systems” Curr Opin Microbiol 14(3):321-7.doi:10.1016 / j.mib.2011.03.005.PMC 3119747.PMID 21531607, Jinek M, Chylinski K,Fonfara I,Hauer See M, Doudna JA, Charpentier E (2012) "A Programmable Dual-RNA-Guided DNA Endonuclease in Adaptive Bacterial Immunity" Science 337(6096):816-21. doi:10.1126 / science.1225829. PMID 22745249, and Brouns SJ (2012) "A swiss army knife of immunity" Science 337(6096):808-9. doi:10.1126 / science.1227253. PMID 22904002. The system has been adapted to guide targeted gene mutagenesis in eukaryotic cells.See, e.g., Wenzhi Jiang, Huanbin Zhou, Honghao Bi, Michael Fromm, Bing Yang, and Donald P. Weeks (2013) "Demonstration of CRISPR / Cas9 / sgRNA-mediated targeted gene modification in Arabidopsis, tobacco, sorghum, and rice," Nucleic Acids Res. Nov 2013;41(20):e188, Published online Aug 31, 2013. doi:10.1093 / nar / gkt780, and references therein.
[0107] As contemplated herein, in some aspects, guide RNAs are used to provide sequence specificity to DNA endonucleases, such as Cas9 endonuclease. In these aspects, the guide RNA comprises a hairpin structure that binds to or is bound by an endonuclease, such as Cas9 (other endonucleases are contemplated as alternatives or additions in some embodiments), and the guide RNA further comprises a recognition sequence that binds, specifically binds, or exclusively binds to a sequence to be removed from a sequencing library or sequencing reaction. The length of the recognition sequence in the guide RNA can vary depending on the desired degree of specificity in the sequence exclusion process. Short recognition sequences that contain sequences that occur frequently in a sample or contain differentially abundant sequences (AT abundance in an AT-rich genomic sample or GC abundance in a GC-rich genomic sample) are more likely to identify a relatively large number of sites and therefore are more likely to lead to frequent nucleic acid modifications, such as endonuclease activity, base excision, methylation, or other activities that interfere with at least one DNA polymerase activity. Longer recognition sequences that contain sequences that occur rarely in a sample or contain under-represented base combinations (GC abundance in an AT-rich genomic sample or AT abundance in a GC-rich genomic sample) are more likely to identify a relatively small number of sites and are therefore more likely to lead to rare nucleic acid modifications, such as endonuclease activity, base excision, methylation, or other activities that interfere with at least one DNA polymerase activity. Thus, as disclosed herein, in some embodiments, the frequency of sequence removal may be moderated through modifications to the length or content of the recognition sequence from the sequence reaction.
[0108] Guide RNA can be synthesized by many methods consistent with the disclosures described herein.Standard synthesis techniques can be used to produce large amounts of guide RNA and / or for highly repetitive target regions, which may require only a few guide RNA molecules to target many unwanted loci.Double-stranded DNA molecules can include an RNA site-specific binding sequence, a guide RNA sequence for Cas9 protein, and a T7 promoter site.In some cases, the double-stranded DNA molecule can be less than about 100 bp in length.T7 polymerase can be used to generate single-stranded RNA molecules that can include a target RNA sequence for Cas9 protein and a guide RNA sequence.
[0109] Guide RNA sequence can be designed by many methods.For example, in some embodiments, the non-genetic repeat sequence of human genome is divided into, for example, 100bp sliding window.Double-stranded DNA molecules can be synthesized in parallel on microarrays using photolithography.
[0110] The window may vary in size. 30-mer target sequences can be designed with a short trinucleotide protospacer adjacent motif (PAM) sequence of NGG flanking the 5' end of the target design sequence, which in some cases facilitates cleavage. See, inter alia, Giedrius Gasiunas et al., (2012) "Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria," Proc. Natl. Acad. Sci. USA. Sep 25, 109(39):E2579-E2586, incorporated herein by reference in its entirety. Redundant sequences can be removed, and the remaining sequences can be analyzed using search engines (e.g., BLAST) against the human genome to avoid hybridization to REFSEQ, ENSEMBL, and other genetic databases to avoid nuclease activity at these sites. The universal Cas9 tracer RNA sequence can be added to the guide RNA target sequence and then flanked by a T7 promoter.The upstream sequence of the T7 promoter site can be synthetic.Due to the highly repetitive nature of the target region in the human genome, in many embodiments, a relatively small number of guide RNA molecules will digest a large percentage of NGS library molecules.
[0111] Although it is estimated that only about 50% of protein-encoding genes have exons containing NGG PAM (protospacer adjacent motif) sequences, numerous strategies are provided herein to increase the percentage of genomes that can be targeted using the Cas9 cleavage system. For example, if a PAM sequence is not available in a DNA region, a PAM sequence can be introduced using a combinatorial strategy that uses a guide RNA bound to a helper DNA containing the PAM sequence. The helper DNA can be synthetic and / or single-stranded. The PAM sequence in the helper DNA is not complementary to the gDNA knockout target in the NGS library and therefore may not bind to the target NGS library template, but may bind to the guide RNA. The guide RNA can be designed to hybridize to both the target sequence and the helper DNA containing the PAM sequence to form a hybrid DNA:RNA:DNA complex that is recognized by the Cas9 system.
[0112] The PAM sequence may be represented as a single-stranded overhang or hairpin. The hairpin may optionally contain modified nucleotides that can be optionally degraded. For example, the hairpin may contain uracil, which can be degraded by uracil DNA glycosylase.
[0113] Instead of using DNA containing a PAM sequence, a modified Cas9 protein that does not require a PAM sequence or a modified Cas9 that is less sensitive to the PAM sequence may be used without the need for a helper DNA sequence.
[0114] In other cases, the sequence of the guide RNA used for Cas9 recognition can be elongated and inverted at one end to act as a dual-cutting system for tightly cleaving multiple sites. The guide RNA sequence can produce two cuts on the NGS DNA library target. This can be achieved by designing a single guide RNA for corresponding strands within a limited distance. One end of the guide RNA can bind to the forward strand of the double-stranded DNA library, and the other end can bind to the reverse strand. Each end of the guide RNA can contain a PAM sequence and a Cas9 binding domain. This can result in double double-strand cuts of NGS library molecules derived from the same DNA sequence separated by a predetermined distance.
[0115] An alternative version of the assay includes at least one sequence-specific nuclease, such as at least one restriction endonuclease, having an abundant recognition site in the first nucleic acid, or in some cases, a combination of sequence-specific nucleases. In some cases, the enzyme includes an activity that reacts with a specific sequence to produce a double-stranded break. In some cases, the enzyme includes any nuclease or other enzyme that digests double-stranded nucleic acid material in an RNA / DNA hybrid.
[0116] A nucleic acid probe (e.g., a biotinylated probe) complementary to the second nucleic acid can be hybridized to the second nucleic acid in solution and pulled down using, for example, magnetic streptavidin-coated beads. Unbound nucleic acids can be washed away, and the captured nucleic acids can then be eluted and amplified for sequencing or genotyping.
[0117] In some embodiments, implementation of the methods herein reduces the sequencing time of a sequencing reaction, such that a nucleic acid library is sequenced in less time, or using fewer reagents, or using less computational power. In some embodiments, implementation of the methods herein reduces the sequencing time of a sequencing reaction for a given nucleic acid library to about 90%, 80%, 70%, 60%, 50%, 40%, 33%, 30%, or less than 30% of the time required for sequencing in the absence of implementation of the methods herein.
[0118] In some aspects, in a given sequencing reaction, particular read sequences from particular regions are of particular interest. A means for enabling rapid identification of such particular regions would be beneficial, as it could reduce computational time or reagent requirements, or both computational time and reagent requirements.
[0119] Some embodiments relate to the generation of guide RNA molecules.In some cases, guide RNA molecules are transcribed from DNA template.A number of RNA polymerases can be used, such as T7 polymerase, RNA PolI, RNA PolII, RNA PolIII, organelle RNA polymerase, viral RNA polymerase, or eubacterial or archaeal polymerase.In some cases, polymerase is T7.
[0120] The guide RNA generating template includes a promoter, such as a promoter compatible with transcription directed by T7 polymerase, RNA Pol I, RNA Pol II, RNA Pol III, organellar RNA polymerase, viral RNA polymerase, or eubacterial or archaeal polymerase. In some cases, the promoter is a T7 promoter.
[0121] The guide RNA template may optionally encode a tag sequence. The tag sequence binds to a nucleic acid modifying enzyme such as a methylase, a base-removing enzyme, or an endonuclease. In the context of a larger guide RNA molecule bound to a non-target site, the tag sequence tethers the enzyme to the non-target region of nucleic acid, directing its activity to the non-target site. An exemplary tethered enzyme is an endonuclease such as Cas9.
[0122] The guide RNA template is complementary to a first nucleic acid corresponding to a ribosomal RNA sequence, a sequence encoding a globin protein, a sequence encoding a transposon, a sequence encoding a retroviral sequence, a sequence comprising a telomeric sequence, a sequence comprising a subtelomeric repeat, a sequence comprising a centromeric sequence, a sequence comprising an intron sequence, a sequence comprising an Alu repeat, a sequence comprising a SINE repeat, a sequence comprising a LINE repeat, a sequence comprising a double nucleic acid repeat, a sequence comprising a triple nucleic acid repeat, a sequence comprising a tetranucleolytic repeat, a sequence comprising a polyA repeat, a sequence comprising a polyT repeat, a sequence comprising a polyC repeat, a sequence comprising a polyG repeat, a sequence comprising an AT-rich sequence, or a sequence comprising a GC-rich sequence.
[0123] In many cases, the tag sequence comprises a stem-loop, such as a partial or complete stem-loop structure. The "stem" of the stem-loop structure is sometimes encoded by a palindromic sequence, and is either complete or interrupted to introduce at least one "kink" or turn into the stem. The "loop" of the stem-loop structure is often not involved in stem base pairing. In some cases, the stem-loop is encoded by a tracr sequence, such as the rtracr sequence disclosed in the references incorporated herein. Some stem-loops bind, for example, Cas9 or other endonucleases.
[0124] Guide RNA molecule further comprises a recognition sequence. The recognition sequence is perfectly or imperfectly reverse complementary to the non-target sequence to be removed from the nucleic acid library sequence set. Because RNA can hybridize with base pair combinations that do not occur in DNA-DNA hybrids (for example, G:U base pair), the recognition sequence does not need to be the exact reverse complementary sequence of the non-target sequence to be bound. In addition, slight perturbations from perfect base pairing can be tolerated in some cases.
[0125] End-protection: Protecting the ends of DNA molecules from degradation can be achieved by a variety of techniques, with the ultimate goal being to prevent exonuclease degradation at the adapter attachment site, thereby preventing the formation of adapter-attached fragments. Adapters can be added by ligation, polymerase-mediated amplification, transposase-mediated tagmentation, end modification, or other techniques. Exemplary adapters include hairpin adapters, which, when added to both ends, effectively join the two strands of a double-stranded nucleic acid to form a single-stranded circular molecule. Such molecules lack exposed ends for single- or double-stranded exonuclease degradation unless further cleaved by an endonuclease. Protection can also be achieved by the attachment of oligonucleotides or other molecules that are resistant to exonuclease activity. Examples of exonuclease-resistant adapters include phosphorothioate oligos, 2-O-methyl-modified nucleotide sugars, inverted dT or ddT, phosphorylation, C3 spacers, or other modifications that prevent exonucleases from degrading adjacent nucleic acids. Alternatively, or in combination, in some cases, the "adapter" constitutes a modification (e.g., a chemical modification) to the end of the sample nucleic acid without the ligation of an additional molecule, whereby the modification renders the nucleic acid resistant to exonuclease degradation.
[0126] A particular feature of the adapters herein is that, although they function locally independently of each other, a nucleic acid is not protected from degradation unless both ends are adapter-modified or modified. Otherwise, the adapter-modified end is protected from exonuclease activity, but the opposite end of the nucleic acid is susceptible to degradation, resulting in degradation of the entire molecule. This is the fate of a nucleic acid that is adapter-modified, as contemplated herein, but is subsequently cleaved by a sequence-specific nucleic acid exonuclease, resulting in two exposed, unprotected nucleic acid ends.
[0127] non-host nucleic acid
[0128] The targeting removal method herein results in the removal of the first nucleic acid from sample and the enrichment of the second nucleic acid.The sample can be used to create a library for sequencing, and the sequencing delivers the sequence data that can be extracted mainly from the second nucleic acid.For example, the second nucleic acid can be non-host nucleic acid.
[0129] In certain aspects, provided herein are methods that result in enrichment of microbial pathogens. In some cases, the methods herein allow for identification of the microbial pathogens. In some aspects, the microbial pathogen comprises a bacterial pathogen. In some aspects, the bacterial pathogen is selected from the group consisting of Bacillus, such as Bacillus anthracis or Bacillus cereus; Bartonella, such as Bartonella henselae or Bartonella quintana; Bordetella, such as Bordetella pertussis; Borrelia, such as Borrelia burgdorferi, Borrelia garinii, Borrelia afzelii, Borrelia relapsing fever; Brucella, such as Brucella abortus, Brucella canis, Brucella malta, or Brucella abortus; Campylobacter, such as Campylobacter jejuni; Chlamydia pneumoniae, Chlamydia trachomatis, Chlamydia psittacosis, and the like. Chlamydia or Chlamydophila such as Clostridium botulinum, Clostridium difficile, Clostridium perfringens, Clostridium tetani, Corynebacterium diphtheriae, Enterococcus such as Enterococcus faecalis or Enterococcus faecium, Escherichia such as Escherichia coli, Francisella such as Francisella tularensis, Haemophilus such as Haemophilus influenzae, Helicobacter such as Helicobacter pylori, Legionella neutrophils, Legionella species such as Leptospira mofila, Leptospira species such as Leptospira interrogans, Leptospira santarosai, Leptospira weilii, or Leptospira noguchii, Listeria species such as Listeria monocytogenes, Mycobacterium species such as Mycobacterium leprae, Mycobacterium tuberculosis, or Mycobacterium ulcerans, Mycoplasma species such as Mycoplasma pneumoniae, Neisseria species such as Neisseria gonorrhoeae or Neisseria meningitidis, Pseudomonas aeruginosa, the genus Monas, the genus Rickettsia such as Rickettsia rickettsii, the genus Salmonella such as Salmonella typhi or Salmonella typhimurium, the genus Shigella such as Shigella sonnei, the genus Staphylococcus such as Staphylococcus aureus, Staphylococcus epidermidis, and Staphylococcus saprophyticus, the genus Streptococcus such as Streptococcus agalactiae, Streptococcus pneumoniae, and Streptococcus pyogenes, the genus Treponema such as Treponema pallidum, the genus Vibrio such as Vibrio cholerae, and the genus Yersinia such as Yersinia pestis, Yersinia enterocolitica, or Yersinia pseudotuberculosis.In some embodiments, the microbial pathogen comprises a viral pathogen, which may be Adenoviridae such as Adenovirus, Herpesviridae such as Herpes Simplex 1, Herpes Simplex 2, Varicella-Zoster Virus, Epstein-Barr Virus, Human Cytomegalovirus, Human Herpesvirus 8, Papillomaviridae such as Human Papillomavirus, Polyomaviridae such as BK virus or JC virus, Poxviridae such as Smallpox, Hepadnaviridae such as Hepatitis B virus, Parvoviridae such as Human Bocavirus or Parvovirus, Astroviridae such as Human Astrovirus, Caliciviridae such as Norwalk virus, Picornaviridae such as Coxsackievirus, Hepatitis A virus, Poliovirus, Rhinovirus, Coronaviridae such as Severe Acute Respiratory Syndrome virus or Wuhan coronavirus, Hepatitis C virus, Yellow fever virus, Dengue virus, Western Angus virus, or the like. Flaviviridae (e.g., rabies virus), Togaviridae (e.g., rubella virus), Hepeviridae (e.g., hepatitis E virus), Retroviridae (e.g., human immunodeficiency virus (HIV)), Orthomyxoviridae (e.g., influenza virus), Arenaviridae (e.g., Guanarito virus, Junin virus, Lassa virus, Machupo virus, Sabia virus), Bunyaviridae (e.g., Crimean-Congo hemorrhagic fever virus), Filoviridae (e.g., Ebola virus, Marburg virus), Paramyxoviridae (e.g., measles virus, mumps virus, parainfluenza virus, respiratory syncytial virus, human metapneumovirus, Hendra virus, Nipah virus), Rhabdoviridae (e.g., rabies virus), Hepatitis D virus, or Reoviridae (e.g., rotavirus, orbivirus, coltivirus, or bannavirus pathogens). In some embodiments, the microbial pathogen comprises a fungal pathogen.In some aspects, the fungal pathogen is actinomycosis, allergic bronchopulmonary aspergillosis, asergilloma, aspergillosis, tinea pedis, basidiobolus, Basidiobolus ranarum, sand mites, blastomycosis, Candida krusei, candidiasis, chronic pulmonary aspergillosis, chrysosporium, chytridiomycosis, coccidioidomycosis, conidiobolomycosis, cryptococcosis, gattii cryptococcosis, deep dermatophytoses, dermatophytoses, dermatophytoses, dermatophytoses, endotrichomycoses, entomopathogenic fungi, epizootic lymphangitis, esophageal candidiasis, extratrichomycoses, fungal meningitis, fungemia, Geotrichum, Geotrichum Candidum, histoplasmosis, lobomycosis, Massospora cicadina, gypsum microsporum, silkworm disease, mycosis, tympanomycosis, Neozygites remaudierei, Neozygites slavi, Ochroconis gallopaba, Ophiocordyceps arborescens, Ophiocordyceps coenomyia, Ophiocordyceps macroacicularis macroacicularis, Hemiptera, thrush, paracoccidioidomycosis, pathogenic dimorphic fungi, penicilliosis, sand lichen, Piedraia, Pneumocystis pneumoniae, pseudoalescheriasis, scedosporiosis, sporotrichosis, tinea, tinea barbae, tinea capitis, tinea corporis, tinea cruris, tinea faciei, atypical tinea, tinea nigricans, white pedis, tinea versicolor, vomocytosis, white nose syndrome, zeaspora, or zygomycosis. In some cases, the methods herein result in enrichment of protozoan nucleic acids. In some cases, the methods herein result in enrichment of cancer nucleic acids. In some cases, the methods herein result in enrichment of fetal nucleic acids.
[0130] Use of endonuclease / exonuclease combinations in targeted removal
[0131] The method described herein for removing the first nucleic acid can result in a sequencing library with dramatically reduced complexity. Unwanted sequences are removed, and the remaining sequences can be more easily analyzed by NGS technology. The reduced complexity of the library reduces the sequencer capacity required for clinical deep sequencing and / or reduces the computational requirements for accurate mapping of unique sequences. The enriched sequences can be searched in bioinformatics databases, such as BLAST, to determine the identity of genes. The sequence information of the enriched nucleic acids can be used to determine the type of pathogen.
[0132] The methods described herein may include a step of performing genomic analysis of the second nucleic acid (e.g., the enriched nucleic acid). A genomic sequence database can be searched to find sequences related to the second nucleic acid. Generally, a search can be performed using a computer-implemented search algorithm by comparing the sequence to be searched with sequence information stored in multiple databases available via a communication network, such as the Internet. Examples of such algorithms include the Basic Local Alignment Search Tool (BLAST) algorithm, the PSI-blast algorithm, the Smith-Waterman algorithm, the Hidden Markov Model (HMM) algorithm, and other similar algorithms.
[0133] In some aspects, an endonuclease is configured to target multiple sites in the genome to be removed, after which exonuclease digestion generates nucleic acid molecules or fragments, which can be removed from the nucleic acid molecules, which are ligated to adapters and cloned, from which a library is prepared.
[0134] Thus, provided herein is an improved method for preparing a library comprising selective nucleic acid molecules from a sample comprising a first nucleic acid and a second nucleic acid, the method comprising the steps of: providing a sample comprising the first nucleic acid and the second nucleic acid; subjecting the sample to a treatment that removes nucleic acid fragments smaller than a threshold size from the sample; exposing the first nucleic acid and the second nucleic acid to an endonuclease to form at least one cleaved first nucleic acid, wherein the endonuclease cleaves the first nucleic acid but does not cleave the second nucleic acid; contacting the sample from step (c) with an exonuclease to produce exonuclease-digested nucleic acid molecules; and enriching the nucleic acid molecules remaining after exonuclease digestion by size-selecting nucleic acid molecules larger than a threshold size to produce a library comprising enriched nucleic acid molecules.
[0135] In some aspects, provided herein are improved methods for enriching selective nucleic acid molecules, such as nucleic acid molecules, from contaminated or biological samples. In some aspects, the methods provided herein increase the specificity of the enriched nucleic acids. In some aspects, the methods include additional steps of size exclusion cleaning and enrichment. In some aspects, the methods provided herein increase the yield of enriched nucleic acids. In some aspects, the methods include exclusion for higher yields of purification steps.
[0136] In some embodiments, yields are increased by 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% or more compared to conventional methods.
[0137] Definitions: A partial list of relevant definitions follows:
[0138] As used herein, the term "enriched" is used in a relative sense, whereby a second nucleotide or a population containing the second nucleotide is enriched upon selective removal of a first nucleotide or a population containing the first nucleotide. It does not require an absolute increase in what is enriched. Rather, an absolute or relative increase resulting from the removal or deletion of other nucleic acids can constitute "enrichment" as used herein.
[0139] As used herein, the term "removal" or "removing" is used in a relative sense, whereby a first nucleotide or a population containing the first nucleotide is degraded upon selective preservation of a second nucleotide or a population containing the second nucleotide. It does not require an absolute reduction in what is removed. Rather, an absolute or relative reduction resulting from preservation of other nucleic acids can constitute "removal" as used herein.
[0140] As used herein, "about" a given value is defined as + / - 10% of the given value.
[0141] As used herein, NGS or next generation sequencing may refer to any number of nucleic acid sequencing techniques such as 5.1 massively parallel sequencing (MPSS), polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLi sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, heliscope single molecule sequencing, single molecule real time (SMRT) sequencing, tunneling current DNA sequencing, sequencing by hybridization, sequencing using mass spectrometry, microfluidic Sanger sequencing, microscope-based techniques, RNAP sequencing, and in vitro viral high-magnification sequencing.
[0142] As used herein, "modifying" a nucleic acid can mean making an alteration to the covalent bonds in the nucleic acid, such as methylation, base removal, or cleavage of the phosphodiester backbone.
[0143] As used herein, "direct transcription" can be the provision of a template sequence from which a particular RNA molecule can be transcribed.
[0144] "Amplified nucleic acid" or "amplified polynucleotide" includes nucleic acid or polynucleotide molecules whose amount has been increased by an in vitro nucleic acid amplification or replication method compared to its starting amount. For example, amplified nucleic acids can optionally be obtained from polymerase chain reaction (PCR), which can amplify DNA in an exponential manner (e.g., amplifying to 2n copies in n cycles), in some cases, where most of the product is generated from an intermediate template rather than directly from the sample template. Amplified nucleic acids can alternatively be obtained from linear amplification, where the amount increases linearly over time, which, in some cases, produces products synthesized directly from the sample.
[0145] The term "biological sample" or "sample" generally refers to a sample or portion separated from a biological entity. A biological sample may refer to a whole biological entity, and examples include, but are not limited to, bodily fluids, dissociated tumor specimens, cultured cells, and combinations thereof. A biological sample may be obtained from one or more individuals. One or more biological samples may be obtained from the same individual. In one non-limiting example, a first sample is obtained from the individual's blood and a second sample is obtained from the individual's tumor biopsy. Examples of biological samples include, but are not limited to, blood, serum, plasma, nasal swabs or nasopharyngeal washes, saliva, urine, gastric juice, spinal fluid, tears, stool, mucus, sweat, earwax, oil, glandular secretions, cerebrospinal fluid, tissue, semen, vaginal fluid, interstitial fluid including interstitial fluid from tumor tissue, ocular fluid, spinal fluid, throat swabs, breath, hair, fingernails, skin, biopsy specimens, placental fluid, amniotic fluid, umbilical cord blood, sputum, cavity fluid, sputum, pus, bacterial flora, meconium, breast milk, and / or other secretions. In some cases, blood samples contain circulating tumor cells or cell-free DNA such as tumor DNA or fetal DNA. Samples include nasopharyngeal washes. Examples of tissue samples of interest include, but are not limited to, connective tissue, muscle tissue, nerve tissue, epithelial tissue, cartilage, cancer or tumor samples, or bone. Samples may be obtained from humans or animals. The sample may be obtained from a mammal, including a vertebrate such as a mouse, monkey, human, farm animal, sport animal, or pet. The sample may be obtained from a living or dead subject. The sample may be obtained fresh from the subject or may have undergone some form of pre-processing, storage, or transport.
[0146] As used herein, nucleic acid sample refers to the nucleic acid sample that first nucleic acid is determined, and nucleic acid sample is sometimes extracted from the above-mentioned biological sample.Alternatively, nucleic acid sample is sometimes artificially synthesized, synthetic or de novo synthesized.DNA sample is sometimes genome, while in other cases, DNA sample is derived from reverse-transcribed RNA sample.
[0147] "Body fluid" generally describes a fluid or secretion that comes from the body of a subject. In some instances, the body fluid is a mixture of two or more types of body fluids that are mixed together. Some non-limiting examples of body fluids include, but are not limited to, blood, urine, bone marrow, cerebrospinal fluid, pleural fluid, lymphatic fluid, amniotic fluid, peritoneal fluid, sputum, or a combination thereof.
[0148] "Complementary" or "complementarity," or sometimes more precisely "reverse complementarity," refer to nucleic acid molecules related by base pairing. Complementary nucleotides are generally A and T (or A and U), or C and G (or G and U). Functionally, two single-stranded RNA or DNA molecules are complementary when they form a double-stranded molecule through hydrogen bond-mediated base pairing. Two single-stranded RNA or DNA molecules are said to be substantially complementary when the nucleotides of one strand, optimally aligned and with appropriate nucleotide insertions or deletions, pair with at least about 90% to about 95% or more complementarity, and more preferably about 98% to about 100% complementarity, and even more preferably 100% complementarity. Alternatively, substantial complementarity exists when an RNA or DNA strand hybridizes to its complementary sequence under selective hybridization conditions. Selective hybridization conditions may or may not include stringent hybridization conditions, but are not limited to: Hybridization temperatures are generally at least about 2°C to about 6°C below the melting temperature (Tm).
[0149] "Duplex" refers to two polynucleotide strands annealed by complementary base pairing, such as in reverse complementary orientation, as the case may be.
[0150] "Known oligonucleotide sequence" or "known oligonucleotide" or "known sequence" refers to a known polynucleotide sequence. In some cases, the known oligonucleotide sequence corresponds to a designed oligonucleotide, e.g., a universal primer for a next-generation sequencing platform (e.g., Illumina, 454), a probe, an adapter, a tag, a primer, a molecular barcode sequence, or an identifier. The known sequence optionally includes a portion of a primer. In some cases, the known oligonucleotide sequence is not actually known to a particular user, but is structurally known, for example, by being stored as computer-accessible data. The known sequence is optionally a trade secret not actually known or secret to one or more users, but known to the entity that designed the particular component, kit, device, or software of the experiment the user is using.
[0151] A "library" may refer to a collection of nucleic acids. A library optionally contains one or more target fragments. In some cases, the target fragments include amplified nucleic acids. In other cases, the target fragments include unamplified nucleic acids. A library optionally contains nucleic acids with one or more known oligonucleotide sequences attached to the 3' end, the 5' end, or both the 3' and 5' ends. A library is optionally prepared such that the fragments contain known oligonucleotide sequences that identify the source of the library (e.g., a molecular identifier barcode that identifies a patient or DNA source). In some cases, two or more libraries are pooled to create a library pool. Libraries are optionally generated using other kits and techniques, such as transposon-mediated labeling or "tagmentation" as known in the art. Kits are commercially available. One non-limiting example of a kit is the Illumina NEXTERA kit (Illumina, San Diego, CA).
[0152] The term "polynucleotide" or "nucleic acid" includes, but is not limited to, various DNA and RNA molecules, derivatives, or combinations thereof, including species such as dNTPs, ddNTPs, DNA, RNA, peptide nucleic acids, cDNA, dsDNA, ssDNA, plasmid DNA, cosmid DNA, chromosomal DNA, genomic DNA, viral DNA, bacterial DNA, mtDNA (mitochondrial DNA), mRNA, rRNA, tRNA, nRNA, siRNA, snRNA, snoRNA, scaRNA, microRNA, dsRNA, ribozymes, riboswitches, and viral RNA.
[0153] Before the methods, compositions, and kits of the present invention are described in more detail, it should be understood that the invention is not limited to the particular methods, compositions, or kits described, as such may, of course, vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is limited only by the appended claims as interpreted herein. The examples are set forth so as to provide those of ordinary skill in the art with a more complete disclosure and description of how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention, nor are they intended to represent that the following experiments are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some account for experimental error and deviation. Unless otherwise indicated, parts are parts by weight, molecular weight is average molecular weight, temperature is in degrees Celsius, and pressure is near atmospheric.
[0154] Where ranges of values are provided, each intervening value, to two decimal places of the lower limit, between the upper and lower limits of that range is specifically disclosed unless otherwise specified. Each subrange between any stated or intervening value in any stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of such subranges may independently be included or excluded, and if either or both are included in the subranges, each range is encompassed within the invention, subject to any specifically excluded limits in the stated range. When a stated range includes one or both limits, ranges excluding either or both of those included limits are also included in the invention.
[0155] Integrating library preparation methods into systems and automating reactions: To integrate library preparation methods into workflows from sample processing to sequencing to target sequence diagnosis, or even bench-to-bedside, devices can be designed that handle and transfer liquids in a programmable fashion. The methods described above are designed to minimize liquid transfer steps, allowing multiple steps to be performed within a single vessel, eliminating the need for vessel-to-vessel sample transfer. Several systems are commercially available that allow hands-free sample sequencing. Exemplary NGS-compatible systems include the epMotion® 5075t Automated Liquid Handler and the Perkin Elmer JANUS NGS Express AJS4NGS / D Automated Liquid Handler.
[0156] Systems that integrate liquid handling in the NGS workflow, from the extraction step to the sequencing reaction, may be available. Exemplary systems include, but are not limited to, the Rhoenix workstation.
[0157] The method has been optimized to minimize cleanup and transfer steps and can be integrated into automated systems designed or further optimized from existing systems to achieve a hands-free, scalable workflow suitable for "sample-to-answer" reactions in clinical POC environments.
[0158] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can optionally be used in the practice or testing of the present invention, potential and preferred methods and materials are described herein. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials with which the publications are cited. It will be understood that the present disclosure supersedes any disclosure of an incorporated publication to the extent of any conflict.
[0159] As will be apparent to those skilled in the art upon review of this disclosure, each embodiment described and illustrated herein has other components and features which may be readily separated from or combined with the features of any of the other embodiments without departing from the scope or spirit of the invention. Any recited method is contemplated to be carried out in the order of events recited or in any other order which is logically possible.
[0160] As used in this specification and the appended claims, the singular forms "a," "an," and "the" should be understood to include the plural forms unless the context clearly indicates otherwise. Thus, for example, a reference to "a cell" includes a plurality of such cells, and a reference to "the peptide" includes a reference to one or more peptides and equivalents thereof, such as polypeptides known to those skilled in the art.
[0161] In some aspects, the methods described herein include (i) removing relatively small DNA molecules, such as less than 1 kb, from a nucleic acid sample, (ii) allowing the genomic nucleic acid to be removed, such as human genomic nucleic acid, to be digested to fragment sizes of 1 kb or less, (iii) sorting and selecting the digested nucleic acids based on size, and (iv) creating a library from the selected digested material. This can be done on genomic DNA as well as full-length cDNA.
[0162] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the publication dates provided may be different from the actual publication dates which may need to be independently confirmed. [Example]
[0163] The following examples are given for the purpose of illustrating various aspects of the present invention and are not meant to limit the invention in any manner. The examples, together with the methods described herein, represent presently preferred embodiments and are exemplary and are not intended to limit the scope of the invention. Those skilled in the art will envision modifications thereof and other uses that are encompassed within the spirit of the invention as defined by the scope of the claims.
[0164] Example 1. Generation of 5' and 3' adapter-ligated single-stranded cDNA This example provides a detailed method for preparing first strand cDNA with 5' and 3' adapters ligated for rapid and efficient generation of amplified cDNA libraries.
[0165] Total RNA is isolated from the sample and fragmented, followed by first-strand cDNA synthesis using random primers with universal adapter sequences at the 5' end. Usual precautions are taken to protect the reaction from RNase action. The reaction is brought to a temperature of 50°C and incubated for approximately 30-60 minutes, as determined by one skilled in the art. The reverse transcriptase reaction is completed by a brief incubation at 75°C for approximately 10 minutes.
[0166] At the end of this reaction, Exo1 (exonuclease) and shrimp alkaline phosphatase (rSAP) are added to the reaction mixture. The primers are digested by adding Exo1. Because dNTPs are digested with rSAP, they do not interfere with the subsequent terminal deoxynucleotidyl transferase (TdT) tailing reaction. All of these reactions can be performed consecutively without sample purification or centrifugation.
[0167] In the next step, the cDNA is denatured from the RNA by heating to 95°C for 15 minutes, which also inactivates SAP and Exo1. The Mg++ present in the buffer from the first fragmentation step continues to digest the now-denatured RNA fragments. The intact cDNA strands are available for the TdT reaction. A mixture of ribonucleotides (NTPs) is added to the solution. The TdT enzyme solution is added along with the required amount of TdT buffer. No cobalt-containing buffer is added. The reaction is allowed to proceed for 30 minutes at 37°C. One to four ribonucleotides are added to the 3' end. The ribonucleotides are not guanine repeats. Random incorporation of one, two, three, and four ribonucleotides can generate the UMI code for each cDNA.
[0168] After an incubation period, adapter ligation reagents are added to the reaction. App-B adenylated DNA oligonucleotide adapters are ligated to the 3'-terminal ribonucleotides. T4 RNA ligase 2 truncated enzyme is used for ligation in the presence of approximately 15% PEG 8000. Ligation is carried out at the manufacturer's recommended temperature, e.g., 28°C, for 30 minutes.
[0169] Finally, after ligation, the product is captured with Ampure XP beads (1.8X) by adding the bead solution to the existing ligation mixture and incubating at room temperature for a period of time to allow the RNA / cDNA to bind. The beads are washed twice with 150 μL of 80% EtOH, air-dried, and eluted with 24 μL of 10 mM Tris pH 8. The eluate is transferred to a new PCR tube.
[0170] The schematic workflow is symbolically represented in Figures 2 and 3 (steps 2-5). Using the protocol described above, sample reactions yielded PCR products as shown in Figures 5A-5B. Reactions were performed in a single vessel (one-pot reactions) using a wide range of starting material, from 10 pg to 100 ng. No input quantification was required. No adapter quantification relative to the template was required. The process is amenable to automation, with only eight addition steps.
[0171] Example 2: Exemplary Methods of Conducting Assays This example provides a general protocol for performing the reaction. The reagents used in the reaction are as follows: Hybrid A-N8 primers were used at 50 pmoles per sample and were obtained from Integrated DNA Technologies. The Ns stretch represents the random primer portion, which is mixed mechanically and purified by HPLC. App-B adapters (Integrated DNA Technologies) (50 pmoles per sample) were used in this example. The random portion of the sequence is mixed mechanically and purified by HPLC. RNaseOUT (60 U per sample) (Thermo) Catalog No. 10777019, 5,000 U. Maxima H minus Reverse Transcriptase (100 U per sample) (Thermo) Catalog No. EP0753, 4X 10,000 U. Exonuclease I (20 U per sample) (Thermo) Catalog No. EN0582, 20,000 U. Shrimp alkaline phosphatase (rSAP) (1 U per sample) (New England Biolabs) Catalog No. M0371L, 2,500 U. T4 RNA Ligase 2 Truncated (T4 RNL2 Truncated) (400 U per sample) (New England Biolabs) Catalog No. M0242L, 10,000 U. Terminal deoxynucleotidyl transferase (TdT) (40 U per sample) (New England Biolabs) Catalog No. M0315L, 2,500 U. 2X Equinox mix (50 μL per sample) (Watchmaker Genomics) Catalog No. 7K0014-384, 12 mL (25 μL per sample in the reaction mixture). Cas9 Nuclease Streptococcus pyogenes (20 μM) (46 pmol per sample) Catalog No. M0386M, 2,500 pmol. Proteinase K (1 μL per sample) (Thermo) Catalog No. EO0491, 1 mL or Catalog No. EO0491, 5X1 mL. PCR Primers - Standard Illumina Primers, full length barcoded (small p5 and p7 primers only). Purification of ligation or amplification products is performed using AmpureXP SPRI beads.
[0172] An exemplary optimization protocol for the method steps is detailed below.
[0173] Fragmentation: Combine the following reagents in a tube or well (e.g., container): 2 μL (100 ng human liver RNA) from above 1 μL of Hybrid A-N8 (50 μM) 2 μL of 5X First Strand Buffer (Maxima H minus) 5 μL Incubate the reagent at 95°C for 1-4 minutes and then place on ice for 2 minutes. Add 3 μL of the mixture prepared above to 2 μL of sample (add 1 μL to 2 μL of sample).
[0174] Mix the following reagents for the reverse transcription reaction (Addition 2) → 5 μL each 2.5 μL nuclease-free water 1 μL 0.1 M DTT 0.5 μL of RNaseOUT (40 U / μL) 0.5 μL 10 mM dNTPs 0.5μL of Maxima H minus (200U / μL) 10 μL
[0175] After mixing, incubate at 25°C for 15 minutes, 50°C for 50 minutes, and 75°C for 10 minutes. Add 1 μL of exonuclease I (Exo I) (20 U / μL) and 1 μL of shrimp alkaline phosphatase (rSAP) (1 U / μL), then incubate at 37°C for 30 minutes and 95°C for 15 minutes. (Add 3) → 2 μL each.
[0176] Terminal deoxynucleotidyl transferase (TdT) reaction. Add 4) → 8 μL 12 μL from above 3.5 μL nuclease-free water 2 μL of 10X TdT buffer 0.5 μL, 1 mM total rNTPs (0.25 mM each) 2 μL of terminal transferase enzyme (20 U / μL) 20 μL total volume
[0177] Eight microliters of the TdT mixture is added to the vessel, which is then incubated at 37° C. for 30 minutes.
[0178] Adapter ligation, T4 RNA Ligase 2 shortened type. Add 5 → 20 μL each
[0179] 20 from above 4 μL 10X T4 RNL2 buffer 1 μL nuclease-free water 1 μL of App-B adapter (50 μM) 12 μL of 50% PEG8000 (final reaction: 15% PEG8000) 2 μL of T4 RNL2 truncated form (200 U / μL) 40 μL total volume Incubated at 28°C for 30 minutes.
[0180] Ampure bead purification
[0181] To purify, add 40 μL of nuclease-free water and 145 μL of Ampure Beads XP (1.8X) to the tube, mix, and then allow the RNA / cDNA to bind for 10 minutes at room temperature. Wash the beads twice with 150 μL of 80% EtOH, air dry for 8 minutes, and elute with 24 μL of 10 mM Tris pH 8 (allow to stand for 5 minutes before elution), then transfer 23 μL to a new PCR tube.
[0182] PCR1:
[0183] Add the following to a PCR tube: 23 μL from above 25μL of 2X Equinox (Watchmaker Genomics) 1.25 μL of PCR primer barcoded p7 (25 μM) 1.25 μL of PCR primer barcoded p5 (25 μM) 100 ng input in 50 μL total volume: 6 cycles
[0184] PCR protocol (→PCR protocol) 98°C, 2 minutes →98℃, 20 seconds →60℃, 30 seconds →72℃, 45 seconds 72°C, 2 minutes 4℃ throughout
[0185] Next, purify the PCR product using magnetic beads. Add 30–40 μL of beads (0.6X–0.8X) and allow the DNA to bind for 10 minutes. Wash twice with 150 μL of 80% EtOH (30 seconds for each wash), air dry for 6 minutes, remove from the magnet, add 13 μL of 10 mM Tris pH 8 to elute the PCR product, allow to stand for 5 minutes, return to the magnet, and transfer 12 μL to a new tube (use 1 μL for Qubit).
[0186] Ribo Removal
[0187] →RNP formation 2 μL of 10X Cas9 buffer (buffer(TM) r3.1) 3.9 μL of human ribo-guide RNA 2.3 μL of Cas9 Nuclease Streptococcus pyogenes (20 μM) 1 μL of RNaseOUT (40 U / μL) 9.2 μL total volume
[0188] Incubate the above mixture at room temperature for 10 minutes. Add the mixture to the tube containing the DNA, approximately 11 μL in volume from above, and mix by pipetting up and down or flicking. Incubate at 40°C for 1 hour. Add 1 μL of proteinase K (20 mg / mL) and incubate at 56°C for 10 minutes.
[0189] Purify the reaction product using magnetic beads. Add 30 μL of NFW and 30 μL of beads (0.6X). Allow the DNA to bind for 10 minutes. Wash twice with 150 μL of 80% EtOH (30 seconds for each wash). Remove from the magnet and air dry for 6 minutes. Elute the PCR product by adding 24 μL of 10 mM Tris pH 8. Allow to stand for 5 minutes and return to the magnet. Transfer 23 μL to a new tube.
[0190] PCR2:
[0191] Add the following reagents to a PCR tube (ensure to use the same barcode as PCR1): 23 μL from above 25μL of 2X Equinox (Watchmaker Genomics) 1.25 μL of PCR primer p5 (25 μM) - use p5 after the barcode (bc) 1.25 μL of PCR primer p7 (25 μM) (use p7 after bc) 50 μL total volume 100ng input: 7-8 cycles
[0192] PCR protocol →
[0193] Maintain 98°C for 2 minutes →98℃ for 20 seconds → 60℃ for 30 seconds →72℃ for 45 seconds Maintain 72°C for 2 minutes, then maintain 4°C throughout Perform magnetic bead cleanup in PCR reactions. 50 μL of beads (1X)
[0194] This step involves binding the DNA for 10 minutes, washing twice with 200 µL of 80% EtOH (30 seconds between each wash), air drying for approximately 6 minutes, removing from the magnet, and adding 26 µL of 10 mM Tris pH 8 to elute the PCR product. After 5 minutes of incubation, return to the magnet, then transfer 25 µL to a new tube. Read the Qubit on a high-sensitivity tape run or BioA (load 2 ng per sample). Pool the samples (perform a second 0.6X-0.8X cleanup if necessary to remove any dimer products).
[0195] Example 3: cDNA Libraries and Depletion Products - Qualitative Assessment, Comparison 1 In this example, a cDNA library prepared using the protocol described herein is subjected to removal of unwanted non-target sequences for further downstream applications, and the quality of the product is then assessed to determine whether the efficient and rapid process for cDNA preparation described herein impacts the quality of the final product.
[0196] A direct comparison of the disclosed sample preparation protocol with other protocols (kits) from commercial sources showed that the protocol performed equally well and in many ways significantly better. Figure 7 shows the yields for each run, based on the starting concentrations indicated below the graph. Preparation A (using the present protocol) was compared with third-party commercial protocols B and C. Gel electrophoresis was performed using the prepared and stripped libraries as indicated. Yields for A from 100 ng, 50 ng, 10 ng, 1 ng, 0.25 ng, 0.1 ng, and 0.01 ng were 73 ng, 88 ng, 80 ng, 85 ng, 58 ng, 45 ng, and 195 ng, respectively. The average library size was approximately 520 bp, while the library size for C was approximately 400 bp. The respective yields for comparison case C, using the same starting concentrations as A and B, were 83 ng, 115 ng, 170 ng, 232 ng, 67 ng, 37 ng, and 24 ng. As is evident in the electrophoresis data on the right, primer dimerization is a problem for the lower concentration samples (last two), thereby compromising yield in C.
[0197] Figure 8 shows the percent of ribosomal RNA remaining in a sample after cDNA library preparation and removal using the present Protocol A described herein, compared to other commercially available protocols C and D. Figure 8 also shows that Protocol A (the present Protocol) left significantly less ribosomal DNA after removal than B, C, or D, demonstrating that the cDNA preparation and removal protocol A is more effective than other competing protocols.
[0198] Figures 9A and 9B demonstrate the level of sample alignment and % overlap after the removal reaction, comparing Protocol C with Protocol A. Notably, for high-input RNA, the present protocol described herein achieved much higher sequence alignment rates and much lower % overlap at high concentrations with Protocol A compared to Protocol C (Figure 9A). For low-input RNA, the % at least assigned was higher in samples obtained using Protocol A compared to other samples (Figure 9B). Further analysis of the percentages of sequences classified as (a) assigned, (b) unassigned, mapping to multiple regions, (c) unassigned, no feature, and (d) unassigned, ambiguous revealed comparable trends (Figures 10A and 10B). For both high-input and low-input RNA samples, Protocol A was advantageous over the other protocols in that it had a higher percentage of assigned and a lower percentage of unassigned multi-mapped sequences relative to the other compared outputs. Using the Star Alignment program for sequence reads, trends very similar to those in Figures 10A and 10B were obtained, as demonstrated in Figures 11A and 11B, respectively. Furthermore, as shown in Figures 12A and 12B, the number of genes detected in sequence reads confirms the findings shown in Figures 10A and 10B and 11A and 11B, respectively, for high-input and low-input RNA, comparing samples generated using the present Protocol A with samples prepared using other commercially available Protocols B-D.
[0199] Example 4: cDNA Libraries and Depletion Products - Qualitative Assessment, Comparison 2 In a separate comparison set, samples obtained using this protocol were directly compared with different commercially available protocols. A total of 15 samples were generated in this study. The RNA input amounts for three samples (N) were 100 ng, 50 ng, and 10 ng, while the RNA input amounts for 12 samples (A) were 100 ng, 1 ng, 0.25 ng, and 0.01 ng. The respective RNP amounts in Preparation A were 0.6x → 700 bp (fragment size), 0.7x → 500 bp, and 0.8x → 400 bp. Different RNA input amounts were used to compare the performance of Protocol N and to utilize the removal protocol discussed in this disclosure. Protein counts and guide counts were tested for analysis. Overall, the results showed maximum removal with input amounts of 0.6x RNP and 0.01 ng, as well as better performance of the protocol disclosed herein compared to Protocol N.
[0200] When comparing the sequence alignment rate with ribosomal RNA, we found that the protocol A was significantly superior because the depleted library samples from Protocol A contained significantly fewer ribosomal sequences and significantly more other sequences, resulting in a higher ratio of (non-depleted) non-ribosomal sequences to (depleted) ribosomal sequences in A compared to N (Figure 13). The high read percentages of alignments in the two cases remained comparable (Figure 14). In general, at high input RNA concentrations, the percentage of overlap was comparable between the two protocols (Figure 15). Similarly, the average coverage rate was slightly higher for the disclosed protocol (A) at high input RNA concentrations compared to the commercial protocol N (Figure 16). Figures 17 and 18 show library complexity plots at different input RNA levels. In comparison, the N sample had a slightly lower percentage of gene reads during sequencing than the A sample, comparing gene sequences with a copy number of more than 10 in the sample (Figure 17). For preparations using Protocol A, read quality decreased with increasing input RNA concentration, with only slight differences observed between 0.6X (700 bp fragment size), 0.7X (500 bp), and 0.8X (400 bp), as shown in Figure 18. Figure 19 shows that T4 RNA Ligase 1 can be successfully used in place of truncated T4 RNA Ligase 2. T4 RNA Ligase 1, used in combination with phosphorylated adapters, is more cost-effective than truncated T4 RNA Ligase 2 and App-adapters.
[0201] Figures 20-24 show the quality of RNA library preparation from RNA obtained from real biological samples, e.g., liver RNA and cell extract RNA. The results demonstrate that the method described herein yields high-quality sequencing reads in terms of counts and coverage percentages of identified and mapped genes.
[0202] Example 5: Exemplary detailed procedure for generating PCT amplification products Exemplary reactions that encompass some of the processes described herein are described in this section, but should not be construed as limiting the disclosure in any way.
[0203] I. Reverse Transcription Reaction: When working with small amounts of sample (1 ng or less), use a diluted solution of random A-N8 primers [use 1 µL of random A-N8 primers (50 µM) and dilute to 250 µL with nuclease-free water] for any dilutions / transfers. The sequence of an exemplary random A-N8 primer is 5'TCCCTACACGACGCTCTTCCGATCTNNNNNNNN. This acts as a nonspecific blocker to prevent sample loss to the pipette tip or tube wall. To a 2 µL volume of sample (0.1 ng to 100 ng of total RNA), add the following: 0.5 µL of random A-N8 primers (50 µM) and 1 µL of 5X first-strand buffer, bringing the reaction volume to 1.5 µL for a final volume of 3.5 µL.
[0204] The above mixture is incubated at 94°C for 3 minutes, followed by 2 minutes at 4°C or on ice. Then the following is added: 0.5 μL 0.1 M DTT 0.25 μL nuclease-free water 0.25 μL of RNaseOUT (40 U / μL) 0.25 μL 10 mM dNTPs 0.25μL of Maxima H minus (200U / μL) 1.5 μL (5 μL final volume)
[0205] Incubate at 25°C for 15 minutes, 50°C for 50 minutes, and 75°C for 10 minutes, with a final holding temperature of 4°C.
[0206] Add 0.5 µL of exonuclease I (20 U / µL) and 0.5 µL of shrimp alkaline phosphatase (rSAP) (1 U / µL).
[0207] Incubate at 37°C for 30 minutes, then at 95°C for 10 minutes, with a final temperature hold of 22°C. Do not stop here; proceed directly to the next step.
[0208] II. Terminal deoxynucleotidyl transferase (TdT) reaction
[0209] To the 6 μL from above add the following: 1.75 μL nuclease-free water 1 μL 10X TdT buffer 0.25 μL, 1 mM total rNTPs (0.25 mM each of rATP, rCTP, rGTP, and rUTP) 1 μL of terminal deoxynucleotidyl transferase (20 U / μL) 4 μL (10 μL final volume)
[0210] The mixture is incubated at 37°C for 30 minutes.
[0211] III. cDNA adapter ligation reaction
[0212] 10X T4 RNL1 buffer, e.g., 36 μL of 10X T4 RNL buffer).
[0213] Add the following to 10 μL from above and incubate at 28°C for 30 minutes: 2 μL of 10X T4 RNA ligase buffer (containing 10 mM ATP, final 1 mM) 0.5 μL nuclease-free water 0.5 μL of 5'-phosphorylated B adapter (50 μM) 6 μL of 50% PEG8000 (final reaction: 15% PEG8000) 1 μL of T4 RNA ligase 1 (10 U / μL) 10 μL (final volume of 20 μL)
[0214] The reaction from above is incubated for 30 minutes at 28° C. An exemplary 5′-phosphorylated B adapter sequence is / 5Phos / NNNTGGAATTCTCGGGTGCCAAGGAA / 3SpC3 / .
[0215] Ampure Bead Purification:
[0216] To the 20 μL from above, add 30 μL of nuclease-free water and 90 μL of AmpureXP beads (1.8X) and mix by pipetting up and down 10 times. Allow the cDNA / RNA to bind to the AmpureXP beads for 10 minutes at room temperature. Place the tube in a magnet and allow the solution to clear. Once the solution is clear, remove and discard the supernatant.
[0217] Without removing the tube from the magnet, wash the AmpureXP beads twice with 200 μL of 80% ethanol solution. Allow to stand for 30 seconds between washes. Remove and discard the ethanol wash. Care should be taken to remove all of the 80% ethanol solution after the second wash.
[0218] Remove the tube from the magnet and allow the AmpureXP beads to air dry for 8 minutes, leaving the tube cap open.
[0219] To elute the cDNA / RNA, add 24 μL 10 mM Tris pH 8 and incubate for 5 minutes. Place the tube back in the magnet and allow the solution to clear. Transfer 23 μL of the cDNA / RNA to a new PCR tube.
[0220] First PCR: To the 23 μL sample from above, add the following: 25μL 2X Equinox buffer 1.25 μL of PCR primer p7 (barcoded) (25 μM) (full-length primer) 1.25 μL of PCR primer p5 (barcoded) (25 μM) (full-length primer) Approximately 50 μL total volume The number of cycles can vary depending on the input (recommended: 5 cycles for 0.25 ng to 100 ng, 10 cycles for less than 0.25 ng). PCR protocol 98°C for 2 min, 98°C for 20 s, 55°C for 30 s, 68°C for 45 s, 72°C for 2 min, final hold temperature 4°C.
[0221] An exemplary first PCR primer p5-(barcoded) has the following sequence:
[0222] 5'AATGATACGGCGACCACCGAGATCTACACNNNNNNNNACACTCTTTCCCTACACGACGCTCTTCCGATCT (NNNNNNNN represents the barcode sequence). The standard 501-508 8-base barcode or the 10-base 384UDI barcode can be used. An exemplary first PCR primer, p7-(barcoded), is: It has the sequence 5'CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCCTTGGCACCCGAGAATTCCA (where NNNNNNNN represents the barcode sequence). Either the standard 701-712 8-base barcode or the 10-base 384UDI barcode can be used.
[0223] Ampure bead purification after PCR1:
[0224] Add 40 μL of AmpureXP beads (0.8X) to the PCR mixture and pipette up and down 10 times. Allow the PCR product to bind to the AmpureXP beads for 10 minutes at room temperature. Place the tube in a magnet to clarify the solution. Once the solution is clear, remove and discard the supernatant. Without removing the tube from the magnet, wash the AmpureXP beads twice with 150 μL of 80% ethanol solution. Allow to stand for 30 seconds between washes. Discard the ethanol wash solution. Care should be taken to remove all of the 80% ethanol solution after the second wash. Remove the tube from the magnet and, leaving the tube cap open, allow the AmpureXP beads to air dry for 5 minutes.
[0225] To elute the PCR product, add 12 μL* of 10 mM Tris pH 8 and let stand for 5 minutes. Place the tube back in the magnet and allow the solution to clear. At this time, transfer 11 μL* of PCR product to a new PCR tube. If the sample input in step 1 was 1 ng or less, a second AmpureXP bead purification should be performed: elute the PCR product with 51 μL of 10 mM Tris pH 8, transfer 50 μL to a new PCR tube, and perform a second AmpureXP bead wash procedure. Add 40 μL of AmpureXP beads (0.8X) and allow the DNA to bind for 10 minutes. Wash the AmpureXP beads twice with 150 μL of 80% EtOH (30 seconds between each wash) and air dry for 5 minutes. After removal from the magnet, add 12 μL of 10 mM Tris pH 8 to elute the PCR product, let stand for 5 minutes, then return it to the magnet and transfer 11 μL to a new PCR tube).
[0226] Ribo Removal
[0227] In this step, the guide RNAs are incubated at 70°C for 2 minutes, followed by 4°C for 2 minutes, and then these guide RNAs are used directly for RNP formation. →RNP formation 2 μL of 10X Cas9 buffer 3.9 μL of human ribo-guide RNA 2.3 μL Cas9 1 μL RNaseOUT 9.2 μL total volume Incubate at room temperature for 10 minutes. Add the RNP mixture to the tube containing the PCR product, approximately 11 μL in volume from above, and mix by pipetting up and down. Incubate at 40°C for 2 hours* for 1 ng to 100 ng, or for 4 hours for 0.01 ng to 0.25 ng. To quench the reaction, add 1 µL of proteinase K and incubate at 56 °C for 10 min.
[0228] Ampure bead purification after ribosomal removal:
[0229] Add 30 μL of nuclease-free water and 40 μL of AmpureXP beads (0.8X) to the PCR mixture and pipette up and down 10 times. Allow the PCR product to bind to the AmpureXP beads for 10 minutes at room temperature. Place the tube in a magnet and allow the solution to clear. Once the solution is clear, remove and discard the supernatant.
[0230] Without removing the tube from the magnet, wash the AmpureXP beads twice with 150 μL of 80% ethanol solution, with a 30 second interval between washes.
[0231] Remove the tube from the magnet and let it air dry for 5 minutes. To elute the PCR product, add 46 μL of 10 mM Tris pH 8 and let it sit for 5 minutes.
[0232] Place the tube back in the magnet and allow the solution to clear.
[0233] Transfer 45 µL of PCR product to a new PCR tube.
[0234] Second PCR: The second PCR reaction is described below. 45 μL of PCR product from above 50 μL 2X Equinox buffer 2.5 μL of PCR primer p5 (25 μM) (p5 primer only - no barcode) 2.5 μL of PCR primer p7 (25 μM) (p7 primer only - no barcode) 100 μL total volume The number of cycles will vary depending on the input (recommended: 8-9 cycles for 100-50 ng, 10-12 cycles for 40-5 ng, 14-16 cycles for 1-0.25 ng, and 11-14 cycles for <0.25 ng). PCR protocol 98°C for 2 min, 98°C for 20 s, 55°C for 30 s, 68°C for 45 s, 72°C for 2 min, final hold temperature 4°C.
[0235] An exemplary second PCR primer p5 sequence is 5'AATGATACGGCGACCACCGAGATCTACA.
[0236] An exemplary second PCR primer p5 sequence is 5'CAAGCAGAAGACGGCATACGAGA.
[0237] Ampure Bead Purification After PCR2: Add 100 μL of AmpureXP beads (1X) to the PCR mixture and pipette up and down 10 times. Allow the PCR product to bind to the AmpureXP beads for 10 minutes at room temperature. Place the tube in a magnet and allow the solution to clear. Once the solution is clear, discard the supernatant. Without removing the tube from the magnet, wash the AmpureXP beads twice with 200 μL of 80% ethanol solution and allow the AmpureXP beads to air dry for 5 minutes. To elute the PCR product, add 26 μL of 10 mM Tris pH 8, let sit for 5 minutes, and return the tube to the magnet. Once the solution is clear, transfer 25 μL of PCR product to a new PCR tube. Analyze the product using Qubit measurement and Run High Sensitivity Tapestation or BioAnalyzer. If too many dimers are present (with very low input), perform a second 0.7X AmpureXP cleanup.
[0238] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It is understood that various alternatives to the embodiments of the present disclosure described herein may be utilized in implementing the present disclosure. The following claims define the scope of the disclosure, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. A method for preparing a cDNA library, comprising: (A) obtaining a plurality of RNA molecules from a sample; (B) generating first strand cDNAs complementary to the RNA molecules from the plurality of RNA molecules using random primers and single-stranded first adapters; (C) incorporating a ribonucleotide tail onto the 3' end of the first strand cDNA using a mixture of ribonucleotide bases; (D) generating the first strand cDNA with single-stranded adapters at both ends by adding a single-stranded second adapter to the 3' end of the ribonucleotide tail using an enzyme that requires a 3' ribonucleotide terminus as an acceptor for ligation of the single-stranded second adapter; (E) preparing the cDNA library by amplifying the first strand cDNA.
2. The method of claim 1 , wherein the sample is a biological sample.
3. 3. The method of claim 2, wherein the biological sample is a fresh biological sample, a frozen biological sample, or a forensic sample.
4. 2. The method of claim 1, wherein the step of obtaining the plurality of RNA molecules comprises extracting total RNA from the sample.
5. The method of claim 1 , wherein the single-stranded first adaptor is a universal adaptor.
6. 10. The method of claim 1, further comprising the step of removing unused primers with an exonuclease and a phosphatase after generating the first strand cDNA.
7. 2. The method of claim 1, wherein the step of incorporating a ribonucleotide tail comprises incorporating less than 10 ribonucleotides at the 3' end of the first strand cDNA.
8. 8. The method of claim 7, wherein the ribonucleotide tail is incorporated using terminal transferase (TdT).
9. 2. The method of claim 1, wherein adding the single-stranded second adaptor to the 3' end of the ribonucleotide tail comprises adding the single-stranded second adaptor to the 3' terminal ribonucleotide incorporated in step (C).
10. 2. The method of claim 1, wherein amplifying comprises performing a polymerase chain reaction (PCR) using primers that anneal to the first adapter and the second adapter to generate amplified double-stranded cDNA.
11. 2. The method of claim 1, wherein the cDNA library obtained from step (E) is a first cDNA library containing at least one cDNA molecule containing a target sequence and at least one cDNA molecule containing an unwanted non-target sequence, and the method further comprises the step of removing a subset of the amplified first cDNA library containing the unwanted non-target sequence from the amplified first cDNA library.
12. 12. The method of claim 11, wherein the step of removing the subset of the amplified first cDNA library containing the unwanted non-target sequences from the amplified first cDNA library is performed using a nucleic acid-guided endonuclease.
13. 13. The method of claim 12, wherein the nucleic acid-guided endonuclease comprises a CAS endonuclease.
14. 11. The method of claim 10, wherein the target nucleotide sequence comprises a pathogen sequence, and wherein removing a subset of the amplified first cDNA library comprises removing non-pathogen host genomic nucleic acids.
15. 15. The method of claim 14, wherein removing a subset of the amplified first cDNA library comprises removing contaminating human nucleic acids.
16. 16. The method of claim 15, wherein the contaminating human nucleic acid is ribosomal nucleic acid.
17. 16. The method of claim 15, wherein the contaminating human nucleic acid is a repetitive nucleic acid sequence.
18. 11. The method of claim 10, wherein the target nucleic acid comprises fetal nucleic acid, and wherein removing the amplified subset of the first cDNA library comprises removing non-target contaminating maternal nucleic acid.
19. 11. The method of claim 10, wherein the target nucleotide comprises a genomic polymorphism, and removing a subset of the amplified first cDNA library comprises removing wild-type nucleic acid sequences.