Method for preparing a normalized nucleic acid sample, kit and device for use in the method
The method normalizes RNA and cDNA samples using oligonucleotide arrays to address sequencing biases, enhancing the detection of less abundant genes and reducing redundant data, thus improving sequencing efficiency and cost-effectiveness.
Patent Information
- Application Number
- JP2025505406
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2023-08-04
- Publication Date
- 2025-08-01
AI Technical Summary
RNA sequencing is biased towards highly expressed housekeeping genes, making it difficult to detect genes specific to a condition and resulting in redundant data generation, which increases processing time and cost, and requires large sequencing amounts to compensate for inefficiencies.
A method involving oligonucleotide arrays to normalize RNA and cDNA samples by redistributing unique sequences to achieve a uniform distribution, using oligo-dT molecules for reverse transcription and subsequent separation of abundant and less abundant sequences.
The method reduces variability in nucleic acid levels by at least 10-90%, achieving a more uniform sequence distribution, allowing for more efficient sequencing and analysis of less abundant genes, thereby reducing processing time and storage needs.
Smart Images

Figure 2025525100000001 
Figure 2025525100000002 
Figure 2025525100000003
Abstract
Description
Technical Field
[0001] The present invention relates to methods and apparatuses for preparing processed RNA and DNA samples. The present invention also relates to the detection of target nucleic acids. Methods for analyzing biological samples using the processing and detection methods are also provided.
Background Art
[0002] RNA sequencing has become a powerful tool for understanding biology (Stark, R., Grzelak, M. and Hadfield, J., RNA sequencing: the teenage years. Nat. Rev. Genet. 20, pp. 631-656 (2019)). Its applications range from drug development to agricultural improvement. RNA sequencing is typically used to identify differences between biological samples. These samples may be samples from infected and control animals for studying disease resistance, or samples of the same type of sample over time to understand growth and development. The primary results resulting from RNA sequencing are the discovery and quantification of all genes and isoforms expressed in the sample. Most cells and tissues share many of the same highly expressed genes, commonly known as housekeeping genes. These genes typically perform basic cellular functions and thus do not give cells specific characteristics. Since these housekeeping genes typically occupy most of the RNA within the sample, RNA sequencing data is usually biased towards sequencing reads from this uninformative RNA. This phenomenon has two main negative effects on obtaining good results from RNA sequencing projects; firstly, it is difficult to detect genes and isoforms specific to the condition in question, and secondly, the data generated is mostly redundant.
[0003] The first major negative effect has two consequences. First, the amount of sequencing necessary to detect the gene of interest must be large enough to handle the inefficiencies of sampling caused by the low relative abundance of the gene of interest. Second, in some cases, low-abundance target genes may simply be impractical to identify. This can be evidenced by the ongoing efforts to annotate the human genome, where new isoforms and genes are regularly reported even after thousands of sequencing projects, and the complete human transcriptome remains poorly understood. Since eukaryotic transcriptomes derive their complexity from alternative splicing that generates combinatorial permutations, the search for new RNAs will be an ongoing endeavor.
[0004] These two consequences ultimately impede scientific progress by limiting the ability of researchers and others to produce ideal results from their sequencing experiments. These consequences also contribute to the impracticality of applying RNA sequencing for broader use. For example, for use in diagnosis and treatment, the tracking of the amount of sequencing required will have constraints in both time and cost.
[0005] The second major negative effect (generation of redundant data) also has two main consequences. The first is that more data requires longer processing times, increasing the overall cost and time of RNA sequencing experiments. These costs are in terms of the additional computational energy required and the working time of bioinformatics researchers tasked with data processing. The second consequence is that redundant data leads to a greater need for storage devices. As sequencing becomes more widespread, data storage devices are becoming an important issue. More efficient data generation that reduces the need for storage devices is necessary for RNA sequencing technology to play more roles.
[0006] To address the problem of high abundance of housekeeping genes that reduce the sample collection efficiency of target genes, cDNA normalization was developed (Alex S. Shcheglov, Pavel A. Zhulidov, Ekaterina A. Bogdanova, D.A.S. Normalization of cDNA Libraries, Nucleic Acids Hybrid. Chapter 5, (2014)). Since RNA sequencing typically relies on the conversion of RNA to double-stranded cDNA, cDNA normalization utilizes the biochemical properties of cDNA to result in a uniform distribution of unique genes and isoforms within a cDNA library. Theoretically, the maximum non-target sample collection efficiency occurs when all unique RNA sequences are represented at the same relative abundance. Therefore, the purpose of normalization is to redistribute the cDNA library to meet this criterion as closely as possible.
[0007] Two forms of full-length cDNA normalization have been previously developed: the double-stranded specific nuclease (DSN) method (Zhulidov, P. A. et al., Simple cDNA normalization using kamchatka crab duplex-specific nuclease. Nucleic Acids Res. 32, e37 (2004)) and the hydroxyapatite column method (Andrews-Pfannkoch, C., Fadrosh, D. W., Thorpe, J. and Williamson, S. J. Hydroxyapatite-mediated separation of double-stranded DNA, single-stranded DNA, and RNA genomes from natural viral assemblages. Appl. Environ. Microbiol. 76, 5039-5045 (2010)). Both methods rely on the denaturation and re-hybridization of cDNA strands. Since single-stranded cDNA molecules move around in solution, more abundant sequences are more likely to find their complementary sequences, which are the targets for re-hybridization. Thus, re-hybridization reaches its limit, and the remaining single-stranded cDNA represents the normalized sequence library.
[0008] The difference between the two methods lies in their approaches for isolating the single-stranded cDNA library from the re-hybridized double-stranded cDNA molecules.
[0009] In the DSN method, an enzyme that specifically cleaves double-stranded DNA is used to degrade all double-stranded cDNA in solution. Then, the solution is purified and size-selected for cDNA sequences longer than a certain length. These sequences are then amplified using polymerase chain reaction (PCR).
[0010] In the column method, the modified and re-hybridized cDNA library passes through a heated column filled with hydroxyapatite granules. Hydroxyapatite preferentially binds to larger DNA molecules. The size of the DNA to be bound is controlled by the concentration of the phosphate buffer in which the cDNA library is dissolved. Therefore, the concentration of the phosphate buffer must be specifically adjusted for cDNA molecules within a certain range of sequence lengths. The cDNA is eluted through the column using increasing concentrations of phosphate buffer to extract DNA molecules of increasing size. Since single-stranded cDNA is approximately half the size of the re-hybridized cDNA, the elution of the single-stranded fraction can be controlled when the average cDNA sequence length is known. The resulting solution is intended to be rich in single-stranded cDNA and is then amplified using PCR.
[0011] Since the DSN method uses an enzyme that cleaves all double-stranded cDNA, in theory, low-abundance sequences having segments that match high-abundance sequences can be exhausted. This effect can also increase the possibility of forming PCR chimeras. PCR chimeras are formed when incomplete single-stranded cDNA sequences act as primers for other sequences, so that sequences are combined in a manner that does not occur naturally. PCR chimeras represent false positives for new isoforms and require extremely laborious efforts to distinguish from true alternative isoforms. Confirming PCR chimeras typically requires detailed biochemical assays. Due to both the exhaustion of low-abundance sequences and the increased possibility of PCR chimeras, the DSN method is not suitable for many RNA sequencing applications.
[0012] Since the column method only enables the separation of high-abundance and low-abundance fractions within a narrow size range, it has a significant bias towards longer cDNA sequences. This effect is the loss of representation of longer RNA sequences. Due to this effect, the column method is not suitable for many RNA sequencing applications.
[0013] Accordingly, in view of these problems, the present invention has been conceived.
Summary of the Invention
[0014] RNA or cDNA samples typically center on sequences from highly expressed genes that can negatively affect sample analysis. The inventors have developed methods and apparatuses for preparing processed nucleic acid samples with a more uniform sequence distribution. Using a first nucleic acid sample, a probe set is generated based on the abundance of unique sequences in the sample. Abundant sequences generate more probes. When a second nucleic acid sample is applied to the probes, more of the abundant sequences bind to the probes and these sequences can be separated from the sample. In this way, the present invention enables the normalization of full-length RNA and cDNA. This technique is also compatible with methods for extracting RNA and use in detecting specific target sequences. The RNA and DNA processing of the present invention is also beneficial in methods for analyzing biological samples and diagnostic methods.
[0015] It should be borne in mind that various aspects have been conceived to be advantageously combined and all such combinations are envisioned within the scope of the present invention. It should also be recognized that options described with respect to one area of improvement will, with the necessary modifications, also apply to other areas, such as sample type, nucleic acid type, etc., as required.
[0016] In a first aspect, the present invention is a method for processing nucleic acids, comprising: (i) contacting a nucleic acid sample (e.g., an RNA sample) with an oligonucleotide array (e.g., a DNA, optionally cDNA array), wherein one or more nucleic acid molecules from the nucleic acid sample anneal to one or more oligonucleotides of the oligonucleotide array; (ii) extracting the unannealed nucleic acid molecules, thereby yielding a processed nucleic acid; and providing a method.
[0017] In certain embodiments, the array comprises: (i) contacting an RNA sample with a surface, wherein the surface comprises two or more oligo-dT molecules and two or more RNA molecules from the RNA sample anneal to the oligo-dT molecules; (ii) extending two or more of the oligo-dT molecules by reverse transcription using the annealed RNA molecules as templates to produce a DNA array comprising two or more cDNA molecules; (iii) dissociating the annealed RNA molecules from the cDNA molecules; (iv) removing the RNA sample from the surface; and is produced by a method comprising these steps.
[0018] Thus, in these embodiments, the array comprises two or more oligonucleotides having a sequence comprising oligo-dT and a subsequent cDNA sequence.
[0019] In further embodiments, the array is produced by a method comprising: (i) denaturing a cDNA sample comprising double-stranded cDNA molecules to produce single-stranded cDNA molecules; (ii) contacting the cDNA sample with a surface, wherein the surface comprises two or more oligo-dT molecules and two or more cDNA molecules from the cDNA sample anneal to the oligo-dT molecules; (iii) extending two or more of the oligo-dT molecules using the annealed cDNA molecules as templates to produce a DNA array comprising two or more DNA molecules; (iv) dissociating the annealed cDNA molecules from the DNA molecules; (v) removing the cDNA sample from the surface; and is produced by a method comprising these steps.
[0020] Accordingly, in these embodiments, the array comprises two or more oligonucleotides having a sequence comprising oligo dT and a subsequent DNA sequence.
[0021] According to a related aspect of the present invention, a method for processing a nucleic acid, comprising: (i) contacting a first nucleic acid sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and two or more nucleic acid molecules derived from the first nucleic acid sample annealing to the oligonucleotides of the oligonucleotide array; (ii) using the annealed nucleic acid molecules as templates to extend two or more of the oligonucleotides to produce a DNA array comprising two or more DNA molecules; (iii) dissociating the annealed nucleic acid molecules from the DNA molecules; (iv) removing the first nucleic acid sample from the surface; (v) contacting a second nucleic acid sample with the DNA array, and one or more nucleic acid molecules derived from the second nucleic acid sample annealing to the DNA molecules; extracting non-annealed nucleic acid molecules, thereby producing a processed nucleic acid. A method is provided that includes these steps.
[0022] The nucleic acid is not limited according to the present invention. Any suitable nucleic acid molecule can be processed using the devices, kits, and methods of the present invention. The nucleic acid may be double-stranded or single-stranded. According to all aspects of the present invention, when the nucleic acid molecule is double-stranded, the double-stranded nucleic acid molecule is first denatured to produce a single-stranded nucleic acid molecule.
[0023] The nucleic acid may be DNA. The DNA may be genomic DNA, mitochondrial DNA, cDNA, etc. cDNA is preferred. The DNA may be purified from any suitable sample. Types of samples include blood samples (especially plasma-derived and serum-derived), saliva, urine, or other body fluids such as lymph. Other types of samples include solid tissues containing frozen tissue or formalin-fixed paraffin-embedded (FFPE) materials. The DNA molecule may be a double-stranded DNA (dsDNA) molecule. In alternative embodiments, the DNA molecule is a single-stranded DNA (ssDNA) molecule. In some embodiments, the ssDNA is already denatured in situ in the original sample. For example, the ssDNA may be purified from FFPE materials. In further embodiments, the nucleic acid sample may contain both ssDNA and dsDNA molecules. For example, in the case of DNA purified from FFPE materials, the DNA may contain both ssDNA and dsDNA. The DNA may be present in the sample or may be derived from cells in the sample. Alternatively, the DNA may be circulating DNA or "cell-free" DNA (cfDNA). Such DNA can be obtained from a range of body fluids including blood samples (especially plasma-derived and serum-derived), saliva, urine, or other body fluids such as lymph.
[0024] The nucleic acid may also be RNA. The RNA may be obtained from the same types of samples as DNA as described above. The RNA may be messenger RNA (mRNA), microRNA (miRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), long non-coding RNA (lncRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), piwi-interacting RNA (piRNA), tRNA-derived small RNA (tsRNA), small rDNA-derived RNA (srRNA), viral RNA, etc. mRNA is preferred.
[0025] Accordingly, the present invention is a method for processing RNA, (i) Contacting a first RNA sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and two or more RNA molecules derived from the first RNA sample annealing to the oligonucleotides of the oligonucleotide array; (ii) Extending two or more of the oligonucleotides using reverse transcription with the annealed RNA molecules as templates to yield a DNA array comprising two or more cDNA molecules; (iii) Dissociating the annealed RNA molecules from the cDNA molecules; (iv) Removing the first RNA sample from the surface; (v) Contacting a second RNA sample with the DNA array, and one or more RNA molecules derived from the second RNA sample annealing to the cDNA molecules; (vi) Extracting unannealed RNA molecules, thereby yielding processed RNA A method is provided that includes the above steps.
[0026] According to a further aspect of the invention, a method for processing cDNA, comprising: (i) Denaturing a first cDNA sample comprising double-stranded cDNA molecules to generate single-stranded cDNA molecules; (ii) Contacting the first cDNA sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and two or more cDNA molecules derived from the first cDNA sample annealing to the oligonucleotides of the oligonucleotide array; (iii) Extending two or more of the oligonucleotides using the annealed cDNA molecules as templates to yield a DNA array comprising two or more DNA molecules; (iv) dissociating the annealed cDNA molecules from the DNA molecules of the DNA array; (v) removing the first cDNA sample from the surface; (vi) denaturing a second cDNA sample containing double-stranded cDNA molecules to generate single-stranded cDNA molecules; (vii) contacting the second cDNA sample with the DNA array, wherein one or more cDNA molecules from the second cDNA sample anneal to the DNA molecules; (viii) extracting non-annealed cDNA molecules, thereby yielding processed cDNA; A method is provided that includes the above steps.
[0027] An "array" means a collection or arrangement of oligonucleotide (DNA, optionally cDNA) molecules linked or bound to a (solid) surface. A plurality of methods for linking oligonucleotides to a surface are available (e.g., amine-modified oligonucleotides covalently bound to activated carboxylate groups or succinimidyl esters, thiol-modified oligonucleotides covalently bound via alkylating reagents such as iodoacetamide or maleimide, digoxigenin NHS ester, cholesterol-TEG, biotin-modified oligonucleotides captured by immobilized streptavidin, etc.), which are well known to those skilled in the art. The linkage may be covalent or non-covalent. The linkage may be direct or indirect. The DNA array may be a cDNA array.
[0028] In certain embodiments, a method for processing RNA reduces the variability in the levels of RNA (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). The method for processing RNA can achieve a more uniform distribution of the RNA sequences. The difference in abundance between the most abundant RNA and the least abundant RNA can be reduced (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). In certain embodiments, the method for processing RNA reduces the number (copy number) of the most abundant (one or more) RNA molecules (1, 10, 100, 1000, or 10,000) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number (copy number) of the most abundant RNA molecules in a (second) RNA sample is reduced by at least 50% in the processed RNA. In further embodiments, the relative abundance of the least abundant (one or more) RNA molecules (1, 10, 100, 1000, or 10,000) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. Thus, the method for processing RNA can be a method for normalizing RNA.
[0029] In certain embodiments, a method for processing cDNA reduces the variability in cDNA levels (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). A method for processing cDNA can achieve a more uniform distribution of cDNA sequences. The difference in abundance between the most abundant cDNA and the least abundant cDNA can be reduced (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). In certain embodiments, a method for processing cDNA reduces the number (copy number) of the most abundant (one or more) cDNA molecules (1, 10, 100, 1000, or 10,000) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number (copy number) of the most abundant cDNA molecules in a (second) cDNA sample is reduced by at least 50% in the processed cDNA. In further embodiments, the relative abundance of the least abundant (one or more) cDNA molecules (1, 10, 100, 1000, or 10,000) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. Thus, a method for processing cDNA can be a method for normalizing cDNA.
[0030] In certain embodiments, a method for processing nucleic acids reduces variability at the nucleic acid level (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). The method for processing nucleic acids can achieve a more uniform distribution of nucleic acid sequences. The difference in abundance between the most abundant nucleic acids and the least abundant nucleic acids can be reduced (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). In certain embodiments, the method for processing nucleic acids reduces the number of molecules (copy number) of the most abundant (one or more) nucleic acid molecules (1, 10, 100, 1000, or 10,000) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number of molecules (copy number) of the most abundant nucleic acid molecules in the (second) nucleic acid sample is reduced by at least 50% in the processed nucleic acids. In further embodiments, the relative abundance of the least abundant (one or more) nucleic acid molecules (1, 10, 100, 1000, or 10,000) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. Thus, the method for processing nucleic acids can be a method for normalizing nucleic acids.
[0031] Theoretically, the maximum non-target sampling efficiency occurs when all unique nucleic acid sequences are represented at the same relative abundance. Thus, the goal of normalization is to redistribute the nucleic acid sample to satisfy this criterion as closely as possible.
[0032] In certain embodiments, the processed RNA, DNA, or nucleic acid is an RNA, DNA, or nucleic acid that is more readily analyzable. Since the relative representation of less abundant sequences is increased, it can be sequenced more efficiently. Thus, the processed RNA, DNA, or nucleic acid can each be a normalized RNA, DNA, or nucleic acid.
[0033] In certain embodiments, the processed RNA comprises RNA sequences having substantially the same levels. For example, the levels of the sequences of the processed RNA vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The processed RNA can be a processed RNA sample from which at least a portion of the most abundant (one or more) sequences (1, 10, 100, 1000, or 10,000) in a second RNA sample have been removed.
[0034] In certain embodiments, the processed cDNA comprises cDNA sequences having substantially the same levels. For example, the levels of the sequences of the processed cDNA vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The processed cDNA can be a processed cDNA sample from which at least a portion of the most abundant (one or more) sequences (1, 10, 100, 1000, or 10,000) in a second cDNA sample have been removed.
[0035] In certain embodiments, the processed nucleic acid comprises nucleic acid sequences having substantially the same levels. For example, the levels of the sequences of the processed nucleic acid vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The processed nucleic acid can be a processed nucleic acid sample from which at least a portion of the most abundant (one or more) sequences (1, 10, 100, 1000, or 10,000) in a second nucleic acid sample have been removed.
[0036] According to a related aspect of the invention, a method for preparing a normalized RNA, comprising (i) Preparing a first RNA sample; (ii) Contacting the first RNA sample with a surface, wherein the surface comprises one or more oligonucleotides and one or more RNA molecules derived from the first RNA sample anneal to the one or more oligonucleotides; (iii) Using one or more of the oligonucleotides as primers to reverse-transcribe one or more RNA molecules to yield one or more cDNA molecules; (iv) Dissociating the one or more RNA molecules from the one or more cDNA molecules; (v) Removing the first RNA sample; (vi) Preparing a second RNA sample; (vii) Contacting the second RNA sample with the surface, wherein one or more RNA molecules derived from the second RNA sample anneal to the one or more cDNA molecules; (viii) Extracting non-annealed RNA molecules for use as a normalized RNA sample and a method is provided.
[0037] In a further aspect, the present invention is a method for preparing a normalized cDNA, comprising: (i) Preparing a first cDNA sample comprising double-stranded cDNA molecules; (ii) Denaturing the first cDNA sample to generate single-stranded cDNA molecules; (iii) Contacting the first cDNA sample with a surface, wherein the surface comprises one or more oligonucleotides and one or more cDNA molecules derived from the first cDNA sample anneal to the one or more oligonucleotides; (iv) using the one or more oligonucleotides as primers and the one or more cDNA molecules as templates to synthesize one or more probe DNA molecules; (v) dissociating the one or more cDNA molecules from the one or more probe DNA molecules; (vi) removing the first cDNA sample; (vii) preparing a second cDNA sample containing double-stranded cDNA molecules; (viii) denaturing the second cDNA sample to generate single-stranded cDNA molecules; (ix) contacting the second cDNA sample with the surface, wherein one or more cDNA molecules derived from the second cDNA sample anneal with the one or more probe DNA molecules; (x) extracting non-annealed cDNA molecules for use as a normalized cDNA sample A method comprising the above steps is provided.
[0038] Normalizing a nucleic acid sample results in the production of a normalized nucleic acid sample. "Normalized" means that the levels of RNA or cDNA sequences in the sample are more equal. To achieve this, the levels of relatively rare or less abundant sequences can be increased and / or the levels of relatively abundant or more abundant sequences can be decreased. In certain embodiments, the normalized RNA or cDNA contains RNA or cDNA sequences having substantially the same levels. For example, the levels of the sequences of the normalized RNA or DNA vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The normalized RNA or cDNA may be a normalized RNA or cDNA sample from which at least a portion of the 10, 100, 1000, or 10,000 most abundant sequences in a second RNA or cDNA sample have been removed. The methods for processing nucleic acids described herein can be methods for equalizing nucleic acid samples.
[0039] The method of the present invention can be employed with both RNA and DNA. However, the use of double-stranded cDNA requires a denaturation step to produce single-stranded DNA molecules. Strand selection can also be employed as part of the processing of double-stranded cDNA. Oligo dT molecules bind only to cDNA strands containing a poly(A) sequence.
[0040] Therefore, a method for processing cDNA, after the last step (step (viii)), (a) contacting non-annealed cDNA molecules with a second oligonucleotide array, the second oligonucleotide array comprising two or more oligo dT molecules linked to a surface, and one or more of the cDNA molecules annealing to the oligo dT molecules; (b) removing non-annealed cDNA molecules from the surface; (c) dissociating the annealed cDNA molecules from the oligo dT molecules for use as a processed cDNA sample may further comprise.
[0041] According to all aspects of the present invention, in certain embodiments, the (one or more) oligonucleotides are DNA molecules. According to all aspects of the present invention, in certain embodiments, the (one or more) oligonucleotides comprise an oligo dT sequence (optionally 2 - 200, 5 - 200, 2 - 100, 5 - 50, 7 - 25, or 12 - 18 nucleotides in length). Thus, in certain embodiments, the (one or more) oligonucleotides are (one or more) oligo dT molecules. An oligo dT molecule means a molecule containing a stretch of deoxythymidine. The oligo dT molecule can be of any length suitable for binding to the poly(A) tail (sequence of adenine nucleotides) of the second strand of messenger RNA or double-stranded cDNA molecule. In certain embodiments, the (one or more) oligo dT molecules are 2 - 100, 5 - 50, 7 - 25, or 12 - 18 nucleotides in length. In further embodiments, the (one or more) oligo dT molecules are at least 2, at least 5, at least 7, at least 12, at least 18, or at least 25 nucleotides in length.
[0042] (One or more) oligonucleotides may be immobilized on a surface. The surface may be two-dimensional such as a glass slide, or three-dimensional such as microbeads or microspheres. According to all aspects of the present invention, in certain embodiments, the surface is one or more beads or spheres, optionally magnetic beads. The method of the present invention may also be performed in a microfluidic flow cell.
[0043] According to all aspects of the present invention, in certain embodiments, the RNA (first RNA sample and / or second RNA sample) comprises full-length RNA.
[0044] In some embodiments, according to all aspects of the present invention, the surface comprises two or more oligonucleotides, which are optimally spaced such that the DNA molecules to be primed do not interact with each other. Thus, in certain embodiments, the oligonucleotides are optimally spaced such that the DNA (cDNA) molecules of the DNA array do not interact with each other. The optimal spacing for a given sample type may be determined based on the length of the DNA (cDNA) molecules expected to be generated. This may similarly be determined by the (maximum) length of the RNA molecules in the first RNA sample or biological sample or the cDNA molecules in the first cDNA sample or the nucleic acid molecules in the (first) nucleic acid sample. In certain embodiments, the spacing between oligonucleotides is at least 1-fold, at least 1.1-fold, at least 1.2-fold, at least 1.3-fold, at least 1.4-fold, at least 1.5-fold, at least 1.6-fold, at least 1.7-fold, at least 1.8-fold, at least 1.9-fold, at least 2-fold, at least 2.5-fold, at least 3-fold, at least 4-fold, or at least 5-fold the (maximum) length of the RNA molecules in the first RNA sample or biological sample or the cDNA molecules in the first cDNA sample or the nucleic acid molecules in the (first) nucleic acid sample. The spacing between oligonucleotides may be between 1 and 5-fold, between 1.3 and 3.5-fold, between 1.4 and 3-fold, or between 1.5 and 2.5-fold the (maximum) length of the RNA molecules in the first RNA sample or biological sample or the cDNA molecules in the first cDNA sample or the nucleic acid molecules in the (first) nucleic acid sample. In certain embodiments, the spacing between oligonucleotides is 2-fold the (maximum) length of the RNA molecules in the first RNA sample or biological sample or the cDNA molecules in the first cDNA sample or the nucleic acid molecules in the (first) nucleic acid sample. In certain embodiments, the spacing between oligonucleotides is at least 2-fold the (maximum) length of the RNA molecules in the first RNA sample.
[0045] In a further embodiment, the oligonucleotides are optimally spaced when the density of the oligonucleotides (of the oligonucleotide array) is between 0.01 oligonucleotides per square micrometer and 10,000 oligonucleotides per square micrometer, preferably between 0.1 oligonucleotides per square micrometer and 1,000 oligonucleotides per square micrometer, more preferably between 1 oligonucleotide per square micrometer and 100 oligonucleotides per square micrometer.
[0046] In some embodiments, the first RNA sample and the second RNA sample are derived from the same (biological) sample. Similarly, the first cDNA sample and the second cDNA sample may be derived from the same (biological) sample. The first nucleic acid sample and the second nucleic acid sample may be derived from the same (biological) sample. Thus, from a given sample, for example a blood sample (optionally processed to extract RNA), a portion may be removed to form the first RNA sample, and a further portion may be removed to form the second RNA sample. Similarly, from a given sample, for example a blood sample (optionally processed to generate cDNA), a portion may be removed to form the first cDNA sample, and a further portion may be removed to form the second cDNA sample. Further, from a given sample, for example a blood sample (optionally processed to extract nucleic acids), a portion may be removed to form the first nucleic acid sample, and a further portion may be removed to form the second nucleic acid sample. In a further embodiment, the first RNA sample (or the first cDNA sample or the first nucleic acid sample) and the second RNA sample (or the second cDNA sample or the second nucleic acid sample) are derived from the same species, organism, tissue and / or cell type. In certain embodiments, the first RNA sample (or the first cDNA sample or the first nucleic acid sample) and the second RNA sample (the second cDNA sample or the second nucleic acid sample) are derived from blood.
[0047] According to all aspects of the present invention, in certain embodiments, the method further comprises sequencing processed RNA, processed cDNA, or processed nucleic acid. The method for processing nucleic acid (cDNA, RNA) may be a method for preparing nucleic acid (cDNA, RNA) for sequencing. The sequencing may be RNA or DNA sequencing. In certain embodiments, the RNA is reverse transcribed into cDNA prior to sequencing. The sequencing may detect and / or quantify (target) nucleic acid molecules. Such methods include the processing of the present invention and subsequent sequencing of the processed product, optionally using a next-generation sequencing (NGS) platform. Examples of NGS platforms include Illumina sequencing (such as Hi-Seq and Mi-Seq), SMRT sequencing (Pacific Biosciences), nanopore sequencing, SoLID sequencing, pyrosequencing (e.g., Roche 454), and Ion-Torrent (Thermo Fisher), which are well-known to those skilled in the art.
[0048] The present invention also relates to RNA extraction. Thus, (a) contacting a biological sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and one or more RNA molecules from the biological sample annealing to the oligonucleotides of the oligonucleotide array; (b) removing unannealed sample from the surface; (c) dissociating the annealed (one or more) RNA molecules from the oligonucleotides to obtain an RNA sample is provided.
[0049] In related aspects, (a) A step of contacting a biological sample with a surface, wherein the surface contains one or more oligonucleotides, and one or more RNA molecules derived from the biological sample anneal with the one or more oligonucleotides; (b) A step of removing the unannealed sample; (c) A step of dissociating the one or more RNA molecules from the one or more oligonucleotides to obtain an RNA sample. A method comprising the above steps is provided.
[0050] These methods may be combined with other methods of the present invention to provide an RNA sample. These methods may be performed before step (i) of the methods listed above. In certain embodiments, a portion of the obtained RNA sample forms a first RNA sample and a further portion forms a second RNA sample. The RNA sample can be reverse transcribed into cDNA. In certain embodiments, the (one or more) oligonucleotides include one or more oligo dT molecules. In certain embodiments, the (one or more) oligonucleotides are oligo dT molecules. The oligo dT molecules anneal with mRNA molecules having a poly(A) tail. In further embodiments, the (one or more) oligonucleotides include random or specific sequences to capture a range of RNAs in addition to mRNA. (One or more) custom oligonucleotides may be designed to capture specific target RNA molecules (having complementary sequences). If the RNA molecule does not contain a poly(A) tail, it may be polyadenylated after extraction.
[0051] After extracting the processed RNA or cDNA, the method may further include dissociating the annealed RNA molecules from the cDNA molecules (or the annealed cDNA molecules from the DNA molecules). The dissociated molecules may be removed (optionally processed), leaving the surface containing the cDNA molecules (or DNA molecules). The surface may then be used to process additional RNA or cDNA samples. In certain embodiments, a method for processing RNA further includes, after step (vi), dissociating the annealed RNA molecules from the cDNA molecules, removing the dissociated RNA molecules from the surface, and optionally repeating steps (v) and (vi) with additional RNA samples. In certain embodiments, a method for processing cDNA further includes, after step (viii), dissociating the annealed cDNA molecules from the DNA molecules, removing the dissociated cDNA molecules from the surface, and optionally repeating steps (vi), (vii), and (viii) with additional cDNA samples.
[0052] According to all aspects of the present invention, in certain embodiments, the (one or more) oligonucleotides are at least 5 nucleotides in length, at least 10 nucleotides in length, at least 100 nucleotides in length, at least 200 nucleotides in length, or at least 500 nucleotides in length. The (one or more) oligonucleotides may consist of 5 to 200 nucleotides. The oligonucleotide array or surface may comprise at least 10, at least 100, at least 1000, at least 10000, at least 100000, or at least 1 million oligonucleotides. The oligonucleotide array or surface may comprise oligonucleotides that are at least 1.1-fold, at least 1.2-fold, at least 1.3-fold, at least 1.4-fold, at least 1.5-fold, at least 1.6-fold, at least 1.7-fold, at least 1.8-fold, at least 1.9-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 100-fold, or at least 1000-fold more than the RNA molecules in the first and / or second RNA sample, the cDNA molecules in the first and / or second cDNA sample, or the nucleic acid molecules in the nucleic acid sample. The oligonucleotide array or surface may comprise at least 10, at least 100, at least 1000, at least 10000, at least 100000, or at least 1 million oligonucleotides having unique sequences (i.e., no two identical sequences). The (one or more) oligonucleotides may comprise sequences complementary to 10, 20, 50, 100, 1000, or 10000 of the most abundant RNAs (mRNAs) in a given sample, optionally 10, 20, 50, 100, 1000, or 10000 of the most abundant RNAs (mRNAs) in human blood. The (one or more) oligonucleotides may comprise one or more sequences complementary to the mRNAs encoding human serum albumin, one or more α-globulins (e.g., haptoglobin), one or more β-globulins (e.g., plasminogen), and / or one or more γ-globulins.
[0053] According to all aspects of the present invention, in certain embodiments, the amount of RNA molecules in a (first and / or second) RNA sample, or cDNA molecules in a (first and / or second) cDNA sample, or nucleic acid molecules in a (first and / or second) nucleic acid sample does not exceed the number of oligonucleotides in an oligonucleotide array and / or DNA molecules in a DNA array. In certain embodiments, the amount of RNA molecules in the second RNA sample does not exceed the number of cDNA molecules in the DNA array.
[0054] Biological sample and sample are used interchangeably herein. According to all aspects of the present invention, in certain embodiments, a (biological) sample includes a biological fluid, or a fluid or lysate derived from a biological material. The biological fluid may include blood. In certain embodiments, the blood is processed on the same day as collection, within 72 hours after collection, within 2 weeks after collection, within 4 weeks after collection, or within 4 to 12 months after collection. In certain embodiments, the blood is stored at -80°C before processing. Plasma, and similarly serum samples, are contemplated. In certain embodiments, the sample is a human sample. Sample types include other biological fluids such as saliva, urine or lymph fluid. Other sample types include solid tissues, including frozen tissue or formalin-fixed paraffin-embedded (FFPE) material. These samples may be processed to lyse the cells.
[0055] The RNA may be messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), long non-coding RNA (lncRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), piwi-interacting RNA (piRNA), tRNA-derived small RNA (tsRNA), small rDNA-derived RNA (srRNA), microRNA (miRNA) or viral RNA, etc.
[0056] The present invention also relates to a system or apparatus for performing a method as described herein.
[0057] Accordingly, the present invention is an RNA processing apparatus for generating processed RNA from a biological sample, (i) (a) Two or more oligonucleotides capable of annealing to one or more RNA molecules in the biological sample, (b) A first outlet for removing the unannealed sample, (c) A sample outlet through which a portion of the sample containing one or more RNA molecules can flow after dissociation from the oligonucleotide A first module for receiving the biological sample, comprising (ii) (a) Two or more oligonucleotides capable of annealing to one or more RNA molecules in the sample, (b) An outlet for the processed RNA, from which the processed RNA can be obtained, (c) An RNA outlet for removing one or more RNA molecules A second module for receiving one or more RNA molecules from the first module, comprising and relates to an RNA processing apparatus, wherein the first module and the second module together define a flow path through which the sample can flow.
[0058] The oligonucleotide may contain an oligo dT sequence (optionally 2 - 200, 5 - 200, 2 - 100, 5 - 50, 7 - 25 or 12 - 18 nucleotides in length). Thus, in certain embodiments, the oligonucleotide of the first module and / or the oligonucleotide of the second module is an oligo dT molecule. An oligo dT molecule means a molecule containing a continuous stretch of deoxythymidine. The oligo dT molecule can be of any length suitable for binding to the poly(A) tail (sequence of adenine nucleotides) of the second strand of messenger RNA or double-stranded cDNA molecule. In certain embodiments, the (one or more) oligo dT molecules are 2 - 100, 5 - 50, 7 - 25 or 12 - 18 nucleotides in length.
[0059] In a further embodiment, the oligonucleotide contains a random or specific sequence in addition to the mRNA to capture a range of RNAs. Custom oligonucleotides may be designed to capture specific target RNA molecules (having complementary sequences).
[0060] In certain embodiments, the oligonucleotide (oligo dT molecule) of the first module is linked to a first surface, and the oligonucleotide (oligo dT molecule) of the second module is linked to a second surface. The first module may further include a sample inlet through which the biological sample can enter the first module. In certain embodiments, the first module further includes a first reagent inlet through which a reagent can enter the first module, and / or the second module further includes a second reagent inlet through which a reagent can enter the second module. The RNA processing device may further include temperature control means for regulating the temperature of the first module and / or the second module. In a further embodiment, the first module includes a flow cell, and / or the second module includes a flow cell. In certain embodiments, the oligonucleotides are optimally spaced. The optimal spacing is described above. In certain embodiments, the spacing between the oligonucleotides of the first module and / or the second module is at least twice the (maximum) length of the RNA molecules in the biological sample. In a further embodiment, the oligonucleotides are optimally spaced when the density of the oligonucleotides (linked to the first and / or second surfaces) is between 0.01 oligonucleotides per square micrometer and 10,000 oligonucleotides per square micrometer, preferably between 0.1 oligonucleotides per square micrometer and 1,000 oligonucleotides per square micrometer, more preferably between 1 oligonucleotide per square micrometer and 100 oligonucleotides per square micrometer per square micrometer.
[0061] According to all aspects of the present invention, in certain embodiments, the RNA processing apparatus further includes a third module for receiving the processed RNA, and the third module includes reagents for preparing the processed RNA for sequencing. The RNA processing apparatus may further include a fourth module for receiving the RNA prepared for sequencing, and the fourth module includes sequencing reagents.
[0062] A further aspect of the present invention provides for the use of an RNA processing apparatus as described herein in a method of normalizing RNA.
[0063] The method for processing a nucleic acid may be a method for removing the nucleic acid from a sample. Thus, the method for processing RNA may be a method for removing (abundant) RNA from a second RNA sample. Similarly, the method for processing cDNA may be a method for removing (abundant) cDNA from a second cDNA sample.
[0064] When a surface or an oligonucleotide array includes one or more oligonucleotides having sequences complementary to a target nucleic acid of interest, the target nucleic acid may bind to the one or more oligonucleotides. In this way, the target nucleic acid can be removed from the sample. The target nucleic acid may also be subjected to further processing such as sequencing.
[0065] Accordingly, the present invention provides a method for processing a nucleic acid, the method including the step of contacting a nucleic acid sample with a surface, the surface including one or more oligonucleotides complementary to a target nucleic acid, the one or more oligonucleotides being at least 100 nucleotides in length, and the target nucleic acid annealing to the one or more oligonucleotides.
[0066] In certain embodiments, the (one or more) oligonucleotides are at least 200 nucleotides in length, optionally at least 500 nucleotides in length. In certain embodiments, the surface comprises two or more oligonucleotides. In further embodiments, the (one or more) oligonucleotides are linked to the surface.
[0067] According to all aspects of the invention, in certain embodiments, the (one or more) oligonucleotides complementary to the target nucleic acid are complementary to the full length of the target nucleic acid (or at least 70%, at least 80%, at least 90% of the full length).
[0068] The invention also provides an RNA processing device for generating processed RNA from a biological sample, (i) (a) two or more oligo dT molecules capable of annealing to one or more RNA molecules in the biological sample; (b) a first outlet for removing unannealed sample; (c) a sample outlet through which a portion of the sample containing one or more RNA molecules can flow after dissociation from the two or more oligo dT molecules, comprising a first module for receiving the biological sample; (ii) (a) one or more oligonucleotide molecules complementary to the target RNA in the sample; (b) an outlet for unannealed sample for removing unannealed sample; (c) a target RNA outlet through which target RNA can be obtained after dissociation from the one or more oligonucleotide molecules, comprising a second module for receiving one or more RNA molecules from the first module; comprising, providing an RNA processing device, wherein the first module and the second module together define a flow path through which the sample can flow.
[0069] In certain embodiments, the one or more oligonucleotides complementary to the target RNA in the sample are at least 100 nucleotides in length, preferably at least 200 nucleotides in length, more preferably at least 500 nucleotides in length.
[0070] The target nucleic acid may be derived from an RNA virus. The target nucleic acid may be a bacterial gene (transcribed therefrom) such as an antibiotic resistance gene. The target nucleic acid may be a biomarker for a disease.
[0071] The magnetic beads for use in the claimed method may also be provided in the form of a kit. Accordingly, in a related aspect, the invention provides a kit for processing an RNA sample, (a) one or more magnetic beads to which two or more oligo dT molecules are linked, (b) a hybridization buffer, (c) a reverse transcriptase and provides a kit comprising.
[0072] Any suitable reverse transcriptase may be included in the kit. Suitable buffers are also well known and commercially available.
[0073] A further aspect of the invention provides the use of a kit as described herein in a method of normalizing RNA.
[0074] The invention also provides a kit for processing a DNA sample, (a) one or more magnetic beads to which two or more oligo dT molecules are linked, (b) a hybridization buffer, (c) a DNA polymerase and provides a kit comprising.
[0075] Examples of DNA polymerases include thermostable polymerases such as Taq or Pfu polymerase, and various derivatives of these enzymes. Suitable buffers are also well-known and commercially available.
[0076] A further aspect of the invention provides for the use of a kit as described herein in a method of normalizing cDNA.
[0077] In a further aspect, the invention provides a kit for the detection of a target nucleic acid in a sample, (a) one or more magnetic beads comprising one or more oligonucleotides complementary to the target nucleic acid, wherein the one or more oligonucleotides are at least 100 nucleotides in length, the magnetic beads; (b) a hybridization buffer; and comprising a kit.
[0078] The hybridization buffer may comprise 1M HEPES (pH = 7.5), 5M NaCl and H2O. The kit of the invention may further comprise one or more, up to all, of the dinucleotide triphosphates (dNTPs), MdCl2 and a buffer.
[0079] (One or more) oligonucleotides may comprise sequences complementary to 10, 20, 50, 100, 1000 or 10,000 of the most abundant RNAs (mRNAs) in a given sample, optionally 10, 20, 50, 100, 1000 or 10,000 of the most abundant RNAs (mRNAs) in human blood. (One or more) oligonucleotides may comprise one or more sequences complementary to mRNAs encoding human serum albumin, one or more alpha globulins (e.g. haptoglobin), one or more beta globulins (e.g. plasminogen) and / or one or more gamma globulins.
[0080] The methods of RNA extraction and processing may be combined and incorporated into a pipeline for analyzing biological samples.
[0081] Accordingly, the present invention provides a method for analyzing a biological sample derived from a subject, comprising: (a) extracting RNA from the biological sample; (b) preparing a processed RNA sample; and (c) sequencing the processed RNA. The RNA may be full-length RNA. The biological sample may include a biological fluid, or a fluid or lysate derived from a biological substance. In certain embodiments, the biological sample is a liquid biopsy. In specific embodiments, the biological sample is a blood sample, optionally a human blood sample.
[0082] In certain embodiments, the step of preparing a processed RNA sample includes RNA normalization (reducing the variability in the levels of different RNA sequences in the sample). Accordingly, the processed RNA sample may be a normalized RNA sample. "Normalized" means that the levels of RNA sequences in the sample are more equal. To achieve this, the levels of relatively expressed or less abundant sequences can be increased and / or the levels of relatively expressed or more abundant sequences can be decreased. In certain embodiments, the normalized RNA sample includes RNA sequences having substantially the same level. For example, the levels of the sequences in the normalized RNA sample vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%. The normalized RNA may be a normalized RNA sample in which at least a portion of the 10, 100, 1000, or 10,000 most abundant sequences in the sample have been removed.
[0083]
[0084] The step of preparing the processed RNA sample may include equalizing the RNA sample. Thus, in the processed RNA sample, the relative abundances of all unique RNA sequences can be more equal. For example, the levels of unique sequences in the processed RNA sample can vary by less than 50%, less than 40%, less than 30%, less than 20%, or less than 10%.
[0085] In certain embodiments, the step of preparing the processed RNA sample reduces the variability of the RNA levels (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). The step of preparing the processed RNA sample can achieve a more uniform distribution of RNA sequences. In the processed RNA sample, the difference in abundance between the most abundant RNA and the least abundant RNA can be reduced (e.g., by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%). In certain embodiments, the step of preparing the processed RNA sample reduces the number (copy number) of the most abundant (one or more) RNA molecules (1, 10, 100, 1000, or 10000) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%. In certain embodiments, the number (copy number) of the most abundant RNA molecules in the RNA sample is reduced by at least 50% in the processed RNA. In further embodiments, the relative abundance of the least abundant (one or more) RNA molecules (1, 10, 100, 1000, or 10000) is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% in the processed RNA.
[0086] In certain embodiments, the processed RNA sample is more readily analyzable and can be sequenced more efficiently as the relative representation of rare sequences is increased.
[0087] In certain embodiments, the method includes diagnosing a disease in a subject.
[0088] Accordingly, the present invention provides a method for diagnosing a disease in a subject, (a) extracting RNA from a biological sample derived from the subject; (b) preparing a processed RNA sample; (c) sequencing the processed RNA, comprising the steps.
[0089] Diagnosing means determining that the subject has a disease at the time of testing.
[0090] In further embodiments, the method includes predicting a disease or identifying an increased risk of developing a disease. Accordingly, the present invention provides a method for predicting a disease or identifying an increased risk of developing a disease in a subject, (a) extracting RNA from a biological sample derived from the subject; (b) preparing a processed RNA sample; (c) sequencing the processed RNA, comprising the steps.
[0091] To predict means to determine, at the time of the test, that a subject without the disease has an increased risk of developing the disease. The increased risk can be that the risk is higher than the average risk of the population. The increased risk can be that the risk is higher than a pre-calculated threshold level. The threshold level can be the point at which the benefits from increased monitoring and / or preventive treatment outweigh the refusal of unnecessary interventions. The increased risk can be a percentage of lifetime risk that is higher than 1.5%, 2%, 5%, 10%, 50%, or 75%.
[0092] In still further embodiments, the method includes selecting a treatment for a subject having the disease, predicting the responsiveness of a subject having the disease to a therapeutic agent, and / or determining the clinical prognosis of a subject having the disease.
[0093] Sequencing the processed RNA makes it possible to determine the presence and / or level of one or more RNA molecules. In certain embodiments, the presence of one or more RNA molecules in a processed RNA sample is used to identify whether the subject has the disease. In further embodiments, the level of one or more RNA molecules in a processed RNA sample is used to identify whether the subject has the disease. Comparison with a reference point or reference value may be used to diagnose or predict a clinical condition or outcome. Even a single specific RNA molecule (an RNA molecule having a specific sequence) can be used to diagnose or predict a clinical prognosis or responsiveness to a therapeutic agent, but more RNA molecules (of other specific sequences) may be used to increase specificity and sensitivity or diagnostic or predictive accuracy.
[0094] In certain embodiments, the RNA extracted from the sample includes cell-free RNA.
[0095] In certain embodiments, the step of extracting RNA from the biological sample (a) A step of contacting the biological sample with an oligonucleotide array, wherein the oligonucleotide array comprises two or more oligonucleotides linked to a surface, and one or more RNA molecules derived from the sample anneal to the oligonucleotides of the oligonucleotide array; (b) A step of removing unannealed sample from the surface; (c) A step of dissociating the annealed RNA molecules from the oligonucleotides, thereby producing an RNA sample; comprising.
[0096] The oligonucleotide may comprise one or more oligo dT sequences. In certain embodiments, the oligonucleotide is an oligo dT molecule.
[0097] By extracting RNA from a biological sample, the extracted RNA is produced. In certain embodiments, the step of preparing a processed RNA sample comprises following the step of the (one or more) methods for processing RNA as defined above. In further embodiments, the step of preparing a processed RNA sample comprises taking a portion of the extracted RNA formed by extracting RNA from the biological sample as a first RNA sample, taking a portion of the extracted RNA as a second RNA sample, and following the step of the (one or more) methods for processing RNA as defined above. Thus, in certain embodiments, the step of preparing a processed RNA sample comprises taking a portion of the extracted RNA as a first RNA sample, taking a portion of the extracted RNA as a second RNA sample, and (i) A step of contacting the first RNA sample with an oligonucleotide array, wherein the oligonucleotide array comprises two or more oligonucleotides linked to a surface, and two or more RNA molecules derived from the first RNA sample anneal to the oligonucleotides of the oligonucleotide array; (ii) extending two or more of the oligonucleotides by reverse transcription using the annealed RNA molecule as a template to produce a DNA array comprising two or more cDNA molecules; (iii) dissociating the annealed RNA molecule from the cDNA molecule; (iv) removing the first RNA sample from the surface; (v) contacting a second RNA sample with the DNA array, wherein one or more RNA molecules from the second RNA sample anneal to the cDNA molecules; (vi) extracting unannealed RNA molecules, thereby producing processed RNA; comprising.
[0098] In a further embodiment, the step of preparing a processed RNA sample comprises: (i) contacting the extracted RNA with an oligonucleotide array (e.g., DNA, optionally a cDNA array), wherein one or more RNA molecules from the extracted RNA anneal to one or more oligonucleotides of the oligonucleotide array; (ii) extracting unannealed RNA molecules, thereby producing processed RNA; comprising.
[0099] A method of analyzing a biological sample from a subject may include one or more uses of one or more of the (one or more) RNA processing apparatuses and (one or more) kits of the present invention.
[0100] In certain embodiments, sequencing of the processed RNA includes long-read sequencing.
[0101] A further aspect of the present invention provides for the use of a method, apparatus or kit as described herein in the process of RNA or DNA sequencing, optionally for the discovery of novel RNAs and / or detection of rare RNAs, and further optionally, the sequencing is single cell sequencing.
[0102] A further aspect of the present invention provides for the use of a method, apparatus or kit as described herein in the process of metagenomic sequencing for the discovery of novel microorganisms and / or detection of rare microorganisms.
[0103] A further aspect of the present invention provides for the use of a method, apparatus or kit as described herein in the process of screening a DNA or RNA sample or screening a genetic sample for the presence of an infectious disease.
[0104] A further aspect of the present invention provides for the use of a method, apparatus or kit as described herein in the process of detecting nucleic acid biomarkers, optionally disease biomarkers, and further optionally cancer biomarkers.
[0105] In certain embodiments, according to all aspects of the present invention, the method further includes the step of reporting a result. The result may be in the form of an RNA or DNA sequence, an indication of the presence or absence of a microorganism or disease, and / or an indication of the presence or level of a disease biomarker.
[0106] The above and other aspects of the present invention will now be described in more detail, by way of example only, with reference to the following examples and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0107]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0108] Figure 1 schematically shows an array of oligonucleotides, in this case oligo dT molecules linked to magnetic beads. A first RNA sample is contacted with the magnetic beads, and RNA molecules containing polyA tails anneal to the oligo dT molecules. The oligo dT molecules are extended by reverse transcription using the annealed RNA molecules as templates, resulting in cDNA molecules linked to the beads (DNA array). Abundant RNA molecules (RNA sequences that are more frequently present in the sample) generate more cDNA molecules. The annealed RNA molecules are dissociated from the cDNA molecules, and the first RNA sample is removed from the magnetic beads, leaving the cDNA molecules linked to the beads.
[0109] This stage of the method for generating cDNA molecules linked to beads (DNA array) involves the following steps: (i) adding magnetic beads having oligo dT molecules to a purified RNA solution (the first RNA sample); (ii) heating for 5 seconds to 1 minute between 65°C and 100°C (higher than 70°C, optionally between 70°C and 100°C, between 80°C and 100°C, or between 90°C and 100°C) to remove secondary structure; (iii) cooling to less than 65°C (i.e., 65°C or less, for example, between 30°C and 65°C) for 5 seconds to 1 minute to anneal the oligo dT molecules to the polyA tails of the RNA; (iv) adding reverse transcription substances (reverse transcriptase, dNTPs, etc.); (v) incubating for reverse transcription between 30°C and 80°C (optionally between 30°C and 60°C, for 1 minute to 2 hours); (vi) heating between 80°C and 100°C (for example, about 98°C) for 5 seconds to 1 minute to dissociate the RNA from the cDNA copy; (vii) immobilizing the beads using an external magnet; (viii) removing the liquid solution containing the RNA and.
[0110] Next, the second RNA sample is contacted with the beads containing the ligated cDNA molecules. As shown in Figure 2, RNA molecules from the second RNA sample anneal to the cDNA molecules having complementary sequences. At the stage shown in Figure 1, since abundant RNA molecules generate more cDNA molecules, more of the abundant RNA molecules in the second RNA sample are captured by the cDNA molecules, which also applies to the less abundant RNA molecules. Thus, the RNA molecules that do not anneal to the cDNA molecules have a more uniform distribution of sequences and the few very abundant sequences are no longer central, so the RNA is normalized. Next, the magnetic beads are immobilized and the unannealed RNA molecules are extracted, thereby yielding processed RNA.
[0111] The amount of RNA molecules in the second RNA sample should ideally not exceed the number of DNA molecules in the DNA array (cDNA forest) for each reaction cycle. If the DNA array outnumbers each pass of the RNA, it is ensured that there are sufficient probes to anneal to the highly abundant RNA.
[0112] This stage of the method for yielding processed RNA consists of the following steps: (i) Optionally, adding a fresh aliquot of an RNA solution (the second RNA sample) from the same sample; (ii) Heating between 65°C and 100°C (higher than 70°C, optionally between 70°C and 100°C, between 80°C and 100°C, or between 90°C and 100°C) for 5 seconds to 1 minute to remove secondary structure; (iii) Cooling and incubating between 45°C and 75°C (e.g., about 68°C) for 10 seconds to 8 hours to allow controlled full-length association with the cDNA; (iv) Immobilizing the magnetic beads using an external magnet; (v) Extracting the liquid solution containing the processed (normalized) RNA fraction. (vi) Repeating with a further fixed amount of the RNA solution to obtain more processed RNA (after dissociation of the annealed RNA). This is accompanied by
[0113] The present invention enables normalization of full-length RNA. The advantages of directly analyzing RNA include the fact that there is no need to perform PCR (saving time and reagents and avoiding PCR artifacts), and there is no bias. Nanopore sequencing can directly detect modifications present in RNA (modifications that change the path by which RNA moves through the pore).
[0114] The oligonucleotide array may be linked to any suitable surface, and the present invention is not limited to the use of magnetic beads. For example, the method may also be performed in a microfluidic flow cell.
[0115] Figure 3 schematically shows an array of oligonucleotides, in this case oligo dT molecules (also referred to herein as oligo dT forest) linked to the surface of a flow cell. A first RNA sample flows through the flow cell, and RNA molecules containing polyA tails anneal to the oligo dT molecules. To anneal the oligo dT molecules to the polyA tails of the RNA, the temperature is cooled to less than 65°C (i.e., 65°C or less, optionally between 30°C and 65°C).
[0116] Reverse transcription reagents are added and incubation is carried out. As shown in Figure 4, the oligo dT molecules are extended by reverse transcription using the annealed RNA molecules as templates, resulting in cDNA molecules linked to the surface (DNA array) via the oligo dT sequence. Abundant RNA molecules (RNA sequences that are more frequently present in the sample) generate more cDNA molecules.
[0117] As shown in Figure 5, the annealed RNA molecules are dissociated from the cDNA molecules by heating. Then, as shown in Figure 6, the first RNA sample is flushed out, leaving the cDNA molecules linked to the surface (DNA array or cDNA forest).
[0118] Next, the second RNA sample is contacted with the surface containing the cDNA molecules and incubated at about 68°C. As shown in FIG. 7, RNA molecules from the second RNA sample anneal to the cDNA molecules having complementary sequences. In the step shown in FIG. 4, since abundant RNA molecules generate more cDNA molecules, more of the abundant RNA molecules in the second RNA sample are captured by the cDNA molecules, which also applies to the less abundant RNA molecules. Thus, the RNA molecules that do not anneal to the cDNA molecules have a more uniform distribution of sequences and the few very abundant sequences are no longer central, so the RNA is normalized. The non-annealing RNA molecules flow through, thereby producing processed RNA.
[0119] FIG. 8 illustrates a further step of dissociating the annealed RNA molecules from the cDNA molecules by heating to 98°C. The dissociated RNA is flushed away. The surface containing the cDNA molecules is then reused with further RNA samples to yield more processed RNA.
[0120] As illustrated in FIG. 9, a similar principle can be applied to RNA extraction. A biological sample containing lysed cells, including RNA, DNA, protein, etc., is flowed over an array of oligonucleotides, in this case oligo dT molecules (also referred to herein as an oligo dT forest) linked to the surface of a flow cell. The temperature is cooled to less than 65°C (i.e., 65°C or less, optionally between 30°C and 65°C) so that the oligo dT molecules anneal to the poly A tails of the RNA. The flow cell is flushed and only the annealed RNA remains. The temperature is then raised to between 80°C and 100°C (e.g., 98°C) to dissociate the annealed RNA molecules from the oligonucleotides to obtain an RNA sample. The RNA is then flowed for further processing.
[0121] RNA extraction and RNA processing can be linked through the combination of microfluidic flow cells. One flow cell (also referred to herein as a module or reaction chamber) then extracts the RNA to be processed in a further flow cell (or module). An RNA processing apparatus comprising two flow cells is schematically illustrated in FIG. 10. A biological sample is introduced through the sample inlet of the first flow cell. The biological sample may be a sample of lysed cells containing RNA, DNA, protein, etc. Reagents enter through the reagent inlet and are, for example, buffers and / or RNA stabilizing reagents. The surface of the first flow cell is as shown in FIG. 9, i.e., an array of oligonucleotides, in this case oligo-dT molecules linked to the surface of the flow cell. Both the first and second flow cells include temperature control means (thermal control) for regulating the temperature. In order to anneal the oligonucleotides to the RNA, the temperature can be cooled by the temperature control means to below 65 °C (i.e., 65 °C or less, optionally between 30 °C and 65 °C). The first flow cell includes a first outlet for removing unannealed sample when the first flow cell is flowed through leaving only annealed RNA. By the temperature control means, the temperature can be raised to between 80 °C and 100 °C (e.g., 98 °C) to dissociate the annealed RNA molecules from the oligonucleotides to obtain an RNA sample. The first flow cell includes a sample outlet through which the RNA sample can flow after dissociation from the oligonucleotides. The first flow cell and the second flow cell together define a flow path through which the sample can flow. Thus, the first and second flow cells are connected by a connector, such as a tube, that allows the RNA to flow from the first flow cell to the second flow cell for further processing.
[0122] The RNA sample enters through the RNA sample inlet of the second flow cell. The second flow cell (module) also includes a second reagent inlet through which reagents can enter the second flow cell. The surface of the second flow cell contains an array of oligonucleotides (e.g., oligo dT molecules) linked to the surface of the flow cell. The RNA sample flows through the flow cell and the RNA molecules anneal to the oligonucleotides. Reverse transcription reagents are added through the second reagent inlet. The oligonucleotides are extended by reverse transcription using the annealed RNA molecules as templates, resulting in cDNA molecules linked to the surface (DNA array). The annealed RNA molecules are dissociated from the cDNA molecules by heating using temperature control means. Thus, the temperature control means can heat the RNA molecules up to 98°C. The second flow cell includes an RNA outlet for removing one or more RNA molecules. The RNA outlet allows the RNA sample to flow out, leaving the cDNA molecules linked to the surface.
[0123] Subsequently, a further RNA sample enters the second flow cell, contacts the surface containing the linked cDNA molecules, and is incubated at a temperature between 45°C and 75°C (e.g., about 68°C). RNA molecules from the further RNA sample anneal to the cDNA molecules having complementary sequences. The second flow cell includes an outlet for processed RNA through which unannealed RNA molecules flow, thereby resulting in processed (normalized) RNA.
[0124] When the surface or oligonucleotide array contains one or more oligonucleotides having sequences complementary to the target nucleic acid of interest, the target nucleic acid can bind to the one or more oligonucleotides. In this way, the target nucleic acid can be removed from the sample. The target nucleic acid may also be subjected to further processing such as sequencing. A further flow cell may be included in the above-described RNA processing apparatus that contains one or more oligonucleotides having sequences complementary to the target nucleic acid of interest. Alternatively, the second flow cell may contain one or more oligonucleotides having sequences complementary to the target nucleic acid of interest.
[0125] The target nucleic acid may also be directly extracted from a biological sample. FIG. 11 shows a surface or array containing oligonucleotides (also referred to herein as probe cDNA forest) complementary to the target RNA. The oligonucleotides are at least 100 nucleotides in length. A biological sample having lysed cells, including RNA, DNA, protein, etc., flows over the oligonucleotide array. Incubation between 45° C. and 75° C. (e.g., about 68° C.) enables full-length association of the target RNA with the oligonucleotides. The flow cell is flowed through, leaving only the annealed RNA. The temperature is then raised to between 80° C. and 100° C. (e.g., 98° C.) to dissociate the annealed RNA molecules from the oligonucleotides to obtain the target RNA. The RNA is then flowed for further processing. The remaining RNA, DNA, and protein can then be discarded or further processed.
[0126] Using the methods, devices, and kits (when used for processing RNA) described herein, it is possible to target RNA viruses, bacterial genes such as antibiotic resistance genes, and disease RNA biomarkers. In combination with RNA sequencing, this enables accurate diagnosis. DNA arrays (probe forests) can be reused. Thus, the device can be used as a rapid, reusable screening for viral infections when combined with (nanopore) sequencing or another detection method (such as PCR, LAMP, etc.). The methods, kits, and devices can also be used in AgriTech to monitor crops and livestock for diseases. Sample processing is fast and efficient.
[0127] The methods, devices, and kits described herein can also be used to process DNA. However, the use of double-stranded DNA requires a denaturation step to generate single-stranded DNA molecules. Figure 12 shows a surface or array containing oligonucleotides complementary to the target DNA (also referred to herein as the designed probe cDNA forest). The oligonucleotides are at least 100 nucleotides in length. A biological sample containing lysed cells, including RNA, fragmented DNA, proteins, etc., flows over the oligonucleotide array. The sample is heated between 80°C and 100°C (e.g., 98°C) to dissociate the double-stranded DNA. Then, incubation between 30°C and 75°C (e.g., at about 68°C) allows for full-length association of the target DNA with the oligonucleotides. The flow cell is flushed, leaving only the annealed DNA. The temperature is then raised between 80°C and 100°C (e.g., 98°C) to dissociate the annealed DNA molecules from the oligonucleotides to obtain the target DNA. The target DNA is then flowed for further processing. The remaining RNA, DNA, and proteins are then discarded or further processed.
[0128] Using the methods, devices and kits (when used for processing DNA) described herein, DNA viruses can be targeted for the identification of bacteria (and other microorganisms) for rapid DNA identification as a means for disease and for confirming individual identity. DNA arrays (probe forests) can be reused. Thus, the device can be used as a rapid, reusable screening for virus infection when combined with (nanopore) sequencing or another detection method (PCR, LAMP, etc.). The methods, kits and devices can also be used in Agritech to monitor crops and livestock for disease. Sample processing is fast and efficient.
[0129] When (one or more) oligonucleotides complementary to the target nucleic acid are employed, they may be complementary to the full length of the target nucleic acid (or at least 70%, at least 80%, or at least 90% of the full length). This is different from typical probe-based systems that use only short oligonucleotide sequences to target nucleic acids.
[0130] Within a DNA array (cDNA forest), there is an optimal distance between DNA molecules such that they do not interact with each other. This distance is expected to be optimal if every two points need to be approximately twice the length of the longest cDNA. For example, when the biological sample is (human) blood, the maximum length of RNA is about 5 kb, and the maximum length of cDNA generated therefrom is about 5 kb. Therefore, at least 10 kb (6000 nm) would be the optimal spacing between oligonucleotides in an oligonucleotide array and / or cDNA molecules in a DNA array. When oligonucleotides complementary to the target DNA or RNA are used, the distance between oligonucleotides can be shorter because the oligonucleotide sequences can be designed based on known sequences, thus minimizing interactions. Therefore, when oligonucleotides complementary to the target DNA or RNA are used, the distance between oligonucleotides can be at least 1.1 times, at least 1.2 times, at least 1.3 times, at least 1.4 times, at least 1.5 times, at least 1.6 times, at least 1.7 times, at least 1.8 times, or at least 1.9 times the length of the oligonucleotide.
[0131] As described above, the density of the oligonucleotides in the oligonucleotide array affects the density of the DNA molecules in the DNA array. Thus, one means of preventing the DNA molecules in the DNA array from interacting with each other is to use a certain spacing (i.e., maximum density) of the oligonucleotides in the array as described above. The density of the DNA molecules in the DNA array is also affected by the concentration of the RNA or cDNA molecules in the first RNA or first cDNA sample, respectively. This concentration, if necessary, affects how many oligonucleotides in the oligonucleotide array capture the RNA molecules or cDNA molecules. Similarly, this affects how many DNA molecules are synthesized using the captured RNA or DNA as a template. Thus, the concentration of the RNA or cDNA molecules in the first RNA or first cDNA sample may be adjusted to prevent the DNA molecules in the DNA array from interacting with each other.
[0132] Performance is also based on the ratio between the RNA and the DNA array (cDNA forest). Thermal control and kinetic control are related to optimal performance. A micropump can be used to create laminar or turbulent flow.
[0133] The methods of RNA extraction and processing as described above may be incorporated into a pipeline for analyzing biological samples in combination. The methods (a) extracting RNA from a biological sample; (b) preparing a processed RNA sample; (c) sequencing the processed RNA and
[0134] A blood or other liquid biopsy sample is collected from a subject in a container containing a cell lysis buffer and an RNA stabilization reagent. RNA stabilization reagents are commercially available and include RNAlater® (Sigma-Aldrich) and RNAprotect (Qiagen). Suitable buffers include EDTA, sodium citrate, and ammonium sulfate. Incubation for cell lysis may be, for example, from 1 minute to 3 hours. The sample is then added to the RNA processing device as described above.
[0135] A first RNA extraction (purification) is performed. The first reaction chamber (also referred to herein as a flow cell or module) purifies the solution for RNA using oligonucleotides that can be either oligo dT or random sequences. These oligonucleotides bind to the surface of the chamber to form an oligo-forest. After pouring the sample solution into the first chamber, the chamber is heated somewhere between 30°C and 75°C (optionally between 30°C and 65°C or between 60°C and 65°C) to allow annealing of the RNA to the oligo-forest. After the incubation period, the remaining fluid is drained out. The chamber is then heated above 75°C (e.g., between 80°C and 100°C) to release the remaining annealed RNA. This is then poured into a second reaction chamber.
[0136] The next step is to prepare the processed RNA sample, which in this case is RNA normalization. The second reaction chamber (flow cell, module) has a different oligo forest. The purified RNA is cooled in this chamber to less than 65 °C (i.e., 65 °C or less, for example, between 30 °C and 65 °C), optionally less than 60 °C, to enable annealing to the oligo forest. Then, reverse transcriptase and buffer are added to generate a complementary DNA strand using the oligo forest as a primer. Once reverse transcription is complete, the chamber is heated to between 80 °C and 100 °C (optionally higher than 90 °C) to dissociate the RNA from the cDNA forest. Then, the solution is drained. Then, another sample of purified RNA is poured into the second chamber containing the cDNA forest. The chamber is heated to between 45 °C and 75 °C (optionally between 60 °C and 75 °C) to enable full-length annealing of the RNA to the cDNA.
[0137] After incubation, for further processing (to be sequenced), the unannealed RNA is poured into a collection chamber. Then, the second chamber is heated to between 80 °C and 100 °C (optionally higher than 90 °C) to release the RNA. The dissociated RNA is drained. This process is repeated until an appropriate amount of normalized RNA is generated for sequencing.
[0138] The next stage is the preparation for sequencing. The normalized RNA can then be poured into a further reaction chamber (flow cell, module) for preparing the sequencing library. Depending on the sequencing technology, this may involve ligation of adapters, second-strand synthesis, and / or any other necessary modifications to enable sequencing. Then, the sequencing library is poured into the sequencing chamber for sequencing.
[0139] The next step is sequencing and data processing. During sequencing, raw data can be uploaded to a cloud server for data processing and archiving.
[0140] By thus combining RNA extraction and processing, an apparatus is provided that can be used for the immediate processing of blood and other samples while minimizing the problem of RNA degradation.
[0141] The present invention is not limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention will be apparent to those skilled in the art from the foregoing description and the accompanying drawings. Such modifications are intended to fall within the scope of the appended claims. Further, all embodiments described herein are considered to be widely applicable and can be combined with any other non - conflicting embodiments as necessary.
[0142] Various publications are cited herein, and the disclosures of which are incorporated by reference in their entirety.
Claims
1. A method for processing RNA, comprising: (i) contacting a first RNA sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and two or more RNA molecules derived from the first RNA sample annealing to the oligonucleotides of the oligonucleotide array; (ii) extending two or more of the oligonucleotides using reverse transcription with the annealed RNA molecules as templates to produce a DNA array comprising two or more cDNA molecules; (iii) dissociating the annealed RNA molecules from the cDNA molecules; (iv) removing the first RNA sample from the surface; (v) contacting a second RNA sample with the DNA array, one or more RNA molecules derived from the second RNA sample annealing to the cDNA molecules; (vi) extracting unannealed RNA molecules, thereby producing processed RNA. A method comprising the steps above.
2. A method for processing cDNA, comprising: (i) denaturing a first cDNA sample comprising double-stranded cDNA molecules to produce single-stranded cDNA molecules; (ii) contacting the first cDNA sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and two or more cDNA molecules derived from the first cDNA sample annealing to the oligonucleotides of the oligonucleotide array; (iii) extending two or more of the oligonucleotides using the annealed cDNA molecules as templates to produce a DNA array comprising two or more DNA molecules; (iv) dissociating the annealed cDNA molecules from the DNA molecules of the DNA array; (v) removing the first cDNA sample from the surface; (vi) denaturing a second cDNA sample comprising double-stranded cDNA molecules to produce single-stranded cDNA molecules. (vii) contacting the second cDNA sample with the DNA array, wherein one or more cDNA molecules derived from the second cDNA sample anneal to the DNA molecules; (viii) extracting non-annealed cDNA molecules, thereby yielding processed cDNA; A method comprising the steps of: **Claim 3** After step (viii), (a) contacting the non-annealed cDNA molecules with a second oligonucleotide array, the second oligonucleotide array comprising two or more oligo dT molecules linked to a surface, wherein one or more of the cDNA molecules anneal to the oligo dT molecules; (b) removing the non-annealed cDNA molecules from the surface; (c) dissociating the annealed cDNA molecules from the oligo dT molecules for use as a processed cDNA sample. The method according to claim 2, further comprising the steps of: (ix) contacting the non-annealed cDNA molecules with a second oligonucleotide array, the second oligonucleotide array comprising two or more oligo dT molecules linked to a surface, wherein one or more of the cDNA molecules anneal to the oligo dT molecules; (x) removing the non-annealed cDNA molecules from the surface; (xi) dissociating the annealed cDNA molecules from the oligo dT molecules for use as a processed cDNA sample. (xii) The method according to claim 1, wherein the oligonucleotide comprises an oligo dT sequence. (xiii) The method according to claim 1, wherein the surface is one or more magnetic beads. (xiv) The method according to claim 1, wherein the method is performed in a microfluidic flow cell. (xv) The method according to claim 1, wherein the first RNA sample and / or the second RNA sample comprises full-length RNA. (xvi) The method according to claim 1, wherein the oligonucleotides are optimally spaced so that the DNA (cDNA) molecules of the DNA array do not interact with each other. (xvii) The method according to claim 1, wherein the first RNA sample and the second RNA sample (or the first cDNA sample and the second cDNA sample) are derived from the same sample. (xviii) The method according to claim 1, further comprising sequencing the processed RNA or the processed cDNA. (xix) Before step (i), (xx) fragmenting the first RNA sample and the second RNA sample (or the first cDNA sample and the second cDNA sample). (xxi) The method according to claim 1, wherein the fragmenting is performed by a method selected from the group consisting of sonication, enzymatic digestion, and chemical cleavage. (xxii) The method according to claim 1, wherein the fragmenting results in fragments having an average length of about 100 to about 1000 nucleotides. (xxiii) The method according to claim 1, wherein the fragmenting is performed under conditions that do not substantially degrade the RNA or cDNA. (xxiv) Before step (i), (a) contacting a biological sample with an oligonucleotide array, wherein the oligonucleotide array comprises two or more oligonucleotides linked to a surface, and one or more RNA molecules derived from the biological sample anneal to the oligonucleotides of the oligonucleotide array; (b) removing unannealed sample from the surface; (c) dissociating the annealed RNA molecule(s) from the oligonucleotide to obtain an RNA sample The method according to any one of claims 1 to 10, further comprising.
12. After step (vi), dissociating the annealed RNA molecule from the cDNA molecule, removing the dissociated RNA molecule from the surface, and optionally repeating steps (v) and (vi) using a further RNA sample. The method according to any one of claims 1 or 4 to 11, further comprising.
13. The method according to claim 11 or 12, wherein the biological sample comprises a biological fluid, or a fluid or lysate derived from a biological material.
14. An RNA processing apparatus for generating processed RNA from a biological sample, (i) (a) two or more oligonucleotides capable of annealing to one or more RNA molecules in the biological sample; (b) a first outlet for removing unannealed sample; (c) a sample outlet through which a portion of the sample containing one or more RNA molecules can flow after dissociation from the oligonucleotide A first module for receiving the biological sample, comprising; (ii) (a) two or more oligonucleotides capable of annealing to one or more RNA molecules in the sample; (b) an outlet for the processed RNA, from which the processed RNA can be obtained; (c) an RNA outlet for removing one or more RNA molecules A second module for receiving the one or more RNA molecules from the first module, comprising; comprising An RNA processing apparatus, wherein the first module and the second module together define a flow path through which a sample can flow.
15. The RNA processing apparatus according to claim 14, wherein the oligonucleotide of the first module is linked to a first surface, and the oligonucleotide of the second module is linked to a second surface.
16. The RNA processing apparatus according to claim 14 or 15, wherein the first module further comprises a sample inlet through which the biological sample can enter the first module.
17. The RNA processing apparatus according to any one of claims 14 to 16, wherein the first module further comprises a first reagent inlet through which a reagent can enter the first module, and / or the second module further comprises a second reagent inlet through which a reagent can enter the second module.
18. The RNA processing apparatus according to any one of claims 14 to 17, further comprising temperature control means for adjusting the temperature of the first module and / or the second module.
19. The RNA processing apparatus according to any one of claims 14 to 18, wherein the first module comprises a flow cell, and / or the second module comprises a flow cell.
20. The RNA processing apparatus according to any one of claims 14 to 19, wherein the oligonucleotides are optimally spaced apart.
21. A method for processing a nucleic acid, comprising the step of contacting a nucleic acid sample with a surface, the surface comprising one or more oligonucleotides complementary to a target nucleic acid, the one or more oligonucleotides being at least 100 nucleotides in length, and the target nucleic acid annealing to the one or more oligonucleotides.
22. The method according to claim 21, wherein the one or more oligonucleotides are at least 200 nucleotides in length, optionally at least 500 nucleotides in length.
23. An RNA processing apparatus for generating processed RNA from a biological sample, (i) (a) two or more oligo-dT molecules capable of annealing to one or more RNA molecules in the biological sample; (b) a first outlet for removing unannealed sample; (c) a sample outlet through which a portion of the sample containing one or more RNA molecules can flow after dissociation from the two or more oligo-dT molecules a first module for receiving the biological sample, comprising (ii) (a) one or more oligonucleotide molecules complementary to the target RNA in the sample; (b) an outlet for unannealed sample for removing unannealed sample; (c) a target RNA outlet from which the target RNA can be obtained after dissociation from the one or more oligonucleotide molecules; a second module for receiving the one or more RNA molecules from the first module, comprising comprising an RNA processing device, wherein the first module and the second module together define a flow path through which a sample can flow. **Claim 24** The RNA processing device according to claim 23, wherein the one or more oligonucleotides complementary to the target RNA in the sample are at least 100 nucleotides in length, preferably at least 200 nucleotides in length, more preferably at least 500 nucleotides in length. **Claim 25** A kit for processing an RNA sample, comprising (a) one or more magnetic beads to which two or more oligo-dT molecules are linked; (b) a hybridization buffer; (c) a reverse transcriptase comprising **Claim 26** A kit for processing a DNA sample, comprising (a) one or more magnetic beads to which two or more oligo-dT molecules are linked; (b) a hybridization buffer; (c) a DNA polymerase comprising **Claim 27** A kit for detecting a target nucleic acid in a sample, comprising (a) one or more magnetic beads comprising one or more oligonucleotides complementary to the target nucleic acid, wherein the one or more oligonucleotides are at least 100 nucleotides in length; (b) a hybridization buffer comprising **Claim 28** Dinucleotide triphosphate (dNTP), MdCl 2 The kit according to any one of claims 26 to 28, further comprising one or more, up to all, of buffer. **Claim 29** A method for analyzing a biological sample from a subject, comprising (a) extracting RNA from the biological sample; (b) preparing a processed RNA sample; (c) sequencing the processed RNA comprising **Claim 30** The method according to claim 29, wherein the RNA is full-length RNA. **Claim 31** The method according to claim 29 or 30, wherein the biological sample comprises a biological fluid, or a fluid or lysate derived from a biological substance, and optionally, the biological fluid comprises a blood sample.
32. The method according to any one of claims 29 to 31, comprising the step of diagnosing a disease in the subject.
33. The method according to any one of claims 29 to 32, wherein the RNA extracted from the sample comprises cell-free RNA.
34. The method according to claim 32 or 33, wherein the presence or absence of one or more RNA molecules in the processed RNA sample is used to identify whether the subject has a disease.
35. The step of extracting RNA from the biological sample is (a) contacting the biological sample with an oligonucleotide array, the oligonucleotide array comprising two or more oligonucleotides linked to a surface, and one or more RNA molecules derived from the sample annealing to the oligonucleotides of the oligonucleotide array; (b) removing the unannealed sample from the surface; (c) dissociating the annealed RNA molecules from the oligonucleotides, thereby generating an RNA sample. The method according to any one of claims 29 to 34, comprising the steps.
36. The method according to claim 35, wherein the oligonucleotide comprises one or more oligo-dT sequences.
37. The step of preparing a processed RNA sample comprises taking a portion of the extracted RNA as a first RNA sample, taking a portion of the extracted RNA as a second RNA sample, and following the steps of the method according to any one of claims 1 or 4 to 10. The method according to any one of claims 29 to 36, comprising the steps.
38. The method according to any one of claims 29 to 37, comprising the use of an RNA processing device according to any one of claims 14 to 20, 23 or 24, or a kit according to any one of claims 25 to 28.