Sequencing and analysis of exosome-bound nucleic acids

JP7911872B2Active Publication Date: 2026-08-27EXOSOME DIAGNOSTICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022076235
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-07-25
Filing Date
2022-05-02
Publication Date
2026-08-27
Estimated Expiration
2037-10-23

Smart Images

  • Figure 0007911872000007
    Figure 0007911872000007
  • Figure 0007911872000008
    Figure 0007911872000008
  • Figure 0007911872000009
    Figure 0007911872000009
Patent Text Reader

Abstract

The present invention provides a series of steps to prepare nucleic acids (RNA and / or DNA) isolated from extracellular endoplasmic reticulum for sequencing. This allows a wide variety of RNA and / or DNA to be efficiently detected, which can then be used to recognize various attributes such as gene expression, alternative splicing, and the detection of both somatic and germline mutations, including single nucleotide variants (SNVs) and structural variants (insertion / deletions, fusions, inversions).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 410,974, filed on October 21, 2016, and U.S. Provisional Patent Application No. 62 / 536,545, filed on July 25, 2017, the contents of which are hereby incorporated by reference in their entirety.

[0002] Field of the Invention The present invention is in the technical field of molecular biology. More particularly, the present invention is in the technical field of molecular diagnostics. In molecular biology, molecules in the form of nucleic acids such as RNA and DNA are isolated from human-derived sample materials such as tissues and various biological fluids and further analyzed by various methodologies.

Background Art

[0003] Background A wide range of nucleic acid sequencing, including RNA sequencing, of exosomes and other extracellular vesicles from biological fluids holds great promise for highly sensitive diagnostics and the resulting detection of diseases, patient stratification, and observation of treatment responses. The term "exosome" as used herein refers to any extracellular membrane-bound vesicle released by cells.

[0004] There is also a fundamental lack of understanding regarding the long-chain RNA cargo within exosomes isolated from either in vitro or ex vivo systems. Conventional studies investigating the RNA cargo of exosomes have mainly focused on the small RNA fraction. A relatively small population of annotated long-chain RNAs and low transcript coverage rates have been reported in these studies, and many have concluded that exosomes carry only short fragments of protein-coding and non-coding RNAs, raising questions about their potential functional capabilities in gene expression regulation and intercellular communication through exosomes.

[0005] Current methods for isolating DNA and / or nucleic acids (including at least RNA) from the extracellular endoplasmic reticulum include ultracentrifugation, ultrafiltration using, for example, a 100 kDa filter, polymer precipitation techniques, and / or size-based filtration. However, there is a need for efficient and effective alternative methods for isolating the extracellular endoplasmic reticulum and, optionally, extracting the nucleic acids contained therein, preferably extracellular endoplasmic reticulum RNA, and sequencing the nucleic acids contained therein for use in a variety of applications, including diagnostic purposes.

[0006] Therefore, there is a need for reliable sequencing and analysis of nucleic acids using extracellular endoplasmic reticulum. This disclosure addresses these and other important objectives. [Overview of the project] [Problems that the invention aims to solve]

[0007] Summary of the Invention The present invention relates to a method for sequencing nucleic acids from a biological sample, comprising: preparing the biological sample; contacting a solid capture surface with the biological sample under conditions sufficient to retain cell-free DNA and extracellular vesicles from the biological sample on or within the capture surface; contacting a lysis reagent with the capture surface while the cell-free DNA and extracellular vesicles are present on or within the capture surface, thereby releasing DNA and RNA from the capture surface and producing a homogenate; extracting DNA, RNA, or both from the homogenate; and extracting from the homogenate or the extracted... The present invention provides a method comprising: selectively removing ribosomal DNA or RNA sequences from extracted DNA, extracted RNA, or both; reverse transcribing the RNA into cDNA; constructing a double-stranded DNA library from the extracted DNA, the reverse-transcribed cDNA, or both the extracted DNA and the reverse-transcribed cDNA; optionally amplifying DNA, RNA, or both DNA and RNA from the library; selectively enriching nucleic acid sequences from the cDNA or double-stranded DNA library; and sequencing the library containing the cDNA, double-stranded DNA, or both cDNA and double-stranded DNA. [Means for solving the problem]

[0008] In some embodiments, the method further includes a step of pretreatment of the homogenate, extracted RNA, or extracted DNA and RNA with a DNase, such as DNase I and / or modified DNase I, before or after selectively removing the ribosomal DNA or RNA sequence. In other embodiments, the method further includes selectively removing the ribosomal DNA or RNA sequence from RNA, cDNA, or double-stranded DNA in any step during library construction.

[0009] In some embodiments, the method includes simultaneously sequencing both RNA and DNA from a biological sample.

[0010] In some embodiments, the method further includes adding an exogenous RNA or DNA spike to the homogenate and / or the extracted DNA, extracted RNA, or both the extracted DNA and RNA before or after extracting DNA, RNA, or both from the homogenate. In some embodiments, the method further includes pre-treating the homogenate with a DNase, such as DNase I, before or after extracting DNA, RNA, or both from the homogenate, and subsequently adding an exogenous RNA spike to the homogenate. In some embodiments, the method further includes the step of adding exogenous RNA or DNA spikes to the homogenate and / or extracted RNA, extracted DNA, or both extracted DNA and RNA at dilutions of 1:100, 1:1000, 1:10,000, 1:100,000, 1:1,000,000, and 1:10,000,000, before or after extracting DNA, RNA, or both from the homogenate.

[0011] In some embodiments, selective removal of ribosomal RNA, cDNA, and double-stranded DNA may be achieved by using enzymatic reagents such as RNase H or restriction enzyme digests; hybridization-based biotinylated probe enrichment and streptavidin-conjugated paramagnetic beads may be utilized. In some embodiments, selective enhancement of nucleic acid sequences from RNA, cDNA, and / or double-stranded DNA libraries may be achieved by PCR-based approaches, complementary oligonucleotides, and / or hybridization-based biotinylated probe enrichment and streptavidin-conjugated paramagnetic beads. In some embodiments, RNA, cDNA, and / or double-stranded DNA molecules are tagged with unique molecular markers (unique molecular tags, unique identifiers, random barcodes), which enable template identification, deduplication, error repair, and copy number measurement. Unique molecular markers are added via primer annealing, adapter coupling reactions, and enzymatic approaches.

[0012] In some embodiments, the nucleic acid includes long RNA having more than 200 nucleotides, for example, more than 300 nucleotides, or more than 500 nucleotides.

[0013] In some embodiments, the biological sample is a small volume of about 0.5 mL, such as about 0.5 mL to about 20 mL, about 0.5 mL to about 10 mL, about 0.5 mL to about 5 mL, about 0.5 mL to about 4 mL, or about 0.5 mL to about 2 mL. In some embodiments, the biological sample is selected from the group consisting of blood, plasma, serum, urine, saliva, cerebrospinal fluid, pleural fluid, nipple aspirate, lymph, bodily fluids contained in the respiratory tract, intestines, and urogenital tract, tears, saliva, breast milk, fluids derived from the lymphatic system, semen, cerebrospinal fluid, visceral fluid, ascites, tumor cystic fluid, amniotic fluid, and combinations thereof. In some embodiments, the biological sample is blood, plasma, or serum.

[0014] In some embodiments, the solid-capture surface is a membrane, such as a column membrane, or beads. The solid-capture surface may be a membrane containing regenerated cellulose. The solid-capture surface may be a membrane having a pore size in the range of 2 to 5 μm, or at least 3 μm. The solid-capture surface may include two or more types of membranes, such as at least two types of membranes, or at least three types of membranes. The solid-capture surface may include three types of membranes, where the three types of membranes are directly adjacent to each other.

[0015] In some embodiments, the solid-capturing surface is magnetic. In some embodiments, the solid-capturing surface is an ion-exchange (IEX) bead, or a positively charged or negatively charged bead. The solid-capturing surface may be functionalized with a quaternary amine, sulfate, sulfonate, tertiary amine, or a combination thereof. The solid-capturing surface is a quaternary ammonium R-CH2-N + It may also be functionalized with (CH3)3.

[0016] In some embodiments, the solid-capacity trapping surface includes IEX beads, such as strongly ferromagnetic, high-capacity beads, or magnetic, high-capacity IEX beads such as strongly ferromagnetic, high-capacity, iron oxide-containing magnetic polymers. The solid-capacity trapping surface may also include IEX beads that do not have a surface exposed to a liquid prone to oxidation. The solid-capacity trapping surface may also include IEX beads that have a bead charge with a high probability relative to the exposed surface.

[0017] In some embodiments, the extraction step further includes adding a protein precipitation buffer to the homogenate before extracting DNA, RNA, or both DNA and RNA from the homogenate. In some embodiments, the extraction step further includes enzymatic digestion. The extraction step may include proteinase digestion. In some embodiments, the extraction step is carried out with or without prior elution of material from a solid surface. In some embodiments, the extraction step includes digestion using a proteinase, DNase, RNase, or a combination thereof. In some embodiments, the extraction step further includes a protein precipitation buffer containing a transition metal ion, a buffer, or both a transition metal ion and a buffer.

[0018] In some embodiments, the method further includes processing the biological sample by filtering it, for example by filtering it using a 0.8 μm filter. In some embodiments, the method further includes a centrifugation step after contacting the capture surface with the biological sample. In some embodiments, the method further includes washing the capture surface after contacting the capture surface with the biological sample.

[0019] In some embodiments, the method further includes adding nucleic acid control spike in to the homogenate.

[0020] In some embodiments, the method further comprises binding the eluate in which the protein has precipitated to a silica column; and eluting the extract from the silica column. In some embodiments, the method is used for high-throughput separation of nucleic acids from biological samples. In some embodiments, the method includes using one or more chemicals to enhance the binding of small RNAs to a solid surface, such as, for example, isopropanol, sodium acetate, and glycogen at optimal concentrations. In some embodiments, the method utilizes an optimal combination of binding conditions selected from the group consisting of cation concentration, anion concentration, washing agent, pH, time, and temperature, and any combination thereof.

[0021] The various aspects and embodiments of the present invention are described in detail below. It is understood that modifications to these details can be made without departing from the scope of the present invention. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include singulars.

[0022] All patents, patent applications, and documents specified are hereby expressly incorporated by reference herein for the purpose of describing and disclosing, for example, the methods described in the documents that may be used in connection with the present invention. These documents are provided only if they were disclosed prior to the filing date of the present application. It should in no way be construed that the inventors admit that they are not qualified to precede such disclosures for the sake of prior inventions or any other reason. No. All statements regarding the dates and expressions of the contents of these documents are based on the information available to the applicant and do not constitute any approval of the dates or contents of these documents.

Brief Description of the Drawings

[0023] [Figure 1] Figure 1 is a series of Bioanalyzer plots showing RNA profiles resulting from the amplification and incorporation of long-chain RNA transcripts in an RNA Seq library.

[0024] [Figure 2] Figure 2 is a plot showing the correlation of transcripts under various conditions. The use of synthetic spike-ins shows a correlation between replicates of 0.999, demonstrating excellent correlation and reproducibility between library replicates.

[0025] [Figure 3] Figure 3 illustrates (from left to right) the percentages of reads extracted as intergenic, intronic, other genomic, transcriptomic, and unmapped reads, where the X-axis provides the percentage of reads.

[0026] [Figure 4] Also, Figure 4 illustrates (from left to right) the percentages of reads extracted as ERCC, Ig, miscRNA, ncRNA, protein-coding genes, pseudogenes, rRNA, repeats, small RNAs, and tRNAs, where the X-axis provides the percentage of reads.

[0027] [Figure 5] Figure 5 is a graph plotting the annotation distribution by the percentage of reads excluding ribosomal RNA, where the Y-axis indicates the percentage of reads and the X-axis indicates the RNA type.

[0028] [Figure 6] Figure 6 is a graph plotting the ribosomal genes 28S, 18S, 12S, and 16S as a percentage of reads for both without (left side) and with (right side) ribodepletion. <​​​​​​​​​​Figure 8 plots the annotation distribution by transcript count for all transcripts (left) and expanded non-coding RNA (right).

[0031] [Figure 9] Figure 9 is a graph plotting all 86,799 covered transcripts, with the percentage of covered transcripts (exons only) on the X-axis and the proportion of covered transcripts on the Y-axis.

[0032] [Figure 10] Figure 10 is a graph plotting the number of transcripts per million molecules on the Y axis, also on a log2 scale, against the number of molecules on the X axis, on a log2 scale.

[0033] [Figure 11] Figure 11 shows a table of categories of top gene ontology known to be represented by long-chain RNA in plasma exosomes.

[0034] [Figure 12] Figure 12 is a plot demonstrating the ERCC correlation for various experimental conditions using the improved workflow provided herein.

[0035] [Figure 13] Figure 13 is a plot of the variation in 3'→5' coverage in the transcript (left) and the variation in 3'→5' coverage in ERCC spike-in (right), where the X-axis is the normalized distance along the transcript and the Y-axis is the normalized coverage.

[0036] [Figure 14] Figure 14 is a bioanalyzer plot showing the final library size distribution of plasma exosomes.

[0037] [Figure 15] Figure 15 illustrates a typical algorithm implemented according to the improved workflow provided herein.

[0038] [Figure 16] Figure 16 provides an overview of the RNASeq pipeline.

[0039] [Figure 17] Figure 17 shows plots for mapping indices (Figure 17A), UHR as a percentage of base coverage, base coverage for RNA only and RNA+cfDNA (Figure 17B), and readings per total cancer (Figure 17C).

[0040] [Figure 18] Figure 18 shows the mapping index (Figure 18A), the base coverage of cfDNA and cfDNA+RNA as a percentage of base coverage (Figure 18B), and a plot on the depth of coverage per target (Figure 18C).

[0041] [Figure 19] Figure 19 shows plots for mapping indices (Figure 19A), UHR as a percentage of base coverage, base coverage for RNA only and RNA+cfDNA (Figure 19B), and gene reading coverage (Figure 19C).

[0042] [Figure 20] Figure 20 plots three independent RNA-seq library preparation workflows optimized for liquid biopsy samples.

[0043] [Figure 21] Figure 21 demonstrates various ribosomal RNA removal approaches.

[0044] [Figure 22] Figure 22 demonstrates the library preparation methods for all RNA and all nucleic acids (cfDNA + RNA).

[0045] [Figure 23]Figure 23 is a plot demonstrating the detection limits of an ERCC exogenous RNA spike-in-based whole RNA Seq assay for six independent library duplications constructed from plasma.

[0046] [Figure 24] Figure 24 shows the RNASeq browser displaying QC indicators and analysis results.

[0047] [Figure 25] Figure 25 shows a differential expression browser for displaying and evaluating the results of differential expression analysis.

[0048] [Figure 26] Figure 26 is an illustration comparing two uniform processes (Methods 4 and 5) for exosome samples using an improved process that combines these separate flows into a single method (Method 6).

[0049] [Figure 27] Figure 27 annotates coverage rates for unmapped, other genomes, inter-genes, introns, and transcriptomes as a measure of percentage input reads per sample.

[0050] [Figure 28] Figure 28 annotates each of the following per million units: biotype ERCC, contaminants, rRNA, protein-coding genes, ncRNA, small RNA, tRNA, pseudogenes, miscRNA, and Ig.

[0051] [Figure 29] Figure 29 is a plot showing the length (number of nucleotides) of the insertions against density.

[0052] [Figure 30] Figure 30 plots the gene biotype on the X-axis based on the number of genes on the Y-axis.

[0053] [Figure 31] Figure 31 provides plots on the detection threshold RPM per gene detected at or above the threshold for all genes (top panel), useful transcriptome (middle panel), and mRNA (bottom panel).

[0054] [Figure 32] Figure 32 plots the percentage of covered transcripts (exons only) on the X-axis against the percentage of all transcripts on the Y-axis.

[0055] [Figure 33] Figure 33 demonstrates the detection limit of ERCC transcripts.

[0056] [Figure 34] Figure 34 plots the normalization position (5'→3') relative to the transcript against the normalization coverage.

[0057] [Figure 35] Figure 35 annotates the coverage rates of unmapped, other genomes, inter-genes, introns, and transcriptomes as a measure of the percentage input reads per sample.

[0058] [Figure 36] Figure 36 annotates each of the following per million units: biotype ERCC, contaminants, rRNA, protein-coding genes, ncRNA, small RNA, tRNA, pseudogenes, miscRNA, and Ig.

[0059] [Figure 37] Figure 37 is a plot showing the length (number of nucleotides) of the insertions against density.

[0060] [Figure 38] Figure 38 plots the gene biotype on the X-axis based on the number of genes on the Y-axis.

[0061] [Figure 39] Figure 39 provides plots on the detection threshold RPM per gene detected at or above the threshold for all genes (top panel), useful transcriptome (middle panel), and mRNA (bottom panel).

[0062] [Figure 40] Figure 40 plots the percentage of covered transcripts (exons only) on the X-axis against the percentage of all transcripts on the Y-axis.

[0063] [Figure 41] Figure 41 highlights the size of transcripts with >80% coverage, based on the plotting of transcript length relative to the transcript percentage.

[0064] [Figure 42] Figure 42 demonstrates the detection limit of ERCC transcripts.

[0065] [Figure 43] Figure 43 plots the normalization position (5'→3') relative to the transcript against the normalization coverage. [Modes for carrying out the invention]

[0066] Detailed description of the present invention The present invention provides a series of steps for preparing nucleic acids (RNA and / or DNA) isolated from exosomes for sequencing. This allows for a wide diversity of RNA and / or DNA to be efficiently detected. These can then be used to recognize various attributes, such as gene expression, alternative splicing, fusion transcripts, circular RNA, and the detection of both somatic and germline mutations, including single nucleotide variants (SNVs) and structural variants (insertions / deletions, fusions, inversions).

[0067] In one embodiment, the present invention provides the ability to combine multiple workflows, such as separate processing conditions for RNA and DNA, within a single workflow, enabling more effective analysis of exosome samples.

[0068] In one embodiment, the present invention provides a workflow that specifically increases the sample concentration with respect to a target of interest and enables deeper sequence coverage. The workflow provides the ability to target captured cDNA or dsDNA in a particular sample, for example, by increasing the sample concentration for a subset of genes of interest and / or removing other genes not of interest.

[0069] According to one embodiment, the present invention provides a platform specifically designed to include both short and long RNA transcripts from exosomes in an RNA sequencing workflow. As used herein, the term “long RNA” refers to RNA having more than 200 nucleotides, for example, more than 300 nucleotides, or more than 500 nucleotides, and may also include longer non-coding RNA, mRNA, and circular RNA.

[0070] According to one embodiment, the present invention provides a platform for processing DNA either alone or in a mixture with RNA from exosomes (both short and long RNA transcripts) within a sequencing workflow.

[0071] The volume of the biological fluid serving as input for the sequencing workflow can be as small as ≥0.5 ml, and there is no upper limit (Figure 16).

[0072] Nucleic acids are isolated from exosomes and other cell-free sources by starting with biological samples described herein, such as human plasma, serum, blood, urine, or cerebrospinal fluid. Alternatively, the nucleic acids may be derived from tissue sources such as reference standards or FFPE materials.

[0073] Exosome-derived nucleic acids may include RNA or DNA individually or as a mixture of RNA and DNA, as illustrated in Figures 17 (RNA and RNA+DNA) and 18 (DNA and RNA+DNA), which illustrate the sequencing of these nucleic acid combinations. Exosome-derived nucleic acids may include substances contained within exosomes or bound to the external surface of exosomes. The DNA component may be of exosome origin or of other cell-free origin (cfDNA).

[0074] In one embodiment, methods for isolating exosomes for further purification of extracellular vesicles having the bound nucleic acids described herein include: 1) Ultracentrifugation, often combined with a sucrose density gradient or sucrose cushion to suspend relatively low-density exosomes. Isolation of exosomes by sequential fractional centrifugation combined with sucrose gradient ultracentrifugation can provide high exosome enrichment. 2) Use of a volume exclusion polymer selected from the group consisting of polyethylene glycol, dextran, dextran sulfate, dextran acetate, polyvinyl alcohol, polyvinyl acetate, or polyvinyl sulfate; where the molecular weight of the volume exclusion polymer is 1,000 to 35,000 daltons, and is preferably used in combination with an additional 0 to 1 M of sodium chloride. 3) Size exclusion chromatography, e.g., Sephadex® G200 column matrix. 4) Selective immunoaffinity or charge-based capture using paramagnetic beads (including immunoprecipitation reactions) by using antibodies against surface antigens including, but not limited to, EpCAM, CD326, KSA, and TROP1. The selective antibody may be conjugated to paramagnetic microbeads. 5) Direct precipitation using chaotropic reagents, such as guanidinium thiocyanate.

[0075] Following exosome isolation, the sample is subjected to the described method. Briefly, in one embodiment, the workflow begins with exosomal RNA isolation, followed by DNase treatment for application, in which case DNA may interfere with the analysis.

[0076] In another embodiment, following exosomal RNA isolation, the sample is treated with a DNase or left untreated. In some embodiments, DNase treatment is useful for applications where DNA may interfere with the analysis. In other embodiments, the sample is left untreated if the effect of DNA is of interest in the analysis.

[0077] The DNase treatments intended herein include wild-type DNase I, as well as its protein-modified or otherwise altered forms. Commercially available examples, but not limited to, include ArcticZymes: Heat & Run gDNA Removal Kit; New England Biolabs: DNase I; Sigma Aldrich: DNase I; ThermoFisher Scientific: Turbo DNase; and ThermoFisher Scientific: Ambion DNase I.

[0078] In some embodiments, a spike-in of a synthetic RNA or DNA standard, also referred to herein as a “synthetic spike-in,” is performed as a quality control indicator before or after the DNase step, or at any step before sequencing library preparation. Exogenous substances, such as synthetic nucleic acids, can serve as quantification reagents that are sample quality control reagents, enabling detection limit, dynamic range, and technical reproducibility testing, and / or testing for the detection of specific sequencing.

[0079] Commercially available synthetic spike-ins, though not limited to these, include Dharmacon: Solaris RNA spike-in control kit; Exiqon: RNA spike-in kit; Horizon Diagnostics: reference standard; Lexogen: spike-in RNA variant control mix; Thermo Fisher Scientific: ERCC RNA spike-in control mix; and Qbeta RNA spike-in, yeast, or Arabidopsis RNA.

[0080] In some embodiments, the synthetic spikein is added to the sample at various dilutions. In some embodiments, the dilutions of the spikein added to the sample may be in the range of 1:1000 to 1:10,000,000, including dilutions of 1:1000, 1:10,000, 1:100,000, 1:1,000,000, and 1:10,000,000. The specific dilution of the spikein added to the sample is determined based on the amount and / or quality and / or origin of nucleic acids present in the sample.

[0081] Next, the sample may be subjected to a reverse transcription reaction or may be left untreated. In some embodiments, the RNA in the sample is reverse transcribed when there is interest in converting it to cDNA. In some embodiments, when only single-stranded cDNA is desired, only first-strand synthesis is performed. In some embodiments, when double-stranded DNA is desired, both first-strand and second-strand synthesis are performed. In some embodiments, when there is interest only in investigating the DNA content in the sample, the sample is left untreated. In some embodiments, cDNA processing steps include, but are not limited to, preservation of strand information by treatment with uracil-N-glycosylase or orientation of an NGS adapter sequence, RNA cleavage, RNA fragmentation, incorporation of non-standard nucleotides, annealing or ligation of an adapter sequence, and second-strand synthesis.

[0082] In some embodiments, the sample is subjected to fragmentation or not processed. Fragmentation is achieved using enzymatic or non-enzymatic methods, or by physical shearing of the material containing RNA or dsDNA. In some embodiments, fragmentation of RNA and / or dsDNA is carried out by thermal denaturation in the presence of divalent cations. The fragmentation duration for a particular sample is determined based on the amount and / or quality and / or origin of nucleic acids present in the sample. In some embodiments, the fragmentation duration ranges from 0 to 30 minutes.

[0083] In some embodiments, the sequencing adapter is added to the material using a linkage-based approach, following end repair and polyadenylation. In some embodiments, the sequencing adapter is added to the material using a PCR-based approach. Nucleic acids in a sample having the sequencing adapter, obtained via any of the embodiments described above, are referred to here as the “library” when referring to the total recovered nucleic acid fragments in the sample, or as “library fragments” when referring to the nucleic acid fragments incorporated into the contents of the sequencing adapter. Inclusion of a unique molecular index (UMI), unique identifier, or molecular tag within the adapter sequence provides additional benefits in terms of read-duplicate elimination and a more accurate estimate of the number of nucleic acid molecules in the sample.

[0084] In some embodiments, by using a bead-based separation technique, the library can be subjected to a further modified process to: 1) remove undesirable products (including, but not limited to, residual adapters, primers, buffers, enzymes, and adapter dimers); 2) ensure that the library composition is within a specific particle size range (by changing the ratio of beads or bead buffer reagent to sample, low molecular weight and / or high molecular weight products can be included in or excluded from the sample); and 3) concentrate the sample by elution into a minimum volume. This process is generally referred to as the “cleanup” step, or the sample is “cleaned up,” and is referred to as such herein. The bead-based separation technique may include, but is not limited to, paramagnetic beads. The bead-based cleanup may be performed once or multiple times, as required.

[0085] Examples of commercially available paramagnetic beads useful with the method described herein include, but are not limited to, Beckman Coulter: Agencourt AMPure XP; Beckman Coulter: Agencourt RNAclean XP; Kapa Biosystems: Kapa Pure beads; Omega Biosystems: MagBind TotalPure NGS beads; and ThermoFisher Scientific: Dynabeads.

[0086] In some embodiments, the beads are subjected to a hydration step, where the dried beads are covered with a resuspended hydration solution, such as water, particularly nuclease-free water, and incubated for about 1 to 10 minutes, for example, about 5 to 10 minutes, at a temperature in the range of about 20°C to about 40°C, for example, about 20°C to about 25°C. In a preferred embodiment, the beads can be incubated at room temperature for 5 minutes for rehydration.

[0087] In some embodiments, following a bead-based cleanup, the library is amplified en bloc using general-purpose primers targeting adapter sequences. The number of amplification cycles can be adjusted to produce sufficient product for downstream processing steps. In some embodiments, fewer cycles are used to minimize the introduction of potential bias. In some embodiments, more cycles are used to produce a library with higher concentrations of molecules. In some embodiments, any round of qPCR is performed to determine the optimal number of cycles for PCR amplification of the library. Following library amplification, the bead-based cleanup is repeated as described above.

[0088] Next, the quantity and quality of the library are quantified using fluorescence quantification techniques, such as the Qubit dsDNA HS assay and / or the Agilent Bioanalyzer HS DNA assay, although these are not limited to the above.

[0089] In some embodiments, aliquots of a sample are used in a hybridization-based enrichment process (referring to Figures 17–19). This process utilizes the hybridization of a nucleotide probe complementary to the genomic sequence region of interest contained in the sample, followed by a series of washes using a buffer selected for the sequence of interest, during which unwanted material is washed away. The probe-sequence hybrid may, but is not limited thereto, be selected to utilize a streptavidin-biotin chemical reaction. The process can, but is not limited thereto, be used to enrich any portion or mixture of a genomic sequence or transcriptome sequence, including exon regions, untranslated regions (UTRs), intergenic regions, and intron regions, and can cover the location of all intragenic or extragenic gene coding regions or specific hotspots. Hybridization probe panels can be used to increase the concentration of any target sequence from a small number of targets (1 or 20) to many targets (>1,000), including, but not limited to, the total protein-coding transcriptome containing ~20,000 genes (see Figure 4), the total non-coding transcriptome including long non-codings, long intergeneric non-coding repeats, e.g., Alu, HerV, Line, antisense and small non-coding transcriptomes or any combination of the above, large panels targeting a wide range of diseases or disease-related pathways, pathological conditions with >1,000 genes or fewer genes (see Figures 17-18), and medium panels targeting diseases of interest (e.g., solid tumors) or disease-related pathways with 50-500 genes. In some embodiments, the sample is not concentrated, in which case the entire sample is sequenced (Figures 20-22).Representative commercially available hybridization kits include, for example, Agilent's SureSelect Exome V2; ArcherDx's Comprehensive Solid Tumor; Asuragen's Quantidex NGS Pan Cancer Kit; ClonTech's SMARTer Target RNA Capture; IDT's Pan-Cancer Panel; Illumina's TruSight RNA Pan-Cancer Panel; Illumina's TruSight Tumor 170; Illumina's RNA Access; New England BioLabs' NEBNext Direct Cancer HotSpot Panel; NuGEN's Ovation Fusion Panel Target Enrichment System V2; Roche's SeqCap EZ Exome v3.0 Kit; and Roche's Avenio ctDNA Expanded Kit.

[0090] In some embodiments, when the entire sample is being analyzed, ribosome sequences (cDNA, RNA, or cfDNA) may interfere with the detection of small amounts of transcripts, in which case it is desirable to remove or eliminate the sample of ribosome sequences, which is hereafter referred to as “ribodepression” (see Figure 21). Selective removal of abundant but undesirable sequences, including, but not limited to, ribosome sequences and / or globin gene sequences, may be achieved at the RNA sequence level (appropriate when only RNA is isolated and analyzed) or at the dsDNA (library) level (appropriate when cDNA and / or cfDNA are analyzed). Ribosome sequence-specific removal may, but not limited to, be achieved using RNase H or enzymatic reagents similar to restriction enzyme digestion. Removal may also be achieved by utilizing hybridization-based biotinylated probe enrichment and streptavidin-conjugated paramagnetic beads to specifically capture and remove ribosome sequences.

[0091] In some embodiments, ribosome removal may also include one or more additional cycles relating to primer annealing, such as one additional cycle, two additional cycles, three additional cycles, or five or more additional cycles.

[0092] In some embodiments, following a hybridization-based targeted enrichment or ribosome removal process, the residual sample material is amplified using a general-purpose primer that recognizes a sequencing adapter. PCR-based amplification uses enough cycles to produce a sufficient amount of product for subsequent steps without using excessive cycles that could potentially introduce bias into the material.

[0093] In some embodiments, following a hybridization-based targeted enrichment and / or ribosome removal process, residual sample material is cleaned up using a bead-based paramagnetic approach as previously described. Cleaning may occur before, after, or before / after the additional amplification cycle described previously.

[0094] In some embodiments, this is followed by quantitative analysis of the library's quantity and quality using fluorescence quantification techniques such as the Qubit dsDNA HS assay and / or the Agilent Bioanalyzer HS DNA assay, but is not limited to these. The library is then normalized, multiplexed, and subjected to sequencing by any next-generation sequencing platform.

[0095] In some embodiments, the sequencing data is then, if necessary, multiplexed, and transcript / gene counts are produced by mapping to an existing genome or transcriptome reference sequence, or to a de novo-synthesized genome or transcript (see Figures 16 and 24). The UMI tag for each sequence can then be used to identify fragments resulting from PCR replication. The counts are normalized, particularly with respect to library size, GC bias, sequence bias, and sequencing depth. These counts can then be used to perform differential expression analysis, which is gene expression analysis related to various conditions (e.g., tumor / healthy) to create a list of potential biomarkers, not limited to the aforementioned applications where they can be distinguished between sample types (Figure 25). The aligned reference data can be used to profile sequence variations, not limited to single nucleotide polymorphisms, insertions / deletions, fusions, inversions, and repeat extensions. Sample separation

[0096] The present invention provides a method for sequencing and / or analyzing nucleic acids, including at least RNA from the extracellular endoplasmic reticulum, by capturing cell-free DNA and extracellular endoplasmic reticulum on a surface, then lysing the extracellular endoplasmic reticulum to release nucleic acids, particularly RNA contained in the extracellular endoplasmic reticulum, thereby lysing DNA and / or DNA, as well as nucleic acids including at least RNA, from the captured surface.

[0097] Microvesicles are released outside the cell from eukaryotic cells or budding from the cell membrane. These membrane vesicles are heterogeneous in size, ranging in diameter from approximately 10 nm to 5000 nm. In this specification, all membrane vesicles released from cells with a diameter of less than 0.8 μm are collectively referred to as "extracellular vesicles" or "microvesicles." These extracellular vesicles include microvesicles, microvesicular particles, prostasomes, dexosomes, texosomes, ectosomes, oncosomes, apoptotic bodies, retroviral particles, and human endogenous retrovirus (HERV) particles. In this field, small microvesicles (often with diameters of approximately 10-1000 nm, or 30-200 nm) released by exocytosis (exocytosis) of intracellular polyvesicles are called "microvesicles."

[0098] Exosomes are known to contain RNA types including mRNA (messenger RNA) and miRNA (microRNA). However, there is a fundamental lack of understanding regarding long RNA cargo in exosomes isolated from either in vitro or ex vivo systems. Conventional studies investigating exosomal RNA cargo have focused primarily on the small RNA fraction. Relatively small populations of annotated long RNAs and low transcript coverage have been reported in these studies, leading many to conclude that exosomes carry only short fragments of protein-coding and non-coding RNA, thus raising questions about their potential functional capacity in gene expression regulation and intercellular communication via exosomes.

[0099] As described herein, a wide range of diversity exists in RNA in plasma exosomes. RNA types identified by the methods herein include those identified in Figure 5. In some embodiments, RNA types identified by the methods herein include, but are not limited to, ribosomal RNA, SINE RNA, long scattered nucleotide sequence RNA, Alu RNA, HERVs, globin RNA, and other types of long non-coding RNA, and / or repeat sequences described elsewhere, for example, gencodegenes.org / gencode_biotypes.html.

[0100] The aforementioned method and kit isolate and extract nucleic acids, e.g., DNA and / or DNA and nucleic acids, from a sample, including at least RNA, using the following general techniques: First, nucleic acids in the sample (e.g., DNA and / or DNA and extracellular endoplasmic reticulum fraction) are bound to a capture surface such as a membrane filter, and the capture surface is washed. Next, an elution reagent is used to dissolve the nucleic acids on the membrane and release the nucleic acids, e.g., DNA and / or DNA and RNA, thereby forming an eluate. The eluate is then brought into contact with a protein precipitation buffer containing a transition metal and a buffer. Next, nucleic acids, including cfDNA and / or DNA and at least RNA from the extracellular endoplasmic reticulum, are isolated from the eluate after precipitation of proteins using one of various techniques recognized in the art, such as binding to a silica column followed by washing and elution.

[0101] In some embodiments, the elution buffer includes a denaturant, a surfactant, a buffering agent, and / or a combination thereof to maintain a specified solution pH. In some embodiments, the elution buffer includes a strong denaturant. In some embodiments, the elution buffer includes a strong denaturant and a reducing agent.

[0102] In some embodiments, the elution buffer contains guanidine thiocyanate (GTC), a denaturant that disrupts vesicular membranes, inactivates nucleases, and modulates ionic strength for solid-phase adsorption.

[0103] In some embodiments, the elution buffer contains a surfactant, such as Tween or Triton X-100, to assist in disrupting the extracellular endoplasmic reticulum membrane and to facilitate the effective elution of biomarkers from the capture surface.

[0104] In some embodiments, the elution buffer contains a reducing agent, such as β-mercaptoethanol (BME), to reduce intramolecular disulfide bonds (Cys-Cys) and help denature proteins, particularly RNases, present in the eluate.

[0105] In some embodiments, the elution buffer includes GTC, a surfactant, and a reducing agent.

[0106] In some embodiments, the transition metal ion in the protein precipitation buffer is zinc. In some embodiments, zinc is present in the protein precipitation buffer as zinc chloride.

[0107] In some embodiments, the buffer in the protein precipitation buffer is sodium acetate (NaAc). In some embodiments, the buffer is NaAc with a pH of ≤ 6.0.

[0108] In some embodiments, the protein precipitation buffer contains zinc chloride and a NaAc buffer at pH ≤ 6.0.

[0109] Current methods for isolating nucleic acids, including DNA and / or DNA and at least RNA, from extracellular endoplasmic reticulum (ER) include toxic substances, ultracentrifugation, ultrafiltration using, for example, 100 kD filters, polymer precipitation techniques, and / or particle size-based filtration. However, there is a need for alternative methods that are efficient and effective for isolating ER for use in a variety of applications, including diagnostic purposes, and for extracting nucleic acids contained within the ER, such as, in some embodiments, extracellular ER RNA, as needed.

[0110] The isolation and extraction methods and / or kits provided herein employ a spin column purification process using affinity membranes that bind to cell-free DNA and / or extracellular endoplasmic reticulum. The methods and kits disclosed herein are capable of processing numerous clinical samples in parallel using a single column with volumes of 0.2–4 mL. Cell-free DNA isolated using the procedures provided herein is of very high purity. Isolated RNA is of very high purity and is protected by vesicular membranes until lysis, allowing for the elution of intact vesicles from the membranes. The procedures can utilize substantially all cell-free DNA from plasma input, and the DNA yield is the same as or better than that of commercially available circulating DNA isolation kits. The procedures can utilize substantially all mRNA from plasma input, and the mRNA / miRNA yield is the same as or better than that of ultracentrifugation or direct lysis. Compared to commercially available kits and / or conventional isolation methods, the methods and / or kits are rich in the extracellular endoplasmic reticulum-bound fraction of miRNA, and the amount of input material can be easily increased. This ability to increase the amount allows for the study of smaller amounts of transcripts of interest. Compared to other commercially available products on the market, the methods and kits of this disclosure offer unique capabilities as demonstrated in the examples provided herein.

[0111] The aforementioned method and kit isolate and extract nucleic acids (e.g., DNA and / or DNA and at least RNA) from a biological sample using the following general procedure: First, the sample containing cfDNA and extracellular endoplasmic reticulum fractions is bound to a membrane filter and the filter is washed. Next, lysis is performed on the membrane using a GTC-based reagent to release the nucleic acids (e.g., DNA and / or DNA and RNA). Then, protein precipitation is performed. Next, the nucleic acids (e.g., DNA and / or DNA and RNA) are bound to a silica column, washed, and eluted. The extracted nucleic acids (e.g., DNA and / or DNA and RNA) may be further analyzed by any of the various downstream assays.

[0112] In some embodiments, nucleic acids are isolated according to the following steps: After adding a lysis reagent, a protein precipitation buffer is added to the homogenate and the solution is mixed vigorously for a short time. The solution is then centrifuged at room temperature at 12,000 × g for 3 minutes. The solution may then be treated with any of the various methods recognized in the art for isolating and / or extracting nucleic acids.

[0113] The isolated nucleic acids (e.g., DNA and / or DNA and RNA) may then be further analyzed using one of a variety of downstream assays. In some embodiments, simultaneous detection of DNA and RNA is used to increase sensitivity to possible mutations. Circulating nucleic acids have multiple possible sources of detectable mutations. For example, living tumor cells are a possible source of RNA and DNA isolated from the extracellular endoplasmic reticulum fraction of a sample. Dead tumor cells are a possible source of cell-free DNA (e.g., apoptotic vesicle DNA and cell-free DNA from necrotic tumor cells). Since the frequency of mutated nucleic acids is relatively low in circulating blood, maximizing detection sensitivity becomes very important. Simultaneous isolation of DNA and RNA provides comprehensive clinical information for evaluating disease progression and the patient's response to treatment. However, compared to the methods and kits provided herein, commercially available kits for detecting nucleic acids in the blood can only isolate plasma-derived, i.e., dead-cell-derived cfDNA. As is obvious to those skilled in the art, the more copies of mutations or other biomarkers there are, the higher the sensitivity and accuracy of identifying the mutations and other biomarkers.

[0114] The disclosed methods can be used to isolate all DNA from plasma samples. The disclosed methods can separate RNA and DNA at the same level for the same sample volume, and can separate that RNA and DNA from each other. These disclosed methods capture equivalent or superior cell-free DNA (cfDNA), equivalent or superior mRNA, and significantly more miRNA compared to commercially available separation kits.

[0115] The disclosed method can also be used for the simultaneous purification of RNA and DNA. The disclosed method (also referred to herein as the procedure) can be used to isolate RNA and DNA from exosomes and other extracellular endoplasmic reticulum using plasma or serum in volumes of 0.2–4 mL, e.g., 0.5–4 mL. A list of compatible plasma tubes includes plasma containing the additives EDTA, sodium citrate, and citrate-phosphate-dextrose. Plasma containing heparin may inhibit RT-qPCR.

[0116] Next, the sample, either alone or diluted with binding buffer, is packed into a spin column with a capture membrane and rotated at 500 × g for 1 minute. The pass-through fraction is discarded, and the column is returned to the same recovery tube. Wash buffer is then added, and the column is rotated at 5000 × g for 5 minutes to remove any residue. Note: After centrifugation, remove the recovery tube from the spin column so that the column does not come into contact with the pass-through fraction. The spin column is then transferred to a new recovery tube, and GTC-based elution buffer is added to the membrane. The spin column is then rotated at 5000 × g for 5 minutes to recover the homogenate containing the lysed exosomes. Protein precipitation is then performed.

[0117] The methods provided herein are useful for isolating and detecting DNA from biological samples. Endoplasmic reticulum RNA is thought to originate from living cells, for example, diseased tissue. Cell-free DNA (cfDNA) is thought to originate from dead cells, for example, necrotic cells in diseased tissue. Therefore, cfDNA is useful as an indicator of therapeutic response, while RNA is an indicator of increasing resistance mutations.

[0118] The method described herein is useful for detecting rare mutations in blood, because it provides a sufficiently sensitive method applicable to a sufficient amount of nucleic acid. Since the actual amount of DNA and RNA molecules in biological fluids is very limited, the method provides an isolation method for extracting all molecules from blood, suitable for mutation detection in small volumes sufficient for effective downstream processing and / or analysis.

[0119] In some embodiments, the sample separation and analysis techniques include methods referred to as EXO50 and / or EXO52, as described, for example, in WO2014 / 107571 and WO2016 / 007755 (each incorporated herein by reference as a whole). Commercially available liquid biopsy sample platforms are also intended, marketed under trade names EXOLUTION®, EXOLUTION PLUS®, EXOLUTION® UPREP, EXOLUTION HT®, UPREP®, EXOEASY®, EXORNEASY®, each available from Exosome Diagnostics, Inc., as well as QIAamp Circulating Nucleic Acids Kit, DNeasy Blood & Tissue Kits, AllPrep DNA / RNA Mini Kit, and AllPrep DNA / RNA / Protein Mini Kit, each available from Qiagen.

[0120] As used herein, the term "nucleic acid" refers to DNA and RNA. Nucleic acids may be single-stranded or double-stranded. In some examples, nucleic acids are DNA. In some examples, nucleic acids are RNA. RNA includes, but is not limited to, messenger RNA, transfer RNA, ribosomal RNA, non-coding RNA, microRNA, and HERV elements.

[0121] As used herein, the term "biological sample" refers to a sample containing biological materials (e.g., DNA, RNA, and proteins).

[0122] In some embodiments, the biological sample may preferably contain bodily fluids derived from the subject. These bodily fluids may be liquids isolated from any location on the subject's body, for example, a peripheral site. Examples of such bodily fluids include, but are not limited to, blood, plasma, serum, urine, saliva, cerebrospinal fluid, pleural fluid, nipple aspirate, lymph, bodily fluids contained in the respiratory tract, intestines, and urogenital tract, tears, saliva, breast milk, fluids derived from the lymphatic system, semen, visceral fluids, ascites, tumor cystic fluid, amniotic fluid, supernatants of cell culture media, and combinations thereof. The biological sample may also contain a fecal sample or a cecal sample, or supernatants isolated therefrom.

[0123] In some embodiments, the biological sample may preferably include the supernatant of a cell culture medium.

[0124] In some embodiments, the biological sample may preferably include a tissue sample derived from the subject. The tissue sample may be isolated from any location on the subject's body.

[0125] The preferred sample volume for body fluids is, for example, in the range of about 0.1 ml to about 30 ml. The volume of body fluid may depend on several factors, such as the type of body fluid used. For example, the volume of a serum sample may be about 0.1 ml to about 4 ml, and in some embodiments, for example, it may be about 0.2 ml to 4 ml. The volume of a plasma sample may be about 0.1 ml to about 4 ml, and 0.5 ml to 4 ml is preferred. The volume of a urine sample may be about 10 ml to about 30 ml, and in some embodiments, for example, it may be about 20 ml.

[0126] Although plasma samples were used in the examples provided herein, these methods are applicable to a variety of biological samples, as will be obvious to those skilled in the art.

[0127] The methods and kits disclosed herein are suitable for use with samples derived from human subjects. They are also suitable for use with samples derived from non-human subjects (e.g., rodents, non-human primates, pet animals (e.g., cats, dogs, horses), and / or livestock (e.g., chickens)).

[0128] The term "subject" refers to all particles that have been shown to or are expected to contain nucleic acids. This includes animals. In certain embodiments, the subject is a mammal, human or non-human primate, dog, cat, horse, cattle, other livestock, or rodent (e.g., mouse, rat, guinea pig, etc.). A human subject may be a healthy human being without any observable abnormalities (e.g., disease). A human subject may be a human being with observable abnormalities (e.g., disease). Such observable abnormalities may be observed by the human being themselves or by a medical professional. The terms “subject,” “patient,” and “individual” are interchangeable herein.

[0129] Although a membrane is used as the capture surface in the examples provided herein, it is obvious that the morphology of the capture surface (e.g., beads or a filter, also referred to herein as a membrane) does not affect the ability of the method provided herein to efficiently capture extracellular vesicles from a biological sample.

[0130] A wide range of capture surfaces can capture extracellular endoplasmic reticulum according to the methods provided herein, but not all capture surfaces capture extracellular endoplasmic reticulum (some surfaces capture nothing).

[0131] This disclosure also describes an apparatus for isolating and concentrating extracellular vesicles from biological or clinical samples using disposable plastic components and a centrifuge. For example, the apparatus includes a column containing a capture surface (i.e., a membrane filter), a holder for securing the capture surface between an outer frit and an inner tube, and a recovery tube. The outer frit preferably contains a large mesh structure through which liquid can pass and is located at one end of the column. The inner tube holds the capture surface in place and is preferably slightly conical. The recovery tube may be a commercially available, i.e., 50 ml Falcon tube. The column is preferably suitable for rotation; that is, its size is preferably compatible with standard centrifuges and microcentrifuges.

[0132] In embodiments where the capture surface is a membrane, the apparatus for isolating the extracellular endoplasmic reticulum fraction from a biological sample includes at least one membrane. In some embodiments, the apparatus includes one, two, three, four, five, or six membranes. In some embodiments, the apparatus includes three membranes. In embodiments where the apparatus includes more than one membrane, all of these membranes are directly adjacent to each other on one side of the column. In embodiments where the apparatus includes more than one membrane, all of these membranes are identical to each other, i.e., have the same charge and / or the same functional groups.

[0133] Furthermore, capture by filtration with a pore size smaller than that of the extracellular endoplasmic reticulum is not the primary reaction mechanism for capture in the methods provided herein. Nevertheless, the pore size of the filter is important. For example, mRNA may be trapped and not recovered through a 20 nm filter, while microRNA can be easily eluted. Also, for example, the pore size of the filter is an important parameter for the available surface capture area.

[0134] The methods provided herein utilize any of a variety of capture surfaces. In some embodiments, the capture surface is a film (also referred to herein as a filter or membrane filter). In some embodiments, the capture surface is a commercially available film. In some embodiments, the capture surface is a commercially available electrostatic film. In some embodiments, the capture surface is neutral. Depending on the embodiment, the capture surface may be Mustang® ion exchange membrane (manufactured by PALL Corporation), Vivapure® Q membrane (manufactured by Sartorius AG), Sartobind Q or Vivapure® Q Maxi H, Sartobind® D (manufactured by Sartorius AG), Sartobind(S) (manufactured by Sartorius AG), Sartobind® Q (manufactured by Sartorius AG), Sartobind® IDA (manufactured by Sartorius AG), Sartobind® Aldehyde (manufactured by Sartorius AG), Whatman® DE81 (Sigma), Fast Trap Virus purification column (EMD Millipore), Thermo Scientific * Selected from Pierce Strong Cation and Anion Exchange Spin Columns.

[0135] In embodiments where the capture surface is charged, the capture surface may be a charged filter selected from the group consisting of positively charged Q PES vacuum filtration (Millipore) with a charge of 0.65 μm, positively charged Q RC spin column filtration (Sartorius) with a charge of 3-5 μm, positively charged Q PES homemade spin column filtration (Pall) with a charge of 0.8 μm, positively charged Q PES syringe filtration (Pall) with a charge of 0.8 μm, negatively charged S PES homemade spin column filtration (Pall) with a charge of 0.8 μm, negatively charged S PES syringe filtration (Pall) with a charge of 0.8 μm, and negatively charged nylon syringe filtration (Sterlitech) with a charge of 50 nm. In some embodiments, the charged filter is not housed in the syringe filtration apparatus because it may be more difficult to remove nucleic acids from the filter in these embodiments. In some embodiments, the charged filter is housed at one end of the column.

[0136] In embodiments where the capture surface is a film, the film may be made from a variety of suitable materials. In some embodiments, the film is polyethersulfone (PES) (e.g., Millipore or Pall). In some embodiments, the film is regenerated cellulose (RC) (e.g., Sartorius or Pierce).

[0137] In some embodiments, the capture surface is a positively charged film. In some embodiments, the capture surface is a positively charged film and is a Q film which is an anion exchanger containing a quaternary amine. For example, the Q film is a quaternary ammonium R-CH3-N + It is functionalized with (CH3)3. In some embodiments, the capture surface is a negatively charged film. In some embodiments, the capture surface is a negatively charged S film and a cation exchanger containing sulfonic acid groups. For example, the S film is a sulfonic acid R-CH3-SO3 - It is functionalized with a diethylamine group R-CH3-NH +The D membrane is a weakly basic anion exchanger containing (C2H5)2. In some embodiments, the capture surface is a metal chelate membrane. For example, the membrane is iminodiacetic acid-N(CH3COOH - The IDA membrane is functionalized with )2. In some embodiments, the capture surface is a microporous membrane functionalized with an aldehyde group-CHO. In other embodiments, the membrane is a weakly basic anion exchanger containing diethylaminoethyl (DEAE) cellulose. Not all charged membranes are suitable for use in the methods provided herein. For example, RNA isolated using a Sartorius Vivapure S membrane spin column showed inhibition of RT-qPCR and was therefore unsuitable for PCR-related downstream assays.

[0138] In the embodiment where the capture surface is charged, the extracellular endoplasmic reticulum can be isolated with a positively charged filter.

[0139] In embodiments where the capture surface is charged, the pH during the capture of extracellular endoplasmic reticulum is 7 or less. In some embodiments, the pH is greater than 4 and 8 or less.

[0140] In the embodiment where the capture surface is a positively charged Q filter, the buffer system includes a washing buffer containing 250 mM bistrispropane (pH: 6.5-7.0). In the embodiment where the capture surface is a positively charged Q filter, the lysis buffer is a GTC-based reagent. In the embodiment where the capture surface is a positively charged Q filter, the lysis buffer is present in a volume of 1. In the embodiment where the capture surface is a positively charged Q filter, the lysis buffer is present in a volume greater than 1.

[0141] Depending on the film material, the pore size of the film is in the range of 3 μm to 20 nm. For example, in an embodiment where the capture surface is a commercially available PES film, the film has pore sizes of 20 nm (Exomir), 0.65 μm (Millipore), or 0.8 μm (Pall). In an embodiment where the capture surface is a commercially available RC film, the film has pore sizes in the range of 3 to 5 μm (Sartorius, Pierce).

[0142] The surface charge of the capture surface may be positive, negative, or neutral. In some embodiments, the capture surface is a positively charged bead or a set of beads.

[0143] The methods provided herein include a dissolution reagent. In some embodiments, the reagent used for dissolution on a membrane is a GTC-based reagent. In some embodiments, the dissolution reagent is a high-salt concentration buffer.

[0144] The methods provided herein involve various buffers, including a packing buffer and a wash buffer. The ionic strength of the packing buffer and wash buffer may be high or low. The salt concentration (e.g., NaCl concentration) may be 0 to 2.4 M. The buffer may contain various components. In some embodiments, the buffer contains one or more of the following components: Tris, bis-Tris, bis-Tris-propane, imidazole, citrate, methylmalonic acid, acetate, ethanolamine, diethanolamine, triethanolamine (TEA), and sodium phosphate. In the methods provided herein, the pH of the packing buffer and wash buffer is important. If the pH of the plasma sample is set to 5.5 or lower before packing, the filter tends to clog (the plasma does not rotate at all in the column). Also, if the pH is high, the extracellular endoplasmic reticulum becomes unstable, resulting in a low recovery rate of extracellular endoplasmic reticulum RNA. A neutral pH is optimal for the recovery of RNA from the extracellular endoplasmic reticulum. In some embodiments, the concentration of the buffer used is 1x, 2x, 3x, or 4x. For example, the packing buffer or binding buffer has a concentration of 2, while the washing buffer has a concentration of 1.

[0145] In some embodiments, the method includes, for example, one or more washing steps after contacting a biological sample with a capture surface. In some embodiments, a detergent is added to the washing buffer to facilitate the removal of nonspecific binding (i.e., contaminants, cellular debris, and circulating protein complexes or nucleic acids) to obtain a higher purity extracellular endoplasmic reticulum fraction. Suitable detergents include, but are not limited to, sodium dodecyl sulfate (SDS), Tween-20, Tween-80, Triton X-100, nonidet P-40 (NP-40), Brij-35, Brij-58, octyl glucoside, octyl thioglucoside, CHAPS, or CHAPSO.

[0146] In some embodiments, the capture surface (e.g., membrane) is housed within the apparatus used, such as a spin column in centrifugation, a vacuum filter holder in a vacuum system, or a syringe filter in pressure filtration. In some embodiments, the capture surface is housed within a spin column or vacuum system.

[0147] Isolating extracellular vesicles (ERs) from biological samples before nucleic acid extraction is beneficial for the following reasons: 1) Extracting nucleic acids from ERs provides an opportunity to selectively analyze disease- or tumor-specific nucleic acids obtained by isolating disease- or tumor-specific ERs from other ERs in a body fluid sample. 2) Compared to the yield / integrity of nucleic acid species obtained by directly extracting nucleic acids from a body fluid sample without first isolating ERs, ERs containing nucleic acids produce nucleic acid species with higher integrity in very high yields. 3) The methods described herein allow for increased volume scalability and sensitivity for detecting nucleic acids expressed at low concentrations, for example, by concentrating ERs from larger volumes of sample. 4) Higher purity and higher quality nucleic acids are extracted because proteins, lipids, cell debris, cells, and other potential contaminants, as well as naturally occurring PCR inhibitors in biological samples, are removed before the nucleic acid extraction step. 5) More options are available in the nucleic acid extraction method. This is because the volume of the isolated extracellular endoplasmic reticulum fraction is less than the starting volume of the sample, allowing nucleic acids to be extracted from these fractions or pellets using a small-volume column filter.

[0148] Several methods for isolating extracellular vesicles from biological samples have been described in this field. For example, fractionation centrifugation is described in the papers of Raposo et al. (Raposo et al, 1996), Skog et al. (Skog et al, 2008), and Nilsson et al. (Nilsson et al., 2009). Methods of ion exchange and / or gel permeation chromatography are described in U.S. Patents 6,899,863 and 6,812,023. Sucrose density gradient electrophoresis, i.e., organelle electrophoresis, is described in U.S. Patent 7,198,923. Magnetic activated cell sorting (MACS) is described in the paper of Taylor and Gercel Taylor (Taylor and Gercel-Taylor, 2008). Nanomembrane ultrafiltration concentration is described in the paper of Cheruvanky et al. (Cheruvanky et al, 2007). The Percoll gradient isolation method is described in Miranda et al.'s paper (Miranda et al., 2010). Furthermore, extracellular vesicles may be identified and isolated from the target body fluid using microfluidic devices (Chen et al., 2010). For research and development and the commercial application of nucleic acid biomarkers, it is desirable to extract high-quality nucleic acids from biological samples using a consistent, reliable, and practical method.

[0149] Therefore, an object of the present invention is to provide a method for quickly and easily isolating nucleic acid-containing particles from biological samples (e.g., bodily fluids) and for extracting high-quality nucleic acids from the isolated particles. The method of the present invention may be suitable for adoption and integration into small devices or semi-automated or fully automated equipment used in laboratories, clinical settings, or in clinical settings.

[0150] In some embodiments, the sample is not pretreated before isolating and extracting nucleic acids (e.g., DNA and / or DNA and RNA) from the biological sample.

[0151] In some embodiments, the sample undergoes a pretreatment step before isolation, purification, or concentration of extracellular vesicles are performed to remove undesirable large particles, cells and / or cellular debris, and other contaminants present in the biological sample. The pretreatment step may be achieved by one or more centrifugation steps (e.g., fractional centrifugation), one or more filtration steps (e.g., ultrafiltration), or a combination thereof. If two or more centrifugation pretreatment steps are performed, the biological sample may first be centrifuged at low speed and then at high speed. Further preferred centrifugation pretreatment steps may be performed as needed. The biological sample may be filtered instead of, or in addition to, one or more centrifugation pretreatment steps. For example, the biological sample may first be centrifuged at 20,000 g for 1 hour to remove undesirable large particles, and then the sample may be filtered, for example, through a 0.8 μm filter.

[0152] In some embodiments, the sample is pre-filtered to remove particles larger than 0.8 μm. In some embodiments, the sample contains additives such as EDTA, sodium citrate, and / or citrate-glucose phosphate. In some embodiments, the sample does not contain heparin because heparin can adversely affect RT-qPCR and other nucleic acid analyses. In some embodiments, the sample is mixed with a buffer before purification and / or nucleic acid isolation and / or extraction. In some embodiments, the buffer is a binding buffer.

[0153] In some embodiments, one or more centrifugation steps are performed before and after contacting the biological sample with the capture surface to separate the extracellular vesicles and to concentrate the extracellular vesicles isolated from the biological sample fraction. To remove undesirable large particles, cells and / or cellular debris, the sample may be centrifuged at a low speed of about 100–500 g, for example, about 250–300 g in some embodiments. Alternatively, or in addition, the sample may be centrifuged at a higher speed. The centrifugation speed is about 200,000 g or less, for example, from about 2,000 g to less than about 200,000 g. Centrifugation speeds of over about 15,000 g and less than about 200,000 g, over about 15,000 g and less than about 100,000 g, and over about 15,000 g and less than about 50,000 g are used in some embodiments. The centrifugation speed is preferably about 18,000g to about 40,000g or about 30,000g, with about 18,000g to about 25,000g being more preferable. In some embodiments, the centrifugation speed is about 20,000g. Generally, a suitable time for centrifugation is about 5 minutes to about 2 hours, for example, about 10 minutes to about 1.5 hours, or about 15 minutes to about 1 hour. About 0.5 hours may be used. Centrifuging a biological sample at about 20,000g for about 0.5 hours is useful in some embodiments. However, any combination of the above speeds and times can be suitably used (e.g., about 18,000g to about 25,000g or about 30,000g to about 40,000g, and about 10 minutes to about 1.5 hours, or about 15 minutes or about 1 hour, or about 0.5 hours, etc.). The one or more centrifugal separation steps may be performed at a temperature below ambient temperature, for example, about 0 to 10°C, for example, about 1 to 5°C, for example, about 3°C ​​or about 4°C.

[0154] In some embodiments, one or more filtration steps are performed before and after bringing the biological sample into contact with the capture surface. A filter with a pore size in the range of about 0.1 to about 1.0 μm, for example, about 0.8 μm or 0.22 μm, may be used. Alternatively, continuous filtration may be performed while reducing the porosity of the filter.

[0155] In some embodiments, to reduce the volume of the sample processed in the chromatography step, one or more concentration steps are performed before or after contacting the biological sample with the capture surface. Concentration may be performed by centrifugation of the sample at high speed, for example, 10,000 to 100,000 g, to precipitate the extracellular vesicles. This may consist of a series of fractional centrifugations. The extracellular vesicles in the resulting pellet may be reconstituted in a suitable buffer in a smaller volume for the next step in the process. Alternatively, this concentration step may be performed by ultrafiltration. In fact, this ultrafiltration concentrates the biological sample and further purifies the extracellular vesicle fraction. In other embodiments, the filtration is ultrafiltration, for example, tangential ultrafiltration. Tangential ultrafiltration consists of concentrating and fractionating a solution between two compartments (filtrate and retaining solution) separated by a membrane with a predetermined cutoff threshold. Separation is performed by applying flow in the retaining solution compartment and intermembrane pressure between this compartment and the filtrate compartment. Ultrafiltration may be performed using various systems, such as helical membranes (Millipore, Amicon), flat membranes, or hollow fibers (Amicon, Millipore, Sartorius, Pall, GF, Sepracor). Within the scope of the present invention, it is beneficial to use membranes with a cutoff threshold of less than 1000 kDa, for example, 100 kDa to 1000 kDa in some embodiments, and 100 kDa to 600 kDa in some embodiments.

[0156] In some embodiments, one or more size exclusion chromatography steps or gel permeation chromatography steps are performed before or after contacting the biological sample with the capture surface. In some embodiments, a support selected from silica, acrylamide, agarose, dextran, ethylene glycol-methacrylic acid copolymer, or a mixture thereof (e.g., an agarose-dextran mixture) is used to perform the gel permeation chromatography step. Examples of such supports include, but are not limited to, SUPERDEX® 200HR (Pharmacia), TSK G6000 (TosoHaas), or SEPHACRYL® S (Pharmacia).

[0157] In some embodiments, one or more affinity chromatography steps are performed before or after contacting the biological sample with the capture surface. Furthermore, several extracellular vesicles can be characterized by specific surface molecules. Since microvesicles form from budding of the cell membrane, these microvesicles often share many of the same surface molecules found in the cells from which they originate. As used herein, “surface molecules” collectively refer to antigens, proteins, lipids, carbohydrates, and markers found on, within, or on the membrane surface of microvesicles. Examples of these surface molecules include receptors, tumor-associated antigens, and membrane protein modifications (e.g., glycosylation structures). For example, microvesicles budding from tumor cells often present tumor-associated antigens on their cell surface. Therefore, affinity chromatography or affinity exclusion chromatography may be used in combination with the methods provided herein to isolate, identify, or enrich specific populations of microvesicles from a particular donor cell type (Al-Nedawi et al., 2008; Taylor and Gercel-Taylor, 2008). For example, (malignant or non-malignant) tumor microvesicles carry tumor-associated surface antigens, and these tumor microvesicles may be detected, isolated, and / or concentrated via these tumor-specific tumor-associated surface antigens. In one example, the surface antigen is epithelial cell adhesion molecule (EpCAM). Epithelial cell adhesion molecule is specific to microvesicles from epithelial malignancies originating from the lung, colorectal, thoracic, prostate, head and neck, and liver, but not from blood cells (Balzar et al., 1999; Went et al., 2004). Tumor-specific microvesicles can also be characterized by the absence of specific surface markers (e.g., CD80 and CD86). In these cases, microvesicles possessing these markers may be excluded, for example, by affinity exclusion chromatography, for further analysis of the tumor-specific markers.Affinity chromatography can be achieved, for example, using various supports, resins, beads, antibodies, aptamers, aptamer analogs, molecular template polymers, or other molecules well known in the art that specifically target desired surface molecules of microvesicles.

[0158] In some embodiments, one or more control particles or one or more nucleic acids may be added to the sample before isolation of the extracellular endoplasmic reticulum or extraction of nucleic acids to act as an internal control and evaluate the efficiency or quality of the extracellular endoplasmic reticulum purification and / or nucleic acid extraction. The methods described herein provide efficient isolation and control nucleic acids along with the extracellular endoplasmic reticulum fraction. These control nucleic acids include one or more nucleic acids from Qβ bacteriophages, one or more nucleic acids from viral particles, or any other particles containing a control nucleic acid (e.g., at least one control target gene, which may occur naturally or be genetically engineered using recombinant DNA technology). In some embodiments, the amount of control nucleic acid is known before addition to the sample. The control target gene may be quantified by real-time PCR analysis. The quantification of the control target gene may be used to determine the efficiency or quality of the extracellular endoplasmic reticulum purification step or the nucleic acid extraction step.

[0159] In some embodiments, the control nucleic acid is a nucleic acid from a Qβ bacteriophage (referred to herein as a Qβ control nucleic acid). The Qβ control nucleic acid used in the methods described herein may be a naturally occurring viral control nucleic acid or a recombinant or genetically engineered viral control nucleic acid. Qβ belongs to the Leviviridae family and is characterized by a linear single-stranded RNA genome consisting of four viral proteins: a coating protein, a maturation protein, a lysating protein, and three genes encoding RNA replication enzymes. When Qβ particles themselves are used as a control, their size is similar to that of an average extracellular endoplasmic reticulum, so Qβ can be easily purified from biological samples by the same purification methods used to isolate extracellular endoplasmic reticulum as described herein. Furthermore, the relatively simple single-stranded gene structure of the Qβ virus makes it useful as a control in amplification-based nucleic acid assays. The Qβ particles contain a control target gene or control target sequence that is detected or measured to quantify the amount of Qβ particles in the sample. For example, the control target gene is a Qβ coating protein gene. When Qβ particles themselves are used as a control, the Qβ particles are added to the biological sample, and then the nucleic acids derived from the Qβ particles are extracted together with the nucleic acids derived from the biological sample using the extraction method described herein. When nucleic acids from Qβ, such as RNA from Qβ, are used as a control, the Qβ nucleic acids are extracted together with the nucleic acids derived from the Qβ particles using the extraction method described herein. Detection of the Qβ control target gene can be determined, for example, by RT-PCR analysis simultaneously with (one or more) biomarkers of interest. The copy number may be determined using standard curves at least two, three, or four concentrations in a 10-fold dilution of the control target gene. The quality of the isolation and / or extraction process may be determined by comparing the detected copy number with the amount of added Qβ particles, or the detected copy number with the amount of added Qβ nucleic acid, such as QβRNA.

[0160] In some embodiments, the copy number of Qβ particles or Qβ nucleic acid (e.g., QβRNA) added to a bodily fluid sample is 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 1,000, or 5,000. In some embodiments, the copy number of Qβ particles or Qβ nucleic acid (e.g., QβRNA) added to a bodily fluid sample is 100. When Qβ particles themselves are used as a control, the copy number of Qβ particles may be calculated based on the ability of the Qβ bacteriophage to infect target cells. Thus, the copy number of Qβ particles correlates with the colony-forming unit of the Qβ bacteriophage.

[0161] Optionally, control particles may be added to the sample before extracellular endoplasmic reticulum isolation or nucleic acid extraction to serve as an internal control for evaluating the efficiency or quality of extracellular endoplasmic reticulum purification and / or nucleic acid extraction. The methods described herein provide efficient isolation and control particles associated with the extracellular endoplasmic reticulum fraction. These control particles may include Qβ bacteriophages, viral particles, or other particles containing control nucleic acids (e.g., at least one control target gene) that may occur naturally or be designed by recombinant DNA technology. In some embodiments, the amount of control particles is known before addition to the sample. The control target gene can be quantified using real-time PCR analysis. Quantification of the control target gene can be used to determine the efficiency or quality of the extracellular endoplasmic reticulum purification or nucleic acid extraction step.

[0162] In some embodiments, Qβ particles are added to the urine sample before nuclear extraction. For example, Qβ particles are added after the pretreatment filtration step of the urine sample before ultrafiltration.

[0163] In some embodiments, the methods and kits described herein include one or more in-process controls. In some embodiments, the in-process control is the detection and analysis of a reference gene indicating sample quality (i.e., an indicator of the quality of a biological sample, e.g., a biological fluid sample). In some embodiments, the in-process control is the detection and analysis of a reference gene indicating plasma quality (i.e., an indicator of the quality of a plasma sample). In some embodiments, the (one or more) reference genes are analyzed by additional qPCR.

[0164] In some embodiments, the in-process control is an in-process control relating to reverse transcriptase and / or PCR performance. Examples of these in-process controls include, but are not limited to, a reference RNA (also referred to herein as ref.RNA) that is spiked in after RNA isolation and before reverse transcription. In some embodiments, the ref.RNA is a control such as Qbeta. In some embodiments, the ref.RNA is analyzed by additional PCR. Nucleic acid extraction

[0165] This invention relates to the use of a capture surface to improve the isolation, purification, or concentration of extracellular endoplasmic reticulum (ER). The method disclosed herein provides a highly concentrated ER fraction for extracting high-quality nucleic acids from the ER. Nucleic acid extracts obtained by the method described herein may be useful in a variety of applications where high-quality nucleic acid extracts are required or preferred, such as the diagnosis, prognosis, or monitoring of diseases or medical conditions.

[0166] Recent studies have revealed that nucleic acids within the microendoplasmic reticulum (ER) can serve as biomarkers. For example, WO2009 / 100029 describes the use of nucleic acids extracted from the ER of serum from GBM patients, particularly in medical diagnosis, prognosis, and treatment evaluation. WO2009 / 100029 also describes the use of nucleic acids extracted from the ER in human urine for the same purpose. The use of nucleic acids extracted from the ER is considered to potentially avoid the need for biopsy, highlighting the significant diagnostic potential of ER biology (Skog et al, 2008).

[0167] The quality or purity of isolated extracellular vesicles can directly affect the quality of extracted extracellular vesicle nucleic acids, and thus directly affect the efficiency and sensitivity of biomarker assays for disease diagnosis, prognosis, and / or monitoring. Given the importance of accurate and highly sensitive diagnostic testing in the clinical field, there is a need for a method to isolate a highly concentrated extracellular vesicle fraction from biological samples. To address this need, a method for isolating extracellular vesicles from biological samples to extract high-quality nucleic acids is described herein. As shown herein, a highly concentrated extracellular vesicle fraction is isolated from a biological sample by the method described herein, and high-quality nucleic acids are subsequently extracted from this highly concentrated extracellular vesicle fraction. These extracted high-quality nucleic acids are useful for measuring or evaluating the presence or absence of biomarkers that are adjunct in the diagnosis, prognosis, and / or monitoring of diseases or other medical conditions.

[0168] As used herein when describing nucleic acid extraction, the term "high quality" means an extract in which 18S rRNA and 28S rRNA can be detected, for example, in some embodiments, in a ratio of about 1:1 to about 1:2; and / or, for example, in some embodiments, in a ratio of about 1:2. Ideally, a high-quality nucleic acid extract obtained by the methods described herein has an RNA integrity number (RIN) of 5 or higher for low-protein biological samples (e.g., urine) and 3 or higher for protein biological samples (e.g., serum), and a nucleic acid yield of 50 pg / ml or higher from 20 ml of low-protein biological sample or 1 ml of high-protein biological sample.

[0169] High-quality RNA extracts are desirable because RNA degradation can negatively impact downstream evaluations of the extracted RNA, such as gene expression and mRNA analysis, and analysis of non-coding RNAs (e.g., small RNAs and microRNAs). The novel method described herein allows for the extraction of high-quality nucleic acids from extracellular endoplasmic reticulum isolated from biological samples, enabling accurate analysis of nucleic acids within the extracellular endoplasmic reticulum.

[0170] After isolating the extracellular vesicle from a biological sample, nucleic acids may be extracted from the isolated or concentrated extracellular vesicle fraction. To achieve this, in some embodiments, the extracellular vesicle may first be lysed. Lysis of the extracellular vesicle and extraction of nucleic acids can be achieved by various methods well known in the art. In some embodiments, nucleic acid extraction can be achieved using protein precipitation according to standard procedures and techniques well known in the art. Alternatively, nucleic acids contained within the extracellular vesicle may be captured using a nucleic acid-binding column. The nucleic acids may be bound and then eluted using a buffer or a solution suitable for inhibiting the interaction between the nucleic acid and the binding column. This allows the nucleic acids to be eluted.

[0171] In some embodiments, the nucleic acid extraction method also includes a step of removing or mitigating harmful factors that prevent the extraction of high-quality nucleic acids from a biological sample. Such harmful factors are heterogeneous, as different biological samples may contain a variety of harmful factors. In some biological samples, factors such as excess DNA may affect the quality of the nucleic acid extract from the biological sample. In other samples, factors such as excess endogenous RNases may affect the quality of the nucleic acid extract from the biological sample. These harmful factors can be removed using a variety of reagents and methods. These methods and reagents are collectively referred to herein as “extraction enhancement operations.” In some examples, extraction enhancement operations may include adding nucleic acid extraction enhancement reagents to the biological sample. To remove harmful factors such as endogenous RNases, the extraction-enhancing reagents defined herein include, but are not limited to, RNase inhibitors (e.g., Superase-In (commercially available from Ambion), RNase INplus (commercially available from Promega), or other reagents with similar functions), proteases (which may function as RNase inhibitors), DNases, reducing agents, decoy substrates (e.g., synthetic RNA and / or carrier RNA), soluble receptors capable of binding to RNases, small interfering RNAs (siRNAs), RNA-binding molecules (e.g., anti-RNA antibodies, basic proteins, or chaperone proteins), RNase-denaturing substances (e.g., hyperosmolar solutions, purifying agents), or combinations thereof.

[0172] For example, the extraction enhancement procedure may include adding an RNase inhibitor to the biological sample and / or isolated extracellular endoplasmic reticulum fraction before extracting nucleic acids; for example, in some embodiments, the concentration of the RNase inhibitor is greater than 0.027 AU (1x) for a sample of 1 μl or more by volume, or 0.135 AU (5x) or more for a sample of 1 μl or more by volume, or 0.27 AU (10x) or more for a sample of 1 μl or more by volume, or 0.675 AU (25x) or more for a sample of 1 μl or more by volume, or 1.35 AU (50x) or more for a sample of 1 μl or more by volume. Here, a 1x concentration refers to the enzymatic conditions for treating extracellular endoplasmic reticulum isolated from 1 μl or more of body fluid with an RNase inhibitor of 0.027 AU or more. A 5x concentration refers to the enzymatic conditions for treating extracellular endoplasmic reticulum isolated from 1 μl or more of body fluid with an RNase inhibitor of 0.135 AU or more. A 10-fold protease concentration refers to enzymatic conditions for processing particles isolated from 1 μl or more of body fluid using an RNase inhibitor of 0.27 AU or higher. A 25-fold concentration refers to enzymatic conditions for processing extracellular endoplasmic reticulum isolated from 1 μl or more of body fluid using an RNase inhibitor of 0.675 AU or higher. A 50-fold protease concentration refers to enzymatic conditions for processing particles isolated from 1 μl or more of body fluid using an RNase inhibitor of 1.35 AU or higher. In some embodiments, the RNase inhibitor is a protease. In this case, 1 AU is the protease activity that releases folin-positive amino acids and peptides corresponding to 1 μmol of tyrosine per minute.

[0173] These nucleic acid extraction enhancement reagents, for example, inhibit RNase activity (e.g., RNA extraction They may exert their function in various ways, such as as enzyme degrading inhibitors, universal protein degradation (e.g., proteases), or as chaperone proteins that bind to and protect RNA (e.g., RNA-binding proteins). In all cases, such extraction-enhancing reagents remove, or at least mitigate, some or all of the harmful factors in the biological sample, or harmful factors associated with the isolated particles that would otherwise hinder high-quality nucleic acid extraction from the isolated particles.

[0174] In some embodiments, the quality of nucleic acid extraction may be determined using quantification of 18S rRNA and 28S rRNA. Detection of nucleic acid biomarkers

[0175] In some embodiments, the extracted nucleic acids include DNA and / or DNA and RNA. In embodiments in which the extracted nucleic acids include DNA and RNA, the RNA is reverse transcribed into complementary DNA (cDNA) before further amplification. Such reverse transcription may be performed alone or in combination with the amplification step. An example of a method combining the reverse transcription and amplification steps is reverse transcription polymerase chain reaction (RT-PCR). This may be further modified quantitatively, for example, quantitative RT-PCR as described in U.S. Patent No. 5,639,606 (the teachings of which are incorporated herein by reference). Another example of such a method includes two separate steps: a first reverse transcription step to produce cDNA from RNA, and a second step to quantify the amount of cDNA by qPCR. As shown in the following examples, the RNA extracted from nucleic acid-containing particles by the methods disclosed herein includes a variety of transcript species. Examples of these include, but are not limited to, ribosomal 18S and 28S rRNA, microRNAs, transfer RNAs, disease or condition-associated transcripts, and biomarkers important for the diagnosis, prognosis, and monitoring of disease conditions.

[0176] For example, RT-PCR analysis determines the Ct (cycle threshold) value for each reaction. In RT-PCR, a positive reaction is detected by the accumulation of a fluorescence signal. The Ct value is defined as the number of cycles required for the fluorescence signal to exceed the threshold (i.e., exceed the background value). The Ct value is inversely proportional to the amount of target nucleic acid or control nucleic acid in the sample (i.e., the lower the Ct value, the greater the amount of control nucleic acid in the sample).

[0177] In other embodiments, the copy number of the control nucleic acid can be measured by any of the various techniques recognized in the prior art. Such techniques include, but are not limited to, RT-PCR. The copy number of the control nucleic acid can be determined by methods well known in the art, such as by constructing and using a calibration curve, or standard curve.

[0178] In some embodiments, one or more biomarkers may be one or a set of genetic abnormalities. In this specification, “biomarker” is used to refer to the quantity of nucleic acid and nucleic acid variants within a nucleic acid-containing particle. Specifically, genetic abnormalities include, but are not limited to, overexpression of a gene (e.g., oncogene) or gene panel, underexpression of a gene (e.g., tumor suppressor genes such as p53 or RB) or gene panel, alternative products of splice variants of a gene or gene panel, copy number variants (CNVs) (e.g., double microDNA) (Hahn, 1993), nucleic acid modifications (e.g., methylation, acetylation, and phosphorylation), single nucleotide polymorphisms (SNPs), chromosomal rearrangements (e.g., inversions, deletions, and duplications), and mutations in a gene or gene panel (insertions, deletions, duplications, missense mutations, nonsense mutations, synonymous mutations, or any other nucleotide changes) or any combination thereof, which often ultimately result in the mutation affecting the activity and function of the gene product, leading to a different transcriptional splice variant and / or changes in gene expression levels.

[0179] Analysis of nucleic acids present in isolated particles is quantitative and / or qualitative. In quantitative analysis, the relative or absolute amount (expression level) of a specific target nucleic acid in the isolated particles is measured by methods well known in the art (as described below). In qualitative analysis, whether the species of the specific target nucleic acid in the isolated extracellular vesicle is wild-type or a variant is identified by methods well known in the art.

[0180] The present invention also includes the use of a method for isolating extracellular vesicles from a biological sample and sequencing nucleic acids for the purpose of (i) assisting in the diagnosis of a subject, (ii) monitoring the progression or recurrence of a disease or other medical condition of a subject, or (iii) assisting in the evaluation of the effectiveness of treatment for a subject who is receiving or is expected to receive treatment for a disease or other medical condition, wherein it is determined whether one or more biomarkers are present in the nucleic acid extract obtained by the method, and the one or more biomarkers are associated with the diagnosis, progression or recurrence of a disease or other medical condition, or the effectiveness of treatment, respectively.

[0181] In some embodiments, amplifying the nucleic acids of the extracellular endoplasmic reticulum before analysis may be beneficial or otherwise desirable. Methods for nucleic acid amplification are commonly used and widely known in the art, and many examples thereof are described herein. If desired, amplification can be carried out to be quantitative. Quantitative amplification allows for quantitative measurement of the relative amounts of various nucleic acids, resulting in genetic or expression profiles.

[0182] Nucleic acid amplification methods include, but are not limited to, polymerase chain reaction (PCR) (U.S. Patent No. 5,219,727) and its variations, such as in-situ polymerase chain reaction (U.S. Patent No. 5,538,871), quantitative polymerase chain reaction (U.S. Patent No. 5,219,727), nested polymerase chain reaction (U.S. Patent No. 5,556,773), auto-persistent sequence replication and its variations (Guatelli et al., 1990), transcription amplification systems and their variations (Kwoh et al., 1989), Qβ replicase and its variations (Miele et al., 1983), cold-PCR (Li et al., 2008), beaming (Li et al., 2006), or any other nucleic acid amplification method, wherein the amplified molecules are subsequently detected using techniques well known to those skilled in the art. Detection schemes designed to detect nucleic acid molecules are particularly useful when the amount of such molecules present is very small. The aforementioned references are incorporated herein by their teachings regarding these methods. In other embodiments, the nucleic acid amplification step is not performed. Instead, the extracted nucleic acid is analyzed directly (e.g., by next-generation sequencing).

[0183] The determination of such genetic abnormalities can be carried out by a variety of techniques known to those skilled in the art. For example, nucleic acid expression levels, alternative splice variants, chromosomal rearrangements, and gene copy numbers can be determined by microarray analysis (see, for example, U.S. Patents 6,913,879, 7,364,848, 7,378,245, 6,893,837, and 6,004,755) and quantitative PCR. In particular, copy number variations can be detected by the Illumina Infinium II whole-genome gene typing assay or the Agilent Human Genome CGH microarray (Steemers et al., 2006). Nucleic acid modifications can be assayed by methods described, for example, U.S. Patent 7,186,512 and Japanese Patent Publication WO / 2003 / 023065. In particular, methylation profiles can be determined by the Illumina DNA methylation OMA003 cancer panel.SNPs and mutations are detected by hybridization with allele-specific probes, enzymatic mutation detection, chemical cleavage of mismatched heteroduplexes (Cotton et al., 1988), ribonuclease cleavage of mismatched bases (Myers et al., 1985), mass spectrometry (US Patent Nos. 6,994,960, 7,074,563, and 7,198,893), nucleic acid sequencing, single-stranded higher-order polymorphism (SSCP) (Orita et al., 1989), denaturing gradient gel electrophoresis (DGGE) (Fischer and Lerman, 1979a; Fischer and Lerman, 1979b), temperature gradient gel electrophoresis (TGGE) (Fischer and Lerman, 1979a; Fischer and Lerman, 1979b), restriction fragment length polymorphism (RFLP) (Kan and Genetic abnormalities can be detected by methods such as Dozy, 1978a; Kan and Dozy, 1978b), oligonucleotide ligation assays (OLA), allele-specific PCR (ASPCR) (U.S. Patent No. 5,639,611), ligation linkage (LCR) and its variations (Abravaya et al., 1995; Landegren et al., 1988; Nakazawa et al., 1994), and flow cytometry heteroduplex analysis (WO / 2006 / 113590) and its combinations / modifications. In particular, gene expression levels can be measured by gene expression linkage analysis (SAGE) techniques (Velculescu et al., 1995). In general, methods for analyzing genetic abnormalities have been reported in numerous publications and are available to those skilled in the art, not limited to the methods cited herein. Appropriate analytical methods are considered to depend on the specific objectives of the analysis, the patient's condition / medical history, and the specific cancer, disease, or other medical condition being detected, monitored, or treated. The aforementioned references are incorporated herein by their teaching of these methods.

[0184] Many biomarkers may be associated with the presence or absence of disease or other medical conditions in a subject. Therefore, detecting the presence or absence of such biomarkers in nucleic acid extracts from isolated particles according to the methods disclosed herein may aid in the diagnosis of disease or other medical conditions in a subject.

[0185] Furthermore, many biomarkers may help monitor disease or medical conditions in subjects. Therefore, detecting the presence or absence of such biomarkers in nucleic acid extracts from particles isolated according to the methods disclosed herein may help monitor the progression or recurrence of disease or other medical conditions in subjects.

[0186] Many biomarkers have been found to influence the effectiveness of treatments in specific patients. Therefore, detecting the presence or absence of such biomarkers in nucleic acid extracts from particles isolated according to the methods disclosed herein may help evaluate the effectiveness of a given treatment in a given patient. Identifying these biomarkers in nucleic acids extracted from particles isolated from patient biological samples may guide the selection of treatments for patients.

[0187] In certain embodiments of the above aspects of the present invention, the disease or other medical condition is a neoplastic disease or condition (e.g., cancer or a cell proliferation disorder).

[0188] In some embodiments, extracted nucleic acids, such as exosomal RNA (also referred to herein as “exoRNA”), are further analyzed based on the detection of a biomarker or combination of biomarkers. In some embodiments, further analysis is performed using machine learning-based modeling, data mining methods, and / or statistical analysis. In some embodiments, the data is analyzed to identify or predict the disease outcomes of patients. In some embodiments, the data is analyzed to stratify patients within a patient population. In some embodiments, the data is analyzed to identify or predict whether a patient is resistant to treatment. In some embodiments, the data is used to measure the progression-free survival rate over time for a subject.

[0189] In some embodiments, data is analyzed to select a treatment option if a biomarker or combination of biomarkers is detected. In some embodiments, the treatment option is a procedure involving a combination of treatments. Sequencing technology

[0190] In some embodiments, “next-generation” sequencing (NGS) or high-throughput sequencing experiments are performed according to the methods of the present invention. These sequencing techniques enable the identification of nucleic acids that are present in small or large quantities in a sample, or otherwise undetectable by more conventional hybridization methods. NGS typically incorporates the addition of nucleotides followed by a washing step.

[0191] Commercially available kits for total RNA sequencing that preserve strand information and are suitable for mammalian RNA and very low input RNA are useful in this regard, and include, but are not limited to, the Clontech: SMARTer stranded total RNASeq kit; Clontech: SMARTSeq v4 ultra low input RNASeq kit; Illumina: Truseq stranded total RNA library preparation kit; Kapa Biosystems: Kapa stranded RNASeq library preparation kit; New England Biolabs: NEBNext ultra directional library preparation kit; Nugen: Ovation Solo RNASeq kit; and Nugen: Nugen Ovation RNASeq System v2. [Examples]

[0192] In the examples provided herein, various membranes and apparatuses are used for the purposes of centrifugation and / or filtration, but obviously these methods can be used with any capture surface and / or containment device that can efficiently capture the extracellular endoplasmic reticulum and release nucleic acids, in particular RNA contained in the extracellular endoplasmic reticulum. Example 1

[0193] Sample separation

[0194] Samples are generally obtained from commercial sources and separated by the EXO50 and / or EXO52 methods described, for example, WO2014 / 107571 and WO2016 / 007755.

[0195] Long-chain RNASeq workflow method 1:

[0196] After sample separation, treat the samples with DNase I enzyme and / or modified DNase I enzyme by incubation at temperatures such as approximately 30°C to 40°C, for example, approximately 35°C to 37°C, generally for about 10 minutes to 2 hours, or approximately 10 minutes to 60 minutes, according to the manufacturer's guidelines.

[0197] Following DNase treatment, exogenous synthetic RNA spike-in was added to the sample at a dilution adjusted according to the sample. The synthetic spike-in may be added to the sample either before or after DNase treatment. Subsequently, the sample was subjected to RNA fragmentation using commercially available reagents / protocols, or the sample was left unfragmented.

[0198] Next, the sample is subjected to first-strand cDNA synthesis (reverse transcription) using commercially available reagents, in accordance with the manufacturer's guidelines.

[0199] Next, the Illumina-based NGS adapter is added to the cDNA using PCR-based techniques and commercially available reagents, in accordance with the manufacturer's guidelines.

[0200] Following PCR-based addition of the NGS adapter, the sample is subjected to one or two rounds of paramagnetic bead-based library cleanup using common commercially available reagents, in accordance with the manufacturer's guidelines.

[0201] Following AMPure cleanup, the samples are subjected to ribodepression using commercially available reagents and protocols.

[0202] Next, the sample is subjected to multiple cycles of PCR amplification using commercially available reagents and protocols. The number of PCR-based amplification cycles may range from 10 to 30.

[0203] Following PCR amplification, the samples are subjected to one or two rounds of paramagnetic bead-based library cleanup using commercially available reagents and protocols.

[0204] At this stage, the final NGS library and samples are subjected to standard NGS QC measurements, including BioAnalyzer (fragment size analysis and concentration) and Qubit (concentration). Samples are diluted to concentrations of 1–4 nM and then stored before preparation for sequencing. Standard sequencing preparations include sample denaturation and dilution to pM concentrations for clustering using sequencing instruments.

[0205] Long-chain RNASeq workflow method 3:

[0206] After separation, the sample is treated with DNase I enzyme, generally by incubation at temperatures such as approximately 30°C to 40°C, for example, approximately 35°C to 37°C, for a period of approximately 10 minutes to 2 hours, e.g., approximately 10 minutes to 60 minutes, or approximately 30 minutes, according to the manufacturer's guidelines.

[0207] After DNase treatment, exogenous synthetic RNA spikein is added to the sample at a dilution ratio adjusted according to the sample.

[0208] Next, the samples are subjected to another round of DNase treatment and primer annealing using commercially available reagents and protocols.

[0209] Next, the sample is subjected to first-strand cDNA synthesis (reverse transcription) using commercially available reagents, in accordance with the manufacturer's guidelines.

[0210] Next, the sample is subjected to a series of cDNA processing steps using commercially available NGS reagents.

[0211] Next, the samples are subjected to second-strand cDNA synthesis (reverse transcription), end repair, and adapter ligation reactions using commercially available reagents and guidelines. Following the adapter ligation reaction, the samples are subjected to one or more rounds of paramagnetic bead-based library cleanup using commercially available reagents. The standard protocol was modified to suit our workflow, with changes to the bead hydration time and elution volume, as specified in the examples.

[0212] Next, quantitative PCR (qPCR) is performed using commercially available reagents to determine the optimal amplification cycle for the sample being tested.

[0213] Next, the sample is subjected to PCR amplification using commercially available reagents and protocols, based on the number of cycles determined in the previous step.

[0214] Following PCR amplification, the samples are subjected to one or two rounds of paramagnetic bead-based library cleanup using commercially available reagents and, furthermore, a standard kit protocol modified to incorporate a bead hydration step into our workflow.

[0215] Next, the sample is subjected to standard BioAnalyzer or Qubit analysis to determine the sample concentration. A maximum of 10 ng of library is then used to proceed to the next step in the workflow.

[0216] Next, the samples were subjected to ribodepression using commercially available reagents and protocols. The standard kit protocol was modified for our workflow by adding an additional primer annealing cycle.

[0217] Next, the sample is subjected to a second round of PCR amplification using commercially available reagents and protocols until multiple cycles of PCR amplification are completed.

[0218] Following PCR amplification, the samples are subjected to one or two rounds of paramagnetic bead-based library cleanup using commercially available reagents and protocols. These protocols were modified to vary hydration times for our workflow, as detailed in the examples.

[0219] At this stage, we have the final NGS library and the samples are subjected to standard NGS QC measurements, including BioAnalyzer (fragment size analysis and concentration) and Qubit (concentration). The samples are diluted to concentrations of 1–4 nM and then stored before preparation for sequencing. Standard sequencing preparations involve sample denaturation and dilution to pM concentrations, which are used for clustering with sequencing instruments.

[0220] Concentrated Workflow

[0221] RNA, DNA, or a combination of RNA and DNA was subjected to the following workflow. Samples were either treated with DNase I using commercially available reagents and guidelines, or left untreated.

[0222] Next, the sample was either subjected to fragmentation or left unfragmented. First-strand cDNA synthesis (reverse transcription) was performed on the sample, and if necessary, second-strand cDNA synthesis (i.e., DNA polymerase reaction) was performed on the sample after first-strand cDNA synthesis.

[0223] The samples are subjected to end repair and adapter ligation reactions according to commercially available reagents and guidelines. Following the adapter ligation reactions, the samples are subjected to a one-round cleanup of a paramagnetic bead-based library using commercially available reagents and guidelines.

[0224] The sample is subjected to a first PCR amplification according to commercially available reagents and guidelines. Following PCR, the sample is subjected to hybridization-based targeted enrichment, and / or the final NGS library prepared from Method 1 or Method 3 (RNA, DNA, or RNA and DNA) is subjected to hybridization-based targeted enrichment. The probe is biotinylated with specific sequences having a size of 60 bp, 80 bp, or 120 bp. The tiling density or overlap of the probe at sequence-specific sites may be 1x, 2x, 3x, 4x, and more.

[0225] The probe is hybridized to the specific sequence of interest in the sample using commercially available reagents and guidelines. Following hybridization, the sample is exposed to streptavidin-conjugated paramagnetic beads to capture the specific sequence and remove undesirable sequences. Cleaning of undesirable sequences, buffers, and enzymes is performed by washing, while simultaneously binding the specific sequence of interest in the sample to the streptavidin-conjugated paramagnetic beads.

[0226] Hybridization of the specific sequence and capture by streptavidin-bound paramagnetic beads are performed once or multiple times, using hybridization times ranging from approximately 2 to 24 hours.

[0227] Following hybridization, the sequence-specifically captured sample is subjected to a second PCR amplification using increased amplification cycles, in accordance with the manufacturer's guidelines, but with improved yield.

[0228] In some embodiments, the sample is subjected to a single round of cleanup using a paramagnetic bead-based library or a filter spin column. The sample is eluted with a smaller volume of elution buffer to increase the final concentration (i.e., ng / μl, nM).

[0229] At this stage, we have the final enriched NGS library and the samples are subjected to standard NGS QC measurements, including BioAnalyzer (fragment size analysis and concentration) and Qubit (concentration). The samples are diluted to concentrations of 1–4 nM and then stored before preparation for sequencing. Standard sequencing preparations involve sample denaturation and dilution to pM concentrations, which are used for clustering with sequencing instruments. Example 2

[0230] We developed a novel platform specifically designed to incorporate both short and long RNA transcripts from exosomes within an RNA sequencing workflow. Using ExoLution or ExoLution Plus, available from Exosome Diagnostics, and starting with human plasma, we isolated high-quality whole exosomal RNA and subjected the resulting long RNASeq workflow method 1.

[0231] Simply put, the workflow was initiated using exosomal RNA isolates, followed by DNase treatment for applications where DNA might interfere with the analysis. Spike-in of synthetic RNA standards was sometimes performed before or after the DNase step as a quality control indicator.

[0232] RNA is transcribed using a mixed oligonucleotide and a reverse transcriptase with template-switching activity. Subsequently, a barcoded or unbarcoded DNA oligo adapter is added using PCR. cDNA is cleaned up starting with smaller oligonucleotides using paramagnetic beads. Since ribosome sequences can affect the detection of small amounts of transcript, a ribosome removal step may be used to selectively remove them. After cleavage or removal of ribosome sequences, the uncleaved library molecules are increased in concentration by PCR. This is followed by another cleanup using paramagnetic beads. The library may then be quantified and sequenced.

[0233] The objective of the experiment is to develop a complete long-chain RNA sequencing platform optimized for plasma exosomes.

[0234] Sample: 2 mL of normal human plasma from a pool of 48 individuals with no gender bias. Synthetic spike-in was added to the sample as a control for sensitivity, technical reproducibility, and stranding-ness.

[0235] Exosomal RNA Seq: Identification and optimization of optimal variables / variable combinations for DNase treatment, RNA / cDNA fragmentation, amplification, and ribosomal RNA removal.

[0236] The figures demonstrate the remarkable results of this method. In particular, Figure 1 provides a bioanalyzer scan showing the amplification and incorporation obtained for long RNA transcripts in the RNASeq library, including the exosomal RNA size distribution and the final library size distribution of exosomes. Amplified cDNA can be observed from both small RNA and long RNA fragments.

[0237] Figure 2 shows the excellent correlation and reproducibility between library replications using this method. By identifying appropriate conditions for plasma exosomal RNA, we improved the reproducibility between library replications from 0.7 to 0.97. The technical reproducibility between replications, determined by the correlation of exogenous RNA spike-in, is 0.999. 97% to 99% of transcripts retained correct strand information. Figure 3 shows that variable optimization increased the proportion of reads located in the transcriptome. In particular, changing these parameters allows 40–45% of read maps to be located in the transcriptome. Similarly, Figure 4 shows that variable optimization minimizes read repetitions and increases protein-coding reads.

[0238] Figure 5 shows the broad diversity of RNA in plasma exosomes in terms of RNA type (excluding ribosomal RNA).

[0239] Figure 6 shows highly efficient removal of ribosomal RNA with and without ribodepression for 28S, 18S, 12S, and 16S ribosomal genes.

[0240] Figure 7 shows the transcriptome coverage rate of plasma exosomes.

[0241] Figure 8 shows the diversity of plasma exosomal RNA cargo.

[0242] Figure 9 shows a large number of long RNAs with complete transcript coverage in exosomes, and a bimodal distribution of the transcript coverage in exosomes.

[0243] Figure 10 shows that this method provides highly sensitive detection of molecules in exosomes.

[0244] Additional data generated by this method are provided in Figures 11-15. Figure 12 shows excellent correlation of spike-in between library replications, while Figure 13 demonstrates that the 5-end of the transcript has higher coverage. Amplification and incorporation of long RNA transcripts in the RNASeq library are provided in Figure 14 as the final library size distribution of plasma exosomes.

[0245] In summary, this method provides a novel approach for long-chain RNA sequencing of exosomes. It demonstrates excellent reproducibility (R>0.97) for RNA transcript detection, high sensitivity for transcript detection (LOD@15M reading = 12 molecules), and highly efficient removal of ribosomal transcripts. In addition, a wide diversity of protein-coding and non-coding RNAs is detected in exosomes by this method. This method identifies a large number of transcripts with complete coverage in exosomes. Example 3

[0246] We developed a novel platform specifically designed to incorporate both short and long RNA transcripts from exosomes within an RNA sequencing workflow. We further extended these workflows to also process DNA, either alone or in mixtures with RNA. We further extended these workflows to specifically increase sample concentration with respect to the target of interest, enabling deeper sequence coverage. Using ExoLution®, ExoLution HT®, UPrep®, ExoEasy®, ExoRNeasy®, or ExoLution Plus®, available from Exosome Diagnostics, and starting with human plasma, we isolated high-quality whole exosome nucleic acids and subjected the resulting long RNASeq workflows Method 1 and / or Method 3.

[0247] The sequencing workflow described begins after the nucleic acid is separated from the biological fluid (Figure 16). The volume of the biological fluid serving as input for the sequencing workflow can be as small as ≥0.5 ml, and there is no upper limit (Figure 16). The nucleic acid may originate from exosomes and / or other cell-free sources.

[0248] In some embodiments, aliquots of the sample are used in a hybridization-based enrichment process (referring to Figures 17–19). This process utilizes the hybridization of a nucleotide probe complementary to the genomic sequence region of interest contained in the sample, followed by a series of washes using a buffer selected for the sequence of interest, during which unwanted material is washed away. The probe-sequence hybrid may, but is not limited to, be selected to utilize a streptavidin-biotin chemical reaction. The process can, but is not limited to, be used to enrich any portion or mixture of genomic sequences, including exon and intron regions, and can cover the location of the entire gene-coding region or a specific hotspot within a gene. Hybridization probe panels, though not limited to these, can be used to increase the concentration of any target sequence from a small number of targets (1-20) to many targets (>1,000), including, but not limited to, the entire protein-coding transcriptome containing ~20,000 genes (see Figure 4), large panels targeting a broad range of diseases or disease-related pathways with >1,000 genes (see Figures 17-18), and medium-sized panels targeting a specific disease (e.g., solid tumors) or disease-related pathway with 50-500 genes.

[0249] Figure 17 demonstrates sample enrichment using a pan-cancer panel. Samples were subjected to library preparation and subsequently enriched for a panel of 1,387 cancer-related targets. Samples containing only RNA or a mixture of RNA and cfDNA from liquid biopsy specimens were examined. Commercial RNA (UHR) was included as a control sample. Figure 17A shows the library mapping index, illustrating a very high percentage for the target readouts produced. Figure 17B shows the base coverage index, indicating that the majority of nucleotides in the panel are more than 1× covered across all three samples. Figure 17C plots the number of mapped readouts per target and illustrates the ability to process samples containing only RNA and RNA+cfDNA, as well as the increase in readout counts when both RNA and cfDNA are analyzed in the same sample.

[0250] Figure 18 demonstrates enrichment using a pan-cancer panel. Samples were subjected to library preparation and subsequently enriched for a panel of 1,387 cancer-related targets. Samples containing only cfDNA or a mixture of RNA and cfDNA from liquid biopsy specimens were examined. Figure 18A shows the library mapping metrics, illustrating a very high percentage for the target readouts produced. Figure 18B shows the base coverage metrics, showing that cfDNA+RNA provides superior target coverage compared to cfDNA alone when the same volume of starting plasma is used. Figure 18C plots the number of mapped readouts per target and illustrates the ability to process samples containing only cfDNA and RNA+cfDNA, as well as the increase in readout counts when the RNA contribution is included together with cfDNA.

[0251] Figure 19 demonstrates sample enrichment using the Whole Exome Capture Panel. Samples were subjected to library preparation and subsequently enriched for a panel of ~20,000 targets representing the entire protein encoding the transcriptome. Samples containing only cfDNA or a mixture of RNA and cfDNA from liquid biopsy specimens were investigated. Commercial RNA (UHR) was included as a control sample. Figure 19A shows the library mapping index, illustrating a very high percentage for the produced target readouts. Figure 19B demonstrates the indicated base coverage index. Figure 19C shows the number of mapped readouts per target and illustrates the ability to process samples containing only cfDNA and RNA+cfDNA. Example 4

[0252] If the sample is not concentrated, the entire sample is sequenced (Figures 20-22).

[0253] Figure 20 demonstrates two independent RNA-seq library preparation workflows optimized for exosomal liquid biopsy samples. To minimize variability, replicated RNA extraction was performed from a control plasma pool. Replication (6 times per method) was then carried out using one of the two optimized workflows. Samples were not subjected to ribosome desorption. All samples were subjected to deep sequencing and downsampled to normalize read counts for analysis. Figure 20A shows highly reproducible detection of transcripts between library replications. Figure 20B illustrates a high percentage of reads located in the transcriptome at the target metric. Figure 20C shows the percentage of reads per biotype.

[0254] In some embodiments, when the entire sample is being analyzed, ribosomal sequences (cDNA, RNA, or cfDNA) may interfere with the detection of small amounts of transcripts, in which case it is desirable to remove or eliminate the sample containing ribosomal sequences (see Figure 21). Selective removal of ribosomal sequences may be achieved at the RNA sequence level, the cDNA level, or the dsDNA (library) level. Ribosomal sequence-specific removal may be achieved using RNase H or enzymatic reagents similar to restriction enzyme digestion, but is not limited to these. Removal may also be achieved by utilizing hybridization-based biotinylated probe enrichment and streptavidin-conjugated paramagnetic beads to specifically capture and remove ribosomal sequences.

[0255] Figure 21 demonstrates various ribosomal RNA removal approaches. To minimize variability, replicated RNA extracts were isolated from a control plasma pool. To further avoid variability, all samples were prepared using the same RNA-seq library procedure. Each sample was subjected to one of three commercially available ribosomal sequence removal approaches (also known as Reagent 1, Reagent 2, and Reagent 3), where removal can occur at the RNA or cDNA level. For Reagent 1, Condition A refers to the commercially available protocol, while Condition B refers to the identified optimal protocol. Samples without ribosomal RNA removal were included as untreated controls. All samples were subjected to deep sequencing and downsampled to normalize read counts for analysis. Figure 21A shows how the transcriptome read rates illustrate how the selection of optimal conditions significantly impacts the effectiveness of removal and subsequent recovery of the RNA of interest. Figure 21B provides a comparison of the removed sample to the unremoved sample and illustrates the importance of selecting optimal conditions that have the most effective removal of undesirable ribosomal RNA while maintaining the diversity of the initial library (purple), minimizing loss (pink), and enabling the elucidation of additional RNA (blue) compared to the treatment. Herein, we found that condition B using reagent 1 resulted in the most optimal ribodepression, leading to a 17-fold improvement in protein-coding reads, 93.5% overlap with unremoved RNA, and 96.5% new transcripts.

[0256] Figure 22 demonstrates library preparation methods for both total RNA and total nucleic acids (cfDNA+RNA). Nucleic acids were isolated from control plasma using one of three methods: one to isolate high-quality RNA only, and two different methods to isolate cfDNA in addition to RNA. To minimize variability, samples from each of these libraries were prepared for sequencing using the same approach. Figure 22A shows that highly reproducible libraries are produced using this library method, and that both isolation methods produce highly similar starting materials. Figure 22B provides mapping indices, while Figure 22C provides transcript coverage demonstrating the higher transcript coverage detected by combining RNA and cfDNA compared to RNA alone. Figure 22D provides identified transcripts for RNA only (top panel), cfDNA+RNA for Method 1 (middle panel), and Method 2 (bottom panel). By combining RNA and DNA from samples and subjecting them to the same workflow, the detection of all transcripts (coding and non-coding) increased from 47.4% (RNA only) to 99.6% (RNA + cfDNA), the detection of protein-coding transcripts increased from 66.5% (RNA only) to 99.9% (RNA + cfDNA), and the detection of lincRNA increased from 14.7% (RNA only) to >99% (RNA + cfDNA).

[0257] Figure 23 demonstrates the detection limit of exogenous RNA spike-in based on a whole RNASeq assay in six independent library replicas constructed from plasma. The figure demonstrates consistent detection of RNA down to 10 molecules or less. The dynamic range of this assay spans five orders of magnitude, from 10 to 1.8 million molecules.

[0258] Following library quantification, the library is normalized, multiplexed, and subjected to sequencing using next-generation sequencing platforms. Next, the sequencing data is demultiplexed if necessary, and transcript / gene counts are generated by mapping to existing genomic or transcriptome reference sequences or to de novo synthesized genomes or transcripts (see FIGS. 16 and 24). FIG. 24 provides a representative RNASeq browser for displaying QC metrics and analysis results.

[0259] Next, the UMI tags of each sequence can be used to identify fragments generated by PCR duplication. Counts are normalized particularly with respect to library size, GC bias, sequence bias, and sequencing depth. These counts can then be used to perform differential expression analysis, which is a gene expression analysis related to various conditions (e.g., tumor / normal) for creating a list of biomarkers distinguishable between sample types as shown in FIG. 25, and it provides a representative differential expression browser for displaying and evaluating the results of differential expression analysis.

[0260] The aligned reference data can be used for profiling sequence changes, including but not limited to single nucleotide polymorphisms, insertions / deletions, fusions, inversions, and repeat expansions. Example 5

[0261] An easily available commercial kit for hybridization-based target enrichment, created for tissue analysis including formalin-fixed paraffin-embedded (FFPE) tissues, was used to test its applicability to extracellular vesicle samples (exoRNA and cfDNA). The kit contains targets for the detection of fusions, insertions / deletions, single nucleotide polymorphisms, and copy number variations.

[0262] To test the feasibility of adapting a uniform process for exosome samples, we examined two processes, Method 4 and Method 5, shown in Figure 26. Investigating the targeted enrichment parameters outlined in Table 1, we found that exoRNA and cfDNA were comparable to small inputs of control RNA (generic human reference RNA) and DNA (normal genomic DNA). The expected range of targeted enrichment (readings located at target-specific sequences) is 70%–99% for exoRNA and cfDNA. Table 1 [Table 1]

[0263] For exoRNA and cfDNA, the percentage of reads located in the transcriptome ranges from 30% to 95%, intronic regions from 5% to 60%, and intergenic regions from 0.2% to 10%.

[0264] In exoRNA and cfDNA, the target sequence bases are covered by between >90% in a single read.

[0265] We found that 30,000 times more coverage depth is needed to detect low-frequency mutations. Example 6

[0266] The experimental objective was to investigate the feasibility of exosome samples (combined exoRNA and cfDNA) using a commercially available pan-cancer kit (Method 6 in Figure 26). The kit was designed for the detection of RNA transcripts and fusions in FFPE and cancer samples.

[0267] We investigated targeted enrichment of exoRNA, cfDNA, and combined exoRNA and cfDNA. The results for targeted enrichment of exosome samples were comparable to those obtained using the general-purpose human reference RNA control. Table 2 [Table 2]

[0268] The expected range of target enrichment (readings located at target-specific sequences) is 75%–99% for exoRNA, cfDNA, and combined exoRNA and cfDNA.

[0269] For combined exoRNA and cfDNA, the percentage of transcriptome-located reads ranges from 35% to 95%, 8% to 45% in intronic regions, and 0.4% to 5% in intergenic regions.

[0270] In the combined exoRNA and cfDNA, the target sequence bases are completely covered by a single read with >80% accuracy.

[0271] The combination of exoRNA and cfDNA provides higher read coverage per gene and better target enrichment, as seen in Figure 17. Example 7

[0272] The objectives of this study are to investigate: (1) the effects of various fragmentation times on total RNASeq data; (2) the effects of DNase treatment; (3) the effects of ribodepression; and (4) the effects of synthetic spike-in. The method is the long-chain RNASeq workflow method 1 outlined previously, followed by RNA isolation from 2 mL of normal human plasma sample using the EXO-50 method.

[0273] Typically, the sample is subjected to DNase treatment, followed by synthetic spike-in, and then a single ribodepression step. The samples are shown in Table 3. Table 3 [Table 3]

[0274] The mapping statistics of the analyzed samples are as shown in Figures 27 - 28, indicating consistency between replicates. Regarding Figure 27, it was found that in all samples, approximately 40 - 50% of the reads are located in the transcriptome.

[0275] As is evident from Figure 28, 30 - 40% of the reads are located in ribosomal RNA, and 40 - 50% of the reads are located in RNA encoding proteins.

[0276] The size distribution of the inserts is relatively consistent between replicates of different libraries, as shown in Figure 29.

[0277] The transcriptome coverage rate is also investigated as shown in Figures 30 and 31. The library detects over 10,000 protein - coding genes. Protein - coding genes are most abundant in exosomes, followed by pseudogenes and processed lincRNAs.

[0278] The transcript coverage rate is also relatively consistent across all samples and replicates, as shown in Figure 32.

[0279] Figure 33 demonstrates the detection limit of synthetic spike - in transcripts. The dynamic range of the synthetic spike - in exceeds 5 orders of magnitude.

[0280] Figure 34 demonstrates the 5' - to - 3' transcript coverage rate shown in all libraries. Example 8

[0281] The purpose of this example is to construct an RNASeq library from normal human plasma exosomes using the long - chain RNASeq workflow method 3 and to investigate (1) the number of amplification cycles; (2) the effect of ribodepletion. The method is the above - mentioned long - chain RNASeq workflow method 3 using synthetic spike - ins, followed by RNA isolation from 2 mL of normal human plasma samples using the EXO - 50 method.

[0282] Typically, the sample is subjected to N-ase treatment, ribodepression, addition of synthetic spike-in, reverse transcription, a variable number of amplification cycles, and further processing (including a cleanup step) according to workflow method 3. The sample is identified as shown in Table 4. Table 4 [Table 4]

[0283] The mapping statistics for the analyzed samples are shown in Figures 35-36. As shown in Figure 35, in all samples, the majority of the readings are located in the transcriptome.

[0284] As is evident from Figure 36, the lowest percentage of ribosome reads in the library was observed in samples 1 and 2, and the highest percentage of protein-coding reads and misc RNA reads was also observed in samples 1 and 2.

[0285] The size distribution of the inserts is very consistent between replications and across all samples, as shown in Figure 37.

[0286] Transcriptome coverage is investigated as shown in Figure 38. Overall, transcriptome coverage is consistent between replications and across all samples.

[0287] Figure 39 shows that, overall, gene detection across samples is consistent at various detection thresholds.

[0288] Overall, the transcript coverage is consistent across samples at the various detection thresholds shown in Figure 40.

[0289] Figure 41 highlights the size of transcripts with >80% coverage.

[0290] Figure 42 demonstrates that the ERCC spike-in detection levels observed in samples 1-7 differed from those in samples 8-9.

[0291] Figure 43 shows that a relatively consistent coverage rate of the full-length transcript was observed using the long-chain RNASeq workflow method 3 library, and that this was consistent between replications and samples. Other embodiments

[0292] Although the present invention has been described in conjunction with its detailed description, the above description is intended to illustrate the scope of the present invention as defined by the appended claims, and not to limit its scope. Other aspects, advantages, and modifications are included within the following scope. References [ka] [ka]

Claims

1. A method for sequencing a long non-ribosomal microvesicle RNA transcript from a biological sample, wherein the long non-ribosomal microvesicle RNA transcript contains more than 200 nucleotides and includes long non-coding RNA, mRNA, circular RNA, or any combination thereof, and the method is as follows: (a) The solid-capturing surface and the biological sample are brought into contact under conditions sufficient to retain the extracellular endoplasmic reticulum containing the long-chain non-ribosomal microendoplasmic reticulum RNA transcript from the biological sample on or within the surface; (b) Bringing the lysis reagent and the capture surface into contact while the extracellular vesicle is present on or within the capture surface, thereby releasing the long non-ribosomal microvesicle RNA transcript from the capture surface and producing a homogenate; (c) Extract the long non-ribosomal microendoplasmic reticulum RNA transcript from the homogenate; (d) The long non-ribosomal microendoplasmic reticulum RNA transcript is reverse transcribed into cDNA; (e) Construct a double-stranded DNA library from the reverse-transcribed cDNA; (f) Remove nucleic acid molecules containing ribosomal DNA or RNA sequences from the double-stranded DNA library; (g) selectively increase the concentration of nucleic acid sequences from the double-stranded DNA library; and, (h) Sequence the long non-ribosomal microendoplasmic reticulum RNA transcript by sequencing the nucleic acid sequence selectively concentrated from the double-stranded DNA library; A method comprising, wherein the method further comprises pretreatment of a homogenate or one or more extracted long non-ribosomal microvesicle RNA transcripts with a DNase before or after step (c).

2. The method according to claim 1, further comprising amplifying the double-stranded DNA library after step (e).

3. The method according to any one of claims 1 to 2, further comprising adding an exogenous RNA or DNA spike to the homogenate or to one or more extracted long microendoplasmic reticulum RNA transcripts before or after step (c).

4. The method according to any one of claims 1 to 3, wherein selective removal of nucleic acid molecules containing ribosome sequences is performed using an enzyme reagent, a biotinylated probe, and streptavidin-conjugated paramagnetic beads, or a combination thereof.

5. The method according to claim 4, wherein the enzyme reagent comprises a restriction enzyme.

6. The method according to any one of claims 1 to 5, wherein selective enhancement of nucleic acid sequences from the double-stranded DNA library includes utilizing a PCR-based approach, complementary oligonucleotides, hybridization-based biotinylated probe enrichment and streptavidin-conjugated paramagnetic beads, or a combination thereof.

7. The method according to any one of claims 1 to 6, wherein the cDNA or double-stranded DNA library molecule is tagged using a unique molecular index, which enables template identification, deduplication, error repair, copy number measurement, or a combination thereof.

8. The method according to any one of claims 1 to 7, wherein the one or more long chain microendoplasmic reticulum RNA transcripts contain more than 300 nucleotides.

9. The method according to any one of claims 1 to 8, wherein the one or more long chain microendoplasmic reticulum RNA transcripts contain more than 500 nucleotides.

10. The method according to any one of claims 1 to 9, wherein the biological sample is selected from the group consisting of blood, plasma, serum, urine, saliva, cerebrospinal fluid, pleural fluid, nipple aspirate, lymph, bodily fluids contained in the respiratory tract, intestines, and urogenital tract, tears, saliva, breast milk, lymphatic fluid, semen, cerebrospinal fluid, visceral fluid, ascites, tumor cystic fluid, amniotic fluid, and combinations thereof.

11. The method according to any one of claims 1 to 10, wherein the biological sample comprises blood, plasma, or serum.

12. The method according to any one of claims 1 to 11, wherein the solid object capturing surface includes a film.

13. The method according to any one of claims 1 to 12, wherein the solid object capturing surface comprises two or more types of films.

14. The method according to any one of claims 1 to 13, wherein the solid object capturing surface comprises three types of membranes.

15. The method according to any one of claims 1 to 14, wherein the solid object trapping surface comprises one or more ion exchange (IEX) beads.

16. The method according to any one of claims 1 to 15, wherein the solid-capturing surface is functionalized with a quaternary amine, a sulfate, a sulfonate, a tertiary amine, or a combination thereof.

17. The solid-capturing surface is a quaternary ammonium R-CH 2 -N + (CH 3 ) 3 The method according to any one of claims 1 to 16, wherein the method is functionalized with

18. The method according to any one of claims 1 to 17, wherein step (c) further comprises adding a protein precipitation buffer to the homogenate before extracting one or more long microendoplasmic reticulum RNA transcripts from the homogenate, the protein precipitation buffer containing a transition metal ion, a buffer, or both a transition metal ion and a buffer.

19. The method according to any one of claims 1 to 18, wherein step (c) comprises carrying out digestion using a proteinase, a DNase, or a combination thereof.

20. The method according to any one of claims 1 to 19, wherein step (c) comprises adding isopropanol, sodium acetate, glycogen, or a combination thereof.

Citation Information

Patent Citations

  • Method for isolating microvesicles

    JP2016502862A

  • Enrichment and next-generation sequencing of total nucleic acids, including both genomic DNA and cdan

    JP2016510992A

  • Sequencing and analysis of exosome-associated nucleic acids

    JP2019535307A

  • Methods of depleting a target molecule from an initial collection of nucleic acids, and compositions and kits for practicing the same

    WO2015122967A1