Methods and compositions for simultaneous profiling of genome and transcriptome
ArcDR-seq addresses the limitations of existing DNA and RNA sequencing methods by enabling simultaneous profiling of genome and transcriptome from single cells or low input materials, including degraded samples, enhancing scalability and applicability to archival samples, and improving clinical diagnostics.
Patent Information
- Application Number
- PCT/US2025/038701
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-24
- Filing Date
- 2025-07-22
- Publication Date
- 2026-01-29
AI Technical Summary
Current methods for simultaneous DNA and RNA sequencing from single cells are limited by low throughput, require physical separation of DNA and RNA, and are not compatible with degraded or archival samples, restricting their application and scalability.
A method called ArcDR-seq that uses probe hybridization for RNA capture and tagmentase-based DNA fragmentation, allowing simultaneous labeling and amplification of DNA and RNA without prior separation, compatible with fresh, frozen, or fixed samples, including FFPE, and enabling profiling of thousands of cells in parallel.
Enables scalable, high-throughput profiling of genome and transcriptome from single cells or low input materials, including degraded samples, reducing costs and enhancing understanding of genetic and transcriptional heterogeneity in tissues like cancer, with applications in clinical diagnostics and patient treatment.
Smart Images

Figure IMGF000039_0001 
Figure IMGF000040_0001 
Figure IMGF000041_0001
Abstract
Description
TITLE OF THE INVENTIONMETHODS AND COMPOSITIONS FOR SIMULTANEOUS PROFILING OF GENOME AND TRANSCRIPTOMECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority of U.S. Provisional Appl. Ser. No. 63 / 675,145, filed July 24, 2024, the entire disclosure of which is incorporated herein by reference.INCORPORATION OF SEQUENCE LISTING
[0002] A sequence listing containing the file named “MDCC016WO_ST26.xml” which is 9,114 bytes (measured in MS-Windows®) and created on July 21, 2025, and comprises 9 sequences, is incorporated herein by reference in its entirety.FIELD OF THE INVENTION
[0003] This present disclosure relates to the field of producing an RNA and a DNA sequencing library for the sequencing and analysis of RNA and DNA approximately simultaneously, and more specifically to methods of producing an RNA and a DNA sequencing library that allows for the approximately simultaneous profiling and analysis of RNA and DNA from a small number of cells or a single cell.BACKGROUND OF THE INVENTION
[0004] Understanding how genotype affects phenotype is as a major question in molecular biology. While genomic and transcriptomic sequencing (DNA-seq, RNA-seq) offer profound insights into genome-wide biological systems, bulk sequencing methods primarily profile mixed cell populations, yielding only an average signal across millions of cells. This cellular averaging can obscure the nuanced interplay between genetic variability and the transcriptome. Current methods known in the art, such as DR-seq, G&T-seq, SiDR-seq, and DNTR-seq, which allow DNA and RNA sequencing from the same single cell, have significant limitations. These methods are low throughput, require physical separation of DNA and RNA within a cell prior to amplification, and / or fail to add cell barcodes to DNA and RNA approximately simultaneously during the amplification process. These previous methods limit the throughputto a few hundred cells and involve extended, costly protocols. Furthermore, these previous methods require fresh and frozen samples with high RNA quality and integrity, limiting their application for the analysis of degraded or archival samples. In contrast, the compositions and methods of present disclosure provide a significant advance in the art by providing methods and compositions capable of profiling both the genome and the transcriptome approximately simultaneously from the same low input material, single cell, or thousands of single cells from fresh, frozen, or fixed samples without the need for prior DNA and RNA separation. The methods of the present disclosure provide a scalable solution to profile fresh, frozen, and archival samples, which is a continuing unmet need in the art given the extensive genetic and transcriptional heterogeneity in tissues such as cancer. In certain embodiments, the methods of the present disclosure may be referred to as ArcDR-seq. The methods of the present disclosure, in certain embodiments, utilize a probe hybridization-based method to capture RNA information and a tagmentase-based method to fragment and label DNA. In particular embodiments, the methods of the present disclosure assign unique cell barcodes to the genomic DNA and matched barcodes to the probes hybridized to the RNA, during an amplification process. After cell barcoding, in some embodiments, all RNA and DNA molecules are then pooled and used to prepare DNA and RNA sequencing libraries. The methods of the present disclosure can be used, in certain embodiments, at different sample input scales - from individual tube-based reactions to high-density nanowell reactions to process thousands of samples in parallel. Also, because the methods of the present disclosure use, in some embodiments, a probe hybridization-based strategy to capture RNA molecules, they are compatible with samples that have degraded RNA, such as formalin-fixed paraffin-embedded (FFPE) samples or long-term preserved frozen samples. In particular embodiments of the present disclosure, probes may be designed to capture any RNA molecule. Non-limiting examples of which include mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, noncoding RNA, poly-adenylated RNA, and other RNA classes. In particular embodiments, the methods of the present disclosure link genomic information and RNA phenotype in low input materials, including single cells. The methods of the present disclosure have far-ranging applications in biology and biomedicine to further the understanding of how genomic alterations impact RNA-expression, and how this impact can lead to distinct phenotypes and disease states. Due to the compatibility with archival tissue samples, the compositions and methods of the present disclosure have translational applications for improving clinical diagnostics and patient treatment.2SUMMARY OF THE INVENTION
[0005] In one aspect, the present disclosure provides a method of producing a DNA library and a RNA or cDNA library, the method comprising: a) hybridizing a probe comprising a first label to RNA or to cDNA in a sample; b) performing a labeling reaction to label genomic DNA in the sample with a second label; and c) performing an amplification reaction sufficient to amplify the first labeled RNA or cDNA and the second labeled genomic DNA approximately simultaneously. In one embodiment, the sample comprises RNA or cDNA and DNA obtained from about 1 cell to about 1,000,000,000 cells. In another embodiment, the sample comprises RNA or cDNA and DNA obtained from a single cell. In yet another embodiment, the probe is a linear probe, a circular probe, or a padlock probe. In still yet another embodiment, the method further comprises disrupting the genomic DNA chromatin structure prior to the hybridizing or prior to performing the labeling reaction to label the genomic DNA. The first labeled RNA or cDNA and the second labeled genomic DNA, in one embodiment, are not physically separated prior to performing the amplification reaction. The labeling reaction, in another embodiment, is performed by an insertional enzyme complex. The insertional enzyme complex, in yet another embodiment, comprises a transposase. The method of the present disclosure, in still yet another embodiment, may further comprise at least one DNA sequence variation, DNA modification, or transcriptional variation. In one embodiment, the methods of the present disclosure may further comprise performing a second labeling reaction to label the RNA or cDNA and the genomic DNA with a third label. In another embodiment, the second labeling reaction and the amplification reaction are performed approximately simultaneously. In yet another embodiment, or the second labeling reaction is performed prior to, after, or approximately simultaneously with hybridizing the probe comprising the first label to RNA or to cDNA in the sample. In still yet another embodiment, the hybridizing and the labeling reaction are performed approximately simultaneously. The hybridizing, in one embodiment, is performed before the labeling reaction. The first label or the second label, in another embodiment, is a nucleotide barcode label. In yet another embodiment, the first label comprises a first nucleotide barcode label, the second label comprises a second nucleotide barcode label, and the first nucleotide barcode label and the second nucleotide barcode label are different. The third label, in still yet another embodiment, comprises a third nucleotide barcode label.
[0006] In one embodiment, the sample comprises a plurality of cells or a plurality of nuclei, and the method further comprises separating the plurality of cells or plurality of nuclei into aplurality of compartments, wherein each compartment comprises about one cell. The separating, in another embodiment, is performed prior to hybridizing the probe comprising the first label to the RNA or cDNA. In yet another embodiment, the separating is performed prior to performing the labeling reaction to label the genomic DNA in the sample with the second label. The first labeled RNA or cDNA and the second labeled genomic DNA, in still yet another embodiment, are not physically separated prior to performing the amplification reaction. In one embodiment, the methods of the present disclosure may further comprise performing a sequencing reaction to obtain a sequence of the first labeled RNA or cDNA or the second labeled genomic DNA. In another embodiment, the methods of the present disclosure may further comprise combining the first labeled RNA or cDNA and the second labeled genomic DNA from each compartment of the plurality of compartments, wherein the combining is performed after performing the amplification reaction. In yet another embodiment, the methods of the present disclosure may further comprise performing a second amplification reaction to further amplify the first labeled RNA or cDNA or the second labeled genomic DNA. In still yet another embodiment, the RNA is selected from the group consisting of mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, and polyadenylated RNA. The methods of the present disclosure, in one embodiment, may further comprise performing DNA-seq analysis to detect at least one DNA sequence variation or at least one DNA modification. The methods of the present disclosure, in yet another embodiment, may further comprise performing RNA-seq analysis to detect at least one transcriptome modification.
[0007] In another aspect, the methods of the present disclosure may further comprise contacting the sample with an oligonucleotide conjugated antibody. In yet another aspect, the methods of the present disclosure may further comprise contacting the sample with a sample barcode or a spatial barcode. In one embodiment, the contacting is performed prior to, after, or approximately simultaneously with the step of hybridizing the probe comprising the first label to RNA or to cDNA in the sample.
[0008] In yet another aspect, the methods of the present disclosure may further comprise disrupting the genomic DNA chromatin structure prior to performing an amplification reaction sufficient to amplify the first labeled RNA or cDNA and the second labeled genomic DNA approximately simultaneously. In one embodiment, the labeling reaction labels at least a first portion of open chromatin region DNA comprised within the genomic DNA with the second label. In another embodiment, the methods of the present disclosure may further comprise4 1performing a second labeling reaction on the genomic DNA to label at least a first portion of exposed closed chromatin region DNA comprised within the genomic DNA with a third label. In yet another embodiment, the methods of the present disclosure may further comprise detecting at least one DNA sequence variation or DNA modification in the second labeled open chromatin region DNA or the third labeled exposed closed chromatin region DNA. The methods of the present disclosure, in still yet another embodiment, may further comprise performing a Cleavage Under Targets and Tagmentation (CUT&TAG) analysis of the open chromatin region DNA to measure proteimDNA interactions of the genomic DNA. The methods of the present disclosure, in one embodiment, may further comprise performing an Assay for Transposase-Accessible Chromatin (AT AC) analysis on the open chromatin region DNA to measure chromatin accessibility differences in the genomic DNA. The first labeling reaction or the second labeling reaction, in another embodiment, is performed by an insertional enzyme complex. The insertional enzyme complex, in yet another embodiment, comprises a transposase. In still yet another embodiment, the first labeling reaction is performed by an antibody attached insertional enzyme complex. The method, in one embodiment, may further comprise detecting at least one DNA sequence variation, DNA modification, or transcriptome variation.
[0009] In still yet another aspect, the methods of the present disclosure may further comprise digesting the sample with a restriction enzyme to fragment genomic DNA prior to hybridizing the probe comprising the first label to RNA or to cDNA in the sample. In one embodiment, the methods of the present disclosure may further comprise contacting the fragmented DNA with a DNA ligase. In another embodiment, the labeling reaction labels ligated fragments of genomic DNA. In yet another embodiment, the methods of the present disclosure may further comprise performing a second labeling reaction on the genomic DNA to label at least a first portion of non-ligated DNA comprised within the genomic DNA with a third label. In still yet another embodiment, the methods of the present disclosure may further comprise performing a high-throughput chromosome conformation capture technique analysis to detect DNA:DNA interactions of the genomic DNA. The first labeling reaction or the second labeling reaction, in one embodiment, is performed by an insertional enzyme complex. The insertional enzyme complex, in another embodiment, comprises a transposase.
[0010] In one aspect, the present disclosure provides a method of identifying a DNA sequence variation or a transcriptome variation in a sample, the method comprising: a) hybridizing a probe comprising a first label to RNA or cDNA in a sample; b) performing a labeling reaction5 1to label genomic DNA in the sample with a second label; c) performing an amplification reaction sufficient to amplify the first labeled RNA or cDNA and the second labeled genomic DNA approximately simultaneously; d) performing a sequencing reaction to obtain a sequence of the first labeled RNA or cDNA or the second labeled genomic DNA; and e) identifying the DNA sequence variation or the transcriptome variation in the sample. In one embodiment, the sample comprises RNA or cDNA and DNA obtained from about 1 cell to about 1,000,000,000 cells. In another embodiment, the sample comprises RNA or cDNA and DNA obtained from a single cell. In yet another embodiment, the first labeled RNA or cDNA and the second labeled genomic DNA are not physically separated prior to performing the amplification reaction. In still yet another embodiment, the labeling reaction is performed by an insertional enzyme complex. The insertional enzyme complex, in one embodiment, comprises a transposase. In another embodiment, the methods of the present disclosure may further comprise performing a second labeling reaction to label the RNA or cDNA and the genomic DNA with a third label. The second labeling reaction and the amplification reaction, in yet another embodiment, are performed approximately simultaneously. The hybridizing and the labeling reaction, in still yet another embodiment, are performed approximately simultaneously. The hybridizing is performed before the labeling reaction, in one embodiment. In another embodiment, the first label or the second label is a nucleotide barcode label. In yet another embodiment, the first label comprises a first nucleotide barcode label, the second label comprises a second nucleotide barcode label, and the first nucleotide barcode label and the second nucleotide barcode label are different. In yet another embodiment, the third label comprises a third nucleotide barcode label.
[0011] In one embodiment, the sample comprises a plurality of cells or a plurality of nuclei, and the method further comprises separating the plurality of cells or plurality of nuclei into a plurality of compartments, wherein each compartment comprises about one cell. The separating, in another embodiment, is performed prior to hybridizing the probe comprising the first label to the RNA or cDNA. In yet another embodiment, the method of the present disclosure may further comprise combining the first labeled RNA or cDNA and the second labeled genomic DNA from each compartment of the plurality of compartments, wherein the combining is performed after performing the amplification reaction. In still yet another embodiment, the methods of the present disclosure may further comprise performing a second amplification reaction to further amplify the first labeled RNA or cDNA or the second labeled genomic DNA. The RNA, in certain embodiments, is selected from the group consisting of6 1mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, and polyadenylated RNA. The methods of the present disclosure, in one embodiment, may further comprise disrupting the genomic DNA chromatin structure prior to the hybridizing or prior to performing the labeling reaction to label the genomic DNA. In another embodiment, the methods of the present disclosure may further comprise identifying the DNA sequence variation in the amplified second labeled genomic DNA. In yet another embodiment, the DNA sequence variation is selected from the group consisting of a mutation, a copy number alteration, a single nucleotide polymorphism, an insertion, a deletion, a short tandem repeat, a translocation, an inversion, a structural variation, and a DNA modification. In still yet another embodiment, identifying the transcriptome variation comprises identifying a decreased expression level of a transcript, an increased expression level of a transcript, an increased expression level of a splice variant, a decreased expression level of a splice variant, an increased expression level of a transcript isoform, a decreased expression level of a transcript isoform, a modification of a transcript nucleotide sequence, a post-transcriptional modification, or a post-transcriptional rearrangement. The DNA sequence variation or transcriptome variation, in one embodiment, is associated with a condition selected from the group consisting of cancer, a genetic disease or condition, a developmental disease or condition, or an immunological disease or condition.
[0012] In another aspect, the present disclosure provides a kit comprising: a) a transposase; b) a probe comprising a first label, wherein the probe is capable of hybridizing to RNA or cDNA; c) an adaptor molecule for labeling genomic DNA with a second label; and d) reagents for performing an amplification reaction. The kit, in one embodiment, further comprises a reagent for disrupting chromatin structure. The kit, in another embodiment, further comprises a first set of primers for amplifying the first labeled RNA or cDNA and a second set of primers for amplifying the second labeled genomic DNA. The probe or the adaptor molecule, in yet another embodiment, further comprises a third label comprising an oligonucleotide barcode sequence. The first set of primers or the second set of primers, in still yet another embodiment, further comprises a third label comprising an oligonucleotide barcode sequence.
[0013] In yet another aspect the present disclosure provides an RNA, a cDNA, or a DNA library produced by the methods of the present disclosure. In some embodiments, the present disclosure provides an RNA library, a cDNA library, a DNA library, an open chromatin region DNA library, an ATAC library,, a CUT&TAG library, and / or an DNA library of ligated DNA fragments produced by the methods of the present disclosure.7 1BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0015] FIG. 1 provides an example of a workflow of a method of the present disclosure. The method may be performed, in certain embodiments, in single tubes, 96 / 384 PCR plates, micro / nanowell chips, droplets, or other compartments. In some embodiments, cell suspensions can be obtained from dissociation of fresh tissues, flash-frozen tissues, or fixed tissues including FFPE blocks and other archival samples. Cell suspensions are first permeabilized and / or fixed, then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (Step 1). Then, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nanowells, or microdroplets (Step 2). The dispensed single cell or low input material of each tube / well is lysed to release the DNA and RNA (Step 3), then DNA adaptors are added to the gDNA fragments (using tagmentation reaction with Tn5 transposome in this example protocol) (Step 4). Primers are added that can then amplify DNA and probes (RNA) followed by DNA and probe co-amplification. The primers used for amplifying DNA and probes can have barcode sequences (e.g., bcl-bc4 shown in FIG. 1), which are used as the well / tube / cell barcode for labeling DNA and RNA molecules from the same well / tube / cell (Step 5). To either or both of these modalities (e.g., RNA or DNA), a pre-amplification step that specifically enriches that modality can be performed in this protocol. Barcoded DNA and probes (targeting RNA) from all reaction tubes / wells are then pooled together (Step 6). DNA and probe (RNA) libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA and RNA sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA and probe (RNA) libraries are separated by using specific PCR primers to amplify the DNA and probe (RNA) mixture, respectively.
[0016] FIG. 2 provides an example of a workflow of a method of the present disclosure. The method may be performed, in certain embodiments, in single tubes, 96 / 384 PCR plates, micro / nanowell chips, droplets, or other compartments. In some embodiments, cell suspensions can be obtained from dissociation of fresh tissues, flash-frozen tissues, or fixed tissues including FFPE blocks and other archival samples. Cell suspensions are first permeabilized and / or fixed, then stained with oligonucleotides conjugated antibodies whichcan serve as antibody -derived tags (ADT), sample or spatial barcoding (SB) (Step 1). The stained cells are further hybridized with DNA or RNA probes that specifically bind to RNA (or cDNA) molecules (Step 2). Then, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nanowells, or microdroplets (Step 3). The dispensed single cell or low input material of each tube / well is lysed to release the DNA, RNA, and the ADT or SB (Step 4), then DNA adaptors are added to the gDNA fragments (using tagmentation reaction with Tn5 transposome in this example protocol) (Step 5). Primers are added that can then amplify DNA, probes (RNA), and ADT or SB followed by DNA, probe, and ADT or SB co- amplification (Step 6). The primers used for amplifying DNA, probes, and ADT or SB can have barcode sequences (e.g., bcl-bc6 shown in FIG. 2), which are used as the well / tube / cell barcode for labeling DNA, RNA, and ADT or SB molecules from the same well / tube / cell (Step 6). To any or all of these modalities (e.g., RNA, DNA, ADT, or SB), a pre-amplification step that specifically enriches that modality can be performed in this protocol. Barcoded DNA, probes (targeting RNA), and ADT or SB from all reaction tubes / wells are then pooled together (Step 7). DNA, probe (RNA), and ADT or SB libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA, RNA, and ADT or SB sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA, probe (RNA), and ADT or SB libraries are separated by using specific PCR primers to amplify the DNA, probe (RNA), and ADT or SB mixture, respectively.
[0017] FIG. 3 provides an example of a workflow of a method of the present disclosure. The method may be performed, in certain embodiments, in single tubes, 96 / 384 PCR plates, micro / nanowell chips, droplets, or other compartments. In some embodiments, cell suspensions can be obtained from dissociation of fresh tissues, flash-frozen tissues, or fixed tissues including FFPE blocks and other archival samples. Cell suspensions are first permeabilized and / or fixed, then the open chromatin is labeled with DNA adaptors (through bulk tagmentation reaction with Tn5 transposome in the example protocol) (Step 1). The tagmented cells are further hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (Step 2). Then, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nanowells, or microdroplets (Step 3). The dispensed single cell or low input material of each tube / well is lysed to release the DNA, RNA, and AT AC fragments (Step 4), then DNA adaptors are added to the gDNA fragments (through 2ndround of tagmentation reaction with Tn5 transposome in this example protocol) (Step 5).9 1Primers are added that can then amplify DNA, probes (RNA), and AT AC fragments followed by DNA, probe, and ATAC co-amplification (Step 6). The primers used for amplifying DNA, probes, and ATAC fragments can have barcode sequences (e.g., bcl-bc6 shown in FIG. 3), which are used as the well / tube / cell barcode for labeling DNA, RNA and ATAC molecules from the same well / tube / cell (Step 6). To any or all of these modalities (e.g., RNA or DNA or ATAC), a pre-amplification step that specifically enriches that modality can be performed in this protocol. Barcoded DNA, probes (targeting RNA), and ATAC fragments from all reaction tubes / wells are then pooled together (Step 7). DNA, probe (RNA), and ATAC libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA, RNA, and ATAC sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA, probe (RNA), and ATAC libraries are separated by using specific PCR primers to amplify the DNA, probe (RNA), and ATAC fragments mixture, respectively.
[0018] FIG. 4 provides an example of a workflow of a method of the present disclosure. The method may be performed, in certain embodiments, in single tubes, 96 / 384 PCR plates, micro / nanowell chips, droplets, or other compartments. In some embodiments, cell suspensions can be obtained from dissociation of fresh tissues, flash-frozen tissues, or fixed tissues including FFPE blocks and other archival samples. Cell suspensions are first permeabilized and / or fixed, then the DNA-protein interaction is labeled with DNA adaptors (through bulk antibody -directed tagmentation reaction with pA-Tn5 transposome in the example protocol) (Step 1). The tagmented cells are further hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (Step 2). Then, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nano wells, or microdroplets (Step 3). The dispensed single cell or low input material of each tube / well is lysed to release the DNA, RNA, and CUT&TAG fragments (Step 4), then DNA adaptors are added to the gDNA fragments (through 2ndround of tagmentation reaction with Tn5 transposome in this example protocol) (Step 5). Primers are added that can then amplify DNA, probes (RNA), and CUT&TAG fragments followed by DNA, probe, and CUT&TAG coamplification (Step 6). The primers used for amplifying DNA, probes, and ATAC fragments can have barcode sequences (e.g., bcl-bc6 shown in FIG. 4), which are used as the well / tube / cell barcode for labeling DNA, RNA and CUT&TAG molecules from the same well / tube / cell (Step 6). To any or all of these modalities (e.g., RNA or DNA or CUT&TAG), a pre-amplification step that specifically enriches that modality can be performed in thisIO 1protocol. Barcoded DNA, probes (targeting RNA), and CUT&TAG fragments from all reaction tubes / wells are then pooled together (Step 7). DNA, probe (RNA), and CUT&TAG libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA, RNA, and CUT&TAG sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA, probe (RNA), and CUT&TAG libraries are separated by using specific PCR primers to amplify the DNA, probe (RNA), and CUT&TAG fragments mixture, respectively.
[0019] FIG. 5 provides an example of a workflow of a method of the present disclosure. The method may be performed, in certain embodiments, in single tubes, 96 / 384 PCR plates, micro / nanowell chips, droplets, or other compartments. In some embodiments, cell suspensions can be obtained from dissociation of fresh tissues, flash-frozen tissues, or fixed tissues including FFPE blocks and other archival samples. Cell suspensions are first permeabilized and / or fixed, then the DNA is fragmented (through restriction enzyme digestion in this example protocol) (Step 1) and the linked DNA-DNA interaction loci (Hi-C) is further linked through ligation (Step 2). The cells are then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (Step 3) and the linked DNA-DNA interaction loci (Hi-C) are labeled with DNA adaptors (through bulk tagmentation reaction with Tn5 transposome in the example protocol) (Step 4). Then, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nano wells, or microdroplets (Step 5). The dispensed single cell or low input material of each tube / well is lysed to release the DNA, RNA, and Hi-C fragments (Step 6), then DNA adaptors are added to the gDNA fragments (through 2ndround of tagmentation reaction with Tn5 transposome in this example protocol) (Step 7). Primers are added that can then amplify DNA, probes (RNA), and Hi-C fragments followed by DNA, probe, and Hi-C co-amplification (Step 8). The primers used for amplifying DNA, probes, and Hi-C fragments can have barcode sequences (e.g., bcl-bc6 shown in FIG. 5), which are used as the well / tube / cell barcode for labeling DNA, RNA and Hi-C molecules from the same well / tube / cell (Step 8). To any or all of these modalities (e.g., RNA or DNA or Hi-C), a pre-amplification step that specifically enriches that modality can be performed in this protocol. Barcoded DNA, probes (targeting RNA), and Hi-C fragments from all reaction tubes / wells are then pooled together (Step 9). DNA, probe (RNA), and Hi-C libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA, RNA, and Hi-C sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA, probe (RNA), and Hi-1 1 1C libraries are separated by using specific PCR primers to amplify the DNA, probe (RNA), and Hi-C fragments mixture, respectively.
[0020] FIG. 6 provides exemplary probe designs for capturing RNA and cDNA molecules. FIG. 6, Panel A and FIG. 6, Panel B show probe hybridization to mRNA having a polyA tail and cDNA having a polyA tail, respectively. FIG. 6, Panel C, Panel D, Panel E, and Panel F show the use of exemplary probes which may comprise one or more probe barcodes, cell barcodes, and / or unique molecule identifiers (UMI). The use of these one or more barcodes or UMIs may distinguish cells, probes, and / or RNA or cDNA molecules or serve as adaptors or templates in subsequent reactions steps.
[0021] FIG. 7 provides the results of varying ligation conditions on RNA probe ligation efficiency. FIG. 7, Panel A shows the result of varying temperature conditions and ligase concentrations on the efficiency of RNA probe ligation. FIG. 7, Panel B shows the result of varying ligation reaction time on the efficiency of RNA probe ligation.
[0022] FIG 8 shows the copy number profiles of seven different individual cells. The x-axis is the genome coordinates from 0 to 3GB for each cell. The y-axis is the copy number ratio for each cell.
[0023] FIG. 9 shows a UMAP showing 11 clusters of single cells, representing different cell types from a normal human breast tissue identified by gene expression data following the simultaneous profiling of the genome and transcriptome of cells derived from normal breast tissue. The cell type abbreviations include luminal secretory epithelial cell (LumSec), luminal hormone response epithelial cell (LumHR), fibroblasts (Fibroblast) and endothelial cells (Endothelial).
[0024] FIG. 10 shows the results of approximately simultaneous profiling of the genome and transcriptome of cells dissociated from a flash-frozen human breast cancer tissue. FIG. 10, Panel A shows a UMAP showing 10 clusters of single cells, representing different cell types identified by the gene expression data. FIG. 10, Panel B shows the gene number per cell detected by each cell clusters identified from the RNA data. FIG. 10, Panel C shows a heatmap of DNA copy number aberrations of single cells, with the DNA subclones (subclones) and RNA cell type (RNA_Clone) sidebar showing the matched DNA subclones and RNA clustering results.
[0025] FIG. 11 shows the results of approximately simultaneous profiling of the genome and transcriptome of cells dissociated from a formalin-fixed paraffin embedded (FFPE) human12 1ductal carcinoma in situ (DCIS) tissue. FIG. 11, Panel A shows a UMAP showing 10 clusters of single cells, representing different cell types identified by the gene expression data. FIG. 11 , Panel B shows the number of genes per cell across all detected cell types based on the RNA data. FIG. 11, Panel C shows a heatmap of the DNA copy number aberrations of single cells, with the DNA superclones (Superclones), DNA subclones (subclones), and RNA cell type (RNA_Clone) annotation sidebar showing the matched DNA clones and RNA clustering results.BRIEF DESCRIPTION OF THE SEQUENCES
[0026] SEQ ID NO:1 - A representative mosaic end nucleotide sequence that is recognized by a transposase.
[0027] SEQ ID NO:2 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to probe labeled RNA.
[0028] SEQ ID NO:3 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify tagmented DNA.
[0029] SEQ ID NO:4 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to probe labeled RNA.
[0030] SEQ ID NO:5 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify tagmented DNA.
[0031] SEQ ID NO:6 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify labeled DNA.
[0032] SEQ ID NO:7 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify labeled DNA.
[0033] SEQ ID NO:8 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify labeled RNA.
[0034] SEQ ID NO:9 - The nucleotide sequence of a representative PCR primer that may be used in an amplification reaction to amplify labeled RNA.DETAILED DESCRIPTION OF THE INVENTION
[0035] The present disclosure provides a novel method that for the first time allows for approximately simultaneous profiling of the genome and transcriptome from a single cell or13 1low input material derived from cell lines, or fresh, frozen, fixed, degraded, or archival FFPE samples. Importantly, the methods of the present disclosure are highly scalable and can be used to profile thousands of cells in parallel, which vastly reduces the cost per cell to profile RNA or cDNA and DNA approximately simultaneously. In one embodiment, the methods of the present disclosure may be referred to as ArcDR-seq. The methods of the present disclosure may be used, in particular embodiments, in the fields of cancer genomics, single-cell genomics, developmental biology, prenatal genetic diagnosis, tissue mosaicism, neurological diseases, neuroscience research, microbiology, and pathogenesis studies. In some embodiments, the methods of the present disclosure may be used in early disease diagnosis, disease surveillance, the detection of minimal residual disease, the development of novel predictive and prognostic biomarkers, and the identification of novel actionable therapeutic targets. The methods of the present disclosure can accommodate a wide range of clinical input materials, including small tissue samples, blood drops, buffy coat, various body fluids, swabs, and patient-derived materials such as organoids and patient-derived xenografts (PDX).A. Method Overview
[0036] In some embodiments of the present disclosure, the methods described herein may be referred to as ArcDR-seq. The methods of the present disclosure, in certain embodiments, may use a tagmentase (e.g., Tn5 transposase) to capture DNA and probe-based hybridization to capture RNA. The barcoded DNA and RNA, in other embodiments, may then be amplified approximately simultaneously without physical separation of the DNA and RNA from single cells or low input materials prior to pooling all of the barcoded libraries together (FIG. 1). In some embodiments, the barcoded DNA and RNA libraries are separated after pooling of all of the barcoded libraries from all of the cells to further construct the DNA and RNA sequencing libraries (separately), before loading to sequencers for high throughput sequencing. In some embodiments, cell suspensions are first permeabilized and / or fixed, then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (FIG. 1, Step 1). In some embodiments, cell suspensions are first permeabilized and / or fixed, stained with oligonucleotide conjugated antibody, and then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (FIG. 2, Step 1). Methods for contacting and labeling cells with oligonucleotide conjugated antibodies are known in the art and any such method may be used according to the methods of the present disclosure. For example, Stoeckius et al., Nature Methods 14:865-868, 2017, describes cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq), a method in which oligonucleotide-14 1labeled antibodies are used to integrate cellular protein and transcriptome measurements into a single-cell readout. In certain embodiments, cell suspensions are first permeabilized and / or fixed labeled with a sample barcode or a spatial barcode, and then hybridized with DNA or RNA probes (FIG. 2, Step 1). Methods for labeling cells with sample and / or spatial barcodes are known in the art and any such method may be used according to the methods of the present disclosure. Non- limiting examples of such methods are described in Wirth et al., Nature Communications 14:1523, 2023, Shi et al., JACS Au 4(5): 1723- 1743, 2024, and Stahl et al., Science 353(6294)78-82, 2016. In some embodiments, cell suspensions are first permeabilized and / or fixed, labeled by open chromatin specific DNA adaptors (e.g., through bulk tagmentation with Tn5 transposome), and then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (FIG. 3, Step 1). As used herein the term “open chromatin region,” “open chromatin,” or “open chromatin region DNA” refers to a nucleosome depleted region of DNA that can be accessed by DNA regulatory elements. DNA regulatory elements have limited or no access to regions of nucleosome dense regions of DNA. These regions are referred to herein as a “closed chromatin region,” “closed chromatin,” or “closed chromatin region DNA.” An amplified open chromatin region DNA product may be, in certain embodiments, from a chromatin region comprising acetylated histones or decreased DNA methylation, whereas an amplified closed chromatin region DNA product may be, in particular embodiments, be from a chromatin region comprising deacetylated histones or increased DNA methylation. In some embodiments, an open chromatin region may refer to a chromatin region that is transcriptionally active, whereas a closed chromatin region may refer to a chromatin region that is transcriptionally inactive. Amplified open chromatin region DNA and closed chromatin region DNA products can be used to directly measure the effect of chromatin structure, and modifications thereof, on gene transcription and to study the regulatory function of nucleosome positioning. These amplified products may additionally be used to study DNA binding and regulation of the general transcription machinery, and how certain genomic regions regulate gene expression. In some embodiments, an Assay for Transposase- Accessible Chromatin (AT AC) analysis may be performed on the open chromatin region DNA to measure chromatin accessibility differences in the genomic DNA.
[0037] In some embodiments, cell suspensions are first permeabilized and / or fixed, labeled by DNA-protein interactions specific DNA adaptors (e.g., through antibody-assisted bulk tagmentation with pA-Tn5 transposome), and then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (FIG. 4, Step 1). Methods of tagmentation15 1using antibody-assisted tagmentation are known in the art and any such method may be used according to certain embodiments of the present disclosure. See, for example, the CUT&TAG method described in Kaya-Okur, et al., Nature Communications 10: 1930, 2019. In some embodiments, cell suspensions are first permeabilized and / or fixed, labeled by DNA-DNA interactions specific DNA adaptors (e.g., through restriction enzyme digestion, ligation, and bulk tagmentation with Tn5 transposome), and then hybridized with DNA or RNA probes to specifically bind to RNA (or cDNA) molecules (FIG. 5, Step 1). Such methods are known in the art, any of which may be used according to embodiments of the present disclosure. See, for example, Belton, et al., Methods 58(3):268-276, 2012. Then, in other embodiments, the suspensions are partitioned into single cells, by dispensing into single tubes, plate wells, micro / nano wells, or microdroplets (Fig. 1, Step 2). The dispensed single cell or low input material of each tube / well, in still other embodiments, is lysed to release the DNA and RNA (FIG. 1, Step 3), then, in still other embodiments, the DNA adaptors are added to the gDNA fragments (using tagmentation reaction with Tn5 transposome in this example protocol) (FIG. 1, Step 4). Primers, in some embodiments, are added that can then amplify DNA and probes (RNA) followed by DNA, and probe co-amplification (FIG. 1, Step 5). The primers used for amplifying DNA and probes, in other embodiments, can have barcode sequences (e.g., bcl- bc6 shown in FIG. 1-5), which are used as the well / tube / cell barcode for labeling DNA and RNA molecules from the same well / tube / cell (FIG. 1, Step 5). To any or all of these modalities (e.g., RNA or DNA), in some embodiments, a pre-amplification step that specifically enriches that modality can be performed in this protocol. Barcoded DNA and probes (targeting RNA) from all reaction tubes / wells, in certain embodiments, are then pooled together (FIG. 1, Step 6). In particular embodiments, DNA and probe (RNA) libraries can be separated according to fragment sizes, biotin modifications, or other methods followed by preparing the pooled DNA and RNA sequencing libraries according to the research goals and sequencing instruments. In this example, the DNA and probe (RNA) libraries are separated by using specific PCR primers to amplify the DNA and probe (RNA) mixture, respectively.
[0038] The input materials that may be used according to certain embodiments of the present disclosure include, but are not limited to, a single cell / nucleus, multiple cells / nuclei, and low input materials. The sample type can be, in certain embodiments, genetically engineered, fixed, lysed, permeabilized, tagmented, digested, or antibody attached cells / nuclei, or organoids, small chunks / pieces of tissue, blood drops, buffy coat, body fluids, swabs, or naked DNA / RNA. To achieve the preparation DNA and RNA libraries approximately16 1simultaneously, in some embodiments, the first step may be to either hybridize DNA / RNA probes to RNA molecules to capture RNA information before lysis or to lyse the input material to release DNA and RNA into the reaction before the hybridization. Any lysis procedure known in the art may be used according to the embodiments of the present disclosure. Nonlimiting examples of which include an enzyme-based method (e.g., protease), a chemical-based method (e.g., detergents such as Tween-20 or Triton X-100), a mechanical based method, an acoustic based method, an electrical based method, and any combination of these methods. In some embodiments, RNase inhibitors can be added at this step to further prevent RNA degradation during lysis or during probe hybridization.
[0039] The present disclosure provides methods that allow for the independent labeling of RNA or cDNA and DNA followed by amplification of the labeled or barcoded materials approximately simultaneously, without physical separation of the RNA and DNA prior to amplification. The methods provided by the present disclosure may be used to approximately simultaneously amplify RNA or cDNA and DNA from a single cell obtained from a sample comprising about 1 cell to about 1,000,000,000 cells, including all ranges derivable therebetween. The sample may comprise, for example, about 1 cell to about 1,000,000,000 cells, about 1 cell to about 500,000,000 cells, about 1 cell to about 100,000,000 cells, about 1 cell to about 50,000,000 cells, about 1 cell to about 10,000,000 cells, about 1 cell to about 5,000,000 cells, about 1 cell to about 1,000,000 cells, about 1 cell to about 500,000 cells, about 1 cell to about 100,000 cells, about 1 cell to about 50,000 cells, about 1 cell to about 10,000 cells, about 1 cell to about 5,000 cells, about 1 cell to about 1,000 cells, about 1 cell to about 500 cells, about 1 cell to about 100 cells, about 1 cell to about 50 cells, about 1 cell to about 25 cells, or about 1 cell to about 10 cells, including all ranges derivable therebetween. In some embodiments, RNA or cDNA and DNA from all samples or sub-samples may be pooled together prior to amplification. In certain embodiments, the labeled RNA or cDNA and DNA libraries may be separated after pooling all of the libraries of all samples or sub-samples to construct the RNA or cDNA and DNA libraries individually. The RNA or cDNA and DNA libraries may then, in some embodiments, be separated prior to high-throughput sequencing. In certain embodiments, methods of the present disclosure may include one or more of the steps of: (1) hybridizing the probes to RNA molecules to capture the RNA information; (2) removing chromatin before fragmenting the DNA with tagmentase; (3) using tagmentase to fragment the DNA and add DNA adaptors approximately simultaneously for labeling the DNA molecules; (4) adding cell barcodes to the probes and fragmented DNA molecules from the17 1same cell by PCR amplification or ligation-based methods; (5) pooling of all of molecules with cell barcodes together, and then performing enrichment reactions (e.g. PCR) to enrich the DNA molecules or probes separately to separate the DNA and RNA libraries; or (6) preparing DNA and probe sequencing libraries to perform high throughput sequencing and data analysis, which involves computationally matching the data from DNA and RNA to perform more detailed downstream analysis. In some embodiments, since different barcodes, adaptors, or labels are used to label the RNA or cDNA and DNA, the methods provided herein can be used to enrich one of these modalities, for example the RNA or cDNA, as desired prior to or after exponential co-amplification. This step may be performed, in certain embodiments, when the concentration of one modality concentration is significantly less than the other. In particular embodiments, the analysis of sequencing data may include computationally matching the data from the RNA or cDNA and DNA and performing a more detailed analysis.
[0040] The methods of the present disclosure are able to add a unique sample or cell barcode to RNA or cDNA and DNA using a multiplexing PCR reaction. This allows for the identification of RNA or cDNA and DNA of each cell or sample following amplification. The methods provided herein thus allow for the combined preparation RNA or cDNA and DNA sequencing libraries from all of the samples, sub-samples, or cells together. This significantly reduces the labor and resource input required to prepare individual RNA or cDNA and DNA libraries from each cell, sample, or sub-sample one at time. This feature of the methods of the present disclosure makes them highly scalable to tube format, plate format, or high-throughput platforms such as nanowells, nanochips, or nanodroplets.
[0041] The methods of the present disclosure have broad application in cancer genomics, single cell genomics, pre-natal genetic diagnosis, drug-target discovery, forensics, neurological disease, neuroscience, microbiology, pathogenesis, and development. The methods provided by the present disclosure additionally have broad application in low-input RNA or cDNA and DNA library construction and highly multiplexed RNA or cDNA and DNA preparation from single cells, multiple cells, or cell mixtures. The methods provided by the present disclosure can also be used in many clinical and translational applications such as early disease diagnosis, disease monitoring, minimal residual disease detection, novel predictive and prognostic biomarkers development, and novel actionable target identification using samples that may include, but are not limited to, small chunks or pieces of tissue, blood droplets, buffy coat, body fluids, swabs, and patient-derived materials such as organoids and patient-derived xenografts.18 1B. Samples and Sample Preparation
[0042] The compositions and methods of present disclosure provide a significant advance in the art by providing methods and compositions capable of profiling both the genome and the transcriptome approximately simultaneously from the same low input material, single cell, or thousands of single cells from fresh, frozen, or fixed samples without the need for prior DNA and RNA separation. The methods of the present disclosure provide a scalable solution to profile fresh, frozen, fixed / archival, and degraded samples, which is a continuing unmet need in the art given the extensive genetic and transcriptional heterogeneity in tissues such as cancer.
[0043] The methods of the present disclosure are able to amplify RNA or cDNA and DNA from material derived from single cells, tens of cells, hundreds of cells, thousands of cells, or millions of cells. For example, RNA or cDNA and DNA may be approximately simultaneously amplified from material derived from about 1 cell to about 1,000,000,000 cells, about 1 cell to about 500,000,000 cells, about 1 cell to about 100,000,000 cells, about 1 cell to about 50,000,000 cells, about 1 cell to about 10,000,000 cells, about 1 cell to about 5,000,000 cells, about 1 cell to about 1,000,000 cells, about 1 cell to about 500,000 cells, about 1 cell to about 100,000 cells, about 1 cell to about 50,000 cells, about 1 cell to about 10,000 cells, about 1 cell to about 5,000 cells, about 1 cell to about 1 ,000 cells, about 1 cell to about 500 cells, about 1 cell to about 100 cells, about 1 cell to about 50 cells, about 1 cell to about 25 cells, or about 1 cell to about 10 cells, including all ranges derivable therebetween. In some embodiments, RNA or cDNA and DNA may be approximately simultaneously amplified from extracted, unextracted, purified, or isolated cells or nuclei from fresh, frozen, or fixed samples. In some embodiments, a sample comprising RNA, cDNA, or DNA for use according to certain embodiments of the present disclosure may be derived from cells, nuclei, organoids, small chunks / pieces of tissues, blood drops, buffy coat, body fluids, or swabs. In one embodiment, the sample may comprise naked RNA, cDNA, or DNA. In particular embodiments of the present disclosure, the RNA to be amplified or analyzed according to the methods of the present disclosure may be any RNA molecule. Non-limiting examples of which include mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, poly-adenylated RNA, and other RNA classes. The cell or cells from which RNA / cDNA and DNA are approximately simultaneously amplified may be from any eukaryotic organism. Non-limiting examples of such organisms include eukaryotic single-celled organisms, eukaryotic multicellular organisms, fungi, plants, animals, reptiles, amphibians, insects,19 1mammals, and humans. Samples for use in the methods provided by the present disclosure may be derived from any eukaryotic cell or tissue, non-limiting examples which include chunks or pieces of tissue, blood droplets, huffy coats, body fluids, tissue swabs, organoids, and patient-derived xenografts. As used herein, the terms “approximately simultaneously” or “approximately simultaneous” when used in reference to amplification of RNA, cDNA, and / or DNA refers to the condition where amplification is performed without physically separating the DNA and RNA / cDNA prior to amplification. In particular embodiments, RNA / cDNA and DNA may be amplified approximately simultaneously when the RNA / cDNA and DNA are amplified at approximately the same time. In further embodiments, DNA and RNA / cDNA may be amplified approximately simultaneously when the DNA and RNA / cDNA are amplified within the same amplification reaction mixture.
[0044] In some embodiments, a sample comprising RNA, cDNA, or DNA may comprise a plurality of cells. For example, the sample may comprise at least 2, at least 5, at least 10, at least 25, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, at least 80,000, at least 90,000, at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 600,000, at least 700,000, at least 800,000, at least 900,000, at least 1 ,000,000, or at least 10,000,000 cells, including all ranges derivable therebetween. In particular embodiments, the sample may be separated into a plurality of sub-samples and each sub-sample may comprise about 1 cell to about 50 cells, about 1 cell to about 40 cells, about 1 cell to about 30 cells, about 1 cell to about 20 cells, about 1 cell to about 15 cells, about 1 cell to about 10 cells, or about 1 cell to about 5 cells, including all ranges derivable therebetween. In specific embodiments, each sub-sample may comprise a single cell. Methods of separating cell suspensions are known in the art and any such method may be used according to the methods provided by the present disclosure. Non-limiting examples of such methods include flow cytometry and cell distribution using single cell microdispensers.
[0045] Cells or nuclei for use in the present disclosure may be permeabilized, fixed, antibody- stained, engineered, perturbed, tagmented, modified, digested, or labeled. Methods for cell permeabilization, fixation, antibody-staining, tagmentation, engineering, modifying, enzyme digestion, and labeling are known in the art and any such method may be used according to the methods provided by the present disclosure. For example, cells or nuclei may be permeabilized using organic solvents, non-limiting examples of which include acetone and methanol, or by using detergents, non-limiting examples of which include saponin, Triton X-20 1100, and Tween-20. Cell or nuclei fixation may be performed using chemical or physical methods. For example, cells may be fixed using cross-linking agents such as formaldehyde, glutaraldehyde, and succinimide esters, or by using solvents. In certain embodiments, cells may be fixed using heat, microwaving, or cryopreservation methods. Cells or nuclei for use in the present disclosure may, in some embodiments, be immunostained, engineered to express a protein of interest, or labeled using, for example, a fluorescent dye. In particular embodiments, cell or nuclei for use in the present disclosure may be prepared into a single cell or single nuclei suspension.C. RNA and cDNA Labeling
[0046] To approximately simultaneously prepare RNA or cDNA and DNA libraries, the RNA or cDNA must be labeled with a first set of labels. These labels may be, for example, a set of oligonucleotides or barcodes. As used herein, the term "label" refers to a directly or indirectly detectable oligonucleotide or nucleotide modification that is conjugated directly or indirectly to the composition to be detected. Non-limiting examples of a nucleotide modification that may be present, in certain embodiments, in a label as described herein include DNA methylation, a biotin label, a fluorescent label, or a chemical modification. Methods to introduce oligonucleotide labels or barcodes are known in the art and any such method may be used according to the methods provided by the present disclosure. As used herein the terms “oligonucleotide:” “polynucleotide,” and “nucleic acid” may be used interchangeably and include linear oligomers of natural or modified monomers or linkages. An oligonucleotide may include, for example, deoxyribonucleosides, ribonucleosides, a-anomeric forms thereof, peptide nucleic acids, and the like, capable of specifically binding to a target polynucleotide by way of a regular pattern of monomer- to-monomer interactions, such as Watson-Crick type of base pairing, base stacking, Hoogsteen, or reverse Hoogsteen type base pairing. Monomers may be linked, in some embodiments, by a phosphodiester bond or an analog thereof to form oligonucleotides ranging in size from a few monomeric units to several tens of monomeric units. Whenever an oligonucleotide is represented by a sequence of letters herein, a person of ordinary skill in the art would understand that the nucleotides are in 5' to 3' orientation from left to right. A person of ordinary skill in the art would further understand that if an oligonucleotide is presented as a sequence of letters that “A” denotes adenine, “C” denotes cytosine, “G” denotes guanine, and “T” denotes thymine, and “U” denotes uracil, unless otherwise noted. Analogs of phosphodiester linkages include, but are not limited to, phosphorothioate, phosphorodithioate, phosphoranilidate, and phosphoramidate linkages. It21 1is clear to those skilled in the art when oligonucleotides having natural or non-natural nucleotides may be employed. For example, a person of ordinary skill in the art would understand when processing by enzymes may be employed or when oligonucleotides consisting of natural nucleotides are required.
[0047] In certain embodiments of the present disclosure, RNA or cDNA may be labeled using a probe. As used herein the term “probe” refers to a nucleic acid molecule that is complementary to a strand of a target nucleic acid and useful in hybridization detection or labeling methods. A probe for use in certain embodiments of the present disclosure may be an RNA probe or a DNA probe and may be designed to hybridize with RNA molecules or cDNA molecules (FIG. 6, Panel A and FIG. 6, Panel B). The probes can be DNA or RNA based probes that either hybridize to RNA molecules and / or cDNA molecules. In some embodiments, probes of the present disclosure may be designed to hybridize to total RNA, mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, poly-adenylated RNA, or a specific targeted RNA panel of interest. In some embodiments, probes comprise universal sequence regions in the PCR handle for subsequent PCR steps. Probe hybridization, in other embodiments, may be performed for one round or for multiple rounds. In one embodiment, for example, a first round of probe hybridization may be performed to specifically bind RNA and / or cDNA, and second and / or third rounds of probe hybridization may be used to capture the first-round probes. For each RNA and / or cDNA target one or multiple probes may be used for specific hybridization. In some embodiments, hybridized probes may be used individually or may be ligated together following hybridization, such that the ligated probes function as one large probe. In other embodiments, multiple hybridized probes can be separated on a target, which can be further extended by DNA / RNA polymerase or reverse transcriptase to seal the gaps between the probes. In still other embodiments, probes may be circularized (e.g., padlock) or linear. In some embodiments, the probes may comprise a part of sequence that functions as probe barcode. In other embodiments, the probes may comprise part of sequence that functions as unique molecular identifier (UMI). In yet other embodiments, the probes may comprise part of sequence that functions as a cell barcode for later cell barcoding steps. In still yet other embodiments, the probes may comprise a probe barcode, a UMI and / or a cell barcode. A probe, in some embodiments, can have DNA / RNA modifications, such as biotin-modifications. The DNA or RNA probes, in other embodiments, can hybridize to cDNA molecules rather than RNA molecules. In yet other embodiments, when probes are designed to hybridize to cDNA, RNA are first converted into cDNA22 1(complimentary DNA) by a reverse transcriptase. Non-limiting examples of reverse transcriptases include MMLV reverse transcriptase, AMV reverse transcriptase, and any other RT enzyme known in the art.
[0048] As used herein the term “primer” refers to a nucleic acid molecule that is designed for use in annealing or hybridization methods that involve an amplification reaction. In some embodiments, primers used for reverse transcription (RT) can prime RNA molecules with poly-T tails, or random nucleotide sequences, as well as other targeted gene-specific sequences. In one embodiment, a primer of the present disclosure may also comprise a universal sequence component, a cell barcode, or part of a cell barcode, which serves to distinguish cell identity. In another embodiment, a primer of the present disclosure may have a unique molecular identifier (UMI) sequence, which distinguishes individual transcripts from PCR duplicate reads. A cell barcode of the present disclosure may comprise, in particular embodiments, one, two, or multiple oligonucleotide sequences.
[0049] A nucleic acid for use in the methods of the present disclosure may be a modified or unmodified and may comprise RNA nucleotides or DNA nucleotides. In some embodiments, a nucleic acid molecule for use in the present disclosure is an unmodified oligonucleotide. An unmodified oligonucleotide may be composed, in certain embodiments, of nucleobases, sugars, and covalent internucleoside linkages. The term “oligonucleotide analog” as used herein refers to oligonucleotides that have one or more non-naturally occurring segments. As used herein the term oligonucleotide also includes oligonucleotide analogs. Oligonucleotide analogs may function, in certain embodiments, similarly, to naturally occurring oligonucleotides. A person of ordinary skill in the art would understand that oligonucleotide analogs may, in some embodiments, have desirable properties such as, for example, enhanced cellular uptake, enhanced affinity for other oligonucleotide or nucleic acid targets, or increased stability in the presence of nucleases. In certain embodiments, oligonucleotides for use in the present disclosure may comprise modified or non-naturally occurring internucleoside linkages. Such non-naturally internucleoside linkages may, in particular embodiments, confer desired properties to the oligonucleotide, non-limiting examples of which include, enhanced cellular uptake, enhanced affinity for other oligonucleotide or nucleic acid targets, or increased stability in the presence of nucleases. Types of non-naturally occurring internucleoside linkages are known in the art and any such linkage may be used according to the methods of the present disclosure. Non-limiting examples of such non- naturally occurring or modified internucleoside linkages include internucleoside linkages that23 1retain a phosphorus atom and internucleoside linkages that do not have a phosphorus atom. In certain embodiments, non-naturally occurring or modified internucleoside linkages that do not include a phosphorus atom may comprise intemucleoside linkages that are formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl, or cycloalkyl intemucleoside linkages, or one or more short chain heteroatomic or heterocyclic internucleoside linkages. These include those having amide backbones; and others, including those having mixed N, O, S and CH2 components.
[0050] In particular embodiments, oligonucleotides may also include oligonucleotide mimetics. An oligonucleotide mimetic may include, for example, an oligonucleotide wherein only the furanose ring or both the furanose ring and the intemucleotide linkage are replaced with a novel group. In one embodiment, the furanose ring may be replaced, for example, with a morpholino ring or any other sugar surrogate known in the art. In certain embodiments, an oligonucleotide mimetic may comprise one or more peptide nucleic acids or cyclohexenyl nucleic acids (Wang etal., I. Am. Chem. Soc. 122:8595-8602, 2000). In another embodiment, an oligonucleotide mimetic may be a phosphonomonoester nucleic acid, which incorporates a phosphorus group in the backbone. In yet another embodiment, an oligonucleotide mimetic may include replacement of the furanosyl ring with a cyclobutyl moiety.
[0051] In certain embodiments, oligonucleotides of the present disclosure may comprise one or more modified or substituted sugar moieties. Such modified or substituted sugars may, in some embodiments, improve stability in the presence of nucleases or binding affinity. Nonlimiting examples of modified or substituted sugars include carbocyclic or acyclic sugars, sugars having substitute groups at one or more of their 2', 3' or 4' positions, sugars having substitutes in place of one or more hydrogen atoms of the sugar, and sugars having a linkage between any two other atoms in the sugar. A large number of sugar modifications are known in the art and any such sugar modification may be used according to the present disclosure.
[0052] In some embodiments, oligonucleotides may include one or more nucleobase modifications or substitutions which are structurally distinguishable from, yet functionally interchangeable with, naturally occurring or synthetic unmodified nucleobases. Modified nucleobases may include in certain embodiments, synthetic or natural nucleobases such as 5- methylcytosine (5-me-C), 5 -hydroxymethyl cytosine, 7-deaza-guanine, or 7-deaza-adenine, 2-aminopyridine, 2-pyridone, 5-substituted pyrimidines, 6- azapyrimidines and N-2 substituted purines, N-6 substituted purines, O-6 substituted purines, 2 aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine.24 1
[0053] An oligonucleotide of the present disclosure may be, in certain embodiments, about 10 to about 1000, about 10 to about 900, about 10 to about 800, about 10 to about 700, about 10 to about 600, about 10 to about 500, about 10 to about 400, about 10 to about 350, about 10 to about 300, about 10 to about 250, about 10 to about 200, about 10 to about 150, about 10 to about 100, about 10 to about 50, about 10 to about 40, about 10 to about 30, or about 10 to about 20 nucleotides in length, including all ranges derivable therebetween.
[0054] The oligonucleotides of the disclosure, in certain embodiments, may comprise a barcode region, which can be used to identify a cellular characteristic. The barcode region can be a polynucleotide of at least, at most, about, or exactly 4,5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, or 200 nucleotides in length. The barcode may comprise, in some embodiments, one or more universal PCR regions, adaptors, such as adaptors for making cDNA libraries, linkers, or a combination thereof. The barcode region may also include, in particular embodiments, a molecular index region (MI) which can be used to count how many barcode sequences are delivered into each cell or nucleus. The MI may be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200 or more nucleotides in length, including all ranges derivable therebetween.
[0055] Non-limiting cellular characteristics that can be identified by the barcode region include sample identity, sub-sample identity, cell identity, nucleus identity, and whether the sample comprises RNA, cDNA, and / or DNA. In some embodiments, the barcode may be specific for RNA, cDNA, DNA, a cell, a nucleus, or a population of cells or nuclei, such that isolation of sequencing of the barcode after combining multiple differentially barcoded RNA molecules, cDNA molecules, DNA molecules, cells, samples, sub-samples, or nuclei identifies the cellular characteristic of the RNA molecules, cDNA molecules, DNA molecules, cells, samples, sub-samples, or nuclei. The cellular characteristic, in certain embodiments, can then be associated with other sequencing data or analysis. For example, the analysis may include epigenomic, genomic, or transcriptomic information obtained by single-cell analysis of mRNA, cDNA, or DNA. In one embodiment, the barcode is unique to one cell. In particular embodiments, the barcode is unique to a population of cells, such as about 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 500, 1,000, 5,000, 10,000, 25,000, 50,000, 100,000, 500,000, or 1,000,000 cells, including all ranges derivable therebetween.
[0056] In certain embodiments, a sample, cell, or nuclei comprising RNA or cDNA may be used to label the RNA or cDNA, or a portion thereof, and to perform reactions, such as PCR25 1and / or ligation, or to add the barcode oligonucleotides for distinguishing RNA or cDNA from DNA. The sample, cell, or nuclei suspension with barcode labeled RNA, cDNA, or DNA may then, in certain embodiments, be sorted by flow cytometry or dispensed by single cell microdispensers into a reaction chamber, tube, plate, well, nanowell, or nanochip, or embedded into microdroplets.D. Chromatin Disruption
[0057] Following or prior to labeling of the RNA or cDNA in a sample, in some embodiments, the chromatin structure of the single sample, sub-sample, nucleus, or cell in each reaction chamber may be disrupted to remove or partially remove chromatin from the DNA. In particular embodiments, the cell, cells, nuclei, or nucleus present in reach reaction may be fully or partially lysed to remove or partially remove chromatin or maintain intact chromatin from the DNA. Methods for cell and nuclei lysis are known in the art and any such method may be used according to the methods of the present disclosure. As non-limiting examples, cell or nuclei lysis may be performed using an enzyme-based method, for example using a protease, a chemical-based method, for example using detergents such as Tween-20 or Triton X-100, a mechanical-based method, an acoustic-based method, an electrical-based method, or a combination of any of the aforementioned methods. Chromatin disruption methods provided by the present disclosure digest the cellular and / or nuclear membrane and disrupts chromatin structure to expose the DNA present within the structure.E. Genomic DNA Labeling
[0058] Genomic DNA labeling may be performed, in certain embodiments, prior to, after, or approximately simultaneously with labeling of RNA or cDNA. In one embodiment, a lysis step may be performed to remove or partially remove chromatin prior to DNA labeling. In one embodiment, DNA may be labeled by tagmentation using a transposase, such as the Tn5 transposase. Any transposase known in the art may be used according to certain embodiments of the present disclosure. In particular embodiments, the transposase can be commercial and loaded such as the Tn5 transposome (Illumina) with universal oligonucleotides, or the transposome can be assembled in the laboratory by combining the transposase with transposase recognized DNA oligonucleotides (e.g., Mosaic End (ME) sequences). The oligonucleotides that attach to the transposase, in another embodiment, can have a barcode or part of a barcode sequence for distinguishing different cells. In yet another embodiment, the oligonucleotides may also serve as the identifier to distinguish DNA, RNA, and / or cDNA26 1molecules from the same single cell / nuclei or input material. Following the tagmentation reaction, in some embodiments, the transposase can be deactivated using protein-denaturing agents such as SDS or by applying heat in the presence of inhibitors such as EDTA. In some embodiments, the oligonucleotide barcode sequence can comprise unique molecular identifier (UMI) sequences to distinguish individual molecular reads from PCR duplicate reads.
[0059] In some embodiments, the DNA fragments obtained after disruption of chromatin structure can also be used to perform reactions, such as PCR or ligation, to add the oligonucleotide or barcode sequences in order to distinguish RNA or cDNA and DNA. In certain embodiments, oligonucleotides used to perform ligation reactions may be phosphorylated, for example, at their 5 ’ end.
[0060] In particular embodiments, oligonucleotide labels or barcodes may be introduced through a tagmentation reaction. In summary, methods of tagmentation may include the use of enzymes known as transposases, which randomly cut DNA into short segments, known as “tags.” Adapter nucleotide sequences are the added to either side of the cut points through ligation. The adaptor oligonucleotide molecules may comprise nucleotide barcodes and / or primer binding sites for detection and amplification of the DNA sequences. Methods of tagmentation are known in the art and are further described in Zahn, H., el al., Nature Methods 14:167, 2017, which is incorporated herein by reference. Any transposome known in the art may be used according to the methods provided herein. Non-limiting examples of such transposomes include a Tn5 transposome, a mutated Tn5 transposome, an antibody attached Tn5 transposome, an engineered Tn5 transposome, a Mu transposome, or a sleeping beauty transposome. In one embodiment, the transposome is an engineered or modified Tn5, such as pA-Tn5 used in the CUT&TAG method described in Kaya-Okur, el al., Nature Communications 10:1930, 2019. In particular embodiments, the transposome may he loaded with universal oligonucleotides, or the transposome can be assembled by combining the transposase with transposase recognized DNA oligonucleotides. Transposase recognized oligonucleotides may include, in certain embodiments, oligonucleotides that comprise mosaic end (ME) sequences. Mosaic end sequences are known in the art and any such mosaic end sequence may be used according to the methods of the present disclosure. A non-limiting example of a mosaic end sequence includes the sequence of SEQ ID NO:1 (AGATGTGTATAAGAGACAG). In particular embodiments, the oligonucleotides that attach to the transposase can comprise a nucleotide barcode sequence, or a portion thereof, for distinguishing individual cells. In certain embodiments, the attached oligonucleotide27 1sequence, or the portion thereof, can serve as the identifier for distinguishing the RNA / cDNA and DNA from the same single cell, nuclei, or sample. In some embodiments, the oligonucleotide barcode sequence can comprise unique molecular identifier (UMI) sequences to distinguish individual molecular reads from PCR duplicate reads.F. Amplification Reactions
[0061] The methods provided by the present disclosure utilize different adaptor or oligonucleotide labels to separately label RNA or cDNA, or a portion thereof, and DNA, or a portion thereof. Therefore, primer pairs for the amplifying the RNA or cDNA and primer pairs for amplifying the DNA may be added, in certain embodiments, to the same reaction during co-amplification. As used herein the term “primer” refers to a nucleic acid molecule that is designed for use in annealing or hybridization methods that involve an amplification reaction. An amplification reaction is an in vitro reaction that amplifies template RNA, cDNA, or DNA to produce an amplicon. As used herein, an “amplicon” is a polynucleotide molecule that has been synthesized using amplification techniques. A pair of primers may be used with template polynucleotide, such as a sample of RNA, cDNA, or DNA, in an amplification reaction, such as polymerase chain reaction (PCR), to produce an amplicon, where the amplicon produced would have a nucleotide sequence corresponding to sequence of the template polynucleotide located between the two sites where the primers hybridized to the template. A primer is typically designed to hybridize to a complementary target polynucleotide strand to form a hybrid between the primer and the target polynucleotide strand. The presence of a primer is a point of recognition by a polymerase to begin extension of the primer using as a template the target polynucleotide strand. Primer pairs refer to use of two primers binding opposite strands of a double stranded nucleotide segment for the purpose of amplifying the nucleotide segment between them. Tn certain embodiments, the primers may comprise an oligonucleotide barcode, or a portion thereof, to distinguish individual cells, nuclei, or samples. In particular embodiments, the primers may comprise a modification that may be used to separate RNA or cDNA from DNA following co-amplification. Oligonucleotide modifications that may be used for separating polynucleotide molecules are known in the art, and any such oligonucleotide modification may be used according to the methods of the present disclosure. In some embodiments, the methods of the present disclosure may include pre-amplification of either RNA, cDNA, or DNA prior to coamplification. For example, pre-amplification may be utilized, in particular embodiments, to enrich RNA or cDNA prior to co-amplification. Pre-amplification may, in some28 1embodiments, may be performed by adding the primer pair specific for either RNA, cDNA, or DNA, if either of these modalities would benefit from enrichment prior to co-amplification. In particular embodiments, amplification to enrich either RNA, cDNA, or DNA may be performed prior to (pre-amplification) or after (post-amplification) exponential coamplification. In certain embodiments, the annealing temperature used during coamplification may be favorable to RNA, cDNA, and / or DNA amplification to control the total number of molecules amplified. In certain embodiments, the amount of amplified RNA, cDNA, and / or DNA may be balanced by controlling the favored annealing temperature.
[0062] In certain embodiments of the present disclosure, the methods described herein use adaptor sequences introduced by the tagmentase to label DNA and adaptor sequences of probes to label RNA or cDNA. By mixing primers that are specific to these DNA and RNA or cDNA modality adaptors, respectively, the methods of the present disclosure can amplify DNA fragments and probe molecules at the same time, enabling DNA and RNA or cDNA coamplification. In certain embodiments, the primers may have a cell barcode or part of a cell barcode, which serves to provide a cell identity. In one embodiment, the primers can have modifications (such as biotin) that are used to separate the pool of RNA or cDNA from DNA following co-amplification. In another embodiment, the primer pairs for the DNA assay and the primer pairs for the RNA / cDNA assay may be added into the same reaction during exponential co-amplification. In some embodiments, the methods of the present disclosure can be used to enrich the number of molecules from one of the assays (e.g., RNA / cDNA assay). In certain embodiments, both of the forward and reverse primers may be added to the assay that would benefit from enrichment, while only one of the primers is added to the other assay. In other embodiment, one or two primers (forward or / and reverse) may be added to the modality that needs to be enriched, an no primers are added for the other modality. In some embodiments, the single-modality assay enrichment procedure can be performed before DNA and RNA / cDNA exponential co-amplification (pre- enrichment) or after DNA and RNA / cDNA exponential co-amplification (post-enrichment).G. Sample, Sub-Sample, or Individual Cell Labeling
[0063] In certain embodiments of the present disclosure, samples, sub-samples, or individual cells or nuclei may be labeled with a cell, nucleus, sub-sample, or sample oligonucleotide barcode prior to pooling the RNA or cDNA and DNA fragments from all cells, nuclei, samples, or sub-samples together. In some embodiments, the RNA or cDNA and DNA from the same sample, sub-sample, cell, or nucleus may be labeled with the same set of oligonucleotide29 1barcodes or a different set of oligonucleotide barcodes, for example, using prior knowledge of the RNA / cDNA and DNA barcode correspondence relationship. The RNA or cDNA and DNA, in some embodiments, may be labeled with a single sample, sub-sample, nucleus, or cell barcode or by multiple samples, sub-sample, or cell barcodes. In one embodiment, the RNA or cDNA and DNA may be labeled with a combination of sample, sub-sample, cell, or nucleus barcodes. The sample, sub-sample, cell, or nucleus barcode or barcodes may be added, for example, to one end of RNA or cDNA and DNA fragments or may be added to both ends of the RNA or cDNA and DNA fragments. In particular embodiments, sample, sub-sample, cell, or nucleus barcodes may be added to the 5’ end, the 3 ’end, or to both the 5’ end and 3’ end of the RNA or cDNA and DNA fragments.
[0064] In some embodiments, tagmentation-based chemistry may be used to fragment the DNA. The DNA fragments may then, in particular embodiments, be labeled or barcoded using a transposome and adaptors comprising different oligonucleotide sequences or barcodes. In certain embodiments, the oligonucleotide sequences or barcodes may be added by attaching different oligonucleotide sequences to the mosaic end sequences of the transposase, or the oligonucleotide sequences or barcodes may be added by using PCR primers that comprise the different oligonucleotide sequences, barcodes, or barcode combinations. Oligonucleotide sequences, barcodes, or barcode combinations, in some embodiments, may be added to DNA fragments by combining the approaches of tagmentation and PCR.
[0065] In certain embodiments, barcodes may be added to RNA or cDNA directly through probes or introduced by PCR using distinct PCR primers with varying barcodes or barcode combinations. In one embodiment, this may be achieved by merging the barcode introduced during the hybridization step with those introduced during PCR steps.
[0066] In certain embodiments, after the RNA / cDNA and DNA of the same cell or sample material are labeled with cell, sample, sub-sample, or nuclei barcodes, the libraries from all reactions are pooled together. From this pool of cells, samples, sub-samples, or nuclei, in certain embodiments, a physical separation of the RNA / cDNA and the DNA may be performed. The method of separation may be based, in some embodiments, on fragment sizes. The RNA / cDNA libraries may be separated based on fragment size, for example, when each library exhibits characteristic size profiles. The method of separation may be based, in one embodiment, on a modification used to label either the RNA / cDNA or the DNA, on the different oligonucleotide sequences or barcodes used to distinguish the RNA / cDNA and the DNA, or by other methods known in the art able to distinguish the RNA / cDNA and the DNA.30 1In another embodiment, separation may be based on distinct adaptors employed to distinguish between the DNA and RNA / cDNA. By using unique adaptors for each modality, the DNA and RNA / cDNA can be separated, for examples, during subsequent processing steps. In one embodiment, these subsequent processing steps may include PCR amplification using different PCR primers that specifically bind to DNA and / or RNA / cDNA. Alternatively, various other methods can be employed to differentiate amplified DNA and RNA molecules. These methods may rely on specific properties or markers inherent to each library type. In some embodiments, DNA and RNA / cDNA libraries from each individual cell or low-input sample may be labeled with different adaptors, which allows for differentiation of the DNA and RNA / cDNA libraries after pooling all of the libraries from single cells or sample materials. In one embodiment, this distinction facilitates the separation of the two distinct libraries post-pooling. In another embodiment, at least two of adaptors differ between the DNA and RNA / cDNA libraries to ensure that the process of preparing DNA and RNA / cDNA libraries separately can effectively occur after pooling. In yet another embodiment, one of the adaptors or primers used in the DNA or RNA / cDNA assay may have a specific modification, such as a biotin modification, to aid in the separation of the two assays following pooling.H. Construction of RNA / cDNA and DNA Libraries
[0067] In particular embodiments, following separation of RNA / cDNA and DNA, the amplified fragments may be used for high-throughput sequencing. In one embodiment, the PCR primers used to amplify the RNA / cDNA and DNA fragments comprise sequencing adaptors used for high-throughput sequencing. A person of ordinary skill in the art will recognize that the methods of the present disclosure can be customized and modified as needed to prepare different sequencing libraries in accordance with specific research purposes and the specifications of the chosen sequencing instrument. High-throughput sequencing platforms represent a wide range of technologies, including but not limited to next-generation sequencing, single molecule sequencing, long-read sequencing, and nanopore sequencing. Methods and primers for high-throughput sequencing are known in the art and any such methods or primers may be used according to the methods of the present disclosure. The separated RNA / cDNA or DNA amplification products can also be, in certain embodiments, further enriched using the RNA / cDNA or DNA-specific labels added during previous steps. These enriched amplification products, in some embodiments, may then be used as the input material for high-throughput sequencing.31 1
[0068] In certain embodiments, the amplified DNA can be used to identify a DNA sequence variation. As used herein the term “DNA sequence variation” may refer to any variation or alteration of a DNA nucleotide sequence. Non-limiting examples of types of DNA sequence variations include copy number variations or alterations (CNV / CNA), single nucleotide polymorphisms (SNPs), mutation of one or more nucleotides, insertions, and deletions (indels), short tandem repeats (STRs), translocations, inversions, structure variations (SVs), and any combinations thereof. The amplified DNA can also be used, in some embodiments, to profile targeted genes or gene panels, for probe-based target capture, for exome capture, or for other capture applications. In certain embodiments, the amplified DNA can also be used to investigate DNA rearrangements and markers, to detect frequency of mutations, or for other DNA-related applications. In certain embodiments, the amplified DNA can be used to identify a DNA modification. As used herein the term “DNA modification” may refer to an epigenetic modification, a mutation, an inversion, a duplication, or one or more nucleotide deletions, insertions, or substitutions. As used herein the term “epigenetic modification” refers to any genetic modification that alters gene activity without altering the DNA sequence. Non-limiting examples of epigenetic modifications include any modification that affects chromatin accessibility, histone modifications (for example histone acetylation, methylation, phosphorylation, ubiquitylation, sumoylation, deamination, and proline isomerization), DNA methylation, nucleosome positioning, loss of imprinting, chromatin structure modifications, and any combinations thereof. In certain embodiments, the DNA can be used to study DNA and protein interactions, and other DNA related applications.
[0069] As used herein the term “transcriptome variation” refers to any variation or alteration of the transcriptome. As used herein the term “transcriptome” refers to a set of RNA transcripts that may be present in a cell or sample. Identifying a transcriptome variation, in some embodiments, may comprise identifying a decreased expression level of a transcript, an increased expression level of a transcript, an increased expression level of a splice variant, a decreased expression level of a splice variant, an increased expression level of a transcript isoform, a decreased expression level of a transcript isoform, a modification of a transcript nucleotide sequence, a post-transcriptional modification, or a post-transcriptional rearrangement. The transcriptome may include coding and non-coding RNA transcripts. In some embodiments, the mRNA transcriptome may be analyzed by reverse transcribing the mRNA into cDNA prior to analysis. In particular embodiments, the amplified probe product (targeting RNA or cDNA), can represent various classes of RNA products. These classes32 1include but are not limited to mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, noncoding RNA, and poly-adenylated RNA. In one embodiment, the data obtained from the amplified probe product may be used to quantify gene expression, conduct differential gene expression analysis, perform allele-specific gene expression analysis, study gene regulatory networks, infer gene expression (RNA) trajectories and velocities, quantify noncoding RNA, and analyze noncoding RNA differential expression patterns. In some embodiments, the amplified DNA and RNA products may be used to address a wide spectrum of research needs, including spanning genomic alterations, gene expression profiling, epigenetics, and beyond, making them invaluable tools for various basic science and biomedical studies.I. Kits
[0070] In certain aspects, the present disclosure provides kits that may be used for performing the methods provided by the present disclosure. In some embodiments, such kits may comprise one or more of the following: first transposase, a probe comprising a first label, wherein the probe is capable of hybridizing to RNA or cDNA; an adaptor molecule for labeling genomic DNA with a second label, or reagents for performing an amplification reaction. In one embodiment, the kit may further comprise a reagent for disrupting chromatin structure, a first set of primers for amplifying the first labeled RNA or cDNA, a second set of primers for amplifying the second labeled DNA, dNTPs, a DNA polymerase, an RNA polymerase, a neutralization buffer, or a cell or nuclei lysis buffer. In one embodiment, the probe or the adaptor molecule further comprises a third label comprising an oligonucleotide barcode sequence. In another embodiment, the first set of primers or the second set of primers further comprises a third label comprising an oligonucleotide barcode sequence. In another embodiment, the kit may further comprise instructions for use of the kit.
[0071] The term "about" is used to indicate that a value includes the standard deviation of the mean for the device or method being employed to determine the value. The use of the term "or" in the claims is used to mean "and / or" unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive. When used in conjunction with the word "comprising" or other open language in the claims, the words "a" and "an" denote "one or more," unless specifically noted otherwise. The terms "comprise," "have," and "include" are open-ended linking verbs. Any forms or tenses of one or more of these verbs, such as "comprises," "comprising," "has," "having," "includes," and "including," are also open-ended. For example, any method that "comprises," "has," or "includes" one or more steps is not limited to possessing only those one or more steps and also covers other unlisted steps.33 1Similarly, any system or method that "comprises," "has," or "includes" one or more components is not limited to possessing only those components and covers other unlisted components.
[0072] Other objects, features, and advantages of the present disclosure are apparent from detailed description provided herein. It should be understood, however, that the detailed description and any specific examples provided, while indicating specific embodiments of the disclosure, are given by way of illustration only, since various changes and modifications within the spirit and scope of the disclosure will become apparent to those skilled in the art from this detailed description. Any embodiment of the present disclosure may be used in combination with any other embodiment described herein.
[0073] All references herein are incorporated herein by reference in their entirety.EXAMPLES
[0074] The following examples are included to illustrate embodiments of the present disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent techniques discovered by the inventor to function well in the practice of the disclosure. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the concept, spirit and scope of the disclosure. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the disclosure as defined by the appended claims.Example 1: Sample Preparation
[0075] Single cell suspensions or low input material is obtained by dissociating fresh samples, extracting nuclear suspensions from frozen samples, or dissociating formalin-fixed paraffin- embedded (FFPE) blocks. Freshly dissociated cells or nuclei may be fixed with 4% PFA or other fixatives. Next, fixed cells or nuclei may be hybridized by probes targeting the whole transcriptome overnight. Last, in-situ ligation may be performed by adding ligation mix containing IX SplintR ligase reaction buffer and SplintR ligase when two paired probes for one target are used.34 1Example 2: Analysis of Genome and Transcriptome
[0076] Lysis buffer mix (4 l) containing 0.5X PBS, 2.25% Tween-20, 0.225% TritonX-100, 13.5 mM Tris-HCL, pH 8.0, and 0.068 mAU / pl protease (Qiagen) was added into each tube or well. Fixed probes (e.g., probes from single cell gene expression FLEX reagents from 10X genomics) were hybridized, and ligated single cells were sorted individually into each tube / well using Melody (BD Bioscience) (1 cell / tube or 1 cell / well). Lysis was carried out at 55 °C for 20 min and protease was inactivated at 70 °C for 15 min. Next, 2 pl of tagmentation mix containing 1.95X TD buffer (Illumina), and 0.05 pl TDE1 (Illumina) was added to each tube / well. A tagmentation reaction was carried out at 55 °C for 8 min. Afterwards, 2 pl of neutralization mix containing 25 mM EDTA (Invitrogen), 2 mM dNTPs (Roche), 2X KAPA HiFi Fidelity Buffer (Roche), 10 pM RNA_S5XX primers (AAGCAGTGGTATCAACGCAGAGTACNNNNNNNNTTGCTAGGACCG (SEQ ID NO:2), and 5 pM DNAJS5XX (AATGATACGGCGACCACCGAGATCT ACACNNNNNNNNTCGTCGGCAGCGTC) (SEQ ID NO:3) primers were added to each tube / well. Neutralization was carried out at 50 °C for 30 min. Lastly, 4 pl of PCR mix containing 2X KAPA HiFi Fidelity Buffer (Roche), 8.75 mM MgC12, 5 pM of RNA_N7XX primers (CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCCTT GGCACCCGAGAATTCCA) (SEQ ID NO:4), 2.5 pM of DNA_N7XX_IN primers (CTGAGTCGGAGACACGCA-NNNNNNNN-GTCTCGTGGGCTCGG) (SEQ ID NO:5), and 0.08 U / pl KAPA HiFi HotStart Polymerase (Roche) was added into each well / tube. Coamplification PCR of DNA and probes (targeting RNA) was cycled as follows: 72 °C for 3 min, 98 °C for 3 min, 18 cycles of 98 °C for 15 s, 53 °C for 20 s, 72°C for 60 s, then 72 °C for 1 min.Example 3: High Throughput Analysis of Genome and Transcriptome
[0077] Single cell suspensions from fixed cells, nuclei, or dissociated FFPE samples were first hybridized with RNA-targeting probes and were further diluted to 32,000 cells / ml with 0.5X PBS containing DAPI. Next, the diluted cell suspensions were dispensed into a 350 nl nanowell chip (Takara) using the ICELL8 CX system (Takara). The nanowell chip was scanned with DAPI signal and only nanowells containing single cells were selected for downstream experiments. Next, 35 nl lysis buffer mix containing 4.5% Tween-20, 0.45% TritonX-100, 17 mM Tris-HCL, pH 8.0, and 0.136 mAU / pl protease (Qiagen) was added into each well. Lysis was carried out at 55 °C for 30 min and protease was inactivated at 75 °Cfor 15 min. Next, 35 nl tagmentation mix containing 1.8X TD buffer (Illumina), and 3.5 nl TDE1 (Illumina) was added to each well. Tagmentation reaction was carried out at 55 °C for 8 min. Afterwards, 35 nl of neutralization mix containing 22.4 mM EDTA (Invitrogen), 2 mM dNTPs, 2.24X KAPA HiFi Fidelity Buffer (Roche), 10 pM RNA_S5XX primers and 5 pM DNA_S5XX primers was added to each well. Neutralization reaction was carried out at 50 °C for 30 min. Next, 35 pl DNA / RNA primer mix containing 2.24X KAPA HiFi Fidelity Buffer (Roche), 15.6 mM MgC12, 10 pM RNA_N7XX primers and 5 pM DNA_N7XX_1N primers was directly added into each well. 35 pl PCR mix containing 1.95X KAPA HiFi Fidelity Buffer and 0.2U / pl KAPA HiFi HotStart Polymerase (Roche) was added afterwards. Lastly, DNA / RNA co-amplification PCR was cycled as follows: 72 °C for 8 min, 98 °C for 3 min, 12 cycles of 98 °C for 15 s, 55 °C for 30 s, 72°C for 60 s. Final elongation was performed for 2 min at 72 °C.Example 4: DNA and Probes (Targeting RNA) Library Preparation
[0078] After the DNA / probes co-amplification was completed, the amplified products were pooled into one tube. 100 pl of the pooled sample was first purified twice by 1.8X Ampure beads (Beckman). Next, 100 ng of the purified product was PCR amplified by adding DNA PCR mix containing 3 pM DNA_N7XX_0UT primer (CAAGCAGAAGACGGCATACGAGATNNNNNNNNCTGAGTCGGAGACACGCA) (SEQ ID NO: 6), 3 pM Bioo_F primers (AATGATACGGCGACCACCGAGATCTACAC) (SEQ ID NO:7) and IX KAPA HiFi HotStart ReadyMix (Roche) for DNA library construction. The DNA-enrichment PCR was cycled as follow: 98 °C for 30 s, 5 cycles of 98 °C for 10 s, 63 °C for 30 s, 72°C for 30. Final elongation was performed for 2 min at 72 °C. Simultaneously, RNA-enrichment PCR used 200 ng of purified product by adding RNA PCR mix containing 3 pM DDR_PCR_P5_UDIXXX_v2 primers (AATGATACGGCGACCACC GAGATCTACACNNNNNNNNGCCTGTCCGCGGAAGCAGTGGTATCAACGCAGAGT AC) (SEQ ID NO:8), 3 pM Bio_R primers (CAAGCAGAAGACGGCATACGAGAT) (SEQ ID NO:9), and IX KAPA HiFi HotStart ReadyMix for RNA library construction. The RNA- enrichment PCR was cycled as follow: 95 °C for 30 s, 8 cycles of 98 °C for 10 s, 67 °C for 30 s, 72°C for 30 s. Final elongation was performed for 2 min at 72 °C. Lastly, both enriched DNA PCR products and RNA PCR products were further purified by 1.8X Ampure beads and both libraries were ready for sequencing afterwards.36 1Example 5: Ligation Condition Tests
[0079] Single tube test experiments were performed to optimize ligation conditions using a formalin fixed single cell suspension (from normal breast tissue cells) that hybridized with the RNA-targeting probes. The single cell suspension was incubated with IX SplintR ligase buffer, SplintR ligase. Three temperature conditions (16°C, 25°C, 30°C), three incubation times (Ih, 2h, overnight), and two different enzyme amounts were tested for the ligation reaction of 50K single cells. Following incubation, the reactions were halted using EDTA, and samples were subsequently lysed with protease. This was followed by PCR amplification using probe-specific primers. The results indicated that overnight incubation yielded the highest concentration of PCR products compared to 1-hour and 2-hour incubations. Additionally, the use of 5 units of SplintR ligase resulted in a higher product concentration compared to 2.5 units (FIG. 7, Panel A). Fluorescence activated cell sorting (FACS) was then performed to sort the single cells from 1 h incubation (5U ligase) and overnight incubation (5U ligase) samples into a 96 well plate. Three and four single cells from the 1 h and overnight conditions, respectively, were selected to measure the gene expression (probe, RNA- targeting) and DNA copy number profiles. For the gene expression profiles of the single cells, it was found that the overnight incubation condition detected many more genes (median 2048 genes) per cell compared to the 1 h ligation incubation condition (median 449 genes per cell) (FIG. 7, Panel B). .15 M reads per cell were sequenced on average with a mean PCR duplicates rate of 19.00% for the DNA copy number profiles from the single cells from the 1 h incubation condition. 134 reads per bin at 220 kb genomic resolution was obtained. 2.01 M reads per cell were sequenced on average with a mean PCR duplicates rate of 19.13% for the DNA copy number profiles from the single cells of the overnight incubation condition. 126 reads per bin at 22 Okb genomic resolution was obtained. The log2 copy number ratio of each single cell was calculated (FIG. 7), and the results demonstrate that under both conditions, the method was able to measure flat diploid copy number profiles in 6 single cells, which was as expected since all cells were from normal breast tissue cells with diploid genomes (FIG. 8).Example 6: High Throughput Analysis of Genome and Transcriptome of Normal Human Breast Tissue
[0080] Simultaneous profiling of genome and transcriptome was generally performed as described in Example 3. Briefly, single cell DNA and probe (RNA-targeting) libraries wereprepared using a fresh viable cell suspension dissociated from a human breast tissue sample. The cell suspension was first fixed and hybridized with DNA probes that target the whole RNA transcriptome (~20K targets), then ligation followed was performed by loading into the 5184-well nanowell chip (ICELL8). 1,971 wells with single cells were selected to perform the remainder of the protocol. For the scRNA data, there were 1,760 cells that passed initial RNA QC with a median read of 98,795 and median gene number of 3,024 per cell. Gene expression clustering identified 11 clusters representing luminal secretory epithelial cells (LumSec), two different fibroblast clusters, T cells, luminal hormone response epithelial cells (LumHR), two endothelial populations, macrophages, pericytes, myoepithelial cells, and plasma B cells according to the canonical markers the clusters expressed (FIG. 9). The top 5 differentially expressed genes in each cluster are shown in Table 1.Table 1: Differentially Expressed Genes Following High Throughput Analysis of Genome and Transcriptome of Normal Human Breast Tissue.
[0081] For the DNA data, 1,764 of the single cells passed QC and on average 419k reads per cell with 17.21% PCR duplicate rates were sequenced. 26 reads (median) in each bin at a 220K variable bin genomic resolution was obtained. As expected, the DNA data did not detect38 1copy number changes since this sample was from a normal breast tissue that should harbor a diploid genome. Mapping the cells based on the RNA clustering results, it was found that all cells were uniformly mixed in the clustered DNA copy number heatmaps. These data demonstrate the method can accurately measure single cell DNA copy number profiles and link the scRNA expression programs to identify different cell types in human tissue samples.Example 7: High Throughput Analysis of Genome and Transcriptome of Single Cell Suspensions Disassociated from Frozen Human Breast Cancer Tissue
[0082] The methods of the present disclosure produce excellent results not only for fresh fixed whole cell suspensions, but also for single nuclei suspensions extracted from frozen samples. This allows for the analysis of samples that may have been stored for years or decades. Single cell DNA and probe (targeting RNA) libraries were prepared using nuclei suspensions dissociated from a frozen human breast ductal carcinoma in situ cancer tissue. After nuclei fixation, hybridization, and overnight ligation, the sample was loaded into the 5184-well nanowell chip (TCELL8). 2, 112 wells, each comprising a single cell were selected to perform the remainder of the protocol. In the scRNA data, 1,712 cells passed initial RNA QC with a median read of 57,752 and median gene number of 4,869 per cell. Gene expression clustering identified 10 clusters representing fibroblast, endothelial, myoepithelial, macrophage, LumSec, T, and 4 tumor cell clusters according to the markers those clusters expressed (FIG. 10, Panel A and Panel B). The top 5 differentially expressed genes in each cluster are shown in Table 2.Table 2: Differentially Expressed Genes Following High Throughput Analysis of Genome and Transcriptome of Human Breast Cancer Tissue.
[0083] For the DNA data, 1,674 of the single cells passed QC. On average, 319k reads per cell were sequenced with 16.94% PCR duplicate rates. 20 reads (median) were detected in each bin at 220K variable bin resolution. After filtering out low quality cells and potential doublets, 1,065 single cells were identified to have both DNA and RNA data. The single cell DNA copy number profiles were clustered of those matched cells, which identified 8 different subclones. Next, the DNA results and RNA results were matched from the same cell, and it was found that the normal cell types including fibroblast, endothelial, myoepithelial, LumSec and T cells were matched to the diploid cluster (c3). The two major tumor RNA clusters matched with subclone c5 and c6 (FIG. 10, Panel C). The RNA only data, however, did not distinguish some subclones, such as c4, c7 and c8, which could only be identified by using high resolution DNA data. These data further demonstrate that the methods of the present disclosure can not only resolve detailed DNA copy number structures of single cell genomes but can also resolve different tumor and normal cell types based on the RNA profiles from the same cell.Example 8: High Throughput Analysis of Genome and Transcriptome of Single Cell Suspensions Disassociated from Formalin-Fixed Paraffin Embedded Ductal Carcinoma in Situ (DCIS) Sample
[0084] In contrast to the fresh cell suspensions and frozen nuclei suspensions, the RNA in FFPE tissues is often highly degraded. To demonstrate that the methods of the present disclosure can be performed using samples with degraded RNA, the genome and transcriptome was approximately simultaneously profiled from single cells dissociated from an FFPE block that was stored at room temperature for almost 1 year. The single cell suspension from FFPE tissue was first disassociated using either a commercialized kit (e.g., FFPE Tissue Dissociation Kit from Miltenyi Biotec) or a protease-based method. The single cell suspensions were then used to prepare the single cell DNA and probe (targeting RNA) libraries. In total 1,785 wells with single cell were selected to perform the remainder of the protocol. In the scRNA data, there were 1,447 (81%) cells that passed initial RNA QC with a median read count of 57,231 and median gene number of 2,540 per cell. Gene expression clustering identified 10 clusters representing cancer cells, T cells, myoepithelial cells, macrophages, fibroblasts, LumSec cells, endothelial cells, B cells, plasma cells, and mitotic cells according to the top marker genes those clusters expressed (FIG. 11, Panel A and Panel B). The top 5 differentially expressed genes in each cluster are shown in Table 3.Table 3: Differentially Expressed Genes Following High Throughput Analysis of Genome and Transcriptome of Ductal Carcinoma in Situ (DCIS).
[0085] From the DNA data, 1,284 (72%) of the single cells passed QC. On average, 256k reads per cell were sequenced with 9.93% PCR duplicate rates. 17 reads (median) in each bin were detected at 220K variable bin genomic resolution. After filtering out low quality cells and potential doublets, there were 1,065 single cells with both DNA and RNA data. The single cell DNA copy number profiles of the matched cells were then clustered, which identified 4 different superclones and 11 different subclones. Next, the DNA results and RNA results from the same cell were matched, and it was found that the normal cell types including T cells, myoepithelial cells, macrophages, fibroblasts, LumSec cells, endothelial cells, B cells, and plasma cells mapped to the diploid cluster (c4). The one cancer cell RNA cluster matched with subclone c0-c3, and c6-cl0 (FIG. 11, Panel C). These data further demonstrate that the methods of the present disclosure can profile DNA and RNA of the same single cell with degraded DNA and RNA, such as from FFPE samples.* *
[0086] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of this invention have been described in terms of preferred embodiments or aspects, it will be apparent to those of skill in the art that variations may be applied to the methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit, and scope of the invention. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.42 1
Claims
CLAIMS1. A method of producing a DNA library and a RNA or cDNA library, the method comprising: a) hybridizing a probe comprising a first label to RNA or to cDNA in a sample; b) performing a labeling reaction to label genomic DNA in the sample with a second label; and c) performing an amplification reaction sufficient to amplify said first labeled RNA or cDNA and said second labeled genomic DNA approximately simultaneously.
2. The method of claim 1, wherein said sample comprises RNA or cDNA and DNA obtained from about 1 cell to about 1,000,000,000 cells.
3. The method of claim 2, wherein said sample comprises RNA or cDNA and DNA obtained from a single cell.
4. The method of claim 1, wherein the probe is a linear probe, a circular probe, or a padlock probe.
5. The method of claim 1, the method further comprising disrupting the genomic DNA chromatin structure prior to said hybridizing or prior to performing said labeling reaction to label the genomic DNA.
6. The method of claim 1, wherein the first labeled RNA or cDNA and the second labeled genomic DNA are not physically separated prior to performing said amplification reaction.
7. The method of claim 1, wherein the labeling reaction is performed by an insertional enzyme complex.
8. The method of claim 7, wherein the insertional enzyme complex comprises a transposase.
9. The method of claim 1, further comprising detecting at least one DNA sequence variation, DNA modification, or transcriptional variation.
10. The method of claim 1, the method further comprising performing a second labeling reaction to label said RNA or cDNA and said genomic DNA with a third label.
11. The method of claim 10, wherein: a) the second labeling reaction and the amplification reaction are performed approximately simultaneously; or b) the second labeling reaction is performed prior to, after, or approximately simultaneously with hybridizing the probe comprising the first label to RNA or to cDNA in the sample.
12. The method of claim 1 , wherein: a) the hybridizing and the labeling reaction are performed approximately simultaneously; or b) the hybridizing is performed before the labeling reaction.
13. The method of claim 1, wherein: a) the first label or the second label is a nucleotide barcode label; or b) the first label comprises a first nucleotide barcode label, the second label comprises a second nucleotide barcode label, and the first nucleotide barcode label and the second nucleotide barcode label are different.
14. The method of claim 10, wherein the third label comprises a third nucleotide barcode label.
15. The method of claim 1, wherein the sample comprises a plurality of cells or a plurality of nuclei, and the method further comprises separating the plurality of cells or plurality of nuclei into a plurality of compartments, wherein each compartment comprises about one cell.
16. The method of claim 15, wherein said separating is performed prior to hybridizing the probe comprising the first label to the RNA or cDNA.
17. The method of claim 15, wherein the separating is performed prior to performing the labeling reaction to label the genomic DNA in the sample with the second label.
18. The method of claim 1 , wherein the first labeled RNA or cDNA and the second labeled genomic DNA are not physically separated prior to performing the amplification reaction.4419. The method of claim 1, the method further comprising performing a sequencing reaction to obtain a sequence of the first labeled RNA or cDNA or the second labeled genomic DNA.
20. The method of claim 15, the method further comprising combining the first labeled RNA or cDNA and the second labeled genomic DNA from each compartment of the plurality of compartments, wherein said combining is performed after performing the amplification reaction.
21. The method of claim 1, the method further comprising performing a second amplification reaction to further amplify the first labeled RNA or cDNA or the second labeled genomic DNA.
22. The method of claim 1, wherein the RNA is selected from the group consisting of mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, and polyadenylated RNA.
23. The method of claim 1, the method further comprising performing DNA-seq analysis to detect at least one DNA sequence variation or at least one DNA modification.
24. The method of claim 1, the method further comprising performing RNA-seq analysis to detect at least one transcriptome modification.
25. An RNA, a cDNA, or a DNA library produced by the method of claim 1.
26. The method of claim 1 , the method further comprising contacting the sample with an oligonucleotide conjugated antibody.
27. The method of claim 26, wherein said contacting is performed prior to, after, or approximately simultaneously with the step of hybridizing the probe comprising the first label to RNA or to cDNA in the sample.
28. The method of claim 1, the method further comprising contacting the sample with a sample barcode or a spatial barcode.
29. The method of claim 28, wherein said contacting is performed prior to, after, or approximately simultaneously with the step of hybridizing the probe comprising the first label to RNA or to cDNA in the sample.45 130. The method of claim 1, further comprising disrupting the genomic DNA chromatin structure prior to performing said step (c).
31. The method of claim 30, wherein the labeling reaction labels at least a first portion of open chromatin region DNA comprised within the genomic DNA with the second label.
32. The method of claim 30, further comprising performing a second labeling reaction on said genomic DNA to label at least a first portion of exposed closed chromatin region DNA comprised within the genomic DNA with a third label.
33. The method of claim 31, further comprising detecting at least one DNA sequence variation or DNA modification in the second labeled open chromatin region DNA or the third labeled exposed closed chromatin region DNA.
34. The method of claim 31 , further comprising performing a Cleavage Under Targets and Tagmentation (CUT&TAG) analysis of the open chromatin region DNA to measure proteimDNA interactions of the genomic DNA.
35. The method of claim 31, further comprising performing an Assay for Transposase- Accessible Chromatin (ATAC) analysis on the open chromatin region DNA to measure chromatin accessibility differences in the genomic DNA.
36. The method of claim 32, wherein the first labeling reaction or the second labeling reaction is performed by an insertional enzyme complex.
37. The method of claim 36, wherein the insertional enzyme complex comprises a transposase.
38. The method of claim 36, wherein the first labeling reaction is performed by an antibody attached insertional enzyme complex.
39. The method of claim 31, further comprising detecting at least one DNA sequence variation, DNA modification, or transcriptome variation.
40. The method of claim 1, further comprising digesting the sample with a restriction enzyme to fragment genomic DNA prior to hybridizing the probe comprising the first label to RNA or to cDNA in the sample.46 141. The method of claim 40, further comprising contacting the fragmented DNA with aDNA ligase.
42. The method of claim 41, wherein the labeling reaction labels ligated fragments of genomic DNA.
43. The method of claim 40, further comprising performing a second labeling reaction on said genomic DNA to label at least a first portion of non-ligated DNA comprised within the genomic DNA with a third label.
44. The method of claim 41, further comprising performing a high-throughput chromosome conformation capture technique analysis to detect DNA:DNA interactions of the genomic DNA.
45. The method of claim 43, wherein the first labeling reaction or the second labeling reaction is performed by an insertional enzyme complex.
46. The method of claim 45, wherein the insertional enzyme complex comprises a transposase.
47. A method of identifying a DNA sequence variation or a transcriptome variation in a sample, the method comprising: a) hybridizing a probe comprising a first label to RNA or cDNA in a sample; b) performing a labeling reaction to label genomic DNA in the sample with a second label; c) performing an amplification reaction sufficient to amplify said first labeled RNA or cDNA and said second labeled genomic DNA approximately simultaneously; d) performing a sequencing reaction to obtain a sequence of the first labeled RNA or cDNA or the second labeled genomic DNA; and e) identifying said DNA sequence variation or said transcriptome variation in said sample.
48. The method of claim 47, wherein said sample comprises RNA or cDNA and DNA obtained from about 1 cell to about 1,000,000,000 cells.
49. The method of claim 48, wherein said sample comprises RNA or cDNA and DNA obtained from a single cell.47 150. The method of claim 47, wherein the first labeled RNA or cDNA and the second labeled genomic DNA are not physically separated prior to performing said amplification reaction.
51. The method of claim 47, wherein the labeling reaction is performed by an insertional enzyme complex.
52. The method of claim 51, wherein the insertional enzyme complex comprises a transposase.
53. The method of claim 47, the method further comprising performing a second labeling reaction to label said RNA or cDNA and said genomic DNA with a third label.
54. The method of claim 53, wherein the second labeling reaction and the amplification reaction are performed approximately simultaneously.
55. The method of claim 47, wherein: a) the hybridizing and the labeling reaction are performed approximately simultaneously; or b) the hybridizing is performed before the labeling reaction.
56. The method of claim 47, wherein: a) the first label or the second label is a nucleotide barcode label; or b) the first label comprises a first nucleotide barcode label, the second label comprises a second nucleotide barcode label, and the first nucleotide barcode label and the second nucleotide barcode label are different.
57. The method of claim 53, wherein the third label comprises a third nucleotide barcode label.
58. The method of claim 47, wherein the sample comprises a plurality of cells or a plurality of nuclei, and the method further comprises separating the plurality of cells or plurality of nuclei into a plurality of compartments, wherein each compartment comprises about one cell.
59. The method of claim 58, wherein said separating is performed prior to hybridizing the probe comprising the first label to the RNA or cDNA.
60. The method of claim 58, the method further comprising combining the first labeled RNA or cDNA and the second labeled genomic DNA from each compartment of the plurality48 1of compartments, wherein said combining is performed after performing the amplification reaction.
61. The method of claim 47, the method further comprising performing a second amplification reaction to further amplify the first labeled RNA or cDNA or the second labeled genomic DNA.
62. The method of claim 47, wherein the RNA is selected from the group consisting of mRNA, pre-mRNA, miRNA, siRNA, piRNA, lincRNA, non-coding RNA, and polyadenylated RNA.
63. The method of claim 47, the method further comprising disrupting the genomic DNA chromatin structure prior to said hybridizing or prior to performing said labeling reaction to label the genomic DNA.
64. The method of claim 47, the method comprising identifying said DNA sequence variation in the amplified second labeled genomic DNA.
65. The method of claim 64, the DNA sequence variation is selected from the group consisting of a mutation, a copy number alteration, a single nucleotide polymorphism, an insertion, a deletion, a short tandem repeat, a translocation, an inversion, a structural variation, and a DNA modification.
66. The method of claim 47, wherein identifying said transcriptome variation comprises identifying a decreased expression level of a transcript, an increased expression level of a transcript, an increased expression level of a splice variant, a decreased expression level of a splice variant, an increased expression level of a transcript isoform, a decreased expression level of a transcript isoform, a modification of a transcript nucleotide sequence, a post- transcriptional modification, or a post-transcriptional rearrangement.
67. The method of claim 47, wherein said DNA sequence variation or transcriptome variation is associated with a condition selected from the group consisting of cancer, a genetic disease or condition, a developmental disease or condition, or an immunological disease or condition.
68. A kit comprising: a) a transposase;49 1b) a probe comprising a first label, wherein the probe is capable of hybridizing to RNA or cDNA; c) an adaptor molecule for labeling genomic DNA with a second label; and d) reagents for performing an amplification reaction.
69. The kit of claim 68, wherein said kit further comprises a reagent for disrupting chromatin structure.
70. The kit of claim 68, wherein said kit further comprises a first set of primers for amplifying the first labeled RNA or cDNA and a second set of primers for amplifying the second labeled genomic DNA.
71. The kit of claim 68, wherein the probe or the adaptor molecule further comprises a third label comprising an oligonucleotide barcode sequence.
72. The kit of claim 70, wherein the first set of primers or the second set of primers further comprises a third label comprising an oligonucleotide barcode sequence.50 1
Citation Information
Patent Citations
Single cell whole genome libraries and combinatorial indexing methods of making thereof
WO2018018008A1
Methods for simultaneous amplification of DNA and RNA
WO2024059622A2