Single cell analysis
Primary template-directed amplification (PTA) methods address the limitations of current nucleic acid amplification and sequencing by enhancing sequence representation and accuracy, enabling sensitive and scalable multi-omic analysis of single cells.
Patent Information
- Application Number
- JP2025090040
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-07-31
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-01
AI Technical Summary
Current methods for nucleic acid amplification and sequencing from single cells are limited by reproducibility, sequence representation, uniformity, and accuracy, particularly in the analysis of RNA, DNA, and protein, necessitating improved scalable and efficient techniques for multi-omic analysis.
The implementation of primary template-directed amplification (PTA) methods that utilize terminator nucleotides for strand displacement replication, followed by adapter ligation and sequencing of genomic and cDNA libraries, enabling accurate and scalable amplification of nucleic acids from single cells.
PTA methods enhance the accuracy and sensitivity of sequencing, improving the detection of single-base variants, copy number diversity, and structural diversity, and facilitate applications such as next-generation sequencing, environmental mutagenicity assessment, and cancer treatment response prediction.
Smart Images

Figure 2025143255000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 881,183, filed July 31, 2019, which is incorporated herein by reference in its entirety. [Background technology]
[0002] background Research methods that utilize nucleic acid amplification, such as next-generation sequencing, provide a large amount of information about complex samples, genomes, and other nucleic acid sources. In some cases, these samples are obtained in small amounts from single cells. There is a need for highly accurate, scalable, and efficient nucleic acid amplification and sequencing methods for research, diagnosis, and treatment involving small amounts of samples, especially methods for simultaneous analysis of RNA, DNA, and protein. Summary of the Invention
[0003] overview Provided herein is a method for multi-omic single-cell analysis, the method comprising: (a) isolating a single cell from a population of cells; (b) sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from the single cell; and (c) sequencing the genome of the single cell, wherein sequencing the genome comprises: (i) contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; and (ii) amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iii) ligating the molecules obtained in step (ii) to adapters, thereby generating a genomic DNA library; and (iv) sequencing the genomic DNA library. Further provided herein is a method wherein the mRNA transcripts comprise polyadenylated mRNA transcripts. Further provided herein is a method wherein the mRNA transcripts do not comprise polyadenylated mRNA transcripts. Further provided herein is a method wherein sequencing the cDNA library comprises amplifying the mRNA transcripts using template switching primers. Further provided herein is a method wherein at least some of the polynucleotides in the cDNA library comprise barcodes. Further provided herein is a method wherein the barcodes comprise cell barcodes or sample barcodes. Further provided herein is a method wherein the cDNA library and the genomic DNA library are pooled prior to sequencing. Further provided herein is a method wherein the single cell is a primary cell. Further provided herein is a method wherein the single cell is derived from liver, skin, kidney, blood, or lung. Further provided herein is a method wherein the single cell is isolated by flow cytometry.Further provided herein is a method, wherein the method further comprises removing at least one terminator nucleotide from the terminated amplification products. Further provided herein is a method, wherein the plurality of terminated amplification products comprises an average length of 1,000 to 2,000 bases. Further provided herein is a method, wherein the plurality of terminated amplification products comprises a length of 250 to 1,500 bases. Further provided herein is a method, wherein the plurality of terminated amplification products comprises at least 97% of the genome of the single cell. Further provided herein is a method, wherein at least some of the amplification products comprise a cell barcode or a sample barcode. Further provided herein is a method in which sequencing the cDNA library comprises cytoplasmic lysis of a single cell and reverse transcription. Further provided herein is a method in which the mRNA transcripts are amplified via template-switching reverse transcription. Further provided herein is a method in which the cDNA library comprises at least 10,000 genes. Further provided herein is a method in which sequencing the genome of the single cell further comprises nuclear lysis of the single cell. Further provided herein is a method in which the method further comprises an additional amplification step using PCR. Further provided herein is a method in which at least one mutation is identified in the genome of a cell, the mutation differing from the corresponding position in a reference sequence. Further provided herein is a method in which the at least one mutation occurs in less than 1% of a population of cells. Further provided herein is a method in which the at least one mutation occurs in 0.1% or less of a population of cells. Further provided herein is a method in which the at least one mutation occurs in 0.001% or less of a population of cells. Further provided herein is a method wherein the at least one mutation occurs in 1% or less of the amplification product sequences. Further provided herein is a method wherein the at least one mutation occurs in 0.1% or less of the amplification product sequences. Further provided herein is a method wherein the at least one mutation occurs in 0.001% or less of the amplification product sequences.
[0004] Provided herein is a method for multi-omic single-cell analysis, the method comprising: (a) isolating a single cell from a population of cells; (b) identifying at least one protein on the surface of the single cell; and (c) sequencing the genome of the single cell, wherein sequencing the genome comprises: (i) contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides includes at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; (ii) amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iii) ligating the molecules obtained in step (ii) to adapters, thereby generating a genomic DNA library; and (iv) sequencing the genomic DNA library. Further provided herein is a method, wherein the step of identifying at least one protein on the surface of the cell comprises contacting the cell with a labeled antibody that binds to at least one protein. Further provided herein is a method, wherein the labeled antibody comprises at least one fluorescent label or mass tag. Further provided herein is a method, wherein the labeled antibody comprises at least one nucleic acid barcode.
[0005] Provided herein is a method for multi-omic single-cell analysis, the method comprising: (a) isolating a single cell from a population of cells; (b) sequencing the genome of the single cell, wherein sequencing the genome comprises: (i) digesting the genome with a methylation-sensitive restriction enzyme to generate genomic fragments; and (ii) contacting at least some of the genomic fragments with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides contains at least one terminator nucleotide that terminates nucleic acid replication by the polymerase. (iii) contacting at least some of the genome fragments with a primer containing an adapter; (iii) amplifying at least some of the genome fragments to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iv) amplifying at least some of the genome fragments with methylation-specific PCR; (v) ligating the molecules obtained in steps (iii and iv) to adapters, thereby generating a genomic DNA library and a methylome DNA library; and (vi) sequencing the genomic DNA library and the methylome DNA library.
[0006] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]
[0007] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Figure 1A] Illustrates a summary of the general workflow for isolation and analysis of protein, DNA, and RNA from single cells. [Figure 1B] We illustrate a workflow for isolating and analyzing protein, DNA, and RNA from single cells using sample splitting to minimize cross-contamination. [Figure 1C] Illustrates a workflow for isolation and analysis of protein, DNA, and RNA from single cells using single-tube pre-amplification. [Figure 1D] We illustrate a workflow for the isolation and analysis of protein, DNA, and RNA from single cells using single-tube pre-amplification with terminators to reduce the size of the amplicon. [Figure 1E] 1 illustrates a workflow for the isolation and analysis of protein, DNA, and RNA from single cells using co-amplification. [Figure 1F] 1 illustrates an informatics workflow combining data from the protein / DNA / RNA single-cell experiments described herein. [Figure 1G] 1 illustrates a comparison of the MDA and PTA irreversible terminator methods in relation to mutation propagation. The PTA method produces an increased number of direct copies of the original DNA template. [Figure 2A] The method steps performed after amplification are illustrated, including terminator removal, end repair, and A-tailing prior to adapter ligation. The pooled cell library can then be enriched for all exons or other specific regions of interest via hybridization prior to sequencing. The cell of origin of each read is identified by its cell barcode (shown as a green and blue sequence). [Figure 2B] (GC) Comparison of the GC content of sequenced bases for MDA and PTA experiments. [Figure 2C] The map quality score (e) (mapQ) for single cells mapping to the human genome (p_mapped) after undergoing PTA or MDA is shown. [Figure 2D]The percentage of reads that map to the human genome (p_mapped) after single cells underwent PTA or MDA is shown. [Figure 2E] (PCR) Comparison of the percent of reads that are PCR replicates for 20 million subsampled reads after single cells underwent MDA and PTA. [Figure 2F] Shown is a workflow for RT amplification of single cells for use with PTA. [Figure 2G] Construction of a library from cDNA obtained by RT is shown. [Figure 3A] Map quality scores (c) (mapQ2) for single cell mapping to the human genome after undergoing PTA with reversible or irreversible terminators (p_mapped2) are shown. [Figure 3B] The percentage of reads that map to the human genome (p_mapped2) after single cells underwent PTA with reversible or irreversible terminators is shown. [Figure 3C] A series of boxplots illustrating the reads aligned for the average percent reads overlapping with Alu elements using various methods are shown. PTA had the highest number of reads aligned to the genome. [Figure 3D] A series of boxplots illustrating PCR replicates of the average percent reads overlapping with Alu elements using various methods are shown. [Figure 3E] A series of boxplots illustrating the GC content of reads for the average percent reads overlapping with Alu elements using various methods are shown. [Figure 3F] A series of boxplots illustrating the mapping quality of the average percent reads overlapping Alu elements using various methods are shown. PTA had the highest mapping quality among the methods tested. [Figure 3G] Comparison of SC mitochondrial genome coverage widths by different WGA methods at a fixed 7.5x sequencing depth. [Figure 4A]Figure 1 shows the average coverage depth of 10-kilobase windows across chromosome 1 after selecting high-quality MDA cells (representing ~50% of cells) compared to random-primed PTA-amplified cells after downsampling each cell to 40 million paired reads. The figure demonstrates low MDA uniformity, with many windows having more (Box A) or less (Box C) than twice the average coverage depth. At the centromere, coverage is absent for both MDA and PTA (Box B) due to the high GC content and poor mapping quality of repetitive regions. [Figure 4B] Plots of sequencing coverage versus genome location for MDA and PTA methods are shown (top). Boxplots below show allele frequencies for MDA and PTA methods compared to the bulk sample. [Figure 5A] Plots of percent genome covered versus number of genome reads are shown to assess coverage at increasing sequencing depths for various methods. The PTA method came close to the two bulk samples at all depths, an improvement over the other methods tested. [Figure 5B] Figure 1 shows a plot of the coefficient of variation of genome coverage versus the number of reads to assess coverage uniformity. The PTA method was found to have the highest uniformity among the methods tested. [Figure 5C] A Lorenz plot of cumulative percentage of total reads versus cumulative percentage of genome is shown. The PTA method was found to have the highest uniformity among the methods tested. [Figure 5D] A series of boxplots of the Gini index calculated for each method tested to estimate the deviation of each amplification reaction from perfect uniformity are shown. The PTA method was found to be more reproducibly uniform than the other methods tested. [Figure 5E]A plot of the percentage of bulk variants called versus the number of reads is shown. The variant call percentage for each method was compared to the corresponding bulk sample at increasing sequencing depths. To estimate sensitivity, we calculated the percent of variants called in the corresponding bulk sample subsampled to 650 million reads found in each cell at each sequencing depth (Figure 3A). The improved coverage and uniformity of PTA led to the detection of 30% more variants than the next most sensitive method, Q-MDA. [Figure 5F] A series of boxplots shows the average percent reads overlapping Alu elements. The PTA method significantly reduced allelic skew at these heterozygous sites. The PTA method amplifies the two alleles more evenly within the same cell compared to other methods tested. [Figure 5G] To assess the specificity of variant calling, a plot of variant calling specificity versus the number of reads is shown. Variants found using various methods that were not found in the bulk sample were considered false positives. The PTA method yielded the lowest false positive calls (highest specificity) of the methods tested. [Figure 5H] The rate of false positive base changes for each type of base change across the various methods is shown. Without being bound by theory, it is possible that such patterns may be polymerase dependent. [Figure 5I] A series of boxplots of the average percent reads overlapping Alu elements for false-positive variant calls are shown. The PTA method yielded the lowest allele frequencies for false-positive variant calls. [Figure 6](Part A) shows beads bearing oligonucleotides with cleavable linkers, unique cell barcodes, and random primers. Part B shows a single cell and bead encapsulated in the same droplet, after which the cell is lysed and the primer is cleaved. The droplet can then be fused with another droplet containing a PTA amplification mixture. Part C shows that after amplification, the droplets are disrupted and the amplicons from all cells are pooled. Next, a protocol according to the present disclosure is utilized for terminator removal, end repair, and A-tailing prior to adapter ligation. The library of pooled cells then undergoes hybridization-mediated enrichment of exons of interest prior to sequencing. The cell of origin of each read is then identified using the cell barcode. [Figure 7A] The following illustrates a workflow for multi-omic (or poly-omic) analysis of single cells using PTA: Step A: Cells are contacted with antibodies containing fluorescent labels and oligonucleotide barcode tags. Step B: Cells are sorted based on the fluorescent marker. Step C: Tubes are coated with antibodies that bind to nuclei, cells are lysed, and cytosolic mRNA undergoes reverse transcription, while intact nuclei bind to the walls of the tube. [Figure 7B] 7 illustrates a workflow for multi-omic analysis of single cells using PTA, continuing from step C in Figure 7A. Step D: After reverse transcription, the RT fraction is removed for sequencing analysis. Step E: Nuclei are lysed and the PTA method is performed on genomic DNA. Step F: PTA generates a cDNA pool of short fragments with approximately 1000-fold amplification. [Figure 8A] Illustrates the primers used for reverse transcription and pre-amplification in a multi-omic DNA / RNA single-cell analysis workflow. [Figure 8B] 8 illustrates the reverse transcription and pre-amplification workflow for multi-omic DNA / RNA single-cell analysis workflow. Primers from Figure 8A were used. [Figure 9A]Figure 1 illustrates a graph of the proliferation rate of a parental cell line treated with 2 nM quizartinib for 3 weeks to generate a line of AML cells that proliferates robustly in the presence of a FLT3 inhibitor. Single resistant cells and parental cells (FACS enriched) were then analyzed by RNA sequencing and low-pass DNA sequencing analysis. [Figure 9B] RNA expression from both parental and resistant cultures demonstrated the ability to generate cDNA pools (C) using single-pot RNAseq chemistry, illustrating that genes expressed in these cells generated distinct patterns that allowed visualization of cell populations by gene expression, with an average of approximately 10,000 genes detected per cell. In a separate workflow, single-cell genomes were amplified using the PTA method. [Figure 9C] Illustrates normalized gene expression profiles for an RNAseq-only control experiment. [Figure 9D] Figure 1 illustrates a graph of the amount of DNA amplified by PTA versus different protocols. Transcripts generated during the RT step (R) are not effectively amplified by the PTA reaction compared to DNA, and DNA within single cells is effectively amplified using the combined protocols (SC1-SC8) compared to standard PTA-amplified genomes from single cells (D, RD). NTC = no template control. R = RT step, D = PTA DNA step, RD = dual RT / PTA. [Figure 10A] Figure 1 illustrates mitochondrial chromosome content (%) for two different protocols (dual RNAseq / PTA, standard RNAseq) using a low-pass sequencing protocol (~5 million reads / cell). The estimated genome size was over 3 billion bases. [Figure 10B] The overlap rate is illustrated for two different protocols (dual RNAseq / PTA, standard RNAseq) using a low-pass sequencing protocol (~5 million reads / cell). [Figure 10C] Illustrates estimated genome sizes for two different protocols (dual RNAseq / PTA, standard RNAseq) using a low-pass sequencing protocol (~5 million reads / cell). [Figure 10D] 1 illustrates feature assignment for three scRNAseq datasets from molm13 cells using the dual RNAseq / PTA protocol. [Figure 10E] Figure 1 illustrates a graph of normalized expression profiles for the Sum159 cell line obtained using standard RNAseq protocols. P = parental cells. R = resistant cells. [Figure 10F] Figure 1 illustrates a graph of normalized expression profiles for the Sum159 cell line obtained using the dual RNAseq / PTA protocol. P = parental cells. R = resistant cells. [Figure 11A] Deep sequencing results for seven parental and five resistant molm13 cells, performed at approximately 25x depth (K), are shown. Reads were aligned to Hg38 using bwamem. Quality control and SNV calling were performed using the best case for GATK4. SNVs were not considered if they were restricted to at least two resistant cells, alternative alleles were not called in any parental cells, and at least six parental cells were genotyped. All cells had at least 96% of their genomes covered at 1x coverage and at least 76% covered at 10x coverage. The inset shows that known Flt3 indels in molm13 cells are detected in all cells (four are shown for clarity). [Figure 11B] Illustrates a heat map of gene expression profiles, including the overexpressed gene GAS6, a known mechanism of quizartinib resistance. Gas6 is a ligand for AXL, a clinically relevant mechanism of resistance in relapsed patients who fail quizartinib treatment. [Figure 12A] 1 illustrates a graph of the percentage of exons covered in bulk versus single cell samples. [Figure 12B] 1 illustrates a graph of the proportion of exons with no coverage in bulk versus single cell samples. [Figure 12C]1 illustrates a graph of the percent of selected bases in bulk versus single cell samples. [Figure 12D] 1 illustrates a graph of the percentage of bases covered at 20x in bulk versus single cell samples. [Figure 13A] 1 illustrates a graph of mapped read-based locations within a genome stratified by treatment and shaded by sample type. [Figure 13B] 1 illustrates a graph of sample intensity versus captured insert size. [Figure 14A] 1 illustrates a graph of percent overlap versus percent selected bases for a 12-plex experiment. [Figure 14B] 1 illustrates a graph of number of target bases versus coverage level. DETAILED DESCRIPTION OF THE INVENTION
[0008] Detailed Description of the Invention There is a need to develop new, scalable, accurate, and efficient methods for nucleic acid amplification (including single-cell and multi-cell genome amplification) and sequencing that overcome the limitations of current methods by reproducibly increasing sequence representation, uniformity, and accuracy. Provided herein are compositions and methods for accurate and scalable primary template-directed amplification (PTA) and sequencing. Such methods and compositions facilitate highly accurate amplification of target (or "template") nucleic acids, thereby improving the accuracy and sensitivity of downstream applications such as next-generation sequencing. Also provided herein are methods for determining single-base variants, copy number diversity, structural diversity, clonotyping, and measuring environmental mutagenicity. Measuring genomic diversity by PTA can be used for a variety of applications, including environmental mutagenicity, predicting the safety of gene editing technologies, measuring genomic changes mediated by cancer therapy, measuring the carcinogenicity of compounds or radiation, including genotoxicity studies to determine the safety of new foods or drugs, age estimation, analysis of resistant bacteria, and identifying environmental bacteria for industrial applications. Furthermore, these methods can be used to detect the selection of specific cell populations after changes in environmental conditions, such as exposure to anti-cancer treatments, as well as to predict response to immunotherapy based on mutations and neoantigen load in single cancer cells.
[0009] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these inventions belong.
[0010] Throughout this disclosure, numerical characteristics are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Accordingly, the description of a range should be considered to specifically disclose all possible subranges, as well as individual numerical values to the tenth of the lower limit within that range, unless the context clearly dictates otherwise. For example, description of a range such as 1 to 6 should be considered to specifically disclose subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and individual values within that range, such as 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges and are also included in the invention, subject to any specifically excluded limits in the stated ranges. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention, unless the context clearly dictates otherwise.
[0011] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, it is understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any or all combinations of one or more of the associated listed items.
[0012] As used herein, unless specifically stated or clear from the context, the term "about" in reference to a number or range of numbers is understood to mean the stated number and + / -10% of that number, i.e., 10% below the lowest recited limit and 10% above the highest recited limit for the values recited for a range.
[0013] As used herein, the term "subject" or "patient" or "individual" refers to, for example, humans, veterinary animals (e.g., cats, dogs, cows, horses, sheep, pigs, etc.), and experimental animal models of disease (e.g., mice, rats). In accordance with the present invention, conventional molecular biology, microbiology, and recombinant DNA techniques may be employed that are within the skill of the art. Such techniques are fully explained in the literature. For example, among others, see Sambrook, Fritsch & Maniatis, Molecular Cloning: A Laboratory Manual, Second Edition (1989), Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (referred to herein as "Sambrook et al., 1989"); DNA Cloning: A Practical Approach, Volumes I and II (D.N. Glover ed. 1985); Oligonucleotide Synthesis (M.J. Gait ed. 1984); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985); Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984); Animal Cell Culture (R.I. Freshney, ed. (1986); Immobilized Cells and Enzymes (L.R. Press, (1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); F.M. Ausubel et al. (eds.), Current See Protocols in Molecular Biology, John Wiley & Sons, Inc. (1994).
[0014] The term "nucleic acid" encompasses not only single-stranded molecules but also multi-stranded molecules. In double- or triple-stranded nucleic acids, the nucleic acid strands need not be coextensive (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). The nucleic acid templates described herein can be of any size (from small cell-free DNA fragments to entire genomes) depending on the sample, including, but not limited to, lengths of 50-300 bases, 100-2000 bases, 100-750 bases, 170-500 bases, 100-5000 bases, 50-10,000 bases, or 50-2000 bases. In some examples, the length of the template is at least 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 50,000, 100,000, 200,000, 500,000, 1,000,000 bases, or more than 1,000,000 bases. The methods described herein provide for the amplification of nucleic acids, such as nucleic acid templates. The methods described herein also provide for the production of isolated and at least partially purified nucleic acids and libraries of nucleic acids. Nucleic acids include, but are not limited to, DNA, RNA, circular RNA, mtDNA (mitochondrial DNA), cfDNA (cell-free DNA), cfRNA (cell-free RNA), siRNA (small interfering RNA), cffDNA (cell-free fetal DNA), mRNA, tRNA, rRNA, miRNA (microRNA), synthetic polynucleotides, polynucleotide analogs, any other nucleic acids consistent with the present specification, or any combination thereof. The length of a polynucleotide, when provided, is stated as the number of bases and expressed in abbreviations such as nt (nucleotides), bp (bases), kb (kilobases), or Gb (gigabases).
[0015] As used herein, the term "droplet" refers to a volume of liquid on a droplet actuator. A droplet, in some examples, may be aqueous or non-aqueous, or a mixture or emulsion including aqueous and non-aqueous components. For non-limiting examples of droplet fluids that may be subjected to droplet operations, see, e.g., International Patent Application Publication No. WO 2007 / 120241. Any suitable system for forming and manipulating droplets can be used in the embodiments presented herein. For example, in some examples, a droplet actuator is used. Non-limiting examples of droplet actuators that may be used are described, for example, in U.S. Pat. Nos. 6,911,132, 6,977,033, 6,773,566, 6,565,727, 7,163,612, 7,052,244, 7,328,979, 7,547,380, 7,641,779, U.S. Patent Application Publication Nos. US20060194331, US20030205632, and US20060194331. See US20060164490, US20070023292, US20060039823, US20080124252, US20090283407, US20090192044, US20050179746, US20090321262, US20100096266, US20110048951, International Patent Application Publication No. WO2007 / 120241. In some cases, beads are provided in a droplet, in a droplet operations gap, or on a droplet operations surface. In some cases, the beads are provided in a reservoir located outside the droplet operations gap or away from the droplet operations surface, and the reservoir may be associated with a flow path that allows droplets containing the beads to enter the droplet operations gap or contact the droplet operations surface.Non-limiting examples of droplet actuator technologies for immobilizing magnetically responsive and / or non-magnetically responsive beads and / or for performing droplet manipulation protocols using beads are described in U.S. Patent Application Publication No. US20080053205, International Patent Application Publication Nos. WO2008 / 098236, WO2008 / 134153, WO2008 / 116221, and WO2007 / 120241. Bead properties can be utilized in multiplexed embodiments of the methods described herein. Examples of beads with properties suitable for multiplexing, as well as methods for detecting and analyzing signals emitted from such beads, can be found in U.S. Patent Application Publication Nos. US20080305481, US20080151240, US20070207513, US20070064990, US20060159962, US20050277197, and US20050118574.
[0016] Primers and / or template switching oligonucleotides can also be attached to a solid substrate to facilitate reverse transcription and template switching of mRNA polynucleotides. In this configuration, part of the RT or template switching reaction occurs in the bulk solution of the device, where the second step of the reaction occurs near the surface. In other configurations, the primers for the template switching oligonucleotide can be released from the solid substrate, allowing the entire reaction to occur on the surface of the solution. In polyomic approaches, primers for multistep reactions are sometimes immobilized on solid substrates or combined with beads to achieve multistep primer combinations.
[0017] Certain microfluidic devices also support polyomic approaches. As an example, devices fabricated with PDMS often feature adjacent chambers for each reaction step. Such multi-chamber devices are often separated using microvalve structures that can control pressure with air or fluids such as water or inert hydrocarbons (Fluorinert). In multiomic approaches, each step of a reaction can be isolated and carried out independently. Upon completion of a particular step, the valves between adjacent chambers can be opened on the substrate, allowing subsequent reactions to be added sequentially. As a result, individual cells can be used as input template materials to emulate a series of sequential reactions, such as a multiomic (protein / RNA / DNA / epigenomic) series of reactions. Various microfluidic platforms are available for single-cell analysis. Cells are manipulated, in some cases, through fluid dynamics (droplet microfluidics, inertial microfluidics, vortexing, microvalves, microstructures (microwells, microtraps, etc.)), electrical methods (dielectrophoresis (DEP), electroosmosis), optical methods (optical tweezers, optically induced dielectrophoresis (ODEP), photothermal capillaries), acoustic methods, or magnetic methods. In some cases, the microfluidic platform comprises a microwell. In some cases, the microfluidic platform comprises a PDMS (polydimethylsiloxane)-based device.Non-limiting examples of single cell analysis platforms compatible with the methods described herein include the ddSEQ single cell isolator (Bio-Rad, Hercules, CA, USA, and Illumina, San Diego, CA, USA), Chromium (10x Genomics, Pleasanton, CA, USA), the Rhapsody single cell analysis system (BD, Franklin Lakes, NJ, USA), the Tapestri platform (MissionBio, San Francisco, CA, USA), Nadia Innovate (Dolomite Bio, Royston, UK), C1 and Polaris (Fluidigm, South San Francisco, CA, USA); the ICELL8 single cell system (Takara); MSND (Wafergen); the Puncher platform (Vycap), the CellRaft AIR system (CellMicrosystems), the DEPArray NxT and DEPArray systems (Menarini Silicon Biosystems), and AVISO. CellCelector (ALS), InDrop system (1CellBio), and TrapTx (Celldom).
[0018] As used herein, the term "unique molecular identifier (UMI)" refers to a unique nucleic acid sequence attached to each of a plurality of nucleic acid molecules.When incorporated into nucleic acid molecules, UMIs are sometimes used to correct subsequent amplification bias by directly counting the sequenced UMIs after amplification.The design, incorporation and application of UMIs are described, for example, in International Patent Application Publication No. WO 2012 / 142213, Islam et al. Nat. Methods (2014) 11:163-166, Kivioja, T. et al. Nat. Methods (2012) 9:72-74, Brenner et al. (2000) PNAS 97(4), 1665, and Hollas and Schuler (2003) Conference: 3rd International Workshop on Algorithms in Bioinformatics, Volume: 2812.
[0019] As used herein, the term "barcode" refers to a nucleic acid tag that can be used to identify a sample or source of nucleic acid material. Thus, when nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample are sometimes tagged with different nucleic acid tags so that the source of the sample can be identified. Barcodes are also commonly referred to as indexes, tags, etc., and are well known to those skilled in the art. Any suitable barcode or set of barcodes can be used. For example, see the non-limiting examples provided in U.S. Patent No. 8,053,192 and International Patent Application Publication No. WO2005 / 068656. Barcoding of single cells can be performed, for example, as described in U.S. Patent Application No. 2013 / 0274117.
[0020] The terms "solid surface," "solid support," and other grammatical equivalents herein refer to any material suitable for, or that can be modified to be suitable for, attachment of the primers, barcodes, and sequences described herein. Exemplary substrates include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon, etc.), polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials (e.g., silicon or modified silicon), carbon, metals, inorganic glass, plastics, fiber optic bundles, and various other polymers. In some embodiments, the solid support comprises a patterned surface suitable for immobilizing primers, barcodes, and sequences in an ordered pattern.
[0021] As used herein, the term "biological sample" includes, but is not limited to, tissues, cells, biological fluids, and isolates thereof. Cells or other samples used in the methods described herein may be isolated from human patients, animals, plants, soil, or other samples containing microorganisms such as bacteria, fungi, or protozoa. In some cases, the biological sample is of human origin. In some cases, the biological sample is of non-human origin. Cells may be subjected to the PTA method described herein and sequenced. Variants detected throughout the genome or at specific locations can be compared with all other cells isolated from the subject to track the lineage history of the cell for research or diagnostic purposes. In some cases, the variants are confirmed through additional methods, such as direct PCR sequencing.
[0022] Single-cell analysis Described herein are methods and compositions for the analysis of single cells. Analyzing cells in bulk provides general information about a cell population but often fails to detect low-frequency variants against background. Such variants may possess important traits, such as mutations associated with drug resistance or cancer. In some cases, DNA, RNA, and / or proteins from the same single cell are analyzed in parallel. Analysis may include identification of epigenetic, post-translational modifications (e.g., glycosylation, phosphorylation, acetylation, ubiquitination, histone modifications) and / or post-transcriptional modifications (e.g., methylation, hydroxymethylation). Such methods may include "primary template-directed amplification" (PTA) to obtain libraries of nucleic acids for sequencing. In some cases, PTA is combined with additional steps or methods, such as RT-PCR or proteomic / protein quantification techniques (e.g., mass spectrometry, antibody staining, etc.). In some cases, various components of the cell are physically or spatially separated from each other between individual analytical steps. For example, the workflow in some examples includes the general steps shown in Figure 1A. Proteins are first labeled with antibodies. In some cases, at least some of the antibodies contain tags or markers (e.g., nucleic acid / oligotags, mass tags, or fluorescent tags). In some cases, some of the antibodies contain oligotags. In some cases, some of the antibodies contain fluorescent markers. In some cases, antibodies are labeled with more than one tag or marker. In some cases, some of the antibodies are sorted based on the fluorescent markers. After RT-PCR, first-strand mRNA products are generated and then removed for analysis. Libraries are then generated from the RT-PCR products and barcodes present in the protein-specific antibodies, which are then sequenced. In parallel, genomic DNA from the same cells is subjected to PCR, libraries are generated, and sequenced. Sequencing results from the genome, proteome, and transcriptome are, in some cases, pooled using bioinformatics methods.The methods described herein, in some cases, include any combination of labeling, cell sorting, affinity separation / purification, lysis of specific cellular components (e.g., outer membrane, nucleus, etc.), RNA amplification, DNA amplification (e.g., PTA), or other steps related to the separation or analysis of proteins, RNA, or DNA. In some cases, the methods described herein include one or more enrichment steps, such as exome enrichment.
[0023] Described herein is the first method for single-cell analysis involving the analysis of RNA and DNA from single cells (Figure 1B). This method involves single-cell isolation, single-cell lysis, and reverse transcription (RT). In some cases, reverse transcription is performed using template-switching oligonucleotides (TSOs). In some cases, the TSOs include a molecular tag, such as biotin, that allows for subsequent pull-down of the cDNA RT product and PCR amplification of the RT product to generate a cDNA library. Alternatively, or in combination, centrifugation is used to separate the RNA in the supernatant from the cDNA in the cell pellet. The remaining cDNA is fragmented and removed in some cases using UDG (uracil DNA glycosylase), and alkaline lysis is used to degrade the RNA and denature the genome. After neutralization, addition of primers, and PTA, the amplified products are purified in some cases on SPRI (solid-phase reversible immobilization) beads and ligated to adapters to generate a gDNA library.
[0024] Described herein is a second method for single-cell analysis, involving the analysis of RNA and DNA from single cells (Figure 1C). In some cases, this method involves single-cell isolation, single-cell lysis, and reverse transcription (RT). In some cases, reverse transcription is performed using template-switching oligonucleotides (TSOs). In some cases, the TSOs contain a molecular tag, such as biotin, that allows for subsequent pull-down of cDNA RT products, and PCR amplification of the RT products to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers, and PTA, the amplified products are in some cases purified on SPRI (solid-phase reversible immobilization) beads and ligated to adapters to generate a gDNA library. RT products are in some cases isolated by pull-down, such as pull-down using streptavidin beads.
[0025] Described herein is a third method for single-cell analysis, involving the analysis of RNA and DNA from single cells (Figure 1D). In some cases, this method involves single-cell isolation, single-cell lysis, and reverse transcription (RT). In some cases, reverse transcription is performed using a template-switching oligonucleotide (TSO) in the presence of a terminator nucleotide. In some cases, the TSO contains a molecular tag, such as biotin, that allows for subsequent pull-down of the cDNA RT products, and PCR amplification of the RT products to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers, and PTA, the amplified products are purified on SPRI (solid-phase reversible immobilization) beads and ligated to adapters to generate a DNA library. In some cases, the RT products are isolated by pull-down, such as pull-down using streptavidin beads.
[0026] Described herein is a fourth method for single-cell analysis, involving the analysis of RNA and DNA from single cells (Figure 1E). In some cases, this method involves single-cell isolation, single-cell lysis, and reverse transcription (RT). In some cases, reverse transcription is performed using template-switching oligonucleotides (TSOs). In some cases, the TSOs include a molecular tag, such as biotin, that allows for subsequent pull-down of cDNA RT products and PCR amplification of the RT products to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers, and PTA, the amplified products are subjected to RNase A and cDNA amplification, in some cases using blocked and labeled primers. gDNA is purified with SPRI (solid-phase reversible immobilization) beads and ligated to adapters to generate a gDNA library. RT products are isolated, in some cases, by pull-down, such as pull-down using streptavidin beads.
[0027] Described herein is a fifth method of single-cell analysis, involving the analysis of RNA and DNA from single cells (Figures 7A and 7B). A population of cells is contacted with an antibody library in which the antibodies are labeled. In some cases, the antibodies are labeled with either a fluorescent label, a nucleic acid barcode, or both. The labeled antibodies bind to at least one cell in the population, and the cells are sorted and placed into containers (e.g., tubes, vials, microwells, etc.), one cell per container. In some cases, the container contains a solvent. In some cases, a region on the surface of the container is coated with a capture moiety. In some cases, the capture moiety is a small molecule, antibody, protein, or other agent capable of binding to one or more cells, organelles, or other cellular components. In some cases, at least one cell, or a single cell, or component thereof, binds to the region on the container surface. In some cases, a nucleus binds to the region on the container. In some cases, the outer membrane of the cell dissolves, releasing the mRNA into the solution within the container. In some cases, the nucleus of the cell, containing genomic DNA, binds to the region on the container surface. Next, RT is often performed using mRNA in solution as a template to generate cDNA. In some cases, the template switching primer contains, from 5' to 3', a TSS region (transcription start site), an anchor region, an RNA BC region, and a poly(dT) tail. In some cases, the poly(dT) tail binds to the poly(A) tail of one or more mRNAs. In some cases, the template switching primer contains, from 3' to 5', a TSS region, an anchor region, and a poly(G) region. In some cases, the poly(G) region contains ribo-G. In some cases, the poly(G) region binds to a poly(C) region on the mRNA transcript. In some cases, ribo-G was added to the mRNA transcript by terminal transferase. After removing the RT PCR product for subsequent sequencing, any remaining RNA in the cell is removed by UNG. The nuclei are then lysed, and the released genomic DNA is subjected to PTA using random primers with an isothermal polymerase.In some cases, the primers are 6-9 bases in length. In some cases, PTA generates genomic amplicons with lengths of 100-5000, 200-5000, 500-2000, 500-2500, 1000-3000, or 300-3000 bases. In some cases, PTA generates genomic amplicons with an average length of 100-5000, 200-5000, 500-2000, 500-2500, 1000-3000, or 300-3000 bases. In some cases, PTA generates genomic amplicons with lengths of 250-1500 bases. In some cases, the methods described herein generate cDNA pools of short fragments with amplification of about 500, about 750, about 1000, about 5000, or about 10,000 fold. In some cases, the methods described herein generate cDNA pools of short fragments with amplification of 500-5000, 750-1500, or 250-10,000 fold. The PTA products are optionally subjected to additional amplification and sequencing.
[0028] Single cell sample preparation and isolation The methods described herein may require the isolation of single cells for analysis. Any method of single-cell isolation, such as mouth pipetting, micropipetting, flow cytometry / FACS, microfluidics, nuclear sorting methods (tetraploid or otherwise), or manual dilution, can be used with PTA. Such methods may be assisted by additional reagents and steps, such as antibody-based enrichment (e.g., circulating tumor cells), other small molecule- or protein-based enrichment methods, or fluorescent labeling. In some cases, the methods of multi-omic analysis described herein involve mechanical or enzymatic dissociation of cells from larger tissues.
[0029] Preparation and analysis of cellular components The multi-omic analysis methods described herein, including PTA, can involve one or more methods for processing cellular components such as DNA, RNA, and / or proteins. In some cases, the nucleus (containing genomic DNA) is physically separated from the cytosol (containing mRNA), followed by treatment with a membrane-selective lysis buffer that dissolves membranes but keeps the nucleus intact. The cytosol is then separated from the nucleus using methods including micropipetting, centrifugation, or antibody-conjugated magnetic microbeads. In another example, magnetic beads coated with oligo-dT primers bind to polyadenylated mRNA and separate it from DNA. In another example, DNA and RNA are simultaneously preamplified and separated for analysis. In another example, a cell is divided into two equal parts, and the mRNA from one half is processed and the genomic DNA from the other half is processed.
[0030] Multi-omics The methods described herein (e.g., PTA) can be used as a replacement for any number of other known methods in the art used for single-cell sequencing (such as multi-omics). PTA has the potential to replace genomic DNA sequencing methods such as MDA, PicoPlex, DOP-PCR, MALBAC, or target-specific amplification. In some cases, PTA replaces standard genomic DNA sequencing in multi-omics methods, including DR-seq (Dey et al., 2015), G&T-seq (MacAulay et al., 2015), scMT-seq (Hu et al., 2016), sc-GEM (Cheow et al., 2016), scTrio-seq (Hou et al., 2016), simultaneous multiplexed measurement of RNA and protein (Darmanis et al., 2016), scCOOL-seq (Guo et al., 2017), CITE-seq (Stoeckius et al., 2017), REAP-seq (Peterson et al., 2017), scNMT-seq (Clark et al., 2018), or SIDR-seq (Han et al., 2018). In some cases, the methods described herein include methods of PTA and polyadenylated mRNA transcripts. In some cases, the methods described herein include methods of PTA and non-polyadenylated mRNA transcripts. In some cases, the methods described herein include methods of PTA and total (polyadenylated and non-polyadenylated) mRNA transcripts.
[0031] In some cases, PTA is combined with standard RNA sequencing methods to obtain genomic and transcriptomic data. In some cases, the multi-omic methods described herein include PTA and one of the following: Drop-seq (Macosko et al. 2015), mRNA-seq (Tang et al. 2009), InDrop (Klein et al. 2015), MARS-seq (Jaitin et al. 2014), Smart-seq2 (Hashimshony et al. 2012; Fish et al. 2016), CEL-seq (Jaitin et al. 2014), STRT-seq (Islam et al. 2011), Quartz-seq (Sasagawa et al. 2013), CEL-seq2 (Hashimshony et al. 2016), cytoSeq (Fan et al. 2015), SuPeR-seq (Fan et al. 2015), or cytoSeq (Fan et al. 2015). al., 2011), RamDA-seq (Hayashi, et al. 2018), MATQ-seq (Sheng et al., 2017), or SMARTer (Verboom et al., 2019).
[0032] Various reaction conditions and mixes can be used to generate cDNA libraries for transcriptome analysis. In some cases, a RT reaction mix is used to generate a cDNA library. In some cases, the RT reaction mix includes a crowding reagent, at least one primer, a template switching oligonucleotide (TSO), a reverse transcriptase, and a dNTP mix. In some cases, the RT reaction mix includes an RNAse inhibitor. In some cases, the RT reaction mix includes one or more detergents. In some cases, the RT reaction mix includes Tween-20 and / or Triton-X. In some cases, the RT reaction mix includes betaine. In some cases, the RT reaction mix includes one or more salts. In some cases, the RT reaction mix includes a magnesium salt (e.g., magnesium chloride) and / or tetramethylammonium chloride. In some cases, the RT reaction mix includes gelatin. In some cases, the RT reaction mix includes PEG (PEG1000, PEG2000, PEG4000, PEG6000, PEG8000, or other lengths of PEG).
[0033] The multi-omic methods described herein can provide both genomic and RNA transcription information from a single cell (e.g., combination or dual protocols). In some cases, genomic information from a single cell is obtained from a PTA method, and RNA transcription information is obtained from reverse transcription to generate a cDNA library. In some cases, a total transcription method is used to obtain a cDNA library. In some cases, 3' or 5' end counting is used to obtain a cDNA library. In some cases, a UMI may not be used to obtain a cDNA library. In some cases, a multi-omic method provides RNA transcription information from a single cell for at least 500, 1000, 2000, 5000, 8000, 10,000, 12,000, or at least 15,000 genes. In some cases, a multi-omic method provides RNA transcription information from a single cell for about 500, 1000, 2000, 5000, 8000, 10,000, 12,000, or about 15,000 genes. In some cases, multi-omic methods provide RNA transcription information from a single cell for 100-12,000, 1000-10,000, 2000-15,000, 5000-15,000, 10,000-20,000, 8000-15,000, or 10,000-15,000 genes. In some cases, multi-omic methods provide genomic sequence information for at least 80%, 90%, 92%, 95%, 97%, 98%, or at least 99% of the genome of a single cell. In some cases, multi-omic methods provide genomic sequence information for about 80%, 90%, 92%, 95%, 97%, 98%, or about 99% of the genome of a single cell.
[0034] Multi-omic methods can involve the analysis of single cells from a cell population. In some cases, at least 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, or at least 8000 cells are analyzed. In some cases, approximately 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, or approximately 8000 cells are analyzed. In some cases, 5-100, 10-100, 50-500, 100-500, 100-1000, 50-5000, 100-5000, 500-1000, 500-10,000, 1000-10,000, or 5,000-20,000 cells are analyzed.
[0035] The multi-omic method can generate a yield of genomic DNA from a PTA reaction based on the type of single cell. In some cases, the amount of DNA generated from a single cell is about 0.1, 1, 1.5, 2, 3, 5, or about 10 micrograms. In some cases, the amount of DNA generated from a single cell is about 0.1, 1, 1.5, 2, 3, 5, or about 10 femtograms. In some cases, the amount of DNA generated from a single cell is at least 0.1, 1, 1.5, 2, 3, 5, or at least 10 micrograms. In some cases, the amount of DNA generated from a single cell is at least 0.1, 1, 1.5, 2, 3, 5, or at least 10 femtograms. In some cases, the amount of DNA generated from a single cell is about 0.1-10, 1-10, 1.5-10, 2-20, 2-50, 1-3, or 0.5-3.5 micrograms. In some cases, the amount of DNA produced from a single cell is about 0.1-10, 1-10, 1.5-10, 2-20, 2-4, 1-3, or 0.5-4 femtograms.
[0036] Methylome analysis Described herein are methods involving PTA, in which methylated DNA sites in a single cell are determined using the PTA method. In some cases, these methods further include parallel analysis of the transcriptome and / or proteome of the same cell. Methods for detecting methylated genomic bases include selective restriction with a methylation-sensitive endonuclease followed by treatment with the PTA method. The sites cleaved by such enzymes are determined by sequencing, and methylated bases are identified. In another example, bisulfite treatment of a genomic DNA library converts unmethylated cytosines to uracils. The library is then amplified, in some cases, with methylation-specific primers that selectively anneal to methylated sequences. Alternatively, non-methylation-specific PCR is performed, followed by one or more methods for distinguishing bisulfite-reactive bases, including direct pyrosequencing, MS-SnuPE, HRM, COBRA, MS-SSCA, or base-specific cleavage / MALDI-TOF. In some cases, the genomic DNA sample is split for parallel analysis of the genome (or enriched portions thereof) and methylome analysis. In some cases, the analysis of the genome and methylome includes enrichment of genomic fragments (e.g., exomes, or other targets) or whole genome sequencing.
[0037] Bioinformatics Data obtained from single-cell analysis methods utilizing the PTA described herein can be compiled into a database. Described herein are methods and systems for bioinformatics data integration. Data from proteomes, genomes, transcriptomes, methylomes, or other sources may, in some cases, be combined / integrated into a database and analyzed. Bioinformatics data integration methods and systems may, in some cases, include one or more of protein detection (FACS and / or NGS), mRNA detection, and / or genome distribution detection. In some cases, this data is correlated with a disease state or condition. In some cases, data from multiple single cells may be compiled to describe characteristics of a larger population of cells, such as cells from a particular sample, region, organism, or tissue. In some cases, protein data may be obtained from fluorescently labeled antibodies that selectively bind to proteins on cells. In some cases, protein detection methods may include grouping cells based on fluorescent markers and reporting the location of the samples after sorting. In some cases, protein detection methods may include detecting sample barcodes, detecting protein barcodes, comparing them to designed sequences, and grouping cells based on barcodes and copy number. In some cases, protein data is obtained from barcoded antibodies that selectively bind to proteins on cells. In some cases, transcriptome data is obtained from sample and RNA-specific barcodes. In some cases, methods for mRNA detection include sample and RNA-specific barcode detection, alignment to the genome, alignment to RefSeq / Encode, reporting of exon / intron / intergenic sequences, analysis of exon-exon junctions, grouping of cells based on barcodes and expression variance, and clustering analysis of variance and top variable genes. In some cases, genomic data is obtained from sample and DNA-specific barcodes.In some cases, methods for genome variance detection include detection of sample and DNA-specific barcodes, alignment to the genome, determination of genome recovery and SNV mapping rates, filtering of reads at exon-exon junctions, generation of variant call files (VCFs), and analysis of variance and clustering analysis of top variable variants.
[0038] mutation In some cases, the methods described herein (e.g., multi-omic PTA) result in higher detection sensitivity and / or lower false positive rates for mutation detection. In some cases, the mutation is a difference between the analyzed sequence (e.g., using the methods described herein) and a reference sequence. The reference sequence may be obtained from other organisms, other individuals of the same or similar species, populations of organisms, or other regions of the same genome. In some cases, the mutation is identified on a plasmid or chromosome. In some cases, the mutation is an SNV (single nucleotide change), SNP (single nucleotide polymorphism), or CNV (copy number variation, or CNA / copy number abnormality). In some cases, the mutation is a base substitution, insertion, or deletion. In some cases, the mutation is a transition, transversion, nonsense mutation, silent mutation, synonymous or nonsynonymous mutation, non-pathogenic mutation, missense mutation, or frameshift mutation (deletion or insertion). In some cases, PTA yields higher detection sensitivity and / or lower false positive rates when compared to methods such as in silico prediction, ChIP-seq, GUIDE-seq, Circle-seq, HTGTS (high-throughput genome-wide translocation sequencing), IDLV (integration-deficient lentivirus), Digenome-seq, FISH (fluorescence in situ hybridization), or DISCOVER-seq.
[0039] Primary template-directed amplification Described herein are nucleic acid amplification methods, such as "primary template-directed amplification (PTA)." In some cases, PTA is combined with other workflows for multi-omic analysis. For example, one embodiment of the PTA method described herein is schematically represented in Figure 1G. In the PTA method, a polymerase (e.g., a strand-displacing polymerase) is used to preferentially generate amplicons from a primary template ("direct copy"). As a result, errors are propagated from daughter amplicons at a lower rate during subsequent amplification compared to MDA. This results in an easily implemented method that can accurately and reproducibly amplify low DNA inputs, including the genomes of single cells, with high coverage breadth and uniformity, unlike existing WGA protocols. Furthermore, terminated amplification products can undergo directional ligation after terminator removal, allowing cell barcodes to be attached to amplification primers, allowing products from all cells to be pooled after undergoing parallel amplification reactions. In some cases, the template nucleic acid is not attached to a solid support. In some cases, the direct copy of the template nucleic acid is not attached to a solid support. In some cases, one or more primers are not bound to a solid support. In some cases, no primers are not bound to a solid support. In some cases, the primers are attached to a first solid support, and the template nucleic acid is attached to a second solid support, and the first and second solid supports are not the same. In some cases, PTA is used to analyze a single cell from a larger cell population. In some cases, PTA is used to analyze more than one cell from a larger cell population, or an entire cell population.
[0040] Described herein are methods for using nucleic acid polymerases with strand displacement activity for amplification. In some cases, such polymerases have strand displacement activity and a low error rate. In some cases, such polymerases have strand displacement activity and proofreading exonuclease activity, such as 3' to 5' proofreading activity. In some cases, nucleic acid polymerases are used in combination with other components, such as reversible or irreversible terminators or additional strand displacement factors. In some cases, the polymerase has strand displacement activity but does not have exonuclease proofreading activity. For example, in some instances, such polymerases include bacteriophage phi29 (Φ29) polymerase, which also has a very low error rate resulting from its 3' to 5' proofreading exonuclease activity (see, e.g., U.S. Patent Nos. 5,198,543 and 5,001,050). In some cases, non-limiting examples of strand-displacing nucleic acid polymerases include, for example, genetically engineered phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I (Jacobsen et al., Eur. J. BioChem. 45:623-627 (1974)), phage M2 DNA polymerase (Matsumoto et al., Gene 84:247 (1989)), phage phi PRD1 DNA polymerase (Jung et al., Proc. Natl. Acad. Sci. USA 84:8287 (1987); Zhu and Ito, Biochim. Biophys. Acta. 1219:267-276 (1994)), Bst DNA polymerase (e.g., Bst large fragment DNA polymerase (exo(-)Bst; Aliotta et al. al., Genet. Anal. (Netherlands) 12:185-195 (1996)), exo(-) Bca DNA polymerase (Walker and Linn, Clinical Chemistry 42:1604-1608 (1996)), Bsu DNA polymerase, Vent R Vent containing (exo-)DNA polymeraseRExamples of strand-displacing nucleic acid polymerases include DNA polymerases (Kong et al., J. Biol. Chem. 268:1965-1975 (1993)), Deep Vent DNA polymerases, including Deep Vent (exo-) DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase (Chatterjee et al., Gene 97:13-19 (1991)), Sequenase (USBiochemicals), T7 DNA polymerase, T7-Sequenase, T7 gp5 DNA polymerase, PRDI DNA polymerase, and T4 DNA polymerase (Kaboord and Benkovic, Curr. Biol. 5:149-157 (1995)). Additional strand-displacing nucleic acid polymerases are also compatible with the methods described herein. The ability of a given polymerase to perform strand displacement replication can be determined, for example, by using the polymerase in a strand displacement replication assay (e.g., as disclosed in U.S. Pat. No. 6,977,148). Such assays are sometimes performed at temperatures appropriate for the optimal activity of the enzyme used, such as 32°C for Phi29 DNA polymerase, 46°C to 64°C for exo(-)Bst DNA polymerase, or 60°C to 70°C for enzymes from hyperthermophilic organisms. Another useful assay for selecting polymerases is the primer blocking assay described in Kong et al., J. Biol. Chem. 268:1965-1975 (1993). This assay consists of a primer extension assay using an M13 ssDNA template in the presence or absence of an oligonucleotide that hybridizes upstream of the extension primer and blocks its progression. Other enzymes capable of displacing the blocking primer in this assay may, in some cases, be useful for the disclosed method. In some cases, the polymerase incorporates dNTPs and terminators in approximately equal proportions.In some cases, the ratio of dNTP and terminator incorporation rates for polymerases described herein is about 1:1, about 1.5:1, about 2:1, about 3:1, about 4:1, about 5:1, about 10:1, about 20:1, about 50:1, about 100:1, about 200:1, about 500:1, or about 1000:1. In some cases, the ratio of dNTP and terminator incorporation rates for polymerases described herein is 1:1 to 1000:1, 2:1 to 500:1, 5:1 to 100:1, 10:1 to 1000:1, 100:1 to 1000:1, 500:1 to 2000:1, 50:1 to 1500:1, or 25:1 to 1000:1.
[0041] Described herein is an amplification method that can promote strand displacement through the use of strand displacement factors, such as helicases.In some cases, such factors are used in combination with additional amplification components, such as polymerases, terminators, or other components.In some cases, strand displacement factors are used with polymerases that do not have strand displacement activity.In some cases, strand displacement factors are used with polymerases that have strand displacement activity.Without being bound by theory, strand displacement factors can increase the rate at which smaller double-stranded amplicons are reprimed.In some cases, any DNA polymerase that can perform strand displacement replication in the presence of strand displacement factors is suitable for use in PTA methods, even if the DNA polymerase does not perform strand displacement replication in the absence of such factors.Strand displacement factors useful in strand displacement replication include, in some cases, the BMRF1 polymerase accessory subunit (Tsurumi et al., J. Virology 67(12):7648-7653 (1993)), adenovirus DNA binding protein (Zijderveld and van der Vliet, J. Virology 68(2):1158-1164 (1994)), herpes simplex virus protein ICP8 (Boehmer and Lehman, J. Virology 67(2):711-715 (1993); Skaliter and Lehman, Proc. Natl. Acad. Sci. USA 91(22):10665-10669 (1994)); single-stranded DNA binding protein (SSB; Rigler and Romano, J. Biol. Chem. 270:8910-8919 (1995)); phage T4 gene 32 protein (Villemain and Giedroc, Biochemistry 35:14395-14404 (1996)); T7 helicase-primase; T7 gp2.5 SSB protein; Tte-UvrD (from Thermoanaerobacter tengcongensis), calf thymus helicase (Siegel et al., J. Biol. Chem. 267:13629-13635 (1992)); bacterial SSB (e.g., E. coli SSB), replication protein A (RPA) in eukaryotes, human mitochondrial SSB (mtSSB), and recombinases (e.g., recombinase A (RecA) family proteins, T4 UvsX, T4 Examples of suitable factors that promote strand displacement and priming include, but are not limited to, UvsY, Sak4 from phage HK620, Rad51, Dmc1, or Radd. Combinations of factors that promote strand displacement and priming are also consistent with the methods described herein. For example, a helicase is used in conjunction with a polymerase.In some cases, the PTA method involves the use of a single-stranded DNA binding protein (SSB, T4 gp32, or other single-stranded DNA binding protein), a helicase, and a polymerase (e.g., Sau DNA polymerase, Bsu polymerase, Bst2.0, GspM, GspM2.0, GspSSD, or other suitable polymerase). In some cases, a reverse transcriptase is used in combination with a strand displacement factor described herein. A reverse transcriptase is used in combination with a strand displacement factor described herein. In some cases, amplification is performed using a polymerase and a nicking enzyme (e.g., "NEAR") as described in U.S. Pat. No. 9,617,586. In some cases, the nicking enzyme is Nt.BspQI, Nb.BbvCi, Nb.BsmI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nt.BbvCI, Nt.BstNBI, Nt.CviPII, Nb.Bpu10I, or Nt.Bpu10I.
[0042] Described herein is an amplification method that includes the use of terminator nucleotides, polymerase, and additional factors or conditions. For example, in some cases, such factors are used to fragment nucleic acid templates or amplicons during amplification. In some cases, such factors include endonucleases. In some cases, factors include transposases. In some cases, mechanical shearing is used to fragment nucleic acids during amplification. In some cases, nucleotides are added during amplification, and these can be fragmented by adding additional proteins or conditions. For example, uracil is incorporated into an amplicon, and treatment with uracil D-glycosylase fragments the nucleic acid at positions containing uracil. Additional systems for selective nucleic acid fragmentation are also utilized in some cases, such as engineered DNA glycosylases that cleave modified cytosine-pyrene base pairs (Kwon, et al. Chem Biol. 2003, 10(4), 351).
[0043] Described herein are amplification methods that involve the use of terminator nucleotides, which terminate nucleic acid replication and thus reduce the size of the amplified product. Such terminators are sometimes used in combination with polymerases, strand displacement factors, or other amplification components described herein. In some cases, terminator nucleotides reduce or decrease the efficiency of nucleic acid replication. In some cases, such terminators reduce the extension rate by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. In some cases, such terminators reduce the extension rate by 50% to 90%, 60% to 80%, 65% to 90%, 70% to 85%, 60% to 90%, 70% to 99%, 80% to 99%, or 50% to 80%. In some cases, terminators reduce the average amplicon product length by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. In some cases, terminators reduce the average amplicon length by 50%-90%, 60%-80%, 65%-90%, 70%-85%, 60%-90%, 70%-99%, 80%-99%, or 50%-80%. In some cases, amplicons containing terminator nucleotides form loops or hairpins, which reduce the ability of polymerases to use such amplicons as templates. The use of terminators, in some cases, slows the rate of amplification at initial amplification sites through the incorporation of terminator nucleotides (e.g., dideoxynucleotides modified to be exonuclease-resistant to stop DNA extension), resulting in smaller amplification products. By generating smaller amplification products than currently used methods (e.g., an average length of 50-2,000 nucleotides for the PTA method compared to an average product length of >10,000 nucleotides for the MDA method), PTA amplification products can, in some cases, undergo direct ligation of adapters without the need for fragmentation, allowing for the efficient incorporation of cell barcodes and unique molecular identifiers (UMIs) (see Figure 2A).
[0044] Terminator nucleotides are present at various concentrations depending on factors such as polymerase, template, or other factors. For example, the amount of terminator nucleotides is sometimes expressed as the ratio of non-terminator nucleotides to terminator nucleotides in the methods described herein. Such concentrations can sometimes control the length of amplicons. In some cases, the ratio of terminator to non-terminator nucleotides is changed depending on the amount of template present or the size of the template. In some cases, the ratio of terminator to non-terminator nucleotides becomes smaller as the sample size becomes smaller (e.g., in the femtogram and picogram ranges). In some cases, the ratio of non-terminator nucleotides to terminator nucleotides is about 2:1, 5:1, 7:1, 10:1, 20:1, 50:1, 100:1, 200:1, 500:1, 1000:1, 2000:1, or 5000:1. In some cases, the ratio of non-terminator to terminator nucleotides is 2:1 to 10:1, 5:1 to 20:1, 10:1 to 100:1, 20:1 to 200:1, 50:1 to 1000:1, 50:1 to 500:1, 75:1 to 150:1, or 100:1 to 500:1. In some cases, at least one of the nucleotides present during amplification using the methods described herein is a terminator nucleotide. Each terminator need not be present at approximately the same concentration; in some cases, the ratio of each terminator present in the methods described herein is optimized for a particular set of reaction conditions, sample type, or polymerase. Without being bound by theory, each terminator may have different efficiencies for incorporation into the growing polynucleotide strand of an amplicon in response to pairing with the corresponding nucleotide on the template strand. For example, in some cases, cytosine-pairing terminators are present at a concentration that is about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration.In some cases, terminators paired with thymine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with guanine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with adenine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with uracil are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, any nucleotide that can terminate nucleic acid elongation by nucleic acid polymerase is used as terminator nucleotide in the method described herein.In some cases, reversible terminator is used to terminate nucleic acid replication.In some cases, irreversible terminator is used to terminate nucleic acid replication.In some cases, non-limiting examples of terminator include reversible and irreversible nucleic acids and nucleic acid analogs, such as 3'-blocked reversible terminator that comprises nucleotides, 3'-unblocked reversible terminator that comprises nucleotides, terminator that comprises 2'-modified deoxynucleotides, terminator that comprises modified nitrogen base of deoxynucleotides, or any combination thereof.In one embodiment, terminator nucleotide is dideoxynucleotide. Other nucleotide modifications that terminate nucleic acid replication and are suitable for practicing the present invention include, but are not limited to, any modification of the r group of the 3' carbon of deoxyribose, such as reverse dideoxynucleotides, 3' biotinylated nucleotides, 3' amino nucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3' carbon spacer nucleotides, including 3' C3 spacer nucleotides, 3' C18 nucleotides, 3' hexanediol spacer nucleotides, acyclonucleotides, and combinations thereof.In some cases, the terminator is a polynucleotide containing 1, 2, 3, 4, or more bases in length. In some cases, the terminator does not contain a detectable moiety or tag (e.g., a mass tag, a fluorescent tag, a dye, a radioactive atom, or other detectable moiety). In some cases, the terminator does not contain a chemical moiety that allows for the attachment of a detectable moiety or tag (e.g., a "click" azide / alkyne, a conjugate addition partner, or other chemical handle for tag attachment). In some cases, all terminator nucleotides contain the same modification that reduces amplification in a region of the nucleotide (e.g., the sugar moiety, the base moiety, or the phosphate moiety). In some cases, at least one terminator has a different modification that reduces amplification. In some cases, all terminators have substantially similar fluorescence excitation or emission wavelengths. In some cases, terminators without modifications to the phosphate group are used with polymerases that do not have exonuclease proofreading activity. When used with a polymerase that has 3' to 5' proofreading exonuclease activity (e.g., Phi29) that can remove terminator nucleotides, terminators are, in some cases, further modified to render them exonuclease-resistant. For example, dideoxynucleotides are modified with alpha-thio groups that create phosphorothioate linkages that render these nucleotides resistant to the 3' to 5' proofreading exonuclease activity of nucleic acid polymerases. Such modifications, in some cases, reduce the exonuclease proofreading activity of the polymerase by at least 99.5%, 99%, 98%, 95%, 90%, or at least 85%.Non-limiting examples of other terminator nucleotide modifications that confer resistance to 3'->5' exonuclease activity include, in some cases, nucleotides with modifications to the alpha group, such as alpha-thiodideoxynucleotides that create phosphorothioate linkages, C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2'-fluoro bases, 3' phosphorylation, 2'-O-methyl modifications (or other 2'-O-alkyl modifications), propyne-modified bases (e.g., deoxycytosine, deoxyuridine), L-DNA nucleotides, L-RNA nucleotides, nucleotides with inverted linkages (e.g., 5'-5' or 3'-3'), 5'-inverted bases (e.g., 5'-inverted 2', 3'-dideoxydT), methylphosphonate backbones, and trans nucleic acids. In some cases, modified nucleotides include base-modified nucleic acids containing a free 3'OH group (e.g., bases containing modifications with bulky chemical groups such as 2-nitrobenzyl alkylated HOMedU triphosphate, solid supports, or other bulky moieties). In some cases, polymerases with strand displacement activity but no 3' to 5' exonuclease proofreading activity are used with terminator nucleotides, with or without modifications to make them exonuclease resistant. Such nucleic acid polymerases include, but are not limited to, Bst DNA polymerase, Bsu DNA polymerase, Deep Vent (exo-) DNA polymerase, Klenow fragment (exo-) DNA polymerase, Therminator DNA polymerase, and Vent. R (exo-) is included.
[0045] Primers and amplicon libraries Described herein are amplicon libraries resulting from the amplification of at least one target nucleic acid molecule. Such libraries are, in some cases, generated using methods described herein, such as those using terminators. Such methods include the use of strand-displacing polymerases or agents, terminator nucleotides (reversible or irreversible), or other features and embodiments described herein. In some cases, the amplicon library generated by the use of terminators described herein is further amplified in a subsequent amplification reaction (e.g., PCR). In some cases, the subsequent amplification reaction does not include a terminator. In some cases, the amplicon library includes polynucleotides, and at least 50%, 60%, 70%, 80%, 90%, 95%, or at least 98% of the polynucleotides include at least one terminator nucleotide. In some cases, the amplicon library includes the target nucleic acid molecule from which the amplicon library was derived. An amplicon library comprises a plurality of polynucleotides, at least some of which are direct copies (e.g., directly replicated from a target nucleic acid molecule, such as genomic DNA, RNA, or other target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 95% or more of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 5% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 10% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 15% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 20% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 50% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule.In some cases, 3% to 5%, 3% to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50%, or 15% to 75% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least some of the polynucleotides are direct copies of the target nucleic acid molecule or descendants of daughters (initial copies of the target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more than 95% of the amplicon polynucleotides are direct copies or descendants of daughters of at least one target nucleic acid molecule. In some cases, at least 5% of the amplicon polynucleotides are direct copies or descendants of daughters of at least one target nucleic acid molecule. In some cases, at least 10% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, at least 20% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, at least 30% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, 3% to 5%, 3% to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50%, or 15% to 75% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, the direct copies of the target nucleic acid are 50-2500, 75-2000, 50-2000, 25-1000, 50-1000, 500-2000, or 50-2000 bases in length. In some cases, the lengths of the daughter progeny are 1000-5000, 2000-5000, 1000-10,000, 2000-5000, 1500-5000, 3000-7000, or 2000-7000 bases in length. In some cases, the average length of the PTA amplification products is 25-3000 nucleotides in length, 50-2500, 75-2000, 50-2000, 25-1000, 50-1000, 500-2000, or 50-2000 bases in length.In some cases, amplicons generated from PTA are 5,000, 4,000, 3,000, 2,000, 1,700, 1,500, 1,200, 1,000, 700, 500 bases or less, or 300 bases or less. In some cases, amplicons generated from PTA are 1,000-5,000, 1,000-3,000, 200-2,000, 200-4,000, 500-2,000, 750-2,500, or 1,000-2,000 bases in length. In some cases, amplicon libraries generated using the methods described herein include at least 1,000, 2,000, 5,000, 10,000, 100,000, 200,000, 500,000, or more than 500,000 amplicons comprising unique sequences. In some cases, the library comprises at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 2000, 2500, 3000, or at least 3500 amplicons. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of less than 1000 bases are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 2000 bases or less are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 3000 to 5000 bases are direct copies of at least one target nucleic acid molecule. In some cases, the ratio of direct copy amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or more than 10,000,000:1.In some cases, the ratio of direct copy amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, where the direct copy amplicons are 700-1200 bases in length or less. In some cases, the ratio of direct copy amplicons and daughter amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1. In some cases, the ratio of direct copy amplicons and daughter amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, wherein the direct copy amplicons are 700-1200 bases in length and the daughter amplicons are 2500-6000 bases in length. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule or daughter amplicons. The number of direct copies can, in some cases, be controlled by the number of PCR amplification cycles. In some cases, up to 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or 3 PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, about 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or about 3 PCR cycles are used to generate copies of the target nucleic acid molecule.In some cases, 3, 4, 5, 6, 7, or 8 PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, 2-4, 2-5, 2-7, 2-8, 2-10, 2-15, 3-5, 3-10, 3-15, 4-10, 4-15, 5-10, or 5-15 PCR cycles are used to generate copies of the target nucleic acid molecule. The amplicon libraries generated using the methods described herein are sometimes subjected to additional steps, such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed before the sequencing step.
[0046] The methods described herein may further include one or more enrichment or purification steps. In some cases, one or more polynucleotides (such as cDNAs, PTA amplicons, or other polynucleotides) are enriched during the methods described herein. In some cases, polynucleotide probes are used to capture one or more polynucleotides. In some cases, the probes are configured to capture one or more genomic exons. In some cases, the library of probes comprises at least 1,000, 2,000, 5,000, 10,000, 50,000, 100,000, 200,000, 500,000, or more than 1 million different sequences. In some cases, the library of probes comprises sequences capable of binding to at least 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, or more than 10,000 genes. In some cases, the probes comprise a moiety for capture by a solid support, such as biotin. In some cases, the enrichment step is performed after the PTA step. In some cases, the enrichment step is performed before the PTA step. In some cases, the probes are configured to bind to a genomic DNA library. In some cases, the probes are configured to bind to a cDNA library.
[0047] In some cases, polynucleotide amplicon libraries generated from the PTA methods and compositions (terminators, polymerases, etc.) described herein exhibit increased uniformity. Uniformity may be described using a Lorenz curve (e.g., FIG. 5C) or other such methods. Such an increase may result in fewer sequencing reads than are required for desired coverage of a target nucleic acid molecule (e.g., genomic DNA, RNA, or other target nucleic acid molecule). For example, 50% or less of the cumulative percentage of polynucleotides contain at least 80% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 60% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 70% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 90% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, uniformity is described using a Gini index (where an index of 0 represents perfect equality of the libraries and an index of 1 represents perfect inequality). In some cases, the amplicon libraries described herein have a Gini index of 0.55, 0.50, 0.45, 0.40, or 0.30 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.40 or less. In some cases, such a uniformity metric depends on the number of reads obtained. For example, 100, 200, 300, 400, or 500 million reads or less are obtained. In some cases, the read length is about 50, 75, 100, 125, 150, 175, 200, 225, or about 250 bases long. In some cases, the uniformity metric depends on the coverage depth of the target nucleic acid. For example, the average depth of coverage is about 10x, 15x, 20x, 25x, or about 30x.In some cases, the average coverage depth is 10-30x, 20-50x, 5-40x, 20-60x, 5-20x, or 10-20x. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less, wherein the average depth of sequencing coverage is 15-fold or less.In some cases, the Gini index of the amplicon library described herein is 0.50 or less, and the average depth of sequencing coverage is 15 times or less. In some cases, the Gini index of the amplicon library described herein is 0.45 or less, and the average depth of sequencing coverage is 15 times or less. In some cases, the uniform amplicon library generated using the method described herein is subjected to additional steps, such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed before the sequencing step.
[0048] Primers include nucleic acids used to prime the amplification reactions described herein. Such primers include, but are not limited to, random deoxynucleotides of any length, with or without modifications to make them exonuclease resistant, random ribonucleotides of any length, with or without modifications to make them exonuclease resistant, modified nucleic acids such as locked nucleic acids, or DNA or RNA primers targeting specific genomic regions and reactions primed with enzymes such as primase. For whole genome PTA, it is preferred to use a set of primers with random or partially random nucleotide sequences. In nucleic acid samples of significant complexity, the specific nucleic acid sequences present in the sample do not need to be known, and primers do not need to be designed to be complementary to specific sequences. Rather, the complexity of nucleic acid samples results in a large number of different hybridization target sequences in the sample, which are complementary to various primers with random or partially random sequences. The complementary portions of primers for use in PTA may in some cases be completely randomized, contain only randomized portions, or be selectively randomized. In some cases, the number of random base positions in the complementary portion of the primer is, for example, 20% to 100% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is 10% to 90%, 15% to 95%, 20% to 100%, 30% to 100%, 50% to 100%, 75% to 100%, or 90% to 95% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the total number of nucleotides in the complementary portion of the primer.A set of primers having random or partially random sequences is synthesized using standard techniques, in some cases by allowing the addition of any nucleotide at each position to be randomized. In some cases, the set of primers is composed of primers of similar length and / or hybridization characteristics. In some cases, the term "random primer" refers to a primer that can exhibit 4-fold degeneracy at each position. In some cases, the term "random primer" refers to a primer that can exhibit 3-fold degeneracy at each position. The random primers used in the methods herein, in some cases, contain random sequences that are 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more bases in length. In some cases, the primers contain random sequences that are 3-20, 5-15, 5-20, 6-12, or 4-10 bases in length. Primers can also contain non-extendable elements that limit subsequent amplification of the amplicon generated therefrom. For example, a primer having a non-extendable element may, in some cases, include a terminator. In some cases, the primer may include a terminator nucleotide, such as 1, 2, 3, 4, 5, 10, or more than 10 terminator nucleotides. Primers need not be limited to components added externally to an amplification reaction. In some cases, primers are generated in situ by the addition of nucleotides and proteins that facilitate priming. For example, a primase-like enzyme in combination with nucleotides may, in some cases, be used to generate random primers for the methods described herein. The primase-like enzyme may, in some cases, be a member of the DnaG or AEP enzyme superfamily. In some cases, the primase-like enzyme may be TthPrimPol. In some cases, the primase-like enzyme may be T7gp4 helicase-primase. Such primases may, in some cases, be used in conjunction with a polymerase or strand displacement factor described herein.In some cases, primases initiate priming using deoxyribonucleotides. In some cases, primases initiate priming using ribonucleotides.
[0049] Following PTA amplification, a specific subset of amplicons can be selected. Such selection, in some cases, relies on size, affinity, activity, hybridization to probes, or other selection factors known in the art. In some cases, selection is performed before or after additional steps described herein, such as adapter ligation and / or library amplification. In some cases, selection is performed based on amplicon size (length). In some cases, smaller amplicons that are less likely to have undergone exponential amplification are selected, which further converts amplification from an exponential to a sublinear amplification process while enriching for products derived from the primary template (Figure 1A). In some cases, amplicons with lengths of 50-2000, 25-5000, 40-3000, 50-1000, 200-1000, 300-1000, 400-1000, 400-600, 600-2000, or 800-1000 bases are selected. In some cases, size selection is performed, for example, by using a protocol that utilizes solid-phase reversible immobilization (SPRI) on carboxylated paramagnetic beads to enrich for nucleic acid fragments of a specific size, or other protocols known to those skilled in the art. Optionally, or in combination, selection is performed during sequencing library preparation through preferential ligation and amplification of small fragments during PCR, and through the preferential formation of clusters from smaller sequencing library fragments during sequencing (e.g., sequencing by synthesis, nanopore sequencing, or other sequencing methods). Other strategies for selecting smaller fragments are also consistent with the methods described herein, including, but not limited to, isolating nucleic acid fragments of a specific size after gel electrophoresis, using silica columns that bind to nucleic acid fragments of a specific size, and other PCR strategies that more strongly enrich for smaller fragments. Any number of library preparation protocols can be used with the PTA method described herein.In some cases, the amplicons generated by PTA are ligated to adapters (optionally with removal of terminator nucleotides). In some cases, the amplicons generated by PTA contain regions of homology generated from transposase-based fragmentation that are used as priming sites. In some cases, libraries are prepared by mechanically or enzymatically fragmenting nucleic acids. In some cases, libraries are prepared using transposome-mediated tagging. In some cases, libraries are created by ligation of adapters, such as Y-adapters, universal adapters, or circular adapters.
[0050] The non-complementary portions of primers used in PTA can contain sequences that can be used to further manipulate and / or analyze the amplified sequences. An example of such a sequence is a "detection tag." Detection tags have sequences complementary to detection probes and are detected using their cognate detection probes. A primer may have one, two, three, four, or more than four detection tags. There is no fundamental limit to the number of detection tags that can be present on a primer, except for the size of the primer. In some cases, a primer has a single detection tag. In some cases, a primer has two detection tags. When multiple detection tags are present, they may have the same sequence or they may have different sequences, with each different sequence being complementary to a different detection probe. In some cases, multiple detection tags have the same sequence. In some cases, multiple detection tags have different sequences.
[0051] Another example of a sequence that can be included in the non-complementary portion of a primer is an "address tag," which can encode other details of the amplicon, such as its location within a tissue section. In some cases, the cellular barcode includes an address tag. The address tag has a sequence complementary to an address probe. The address tag is incorporated into the end of the amplified strand. If present, there can be one or more address tags in a primer. There is no fundamental limit to the number of address tags that can be present in a primer, except for the size of the primer. If multiple address tags are present, they can have the same sequence or different sequences, with each different sequence complementary to a different address probe. The address tag portion can be any length that supports specific and stable hybridization between the address tag and the address probe. In some cases, nucleic acids from more than one source can incorporate a variable tag sequence. This tag sequence can be up to 100 nucleotides long, preferably 1 to 10 nucleotides long, and most preferably 4, 5, or 6 nucleotides long, and can contain any combination of nucleotides. In some cases, the tag sequence is 1-20, 2-15, 3-13, 4-12, 5-12, or 1-10 nucleotides in length. For example, six base pairs are selected to form the tag, and four different nucleotide permutations are used, which then creates a total of 4096 nucleic acid anchors (e.g., hairpins), each with a unique six base pairs.
[0052] The primers described herein can be in solution or immobilized on a solid support. In some cases, primers bearing sample barcodes and / or UMI sequences can be immobilized on a solid support. The solid support can be, for example, one or more beads. In some cases, to identify individual cells, individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences. In some cases, lysates from individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. In some cases, to identify nucleic acids extracted from individual cells, nucleic acids extracted from individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences. The beads can be manipulated by any suitable method known in the art, for example, using a droplet actuator described herein. The beads can be of any suitable size, including, for example, microbeads, microparticles, nanobeads, and nanoparticles. In some embodiments, the beads are magnetically responsive, while in other embodiments, the beads are not significantly magnetically responsive.Non-limiting examples of suitable beads include flow cytometry microbeads, polystyrene microparticles and nanoparticles, functionalized polystyrene microparticles and nanoparticles, coated polystyrene microparticles and nanoparticles, silica microbeads, fluorescent microspheres and nanospheres, functionalized fluorescent microspheres and nanospheres, coated fluorescent microspheres and nanospheres, colored microparticles and nanoparticles, magnetic microparticles and nanoparticles, superparamagnetic microparticles and nanoparticles (e.g., DYNABEADS® available from Invitrogen Group, Carlsbad, CA), fluorescent microparticles and nanoparticles, coated magnetic microparticles and nanoparticles, ferromagnetic microparticles and nanoparticles, coated ferromagnetic microparticles and nanoparticles, and those described in U.S. Patent Application Publication Nos. US20050260686, US20030132538, US20050118574, 20050277197, 20060159962. The beads may be pre-bound with antibodies, proteins or antigens, DNA / RNA probes, or any other molecules with affinity for the desired target. In some embodiments, primers with sample barcodes and / or UMI sequences may be in solution. In certain embodiments, multiple droplets may be presented, each droplet within the multiple droplets having a sample barcode unique to the droplet and a UMI unique to the molecule, such that the UMI is repeated multiple times within the collection of droplets. In some embodiments, individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cells. In some embodiments, lysates from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. In some embodiments, nucleic acids extracted from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates.
[0053] PTA primers may include sequence-specific or random primers, cell barcodes, and / or unique molecular identifiers (UMIs) (see, e.g., Figures 10A (linear primers) and 10B (hairpin primers)). In some cases, primers include sequence-specific primers. In some cases, primers include random primers. In some cases, primers include cell barcodes. In some cases, primers include sample barcodes. In some cases, primers include unique molecular identifiers. In some cases, primers include more than one cell barcode. Such barcodes, in some cases, identify a unique sample source or a unique workflow. Such barcodes or UMIs are, in some cases, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or greater than 30 bases in length. In some cases, primers are at least 1000, 10,000, 50,000, 100,000, 250,000, 500,000, 10 6 , 10 7 , 10 8 , 10 9 , or at least 10 10Each primer contains 8, 16, 96, or 384 unique barcodes or UMIs. In some cases, primers contain at least 8, 16, 96, or 384 unique barcodes or UMIs. In some cases, standard adapters are ligated to the amplification products before sequencing, and after sequencing, reads are first assigned to specific cells based on their cellular barcodes. Suitable adapters that can be used with the PTA method include, for example, the xGen® Dual Index UMI adapters available from Integrated DNA Technologies (IDT). Reads from each cell are then grouped using the UMI, and reads with the same UMI are collapsed into a consensus read. The use of cellular barcodes allows all cells to be pooled before library preparation, as they can later be identified by their cellular barcodes. The use of UMIs to form consensus reads, in some cases, corrects for PCR bias and improves the detection of copy number variation (CNV) (Figures 11A and 11B). Additionally, sequencing errors can be corrected by requiring a fixed percentage of reads from the same molecule to have the same base change detected at each position. This approach has been utilized to improve CNV detection and correct sequencing errors in bulk samples. In some cases, UMIs are used in conjunction with the methods described herein; for example, U.S. Patent No. 8,835,358 discloses the principle of digital counting after randomly attaching amplifiable barcodes. Schmitt et al. and Fan et al. disclose similar methods for correcting sequencing errors. In some cases, libraries are generated for sequencing using primers. In some cases, libraries are composed of fragments 200-700 bases, 100-1000 bases, 300-800 bases, 300-550 bases, 300-700 bases, or 200-800 bases in length. In some cases, libraries include fragments at least 50, 100, 150, 200, 300, 500, 600, 700, 800, or at least 1000 bases in length.In some cases, the library comprises fragments that are about 50, 100, 150, 200, 300, 500, 600, 700, 800, or about 1000 bases in length.
[0054] The methods described herein may further include additional steps, including steps performed on a sample or template. Such a sample or template may be subjected to one or more steps prior to PTA. In some cases, a sample containing cells is subjected to a pretreatment step. For example, cells are subjected to lysis and proteolysis using a combination of freeze-thaw, Triton X-100, Tween 20, and Proteinase K to increase chromatin accessibility. Other lysis strategies are also suitable for practicing the methods described herein. Such strategies include, but are not limited to, lysis using detergents and / or lysozyme and / or protease treatment and / or physical disruption of cells, such as sonication, and / or other combinations of alkaline lysis and / or hypotonic lysis. In some cases, the primary template or target molecule is subjected to a pretreatment step. In some cases, the primary template (or target) is denatured using sodium hydroxide, followed by neutralization of the solution. Other denaturation strategies may also be suitable for practicing the methods described herein. Such strategies include, but are not limited to, combining alkaline lysis with other basic solutions, increasing the temperature of the sample and / or changing the salt concentration in the sample, adding additives such as solvents or oils, other modifications, or any combination thereof. In some cases, additional steps include sorting, filtering, or separating the sample, template, or amplicon by size. For example, after amplification using the methods described herein, the amplicon library is enriched for amplicons having a desired length. In some cases, the amplicon library is enriched for amplicons having a length of 50-2000, 25-1000, 50-1000, 75-2000, 100-3000, 150-500, 75-250, 170-500, 100-500, or 75-2000 bases.In some cases, the amplicon library is enriched for amplicons having a length of 75, 100, 150, 200, 500, 750, 1000, 2000, 5000, or 10,000 bases or less. In some cases, the amplicon library is enriched for amplicons having a length of at least 25, 50, 75, 100, 150, 200, 500, 750, 1000, or at least 2000 bases. .
[0055] The methods and compositions described herein may include buffers or other formulations. Such buffers are, in some cases, used for PTA, RT, or other methods described herein. In some cases, such buffers include surfactants / detergents or denaturants (Tween-20, DMSO, DMF, PEGylated polymers containing hydrophobic groups, or other surfactants), salts (potassium phosphate or sodium phosphate (monobasic or dibasic), sodium chloride, potassium chloride, TrisHCl, magnesium chloride or sulfate, ammonium salts such as phosphate, nitrate, sulfate, EDTA), reducing agents (DTT, THP, DTE, beta-mercaptoethanol, TCEP, or other reducing agents), or other components (glycerol, hydrophilic polymers such as PEG). In some cases, buffers are used in combination with components such as polymerases, strand displacement factors, terminators, or other reaction components described herein. In some cases, buffers are used in combination with components such as polymerases, strand displacement factors, terminators, or other reaction components described herein. Buffers may include one or more crowding agents. In some cases, the crowding reagent includes a polymer. In some cases, the crowding reagent includes a polymer such as a polyol. In some cases, the crowding reagent includes a polyethylene glycol polymer (PEG). In some cases, the crowding reagent includes a polysaccharide. Non-limiting examples of crowding reagents include Ficoll (e.g., Ficoll PM400, Ficoll PM70, or other molecular weight Ficoll), PEG (e.g., PEG1000, PEG2000, PEG4000, PEG6000, PEG8000, or other molecular weight PEG), and dextran (dextran 6, dextran 10, dextran 40, dextran 70, dextran 6000, dextran 138k, or other molecular weight dextran).
[0056] Nucleic acid molecules amplified according to the methods described herein can be sequenced and analyzed using methods known to those skilled in the art. Non-limiting examples of sequencing methods used in some cases include, for example, sequencing by hybridization (SBH), sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads (U.S. Patent No. 7,425,431), wobble sequencing (International Patent Application Publication No. WO2006 / 073504), multiplex sequencing (U.S. Patent Application Publication No. US2008 / 0269068; Porreca et al., 2007, Nat. Methods 4:931), polymerized colony (POLONY) sequencing (U.S. Pat. Nos. 6,432,360, 6,485,944, and 6,511,803, and International Patent Application Publication No. WO 2005 / 082098), nanogrid rolling circle sequencing (ROLONY) (U.S. Pat. No. 9,624,538), allele-specific oligo ligation assays (e.g., oligo ligation assay (OLA), single template molecule OLA using ligated linear probes and rolling circle amplification (RCA) readout, ligated padlock probes, and / or single template molecule OLA using ligated circular padlock probes and rolling circle amplification (RCA) readout), e.g., Roche 454, Illumina These include high-throughput sequencing methods such as those using the Solexa, AB-SOLiD, Helicos, Polonator platforms, and light-based sequencing technologies (Landegrenetal. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmacogenomics 1:95-100; and Shi (2001) Clin. Chem. 47:164-172).In some cases, the amplified nucleic acid molecules are shotgun sequenced. Sequencing of the sequencing library is, in some cases, performed using a suitable sequencing technology, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis (array / colony-based or nanoball-based).
[0057] The sequencing library generated using the method described herein (e.g., PTA or RNAseq) can be sequenced to obtain a desired number of sequencing reads. In some cases, the library is generated from a single cell or a sample containing a single cell (either alone or as part of a multi-omics workflow). In some cases, the library is sequenced to obtain at least 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or at least 10 million reads. In some cases, the library is sequenced to obtain no more than 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or 10 million reads. In some cases, the library is sequenced to obtain about 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or about 10 million reads. In some cases, the library is sequenced to obtain 0.1-10, 0.1-5, 0.1-1, 0.2-1, 0.3-1.5, 0.5-1, 1-5, or 500-5 million reads per sample. In some cases, the number of reads depends on the size of the genome. In some cases, a sample containing a bacterial genome is sequenced to obtain 500-1 million reads. In some cases, the library is sequenced to obtain at least 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or at least 900 million reads. In some cases, the library is sequenced to obtain no more than 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or 900 million reads. In some cases, the library is sequenced to obtain about 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or about 900 million reads. In some cases, a sample comprising a mammalian genome is sequenced to obtain 500 to 600 million reads.In some cases, the type of sequencing library (cDNA library or genomic library) is identified during sequencing. In some cases, cDNA and genomic libraries are identified during sequencing using unique barcodes.
[0058] The term "cycle" when used in reference to a polymerase-mediated amplification reaction is used herein to describe the steps of dissociating at least a portion of the double-stranded nucleic acid (e.g., denaturing the template from the amplicon, or the double-stranded template), hybridizing (annealing) at least a portion of the primer to the template, and extending the primer to generate an amplicon. In some cases, the temperature remains constant during the cycle of amplification (such as an isothermal reaction). In some cases, the number of cycles directly correlates with the number of amplicons generated. In some cases, the number of cycles in an isothermal reaction is controlled by the length of time the reaction is allowed to proceed.
[0059] Methods and Applications Described herein are methods for identifying mutations in cells using multi-omic analysis, such as single-cell analysis, using PTA. The use of PTA, in some cases, provides an improvement over known methods, such as MDA. PTA, in some cases, results in lower false-positive and false-negative variant call rates than MDA. Genomes, such as the NA12878 Platinum genome, are used to determine whether greater genome coverage and uniformity of PTA result in lower false-negative variant call rates. Without being bound by theory, it may be determined that the lack of error propagation in PTA reduces the false-positive variant call rate. The amplification balance between alleles from the two methods is estimated in some cases by comparing the allele frequencies of heterozygous mutation calls at known positive loci. In some cases, the amplicon library generated using PTA is further amplified by PCR. In some cases, PTA is used in a workflow that employs additional analytical methods, such as RNA sequencing, methylome analysis, or other methods described herein.
[0060] In some cases, the cells analyzed using the methods described herein include tumor cells. For example, circulating tumor cells can be isolated from fluids collected from patients, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, or aqueous humor. The cells are then subjected to the methods described herein (e.g., PTA) and sequencing to determine the mutation load and combination of mutations in each cell. These data can be used in some cases as a tool for diagnosing specific diseases or predicting treatment response. Similarly, in some cases, cells of unknown malignant potential are isolated from bodily fluids collected from patients, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, aqueous humor, blastocyst fluid, or collection medium surrounding cells in culture. In some cases, samples are obtained from collection medium surrounding embryonic cells. After utilizing the methods described herein and sequencing, such methods can be further used to determine the mutation load and combination of mutations in each cell. These data are sometimes used to diagnose specific diseases or as a tool to predict the progression of premalignant states to overt malignancies. In some cases, cells can be isolated from primary tumor samples. Cells are then subjected to PCR and sequencing to determine the mutational burden and combination of mutations in each cell. In some cases, these data can be used to diagnose specific diseases or as a tool to predict the probability that a patient's malignancy will be resistant to available anticancer drugs. By exposing samples to different chemotherapeutic agents, it was found that major and minor clones have differential sensitivity to specific drugs, which does not necessarily correlate with the presence of known "driver mutations," suggesting that the combination of mutations within a clonal population determines its sensitivity to a particular chemotherapeutic drug. Without being bound by theory, these findings suggest that it may be easier to eradicate malignancies if premalignant lesions are detected that evolve into clones with an increasing number of genomic modifications that may be more resistant to treatment.See Ma et al., 2018, "Pan-cancer genome and transcriptome analyses of 1,699 pediatric leukemias and solid tumors." Single-cell genomics protocols are sometimes used to detect combinations of somatic genetic variants in single cancer cells, or clonotypes, within a mixture of normal and malignant cells isolated from patient samples. This technology is sometimes further utilized to identify clonotypes that undergo positive selection after drug exposure both in vitro and / or in patients. As shown in Figure 6A, by comparing surviving clones exposed to chemotherapy with clones identified at diagnosis, a catalog of cancer clonotypes documenting resistance to specific drugs can be created. PTA methods can sometimes detect the sensitivity of specific clones, as well as combinations of them, to existing or novel drugs within samples composed of multiple clonotypes, where the method can detect the sensitivity of specific clones to drugs. This approach can, in some cases, indicate the effectiveness of a drug against a specific clone that may not be detected using current drug sensitivity measures that consider the sensitivity of all cancer clones together in a single measurement. When the PTA described herein is applied to patient samples collected at the time of diagnosis to detect cancer clonotypes in a given patient's cancer, it searches for those clones using a catalog of drug sensitivities, thereby providing oncologists with information on which drugs or drug combinations will not work and which drugs or drug combinations are likely to be most effective against that patient's cancer. PTA can be used to analyze samples containing groups of cells. In some cases, the sample contains neurons or glial cells. In some cases, the sample contains nuclei.
[0061] Described herein are methods for measuring changes in gene expression in combination with the mutagenicity of environmental factors. For example, cells (single or population) are exposed to potential environmental conditions. For example, cells derived from organs (liver, pancreas, lung, colon, thyroid, or other organs), tissues (skin, or other tissues), blood, or other biological sources are used in some cases. In some cases, the environmental conditions include heat, light (e.g., ultraviolet light), radiation, chemicals, or any combination thereof. In some cases, after exposure to the environmental conditions for minutes, hours, days, or longer, single cells are isolated and subjected to PTA. In some cases, samples are tagged using molecular barcodes and unique molecular identifiers. The samples are sequenced and then analyzed to identify changes in gene expression resulting from mutations resulting from exposure to the environmental conditions. In some cases, such mutations are compared to a control environmental condition, such as a known non-mutagenic agent, vehicle / solvent, or the absence of the environmental condition. Such analyses, in some cases, provide not only the total number of mutations caused by an environmental condition, but also the location and nature of those mutations. Patterns can, in some cases, be identified from the data and used to diagnose a disease or condition. In some cases, patterns can be used to predict future disease states or conditions. In some cases, the methods described herein measure the mutation load, location, and pattern in cells after exposure to an environmental agent, such as a potential mutagen or teratogen. This approach, in some cases, is used to evaluate the safety of a given drug, including its potential to induce mutations that may contribute to disease development. For example, this method can be used to predict the carcinogenic or teratogenic potential of a drug for a particular cell type after exposure to a particular concentration of the drug.
[0062] Described herein is a method for identifying gene expression changes combined with mutations in animal, plant, or microbial cells that have undergone genome editing (e.g., using CRISPR technology). In some cases, such cells can be isolated and subjected to PCR and sequencing to determine the mutation load and combination of mutations in each cell. The mutation rate and location of mutations per cell resulting from a genome editing protocol can, in some cases, be used to evaluate the safety of a given genome editing method.
[0063] Described herein are methods for determining changes in gene expression in combination with mutations in cells used for cell therapy, such as, but not limited to, transplantation of induced pluripotent stem cells, transplantation of unmanipulated hematopoietic or other cells, or transplantation of hematopoietic or other cells that have undergone genome editing. The cells can then undergo PTA and sequencing to determine the mutation load and combination of mutations in each cell. The mutation rate and location of mutations per cell in a cell therapy product can be used to assess the safety and potential efficacy of the product.
[0064] Cells for use with the PTA method can be fetal cells, such as embryonic cells. In some embodiments, PTA is used in conjunction with non-invasive preimplantation genetic testing (NIPGT). In further embodiments, cells can be isolated from blastomeres produced by in vitro fertilization. The cells can then undergo PTA and sequencing to determine the burden and combination of potential disease-predisposing genetic mutations in each cell. Changes in gene expression combined with the mutation profile of the cells can then be used to extrapolate the genetic predisposition of the blastomere to specific diseases before implantation. In some cases, embryos in culture release nucleic acids that are used to assess the health of the embryo using low-pass genomic sequencing. In some cases, the embryos are frozen and thawed. In some cases, nucleic acids are obtained from blastocyst culture conditioned medium (BCCM), blastocoelic fluid (BF), or a combination thereof. In some cases, PTA analysis of fetal cells is used to detect chromosomal abnormalities, such as fetal aneuploidy. In some cases, PTA is used to detect diseases such as Down syndrome or Patau syndrome. In some cases, frozen blastocysts are thawed and cultured for a period of time before obtaining nucleic acid for analysis (e.g., medium, BF, or cell biopsy). In some cases, the blastocysts are cultured for no more than 4, 6, 8, 12, 16, 24, 36, 48, or 64 hours before obtaining nucleic acid for analysis.
[0065] In another embodiment, microbial cells (e.g., bacteria, fungi, protozoa) can be isolated from plants or animals (e.g., from microbiota samples (e.g., GI microbiota, skin microbiota, etc.) or from bodily fluids such as, for example, blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural fluid, pericardial fluid, ascites, or aqueous humor). Additionally, microbial cells may be isolated from indwelling medical devices such as, but not limited to, intravenous catheters, urinary catheters, cerebrospinal shunts, artificial valves, artificial joints, or endotracheal tubes. Cells can then undergo PTA and sequencing to determine the identity of specific microorganisms and to detect the presence of microbial genetic variants that predict response (or resistance) to specific antimicrobial agents. These data can be used for the diagnosis of specific infections and / or as a tool for predicting treatment response.
[0066] Described herein are methods for generating amplicon libraries from samples containing short nucleic acids using the PTA methods described herein. In some cases, PTA results in improved fidelity and uniformity of amplification of shorter nucleic acids. In some cases, the nucleic acids are 2,000 bases or less in length. In some cases, the nucleic acids are 1,000 bases or less in length. In some cases, the nucleic acids are 500 bases or less in length. In some cases, the nucleic acids are 200, 400, 750, 1,000, 2,000, or 5,000 bases or less in length. In some cases, samples containing short nucleic acid fragments include, but are not limited to, ancient DNA (hundreds, thousands, millions, or even billions of years old), FFPE (formalin-fixed, paraffin-embedded) samples, cell-free DNA, or other samples containing short nucleic acids.
[0067] Embodiment Described herein are methods for amplifying a target nucleic acid molecule, the methods comprising: a) contacting a sample containing the target nucleic acid molecule, one or more amplification primers, a nucleic acid polymerase, and a mixture of nucleotides including one or more terminator nucleotides that terminate nucleic acid replication by the polymerase; and b) incubating the sample under conditions that promote replication of the target nucleic acid molecule to obtain a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product between about 50 and about 2000 nucleotides in length. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product between about 400 and about 600 nucleotides in length. In one embodiment of any of the above methods, the method further comprises c) repairing the ends and A-tailing, and d) ligating the molecules obtained in step (c) to adapters, thereby generating a library of amplification products. In some embodiments, the method further comprises removing terminator nucleotides from the terminated amplification product. In one embodiment of any of the above methods, the method further comprises sequencing the amplification product. In one embodiment of any of the above methods, the amplification is carried out under substantially isothermal conditions. In one embodiment of any of the above methods, the nucleic acid polymerase is a DNA polymerase.
[0068] In one embodiment of any of the above methods, the DNA polymerase is a strand-displacing DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase is selected from the group consisting of bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R The nucleic acid polymerase is selected from (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase has 3' to 5' exonuclease activity, and the terminator nucleotide inhibits such 3' to 5' exonuclease activity. In one specific embodiment, the terminator nucleotide is selected from nucleotides with modifications at the alpha group (e.g., alpha-thiodideoxynucleotides), C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2' fluoronucleotides, 3' phosphorylated nucleotides, 2'-O-methyl modified nucleotides, and trans nucleic acids. In one embodiment of any of the above methods, the nucleic acid polymerase does not have 3' to 5' exonuclease activity. In one particular embodiment, the polymerase is Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent RThe terminator nucleotide is selected from (exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, and Therminator DNA polymerase. In one specific embodiment, the terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. In one specific embodiment, the terminator nucleotide is selected from a 3'-blocked reversible terminator comprising a nucleotide, a 3'-unblocked reversible terminator comprising a nucleotide, a terminator comprising a 2'-modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one specific embodiment, the terminator nucleotide is selected from a dideoxynucleotide, a reverse dideoxynucleotide, a 3'-biotinylated nucleotide, a 3'-amino nucleotide, a 3'-phosphorylated nucleotide, a 3'-O-methyl nucleotide, a 3'-C3 spacer nucleotide, a 3'-C18 nucleotide, a 3'-carbon spacer nucleotide comprising a 3'-hexanediol spacer nucleotide, an acyclonucleotide, and combinations thereof. In one embodiment of any of the above methods, the amplification primers are 4 to 70 nucleotides in length. In one embodiment of any of the above methods, the amplification product is about 50 to about 2000 nucleotides in length. In one embodiment of any of the above methods, the target nucleic acid is DNA (e.g., cDNA or genomic DNA). In one embodiment of any of the above methods, the amplification primers are random primers. In one embodiment of any of the above methods, the amplification primers comprise a barcode. In one particular embodiment, the barcode comprises a cell barcode. In one particular embodiment, the barcode comprises a sample barcode. In one embodiment of any of the above methods, the amplification primers comprise a unique molecular identifier (UMI). In one embodiment of any of the above methods, the method comprises denaturing the target nucleic acid or genomic DNA prior to initial primer annealing. In one particular embodiment, denaturation is performed under alkaline conditions followed by neutralization.In one embodiment of any of the above methods, the sample, amplification primers, nucleic acid polymerase, and nucleotide mixture are contained in a microfluidic device. In one embodiment of any of the above methods, the sample, amplification primers, nucleic acid polymerase, and nucleotide mixture are contained in a droplet. In one embodiment of any of the above methods, the sample is a tissue sample, a cell, a biological fluid sample (e.g., blood, urine, saliva, lymph, cerebrospinal fluid (CSF), amniotic fluid, pleural fluid, pericardial fluid, peritoneal fluid, aqueous humor), a bone marrow sample, a semen sample, a biopsy sample, a cancer sample, a tumor sample, a cell lysate sample, a forensic sample, an archaeological sample, a paleontological sample, an infection sample, a production sample, a whole plant, a plant part, a microbiota sample, a virus preparation, a soil sample, a marine sample, a freshwater sample, a household or industrial sample, and combinations and isolates thereof. In one embodiment of any of the above methods, the sample is a cell (e.g., an animal cell [e.g., a human cell], a plant cell, a fungal cell, a bacterial cell, or a protozoan cell). In a particular embodiment, the cells are lysed prior to replication. In a particular embodiment, cell lysis involves proteolysis. In a particular embodiment, the cells are cells from preimplantation embryos, stem cells, fetal cells, tumor cells, suspected cancer cells, cancer cells, cells subjected to gene editing procedures, cells from pathogenic organisms, cells obtained from forensic samples, cells obtained from archaeological samples, and cells obtained from paleontological samples. In one embodiment of any of the above methods, the sample is a cell from a preimplantation embryo (e.g., a blastomere (e.g., a blastomere obtained from an 8-cell embryo generated by in vitro fertilization)). In a particular embodiment, the method further comprises determining the presence of a disease-predisposing germline or somatic mutation in the embryonic cells. In one embodiment of any of the above methods, the sample is a cell from a pathogenic organism (e.g., a bacterium, fungus, protozoan).In one particular embodiment, the pathogenic organism cells are obtained from fluid obtained from a patient, a microbiota sample (e.g., a GI microbiota sample, a vaginal microbiota sample, a skin microbiota sample, etc.), or from an indwelling medical device (e.g., an intravenous catheter, a urinary catheter, a cerebrospinal shunt, an artificial valve, an artificial joint, an endotracheal tube, etc.). In one particular embodiment, the method further comprises determining the identity of the pathogenic organism. In one particular embodiment, the method further comprises determining the presence of a genetic variant responsible for the pathogenic organism's resistance to the treatment. In one embodiment of any of the above methods, the sample is a tumor cell, a cell suspected of cancer, or a cancer cell. In one particular embodiment, the method further comprises determining the presence of one or more diagnostic or prognostic mutations. In one particular embodiment, the method further comprises determining the presence of a germline or somatic variant responsible for resistance to the treatment. In one embodiment of any of the above methods, the sample is a cell to be subjected to a gene editing procedure. In one particular embodiment, the method further comprises determining the presence of unplanned mutations caused by the gene editing process. In one embodiment of any of the above methods, the method further comprises determining the history of the cell lineage. In a related aspect, the invention provides the use of any of the above methods to identify low frequency sequence variants (e.g., variants that constitute ≧0.01% of the total sequence).
[0069] In a related aspect, the present invention provides a kit comprising a nucleic acid polymerase, one or more amplification primers, a mixture of nucleotides including one or more terminator nucleotides, and optionally instructions for use. In one embodiment of the kit of the present invention, the nucleic acid polymerase is a strand-displacing DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase is selected from the group consisting of bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R The nucleic acid polymerase is selected from (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase has 3' to 5' exonuclease activity, and the terminator nucleotide inhibits such 3' to 5' exonuclease activity (e.g., nucleotides with modifications at the alpha group [e.g., alpha-thiodideoxynucleotides], C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2' fluoronucleotides, 3' phosphorylated nucleotides, 2'-O-methyl modified nucleotides, trans nucleic acids). In one embodiment of the kit of the present invention, the nucleic acid polymerase does not have 3' to 5' exonuclease activity (e.g., Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, Therminator DNA polymerase). In one specific embodiment, the terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. In one specific embodiment, the terminator nucleotide is selected from a 3'-blocked reversible terminator comprising nucleotides, a 3'-unblocked reversible terminator comprising nucleotides, a terminator comprising a 2'-modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one specific embodiment, the terminator nucleotide is selected from a dideoxynucleotide, a reversed dideoxynucleotide, a 3'-biotinylated nucleotide, a 3'-amino nucleotide, a 3'-phosphorylated nucleotide, a 3'-O-methyl nucleotide, a 3'-C3 spacer nucleotide, a 3'-C18 nucleotide, a 3'-carbon spacer nucleotide comprising a 3'-hexanediol spacer nucleotide, an acyclonucleotide, and combinations thereof.
[0070] Described herein are methods for amplifying a genome, comprising: a) a genome, a plurality of amplification primers (e.g., two or more primers), a nucleic acid polymerase, and nucleotides containing one or more terminator nucleotides that terminate nucleic acid replication by the polymerase; and b) incubating the sample under conditions that promote replication of the genome to obtain a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product that is between about 50 and about 2000 nucleotides in length. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product that is between about 400 and about 600 nucleotides in length. In one embodiment of any of the above methods, the method further comprises c) repairing the ends and A-tailing, and d) ligating the molecules obtained in step (c) to adapters, thereby generating a library of amplification products. In one embodiment of any of the above methods, the method further comprises sequencing the amplification products. In one embodiment of any of the above methods, the amplification is carried out under substantially isothermal conditions. In one embodiment of any of the above methods, the nucleic acid polymerase is a DNA polymerase.
[0071] In one embodiment of the above method, the DNA polymerase is a strand-displacing DNA polymerase. In one embodiment of the above method, the nucleic acid polymerase is selected from the group consisting of bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent RThe nucleic acid polymerase is selected from (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase has 3' to 5' exonuclease activity, and the terminator nucleotide inhibits such 3' to 5' exonuclease activity. In one specific embodiment, the terminator nucleotide is selected from a nucleotide having a modification in the alpha group (e.g., an alpha-thiodideoxynucleotide that creates a phosphorothioate bond), a C3 spacer nucleotide, a locked nucleic acid (LNA), an inverted nucleic acid, a 2' fluoronucleotide, a 3' phosphorylated nucleotide, a 2'-O-methyl modified nucleotide, and a trans nucleic acid. In one embodiment of any one of the above methods, the nucleic acid polymerase does not have 3' to 5' exonuclease activity. In one particular embodiment, the polymerase is selected from the group consisting of Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent RThe terminator nucleotide is selected from (exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, and Therminator DNA polymerase. In one specific embodiment, the terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. In one specific embodiment, the terminator nucleotide is selected from a 3'-blocked reversible terminator comprising a nucleotide, a 3'-unblocked reversible terminator comprising a nucleotide, a terminator comprising a 2'-modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one specific embodiment, the terminator nucleotide is selected from a dideoxynucleotide, a reverse dideoxynucleotide, a 3'-biotinylated nucleotide, a 3'-amino nucleotide, a 3'-phosphorylated nucleotide, a 3'-O-methyl nucleotide, a 3'-C3 spacer nucleotide, a 3'-C18 nucleotide, a 3'-carbon spacer nucleotide comprising a 3'-hexanediol spacer nucleotide, an acyclonucleotide, and combinations thereof. In one embodiment of any of the above methods, the amplification primers are 4 to 70 nucleotides in length. In one embodiment of any of the above methods, the amplification product is about 50 to about 2000 nucleotides in length. In one embodiment of any of the above methods, the target nucleic acid is DNA (e.g., cDNA or genomic DNA). In one embodiment of any of the above methods, the amplification primers are random primers. In one embodiment of any of the above methods, the amplification primers comprise a barcode. In one particular embodiment, the barcode comprises a cell barcode. In one particular embodiment, the barcode comprises a sample barcode. In one embodiment of any of the above methods, the amplification primers comprise a unique molecular identifier (UMI). In one embodiment of any of the above methods, the method comprises denaturing the target nucleic acid or genomic DNA prior to initial primer annealing. In one particular embodiment, denaturation is performed under alkaline conditions followed by neutralization.In one embodiment of any of the above methods, the sample, amplification primers, nucleic acid polymerase, and nucleotide mixture are contained in a microfluidic device. In one embodiment of any of the above methods, the sample, amplification primers, nucleic acid polymerase, and nucleotide mixture are contained in a droplet. In one embodiment of any of the above methods, the sample is selected from a tissue sample, a cell, a body fluid sample (e.g., blood, urine, saliva, lymph, cerebrospinal fluid (CSF), amniotic fluid, pleural fluid, pericardial fluid, peritoneal fluid, aqueous humor), a bone marrow sample, a semen sample, a biopsy sample, a cancer sample, a tumor sample, a cell lysate sample, a forensic sample, an archaeological sample, a paleontological sample, an infection sample, a production sample, a whole plant, a plant part, a microbiota sample, a virus preparation, a soil sample, a seawater sample, a freshwater sample, a household or industrial sample, and combinations and isolates thereof. In one embodiment of any of the above methods, the sample is a cell (e.g., an animal cell [e.g., a human cell], a plant cell, a fungal cell, a bacterial cell, or a protozoan cell). In a particular embodiment, the cells are lysed prior to replication. In a particular embodiment, cell lysis involves proteolysis. In a particular embodiment, the cells are selected from cells from a preimplantation embryo, stem cells, fetal cells, tumor cells, suspected cancer cells, cancer cells, cells that have undergone a gene editing procedure, cells from a pathogen, cells obtained from a forensic sample, cells obtained from an archaeological sample, and cells obtained from a paleontological sample. In one embodiment of any of the above methods, the sample is a cell of a preimplantation embryo (e.g., a blastomere (e.g., a blastomere obtained from an 8-cell stage embryo generated by in vitro fertilization)). In one particular embodiment, the method further comprises determining the presence of a disease-predisposing germline or somatic variant in the embryonic cell. In one embodiment of any of the above methods, the sample is a cell of a pathogen (e.g., a bacterium, fungus, protozoan).In one particular embodiment, the pathogen cells are obtained from bodily fluids collected from a patient, a microbiota sample (e.g., a GI microbiota sample, a vaginal microbiota sample, a skin microbiota sample, etc.), or an indwelling medical device (e.g., an intravenous catheter, a urinary catheter, a cerebrospinal shunt, an artificial valve, an artificial joint, an endotracheal tube, etc.). In one particular embodiment, the method further comprises determining the identity of the pathogen. In one particular embodiment, the method further comprises determining the presence of a genetic variant that causes resistance of the pathogen to the treatment. In one embodiment of any of the above methods, the sample is a tumor cell, a suspected cancer cell, or a cancer cell. In one particular embodiment, the method further comprises determining the presence of one or more diagnostic or prognostic mutations. In one particular embodiment, the method further comprises determining the presence of a germline or somatic variant that causes resistance to the treatment. In one embodiment of any of the above methods, the sample is a cell that has undergone a gene editing procedure. In one particular embodiment, the method further comprises determining the presence of an unplanned mutation caused by the gene editing process. In one embodiment of any of the above methods, the method further comprises determining the lineage history of the cell. In a related aspect, the invention provides use of any of the above methods to identify low frequency sequence variants (e.g., variants that constitute > 0.01% of the total sequence).
[0072] In a related aspect, the present invention provides a kit comprising a reverse transcriptase, a nucleic acid polymerase, one or more amplification primers, a mixture of nucleotides including one or more terminator nucleotides, and optionally instructions for use. In one embodiment of the kit of the present invention, the nucleic acid polymerase is a strand-displacing DNA polymerase. In some cases, the reverse transcriptase performs template switching. In some cases, the reverse transcriptase is a variant of MMLV (Moloney murine leukemia virus), HIV-1, AMV (avian myeloblastosis virus), telomerase RT, FIV (feline immunodeficiency virus), or XMRV (xenotropic murine leukemia virus-related virus). Non-limiting examples of reverse transcriptases include SuperScript I (Thermo), SuperScript II (Thermo), SuperScript III (Thermo), SuperScript IV (Thermo), OmniScript (Qiagen), SensiScript (Qiagen), PrimeScript (Takara), MaximaH- (Thermo), AcuuScript Hi-Fi (Agilent), iScript (Bio-Rad), eAMV (Merck KGaA), qScript (Quanta Biosciences), SmartScribe (Clontech), or GoScript (Promega). In one embodiment of the kit of the present invention, the nucleic acid polymerase is selected from the group consisting of bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent RThe nucleic acid polymerase is selected from (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase has 3' to 5' exonuclease activity, and the terminator nucleotide inhibits such 3' to 5' exonuclease activity (e.g., nucleotides with modifications at the alpha group [e.g., alpha-thiodideoxynucleotides], C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2' fluoronucleotides, 3' phosphorylated nucleotides, 2'-O-methyl modified nucleotides, trans nucleic acids). In one embodiment of the kit of the present invention, the nucleic acid polymerase does not have 3' to 5' exonuclease activity (e.g., Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, Therminator DNA polymerase). In one specific embodiment, the terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. In one specific embodiment, the terminator nucleotide is selected from a 3'-blocked reversible terminator comprising nucleotides, a 3'-unblocked reversible terminator comprising nucleotides, a terminator comprising a 2'-modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one specific embodiment, the terminator nucleotide is selected from a dideoxynucleotide, a reversed dideoxynucleotide, a 3'-biotinylated nucleotide, a 3'-amino nucleotide, a 3'-phosphorylated nucleotide, a 3'-O-methyl nucleotide, a 3'-C3 spacer nucleotide, a 3'-C18 nucleotide, a 3'-carbon spacer nucleotide comprising a 3'-hexanediol spacer nucleotide, an acyclonucleotide, and combinations thereof. In some cases, the kit includes at least one enzyme stabilizer, a neutralization buffer, a denaturation buffer, or a combination thereof. In some cases, the kit includes one or more modules. In some cases, the kit includes a genome module and a transcriptome module.
[0073] Numbered Embodiments Described herein are the following numbered embodiments 1-46. 1. Described herein are embodiments of methods for multi-omic single-cell analysis, the methods including: a. isolating a single cell from a population of cells; b. sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from the cell; and c. sequencing the genome of the cell, wherein sequencing the genome of the cell includes: i. providing a genome from the single cell; ii. contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; and iii. amplifying at least some of the genome to produce a plurality of terminated amplicons, wherein replication proceeds by strand displacement replication; iv. ligating molecules obtained in step (iii) to adapters, thereby producing a genomic DNA library; and v. sequencing the genomic DNA library. 2. Further provided herein is the method of embodiment 1, wherein the method further comprises identifying at least one protein on the surface of the cells. 3. Further provided herein is the method of embodiment 1, wherein the mRNA transcripts comprise polyadenylated mRNA transcripts. 4. Further provided herein is the method of embodiment 1, wherein the mRNA transcripts do not comprise polyadenylated mRNA transcripts. 5. Further provided herein is the method of any one of embodiments 1-4, wherein sequencing the cDNA library comprises amplifying mRNA transcripts using template switching primers. 6. Further provided herein is the method of any one of embodiments 1-4, wherein at least some of the polynucleotides in the cDNA library comprise barcodes.7. Further provided herein is the method of any one of embodiments 1-4, wherein at least some of the polynucleotides of the cDNA library comprise at least two barcodes. 8. Further provided herein is the method of embodiment 6 or 7, wherein the barcodes comprise cell barcodes. 9. Further provided herein is the method of embodiment 6 or 7, wherein the barcodes comprise sample barcodes. 10. A method for multi-omic single-cell analysis comprising: a. isolating a single cell from a population of cells; b. identifying at least one protein on the surface of the cell; and c. sequencing the genome of the cell, wherein sequencing the genome of the cell comprises: i. providing a genome from the single cell; ii. contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; iii. amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iv. ligating molecules obtained in step (iii) to adapters, thereby generating a genomic DNA library; and v. sequencing the genomic DNA library. 11. Further provided herein is the method of embodiment 10, wherein identifying at least one protein on the surface of the cell comprises contacting the cell with a labeled antibody that binds to the at least one protein. 12. Further provided herein is the method of embodiment 11, wherein the labeled antibody comprises at least one fluorescent label. 13. Further provided herein is the method of embodiment 11, wherein the labeled antibody comprises at least one mass tag. 14. Further provided herein is the method of embodiment 11, wherein the labeled antibody comprises at least one nucleic acid barcode.15. A method for multi-omic single cell analysis, the method comprising: a. isolating a single cell from a population of cells; b. sequencing the genome of said cell, wherein sequencing the genome of said cell comprises: i. providing a genome from the single cell; ii. digesting said genome with a methylation-sensitive restriction enzyme to generate genomic fragments; iii. contacting at least some of said genomic fragments with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein said mixture of nucleotides contains at least one terminator that terminates nucleic acid replication by said polymerase. iv. amplifying at least some of the genome fragments to generate a plurality of termination amplification products, wherein replication proceeds by strand displacement replication; v. amplifying at least some of the genome fragments with methylation-specific PCR; vi. ligating the molecules obtained in steps (iv) and (v) to adapters, thereby generating a genomic DNA library and a methylome DNA library; and vii. sequencing the genomic DNA library and the methylome DNA library. 16. Further provided herein is the method of embodiment 15, wherein the step of identifying at least one protein on the surface of the cells comprises contacting the cells with a labeled antibody that binds to at least one protein. 17. Further provided herein is the method of embodiment 16, wherein the labeled antibody comprises at least one fluorescent label. 18. Further provided herein is the method of embodiment 16, wherein the labeled antibody comprises at least one mass tag. 19. Further provided herein is the method of embodiment 16, wherein the labeled antibody comprises at least one nucleic acid barcode. 20. Further provided herein is the method of any one of embodiments 1-19, wherein the single cell is a mammalian cell. 21. Further provided herein is the method of any one of embodiments 1-19, wherein the single cell is a human cell.22. Further provided herein is the method of any one of embodiments 1 to 19, wherein the single cell is derived from liver, skin, kidney, blood, or lung. 23. Further provided herein is the method of any one of embodiments 1 to 19, wherein the single cell is a primary cell. 24. Further provided herein is the method of any one of embodiments 1 to 23, wherein the method further comprises removing at least one terminator nucleotide from the terminated amplification products. 25. Further provided herein is the method of any one of embodiments 1 to 23, wherein at least some of the amplification products comprise a barcode. 26. Further provided herein is the method of any one of embodiments 1 to 23, wherein at least some of the amplification products comprise at least two barcodes. 27. Further provided herein is the method of embodiment 24 or 26, wherein the barcode comprises a cell barcode. 28. Further provided herein is the method of embodiment 24 or 26, wherein the barcode comprises a sample barcode. 29. Further provided herein is the method of any one of embodiments 1-28, wherein at least some of the amplification primers comprise a unique molecular identifier (UMI). 30. Further provided herein is the method of any one of embodiments 1-28, wherein at least some of the amplification primers comprise at least two unique molecular identifiers (UMI). 31. Further provided herein is the method of any one of embodiments 1-30, wherein the method further comprises an additional amplification step using PCR. 32. Further provided herein is the method of any one of embodiments 1-30, wherein at least one mutation is identified in the genome of the cell, and wherein the mutation differs from the corresponding position in the reference sequence. 33. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in less than 50% of the population of cells. 34. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in less than 25% of the population of cells.35. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in less than 1% of the population of cells. 36. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.1% or less of the population of cells. 37. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.01% or less of the population of cells. 38. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.001% or less of the population of cells. 39. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.0001% or less of the population of cells. 40. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 50% or less of the amplification product sequences. 41. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 25% or less of the amplification product sequences. 42. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 1% or less of the amplified product sequences. 43. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.1% or less of the amplified product sequences. 44. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.01% or less of the amplified product sequences. 45. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.001% or less of the amplified product sequences. 46. Further provided herein is the method of embodiment 32, wherein the at least one mutation occurs in 0.0001% or less of the amplified product sequences. [Example]
[0074] The following examples are presented to more clearly illustrate to those skilled in the art the principles and practice of the embodiments disclosed herein, and should not be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are by weight.
[0075] Example 1: Primary template-directed amplification (PTA) Although PTA can be used for any nucleic acid amplification, it is particularly useful for whole genome amplification because it allows for capturing a greater percentage of the cellular genome in a more uniform and reproducible manner and with a lower error rate than currently used methods, such as multiple displacement amplification (MDA), while avoiding the drawbacks of currently used methods, such as exponential amplification where the polymerase initially extends random primers, which results in random overrepresentation of loci and alleles and the propagation of mutations (see Figure 1G). PTA can also be used in conjunction with other analytical techniques, such as transcriptome analysis.
[0076] cell culture Human NA12878 (Coriell Institute) cells were maintained in RPMI medium supplemented with 15% FBS and 2 mM L-glutamine, as well as 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B (Gibco, Life Technologies). Cells were cultured at 3.5 × 10 5 Cells were seeded at a density of 1000 cells / ml. Cultures were split every 3 days and maintained in a humidified incubator at 37°C with 5% CO2.
[0077] Single cell isolation and WTA The general protocol for WTA (whole transcriptome analysis) is shown in Figure 2F. Cells were resuspended at a concentration of 150–500 cells / µL. This cell suspension was stained with 20 µL of freshly prepared staining buffer (2.5 µL ethidium homodimer-1 and 0.625 µL calcein AM from LifeTechnology's LIVE / DEAD® Viability / Cytotoxicity Kit added to 1.25 mL of cell buffer containing 1x PBS and 0.05% Tween-20). Cells were then sorted using a FACS Aria III sorting instrument and deposited into each of 96 wells. A reaction mix containing 5x RT buffer, PEG 4000, RT primer (100 μM), TS oligo (20 μM), reverse transcriptase, RNAse inhibitor, gelatin, Tween-20, Triton-X, dNTP mix, TMAC (1 M), betaine (5 M), MgCl2 (50 mM), and ERCC spike was added to each well. The samples were then placed in a thermal cycler for 90 minutes at 42°C and 30 minutes at 50°C. The samples were then held at 4°C until ready for preamplification. After RT thermal cycling, the samples were processed for DNA amplification or for preamplification of the first-strand cDNA resulting from the RT reaction. Sample preamplification was achieved using a single primer (semi-suppressive PCR) using the following protocol to amplify the cDNA product. Briefly, 5 μL of the RT reaction was added to a 30 microliter reaction containing 2x master mix, 1 micromolar primers, and 5x preamplification buffer, and thermal cycling conditions were used: 95°C for 1 minute, followed by 21 cycles of 95°C for 15 seconds, 60°C for 30 seconds, and 68°C for 4 minutes, followed by a 10-minute hold at 72°C. Samples were then converted into sequencing libraries using the Nextera XT Library Prep Kit using the manufacturer's instructions (Figure 2G). Results of the RT experiment are shown in Table 1 for six samples.
[0078] [Table 1]
[0079] Single cell isolation and WGA 3.5×10 5 After seeding at a density of 1000 cells / ml, NA12878 cells were cultured for a minimum of 3 days, after which 3 mL of cell suspension was pelleted at 300 × g for 10 min. The medium was then discarded, and the cells were washed with 1 mL of cell wash buffer (Mg 2+ or Ca 2+ The cells were washed three times with 1x PBS containing 2% FBS (without ATP) by spinning at 300 x g, 200 x g, and finally 100 x g for 5 minutes. The cells were then resuspended in 500 μL of cell wash buffer. Following this, they were stained with 100 nM calcein AM (Molecular Probes) and 100 ng / ml propidium iodide (PI; Sigma-Aldrich) to differentiate the live cell population. The cells were loaded onto a BD FACScan flow cytometer (FACSAria II) (BD Biosciences) thoroughly washed with ELIMINase (Decon Labs) and calibrated using Accudrop fluorescent beads (BD Biosciences) for cell sorting. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS containing 0.2% Tween 20, followed by PTA (Sigma-Aldrich). Several wells were intentionally left empty to serve as no template controls (NTCs). Immediately after sorting, plates were briefly centrifuged and placed on ice. Cells were then frozen at -20°C for a minimum of overnight. The following day, WGA reactions were assembled in a pre-PCR workstation that provided constant positive pressure of HEPA-filtered air and was decontaminated with UV light for 30 minutes before each experiment.
[0080] MDA was performed using a modification previously shown to improve amplification uniformity. Specifically, exonuclease-resistant random primers (ThermoFisher) were added to the lysis buffer / mix to a final concentration of 125 μM. 4 μL of the resulting lysis / denaturation mix was added to the tube containing the single cells, vortexed, briefly spun, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, vortexed, briefly centrifuged, and placed at room temperature. 40 μL of amplification mix was then added, followed by incubation at 30°C for 8 hours, followed by heating to 65°C for 3 minutes to terminate the amplification.
[0081] PTA was performed by first further lysing the cells after freeze-thawing by adding 2 μl of a pre-chilled solution of a 1:1 mixture of 5% Triton X-100 (Sigma-Aldrich) and 20 mg / ml proteinase K (Promega). The cells were then vortexed, briefly centrifuged, and placed at 40°C for 10 minutes. Next, 4 μl of lysis buffer / mix and 1 μl of 500 μM exonuclease-resistant random primers were added to the lysed cells to denature the DNA, followed by vortexing, rotation, and placement at 65°C for 15 minutes. Next, 4 μl of room-temperature quenching buffer was added, and the sample was vortexed and spun down. 56 μl of amplification mix (primers, dNTPs, polymerase, buffer) contained equal ratios of alpha-thio-ddNTPs at a concentration of 1200 μM in the final amplification reaction. The sample was then placed at 30°C for 8 hours, after which amplification was stopped by heating at 65°C for 3 minutes.
[0082] After the amplification step, DNA from both the MDA and PTA reactions was purified using AMPure XP magnetic beads (Beckman Coulter) at a 2:1 ratio of beads to sample, and yields were measured using a Qubit dsDNA HS assay kit according to the manufacturer's instructions (Life Technologies) using a Qubit 3.0 fluorometer.
[0083] Library preparation The MDA reaction yielded 40 μg of amplified DNA. 1 μg of product was fragmented for 30 minutes according to standard procedures. The sample then underwent standard library preparation using 15 μM dual-index adapters (T4 polymerase, T4 polynucleotide kinase, and end repair with Taq polymerase for A-tailing) and four cycles of PCR. Each PTA reaction generated 40–60 ng of material that was used in its entirety for standard DNA sequencing library preparation without fragmentation. 2.5 μM adapters with UMI and dual indexes were used for ligation with T4 ligase, and 15 cycles of PCR (hot-start polymerase) were used in the final amplification. The library was then cleaned up using two-sided SPRI using ratios of 0.65× and 0.55× for right- and left-hand selection, respectively. The final library was quantified using the Qubit dsDNABR Assay Kit and the 2100 Bioanalyzer (Agilent Technologies) and subsequently sequenced on the Illumina NextSeq platform. All Illumina sequencing platforms, including NovaSeq, are also compatible with this protocol.
[0084] Data analysis Sequencing reads were demultiplexed based on cell barcodes using Bcl2fastq. Reads were then trimmed using trimmomatic and aligned to hg19 using BWA. Reads were duplicate-marked using Picard, followed by local realignment and base recalibration using GATK4.0. All files used to calculate quality metrics were downsampled to 20 million reads using Picard DownSampleSam. Quality metrics were obtained from the final bam files using qualimap, PicardAlignmentSummaryMetrics, and CollectWgsMetrics. Total genome coverage was also estimated using Preseq.
[0085] Variant calling Single nucleotide variants and indels were called using the GATK Unified Genotyper from GATK 4.0. Standard filtering criteria using GATK best practices were used throughout the process (https: / / software.broadinstitute.org / gatk / best-practices / ). Copy number variants were called using Control-FREEC (Boeva et al., Bioinformatics, 2012, 28(3):423-5). Structural variants were also detected using CREST (Wang et al., Nat Methods, 2011, 8(8):652-4).
[0086] result As shown in Figure 3A and Figure 3B, the mapping rate and mapping quality score for amplification using only dideoxynucleotides ("reversible") are 15.0 + / - 2.2 and 0.8 + / - 0.08, respectively, while incorporation of exonuclease-resistant alpha-thiodideoxynucleotide terminators ("irreversible") yields mapping rates and quality scores of 97.9 + / - 0.62 and 46.3 + / - 3.18, respectively. Experiments were also performed using reversible ddNTPs and various concentrations of terminators (Figure 2A, bottom).
[0087] Figures 2B–2E show comparative data generated from NA12878 human single cells subjected to MDA (according to Dong, X. et al., Nat Methods. 2017, 14(5):491–493) or PTA. While both protocols produced comparable low PCR duplication rates (MDA 1.26% + / - 0.52 vs. PTA 1.84% + / - 0.99) and GC% (MDA 42.0 + / - 1.47 vs. PTA 40.33 + / - 0.45), PTA produced smaller amplicon sizes. The percentage of mapped reads and mapping quality score were also significantly higher with PTA compared to MDA (PTA 97.9 + / - 0.62 vs. MDA 82.13 + / - 0.62 and PTA 46.3 + / - 3.18 vs. MDA 43.2 + / - 4.21, respectively). Overall, PTA generates more usable mapped data when compared to MDA. Figure 4A shows that compared to MDA, PTA significantly improves amplification uniformity, resulting in broader coverage and fewer regions with near-zero coverage. The use of PTA can identify low-frequency sequence variants within a population of nucleic acids containing the variant, which constitutes 0.01% or more of the total sequence. PTA can be successfully used for amplification of single-cell genomes.
[0088] Example 2: Comparative Analysis of PTA Benchmarking the maintenance and isolation of PTA and SCMDA cells Lymphoblastoid cells from the 1000 Genomes Project target NA12878 (Coriell Institute, Camden, NJ, USA) were maintained in RPMI medium supplemented with 15% FBS, 2 mM L-glutamine, 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B. Cells were cultured at 3.5 × 10 5 Cells were seeded at a density of 1000 cells / ml and split every 3 days. They were maintained in a humidified incubator at 37°C with 5% CO2. Before isolating single cells, 3 mL of cell suspension expanded over the past 3 days was spun at 300 x g for 10 minutes. Pelleted cells were washed with 1 mL of cell wash buffer (Mg 2+ or Ca 2+ The cells were washed three times with 1x PBS containing 2% FBS without ATP and spun consecutively at 300 x g, 200 x g, and finally 100 x g for 5 minutes to remove dead cells. Next, the cells were resuspended in 500 μL of cell wash buffer and subsequently stained with 100 nM calcein AM and 100 ng / ml propidium iodide (PI) to differentiate the live cell population. The cells were thoroughly washed with ELIMINase and loaded onto a calibrated BD FACScan flow cytometer (FACSAria II) using Accudrop fluorescent beads. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS with 0.2% Tween 20. Several wells were intentionally left empty to serve as no-template controls. Immediately after sorting, the plate was briefly centrifuged and placed on ice. The cells were then frozen at -80°C for at least overnight.
[0089] PTA and SCMDA experiments WGA reactions were assembled on a pre-PCR workstation, which provided constant positive pressure with HEPA-filtered air and was decontaminated with UV light for 30 minutes before each experiment. MDA was performed according to the SCMDA methodology using a published protocol (Dong et al. Nat. Meth. 2017, 14, 491-493). Specifically, exonuclease-resistant random primers were added to the lysis buffer at a final concentration of 12.5 μM. 4 μL of the resulting lysis mix was added to the tube containing the single cells, mixed by pipetting three times, briefly spun, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, mixed by pipetting three times, briefly centrifuged, and placed on ice. Following this, 40 μL of amplification mix was added, followed by incubation at 30°C for 8 hours, and then amplification was terminated by heating to 65°C for 3 minutes. PTA was performed by further lysing the cells after freeze-thawing by adding 2 μL of a pre-chilled 1:1 mixture of 5% Triton X-100 and 20 mg / ml proteinase K. The cells were then vortexed, briefly centrifuged, and then placed at 40°C for 10 minutes. Next, 4 μL of denaturation buffer and 1 μL of 500 μM exonuclease-resistant random primers were added to the lysed cells to denature the DNA, followed by vortexing, rotation, and placement at 65°C for 15 minutes. Next, 4 μL of room-temperature quenching solution was added, and the sample was vortexed and spun down. In the final amplification reaction, 56 μL of amplification mix contained an equal ratio of alpha-thio-ddNTPs at a concentration of 1200 μM. The sample was then placed at 30°C for 8 hours, after which amplification was terminated by heating at 65°C for 3 minutes. After SCMDA or PTA amplification, DNA was purified using AMPure XP magnetic beads at a 2:1 bead-to-sample ratio, and yields were measured using a Qubit 3.0 fluorometer with the Qubit dsDNA HS assay kit according to the manufacturer's instructions.
[0090] Library preparation After adding the preparation solution, 1 μg of SCMDA product was fragmented for 30 minutes according to the HyperPlus protocol. The sample then underwent standard library preparation using 15 μM unique dual-index adapters and four cycles of PCR. The entire product of each PTA reaction was used for DNA sequencing library preparation using the standard amplification protocol, without fragmentation. 2.5 μM unique dual-index adapters were used in the ligation, and 15 cycles of PCR were used in the final amplification. SCMDA and PTA libraries were then visualized on a 1% agarose E-Gel. 400-700 bp fragments were excised from the gel and recovered using the Gel DNA Recovery Kit. The final libraries were quantified using the Qubit dsDNA BR Assay Kit and the Agilent 2100 Bioanalyzer before sequencing on the NovaSeq 6000.
[0091] Data analysis Data were trimmed using trimmomatic and then aligned to hg19 using BWA. Reads were marked for duplicates using Picard, followed by local realignment and base recalibration using best practices in GATK3.5. All files were downsampled to the specified number of reads using PicardDownSampleSam. Quality metrics were obtained from the final bam files using qualimap and Picard AlignmentMetricsAnumary and CollectWgsMetrics. Lorenz curves were plotted and Gini indices were calculated using htSeqTools. SNV calling was performed using UnifiedGenotyper, which was then filtered using standard recommended criteria (QD<2.0 ||FS>60.0 ||MQ<40.0 ||SOR>4.0 ||MQRankSum<-12.5 ||ReadPosRankSum<-8.0). No regions were excluded from the analysis, and no other data normalization or manipulation was performed. Sequencing metrics for the tested methods are shown in Table 2.
[0092] [Table 2]
[0093] Breadth and uniformity of genome coverage We performed a comprehensive comparison of PTA with all common single-cell WGA methods. To achieve this, we performed PTA and an improved version of MDA called single-cell MDA (Dong et al. Nat. Meth. 2017, 14, 491-493) (SCMDA) on 10 NA12878 cells each. Furthermore, these results were compared for cells subjected to amplification using DOP-PCR (Zhang et al. PNAS 1992, 89, 5847-5851), MDA Kit 1 (Dean et al. PNAS 2002, 99, 5261-5266), MDA Kit 2, MALBAC (Zong et al. Science 2012, 338, 1622-1626), LIANTI (Chen et al., Science 2017, 356, 189-194), or PicoPlex (Langmore, Pharmacogenomics 3, 557-560 (2002)) using data generated as part of the LIANTI study.
[0094] To normalize across samples, raw data from all samples were aligned and preprocessed for variant calling using the same pipeline. The bam files were then subsampled to 300 million reads each before comparisons were performed. Importantly, PTA and SCMDA products were not screened before further analysis, while all other methods were screened for genome coverage and uniformity before selecting the highest-quality cells used in subsequent analyses. Notably, SCMDA and PTA were compared to bulk diploid NA12878 samples, while all other methods were compared to bulk BJ1 diploid fibroblasts used in the LIANTI study. As seen in Figures 3C-3F, PTA had the highest percent of reads aligned to the genome and the highest mapping quality. PTA, LIANTI, and SCMDA had similar GC content, all of which were lower than the other methods. PCR replication rates were similar across all methods. Furthermore, the PTA method allowed smaller templates, such as mitochondrial genomes, to give higher coverage rates (similar to larger standard chromosomes) compared to other methods tested (Figure 3G).
[0095] Next, we compared the coverage breadth and uniformity of all methods. Example coverage plots across chromosome 1 are shown for SCMDA and PTA, showing that PTA significantly improved coverage uniformity and allele frequency (Figure 4B). Next, we calculated the coverage percentage for all methods using increasing read counts. PTA approached two bulk samples at all depths, a significant improvement over all other methods (Figure 5A). Next, we measured coverage uniformity using two strategies. The first approach was to calculate the coefficient of variation of coverage at increasing sequencing depths, where PTA was found to be more uniform than all other methods (Figure 5B). The second strategy was to calculate Lorenz curves for each subsampled bam file, where PTA was again found to have the greatest uniformity (Figure 5C). To measure the reproducibility of amplification uniformity, the Gini index was calculated to estimate the difference of each amplification reaction from perfect uniformity (de Bourcy et al., PloS One 9, e105585 (2014)). PTA was again shown to be more reproducibly uniform than the other methods (Figure 5D).
[0096] SNV sensitivity To determine the impact of these differences on the performance of amplification methods for SNV calling, we compared the variant call rates of each method relative to the corresponding bulk sample at increasing sequencing depths. To estimate sensitivity, we compared the percentage of variants called in the corresponding bulk sample, subsampled to 650 million reads found in each cell at each sequencing depth (Figure 5E). The improved coverage and uniformity of PTA resulted in the detection of 45.6% more variants than the next most sensitive method, MDA Kit 2. Examination of sites called as heterozygous in the bulk sample showed that PTA significantly reduced allele skewing at those heterozygous sites (Figure 5F). This finding supports the assertion that PTA not only amplifies more uniformly across the genome, but also more uniformly amplifies two alleles within the same cell.
[0097] SNV specificity To estimate the specificity of variant calling, variants called in each single cell that were not found in the corresponding bulk sample were considered false positives. Cryolysis of SCMDA significantly reduced the number of false positive variant calls (Figure 5G). Methods using thermostable polymerases (MALBAC, PicoPlex, and DOP-PCR) showed that the specificity of SNV calling further decreased with increasing sequencing depth. Without being bound by theory, this may be the result of a significantly increased error rate for these polymerases compared to Phi29 DNA polymerase. Furthermore, the base change patterns observed in the false positive calls also appear to be polymerase-dependent (Figure 5H). As seen in Figure 5G, the model of suppressed error propagation in PTA is supported by the lower false positive SNV call rate in PTA compared to standard MDA protocols. Furthermore, PTA had the lowest allele frequency for false positive variant calls, which is also consistent with the model of suppressed error propagation by PTA (Figure 5I).
[0098] Example 3: Massively parallel single-cell DNA sequencing A protocol for massively parallel DNA sequencing using PTA is established. First, cell barcodes are added to random primers. Two strategies are employed to minimize any amplification bias introduced by the cell barcodes: 1) increasing the size of the random primers and / or 2) creating primers that loop back on themselves to prevent the cell barcode from binding to the template (Figure 10B). Once the optimal primer strategy is established, sorting is scaled up to 384 sorted cells using, for example, the Mosquito HTS liquid handler, which can pipette even viscous liquids up to 25 nL with high precision. This liquid handler reduces reagent costs by approximately 50-fold by using a 1 μL PTA reaction instead of the standard 50 μL reaction volume.
[0099] The amplification protocol is transferred to the droplets by delivering primers bearing cell barcodes to the droplets. Solid supports, such as beads created using a split-and-pool strategy, are optionally used. Suitable beads are available, for example, from ChemGenes. The oligonucleotides, in some cases, include random primers, cell barcodes, unique molecular identifiers, and cleavable sequences or spacers for releasing the oligonucleotides after the beads and cells are encapsulated in the same droplet. During this process, the concentrations of template, primers, dNTPs, alpha-thio-ddNTPs, and polymerase in the droplets are optimized for low nanoliter volumes. Optimization, in some cases, involves the use of larger droplets to increase reaction volume. As shown in Figure 9, this process requires two sequential reactions to lyse the cells, followed by WGA. The first droplet containing the lysed cells and beads is combined with the second droplet containing the amplification mix. Alternatively, or in combination, cells can be encapsulated in hydrogel beads before lysis, and then both beads are added to the oil droplet. See Lan, F. et al., Nature Biotechnol., 2017, 35:640-646.
[0100] Additional methods include the use of microwells, which in some cases capture 140,000 single cells in a 20 picoliter reaction chamber on a device the size of a 3" x 2" microscope slide. Similar to droplet-based methods, these wells combine cells with beads containing cell barcodes to enable massively parallel processing. See Gole et al., Nature Biotechnol., 2013, 31:1126-1132.
[0101] Example 4: Parallel analysis of genome and transcriptome in single cells Single cells from a cell population are sorted and placed one cell per well. Each well contains an antibody immobilized on a surface region that binds to the cell nucleus. The cell's outer membrane is lysed, releasing the mRNA into the solution within the well, while the nuclease remains intact and bound to the well region. RT is performed using the mRNA in solution as a template to generate cDNA using the primers shown in Figure 8A. Optionally, an rRNA (ribosomal RNA) depletion step is performed. RT PCR is performed using a first template switching primer containing the TSS region (transcription start site), anchor region, RNA BC region, and poly(dT) tail at the 5' to 3' end, and a second template switching primer containing the TSS region, anchor region, and poly(G) tail at the 3' to 5' end. After removing the RT PCR product (cDNA library) for subsequent sequencing, any remaining RNA in the cells is removed by UNG. An RNA library is prepared using Nextera / transposon-based sequencing methods and reagents (Figure 8B). The cDNA library contains short cDNAs that have been amplified approximately 1000-fold. The nuclei are then lysed, and the released genomic DNA is subjected to random primer PTA using an isothermal polymerase with random primers 6-9 bases in length. Amplification conditions for PTA are selected to generate amplicons 250-1500 bases in length. The PTA products are optionally subjected to additional amplification and then sequenced. RNA and DNA sequencing data are compiled into a database for analysis.
[0102] Example 5: Single-cell multi-omic analysis A population of cells is contacted with an antibody library in which the antibodies are labeled with either a fluorescent label, a nucleic acid barcode, or both. The labeled antibody binds to at least one cell in the population, and the cells are sorted, with one cell per well. Some labeled antibodies provide specific information about cell surface protein markers after binding, which is obtained either by fluorescence microscopy or by reading the barcode tagged to the antibody. Each well contains an antibody immobilized on a region of the surface, which binds to the cell nucleus. The outer membrane of the cell is lysed, releasing the mRNA into the solution within the well, while the nuclease remains intact and bound to the region of the well. Optionally, an rRNA (ribosomal RNA) depletion step is performed. RT is then performed using the mRNA in the solution as a template to generate cDNA. A first template switching primer containing a 5' to 3' region of the TSS region (transcription start site), an anchor region, an RNA BC region, and a poly(dT) tail, and a second template switching primer containing a 3' to 5' region of the TSS region, an anchor region, and a poly(G) tail, are used for RT-PCR. After removing the RT-PCR products (cDNA library) for subsequent sequencing, any remaining RNA in the cells is removed by UNG. The cDNA library contains short cDNAs amplified approximately 1000-fold. Next, nuclei are lysed, and the released genomic DNA is subjected to random primer PTA using an isothermal polymerase with random primers 6-9 bases in length. Amplification conditions for PTA are selected to generate amplicons 250-1500 bases long. The PTA products are optionally subjected to additional amplification and sequenced. RNA and DNA sequencing data are compiled into a database for analysis.
[0103] Example 6: Single-cell analysis of the methylome and transcriptome Single cells from a cell population are sorted and placed one cell per well. Each well contains an antibody immobilized on a surface region that binds to the cell nucleus. The cell's outer membrane is lysed, releasing the mRNA into the solution within the well, while the nuclease remains intact and bound to the well region. The mRNA transcript is then contacted with terminal transferase, which adds riboguanine to the 5' end of the mRNA strand. RT is then performed using the mRNA in solution as a template to generate cDNA. A first template switching primer containing a TSS region (transcription start site), anchor region, RNA BC region, and poly(dT) tail at the 5' to 3' end, and a second template switching primer containing a TSS region, anchor region, and poly(G) tail at the 3' to 5' end, are used for RT PCR. After removing the RT PCR products (cDNA library) for subsequent sequencing, any remaining RNA in the cells is removed by UNG. The cDNA library contains short cDNAs that have been amplified approximately 1000-fold. The nuclei are then lysed, and the released genomic DNA is fragmented using a methylation-sensitive endonuclease. The genomic fragments are then subjected to random primer PTA using an isothermal polymerase with 6- to 9-base-long random primers. Amplification conditions for PTA are selected to generate amplicons 250 to 1500 bases long. The PTA products are optionally subjected to additional amplification and then sequenced. The RNA and DNA sequencing data are compiled into a database for analysis, and methylation-sensitive endonuclease cleavage sites are identified. These sites are used to map the location of methylation on the original genomic DNA.
[0104] Example 7: Single-cell analysis of the methylome and genome. Single cells from a cell population are sorted and placed one cell per well. Each well contains an antibody immobilized on a surface region that binds to the cell nucleus. The cells are lysed using a methylation-sensitive enzyme, and the genome is subjected to random primer PCR using an isothermal polymerase with random primers 6 to 9 bases in length. The PTA amplification conditions are selected to generate amplicons 250 to 1500 bases in length. The reaction mixture is split, and half of the mixture is subjected to exome enrichment, whole genome sequencing, or other targeted sequencing methods. The other half of the reaction mixture is subjected to methylation-sensitive PCR conditions. Methylation and DNA sequencing data are compiled into a database for analysis.
[0105] Example 8: Single-cell analysis of the surface proteome and genome. Cells from a sample containing a population of cells are contacted with a library of baits, such as antibodies, polynucleotides, or other small molecules. In some cases, the baits are barcoded (e.g., barcoded antibodies) to allow for pull-down and identification of bait binding to cell surface proteins. Alternatively, or in combination, the baits are labeled with other labels, such as fluorescent labels or mass tags. Single cells from the cell population are sorted, with one cell per well. Optionally, the cell surface-bound baits are removed for sequencing or identification before preparation of a genomic library. The cells are lysed, releasing the genome into solution, and fragments are generated. The genomic fragments are subjected to random primer PTA using an isothermal polymerase with random primers 6-9 bases long. Alternatively, the genome is not fragmented prior to PTA amplification. PTA amplification conditions are selected to generate amplicons 250-1500 bases long. The PTA products are optionally further amplified and sequenced. The cell surface protein and DNA sequencing data are compiled into a database for analysis.
[0106] Example 9: Multiomics for measuring drug resistance Monotherapy using small molecule inhibitors targeting FLT3 in acute myeloid leukemia (AML) has shown clinical benefit, but resistance invariably develops. The FLT3 inhibitor quizartinib (AC220) is one such inhibitor, resulting in approximately 50% combined complete remission in patients with relapsed or refractory AML. Despite this success, secondary FLT3 mutations in the activation loop (D835) and gatekeeper residue F691 have been identified in patients with FLT3-ITD who relapsed on quizartinib therapy. Clinical resistance to the multikinase inhibitor PKC412 was determined to be the result of secondary mutations in the FLT3 kinase domain. Additional FLT3-independent modes of resistance to targeted therapy have been identified in FLT3-ITD AML, including AXL bypass pathway activation, as well as NRAS, TET2, and IDH1 / 2 mutations. Mutations in epigenetic modifying enzymes and transcription factors have also been observed, highlighting the complexity and diversity of mechanisms of resistance to FLT3 inhibition.
[0107] Quizartinib-resistant, matched parental MOLM-13 AML cell lines and cell lines harboring heterozygous FLT3-ITD mutations were generated. The PTA method, combined with RNA-seq chemistry, was used to genomically and transcriptionally probe these drug-resistant single cells to gain insight into the mechanisms of resistance following FLT3 inhibition in AML. Briefly, the workflow consisted of: (1) resistant cell generation, (2) resistant cell isolation, (3) cytoplasmic lysis to release mRNA, (4) reverse transcription to generate cDNA from mRNA, (5) nuclear lysis to release genomic DNA, (6) PTA amplification, (7) separate DNA / RNA enrichment, (8) cDNA PreAMP of enriched mRNA, (9) library preparation, QC, and pooling, (10) next-generation sequencing, and (11) data analysis.
[0108] Cell culture. MOLM-13 acute myeloid leukemia cells harboring a heterozygous FLT3 internal tandem duplication (ITD) 1 were obtained from the DSMZ-German Collection of Microorganisms and Cell Cultures (ACC 554). Cells were maintained in RPMI 1640 (Gibco 11875-093) supplemented with 10% FBS and penicillin / streptomycin and subcultured every 2–3 days, maintaining a density range of 2.5E5–1.5E6 cells / ml. For the generation of quizartinib-resistant MOLM-13 lines, cells were continuously treated with 2 nM quizartinib, supplemented with drug at each subculture until resistant clones emerged after 5 weeks of culture (Figure 9A). Genomic DNA or total RNA was isolated from quizartinib-resistant matched parental MOLM-13 cells during FACS sorting to generate bulk sequencing control libraries for comparison with single-cell datasets.
[0109] For single-cell analysis, ~2.0 E6 MOLM-13 quizartinib-resistant or corresponding parental cells were rinsed twice with calcium- and magnesium-free Dulbecco's phosphate-buffered saline (Gibco) supplemented with 2% FBS and kept on ice until BD FACSAriaIII FACS sorting. Following calcein AM, propidium iodide, and DAPI staining, live cell gating was established (DAPI / PI negative, top 70% calcein AM positive). Single cells were sorted (130 micron nozzle assembly) into a low-binding 96-well PCR plate (semi-skirted) containing cell buffer. After brief vortexing and centrifugation, they were immediately frozen on dry ice.
[0110] Combined genomic / transcriptomic analysis. First, biotin-conjugated oligo-dT primers were used in template-switching reverse transcription reactions to generate first-strand cDNA from single MOLM-13 parental or quizartinib-resistant cells. Primary template-directed amplification (PTA) was performed sequentially following reverse transcription. Next, the first-strand cDNA was affinity-purified using streptavidin M-280 beads and subjected to two high-salt washes followed by one low-salt wash. Twenty cycles of preamplification were performed to generate second-strand cDNA, and RNA sequencing libraries were prepared using the Nextera DNA Flex Library Preparation Kit. For PTA library preparation, PTA products not bound to streptavidin beads were purified using beads and ligated to TruSeq adapters. The amplification products from the PTA reaction were first purified by bead cleanup, measured by Qubit, and analyzed by electrophoresis. Typical yields from mammalian cells (~6 pg DNA) were 1–3 μg, while a single bacterial genome (2–4 fg) yielded up to 50 ng. The amplicon product sizes of PTA-amplified samples ranged from 0.2–4 kB (average 1.5 kb). PTA libraries were prepared for WGS without fragmentation, resulting in yields of approximately 500 ng with a size range of 300–550 bases. Whole genomes from mammalian cells were analyzed by NovaSeq, targeting ~550 million reads. Sequencing files were then transferred for trimming, alignment, and VCF file creation, which were then analyzed by the Trailblazer™ cloud-based bioinformatics platform solution. QC and library preparation times were 4–6 hours. A parallel experiment using RNASeq alone was performed for comparison.
[0111] Results. RNA expression from both parental and resistant cultures demonstrated the ability to create cDNA pools (Figure 9B) using single-pot RNA sequencing chemistry. Genes expressed in these cells created unique patterns that allowed visualization of cell populations by gene expression, with an average of ~10,000 genes detected per cell. In a separate workflow, single-cell genomes were amplified using the PTA method. The two protocols were then combined (yields in Figure 9D) to generate combined transcriptome and genome cDNA pools from each cell. Low-pass (~5 million reads / cell) demonstrate effective amplification and library preparation of both the resistant and parental strains, which have low mitochondrial chromosome content and high complete PreSeq genome estimates (Figure 10A-10C). The data demonstrated that transcripts generated during the RT step were not effectively amplified by the PTA reaction compared to DNA, and that DNA within single cells was effectively amplified using the combined protocol compared to standard PTA-amplified genomes from single cells (Figure 9D). The combined RNASeq / PTA method produced similar results (Figure 10A) as the standard PTA protocol, where ChrM and overlap percentages were typically less than 2%, and the estimated genome size exceeded 3 billion bases (Figures 10A-10C). Genomic assessment revealed greater than 90% mapping and coverage, and greater than 75% specific calling of single nucleotide variants within each cell. Compared to standard PTA genomic chemistry, more variability was observed with the dual protocol. For the transcriptome, the prototype chemistry appeared to detect approximately 3,000–5,000 genes that contained exon-exon junctions. Compared to the RNAseq-only protocol (Figure 9C), the dual protocol (Figure 10D) detected ~30% of genes. Furthermore, the dual / combined RNASeq / PTA protocol was used with a second resistant cell line, SUM159 (a triple-negative breast cancer cell line). RNAseq data run with both protocols produced similar PCA distributions.This indicates that the combined chemicals can detect differential gene expression that is not restricted to parental and resistant cells of a single cell type (Figures 10E-10F).
[0112] Deep sequencing of seven parental and five resistant Molm13 cells was performed to a depth of approximately 25x (Figure 11). Reads were aligned to Hg38 using bwa mem. Quality control and SNV calling were performed using the best case of GATK4. SNVs were restricted to at least two resistant cells, and were only considered if alternative alleles were not called in any parental cells and at least six parental cells were genotyped. All cells had at least 96% of their genomes covered at 1x coverage and at least 76% covered at 10x coverage. The inset shows that known Flt3 indels in Molm13 cells were detected in all cells (four are shown for clarity).
[0113] RNAseq and PTA methods were generally comparable, with both mapping and coverage exceeding 95%, and ChrM and PCR overlap generally below 2.0%. Furthermore, greater than 95% of the genomes of selected samples of both 159 parental and resistant cell lines were recovered. In the Molm13 cell line, an overexpressed gene, GAS6(L), was identified, which is a known mechanism of quizartinib resistance. Gas6 is a ligand for AXL, which is a clinically relevant resistance mechanism in relapsed patients who fail quizartinib treatment (Figure 11B). Deep genome sequencing of both parental and resistant MOLM13 cell lines from the dual protocol detected mutations distributed across all chromosomes. Collectively, 5,675 SNVs unique to the quizartinib-resistant population were identified among all single-cell analyses. While mutations in coding sequences were detected, the majority of observed variants were in the intergenic space. Without being bound by theory, passenger mutations are undoubtedly present in this variant cohort, suggesting that regulation of gene expression at the enhancer or promoter level contributes to resistance and potentially regulation of non-coding RNAs. Dual mRNA-seq transcriptome chemistry / PTA has the ability to detect over 10,000 genes in single cells that can be enriched by FACS. The PTA method has the ability to recover over 97% of the complete genome of an individual cell. The ability to recover both the transcriptome and genome does not significantly impact sensitivity compared to the ability to recover large portions of the genome. When comparing transcriptome-only or combined transcriptome / genome amplification chemistries, over 70% of expressed genes can be detected in many cells.
[0114] Example 10: PTA single cell analysis with exome capture. The general PTA method of Example 3 was used with a modification: an additional exome capture step was utilized to enrich the PTA-generated amplicons. Sixty million reads were obtained from both single-cell samples (27 samples) and bulk samples (112 samples). The results of exome capture sequencing from single cells were compared with those from bulk samples (Figures 12A-12D, 13A, 14A, and 14B). Sequencing results were consistent across multiple samples (Figure 13A), and the average size of the captured amplicons was 623 bases (Figure 13B).
[0115] Example 11: Exome capture + multiomics The general method of any of Examples 5-8 is used with a modification: an additional capture step is utilized to enrich for PTA-generated amplicons generated from genomic DNA. The capture step includes either an exome panel or other panel targeting specific genes. In some cases, such panels are directed to cancer hotspots, viral genomes, or mitochondrial DNA.
[0116] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. 1. A method for multi-omic single cell analysis, said method comprising: a. isolating a single cell from a population of cells; b. sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from said single cell; and c. sequencing the genome of said single cell, wherein said sequencing the genome comprises: i. contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides includes at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; ii. amplifying at least some of the genomes to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iii. Ligating the molecules obtained in step (ii) to adapters, thereby generating a genomic DNA library; and iv. Sequencing the genomic DNA library a sequencing step comprising: A method comprising:
2. The method of claim 1, wherein the mRNA transcript comprises a polyadenylated mRNA transcript.
3. The method of claim 1 , wherein the mRNA transcripts do not include polyadenylated mRNA transcripts.
4. 10. The method of claim 1, wherein the step of sequencing the cDNA library comprises amplifying mRNA transcripts using template-switching primers.
5. 2. The method of claim 1, wherein at least some of the polynucleotides in the cDNA library comprise a barcode.
6. 6. The method of claim 5, wherein the barcode comprises a cell barcode or a sample barcode.
7. 2. The method of claim 1, wherein the cDNA library and the genomic DNA library are pooled prior to said sequencing.
8. The method of claim 1 , wherein the single cell is a primary cell.
9. 10. The method of claim 1, wherein the single cell is derived from liver, skin, kidney, blood, or lung.
10. 10. The method of claim 1, wherein the single cell is a cancer cell, a neuron, a glial cell, or a fetal cell.
11. The method of claim 1 , wherein the single cells are isolated by flow cytometry.
12. 2. The method of claim 1, wherein the method further comprises removing at least one terminator nucleotide from the terminated amplification product.
13. 2. The method of claim 1, wherein the plurality of terminated amplification products comprises an average of 1000 to 2000 bases in length.
14. 2. The method of claim 1, wherein the plurality of terminated amplification products are 250 to 1500 bases in length.
15. 2. The method of claim 1, wherein the plurality of terminated amplification products comprises at least 97% of the genome of the single cell.
16. 10. The method of claim 1, wherein at least some of the amplification products comprise a cell barcode or a sample barcode.
17. 2. The method of claim 1, wherein the step of sequencing the cDNA library comprises single-cell cytoplasmic lysis and reverse transcription.
18. 2. The method of claim 1, wherein the mRNA transcripts are amplified via template-switching reverse transcription.
19. 2. The method of claim 1, wherein the cDNA library comprises at least 10,000 genes.
20. 10. The method of claim 1, wherein sequencing the genome of the single cell further comprises nuclear lysis of the single cell.
21. The method of claim 1 , wherein the method further comprises an additional amplification step using PCR.
22. 2. The method of claim 1, wherein the at least one mutation is identified in the genome of a cell, and the mutation differs from the corresponding position in a reference sequence.
23. 2. The method of claim 1, wherein the at least one mutation occurs in less than 1% of the population of cells.
24. 2. The method of claim 1, wherein the at least one mutation occurs in 0.1% or less of the population of cells.
25. 2. The method of claim 1, wherein the at least one mutation occurs in 0.001% or less of the population of cells.
26. 2. The method of claim 1, wherein the at least one mutation occurs in 1% or less of the amplified product sequences.
27. 2. The method of claim 1, wherein the at least one mutation occurs in 0.1% or less of the amplified product sequences.
28. 2. The method of claim 1, wherein the at least one mutation occurs in 0.001% or less of the amplification product sequences.
29. 1. A method for multi-omic single cell analysis, said method comprising: a. isolating a single cell from a population of cells; b. identifying at least one protein on the surface of said single cell; and c. sequencing the genome of said single cell, wherein said sequencing the genome comprises: i. contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides includes at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; ii. amplifying at least some of the genomes to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iii. Ligating the molecules obtained in step (ii) to adapters, thereby generating a genomic DNA library; and iv. Sequencing the genomic DNA library a sequencing step comprising: A method comprising:
30. 30. The method of claim 29, wherein the step of identifying at least one protein on the surface of the cell comprises contacting the cell with a labeled antibody that binds to the at least one protein.
31. 31. The method of claim 30, wherein the labeled antibody comprises at least one fluorescent label or mass tag.
32. 31. The method of claim 30, wherein the labeled antibody comprises at least one nucleic acid barcode.
33. 1. A method for multi-omic single cell analysis, said method comprising: a. isolating a single cell from a population of cells; b. sequencing the genome of said single cell, said step of sequencing said genome comprising: i. digesting the genome with a methylation-sensitive restriction enzyme to generate genomic fragments; ii. contacting at least some of the genomic fragments with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides includes at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; iii. Amplifying at least some of the genomes to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iv. amplifying at least a portion of said genomic fragment by methylation-specific PCR; v. Ligating the molecules obtained in steps (iii and iv) to adapters, thereby generating a genomic DNA library and a methylome DNA library; and vi. Sequencing the genomic DNA library and the methylome DNA library. a sequencing step comprising: A method comprising:
Citation Information
Patent Citations
DNA sequence coding non-human carbonyl hydrolase variant, vector and host cell transformed with the vector
JP1998042876A
Methods for screening nucleic acids for nucleotide mutations
JP2002523063A
Identification of polynucleotides related to the sample
JP2014518618A
Assays to identify genetic elements that influence phenotype
JP2018525024A
Methods for capturing nascent proteins
US20100075374A1