Single-cell analysis
The method for multi-omic single cell analysis using primary template-directed amplification (PTA) addresses the challenges of current nucleic acid amplification and sequencing by providing accurate and efficient analysis of RNA, DNA, and proteins from single cells.
Patent Information
- Application Number
- JP2022506428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-07-31
- Filing Date
- 2020-07-30
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-07-30
AI Technical Summary
Current nucleic acid amplification and sequencing methods struggle with accuracy, scalability, and efficiency, particularly when analyzing small samples from single cells, and fail to simultaneously analyze RNA, DNA, and proteins effectively.
A method for multi-omic single cell analysis involving the isolation of single cells, sequencing of a cDNA library from amplified mRNA transcripts, and sequencing of the genome using primary template-directed amplification (PTA) with strand displacement replication and terminator nucleotides.
This method provides highly accurate and efficient amplification and sequencing of nucleic acids from single cells, enabling simultaneous analysis of RNA, DNA, and proteins, and achieving improved representation, uniformity, and accuracy in a reproducible manner.
Smart Images

Figure 0007691975000003 
Figure 0007691975000004 
Figure 0007691975000005
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 881,183, filed Jul. 31, 2019, which is hereby incorporated by reference in its entirety.
Background Art
[0002] Background Research methods that utilize nuclear amplification, such as next-generation sequencing, provide a wealth of information regarding complex samples, genomes, and other nucleic acid sources. In some cases, these samples are obtained in small amounts from single cells. There is a need for highly accurate, scalable, and efficient nucleic acid amplification and sequencing methods for research, diagnosis, and treatment involving small samples, particularly methods for the simultaneous analysis of RNA, DNA, and proteins.
Summary of the Invention
[0003] Summary Provided herein is a method for multi-omic single cell analysis, the method comprising: (a) isolating single cells from a population of cells; (b) sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from the single cells; and (c) sequencing the genome of the single cells, wherein the step of sequencing the genome comprises: (i) contacting the genome with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; (ii) amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iii) ligating the molecules obtained in step (ii) to adapters to thereby generate a genomic DNA library; and (iv) sequencing the genomic DNA library. Further provided herein is a method wherein the mRNA transcripts comprise polyadenylated mRNA transcripts. Further provided herein is a method wherein the mRNA transcripts do not comprise polyadenylated mRNA transcripts. Further provided herein is a method wherein the step of sequencing the cDNA library comprises amplification of mRNA transcripts using a template switching primer. Further provided herein is a method wherein at least some of the polynucleotides of the cDNA library comprise barcodes. Further provided herein is a method wherein the barcode comprises a cell barcode or a sample barcode. Further provided herein is a method wherein the cDNA library and the genomic DNA library are pooled prior to the sequencing. Further provided herein is a method wherein the single cell is a primary cell. Further provided herein is a method wherein the single cell is derived from liver, skin, kidney, blood, or lung. Further provided herein is a method wherein the single cell is isolated by flow cytometry.Also provided herein is a method wherein the above method further comprises the step of removing at least one terminator nucleotide from the above termination amplification product. Also provided herein is a method wherein the above plurality of termination amplification products contain bases with an average length of 1000-2000. Also provided herein is a method wherein the above plurality of termination amplification products have a length of 250-1500 bases. Also provided herein is a method wherein the above plurality of termination amplification products contain at least 97% of the genome of the above single cell. Also provided herein is a method wherein at least some of the above amplification products contain a cell barcode or a sample barcode. A method. Also provided herein is a method wherein the step of sequencing the above cDNA library includes single cell lysis and reverse transcription. Also provided herein is a method wherein the above mRNA transcript is amplified via template switching reverse transcription. Also provided herein is a method wherein the above cDNA library contains at least 10,000 genes. Also provided herein is a method wherein the step of sequencing the genome of the above single cell further includes nuclear lysis of the above single cell. Also provided herein is a method wherein the above method further comprises an additional amplification step using PCR. Also provided herein is a method wherein at least one mutation is identified in the genome of a cell, and the above mutation is different from the corresponding position in the reference sequence. Also provided herein is a method wherein the above at least one mutation occurs in less than 1% of a population of cells. Also provided herein is a method wherein the above at least one mutation occurs in 0.1% or less of a population of cells. Also provided herein is a method wherein the above at least one mutation occurs in 0.001% or less of a population of cells. Also provided herein is a method wherein the above at least one mutation occurs in 1% or less of the amplified product sequence. Also provided herein is a method wherein the above at least one mutation occurs in 0.1% or less of the amplified product sequence. Also provided herein is a method wherein the above at least one mutation occurs in 0.001% or less of the amplified product sequence.
[0004] Provided herein is a method for multi-omic single cell analysis, the method comprising: (a) isolating single cells from a population of cells; (b) identifying at least one protein on the surface of the single cells; and (c) sequencing the genome of the single cells, wherein the step of sequencing the genome comprises: (i) contacting the genome with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; (ii) amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iii) ligating the molecules obtained in step (ii) to adapters to thereby generate a genomic DNA library; and (iv) sequencing the genomic DNA library. Further provided herein is a method wherein the step of identifying at least one protein on the surface of the cells comprises contacting the cells with a labeled antibody that binds to the at least one protein. Further provided herein is a method wherein the labeled antibody comprises at least one fluorescent label or mass tag. Further provided herein is a method wherein the labeled antibody comprises at least one nucleic acid barcode.
[0005] Provided herein is a method for multi-omic single cell analysis, the method comprising: (a) isolating single cells from a population of cells; (b) sequencing the genome of the single cells, the genome sequencing step comprising: (i) digesting the genome with a methylation-sensitive restriction enzyme to generate genomic fragments; (ii) contacting at least some of the genomic fragments with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the nucleotide mixture comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; (iii) amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; (iv) amplifying at least a portion of the genomic fragments by methylation-specific PCR; (v) ligating the molecules obtained in steps (iii) and (iv) to adapters, thereby generating a genomic DNA library and a methylome DNA library; and (vi) sequencing the genomic DNA library and the methylome DNA library.
[0006] Incorporation by reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application had been specifically and individually indicated to be incorporated by reference.
Brief Description of the Drawings
[0007] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description which illustrates exemplary embodiments in which the principles of the invention are utilized, and from the appended drawings thereof.
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 1F
Figure 1G
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 2E
Figure 2F
Figure 2G
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 3F
Figure 3G
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 5F
Figure 5G
Figure 5H
Figure 5I
Figure 6
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 10A
Figure 10B
Figure 10C
Figure 10D
Figure 10E
Figure 10F
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 12C
Figure 12D
Figure 13A
Figure 13B
Figure 14A
Figure 14B
DETAILED DESCRIPTION OF THE INVENTION
[0008] Detailed Description of the Invention There is a need to develop new expandable, accurate, and efficient methods for nucleic acid amplification (including single-cell and multi-cell genome amplification) and sequencing that overcome the limitations of current methods by enhancing representation, uniformity, and accuracy in a reproducible manner. Provided herein are compositions and methods for providing accurate and expandable primary template-directed amplification (PTA) and sequencing. Such methods and compositions facilitate highly accurate amplification of target (or "template") nucleic acids, thereby improving the accuracy and sensitivity of downstream applications such as next-generation sequencing. Also provided herein are methods for determination of single nucleotide variants, copy number diversity, structural diversity, chronotyping, and measurement of environmental mutagenicity. Measurement of genomic diversity by PTA can be used for a variety of applications including measurement of carcinogenicity of compounds or radiation, including measurement of environmental mutagenicity, prediction of the safety of gene editing technologies, measurement of genomic changes through cancer therapy, genotoxicity studies to determine the safety of new foods or drugs, estimation of age, analysis of resistant bacteria, and identification of bacteria in the environment for industrial applications. Furthermore, these methods can be used to detect selection of specific cell populations after changes in environmental conditions such as exposure to anti-cancer therapy, and to predict response to immunotherapy based on mutations and neoantigen load in single cancer cells.
[0009] Definitions Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these inventions belong.
[0010] Throughout this disclosure, numerical features are presented in range format. It should be understood that descriptions in range format are for convenience and brevity only and should not be construed as an inflexible limitation on the scope of any embodiment. Thus, a recitation of a range should be considered to specifically disclose all possible sub-ranges, as well as individual numerical values down to one tenth of the unit of the lower limit of that range, unless the context clearly dictates otherwise. For example, a description of a range such as 1 to 6 should be considered to specifically disclose sub-ranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, and individual values within that range, such as 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the width of the range. The upper and lower limits of these intervening ranges may independently be included in a smaller range and are included in the present invention, subject to the limits specifically excluded in the recited range. If the recited range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the present invention, unless the context clearly dictates otherwise.
[0011] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, the terms "comprising" and / or "comprises" as used herein, when used, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0012] As used herein, unless specifically stated otherwise or apparent from the context, the term "about" with respect to a number or range of numbers means the stated number and numbers + / − 10% thereof, i.e., for values listed for a certain range, numbers that are 10% below the lowest listed limit and 10% above the highest listed limit.
[0013] As used herein, the terms "subject", "patient", or "individual" refer to, for example, humans, veterinary animals (e.g., cats, dogs, cows, horses, sheep, pigs, etc.), and experimental animal models of diseases (e.g., mice, rats). According to the present invention, within the scope of the skills of those skilled in the art, conventional molecular biology, microbiology, and recombinant DNA techniques can be used. Such techniques are fully described in the literature. For example, among others, Sambrook, Fritsch & Maniatis, Molecular Cloning: A Laboratory Manual, Second Edition (1989) Cold Spring Harbour Laboratory Press, Cold Spring Harbour, New York (referred to herein as "Sambrook et al., 1989"); DNA Cloning: A practical Approach, Volumes I and II (D.N. Glover ed. 1985); Oligonucleotide Synthesis (MJ. Gait ed. 1984); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985); Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984); Animal Cell Culture (R.I. Freshney, ed. (1986); Immobilized Cells and Enzymes (lRL Press, (1986); B. Perbal, A practical Guide To Molecular Cloning (1984); F.M. Ausubel et al. (eds.), Current Protocols in Molecular Biology, John Wiley & Sons, Inc. (1994).
[0014] The term "nucleic acid" encompasses not only single-stranded molecules but also multi-stranded molecules. In double-stranded or triple-stranded nucleic acids, the nucleic acid strands need not have the same extent (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). The nucleic acid templates described herein can be of any size (from small cell-free DNA fragments to the entire genome), depending on the sample, and include, but are not limited to, lengths of 50-300 bases, 100-2000 bases, 100-750 bases, 170-500 bases, 100-5000 bases, 50-10,000 bases, or 50-2000 bases. In some examples, the length of the template is at least 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 50,000, 100,000, 200,000, 500,000, 1,000,000 bases, or longer than 1,000,000 bases. The methods described herein provide for the amplification of nucleic acids such as nucleic acid templates. The methods described herein further provide for the generation of isolated, at least partially purified nucleic acids and libraries of nucleic acids. Nucleic acids include, but are not limited to, DNA, RNA, circular RNA, mtDNA (mitochondrial DNA), cfDNA (cell-free DNA), cfRNA (cell-free RNA), siRNA (small interfering RNA), cffDNA (cell-free fetal DNA), mRNA, tRNA, rRNA, miRNA (microRNA), synthetic polynucleotides, polynucleotide analogs, any other nucleic acids consistent with this specification, or any combination thereof. When provided, the length of a polynucleotide is described as the number of bases and is represented in abbreviated form such as nt (nucleotide), bp (base pair), kb (kilobase), or Gb (gigabase).
[0015] As used herein, the term "droplet" refers to the volume of liquid on a droplet actuator. The droplet can, in some instances, be aqueous or non-aqueous, or a mixture or emulsion containing aqueous and non-aqueous components. For non-limiting examples of droplet fluids that can be subjected to droplet manipulation, see, for example, International Patent Application Publication No. WO2007 / 120241. Any suitable system for forming and manipulating droplets can be used in the embodiments presented herein. For example, in some instances, a droplet actuator is used. For non-limiting examples of droplet actuators that can be used, see, for example, U.S. Patent Nos. 6,911,132; 6,977,033; 6,773,566; 6,565,727; 7,163,612; 7,052,244; 7,328,979; 7,547,380; 7,641,779; U.S. Patent Application Publication Nos. US20060194331; US20030205632; US20060164490; US20070023292; US20060039823; US20080124252; US20090283407; US20090192044; US20050179746; US20090321262; US20100096266; US20110048951; International Patent Application Publication No. WO2007 / 120241. In some cases, beads are provided within the droplet, within the droplet manipulation gap, or on the droplet manipulation surface. In some cases, the beads are provided outside the droplet manipulation gap or in a reservoir located away from the droplet manipulation surface, and this reservoir can be associated with a flow path that enables a droplet containing the beads to enter the droplet manipulation gap or contact the droplet manipulation surface.Non-limiting examples of droplet actuator technologies for immobilizing magnetic-responsive beads and / or non-magnetic-responsive beads and / or for performing droplet manipulation protocols using beads are described in U.S. Patent Application Publication No. US20080053205, International Patent Application Publication Nos. WO2008 / 098236, WO2008 / 134153, WO2008 / 116221, WO2007 / 120241. Bead properties can be utilized in multiplexed embodiments of the methods described herein. Examples of beads having properties suitable for multiplexing, as well as methods for detecting and analyzing signals emitted from such beads, can be found in U.S. Patent Application Publication Nos. US20080305481, US20080151240, US20070207513, US20070064990, US20060159962, US20050277197, US20050118574.
[0016] Primers and / or template-switching oligonucleotides can also be attached to a solid substrate to facilitate reverse transcription and template switching of mRNA polynucleotides. In this arrangement, part of the RT or template-switching reaction occurs in the bulk solution of the device, where the second step of the reaction occurs near the surface. In other arrangements, the primer of the template-switching oligonucleotide can be released from the solid substrate to cause the entire reaction to occur on the surface of the solution. In a polyomic approach, primers for multi-step reactions are, in some cases, immobilized on a solid substrate or combined with beads to achieve a combination of multi-step primers.
[0017] Certain microfluidic devices also support a polyomic approach. As an example, devices fabricated in PDMS often have chambers adjacent to each reaction step. Such multi-chamber devices are often separated using microvalve structures that can control pressure with air or fluids such as water or inert hydrocarbons (fluorinated). In the polyomic approach, each stage of the reaction can be isolated and run individually. At the completion of a particular stage, the valve between adjacent chambers can be released onto the substrate to continuously add subsequent reactions. As a result, individual cells can be used as input template materials to emulate a series of consecutive reactions, such as a series of polyomic (protein / RNA / DNA / epigenomic) reactions. For single-cell analysis, various microfluidics platforms can be used. Cells are manipulated in some cases through hydrodynamic (droplet microfluidics, inertial microfluidics, vortices, microvalves, microstructures (microwells, microtraps, etc.)), electrical methods (dielectrophoresis (DEP), electroosmosis), optical methods (optical tweezers, optically induced dielectrophoresis (ODEP), photothermal capillary), acoustic methods, or magnetic methods. In some cases, the microfluidics platform includes microwells. In some cases, the microfluidics platform includes PDMS (polydimethylsiloxane)-based devices.Non-limiting examples of single cell analysis platforms compatible with the methods described herein are the ddSEQ Single Cell Isolator (Bio-Rad, Hercules, CA, USA, and Illumina, San Diego, CA, USA), Chromium (10x Genomics, Pleasanton, CA, USA), Rhapsody Single Cell Analysis System (BD, Franklin Lakes, NJ, USA), Tapestri Platform (MissionBio, San Francisco, CA, USA), Nadia Innovate (Dolomite Bio, Royston, UK), C1 and Polaris (Fluidigm, South San Francisco, CA, USA); ICELL8 Single Cell System (Takara); MSND (Wafergen); Puncher Platform (Vycap), CellRaft AIR System (CellMicrosystems), DEPArray NxT and DEPArray System (Menarini Silicon Biosystems), AVISO CellCelector (ALS), InDrop System (1CellBio), and TrapTx (Celldom).
[0018] As used herein, the term "unique molecular identifier (UMI)" refers to a unique nucleic acid sequence attached to each of a plurality of nucleic acid molecules. When incorporated into a nucleic acid molecule, the UMI is, in some cases, used to correct for subsequent amplification bias by directly counting the UMI sequenced after amplification. The design, incorporation, and application of UMIs are described, for example, in International Patent Application Publication No. WO 2012 / 142213, Islam et al. Nat. Methods (2014) 11:163-166, Kivioja, T. et al. Nat. Methods (2012) 9:72-74, Brenner et al. (2000) PNAS 97(4), 1665, and Hollas and Schuler (2003) Conference: 3rd International Workshop on Algorithms in Bioinformatics, Volume: 2812.
[0019] As used herein, the term "barcode" refers to a nucleic acid tag that can be used to identify a sample or source of nucleic acid material. Thus, when a nucleic acid sample is derived from multiple sources, the nucleic acids in each nucleic acid sample are, in some cases, tagged with different nucleic acid tags so as to be able to identify the source of the sample. Barcodes are commonly also referred to as indexes, tags, etc. and are well known to those skilled in the art. Any suitable barcode or set of barcodes can be used. See, for example, the non-limiting examples provided in U.S. Patent No. 8,053,192 and International Patent Application Publication No. WO2005 / 068656. Barcoding of single cells can be carried out, for example, as described in U.S. Patent Application No. 2013 / 0274117.
[0020] As used herein, the terms "solid surface", "solid support" and other grammatical equivalents refer to any material that is suitable for, or can be modified to be suitable for, the attachment of the primers, barcodes and sequences described herein. Exemplary substrates include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylic, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon®, etc.), polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials (e.g., silicon or modified silicon), carbon, metals, inorganic glass, plastics, fiber optic bundles, and various other polymers. In some embodiments, the solid support includes a patterned surface suitable for immobilizing primers, barcodes, and sequences in a regular pattern.
[0021] As used herein, the term "biological sample" includes, but is not limited to, tissues, cells, biological fluids, and isolates thereof. The cells or other samples used in the methods described herein are, in some cases, isolated from a human patient, animal, plant, soil, or other sample including microorganisms such as bacteria, fungi, protozoa. In some cases, the biological sample is of human origin. In some cases, the biological sample is of non-human origin. Cells, in some cases, are subjected to the PTA method and sequencing described herein. Variants detected across the genome or at specific loci can be compared to all other cells isolated from the subject for the purpose of tracking the lineage of the cell line for research or diagnostic purposes. In some cases, variants are confirmed through additional methods such as direct PCR sequencing.
[0022] Single cell analysis Described herein are methods and compositions for the analysis of single cells. When cells are analyzed in bulk, general information about the cell population is obtained, but it is often the case that low-frequency variants relative to the background cannot be detected. Such variants may possess important properties such as drug resistance or mutations associated with cancer. In some cases, DNA, RNA, and / or protein from the same single cell are analyzed in parallel. The analysis may include the identification of epigenetic post-translational modifications (e.g., glycosylation, phosphorylation, acetylation, ubiquitination, histone modifications) and / or post-transcriptional modifications (e.g., methylation, hydroxymethylation). Such methods may include "primary template-directed amplification" (PTA) to obtain a library of nucleic acids for sequencing. In some cases, PTA is combined with additional steps or methods such as RT-PCR or proteome / protein quantification techniques (e.g., mass spectrometry, antibody staining, etc.). In some cases, the various components of the cell are physically or spatially separated from each other during the individual analysis steps. For example, the workflow in some examples includes the general steps of FIG. 1A. Proteins are first labeled with antibodies. In some cases, at least some of the antibodies include tags or markers (e.g., nucleic acid / oligotags, mass tags, or fluorescent tags). In some cases, a portion of the antibody includes an oligotag. In some cases, a portion of the antibody includes a fluorescent marker. In some cases, the antibody is labeled with two or more tags or markers. In some cases, a portion of the antibody is sorted based on a fluorescent marker. After RT-PCR, the first-strand mRNA products are generated and then removed for analysis. Next, libraries are generated from the RT-PCR products and barcodes present in the protein-specific antibodies, and then these are sequenced. In parallel, genomic DNA from the same cell is subjected to PTA, a library is generated, and then sequenced. The sequencing results arising from the genome, proteome, and transcriptome are, in some cases, pooled using bioinformatics methods.The methods described herein, in some cases, include any combination of labeling, cell sorting, affinity separation / purification, lysis of specific cell components (e.g., outer membrane, nucleus, etc.), RNA amplification, DNA amplification (e.g., PTA), or other processes related to the separation or analysis of proteins, RNA, or DNA. In some cases, the methods described herein include one or more enrichment steps, such as exosome enrichment.
[0023] Described herein is a first method of single cell analysis, including the analysis of RNA and DNA from single cells (Figure 1B). This method includes the isolation of single cells, the lysis of single cells, and reverse transcription (RT). In some cases, reverse transcription is performed using a template switching oligonucleotide (TSO). In some cases, the TSO includes a molecular tag, such as biotin, that enables subsequent pull-down of the cDNA RT product, and PCR amplification of the RT product to generate a cDNA library. Alternatively, or in combination, centrifugation is used to separate RNA in the supernatant from cDNA in the cell pellet. The remaining cDNA is, in some cases, fragmented and removed using UDG (uracil DNA glycosylase), RNA is degraded using alkaline lysis, and the genome is denatured. After neutralization, addition of primers and PTA, the amplification products are, in some cases, purified on SPRI (solid phase reversible immobilization) beads and ligated to adapters to generate a gDNA library.
[0024] Described herein is a second method of single cell analysis that includes the analysis of RNA and DNA from a single cell (Figure 1C). In some cases, the method includes the isolation of a single cell, the lysis of the single cell, and reverse transcription (RT). In some cases, reverse transcription is performed using a template switching oligonucleotide (TSO). In some cases, the TSO includes a molecular tag such as biotin that enables subsequent pull-down of the cDNA RT product, and PCR amplification of the RT product to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers and PTA, the amplification products are in some cases purified on SPRI (solid phase reversible immobilization) beads and ligated to adapters to generate a gDNA library. The RT products are in some cases isolated by a pull-down such as a pull-down using streptavidin beads.
[0025] Described herein is a third method of single cell analysis that includes the analysis of RNA and DNA from a single cell (Figure 1D). In some cases, the method includes the isolation of a single cell, the lysis of the single cell, and reverse transcription (RT). In some cases, reverse transcription is performed using a template switching oligonucleotide (TSO) in the presence of terminator nucleotides. In some cases, the TSO includes a molecular tag such as biotin that enables subsequent pull-down of the cDNA RT product, and PCR amplification of the RT product to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers and PTA, the amplification products are in some cases purified on SPRI (solid phase reversible immobilization) beads and ligated to adapters to generate a DNA library. The RT products are in some cases isolated by a pull-down such as a pull-down using streptavidin beads.
[0026] Described herein is a fourth method of single cell analysis, including the analysis of RNA and DNA from a single cell (Figure 1E). In some cases, the method includes isolation of a single cell, lysis of the single cell, and reverse transcription (RT). In some cases, reverse transcription is performed using a template switching oligonucleotide (TSO). In some cases, the TSO includes a molecular tag such as biotin that enables subsequent pull-down of the cDNA RT product, and PCR amplification of the RT product to generate a cDNA library. In some cases, alkaline lysis is then used to degrade the RNA and denature the genome. After neutralization, addition of random primers and PTA, the amplification products are, in some cases, subjected to RNase and cDNA amplification using block and labeled primers. gDNA is purified with SPRI (solid phase reversible immobilization) beads and ligated to adapters to generate a gDNA library. The RT products are, in some cases, isolated by pull-down such as pull-down using streptavidin beads.
[0027] Described herein is a fifth method of single cell analysis, including the analysis of RNA and DNA from a single cell (Figs. 7A and 7B). A population of cells is contacted with an antibody library in which the antibodies are labeled. In some cases, the antibodies are labeled with either a fluorescent label, a nucleic acid barcode, or both. The labeled antibodies bind to at least one cell within the population, and such cells are sorted such that one cell is placed per container (e.g., tube, vial, microwell, etc.). In some cases, the container contains a solvent. In some cases, the surface area of the container is coated with a capture moiety. In some cases, the capture moiety is a small molecule, antibody, protein, or other factor that can bind to one or more cells, organelles, or other cellular components. In some cases, at least one cell, or a single cell, or a component thereof binds to the surface area of the container. In some cases, the nucleus binds to the area of the container. In some cases, the outer membrane of the cell lyses and the mRNA is released into the solution within the container. In some cases, the nucleus of the cell containing genomic DNA binds to the surface area of the container. Next, RT is often performed using the mRNA in the solution as a template to generate cDNA. In some cases, the template-switching primer contains, from 5' to 3', a TSS region (transcription start site), an anchor region, an RNA BC region, and a poly dT tail. In some cases, the poly dT tail binds to the poly A tail of one or more mRNAs. In some cases, the template-switching primer contains, from 3' to 5', a TSS region, an anchor region, and a poly G region. In some cases, the poly G region contains ribo G. In some cases, the poly G region binds to the poly C region on the mRNA transcript. In some cases, ribo G is added to the mRNA transcript by terminal transferase. After removing the RT PCR product for subsequent sequencing, any RNA remaining within the cell is removed by UNG. Next, the nucleus is lysed and the released genomic DNA is subjected to the PTA method using random primers with an isothermal polymerase.In some cases, the primer length is 6 - 9 bases. In some cases, PTA generates genomic amplicons that are 100 - 5000, 200 - 5000, 500 - 2000, 500 - 2500, 1000 - 3000, or 300 - 3000 bases in length. In some cases, PTA generates genomic amplicons with an average length of 100 - 5000, 200 - 5000, 500 - 2000, 500 - 2500, 1000 - 3000, or 300 - 3000 bases. In some cases, PTA generates genomic amplicons that are 250 - 1500 bases in length. In some cases, the methods described herein generate a cDNA pool of short fragments having an amplification of about 500, about 750, about 1000, about 5000, or about 10,000 - fold. In some cases, the methods described herein generate a cDNA pool of short fragments having an amplification of 500 - 5000, 750 - 1500, or 250 - 10,000 - fold. The PTA products are optionally subjected to additional amplification and sequencing.
[0028] Single - cell sample preparation and isolation The methods described herein may require the isolation of single cells for analysis. Any method of single - cell isolation, such as mouth pipetting, micropipetting, flow cytometry / FACS, microfluidics, nuclear sorting methods (tetraploid or otherwise), or manual dilution, can be used with PTA. Such methods are assisted by additional reagents and procedures, such as antibody - based enrichment (e.g., circulating tumor cells), other small - molecule or protein - based enrichment methods, or fluorescence labeling. In some cases, the multi - omic analysis methods described herein include the mechanical or enzymatic dissociation of cells from larger tissues.
[0029] Cell component preparation and analysis The multi-omic analysis method involving PTA described in this specification may include one or more methods for processing cellular components such as DNA, RNA, and / or proteins. In some cases, the nucleus (containing genomic DNA) is physically separated from the cytosol (containing mRNA), followed by treatment with a membrane-selective lysis buffer that lyses the membrane while keeping the nucleus intact. Next, the cytosol is separated from the nucleus using methods including micropipetting, centrifugation, or methods involving antibody-conjugated magnetic microbeads. In another example, magnetic beads coated with oligo dT primers bind to polyadenylated mRNA and separate it from DNA. In another example, DNA and RNA are pre-amplified simultaneously and separated for analysis. In another example, one cell is divided into two equal parts, mRNA from one half is processed, and genomic DNA from the remaining half is processed.
[0030] Multi-omics The methods described herein (e.g., PTA) can be used as an alternative to any number of other known methods in the art for single cell sequencing (such as multiomics). PTA may replace genomic DNA sequencing methods such as MDA, PicoPlex, DOP-PCR, MALBAC, or target-specific amplification. In some cases, PTA replaces the standard genomic DNA sequencing method in multiomics methods including DR-seq (Dey et al., 2015), G&T seq (MacAulay et al., 2015), scMT-seq (Hu et al., 2016), sc-GEM (Cheow et al., 2016), scTrio-seq (Hou et al., 2016), simultaneous multiplex measurement of RNA and protein (Darmanis et al., 2016), scCOOL-seq (Guo et al., 2017), CITE-seq (Stoeckius et al., 2017), REAP-seq (Peterson et al., 2017), scNMT-seq (Clark et al., 2018), or SIDR-seq (Han et al., 2018). In some cases, the methods described herein include PTA and methods for polyadenylated mRNA transcripts. In some cases, the methods described herein include PTA and methods for non-polyadenylated mRNA transcripts. In some cases, the methods described herein include PTA and methods for total (polyadenylated and non-polyadenylated) mRNA transcripts.
[0031] In some cases, PTA is combined with standard RNA sequencing methods to obtain genomic and transcriptome data. In some cases, the multi-omic methods described herein include PTA and one of the following: Drop-seq (Macosko, et al. 2015), mRNA-seq (Tang et al., 2009), InDrop (Klein et al., 2015), MARS-seq (Jaitin et al., 2014), Smart-seq2 (Hashimshony, et al., 2012; Fish et al., 2016), CEL-seq (Jaitin et al., 2014), STRT-seq (Islam, et al., 2011), Quartz-seq (Sasagawa et al., 2013), CEL-seq2 (Hashimshony, et al. 2016), cytoSeq (Fan et al., 2015), SuPeR-seq (Fan et al., 2011), RamDA-seq (Hayashi, et al. 2018), MATQ-seq (Sheng et al., 2017), or SMARTer (Verboom et al., 2019).
[0032] To generate a cDNA library for transcriptome analysis, various reaction conditions and mixes can be used. In some cases, an RT reaction mix is used to generate the cDNA library. In some cases, the RT reaction mixture contains a congestion reagent, at least one primer, a template switching oligonucleotide (TSO), reverse transcriptase, and a dNTP mix. In some cases, the RT reaction mix contains an RNAse inhibitor. In some cases, the RT reaction mix contains one or more surfactants. In some cases, the RT reaction mix contains Tween-20 and / or Triton-X. In some cases, the RT reaction mix contains betaine. In some cases, the RT reaction mix contains one or more salts. In some cases, the RT reaction mixture contains a magnesium salt (e.g., magnesium chloride) and / or tetramethylammonium chloride. In some cases, the RT reaction mix contains gelatin. In some cases, the RT reaction mix contains PEG (PEG1000, PEG2000, PEG4000, PEG6000, PEG8000, or PEG of other lengths).
[0033] The multi-omic methods described herein can provide both genomic and RNA transcription information from a single cell (e.g., combinatorial or dual protocols). In some cases, genomic information from a single cell is obtained from the PTA method, and RNA transcription information is obtained from reverse transcription to generate a cDNA library. In some cases, the full transcription method is used to obtain a cDNA library. In some cases, 3' or 5' end counting is used to obtain a cDNA library. In some cases, UMI may not be used to obtain a cDNA library. In some cases, the multi-omic method provides RNA transcription information from a single cell of at least 500, 1000, 2000, 5000, 8000, 10,000, 12,000, or at least 15,000 genes. In some cases, the multi-omic method provides RNA transcription information from a single cell of about 500, 1000, 2000, 5000, 8000, 10,000, 12,000, or about 15,000 genes. In some cases, the multi-omic method provides RNA transcription information from a single cell of 100 - 12,000, 1000 - 10,000, 2000 - 15,000, 5000 - 15,000, 10,000 - 20,000, 8000 - 15,000, or 10,000 - 15,000 genes. In some cases, the multi-omic method provides genomic sequence information of at least 80%, 90%, 92%, 95%, 97%, 98%, or at least 99% of the genome of a single cell. In some cases, the multi-omic method provides genomic sequence information of about 80%, 90%, 92%, 95%, 97%, 98%, or about 99% of the genome of a single cell.
[0034] The multi-omic method can include the analysis of single cells from a cell population. In some cases, at least 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, or at least 8000 cells are analyzed. In some cases, about 5, 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, or about 8000 cells are analyzed. In some cases, 5 - 100, 10 - 100, 50 - 500, 100 - 500, 100 - 1000, 50 - 5000, 100 - 5000, 500 - 1000, 500 - 10000, 1000 - 10000, or 5000 - 20,000 cells are analyzed.
[0035] The multi-omic method can generate the yield of genomic DNA from the PTA reaction based on the type of single cells. In some cases, the amount of DNA generated from a single cell is about 0.1, 1, 1.5, 2, 3, 5, or about 10 micrograms. In some cases, the amount of DNA generated from a single cell is about 0.1, 1, 1.5, 2, 3, 5, or about 10 femtograms. In some cases, the amount of DNA generated from a single cell is at least 0.1, 1, 1.5, 2, 3, 5, or at least 10 micrograms. In some cases, the amount of DNA generated from a single cell is at least 0.1, 1, 1.5, 2, 3, 5, or at least 10 femtograms. In some cases, the amount of DNA generated from a single cell is about 0.1 - 10, 1 - 10, 1.5 - 10, 2 - 20, 2 - 50, 1 - 3, or 0.5 - 3.5 micrograms. In some cases, the amount of DNA generated from a single cell is about 0.1 - 10, 1 - 10, 1.5 - 10, 2 - 20, 2 - 4, 1 - 3, or 0.5 - 4 femtograms.
[0036] Methylome analysis Described herein are methods that include PTA, and sites of methylated DNA in single cells are determined using the PTA method. In some cases, these methods further include parallel analysis of the transcriptome and / or proteome of the same cells. Methods for detecting methylated genomic bases include selective restriction by a methylation-sensitive endonuclease, followed by treatment with the PTA method. Sites cleaved by such enzymes are determined from sequencing, and methylated bases are identified. In another example, bisulfite treatment of a genomic DNA library converts unmethylated cytosine to uracil. The library is then amplified, in some cases, with methylation-specific primers that selectively anneal to methylated sequences. Alternatively, non-methylation-specific PCR is performed, followed by one or more methods for distinguishing bisulfite reaction bases, including direct pyrosequencing, MS-SnuPE, HRM, COBRA, MS-SSCA, or base-specific cleavage / MALDI-TOF. In some cases, genomic DNA samples are divided for parallel analysis of the genome (or an enriched portion thereof) and methylome analysis. In some cases, analysis of the genome and methylome includes enrichment of genomic fragments (e.g., exome, or other targets) or whole-genome sequencing.
[0037] Bioinformatics Data obtained from the single-cell analysis using PTA described herein can be compiled into a database. Described herein are methods and systems for bioinformatics data integration. Data from proteomes, genomes, transcriptomes, methylomes, or other data are, in some cases, combined / integrated into a database and analyzed. Bioinformatics data integration methods and systems include, in some cases, one or more of protein detection (FACS and / or NGS), mRNA detection, and / or genomic variant detection. In some cases, this data is correlated with a disease state or condition. In some cases, data from multiple single cells are compiled to characterize a larger cell population, such as cells from a particular sample, region, organism, or tissue. In some cases, protein data is obtained from fluorescently labeled antibodies that selectively bind to proteins on the cell. In some cases, the method for protein detection includes the steps of grouping cells based on a fluorescent marker and reporting the position of the sample after sorting. In some cases, the method for protein detection includes the detection of sample barcodes, the detection of protein barcodes, comparison with a designed sequence, and grouping of cells based on barcodes and copy number. In some cases, protein data is obtained from barcoded antibodies that selectively bind to proteins on the cell. In some cases, transcriptome data is obtained from sample and RNA-specific barcodes. In some cases, the method for mRNA detection includes the detection of sample and RNA-specific barcodes, alignment to the genome, alignment to RefSeq / Encode, reporting of exon / intron / intergenic sequences, analysis of exon-exon junctions, grouping of cells based on barcodes and expression variance, and clustering analysis of variance and top variable genes. In some cases, genomic data is obtained from sample and DNA-specific barcodes.In some cases, the method for genomic variant detection includes detection of the sample and DNA-specific barcodes, alignment to the genome, determination of genome recovery and SNV mapping rate, filtering of reads at exon-exon junctions, generation of a variant call file (VCF), and variant analysis and clustering analysis of top variable mutations.
[0038] variant In some cases, the methods described herein (e.g., multi-omic PTA) result in higher detection sensitivity and / or lower false detection rate for the detection of variants. In some cases, a variant is a difference between an analyzed sequence (e.g., using the methods described herein) and a reference sequence. The reference sequence is, in some cases, obtained from another organism, another individual of the same or a similar species, a population of organisms, or another region of the same genome. In some cases, a variant is identified on a plasmid or chromosome. In some cases, a variant is an SNV (single nucleotide change), SNP (single nucleotide polymorphism), or CNV (copy number variation, or CNA / copy number aberration). In some cases, a variant is a substitution, insertion, or deletion of a base. In some cases, a variant is a transition, transversion, nonsense mutation, silent mutation, synonymous or non-synonymous mutation, non-pathogenic mutation, missense mutation, or frameshift mutation (deletion or insertion). In some cases, PTA results in higher detection sensitivity and / or lower false detection ratio when compared to methods such as in silico prediction, ChIP-seq, GUIDE-seq, circle-seq, HTGTS (high-throughput genome-wide translocation sequencing), IDLV (integrated defective lentivirus), Digenome-seq, FISH (fluorescent in situ hybridization), or DISCOVER-seq.
[0039] primed template-directed amplification Described herein are nucleic acid amplification methods such as "primary template-directed amplification (PTA)". In some cases, PTA is combined with other workflows for multi-omic analysis. For example, one embodiment of the PTA method described herein is schematically represented in FIG. 1G. In the PTA method, a polymerase (e.g., a strand-displacing polymerase) is used to preferentially generate amplicons from a primary template ("direct copy"). As a result, errors are propagated at a lower rate from daughter amplicons during subsequent amplification compared to MDA. As a result, unlike existing WGA protocols, a simple and easily executable method is obtained that can amplify low DNA inputs containing single-cell genomes with high coverage width and uniformity in an accurate and reproducible manner. Furthermore, the terminated amplification products can undergo directional ligation after removal of the terminator, and cell barcodes can be attached to the amplification primers, so that the products from all cells can be pooled after undergoing parallel amplification reactions. In some cases, the template nucleic acid is not bound to a solid support. In some cases, the direct copy of the template nucleic acid is not bound to a solid support. In some cases, one or more primers are not bound to a solid support. In some cases, there are no primers not bound to a solid support. In some cases, the primer attaches to a first solid support and the template nucleic acid attaches to a second solid support, and the first and second solid supports are not the same. In some cases, PTA is used to analyze single cells from larger cell populations. In some cases, PTA is used to analyze more than one cell, or the entire cell population, from a larger cell population.
[0040] Described herein are methods of using nucleic acid polymerases having strand displacement activity for amplification. In some cases, such polymerases include strand displacement activity and a low error rate. In some cases, such polymerases include strand displacement activity and proofreading exonuclease activity, such as 3’->5’ proofreading activity. In some cases, the nucleic acid polymerase is used in combination with other components, such as reversible or irreversible terminators, or additional strand displacement factors. In some cases, the polymerase has strand displacement activity but does not have exonuclease proofreading activity. For example, in some instances, such polymerases include bacteriophage phi29 (Φ29) polymerase, which also has a very low error rate as a result of 3’->5’ proofreading exonuclease activity (see, e.g., U.S. Patent Nos. 5,198,543 and 5,001,050). In some cases, non-limiting examples of strand displacement nucleic acid polymerases include, for example, genetically modified phi29 (Φ29) DNA polymerase, the Klenow fragment of DNA polymerase I (Jacobsen et al., Eur. J. Biochem. 45:623-627 (1974)), phage M2 DNA polymerase (Matsumoto et al., Gene 84:247 (1989)), phage phiPRD1 DNA polymerase (Jung et al., Proc. Natl. Acad. Sci. USA 84:8287 (1987); Zhu and Ito, Biochim. Biophys. Acta. 1219:267-276 (1994)), Bst DNA polymerase (e.g., Bst large fragment DNA polymerase (exo(-)Bst; Aliotta et al., Genet. Anal. (Netherlands) 12:185-195 (1996)), exo(-)Bca DNA polymerase (Walker and Linn, Clinical Chemistry 42:1604-1608 (1996)), Bsu DNA polymerase, Vent R (exo-) DNA polymerase, including VentRDNA polymerases (Kong et al., J. Biol. Chem. 268:1965 - 1975 (1993)), Deep Vent DNA polymerase including Deep Vent (exo -) DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase (Chatterjee et al., Gene 97:13 - 19 (1991)), Sequenase (U.S. Biochemicals), T7 DNA polymerase, T7 - Sequenase, T7 gp5 DNA polymerase, PRDI DNA polymerase, T4 DNA polymerase (Kaboord and Benkovic, Curr. Biol. 5:149 - 157 (1995)) are included. Additional strand - displacing nucleic acid polymerases are also compatible with the methods described herein. The ability of a given polymerase to perform strand - displacement replication can be determined, for example, by using the polymerase in a strand - displacement replication assay (e.g., as disclosed in U.S. Patent No. 6,977,148). Such assays are, in some cases, performed at a temperature suitable for the optimal activity of the enzyme being used, which is, for example, 32°C for phi29 DNA polymerase, 46°C - 64°C for exo(-)Bst DNA polymerase, or 60°C - 70°C for enzymes from hyperthermophilic organisms. Another useful assay for selecting a polymerase is the primer - block assay described in Kong et al., J. Biol. Chem. 268:1965 - 1975 (1993). This assay consists of a primer - extension assay using an M13 ssDNA template in the presence or absence of an oligonucleotide that hybridizes upstream of the extending primer and blocks its progression. Other enzymes that can displace the blocking primer in this assay are, in some cases, useful for the disclosed methods. In some cases, the polymerase incorporates dNTPs and terminators in approximately equal ratios.In some cases, the ratio of the incorporation rate of dNTPs and terminators for the polymerase described herein is about 1:1, about 1.5:1, about 2:1, about 3:1, about 4:1, about 5:1, about 10:1, about 20:1, about 50:1, about 100:1, about 200:1, about 500:1, or about 1000:1. In some cases, the ratio of the incorporation rate of dNTPs and terminators for the polymerase described herein is from 1:1 to 1000:1, from 2:1 to 500:1, from 5:1 to 100:1, from 10:1 to 1000:1, from 100:1 to 1000:1, from 500:1 to 2000:1, from 50:1 to 1500:1, or from 25:1 to 1000:1.
[0041] Described herein are methods of amplification in which strand displacement can be facilitated through the use of a strand displacement factor such as, for example, a helicase. Such factors are, in some cases, used in combination with additional amplification components such as a polymerase, a terminator, or other components. In some cases, the strand displacement factor is used with a polymerase that does not have strand displacement activity. In some cases, the strand displacement factor is used with a polymerase that has strand displacement activity. Without being bound by theory, the strand displacement factor may increase the rate at which smaller double-stranded amplicons are re-primed. In some cases, any DNA polymerase capable of performing strand displacement replication in the presence of a strand displacement factor is suitable for use in the PTA method, even if the DNA polymerase does not perform strand displacement replication in the absence of such a factor.In some cases, strand displacement factors useful in strand displacement replication include the BMRF1 polymerase accessory subunit (Tsurumi et al., J. Virology 67(12):7648-7653(1993)), the adenovirus DNA binding protein (Zijderveld and van der Vliet, J. Virology 68(2):1158-1164(1994)), the herpes simplex virus protein ICP8 (Boehmer and Lehman, J. Virology 67(2):711-715(1993); Skaliter and Lehman, Proc. Natl. Acad. Sci. USA 91(22):10665-10669(1994)); single-stranded DNA binding protein (SSB; Rigler and Romano, J. Biol. Chem. 270:8910-8919(1995)); phage T4 gene 32 protein (Villemain and Giedroc, Biochemistry 35:14395-14404(1996); T7 helicase-primase; T7 gp2.5 SSB protein; Tte-UvrD (from Thermoanaerobacter tengcongensis), calf thymus helicase (Siegel et al., J. Biol. Chem. 267:13629-13635(1992)); bacterial SSB (e.g., E. coli SSB), replication protein A (RPA) in eukaryotes, human mitochondrial SSB (mtSSB), and recombinases (e.g., recombinase A (RecA) family proteins, T4 UvsX, T4 UvsY, Sak4 of phage HK620, Rad51, Dmc1, or Radb) (but not limited to these). Combinations of factors that promote strand displacement and priming are also consistent with the methods described herein. For example, helicases are used with polymerases.In some cases, the PTA method involves the use of single-stranded DNA binding proteins (SSB, T4 gp32, or other single-stranded DNA binding proteins), helicases, and polymerases (e.g., SauDNA polymerase, Bsu polymerase, Bst2.0, GspM, GspM2.0, GspSSD, or other suitable polymerases). In some cases, reverse transcriptase is used in combination with the strand displacement factors described herein. Reverse transcriptase is used in combination with the strand displacement factors described herein. In some cases, amplification is performed using a polymerase and a nicking enzyme (e.g., "NEAR") as described in U.S. Patent No. 9,617,586. In some cases, the nicking enzyme is Nt.BspQI, Nb.BbvCi, Nb.BsmI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nt.BbvCI, Nt.BstNBI, Nt.CviPII, Nb.Bpu10I, or Nt.Bpu10I.
[0042] Described herein are amplification methods that include the use of terminator nucleotides, polymerases, and additional factors or conditions. For example, such factors are in some cases used to fragment nucleic acid templates or amplicons during amplification. In some cases, such factors include endonucleases. In some cases, the factors include transposases. In some cases, mechanical shearing is used to fragment nucleic acids during amplification. In some cases, nucleotides are added during amplification and these may be fragmented by the addition of further proteins or conditions. For example, uracil is incorporated into the amplicon and treatment with uracil D-glycosylase fragments the nucleic acid at positions containing uracil. Additional systems for selective nucleic acid fragmentation are also utilized in some examples, such as engineered DNA glycosylases that cleave modified cytosine-pyrene base pairs (Kwon, et al. Chem Biol. 2003, 10(4), 351).
[0043] Described herein is an amplification method that includes the use of terminator nucleotides, which terminate nucleic acid replication and thus reduce the size of the amplification product. Such terminators are, in some cases, used in combination with a polymerase, a strand displacement factor, or other amplification components described herein. In some cases, the terminator nucleotides reduce or decrease the efficiency of nucleic acid replication. Such terminators, in some cases, decrease the elongation rate by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. Such terminators, in some cases, decrease the elongation rate by 50% - 90%, 60% - 80%, 65% - 90%, 70% - 85%, 60% - 90%, 70% - 99%, 80% - 99%, or 50% - 80%. In some cases, the terminator decreases the average amplicon product length by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. The terminator, in some cases, decreases the average amplicon length by 50% - 90%, 60% - 80%, 65% - 90%, 70% - 85%, 60% - 90%, 70% - 99%, 80% - 99%, or 50% - 80%. In some cases, the amplicon containing the terminator nucleotide forms a loop or a hairpin, which decreases the ability of the polymerase using such an amplicon as a template. The use of the terminator, in some cases, slows the amplification rate at the initial amplification site through the incorporation of terminator nucleotides (e.g., dideoxynucleotides modified to be exonuclease-resistant to stop DNA elongation), resulting in smaller amplification products. By generating amplification products that are smaller than those of currently used methods (e.g., the average product length of the PTA method of 50 - 2000 nucleotides compared to the MDA method >10,000 nucleotides average product length), the PTA amplification products, in some cases, undergo direct ligation of adapters without the need for fragmentation, allowing for efficient incorporation of cell barcodes and unique molecular identifiers (UMIs) (see Figure 2A).
[0044] Terminator nucleotides are present at various concentrations depending on factors such as polymerase, template, or other factors. For example, the amount of terminator nucleotides is, in some cases, expressed as a ratio of non-terminator nucleotides to terminator nucleotides in the methods described herein. Such concentrations, in some cases, allow for control of amplicon length. In some cases, the ratio of terminator to non-terminator nucleotides is altered depending on the amount or size of the template present. In some cases, the ratio of terminator to non-terminator nucleotides decreases as the sample size decreases (e.g., in the femtogram to picogram range). In some cases, the ratio of non-terminator nucleotides to terminator nucleotides is about 2:1, 5:1, 7:1, 10:1, 20:1, 50:1, 100:1, 200:1, 500:1, 1000:1, 2000:1, or 5000:1. In some cases, the ratio of non-terminator to terminator nucleotides is 2:1 to 10:1, 5:1 to 20:1, 10:1 to 100:1, 20:1 to 200:1, 50:1 to 1000:1, 50:1 to 500:1, 75:1 to 150:1, or 100:1 to 500:1. In some cases, at least one of the nucleotides present during amplification using the methods described herein is a terminator nucleotide. Each terminator does not need to be present at approximately the same concentration, and in some cases, the ratio of each terminator present in the methods described herein is optimized for a particular set of reaction conditions, sample type, or polymerase. Without being bound by theory, each terminator may have different efficiencies for incorporation into the growing polynucleotide chain of the amplicon in response to pairing with the corresponding nucleotide on the template strand. For example, in some cases, the terminator that pairs with cytosine is present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration.In some cases, the terminator paired with thymine is present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, the terminator paired with guanine is present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, the terminator paired with adenine is present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, the terminator paired with uracil is present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, any nucleotide capable of terminating nucleic acid elongation by a nucleic acid polymerase is used as a terminator nucleotide in the methods described herein. In some cases, reversible terminators are used to terminate nucleic acid replication. In some cases, irreversible terminators are used to terminate nucleic acid replication. In some cases, non-limiting examples of terminators include reversible and irreversible nucleic acids and nucleic acid analogs, such as 3'-blocked reversible terminators containing nucleotides, 3'-unblocked reversible terminators containing nucleotides, terminators containing 2'-modifications of deoxynucleotides, terminators containing modifications to the nitrogenous bases of deoxynucleotides, or any combination thereof. In one embodiment, the terminator nucleotide is a dideoxynucleotide. Other nucleotide modifications suitable for terminating nucleic acid replication and practicing the present invention include, without limitation, reverse dideoxynucleotides, 3'-biotinylated nucleotides, 3'-aminonucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3'C3 spacer nucleotides, 3'C18 nucleotides, 3'-hexanediol spacer nucleotides including 3'-carbon spacer nucleotides, acyclonucleotides, and any modification of the r group of the 3'-carbon of deoxyribose such as combinations thereof.In some cases, the terminator is a polynucleotide comprising 1, 2, 3, 4, or more bases in length. In some cases, the terminator does not contain a detectable moiety or tag (e.g., a mass tag, a fluorescent tag, a dye, a radioactive atom, or other detectable moiety). In some cases, the terminator does not contain a chemical moiety that allows attachment of a detectable moiety or tag (e.g., a "click" azide / alkyne, a conjugate addition partner, or other chemical handle for attachment of a tag). In some cases, all terminator nucleotides contain the same modification that reduces amplification in the region of the nucleotide (e.g., the sugar moiety, the base moiety, or the phosphate moiety). In some cases, at least one terminator has a different modification that reduces amplification. In some cases, all terminators have substantially similar fluorescence excitation or emission wavelengths. In some cases, terminators with unmodified phosphate groups are used with polymerases that do not have exonuclease proofreading activity. When terminators are used with a polymerase having 3'->5' proofreading exonuclease activity (e.g., phi29) that can remove terminator nucleotides, in some cases they are further modified to be exonuclease resistant. For example, dideoxynucleotides are modified with alpha-thio groups that create phosphorothioate linkages that are resistant to the 3'->5' proofreading exonuclease activity of nucleic acid polymerases. Such modifications reduce the exonuclease proofreading activity of the polymerase by at least 99.5%, 99%, 98%, 95%, 90%, or at least 85% in some cases.Non-limiting examples of other terminator nucleotide modifications that provide resistance to 3’->5’ exonuclease activity include, in some instances, nucleotides having modifications to the alpha group, such as alpha-thiodideoxynucleotides that create phosphorothioate linkages, C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2’ fluoro bases, 3’ phosphorylations, 2’-O-methyl modifications (or other 2’-O-alkyl modifications), propynyl-modified bases (e.g., deoxycytosine, deoxyuridine), L-DNA nucleotides, L-RNA nucleotides, nucleotides having reverse linkages (e.g., 5’-5’ or 3’-3’), 5’ reverse bases (e.g., 5’ reverse 2’,3’-dideoxy dT), methylphosphonate backbones, and trans nucleic acids. In some cases, nucleotides having modifications include base-modified nucleic acids having a free 3’ OH group (e.g., bases containing modifications with large chemical groups such as 2-nitrobenzyl alkylated HOMedU triphosphate, solid supports, or other large moieties). In some cases, polymerases that have strand displacement activity but do not have 3’->5’ exonuclease proofreading activity are used with terminator nucleotides, with or without modifications for exonuclease resistance. Such nucleic acid polymerases include, without limitation, Bst DNA polymerase, Bsu DNA polymerase, Deep Vent (exo-) DNA polymerase, Klenow fragment (exo-) DNA polymerase, Therminator DNA polymerase, and Vent. R (exo-) is included.
[0045] Primers and amplicon libraries Described herein is an amplicon library resulting from the amplification of at least one target nucleic acid molecule. Such libraries are, in some cases, generated using the methods described herein, such as those that use terminators. Such methods include the use of strand-displacing polymerases or factors, terminator nucleotides (reversible or irreversible), or other features and embodiments described herein. In some cases, an amplicon library generated by the use of terminators described herein is further amplified in a subsequent amplification reaction (e.g., PCR). In some cases, the subsequent amplification reaction does not include a terminator. In some cases, the amplicon library contains polynucleotides, and at least 50%, 60%, 70%, 80%, 90%, 95%, or at least 98% of the polynucleotides contain at least one terminator nucleotide. In some cases, the amplicon library contains the target nucleic acid molecule from which the amplicon library is derived. The amplicon library contains a plurality of polynucleotides, and at least some of the polynucleotides are direct copies (e.g., directly replicated from a target nucleic acid molecule such as genomic DNA, RNA, or other target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 95% or more of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 5% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 10% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 15% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 20% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 50% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule.In some cases, 3% to 5%, 3 to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50%, or 15% to 75% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least some of the polynucleotides are direct copies of the target nucleic acid molecule or descendants of the daughter (the first copy of the target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 95% or more of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, at least 5% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, at least 10% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, at least 20% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, at least 30% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, 3% to 5%, 3% to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50 or 15% to 75% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule or descendants of the daughter. In some cases, the direct copy of the target nucleic acid is 50 to 2500, 75 to 2000, 50 to 2000, 25 to 1000, 50 to 1000, 500 to 2000, or 50 to 2000 bases in length. In some cases, the length of the descendants of the daughter is 1000 to 5000, 2000 to 5000, 1000 to 10,000, 2000 to 5000, 1500 to 5000, 3000 to 7000, or 2000 to 7000 bases in length. In some cases, the average length of the PTA amplification product is 25 to 3000 nucleotide lengths, 50 to 2500, 75 to 2000, 50 to 2000, 25 to 1000, 50 to 1000, 500 to 2000, or 50 to 2000 base lengths.In some cases, the amplicons generated from PTA are 5000, 4000, 3000, 2000, 1700, 1500, 1200, 1000, 700, 500 bases or less, or 300 bases or less. In some cases, the amplicons generated from PTA are 1000 - 5000, 1000 - 3000, 200 - 2000, 200 - 4000, 500 - 2000, 750 - 2500, or 1000 - 2000 bases in length. In some cases, the amplicon library generated using the methods described herein contains at least 1000, 2000, 5000, 10,000, 100,000, 200,000, 500,000, or more than 500,000 amplicons containing unique sequences. In some cases, the library contains at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 2000, 2500, 3000, or at least 3500 amplicons. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of less than 1000 bases are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 2000 bases or less are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 3000 - 5000 bases are direct copies of at least one target nucleic acid molecule. In some cases, the ratio of direct copy amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or more than 10,000,000:1.In some cases, the ratio of the direct copy amplicon to the target nucleic acid molecule is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, where the direct copy amplicon is 700 to 1200 bases in length or less. In some cases, the ratio of the direct copy amplicon and the daughter amplicon to the target nucleic acid molecule is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1. In some cases, the ratio of the direct copy amplicon and the daughter amplicon to the target nucleic acid molecule is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, where the direct copy amplicon is 700 to 1200 bases in length and the daughter amplicon is 2500 to 6000 bases in length. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule or the daughter amplicon. The number of direct copies can, in some cases, be controlled by the number of PCR amplification cycles. In some cases, 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or 3 or fewer PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, about 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or about 3 PCR cycles are used to generate copies of the target nucleic acid molecule.In some cases, 3, 4, 5, 6, 7, or 8 PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, 2 - 4, 2 - 5, 2 - 7, 2 - 8, 2 - 10, 2 - 15, 3 - 5, 3 - 10, 3 - 15, 4 - 10, 4 - 15, 5 - 10, or 5 - 15 PCR cycles are used to generate copies of the target nucleic acid molecule. The amplicon libraries generated using the methods described herein are, in some cases, subjected to additional steps such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed prior to the sequencing step.
[0046] The methods described herein may further comprise one or more enrichment or purification steps. In some cases, one or more polynucleotides (such as cDNA, PTA amplicons, or other polynucleotides) are enriched during the methods described herein. In some cases, polynucleotide probes are used to capture one or more polynucleotides. In some cases, the probes are configured to capture one or more genomic exons. In some cases, the library of probes comprises at least 1000, 2000, 5000, 10,000, 50,000, 100,000, 200,000, 500,000, or more than 1 million different sequences. In some cases, the library of probes comprises sequences capable of binding to at least 10, 20, 50, 100, 200, 500, 1000, 2000, 5000, 10,000, or more than 10,000 genes. Optionally, the probe comprises a moiety for capture by a solid support such as biotin. In some cases, an enrichment step is performed after the PTA step. In some cases, an enrichment step is performed before the PTA step. In some cases, the probe is configured to bind to a genomic DNA library. In some cases, the probe is configured to bind to a cDNA library.
[0047] In some cases, the amplicon libraries of polynucleotides generated from the PTA methods and compositions (terminators, polymerases, etc.) described herein have increased uniformity. Uniformity is described in some cases using a Lorenz curve (e.g., FIG. 5C) or other such methods. Such an increase results in some cases in lower sequencing reads required for desired coverage of a target nucleic acid molecule (e.g., genomic DNA, RNA, or other target nucleic acid molecule). For example, 50% or less of the cumulative percentage of polynucleotides contains at least 80% of the sequence of the cumulative percentage of the sequence of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contains at least 60% of the sequence of the cumulative percentage of the sequence of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contains at least 70% of the sequence of the cumulative percentage of the sequence of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contains at least 90% of the sequence of the cumulative percentage of the sequence of the target nucleic acid molecule. In some cases, uniformity is described using the Gini index (where an index of 0 represents perfect equality of the library and an index of 1 represents perfect inequality). In some cases, the amplicon libraries described herein have a Gini index of 0.55, 0.50, 0.45, 0.40, or 0.30 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.40 or less. In some cases, such a measure of uniformity depends on the number of reads obtained. For example, 100 million, 200 million, 300 million, 400 million, or 500 million or fewer reads are obtained. In some cases, the read length is about 50, 75, 100, 125, 150, 175, 200, 225, or about 250 bases in length. In some cases, the measure of uniformity depends on the coverage depth of the target nucleic acid. For example, the average depth of coverage is about 10-fold, 15-fold, 20-fold, 25-fold, or about 30-fold.In some cases, the average coverage depth is 10 to 30 times, 20 to 50 times, 5 to 40 times, 20 to 60 times, 5 to 20 times, or 10 to 20 times. In some cases, the amplicon library described herein has a Gini index of 0.55 or less and approximately 300 million reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.50 or less and approximately 300 million reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.45 or less and approximately 300 million reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.55 or less and 300 million or fewer reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.50 or less and 300 million or fewer reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.45 or less and 300 million or fewer reads were obtained. In some cases, the amplicon library described herein has a Gini index of 0.55 or less and the average depth of sequencing coverage is approximately 15 times. In some cases, the amplicon library described herein has a Gini index of 0.50 or less and the average depth of sequencing coverage is approximately 15 times. In some cases, the amplicon library described herein has a Gini index of 0.45 or less and the average depth of sequencing coverage is approximately 15 times. In some cases, the amplicon library described herein has a Gini index of 0.55 or less and the average depth of sequencing coverage is at least 15 times. In some cases, the amplicon library described herein has a Gini index of 0.50 or less and the average depth of sequencing coverage is at least 15 times. In some cases, the amplicon library described herein has a Gini index of 0.45 or less and the average depth of sequencing coverage is at least 15 times. In some cases, the amplicon library described herein has a Gini index of 0.55 or less, where the average depth of sequencing coverage is 15 times or less.In some cases, the Gini index of the amplicon library described herein is 0.50 or less, and the average depth of sequencing coverage is 15-fold or less. In some cases, the Gini index of the amplicon library described herein is 0.45 or less, and the average depth of sequencing coverage is 15-fold or less. The uniform amplicon library generated using the methods described herein is, in some cases, subjected to additional steps such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed prior to the sequencing step.
[0048] Primers include nucleic acids used to prime the amplification reactions described herein. Such primers include, in some cases, random deoxynucleotides of any length, with or without modifications for exonuclease resistance, random ribonucleotides of any length, with or without modifications for exonuclease resistance, modified nucleic acids such as locked nucleic acids, or DNA or RNA primers that target reactions primed by specific genomic regions and enzymes such as primase, but are not limited thereto. For whole genome PTA, it is preferred that a set of primers with random or partially random nucleotide sequences be used. In nucleic acid samples of significant complexity, the specific nucleic acid sequences present in the sample need not be known, and the primers need not be designed to be complementary to a specific sequence. Rather, the complexity of the nucleic acid sample results in a large number of different hybridization target sequences in the sample, which are complementary to various primers of random or partially random sequences. The complementary portions of the primers for use in PTA are, in some cases, completely randomized, contain only randomized portions, or are otherwise selectively randomized. In some cases, the number of random base positions in the complementary portion of the primer is, for example, 20% to 100% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is 10% to 90%, 15 to 95%, 20% to 100%, 30% to 100%, 50% to 100%, 75 to 100%, or 90 to 95% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the total number of nucleotides in the complementary portion of the primer.A set of primers having a random or partially random array can, in some cases, be synthesized using standard techniques by allowing the addition of any nucleotide at each position to be randomized. In some cases, the set of primers is composed of primers of similar length and / or hybridization characteristics. In some cases, the term "random primer" refers to a primer that can exhibit a four-fold degeneracy at each position. In some cases, the term "random primer" refers to a primer that can exhibit a three-fold degeneracy at each position. Random primers used in the methods described herein can, in some cases, include a random array of bases that is 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more bases in length. In some cases, the primer includes a random array that is 3-20, 5-15, 5-20, 6-12, or 4-10 bases in length. The primer can also include a non-extendable element that limits the subsequent amplification of the amplicon generated therefrom. For example, a primer having a non-extendable element can, in some cases, include a terminator. In some cases, the primer includes terminator nucleotides such as 1, 2, 3, 4, 5, 10, or more than 10 terminator nucleotides. The primer need not be limited to components added externally to the amplification reaction. In some cases, the primer is generated in situ by the addition of nucleotides and proteins that promote priming. For example, a primase-like enzyme in combination with nucleotides can, in some cases, be used to generate random primers for the methods described herein. The primase-like enzyme can, in some cases, be a member of the DnaG or AEP enzyme superfamily. In some cases, the primase-like enzyme is TthPrimPol. In some cases, the primase-like enzyme is T7gp4 helicase-primase. Such primases can, in some cases, be used in conjunction with the polymerases or strand displacement factors described herein.In some cases, primase initiates priming using deoxyribonucleotides. In some cases, primase initiates priming using ribonucleotides.
[0049] Following PTA amplification, a specific subset of amplicons can be selected. Such selection depends, in some cases, on size, affinity, activity, hybridization to a probe, or other known selection factors in the art. In some cases, the selection is performed before or after additional steps described herein, such as adapter ligation and / or library amplification. In some cases, the selection is based on the size (length) of the amplicon. In some cases, smaller amplicons that are less likely to have undergone exponential amplification are selected, which enriches for products derived from the primary template while further converting the amplification from an exponential amplification process to a quasi-linear amplification process (Figure 1A). In some cases, amplicons with lengths of 50 - 2000, 25 - 5000, 40 - 3000, 50 - 1000, 200 - 1000, 300 - 1000, 400 - 1000, 400 - 600, 600 - 2000, or 800 - 1000 bases are selected. Size selection is performed, in some cases, by use of a protocol that utilizes solid-phase reversible immobilization (SPRI) on carboxylated paramagnetic beads, for example, to enrich for nucleic acid fragments of a particular size, or by use of other protocols known to those of skill in the art. Optionally, or in combination, selection is performed through preferential ligation and amplification of smaller fragments during PCR while preparing a sequencing library, and through preferential formation of clusters from smaller sequencing library fragments during sequencing (e.g., sequencing by synthesis, nanopore sequencing, or other sequencing methods). Other strategies for selecting smaller fragments are also consistent with the methods described herein and include, without limitation, isolation of nucleic acid fragments of a particular size after gel electrophoresis, use of silica columns that bind nucleic acid fragments of a particular size, and other PCR strategies that more strongly enrich for smaller fragments. Any number of library preparation protocols can be used with the PTA method described herein.The amplicons generated by PTA are, in some cases, ligated to adapters (optionally with removal of terminator nucleotides). In some cases, the amplicons generated by PTA contain regions of homology generated from transposase-based fragmentation used as priming sites. In some cases, libraries are prepared by mechanically or enzymatically fragmenting nucleic acids. In some cases, libraries are prepared using transposome-mediated tagging. In some cases, libraries are created by ligation of adapters such as Y-adapters, universal adapters, circular adapters.
[0050] The non-complementary portions of the primers used in PTA can include sequences that can be used to further manipulate and / or analyze the amplified sequences. Examples of such sequences are "detection tags". Detection tags have sequence determinations complementary to detection probes and are detected using their cognate detection probes. There can be one, two, three, four, or more than four detection tags on a primer. There is no fundamental limit to the number of detection tags that can be present on a primer, except for the size of the primer. In some cases, there is a single detection tag on a primer. In some cases, there are two detection tags on a primer. When multiple detection tags are present, they can have the same sequence or they can have different sequences, each different sequence being complementary to a different detection probe. In some cases, multiple detection tags have the same sequence. In some cases, multiple detection tags have different sequences.
[0051] Another example of a sequence that can be included in the non-complementary portion of a primer is an "address tag" that can encode other details of the amplicon, such as its location within a tissue section. In some cases, the cell barcode includes an address tag. The address tag has a sequence complementary to an address probe. The address tag is incorporated at the end of the amplified strand. If present, there may be one or more address tags in the primer. Other than the size of the primer, there is no fundamental limit to the number of address tags that can be present in the primer. If multiple address tags are present, they may have the same sequence or different sequences, and each different sequence is complementary to a different address probe. The address tag portion can be of any length that supports specific and stable hybridization between the address tag and the address probe. In some cases, nucleic acids from more than one source can incorporate a variable tag sequence. This tag sequence can be up to 100 nucleotides in length, preferably 1 to 10 nucleotides in length, most preferably 4, 5, or 6 nucleotides in length, and includes combinations of nucleotides. In some cases, the tag sequence is 1 to 20, 2 to 15, 3 to 13, 4 to 12, 5 to 12, or 1 to 10 nucleotides in length. For example, six base pairs are selected to form a tag, and the permutations of four different nucleotides are used, and then a total of 4096 nucleic acid anchors (such as hairpins) each having a unique six-base pair can be created.
[0052] The primers described herein can be present in solution or immobilized on a solid support. In some cases, primers having sample barcodes and / or UMI sequences can be immobilized on a solid support. The solid support can be, for example, one or more beads. In some cases, to identify individual cells, individual cells are contacted with one or more beads having a unique set of sample barcodes and / or UMI sequences. In some cases, lysates from individual cells are contacted with one or more beads having a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. In some cases, nucleic acids extracted from individual cells are contacted with one or more beads having a unique set of sample barcodes and / or UMI sequences to identify nucleic acids extracted from individual cells. The beads can be manipulated by any suitable method known in the art, for example, using the droplet actuators described herein. The beads can be of any suitable size, including, for example, microbeads, microparticles, nanobeads, and nanoparticles. In some embodiments, the beads are magnetically responsive, and in other embodiments, the beads are not significantly magnetically responsive.Non-limiting examples of suitable beads include flow cytometry microbeads, polystyrene microparticles and nanoparticles, functionalized polystyrene microparticles and nanoparticles, coated polystyrene microparticles and nanoparticles, silica microbeads, fluorescent microspheres and nanospheres, functionalized fluorescent microspheres and nanospheres, coated fluorescent microspheres and nanospheres, colored microparticles and nanoparticles, magnetic microparticles and nanoparticles, superparamagnetic microparticles and nanoparticles (e.g., DYNABEADS® available from Invitrogen Group, Carlsbad, CA), fluorescent microparticles and nanoparticles, coated magnetic microparticles and nanoparticles, ferromagnetic microparticles and nanoparticles, coated ferromagnetic microparticles and nanoparticles, and those described in U.S. Patent Application Publication Nos. US20050260686, US20030132538, US20050118574, 20050277197, 20060159962. The beads can be pre-bound with an antibody, protein or antigen, DNA / RNA probe, or any other molecule having an affinity for a desired target. In some embodiments, primers having sample barcodes and / or UMI sequences can be in solution. In certain embodiments, a plurality of droplets can be presented, and each droplet in the plurality of droplets has a sample barcode unique to the droplet and a UMI unique to the molecule such that the UMI is repeated multiple times in the collection of droplets. In some embodiments, individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify the individual cells. In some embodiments, lysates from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify the individual cell lysates. In some embodiments, nucleic acids extracted from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify the nucleic acids extracted from the individual cells.
[0053] The PTA primer may include sequence-specific or random primers, cell barcodes, and / or unique molecular identifiers (UMIs) (see, e.g., FIGS. 10A (linear primer) and 10B (hairpin primer)). In some cases, the primer includes a sequence-specific primer. In some cases, the primer includes a random primer. In some cases, the primer includes a cell barcode. In some cases, the primer includes a sample barcode. In some cases, the primer includes a unique molecular identifier. In some cases, the primer includes two or more cell barcodes. Such barcodes identify, in some cases, a unique sample source or a unique workflow. Such barcodes or UMIs are, in some cases, longer than 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or 30 bases. In some cases, the primer is at least 1000, 10,000, 50,000, 100,000, 250,000, 500,000, 10 6 、10 7 、10 8 、10 9 、or at least 10 10Contains individual unique barcodes or UMIs. In some cases, the primers contain at least 8, 16, 96, or 384 unique barcodes or UMIs. In some cases, standard adapters are ligated to the amplification products prior to sequencing, and after sequencing, reads are first assigned to specific cells based on the cell barcodes. Suitable adapters available for use with the PTA method include, for example, the xGen® Dual Index UMI Adapter available from Integrated DNA Technologies (IDT). Next, reads from each cell are grouped using the UMIs, and reads associated with the same UMI are collapsed into a consensus read. The use of cell barcodes allows all cells to be pooled prior to library preparation because they can be identified later by the cell barcodes. The use of UMIs to form consensus reads is, in some cases, corrected for PCR bias and improves the detection of copy number variation (CNV) (Figures 11A and 11B). In addition, the percentage of reads from the same molecule that are required to have the same base change detected at each position can correct for sequencing errors. This approach has been used to improve CNV detection and correct sequencing errors in bulk samples. In some cases, UMIs are used with the methods described herein. For example, U.S. Patent No. 8,835,358 discloses the principle of digital counting after attaching randomly amplifiable barcodes. Schmitt et al. and Fan et al. have disclosed similar methods for correcting sequencing errors. In some cases, libraries are generated for sequencing using primers. In some cases, libraries are composed of fragments that are 200-700 bases in length, 100-1000, 300-800, 300-550, 300-700, or 200-800 bases. In some cases, libraries contain fragments that are at least 50, 100, 150, 200, 300, 500, 600, 700, 800, or at least 1000 bases in length.In some cases, the library contains fragments that are approximately 50, 100, 150, 200, 300, 500, 600, 700, 800, or approximately 1000 bases in length.
[0054] The methods described herein may further include additional steps, including steps performed on a sample or template. Such a sample or template may be subjected to one or more steps prior to PTA. In some cases, a sample containing cells is subjected to a pretreatment step. For example, cells are lysed and proteolyzed using a combination of freeze-thaw, Triton X-100, Tween 20, and Proteinase K to increase chromatin accessibility. Other lysis strategies are also suitable for practicing the methods described herein. Such strategies include, but are not limited to, lysis using detergents and / or lysozyme and / or protease treatment and / or physical disruption of cells such as sonication and / or other combinations of alkaline lysis and / or hypotonic lysis. In some cases, the primary template or target molecule is subjected to a pretreatment step. In some cases, the primary template (or target) is denatured using sodium hydroxide and then the solution is neutralized. Other denaturation strategies may also be appropriate for practicing the methods described herein. Such strategies include, but are not limited to, combinations of alkaline lysis with other basic solutions, increasing the temperature of the sample and / or changing the salt concentration in the sample, adding additives such as solvents or oils, other modifications, or any combination thereof. In some cases, additional steps include sorting, filtering, or separating the sample, template, or amplicon by size. For example, after amplification using the methods described herein, the amplicon library is enriched for amplicons having a desired length. In some cases, the amplicon library is enriched for amplicons having a length of 50-2000, 25-1000, 50-1000, 75-2000, 100-3000, 150-500, 75-250, 170-500, 100-500, or 75-2000 bases.In some cases, the amplicon library is enriched for amplicons having a length of 75, 100, 150, 200, 500, 750, 1000, 2000, 5000, or 10,000 bases or less. In some cases, the amplicon library is enriched for amplicons having a length of at least 25, 50, 75, 100, 150, 200, 500, 750, 1000, or at least 2000 bases. 。
[0055] The methods and compositions described herein may include buffers or other formulations. Such buffers are, in some cases, used for PTA, RT, or other methods described herein. Such buffers may, in some cases, include surfactants / detergents or denaturants (Tween-20, DMSO, DMF, pegylated polymers containing hydrophobic groups, or other surfactants), salts (potassium phosphate or sodium phosphate (monobasic or dibasic), sodium chloride, potassium chloride, TrisHCl, magnesium chloride or sulfate, ammonium salts such as phosphate, nitrate, sulfate, EDTA), reducing agents (DTT, THP, DTE, beta-mercaptoethanol, TCEP, or other reducing agents) or other components (hydrophilic polymers such as glycerol, PEG). In some cases, the buffer is used in combination with components such as polymerase, strand displacement factor, terminator, or other reaction components described herein. In some cases, the buffer is used in combination with components such as polymerase, strand displacement factor, terminator, or other reaction components described herein. The buffer may include one or more crowding agents. In some cases, the crowding reagent includes a polymer. In some cases, the crowding reagent includes a polymer such as a polyol. In some cases, the crowding reagent includes polyethylene glycol (PEG). In some cases, the crowding reagent includes a polysaccharide. By way of non-limiting example, crowding reagents include ficoll (e.g., ficoll PM400, ficoll PM70, or ficoll of other molecular weights), PEG (e.g., PEG1000, PEG2000, PEG4000, PEG6000, PEG8000, or PEG of other molecular weights), dextran (dextran 6, dextran 10, dextran 40, dextran 70, dextran 6000, dextran 138k, or dextran of other molecular weights).
[0056] The nucleic acid molecules amplified according to the methods described herein can be sequenced and analyzed using methods known to those of skill in the art. Non-limiting examples of sequencing methods used in some cases include, for example, sequencing by hybridization (SBH), sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728), quantitative incremental fluorescence nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads (U.S. Patent No. 7,425,431), wobble sequencing (International Patent Application Publication No. WO2006 / 073504), multiplex sequencing (U.S. Patent Application Publication No. US2008 / 0269068; Porreca et al., 2007, Nat. Methods 4:931), polymerase colony (POLONY) sequencing (U.S. Patents Nos. 6,432,360, 6,485,944, and 6,511,803, and International Patent Application Publication No. WO2005 / 082098), nanogrid rolling circle sequencing (ROLONY) (U.S. Patent No. 9,624,538), allele-specific oligoligation assays (e.g., oligoligation assay (OLA), single template molecule OLA using ligated linear probes and rolling circle amplification (RCA) readout, ligated padlock probes, and / or single template molecule OLA using ligated circular padlock probes and rolling circle amplification (RCA) readout), for example, high-throughput sequencing methods such as those using Roche 454, Illumina Solexa, AB-SOLiD, Helicos, Polonator platforms, etc., and methods using light-based sequencing technologies (Landegren et al. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmacogenomics 1:95-100; and Shi (2001) Clin. Chem. 47:164-172).In some cases, the amplified nucleic acid molecules are shotgun sequenced. Sequencing of the sequencing library is performed, in some cases, using a suitable sequencing technology including, but not limited to, single molecule real time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis (array / colony-based or nanoball-based).
[0057] Sequencing a sequencing library generated using the methods described herein (e.g., PTA or RNAseq) can obtain a desired number of sequencing reads. In some cases, the library is generated from a single cell or a sample containing single cells (either alone or as part of a multi-omics workflow). In some cases, the library is sequenced to obtain at least 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or at least 10 million reads. In some cases, the library is sequenced to obtain 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or 10 million or fewer reads. In some cases, the library is sequenced to obtain approximately 0.1, 0.2, 0.4, 0.5, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.5, 2, 5, or approximately 10 million reads. In some cases, the library is sequenced to obtain 0.1 - 10, 0.1 - 5, 0.1 - 1, 0.2 - 1, 0.3 - 1.5, 0.5 - 1, 1 - 5, or 50 - 5 million reads per sample. In some cases, the number of reads depends on the size of the genome. In some cases, a sample containing a bacterial genome is sequenced to obtain 50 - 1 million reads. In some cases, the library is sequenced to obtain at least 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or at least 900 million reads. In some cases, the library is sequenced to obtain 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or 900 million or fewer reads. In some cases, the library is sequenced to obtain approximately 2, 4, 10, 20, 50, 100, 200, 300, 500, 700, or approximately 900 million reads. In some cases, a sample containing a mammalian genome is sequenced to obtain 500 million to 600 million reads.In some cases, the type of sequencing library (cDNA library or genomic library) is identified during sequencing. In some cases, cDNA libraries and genomic libraries are identified during sequencing using unique barcodes.
[0058] As used herein, the term "cycle" with respect to polymerase-mediated amplification reactions refers to the dissociation of at least a portion of a double-stranded nucleic acid (e.g., template from an amplicon, or double-stranded template, denaturation), hybridization (annealing) of at least a portion of a primer to a template, and extension of the primer to generate an amplicon. In some cases, during cycles of amplification (such as isothermal reactions), the temperature remains constant. In some cases, the number of cycles is directly correlated with the number of amplicons generated. In some cases, the number of cycles of an isothermal reaction is controlled by the length of time the reaction is allowed to proceed.
[0059] Methods and Applications Described herein is a method for identifying mutations in a cell using a method of multi-omic analysis PTA such as single cells. The use of the PTA method results in an improvement over known methods such as the MDA method in some cases. PTA has a lower false positive and false negative variant calling rate than the MDA method in some cases. Genomes such as the NA12878 platinum genome are used in some cases to determine whether the false negative variant calling rate decreases as the genome coverage (scope of application) and the uniformity of PTA increase. Although not bound by theory, the lack of error propagation in PTA may be determined to reduce the false positive variant calling rate. The amplification balance between alleles by the two methods is estimated in some cases by comparing the allele frequencies of heterozygous variant calls at known positive loci. In some cases, the amplicon library generated using PTA is further amplified by PCR. In some cases, PTA is used in a workflow using additional analytical methods such as RNAseq, methylome analysis, or other methods described herein.
[0060] Cells analyzed using the methods described herein may, in some cases, include tumor cells. For example, circulating tumor cells can be isolated from fluids collected from a patient, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, or aqueous humor. The cells are then subjected to the methods described herein (e.g., PTA) and sequencing to determine the mutation load and combination of mutations in each cell. These data are, in some cases, used as tools for diagnosing a particular disease or predicting a treatment response. Similarly, in some cases, cells of unknown malignancy potential are isolated from body fluids collected from a patient, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, aqueous humor, cleavage cavity fluid, or the collection medium surrounding cells in culture. In some cases, the sample is obtained from the collection medium surrounding embryonic cells. After utilizing the methods and sequencing described herein, such methods are further used to determine the mutation load and combination of mutations in each cell. These data are, in some cases, used as tools for diagnosing a particular disease or predicting progression from a pre-cancerous state to overt malignancy. In some cases, cells can be isolated from a primary tumor sample. The cells are then subjected to PTA and sequencing to determine the mutation load and combination of mutations in each cell. These data are, in some cases, used as tools for diagnosing a particular disease or predicting the probability that a patient's malignancy is resistant to available anticancer agents. By exposing samples to different chemotherapeutic agents, it has been found that major and minor clones have differential sensitivity to a particular agent and do not necessarily correlate with the presence of known "driver mutations", suggesting that the combination of mutations within a clonal population determines its sensitivity to a particular chemotherapeutic drug. Without being bound by theory, these findings suggest that if pre-cancerous lesions are detected that have not yet expanded and may evolve into clones that become more resistant to treatment as the number of genomic modifications increases, it may be easier to eradicate the malignancy.See Ma et al., 2018, 「Pan-cancer genome and transcriptome analyses of 1,699 pediatric leukemias and solid tumors.」 Single-cell genomics protocols are used in some cases to detect combinations of somatic genetic variants in single cancer cells or clonotypes within a mixture of normal and malignant cells isolated from a patient's sample. This technique is further utilized in some cases to identify clonotypes that undergo positive selection after exposure to drugs both in vitro and / or in patients. As shown in Figure 6A, by comparing the surviving clones exposed to chemotherapy with the clones identified at the time of diagnosis, a catalog of cancer clonotypes documenting resistance to specific agents can be created. The PTA method can detect, in some cases, the sensitivity of specific clones, as well as combinations thereof, within a sample composed of multiple clonotypes to existing or new drugs, where this method can detect the sensitivity of specific clones to drugs. This approach can, in some cases, demonstrate the efficacy of a drug against specific clones that may not be detected using current drug sensitivity measurements that consider the sensitivity of all cancer clones together in a single measurement. When the PTA described herein is applied to a patient sample collected at the time of diagnosis to detect cancer clonotypes in a given patient's cancer, the drug sensitivity catalog is used to search for those clones, thereby providing the oncologist with information on which drugs or combinations of drugs are non-functional and which drugs or combinations of drugs are most likely to be effective against that patient's cancer. PTA can be used for the analysis of samples containing populations of cells. In some cases, the sample contains neurons or glial cells. In some cases, the sample contains nuclei.
[0061] Described herein is a method for measuring changes in gene expression in combination with the mutagenicity of environmental factors. For example, cells (single or population) are exposed to potential environmental conditions. For example, cells derived from organs (liver, pancreas, lung, colon, thyroid, or other organs), tissues (skin, or other tissues), blood, or other biological sources are used in this method in some cases. In some cases, the environmental conditions include heat, light (e.g., ultraviolet light), radiation, chemicals, or any combination thereof. In some cases, after exposure to environmental conditions for a length of minutes, hours, days, or more, single cells are isolated and subjected to the PTA method. In some cases, samples are tagged using molecular barcodes and unique molecular identifiers. The samples are sequenced and then analyzed to identify changes in gene expression resulting from mutations arising from exposure to environmental conditions. In some cases, such mutations are compared to control environmental conditions such as known non-mutagenic substances, vehicle / solvent, or absence of environmental conditions. Such analysis provides, in some cases, not only the total number of mutations caused by environmental conditions, but also the location and nature of such mutations. Patterns are identified in some cases from the data and can be used for the diagnosis of diseases or conditions. In some cases, patterns can be used to predict future medical conditions or states. In some cases, the methods described herein measure the mutation load, location, and pattern in cells after exposure to environmental factors such as potential mutagens or teratogens. This approach is used in some cases to evaluate the safety of a given drug, including its potential ability to induce mutations that may contribute to the development of disease. For example, this method can be used to predict the carcinogenic or teratogenic potential of a drug on a specific cell type after exposure to a specific concentration of the drug.
[0062] Disclosed herein is a method for identifying gene expression changes in combination with mutations in animal, plant, or microbial cells that have undergone genome editing (e.g., using CRISPR technology). In some cases, such cells are isolated and subjected to PTA and sequencing to determine the mutation load and combination of mutations in each cell. The mutation rate and location of mutations for each cell resulting from a genome editing protocol are used in some cases to assess the safety of a given genome editing method.
[0063] Disclosed herein is a method for determining changes in gene expression in combination with mutations in cells used for cell therapy, such as, but not limited to, transplantation of induced pluripotent stem cells, transplantation of unmanipulated hematopoietic or other cells, or transplantation of genome-edited hematopoietic or other cells. The cells are then subjected to PTA and sequencing to determine the mutation load and combination of mutations in each cell. The mutation rate and location of mutations for each cell in a cell therapy product can be used to assess the safety and potential efficacy of the product.
[0064] Cells for use with the PTA method can be fetal cells such as embryonic cells. In some embodiments, PTA is used in combination with non-invasive preimplantation genetic testing (NIPGT). In further embodiments, the cells can be isolated from blastomeres created by in vitro fertilization. The cells can then be subjected to PTA and sequencing to determine the burden and combination of genetic mutations that are potential disease predispositions in each cell. Next, changes in gene expression combined with the mutation profile of the cells can be used to extrapolate the genetic predisposition of the blastomeres to specific diseases prior to implantation. In some cases, embryos in culture release nucleic acids that are used to assess the health of the embryo using low-pass genome sequencing. In some cases, the embryos are cryopreserved. In some cases, the nucleic acids are obtained from undifferentiated germ cell culture-conditioned medium (BCCM), blastocoel fluid (BF), or combinations thereof. In some cases, PTA analysis of fetal cells is used to detect chromosomal abnormalities such as aneuploidy in the fetus. In some cases, PTA is used to detect diseases such as Down syndrome or Patau syndrome. In some cases, frozen undifferentiated germ cells are thawed and cultured for a period of time before nucleic acids for analysis are obtained (e.g., medium, BF, or cell biopsy). In some cases, the undifferentiated germ cells are cultured for within 4, 6, 8, 12, 16, 24, 36, 48 hours, or within 64 hours, before nucleic acids for analysis are obtained.
[0065] In another embodiment, microbial cells (e.g., bacteria, fungi, protozoa) can be isolated from plants or animals (e.g., from a microbiome sample [e.g., gut microbiome, skin microbiome, etc.], or from body fluids such as, for example, blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural fluid, pericardial fluid, ascites, or aqueous humor). Further, the microbial cells may be isolated from indwelling medical devices such as, but not limited to, intravenous catheters, urinary catheters, cerebrospinal shunts, artificial valves, artificial joints, or endotracheal tubes. Next, the cells can undergo PTA and sequencing to determine the identity of the specific microorganism and to detect the presence of genetic variants of the microorganism that predict response (or resistance) to specific antimicrobial agents. These data can be used as tools for the diagnosis of specific infectious diseases and / or for predicting treatment response.
[0066] Described herein is a method for generating an amplicon library from a sample containing short nucleic acids using the PTA method described herein. In some cases, PTA results in improved fidelity and uniformity of amplification of shorter nucleic acids. In some cases, the length of the nucleic acid is 2000 bases or less. In some cases, the length of the nucleic acid is 1000 bases or less. In some cases, the length of the nucleic acid is 500 bases or less. In some cases, the length of the nucleic acid is 200, 400, 750, 1000, 2000, or 5000 bases or less. In some cases, samples containing short nucleic acid fragments include, but are not limited to, ancient DNA (hundreds, thousands, millions, or even billions of years old), FFPE (formalin-fixed paraffin-embedded) samples, cell-free DNA, or other samples containing short nucleic acids.
[0067] Embodiment Described herein is a method for amplifying a target nucleic acid molecule, the method comprising: a) contacting a sample comprising a mixture of nucleotides comprising the target nucleic acid molecule, one or more amplification primers, a nucleic acid polymerase, and one or more terminator nucleotides that terminate nucleic acid replication by the polymerase; and b) incubating the sample under conditions that promote replication of the target nucleic acid molecule to obtain a plurality of terminated amplification products, wherein the replication proceeds by strand displacement replication. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product having a length between about 50 and about 2000 nucleotides. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product having a length between about 400 and about 600 nucleotides. In one embodiment of any of the above methods, the method further comprises: c) repairing the termini and A-tailing; and d) ligating the molecules obtained in step (c) to an adapter, thereby generating a library of amplification products. In some embodiments, the method further comprises removing terminator nucleotides from the terminated amplification products. In one embodiment of any of the above methods, the method further comprises sequencing the amplification products. In one embodiment of any of the above methods, the amplification is performed under substantially isothermal conditions. In one embodiment of any of the above methods, the nucleic acid polymerase is a DNA polymerase.
[0068] In one embodiment of any of the above methods, the DNA polymerase is a strand displacement DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase is bacteriophage phi29 (Φ29) polymerase, a genetically modified phi29 (Φ29) DNA polymerase, the Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phiPRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase has 3'->5' exonuclease activity, and the terminator nucleotide inhibits such 3'->5' exonuclease activity. In one particular embodiment, the terminator nucleotide is a nucleotide having a modification on the alpha group (e.g., an alpha-thiodeoxynucleotide), a C3 spacer nucleotide, a locked nucleic acid (LNA), an inverted nucleic acid, a 2'-fluoro nucleotide, a 3'-phosphorylated nucleotide, a 2'-O-methyl modified nucleotide, and a trans nucleic acid. In one embodiment of any of the above methods, the nucleic acid polymerase does not have 3'->5' exonuclease activity. In one particular embodiment, the polymerase is Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(Exo-)DNA polymerase, Deep Vent (Exo-)DNA polymerase, Klenow fragment (Exo-)DNA polymerase, and Therminator DNA polymerase are selected. In one particular embodiment, the terminator nucleotide comprises a modification of the r group at the 3'-carbon of deoxyribose. In one particular embodiment, the terminator nucleotide is selected from 3'-blocked reversible terminators containing nucleotides, 3'-unblocked reversible terminators containing nucleotides, terminators containing a 2'-modification of deoxynucleotides, terminators containing a modification to the nitrogenous base of deoxynucleotides, and combinations thereof. In one particular embodiment, the terminator nucleotide is selected from dideoxynucleotides, reverse dideoxynucleotides, 3'-biotinylated nucleotides, 3'-aminonucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3'C3 spacer nucleotides, 3'C18 nucleotides, 3'-carbon spacer nucleotides containing 3'-hexanediol spacer nucleotides, acyclonucleotides, and combinations thereof. In one embodiment of any of the above methods, the amplification primer is 4 to 70 nucleotides in length. In one embodiment of any of the above methods, the amplification product is about 50 to about 2000 nucleotides in length. In one embodiment of any of the above methods, the target nucleic acid is DNA (e.g., cDNA or genomic DNA). In one embodiment of any of the above methods, the amplification primer is a random primer. In one embodiment of any of the above methods, the amplification primer contains a barcode. In one particular embodiment, the barcode contains a cell barcode. In one particular embodiment, the barcode contains a sample barcode. In one embodiment of any of the above methods, the amplification primer contains a unique molecular identifier (UMI). In one embodiment of any of the above methods, the method includes a step of denaturing the target nucleic acid or genomic DNA prior to the first primer annealing. In one particular embodiment, the denaturation is performed under alkaline conditions and subsequently neutralized.In one embodiment of any of the above methods, the mixture of sample, amplification primers, nucleic acid polymerase, and nucleotides is contained in a microfluidic device. In one embodiment of any of the above methods, the mixture of sample, amplification primers, nucleic acid polymerase, and nucleotides is contained in droplets. In one embodiment of any of the above methods, the sample is a tissue sample, cells, a biological fluid sample (e.g., blood, urine, saliva, lymph, cerebrospinal fluid (CSF), amniotic fluid, pleural fluid, pericardial fluid, ascites, aqueous humor), a bone marrow sample, a semen sample, a biopsy sample, a cancer sample, a tumor sample, a cell lysate sample, a forensic sample, an archaeological sample, a paleontological sample, an infectious sample, a production sample, a whole plant, a plant part, a microbiota sample, a virus preparation, a soil sample, a marine sample, a fresh water sample, a household or industrial sample, and combinations and isolates thereof. In one embodiment of any of the above methods, the sample is cells (e.g., animal cells [e.g., human cells], plant cells, fungal cells, bacterial cells, and protozoan cells). In one particular embodiment, the cells are lysed prior to replication. In one particular embodiment, cell lysis involves proteolysis. In one particular embodiment, the cells are cells from a preimplantation embryo, stem cells, fetal cells, tumor cells, cells suspected of having cancer, cancer cells, cells subjected to a gene editing procedure, cells from a pathogenic organism, cells obtained from a forensic sample, cells obtained from an archaeological sample, and cells obtained from a paleontological sample. In one embodiment of any of the above methods, the sample is cells from a preimplantation embryo (e.g., blastomeres [e.g., blastomeres obtained from an 8-cell stage embryo generated by in vitro fertilization]). In one particular embodiment, the method further comprises determining the presence of germline or somatic variants predisposing to disease in the embryonic cells. In one embodiment of any of the above methods, the sample is cells from a pathogenic organism (e.g., bacteria, fungi, protozoa).In one particular embodiment, the pathogenic biological cell is obtained from a patient, a microbiota sample (e.g., a GI microbiota sample, a vaginal microbiota sample, a skin microbiota sample, etc.) or a liquid obtained from an indwelling medical device (e.g., an intravenous catheter, a urinary catheter, a cerebrospinal shunt, an artificial valve, an artificial joint, an endotracheal tube, etc.). In one particular embodiment, the method further includes the step of determining the identity of the pathogenic organism. In one particular embodiment, the method further includes determining the presence of genetic variant mutants involved in the resistance of the pathogenic organism to treatment. In one embodiment of any of the above methods, the sample is a tumor cell, a cell suspected of cancer, or a cancer cell. In one particular embodiment, the method further includes determining the presence of one or more diagnostic or prognostic mutations. In one particular embodiment, the method further includes determining the presence of germline or somatic variants that cause resistance to treatment. In one embodiment of any of the above methods, the sample is a cell subjected to a gene editing procedure. In one particular embodiment, the method further includes determining the presence of unintended mutations caused by the gene editing process. In one embodiment of any of the above methods, the method further includes determining the history of the cell line. In a related aspect, the invention provides the use of any of the above methods for identifying low-frequency sequence variants (e.g., variants that constitute ≧0.01% of the entire sequence).
[0069] In related aspects, the present invention provides a kit comprising a nucleic acid polymerase, one or more amplification primers, a mixture of nucleotides comprising one or more terminator nucleotides, and optionally instructions for use. In one embodiment of the kit of the present invention, the nucleic acid polymerase is a strand displacement DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase is bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, the Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phiPRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase has 3'->5' exonuclease activity and the terminator nucleotide inhibits such 3'->5' exonuclease activity (e.g., a nucleotide having a modification on the alpha group [e.g., alpha-thiodideoxynucleotide], a C3 spacer nucleotide, a locked nucleic acid (LNA), an inverted nucleic acid, a 2'-fluoro nucleotide, a 3'-phosphorylated nucleotide, a 2'-O-methyl modified nucleotide, a trans nucleic acid). In one embodiment of the kit of the present invention, the nucleic acid polymerase does not have 3'->5' exonuclease activity (e.g., Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(Exo-)DNA polymerases, Deep Vent (Exo-)DNA polymerase, Klenow fragment (Exo-)DNA polymerase, Therminator DNA polymerase). In one particular embodiment, the terminator nucleotide comprises a modification of the r group at the 3' carbon of deoxyribose. In one particular embodiment, the terminator nucleotide is selected from 3'-blocked reversible terminators comprising nucleotides, 3'-unblocked reversible terminators comprising nucleotides, terminators comprising a 2'-modification of a deoxynucleotide, terminators comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one particular embodiment, the terminator nucleotide is selected from dideoxynucleotides, reverse dideoxynucleotides, 3'-biotinylated nucleotides, 3'-aminonucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3'C3 spacer nucleotides, 3'C18 nucleotides, 3'-hexanediol spacer nucleotides comprising 3'-carbon spacer nucleotides, acyclonucleotides, and combinations thereof.
[0070] Described herein is a method for amplifying a genome, the method comprising: a) nucleotides comprising a genome, a plurality of amplification primers (e.g., two or more primers), a nucleic acid polymerase, and one or more terminator nucleotides that terminate nucleic acid replication by the polymerase; and b) incubating a sample under conditions that promote replication of the genome to obtain a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product having a length between about 50 and about 2000 nucleotides. In one embodiment of any of the above methods, the method further comprises isolating from the plurality of terminated amplification products a product having a length between about 400 and about 600 nucleotides. In one embodiment of any of the above methods, the method further comprises: c) repairing the termini and A-tailing; and d) ligating the molecules obtained in step (c) to an adapter, thereby generating a library of amplification products. In one embodiment of any of the above methods, the method further comprises sequencing the amplification products. In one embodiment of any of the above methods, the amplification is performed under substantially isothermal conditions. In one embodiment of any of the above methods, the nucleic acid polymerase is a DNA polymerase.
[0071] In one embodiment of any of the above methods, the DNA polymerase is a strand displacement DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase is bacteriophage phi29 (Φ29) polymerase, a genetically modified phi29 (Φ29) DNA polymerase, the Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phiPRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R(Exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (Exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of any of the above methods, the nucleic acid polymerase has 3'->5' exonuclease activity, and the terminator nucleotide inhibits such 3'->5' exonuclease activity. In one particular embodiment, the terminator nucleotide is a nucleotide having a modification on the alpha group (e.g., an alpha-thiodideoxynucleotide that creates a phosphorothioate bond), a C3 spacer nucleotide, a locked nucleic acid (LNA), an inverted nucleic acid, a 2'-fluoro nucleotide, a 3'-phosphorylated nucleotide, a 2'-O-methyl modified nucleotide, and a trans nucleic acid. In one embodiment of any of the above methods, the nucleic acid polymerase does not have 3'->5' exonuclease activity. In one particular embodiment, the polymerase is Bst DNA polymerase, Exo(-)Bst polymerase, Exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(Exo-)DNA polymerase, Deep Vent (Exo-)DNA polymerase, Klenow fragment (Exo-)DNA polymerase, and Therminator DNA polymerase. In one particular embodiment, the terminator nucleotide comprises a modification of the r group at the 3' carbon of deoxyribose. In one particular embodiment, the terminator nucleotide is selected from 3'-blocked reversible terminators containing nucleotides, 3'-unblocked reversible terminators containing nucleotides, terminators containing a 2'-modification of deoxynucleotides, terminators containing a modification to the nitrogenous base of deoxynucleotides, and combinations thereof. In one particular embodiment, the terminator nucleotide is selected from dideoxynucleotides, reverse dideoxynucleotides, 3'-biotinylated nucleotides, 3'-aminonucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3'C3 spacer nucleotides, 3'C18 nucleotides, 3'-hexanediol spacer nucleotides as 3'-carbon spacer nucleotides, acyclonucleotides, and combinations thereof. In one embodiment of any of the above methods, the amplification primer is 4 to 70 nucleotides in length. In one embodiment of any of the above methods, the amplification product is about 50 to about 2000 nucleotides in length. In one embodiment of any of the above methods, the target nucleic acid is DNA (e.g., cDNA or genomic DNA). In one embodiment of any of the above methods, the amplification primer is a random primer. In one embodiment of any of the above methods, the amplification primer contains a barcode. In one particular embodiment, the barcode contains a cell barcode. In one particular embodiment, the barcode contains a sample barcode. In one embodiment of any of the above methods, the amplification primer contains a unique molecular identifier (UMI). In one embodiment of any of the above methods, the method includes a step of denaturing the target nucleic acid or genomic DNA prior to the first primer annealing. In one particular embodiment, the denaturation is performed under alkaline conditions and then neutralized.In one embodiment of any of the above methods, the mixture of sample, amplification primers, nucleic acid polymerase, and nucleotides is contained in a microfluidic device. In one embodiment of any of the above methods, the mixture of sample, amplification primers, nucleic acid polymerase, and nucleotides is contained in droplets. In one embodiment of any of the above methods, the sample is selected from tissue samples, cells, body fluid samples (e.g., blood, urine, saliva, lymph fluid, cerebrospinal fluid (CSF), amniotic fluid, pleural effusion, pericardial fluid, ascites, aqueous humor), bone marrow samples, semen samples, biopsy samples, cancer samples, tumor samples, cell lysate samples, forensic samples, archaeological samples, paleontological samples, infectious samples, production samples, whole plants, plant parts, microbiota samples, virus preparations, soil samples, seawater samples, freshwater samples, household or industrial samples, and combinations and isolates thereof. In one embodiment of any of the above methods, the sample is a cell (e.g., animal cells [e.g., human cells], plant cells, fungal cells, bacterial cells, and protozoan cells). In one particular embodiment, the cells are lysed prior to replication. In one particular embodiment, protein degradation is involved in cell lysis. In one particular embodiment, the cells are selected from cells from pre-implantation embryos, stem cells, fetal cells, tumor cells, suspected cancer cells, cancer cells, cells that have undergone gene editing procedures, cells from pathogens, cells obtained from forensic samples, cells obtained from archaeological samples, and cells obtained from paleontological samples. In one embodiment of any of the above methods, the sample is a cell of a pre-implantation embryo (e.g., a blastomere [e.g., a blastomere obtained from an 8-cell stage embryo generated by in vitro fertilization]). In one particular embodiment, the above method further comprises determining the presence of germline variants or somatic variants predisposing to a disease in the embryonic cells. In one embodiment of any of the above methods, the sample is a cell of a pathogen (e.g., bacteria, fungi, protozoa).In one particular embodiment, the pathogen cell is obtained from a patient, a microbiota sample (e.g., a GI microbiota sample, a vaginal microbiota sample, a skin microbiota sample, etc.), or a body fluid taken from an indwelling medical device (e.g., an intravenous catheter, a urinary catheter, a cerebrospinal shunt, an artificial valve, an artificial joint, an endotracheal tube, etc.). In one particular embodiment, the method further includes determining the identity of the pathogen. In one particular embodiment, the method further includes determining the presence of gene variants that cause resistance of the pathogen to treatment. In any one embodiment of the method, the sample is a tumor cell, a suspected cancer cell, or a cancer cell. In one particular embodiment, the method further includes determining the presence of one or more diagnostic or prognostic mutations. In one particular embodiment, the method further includes determining the presence of germline or somatic variants that cause resistance to treatment. In any one embodiment of the method, the sample is a cell that has undergone a gene editing procedure. In one particular embodiment, the method further includes determining the presence of unintended mutations caused by the gene editing process. In any one embodiment of the method, the method further includes determining the history of the cell line. In a related aspect, the invention provides the use of any of the above methods for identifying low-frequency sequence variants (e.g., variants that constitute ≧0.01% of the entire sequence).
[0072] In related aspects, the present invention provides a kit comprising a reverse transcriptase, a nucleic acid polymerase, one or more amplification primers, a mixture of nucleotides comprising one or more terminator nucleotides, and optionally instructions for use. In one embodiment of the kit of the present invention, the nucleic acid polymerase is a strand displacement DNA polymerase. In some cases, the reverse transcriptase performs template switching. In some cases, the reverse transcriptase is a variant of MMLV (Moloney murine leukemia virus), HIV-1, AMV (avian myeloblastosis virus), telomerase RT, FIV (feline immunodeficiency virus), or XMRV (xenotropic murine leukemia virus-related virus). Non-limiting examples of reverse transcriptases include SuperScript I (Thermo), SuperScript II (Thermo), SuperScript III (Thermo), SuperScript IV (Thermo), OmniScript (Qiagen), SensiScript (Qiagen), PrimeScript (Takara), MaximaH- (Thermo), AcuuScript Hi-Fi (Agilent), iScript (Bio-Rad), eAMV (Merck KGaA), qScript (Quanta Biosciences), SmartScribe (Clontech), or GoScript (Promega). In one embodiment of the kit of the present invention, the nucleic acid polymerase is bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, the Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phiPRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R(Exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (Exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase has 3'->5' exonuclease activity, and the terminator nucleotide inhibits such 3'->5' exonuclease activity (e.g., a nucleotide having a modification in the alpha group [e.g., alpha-thiodeoxynucleotide], C3 spacer nucleotide, locked nucleic acid (LNA), inverted nucleic acid, 2'-fluoro nucleotide, 3'-phosphorylated nucleotide, 2'-O-methyl modified nucleotide, trans nucleic acid). In one embodiment of the kit of the present invention, the nucleic acid polymerase does not have 3'->5' exonuclease activity (e.g., Bst DNA polymerase, Exo(-)Bst polymerase, Exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(Exo-)DNA polymerase, Deep Vent (Exo-)DNA polymerase, Klenow fragment (Exo-)DNA polymerase, Therminator DNA polymerase). In one particular embodiment, the terminator nucleotide comprises a modification of the r group of the 3'-carbon of deoxyribose. In one particular embodiment, the terminator nucleotide is selected from 3'-blocked reversible terminators comprising a nucleotide, 3'-unblocked reversible terminators comprising a nucleotide, terminators comprising a 2'-modification of a deoxynucleotide, terminators comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one particular embodiment, the terminator nucleotide is selected from dideoxynucleotides, inverse dideoxynucleotides, 3'-biotinylated nucleotides, 3'-aminonucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3'C3 spacer nucleotides, 3'C18 nucleotides, 3'-hexanediol spacer nucleotides comprising a 3'-carbon spacer nucleotide, acyclonucleotides, and combinations thereof. In some cases, the kit comprises at least one enzyme stabilizer, neutralizing buffer, denaturing buffer, or combinations thereof. In some cases, the kit comprises one or more modules. In some cases, the kit comprises a genomic module and a transcriptome module.
[0073] Numbered embodiments The following numbered embodiments 1 to 46 are described herein. 1. Described herein is a method for multi-omic single cell analysis, the method comprising: a. isolating single cells from a population of cells; b. sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from said cells; and c. sequencing the genome of said cells, wherein the step of sequencing the genome of said cells comprises: i. providing a genome from a single cell; ii. contacting said genome with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase, contacting; and iii. amplifying at least some of said genome to produce a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication, producing; iv. ligating the molecules obtained in step (iii) to an adapter, thereby generating a genomic DNA library; and v. sequencing said genomic DNA library, the method comprising a sequencing step. 2. Further provided herein is the method according to embodiment 1, wherein the method further comprises the step of identifying at least one protein on the surface of said cells. 3. Further provided herein is the method according to embodiment 1, wherein said mRNA transcripts comprise polyadenylated mRNA transcripts. 4. Further provided herein is the method according to embodiment 1, wherein said mRNA transcripts do not comprise polyadenylated mRNA transcripts. 5. Further provided herein is the method according to any one of embodiments 1 to 4, wherein the step of sequencing said cDNA library comprises amplification of mRNA transcripts using a template switching primer. 6. Further provided herein is the method according to any one of embodiments 1 to 4, wherein at least some of the polynucleotides of said cDNA library comprise barcodes.7. Also provided herein is the method according to any one of Embodiments 1 to 4, wherein at least some of the polynucleotides of the cDNA library contain at least two barcodes. 8. Also provided herein is the method according to Embodiment 6 or 7, wherein the barcode contains a cell barcode. 9. Also provided herein is the method according to Embodiment 6 or 7, wherein the barcode contains a sample barcode. 10. A method for multi-omic single-cell analysis, the method comprising: a. isolating single cells from a population of cells; b. identifying at least one protein on the surface of the cells; and c. sequencing the genome of the cells, wherein the step of sequencing the genome of the cells comprises: i. providing a genome from a single cell; ii. contacting the genome with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides contains at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; iii. amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iv. ligating the molecules obtained in step (iii) to an adapter, thereby generating a genomic DNA library; and v. sequencing the genomic DNA library. 11. Also provided herein is the method according to Embodiment 10, wherein the step of identifying at least one protein on the surface of the cells comprises contacting the cells with a labeled antibody that binds to the at least one protein. 12. Also provided herein is the method according to Embodiment 11, wherein the labeled antibody contains at least one fluorescent label. 13. Also provided herein is the method according to Embodiment 11, wherein the labeled antibody contains at least one mass tag. 14. Also provided herein is the method according to Embodiment 11, wherein the labeled antibody contains at least one nucleic acid barcode.15. A method of multi-omic single-cell analysis, the method comprising: a. isolating single cells from a population of cells; b. sequencing the genome of the cells, wherein the step of sequencing the genome of the cells comprises: i. providing a genome from a single cell; ii. digesting the genome with a methylation-sensitive restriction enzyme to generate genomic fragments; iii. contacting at least some of the genomic fragments with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; iv. amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; v. amplifying at least a portion of the genomic fragments by methylation-specific PCR; vi. ligating the molecules obtained in steps (iv) and (v) to an adapter, thereby generating a genomic DNA library and a methylome library; and vii. sequencing the genomic DNA library and the methylome DNA library. 16. Further provided herein is the method according to embodiment 15, wherein the step of identifying at least one protein on the surface of the cells comprises contacting the cells with a labeled antibody that binds to the at least one protein. 17. Further provided herein is the method according to embodiment 16, wherein the labeled antibody comprises at least one fluorescent label. 18. Further provided herein is the method according to embodiment 16, wherein the labeled antibody comprises at least one mass tag. 19. Further provided herein is the method according to embodiment 16, wherein the labeled antibody comprises at least one nucleic acid barcode. 20. Further provided herein is the method according to any one of embodiments 1 to 19, wherein the single cell is a mammalian cell. 21. Further provided herein is the method according to any one of embodiments 1 to 19, wherein the single cell is a human cell.22. Also provided herein is the method according to any one of Embodiments 1-19, wherein the single cell is derived from the liver, skin, kidney, blood, or lung. 23. Also provided herein is the method according to any one of Embodiments 1-19, wherein the single cell is a primary cell. 24. Also provided herein is the method according to any one of Embodiments 1-23, wherein the method further comprises removing at least one terminator nucleotide from the above-mentioned end amplification product. 25. Also provided herein is the method according to any one of Embodiments 1-23, wherein at least some of the above-mentioned amplification products contain barcodes. 26. Also provided herein is the method according to any one of Embodiments 1-23, wherein at least some of the above-mentioned amplification products contain at least two barcodes. 27. Also provided herein is the method according to Embodiment 24 or 26, wherein the barcode contains a cell barcode. 28. Also provided herein is the method according to Embodiment 24 or 26, wherein the barcode contains a sample barcode. 29. Also provided herein is the method according to any one of Embodiments 1-28, wherein at least some of the amplification primers contain unique molecular identifiers (UMIs). 30. Also provided herein is the method according to any one of Embodiments 1-28, wherein at least some of the amplification primers contain at least two unique molecular identifiers (UMIs). 31. Also provided herein is the method according to any one of Embodiments 1-30, wherein the method further comprises an additional amplification step using PCR. 32. Also provided herein is the method according to any one of Embodiments 1-30, wherein at least one mutation is identified in the genome of the cell, and the mutation is different from the corresponding position in the reference sequence. 33. Also provided herein is the method according to Embodiment 32, wherein the at least one mutation occurs in less than 50% of the cell population. 34. Also provided herein is the method according to Embodiment 32, wherein the at least one mutation occurs in less than 25% of the cell population.35. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in less than 1% of the cell population. 36. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.1% or less of the cell population. 37. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.01% or less of the cell population. 38. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.001% or less of the cell population. 39. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.0001% or less of the cell population. 40. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 50% or less of the amplified product sequence. 41. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 25% or less of the amplified product sequence. 42. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 1% or less of the amplified product sequence. 43. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.1% or less of the amplified product sequence. 44. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.01% or less of the amplified product sequence. 45. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.001% or less of the amplified product sequence. 46. Also provided herein is the method according to embodiment 32, wherein the at least one mutation occurs in 0.0001% or less of the amplified product sequence.
Example
[0074] The following examples are described to clearly illustrate to those skilled in the art the principles and implementation of the embodiments disclosed herein and should not be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are by weight.
[0075] Example 1: Primary Template Directed Amplification (PTA) PTA can be used for any nucleic acid amplification, but it avoids the drawbacks of currently used methods, such as exponential amplification at the location where the polymerase first elongates the random primer, which results in random overrepresentation of loci and alleles as well as propagation of mutations. It enables capturing a larger percentage of the cell genome in a more uniform and reproducible manner and with a lower error rate than currently used methods such as Multiple Displacement Amplification (MDA), for example, and is thus particularly useful for whole genome amplification (see Figure 1G). PTA is also used in conjunction with other analytical techniques such as transcriptome analysis.
[0076] Cell culture Human NA12878 (Coriell Institute) cells were maintained in RPMI medium supplemented with 15% FBS and 2 mM L-glutamine, as well as 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B (Gibco, Life Technologies). The cells were seeded at a density of 3.5×10 5 cells / ml. The cultures were split every 3 days and maintained in a humidified incubator at 37 °C using 5% CO 2 2.
[0077] Single cell isolation and WTA A general protocol for WTA (whole transcriptome analysis) is shown in Figure 2F. Cells were resuspended at a concentration of 150 - 500 cells / μL. This cell suspension was stained with 20 μL of freshly prepared staining buffer (2.5 μL of ethidium homodimer-1 and 0.625 μL of calcein AM from Life Technology's LIVE / DEAD® Viability / Cytotoxicity Kit added to 1.25 mL of cell buffer containing 1×PBS and 0.05% tween-20). Next, the cells were sorted using a FACS Aria III sorting device and the cells were deposited into each of the 96 wells. A reaction mix containing 5×RT buffer, PEG4000, RT primer (100 μM), TS oligo (20 μM), reverse transcriptase, RNase inhibitor, gelatin, Tween-20, Triton-X, dNTP mix, TMAC (1 M), betaine (5 M), MgCl 2 (50 mM), and ERCC spike were added to each well. Then, the samples were placed in a thermal cycler at 42 °C for 90 minutes, 50 °C for 30 minutes, and then the samples were held at 4 °C until they could be processed for pre-amplification. After the thermal cycle of RT, the samples are either processed for DNA amplification or processed for pre-amplification of the first-strand cDNA resulting from the RT reaction. Pre-amplification of the samples was achieved using a single primer using the following protocol (semi-suppressive PCR) to amplify the cDNA product. Briefly, 5 μL of the RT reaction was added to a 30 μL reaction containing 2× master mix, 1 micromolar concentration primer, and 5× pre-amplification buffer, with thermal cycling conditions of 95 °C for 1 minute, then 21 cycles of 95 °C - 15 seconds, 60 °C - 30 seconds, 68 °C - 4 minutes, followed by holding at 72 °C for 10 minutes. Next, the samples were converted into sequencing libraries using the Nextera XT library preparation kit using the manufacturer's instructions (Figure 2G). The results of the RT experiments are shown in Table 1 for six samples.
[0078]
Table 1
[0079] Single cell isolation and WGA 3.5×10 5 NA12878 cells were cultured for at least 3 days after seeding at a density of cells / ml, and then 3 mL of the cell suspension was pelleted at 300×g for 10 minutes. Next, the medium was discarded, and the cells were washed three times by rotating at 300×g, 200×g, and finally 100×g for 5 minutes with 1 mL of cell wash buffer (1×PBS containing 2% FBS without Mg 2+ or Ca 2+ ). Next, the cells were resuspended in 500 μL of cell wash buffer. Subsequently, the cells were stained with 100 nM calcein AM (Molecular Probes) and 100 ng / ml propidium iodide (PI; Sigma-Aldrich) to distinguish the viable cell population. The cells were loaded onto a BD FACScan flow cytometer (FACSAria II) (BD Biosciences) that had been thoroughly washed with ELIMINase (Decon Labs) and calibrated using Accudrop fluorescent beads (BD Biosciences) for cell sorting. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS with 0.2% Tween20 in cells that received PTA (Sigma-Aldrich). Multiple wells were intentionally left empty for use as no-template controls (NTCs). Immediately after sorting, the plate was centrifuged briefly and placed on ice. Next, the cells were frozen at -20°C for at least one night. The next day, the WGA reaction was assembled in a pre-PCR workstation that provided a constant positive pressure of HEPA-filtered air and had been decontaminated with UV light for 30 minutes before each experiment.
[0080] MDA was performed using modifications that have previously been shown to improve amplification uniformity. Specifically, exonuclease-resistant random primers (ThermoFisher) were added to the lysis buffer / mix to a final concentration of 125 μM. The resulting 4 μL of lysis / denaturation mix was added to a tube containing single cells, vortexed, briefly spun, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, mixed by vortexing, briefly centrifuged, and placed at room temperature. Then, 40 μl of amplification mix was added and the sample was incubated at 30 °C for 8 hours, after which amplification was terminated by heating at 65 °C for 3 minutes.
[0081] PTA was first performed by further lysing the cells after freeze-thawing by adding 2 μl of a pre-chilled solution of a 1:1 mixture of 5% Triton X-100 (Sigma-Aldrich) and 20 mg / ml proteinase K (Promega). Next, the cells were vortexed, briefly centrifuged, and then placed at 40 °C for 10 minutes. Next, 4 μl of lysis buffer / mix and 1 μl of 500 μM exonuclease-resistant random primer were added to the lysed cells to denature the DNA, after which the sample was vortexed, spun, and placed at 65 °C for 15 minutes. Next, 4 μl of room temperature quenching buffer was added, the sample was vortexed, and spun down. The 56 μl of amplification mix (primers, dNTPs, polymerase, buffer) contained equimolar amounts of alpha-thio-ddNTPs at a concentration of 1200 μM in the final amplification reaction. Next, the sample was placed at 30 °C for 8 hours, after which amplification was stopped by heating at 65 °C for 3 minutes.
[0082] After the amplification step, DNA from both the MDA and PTA reactions was purified using AMPure XP magnetic beads (Beckman Coulter) at a bead-to-sample ratio of 2:1, and the yield was measured using a Qubit 3.0 fluorometer according to the manufacturer's (Life Technologies) instructions with the Qubit dsDNA HS assay kit.
[0083] Library preparation The MDA reaction yielded 40 μg of amplified DNA. Following standard procedures, 1 μg of the product was fragmented for 30 minutes. Next, the samples underwent standard library preparation using 15 μM dual-index adapters (end repair with T4 polymerase, T4 polynucleotide kinase, and Taq polymerase for A-tailing) and 4 cycles of PCR. Each PTA reaction generated 40 - 60 ng of material that was used in its entirety, without fragmentation, for the preparation of a standard DNA sequencing library. A 2.5 μM adapter with UMI and dual-index was used for ligation with T4 ligase, and 15 cycles of PCR (hot-start polymerase) were used for final amplification. The library was then cleaned up using both-sided SPRI, using ratios of 0.65× and 0.55× for right- and left-sided selection, respectively. The final library was quantified using the Qubit dsDNA BR assay kit and the 2100 Bioanalyzer (Agilent Technologies), and then sequenced on the Illumina NextSeq platform. All Illumina sequencing platforms, including NovaSeq, are also compatible with this protocol.
[0084] Data analysis Array determination reads were demultiplexed based on cell barcodes using Bcl2fastq. Reads were then trimmed using trimmomatic and subsequently aligned to hg19 using BWA. Reads were subjected to duplicate marking by Picard, followed by local realignment and base recalibration using GATK4.0. All files used to calculate quality metrics were downsampled to 20 million reads using Picard DownSampleSam. Quality metrics were obtained from the final bam file using qualimap, as well as PicardAlignmentSummaryMetrics and CollectWgsMetrics. Total genomic coverage was also estimated using Preseq.
[0085] Variant calling Single nucleotide variants and indels were called using GATK UnifiedGenotyper from GATK4.0. Standard filtering criteria using the best practices of GATK were used in all steps of the process (https: / / software.broadinstitute.org / gatk / best-practices / ). Copy number variants were called using Control-FREEC (Boeva et al., Bioinformatics, 2012, 28(3):423-5). Structural variants were also detected using CREST (Wang et al., Nat Methods, 2011, 8(8):652-4).
[0086] Results As shown in FIGS. 3A and 3B, the mapping rate and mapping quality score of amplification using only dideoxynucleotides (“reversible”) are 15.0+ / -2.2 and 0.8+ / -0.08, respectively, while the incorporation of exonuclease-resistant alpha-thio-dideoxynucleotide terminators (“irreversible”) results in mapping rates and quality scores of 97.9+ / -0.62 and 46.3+ / -3.18, respectively. Experiments were also performed using reversible ddNTPs and terminators at various concentrations (FIG. 2A, bottom).
[0087] FIGS. 2B - 2E show comparative data generated from NA12878 human single cells that received MDA (according to Dong, X. et al., Nat Methods. 2017, 14(5):491 - 493) or PTA. Both protocols generated comparable low PCR duplication rates (MDA 1.26%+ / -0.52 vs PTA 1.84%+ / -0.99) and GC% (MDA 42.0+ / -1.47 vs PTA 40.33+ / -0.45), but PTA generated smaller amplicon sizes. The proportion of mapped reads and mapping quality scores were also significantly higher with PTA compared to MDA (PTA 97.9+ / -0.62 vs MDA 82.13+ / -0.62 and PTA 46.3+ / -3.18 vs MDA 43.2+ / -4.21, respectively). Overall, PTA generates more usable mapped data when compared to MDA. FIG. 4A shows that PTA significantly improves the uniformity of amplification, has a wider coverage width, and fewer regions with coverage close to 0 compared to MDA. The use of PTA can identify low-frequency sequence variants within a population of nucleic acids containing variants, which constitute more than 0.01% of the total sequence. PTA can be successfully used for the amplification of single-cell genomes.
[0088] Example 2: Comparative Analysis of PTA Benchmarking the Maintenance and Isolation of PTA and SCMDA Cells Lymphoblastoid cells from the 1000 Genomes Project target NA12878 (Coriell Institute, Camden, NJ, USA) were maintained in RPMI medium supplemented with 15% FBS, 2 mM L-glutamine, 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B. Cells were seeded at a density of 3.5×10 5 cells / ml and split every 3 days. They were maintained in a humidified incubator at 37 °C containing 5% CO 2 . Prior to separating single cells, a 3 mL cell suspension that had been expanded over the past 3 days was rotated at 300×g for 10 minutes. The pelleted cells were washed 3 times with 1 mL of cell wash buffer (1×PBS containing 2% FBS without Mg 2+ or Ca 2+ ) and successively rotated at 300×g, 200×g, and finally 100×g for 5 minutes to remove dead cells. Next, the cells were resuspended in 500 μL of cell wash buffer and subsequently stained with 100 nM calcein AM and 100 ng / ml propidium iodide (PI) to distinguish the live cell population. The cells were washed thoroughly with ELIMINase and loaded onto a BD FACScan flow cytometer (FACSAria II) calibrated using Accudrop fluorescent beads. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS with 0.2% Tween20. Multiple wells were intentionally left empty and used as template-free controls. Immediately after sorting, the plate was centrifuged briefly and placed on ice. Next, the cells were frozen at -80 °C for at least one night.
[0089] PTA and SCMDA experiments The WGA reaction was assembled on a pre-PCR workstation that provided a constant positive pressure with air passing through a HEPA filter and was decontaminated with UV light for 30 minutes before each experiment. MDA was performed according to the SCMDA methodology using a published protocol (Dong et al. Nat. Meth. 2017, 14, 491 - 493). Specifically, exonuclease-resistant random primers were added to the lysis buffer at a final concentration of 12.5 μM. The resulting 4 μL of lysis mix was added to a tube containing single cells, mixed by pipetting three times, rotated briefly, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, mixed by pipetting three times, centrifuged briefly, and placed on ice. Subsequently, 40 μL of amplification mix was added, followed by incubation at 30 °C for 8 hours, and then amplification was terminated by heating at 65 °C for 3 minutes. PTA was performed by further lysing the cells by adding 2 μL of a pre-cooled solution of a 1:1 mixture of 5% Triton X-100 and 20 mg / ml proteinase K after freeze-thawing. Next, the cells were vortexed, centrifuged briefly, and then placed at 40 °C for 10 minutes. Next, 4 μL of denaturation buffer and 1 μl of 500 μM exonuclease-resistant random primer were added to the lysed cells to denature the DNA, followed by vortexing, rotation, and placement at 65 °C for 15 minutes. Next, 4 μL of room temperature quenching solution was added, and the sample was vortexed and spun down. In the final amplification reaction, 56 μL of amplification mix contained equimolar alpha-thio-ddNTP at a concentration of 1200 μM. Next, the sample was placed at 30 °C for 8 hours, and then amplification was terminated by heating at 65 °C for 3 minutes. After SCMDA or PTA amplification, the DNA was purified using AMPure XP magnetic beads with a bead-to-sample ratio of 2:1, and the yield was measured using a Qubit 3.0 fluorometer with a Qubit dsDNA HS assay kit according to the manufacturer's instructions.
[0090] Library Preparation After the addition of the preparation solution, 1 μg of the SCMDA product was fragmented for 30 minutes according to the HyperPlus protocol. Next, the samples underwent preparation of a standard library using a unique dual-index adapter at 15 μM and 4 cycles of PCR. The entire product of each PTA reaction was used for the preparation of a DNA sequencing library using a standard amplification protocol without fragmentation. A unique dual-index adapter at 2.5 μM was used in ligation, and 15 cycles of PCR were used in the final amplification. Next, the libraries of SCMDA and PTA were visualized on a 1% agarose E-Gel. Fragments of 400 - 700 bp were excised from the gel and recovered using a Gel DNA Recovery Kit. The final library was quantified using a Qubit dsDNA BR assay kit and an Agilent 2100 Bioanalyzer before sequencing on a NovaSeq 6000.
[0091] Data analysis The data was trimmed using trimmomatic and then aligned to hg19 using BWA. The reads underwent duplicate marking by Picard and then local realignment and base recalibration using the best practices of GATK3.5. All files were downsampled to the specified number of reads using PicardDownSampleSam. Quality metrics were obtained from the final bam file using qualimap, and Picard AlignmentMetricsAummary and CollectWgsMetrics. Lorenz curves were drawn and the Gini index was calculated using htSeqTools. SNV calling was performed using UnifiedGenotyper, which was then filtered using standard recommended criteria (QD < 2.0 || FS > 60.0 || MQ < 40.0 || SOR > 4.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0). There were no regions excluded from the analysis, and no normalization or manipulation of other data was performed. The sequencing metrics for the tested methods are shown in Table 2.
[0092]
Table 2
[0093] Width and uniformity of genomic coverage An inclusive comparison of PTA with all common single-cell WGA methods was performed. To achieve this, PTA and a modified version of MDA called single-cell MDA (Dong et al. Nat. Meth. 2017, 14, 491-493) (SCMDA) were each performed on 10 NA12878 cells. Further, these results for cells amplified using DOP-PCR (Zhang et al. PNAS 1992, 89, 5847-5851), MDA kit 1 (Dean et al. PNAS 2002, 99, 5261-5266), MDA kit 2, MALBAC (Zong et al. Science 2012, 338, 1622-1626), LIANTI (Chen et al., Science 2017, 356, 189-194), or PicoPlex (Langmore, Pharmacogenomics 3, 557-560 (2002)) were compared using data generated as part of the LIANTI study.
[0094] To normalize across samples, raw data from all samples were aligned and underwent preprocessing for variant calling using the same pipeline. Next, the bam files were subsampled to 300 million reads each before performing the comparisons. Importantly, the PTA and SCMDA products were not screened prior to further analysis, while all other methods underwent screening for genomic coverage and uniformity before selecting the highest quality cells for subsequent analysis. Notably, SCMDA and PTA were compared to the bulk diploid NA12878 sample, while all other methods were compared to the bulk BJ1 diploid fibroblasts used in the LIANTI study. As seen in Figures 3C - 3F, PTA had the highest percentage of reads aligned to the genome and the highest mapping quality. PTA, LIANTI, and SCMDA had similar GC contents, all of which were lower than the other methods. The PCR duplication rate was similar across all methods. Additionally, the PTA method enabled smaller templates, such as the mitochondrial genome, to give a higher coverage rate (similar to larger standard chromosomes) compared to the other methods tested (Figure 3G).
[0095] Next, the coverage width and uniformity of all methods were compared. Examples of coverage plots across the entire first chromosome are shown for SCMDA and PTA, where PTA is shown to have significantly improved coverage uniformity and allele frequency (Figure 4B). Next, the coverage rate for all methods was calculated using increased read counts. PTA approaches the two bulk samples at all depths, which is a significant improvement over all other methods (Figure 5A). Next, the inventors measured the coverage uniformity using two strategies. The first approach was to calculate the coefficient of variation of coverage at increasing sequencing depths, where PTA was found to be more uniform than all other methods (Figure 5B). The second strategy was to calculate the Lorenz curve for each subsampled bam file, where PTA was again found to have the greatest uniformity (Figure 5C). To measure the reproducibility of amplification uniformity, the Gini index was calculated to estimate the difference of each amplification reaction from perfect uniformity (de Bourcy et al., PloS one 9, e105585 (2014)). PTA was again shown to be more reproducibly uniform than other methods (Figure 5D).
[0096] SNV sensitivity To determine the impact of these differences in the performance of amplification methods for SNV calls, the respective variant call rates for the corresponding bulk samples were compared at increasing sequencing depths. To estimate sensitivity, the percentage of variants called in the corresponding bulk samples, subsampled to 650 million reads found in each cell at each sequencing depth, was compared (Figure 5E). The improved coverage and uniformity of PTA resulted in the detection of 45.6% more variants than MDA Kit 2, the next most sensitive method. Examination of sites called as heterozygosity in the bulk samples showed that PTA significantly reduced allelic skewing at these heterozygous sites (Figure 5F). This finding supports the claim that PTA not only amplifies more uniformly across the genome but also amplifies the two alleles within the same cell more uniformly.
[0097] SNV Specificity To estimate the specificity of variant calls, variants called in each single cell that were not found in the corresponding bulk sample were considered false positives. The cryolysis of SCMDA significantly reduced the number of false positive variant calls (Figure 5G). Methods using thermostable polymerases (MALBAC, PicoPlex, and DOP-PCR) showed a further decrease in the specificity of SNV calls with increasing sequencing depth. Although not bound by theory, this may be the result of a significantly increased error rate of these polymerases compared to phi29 DNA polymerase. Additionally, the base change patterns seen in false positive calls also appear to be polymerase-dependent (Figure 5H). As seen in Figure 5G, the model of suppressed error propagation in PTA is supported by the lower false positive SNV call rate in PTA compared to the standard MDA protocol. Furthermore, PTA has the lowest allelic frequency of false positive variant calls, which is also consistent with the model of suppressed error propagation by PTA (Figure 5I).
[0098] Example 3: Ultra-Parallel Single-Cell DNA Sequencing Using PTA, a protocol for ultra-parallel DNA sequencing is established. First, cell barcodes are added to random primers. Two strategies are employed to minimize any bias in amplification introduced by the cell barcodes, which are 1) increasing the length of the random primers and / or 2) creating primers that loop back on themselves to prevent the cell barcodes from binding to the template (Figure 10B). Once the optimal primer strategy is established, up to 384 sorted cells are scaled using, for example, a Mosquito HTS liquid handler that can pipette with high precision down to a volume of 25 nL, even with a viscous liquid. This liquid handler reduces reagent costs by approximately 1 / 50 by using a 1 μL PTA reaction instead of the standard 50 μL reaction volume.
[0099] The amplification protocol is transferred to droplets by delivering primers with cell barcodes to the droplets. Solid supports such as beads made using a split and pool strategy are optionally used. Suitable beads are available, for example, from ChemGenes. In some cases, the oligonucleotides include random primers, cell barcodes, unique molecular identifiers, and a cleavable sequence or spacer for releasing the oligonucleotide after the bead and cell are encapsulated in the same droplet. During this process, the low-nanoliter volume template, primer, dNTP, alpha-thio-ddNTP, and polymerase concentrations in the droplets are optimized. Optimization includes, in some cases, the use of large droplets to increase the reaction volume. As shown in FIG. 9, this process requires two consecutive reactions to lyse the cells, followed by WGA. The first droplet containing the lysed cells and beads is combined with a second droplet having the amplification mix. Alternatively, or in combination, the cells can be encapsulated in hydrogel beads prior to lysis and then both beads can be added to the oil droplets. See Lan, F. et al., Nature Biotechnol., 2017, 35:640-646.
[0100] Additional methods include the use of microwells, which, in some cases, capture 140,000 single cells in 20 picoliter reaction chambers on a device the size of a 3” x 2” microscope slide. Similar to the droplet-based methods, these wells combine the cells with beads containing cell barcodes to enable ultra-parallel processing. See Gole et al., Nature Biotechnol., 2013, 31:1126-1132.
[0101] Example 4: Parallel Analysis of Genomes and Transcriptomes in Single Cells Sort single cells from a cell population and place one cell per well. Each well contains an antibody immobilized on the surface area, and this antibody binds to the cell nucleus. The outer membrane of the cell is lysed, releasing mRNA into the solution in the well, while the nuclease remains intact and bound to the well area. Reverse transcription (RT) is performed using the mRNA in the solution as a template, and cDNA is generated using the primers of FIG. 8A. Optionally, a ribosomal RNA (rRNA) depletion step is performed. A first template-switching primer containing a transcription start site (TSS) region, an anchor region, an RNA BC region, and a poly dT tail from 5' to 3', and a second template-switching primer containing a TSS region, an anchor region, and a poly G region from 3' to 5' are used for RT-PCR. After removing the RT-PCR product (cDNA library) for subsequent sequencing, any remaining RNA in the cell is removed by UNG. The RNA library is prepared using a Nextera / transposon-based sequencing method and reagents (FIG. 8B). The cDNA library contains short cDNAs amplified approximately 1000-fold. Next, the nucleus is lysed, and the released genomic DNA is subjected to the PTA method using a random primer with an isothermal polymerase using a random primer with a length of 6 to 9 bases. The amplification conditions for PTA are selected to generate amplicons with a length of 250 to 1500 bases. The PTA product is optionally subjected to additional amplification and sequenced. The RNA sequencing data and the DNA sequencing data are compiled into a database for analysis.
[0102] Example 5: Single-Cell Multi-Omic Analysis A population of cells is contacted with an antibody library in which the antibodies are labeled, and the antibodies are labeled with either a fluorescent label, a nucleic acid barcode, or both. The labeled antibodies bind to at least one cell within the population, and such cells are sorted and one cell is placed per well. Some of the labeled antibodies provide specific information regarding cell surface protein markers after binding, which is obtained by either fluorescence microscopy or reading the barcode tagged to the antibody. Each well contains an antibody immobilized on the surface region, which binds to the cell nucleus. The outer membrane of the cell is lysed, releasing mRNA into the solution within the well, while the nuclease remains intact and bound to the region of the well. Optionally, an rRNA (ribosomal RNA) depletion step is performed. Next, RT is performed using the mRNA in the solution as a template to generate cDNA. A first template-switching primer containing a TSS region (transcription start site), an anchor region, an RNA BC region, and a poly dT tail from 5' to 3', and a second template-switching primer containing a TSS region, an anchor region, and a poly G region from 3' to 5' are used for RT-PCR. After removing the RT-PCR product (cDNA library) for subsequent sequencing, any RNA remaining within the cell is removed by UNG. The cDNA library contains short cDNAs amplified approximately 1000-fold. Next, the nucleus is lysed, and the released genomic DNA is subjected to the PTA method using a random primer with an isothermal polymerase using a random primer of 6-9 base length. The amplification conditions for PTA are selected to generate amplicons of 250-1500 base length. The PTA product is optionally subjected to additional amplification and sequenced. The RNA sequencing data and the DNA sequencing data are compiled into a database for analysis.
[0103] Example 6: Single-cell analysis of the methylome and transcriptome Sort single cells from a cell population and place one cell per well. Each well contains an antibody immobilized on the surface area, and this antibody binds to the cell nucleus. The outer membrane of the cell is lysed to release mRNA into the solution in the well, while the nuclease remains intact and bound to the well area. The mRNA transcript is contacted with terminal transferase to add riboguanine to the 5' end of the mRNA strand. Next, RT is performed using the mRNA in the solution as a template to generate cDNA. A first template-switching primer containing the TSS region (transcription start site), anchor region, RNA BC region, and poly dT tail from 5' to 3', and a second template-switching primer containing the TSS region, anchor region, and poly G region from 3' to 5' are used for RT-PCR. After removing the RT-PCR product (cDNA library) for subsequent sequencing, any remaining RNA in the cell is removed by UNG. The cDNA library contains short cDNAs amplified approximately 1000-fold. Next, the nucleus is lysed, and the released genomic DNA is fragmented using a methylation-sensitive endonuclease. The genomic fragments are subjected to the PTA method using a random primer with an isothermal polymerase using a random primer of 6-9 base length. The amplification conditions for PTA are selected to generate amplicons of 250-1500 base length. The PTA product is optionally subjected to additional amplification and sequenced. The RNA sequencing data and DNA sequencing data are compiled into a database for analysis, and the methylation-sensitive endonuclease cleavage sites are identified. These sites are used to map the positions of methylation on the original genomic DNA.
[0104] Example 7: Methyolom and genomic single-cell analysis. Sort single cells from a cell population and place one cell per well. Each well contains an antibody immobilized on the surface area, and this antibody binds to the cell nucleus. The cells are lysed using methylation-sensitive enzymes, and the genome is subjected to the PTA method using random primers with an isothermal polymerase using random primers of 6-9 bases in length. The amplification conditions for PTA are selected to generate amplicons 250-1500 bases in length. The reaction mixture is split, and half of the mixture is subjected to exosome enrichment, whole-genome sequencing, or other targeted sequencing methods. The other half of the reaction mixture is subjected to methylation-sensitive PCR conditions. The methylation and DNA sequencing data are compiled into a database for analysis.
[0105] Example 8: Single-cell analysis of surface proteome and genome. Contact cells from a sample containing a population of cells with a library of baits such as antibodies, polynucleotides, or other small molecules. In some cases, the baits are barcoded (such as barcoded antibodies) to enable pull-down and identification of the binding of the bait to proteins on the cell surface. Alternatively, or in combination, the baits are labeled with other labels such as fluorescent labels or mass tags. Sort single cells from the cell population and place one cell per well. Optionally, the bait bound to the cell surface is removed for sequencing or identification prior to the preparation of the genomic library. The cells are lysed, the genome is released into solution, and fragments are generated. The genomic fragments are subjected to the PTA method using random primers with an isothermal polymerase using random primers of 6-9 bases in length. Alternatively, the genome is not fragmented prior to amplification with PTA. The amplification conditions for PTA are selected to generate amplicons 250-1500 bases in length. The PTA product is optionally subjected to additional amplification and sequencing. The cell surface protein and DNA sequencing data are compiled into a database for analysis.
[0106] Example 9: Multi-omics for measuring drug resistance Monotherapy with small molecule inhibitors targeting FLT3 in AML (acute myeloid leukemia) has shown clinical benefit, but resistance always occurs. Quizartinib (AC220), an FLT3 inhibitor, is one such inhibitor that has brought about a composite complete remission of approximately 50% in patients with relapsed or refractory AML. Despite this success, secondary FLT3 mutations in the activation loop (D835) and gatekeeper residue F691 have been identified in FLT3-ITD patients who relapsed on quizartinib therapy. Clinical resistance to the multi-kinase inhibitor PKC412 was determined to be the result of secondary mutations in the FLT3 kinase domain. Additional modes of resistance independent of FLT3 to targeted therapy have been identified in FLT3-ITD AML and include bypass pathway activation of AXL, as well as NRAS, TET2, and IDH1 / 2 mutations. Mutations in epigenetic modifying enzymes and transcription factors have also been observed, highlighting the complexity and diversity of the mechanisms of resistance to FLT3 inhibition.
[0107] Quizartinib-resistant, matching parental MOLM-13 AML cell lines, and cell lines with heterozygous FLT3-ITD mutations were generated. The PTA method was used to genomically and transcriptionally probe these drug-resistant single cells by combining RNAseq chemistry and gaining insights into the mechanisms of resistance after FLT3 inhibition in AML. Briefly, the workflow consisted of (1) generation of resistant cells, (2) isolation of resistant cells, (3) lysis of cells to release mRNA, (4) reverse transcription to generate cDNA from mRNA, (5) nuclear lysis to release genomic DNA, (6) PTA amplification, (7) separate DNA / RNA enrichment, (8) cDNA PreAMP of enriched mRNA, (9) library preparation, QC, and pooling, (10) next-generation sequencing, and (11) data analysis.
[0108] Cell culture. MOLM-13 acute myeloid leukemia cells with heterozygous FLT3 internal tandem duplication (ITD) 1 were obtained from DSMZ - German Collection of Microorganisms and Cell Cultures (ACC 554). The cells were maintained in RPMI 1640 (Gibco 11875 - 093) supplemented with 10% FBS and penicillin / streptomycin, and subcultured every 2 - 3 days while maintaining a density range of 2.5 E5 - 1.5E6 cells / ml. For the generation of the quizartinib-resistant MOLM-13 line, cells were continuously treated with 2 nM quizartinib and the drug was replenished in each subculture until resistant clones appeared after 5 weeks of culture (Figure 9A). Genomic DNA or total RNA was isolated from quizartinib-resistant and corresponding parental MOLM-13 cells that matched at the time of FACS sorting to generate bulk sequencing control libraries for comparison with single-cell datasets.
[0109] FACS. For single-cell analysis, ~2.0 E6 MOLM-13 quizartinib-resistant or corresponding parental cells were rinsed twice with calcium- and magnesium-free Dulbecco's phosphate-buffered saline (Gibco) supplemented with 2% FBS and kept on ice until BD FACSAriaIII FACS sorting. Following calcein AM, propidium iodide, and DAPI staining, live-cell gating was established (DAPI / PI negative, top 70% calcein-AM positive), and single cells were sorted into low-binding 96-well PCR plates (semi-skirted) containing cell buffer (130 micron nozzle assembly), vortexed briefly and centrifuged, and immediately frozen on dry ice.
[0110] Combination of genomic / transcriptomic analysis. First, biotinylated oligo dT primers were used in a template-switching reverse transcription reaction to generate first-strand cDNA from single MOLM-13 parental cells or crizotinib-resistant cells. Primary template-directed amplification (PTA) was performed continuously following reverse transcription. Next, streptavidin M-280 beads were used to affinity purify the first-strand cDNA, which was subjected to two high-salt washes and then one low-salt wash. Twenty cycles of pre-amplification were performed to generate second-strand cDNA, and an RNA sequencing library was prepared using the Nextera DNA Flex Library Preparation Kit. For the preparation of the PTA library, the PTA products not bound to streptavidin beads were purified using beads and ligated to TruSeq adapters. The amplification products from the PTA reaction were first purified by bead cleanup, measured by Qubit, and analyzed by electrophoresis. The typical yield of mammalian cells (~6 pg DNA) was 1 - 3 μg, and a single bacterial genome (2 - 4 fg) produced up to 50 ng. The amplicon product size of the samples amplified by PTA was 0.2 - 4 kB (average 1.5 Kb). The PTA library was prepared without fragmentation for the WGS method, and a yield of approximately 500 ng was obtained in the size range of 300 - 550 bases. The whole genome from mammalian cells was analyzed by NovaSeq targeting ~550 million reads. Next, the sequencing files were transferred for trimming alignment and VCF file creation and analyzed by the Trailblazer™ cloud-based bioinformatics platform solution. The QC and library preparation time was 4 - 6 hours. Parallel experiments were performed using only RNA Seq for comparison.
[0111] Results. RNA expression from both parental and resistant cultures demonstrated the ability to create cDNA pools (Figure 9B) using single - spot RNA sequencing chemistry, and the genes expressed in these cells created a unique pattern of gene expression that enabled visualization of cell populations over an average of ~10,000 genes detected per cell. In another workflow, single - cell genomes were amplified using the PTA method. Next, by combining the two protocols (yield in Figure 9D), combined transcriptome and genomic cDNA pools were generated from each cell. Low - pass (~5 million reads / cell) demonstrated effective amplification and library preparation for both the resistant and parental strains with low mitochondrial chromosomal content and high complete PreSeq genome estimates (Figures 10A - 10C). This data demonstrated that transcripts generated during the RT step were not effectively amplified by the PTA reaction compared to DNA, and that DNA within single cells was effectively amplified using the combined protocol compared to standard PTA - amplified genomes from single cells (Figure 9D). The combined RNASeq / PTA approach generated results similar to the standard PTA protocol (Figure 10A), where ChrM and the percent duplication were typically less than 2% and the estimated genome size exceeded 3 billion bases (Figures 10A - 10C). Genome evaluation revealed mapping and coverage exceeding 90% and specific calls for over 75% of single - nucleotide variants within each cell. More variability was observed with the dual protocol compared to the standard PTA genome chemistry. For the transcriptome, the prototype chemistry appeared to detect approximately 3000 - 5000 genes that included exon - exon junctions. Approximately 30% more genes were detected with the dual protocol (Figure 10D) compared to the RNAseq - only protocol (Figure 9C). Additionally, the dual / combined RNASeq / PTA protocol was used with a second resistant cell line, SUM159 (triple - negative breast cancer cell line). RNAseq data run with both protocols generated similar PCA distributions.This shows that the combined chemical substances can detect differential gene expression not limited to the parental and resistant cells of a single cell type (Figures 10E - 10F).
[0112] Deep sequencing of 7 parental cells and 5 resistant molm13 cells was performed to approximately 25-fold depth (Figure 11). Reads were aligned to Hg38 using bwa mem. Quality control and SNV calling were performed using the best practices of GATK4. SNVs were restricted to at least 2 resistant cells, alternative alleles were not called in any parental cells, and were only considered if the genotypes of at least 6 parental cells were determined. All cells had at least 96% of the genome covered at 1x coverage and at least 76% covered at 10x coverage. The inset shows that all cells detected the known Flt3 indel in molm13 cells (4 are shown for clarity).
[0113] The RNAseq and PTA methods were generally equivalent, with both mapping and coverage exceeding 95%, and ChrM and PCR duplicates generally being below 2.0%. Furthermore, more than 95% of the genomes of selected samples from both 159 parental and resistant cell lines were recovered. In the Molm13 cell line, the overexpressed gene GAS6(L) was identified, which is a known mechanism of crizotinib resistance. Gas6 is a ligand of AXL, which is a clinically relevant resistance mechanism in relapsed patients who have failed crizotinib treatment (Figure 11B). Deep genomic sequencing of both the parental and resistant MOLM13 cell lines from the dual protocol detected mutations distributed across all chromosomes. In summary, among all single cells, 5675 SNVs specific to the crizotinib-resistant population were identified. Mutations in the coding sequence were detected, however, most of the observed variants were in the intergenic space. Although not bound by theory, passenger mutations are undoubtedly present in this variant cohort, suggesting that regulation of gene expression at the enhancer or promoter level contributes to resistance and potentially the regulation of non-coding RNAs. The dual mRNAseq transcriptome chemistry / PTA has the ability to detect over 10K genes within a single cell, which can be enriched by FACS. The PTA method has the ability to recover over 97% of the complete genome of individual cells. The ability to recover both the transcriptome and genome does not significantly affect the sensitivity of the ability to recover most of the genome. Comparing the transcriptome only or combinations of transcriptome / genome amplification chemistries, over 70% of the expressed genes can be detected in many cells.
[0114] Example 10: PTA Single-Cell Analysis with Exosome Capture. The general PTA method of Example 3 was used with modifications: an additional exome capture step was utilized to enrich the PTA-generated amplicons. Sixty million reads were obtained in both single-cell samples (27 samples) and bulk samples (112 samples). The results of exome capture sequencing from single cells were compared to the results of bulk samples (Figures 12A - 12D, 13A, 14A, and 14B). The sequencing results were consistent among multiple samples (Figure 13A), and the average size of the captured amplicons was 623 bases (Figure 13B).
[0115] Example 11: Exome Capture + Multiomics The general methods of any of Examples 5 - 8 are used with modifications: an additional capture step is utilized to enrich the PTA-generated amplicons generated from genomic DNA. The capture step includes either an exome panel or another panel targeting specific genes. In some cases, such panels are directed to cancer hotspots, viral genomes, or mitochondrial DNA.
[0116] Although the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within these claims, and their equivalents, be covered thereby.
Claims
1. A method for multi-omic single cell analysis, the method comprising: a. isolating single cells from a population of cells; b. sequencing a cDNA library comprising polynucleotides amplified from mRNA transcripts from said single cells; and c. sequencing the genome of said single cell, wherein the step of sequencing the genome comprises: i. contacting the genome with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by said polymerase; ii. amplifying at least some of the genome to produce a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iii. ligating the molecules obtained in step (ii) to adapters to thereby produce a genomic DNA library; and iv. sequencing said genomic DNA library comprising the step of sequencing. A method comprising.
2. The method according to claim 1, wherein at least some of the polynucleotides of said cDNA library comprise barcodes.
3. The method according to claim 1, wherein said cDNA library and said genomic DNA library are pooled prior to said sequencing.
4. The method according to claim 1, further comprising the step of removing at least one terminator nucleotide from said terminated amplification products.
5. The method according to claim 1, wherein said plurality of terminated amplification products comprise bases having an average length of 1000-2000.
6. The method according to claim 1, wherein said plurality of terminated amplification products have a length of 250-1500 bases.
7. The method according to claim 1, wherein said plurality of terminated amplification products comprise at least 97% of the genome of said single cell.
8. The method according to claim 1, wherein the step of sequencing the genome of said single cell further comprises nuclear lysis of said single cell.
9. At least one mutation is identified in the genome of a cell, and said at least one mutation occurs in less than 1% of said population of cells. The method according to claim 1.
10. The step of identifying at least one protein on the surface of said single cell The method according to claim 1, further comprising
11. The method according to claim 10, wherein the step of identifying at least one protein on the surface of the single cell comprises contacting the single cell with a labeled antibody that binds to the at least one protein.
12. A method for multi-omic single cell analysis, the method comprising: a. isolating a single cell from a population of cells; b. sequencing the genome of the single cell, wherein the step of sequencing the genome of the cell comprises: i. digesting the genome with a methylation-sensitive restriction enzyme to generate genomic fragments; ii. contacting at least some of the genomic fragments with a mixture of at least one amplification primer, at least one nucleic acid polymerase, and nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; iii. amplifying at least some of the genome to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; iv. amplifying at least a portion of the genomic fragments by methylation-specific PCR; v. ligating the molecules obtained in steps (iii and iv) to an adapter, thereby generating a genomic DNA library and a methylome DNA library; and vi. sequencing the genomic DNA library and the methylome DNA library comprising the step of sequencing comprising the method.
13. The method according to any one of claims 1 to 12, wherein the terminator nucleotide comprises an irreversible terminator nucleotide.
14. The method according to claim 13, wherein the irreversible terminator nucleotide comprises an alpha-thio-dideoxynucleotide.
Citation Information
Patent Citations
Methods, compositions, and kits for generating libraries of stranded RNA or DNA
JP2016511007A
Nucleic acid sequencing method and system
JP2018527020A
Methods for nucleic acid amplification
JP2021511794A
Gene mutation analysis
JP2022543375A
Methods for closed chromatin mapping and DNA methylation analysis for single cells
US20150368694A1