Gene mutation analysis

Primary template-directed amplification with strand displacement replication and terminator nucleotides addresses the limitations of current methods, enhancing mutation detection accuracy and sensitivity in small samples, particularly in single-cell analysis.

JP7819093B2Active Publication Date: 2026-02-24BIOSKRYB GENOMICS INC +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022506476
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-31
Filing Date
2020-07-30
Publication Date
2026-02-24
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

Current nucleic acid amplification and sequencing methods are inadequate for accurately detecting mutations in small samples exposed to mutagenic conditions, lacking scalability, efficiency, and reproducibility, especially in single-cell analysis.

Method used

A method involving primary template-directed amplification (PTA) using strand displacement replication with terminator nucleotides, followed by sequencing and comparison to reference sequences to identify mutations, and incorporating gene editing techniques like CRISPR for targeted analysis.

Benefits of technology

Enhances the accuracy and sensitivity of mutation detection in small samples, improving sequence representation and uniformity, enabling precise identification of mutations and structural variants in single cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819093000002
    Figure 0007819093000002
  • Figure 0007819093000003
    Figure 0007819093000003
  • Figure 0007819093000004
    Figure 0007819093000004
Patent Text Reader

Abstract

Provided herein are compositions and methods for accurate and scalable primary template-directed amplification (PTA) nucleic acid amplification and sequencing, and their applications for mutation analysis in research, diagnostics, and therapeutics. Such methods and compositions facilitate highly accurate amplification of target (or "template") nucleic acids, which improves the accuracy and sensitivity of downstream applications such as next-generation sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 881,180, filed July 31, 2019, which is incorporated herein by reference in its entirety. [Background technology]

[0002] background Research methods that utilize nucleic acid amplification, such as next-generation sequencing, provide a large amount of information about complex samples, genomes, and other nucleic acid sources. In some cases, these samples have been exposed to mutagenic conditions in the environment or through gene editing techniques. Highly accurate, scalable, and efficient nucleic acid amplification and sequencing methods are needed for research, diagnosis, and treatment involving small samples, such as samples exposed to mutagenic conditions. Summary of the Invention

[0003] overview Described herein are methods for detecting mutations in a sample, genome, or other source of nucleic acid.

[0004] (c) providing a cell lysate from the single cell; (d) contacting the cell lysate with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; (d) amplifying a target nucleic acid molecule to generate a plurality of terminated amplification products, wherein the replication proceeds by strand displacement replication; (e) ligating the molecules obtained in step (e) to adapters, thereby generating a library of amplification products; and (f) sequencing the library of amplification products and comparing the sequences of the amplification products to at least one reference sequence to identify at least one mutation. Further described herein is a method in which at least one mutation is present in the target sequence. Further described herein is a method in which at least one mutation is absent from the target sequence. Further described herein is a gene editing method comprising the use of CRISPR, TALEN, ZFN, recombinase, meganuclease, or viral integration (intentional or unintentional). Further described herein is a method in which the gene editing technique comprises the use of CRISPR. Further described herein is a method in which the gene editing technique comprises the use of gene therapy. Further described herein is a method in which the gene therapy is not configured to modify the somatic or germline DNA of a cell. Further described herein is a method in which the reference sequence is a genome. Further described herein is a method in which the reference sequence is a specificity-determining sequence, wherein the specificity-determining sequence is configured to bind to the target sequence. Further described herein is a method in which the at least one mutation is present in a region of the sequence that differs from the specificity-determining sequence by at least one base.Further described herein is a method wherein the at least one mutation is present in a region of the sequence that differs from the specificity-determining sequence by at least two bases. Further described herein is a method wherein the at least one mutation is present in a region of the sequence that differs from the specificity-determining sequence by at least three bases. Further described herein is a method wherein the at least one mutation is present in a region of the sequence that differs from the specificity-determining sequence by at least five bases. Further described herein is a method wherein the at least one mutation comprises an insertion, deletion, or substitution. Further described herein is a method wherein the reference sequence is a CRISPR RNA (crRNA) sequence. Further described herein is a method wherein the reference sequence is a single guide RNA (sgRNA) sequence. Further described herein is a method wherein the at least one mutation is present in a region of a sequence that binds to catalytically active Cas9. Further described herein is a method wherein the single cell is a mammalian cell. Further described herein is a method wherein the single cell is a human cell. Further described herein is a method wherein the single cell is derived from liver, skin, kidney, blood, or lung. Further described herein is a method wherein the single cell is a primary cell. Further described herein is a method wherein the single cell is a stem cell. Further described herein is a method wherein at least some of the amplification products comprise a barcode. Further described herein is a method wherein at least some of the amplification products comprise at least two barcodes. Further described herein is a method wherein the barcode comprises a cell barcode. Further described herein is a method wherein the barcode comprises a sample barcode. Further described herein is a method wherein at least some of the amplification primers comprise a unique molecular identifier (UMI). Further described herein is a method wherein at least some of the amplification primers comprise at least two unique molecular identifiers (UMI). Further described herein is a method wherein the method further comprises an additional amplification step using PCR.Further described herein is a method further comprising removing at least one terminator nucleotide from the terminated amplification product prior to ligation to an adapter. Further described herein is a method of isolating a single cell from the population using a method comprising a microfluidic device. Further described herein is a method wherein the at least one mutation occurs in less than 50% of the population of cells. Further described herein is a method wherein the at least one mutation occurs in less than 25% of the population of cells. Further described herein is a method wherein the at least one mutation occurs in less than 1% of the population of cells. Further described herein is a method wherein the at least one mutation occurs in 0.1% or less of the population of cells. Further described herein is a method wherein the at least one mutation occurs in 0.01% or less of the population of cells. Further described herein is a method wherein the at least one mutation occurs in 0.001% or less of the population of cells. Further described herein is a method wherein the at least one mutation occurs in 0.0001% or less of the population of cells. Further described herein is a method wherein the at least one mutation occurs in 25% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation occurs in 1% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation occurs in 0.1% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation occurs in 0.01% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation occurs in 0.001% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation occurs in 0.0001% or less of the amplified product sequences. Further described herein is a method wherein the at least one mutation is present in a region of a sequence correlated with a genetic disease or condition. Further described herein are methods, wherein the at least one mutation is in a region of the sequence that is not correlated with binding of a DNA repair enzyme.Further described herein is a method, wherein the at least one mutation is present in a region of the sequence that is not correlated with the binding of MRE11. Further described herein is a method, further comprising identifying the false positive mutations that have been previously sequenced by alternative off-target detection methods. Further described herein is a method, wherein the off-target detection method is in silico prediction, ChIP-seq, GUIDE-seq, circle-seq, HTGTS (high-throughput genome-wide translocation sequencing), IDLV (integration-deficient lentivirus), Digenome-seq, FISH (fluorescence in situ hybridization) or DISCOVER-seq.

[0005] Described herein is a method for identifying a specificity-determining sequence, the method comprising: (a) providing a library of nucleic acids, wherein at least some of the nucleic acids comprise a specificity-determining sequence; (b) performing a gene editing method on at least one cell, wherein the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; (c) sequencing the genome of the at least one cell using the method described herein, wherein the specificity-determining sequence contacted with the at least one cell is identified; and (d) identifying at least one specificity-determining sequence that provides the fewest off-target mutations. Further described herein is a method in which the off-target mutations are synonymous or non-synonymous mutations. Further described herein is a method in which the off-target mutations are located outside the coding region of a gene.

[0006] Described herein is a method for in vivo mutation analysis, the method comprising: (a) performing a gene editing method on at least one cell in an organism, wherein the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; (b) isolating at least one cell from the organism; and (d) sequencing the genome of the at least one cell using the method described herein. Further described herein is a method comprising at least two cells. Further described herein is a method further comprising identifying mutations by comparing the genome of a first cell with the genome of a second cell. Further described herein is a method wherein the first cell and the second cell are from different tissues.

[0007] Described herein are methods for predicting the age of a subject, the methods comprising: (a) providing at least one sample from the subject, wherein the at least one sample comprises a genome; (b) sequencing the genome using a method described herein to identify mutations; (c) comparing the mutations obtained in step b to a standard reference curve, wherein the standard reference curve correlates the number and location of mutations with a verified age; and (d) predicting the age of the subject based on comparison of the mutations to the standard reference curve. Further described herein are methods wherein the standard reference curve is specific to the gender of the subject. Further described herein are methods wherein the standard reference curve is specific to the ethnicity of the subject. Further described herein are methods wherein the standard reference curve is specific to the geographic location of the subject where the subject spent a period of their lifetime. Further described herein are methods wherein the subject is under 50 years of age. Further described herein are methods wherein the subject is under 18 years of age. Further described herein is a method wherein the subject is under 15 years old. Further described herein is a method wherein the at least one sample is more than 10 years old. Further described herein is a method wherein the at least one sample is more than 100 years old. Further described herein is a method wherein the at least one sample is more than 1000 years old. Further described herein is a method wherein at least two samples are sequenced. Further described herein is a method wherein at least five samples are sequenced. Further described herein is a method wherein the at least two samples are from different tissues.

[0008] Described herein are methods for sequencing microbial or viral genomes, comprising: (a) obtaining a sample containing one or more genomes or genome fragments; (b) sequencing the sample using the methods described herein to obtain a plurality of sequencing reads; and (c) assembling and sorting the sequencing reads to generate a microbial or viral genome from a single bacterial cell or even a single viral particle. Further described herein are methods wherein the sample contains genomes from at least two organisms. Further described herein are methods wherein the sample contains genomes from at least 10 organisms. Further described herein are methods wherein the sample contains genomes from at least 100 organisms. Further described herein are methods wherein the sample originates from an environment, including a deep-sea vent, ocean, mine, stream, lake, meteorite, glacier, or volcano. Further described herein are methods further comprising identifying at least one gene in the microbial genome. Further described herein are methods wherein the microbial genome corresponds to an unculturable organism. Further described herein is a method wherein the microbial genome corresponds to a symbiont. Further described herein is a method further comprising cloning at least one gene in a recombinant host organism. Further described herein is a method wherein the recombinant host organism is a bacterium. Further described herein is a method wherein the recombinant host organism is Escherichia, Bacillus, or Streptomyces. Further described herein is a method wherein the recombinant host organism is a eukaryotic cell. Further described herein is a method wherein the recombinant host organism is a yeast cell. Further described herein is a method wherein the recombinant host organism is Saccharomyces or Pichia.

[0009] The present invention provides a kit for nucleic acid sequencing, comprising at least one amplification primer, at least one nucleic acid polymerase, a mixture of at least two nucleotides, the mixture comprising at least one terminator nucleotide that terminates nucleic acid replication by the polymerase, and instructions for using the kit to perform nucleic acid sequencing.The present invention also provides a kit in which at least one amplification primer is a random primer.The present invention also provides a kit in which the nucleic acid polymerase is a DNA polymerase.The present invention also provides a kit in which the DNA polymerase is a strand-displacing DNA polymerase. Further described herein is a nucleic acid polymerase, wherein the nucleic acid polymerase is selected from the group consisting of bacteriophage phi 29 (Φ29) polymerase, genetically modified phi 29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent RThe kit is an exo-DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, terminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, or T4 DNA polymerase. Further described herein is a kit in which the nucleic acid polymerase comprises 3' to 5' exonuclease activity and at least one terminator nucleotide inhibits the 3' to 5' exonuclease activity. Further described herein is a kit in which the nucleic acid polymerase does not comprise 3' to 5' exonuclease activity. Further described herein is a kit in which the polymerase is Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, or terminator DNA polymerase. Further described herein are kits in which at least one terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. Further described herein are kits in which at least one terminator nucleotide is selected from the group consisting of 3'-blocked reversible terminator containing nucleotides, 3'-unblocked reversible terminator containing nucleotides, terminators comprising a 2'-modification of a deoxynucleotide, terminators comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. Further described herein are kits in which at least one terminator nucleotide is selected from the group consisting of dideoxynucleotides, inverted dideoxynucleotides, 3' biotinylated nucleotides, 3' amino nucleotides, 3' phosphorylated nucleotides, 3'-O-methyl nucleotides, 3' C3 spacer nucleotides, 3' C18 nucleotides, 3' carbon spacer nucleotides including 3' hexanediol spacer nucleotides, acyclonucleotides, and combinations thereof. Further described herein are kits in which at least one terminator nucleotide is selected from the group consisting of nucleotides with a modification in an alpha group, C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2' fluoronucleotides, 3' phosphorylated nucleotides, 2'-O-methyl modified nucleotides, and trans nucleic acids. Further described herein are kits in which the nucleotide with a modification in an alpha group is an alpha-thiodideoxynucleotide. Further described herein are kits in which the amplification primers are 4 to 70 nucleotides in length. Further described herein are kits wherein at least one amplification primer is 4 to 20 nucleotides in length. Further described herein are kits wherein at least one amplification primer comprises a randomized region.Further described herein is a kit wherein the randomized region is 4 to 20 nucleotides in length. Further described herein is a kit wherein the randomized region is 8 to 15 nucleotides in length. Further described herein is a kit further comprising a library preparation kit. Further described herein is a kit wherein the library preparation kit comprises one or more of at least one polynucleotide adaptor, at least one high-fidelity polymerase, at least one ligase, a reagent for nucleic acid shearing, and at least one primer. Further described herein is a kit further comprising reagents configured for gene editing.

[0010] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]

[0011] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Figure 1A] We illustrate a workflow for detecting mutations using PTA, single-cell sequencing, and alignment: Edited and unedited control cells are amplified using PTA, sequenced using short read sequencing, and aligned to a reference genome. [Figure 1B]Illustrates the detection of small indels. Indels (black ovals) are identified by comparing aligned sequence data to a reference genome using variant calling software. Potential candidate indels for CRISPR editing events are identified by comparing indels between edited and unedited control cells and restricting the search space to regions of the genome that show sequence similarity to the gRNA target site. Evidence for candidate editing events includes 1) indels located 3–4 bases upstream from the putative PAM sequence in the genomic region that shows similarity to the target site, and 2) the restriction of these indels to edited cells with no evidence in unedited control cells. [Figure 1C] The detection of translocations and large deletions is illustrated. Comparison of read pair mapping patterns between edited and unedited cells allows for the identification of CRISPR-induced structural variants, including inter- and intrachromosomal translocations, inversions, and large deletions. CRISPR-induced translocations are identified by read pair alignment in edited cells, where at least two regions of the read pair align to different chromosomes and the breakpoints are located in regions that show similarity to the gRNA target sequence. These discordant read pairs should not be present in the alignment in unedited cells (Figure 1C). Large deletions are identified by read pairs that show the correct orientation but contain regions that align to distant parts of the reference genome (Figure 1D). [Figure 1D]The detection of translocations and large deletions is illustrated. Comparison of read pair mapping patterns between edited and unedited cells allows for the identification of CRISPR-induced structural variants, including inter- and intrachromosomal translocations, inversions, and large deletions. CRISPR-induced translocations are identified by read pair alignment in edited cells, where at least two regions of the read pair align to different chromosomes and the breakpoints are located in regions that show similarity to the gRNA target sequence. These discordant read pairs should not be present in the alignment in unedited cells (Figure 1C). Large deletions are identified by read pairs that show the correct orientation but contain regions that align to distant parts of the reference genome (Figure 1D). [Figure 1E] A comparison of the conventional multiple displacement amplification (MDA) method with one embodiment of the primary template-directed amplification (PTA) method, namely the PTA irreversible terminator method, is illustrated. [Figure 1F] 1 illustrates a comparison of the PTA irreversible terminator method with a different embodiment, namely the PTA reversible terminator method. [Figure 1G] Illustrates a comparison of the MDA and PTA irreversible terminator methods in relation to mutation propagation. [Figure 1H] The method steps performed after amplification are illustrated, including terminator removal, end repair, and A-tailing prior to adapter ligation. The pooled cell library can then be enriched for all exons or other specific regions of interest via hybridization prior to sequencing. The cell of origin of each read is identified by its cell barcode (shown as a green and blue sequence). [Figure 2A] The size distribution of amplicons after PTA with increasing concentrations of terminators (top gel) is shown. The bottom gel shows the size distribution of amplicons after PTA with the addition of increasing concentrations of reversible terminators or increasing concentrations of irreversible terminators. [Figure 2B](GC) Comparison of the GC content of the sequenced bases of MDA and PTA is shown. [Figure 2C] The map quality score (e) (mapQ) for single cells mapping to the human genome (p_mapped) after undergoing PTA or MDA is shown. [Figure 2D] The percentage of reads that map to the human genome (p_mapped) after single cells underwent PTA or MDA is shown. [Figure 2E] (PCR) Comparison of the percent of reads that are PCR replicates for 20 million subsampled reads after single cells underwent MDA and PTA. [Figure 2F] Amplification kinetics as amplicon yield versus time (hr) for MDA, MDA no template control (NTC), PTA, and PTA no template control (NTC) are shown. [Figure 3A] Map quality scores (c) (mapQ2) for single cells mapping to the human genome after undergoing PTA with reversible or irreversible terminators (p_mapped2) are shown. [Figure 3B] The percentage of reads that map to the human genome (p_mapped2) after single cells underwent PTA with reversible or irreversible terminators is shown. [Figure 3C] A series of boxplots illustrating the reads aligned for the average percent reads overlapping with Alu elements using various methods are shown. PTA had the highest number of reads aligned to the genome. [Figure 3D] A series of boxplots illustrating PCR replicates of the average percent reads overlapping with Alu elements using various methods are shown. [Figure 3E] A series of boxplots illustrating the GC content of reads for the average percent reads overlapping with Alu elements using various methods are shown. [Figure 3F]A series of boxplots illustrating the mapping quality of the average percent reads overlapping Alu elements using various methods are shown. PTA had the highest mapping quality among the methods tested. [Figure 3G] Comparison of SC mitochondrial genome coverage widths by different WGA methods at a fixed 7.5x sequencing depth. [Figure 4A] Figure 1 shows the average coverage depth of 10-kilobase windows across chromosome 1 after selecting high-quality MDA cells (representing ~50% of cells) compared to random-primed PTA-amplified cells after downsampling each cell to 40 million paired reads. The figure demonstrates low MDA uniformity, with many windows having more (Box A) or less (Box C) than twice the average coverage depth. At the centromere, coverage is absent for both MDA and PTA (Box B) due to the high GC content and poor mapping quality of repetitive regions. [Figure 4B] Plots of sequencing coverage versus genome location for MDA and PTA methods are shown (top). Boxplots below show allele frequencies for MDA and PTA methods compared to the bulk sample. [Figure 5A] Plots of percent genome covered versus number of genome reads are shown to assess coverage at increasing sequencing depths for various methods. The PTA method came close to the two bulk samples at all depths, an improvement over the other methods tested. [Figure 5B] Figure 1 shows a plot of the coefficient of variation of genome coverage versus the number of reads to assess coverage uniformity. The PTA method was found to have the highest uniformity among the methods tested. [Figure 5C] A Lorenz plot of cumulative percentage of total reads versus cumulative percentage of genome is shown. The PTA method was found to have the highest uniformity among the methods tested. [Figure 5D]A series of boxplots of the Gini index calculated for each method tested to estimate the deviation of each amplification reaction from perfect uniformity are shown. The PTA method was found to be more reproducibly uniform than the other methods tested. [Figure 5E] A plot of the percentage of bulk variants called versus the number of reads is shown. The variant call percentage for each method was compared to the corresponding bulk sample at increasing sequencing depths. To estimate sensitivity, we calculated the percent of variants called in the corresponding bulk sample subsampled to 650 million reads found in each cell at each sequencing depth (Figure 5A). The improved coverage and uniformity of PTA led to the detection of 30% more variants than the Q-MDA method, which was the next most sensitive method. [Figure 5F] A series of boxplots show the average percent reads overlapping Alu elements. The PTA method significantly reduced allelic skew at these heterozygous sites. The PTA method amplifies two alleles more evenly within the same cell compared to other methods tested. [Figure 5G] To assess the accuracy of mutation calling, a plot of variant calling accuracy versus the number of reads is shown. Variants found using various methods that were not found in the bulk sample were considered false positives. The PTA method produced the lowest false positive calls (highest accuracy) of the methods tested. [Figure 5H] The rate of false positive base changes for each type of base change across the various methods is shown. Without being bound by theory, it is possible that such patterns may be polymerase dependent. [Figure 5I] A series of boxplots of the average percent reads overlapping Alu elements for false-positive variant calls are shown. The PTA method yielded the lowest allele frequencies for false-positive variant calls. [Figure 5J]Shown are the mean coefficients of variation (CV) of coverage at increasing bin sizes in primary leukemia samples using a commercial kit as an estimate of CNV calling accuracy. [Figure 5K] CNV profiles of PTA products from single cells are shown for chromosomes where CNVs were called in the bulk sample (shaded arrows). Unshaded arrows represent regions where subclonal CNVs were suggested but not called in the bulk sample; two of five cells were found to have the same alteration. Regions of the karyogram with reduced CNV detection represent centromeres, indicating reduced coverage in PTA-amplified cells. (For dot and line plots, error bars represent 1 SD; for box plots, the center line is the median. Box limits represent upper and lower quartiles, and whiskers represent the 1.5x interquartile range. Dots indicate outliers.) [Figure 6A] 1 shows a schematic representation of a catalog of clonotype drug susceptibilities according to the present disclosure. By identifying the drug susceptibilities of different clonotypes, a catalog can be created that allows oncologists to translate the clonotypes identified in a patient's tumor into a list of drugs that best target the resistant population. [Figure 6B] Figure 1 shows the change in the number of leukemic clones with increasing numbers of leukemic cells per clone after 100 simulations. Using the mutation rate per cell, the simulation predicts the enormous diversity of smaller clones created when a single cell expands to 10-100 billion cells (Box A). Using current sequencing methods, only the most frequent 1-5 clones (Box C) are detected. In one embodiment of the present invention, a method is provided for determining drug resistance of hundreds of clones (Box B), which are just below the detection level of current methods. [Figure 7] Illustrates an exemplary embodiment of the present disclosure. Compared to the diagnostic samples in the bottom row, clones containing activating KRAS mutations (red boxes, bottom right corner) were selected for culture without chemotherapy. Conversely, the clones were killed by prednisolone or daunorubicin (green boxes, top right corner), while low-frequency clones underwent positive selection (dashed boxes). [Figure 8] FIG. 1 is an overview of one embodiment of the present disclosure, namely, an experimental design for quantifying the relative sensitivity of clones with specific genotypes to a particular drug. [Figure 9] (Part A) shows beads bearing oligonucleotides with cleavable linkers, unique cell barcodes, and random primers. Part B shows a single cell and bead encapsulated in the same droplet, after which the cell is lysed and the primer is cleaved. The droplet can then be fused with another droplet containing a PTA amplification mixture. Part C shows that after amplification, the droplets are disrupted and the amplicons from all cells are pooled. Next, a protocol according to the present disclosure is utilized for terminator removal, end repair, and A-tailing prior to adapter ligation. The library of pooled cells then undergoes hybridization-mediated enrichment of exons of interest prior to sequencing. The cell of origin of each read is then identified using the cell barcode. [Figure 10A] 1 shows the incorporation of a cell barcode and / or a unique molecular identifier into a PTA reaction using primers containing the cell barcode and / or a unique molecular identifier. [Figure 10B] 1 shows the incorporation of a cell barcode and / or a unique molecular identifier into a PTA reaction using a hairpin primer containing the cell barcode and / or a unique molecular identifier. [Figure 11A] (PTA_UMI) indicates that the incorporation of unique molecular identifiers (UMIs) enables the creation of consensus reads and reduces the false discovery rate caused by sequencing and other errors, leading to increased sensitivity when performing germline or somatic variant calling. [Figure 11B] We show that collapsing reads at the same UMI allows for the correction of amplification and other biases that can result in false positives or limited sensitivity when calling copy number variants. [Figure 12A]Plot of number of mutations versus treatment group for direct measurement of environmental mutagenicity experiments. Single human cells were exposed to vehicle (VHC), mannose (MAN), or the direct mutagen N-ethyl-N-nitrosourea (ENU) at different treatment levels, and the number of mutations was measured. [Figure 12B] A series of plots of the number of mutations versus different treatment groups and levels is shown, further separated by type of base mutation. [Figure 12C] Shown is a representation of mutation patterns in a trinucleotide context. The base on the y-axis is at position n-1, and the base on the x-axis is at position n+1. Dark areas indicate low mutation frequency, and light areas indicate high mutation frequency. The solid black box in the top row (cytosine mutations) indicates a low frequency of cytosine mutagenesis when cytosine is followed by guanine. The dashed black box in the bottom row (thymine mutations) indicates that most thymine mutations occur at positions where an adenine immediately precedes the thymine. [Figure 12D] Graph comparing the locations of known DNase I hypersensitive sites in CD34+ cells with corresponding locations from cells treated with N-ethyl-N-nitrosourea. No significant enrichment of cytosine variants was observed. [Figure 12E] Figure 1 shows the percentage of ENU-induced mutations at DNase I hypersensitive (DH) sites. DH sites in CD34+ cells previously cataloged by the Roadmap Epigenomics Project were used to investigate whether ENU mutations are more prevalent at DH sites, which represent sites of open chromatin. No significant enrichment for variant positions at DH sites was identified, and no enrichment for variants restricted to cytosines was observed at DH sites. [Figure 12F] A series of boxplots of the proportion of ENU-induced mutations at genomic locations with specific annotations are shown. No specific enrichment was observed for specific annotations of variants in each cell (left boxes) compared to the proportion of the genome each annotation contains (right boxes). [Figure 13A]Shown are indel counts in edited versus non-edited cells within a Hamming distance of 7 of the target site after genome editing experiments and PTA. [Figure 13B] Shown are structural variant counts in edited versus non-edited cells within a Hamming distance of 6 of the target site in a genome editing experiment and after PTA. [Figure 14A] 1 shows detection of CRISPR-induced editing in two edited single cells using PTA. [Figure 14B] 1 shows detection of large (>1 KB) deletions resulting from CRISPR-induced editing restricted to cell #1, which was edited using PTA. [Figure 14C] Figure 1 shows the detection of an interchromosomal translocation between chromosome 2 position 241,275,213 and chromosome 4 position 38,536,006 in cell #1 edited using PTA. [Figure 15A] Alignment and SNV calling metrics in primary leukemia cells with increasing sequencing depth of coverage width are shown (n=5 for each method, error bars represent 1 SD). [Figure 15B] Alignment and SNV calling metrics in primary leukemia cells upon increasing sequencing depth of CV coverage (n=5 for each method, error bars represent 1 SD). [Figure 15C] Alignment and SNV calling metrics in primary leukemia cells are shown with increasing sequencing depth for calling sensitivity (n=5 for each method, error bars represent 1 SD). [Figure 15D] Alignment and SNV calling metrics in primary leukemia cells with increasing sequencing depth for SNV calling accuracy (n=5 for each method, error bars represent 1 SD). [Figure 16A] An overview of a kinship cell experiment is shown, in which single cells are plated and cultured prior to individual cell re-isolation, PTA, and sequencing. [Figure 16B]We present a method for classifying variant types by comparing bulk and single-cell data. [Figure 16C] The sensitivity and accuracy of SNV calling for each cell using the bulk as a standard are shown. [Figure 16D] The percentage of variants called heterozygous for the different variant classes is shown. [Figure 16E] False positive and somatic mutation rates measured in single CD34+ human cord blood cells are shown. [Figure 17A] An overview of the number of mutations in each sample for all variants is shown. [Figure 17B] Shows an overview of the number of mutations in each sample for somatic variants. [Figure 17C] An overview of the number of mutations in each sample for false positive variants is shown. [Figure 18A] An overview of allele frequency distributions for germline variants is shown. [Figure 18B] An overview of allele frequency distributions for somatic variants is shown. [Figure 18C] An overview of the allele frequency distribution for false positive variants is shown. [Figure 19] Figure 1 shows the density of homozygous or heterozygous false positive variant calls across chromosome 14, which had the highest number of false positive calls. The average GC content in 100 Kb intervals is below the karyotype. [Figure 20A] We present experimental and computational methods for measuring the off-target activity of genome editing strategies at single-cell resolution, where single-edited cells are sequenced and indel calling is limited to sites with up to five mismatches to the protospacer. [Figure 20B]The number of indel calls per cell is shown. Each control or experimental cell type received an indel call if the target region had a maximum of 5-base mismatch with either the VEGFA or EMX1 protospacer sequence. The gRNA or control listed in the key identifies which gRNA the cell received. If an indel is called in a genomic region that does not match the gRNA received by the cell, it is presumed to be a false positive. [Figure 20C] A table of the total number of off-target indel positions called that were either unique to one cell or found in multiple cells is shown. [Figure 20D] Genomic locations of recurrent indels with EMX1 or VEGFA gRNAs are shown. On-target sites are annotated in gray. [Figure 20E] Circosmotic plots of SVs identified in each cell type that received either EMX1 or VEGFA gRNA are shown, with sites that contained at least one recurrent breakpoint seen across cell types in green or only in that cell type in red. The number of SVs detected per cell is plotted on the right (for boxplots, the center line is the median, box limits represent the upper and lower quartiles, and whiskers represent the 1.5x interquartile range; dots indicate outliers). [Figure 21] This experiment demonstrates how removing non-recurrent single-base pair insertions improved the accuracy of off-target detection. Each control or experimental cell type received an indel call requiring five or fewer mismatches to either the VEGFA or EMX1 guide RNA sequence. Off-target events identify the genomic region to which the gRNA must match, while the gRNA or control listed in the key identifies which gRNA the cell received. If an indel is called in a genomic region that does not match the gRNA received by the cell, it is presumed to be a false positive. [Figure 22A] The longest contig length of any bacterial sample analyzed using the PTA method is shown. [Figure 22B]A graph for each sample containing the ratio of cumulative length to cumulative contig length is shown, as well as the closest hit genus for each sample based on sequence alignment to the genome. [Figure 22C] A graph of the 10 bacterial samples for the ratio of cumulative length to cumulative contig length is shown, along with the closest hit genera for each sample based on sequence alignment to the genomes of Haemophilus and Streptococcus. [Figure 22D] For each bacterial sample tested, the read pairs aligning to human chromosomes are shown. [Figure 22E] A scheme for assigning reads as of human origin is shown. [Figure 22F] Read pair mapping positions of all pairs with at least one human-mapped read are shown for all bacterial samples tested. [Figure 22G] Taxonomic ranks for the assignment of contigs belonging to bacterial sample 10 are shown. DETAILED DESCRIPTION OF THE INVENTION

[0012] Detailed Description of the Invention There is a need to develop new, scalable, accurate, and efficient methods for nucleic acid amplification (including single-cell and multi-cell genome amplification) and sequencing that overcome the limitations of current methods by reproducibly increasing sequence representation, uniformity, and accuracy. Provided herein are compositions and methods for accurate and scalable primary template-directed amplification (PTA) and sequencing. Such methods and compositions facilitate highly accurate amplification of target (or "template") nucleic acids, thereby improving the accuracy and sensitivity of downstream applications such as next-generation sequencing. Also provided herein are methods for determining single-base variants, copy number diversity, structural diversity, clonotyping, and measuring environmental mutagenicity. Measuring genomic diversity by PTA can be used for a variety of applications, including environmental mutagenicity, predicting the safety of gene editing technologies, measuring genomic changes mediated by cancer therapy, measuring the carcinogenicity of compounds or radiation, including genotoxicity studies to determine the safety of new foods or drugs, age estimation, analysis of resistant bacteria, and identifying environmental bacteria for industrial applications. Furthermore, these methods can be used to detect the selection of specific cell populations after changes in environmental conditions, such as exposure to anti-cancer treatments, as well as to predict response to immunotherapy based on mutations and neoantigen load in single cancer cells.

[0013] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these inventions belong.

[0014] Throughout this disclosure, numerical characteristics are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges, as well as individual numerical values ​​to the tenth of the lower limit within that range, unless the context clearly dictates otherwise. For example, description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual values ​​within that range (such as 1.1, 2, 2.3, 5, and 5.9). This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges and are also included in the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention, unless the context clearly dictates otherwise.

[0015] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiment. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, it is understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any or all combinations of one or more of the associated listed items.

[0016] As used herein, unless specifically stated or clear from the context, the term "about" in reference to a number or range of numbers is understood to mean the stated number and + / -10% of that number, i.e., 10% below the lowest recited limit and 10% above the highest recited limit for the values ​​recited for a range.

[0017] As used herein, the term "subject" or "patient" or "individual" refers to, for example, humans, veterinary animals (e.g., cats, dogs, cows, horses, sheep, pigs, etc.), and experimental animal models of disease (e.g., mice, rats). In accordance with the present invention, conventional molecular biology, microbiology, and recombinant DNA techniques may be employed that are within the skill of the art. Such techniques are fully explained in the literature. For example, among others, see Sambrook, Fritsch & Maniatis, Molecular Cloning: A Laboratory Manual, Second Edition (1989), Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (referred to herein as "Sambrook et al., 1989"); DNA Cloning: A Practical Approach, Volumes I and II (D.N. Glover ed. 1985); Oligonucleotide Synthesis (M.J. Gait ed. 1984); Nucleic Acid Hybridization (B.D. Hames & S.J. Higgins eds. (1985); Transcription and Translation (B.D. Hames & S.J. Higgins, eds. (1984); Animal Cell Culture (R.I. Freshney, ed. (1986); Immobilized Cells and Enzymes (L.R. Press, (1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); F.M. Ausubel et al. (eds.), Current See Protocols in Molecular Biology, John Wiley & Sons, Inc. (1994).

[0018] The term "nucleic acid" encompasses not only single-stranded molecules but also multi-stranded molecules. In double- or triple-stranded nucleic acids, the nucleic acid strands need not be coextensive (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). The nucleic acid templates described herein can be of any size (from small cell-free DNA fragments to entire genomes) depending on the sample, including, but not limited to, lengths of 50-300 bases, 100-2000 bases, 100-750 bases, 170-500 bases, 100-5000 bases, 50-10,000 bases, or 50-2000 bases. In some examples, the length of the template is at least 50, 100, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 50,000, 100,000, 200,000, 500,000, 1,000,000 bases, or more than 1,000,000 bases. The methods described herein provide for the amplification of nucleic acids, such as nucleic acid templates. The methods described herein also provide for the production of isolated and at least partially purified nucleic acids and libraries of nucleic acids. Nucleic acids include, but are not limited to, DNA, RNA, circular RNA, mtDNA (mitochondrial DNA), cfDNA (cell-free DNA), cfRNA (cell-free RNA), siRNA (small interfering RNA), cffDNA (cell-free fetal DNA), mRNA, tRNA, rRNA, miRNA (microRNA), synthetic polynucleotides, polynucleotide analogs, any other nucleic acids consistent with the present specification, or any combination thereof. The length of a polynucleotide, when provided, is stated as the number of bases and expressed in abbreviations such as nt (nucleotides), bp (bases), kb (kilobases), or Gb (gigabases).

[0019] As used herein, the term "droplet" refers to a volume of liquid on a droplet actuator. A droplet, in some examples, may be aqueous or non-aqueous, or a mixture or emulsion including aqueous and non-aqueous components. For non-limiting examples of droplet fluids that may be subjected to droplet operations, see, e.g., International Patent Application Publication No. WO 2007 / 120241. Any suitable system for forming and manipulating droplets can be used in the embodiments presented herein. For example, in some examples, a droplet actuator is used. Non-limiting examples of droplet actuators that may be used are described, for example, in U.S. Pat. Nos. 6,911,132, 6,977,033, 6,773,566, 6,565,727, 7,163,612, 7,052,244, 7,328,979, 7,547,380, 7,641,779, U.S. Patent Application Publication Nos. US20060194331, US20030205632, and US20060194331. See US20060164490, US20070023292, US20060039823, US20080124252, US20090283407, US20090192044, US20050179746, US20090321262, US20100096266, US20110048951, International Patent Application Publication No. WO2007 / 120241. In some cases, beads are provided in a droplet, in a droplet operations gap, or on a droplet operations surface. In some cases, the beads are provided in a reservoir located outside the droplet operations gap or away from the droplet operations surface, and the reservoir may be associated with a flow path that allows droplets containing the beads to enter the droplet operations gap or contact the droplet operations surface.Non-limiting examples of droplet actuator technologies for immobilizing magnetically responsive and / or non-magnetically responsive beads and / or for performing droplet manipulation protocols using beads are described in U.S. Patent Application Publication No. US20080053205, International Patent Application Publication Nos. WO2008 / 098236, WO2008 / 134153, WO2008 / 116221, and WO2007 / 120241. Bead properties can be utilized in multiplexed embodiments of the methods described herein. Examples of beads with properties suitable for multiplexing, as well as methods for detecting and analyzing signals emitted from such beads, can be found in U.S. Patent Application Publication Nos. US20080305481, US20080151240, US20070207513, US20070064990, US20060159962, US20050277197, and US20050118574.

[0020] As used herein, the term "unique molecular identifier (UMI)" refers to a unique nucleic acid sequence attached to each of a plurality of nucleic acid molecules.When incorporated into nucleic acid molecules, UMIs are sometimes used to correct subsequent amplification bias by directly counting the sequenced UMIs after amplification.The design, incorporation and application of UMIs are described, for example, in International Patent Application Publication No. WO 2012 / 142213, Islam et al. Nat. Methods (2014) 11:163-166, Kivioja, T. et al. Nat. Methods (2012) 9:72-74, Brenner et al. (2000) PNAS 97(4), 1665, and Hollas and Schuler (2003) Conference: 3rd International Workshop on Algorithms in Bioinformatics, Volume: 2812.

[0021] As used herein, the term "barcode" refers to a nucleic acid tag that can be used to identify a sample or source of nucleic acid material. Thus, when nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample are sometimes tagged with different nucleic acid tags so that the source of the sample can be identified. Barcodes are also commonly referred to as indexes, tags, etc., and are well known to those skilled in the art. Any suitable barcode or set of barcodes can be used. For example, see the non-limiting examples provided in U.S. Patent No. 8,053,192 and International Patent Application Publication No. WO2005 / 068656. Barcoding of single cells can be performed, for example, as described in U.S. Patent Application No. 2013 / 0274117.

[0022] The terms "solid surface," "solid support," and other grammatical equivalents herein refer to any material suitable for, or that can be modified to be suitable for, attachment of the primers, barcodes, and sequences described herein. Exemplary substrates include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon, etc.), polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials (e.g., silicon or modified silicon), carbon, metals, inorganic glass, plastics, fiber optic bundles, and various other polymers. In some embodiments, the solid support comprises a patterned surface suitable for immobilizing primers, barcodes, and sequences in an ordered pattern.

[0023] As used herein, the term "biological sample" includes, but is not limited to, tissues, cells, biological fluids, and isolates thereof. Cells or other samples used in the methods described herein may be isolated from human patients, animals, plants, soil, or other samples containing microorganisms such as bacteria, fungi, and protozoa. In some cases, the biological sample is of human origin. In some cases, the biological sample is of non-human origin. Cells may be subjected to the PTA method and sequencing described herein. Variants detected throughout the genome or at specific locations can be compared with all other cells isolated from the subject to trace the lineage history of the cell for research or diagnostic purposes.

[0024] The terms "accuracy" and "specificity" are sometimes used synonymously. In some cases, accuracy (or positive predictive value) is defined as the number of true positive hits divided by the total number of positive hits identified (true positives + false positives).

[0025] The term "cycle" when used in reference to a polymerase-mediated amplification reaction is used herein to describe the steps of dissociating at least a portion of the double-stranded nucleic acid (e.g., denaturing the template from the amplicon, or the double-stranded template), hybridizing (annealing) at least a portion of the primer to the template, and extending the primer to generate an amplicon. In some cases, the temperature remains constant during the cycle of amplification (such as an isothermal reaction). In some cases, the number of cycles directly correlates with the number of amplicons generated. In some cases, the number of cycles in an isothermal reaction is controlled by the length of time the reaction is allowed to proceed.

[0026] Methods and Applications Described herein are methods for identifying mutations in cells using PTA. The use of PTA, in some cases, provides an improvement over known methods such as MDA. PTA, in some cases, results in lower false-positive and false-negative variant call rates than MDA. Genomes such as the NA12878 Platinum genome are used in some cases to determine whether greater genome coverage and PTA uniformity result in lower false-negative variant call rates. Without being bound by theory, it may be determined that the lack of error propagation in PTA reduces the false-positive variant call rate. The amplification balance between alleles by the two methods is, in some cases, estimated by comparing the allele frequencies of heterozygous mutation calls at known positive loci. In some cases, the amplicon library generated using PTA is further amplified by PCR. In some cases, the PTA method identifies mutations present in a single cell of a population, and the mutations detected by PTA occur in less than 2%, 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, 0.01%, 0.001%, 0.0001%, or 0.00001% of the cells in the population. In some cases, the PTA method identifies mutations in less than 2%, 1%, 0.5%, 0.2%, 0.1%, 0.05%, 0.02%, 0.01%, 0.001%, 0.0001%, or 0.00001% of the sequencing reads for a given base or region.

[0027] Gene editing safety The continued development of genome editing tools shows great promise for improving human health, from correcting genes that result in or contribute to the formation of diseases (such as sickle cell anemia and many other disorders) to eradicating currently incurable infectious diseases. However, the safety of these interventions remains unknown as a result of our incomplete understanding of how these tools interact with and permanently alter other locations within the genome of edited cells. While methods have been developed to estimate the off-target rate of genome editing strategies, tools developed to date investigate groups of cells together, making it impossible to measure cell-by-cell off-target rates and cell-to-cell variation in off-target activity, as well as to detect rare editing events occurring in small numbers of cells. These suboptimal strategies for measuring genome editing fidelity have resulted in a limited ability to determine the sensitivity and precision of a given genome editing approach.

[0028] Gene therapy approaches can involve modifying mutated disease-causing genes, knocking out disease-causing genes, or introducing new genes into cells. In some cases, these approaches involve modifying genomic DNA. In other cases, viruses or other delivery systems are configured so that they do not integrate or modify genomic DNA within cells. However, such systems may still result in undesired or unexpected modifications to somatic or germline DNA. Taking advantage of the improved variant calling sensitivity and accuracy of PTA in single cells, in some cases, quantitative measurement of the unintended insertion rate of gene therapy approaches can be performed with high sensitivity in single cells. In some cases, this method detects the insertion of specific sequences at undesired locations by detecting surrounding sequences to determine whether the gene therapy approach causes insertion or modification of the host genome.

[0029] Described herein are methods for identifying mutations and structural modifications (i.e., translocations, insertions, and deletions) in animal, plant, or microbial cells that have undergone genome editing (e.g., CRISPR (clustered regularly interspaced short palindromic repeats), TALEN (transcription activator-like effector nucleases), ZFN (zinc finger nucleases), recombinases, meganucleases, viral integration, or other genome editing techniques). In some embodiments, the genome editing is unintentional or is a secondary effect of another process. In some cases, the genome editing involves site-specific or targeted genome editing. Such cells, in some cases, can be isolated, subjected to PTA, and sequenced to determine the mutation load, mutation combinations, and structural variations in each cell. The mutation rate and location of mutations per cell resulting from a genome editing protocol may, in some cases, be used to evaluate the safety and / or efficiency of a given genome editing method. Identifying mutations, in some cases, involves comparing sequence data obtained using PTA methods to a reference sequence. In some cases, the reference sequence is a genome. In some cases, after the gene editing process, at least one mutation is identified by PTA. In some cases, the reference sequence is a specificity-determining sequence that promotes the introduction of mutations into the target sequence of the nucleic acid. In some cases, after the gene editing process, at least one mutation is identified by PTA, and this mutation is located in the target sequence. In some cases, the off-target mutation rate is analyzed by identifying at least one mutation that is not in the target sequence. Some regions of the nucleic acid may be predicted to suffer from off-target mutations based on sequence homology to the target sequence, but regions with lower homology may also have off-target mutations. In some cases, the PTA method identifies mutations in off-target regions of the sequence that contain at least 3, 4, 5, 6, 7, or 8 base mismatches with the target sequence or its reverse complement. In some cases, a single cell is analyzed by PTA. In some cases, a population of cells is analyzed by PTA.

[0030] Many current methods for mutation analysis obtain sequencing data on bulk cell populations. However, such approaches provide limited information about the actual frequency of mutations in a population. Single-cell analysis using PTA, in some cases, provides much higher resolution of off-target insertion rates, strand breaks (leading to mutations), and translocations as the number of cells (i.e., single cells) is also known. PTA has a known mutation detection rate in a known number of single cells, allowing this method to accurately determine the frequency and combination of mutations per cell in a cell population. In some cases, at least 10, 100, 1000, 10,000, 100,000, or more than 100,000 single cells are analyzed by PTA to establish a mutation rate. In some cases, no more than 10, 100, 1000, 10,000, 100,000, or 100,000 single cells are analyzed by PTA to establish a mutation rate. In some cases, 10-1000, 50-5000, 100-100,000, 100-100,000, 100-1,000,000, or 100-10,000 single cells are analyzed by PTA to establish a variability rate. In some cases, mutations identified by analysis of one or more single cells are not identified or detected from bulk sequencing of a population of cells.

[0031] CRISPR can be used to introduce mutations into one or more cells, such as mammalian cells, which are then analyzed by PTA. In some cases, the specificity-determining sequence is present in the CRISPR RNA (crRNA) or single guide RNA (sgRNA). In some cases, the mammalian cells are human cells. In some cases, the cells are derived from liver, skin, kidney, blood, or lung. In some cases, the cells are primary cells. In some cases, the cells are stem cells. Previously reported methods for identifying off-target mutations generated from CRISPR include pulldown of sequences that bind to catalytically active Cas9, but this can lead to false positives because mutations are not introduced at all Cas9 binding sites. In some cases, the PTA method identifies at least one mutation present in the region of the sequence that binds catalytically active Cas9. In some cases, the PTA method is less likely to produce false positives for at least one mutation present in the region of the sequence that binds catalytically active Cas9.

[0032] Described herein are methods for identifying mutations in animal, plant, or microbial cells that have undergone genome editing (e.g., CRISPR, TALEN, ZFN, recombinase, meganuclease, viral integration, or other techniques), the method comprising amplifying a genome or a fragment thereof in the presence of at least one terminator nucleotide. In some cases, amplification using a terminator occurs in solution. In some cases, either at least one primer or at least one genome fragment is attached to a surface. In some cases, at least one primer is attached to a first solid support, at least one genome fragment is attached to a second solid support, and the first and second solid supports are not connected. In some cases, at least one primer is attached to a first solid support, at least one genome fragment is attached to a second solid support, and the first and second solid supports are not the same solid support. In some cases, the method includes amplifying a genome or a fragment thereof in the presence of at least one terminator nucleotide, and the number of amplification cycles is less than 12, 10, 9, 8, 7, 6, 5, 4, or 3. In some cases, the average length of the amplification products is 100-1000, 200-500, 200-700, 300-700, 400-1000, or 500-1200 bases. In some cases, the method includes amplifying a genome or a fragment thereof in the presence of at least one terminator nucleotide, and the number of amplification cycles is 6 or less. In some cases, at least one terminator nucleotide comprises a detectable label or tag. In some cases, the amplification comprises two, three, or four terminator nucleotides. In some cases, at least two of the terminator nucleotides comprise different bases. In some cases, at least three of the terminator nucleotides comprise different bases. In some cases, each of the four terminator nucleotides comprises a different base.

[0033] Described herein is a method for determining the safety of gene therapy.In some cases, cell function is modified by gene editing or other expression methods.In some cases, viral delivery systems for changing cell function are configured so that they are not integrated into the cell's genome.In some cases, PTA method is used to identify unexpected or undesirable changes to cell genome.In some cases, PTA is used to identify the mutations in somatic or germline DNA that result from gene therapy.

[0034] Clonal analysis of tumor cells In some cases, the cells analyzed using the methods described herein include tumor cells. For example, circulating tumor cells can be isolated from fluids collected from patients, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, or aqueous humor. The cells are then subjected to the methods described herein (e.g., PTA) and sequencing to determine the mutation load and combination of mutations in each cell. These data can be used in some cases as a tool for diagnosing specific diseases or predicting treatment response. Similarly, in some cases, cells of unknown malignant potential are isolated from bodily fluids collected from patients, such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural effusion, pericardial effusion, ascites, aqueous humor, blastocyst fluid, or collection medium surrounding cells in culture. In some cases, samples are obtained from collection medium surrounding embryonic cells. After utilizing the methods described herein and sequencing, such methods can be further used to determine the mutation load and combination of mutations in each cell. These data are sometimes used to diagnose specific diseases or as a tool to predict the progression of premalignant states to overt malignancies. In some cases, cells can be isolated from primary tumor samples. Cells are then subjected to PCR and sequencing to determine the mutational burden and combination of mutations in each cell. In some cases, these data can be used to diagnose specific diseases or as a tool to predict the probability that a patient's malignancy will be resistant to available anticancer drugs. By exposing samples to different chemotherapeutic agents, it was found that major and minor clones have differential sensitivity to specific drugs, which does not necessarily correlate with the presence of known "driver mutations," suggesting that the combination of mutations within a clonal population determines its sensitivity to a particular chemotherapeutic drug. Without being bound by theory, these findings suggest that it may be easier to eradicate malignancies if premalignant lesions are detected that evolve into clones with an increasing number of genomic modifications that may be more resistant to treatment.See Ma et al., 2018, "Pan-cancer genome and transcriptome analyses of 1,699 pediatric leukemias and solid tumors." Single-cell genomics protocols are sometimes used to detect combinations of somatic genetic variants in single cancer cells or clonotypes within a mixture of normal and malignant cells isolated from patient samples. This technology is sometimes further utilized to identify clonotypes that undergo positive selection after drug exposure both in vitro and / or in patients. As shown in Figure 6A, by comparing surviving clones exposed to chemotherapy with clones identified at diagnosis, a catalog of cancer clonotypes documenting their resistance to specific drugs can be created. PTA methods sometimes detect the sensitivity of specific clones, as well as combinations of them, to existing or novel drugs within samples composed of multiple clonotypes, where the method can detect the sensitivity of specific clones to drugs. This approach sometimes indicates the effectiveness of a drug against a specific clone that may not be detected using current drug sensitivity measures that consider the sensitivity of all cancer clones together in a single measurement. When the PTA described herein is applied to patient samples collected at the time of diagnosis to detect cancer clonal types in a given patient's cancer, it searches for those clones using a catalog of drug sensitivities, thereby providing oncologists with information on which drugs or drug combinations will not work and which drugs or drug combinations are likely to be most effective against that patient's cancer. PTA can be used to analyze samples containing groups of cells. In some cases, the sample contains neurons or glial cells. In some cases, the sample contains nuclei.

[0035] Clinical and environmental mutagenesis Described herein are methods for measuring the mutagenicity of environmental factors. For example, cells (single or population) are exposed to a potential environmental condition. For example, cells derived from an organ (liver, pancreas, lung, colon, thyroid, or other organ), tissue (skin, or other tissue), blood, or other biological source may be used in this method. In some cases, the environmental condition includes heat, light (e.g., ultraviolet light), radiation, chemicals, or any combination thereof. In some cases, after exposure to the environmental condition for minutes, hours, days, or longer, single cells are isolated and subjected to the PTA method. In some cases, samples are tagged with molecular barcodes and unique molecular identifiers. The samples are sequenced and then analyzed to identify mutations resulting from exposure to the environmental condition. In some cases, such mutations are compared to a control environmental condition, such as a known non-mutagenic agent, vehicle / solvent, or the absence of the environmental condition. Such analysis, in some cases, provides not only the total number of mutations caused by the environmental condition, but also the location and nature of such mutations. In some cases, patterns can be identified from data and used to diagnose a disease or condition. In some cases, patterns can be used to predict future disease states or conditions. In some cases, the methods described herein measure mutation load, location, and patterns in cells after exposure to an environmental agent, such as a potential mutagen or teratogen. This approach, in some cases, is used to evaluate the safety of a given agent, including its potential to induce mutations that may contribute to disease development. For example, this method can be used to predict the carcinogenic or teratogenic potential of an agent for a particular cell type after exposure to a particular concentration of the agent. In some cases, the agent is a pharmaceutical or drug. In some cases, the agent is a food product. In some cases, the agent is a genetically modified food product. In some cases, the agent is a pesticide or other agricultural chemical. In some cases, the location and rate of mutations are used to predict the age of an organism.Such methods are sometimes performed on samples that are hundreds, thousands, or tens of thousands of years old. The mutation patterns are sometimes compared to other data methods, such as radiocarbon dating, to generate a standard curve. In some cases, the age of a person is determined by comparing the number and pattern of mutations from samples.

[0036] Described herein are methods for determining mutations in cells used for cell therapy, such as, but not limited to, transplantation of induced pluripotent stem cells, transplantation of unmanipulated hematopoietic or other cells, or transplantation of hematopoietic or other cells that have undergone genome editing. The cells can then undergo PTA and sequencing to determine the mutation load and combination of mutations in each cell. The mutation rate and location of mutations per cell in a cell therapy product can be used to assess the safety and potential efficacy of the product, including measuring neoantigen load.

[0037] Microbial samples Described herein are methods for analyzing microbial samples. In another embodiment, microbial cells (e.g., bacteria, fungi, protozoa) can be isolated from plants or animals (e.g., from microbiota samples (e.g., GI microbiota, skin microbiota, etc.) or from bodily fluids such as blood, bone marrow, urine, saliva, cerebrospinal fluid, pleural fluid, pericardial fluid, ascites, or aqueous humor). Furthermore, microbial cells may be isolated from indwelling medical devices, such as, but not limited to, intravenous catheters, urinary catheters, cerebrospinal shunts, artificial valves, artificial joints, or endotracheal tubes. Cells can then undergo PTA and sequencing to determine the identity of specific microorganisms and detect the presence of microbial genetic variants that predict response (or resistance) to specific antimicrobial agents. These data can be used for the diagnosis of specific infections and / or as a tool for predicting treatment response. In some cases, single microbial cells are analyzed for mutations. In one embodiment, PTA is used to identify microorganisms of high value for industrial applications, such as biofuel production or environmental remediation (oil spill cleanup, CO2 sequestration / removal). In some cases, microbial samples are obtained from extreme environments, such as deep-sea vents, oceans, mines, streams, lakes, meteorites, glaciers, and volcanoes. In some cases, microbial samples include strains of microorganisms that are "unculturable" in the laboratory under standard conditions. Sequencing microbial samples prepared using PTA, in some cases, includes obtaining sequencing reads for assembly into contigs. In some cases, no more than 0.1, 0.5, 1, 1.5, 2, 3, 5, 8, or 10 million reads are obtained. Analysis and identification of microbial samples, in some cases, includes comparing the assembled contigs to known microbial genome reference sequences. In some cases, the largest assembled contig is used for comparison with the reference sequence. In some cases, reads that map to one or more genes in human genomic DNA are filtered. In some cases, filtering is performed if both (forward and reverse) reads map to a human gene.In some cases, filtering is performed if at least one read (forward or reverse) maps to a human gene. In some cases, the human gene is GRCh38. In some cases, assembly-free identification methods involving PTA are used. In some cases, assembly-free methods such as Kraken are used. In some cases, assembly-free methods involve assigning reads to taxa based on k-mers using a reference database. .

[0038] fetal cells Cells for use with the PTA method can be fetal cells, such as embryonic cells. In some embodiments, PTA is used in conjunction with non-invasive preimplantation genetic testing (NIPGT). In further embodiments, cells can be isolated from blastomeres or blastocysts produced by in vitro fertilization. The cells can then undergo PTA (e.g., nucleic acids within the cells are amplified during PTA) and sequencing to determine the burden and combination of potential disease-predisposing genetic mutations in each cell. The mutation profile of the cells can then be used to extrapolate the genetic predisposition of the blastomere to a specific disease before implantation. In some cases, embryos in culture release nucleic acids that are used to assess the health of the embryo using low-pass genomic sequencing. In some cases, the embryos are frozen and thawed. In some cases, nucleic acids are obtained from blastocyst culture-conditioned medium (BCCM), blastocoelic fluid (BF), or a combination thereof. In some cases, PTA analysis of fetal cells is used to detect chromosomal abnormalities, such as fetal aneuploidy. In some cases, PTA is used to detect diseases such as Down syndrome or Patau syndrome. In some cases, frozen blastocysts are thawed and cultured for a period of time before obtaining nucleic acid for analysis (e.g., medium, BF, or cell biopsy). In some cases, the blastocysts are cultured within 4, 6, 8, 12, 16, 24, 36, 48, or 64 hours before obtaining nucleic acid for analysis.

[0039] mutation In some cases, the methods described herein (e.g., PTA) result in higher detection sensitivity and / or lower false positive rates for mutation detection. In some cases, the mutation is a difference between the analyzed sequence (e.g., using the methods described herein) and a reference sequence. The reference sequence may be obtained from other organisms, other individuals of the same or similar species, a population of organisms, or other regions of the same genome. In some cases, the mutation is identified on a plasmid or chromosome. In some cases, the mutation is an SNV (single nucleotide change), SNP (single nucleotide polymorphism), or CNV (copy number variation, or CNA / copy number abnormality). In some cases, the mutation is a base substitution, insertion, or deletion. In some cases, the mutation is a transition, transversion, nonsense mutation, silent mutation, synonymous or nonsynonymous mutation, nonpathogenic mutation, missense mutation, or frameshift mutation (deletion or insertion). In some cases, PTA yields higher detection sensitivity and / or lower false positive rates when compared to methods such as in silico prediction, ChIP-seq, GUIDE-seq, Circle-seq, HTGTS (high-throughput genome-wide translocation sequencing), IDLV (integration-deficient lentivirus), Digenome-seq, FISH (fluorescence in situ hybridization), or DISCOVER-seq.

[0040] Primary template-directed amplification Described herein are nucleic acid amplification methods such as "primary template-directed amplification (PTA)." For example, the PTA method described herein is represented schematically in Figures 1A-1H. In the PTA method, a polymerase (e.g., a strand-displacing polymerase) is used to preferentially generate amplicons from a primary template ("direct copy"). As a result, errors are propagated from daughter amplicons at a lower rate during subsequent amplification compared to MDA. The result is an easily performed method that can accurately and reproducibly amplify low DNA inputs, including the genomes of single cells, with high coverage breadth and uniformity, unlike existing WGA protocols. Furthermore, terminated amplification products can undergo directional ligation after terminator removal, allowing cell barcodes to be attached to amplification primers, resulting in pooling of products from all cells after undergoing parallel amplification reactions (Figure 1F). In some cases, terminator removal prior to amplification and / or adapter ligation is not required.

[0041] Described herein are methods for using nucleic acid polymerases with strand displacement activity for amplification. In some cases, such polymerases have strand displacement activity and a low error rate. In some cases, such polymerases have strand displacement activity and proofreading exonuclease activity, such as 3' to 5' proofreading activity. In some cases, nucleic acid polymerases are used in combination with other components, such as reversible or irreversible terminators or additional strand displacement factors. In some cases, the polymerase has strand displacement activity but does not have exonuclease proofreading activity. For example, in some instances, such polymerases include bacteriophage phi29 (Φ29) polymerase, which also has a very low error rate resulting from its 3' to 5' proofreading exonuclease activity (see, e.g., U.S. Patent Nos. 5,198,543 and 5,001,050). In some cases, non-limiting examples of strand-displacing nucleic acid polymerases include, for example, genetically engineered phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I (Jacobsen et al., Eur. J. BioChem. 45:623-627 (1974)), phage M2 DNA polymerase (Matsumoto et al., Gene 84:247 (1989)), phage phi PRD1 DNA polymerase (Jung et al., Proc. Natl. Acad. Sci. USA 84:8287 (1987); Zhu and Ito, Biochim. Biophys. Acta. 1219:267-276 (1994)), Bst DNA polymerase (e.g., Bst large fragment DNA polymerase (exo(-)Bst; Aliotta et al. al., Genet. Anal. (Netherlands) 12:185-195 (1996)), exo(-) Bca DNA polymerase (Walker and Linn, Clinical Chemistry 42:1604-1608 (1996)), Bsu DNA polymerase, Vent R Vent containing (exo-)DNA polymeraseRExamples of strand-displacing nucleic acid polymerases include DNA polymerases (Kong et al., J. Biol. Chem. 268:1965-1975 (1993)), Deep Vent DNA polymerases, including Deep Vent (exo-) DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase (Chatterjee et al., Gene 97:13-19 (1991)), Sequenase (USBiochemicals), T7 DNA polymerase, T7-Sequenase, T7 gp5 DNA polymerase, PRDI DNA polymerase, and T4 DNA polymerase (Kaboord and Benkovic, Curr. Biol. 5:149-157 (1995)). Additional strand-displacing nucleic acid polymerases are also compatible with the methods described herein. The ability of a given polymerase to perform strand displacement replication can be determined, for example, by using the polymerase in a strand displacement replication assay (e.g., as disclosed in U.S. Pat. No. 6,977,148). Such assays are sometimes performed at temperatures appropriate for the optimal activity of the enzyme used, such as 32°C for Phi29 DNA polymerase, 46°C to 64°C for exo(-)Bst DNA polymerase, or 60°C to 70°C for enzymes from hyperthermophilic organisms. Another useful assay for selecting polymerases is the primer blocking assay described in Kong et al., J. Biol. Chem. 268:1965-1975 (1993). This assay consists of a primer extension assay using an M13 ssDNA template in the presence or absence of an oligonucleotide that hybridizes upstream of the extension primer and blocks its progression. Other enzymes capable of displacing the blocking primer in this assay may, in some cases, be useful for the disclosed method. In some cases, the polymerase incorporates dNTPs and terminators in approximately equal proportions.In some cases, the ratio of dNTP and terminator incorporation rates for polymerases described herein is about 1:1, about 1.5:1, about 2:1, about 3:1, about 4:1, about 5:1, about 10:1, about 20:1, about 50:1, about 100:1, about 200:1, about 500:1, or about 1000:1. In some cases, the ratio of dNTP and terminator incorporation rates for polymerases described herein is 1:1 to 1000:1, 2:1 to 500:1, 5:1 to 100:1, 10:1 to 1000:1, 100:1 to 1000:1, 500:1 to 2000:1, 50:1 to 1500:1, or 25:1 to 1000:1.

[0042] Described herein is an amplification method that can promote strand displacement through the use of strand displacement factors, such as helicases.In some cases, such factors are used in combination with additional amplification components, such as polymerases, terminators, or other components.In some cases, strand displacement factors are used with polymerases that do not have strand displacement activity.In some cases, strand displacement factors are used with polymerases that have strand displacement activity.Without being bound by theory, strand displacement factors can increase the rate at which smaller double-stranded amplicons are reprimed.In some cases, any DNA polymerase that can perform strand displacement replication in the presence of strand displacement factors is suitable for use in PTA methods, even if the DNA polymerase does not perform strand displacement replication in the absence of such factors.Strand displacement factors useful in strand displacement replication include, in some cases, the BMRF1 polymerase accessory subunit (Tsurumi et al., J. Virology 67(12):7648-7653 (1993)), adenovirus DNA binding protein (Zijderveld and van der Vliet, J. Virology 68(2):1158-1164 (1994)), herpes simplex virus protein ICP8 (Boehmer and Lehman, J. Virology 67(2):711-715 (1993); Skaliter and Lehman, Proc. Natl. Acad. Sci. USA 91(22):10665-10669 (1994)); single-stranded DNA binding protein (SSB; Rigler and Romano, J. Biol. Chem. 270:8910-8919 (1995)); phage T4 gene 32 protein (Villemain and Giedroc, Biochemistry 35:14395-14404 (1996)); T7 helicase-primase; T7 gp2.5 SSB protein; Tte-UvrD (from Thermoanaerobacter tengcongensis), calf thymus helicase (Siegel et al., J. Biol. Chem. 267:13629-13635 (1992)); bacterial SSB (e.g., E. coli SSB), replication protein A (RPA) in eukaryotes, human mitochondrial SSB (mtSSB), and recombinases (e.g., recombinase A (RecA) family proteins, T4 UvsX, T4 Examples of suitable factors that promote strand displacement and priming include, but are not limited to, UvsY, Sak4 from phage HK620, Rad51, Dmc1, or Radd. Combinations of factors that promote strand displacement and priming are also consistent with the methods described herein. For example, a helicase is used in conjunction with a polymerase.In some cases, the PTA method involves the use of a single-stranded DNA binding protein (SSB, T4 gp32, or other single-stranded DNA binding protein), a helicase, and a polymerase (e.g., Sau DNA polymerase, Bsu polymerase, Bst2.0, GspM, GspM2.0, GspSSD, or other suitable polymerase). In some cases, a reverse transcriptase is used in combination with a strand displacement factor described herein. In some cases, amplification is performed using a polymerase and a nicking enzyme (e.g., "NEAR") as described in U.S. Patent No. 9,617,586. In some cases, the nicking enzyme is Nt.BspQI, Nb.BbvCi, Nb.BsmI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nt.BbvCI, Nt.BstNBI, Nt.CviPII, Nb.Bpu10I, or Nt.Bpu10I.

[0043] Described herein is an amplification method that includes the use of terminator nucleotides, polymerase, and additional factors or conditions. For example, in some cases, such factors are used to fragment nucleic acid templates or amplicons during amplification. In some cases, such factors include endonucleases. In some cases, factors include transposases. In some cases, mechanical shearing is used to fragment nucleic acids during amplification. In some cases, nucleotides are added during amplification, and these can be fragmented by adding additional proteins or conditions. For example, uracil is incorporated into an amplicon, and treatment with uracil D-glycosylase fragments the nucleic acid at positions containing uracil. Additional systems for selective nucleic acid fragmentation are also utilized in some cases, such as engineered DNA glycosylases that cleave modified cytosine-pyrene base pairs (Kwon, et al. Chem Biol. 2003, 10(4), 351).

[0044] Described herein are amplification methods that involve the use of terminator nucleotides, which terminate nucleic acid replication and thus reduce the size of the amplified product. Such terminators are sometimes used in combination with polymerases, strand displacement factors, or other amplification components described herein. In some cases, terminator nucleotides reduce or decrease the efficiency of nucleic acid replication. In some cases, such terminators reduce the extension rate by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. In some cases, such terminators reduce the extension rate by 50% to 90%, 60% to 80%, 65% to 90%, 70% to 85%, 60% to 90%, 70% to 99%, 80% to 99%, or 50% to 80%. In some cases, the terminator reduces the average amplicon product length by at least 99.9%, 99%, 98%, 95%, 90%, 85%, 80%, 75%, 70%, or at least 65%. In some cases, the terminator reduces the average amplicon length by 50% to 90%, 60% to 80%, 65% to 90%, 70% to 85%, 60% to 90%, 70% to 99%, 80% to 99%, or 50% to 80%. In some cases, amplicons containing terminator nucleotides form loops or hairpins, which reduces the ability of a polymerase to use such amplicons as templates. The use of terminators, in some cases, slows the rate of amplification at the initial amplification site through the incorporation of terminator nucleotides (e.g., dideoxynucleotides modified to be exonuclease resistant to stop DNA extension), resulting in smaller amplification products.By generating smaller amplification products than currently used methods (e.g., an average length of 50-2,000 nucleotides for the PTA method compared to an average product length of >10,000 nucleotides for the MDA method), PTA amplification products can, in some cases, undergo direct ligation of adapters without the need for fragmentation, allowing for the efficient incorporation of cell barcodes and unique molecular identifiers (UMIs) (see Figures 1H, 2B-3E, 9, 10A, and 10B).

[0045] Terminator nucleotides are present at various concentrations depending on factors such as polymerase, template, or other factors. For example, the amount of terminator nucleotides is sometimes expressed as the ratio of non-terminator nucleotides to terminator nucleotides in the methods described herein. Such concentrations can sometimes control the length of amplicons. In some cases, the ratio of terminator to non-terminator nucleotides is changed depending on the amount of template present or the size of the template. In some cases, the ratio of terminator to non-terminator nucleotides becomes smaller as the sample size becomes smaller (e.g., in the femtogram and picogram ranges). In some cases, the ratio of non-terminator nucleotides to terminator nucleotides is about 2:1, 5:1, 7:1, 10:1, 20:1, 50:1, 100:1, 200:1, 500:1, 1000:1, 2000:1, or 5000:1. In some cases, the ratio of non-terminator to terminator nucleotides is 2:1 to 10:1, 5:1 to 20:1, 10:1 to 100:1, 20:1 to 200:1, 50:1 to 1000:1, 50:1 to 500:1, 75:1 to 150:1, or 100:1 to 500:1. In some cases, at least one of the nucleotides present during amplification using the methods described herein is a terminator nucleotide. Each terminator need not be present at approximately the same concentration; in some cases, the ratio of each terminator present in the methods described herein is optimized for a particular set of reaction conditions, sample type, or polymerase. Without being bound by theory, each terminator may have different efficiencies for incorporation into the growing polynucleotide strand of an amplicon in response to pairing with the corresponding nucleotide on the template strand. For example, in some cases, cytosine-pairing terminators are present at a concentration that is about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration.In some cases, terminators paired with thymine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with guanine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with adenine are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, terminators paired with uracil are present at a concentration about 3%, 5%, 10%, 15%, 20%, 25%, or 50% higher than the average terminator concentration. In some cases, any nucleotide that can terminate nucleic acid elongation by nucleic acid polymerase is used as terminator nucleotide in the method described herein.In some cases, reversible terminator is used to terminate nucleic acid replication.In some cases, irreversible terminator is used to terminate nucleic acid replication.In some cases, non-limiting examples of terminator include reversible and irreversible nucleic acids and nucleic acid analogs, such as 3'-blocked reversible terminator that comprises nucleotides, 3'-unblocked reversible terminator that comprises nucleotides, terminator that comprises 2'-modified deoxynucleotides, terminator that comprises modified nitrogen base of deoxynucleotides, or any combination thereof.In one embodiment, terminator nucleotide is dideoxynucleotide. Other nucleotide modifications that terminate nucleic acid replication and are suitable for practicing the present invention include, but are not limited to, any modification of the r group of the 3' carbon of deoxyribose, such as reverse dideoxynucleotides, 3' biotinylated nucleotides, 3' amino nucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3' carbon spacer nucleotides, including 3' C3 spacer nucleotides, 3' C18 nucleotides, 3' hexanediol spacer nucleotides, acyclonucleotides, and combinations thereof.In some cases, the terminator is a polynucleotide containing 1, 2, 3, 4, or more bases in length. In some cases, the terminator does not contain a detectable moiety or tag (e.g., a mass tag, a fluorescent tag, a dye, a radioactive atom, or other detectable moiety). In some cases, the terminator does not contain a chemical moiety that allows for the attachment of a detectable moiety or tag (e.g., a "click" azide / alkyne, a conjugate addition partner, or other chemical handle for tag attachment). In some cases, all terminator nucleotides contain the same modification that reduces amplification in a region of the nucleotide (e.g., the sugar moiety, the base moiety, or the phosphate moiety). In some cases, at least one terminator has a different modification that reduces amplification. In some cases, all terminators have substantially similar fluorescence excitation or emission wavelengths. In some cases, terminators without modifications to the phosphate group are used with polymerases that do not have exonuclease proofreading activity. When used with a polymerase that has 3' to 5' proofreading exonuclease activity (e.g., Phi29) that can remove terminator nucleotides, terminators are, in some cases, further modified to render them exonuclease-resistant. For example, dideoxynucleotides are modified with alpha-thio groups that create phosphorothioate linkages that render these nucleotides resistant to the 3' to 5' proofreading exonuclease activity of nucleic acid polymerases. Such modifications, in some cases, reduce the exonuclease proofreading activity of the polymerase by at least 99.5%, 99%, 98%, 95%, 90%, or at least 85%.Non-limiting examples of other terminator nucleotide modifications that confer resistance to 3'->5' exonuclease activity include, in some cases, nucleotides with modifications to the alpha group, such as alpha-thiodideoxynucleotides that create phosphorothioate linkages, C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2'-fluoro bases, 3' phosphorylation, 2'-O-methyl modifications (or other 2'-O-alkyl modifications), propyne-modified bases (e.g., deoxycytosine, deoxyuridine), L-DNA nucleotides, L-RNA nucleotides, nucleotides with inverted linkages (e.g., 5'-5' or 3'-3'), 5'-inverted bases (e.g., 5'-inverted 2', 3'-dideoxydT), methylphosphonate backbones, and trans nucleic acids. In some cases, modified nucleotides include base-modified nucleic acids containing a free 3'OH group (e.g., bases containing modifications with bulky chemical groups such as 2-nitrobenzyl alkylated HOMedU triphosphate, solid supports, or other bulky moieties). In some cases, polymerases with strand displacement activity but no 3' to 5' exonuclease proofreading activity are used with terminator nucleotides, with or without modifications to make them exonuclease resistant. Such nucleic acid polymerases include, but are not limited to, Bst DNA polymerase, Bsu DNA polymerase, Deep Vent (exo-) DNA polymerase, Klenow fragment (exo-) DNA polymerase, Therminator DNA polymerase, and Vent. R (exo-) is included.

[0046] Primers and amplicon libraries Described herein are amplicon libraries resulting from the amplification of at least one target nucleic acid molecule. Such libraries are, in some cases, generated using methods described herein, such as those using terminators. Such methods include the use of strand-displacing polymerases or agents, terminator nucleotides (reversible or irreversible), or other features and embodiments described herein. In some cases, the amplicon library generated by the use of terminators described herein is further amplified in a subsequent amplification reaction (e.g., PCR). In some cases, the subsequent amplification reaction does not include a terminator. In some cases, the amplicon library includes polynucleotides, and at least 50%, 60%, 70%, 80%, 90%, 95%, or at least 98% of the polynucleotides include at least one terminator nucleotide. In some cases, the amplicon library includes the target nucleic acid molecule from which the amplicon library was derived. An amplicon library comprises a plurality of polynucleotides, at least some of which are direct copies (e.g., directly replicated from a target nucleic acid molecule, such as genomic DNA, RNA, or other target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 95% or more of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 5% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 10% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 15% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 20% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least 50% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule.In some cases, 3% to 5%, 3% to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50%, or 15% to 75% of the amplicon polynucleotides are direct copies of at least one target nucleic acid molecule. In some cases, at least some of the polynucleotides are direct copies of the target nucleic acid molecule or descendants of daughters (initial copies of the target nucleic acid). For example, at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more than 95% of the amplicon polynucleotides are direct copies or descendants of daughters of at least one target nucleic acid molecule. In some cases, at least 5% of the amplicon polynucleotides are direct copies or descendants of daughters of at least one target nucleic acid molecule. In some cases, at least 10% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, at least 20% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, at least 30% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, 3% to 5%, 3% to 10%, 5% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 5% to 30%, 10% to 50%, or 15% to 75% of the amplicon polynucleotides are direct copies or daughter progeny of at least one target nucleic acid molecule. In some cases, the direct copies of the target nucleic acid are 50-2500, 75-2000, 50-2000, 25-1000, 50-1000, 500-2000, or 50-2000 bases in length. In some cases, the lengths of the daughter progeny are 1000-5000, 2000-5000, 1000-10,000, 2000-5000, 1500-5000, 3000-7000, or 2000-7000 bases in length. In some cases, the average length of the PTA amplification products is 25-3000 nucleotides in length, 50-2500, 75-2000, 50-2000, 25-1000, 50-1000, 500-2000, or 50-2000 bases in length.In some cases, amplicons generated from PTA are 5,000, 4,000, 3,000, 2,000, 1,700, 1,500, 1,200, 1,000, 700, 500 bases or less, or 300 bases or less. In some cases, amplicons generated from PTA are 1,000-5,000, 1,000-3,000, 200-2,000, 200-4,000, 500-2,000, 750-2,500, or 1,000-2,000 bases in length. In some cases, amplicon libraries generated using the methods described herein include at least 1,000, 2,000, 5,000, 10,000, 100,000, 200,000, 500,000, or more than 500,000 amplicons comprising unique sequences. In some cases, the library comprises at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 2000, 2500, 3000, or at least 3500 amplicons. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of less than 1000 bases are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 2000 bases or less are direct copies of at least one target nucleic acid molecule. In some cases, at least 5%, 10%, 15%, 20%, 25%, 30%, or more than 30% of the amplicon polynucleotides having a length of 3000 to 5000 bases are direct copies of at least one target nucleic acid molecule. In some cases, the ratio of direct copy amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or more than 10,000,000:1.In some cases, the ratio of direct copy amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, where the direct copy amplicons are 700-1200 bases in length or less. In some cases, the ratio of direct copy amplicons and daughter amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1. In some cases, the ratio of direct copy amplicons and daughter amplicons to target nucleic acid molecules is at least 10:1, 100:1, 1000:1, 10,000:1, 100,000:1, 1,000,000:1, 10,000,000:1, or greater than 10,000,000:1, wherein the direct copy amplicons are 700-1200 bases in length and the daughter amplicons are 2500-6000 bases in length. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule. In some cases, the library contains about 50 to 10,000, about 50 to 5,000, about 50 to 2500, about 50 to 1000, about 150 to 2000, about 250 to 3000, about 50 to 2000, about 500 to 2000, or about 500 to 1500 amplicons that are direct copies of the target nucleic acid molecule or daughter amplicons. The number of direct copies can, in some cases, be controlled by the number of PCR amplification cycles. In some cases, up to 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or 3 PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, about 30, 25, 20, 15, 13, 11, 10, 9, 8, 7, 6, 5, 4, or about 3 PCR cycles are used to generate copies of the target nucleic acid molecule.In some cases, 3, 4, 5, 6, 7, or 8 PCR cycles are used to generate copies of the target nucleic acid molecule. In some cases, 2-4, 2-5, 2-7, 2-8, 2-10, 2-15, 3-5, 3-10, 3-15, 4-10, 4-15, 5-10, or 5-15 PCR cycles are used to generate copies of the target nucleic acid molecule. The amplicon libraries generated using the methods described herein are sometimes subjected to additional steps, such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed before the sequencing step.

[0047] In some cases, polynucleotide amplicon libraries generated from the PTA methods and compositions (terminators, polymerases, etc.) described herein exhibit increased uniformity. Uniformity may be described using a Lorenz curve (e.g., FIG. 5C) or other such methods. Such an increase may result in fewer sequencing reads than are required for desired coverage of a target nucleic acid molecule (e.g., genomic DNA, RNA, or other target nucleic acid molecule). For example, 50% or less of the cumulative percentage of polynucleotides contain at least 80% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 60% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 70% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, 50% or less of the cumulative percentage of polynucleotides contain at least 90% of the sequence of the cumulative percentage of sequences of the target nucleic acid molecule. In some cases, uniformity is described using a Gini index (where an index of 0 represents perfect equality of the libraries and an index of 1 represents perfect inequality). In some cases, the amplicon libraries described herein have a Gini index of 0.55, 0.50, 0.45, 0.40, or 0.30 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less. In some cases, the amplicon libraries described herein have a Gini index of 0.40 or less. In some cases, such a uniformity metric depends on the number of reads obtained. For example, 100, 200, 300, 400, or 500 million reads or less are obtained. In some cases, the read length is about 50, 75, 100, 125, 150, 175, 200, 225, or about 250 bases long. In some cases, the uniformity metric depends on the coverage depth of the target nucleic acid. For example, the average depth of coverage is about 10x, 15x, 20x, 25x, or about 30x.In some cases, the average coverage depth is 10-30x, 20-50x, 5-40x, 20-60x, 5-20x, or 10-20x. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and have yielded approximately 300 million reads. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and an average depth of sequencing coverage of about 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.50 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.45 or less and an average depth of sequencing coverage of at least 15-fold. In some cases, the amplicon libraries described herein have a Gini index of 0.55 or less, wherein the average depth of sequencing coverage is 15-fold or less.In some cases, the Gini index of the amplicon library described herein is 0.50 or less, and the average depth of sequencing coverage is 15 times or less. In some cases, the Gini index of the amplicon library described herein is 0.45 or less, and the average depth of sequencing coverage is 15 times or less. In some cases, the uniform amplicon library generated using the methods described herein is subjected to additional steps, such as adapter ligation and further PCR amplification. In some cases, such additional steps are performed before the sequencing step.

[0048] Primers include nucleic acids used to prime the amplification reactions described herein. Such primers include, but are not limited to, random deoxynucleotides of any length, with or without modifications to make them exonuclease resistant, random ribonucleotides of any length, with or without modifications to make them exonuclease resistant, modified nucleic acids such as locked nucleic acids, or DNA or RNA primers targeting specific genomic regions and reactions primed with enzymes such as primase. For whole genome PTA, it is preferred to use a set of primers with random or partially random nucleotide sequences. In nucleic acid samples of significant complexity, the specific nucleic acid sequences present in the sample do not need to be known, and primers do not need to be designed to be complementary to specific sequences. Rather, the complexity of nucleic acid samples results in a large number of different hybridization target sequences in the sample, which are complementary to various primers with random or partially random sequences. The complementary portions of primers for use in PTA may in some cases be completely randomized, contain only randomized portions, or be selectively randomized. In some cases, the number of random base positions in the complementary portion of the primer is, for example, 20% to 100% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is 10% to 90%, 15% to 95%, 20% to 100%, 30% to 100%, 50% to 100%, 75% to 100%, or 90% to 95% of the total number of nucleotides in the complementary portion of the primer. In some cases, the number of random base positions in the complementary portion of the primer is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the total number of nucleotides in the complementary portion of the primer.A set of primers having random or partially random sequences is synthesized using standard techniques, in some cases by allowing the addition of any nucleotide at each position to be randomized. In some cases, the set of primers is composed of primers of similar length and / or hybridization characteristics. In some cases, the term "random primer" refers to a primer that can exhibit 4-fold degeneracy at each position. In some cases, the term "random primer" refers to a primer that can exhibit 3-fold degeneracy at each position. The random primers used in the methods herein, in some cases, contain random sequences that are 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more bases in length. In some cases, the primers contain random sequences that are 3-20, 5-15, 5-20, 6-12, or 4-10 bases in length. Primers can also contain non-extendable elements that limit subsequent amplification of the amplicon generated therefrom. For example, a primer having a non-extendable element may, in some cases, include a terminator. In some cases, the primer may include a terminator nucleotide, such as 1, 2, 3, 4, 5, 10, or more than 10 terminator nucleotides. Primers need not be limited to components added externally to an amplification reaction. In some cases, primers are generated in situ by the addition of nucleotides and proteins that facilitate priming. For example, a primase-like enzyme in combination with nucleotides may, in some cases, be used to generate random primers for the methods described herein. The primase-like enzyme may, in some cases, be a member of the DnaG or AEP enzyme superfamily. In some cases, the primase-like enzyme may be TthPrimPol. In some cases, the primase-like enzyme may be T7gp4 helicase-primase. Such primases may, in some cases, be used in conjunction with a polymerase or strand displacement factor described herein.In some cases, primases initiate priming using deoxyribonucleotides. In some cases, primases initiate priming using ribonucleotides.

[0049] Following PTA amplification, a specific subset of amplicons can be selected. Such selection, in some cases, relies on size, affinity, activity, hybridization to probes, or other selection factors known in the art. In some cases, selection is performed before or after additional steps described herein, such as adapter ligation and / or library amplification. In some cases, selection is performed based on amplicon size (length). In some cases, smaller amplicons that are less likely to have undergone exponential amplification are selected, which further converts amplification from an exponential to a sublinear amplification process while enriching for products derived from the primary template (Figure 1A). In some cases, amplicons with lengths of 50-2000, 25-5000, 40-3000, 50-1000, 200-1000, 300-1000, 400-1000, 400-600, 600-2000, or 800-1000 bases are selected. In some cases, size selection is performed, for example, by using a protocol that utilizes solid-phase reversible immobilization (SPRI) on carboxylated paramagnetic beads to enrich for nucleic acid fragments of a specific size, or other protocols known to those skilled in the art. Optionally, or in combination, selection is performed during sequencing library preparation through preferential ligation and amplification of small fragments during PCR, and through the preferential formation of clusters from smaller sequencing library fragments during sequencing (e.g., sequencing by synthesis, nanopore sequencing, or other sequencing methods). Other strategies for selecting smaller fragments are also consistent with the methods described herein, including, but not limited to, isolating nucleic acid fragments of a specific size after gel electrophoresis, using silica columns that bind to nucleic acid fragments of a specific size, and other PCR strategies that more strongly enrich for smaller fragments. Any number of library preparation protocols can be used with the PTA method described herein.In some cases, the amplicons generated by PTA are ligated to adapters (optionally with removal of terminator nucleotides). In some cases, the amplicons generated by PTA contain regions of homology generated from transposase-based fragmentation that are used as priming sites. In some cases, libraries are prepared by mechanically or enzymatically fragmenting nucleic acids. In some cases, libraries are prepared using transposome-mediated tagging. In some cases, libraries are created by ligation of adapters, such as Y-adapters, universal adapters, or circular adapters.

[0050] The non-complementary portions of primers used in PTA can contain sequences that can be used to further manipulate and / or analyze the amplified sequences. An example of such a sequence is a "detection tag." Detection tags have sequences complementary to detection probes and are detected using their cognate detection probes. A primer may have one, two, three, four, or more than four detection tags. There is no fundamental limit to the number of detection tags that can be present on a primer, except for the size of the primer. In some cases, a primer has a single detection tag. In some cases, a primer has two detection tags. When multiple detection tags are present, they may have the same sequence or they may have different sequences, with each different sequence being complementary to a different detection probe. In some cases, multiple detection tags have the same sequence. In some cases, multiple detection tags have different sequences.

[0051] Another example of a sequence that can be included in the non-complementary portion of a primer is an "address tag," which can encode other details of the amplicon, such as its location within a tissue section. In some cases, the cellular barcode includes an address tag. The address tag has a sequence complementary to an address probe. The address tag is incorporated into the end of the amplified strand. If present, there can be one or more address tags in a primer. There is no fundamental limit to the number of address tags that can be present in a primer, except for the size of the primer. If multiple address tags are present, they can have the same sequence or different sequences, with each different sequence complementary to a different address probe. The address tag portion can be any length that supports specific and stable hybridization between the address tag and the address probe. In some cases, nucleic acids from more than one source can incorporate a variable tag sequence. This tag sequence can be up to 100 nucleotides long, preferably 1 to 10 nucleotides long, and most preferably 4, 5, or 6 nucleotides long, and can contain any combination of nucleotides. In some cases, the tag sequence is 1-20, 2-15, 3-13, 4-12, 5-12, or 1-10 nucleotides in length. For example, six base pairs are selected to form the tag, and four different nucleotide permutations are used, which then creates a total of 4096 nucleic acid anchors (e.g., hairpins), each with a unique six base pairs.

[0052] The primers described herein can be in solution or immobilized on a solid support. In some cases, primers bearing sample barcodes and / or UMI sequences can be immobilized on a solid support. The solid support can be, for example, one or more beads. In some cases, to identify individual cells, individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences. In some cases, lysates from individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. In some cases, to identify nucleic acids extracted from individual cells, nucleic acids extracted from individual cells are contacted with one or more beads bearing a unique set of sample barcodes and / or UMI sequences. The beads can be manipulated by any suitable method known in the art, for example, using a droplet actuator described herein. The beads can be of any suitable size, including, for example, microbeads, microparticles, nanobeads, and nanoparticles. In some embodiments, the beads are magnetically responsive, while in other embodiments, the beads are not significantly magnetically responsive.Non-limiting examples of suitable beads include flow cytometry microbeads, polystyrene microparticles and nanoparticles, functionalized polystyrene microparticles and nanoparticles, coated polystyrene microparticles and nanoparticles, silica microbeads, fluorescent microspheres and nanospheres, functionalized fluorescent microspheres and nanospheres, coated fluorescent microspheres and nanospheres, colored microparticles and nanoparticles, magnetic microparticles and nanoparticles, superparamagnetic microparticles and nanoparticles (e.g., DYNABEADS® available from Invitrogen Group, Carlsbad, CA), fluorescent microparticles and nanoparticles, coated magnetic microparticles and nanoparticles, ferromagnetic microparticles and nanoparticles, coated ferromagnetic microparticles and nanoparticles, and those described in U.S. Patent Application Publication Nos. US20050260686, US20030132538, US20050118574, 20050277197, 20060159962.

[0053] The beads may be pre-bound with antibodies, proteins or antigens, DNA / RNA probes, or any other molecules with affinity for the desired target. In some embodiments, primers with sample barcodes and / or UMI sequences may be in solution. In certain embodiments, multiple droplets may be presented, each droplet within the multiple droplets having a sample barcode unique to the droplet and a UMI unique to the molecule, such that the UMI is repeated multiple times within the collection of droplets. In some embodiments, individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cells. In some embodiments, lysates from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. In some embodiments, nucleic acids extracted from individual cells are contacted with droplets having a unique set of sample barcodes and / or UMI sequences to identify individual cell lysates. Various microfluidics platforms can be used for single-cell analysis. Cells are manipulated, in some cases, through fluid dynamics (droplet microfluidics, inertial microfluidics, vortexing, microvalves, microstructures (microwells, microtraps, etc.)), electrical methods (dielectrophoresis (DEP), electroosmosis), optical methods (optical tweezers, optically induced dielectrophoresis (ODEP), photothermal capillaries), acoustic methods, or magnetic methods. In some cases, the microfluidic platform comprises a microwell. In some cases, the microfluidic platform comprises a PDMS (polydimethylsiloxane)-based device.Non-limiting examples of single cell analysis platforms compatible with the methods described herein include the ddSEQ single cell isolator (Bio-Rad, Hercules, CA, USA, and Illumina, San Diego, CA, USA), Chromium (10x Genomics, Pleasanton, CA, USA), the Rhapsody single cell analysis system (BD, Franklin Lakes, NJ, USA), the Tapestri platform (MissionBio, San Francisco, CA, USA), Nadia Innovate (Dolomite Bio, Royston, UK), C1 and Polaris (Fluidigm, South San Francisco, CA, USA); the ICELL8 single cell system (Takara); MSND (Wafergen); the Puncher platform (Vycap), the CellRaft AIR system (CellMicrosystems), the DEPArray NxT and DEPArray systems (Menarini Silicon Biosystems), and AVISO. CellCelector (ALS), InDrop system (1CellBio), and TrapTx (Celldom).

[0054] PTA primers may include sequence-specific or random primers, address tags, cell barcodes, and / or unique molecular identifiers (UMIs) (see, e.g., Figures 10A (linear primers) and 10B (hairpin primers)). In some cases, primers include sequence-specific primers. In some cases, primers include random primers. In some cases, primers include cell barcodes. In some cases, primers include sample barcodes. In some cases, primers include unique molecular identifiers. In some cases, primers include more than one cell barcode. Such barcodes, in some cases, identify a unique sample source or a unique workflow. Such barcodes or UMIs are, in some cases, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or greater than 30 bases in length. In some cases, primers are at least 1000, 10,000, 50,000, 100,000, 250,000, 500,000, 10 6 , 10 7 , 10 8 , 10 9 , or at least 10 10Each primer contains 8, 16, 96, or 384 unique barcodes or UMIs. In some cases, primers contain at least 8, 16, 96, or 384 unique barcodes or UMIs. In some cases, standard adapters are ligated to the amplification products before sequencing, and after sequencing, reads are first assigned to specific cells based on their cellular barcodes. Suitable adapters that can be used with the PTA method include, for example, the xGen® Dual Index UMI adapters available from Integrated DNA Technologies (IDT). Reads from each cell are then grouped using the UMI, and reads with the same UMI are collapsed into a consensus read. The use of cellular barcodes allows all cells to be pooled before library preparation, as they can later be identified by their cellular barcodes. The use of UMIs to form consensus reads, in some cases, corrects for PCR bias and improves the detection of copy number variation (CNV) (Figures 11A and 11B). Additionally, sequencing errors can be corrected by requiring a fixed percentage of reads from the same molecule to have the same base change detected at each position. This approach has been utilized to improve CNV detection and correct the sequencing error of bulk samples.In some cases, UMI is used together with the method described herein, for example, US Patent No. 8,835,358 discloses the principle of digital counting after randomly attaching amplifiable barcodes.Schmitt et al. and Fan et al. (see above) disclose similar methods for correcting sequencing error.

[0055] The methods described herein may further include additional steps, including steps performed on a sample or template. Such a sample or template may be subjected to one or more steps prior to PTA. In some cases, a sample containing cells is subjected to a pretreatment step. For example, cells are subjected to lysis and proteolysis using a combination of freeze-thaw, Triton X-100, Tween 20, and Proteinase K to increase chromatin accessibility. Other lysis strategies are also suitable for practicing the methods described herein. Such strategies include, but are not limited to, lysis using other combinations of detergent and / or lysozyme and / or protease treatment and / or physical disruption of cells, such as sonication, alkaline lysis, and / or hypotonic lysis. In some cases, cells are lysed mechanically (e.g., high-pressure homogenizer, bead milling) or non-mechanically (physical, chemical, or biological). In some cases, physical lysis methods include heating, osmotic shock, and / or cavitation. In some cases, chemical lysis includes alkali and / or detergent. In some cases, biological lysis includes the use of enzymes. Combinations of lysis methods are also compatible with the methods described herein. Non-limiting examples of lysis enzymes include recombinant lysozyme, serine proteases, and bacterial lysine. In some cases, enzymatic lysis involves the use of lysozyme, lysostaphin, zymolase, cellulose, proteases, or glycanases. In some cases, the primary template or target molecule is subjected to a pretreatment step. In some cases, the primary template (or target) is denatured using sodium hydroxide, followed by neutralization of the solution. Other denaturation strategies may also be suitable for carrying out the methods described herein. Such strategies include, but are not limited to, combining alkaline lysis with other basic solutions, increasing the temperature and / or changing the salt concentration in the sample, adding additives such as solvents or oils, other modifications, or any combination thereof.In some cases, additional steps include sorting, filtering, or separating the sample, template, or amplicon by size. For example, after amplification using the methods described herein, the amplicon library is enriched for amplicons having a desired length. In some cases, the amplicon library is enriched for amplicons having a length of 50 to 2,000, 25 to 1,000, 50 to 1,000, 75 to 2,000, 100 to 3,000, 150 to 500, 75 to 250, 170 to 500, 100 to 500, or 75 to 2,000 bases. In some cases, the amplicon library is enriched for amplicons having a length of 75, 100, 150, 200, 500, 750, 1,000, 2,000, 5,000, or 10,000 bases or less. In some cases, the amplicon library is enriched for amplicons having a length of at least 25, 50, 75, 100, 150, 200, 500, 750, 1000, or at least 2000 bases.

[0056] The methods and compositions described herein may include buffers or other formulations. Such buffers, in some cases, include surfactants / detergents or denaturants (Tween-20, DMSO, DMF, PEGylated polymers containing hydrophobic groups, or other surfactants), salts (potassium phosphate or sodium phosphate (monobasic or dibasic), sodium chloride, potassium chloride, TrisHCl, magnesium chloride or sulfate, ammonium salts such as phosphate, nitrate, sulfate, EDTA), reducing agents (DTT, THP, DTE, beta-mercaptoethanol, TCEP, or other reducing agents), or other components (glycerol, hydrophilic polymers such as PEG). In some cases, buffers are used in combination with components such as polymerases, strand displacement factors, terminators, or other reaction components described herein. Buffers may include one or more crowding agents. In some cases, the crowding reagent includes a polymer. In some cases, the crowding reagent includes a polymer such as a polyol. In some cases, the crowding reagent includes a polyethyleneglycol polymer (PEG). In some cases, the crowding reagent includes a polysaccharide. Non-limiting examples of crowding reagents include Ficoll (e.g., Ficoll PM400, Ficoll PM70, or other molecular weight Ficoll), PEG (e.g., PEG1000, PEG2000, PEG4000, PEG6000, PEG8000, or other molecular weight PEG), and dextran (dextran 6, dextran 10, dextran 40, dextran 70, dextran 6000, dextran 138k, or other molecular weight dextran). .

[0057] Nucleic acid molecules amplified according to the methods described herein can be sequenced and analyzed using methods known to those skilled in the art. Non-limiting examples of sequencing methods used in some cases include, for example, sequencing by hybridization (SBH), sequencing by ligation (SBL) (Shendure et al. (2005) Science 309:1728), quantitative incremental fluorescent nucleotide addition sequencing (QIFNAS), stepwise ligation and cleavage, fluorescence resonance energy transfer (FRET), molecular beacons, TaqMan reporter probe digestion, pyrosequencing, fluorescence in situ sequencing (FISSEQ), FISSEQ beads (U.S. Patent No. 7,425,431), wobble sequencing (International Patent Application Publication No. WO2006 / 073504), multiplex sequencing (U.S. Patent Application Publication No. US2008 / 0269068; Porreca et al., 2007, Nat. Methods 4:931), polymerized colony (POLONY) sequencing (U.S. Pat. Nos. 6,432,360, 6,485,944, and 6,511,803, and International Patent Application Publication No. WO 2005 / 082098), nanogrid rolling circle sequencing (ROLONY) (U.S. Pat. No. 9,624,538), allele-specific oligo ligation assays (e.g., oligo ligation assay (OLA), single template molecule OLA using ligated linear probes and rolling circle amplification (RCA) readout, ligated padlock probes, and / or single template molecule OLA using ligated circular padlock probes and rolling circle amplification (RCA) readout), e.g., Roche 454, Illumina These include high-throughput sequencing methods such as those using the Solexa, AB-SOLiD, Helicos, Polonator platforms, etc., and light-based sequencing technologies (Landegrenetal. (1998) Genome Res. 8:769-76; Kwok (2000) Pharmacogenomics 1:95-100; and Shi (2001) Clin. Chem. 47:164-172).In some cases, the amplified nucleic acid molecules are shotgun sequenced. Sequencing of the sequencing library is, in some cases, performed using a suitable sequencing technology, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis (array / colony-based or nanoball-based).

[0058] Described herein are methods for generating amplicon libraries from samples containing short nucleic acids using the PTA methods described herein. In some cases, PTA results in improved fidelity and uniformity of amplification of shorter nucleic acids. In some cases, the nucleic acids are 2,000 bases or less in length. In some cases, the nucleic acids are 1,000 bases or less in length. In some cases, the nucleic acids are 500 bases or less in length. In some cases, the nucleic acids are 200, 400, 750, 1,000, 2,000, or 5,000 bases or less in length. In some cases, samples containing short nucleic acid fragments include, but are not limited to, ancient DNA (hundreds, thousands, millions, or even billions of years old), FFPE (formalin-fixed, paraffin-embedded) samples, cell-free DNA, or other samples containing short nucleic acids.

[0059] kit Described herein are kits that facilitate the implementation of PTA methods. Various combinations of components described above with respect to exemplary reaction mixtures and reaction methods can be provided in kit form. Kits may include individual components separated from one another, for example, carried in separate containers or packages. Kits may, in some cases, include one or more subcombinations of components described herein, where one or more subcombinations are separated from the other components of the kit. In some cases, subcombinations can be combined to create reaction mixtures described herein (or combined to carry out reactions described herein). In certain embodiments, the subcombinations of components present in individual containers or packages are insufficient to carry out the reactions described herein. However, the kit as a whole may, in some cases, include a collection of containers or packages, the contents of which can be combined to carry out the reactions described herein.

[0060] The kit may include suitable packaging materials for housing the contents of the kit. The packaging materials, in some cases, are preferably constructed by well-known methods to provide a sterile, contaminant-free environment. Packaging materials as used herein include those conventionally utilized in commercially available kits sold for use with, for example, nucleic acid sequencing systems. Exemplary packaging materials include, but are not limited to, glass, plastic, paper, foil, and the like, capable of retaining the components described herein within a fixed range. The packaging materials may include a label indicating a specific use of the components. The use of the kit indicated by the label may, in some cases, be one or more of the methods described herein that are appropriate for the particular combination of components present in the kit. For example, in some cases, the label may indicate that the kit is useful for a method of detecting mutations in a nucleic acid sample using a PTA method. Instructions for use of the packaged reagents or components are also included in the kit. The instructions typically include specific wording describing reaction parameters, such as the relative amounts of kit components and sample to be mixed, the duration of the reagent / sample mixture, temperature, buffer conditions, etc. It is understood that not all components required for a particular reaction need be present in a particular kit. Rather, one or more additional components are provided from other sources in some cases. In some cases, the instructions provided with the kit identify the additional components provided and where they can be obtained. In one embodiment, the kit provides at least one amplification primer, at least one nucleic acid polymerase, and a mixture of at least two nucleotides, the mixture of nucleotides including at least one terminator nucleotide that terminates nucleic acid replication by the polymerase, and instructions for using the kit. In some cases, the kit provides reagents for carrying out the methods described herein, such as PTA. In some cases, the kit further includes reagents configured for gene editing (e.g., Crispr / cas9 or other methods described herein).

[0061] In a related aspect, the present invention provides a kit comprising a reverse transcriptase, a nucleic acid polymerase, one or more amplification primers, a mixture of nucleotides including one or more terminator nucleotides, and optionally instructions for use. In one embodiment of the kit of the present invention, the nucleic acid polymerase is a strand-displacing DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase is selected from the group consisting of bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R DNA polymerase, Vent R The nucleic acid polymerase is selected from (exo-)DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-)DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, and T4 DNA polymerase. In one embodiment of the kit of the present invention, the nucleic acid polymerase has 3' to 5' exonuclease activity, and the terminator nucleotide inhibits such 3' to 5' exonuclease activity (e.g., nucleotides with modifications at the alpha group [e.g., alpha-thiodideoxynucleotides], C3 spacer nucleotides, locked nucleic acids (LNAs), inverted nucleic acids, 2' fluoronucleotides, 3' phosphorylated nucleotides, 2'-O-methyl modified nucleotides, trans nucleic acids). In one embodiment of the kit of the present invention, the nucleic acid polymerase does not have 3' to 5' exonuclease activity (e.g., Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, Vent R(exo-)DNA polymerase, Deep Vent (exo-)DNA polymerase, Klenow fragment (exo-)DNA polymerase, Therminator DNA polymerase). In one specific embodiment, the terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. In one specific embodiment, the terminator nucleotide is selected from a 3'-blocked reversible terminator comprising nucleotides, a 3'-unblocked reversible terminator comprising nucleotides, a terminator comprising a 2'-modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. In one specific embodiment, the terminator nucleotide is selected from a dideoxynucleotide, a reversed dideoxynucleotide, a 3'-biotinylated nucleotide, a 3'-amino nucleotide, a 3'-phosphorylated nucleotide, a 3'-O-methyl nucleotide, a 3'-C3 spacer nucleotide, a 3'-C18 nucleotide, a 3'-carbon spacer nucleotide comprising a 3'-hexanediol spacer nucleotide, an acyclonucleotide, and combinations thereof.

[0062] Numbered Embodiments Described herein are the following numbered embodiments 1-104. 1. Provided herein is a method of determining mutations, the method comprising: a. exposing a population of cells to a gene editing method, wherein the gene editing method utilizes a reagent configured to introduce mutations in a target sequence; b. isolating a single cell from the population; c. providing a cell lysate from the single cell; d. contacting the cell lysate with at least one amplification primer, at least one nucleic acid polymerase, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; and e. amplifying a target nucleic acid molecule to generate a plurality of terminated amplicons, wherein replication proceeds by strand displacement replication; f. ligating the molecules obtained in step (e) to adapters, thereby generating a library of amplicons; g. sequencing the library of amplicons; and h. comparing the sequences of the amplicons to at least one reference sequence to identify at least one mutation. 2. Further provided herein is the method of embodiment 1, wherein at least one mutation is present in the target sequence. 3. Further provided herein is the method of embodiment 1, wherein at least one mutation is absent from the target sequence. 4. Further provided herein is the method of embodiment 1 or 2, comprising the use of CRISPR, TALEN, ZFN, recombinase, or meganuclease. 5. Further provided herein is the method of embodiment 1 or 2, wherein the gene editing technique comprises the use of CRISPR. 6. Further provided herein is the method of embodiment 1 or 2, wherein the gene editing technique comprises the use of gene therapy. 7. Further provided herein is the method of embodiment 6, wherein the gene therapy method is not configured to modify somatic or germline DNA of the cell. 8. Further provided herein is the method of embodiment 5, wherein the reference sequence is genomic.9. Further provided herein is the method of embodiment 5, wherein the reference sequence is a specificity-determining sequence, wherein the specificity-determining sequence is configured to bind to the target sequence. 10. Further provided herein is the method of embodiment 9, wherein the at least one mutation is present in a region of sequence that differs from the specificity-determining sequence by at least one base. 11. Further provided herein is the method of embodiment 9, wherein the at least one mutation is present in a region of sequence that differs from the specificity-determining sequence by at least two bases. 12. Further provided herein is the method of embodiment 9, wherein the at least one mutation is present in a region of sequence that differs from the specificity-determining sequence by at least three bases. 13. Further provided herein is the method of embodiment 9, wherein the at least one mutation is present in a region of sequence that differs from the specificity-determining sequence by at least five bases. 14. Further provided herein is the method of embodiment 1, wherein the at least one mutation comprises an insertion, deletion, or substitution. 15. Further provided herein is the method of embodiment 5, wherein the reference sequence is the sequence of a CRISPR RNA (crRNA). 16. Further provided herein is the method of embodiment 5, wherein the reference sequence is the sequence of a single guide RNA (sgRNA). 17. Further provided herein is the method of embodiment 5, wherein the at least one mutation is in a region of a sequence that binds catalytically active Cas9. 18. Further provided herein is the method of embodiment 1, wherein the single cell is a mammalian cell. 19. Further provided herein is the method of embodiment 1, wherein the single cell is a human cell. 20. Further provided herein is the method of any one of embodiments 1-19, wherein the single cell is derived from liver, skin, kidney, blood, or lung. 21. Further provided herein is the method of any one of embodiments 1-20, wherein the single cell is a primary cell. 22. Further provided herein is the method of any one of embodiments 1-20, wherein the single cell is a stem cell.23. Further provided herein is a method according to any one of embodiments 1 to 20, wherein at least some of the amplification products comprise a barcode. 24. Further provided herein is a method according to any one of embodiments 1 to 20, wherein at least some of the amplification products comprise at least two barcodes. 25. Further provided herein is a method according to embodiment 23, wherein the barcode comprises a cell barcode. 26. Further provided herein is a method according to embodiment 23 or 25, wherein the barcode comprises a sample barcode. 27. Further provided herein is a method according to any one of embodiments 1 to 26, wherein at least some of the amplification primers comprise a unique molecular identifier (UMI). 28. Further provided herein is a method according to any one of embodiments 1 to 27, wherein at least some of the amplification primers comprise at least two unique molecular identifiers (UMI). 29. Further provided herein is a method according to any one of embodiments 1 to 27, wherein the method further comprises an additional amplification step using PCR. 30. Further provided herein is the method of any one of embodiments 1-29, wherein the method further comprises removing at least one terminator nucleotide from the terminated amplification product prior to ligation to the adapter. 31. Further provided herein is the method of any one of embodiments 1-30, wherein a single cell is isolated from the population using a method comprising a microfluidic device. 32. Further provided herein is the method of any one of embodiments 1-31, wherein the at least one mutation occurs in less than 50% of the population of cells. 33. Further provided herein is the method of any one of embodiments 1-31, wherein the at least one mutation occurs in less than 25% of the population of cells. 34. Further provided herein is the method of any one of embodiments 1-31, wherein the at least one mutation occurs in less than 1% of the population of cells. 35. Further provided herein is the method of any one of embodiments 1-31, wherein the at least one mutation occurs in 0.1% or less of the population of cells.36. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.01% or less of the population of cells. 37. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.001% or less of the population of cells. 38. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.0001% or less of the population of cells. 39. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 25% or less of the amplification product sequences. 40. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 1% or less of the amplification product sequences. 41. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.1% or less of the amplification product sequences. 42. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.01% or less of the amplified product sequences. 43. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.001% or less of the amplified product sequences. 44. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation occurs in 0.0001% or less of the amplified product sequences. 45. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation is present in a region of the sequence that is correlated with a genetic disease or condition. 46. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation is present in a region of the sequence that is not correlated with the binding of a DNA repair enzyme. 47. Further provided herein is a method according to any one of embodiments 1 to 31, wherein the at least one mutation is present in a region of the sequence that is not correlated with the binding of MRE11.48. Further provided herein is the method of any one of embodiments 1 to 31, further comprising identifying previously sequenced false-positive mutations by an alternative off-target detection method. 49. Further provided herein is the method of embodiment 48, wherein the off-target detection method is in silico prediction, ChIP-seq, GUIDE-seq, circle-seq, HTGTS (high-throughput genome-wide translocation sequencing), IDLV (integration-deficient lentivirus), Digenome-seq, FISH (fluorescence in situ hybridization), or DISCOVER-seq. 50. Described herein is a method for identifying a specificity-determining sequence, the method comprising: a. providing a library of nucleic acids, wherein at least some of the nucleic acids comprise a specificity-determining sequence; b. performing a gene editing method on at least one cell, wherein the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; c. sequencing the genome of the at least one cell using the method of embodiments 1 to 38, wherein the specificity-determining sequence contacted with the at least one cell is identified; and d. identifying at least one specificity-determining sequence that provides the fewest off-target mutations. 51. Further provided herein is a method according to embodiment 50, wherein the off-target mutation is a silent mutation. 52. Further provided herein is a method according to embodiment 50, wherein the off-target mutation is located outside a gene coding region. 53. Described herein is a method for in vivo mutation analysis, comprising: a. performing a gene editing method on at least one cell in an organism, wherein the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; b. isolating at least one cell from the organism; and c. sequencing the genome of the at least one cell using the methods of embodiments 1-49.54. Further provided herein is the method of embodiment 53, wherein the method comprises at least two cells. 55. Further provided herein is the method of embodiment 154, further comprising identifying mutations by comparing the genome of the first cell with the genome of the second cell. 56. Further provided herein is the method of embodiment 54 or 55, wherein the first cell and the second cell are from different tissues. 57. Described herein is a method for predicting the age of a subject, the method comprising: a. providing at least one sample from the subject, wherein the at least one sample comprises a genome; b. providing any of embodiments 1-38 to identify mutations.

[0042] The method includes: sequencing the genome using a method according to any one of embodiments 57 to 60; c. comparing the mutations obtained in step b to a standard reference curve, wherein the standard reference curve correlates the number and location of mutations with a verified age; and d. predicting the age of the subject based on comparison of the mutations to the standard reference curve. 58. Further provided herein is the method of embodiment 57, wherein the standard reference curve is specific to the gender of the subject. 59. Further provided herein is the method of embodiment 57, wherein the standard reference curve is specific to the ethnicity of the subject. 60. Further provided herein is the method of embodiment 57, wherein the standard reference curve is specific to the geographic location of the subject where the subject spent a period of their lifetime. 61. Further provided herein is the method of any one of embodiments 57 to 60, wherein the subject is under 50 years of age. 62. Further provided herein is the method of any one of embodiments 57 to 60, wherein the subject is under 18 years of age. 63. Further provided herein is the method of any one of embodiments 57-60, wherein the subject is under 15 years old. 64. Further provided herein is the method of any one of embodiments 57-63, wherein the at least one sample is more than 10 years old. 65. Further provided herein is the method of any one of embodiments 57-63, wherein the at least one sample is more than 100 years old. 66. Further provided herein is the method of any one of embodiments 57-63, wherein the at least one sample is more than 1000 years old. 67. Further provided herein is the method of any one of embodiments 57-66, wherein at least two samples are sequenced. 68. Further provided herein is the method of any one of embodiments 57-66, wherein at least five samples are sequenced. 69. Further provided herein is the method of embodiment 67, wherein the at least two samples are from different tissues.70. Described herein is a method for sequencing a microbial or viral genome, comprising: a. obtaining a sample containing one or more genomes or genome fragments; b. sequencing the sample using a method described in any one of embodiments 1-38 to obtain a plurality of sequencing reads; and c. assembling and sorting the sequencing reads to generate a microbial or viral genome. 71. Further provided herein is the method of embodiment 70, wherein the sample comprises genomes from at least two organisms. 72. Further provided herein is the method of embodiment 70, wherein the sample comprises genomes from at least 10 organisms. 73. Further provided herein is the method of embodiment 70, wherein the sample comprises genomes from at least 100 organisms. 74. Further provided herein is the method of any one of embodiments 70-73, wherein the sample originates from an environment, including a deep-sea vent, ocean, mine, stream, lake, meteorite, glacier, or volcano. 75. Further provided herein is the method of any one of embodiments 70 to 74, further comprising identifying at least one gene in the microbial genome. 76. Further provided herein is the method of any one of embodiments 70 to 75, wherein the microbial genome corresponds to an unculturable organism. 77. Further provided herein is the method of embodiment 76, wherein the microbial genome corresponds to a symbiont. 78. Further provided herein is the method of any one of embodiments 70 to 77, further comprising cloning at least one gene in a recombinant host organism. 79. Further provided herein is the method of embodiment 78, wherein the recombinant host organism is a bacterium. 80. Further provided herein is the method of embodiment 79, wherein the recombinant host organism is Escherichia, Bacillus, or Streptomyces. 81. Further provided herein is the method of embodiment 78, wherein the recombinant host organism is a eukaryotic cell. 82. Further provided herein is the method of embodiment 81, wherein the recombinant host organism is a yeast cell.83. Further provided herein is the method of embodiment 82, wherein the recombinant host organism is Saccharomyces or Pichia. 84. Described herein is a kit for nucleic acid sequencing, the kit comprising: a. at least one amplification primer; b. at least one nucleic acid polymerase; c. a mixture of at least two nucleotides, the mixture of nucleotides comprising at least one terminator nucleotide that terminates nucleic acid replication by the polymerase; and d. instructions for use of the kit to perform nucleic acid sequencing. 85. Further provided herein is the kit of embodiment 84, wherein at least one amplification primer is a random primer. 86. Further provided herein is the kit of embodiment 84, wherein the nucleic acid polymerase is a DNA polymerase. 87. Further provided herein is the kit of embodiment 86, wherein the DNA polymerase is a strand-displacing DNA polymerase. 88. Further provided herein is the kit of any of embodiments 84-87, wherein the nucleic acid polymerase is bacteriophage phi29 (Φ29) polymerase, genetically modified phi29 (Φ29) DNA polymerase, Klenow fragment of DNA polymerase I, phage M2 DNA polymerase, phage phi PRD1 DNA polymerase, Bst DNA polymerase, Bst large fragment DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, VentR DNA polymerase, VentR (exo-) DNA polymerase, Deep Vent DNA polymerase, Deep Vent (exo-) DNA polymerase, IsoPol DNA polymerase, DNA polymerase I, Therminator DNA polymerase, T5 DNA polymerase, Sequenase, T7 DNA polymerase, T7-Sequenase, or T4 DNA polymerase.89. Further provided herein is a kit according to any of embodiments 84-88, wherein the nucleic acid polymerase comprises 3' to 5' exonuclease activity and at least one terminator nucleotide inhibits the 3' to 5' exonuclease activity. 90. Further provided herein is a kit according to any of embodiments 84-88, wherein the nucleic acid polymerase does not comprise 3' to 5' exonuclease activity. 91. Further described herein is a kit according to any of embodiments 84-88, wherein the polymerase is Bst DNA polymerase, exo(-)Bst polymerase, exo(-)Bca DNA polymerase, Bsu DNA polymerase, VentR (exo-) DNA polymerase, Deep Vent (exo-) DNA polymerase, Klenow fragment (exo-) DNA polymerase, or Therminator DNA polymerase. 92. Further provided herein is a kit according to any one of embodiments 84 to 92, wherein at least one terminator nucleotide comprises a modification of the r group of the 3' carbon of deoxyribose. 93. Further provided herein is a kit according to any one of embodiments 84 to 92, wherein at least one terminator nucleotide is selected from the group consisting of a 3' blocked reversible terminator comprising nucleotides, a 3' unblocked reversible terminator comprising nucleotides, a terminator comprising a 2' modification of a deoxynucleotide, a terminator comprising a modification to the nitrogenous base of a deoxynucleotide, and combinations thereof. 94. Further described herein is a kit according to any one of embodiments 84 to 93, wherein the at least one terminator nucleotide is selected from the group consisting of dideoxynucleotides, inverted dideoxynucleotides, 3' biotinylated nucleotides, 3' amino nucleotides, 3'-phosphorylated nucleotides, 3'-O-methyl nucleotides, 3' C3 spacer nucleotides, 3' C18 nucleotides, 3' carbon spacer nucleotides including 3' hexanediol spacer nucleotides, acyclonucleotides, and combinations thereof.95. Further described herein is a kit according to any one of embodiments 84-94, wherein at least one terminator nucleotide is selected from the group consisting of a nucleotide with a modification in an alpha group, a C3 spacer nucleotide, a locked nucleic acid (LNA), an inverted nucleic acid, a 2'-fluoronucleotide, a 3'-phosphorylated nucleotide, a 2'-O-methyl modified nucleotide, and a trans nucleic acid. 96. Further described herein is a kit according to any one of embodiments 84-95, wherein the nucleotide with a modification in an alpha group is an alpha-thiodideoxynucleotide. 97. Further described herein is a kit according to any one of embodiments 84-96, wherein the amplification primers are 4 to 70 nucleotides in length. 98. Further described herein is a kit according to any one of embodiments 84-97, wherein at least one amplification primer is 4 to 20 nucleotides in length. 99. Further described herein is a kit according to any one of embodiments 84-98, wherein at least one amplification primer comprises a randomized region. 100. Further provided herein is the kit of embodiment 99, wherein the randomized region is 4 to 20 nucleotides in length. 101. Further provided herein is the kit of embodiment 99 or 100, wherein the randomized region is 8 to 15 nucleotides in length. 102. Further provided herein is the kit of any one of embodiments 84 to 101, wherein the kit further comprises a library preparation kit. 103. Further provided herein is the library preparation kit comprising one or more of: a. at least one polynucleotide adaptor; b. at least one high-fidelity polymerase; c. at least one ligase; d. reagents for nucleic acid shearing; and e. at least one primer, wherein the primer is configured to bind to the adaptor. 104. Further provided herein is the kit of any one of embodiments 84 to 103, wherein the kit further comprises reagents configured for gene editing. [Example]

[0063] The following examples are presented to more clearly illustrate to those skilled in the art the principles and practice of the embodiments disclosed herein, and should not be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are by weight.

[0064] Example 1: Primary template-directed amplification (PTA) Although PTA can be used for any nucleic acid amplification, it is particularly useful for whole genome amplification because it allows for capturing a greater percentage of the cellular genome in a more uniform and reproducible manner, and with a lower error rate, than currently used methods, such as Multiple Displacement Amplification (MDA), while avoiding the drawbacks of currently used methods, such as exponential amplification where the polymerase initially extends random primers, which results in random over-representation of loci and alleles and propagation of mutations (see Figures 1A-1C).

[0065] cell culture Human NA12878 (Coriell Institute) cells were maintained in RPMI medium supplemented with 15% FBS and 2 mM L-glutamine, as well as 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B (Gibco, Life Technologies). Cells were cultured at 3.5 × 10 5 Cells were seeded at a density of 1000 cells / ml. Cultures were split every 3 days and maintained in a humidified incubator at 37°C with 5% CO2.

[0066] Single cell isolation and WGA 3.5×10 5After seeding at a density of 1000 cells / ml, NA12878 cells were cultured for a minimum of 3 days, after which 3 mL of cell suspension was pelleted at 300 × g for 10 min. The medium was then discarded, and the cells were washed with 1 mL of cell wash buffer (Mg 2+ or Ca 2+ The cells were washed three times with 1x PBS containing 2% FBS (without ATP) by spinning at 300 x g, 200 x g, and finally 100 x g for 5 minutes. The cells were then resuspended in 500 μL of cell wash buffer. Following this, they were stained with 100 nM calcein AM (Molecular Probes) and 100 ng / ml propidium iodide (PI; Sigma-Aldrich) to differentiate the live cell population. The cells were loaded onto a BD FACScan flow cytometer (FACSAria II) (BD Biosciences) thoroughly washed with ELIMINase (Decon Labs) and calibrated using Accudrop fluorescent beads (BD Biosciences) for cell sorting. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS containing 0.2% Tween 20, followed by PTA (Sigma-Aldrich). Several wells were intentionally left empty to serve as no template controls (NTCs). Immediately after sorting, plates were briefly centrifuged and placed on ice. Cells were then frozen at -20°C for a minimum of overnight. The following day, WGA reactions were assembled in a pre-PCR workstation that provided constant positive pressure of HEPA-filtered air and was decontaminated with UV light for 30 minutes before each experiment.

[0067] MDA was performed using a modification previously shown to improve amplification uniformity. Specifically, exonuclease-resistant random primers were added to the lysis buffer / mix to a final concentration of 125 μM. 4 μL of the resulting lysis / denaturation mix was added to the tube containing the single cells, vortexed, briefly spun, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, vortexed, briefly centrifuged, and placed at room temperature. 40 μL of amplification mix was then added, followed by incubation at 30°C for 8 hours, followed by heating to 65°C for 3 minutes to terminate the amplification.

[0068] PTA was performed by first further lysing the cells after freeze-thawing by adding 2 μl of a pre-chilled solution of a 1:1 mixture of 5% Triton X-100 (Sigma-Aldrich) and 20 mg / ml proteinase K (Promega). The cells were then vortexed, briefly centrifuged, and placed at 40°C for 10 minutes. Next, 4 μl of lysis buffer / mix and 1 μl of 500 μM exonuclease-resistant random primers were added to the lysed cells to denature the DNA, followed by vortexing, rotation, and placement at 65°C for 15 minutes. Next, 4 μl of room-temperature quenching buffer was added, and the sample was vortexed and spun down. 56 μl of amplification mix (primers, dNTPs, polymerase, buffer) contained equal ratios of alpha-thio-ddNTPs at a concentration of 1200 μM in the final amplification reaction. The sample was then placed at 30°C for 8 hours, after which amplification was stopped by heating at 65°C for 3 minutes.

[0069] After the amplification step, DNA from both the MDA and PTA reactions was purified using AMPure XP magnetic beads (Beckman Coulter) at a 2:1 ratio of beads to sample, and yields were measured using a Qubit dsDNA HS Assay Kit according to the manufacturer's instructions (Life Technologies) using a Qubit 3.0 fluorometer.

[0070] Library preparation The MDA reaction resulted in the generation of 40 μg of amplified DNA. 1 μg of product was enzymatically fragmented for 30 minutes according to standard procedures. The sample then underwent standard library preparation using 15 μM dual-index adapters (T4 polymerase, T4 polynucleotide kinase, and end repair with Taq polymerase for A-tailing) and four cycles of PCR. Each PTA reaction generated 40–60 ng of material used for standard DNA sequencing library preparation. 2.5 μM adapters with UMI and dual index were used for ligation with T4 ligase, and 15 cycles of PCR (hot-start polymerase) were used in the final amplification. The library was then cleaned up using two-sided SPRI using ratios of 0.65× and 0.55× for right- and left-hand selection, respectively. The final library was quantified using the Qubit dsDNA BR Assay Kit and a 2100 Bioanalyzer (Agilent Technologies) and subsequently sequenced on the Illumina NextSeq platform. All Illumina sequencing platforms, including NovaSeq, are also compatible with this protocol.

[0071] Data analysis Sequencing reads were demultiplexed based on cell barcodes using Bcl2fastq. Reads were then trimmed using trimmomatic and aligned to hg19 using BWA. Reads were duplicate-marked using Picard, followed by local realignment and base recalibration using GATK4.0. All files used to calculate quality metrics were downsampled to 20 million reads using Picard DownSampleSam. Quality metrics were obtained from the final bam files using qualimap and Picard AlignmentSummaryMetrics and CollectWgsMetrics. Total genome coverage was also estimated using Preseq.

[0072] Variant calling Single nucleotide variants and indels were called using the GATK Unified Genotyper from GATK 4.0. Standard filtering criteria using GATK best practices were used throughout the process (https: / / software.broadinstitute.org / gatk / best-practices / ). Copy number variants were called using Control-FREEC (Boeva ​​et al., Bioinformatics, 2012, 28(3):423-5). Structural variants were also detected using CREST (Wang et al., Nat Methods, 2011, 8(8):652-4).

[0073] result As shown in Figure 3A and Figure 3B, the mapping rate and mapping quality score for amplification using only dideoxynucleotides ("reversible") are 15.0 + / - 2.2 and 0.8 + / - 0.08, respectively, while incorporation of exonuclease-resistant alpha-thiodideoxynucleotide terminators ("irreversible") yields mapping rates and quality scores of 97.9 + / - 0.62 and 46.3 + / - 3.18, respectively. Experiments were also performed using reversible ddNTPs and various concentrations of terminators (Figure 2A, bottom).

[0074] Figures 2B–2E show comparative data generated from NA12878 human single cells subjected to MDA (according to Dong, X. et al., Nat Methods. 2017, 14(5):491–493) or PTA. Both protocols produced comparable low PCR duplication rates (MDA 1.26% + / - 0.52 vs. PTA 1.84% + / - 0.99) and GC% (MDA 42.0 + / - 1.47 vs. PTA 40.33 + / - 0.45), but PTA produced smaller amplicon sizes. The percentage of mapped reads and mapping quality score were also significantly higher with PTA compared to MDA (PTA 97.9 + / - 0.62 vs. MDA 82.13 + / - 0.62 and PTA 46.3 + / - 3.18 vs. MDA 43.2 + / - 4.21, respectively). Overall, PTA generates more usable mapped data when compared to MDA. Figure 4A shows that compared to MDA, PTA significantly improves amplification uniformity, resulting in broader coverage and fewer regions with near-zero coverage. The use of PTA can identify low-frequency sequence variants within a population of nucleic acids containing the variant, which constitutes 0.01% or more of the total sequence. PTA can be successfully used for amplification of single-cell genomes.

[0075] Example 2: Comparative Analysis of PTA Benchmarking the maintenance and isolation of PTA and SCMDA cells Lymphoblastoid cells from the 1000 Genomes Project target NA12878 (Coriell Institute, Camden, NJ, USA) were maintained in RPMI medium supplemented with 15% FBS, 2 mM L-glutamine, 100 units / mL penicillin, 100 μg / mL streptomycin, and 0.25 μg / mL amphotericin B. Cells were cultured at 3.5 × 10 5 Cells were seeded at a density of 1000 cells / ml and split every 3 days. They were maintained in a humidified incubator at 37°C with 5% CO2. Before isolating single cells, 3 mL of cell suspension expanded over the past 3 days was spun at 300 x g for 10 min. Pelleted cells were washed with 1 mL of cell wash buffer (Mg 2+ or Ca 2+ The cells were washed three times with 1x PBS containing 2% FBS without ATP and spun consecutively at 300 x g, 200 x g, and finally 100 x g for 5 minutes to remove dead cells. Next, the cells were resuspended in 500 μL of cell wash buffer and subsequently stained with 100 nM calcein AM and 100 ng / ml propidium iodide (PI) to differentiate the live cell population. The cells were thoroughly washed with ELIMINase and loaded onto a calibrated BD FACScan flow cytometer (FACSAria II) using Accudrop fluorescent beads. Single cells from the calcein AM-positive, PI-negative fraction were sorted into each well of a 96-well plate containing 3 μL of PBS with 0.2% Tween 20. Several wells were intentionally left empty to serve as no-template controls. Immediately after sorting, the plate was briefly centrifuged and placed on ice. The cells were then frozen at -80°C for at least overnight.

[0076] PTA and SCMDA experiments WGA reactions were assembled on a pre-PCR workstation, which provided constant positive pressure with HEPA-filtered air and was decontaminated with UV light for 30 minutes before each experiment. MDA was performed according to the published protocol (Dong et al. Nat. Meth. 2017, 14, 491-493) using SCMDA. Specifically, exonuclease-resistant random primers were added to the lysis buffer at a final concentration of 12.5 μM. 4 μL of the resulting lysis mix was added to the tube containing the single cells, mixed by pipetting three times, briefly spun, and incubated on ice for 10 minutes. The cell lysate was neutralized by adding 3 μL of quenching buffer, mixed by pipetting three times, briefly centrifuged, and placed on ice. Following this, 40 μL of amplification mix was added, followed by incubation at 30°C for 8 hours, and then amplification was terminated by heating to 65°C for 3 minutes. PTA was performed by first lysing the cells after freeze-thawing by adding 2 μL of a pre-chilled 1:1 mixture of 5% Triton X-100 and 20 mg / ml proteinase K. The cells were then vortexed, briefly centrifuged, and then placed at 40°C for 10 minutes. Next, 4 μL of denaturation buffer and 1 μL of 500 μM exonuclease-resistant random primers were added to the lysed cells to denature the DNA, followed by vortexing, rotation, and placing at 65°C for 15 minutes. Next, 4 μL of room-temperature quenching solution was added, and the sample was vortexed and spun down. In the final amplification reaction, 56 μL of amplification mix contained an equal ratio of alpha-thio-ddNTPs at a concentration of 1200 μM. The sample was then placed at 30°C for 8 hours, after which amplification was terminated by heating at 65°C for 3 minutes. After SCMDA or PTA amplification, DNA was purified using AMPure XP magnetic beads at a 2:1 bead-to-sample ratio, and yields were measured using the Qubit 3.0 fluorometer with the Qubit dsDNA HS Assay Kit according to the manufacturer's instructions. PTA experiments were also performed using reversible ddNTPs and various concentrations of terminators (Figure 2A, top).

[0077] Library preparation Following standard protocols, 1 μg of SCMDA product was enzymatically fragmented for 30 minutes. The sample then underwent standard library preparation using 15 μM unique dual-index adapters and four cycles of PCR. The entire product of each PTA reaction was used for DNA sequencing library preparation without fragmentation. 2.5 μM unique dual-index adapters were used in ligation, and 15 cycles of PCR were used in the final amplification. SCMDA and PTA libraries were then visualized on a 1% agarose E-Gel. 400-700 bp fragments were excised from the gel and recovered using the Gel DNA Recovery Kit. The final libraries were quantified using the Qubit dsDNA BR Assay Kit and Agilent 2100 Bioanalyzer before sequencing on the NovaSeq 6000.

[0078] Data analysis Data were trimmed using trimmomatic and then aligned to hg19 using BWA. Reads were marked for duplicates using Picard, followed by local realignment and base recalibration using best practices in GATK3.5. All files were downsampled to the specified number of reads using Picard DownSampleSam. Quality metrics were obtained from the final bam files using qualimap and Picard AlignmentMetricsAnumary and CollectWgsMetrics. Lorenz curves were plotted and Gini indices were calculated using htSeqTools. SNV calling was performed using UnifiedGenotyper, which was then filtered using standard recommended criteria (QD<2.0 ||FS>60.0 ||MQ<40.0 ||SOR>4.0 ||MQRankSum<-12.5 ||ReadPosRankSum<-8.0). No regions were excluded from the analysis, and no other data normalization or manipulation was performed. Sequencing metrics for the tested methods are shown in Table 1.

[0079] [Table 1]

[0080] Breadth and uniformity of genome coverage We performed a comprehensive comparison of PTA with all common single-cell WGA methods. To achieve this, we performed PTA and an improved version of MDA called single-cell MDA (Dong et al. Nat. Meth. 2017, 14, 491-493) (SCMDA) on 10 NA12878 cells each. Furthermore, these results were compared for cells subjected to amplification using DOP-PCR (Zhang et al. PNAS 1992, 89, 5847-5851), MDA Kit 1 (Dean et al. PNAS 2002, 99, 5261-5266), MDA Kit 2, MALBAC (Zong et al. Science 2012, 338, 1622-1626), LIANTI (Chen et al., Science 2017, 356, 189-194), or PicoPlex (Langmore, Pharmacogenomics 3, 557-560 (2002)) using data generated as part of the LIANTI study.

[0081] To normalize across samples, raw data from all samples were aligned and preprocessed for variant calling using the same pipeline. The bam files were then subsampled to 300 million reads each before comparisons were performed. Importantly, PTA and SCMDA products were not screened before further analysis, while all other methods were screened for genome coverage and uniformity before selecting the highest-quality cells used in subsequent analyses. Notably, SCMDA and PTA were compared to bulk diploid NA12878 samples, while all other methods were compared to bulk BJ1 diploid fibroblasts used in the LIANTI study. As seen in Figures 3C-3F, PTA had the highest percent of reads aligned to the genome and the highest mapping quality. PTA, LIANTI, and SCMDA had similar GC content, all of which were lower than the other methods. PCR replication rates were similar across all methods. Furthermore, the PTA method allowed smaller templates, such as mitochondrial genomes, to give higher coverage rates (similar to larger standard chromosomes) compared to other methods tested (Figure 3G).

[0082] Next, we compared the coverage breadth and uniformity of all methods. Example coverage plots across chromosome 1 are shown for SCMDA and PTA, showing that PTA significantly improved coverage uniformity and allele frequency (Figure 4B). Next, we calculated the coverage percentage for all methods using increasing read counts. PTA approached two bulk samples at all depths, a significant improvement over all other methods (Figure 5A). Next, we measured coverage uniformity using two strategies. The first approach was to calculate the coefficient of variation of coverage at increasing sequencing depths, where PTA was found to be more uniform than all other methods (Figure 5B). The second strategy was to calculate Lorenz curves for each subsampled bam file, where PTA was again found to have the greatest uniformity (Figure 5C). To measure the reproducibility of amplification uniformity, the Gini index was calculated to estimate the difference of each amplification reaction from perfect uniformity (de Bourcy et al., PloS One 9, e105585 (2014)). PTA was again shown to be more reproducibly uniform than the other methods (Figure 5D).

[0083] SNV sensitivity To determine the impact of these differences on the performance of amplification methods for SNV calling, we compared the variant call rates of each method relative to the corresponding bulk sample at increasing sequencing depths. To estimate sensitivity, we compared the percentage of variants called in the corresponding bulk sample, subsampled to 650 million reads found in each cell at each sequencing depth (Figure 5E). The improved coverage and uniformity of PTA resulted in the detection of 45.6% more variants than the next most sensitive method, MDA Kit 2. Examination of sites called as heterozygous in the bulk sample showed that PTA significantly reduced allele skewing at those heterozygous sites (Figure 5F). This finding supports the assertion that PTA not only amplifies more uniformly across the genome, but also more uniformly amplifies two alleles within the same cell.

[0084] SNV accuracy To estimate the accuracy of variant calling, variants called in each single cell that were not found in the corresponding bulk sample were considered false positives. Cryolysis of SCMDA significantly reduced the number of false positive variant calls (Figure 5G). Methods using thermostable polymerases (MALBAC, PicoPlex, and DOP-PCR) showed a further decrease in SNV calling accuracy with increasing sequencing depth. Without being bound by theory, this may be the result of a significantly increased error rate for these polymerases compared to Phi29 DNA polymerase. Furthermore, the base change patterns observed in the false positive calls also appear to be polymerase-dependent (Figure 5H). As seen in Figure 5G, the model of suppressed error propagation in PTA is supported by the lower false positive SNV call rate in PTA compared to standard MDA protocols. Furthermore, PTA had the lowest allele frequency for false positive variant calls, which is also consistent with the model of suppressed error propagation by PTA (Figure 5I).

[0085] Example 3: Direct Measurement of Environmental Mutagenicity (DMEM) Using PTA, we developed a novel mutagenicity assay that provides a framework for conducting high-resolution, genome-wide human toxicogenomics studies. Previous studies, such as the Ames test, rely on bacterial genetics to generate measurements assumed to be representative of human cells but provide limited information about the number and patterns of mutations induced in each exposed cell. To overcome these limitations, we developed a human mutagenesis system, the Direct Measurement of Environmental Mutagenicity (DMEM), in which single human cells are exposed to environmental compounds, isolated as single cells, and subjected to single-cell sequencing to identify new mutations induced in each cell.

[0086] Umbilical cord blood cells expressing the stem / progenitor marker CD34 were exposed to increasing concentrations of the direct mutagen N-ethyl-N-nitrosourea (ENU). ENU is known to have a relatively low Swain-Scott substrate constant and has been shown to act primarily via a two-step SN1 mechanism, resulting in preferential alkylation of O4-thymine, O2-thymine, and O2-cytosine. Through limited sequencing of target genes, ENU has also been shown to preferentially convert T to A (A to T), T to C (A to G), and C to T (G to A) into amino acids in mice, a pattern significantly different from that seen in E. coli.

[0087] Isolation and expansion of cord blood cells for mutagenicity experiments ENU (CAS 759-73-9) and D-mannitol (CAS 69-65-8) were placed in solution at their maximum solubility. Fresh anticoagulated umbilical cord blood (CB) was obtained from the St. Louis Cord Blood Bank. CB was diluted 1:2 with PBS, and mononuclear cells (MNCs) were isolated by density gradient centrifugation on Ficoll-Paque Plus according to the manufacturer's instructions. CD34-expressing CB MNCs were then immunomagnetically selected using a human CD34 microbead kit and a magnetic cell sorting (MACS) system according to the manufacturer's instructions. Cell counting and viability were assessed using a Luna FL cell counter. CB CD34+ cells were cultured at 2.5 x 10 in StemSpan SFEM supplemented with 1x CD34+ expansion supplement, 100 units / mL penicillin, and 100 μg / mL streptomycin. 4 cells / mL and allowed to expand for 96 h before proceeding to mutagen exposure.

[0088] Direct measurement of environmental mutagenicity (DMEM) Expanded cord blood CD34+ cells were cultured in StemSpan SFEM supplemented with 1x CD34+ proliferation supplement, 100 units / mL penicillin, and 100 μg / mL streptomycin. Cells were exposed to ENU at concentrations of 8.54, 85.4, and 854 μM, D-mannitol at 1152.8 and 11528 μM, or 0.9% sodium chloride (vehicle control) for 40 hours. Single-cell suspensions from drug-treated and vehicle control samples were harvested and stained for viability as described above. Single-cell sorting was performed as described above. PTA was performed and libraries were prepared using the general method described herein, as well as a simplified and improved protocol according to Example 2.

[0089] Analysis of DMEM data Data from cells in the DMEM experiment were trimmed using Trimmomatic, aligned to GRCh38 using BWA, and further processed using best practices in GATK 4.0.1 without deviating from recommended parameters. Genotyping was performed using HaplotypeCaller, and joint genotypes were filtered again using standard parameters. Variants were considered to be the result of a mutagen only if they had a Phred quality score of at least 100 and were detected in a single cell but not in the bulk sample. The trinucleotide context of each SNV was determined by extracting surrounding bases from the reference genome using bedtools. The number and context of mutations were visualized using ggplot2 and heatmap2 in R.

[0090] To determine whether mutations are enriched in DNase I hypersensitive sites (DHSs) in CD34+ cells, we calculated the proportion of SNVs in each sample that overlapped with DHS sites from 10 CD34+ primary cell datasets generated by the Roadmap Epigenomics Project. DHS sites spanned two nucleosomes or 340 bases in either direction. Each DHS dataset was paired with a single cell sample, where we determined the proportion of the human genome with at least 10-fold coverage within that cell that overlapped with the DHS and compared this to the proportion of SNVs found within the covered DHS sites.

[0091] result Consistent with these studies, we observed a dose-dependent increase in the number of mutations in each cell, with similar numbers detected at the lowest dose of ENU compared to either the vehicle control or a toxic dose of mannitol (Figure 12A). Also consistent with previous studies in mice using ENU, the most common mutations were T to A (A to T), T to C (A to G), and C to T (G to A). Although C to G (G to C) transversions appear to be rare, the other three types of base changes were also observed (Figure 12B). Examination of the trinucleotide context of SNVs illustrates two distinct patterns (Figure 12C). First, cytosine mutagenesis appears to be rare when cytosine is followed by guanine. Cytosine followed by guanine is typically methylated at the fifth carbon position in the human genome, a marker of heterochromatin. Without being bound by theory, we hypothesized that 5-methylcytosine is not subject to ENU alkylation due to its inaccessibility to heterochromatin or as a result of unfavorable reaction conditions with 5-methylcytosine compared to cytosine. To test the former hypothesis, we compared the locations of the mutation sites to known DNase I hypersensitive sites in CD34+ cells cataloged by the Roadmap Epigenomics Project. As seen in Figure 12D, no enrichment of cytosine variants at DNase I hypersensitive sites was observed. Furthermore, no enrichment of variants restricted to cytosine was observed at DH sites (Figure 12E). Furthermore, most thymine variants occur where an adenine precedes the thymine. The genomic feature annotation of the variants did not significantly differ from the annotation of their function within the genome (Figure 12F).

[0092] Example 4: Massively parallel single-cell DNA sequencing A protocol for massively parallel DNA sequencing using PTA is established. First, cell barcodes are added to random primers. Two strategies are employed to minimize any amplification bias introduced by the cell barcodes: 1) increasing the size of the random primers and / or 2) creating primers that loop back on themselves to prevent the cell barcode from binding to the template (Figure 10B). Once the optimal primer strategy is established, sorting is scaled up to 384 sorted cells using, for example, the Mosquito HTS liquid handler, which can pipette even viscous liquids up to 25 nL with high precision. This liquid handler reduces reagent costs by approximately 50-fold by using a 1 μL PTA reaction instead of the standard 50 μL reaction volume.

[0093] The amplification protocol is transferred to the droplets by delivering primers bearing cell barcodes to the droplets. Solid supports, such as beads created using a split-and-pool strategy, are optionally used. Suitable beads are available, for example, from ChemGenes. The oligonucleotides, in some cases, include random primers, cell barcodes, unique molecular identifiers, and cleavable sequences or spacers for releasing the oligonucleotides after the beads and cells are encapsulated in the same droplet. During this process, the concentrations of template, primers, dNTPs, alpha-thio-ddNTPs, and polymerase in the droplets are optimized for low nanoliter volumes. Optimization, in some cases, involves the use of larger droplets to increase reaction volume. As shown in Figure 9, this process requires two sequential reactions to lyse the cells, followed by WGA. The first droplet containing the lysed cells and beads is combined with the second droplet containing the amplification mix. Alternatively, or in combination, cells can be encapsulated in hydrogel beads before lysis, and then both beads are added to the oil droplet. See Lan, F. et al., Nature Biotechnol., 2017, 35:640-646.

[0094] Additional methods include the use of microwells, which in some cases capture 140,000 single cells in a 20 picoliter reaction chamber on a device the size of a 3" x 2" microscope slide. Similar to droplet-based methods, these wells combine cells with beads containing cell barcodes to enable massively parallel processing. See Gole et al., Nature Biotechnol., 2013, 31:1126-1132.

[0095] Example 5: Application of PTA to childhood acute lymphoblastic leukemia (ALL) Single-cell exome sequencing of individual leukemia cells harboring the ETV6-RUNX1 translocation measured approximately 200 coding mutations per cell, of which only 25 were present in enough cells to be detected using standard bulk sequencing for that patient. The mutation load per cell was then incorporated into other known characteristics of this type of leukemia, such as the replication-associated mutation rate (1 coding mutation / 300 cell divisions), time from onset to diagnosis (4.2 years), and population size at diagnosis (100 billion cells) to create an in silico simulation of disease development. Even in what was thought to be a genetically simple cancer, such as childhood ALL, we unexpectedly discovered an estimated 330 million clones with distinct coding mutation profiles at the time of diagnosis for that patient. Interestingly, as seen in Figure 6B, standard bulk sequencing detected only 1–5 of the most abundant clones (Box C), leaving tens of millions of clones (Box A) that are likely to be clinically significant because they comprise a small number of cells. Thus, methods are provided to enhance the sensitivity of detection, allowing the detection of clones (Box B) that comprise at least 0.01% (1:10,000) of cells, as this is the hypothesized stratum in which the most resistant disease leading to relapse resides.

[0096] Given the genetic diversity of such large populations, it is hypothesized that within a given patient, clones exist that are more resistant to treatment. To test this hypothesis, samples are cultured and leukemia cells are exposed to increasing concentrations of standard ALL chemotherapy drugs. As seen in Figure 7, clones with activating KRAS mutations continued to expand in control samples and in samples treated with the lowest dose of asparaginase. However, this clone proved more sensitive to prednisolone and daunorubicin, while other previously undetectable clones became more clearly detectable after treatment with these drugs (Figure 7, dashed box). This approach also employed bulk sequencing of the treated samples. The use of single-cell DNA sequencing allows, in some cases, the determination of the diversity and clonal type of the expanding population.

[0097] Catalogue of ALL clonotype drug sensitivity As shown in Figure 8, to create a catalog of ALL clonotype drug sensitivities, an aliquot of the diagnostic sample is taken and single-cell sequencing of 10,000 cells is performed to determine the abundance of each clonotype. In parallel, diagnostic leukemia cells are exposed in vitro to standard ALL drugs (vincristine, daunorubicin, mercaptopurine, prednisolone, and asparaginase) and a group of targeted drugs (ibrutinib, dasatanib, and ruxolitinib). Viable cells are selected, and single-cell DNA sequencing is performed on at least 2,500 cells per drug exposure. Finally, bone marrow samples from the same patients after completing 6 weeks of treatment are classified for viable, residual pre-leukemia and leukemia using established protocols for bulk sequencing studies. Next, PTA is used to perform single-cell DNA sequencing of tens of thousands of cells in a scalable, efficient, and cost-effective manner, thereby achieving the following goals:

[0098] From clonal types to drug susceptibility catalogues Once sequencing data is obtained, the clonal type of each cell is established. To achieve this, variants are called and the clonal type is determined. Utilizing PTA limits allelic dropout and coverage bias introduced during currently used WGA methods. A systematic comparison of tools for calling variants from single cells undergoing MDA was performed, and a recently developed tool, Monovar, was found to have the highest sensitivity and accuracy (Zafar et al., Nature Methods, 2016, 13:505-507). Once variant calling is performed, it is determined whether two cells have the same clonal type, even if some variant calls are missing due to allelic dropout. To achieve this, a multivariate Bernoulli mixture model can be used (Gawad et al., Proc. Natl. Acad. Sci. USA, 2014, 111(50):17947-52). After establishing that the cells have the same clonal type, it is then determined which variants to include in the catalog. Genes meeting any of the following criteria were included: 1) they are nonsynonymous or loss-of-function variants (frameshift, nonsense, or splicing) detected in any of the mutational hotspots present in known tumor suppressor genes identified in large-scale pediatric cancer genome sequencing projects; 2) they are variants repeatedly detected in recurrent cancer samples; and 3) because ALL patients receive 6 weeks of treatment, they are recurrent variants that undergo positive selection in the current bulk sequencing study of residual disease. Clones that do not have at least two variants that meet these criteria are not included in the catalog. As more genes associated with treatment resistance or disease recurrence are identified, clones may be "rescued" and included in the catalog. To determine whether a clonotype underwent positive or negative selection between control and drug treatment, Fisher's exact test was used to identify clones that significantly differed from controls.Clones are added to the catalog only if at least two matching combinations of mutations are shown to have the same correlation with exposure to a particular drug. Known activating mutations in oncogenes or loss-of-function mutations in tumor suppressors of the same gene are considered equivalent between clones. If clonotypes are not exactly matched, the common mutations are entered into the catalog. For example, if clonotype 1 is A+B+C and clonotype 2 is B+C+D, the B+C clonotype is entered into the catalog. If recurrently mutated genes in resistant cells with a limited number of co-occurring mutations are identified, the clones may be collapsed into functionally equivalent clonotypes.

[0099] Example 6. Measuring the rate and location of CRISPR off-target activity in single human cells Utilizing the improved variant calling sensitivity and accuracy of PTA in single cells, carry out quantitative measurement of CRISPR-mediated genome editing using specific guide RNA with high sensitivity in single cells.Single cells are subjected to the general PTA method of Example 4.Comparing cellular indel and SV counts for both non-edited cells and edited cells (Figure 13A and Figure 13B).

[0100] The types of structural variation these genome editing methods can induce in single human cells were also investigated, and the results are shown in Figures 14A-14C. As shown in Figure 14A, the target region is indicated at the bottom (a) and is found on chromosome 6 between positions 43,770,818 and 43,770,841 (b). Sequencing data in the form of paired-end reads (small horizontal bars without dashes) indicate a match between the single-cell sequencing data and the target genome (c). Dashes within the reads indicate a deletion in the genome relative to the reference genome (d). In this example, both edited cells show a deletion (d) that overlaps with the target site (a). In contrast, the two unedited cells contain reads indicating a match to the reference genome at this location, and therefore no editing occurs. Figure 14B shows the detection of a large (>1KB) deletion resulting from CRISPR-induced editing that is restricted to edited cell #1. The target region is shown at the bottom (a) and is found between positions 23,779,588 and 23,779,611 on chromosome 18 (b). Sequencing data in the form of reads (small colored horizontal bars, usually gray) indicates concordance between the single-cell sequencing data and the target genome (c). Regions with abrupt drops in aligned reads indicate deviations from the reference genome at these locations. In this case, the sudden loss of read coverage between positions 23,778,472 and 23,779,607 on chromosome 18 indicates a large deletion in edited cell #1 (d). The breakpoint at the far right of the diagram overlaps with a region of the genome that is highly similar to the target site (a), and because the deletion is absent in unedited cells, this deletion is identified as a CRISPR-mediated deletion. Lowercase letters in (a) indicate bases that differ from the target site. Figure 14C shows the detection of an interchromosomal translocation between position 241,275,213 on chromosome 2 and position 38,536,06 on chromosome 4 in edited cell #1. The translocation breakpoint resembles the gRNA target site and overlaps with the gRNA off-target region of each chromosome shown at the bottom [(a) and (b)]. The left panel represents reads aligned to the chromosome 2 region containing the breakpoint, and the right panel represents reads aligned to the chromosome 4 region containing the breakpoint.Edited cell #1 is split into two views: (c) a view in which all reads align to the region surrounding the breakpoint, and (d) a view of the same region showing only the read pairs that are evidence of a translocation. For read pairs that support a translocation, one read of the pair aligns to chromosome 2 with a sharp drop in coverage at the breakpoint, while the other read aligns to chromosome 4, also with a sharp drop in read coverage at the breakpoint (e). This translocation is identified as a CRISPR-induced translocation because at least one of the translocation breakpoints overlaps a region of the genome that is highly similar to the target sites (two in this case: a and b) in edited cells, and there is no evidence of a translocation in unedited cells. Lowercase letters in (a) and (b) indicate bases that differ from the target site.

[0101] To confirm putative off-target sites and to assess the accuracy of variant calling with increasing numbers of guide RNA genomic mismatches, we also performed microfluidic high-throughput PCR-based resequencing of putative off-target sites in all cells (data not shown).

[0102] Example 7: Age Estimation Data is collected for a population of at least 1,000 subjects, including geographic location (where they spent the most time), gender, age, ethnicity, and the frequency and location of genomic mutations established using the PTA method. Samples are run in duplicate and obtained from one or more tissues for each subject. A standard curve is generated by correlating variables such as geographic location (where they spent the most time), gender, age, ethnicity, mutation frequency, mutation location, or other data obtained, with the subject's age. Genomes from samples of subjects of unknown age are sequenced using the PTA method, and the standard curve is used to determine the individual's age. If additional information about the subject (ethnicity, geographic location) is known, this can be used to further refine the prediction.

[0103] Example 8: Identification and diagnosis of clinical bacterial samples. Cell samples from subjects suspected of having a bacterial infection are obtained and subjected to single-cell genome sequencing using PTA. Mutations identified by PTA are compared with known antibiotic resistance mutations or used to identify the bacterial strain. This information is used to select appropriate treatments, such as effective antibiotics.

[0104] Example 9: Identification of microbial species and genes Water samples are collected from various sources, such as deep-sea vents, oceans, mines, streams, lakes, meteorites, glaciers, or volcanoes. The samples are passed through a 20-micron prefilter to remove particles and then fractionated into size groups such as 3-20 microns, 0.8-3 microns, 0.1-0.8 microns, and 50 kDa-0.1 microns. The samples are then processed to isolate individual cells or, optionally, processed in bulk. Genomic, plasmid, or other DNA is isolated using standard techniques, subjected to PCR, and then sequenced. After reassembly of the genome sequence, known species are identified, and unknown species and / or genes are characterized for potential industrial applications.

[0105] Example 10. Measuring the unintended insertion rate of gene therapy approaches Utilizing the improved variant calling sensitivity and accuracy of single-cell PTA, we quantitatively measure the unintended insertion rate of gene therapy approaches with high sensitivity in single cells. This method can detect the insertion of specific sequences into undesired locations by detecting surrounding sequences to determine whether the gene therapy approach causes insertion or modification of the host genome. Nucleic acids encoding proteins are introduced into viral carrier vectors and then delivered to one or more cells in an organism or in vitro. The virus delivers the nucleic acid to the nucleus, where it is transcribed into mRNA. After translation of the mRNA, the protein is produced. Cells modified by this gene therapy are sequenced using the general PTA method described in Example 4 to detect mutations (mutation frequency and location / pattern) caused by the gene therapy approach.

[0106] Example 11. Calling CNVs using PTA in primary cancer cells Further validation studies of the PTA protocol for SNV and copy number variation (CNV) calling were conducted using primary leukemia cells, following the general method described in Example 1, compared with MDA and recently developed or improved commercial kits. The PTA protocol demonstrated further increases in coverage width and remained the most consistent method based on CV calculations at base-pair resolution (Figure 19). PTA also remained the most sensitive method for SNV calling at all sequencing depths, and switching to low-temperature lysis resulted in the highest specificity for SNV calling. PCR-dependent methods (WGA Kit 3, PicoPlex Gold) also continued to show a decrease in specificity with increasing sequencing depth, but the specificity decrease was significantly improved over MALBAC and previous versions of PicoPlex.

[0107] To estimate the accuracy of calling CNVs of different sizes for each method, each bam file was sampled to 300 million reads, and the CV was measured at increasing bin sizes (Figure 5J). PTA was found to have the lowest CV compared to all other WGA methods across all bins (Figure 5J). WGA Kit 2 and PicoPlex Gold showed a sharp decline in CV values ​​with increasing depth. This particular leukemia sample had no known CNVs on 5q or 11q. As expected, a single-copy X chromosome was detected in all bulk samples and single cells. CNV analysis revealed that the 5q deletion was clonal, while the 11q alteration was only seen in a subset of cells (Figure 5K, shaded arrow). Bulk data suggested a possible deletion on 12p, but it was not called in the bulk sample. Two out of five single cells were found to have a CNV at the same location, suggesting that single-cell CNV profiling is more sensitive and a better strategy for estimating the percentage of cells in a tissue with a given copy number alteration.

[0108] Example 12. Determination of SNV rates in related cells. Kinship cell studies were performed by plating single CD34+ CB cells into a single well and subsequently expanding them for 5 days (Figure 16A). Next, single cells were reisolated from the culture to compare variant calling of nearly genetically identical cells. Furthermore, we used the bulk as a reference to distinguish between germline, false positive, and somatic variant calls (Figure 16B). Using this approach, and also using the bulk sample as ground truth, we determined that variant calling accuracy increased to 99.9% using a cryoprotocol with GATK4 genotyping (Figure 16C). Furthermore, most of these primary cells had similar or improved variant detection sensitivity. However, there was one cell with significantly lower variant calling sensitivity, which, without being bound by theory, may be the result of manually manipulating fragile primary cells. Furthermore, the two cells with higher variant calling sensitivity had fewer homozygous somatic variant calls, which may be the result of reduced allele dropout (Figure 15B). These false-positive variants are skewed to lower allele frequencies. Without being bound by theory, this can be explained by the fact that these rapidly dividing cells are tetraploid in the late S or G2 / M phase of the cell cycle, where only one of the four alleles acquires a copy error (Figures 17A-17C). Homozygous false-positive calls were observed to cluster at specific locations, whereas heterozygous calls were not. Without being bound by theory, this may be the result of loss or absence of template that degenerates at one allele at those locations during amplification, which does not appear to depend on the GC content of the genomic region (Figures 18A-18C). Most false-positive and somatic variants were called heterozygous, which is consistent with a model in which only one allele mutates, either as a result of a copy error or during development (Figure 16D). False-positive and somatic mutation rates were measured in neonatal CD34+ hematopoietic cells and were estimated to be 0.9 and 1.4 per Mb of genome, respectively.

[0109] Example 13: Measuring the rate and location of CRISPR off-target activity in single human cells The continued development of genome-editing tools shows great promise for improving human health, from correcting genes that cause or contribute to disease formation to eradicating currently incurable infectious diseases. However, the safety of these interventions remains unclear as a result of our incomplete understanding of how these tools interact with and permanently alter other locations within the genome of edited cells. While methods have been developed to estimate the off-target rate of genome-editing strategies, all of the tools developed to date investigate groups of cells together, making it impossible to measure the off-target rate per cell and the variation in off-target activity between cells, as well as to detect rare editing events that occur in small numbers of cells. Single-cell cloning of edited cells has been performed, but while it can select against cells that acquire lethal off-target editing events, it is impractical for many types of primary cells.

[0110] Taking advantage of the improved variant calling sensitivity and specificity of PTA, we obtained quantitative measurements of CRISPR-mediated genome editing using specific guide RNAs (gRNAs) in single cells (Figure 20A). These studies utilized three cell types: the U20S osteosarcoma cell line, primary hematopoietic CD34+ CB cells, and embryonic stem (ES) cells. Additionally, we used two previously described gRNAs, one known to be accurate (EMX1) and the other known to have high levels of off-target activity (VEGFA). To identify indels with high specificity, variant calling was restricted to genomic locations with perfect agreement with PAM sequencing and up to five mismatches with the protospacer (Figure 16A).

[0111] Compared to control cells that received either Cas9 alone or mock transfection, VEGFA-edited cells contained many more off-target indels, demonstrating wide cell-to-cell variability, whereas only a small number of off-target EMX1 editing events were detected (Figure 20B). It was noted that most of the putative false-positive edits seen in control cells were single base pair insertions. Removal of non-recurrent single base pair insertions further improved the specificity of indel calling (Figure 21). Most, but not all, recurrent off-target sites were cell type-specific, further supporting the finding that the general chromatin structure of a cell type influences off-target genomic locations (Figure 20D). Structural variant (SV) calling was performed to identify SVs induced by genome editing, requiring the regions around both breakpoints to perfectly match the PAM sequence and tolerate up to five mismatches with the protospacer. The increased number of SVs was measured using VEGFA guide RNA; only one SV was detected in EMX1-edited cells, whereas no SVs were detected in control cells (Figure 20E). Recurrent VEGFA-mediated SVs were detected, some of which were cell type-specific, with larger SVs detected in ES cells (Fig. 20C).

[0112] Example 14: Bacterial genome assembly using PTA Buccal swabs were obtained and cultured overnight in LB medium. Single bacterial colonies were sorted into 96-well plates as individual samples, and the general PTA method described in Example 1 was performed on each well to prepare each sample for sequencing. Between 1 and 1 million reads were obtained per sample, and the reads were assembled using SPAdes (a contig-based approach). Data for the longest contigs of 10 different bacterial samples are shown in Figure 22A. In silico analysis of the sequencing data, contigs from each sample were added sequentially in descending order of length (Figure 22B). Data for bacterial sample 10 are shown in Figure 22C. Next, the percentage of the total assembly assigned to each genus was determined. Contaminating sequences, accompanied by small fragments of genomic DNA, are present. These can be identified as small contigs (>5 KB, Figure 22D) in the dataset. Read pairs were considered human if both reads aligned to GRCh38 in the joint GRCh38-contig reference (Figures 22E-22F). Alternatively, we used an assembly-free approach for all samples (e.g., Kraken) by assigning reads to taxa using k-mers from the reference database. Results from the read-based approach for bacterial sample 10 are shown in Figure 22G1 and were consistent with the contig-based approach.

[0113] Example 15: Preimplantation genetic testing using PTA Non-invasive preimplantation genetic screening (NIPGS) is performed by preparing 20 cultured embryos (frozen or fresh) according to the general method described in Kuznyetsov et al. (2018) PLoS ONE, 13(5):e0197262. Briefly, each embryo is transferred to fresh Global HP medium containing HSA on day 4 of culture and cultured in oil until it reaches the blastocyst stage (day 5 or 6). Upon reaching a fully expanded blastocyst stage, each blastocyst undergoes laser-assisted trophectoderm biopsy, followed by laser disruption, allowing the BF to mix with the BCCM. The embryos are then transferred to cryopreservation medium and frozen by vitrification. After embryo removal, combined BCCM and BF samples are collected and frozen at -80°C until testing. After nucleic acid extraction from the BCCM / BF samples, the nucleic acids are subjected to the general PTA method described in Example 1. The resulting genomic DNA library generated from PTA is then analyzed for genetic mutations, such as chromosomal abnormalities.

[0114] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. 1. A method for determining mutations, said method comprising: a. exposing a population of cells to a gene editing method, wherein said gene editing method utilizes a reagent configured to introduce a mutation in a target sequence; b. isolating a single cell from said population; c. Providing a cell lysate from a single cell; d. contacting the cell lysate with at least one amplification primer, at least one nucleic acid polymerase comprising 3'-5' exonuclease activity, and a mixture of nucleotides, wherein the mixture of nucleotides comprises at least one terminator nucleotide that terminates nucleic acid replication by the polymerase, and wherein the at least one terminator nucleotide is an alpha-thiodideoxynucleotide; e. amplifying the target nucleic acid molecule to generate a plurality of terminated amplification products, wherein replication proceeds by strand displacement replication; f. ligating the molecules obtained in step (e) to adapters, thereby generating a library of amplification products; g. sequencing the library of amplification products, and h. Comparing the sequence of the amplified product to at least one reference sequence to identify at least one mutation. A method comprising:

2. 10. The method of claim 1, wherein the gene editing method comprises the use of CRISPR, TALEN, ZFN, recombinase, meganuclease, or viral integration.

3. 2. The method of claim 1, wherein the at least one mutation comprises an insertion, deletion, or substitution.

4. 1. A method for identifying a specificity determining sequence, said method comprising: a. providing a library of nucleic acids, wherein at least some of the nucleic acids comprise a specificity-determining sequence; b. performing a gene editing method on at least one cell, wherein the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; c. Sequencing the genome of said at least one cell using the method of claim 1, wherein a specificity-determining sequence that contacted said at least one cell is identified; and d. identifying at least one specificity-determining sequence that provides the fewest off-target mutations; A method comprising:

5. 5. The method of claim 4, wherein the off-target mutation is a synonymous or non-synonymous mutation.

6. 5. The method of claim 4, wherein the off-target mutation is located outside a gene coding region.

7. 1. A method for in vivo mutation analysis, said method comprising: a. performing a gene editing method on at least one cell in an organism, wherein the organism is a non-human subject, and the gene editing method comprises contacting the cell with a reagent comprising at least one specificity-determining sequence; b. isolating at least one cell from said organism; c. Sequencing the genome of said at least one cell using the method of claim 1. A method comprising:

8. The method of claim 7 , wherein the method comprises at least two cells.

9. 9. The method of claim 8, further comprising identifying mutations by comparing the genome of the first cell with the genome of the second cell.

10. 1. A method for predicting the age of a subject, said method comprising: a. providing at least one sample from said subject, wherein said at least one sample comprises a genome; b. sequencing the genome using the method of claim 1 to identify mutations; c. Comparing the mutations obtained in step b to a standard reference curve, wherein the standard reference curve correlates the number and location of mutations with verified age; and d. predicting the age of the subject based on a comparison of the variation to the standard reference curve A method comprising:

11. 11. The method of claim 10, wherein the subject is under 15 years of age.

12. 11. The method of claim 10, wherein the at least one sample is more than 1000 years old.

13. 1. A method for sequencing a microbial or viral genome, said method comprising: a. obtaining a sample containing one or more genomes or genome fragments; b. sequencing the sample using the method of claim 1 to obtain a plurality of sequencing reads; and c. Assembling and sorting the sequencing reads to generate the microbial or viral genome. A method comprising:

14. 14. The method of claim 13, wherein the sample comprises genomes from at least 10 organisms.

15. 14. The method of claim 13, wherein the microbial genome corresponds to an unculturable organism.

Citation Information

Patent Citations

  • dna amplification and sequencing using dna molecules generated by random fragmentation

    JP2005535283A

  • Single-cell analysis

    JP2022543051A

  • Gene mutation analysis

    JP2022543375A

  • Methods of Amplifying Whole Genome of a Single Cell

    US20140200146A1

  • Methods of assessing nuclease cleavage

    WO2018129368A2