Compositions and methods for characterizing alterations in polynucleotide sequences

JP2024516637A5Inactive Publication Date: 2025-05-07THE BRIGHAM & WOMEN S HOSPITAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023565493
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-04-26
Filing Date
2022-04-25
Publication Date
2025-05-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current technologies lack the ability to simultaneously analyze genome and transcriptome at the single-cell level, particularly in primary human cells, and fail to effectively characterize the results and phenotypes of CRISPR editing.

Method used

A method involving labeling cells with detectable antibodies, index sorting, and using unique molecular identifiers (UMIs) to characterize genomic DNA and mRNA, combined with oligoconjugated antibodies and reverse transcriptase to generate cDNA libraries, allowing for simultaneous amplification and sequencing of genomic, cDNA, and antibody-derived tag (ADT) libraries.

Benefits of technology

Enables robust and scalable analysis of genomic DNA and mRNA changes in CRISPR-edited cells, correlating DNA modifications with mRNA and protein expression, suitable for primary human cells, and providing insights into functional consequences of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
  • Figure 00000000_0001_ABST
    Figure 00000000_0001_ABST
Patent Text Reader

Abstract

The present disclosure provides compositions and methods for characterizing genomes and transcriptomes at single cell level.In some embodiments, the methods provide characterization of the results and phenotypes of CRISPR editing, and other changes in polynucleotide sequence, particularly in primary human cells. TIFF2024516637000003.tif75170
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 179,921, filed April 26, 2021, the entire contents of which are incorporated herein by reference.

[0002] STATEMENT OF RIGHTS TO INVETIONS MADE UNDER FEDERALLY SPONSORED RESEARCH This invention was made with Government support under Grant No. AR063759 awarded by the National Institutes of Health. The United States Government has certain rights in the invention. [Background technology]

[0003] 2. Background of the Invention Simultaneous sequencing of genomes and transcriptomes at the single-cell level is a powerful tool to characterize and correlate genome and transcriptome variations. However, analyzing both genomes and transcriptomes in the same cell remains technically challenging. Currently, there is a lack of technologies that allow for the analysis of CRISPR editing outcomes and phenotypes, especially in primary human cells. Summary of the Invention

[0004] As described below, the present disclosure features compositions and methods for characterizing genomes and transcriptomes at single cell level.In some aspects, the method provides for characterizing the results and phenotypes of CRISPR editing, for example, by using antibodies for sequencing and hashing from flow cytometry.Similar methods are provided for characterizing other changes in polynucleotide sequence.

[0005] In one aspect, the disclosed invention features a method for simultaneously characterizing genomic DNA and mRNA of a single cell. The method includes (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest. The method also includes (b) incubating the detectably labeled cells of step (a) with an oligo-conjugated antibody. The method further includes (c) index sorting the cells into a single well, characterizing the expression of the cell surface marker for each cell, and lysing the cells in the presence of dNTPs and a well-specific barcoded oligo-DT primer that includes a unique molecular identifier (UMI) and a PCR handle. The method also includes (d) incubating the product of step (c) with a reverse transcriptase, a custom template switch oligo (TSO) that includes one member of a binding pair, under conditions that allow for the generation of cDNA. The method further comprises (e) incubating the product of step (d) with a genomic primer that specifically binds to a region of interest (ROI), a cDNA amplification primer that specifically binds to a PCR handle and a cDNA amplification primer that specifically binds to a TSO, an antibody derived tag (ADT) specific primer, dNTPs, and a polymerase under conditions that support amplification, thereby simultaneously amplifying gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library. The method further comprises (f) incubating at least a portion of the genomic DNA from each well of step (e) with dNTPs, a polymerase, and nested primers that specifically bind to a region of interest to obtain a gDNA library. At least one of the nested primers comprises i) a well-specific barcode, a UMI, and a PCR handle; or ii) a capture sequence. If the nested primer comprises a capture sequence, step (e) further comprises incubating the product of step (e) with an exonuclease and a capture oligo.The capture oligo comprises a capture sequence, a well-specific barcode, an exonuclease blocking agent, and a UMI. The capture oligo binds to the amplicons generated using the nested primers, effectively labeling the products with barcodes during the PCR reaction. The method also comprises (g) pooling at least a portion of the samples from each well after step (e) or step (f), followed by separating at least two of the cDNA library, the ADT library, and the gDNA library. The method also comprises (h) preparing the gDNA library, the cDNA library, and the ADT library for sequencing by amplifying each library in the presence of a sequencing primer.

[0006] In another aspect, the present invention features a method for simultaneously characterizing DNA amplicons, 3'mRNA transcripts, antibody-derived tags (ADTs), and index flow sorting information from a cell sample. The method includes (a) labeling a plurality of cells with a detectable antibody that specifically binds to a cell surface marker of interest, and single-cell index sorting the cells into individual wells. The method also includes (b) lysing the cells in the presence of a reverse transcriptase, a template switch oligo, a well-specific barcode, a primer comprising an oligoDT primer that comprises a unique molecular identifier (UMI) and a PCR handle, and an ADT under conditions that allow reverse transcription to obtain cDNA. The method further includes (c) amplifying the cDNA, ADT, and specific genomic DNA in a single pool containing a genomic primer that specifically binds to the region of interest, a cDNA amplification primer that specifically binds to the PCR handle, and a cDNA amplification primer that specifically binds to the TSO, an ADT-specific primer, dNTPs, and Taq polymerase, thereby simultaneously amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library. The method further includes (d) using at least a portion of the product of step (c) to further amplify the genomic ROI using nested primers to obtain a gDNA library. At least one of the nested primers includes i) a well-specific barcode, a UMI, and a PCR handle; or ii) a capture sequence. If the nested primer includes a capture sequence, step (d) further includes incubating the product of step (c) with an exonuclease and a capture oligo. The capture oligo includes a capture sequence, a well-specific barcode, an exonuclease blocking factor, and a UMI. The capture oligos bind to amplicons generated with the nested primers, effectively labeling the products with barcodes during the PCR reaction.The method further includes (e) pooling at least a portion of each well, followed by separating at least two of the gDNA library, the cDNA library, and the ADT library.The method also includes (f) preparing a gDNA library, a cDNA library, and an ADT library for sequencing. The step of preparing the libraries for sequencing includes amplifying the ADT library with a sequencing primer, tagmenting the cDNA library and preferentially amplifying the 3' end with a sequencing primer, and amplifying the gDNA library with a sequencing primer.

[0007] In another aspect, the disclosed invention features a method for simultaneously characterizing genomic DNA and mRNA of a single cell. The method includes (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest. The method also includes (b) incubating the detectably labeled cells of step (a) with an oligo-conjugated antibody. The method further includes (c) index sorting the cells into a single well, characterizing the expression of cell surface markers for each cell, and lysing the cells in the presence of dNTPs, a well-specific barcoded oligo-DT primer that includes a unique molecular identifier (UMI) and a PCR handle, and a capture oligo that includes a capture sequence, a well-specific barcode, an exonuclease blocking agent, and a unique molecular identifier. The method further includes (d) incubating the product of step (c) with a reverse transcriptase, a custom template switch oligo (TSO) that includes one member of a binding pair, and a reverse transcriptase under conditions that allow for the generation of cDNA. The method also includes (e) incubating the products of step (d) with a genomic primer that specifically binds to the region of interest (ROI), a cDNA amplification primer that specifically binds to the PCR handle and a cDNA amplification primer that specifically binds to the TSO, an antibody-derived tag (ADT)-specific primer, dNTPs, and a polymerase under conditions that support amplification, thereby simultaneously amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library. The method also includes (f) contacting the products of step (e) with an exonuclease to degrade unconsumed primers. The method further includes (g) incubating at least a portion of the genomic ROI library from each well of step (f) with dNTPs, a polymerase, and nested primers that can specifically amplify a region within the genomic ROI library. At least one of the nested primers includes a capture sequence. The capture oligo binds to the amplicons generated using the nested primers, effectively labeling the products with barcodes during the PCR reaction, resulting in a gDNA library.The method also includes (g) pooling at least a portion of the samples from each well, followed by separating the gDNA, cDNA, and ADT libraries. The method further includes (h) preparing the gDNA, cDNA, and ADT libraries for sequencing by amplifying each library in the presence of a sequencing primer.

[0008] In another aspect, the disclosed invention provides a method for simultaneously characterizing genomic DNA and mRNA of a single cell. The method includes (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest. The method further includes (b) incubating the detectably labeled cells of step (a) with an oligo-conjugated antibody. The method also includes (c) index sorting the cells into a single well, characterizing the expression of cell surface markers for each cell, and lysing the cells in the presence of dNTPs and a well-specific barcoded oligoDT primer that includes a unique molecular identifier (UMI) and a PCR handle. The method further includes (d) incubating the product of step (c) with a reverse transcriptase and a custom template switch oligo (TSO) that includes one member of a binding pair under conditions that allow for the generation of cDNA. The method further includes (e) incubating the product of step (d) with a genomic primer that specifically binds to the region of interest (ROI), a cDNA amplification primer that specifically binds to the PCR handle and a cDNA amplification primer that specifically binds to the TSO, an antibody-derived tag (ADT)-specific primer, dNTPs, and a polymerase under conditions that support amplification, thereby simultaneously amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library. The method also includes (f) pooling at least a portion of the samples from each well, followed by separating the cDNA library and the ADT library. The method also includes (g) incubating at least a portion of the genomic DNA from each well of step (e) with dNTPs, a polymerase, and a nested primer that specifically binds to the region of interest to obtain a gDNA library. The nested primer includes a well-specific barcode, a UMI, and a PCR handle. The method includes (h) preparing the gDNA library, the cDNA library, and the ADT library for sequencing by amplifying each library in the presence of a sequencing primer.

[0009] In any of the above aspects or embodiments thereof, the method further comprises sequencing the library.

[0010] In any of the above aspects or embodiments thereof, the method further comprises adding a capture oligo prior to the initial amplification of the gDNA, cDNA, and ADT.

[0011] In any of the above aspects or embodiments thereof, the exonuclease is ExoI. In any of the above aspects or embodiments thereof, the blocking agent is a phosphoryl group or an acetyl group. In any of the above aspects or embodiments thereof, the blocking agent is linked to the 3'OH group of the capture oligomer.

[0012] In any of the above aspects or embodiments thereof, all amplifications before preparing gDNA library, cDNA library, and ADT library are performed in the same well.In any of the above aspects or embodiments thereof, the formation of cDNA library, genomic ROI library, and ADT library is performed in a first well, and gDNA library is prepared in another well.

[0013] In any of the above aspects or embodiments thereof, the gDNA library, the cDNA library, and / or the ADT library are separated using Solid Phase Reversible Immobilization (SPRI) beads.

[0014] In any of the above aspects or embodiments, the separation comprises first separating the gDNA library from the cDNA library and the ADT library using SPRI beads, and then separating the cDNA library from the ADT library using SPRI beads.In any of the above aspects or embodiments, the separation of the cDNA library from the ADT library comprises separating the amplicons that are greater than 500 bp in length and the amplicons that are less than 500 bp in length from each other.In any of the above aspects or embodiments, the separation of the cDNA library and the ADT library is carried out before or in parallel with the preparation of the gDNA library.

[0015] In any of the above aspects or embodiments, one or more of the cells comprise changes in genomic DNA sequence compared to the sequence of reference genome.In some embodiments, the changes are introduced by using genome editing technology.In some embodiments, the genome editing technology comprises base-editing (BE) or homology-directed recombination (HDR) editing.

[0016] In any of the above aspects or embodiments thereof, one or more of the cells comprises an alteration in mRNA expression relative to the mRNA expression of a reference cell.In any of the above aspects or embodiments thereof, one or more of the cells comprises an alteration in expression of a cell surface marker relative to the reference cell.

[0017] In any of the above aspects or embodiments, the cell is edited with CRISPR before characterization.In any of the above aspects or embodiments, the cell is a primary cell.In any of the above aspects or embodiments, the cell is an immune cell.In any of the above aspects or embodiments, the cell is a mammalian cell.In any of the above aspects or embodiments, the cell is a human cell.

[0018] In any of the above aspects or embodiments thereof, the cells are sorted using a FACS sorter.In any of the above aspects or embodiments thereof, at least about 500,000 to more than 10 million cells are characterized.In any of the above aspects or embodiments thereof, the cell surface marker is CD45, CD81, or MHC class 1.

[0019] In any of the above aspects or embodiments thereof, the polymerase is Taq polymerase. In some embodiments, the Taq polymerase is KAPA HiFI Taq polymerase or Q5 Taq polymerase.

[0020] In any of the above aspects or embodiments thereof, after the incubation or amplification step, the products of the incubation or amplification are cleaned. In one embodiment, the cleaning is performed using solid phase reversible immobilization (SPRI) beads.

[0021] In any of the above aspects or embodiments thereof, the detectable antibody comprises a fluorophore. In any of the above aspects or embodiments thereof, the oligoconjugated antibody comprises a polyA sequence.

[0022] In any of the above aspects or embodiments thereof, the sequencing primers are Illumina primers P5 and P7.

[0023] In any of the above aspects or embodiments thereof, steps (c) to (e) occur simultaneously or sequentially.

[0024] The compositions and articles defined in the present invention have been isolated or prepared in conjunction with the examples provided below. Other features and advantages of the invention will become apparent from the detailed description and claims.

[0025] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in the present invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). The following terms used herein have the meanings assigned below, unless otherwise specified.

[0026] The term "adapter" refers to a sequence that is added to a nucleic acid, for example, by ligation. The adapter may be about 5 to about 100 bases in length and may provide a primer binding site for sequencing (e.g., an amplification primer binding site) and a molecular barcode, such as a sample identifier sequence or a molecular identifier sequence, preferably a unique identifier sequence. An adapter may be added to 1) the 5' end, 2) the 3' end, or 3) both ends of a nucleic acid molecule. A double-stranded adapter comprises a double-stranded end that is linked to the nucleic acid. The adapter may have an overhang and may be blunt-ended. As described in more detail below, a double-stranded adapter may be added to a fragment by ligating only one strand of the adapter to the fragment. The sequence of the unligated strand of the adapter may be added to the fragment using a polymerase. Types of double-stranded adapters include Y-shaped adapters and looped adapters.

[0027] "Alteration" refers to a change (increase or decrease) in the structure, expression level or activity of a gene or polypeptide, as detected by standard methods known in the art, as described herein. In one embodiment, sequence alterations (i.e., insertions, deletions, point mutations, copy number alterations (CNAs), or loss of heterozygosity (LOH)) are determined relative to a reference sequence, reference exome, and / or reference genome. In some embodiments, the alteration is a change in the sequence of a polynucleotide, e.g., an alteration associated with CRISPR editing. As used herein, alteration includes a 10% change in expression level, preferably a 25% change, more preferably a 40% change, and most preferably a 50% or more change in expression level.

[0028] By "amplicon" is meant a fragment of a nucleic acid, e.g., DNA or RNA, that is the source and / or product of amplification or replication.

[0029] The term "antisense strand" as used herein refers to a polynucleotide that is substantially or 100% complementary to the target nucleic acid of interest.For example, antisense strand can be fully or partially complementary to mRNA (messenger RNA) molecule, non-mRNA RNA sequence (e.g., microRNA, piwiRNA, tRNA, rRNA, hnRNA), or coding or non-coding DNA sequence.The terms "antisense strand" and "guide strand" are used interchangeably herein.

[0030] As used herein, "biological sample" refers to a sample obtained from a biological subject, such as a sample from a living tissue or body fluid obtained, achieved, or collected in vivo or in situ, that contains or is suspected to contain a polynucleotide. Biological samples also include samples from areas of a biological subject that contain immune cells, pre-cancerous or cancerous cells or tissues. Such samples can be, but are not limited to, organs, tissues, fractions, and cells isolated from mammals, including human patients, mice, and rats. Biological samples also include sections of biological samples such as tissues, such as frozen sections taken for histological purposes.

[0031] "Barcode" refers to a degenerate or semi-degenerate nucleic acid sequence that is different for each plasmid or for each genome. Barcode sequences can be distinct degenerate or semi-degenerate sequences. For example, barcodes can contain distinct degenerate sequences with several possible bases at any position of the nucleic acid sequence. Barcodes can uniquely label or detect a single cell. Barcodes may also be used in sequencing methods to identify genomes.

[0032] "Complementary" means capable of pairing to form a double-stranded nucleic acid molecule or a portion thereof. Complementarity need not be perfect and may contain mismatches at one, two, three or more nucleotides.

[0033] In this disclosure, "comprises," "comprising," "containing," "having," and the like can have the meanings given them in U.S. Patent Law and can mean "includes," "including," and the like; "consisting essentially of" or "consists essentially" likewise has the meaning given to it in U.S. Patent Law, and the term is open-ended, permitting the presence of more than what is recited, but excluding prior art aspects, so long as the basic or novel characteristics of what is recited are not altered by the presence of more than what is recited.

[0034] "Detecting" refers to determining the presence, absence, or amount of the analyte being detected. In some embodiments, the analyte is a sequence variation.

[0035] "Detectable label" refers to a composition that, when linked to a molecule of interest, allows the molecule to be detected through spectroscopic, photochemical, biochemical, immunochemical, or chemical means.For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (such as those commonly used in ELISA), biotin, digoxigenin, or haptens.

[0036] "Exonuclease" refers to an enzyme that cleaves a polynucleotide chain by removing nucleotides one by one from the end of the polynucleotide chain. In one embodiment, an exonuclease useful for selectively degrading linear DNA as opposed to circular DNA is RecBCD.

[0037] The term "expression" or "expressed" as used herein in relation to a gene refers to the transcription and / or translation product of that gene. The expression level of a DNA molecule in a cell can be determined based on either the amount of corresponding mRNA present in the cell or the amount of protein encoded by the DNA produced by the cell (Sambrook et al., 1989 Molecular Cloning: A Laboratory Manual, 18.1-18.88). Expression of a transfected gene can occur transiently or stably in a cell. In "transient expression", the transfected gene is not transferred to daughter cells during cell division. Since its expression is limited to the transfected cell, the expression of the gene is lost over time. In contrast, stable expression of a transfected gene can occur when the gene is co-transfected with another gene that confers a selective advantage to the transfected cell. Such a selective advantage can be resistance to a particular toxin exhibited by the cell.

[0038] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. The portion preferably comprises at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. A fragment can comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0039] The term "gene" refers to a segment of DNA involved in producing a protein; it includes the region preceding the coding region (leader) and following the coding region (trailer), as well as intervening sequences (introns) between individual coding segments (exons). Leaders, trailers, and introns contain the regulatory elements utilized during transcription and translation of a gene. Additionally, a "protein gene product" refers to the protein expressed from a particular gene.

[0040] By "genomic library" is meant the entire genome of an organism, virus, bacterium, plant, or cell, or a collection of cloned DNA molecules consisting of at least one copy of every gene from a particular organism or cell.

[0041] "High throughput sequencing" refers to a sequencing technology that allows for the sequencing of large amounts of nucleic acids.

[0042] "Hybridization" means hydrogen bonding between complementary nucleobases, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.

[0043] The terms "isolated," "purified," or "biologically pure" refer to a material that is free, to a greater or lesser extent, from components that normally accompany it as found in its natural state. "Isolated" means separated to some degree from the original source or surrounding environment. "Purified" means separated to a greater degree than isolated. A "purified" or "biologically pure" protein is one in which other materials have been sufficiently removed so that the impurities do not significantly affect the biological properties of the protein or cause other deleterious consequences. That is, a nucleic acid or peptide of the invention is purified if it is substantially free of cellular material, viral material, or culture medium if produced by recombinant DNA technology, or chemical precursors or other chemicals if chemically synthesized. Purity and homogeneity are usually confirmed using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. In the case of proteins that can be modified, such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins, which can be purified separately.

[0044] "Isolated polynucleotide" refers to a nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, the term includes recombinant DNA that is incorporated, for example, into a vector, an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryotic or eukaryotic organism; or exists as a separate molecule independent of other sequences (e.g., cDNA, or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). In addition, the term also includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene that codes for additional polypeptide sequences.

[0045] By "isolated polypeptide" is meant a polypeptide of the invention separated from components that naturally accompany it. Generally, a polypeptide is isolated when it is at least 60%, by weight, free from proteins and naturally occurring organic molecules with which it is naturally associated. Preferably, a preparation is at least 75%, more preferably at least 90%, and most preferably at least 99%, by weight, a polypeptide of the invention. An isolated polypeptide of the invention can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding the polypeptide, or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0046] By "marker" is meant any protein or polynucleotide that has an alteration in expression levels or activity that is associated with an alteration in the genome of a cell or that is associated with a disease or disorder.

[0047] As used herein, "obtaining," such as "obtaining a substance," includes synthesizing, purchasing, or otherwise acquiring the substance.

[0048] "Primer set" refers to a set of oligonucleotides that can be used, for example, in PCR. A primer set consists of at least 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, 80, 100, 200, 250, 300, 400, 500, 600, or more primers.

[0049] By "reduce" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0050] "Reference" means a standard or control condition.

[0051] A "reference genome" is a defined genome used as a basis for genome comparison or for alignment of sequencing reads to it. A reference genome can be a subset or the entirety of a specified genome, e.g., a subset of a genome sequence, such as an exome sequence, or a complete genome sequence.

[0052] "Reference sequence" is a defined sequence used as the basis for sequence comparison. Reference sequence can be a subset or the whole of a specified sequence, for example, a full-length cDNA or gene sequence, i.e., a segment of a complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, preferably at least about 20 amino acids, more preferably at least about 25 amino acids, even more preferably about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, preferably at least about 60 nucleotides, more preferably at least about 75 nucleotides, even more preferably about 100 nucleotides or about 300 nucleotides, or any integer number therebetween.

[0053] Nucleic acid molecules useful in the method of the present invention include any nucleic acid molecule that encodes a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules do not need to be 100% identical to an endogenous nucleic acid sequence, but will usually show substantial identity. A polynucleotide that has "substantial identity" to an endogenous sequence can usually hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the method of the present invention include any nucleic acid molecule that encodes a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules do not need to be 100% identical to an endogenous nucleic acid sequence, but will usually show substantial identity. A polynucleotide that has "substantial identity" to an endogenous sequence can usually hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridize" means pairing to form a double-stranded molecule between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various conditions of stringency (see, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).

[0054] For example, stringent salt concentrations are usually less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvents, such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions usually include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters, such as hybridization time, concentration of detergent such as sodium dodecyl sulfate (SDS), and inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as necessary. In a preferred embodiment, hybridization is performed in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS at 30° C. In a more preferred embodiment, hybridization is performed in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37° C. In a most preferred embodiment, hybridization is performed in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, 200 μg / ml ssDNA at 42° C. Beneficial modifications to these conditions will be readily apparent to those of skill in the art.

[0055] In most applications, the washing steps following hybridization may also vary in stringency. Wash stringency conditions may be defined by salt concentration and temperature. As mentioned above, washing stringency can be increased by decreasing salt concentration or increasing temperature. For example, stringent salt concentrations for the washing steps are preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the washing steps usually include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In a preferred embodiment, the washing steps are performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing steps are performed at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing steps are performed at 68° C. in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further modifications to these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0056] "RNA-seq" refers to RNA sequencing to detect and quantify messenger RNA molecules (mRNA) in biological samples, e.g., used to study cellular responses. The related term "scRNA-seq" is single-cell RNA sequencing, which can be, for example, droplet-based single-cell RNA-seq or "Drop-seq," a sequencing technology for analyzing RNA expression in at least hundreds of thousands of individual cells in embodiments of the present invention, but alternatively other high-throughput sequencing platforms can be used.

[0057] By "substantially identical" is meant a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or a reference nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). Preferably, such a sequence is at least 60%, more preferably 80% or 85%, and even more preferably 90%, 95%, or 99% identical at the amino acid or nucleic acid level to the sequence used for comparison.

[0058] Sequence identity is typically measured using sequence analysis software (e.g., the Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program can be used, e.g.,-3 and e -100 Probability scores between 0 and 1 indicate closely related sequences.

[0059] By "subject" is meant a mammal, including, but not limited to, a human or a non-human mammal, such as a cow, a horse, a dog, a sheep, a cat, etc.

[0060] The term "tagmentation" refers to a step in the Assay for Transposase Accessible Chromatin using sequencing (ATAC-seq) as described. (See Buenrostro, JD, Giresi, PG, Zaba, LC, Chang, HY, Greenleaf, WJ, Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nature methods 2013; 10 (12): 1213-1218). Specifically, hyperactive Tn5 transposase loaded in vitro with adapters for high-throughput DNA sequencing can simultaneously fragment and tag the genome with adapters for sequencing. In one embodiment, such adapters are compatible with the methods described herein.

[0061] Single-cell ATAC-seq detects open chromatin in individual cells. ATAC-seq (assay for transposase-accessible chromatin) identifies regions of open chromatin using the highly active prokaryotic Tn5 transposase, which preferentially inserts into accessible chromatin and tags the sites with sequencing adapters (Buenrostro JD, Giresi PG, Zaba LC, Chang HY, Greenleaf WJ. Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nat Methods. 2013;10:1213-128). The protocol is simple, robust, and widely available. Until now, ATAC-seq and other methods for identifying open chromatin have required large pools of cells (Buenrostro, 2013; Thurman RE, Rynes E, Humbert R, Vierstra J, Maurano MT, Haugen E, et al. The accessible chromatin landscape of the human genome. Nature. 2012;488:75-82); this means that the data collected reflects the cumulative accessibility of all cells in the pool.Independent studies have refined the ATAC-seq protocol for single-cell applications (scATAC-seq) (Buenrostro JD, Wu B, Litzenburger UM, Ruff D, Gonzales ML, Snyder MP, et al. Single-cell chromatin accessibility reveals principles of regulatory variation. Nature. 2015;523:486-90; and Cusanovich DA, Daza R, Adey A, Pliner HA, Christiansen L, Gunderson KL, et al. Epigenetics. Multiplex single-cell profiling of chromatin accessibility by combinatorial cellular indexing. Science. 2015;348:910-4). These studies provide data on hundreds (Buenrostro, 2015) or thousands (Cusanovich, 2015) of single cells in parallel. Both methods are limited in the number of cells analyzed or in the per-cell coverage.

[0062] By "transcriptome" is meant all the messenger RNA (mRNA) molecules expressed from an organism's RNA genes.

[0063] "Unique molecular identifier" or "UMI" refers to a short nucleic acid sequence that is identifiable in high-throughput sequencing techniques, such as, but not limited to, single-cell RNA-seq. UMIs can be used for detection as well as quantification.

[0064] Ranges provided herein are understood to be shorthand for all values ​​within that range. For example, the range 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0065] As used herein, the term "or" is understood to be inclusive unless otherwise stated or clear from context. As used herein, the terms "a," "an," and "the" are understood to be singular or plural unless otherwise stated or clear from context.

[0066] Unless otherwise specified or clear from the context, the term "about" as used herein is understood to mean within the normal tolerance in the art, for example, within 2 standard deviations of the mean. About can be understood as within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise clear from the context, all numerical values ​​provided herein are modified with the term about.

[0067] The recitation of a listing of a chemical group in a definition of a variable herein includes definitions of that variable as a single group or as any combination of the listed groups. The recitation of an embodiment of a variable or aspect herein includes that embodiment as a single embodiment or in combination with other embodiments or portions thereof.

[0068] Any composition or method provided herein can be combined with any one or more of the other compositions and methods provided herein. [Brief description of the drawings]

[0069] [Figure 1A] Figures 1A-1G provide schematics, boxplots, plots, and heatmaps showing that single-cell multi-omics analysis of genomic DNA and mRNA from CRISPR-edited HH cells revealed a strong correlation between induced deletion size and HLADQB1 expression. Figure 1A provides a representative schematic of multi-omics single-cell editing of HH cells. Figure 1B provides a representative plot of exon and sgRNA locations in HLADQB1. Alignment of an exemplary amplicon from a single HH cell, created with CRISPResso2, to a reference sequence. Figure 1C provides a heatmap of single-cell DNA editing, with each row representing one cell and each column representing one nucleotide. Cells are shaded by the percentage of reads with modifications (insertion, deletion, or substitution) at that nucleotide, and colors represent different cell clusters. Cell groupings were identified by k-means clustering as defined. Figure 1D provides a boxplot of single-cell HLADQB1 gene expression per DNA cluster defined in Figure 1C. Figure 1E-1F provide plots showing the correlation between average deletion size and gene expression of HLADQB1 and HLADRB1. Correlations were calculated using a linear regression model and p-values ​​for gene coefficients are shown. Figure 1G provides a Manhattan plot of genome-wide differential gene (non-zero in 30% of cells) expression analysis performed using DESeq2 with deletion size as the response variable. The red line represents a Bonferroni corrected p-value of 0.01. Each dot represents a single cell. Gene expression values ​​have been scaled and normalized using Seurat. [Figure 1B] See legend to Figure 1A. [Figure 1C] See legend to Figure 1A. [Figure 1D] See legend to Figure 1A. [Figure 1E] See legend to Figure 1A. [Figure 1F] See legend to Figure 1A. [Figure 1G] See legend to Figure 1A. [Figure 2A-1]Figures 2A-2K provide schematics, flow cytometry plots, diagrams, box plots, volcano plots, and scatter plots showing that single-cell multi-omics sequencing of PTRPC-edited primary CD4 T cells revealed robust correlations between various genotypes and protein expression. Figure 2A provides a schematic of base editing a premature stop codon during PTRPC in primary CD4 T cells. Grey circles represent the bulk data collection endpoint. Black circles represent the single-cell collection endpoint. Figure 2B provides representative flow cytometry plots and summary analysis of a non-targeting sample (N) and a base-edited knockout sample (BE). Red lines indicate samples used for single-cell MINECRAFTseq processing. Significant differences were measured using paired Mann-Whitney U tests (** p<0.01). Figure 2C provides a diagram showing bulk DNA editing results from three healthy individual samples, with targeted nucleotides highlighted. Arrows indicate the position of the sgRNA away from the PAM. Figure 2D provides a box plot (left panel) showing bulk mRNA expression of PTPRC from four healthy individuals. Gene expression values ​​are scaled and normalized as logUMI+1. Figure 2D also provides a volcano plot (right panel) of differential gene expression, with each dot representing one tested gene. Solid lines represent Bonferroni-corrected p-values ​​of 0.01. Figure 2E provides a diagram. Ten plates from one healthy individual (light grey line in Figure 2B) were processed with the single-cell MINECRAFTseq protocol. Common genotypes (>4 cells) recovered. Rare genotypes (<4 cells) are not shown. Histograms and numbers on the right represent the number of cells from each genotype. Arrows indicate the sequence and position of the sgRNA, pointing away from the PAM site. Figure 2F provides a box plot of the corresponding expression of CD45-FITC measured by index flow cytometry from each genotype in Figure 2E , as well as a box plot of the biexponentially scaled or CLR normalized ADT counts of CD45 from each genotype in Figure 2E ( Figure 2G ).Figure 2H provides a plot showing the Uniform Manifold Approximation and Projection (UMAP) of all 33 measured ADT markers colored by CD45 expression. Figure 2I provides a boxplot showing all significant changes of ADT markers correlating with dosage at the target base (compare genotypes A, C, and B). All other genotypes were excluded from the analysis. CLR normalized counts for each marker. ADT markers are shown in order of mean expression. Figure 2J provides a plot showing the Uniform Manifold Approximation and Projection (UMAP) of variable gene expression from single cells, with colors representing the scaled and normalized expression of PTPRCs. Figure 2K provides a boxplot of gene expression of PTPRCs by genotype, with genotypes D-H grouped into one category. Each dot represents one cell unless otherwise specified. [Figure 2A-2] See description of Figure 2A-1. [Figure 2A-3] See description of Figure 2A-1. [Figure 2B] See description of Figure 2A-1. [Figure 2C] See description of Figure 2A-1. [Figure 2D] See description of Figure 2A-1. [Figure 2E] See description of Figure 2A-1. [Figure 2F] See description of Figure 2A-1. [Figure 2G] See description of Figure 2A-1. [Figure 2H] See description of Figure 2A-1. [Figure 2I] See description of Figure 2A-1. [Figure 2J] See description of Figure 2A-1. [Figure 2K] See description of Figure 2A-1. [Figure 3A]Figures 3A-3K provide schematics, diagrams, heatmaps, volcano plots, and boxplots showing genome editing of four variants in the UBASH3A autoimmune locus, where causal variants were identified by single-cell MINECRAFTseq. Three healthy individuals (all references for the four targeted variants) were recruited, CD4 T cells were isolated, and genome editing was performed with individual variants of interest. HDR or BE samples and controls were indexed and pooled to prepare multi-omics single-cell libraries, as shown in Figures 8A and 8B. In some embodiments, libraries are prepared as shown in Figures 19A and 19B. Figure 3A provides a schematic of the UBASH3A locus with variants of interest highlighted, along with the CRISPR-Cas editing technology used in the study. Recovered common genotypes (>4 cells) when targeting (Figure 3B) rs80054410 and (Figure 3C) rs11203203 are shown along with sgRNA sequence and cell counts. (Figure 3D) Heatmap of editing at the rs9981624 locus and (Figure 3E) rs11203202 locus with variant and sgRNA location indicated (arrows point away from the PAM). Each row represents one cell and each column represents one nucleotide in the amplicon sequence. Shading represents the percentage of reads edited (inserted, deleted, or substituted) at that nucleotide. DNA editing clusters were defined by k-means clustering assuming four clusters. Volcano plots of differential gene (>30% non-zero) expression versus gene dosage (0,1,2) at (Figure 3F) rs80054410 and (Figure 3G) rs11203203 considering plates. Each dot represents one gene. Figure 3H provides a plot of scaled and normalized gene expression of RIPK1 per genotype identified in Figure 3C. Volcano plot of differential gene (>30% non-zero) expression versus average deletion size in (Figure 3I) rs9981624 and (Figure 3J) rs11203202 considering plates. Each dot represents one gene.Figure 3K provides plots showing scaled and normalized gene expression of IL2RA per DNA cluster identified in Figure 3E. Differential gene expression was performed on unscaled and unnormalized values ​​using DESeq2. The solid line on the Volcano plot is the Bonferroni corrected p-value of 0.05. In H and K, each dot represents one cell. Scaled and normalized gene expression was calculated using Seurat. [Figure 3B] See legend to Figure 3A. [Figure 3C] See legend to Figure 3A. [Figure 3D] See legend to Figure 3A. [Figure 3E] See legend to Figure 3A. [Figure 3F] See legend to Figure 3A. [Figure 3G] See legend to Figure 3A. [Figure 3H] See legend to Figure 3A. [Figure 3I] See legend to Figure 3A. [Figure 3J] See legend to Figure 3A. [Figure 3K] See legend to Figure 3A. [Figure 4A]Figures 4A-4I provide schematics, diagrams, box plots, and volcano plots showing CRISPR-Cas base editing of three variants in IL2RA, confirming the causal role of rs61839660 and its neighboring nucleotides in regulating CD25 expression. Three healthy individuals with different genotypes at the three target variants were recruited, CD4 T cells were isolated, and genome editing was performed with individual variants of interest or as a large multiplex pool. All conditions were indexed and pooled to prepare multi-omics single cell libraries as shown in Figures 8A and 8B. In some embodiments, libraries are prepared as shown in Figures 19A and 19B. Sequences from each variant in each cell were generated. Figure 4A provides a schematic of the IL2RA locus with the variants of interest highlighted. Figure 4B provides a diagram showing the conditions used in this experiment with different CRISPR-Cas base editors. Figure 4C provides a diagram showing the highly common genotypes recovered (>20 cells for brevity), with sgRNA sequences and cell counts shown for all three target regions (right). Location of variants of interest along with the name "multiplex SNP". The derived single nucleotide polymorphism (SNP) IDs (SNPs 1-18) at the bottom, with names at the bottom, represent the location of the variants identified in this study and will be used in follow-up analyses. Full details on genotypes per individual and per condition can be found in Figures 18A and 18B. Figure 4D provides a box plot showing the CLR-normalized counts of gene dosage significant ADT markers at the identified multiplex SNPs. Figure 4E provides a box plot showing the CLR-normalized counts of gene dosage significant ADT markers at rs61839660 conditional on the multiplex SNPs. FIG. 4F provides a volcano plot of differential gene (>30% non-zero) expression versus gene dosage at rs61839660, corrected for gene dosage at plate and multiplex SNPs.Figure 4H provides a volcano plot showing scaled and normalized gene expression of RORA per gene dosage at rs61839660, faceted by individual, regardless of genotype. Volcano plot of differential gene (>30% non-zero) expression versus gene dosage at multiplex SNPs considering plate. Figure 4I provides a volcano plot showing scaled and normalized gene expression of MAPK6 per gene dosage at multiplex SNPs, faceted by individual, regardless of genotype. Differential gene expression was performed on unscaled and unnormalized values ​​using DESeq2. The solid line on the volcano plot is the Bonferroni corrected p-value of 0.05. In the volcano plot, each dot represents one gene. In Figures 4G and 4I, each dot represents one cell. Scaled and normalized gene expression was calculated using Seurat. [Figure 4B] See legend to Figure 4A. [Figure 4C-1] See legend to Figure 4A. [Figure 4C-2] See legend to Figure 4A. [Figure 4D] See legend to Figure 4A. [Figure 4E] See legend to Figure 4A. [Figure 4F] See legend to Figure 4A. [Figure 4G] See legend to Figure 4A. [Figure 4H] See legend to Figure 4A. [Figure 4I] See legend to Figure 4A. [Figure 5A]Figures 5A-5E provide violin plots, Uniform Manifold Approximation and Projection (UMAP) plots, heat maps, and box plots showing metrics of genomic DNA amplicons from HH edited cells. Figure 5A provides a violin plot of all edits (substitutions, insertions, and deletions) as a proportion of edited reads from 0 to 1, summed across all cells examined and graphed per nucleotide across the amplicon. Each dot represents one nucleotide within the amplicon, and peaks indicate the center of CRISPR editing and the most likely regions that are mutated. Values ​​were extracted from CRISPR Resso2 analysis as described in the methods. Figure 5D provides a heat map showing the number of reads aligned to the indicated amplicons per cell. Figures 5B and 5C provide Uniform Manifold Approximation and Projection (UMAP) plots. Figure 5E provides a box plot showing HLA-DQB1 gene expression. [Figure 5B] See legend to Figure 5A. [Figure 5C] See legend to Figure 5A. [Figure 5D] See legend to Figure 5A. [Figure 5E] See legend to Figure 5A. [Figure 6A-1]Figures 6A-6G provide schematics, box plots, and diagrams showing the optimization of a multi-omics single-cell protocol for capturing genomic DNA, ADT, and mRNA from Jurkat cells base-edited with variant rs61839660. Figure 6A provides a schematic of the experimental overview and a schematic of the IL2RA locus and the targeted variant (rs61839660). Figure 6B provides a box plot of the total genomic DNA reads recovered per cell and the percentage of targeted base-edited reads per cell per condition defined in A. Figure 6C provides a box plot of the total antibody-derived tag (ADT) unique molecular identifiers (UMIs) recovered per cell and the distribution of count log ratio normalized counts for each antibody. Figure 6D provides a box plot of the UMIs per cell and the total number of genes recovered per cell per condition. In Figures 6A-6D, all comparisons between conditions are significant using Kruskal-Wallis tests and Dunn's post-test comparisons. Figure 6E provides a diagram showing the common genotypes (greater than 4 cells) recovered. Rare genotypes (four cells or less) are not shown. The histograms and numbers on the right represent the number of cells from each genotype. Arrows indicate the sequence and position of the sgRNA and point away from the PAM site. Figure 6F provides box plots showing gene expression of IL2RA and CLR-normalized counts of ADT CD25 in G1, G2, G3, and G4+ based on the genotypes in E. G4+ represents genotype G4 and all other rare genotypes with less than 4 cells. Figure 6G provides volcano plots and box plots showing differential gene (non-zero in 30% of cells) expression versus gene dosage in targeted variants (G1, G2, G3), excluding all rare (G4+) genotypes. In Figure 6G, each dot represents one gene. The dotted line is a Bonferroni corrected p-value of 0.05. The expression of significant genes in all four genotypes is shown. Each dot represents one cell. Gene expression values ​​were scaled and normalized using Seurat. [Figure 6A-2] See description of Figure 6A-1. [Figure 6B]See description of Figure 6A-1. [Figure 6C] See description of Figure 6A-1. [Figure 6D] See description of Figure 6A-1. [Figure 6E] See description of Figure 6A-1. [Figure 6F] See description of Figure 6A-1. [Figure 6G] See description of Figure 6A-1. [Figure 7A] Figures 7A-7D provide plots and heat maps showing RNA clustering of CRISPR-Cas edited Jurkat cells. Figure 7A provides plots showing analysis of RNA from rs61839660 edited Jurkat cells described in Figures 2A-2K using Seurat (six clusters were identified). Figure 7B provides plots showing that RNA clustering did not reveal condition bias after running Harmony. Figure 7C provides plots showing that IL2RA gene expression was not significantly different per cluster. Figure 7D provides a heat map showing logFC values ​​of genes identified in differential gene expression analysis using Poisson modeling for RNA clusters. Each dot represents one cell. [Figure 7B] See legend to Figure 7A. [Figure 7C] See legend to Figure 7A. [Figure 7D] See legend to Figure 7A. [Figure 8A-1] Figures 8A and 8B provide schematics of single-cell MINECRAFTseq. Figure 8A provides a schematic outlining CD4 T cell isolation, CRISPR editing, indexing, ADT staining, and cell sorting prior to library generation. Figure 8B provides a schematic outlining library generation for sequencing from each of the three single-cell modalities: genomic DNA (top right corner of figure), mRNA (middle right corner of figure), and antibody-derived tag (ADT, bottom right corner of figure). [Figure 8A-2] See description of Figure 8A-1. [Figure 8A-3]See description of Figure 8A-1. [Figure 8A-4] See description of Figure 8A-1. [Figure 8B-1] See description of Figure 8A-1. [Figure 8B-2] See description of Figure 8A-1. [Figure 8B-3] See description of Figure 8A-1. [Figure 8B-4] See description of Figure 8A-1. [Figure 9A] Figures 9A-9C provide plots and histograms showing correlation between ADT metrics and index flow cytometry information from PTRPC-edited primary CD4 T cells. Figure 9A shows that the Uniform Manifold Approximation and Projection (UMAP) of ADT markers are well mixed per plate. Figure 9B provides plots of index flow staining of CD45-FITC (biexponentially transformed values ​​on the y-axis) and CLR-normalized counts of ADT UMIs (x-axis) showing strong correlation, identifying knockout, heterozygote, and wild-type cells. Genotypes (A, B, C) of cells are defined in Figures 3B and 3C. Figure 9C provides histograms showing CLR-normalized counts of all 33 ADT markers used in the experiment from all cells. Each dot represents one cell. [Figure 9B] See legend to Figure 9A. [Figure 9C] See legend to Figure 9A. [Figure 10A]Figures 10A-10D provide violin plots, plots, and heat maps showing RNA metrics and clustering of PTRPC-edited primary CD4 T cells. Figure 10A provides violin plots showing the percentage of mitochondrial reads detected per cell, the number of unique molecular identifiers (UMIs), and the total number of genes. Cells were not filtered by any criteria before plotting. Figure 10B provides a Uniform Manifold Approximation and Projection (UMAP) plot based on the variable gene mRNA PCs where clusters were identified in Seurat. Figure 10C provides an RNA UMAP plot where plate identity was plotted. Figure 10D provides a heat map showing the logFC values ​​of genes identified in differential gene expression analysis with Poisson modeling for RNA clusters. Each dot represents a single cell. RNA analysis was performed in Seurat. [Figure 10B] See legend to Figure 10A. [Figure 10C] See legend to Figure 10A. [Figure 10D] See legend to Figure 10A. [Figure 11A] Figures 11A-11D provide volcano plots and box plots showing differential gene expression of PTRPC-edited primary CD4 T cells. Figure 11A provides a volcano plot of differential gene expression versus gene dosage at the targeted nucleotide. Only genotypes A, B, and C, as defined in Figures 3B and 3C, were used for the analysis. Genes for analysis were selected based on greater than 30% non-zero expression. The dotted lines on the volcano plots are Bonferroni-corrected p-values ​​of 0.05. Each dot represents one gene. Figures 11B and 11D provide box plots showing scaled and normalized gene expression of the top three identified genes in Figure 11A. Each dot represents one cell. Scaled and normalized gene expression was calculated using Seurat. [Figure 11B] See legend to Figure 11A. [Figure 11C]See legend to Figure 11A. [Figure 11D] See legend to Figure 11A. [Figure 12A] Figures 12A-12J provide diagrams, box plots, and plots showing bulk RNA, DNA, and flow cytometry data from editing at the UBASH3A locus. Figures 12A-12D provide diagrams showing bulk DNA editing results from three healthy individuals with targeted nucleotides or regions highlighted, generated using CRISPResso2. Arrows indicate the position of the sgRNA away from the PAM. N is non-targeting sample, HDR is homology directed repair sample, and BE is base-edited sample. Numbers indicate the percentage of modified reads, with black bars representing deletions. Figures 12E-12J provide box plots and plots showing bulk mRNA expression of UBASH3A from the same healthy individuals. Gene expression values ​​are scaled and normalized as logUMI+1. (G and H) Bulk flow cytometry from healthy individual samples shown as median fluorescence intensity of key immune markers on CD4 T cells. Flow cytometry values ​​were calculated in FlowJo and plotted using GraphPad. Points connected by lines are from paired samples. [Figure 12B] See legend to Figure 12A. [Figure 12C] See legend to Figure 12A. [Figure 12D] See legend to Figure 12A. [Figure 12E] See legend to Figure 12A. [Figure 12F] See legend to Figure 12A. [Figure 12G] See legend to Figure 12A. [Figure 12H] See legend to Figure 12A. [Figure 12I] See legend to Figure 12A. [Figure 12J] See legend to Figure 12A. [Figure 13A]Figures 13A-13D provide diagrams and bar graphs showing sequence data for HDR corrected cells from rs11203202 and rs9981624 editing conditions. Figures 13A and 13B provide diagrams showing the recovered corrected genotypes, with sgRNA sequences and cell numbers indicated (right). Number of cells with a specific insertion (- value) or deletion (+ value) for (Figure 13C) rs11203202 HDR edited samples or (Figure 13D) rs9981624 HDR edited samples. Most of the cells edited from rs11203202 contained a single insertion, as evident from the bulk data. [Figure 13B] See legend to Figure 13A. [Figure 13C] See legend to Figure 13A. [Figure 13D] See legend to Figure 13A. [Figure 14A]Figures 14A-14L provide heatmaps, violin plots, and plots showing single-cell RNA metrics and clustering from variant editing at the UBASH3A locus. Figure 14D provides violin plots showing the percent of mitochondrial reads, number of unique molecular identifiers (UMIs), and total number of genes per cell detected from base-edited cells including non-targeted control, rs80054410, and rs11203203 conditions. Figure 14E provides a Uniform Manifold Approximation and Projection (UMAP) plot based on the variable gene mRNA PCs where clusters were identified with Seurat from base-edited cells. Figure 14F provides an RNA UMAP where expression of UBASH3A was plotted from base-edited cells. Figure 14A provides a heatmap showing logFC values ​​of genes identified in differential gene expression analysis using Poisson modeling for RNA clusters from base-edited cells. Figure 14J provides violin plots showing the percent of mitochondrial reads, number of unique molecular identifiers (UMIs), and total number of genes per cell detected from HDR edited cells including non-targeted control, rs11203202, and rs9981624 conditions. Figure 14K provides a UMAP plot based on variable gene mRNA PCs where clusters were identified in Seurat from HDR edited cells. Figure 14L provides an RNA UMAP where expression of UBASH3A was plotted from HDR edited cells. Figure 14G provides a heatmap showing logFC values ​​of genes identified in differential gene expression analysis using Poisson modeling for RNA clusters from HDR edited cells. Figures 14B, 14C (both for rs11203203) and Figures 14H, 14I (both for rs9981624) provide violin plots showing #UMIs and #genes. Each dot represents a single cell. RNA analysis was performed in Seurat. [Figure 14B] See legend to Figure 14A. [Figure 14C] See legend to Figure 14A. [Figure 14D] See legend to Figure 14A. [Figure 14E] See legend to Figure 14A. [Figure 14F] See legend to Figure 14A. [Figure 14G] See legend to Figure 14A. [Figure 14H] See legend to Figure 14A. [Figure 14I] See legend to Figure 14A. [Figure 14J] See legend to Figure 14A. [Figure 14K] See legend to Figure 14A. [Figure 14L] See legend to Figure 14A. [Figure 15-1] 15A-15N provide histograms and plots showing that editing variants at the UBASH3A locus did not affect cell surface protein expression. FIG. 15A provides histograms showing the CLR-normalized distribution of all ADT markers measured from base-edited samples. FIG. 15B-15D provide plots showing that expression of HLA-DR, CD27, and CD45RO defines distinct clusters of CD4 T cells in base-edited samples. FIG. 15E-15G provide plots showing that cells were evenly mixed per donor, plate, and condition in base-edited samples. FIG. 15H provides histograms showing the CLR-normalized distribution of all ADT markers measured from HDR-edited samples. FIG. 15I-15K provide plots showing that expression of HLA-DR, CD27, and CD45RO defines distinct clusters of CD4 T cells in HDR-edited samples. Figures 15L-15N provide plots showing that in the base-edited samples, cells were evenly mixed per donor, plate, and condition. Each dot represents a single cell. [Figure 15-2] See description of Figure 15-1. [Figure 15-3] See description of Figure 15-1. [Figure 15-4] See description of Figure 15-1. [Figure 15-5]See description of Figure 15-1. [Figure 15-6] See description of Figure 15-1. [Figure 16A-1] Figures 16A-16C provide diagrams, histograms, and plots showing bulk RNA, DNA, and flow cytometry data from edits at the Il2RA locus. Figure 16A provides a diagram showing bulk DNA editing results from three healthy individuals with targeted nucleotides or regions highlighted, generated using CRISPResso2. Arrows indicate the position of the sgRNA away from the PAM. N is non-targeted sample, BE is individually base-edited sample, and multiplex is simultaneous edits at all three variants. Numbers indicate the percentage of modified reads. Figure 16B provides an overlay of flow cytometry histograms and plots showing representative bulk flow cytometry from edited samples and median fluorescence intensity versus controls. Flow cytometry values ​​were extracted with FlowJo and plotted using GraphPad. Figure 16C provides plots showing bulk mRNA expression of IL2RA from the same healthy individuals. Gene expression values ​​are scaled and normalized as logUMI+1. Dots connected by lines represent paired samples. [Figure 16A-2] See description of Figure 16A-1. [Figure 16A-3] See description of Figure 16A-1. [Figure 16B] See description of Figure 16A-1. [Figure 16C] See description of Figure 16A-1. [Figure 17-1] Figure 17 provides a diagram showing single cell DNA genotypes from each healthy individual per condition from edits at the IL2RA locus. Zoomed in view of common (>4 cell) genotypes identified per healthy individual and per condition. Cell counts per genotype are shown to the right of each plot. Healthy individuals are represented in columns and conditions in rows. [Figure 17-2] See description of Figure 17-1. [Figure 17-3] See description of Figure 17-1. [Figure 18A] Figures 18A and 18B provide plots showing that linear modeling of ADT counts identifies multiplex single nucleotide polymorphisms (SNPs) and rs61939660 as correlates of CD25 expression. Figure 18A provides plots showing that linear regression modeling was performed to evaluate which variant nucleotides correlate with CLR-normalized CD25 ADT expression taking into account plate. Figure 18B provides plots showing that, conditional on gene dosage at SNP3, linear regression was performed again taking into account plate. Nominal p-values ​​are plotted, and dotted lines represent Bonferroni-corrected p-values ​​of 0.05. SNP identities are defined in Figure 4C. [Figure 18B] See legend to Figure 18A. [Figure 19A-1] The combination of Figures 19A and 19B provides a schematic diagram of an improved version of MINECRAFTseq. Figure 19A provides a schematic overview of cell preparation and sorting prior to library creation. Figure 19B provides a schematic overview of library creation. [Figure 19A-2] See description of Figure 19A-1. [Figure 19A-3] See description of Figure 19A-1. [Figure 19A-4] See description of Figure 19A-1. [Figure 19B-1] See description of Figure 19A-1. [Figure 19B-2] See description of Figure 19A-1. [Figure 19B-3] See description of Figure 19A-1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0070] Detailed Description of the Invention The present invention features compositions and methods useful for characterizing alterations in polynucleotides compared to a reference sequence.

[0071] The present invention is based, at least in part, on the discovery of a technique that provides for the investigation of polynucleotide sequence changes, including those associated with CRISPR editing. It can be applied to a wide variety of cells, including cell lines and primary cells. The technique uses flow-assisted sorting of single cells into plates to capture DNA amplicons, total 3'mRNA, and antibody-derived tags (ADTs) from CRISPR-edited cells, correlating genomic editing of targeted regions with protein expression and mRNA outcomes. This novel approach utilizes a 3'mRNA capture approach and extensive multiplexing to allow robust and relatively inexpensive analysis of tens of thousands, or even hundreds of thousands, of cells. It provides for simultaneous analysis of DNA genomic editing, RNA expression changes (including characterizing global expression changes of genes of interest), and can use antibody-derived tags (ADTs) to characterize the phenotype of specific cells of interest.

[0072] Genetic studies have identified thousands of disease-causing coding and non-coding alleles, but identifying the function of these alleles has proven to be a significant bottleneck. CRISPR-Cas gene editing technology has enabled targeted modification of DNA. However, these approaches are limited due to heterogeneous outcomes of successful DNA modification. In some embodiments, the disclosed invention provides a scalable, plate-based, single-cell approach that simultaneously captures genomic DNA amplicons, mRNA transcriptome, and ADT expression. As further described in the Examples herein, this novel multi-omics approach was used in combination with a range of genome editing techniques to interrogate coding and regulatory alleles of HLADQB1, IL2RA, PTPRC, and UBASH3A in cell lines and primary human CD4 T cells. This combination of single-cell editing led to robust detection of functional outcomes, as shown in the Examples.

[0073] An effective method for quickly evaluating the effect of genome editing is to capture the target DNA information of single cells in parallel with reading mRNA and cell surface expression.This approach provided in the inventive aspect of the present disclosure has the advantage that it allows the analysis of primary cells and allows the powerful comparison of edited cells and non-edited cells in the same experiment.In one embodiment, the method provided herein is suitable for the analysis of primary immune cells or CRISPR edited samples.

[0074] Limitations of current approaches Current approaches to isolate and sequence single-cell DNA and mRNA include simultaneous isolation of genomic DNA and total RNA ( S imultaneous I Solution of genomics D NA and total R These include "NA:SIDR", "G&T", "DR-seq" and "TARGET-seq". All of these approaches capture DNA amplicons or whole genomes along with mRNA using various techniques. However, these approaches are still limited in terms of scalability, reliability and efficiency. Furthermore, none of these techniques incorporate antibody-derived tags into their protocols and importantly, none of them have been applied to CRISPR editing. Other single-cell approaches focusing on CRISPR-edited cells rely entirely on the expression of sgRNA as a proxy for DNA editing. This is a major limitation for deciphering the diverse and heterogeneous outcomes of editing on mRNA expression. Moreover, these applications rely on DNA integration into the genome, making them unstable and usually can only be performed in cell lines and not primary cells.

[0075] Multiomic Investigation of Nucleotide Editing by CRISPR with ADT, Flow Cytometry and Transcriptome sequencing (MINECRAFTseq) addresses all of these issues by capturing up to four modalities: A) flow index information, B) mRNA, C) DNA amplicons, and D) ADT from cell lines and primary cells edited with CRISPR reagents. Such methods are useful, for example, for VDJ sequencing of TCR clonotypes (in addition to other modalities), telomere sequencing to understand the relationship between cell age and immunity, characterizing splice isoforms to precisely measure the impact of autoimmune variants on differential isoform usage, and characterizing cancer heterogeneity.

[0076] Single-cell analysis MINECRAFTseq provides multi-omics analysis of single cells. Single cells can be isolated using microfluidic devices. Microfluidic technology includes microscale devices that handle small volumes of fluids. Microfluidic technology can accurately and reproducibly control and dispense small volumes of fluids, especially volumes less than 1 μl, so the application of microfluidic technology can result in significant cost reductions. The use of microfluidic technology reduces cycle times, shortening time to results and increasing throughput. Small-volume microfluidic technology improves the amplification and construction of DNA libraries made from single cells and single isolated cellular component aggregates. Furthermore, the incorporation of microfluidic technology promotes system integration and automation.

[0077] The single cell of the present invention can be divided into single droplets using a microfluidic device. The single cell in such droplets can be further labeled with barcode. In this regard, see Macosko et al., 2015, "Highly Parallel Genome-wide Expression Profiling of Individual Cells Using Nanoliter Droplets" Cell 161, 1202-1214 and Klein et al., 2015, "Droplet Barcoding for Single-Cell Transcriptomics Applied to Embryonic Stem Cells" Cell 161, 1187-1201; the entire contents and disclosures of each are incorporated herein by reference in their entirety.

[0078] Microfluidic reactions are generally carried out in microdroplets. The ability to carry out reactions in microdroplets depends on the ability to merge different microdroplets with different sample fluids. See, for example, U.S. Patent Application Publication No. 20120219947 and PCT Publication No. WO2014085802 A1.

[0079] Droplet microfluidic technologies (e.g., 10X, DROPSEQ, InDrop) offer great advantages in conducting high-throughput screens and highly sensitive assays. Droplets allow for a significant reduction in sample volume while at the same time reducing costs. Operation and measurement at kilohertz speeds allows for up to 10 8 Droplet microfluidics technology allows for the screening of samples of up to 10 ...

[0080] The manipulation of fluids to form desired shaped fluid streams, discontinuous fluid streams, droplets, particles, dispersions, etc. for purposes of fluid delivery, product manufacturing, analysis, etc. is a relatively well-studied field. Microfluidic systems have been described in a variety of fields, commonly in the field of miniaturized laboratory (e.g., clinical) analysis; other applications have also been described. See, for example, WO 2001 / 89788; WO 2006 / 040551; U.S. Patent Application Publication No. 2009 / 0005254; WO 2006 / 040554; U.S. Patent Application Publication No. 2007 / 0184489; WO 2004 / 002627; U.S. Patent No. 7,708,949; WO 2008 / 063227; U.S. Patent Application Publication No. 2008 / 0003142; WO 2004 / 091763; U.S. Patent Application Publication No. 2006 / 0163385; WO 2005 / 021151; U.S. Patent Application Publication No. 2007 / 0003442; WO 2006 / 096571; US Patent Application Publication No. 2009 / 0131543; WO 2007 / 089541; US ​​Patent Application Publication No. 2007 / 0195127; WO 2007 / 081385; US Patent Application Publication No. 2010 / 0137163; WO 2007 / 133710; US Patent Application Publication No. 2008 / 0014589; US Patent Application Publication No. 2014 / 0256595; and WO 2011 / 079176. In a preferred embodiment, single molecule analysis is performed in the droplet using the method described in WO 2014085802. Each of these patents and publications is incorporated herein by reference in its entirety for all purposes.

[0081] Single cell isolation and lysis The present disclosure includes a process for isolating individual cells from a sample, where the cells are separated and isolated into individual compartments. The method used for cell separation depends, in part, on the origin and type of sample used. For example, separation of individual cells from single cell suspensions of blood or tissue can be performed by methods routinely performed in the art, such as flow cytometry or microfluidic techniques (e.g., single cell sorting using fluorescence activated cell sorting (FACS) techniques).

[0082] In certain embodiments, single cells obtained or isolated from tissue are isolated into individual compartments, for example, by being placed into individual wells of a tissue culture plate or into microfluidic droplets. In certain embodiments, individual cells are encapsulated within individual gel beads. In certain aspects, the beads are plastic, glass, silica or metal beads, and the target biomolecule is released from the beads by chemical or enzymatic reaction.

[0083] In certain embodiments, individual cells are encapsulated within individual oil droplets. In some embodiments, the oil droplets are an aqueous solution surrounded by oil. In certain embodiments, the oil is immiscible with water. In certain embodiments, the oil is transparent. In certain embodiments, the oil droplets have a volume between 1 pL and 100 nL. In certain embodiments, the aqueous solution surrounded by the oil comprises a buffer. In certain embodiments, a surfactant is added to the oil droplets.

[0084] The method includes lysing individual cells to expose target biomolecules for detection. The lysis protocol of cells depends, in part, on the nature and subcellular location of the target biomolecule to be detected. Any method known in the art for lysing membranes and / or extracting target biomolecules from cells can be used. Examples of lysing agents include, but are not limited to, detergents (e.g., NP-40 (nonylphenoxypolyethoxylethanol)), surfactants (e.g., non-ionic surfactants such as Triton X-100 and Tween 20, or ionic surfactants such as Sarkosyl and sodium dodecyl sulfate), or lytic enzymes (e.g., lysozyme). In certain embodiments, the lysing agent disrupts cell membranes but not oil droplets. In other embodiments, non-reagent-based lysis systems can be used, including, but not limited to, heat, electroporation, mechanical disruption, and acoustic disruption (e.g., sonication). In one embodiment, cells are lysed with a solution containing at least one detergent, surfactant, or lytic enzyme. In certain embodiments, cells are lysed using a combination of lysis reagents and techniques. In certain embodiments, the surfactant is Triton X-100. In another embodiment, the detergent is NP-40 (nonylphenoxypolyethoxylethanol). In one embodiment, cells are lysed with a buffer containing sodium dodecyl sulfate. In certain embodiments, the cellular material released from the lysed cells includes cellular proteins. In certain aspects, lysis of cells is performed within individual single-cell compartments.

[0085] In certain embodiments, RNA, DNA, and proteins from cells can be extracted separately from individual cells, allowing for multiplexed analysis of the transcriptome, genome, and / or proteome from each cell. In one aspect, RNA, DNA, and proteins can be extracted using extraction reagents that allow for simultaneous isolation of RNA, DNA, and proteins.

[0086] Labeling cells with a detectable marker The cells described herein can be labeled for sorting and identification. Detectable markers can be any molecule that can generate a signal for detecting target biomolecules. For example, the detectable marker for cell identification is a fluorescent marker. Detectable markers for cell identification include, but are not limited to, fluorescent molecules, chemiluminescent molecules, chromophores, enzymes, enzyme substrates, enzyme cofactors, enzyme inhibitors, dyes, metal ions, metal sols, ligands (e.g., biotin, avidin, streptavidin, or haptens), radioisotopes, molecules designed for electron / ion detection (e.g., by ISFET), and the like, and combinations thereof.

[0087] Detectable markers can be chemically and / or covalently bound to the appropriate region of cell discrimination probe. In some embodiments, detectable markers are fluorescent molecules. Fluorescent molecules can be fluorescent proteins or reactive derivatives of fluorescent molecules known as fluorophores. Fluorophores are fluorescent compounds that emit light upon light excitation. In some embodiments, fluorophores can selectively bind to specific regions or functional groups on target molecules and be chemically or biologically attached.

[0088] Examples of usable labels include those known to those skilled in the art, such as fluorescent dyes, enzymes, coenzymes, chemiluminescent substances, and radioactive substances, so long as the labels detect double-stranded nucleic acids. Specific examples include radioisotopes (e.g., 32 P, 14 C. 125 I, 3 H, and 131I), fluorescein, rhodamine, dansyl chloride, umbelliferone, luciferase, peroxidase, alkaline phosphatase, β-galactosidase, β-glucosidase, horseradish peroxidase, glucoamylase, lysozyme, saccharide oxidase, microperoxidase, biotin, and ruthenium. When biotin is used as a labeling substance, preferably, streptavidin bound to an enzyme (e.g., peroxidase) is further added after the addition of the biotin-labeled antibody. Advantageously, the label intercalates into double-stranded DNA, such as ethidium bromide.

[0089] Advantageously, the label is a fluorescent label. This dye may be the Evagreen dye or the ROX dye. Examples of fluorescent labels include, but are not limited to: Atto dyes, 4-acetamido-4'-isothiocyanatostilbene-2,2'-disulfonic acid; acridine and derivatives: acridine, acridine isothiocyanate; 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS); 4-amino-N-[3-vinylsulfonyl)phenyl]naphthalimide-3,5-disulfonate; N-(4-anilino-1-naphthyl)maleimide; anthranilamide; BODIPY; brilliant yellow; coumarin and derivatives: coumarin, 7-amino-4-methylcoumarin (AMC, coumarin 120), 7-amino-4-trifluoromethylcoumarin (coumaran 151); cyanine dyes; cyanosine; 4',6-diaminidino-2-phenylindole (DAPI); 5'5''-dibromopyrogallol-sulfonaphthalene (Bromopyrogallol Red);7-Diethylamino-3-(4'-isothiocyanatophenyl)-4-methylcoumarin;Diethylenetriaminepentaacetate;4,4'-Diisothiocyanatodihydro-stilbene-2,2'-disulfonic acid;4,4'-Diisothiocyanatostilbene-2,2'-disulfonic acid;5-[Dimethylamino]naphthalene-1-sulfonyl chloride (DNS, dansyl chloride);4-Dimethylaminophenylazophenyl-4'-isothiocyanate (DABITC);Eosin and derivatives: Eosin, Eosin isothiocyanate, Erythrosine and derivatives: Erythrosine B, Erythrosine, isothiocyanate;Ethidium;Fluorescein and derivatives: 5-Carboxyfluorescein (FAM), 5-(4,6-Dichlorotriazin-2-yl)aminofluorescein (DTAF), 2',7'-dimethoxy-4'5'-dichloro-6-carboxyfluorescein, fluorescein, fluorescein isothiocyanate, QFITC ​​(XRITC); fluorescamine; IR144; IR1446; malachite green isothiocyanate; 4-methylumbelliferone orthocresolphthalein; nitrotyrosine; pararosaniline; phenol red; B-phycoerythrin; o-phthaldialdehyde;Pyrene and derivatives: Pyrene, pyrene butyrate, succinimidyl 1-pyrene; butyrate quantum dots; Reactive Red 4 (Cibacron Brilliant Red 3B-A); Rhodamine and derivatives: 6-carboxy-X-rhodamine (ROX), 6-carboxyrhodamine (R6G), Lissamine rhodamine B sulfonyl chloride rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine X isothiocyanate, sulforhodamine B, sulforhodamine 101, sulfonyl chloride derivative of sulforhodamine 101 (Texas Red); N,N,N',N' tetramethyl-6-carboxyrhodamine (TAMRA); tetramethylrhodamine; tetramethylrhodamine isothiocyanate (TRITC); riboflavin; rosolic acid; terbium chelate derivatives; Cy3; Cy5; Cy5.5; Cy7; IRD 700; IRD 800; La Jolta Blue; phthalocyanine; and naphthalocyanine.

[0090] The fluorescent label can be a fluorescent protein, such as blue fluorescent protein, cyan fluorescent protein, green fluorescent protein, red fluorescent protein, yellow fluorescent protein, or any photoconvertible protein. Colorimetric labeling, bioluminescent labeling, and / or chemiluminescent labeling can further achieve labeling. Labeling can further include energy transfer between molecules in a hybridization complex by perturbation analysis, quenching, or electron transport between donor and acceptor molecules, the latter of which can be facilitated by double-stranded match hybridization complexes. The fluorescent label can be perylene or terrylene. Alternatively, the fluorescent label can be a fluorescent barcode.

[0091] The label may be a fluorescent label, advantageously fluorescein or rhodamine. In another embodiment, the label may be an organic label. In some embodiments, fluorescent tags useful in the methods of the present disclosure include, but are not limited to, green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), fluorescein, fluorescein isothiocyanate (FITC), tetramethylrhodamine isothiocyanate (TRITC), cyanine (Cy3), phycoerythrin (R-PE), 5,6-carboxymethylfluorescein, (5-carboxyfluorescein-N-hydroxysuccinimide ester), Texas Red, nitrobenzo-2-oxa-1,3-diazol-4-yl (NBD), coumarin, dansyl chloride, and rhodamine (5,6-tetramethylrhodamine).

[0092] In certain embodiments, the detectable marker is configured for electronic detection, for example, the detectable marker can release ions upon subsequent reaction, altering the pH of its environment in a reliably detectable manner.

[0093] Polynucleotide barcoding In the context of polynucleotides, a barcode refers to a unique, non-naturally occurring nucleic acid sequence that can be used to identify the origin of a nucleic acid fragment. Such barcodes include: The sequences may include, but are not limited to, TIFF2024516637000002.tif11135. Although it is not necessary to understand the mechanism of an invention, it is believed that the barcode sequences provide high quality individual reads of the barcodes associated with a particular polynucleotide (e.g., labeled ligand, shRNA, sgRNA, or cDNA) so that multiple species can be sequenced together. Furthermore, it is believed that these putative barcode loci are short enough that they can be easily sequenced with current technology. Kress et al., "DNA barcodes: Genes, genomics, and bioinformatics" PNAS 105(8):2761-2762 (2008).

[0094] Software for DNA barcoding requires integration with a field information management system (FIMS), a laboratory information management system (LIMS), sequence analysis tools, workflow tracking to link field and laboratory data, database deposition tools, and pipeline automation to scale up to ecosystem-wide projects. Geneious Pro can be used for the sequence analysis component, and two plugins made freely available through the Moorea Biocode Project, the Biocode LIMS and Genbank Submission plugins, handle integration with FIMS, LIMS, workflow tracking, and database deposition.

[0095] Additionally, other barcoding designs and tools have been described (see, e.g., Birrell et al., (2001) Proc. Natl Acad. Sci. USA 98, 12608-12613; Giaever, et al., (2002) Nature 418, 387-391; Winzeler et al., (1999) Science 285, 901-906; and Xu et al., (2009) Proc Natl Acad Sci US A. Feb 17;106(7):2289-94).

[0096] The cell identifier oligonucleotide barcode can be of any length that allows for efficient binding to the target sequence. In certain aspects, the cell identifier oligonucleotide barcode is less than 200 nucleotides long, less than 100 nucleotides long, less than 80 nucleotides long, less than 50 nucleotides long, less than 40 nucleotides long, less than 30 nucleotides long, or less than 20 nucleotides long. The complementarity of the cell identifier oligonucleotide barcode to the cell identifier probe oligonucleotide is an exact match between the nucleic acid sequences, e.g., between the cell identifier probe oligonucleotide sequence and the cell identifier oligonucleotide barcode sequence of interest (e.g., nucleotide sequence variant), such that stable and specific binding occurs. It is understood that the sequence of a nucleic acid need not be 100% complementary to the sequence of its target or complement. In some cases, the sequence is complementary to the other sequence except for one or two mismatches. In some cases, the sequence is complementary except for one mismatch. In some cases, the sequence is complementary except for two mismatches. In some cases, the sequence is complementary except for three mismatches. In still other cases, the sequences are complementary except for 4, 5, 6, 7, 8, 9 or more mismatches. In certain aspects, the number of mismatches is 20% or less, 10% or less, 5% or less, or 2% or less of the number of nucleotides present in the cell identifier oligonucleotide barcode. In certain aspects, the cell identifier oligonucleotide barcode and the cell identification probe oligonucleotide are complementary to at least 18, at least 17, at least 16, at least 15, at least 14, at least 13, at least 12, at least 11, at least 1, at least 9, at least 8, at least 7, at least 6, or at least 5 nucleotides of the target nucleotide sequence. In certain aspects, the tag is complementary to one or more individual probes. In certain aspects, the tag does not bind to alternative sequences because mismatches in the sequence lead to loss of complementarity.

[0097] In certain embodiments, the cell identification tag is conjugated or attached to the target biomolecule using enzymatic conjugation.

[0098] Methods for synthesizing barcodes include, in certain embodiments, randomly adding mixed bases during nucleic acid synthesis to generate sequences that can be used to identify specific oligonucleotide molecules through analysis of sequencing data. In certain embodiments, the synthesis of barcodes includes controlling the addition of bases to generate known sequences. In certain embodiments, barcode sequences can be verified by sequencing. In certain aspects, barcodes can be synthesized and extended using polymerase to attach barcodes to oligonucleotides on probes and tags, such as cell identification probes, target detection probes, cell identification tags, and target identification tags. In other embodiments, barcode sequences can be synthesized without probes and ligated or annealed to probes in a separate step.

[0099] Oligonucleotide Conjugates In certain embodiments, the assays described herein involve contacting cellular material (e.g., DNA, RNA) from a single cell with an oligonucleotide conjugated to an antibody. The oligonucleotide can be conjugated to the antibody by many methods known in the art (Kozlov et al., "Efficient strategies for the conjugation of oligonucleotides to antibodies enabling highly sensitive protein detection"; Biopolymers; 73(5); Apr. 5, 2004; pp. 621-630). Aldehydes can be introduced into the antibody by modification of primary amines or oxidation of carbohydrate residues. Aldehyde or hydrazine modified oligonucleotides are prepared during phosphoramidite synthesis or by post-synthetic derivatization. Conjugation of the modified oligonucleotide with the antibody results in the formation of a hydrazone bond that is stable for long periods of time under physiological conditions. Oligonucleotides can also be conjugated to antibodies by generating chemical handles via thiol / maleimide chemistry, azide / alkyne chemistry, tetrazine / cyclooctyne chemistry and other click chemistries, which are prepared during or after phosphoramidite synthesis.

[0100] In one embodiment, the oligonucleotide-antibody conjugates are designed for use in single-cell sequencing platforms that rely on poly-dT oligonucleotides as an mRNA capture method (scRNA-seq). The antibodies are integrated into the scRNA-seq workflow by mimicking native mRNA thanks to the poly-A tail sequence in the conjugated oligonucleotide. The oligonucleotides also contain a barcode that permanently labels a specific clone, and a PCR handle that makes them compatible with Illumina® sequencing reagents and instruments.

[0101] Cell Hashing In some embodiments, oligonucleotide-tagged antibodies are used to convert the detection of cell surface proteins into sequenceable readouts together with scRNA-seq. A distinct set of oligo-tagged antibodies against ubiquitous surface proteins is used to uniquely label different experimental samples. This allows these samples to be pooled together. The signal of the barcoded antibodies is used as a fingerprint for reliable demultiplexing. This approach is called cell hashing, and it is based on the concept of a hash function in computer science to index a dataset with certain characteristics; our set of oligo-derived hashtags equally defines a "lookup table" to assign each multiplexed cell to its original sample.

[0102] In some embodiments, cell hashing involves the use of oligo-tagged antibodies against ubiquitously expressed surface proteins, which can be used to uniquely label cells from different samples before pooling the samples. By sequencing these tags in parallel with the cell's transcriptome, each cell is assigned to its original sample, and cross-sample multiplets are robustly identified and "super-loaded" into commercially available droplet-based systems for significant cost reduction. Hashing can generalize the benefits of single cell multiplexing to diverse samples and experimental designs.

[0103] Sequencing The gDNA, cDNA and ADT libraries described herein are suitable for virtually any sequencing method known in the art. In some embodiments, sequencing is performed using next-generation sequencing technology (NGS), which allows for massively parallel sequencing. In certain embodiments, clonally amplified DNA templates or single DNA molecules are sequenced in a massively parallel manner in a flow cell (e.g., as described in Volkerding et al. Clin Chem 55:641-658

[2009] ; Metzker M Nature Rev 11:31-46

[2010] ). Sequencing techniques for NGS include, but are not limited to, pyrosequencing, sequencing-by-synthesis using reversible dye terminators, sequencing by ligation of oligonucleotide probes, and ion semiconductor sequencing. DNA from each sample can be sequenced individually (i.e., singleplex sequencing), or DNA from multiple samples can be pooled and sequenced as indexed genomic molecules in a single sequencing run (i.e., multiplex sequencing), generating up to hundreds of millions of reads of DNA sequence. Examples of sequencing technologies that can be used to obtain sequence information according to the present methods are described further herein.

[0104] Several sequencing technologies are commercially available, such as the sequencing-by-hybridization platform from Affymetrix (Sunnyvale, Calif.), sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (Hayward, Calif.), and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.), as described below. In addition to single molecule sequencing performed using Helicos Biosciences' sequencing-by-synthesis, other single molecule sequencing technologies include, but are not limited to, Pacific Biosciences' SMRT™ technology, ION TORRENT technology, and nanopore sequencing, e.g., developed by Oxford Nanopore Technologies.

[0105] Although the automated Sanger method is considered to be a "first generation" technology, Sanger sequencing methods, including automated Sanger sequencing, can also be used in the methods described herein. Additional suitable sequencing methods include, but are not limited to, nucleic acid imaging techniques, such as atomic force microscopy (AFM) or transmission electron microscopy (TEM). Exemplary sequencing techniques are described in more detail below.

[0106] In some embodiments, the methods provided herein include obtaining sequence information of nucleic acids in a test sample by massively parallel sequencing of millions of DNA fragments using Illumina's sequencing-by-synthesis and reversible terminator-based sequencing chemistry (e.g., as described in Bentley et al., Nature 6:53-59

[2009] ). The template DNA can be genomic DNA, e.g., cellular DNA or cDNA. In some embodiments, genomic DNA from isolated cells is used as a template, which is fragmented to lengths of several hundred base pairs. In some embodiments, Illumina's sequencing technology relies on attachment of the fragmented genomic DNA to a flat, optically transparent surface to which an oligonucleotide anchor is attached. The template DNA is end-repaired to generate 5' phosphorylated blunt ends, and a single A base is added to the 3' ends of the blunt phosphorylated DNA fragments using the polymerase activity of the Klenow fragment. This addition prepares the DNA fragments for ligation to an oligonucleotide adaptor, which has a single T base overhang at its 3' end to increase ligation efficiency. The adaptor oligonucleotide is complementary to the anchor oligo of the flow cell. Under limiting dilution conditions, the adaptor-modified single-stranded template DNA is added to the flow cell and immobilized by hybridization to the anchor oligo. The attached DNA fragments are extended and bridge-amplified to create an ultra-high density sequencing flow cell containing hundreds of millions of clusters, each containing approximately 1,000 copies of the same template.

[0107] In one embodiment, randomly fragmented library DNA (e.g., genomic DNA, cDNA) is amplified using PCR before being subjected to cluster amplification. Alternatively or additionally, amplification-free genomic library preparations are used, and randomly fragmented genomic DNA or other polynucleotides are enriched using cluster amplification alone (Kozarewa et al., Nature Methods 6:291-295

[2009] ). In some applications, templates are sequenced using a robust four-color DNA sequencing-by-synthesis technique that uses reversible terminators with removable fluorescent dyes. Highly sensitive fluorescent detection is achieved using laser excitation and total internal reflection optics. Short sequence reads of tens to hundreds of base pairs are aligned to a reference genome, and unique mappings of the short sequence reads to the reference genome are identified using specially developed data analysis pipeline software. After completion of the first read, the template can be regenerated in situ to allow a second read from the opposite end of the fragment. Thus, single-end or paired end sequencing of DNA fragments can be used.

[0108] In various embodiments of the present disclosure, sequencing by synthesis can be used, which allows paired-end sequencing. In some embodiments, Illumina's sequencing by synthesis platform includes clustering of fragments. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, as in the examples described herein, the fragments have two different adapters attached to both ends, which allow the fragments to hybridize to two different oligos on the surface of the flow cell lane. The fragments further include or are connected to two index sequences at both ends, which provide labels to distinguish different samples in multiplex sequencing. In some sequencing platforms, the fragments that are sequenced from both ends are also called inserts.

[0109] In some embodiments, the flow cell for clustering on the Illumina platform is a glass slide with multiple lanes. Each lane is a glass channel coated on one side with two oligos (e.g., P5 and P7' oligos). Hybridization is enabled by the first of the two oligos on the surface. This oligo is complementary to the first adaptor at one end of the fragment. A polymerase generates the complementary strand of the hybridized fragment. This double-stranded molecule is denatured and the original template strand is washed away. The remaining strand is clonally amplified by bridge application in parallel with many other remaining strands.

[0110] In bridge amplification and other sequencing methods, including clustering, the strands fold over and a second adapter region at the other end of the strand hybridizes to the other type of oligo on the surface of the flow cell. A polymerase creates a complementary strand, forming a double-stranded bridge molecule. This double-stranded molecule is denatured, resulting in two single-stranded molecules tethered to the flow cell via two different oligos. This process is then repeated many times, simultaneously for millions of clusters, resulting in clonally amplified versions of all the fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3' end is blocked to prevent unwanted priming.

[0111] After clustering, sequencing begins by extending the first sequencing primer to generate the first read. In each cycle, fluorescently labeled nucleotides compete for addition to the growing strand. Based on the sequence of the template, only one is incorporated. After each nucleotide is added, the cluster is excited with a light source, which gives off a characteristic fluorescent signal. The number of cycles determines the length of the read. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical strands are read simultaneously. Hundreds of millions of clusters are sequenced in a massively parallel manner. After the first read is completed, the read products are washed away.

[0112] In the next step of the protocol, which includes two index primers, an index 1 primer is introduced to hybridize to the index 1 region on the template. The index region provides fragment discrimination, which is useful for demultiplexing samples in a multiplex sequencing process. The index 1 read is generated similarly to the first read. After completion of the index 1 read, the read product is washed away and the 3' end of the strand is deprotected. The template strand then folds back and binds to the second oligo on the flow cell. The index 2 sequence is read in the same manner as index 1. The index 2 read product is then washed away upon completion of this step.

[0113] After two index reads, read 2 is initiated, using polymerase to extend the second flow cell oligo to form a double-stranded bridge. This double-stranded DNA is denatured and the 3' end is blocked. The original forward strand is cleaved and washed away, leaving the reverse strand. Read 2 begins with the introduction of the read 2 sequencing primer. As with read 1, the sequencing steps are repeated until the desired length is achieved. The products of read 2 are washed away. This entire process generates millions of reads representing all fragments. Sequences from the pooled sample libraries are separated based on unique indexes introduced during sample preparation. For each sample, reads of similar stretches of base calls are locally clustered. Forward and reverse reads are paired to create contiguous sequences. These contiguous sequences are aligned to the reference genome for variant identification.

[0114] Sequencing by synthesis involves paired end reads. Paired-end sequencing involves two reads from both ends of a fragment. Paired-end reads are used to resolve ambiguous alignments. Paired-end sequencing allows the user to select the length of the insert (or the fragment to be sequenced) and sequence both ends of the insert, producing high-quality, alignable sequence data. Because the distance between each paired read is known, alignment algorithms can use this information to more accurately map reads across repetitive regions. This results in better alignment of reads, especially across repetitive regions of the genome that are difficult to sequence. Paired-end sequencing allows detection of rearrangements, such as insertions and deletions (indels) and inversions.

[0115] Paired-end reads can use inserts of different lengths (i.e., different fragment sizes to be sequenced). As a default in the sense of this disclosure, paired-end reads are used to refer to reads obtained from various insert lengths. In some cases, to distinguish short-insert paired-end reads from long-insert paired-end reads, the latter are specifically referred to as mate pair reads. In some embodiments involving mate pair reads, two biotin junction adapters are attached to both ends of a relatively long insert (e.g., several kb). The biotin junction adapters then link both ends of the insert to form a circular molecule. The circular molecule is then further fragmented to obtain subfragments that contain the biotin junction adapters. The subfragments that contain both ends of the original fragment in reverse sequence order can then be sequenced using the same procedure as the paired-end sequencing of short inserts described above. Further details of mate pair sequencing using the Illumina platform are provided in an online publication at the following address, which is incorporated herein by reference in its entirety: res.illumina.com / documents / products / technotes / technote_nextera_matepair_d-ata_processing.pdf

[0116] After sequencing of DNA fragments, sequence reads of a given length, for example 100 bp, are localized by mapping (aligning) to a known reference genome. The mapped reads and their corresponding positions on the reference sequence are also called tags. In another embodiment of this procedure, localization is achieved by k-mer sharing and read-to-read alignment. In the analysis of many embodiments disclosed herein, aligned reads (tags) as well as poorly aligned or unalignable reads are used. In one embodiment, the reference genome sequence is the NCBI36 / hg18 sequence, which is available on the World Wide Web (WWW) at genome.ucsc.edu / cgi-bin / hgGateway?org=Human&db=hg18&hgsid=166260105). Alternatively, the reference genome sequence is GRCh37 / hg19 or GRCh38, which is available on the World Wide Web at genome.ucsc.edu / cgi-bin / hgGateway. Other public sequence sources include GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (The DNA Databank of Japan). Many computer algorithms are available for sequence alignment, including but not limited to BLAST (Altschul et al., 1990), BLITZ (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al., Genome Biology 10:R25.1-R25.10

[2009] ), or ELAND (Illumina, Inc., San Diego, Calif., USA).In one embodiment, one end of the clonally amplified copies of the plasma cfDNA molecules are sequenced and processed by bioinformatics alignment analysis on an Illumina Genome Analyzer using Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) software.

[0117] In one exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information of nucleic acids in a test sample using Helicos' True Single Molecule Sequencing (tSMS) technology, a single molecule sequencing technology (e.g., as described in Harris TD et al., Science 320:106-109

[2008] ). In the tSMS method, a DNA sample is sheared into strands of about 100-200 nucleotides and a polyA sequence is added to the 3' end of each DNA strand. Each strand is labeled by the addition of a fluorescently labeled adenosine nucleotide. The DNA strands are then hybridized to a flow cell, which has millions of oligo-T capture sites immobilized on its surface. In certain embodiments, the templates are approximately 100 million templates / cm. 2The flow cell is then loaded into an instrument, such as a HeliScope™ sequencer, and a laser is shone on the surface of the flow cell to reveal the location of each template. A CCD camera can map the location of the templates on the flow cell surface. The fluorescent labels of the templates are then cleaved and washed away. The sequencing reaction begins by introducing DNA polymerase and fluorescently labeled nucleotides. Oligo-T nucleic acids serve as primers. The polymerase incorporates the labeled nucleotides into the primers as directed by the template. The polymerase and unincorporated nucleotides are removed. Templates that have been directed to incorporate fluorescently labeled nucleotides are identified by imaging the flow cell surface. After imaging, a cleavage step removes the fluorescent label, and the process is repeated with other fluorescently labeled nucleotides until the desired read length is achieved. Sequence information is collected at each nucleotide addition step. Whole genome sequencing by single molecule sequencing techniques precludes or commonly omits PCR-based amplification in the preparation of sequencing libraries; this method allows for a direct measurement of the sample, rather than a measurement of copies of the sample.

[0118] In another exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information of nucleic acids in a test sample using 454 sequencing (Roche) (e.g., as described in Margulies, M. et al. Nature 437:376-380

[2005] ). 454 sequencing typically involves two steps. In the first step, DNA is cleaved into fragments of about 300-800 base pairs and the fragments are blunt-ended. Oligonucleotide adapters are then ligated to the ends of the fragments. The adapters serve as primers for amplification and sequencing of the fragments. The fragments can be attached to DNA capture beads, e.g., streptavidin-coated beads, e.g., using adapter B, which contains a 5'-biotin tag. The fragments attached to the beads are PCR amplified within droplets of an oil-water emulsion. The result is multiple copies of the clonally amplified DNA fragments on each bead. In the second step, the beads are captured in wells (e.g., picoliter-sized wells). Pyrosequencing is performed in parallel on each DNA fragment. The addition of one or more nucleotides generates a light signal, which is recorded by the CCD camera of the sequencing instrument. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing utilizes pyrophosphate (PPi), which is released upon addition of a nucleotide. PPi is converted to ATP by ATP sulfurylase in the presence of adenosine 5' phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, a reaction that generates light, which is measured and analyzed.

[0119] In another exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information of nucleic acids in a test sample using SOLiD™ technology (Applied Biosystems). In SOLiD™ sequencing-by-ligation, genomic DNA is cleaved into fragments and adapters are attached to the 5' and 3' ends of the fragments to create a fragment library. Alternatively, to introduce internal adapters, adapters are ligated to the 5' and 3' ends of the fragments, the fragments are circularized, the circularized fragments are digested to generate internal adapters, and the resulting fragments are fitted with adapters to the 5' and 3' ends to create a mate pair library. A clonal bead population is then prepared in a microreactor containing beads, primers, templates, and PCR components. After PCR, the templates are denatured and the beads are enriched to separate beads with extended templates. The templates on selected beads undergo a 3' modification that allows for attachment to a glass slide. The sequence can be determined by sequential hybridization and ligation of partially random oligonucleotides that contain a central defined base (or pair of bases) that is recognized by a specific fluorophore. After recording the color, the ligated oligonucleotide is cleaved and removed, and the process is then repeated.

[0120] In another exemplary, but non-limiting embodiment, the method described herein includes obtaining sequence information of nucleic acids in a test sample using Pacific Biosciences' Single Molecule Real-Time (SMRT™) sequencing technology. In SMRT sequencing, the sequential incorporation of dye-labeled nucleotides is imaged during DNA synthesis. Single DNA polymerase molecules are attached to the bottom of individual zero mode wavelength detectors (ZMW detectors) to obtain sequence information while phospholinked nucleotides are incorporated into the growing primer strand. The ZMW detectors contain confinement structures that allow observation of the incorporation of single nucleotides by the DNA polymerase against a background of fluorescent nucleotides that rapidly (e.g., in microseconds) diffuse out of the ZMW. Incorporation of a nucleotide into the growing strand typically takes several milliseconds. During this time, the fluorescent label is excited to emit a fluorescent signal and the fluorescent tag is cleaved. Measuring the fluorescence of the corresponding dye tells which base has been incorporated. Repeating this process provides the sequence.

[0121] In another exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information of nucleic acids in a test sample using nanopore sequencing (e.g., as described in Soni GV and Meller A. Clin Chem 53: 1996-2001

[2007] ). Nanopore sequencing DNA analysis technology has been developed by a number of companies, including, for example, Oxford Nanopore Technologies (Oxford, UK), Sequenom, and NABsys. Nanopore sequencing is a single molecule sequencing technology in which a single molecule of DNA is directly sequenced as it passes through a nanopore. A nanopore is a small hole, typically around 1 nanometer in diameter. When the nanopore is immersed in a conducting liquid and a potential (voltage) is applied across it, a small current is generated by the conduction of ions through the nanopore. The amount of current that flows is affected by the size and shape of the nanopore. As a DNA molecule passes through the nanopore, each nucleotide on the DNA molecule blocks the nanopore to a different extent and changes the magnitude of the electrical current through the nanopore to a different extent. This change in electrical current as the DNA molecule passes through the nanopore thus provides a readout of the DNA sequence.

[0122] In another exemplary, but non-limiting embodiment, the methods described herein include obtaining sequence information of nucleic acids in a test sample using a chemical-sensitive field effect transistor (chemFET) array (e.g., as described in U.S. Patent Application Publication No. 2009 / 0026082). In one example of this technique, DNA molecules are placed in a reaction chamber and a template molecule is hybridized to a sequencing primer bound to a polymerase. The incorporation of one or more triphosphates into the new nucleic acid strand at the 3' end of the sequencing primer is identified by the chemFET as a change in current. The array can have multiple chemFET sensors. In another example, a single nucleic acid can be bound to a bead, the nucleic acid amplified on the bead, and individual beads transferred to individual reaction chambers on a chemFET array (each chamber has a chemFET sensor) to determine the sequence of the nucleic acid.

[0123] Ion Torrent PGM™ sequencers (Life Technologies) and Ion Torrent Proton™ sequencers (Life Technologies) are ion-based sequencing systems that determine the sequence of a nucleic acid template by detecting ions produced as a by-product of nucleotide incorporation. Typically, hydrogen ions are released as a by-product of nucleotide incorporation during template-dependent nucleic acid synthesis by a polymerase. Ion Torrent PGM™ sequencers and Ion Torrent Proton™ sequencers detect nucleotide incorporation by detecting hydrogen ions, which are a by-product of nucleotide incorporation. Ion Torrent PGM™ sequencers and Ion Torrent Proton™ sequencers contain multiple nucleic acid templates to be sequenced, each template being placed in a respective sequencing reaction well of an array. Each well of the array is coupled to at least one ion sensor that can detect the release of H+ ions or a change in solution pH that occurs as a by-product of nucleotide incorporation. The ion sensor includes a field effect transistor (FET) coupled to an ion-sensitive detection layer that can sense the presence of H+ ions or a change in solution pH. The ion sensor provides an output signal indicative of nucleotide incorporation, which can be expressed as a voltage change, and the magnitude of this signal correlates with the H+ ion concentration in the respective well or reaction chamber. As different types of nucleotides are sequentially flowed into the reaction chamber, the polymerase incorporates the nucleotide into the extending primer (or polymerization site) in an order determined by the sequence of the template. The incorporation of each nucleotide is accompanied by the release of an H+ ion in the reaction well and a concomitant change in the local pH. The release of the H+ ion is registered by the sensor's FET, which emits a signal indicating that nucleotide incorporation has occurred. Nucleotides not incorporated during the flow of a particular nucleotide do not emit a signal.The amplitude of the signal from the FET can also be correlated with the number of specific types of nucleotides incorporated into the elongating nucleic acid molecule, thereby elucidating homopolymeric regions. Thus, by flowing multiple nucleotides into the reaction chamber and monitoring the incorporation across multiple wells or reaction chambers during sequencer operation, the instrument is capable of simultaneously analyzing the sequence of many nucleic acid templates. Further details of the construction, design and operation of the Ion Torrent PGM™ sequencer can be found, for example, in U.S. Patent Application No. 12 / 002,781, now published as U.S. Patent Publication No. 2009 / 0026082; U.S. Patent Application No. 12 / 474,897, now published as U.S. Patent Publication No. 2010 / 0137143; and U.S. Patent Application No. 12 / 492,844, now published as U.S. Patent Publication No. 2010 / 0282617; all of which applications are incorporated herein by reference in their entirety. In some embodiments, the amplicons can be manipulated or amplified via bridge amplification or emPCR to generate multiple clonal templates suitable for various downstream processes, including nucleic acid sequencing. In one embodiment, the nucleic acid templates sequenced using the Ion Torrent PGM™ or Ion Proton PGM™ system can be prepared from a population of nucleic acid molecules using one or more target-specific amplification techniques outlined herein. Optionally, the target-specific amplification can be followed by a secondary and / or tertiary amplification process, including but not limited to a library amplification step and / or a clonal amplification step such as emPCR. The use of such next-generation sequencers is contemplated herein for rapid characterization of the changes in gDNA libraries, cDNA libraries, and ADT libraries compared to reference sequences at the single-cell level.

[0124] In another embodiment, the method includes obtaining sequence information of nucleic acids in a test sample using sequencing by hybridization. Sequencing by hybridization includes contacting a plurality of polynucleotide sequences with a plurality of polynucleotide probes, where each of the plurality of polynucleotide probes can be optionally tethered to a substrate. The substrate can be a flat surface that includes an array of known nucleotide sequences. The pattern of hybridization to the array can be used to determine the sequence of polynucleotides present in the sample. In other embodiments, each probe is tethered to a bead, such as a magnetic bead. Hybridization to the beads can be confirmed and used to identify the plurality of polynucleotide sequences in the sample.

[0125] In some embodiments of the methods described herein, the sequence reads are about 20bp, about 25bp, about 30bp, about 35bp, about 40bp, about 45bp, about 50bp, about 55bp, about 60bp, about 65bp, about 70bp, about 75bp, about 80bp, about 85bp, about 90bp, about 95bp, about 100bp, about 110bp, about 120bp, about 130bp, about 140bp, about 150bp, about 200bp, about 250bp, about 300bp, about 350bp, about 400bp, about 450bp, or about 500bp. Advances in technology are expected to allow single-end reads of more than 500bp, and reads of more than about 1000bp when paired-end reads are generated. In some embodiments, paired-end reads include sequence reads of about 20bp to 1000bp, about 50bp to 500bp, or 80bp to 150bp, and are used to determine the sequence of interest. In various embodiments, paired-end reads are used to evaluate the sequence of interest. The sequence of interest is longer than the read. In some embodiments, the sequence of interest is longer than about 100bp, 500bp, 1000bp, or 4000bp. Mapping of sequence reads is accomplished by comparing the sequence of the read to the sequence of a reference to reveal the chromosomal origin of the sequenced nucleic acid molecule, and no specific gene sequence information is required. A small amount of mismatch (0 to 2 mismatches per read) can be tolerated to account for minor polymorphisms that may exist between the reference genome and the genome in the mixed sample. In some embodiments, reads that align to the reference sequence are used as anchor reads, and reads that align to the anchor read but cannot align or are poorly aligned to the reference sequence are used as anchored reads. In some embodiments, poorly aligned reads can have a relatively high percentage of mismatches per read, e.g., at least about 5%, at least about 10%, at least about 15%, or at least about 20% mismatches per read.

[0126] Multiple sequence tags (i.e., reads aligned to a reference sequence) are typically obtained for each sample. In some embodiments, at least about 3×10 6 sequence tags, at least about 5 × 10 6 sequence tags, at least about 8 × 10 6 sequence tags, at least about 10 × 10 6 sequence tags, at least about 15 × 10 6 sequence tags, at least about 20 × 10 6 sequence tags, at least about 30 × 10 6 sequence tags, at least about 40 × 10 6 sequence tags, or at least about 50 × 10 6 The sequence tags are obtained by mapping the reads to the reference genome for each sample. In some embodiments, all sequence reads are mapped to all regions of the reference genome to provide genome-wide reads. In other embodiments, the reads are mapped to the sequence of interest.

[0127] Hardware implementation In various aspects, the methods described herein are carried out using a computer-based system configured to execute machine-readable instructions that, when executed by the processor of the system, cause the system to perform steps including determining the identity, size, nucleotide sequence or other measurable characteristics of the amplicons generated by the methods of the invention. One or more features of any one or more of the teachings and / or exemplary embodiments described above may be performed or implemented using appropriately configured and / or programmed hardware and / or software elements. The decision of whether an embodiment is implemented using hardware and / or software elements may be based on a number of factors, such as desired computation speed, power levels, thermal tolerance, amount of processing cycles, input data rate, output data rate, memory resources, data bus speed, etc., and other design or performance constraints.

[0128] Examples of hardware elements may include processors, microprocessors, input(s) and / or output(s) (I / O) devices (or peripherals) communicatively coupled via local interface circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. The local interface may include, for example, one or more buses or other wired or wireless connections, controllers, buffers (caches), drivers, repeaters and receivers, etc., to enable appropriate communication between the hardware components. A processor is a hardware device for executing software, particularly software stored in a memory. A processor may be any custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with a computer, a semiconductor-based microprocessor (e.g., in the form of a microchip or chipset), a microprocessor, or generally any device for executing software instructions. A processor may also represent a distributed processing architecture. I / O devices can include input devices, such as keyboards, mice, scanners, microphones, touch screens, interfaces for various medical and / or laboratory instruments, barcode readers, styluses, laser readers, radio-frequency device readers, etc. Additionally, I / O devices can also include output devices, such as printers, barcode printers, displays, etc. Finally, I / O devices can further include devices that communicate as both inputs and outputs, such as modulators / demodulators (modems; for accessing other devices, systems, or networks), radio frequency (RF) or other transceivers, telephone interfaces, bridges, routers, etc.

[0129] Examples of software include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. The software in memory can include one or more separate programs, which may include an ordered list of executable instructions to implement logical functions. The software in memory can include a system for identifying data streams according to the present teachings, and a suitable operating system (O / S), custom or commercially available, that controls the execution of other computer programs such as the system, and provides scheduling, input / output control, file and data management, memory management, communication control, and the like.

[0130] According to various exemplary embodiments, one or more features of any one or more of the above teachings and / or exemplary embodiments may be performed or implemented at least in part using distributed, cluster, remote, or cloud computing resources.

[0131] According to various exemplary embodiments, one or more features of any one or more of the teachings and / or exemplary embodiments described above may be performed or implemented using a source program, an executable program (object code), a script, or other entity that includes a set of instructions to be executed. When using a source program, the program may be translated via a compiler, assembler, interpreter, etc., which may or may not be included in memory, to operate appropriately in conjunction with an O / S. These instructions may be written using (a) an object-oriented programming language with classes of data and methods, or (b) a procedural programming language (including, for example, C, C++, Pascal, Basic, Fortran, Cobol, Pert, Java, and Ada) with routines, subroutines, and / or functions.

[0132] According to various exemplary embodiments, one or more of the exemplary embodiments described above may include transmitting, displaying, storing, printing, or outputting to a user interface device, a computer-readable storage medium, a local computer system, or a remote computer system any information, signals, data, and / or information related to intermediate or final results that may have been generated, accessed, or used by such exemplary embodiments. Such transmitted, displayed, stored, printed, or outputted information may take the form of, for example, searchable and / or filterable lists of runs and reports, images, tables, charts, graphs, spreadsheets, correlations, sequences, and combinations thereof.

[0133] The practice of the present invention employs, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry and immunology well within the skill of the art. Such techniques are fully explained in such references as: "Molecular Cloning: A Laboratory Manual", second edition (Sambrook, 1989); "Oligonucleotide Synthesis" (Gait, 1984); "Animal Cell Culture" (Freshney, 1987); "Methods in Enzymology" "Handbook of Experimental Immunology" (Weir, 1996); "Gene Transfer Vectors for Mammalian Cells" (Miller and Calos, 1987); "Current Protocols in Molecular Biology" (Ausubel, 1987); "PCR: The Polymerase Chain Reaction", (Mullis, 1994); "Current Protocols in Immunology" (Coligan, 1991). These techniques are applicable to the production of the polynucleotides and polypeptides of the invention and therefore may be considered in making and practicing the invention. Techniques that are particularly useful for particular embodiments will be described in the following sections.

[0134] The following examples are intended to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the assays, screening and treatment methods of the present invention, and are not intended to limit the scope of what the inventors regard as their invention. EXAMPLES

[0135] MINECRAFTseq is a single-cell multi-omics approach that captures DNA amplicons, 3'mRNA transcripts, antibody-derived tags (ADTs), and index flow sorting information from CRISPR-edited and sorted cells. The method can be applied to cell lines and primary blood cells, particularly B and T cells, to simultaneously examine the effects of CRISPR editing and the consequences on RNA and cell surface expression. It is a highly adaptable technology that can be used with any cell line and primary edited cells. It is also highly modular and scalable with automated liquid handling.

[0136] The technique relies on sorted and pooled cells to control for plate and sample effects. The technique can also be applied to low inputs of less than 1000 cells for bulk multi-omics estimation. Importantly, previous techniques capturing DNA and mRNA have not been applied to CRISPR-edited cells, but rather focused on cancer-related heterogeneity and applications.

[0137] As described in the Examples below, MINECRAFTseq was used to characterize CRISPR editing in: a. CRISPR-Cas9 cleavage in cell lines can be used to dissect regulatory regions. b. CRISPR-Cas base editing (BE4) can be performed in cell lines to test for functional variants. c. CRISPR-Cas base editing can be performed in primary cells to study gene knockout. d. CRISPR-Cas base editing can be performed in primary cells to investigate autoimmune variants. e. CRISPR-Cas base editing can be multiplexed in primary cells. f. CRISPR-Cas HDR can be performed in primary cells.

[0138] The MINECRAFTseq method can start with CRISPR-edited cell lines or primary cells and relies on sorting single cells using a FACS sorter such as the ARIA II (BD) into either 96- or 384-well plates for further processing into sequencing libraries. Plate processing can be automated using liquid handling platforms to reduce volume.

[0139] The cell lines obtained from ATCC are nucleofected with CRISPR reagents to induce genomic perturbation. The exact conditions for these nucleofections are highly dependent on the cells selected and the reagents used. To date, the inventors have successfully performed genome editing in HH and Jurkat T cell lines using CRISPR-Cas RNPs (ribonucleoproteins) in the presence or absence of single-stranded oligo donors for HDR repair and base-editing Cas mRNA. For primary immune cells, the inventors used the same set of reagents to perform genome editing in primary T cells.

[0140] After the cells are genome edited under various conditions, they are subjected to MINECRAFTseq (a protocol for single-cell sorting to prepare single-cell libraries for sequencing). This protocol is divided into six sections.

[0141] Briefly, cells are labeled with antibodies, single cell index sorted onto plates, and lysed in the presence of proteases. Reverse transcription with template switch oligos is performed to convert the mRNA to cDNA and append well-specific barcodes and UMIs. At this stage, the cDNA is amplified along with the ADT and specific genomic DNA in one large pool per well. After amplification, a sample of the product is used for further amplification with barcoded nested primers to append well-specific identifiers. At this point, the DNA products are pooled, cleaned up, amplified using barcoded Illumina-specific P5 / P7 primers per plate, pooled, and ready for sequencing. The remainder of the cDNA / ADT / DNA amplified products can be used to isolate the cDNA and ADT using solid-phase reversible immobilization (SPRI) cell size exclusion. The ADTs are then amplified once more using barcoded Illumina-specific P5 / P7 primers per plate, pooled, and ready for sequencing. The cDNA is first tagmented using NexteraXT Tn5, and then the 3' ends are preferentially amplified using custom Illumina-specific P5 / P7 barcoded primers per plate, pooled, and ready for sequencing.

[0142] Example 1: Multiomic Investigation of Nucleotide Editing by CRISPR with ADT, Flow cytometry and Transcriptome sequencing (MINECRAFT-seq) To highlight the utility of high-dimensional multimodal targeting approaches in analyzing the outcomes of CRISPR editing, we applied our modified TARGETseq protocol to a single plate of 96 HH cells edited with a previously validated CRISPR-Cas9 nuclease targeting the upstream regulatory region of HLADQB1 (Figure 1A). In these 96 cells, we recovered genomic DNA and mRNA pairs from 68 samples, filtering for at least 10 aligned genomic DNA reads per cell, more than 300 mRNA genes per cell, and less than 10% mitochondrial gene reads. When we analyzed genomic DNA amplicons around the target sites from these single cells, we observed a very large heterogeneity of genome editing. A total of 29 unique genotypes were observed, which could be grouped into five distinct clusters (Figure 1B and 1C). We then investigated mRNA expression levels in individual cells by calculating the number of unique molecular identifiers per gene and using STARSolo to enable barcode error correction. HLADQB1 expression normalized per cell and scaled across all cells differed for different clusters of DNA editing (p=4.05e-05, ANOVA). After adjusting for 21815 tested genes, HLADQB1 was observed to correlate with the average deletion size (p<4.58e-07=0.01 / 21815, Figures 1D-1G and 5A-5E). None of the other genes were observed to correlate with deletion size, demonstrating that this single-cell assay can accurately identify target genes in a limited number of cells.

[0143] Cell surface protein antibody assays are recognized as essential to understand protein expression and characterize cellular states. In the case of immune cell populations, phenotyping cells using antibody-derived tags has proven essential to unravel single-cell immune populations. To improve current plate-based methodologies and incorporate this crucial modality, several multi-omics protocols were tested (Figures 6A-6G). Using a CRISPR-Cas base editor (BE4-NG), we targeted variant rs61839660 in IL2RA and investigated the recoverability of all three modalities: genomic DNA amplicon, 3' mRNA, and antibody-derived tag (ADT) expression. Optimization showed that MaximaH, KapaHIfI, and a Smart-seq3-like approach using increased template switch oligo (TSO) concentrations resulted in good recoverability of all three modalities, as measured by the number of recovered and aligned genomic DNA reads per cell, unique molecular identifiers (UMIs) of ADTs per cell, and UMIs of genes per cell (Figures 6A-6G). When all data were pooled together, no correlation was observed between editing at rs61839660 and expression of IL2RA or CD25 ADTs (Figures 6A-6G and 7A-7D).

[0144] Interrogation of disease loci and variants in primary cells of interest represents the next frontier in functional genomics. To investigate the effects of CRISPR editing in primary immune cells, we integrated all three single-cell modalities, genomic DNA, ADT, and 3'mRNA, with index sorting to create a plate-based single-cell multi-omics approach called MINECRAFT-seq (Multiomic Investigation of Nucleotide Editing by CRISPR with ADT, Flow cytometry and Transcriptome sequencing) (Figures 2A, 8A, and 8B). Critically, this methodology allows for mixing of conditions, reducing batch effects, and does not require expensive droplet generation equipment.

[0145] Example 2: Multiomic Investigation of Nucleotide Editing by CRISPR with ADT, Flow cytometry and Transcriptome sequencing (MINECRAFT-seq) applied to CD4 T cells To demonstrate the utility of this approach in primary CD4 T cells, we used a CRISPR-Cas base editor to induce a premature stop codon in the PTPRC and processed 960 cells from one healthy individual using single-cell MINECRAFT-seq (Figure 2A-2K). Genomic DNA was filtered for at least 10 reads per cell and aligned to a reference amplicon sequence using CRISPR Resso2. We aligned and calculated mRNA counts using STARSolo and ADT counts using kallisto KITE. For comparison, bulk analysis from an additional healthy individual was also performed. We observed that CRISPR base editing in the PTRPC caused a substantial knockout of CD45, skewing the cells towards late activation, but did not cause changes in global gene expression (Figure 2B-2D). Using the single-cell multi-omics approach, eight unique genotypes (comprising at least four cells) were identified, including bystander outcomes (Figure 2E). Of the edited cells, some had no edits introduced and some had the intended GGG genotypic change. These unique genotypes correlated with varying levels of CD45 protein expression, confirming a clear heterozygote effect (p<2e-16, ANOVA, Figures 2F and 2G). Analysis of all ADT markers revealed clearly distinct knockout clusters (Figures 2H, 2I, and Figures 9A-9C). Furthermore, a strong correlation was shown between flow-based cell surface expression and sequencing reads (Figures 9A-9C). Clustering of mRNA, unlike ADT, did not identify unique knockout clusters, which were only supported by a small, non-significant decrease in PTPRC (Figures 2J-2K and Figures 10A-10D). Differential gene expression at gene doses of the targeted base (comparing genotypes A, C, and B) revealed more widespread expression changes, suggesting subtle changes in cell state that were not identified in the bulk data (Figures 11A-11D).

[0146] Example 3: Application of Multiomic Investigation of Nucleotide Editing by CRISPR with ADT, Flow cytometry and Transcriptome sequencing (MINECRAFT-seq) to investigate disease-causing variants Capturing gDNA, cell surface protein expression, and mRNA from single cells was useful in revealing heterogeneity in CRISPR editing and inferring phenotypic outcomes. This provided a great opportunity to study disease-associated variants directly in the primary cells of interest. For many autoimmune diseases, that cell type is CD4 T cells. In a recent study, fine mapping of autoimmune loci identified potential causative variants common to type 1 diabetes and rheumatoid arthritis, specifically in two loci, UBASH3A and IL2RA. Four variants in UBASH3A and three variants in IL2RA were selected for functional follow-up using single-cell MINECRAFT-seq in primary CD4 T cells.

[0147] UBASH3A is a ubiquitin-associated protein that appears to regulate T cell simulta- tion via the T cell receptor (TCR). Knocking out Ubash3a enhances signaling capacity and increases proliferation and IL-2 expression. To investigate all four potential causative variants in UBASH3A, CRISPR-Cas base editing and HDR repair tools were applied (Figure 3A, Figure 8A and 8B).

[0148] Analysis of bulk genomic DNA, mRNA, and flow cytometry data confirmed evidence of differences in editing efficiency and predominant indels (insertions and deletions) in HDR-targeted samples (Figures 12A-12J). Targeting only rs11203203 and not other variants resulted in a nominal increase in UBASH3A expression (FC=0.83, p=0.1791 Welch t-test), but this was not significant after genome-wide correction (Figures 12A-12J).

[0149] Single-cell MINECRAFTseq provided a clearer picture of editing efficacy in both base-edited and HDR-edited variants. Single-cell genomic DNA sequencing identified substantial bystander editing in base-edited cells and clearly distinct clusters of indels in HDR-edited cells (Figure 3B-3E). HDR editing was successful at both rs9981624 and rs11203202, but was incredibly rare, with insertions dominating the editing around rs11203202 and deletions dominating the editing at rs9981624 (Figure 13A-13D). The advantage of this method is that we can take advantage of this diverse and heterogeneous editing to determine the effects on gene and protein expression. Differential gene expression in base-edited cells, modeling gene dosage of targeted single nucleotide polymorphisms (SNPs) along with plates, revealed expression changes of RIPK1 and NPAS1 associated with rs11203203 (Figure 3F-3H). Similarly, in HDR edited samples, editing near rs11203202, but not rs9981624, caused upregulation of several genes, including IL2RA (Figure 3I-3K), providing a possible link to proliferation and IL-2 signaling. There was no effect on mRNA clustering or ADT expression (Figure 14A-14L and 15A-15N). Overall, using this single-cell multi-omics approach with limited numbers of samples and cells, we found that two of the variants in UBASH3A were causative and could affect cell proliferation and IL2 signaling. Finally, we applied single-cell MINECRAFTseq at the IL2RA locus in primary CD4 T cells. Previous computational work identified three potentially causative variants, two of which (rs706778 and rs3118470) had high TBET IMPACT scores and were selected for functional follow-up. Another previously validated variant, rs61839660, was also selected for investigation as it has been shown to have potent activation-dependent effects in regulating T cell differentiation.We recruited three individuals with unique genotypes for the three variants of interest and targeted each variant individually or as one large multiplexed pool with CRISPR-Cas base editors (Figures 4A and 4B). We selected combinations of base editors with different genotypes to examine the effects on heterozygotes and non-targeted edited sites.

[0150] Flow cytometry, genomic DNA, and bulk analysis of mRNA revealed that multiplex editing with various base editors is possible, but that the A to T base editor is preferentially used (Figures 16A-16C). Bulk flow cytometry data showed consistent increases in rs61839660 and CD25 expression from multiplex samples, whereas bulk mRNA sequencing did not (Figures 16A-16C).

[0151] Single-cell MINECRAFTseq identified many unique genotypes in various combinations (Figures 4C and 17). As expected, targeting heterozygous individuals could be used to convert them to homozygotes. Considering the wide range of induced mutations, all of the targeted nucleotides fell within the region of interest (labeled SNP1–SNP18) (Figure 4C). Using these labels, we investigated which targeted nucleotides correlated with CD25 ADT expression using a linear regression framework that considered plate effects (Figures 18A and 18B). We found that targeting SNP3 (hereafter named multiplex SNP), but not any of the variants examined, had the strongest effect on CD25 expression. When this analysis was performed conditional on SNP3, a quadratic effect at rs61839660 was confirmed (Figures 18A and 18B). Using additional linear models, significant changes in CD25, CD27, CD38, and CD48 were found to be associated with the multiplex SNPs, with only CD25 associated with rs61839660 (Figures 4D and 4E). Differential gene expression was performed using pooled single-cell genomic DNA and mRNA to confirm correlations of CTIF and RORA with gene dosage at rs61839660 (conditional on the multiplex SNPs), suggesting a possible link to Tregs (Figures 4F and 4G). Similarly, gene dosage at the multiplex SNPs correlated with expression of MAPK6, ACYP2, and DARS2 genes (Figures 4H and 4I). In summary, using a single-cell multi-omics approach, we experimentally interrogated three variants in primary CD4 T cells and revealed clear effects at rs61839660 and its neighboring nucleotides, helping to further elucidate this locus and its effects.

[0152] As the list of loci and variants implicated in disease continues to grow, scalable and flexible methods must be used to investigate potential causal relationships. Herein, a highly customizable multi-omics approach is provided that applies to numerous genome editing techniques at several disease loci in cell lines and primary immune cells to find causal variants, regions, and genes. By capturing and integrating four single-cell modalities: index flow cytometry, genomic DNA amplicons, 3' mRNA sequencing, and ADT sequencing, single-cell MINECRAFTseq presents a holistic view of genome editing and paves the way for future studies linking variants to function in primary immune cells.

[0153] In the above examples, the following materials and methods were used.

[0154] Cell line cultivation and genome editing HH skin T cell line (ATCC: CRL-2105) and Jurkat E6-1 (ATCC: TIB-152) were cultured in complete RPMI, which is RPMI 1640 supplemented with 10% heat-inactivated FBS, 1% non-essential amino acids, sodium pyruvate, HEPES, L-glutamine, penicillin-streptomycin (Penn-Strep), and 0.1% β-mercaptoethanol. To investigate the regulatory region around HLADQB1, the sgRNA closest to the validated single nucleotide polymorphism (SNP) of interest was selected and HH cells were genomically modified as previously described (cited Nat Gen, Maria and Yuriy). Briefly, 40 μM Cas9 protein (QB3 Microlabs) was mixed with an equal volume of 40 μM modified sgRNA (Synthego) and incubated at 37°C for 15 min to allow ribonucleoprotein (RNP) complex formation. 2 μL of RNP was nucleofected into HH cells using an Amaxa 4D nucleofector (SE protocol: CL-120). The cells were immediately transferred to a 24-well plate with pre-warmed medium and cultured. After 10 days, the cells were single-cell sorted on a BD FACS ARIA II into 96-well plates for processing according to the modified TARGETseq protocol. For Jurkat cells targeting the single nucleotide polymorphism (SNP) rs61839660, 1 μl of mRNA (2 μg / μl) encoding the base editor BE4-NG was mixed with 1 μl of 40 μM modified sgRNA (Synthego) targeting the variant of interest. 2 μl of the mRNA / sgRNA mixture was then nucleofected into Jurkat cells using an Amaxa 4D nucleofector (SE protocol: CL-120). Cells were incubated in 24-well plates as above for 7 days and then stimulated with anti-CD3 / anti-CD28 microbeads (ThermoFisher) at a 1 bead:1 cell ratio for 18 h. After stimulation, Jurkat cells were stained with ADT antibody, single-cell sorted, and processed with one of four optimized protocols. ADT staining of Jurkat cells was performed similarly to the staining of primary CD4 T cells described below.

[0155] Recruitment of healthy volunteers, isolation of PBMCs, and magnetic sorting of CD4 T cells To investigate selected variants using the single-cell MINECRAFTseq protocol, healthy subjects were recruited and 40–50 ml of peripheral blood was processed according to an IRB-approved protocol (IRB# 2008P000427). PBMCs were isolated by layering Ficoll Paque (Sigma-Aldrich) under blood diluted 1:1 with PBS, followed by centrifugation. The buffy coat layer was extracted, washed with PBS, and then resuspended in XVIVO15 medium (Lonza) supplemented with 5% FBS (Gemini Bio), 55 μM 2-mercaptoethanol (Sigma), and 10 mM N-acetyl-L-cysteine ​​(Sigma); hereafter this medium is referred to as cVIVO15. At this stage, 500,000 cells were harvested for DNA isolation and Sanger sequencing of selected variants. The rest of the cells were stored by adding an equal volume of freezing medium (10% DMSO and 50% FBS in xVIVO15) and frozen in liquid nitrogen until use. All healthy subjects were aged 20-40 years and had no reported autoimmune diseases. To isolate CD4 T cells, frozen PBMCs were rapidly thawed and immediately placed in warm xVIVO15. Cells were washed twice and isolated using a magnetic negative selection kit (Miltenyi, CD4 + Total CD4 T cells were isolated using the CD4 T cell Isolation Kit (human) according to the manufacturer's protocol.

[0156] The isolated cells were then placed in 96-well U-bottom plates at a concentration of 2.5 million / 250 μl of cXVIVO 15 mL overnight until use.

[0157] Sanger sequencing of selected variants Before CRISPR editing, healthy individuals were genotyped at the targeted variants by Sanger sequencing. Genomic DNA was isolated using a Qiagen DNA extraction kit according to the manufacturer's protocol. Then, 200 bp–1 kb fragments surrounding the variants of interest were amplified with custom PCR primers and Sanger sequenced (Eurofins Genomics). Sequence chromatograms were analyzed in SNAPGENE (v4.3.6) and genotyped based on the distribution of the variants of interest.

[0158] CRISPR editing of primary CD4 T cells Isolated CD4 T cells were rested overnight and then stimulated with anti-CD3 and anti-CD28 Dynabeads (ThermoFisher) at a 1:1 (cells:Dynabeads) ratio in 48-well plates in the presence of 5ng / ml rhIL-2 (Biolegend). After 2 days, cells were removed and the Dynabeads were removed using a magnet before proceeding with CRISPR editing. To interrogate regions and variants of interest, we used the CRISPR-Cas9 C to T base editor (BE4-NG), A to G base editor (ABE8e-NG), or CRISPR-Cas9-mediated HDR repair. For base editors, 1 μl of mRNA (2 μg / μl) encoding the modified Cas9 protein complexed with 1 μl of sgRNA (40 μM, Synthego) was nucleofected into 500,000 stimulated CD4 T cells using an Amaxa 4D nucleofector (P3 protocol: EH-115). For HDR repair induction, 2 μl of Cas9 RNP and 1 μl of asymmetric ssDNA donor were nucleofected into 500,000 CD4 T cells using an Amaxa 4D nucleofector (P3 protocol: EH-115). After nucleofection, cells were transferred to 48-well plates and cultured in cXVIVO15 medium supplemented with 5 ng / ml rhIL-2 until use.

[0159] Cell staining with indexing and oligoconjugated antibodies Nucleofected samples were cultured for an additional 7 days before cells were detached, washed, and resuspended in FACS buffer (2% FBS in PBS with EDTA). An aliquot of cells was then used for bulk RNA, DNA, and flow cytometry analysis. The remainder of the sample (200,000 cells) was stained with fluorophore-conjugated antibodies for 20 min on ice. After staining, cells were washed, counted, and different conditions per sample were pooled together. Samples were then spun down and stained with oligo-conjugated antibodies as described previously. Briefly, 1 million cells were resuspended in FC blocking solution (Biolegend) for 5 min at room temperature, followed by staining with room temperature antibodies for an additional 15 min. Samples were then stained with cold antibody mix for 20 min on ice and washed before proceeding to single cell sorting.

[0160] Bulk RNA and DNA isolation Cells were pelleted, resuspended in RLT+ buffer (Qiagen), flash frozen on dry ice, and stored at -80 until processing. For RNA / DNA isolation, samples were thawed, vortexed, and incubated at room temperature for 5 min before proceeding with RNA / DNA isolation using the Qiagen RNA / DNA extraction kit according to the manufacturer's protocol. After isolation, RNA and DNA concentrations were measured by spectrophotometer (Nanovue) and stored at -20 until use. For bulk mRNA sequencing, samples were converted to cDNA using custom oligo DT and TSO primers in an optimized reverse transcription reaction, followed by amplification with 10 cycles of PCR. Next, tagmentation was performed on full-length cDNA using Nextera XT reagents, and custom Illumina adapters were added to amplify 3' mRNA products for sequencing. For DNA, sequential nested PCR reactions were used to amplify 200-400 bp genomic regions of interest and custom Illumina adapters were added for sequencing. Bulk samples were sequenced in parallel with the single-cell libraries.

[0161] Bulk Flow Cytometry Stimulated and genome-edited CD4 T cells were assayed by flow cytometry for expression of key protein markers using a panel of fluorophore-conjugated antibodies 7 days after nucleofection. For all samples, cells were detached, washed twice with PBS, and Fc receptors were blocked with FcX True Stain (Biolegend) for 15 minutes on ice, followed by staining with directly-conjugated antibodies for 30 minutes on ice. Cells were then washed and samples were analyzed on a BD LSR Fortessa. All data were processed using FlowJo and analyzed in GraphPad PRISM.

[0162] Improved TARGETseq - sequencing of genomic DNA and mRNA from HH cells HH cells to be genome edited were processed with a modified TARGETseq approach that allows for increased multiplexing and capture of genomic DNA amplicons and mRNA. Cells were edited as above, washed twice, filtered through 40 μM, and single-cell sorted using a FACS ARIA II into 96-well plates containing lysis buffer and well / cell barcoded oligo-DT primers. After sorting, plates were spun down and incubated at room temperature for 5 min before flash freezing on dry ice and storing at -80 until use. After thawing, plates were incubated at 72°C for proteinase inactivation and cDNA synthesis was performed. After synthesis, cDNA was amplified for 22 cycles using SeqAMP PCR reagents in the presence of genomic DNA-specific primers targeting the targeted HLADQB1 region. After amplification, 1 μl of the product was taken and genomic DNA was further amplified using nested HLADQB1 primers with well / cell specific barcodes. The remaining cDNA amplification products were subjected to solid-phase reversible immobilization (SPRI) clean at 0.65X (beads:sample) to purify full-length cDNA, followed by tagmentation using Nextera XT reagents and the addition of custom Illumina adapters to amplify 3' mRNA products for sequencing. After nested PCR of genomic DNA, samples were pooled per column (8 reactions) and further amplified with custom Illumina primers with indexing barcodes for sequencing. 3' mRNA and DNA libraries were SPRI cleaned at 1X, concentration was measured using QuBit (ThermoFisher) with the 1X HS DNA kit, distribution was examined with a D1000 Agilent TScreenTape, and sequencing was performed at either the Broad Institute's Genomic Platform or the Dana-Farber Cancer Institute's (DFCI) Molecular Biology Core Facilities (MBCF).

[0163] optimization Jurkat cells to be genome edited were processed using four different protocols. As before, cells were stained with edited oligo- and fluorophore-conjugated antibodies, sorted into PCR plates containing lysis buffer, and stored until use. For library generation, plates were incubated at 72°C to inactivate proteinase and perform cDNA synthesis. After synthesis, cDNA was amplified for 20 cycles using one of four reaction conditions in the presence of gDNA-specific primers targeting the IL2RA region. After amplification, an aliquot of the product was taken and genomic DNA was further amplified using nested IL2RA primers with well / cell-specific barcodes. The remainder of the product was solid-phase reversible immobilization (SPRI) cleaned at 0.65X (beads:sample) for size selection and full-length cDNA purification. The flow-through was collected and SPRI cleaned again at 2X to isolate the ADT fraction. The full-length cDNA was then tagmented as before and amplified with custom Illumina adapters for sequencing. The ADT fraction was amplified by PCR using custom Illumina adapters for 10 cycles, and library concentration and distribution were determined as before before proceeding to sequencing.

[0164] MINECRAFTseq Primary CD4 T cells to be genome edited were stained with oligo- and fluorophore-conjugated antibodies and sorted into PCR plates containing 2.1 μl of lysis buffer. Plates were stored at -80 °C until use. For library generation, plates were thawed and incubated at 72 °C to inactivate proteinases. An additional 2.9 μl of cDNA synthesis mix was then added to each well containing MaximaH RT enzyme and custom buffer with GTP and PEG (details in Supplementary Table). After first strand synthesis, an additional 7.5 μl of PCR mix was added to amplify cDNA, ADT, and genomic DNA. Specific genomic DNA primers targeting the variant of interest were added. After 20 cycles of amplification, 0.5 μl of product was taken and genomic DNA was further amplified using nested primers containing well / cell specific barcodes. After nested genomic DNA barcode addition, samples were pooled per plate, purified with 1X SPRI, and DNA was quantified with QuBit. 5ng of product per plate was then amplified using custom Illumina compatible primers and cleaned with 1X SPRI before being subjected to sequencing. The remainder of the cDNA products were pooled per plate and SPRI cleaned at 0.65X for size selection to purify full-length cDNA. The flow-through was collected and SPRI cleaned again at 2X to isolate the ADT fraction. cDNA concentration was measured with QuBit and 0.5ng was used for tagmentation using the NexteraXT kit (Illumina). After tagmentation, the 3' ends of the cDNA molecules were amplified using custom Illumina compatible primers. After amplification, the PCR products were cleaned with 1X SPRI reagent before being subjected to sequencing. The ADT fraction was quantified with Qubit and 5ng of product was used for subsequent amplification with custom Illumina primers. Again, the final ADT product was purified with 1X SPRI and quantified before sequencing. In experiments involving multiple conditions per healthy individual, all relevant conditions were indexed with fluorophore-conjugated antibodies and pooled into one sample before sorting.Each plate sorted and processed was a mix of conditions to reduce batch effects. In the UBASH3A experiment, HDR and BE conditions were pooled and processed separately. In the IL2RA experiment, all conditions were pooled and processed together. During amplification of genomic DNA, all regions within the pool were amplified in the same reaction using multiple specific nested primer sets.

[0165] Illumina Deep Sequencing DNA concentration of samples was quantified with QuBit and distribution was measured with TapeStation (Agilent) before sequencing at the Broad Institute Genomics Platform or the Dana-Farber Cancer Institute (DFCI) Molecular Biology Core Facilities (MBCF). All DNA amplicon libraries were sequenced on a MiSeq for 300 cycles. ADT amplicon libraries were sequenced on a MiSeq for 50 cycles or pooled with RNA libraries and sequenced on a NextSeq550 for 75 cycles or on a NovaSeq6000 for 100 cycles.

[0166] Sequencing data processing and statistical analysis I7 and I5 demultiplexed raw fastqs were merged across different lanes and runs using custom bash scripts. For analysis of DNA amplicon data, fastqs were further demultiplexed based on cell barcodes using custom bash scripts. Fastqs were then aligned to reference sequences using CRISPResso2. Deletion statistics, allele usage, and nucleotide modification frequencies were calculated and imported from CRISPResso2 output into a custom R analysis and graphing pipeline. For all DNA analyses, cells were filtered with a minimum of 10 aligned reads. To identify cell genotypes, expressed alleles were filtered with at least 10 reads and at least 10% of the total alleles recovered. Cells with more than two alleles after filtering were excluded from the analysis, assuming dizygosity. At this stage, if only one allele was recovered, the cell was assumed to be homozygous. Heatmaps of DNA editing were generated from nucleotide modification tables created with CRISPResso2. Total nucleotide modification frequencies (including substitutions, deletions, and insertions) per nucleotide were utilized to generate heatmaps using the complexHeatmap package. Frequencies were binned into three groups for visualization: <0.3, 0.3–0.7, and >0.7, including reference (0), heterozygous (0.5), and homozygous (1) edits. In these analyses, insertions were quantified as affecting both nucleotides at the insertion site (Figures 3B–3E). Clustering of DNA edits was performed using supervised k-means clustering. For analysis of ADT, kallisto KITE was used. References were created based on the barcodes used per experiment and the aligned ADT sequences. UMI counts were calculated, imported into R, and CLR normalized using a custom R function.For visualization, PCA was performed on the CLR-normalized and scaled variable ADT, followed by plate correction with Harmony and Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction on the harmonized PCs. Linear modeling of ADT was performed using the lm function in R, and significance was calculated using anova against a null model. For RNA analysis, gene counts were obtained using STARSolo. A reference was created from the human GRCh38 transcriptome, and reads were mapped with custom barcodes and UMI lengths. The resulting count matrices were imported into R and processed with Seurat. In all experiments, cells were filtered for at least 300 genes, 500 UMIs, and less than 10% mitochondrial reads. For visualization, PCA was performed on the variable genes, followed by batch correction with Harmony and dimensionality reduction with UMAP. Differential gene expression was performed for expressed genes (cells >30% with non-zero expression) using DESeq2, which models plate effects. Normalized and scaled counts were used for visualization. In some experiments, we noticed that certain cell barcodes resulted in inaccurate RNA mapping. In UBASH3A experiments, we excluded all cells with cell barcode CTGGTTCTGTTG. In IL2RA experiments, we excluded all cells with cell barcode GCTACTCCAGTT. For index flow cytometry analysis, raw fluorescence values ​​were imported into R and processed using the ggcyto package. Using index flow cytometry information, we also excluded rare cells with the non-reference genotypes quantified above from the control non-targeted condition.

[0167] Protocol Changes In some experiments, MINECRAFTseq was performed as described in Figures 19A and 19B. In the MINECRAFTseq method described in Figures 19A and 19B, cells are lysed in the presence of a capture oligomer (capture oligo) that contains a capture sequence (CS) and a well-specific barcode. Cells are also lysed in the presence of an oligoDT primer. The capture oligomer contains a blocking factor that prevents the capture oligomer from being degraded by ExoI. After amplification of cDNA, ADT, and specific genomic DNA, the unblocked single-stranded DNA oligomer is digested with ExoI. After ExoI digestion, a nested primer for the specific genomic DNA is added to each well and an additional PCR is performed. One of the primers for the specific DNA contains a capture sequence, so that the capture sequence is added to the amplification product; as a result, the capture oligomer is finally ligated to the amplicon generated using the nested primer for the specific genomic DNA. This improved version of the MINECRAFTseq method has the advantage that all PCR amplifications following library preparation can be carried out in a single well, thereby simplifying the MINECRAFTseq method. Non-limiting examples of blocking agents include phosphoryl and acetyl groups. In some embodiments, the blocking agent is covalently attached to the 3'OH group of the capture oligomer. In some embodiments, the capture sequence is a unique sequence that is found less than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 100 times in the genome of the target cell.

[0168] Other Aspects From the foregoing description, it will be apparent that variations and modifications can be made to adapt the invention described herein to various usages and conditions, such variations also falling within the scope of the following claims.

[0169] The recitation of a list of elements in a definition of a variable herein includes definitions of that variable as a single element or as a combination (or subcombination) of the listed elements. The recitation of an embodiment herein includes that embodiment as a single embodiment or in combination with other embodiments or portions thereof.

[0170] All patents and publications mentioned in this specification are herein incorporated by reference to the same extent as if each individual patent or publication was specifically and individually indicated to be incorporated by reference.

Claims

1. 1. A method for simultaneously characterizing genomic DNA and mRNA of a single cell, comprising: (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest; (b) incubating the detectably labeled cells of (a) with an oligoconjugated antibody; (c) index sorting the cells into a single well, characterizing the expression of cell surface markers for each cell, and lysing the cells in the presence of dNTPs and a well-specific barcoded oligoDT primer comprising a unique molecular identifier (UMI) and a PCR handle; (d) incubating the product of (c) with a reverse transcriptase and a custom template switch oligo (TSO) that comprises one member of a binding pair under conditions that allow for the generation of cDNA; (e) incubating the product of step (d) with a genomic primer that specifically binds to a region of interest (ROI), a cDNA amplification primer that specifically binds to the PCR handle and a cDNA amplification primer that specifically binds to the TSO, an antibody derived tag (ADT) specific primer, dNTPs, a capture oligo, and a polymerase under conditions that support amplification, thereby co-amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library; the capture oligo comprises a capture sequence, a well-specific barcode, an exonuclease blocking agent, and a UMI; (f) incubating at least a portion of the genomic DNA from each well of (e) with dNTPs, a polymerase, and nested primers that specifically bind to the regions of interest to obtain a gDNA library; At least one of the nested primers i) Well-specific barcodes, UMIs, and PCR handles, or ii) Capture Sequence Including, If the nested primer comprises the capture sequence, this step further comprises incubating the product of (e) with an exonuclease and a capture oligo; the capture oligo comprises the capture sequence, a well-specific barcode, an exonuclease blocking agent, and a UMI; The capture oligo binds to the amplicon generated using the nested primer, effectively labeling the product with the barcode during the PCR reaction. Process; (g) after step (e) or step (f), pooling at least a portion of the samples from each well, followed by separating at least two of the cDNA library, the ADT library, and the gDNA library; and (h) preparing a gDNA library, a cDNA library, and an ADT library for sequencing by amplifying each library in the presence of a sequencing primer; The method comprising:

2. 1. A method for simultaneously characterizing DNA amplicons, 3' mRNA transcripts, antibody derived tags (ADTs), and index flow sorting information from a cell sample, comprising: (a) labeling a plurality of cells with a detectable antibody that specifically binds to a cell surface marker of interest and single cell index sorting the cells into individual wells; (b) lysing the cells in the presence of reverse transcriptase, template switch oligo, a well-specific barcode, a primer comprising an oligoDT primer comprising a unique molecular identifier (UMI) and a PCR handle, and ADT, under conditions that allow reverse transcription to obtain cDNA; (c) amplifying the cDNA, ADT, and specific genomic DNA in a single pool containing a genomic primer that specifically binds to the region of interest, a cDNA amplification primer that specifically binds to the PCR handle, and a cDNA amplification primer that specifically binds to the TSO, an ADT-specific primer, dNTPs, capture oligos, and Taq polymerase, thereby simultaneously amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library; the capture oligo comprises a capture sequence, a well-specific barcode, an exonuclease blocking agent, and a UMI; (d) using at least a portion of the product of (c) to further amplify the genomic ROI with nested primers to obtain a gDNA library, At least one of the nested primers i) Well-specific barcodes, UMIs, and PCR handles, or ii) Capture Sequence Including, If the nested primer comprises the capture sequence, this step further comprises incubating the product of (c) with an exonuclease and a capture oligo; the capture oligo comprises the capture sequence, a well-specific barcode, an exonuclease blocking agent, and a UMI; said capture oligo binds to the amplicon generated with said nested primer, effectively labeling the product with said barcode during a PCR reaction; (e) pooling at least a portion of each well and then separating at least two of the gDNA library, the cDNA library, and the ADT library; (f) preparing a gDNA library, a cDNA library, and an ADT library for sequencing, comprising amplifying the ADT library with a sequencing primer, tagmenting the cDNA library and preferentially amplifying the 3' end with a sequencing primer, and amplifying the gDNA library with a sequencing primer. The method comprising:

3. 3. The method of claim 1 or 2, wherein the blocking agent is a phosphoryl or acetyl group, and the blocking agent is preferably linked to the 3'OH group of the capture oligomer.

4. All amplifications prior to preparing the gDNA library, the cDNA library, and the ADT library are performed in the same well; or The formation of the cDNA library, the genomic ROI library, and the ADT library is performed in a first well, and the gDNA library is prepared in another well; or The separation of the cDNA library and the ADT library is performed prior to or in parallel with the preparation of the gDNA library; The method according to claim 1 or 2.

5. Sequencing the gDNA library, the cDNA library, and the ADT library.

3. The method of claim 1 or 2, wherein the polymerase is Taq polymerase.

6. 6. The method of claim 5, wherein the Taq polymerase is KAPA HiFI Taq polymerase or Q5 Taq polymerase.

7. The method of claim 1 or 2, wherein the cell surface marker is CD45, CD81, or MHC class 1.

8. the gDNA library, the cDNA library, and / or the ADT library are separated using solid phase reversible immobilization (SPRI) beads; said separating preferably comprises first separating said gDNA library from said cDNA library and said ADT library using SPRI beads, followed by separating said cDNA library from said ADT library using SPRI beads; Separating the cDNA library from the ADT library preferably comprises separating amplicons greater than 500 bp in length from amplicons less than 500 bp in length, respectively, from each other. The method according to claim 1 or 2.

9. one or more of the cells contain an alteration in a genomic DNA sequence compared to a sequence of a reference genome; The change is preferably introduced using genome editing techniques; The genome editing technique preferably includes base editing or homologous recombination (HDR) editing; The method according to claim 1 or 2.

10. The method of claim 1 or 2, wherein one or more of the cells comprises altered mRNA expression compared to the mRNA expression of a reference cell or comprises altered expression of a cell surface marker compared to a reference cell.

11. 3. The method of claim 1 or 2, wherein the cells are edited using CRISPR prior to characterization.

12. the cells are primary cells and / or immune cells, The primary cells and / or immune cells are preferably mammalian cells, The mammalian cell is preferably a human cell. The method according to claim 1 or 2.

13. The cells are sorted using a FACS sorter; and / or At least about 500,000 to more than 10 million cells are characterized; The method according to claim 1 or 2.

14. After the incubation or amplification step, the products of said incubation or amplification are cleaned, The cleaning is preferably carried out using solid phase reversible immobilization (SPRI) beads. The method according to claim 1 or 2.

15. the detectable antibody comprises a fluorophore; and / or the oligoconjugate antibody comprises a polyA sequence; The method according to claim 1 or 2.

16. 3. The method of claim 1 or 2, wherein the sequencing primers are Illumina primers P5 and P7.

17. 1. A method for simultaneously characterizing genomic DNA and mRNA of a single cell, comprising: (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest; (b) incubating the detectably labeled cells of (a) with an oligoconjugated antibody; (c) index sorting the cells into a single well, characterizing the expression of cell surface markers for each cell, and lysing the cells in the presence of dNTPs, a well-specific barcoded oligoDT primer comprising a unique molecular identifier (UMI) and a PCR handle, and a capture oligo comprising a capture sequence, a well-specific barcode, an exonuclease blocking agent, and a unique molecular identifier; (d) incubating the product of (c) with a reverse transcriptase, a custom template switch oligo (TSO) comprising one member of a binding pair, and the reverse transcriptase under conditions that allow for the generation of cDNA; (e) incubating the product of step (d) with a genomic primer that specifically binds to a region of interest (ROI), a cDNA amplification primer that specifically binds to the PCR handle and a cDNA amplification primer that specifically binds to the TSO, an antibody derived tag (ADT) specific primer, dNTPs, and a polymerase under conditions that support amplification, thereby co-amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library; (f) contacting the product of step (e) with an exonuclease to degrade any unconsumed primer; (g) incubating at least a portion of the genomic ROI library from each well of (f) with dNTPs, polymerase, and nested primers capable of specifically amplifying a region within the genomic ROI library to obtain a gDNA library, where at least one of the nested primers comprises a capture sequence and a capture oligo binds to the amplicons generated with the nested primers, effectively labeling the products with a barcode during the PCR reaction; (h) pooling at least a portion of the samples from each well, followed by separating the gDNA library, the cDNA library, and the ADT library; and (i) preparing a gDNA library, a cDNA library, and an ADT library for sequencing by amplifying each library in the presence of a sequencing primer; Including, The method, wherein steps (c) to (e) occur simultaneously or sequentially.

18. The method of claim 14, wherein the exonuclease blocking agent is a phosphoryl group or an acetyl group.

19. 18. The method of claim 1, 2 or 17, wherein the exonuclease is ExoI.

20. 1. A method for simultaneously characterizing genomic DNA and mRNA of a single cell, comprising: (a) labeling a plurality of isolated cells with a detectable antibody that specifically binds to a cell surface marker of interest; (b) incubating the detectably labeled cells of (a) with an oligoconjugated antibody; (c) index sorting the cells into single wells, characterizing the expression of cell surface markers for each cell, and lysing the cells in the presence of dNTPs and a well-specific barcoded oligoDT primer comprising a unique molecular identifier (UMI) and a PCR handle; (d) incubating the product of (c) with a reverse transcriptase and a custom template switch oligo (TSO) that comprises one member of a binding pair under conditions that allow for the generation of cDNA; (e) incubating the product of step (d) with a genomic primer that specifically binds to a region of interest (ROI), a cDNA amplification primer that specifically binds to the PCR handle and a cDNA amplification primer that specifically binds to the TSO, an antibody derived tag (ADT) specific primer, dNTPs, and a polymerase under conditions that support amplification, thereby co-amplifying the gDNA, cDNA, and ADT to form a cDNA library, a genomic ROI library, and an ADT library; (f) pooling at least a portion of the samples from each well, followed by separation of the cDNA library and the ADT library; (g) incubating at least a portion of the genomic DNA from each well of (e) with dNTPs, a polymerase, and a nested primer that specifically binds to a region of interest to obtain a gDNA library, wherein the nested primer comprises a well-specific barcode, a UMI, and a PCR handle; and (h) preparing a gDNA library, a cDNA library, and an ADT library for sequencing by amplifying each library in the presence of a sequencing primer; The method comprising: