Single-cell edit capture sequencing
By employing exogenous nucleases and insert oligonucleotides with bacteriophage promoters to track gene edits, the method addresses the challenge of linking on- and off-target effects in CRISPR-Cas9 systems, enhancing the analysis of gene expression changes and improving the accuracy of gene editing systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Existing gene editing systems, particularly CRISPR-Cas9, face challenges in accurately linking on- and off-target edits to their effects on gene expression due to promiscuity of the Cas9 endonuclease, leading to variable guide expression levels and viral genome recombination, which impairs the effectiveness of methods like Perturb-seq in capturing true perturbations and analyzing non-coding regulatory sequences.
A method involving the use of exogenous nucleases and an exogenous insert oligonucleotide with a bacteriophage promoter sequence operably linked to a barcode sequence is introduced into cells to cleave endogenous nucleic acids, allowing for the insertion of the oligonucleotide at edit sites, followed by in situ transcription to generate RNA transcripts for analysis, enabling accurate tracking of gene edits and their functional consequences.
This approach enhances the ability to analyze both on- and off-target gene editing events by providing precise nucleic acid analytes, including RNA and protein activity changes, thereby improving the accuracy and power of studying gene expression impacts.
Smart Images

Figure US2025047000_26032026_PF_FP_ABST
Abstract
Description
SINGLE-CELL EDIT CAPTURE SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 696,694, filed on September 19, 2024, the entire disclosure of which is incorporated herein by reference.ACKNOWLEDGEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with government support under grant number HG011315 awarded by the National Institutes of Health. The government has certain rights in the invention.INCORPORATION BY REFERENCE OF SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety for all purposes. The XML copy created on September 18, 2025 is referred to as SIBS.P0018WO_SequenceListing.xml and is 94,755 bytes in size.TECHNICAL FIELD
[0004] This description is generally directed towards compositions, systems, and methods for monitoring and analyzing gene edits in cells induced by exogenous gene editing systems. More specifically, the compositions, systems and methods use a unique oligonucleotide insert to track edits as well as their functional consequences to endogenous gene expression.BACKGROUND
[0005] A longstanding barrier in genome engineering has been the inability to directly connect both on- and off-target edits to their effects on gene expression. This is particularly recognized in CRISPR-Cas gene editing, especially CRISPR-Cas9 gene editing, because of the promiscuity of the Cas9 endonuclease, but is a recognized issue across all gene editing systems.
[0006] A key opportunity to address this gap, at least for CRISPR based systems, are single-cell screens (Perturb-seq) that link hundreds or thousands of perturbations to molecularphenotypes by joint sequencing of cellular nucleic acid and viral vector-integrated guides, or “guide capture”. However, Perturb-seq and other similar screens have a key limitation: guide capture poorly correlates with the occurrence of genome perturbations, impairing Perturb-seq effectiveness. This disconnect is caused by variable Cas effector and guide expression levels, variable on-target and off-target guide activity, and viral genome recombination that disrupts guide vector sequences. As a result, guide capture defines guide-positive cell populations that contain a substantial fraction of confounding unperturbed and off-target cells, and it fails to capture many cells with genuine perturbations. This yields noisy, dampened differential analyses and poor power to study the small effects of non-coding regulatory sequences and variants.
[0007] As such, there is a need for systems and methods that can accurately and efficiently analyze the occurrence and functional impact (e.g., on gene expression) for both on and off-target gene editing events in a cell. The present disclosure addresses this and other needs.SUMMARY
[0008] In some aspects, a method for genetically modifying a cell is disclosed. In various embodiments, a method for genetically modifying one or more cells may comprise: (a) contacting the one or more cells with one or more gene editing compositions comprising: (i) one or more exogenous nucleases or one or more encoding nucleic acids thereof; and (ii) an exogenous insert oligonucleotide comprising a bacteriophage promoter sequence operably linked to a barcode sequence, or an encoding nucleic acid thereof; and (b) introducing the one or more exogenous nucleases or one or more encoding nucleic acids thereof and the exogenous insert oligonucleotide into the one or more cells, wherein after introduction, the one or more exogenous nucleases cleave one or more endogenous nucleic acids at one or more edit sites, and wherein the exogenous insert oligonucleotide is inserted into at least one edit site.
[0009] In various aspects, the methods provided herein further comprise analyzing one or more nucleic acid analytes derived from the one or more edit sites. In various aspects, the one or more nucleic acid analytes comprise nucleic acid analytes derived from a single edit site or nucleic acid analytes derived from more than one edit sites. In various aspects, the more than one edit sites comprise edit sites in a single cell and / or more than one cells. In variousaspects, the one or more nucleic acid analytes are derived from an edit site in a target gene locus, an edit site at an off-target gene locus, or any combination thereof.
[0010] In various aspects, each nucleic acid analyte comprises the barcode or a portion thereof and an extension sequence corresponding to a portion of the endogenous nucleic acid near to or adjacent to the edit site. In some aspects, the one or more nucleic acid analytes comprise RNA or a cDNA thereof. For instance, in some aspects, the one or more nucleic acid analytes comprise RNA transcripts of the bacteriophage promoter, or cDNA thereof.
[0011] In accordance with the foregoing, the methods herein may further comprise performing in situ transcription or in vitro transcription of the one or more cells to generate the RNA transcripts of the bacteriophage promoter. In various aspects, the methods comprise performing in situ transcription. In various aspects, performing in situ transcription further comprises (i) fixing the one or more cells and (ii) contacting the one or more fixed cells with a transcription composition comprising an RNA polymerase corresponding to the bacteriophage promoter. In some aspects, the transcription composition further comprises nucleoside triphosphates, magnesium, dithiothreitol (DTT), spermidine, inorganic pyrophosphate or any combination thereof.
[0012] In various aspects, analyzing the one or more nucleic acid analytes comprises determining a sequence or sequences of the one or more nucleic acid analytes, or any reverse complement thereof. In some aspects, the methods comprise analyzing one or more nucleic acid analytes from intact cells or isolated nucleic.
[0013] Any of the methods provided herein may further comprise analyzing endogenous gene expression in the one or more cells.
[0014] In various aspects, analyzing endogenous gene expression in one or more cells comprises analyzing one or more endogenous RNAs of the one or more cells. In some aspects, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are increased relative to a non-edited control cell. In some aspects, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are decreased relative to a non-edited control cell. In various aspects, the endogenous RNAs comprise mRNA, tRNA, miRNA, rRNA, mtRNA, or any combination thereof. In any of these aspects, the one or more endogenous RNAs may be encoded by a target gene locus or an endogenous nucleic acid that is operably linked to thetarget gene locus. In any of these aspects, the one or more endogenous RNAs are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus.
[0015] In various aspects, endogenous gene expression in one or more cells comprises detecting a change in activity of one or more endogenous proteins of the one or more cells. In various aspects, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is increased relative to a non-edited control cell. In various aspects, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is decreased relative to a non-edited control cell. In any of these methods, detecting the change of activity of the one or more endogenous proteins may comprise detecting a change of an expression level of the one or more endogenous proteins. In various aspects, the one or more endogenous proteins are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus. In various aspects, the one or more endogenous proteins are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus.
[0016] In any of the foregoing or related methods, analyzing the one or more nucleic acid analytes, the one or more endogenous RNAs and / or the one or more endogenous proteins comprises combinatorial indexing. In some aspects, combinatorial indexing comprises single cell RNA sequencing. In some aspects combinatorial indexing comprises SPLiT-seq, droplet based RNA sequencing, 10X sequencing, sci-RNA-seq, or any combination thereof. In some aspects, combinatorial indexing comprises SPLiT-Seq.
[0017] In any of the foregoing or related methods, the one or more cells repair the cleavage of the endogenous nucleic acid and inserts the exogenous oligonucleotide at each edit site using a repair mechanism. In some aspects, the repair mechanism comprises nonhom ologous end joining (NHEJ), microhomology mediated end joining (MMEJ), homologous recombination (HR), viral genome integration, transposition, or any combination thereof. In some aspects, the repair mechanism comprises non-homologous end joining (NHEJ).
[0018] In any of the foregoing or related methods, the bacteriophage promoter sequence may comprise a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof. In some aspects, the bacteriophage promoter sequence comprises anucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103. Accordingly, in some aspects, the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8) or (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT) (SEQ ID NO: 9).
[0019] In any of the foregoing or related methods, the one or more exogenous nucleases comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme.
[0020] In various aspects, the one or more exogenous nucleases comprise a CRISPR associated (Cas) endonuclease and the one or more gene editing compositions further comprises one or more guide RNAs or one or more encoding nucleic acids thereof. In some aspects, the one or more gene editing compositions further comprises one or more ribonucleoprotein (RNP) complexes each comprising the Cas endonuclease and at least one gRNA. In some aspects, the one or more gene editing compositions comprise one or more expression constructs encoding the one or more guide RNAs. In various aspects, the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor. In various aspects, the single-strand editing “nickase” editor comprises a prime editor. In various aspects, the Cas endonuclease comprises Cas9, Cas3, CaslO, or Casl2.
[0021] In various aspects, the one or more gene editing compositions comprise one or more expression constructs comprising the one or more nucleic acids encoding the one or more exogenous nucleases, the exogenous oligonucleotide and / or the one or more gRNAs, for example, any expression construct provided herein.
[0022] In any of the foregoing or related methods, introducing the one or more exogenous nucleases or one or more encoding nucleic acids thereof and / or the exogenous insert oligonucleotide into the cell comprises transfection, transduction, electroporation or injection. In various aspects, therefore, the one or more gene editing compositions may further comprise a transfection reagent, a transduction reagent and / or an electroporation reagent. For example, in some aspects, the one or more gene editing compositions comprise anelectrolytic / isotonic buffer, an exogenous “enhancer” oligonucleotide that increases electroporation delivery efficiency, or any combination thereof.
[0023] Also provided herein is a gene editing system comprising: (a) one or more nucleases or one or more encoding nucleic acids thereof; and (b) an exogenous insert oligonucleotide or encoding nucleic acid thereof, the exogenous insert oligonucleotide comprising a bacteriophage promoter sequence and a barcode sequence.
[0024] In various aspects, the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof. For example, in some aspects the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 7 and 103. In further aspects, the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8). In some aspects, the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO: 9).
[0025] In various aspects, the exogenous insert oligonucleotide comprises one or more modified nucleotides. In various aspects, the exogenous insert oligonucleotide is a synthetic oligonucleotide.
[0026] In any of the foregoing aspects, the one or more nucleases comprise CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a homing endonuclease, a transposase, an integrase, or a restriction enzyme, or any combination thereof.
[0027] In some aspects, the one or more exogenous nucleases used in the gene editing system comprise a CRISPR associated (Cas) endonuclease and the gene editing system further comprises (c) one or more guide RNAs or one or more encoding nucleic acids thereof. In some aspects, the Cas endonuclease comprises a double strand editing Cas or a singlestrand editing Cas “nickase” editor (e.g., a prime editor). In some aspects, the Cas endonuclease comprises Cas9, Cas3, CaslO, or Casl2. In various aspects, the gene editingsystem further comprises one or more ribonucleoprotein (RNP) complexes each comprising the Cas endonuclease and at least one gRNA.
[0028] In any of these aspects, the components of the gene editing systems provided herein (e.g., one or more nucleases, exogenous insert oligonucleotide and gRNA, if needed) are comprised in one or more gene editing compositions. In some aspects, the one or more gene editing compositions comprise any nucleic acid and / or expression construct provided herein. In some aspects, the one or more gene editing compositions further comprise at least one reagent for transfection, transduction, electroporation, or injection. In some aspects, the one or more gene editing compositions further comprise a cell culture medium.
[0029] Also provided herein are nucleic acids comprising a bacteriophage promoter operably linked to a barcode sequence. In some aspects, the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof. For example, in various aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103. In still further aspects, the nucleic acid may comprise a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO:8). In still further aspects, the nucleic acid may comprise a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO:9).
[0030] In various aspects, the nucleic acid further comprises one or more modified nucleotides. In some aspects, the one or more modified nucleotides are located at a 5’ end, a 3’ end, or any combination thereof of the nucleic acid. In some aspects, at least one modification on the one or more modified nucleotides comprises a 5’ phosphate and / or a phosphorothioate linkage.
[0031] In any of the foregoing aspects, the nucleic acid may comprise DNA. In further aspects, the nucleic acid is synthetic.
[0032] Also provided herein are expression constructs encoding any nucleic acid described herein. In various aspects, the expression construct further comprises one or more regulatory elements configured to regulate the expression of the nucleic acid.
[0033] In various aspects, the expression constructs provided herein may further comprise (a) a nucleic acid encoding for a nuclease and (b) one or more regulatory elements configured to regulate the expression of (a). In various aspects, the nuclease comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme. For example, in some aspects, the nuclease comprises a Cas endonuclease. In some aspects, the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor. In various aspects, the single-strand editing “nickase” editor comprises a prime editor. In various aspects, the Cas endonuclease comprises a Cas9, Cas3, CaslO, or Casl2 endonuclease.
[0034] In any of the foregoing aspects, an expression construct as provided herein may of further comprise: (c) a nucleic acid encoding for a guide RNA (gRNA); and (d) one or more regulatory elements configured to regulate the expression of (c).
[0035] In any of the foregoing or related aspects, the expression construct may be a plasmid or a viral vector. In some aspects, the viral vector is a lentiviral vector or an adeno- associated viral vector (AAV).
[0036] Also provided herein are genetically modified cells, comprising one or more nucleic acid insertions, each nucleic acid insertion comprising a bacteriophage promoter sequence and a barcode. In various aspects, the bacteriophage promoter sequence may comprise a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof. In various aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103. In further aspects, each nucleic acid insertion comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8). In some aspects, each nucleic acid insertion comprises a sequence having at least 85%, at least 90%,at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO: 9).
[0037] Also provided are any genetically modified cells comprising a nucleic acid and / or an expression vector described herein. Also provided are any genetically modified cells edited using the gene editing system provided herein.
[0038] Any of the genetically modified cells provided herein may be eukaryotic. In some aspects, the genetically modified cell is mammalian. In some aspects, the genetically modified cell is murine or human.
[0039] Also provided herein are kits comprising a nucleic acid comprising a bacteriophage promoter operably linked to a barcode sequence and at least one container. In some aspects, the kit further comprises an expression vector encoding for the nucleic acid. In further aspects, the expression vector further comprises one or more regulatory elements configured to regulate the expression of the nucleic acid.
[0040] In various aspects, the kits further comprise one or more nucleases or a nucleic acid encoding for the one or more nucleases. In some aspects, the kit further comprises an expression vector comprising the nucleic acid encoding for the one or more nucleases. In various aspects, the expression vector further comprises one or more regulatory elements configured to regulate the expression of the one or more nucleases.
[0041] In various aspects, the one or more nucleases comprise comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme or any combination thereof. In various aspects, the one or more nucleases comprise comprises a CRISPR associated (Cas) endonuclease. In various aspects, the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor (e.g., a prime editor). In various aspects, the Cas endonuclease comprises a Cas9, Cas3, CaslO, or Casl2 endonuclease.
[0042] In various aspects, the kits provided herein may further comprise one or more gRNAs or an encoding nucleic acid thereof. In some aspects, the kits comprise an expression vector encoding the one or more gRNAs. In various aspects, the expression vector furthercomprises one or more regulatory elements configured to regulate the expression of the one or more gRNAs.
[0043] In any of the foregoing aspects, each component in a kit provided herein may be comprised in one or more compositions. In various aspects, the one or more compositions further comprise one or more reagent for transfection, transduction, electroporation, or injection.INCORPORATION BY REFERENCE
[0044] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF FIGURES
[0045] The unique features of the technology are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present technology will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the technology are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein). The accompanying drawings are not intended to be drawn to scale. Like reference numbers and designations in the various drawings indicate like elements. For purposes of clarity, not every component can be labeled in every drawing. In the drawings:
[0046] FIG. l is a graphical illustration of a gene editing system according to various embodiments.
[0047] FIG. 2 is a graphical illustration of one or more nucleic acid analytes derived from one or more gene edits induced by a gene editing system described herein according to various embodiments.
[0048] FIG. 3 is a graphical illustration of various nucleic acid analytes that may be derived from edits in a single cell according to various embodiments.
[0049] FIG. 4 is another graphical illustration of various nucleic acid analytes that may be derived from edits in a single cell according to various embodiments.
[0050] FIG. 5 is a graphical illustration of various nucleic acid analytes that may be derived from a plurality of cells according to various embodiments.
[0051] FIG. 6 is a graphical illustration of an in situ transcription method that may be used in the methods of the present disclosure according to various embodiments.
[0052] FIG. 7 is a graphical illustration of general analytical processes that may be used to analyze one or more nucleic acid analytes and endogenous gene expression in cells according to various embodiments.
[0053] FIG. 8 is a flow diagram depicting a method for precisely analyzing gene edit effects on gene expression according to various embodiments.
[0054] FIG. 9 is an overview of a combinatorial indexing method of the present disclosure showing how to detect CRISPR edit sites, allelic edit dosage, and gene expression quantification in single cells according to various embodiments. Illustrative barcode depicted in panel 3 is GGGAGAGTAT (SEQ ID NO: 52).
[0055] FIG. 10 is a graphical illustration depicting a “guide capture” protocol (Perturb- seq) to capture gene edits according to various embodiments.
[0056] FIG. 11 is a graphical illustration and plot showing that guide capture technology poorly measures true perturbations after gene editing in cells according to various embodiments.
[0057] FIG. 12 is a graphical illustration of a method to directly “edit capture” in single cells according to various embodiments.
[0058] FIG. 13 is a graphical illustration and plots showing successful edit capture in K562 cells according to various embodiments.
[0059] FIG. 14 is a graphical illustration of a split-pool RNA-seq method and output according to various embodiments. tRNA reads depicted on right: SEQ ID NOs: 10-16.
[0060] FIG. 15 is a set of plots showing that bacteriophage RNA successfully marks edit sites in different cells according to various embodiments.
[0061] FIG. 16 is a set of data plots showing capture of single-cell edit allele effects by in situ transcript sequencing according to various embodiments.
[0062] FIG. 17 is a plot of functional profiling of genome-wide off-target events according to various embodiments.
[0063] FIG. 18 depicts features and possible applications of single cell edit capture sequencing technology according to various embodiments.
[0064] FIGS. 19A-19H depict homology -free knock-in of phage T7 promoter at targeted genome edits according to various embodiments. FIG. 19A shows the three-step single-cell edit capture sequence process consisting of Cas9 edit labeling with T7 promoters, cell fixation and in situ transcription of edit-marking T7 RNA in fixed cells, and joint combinatorial scRNA-seq of T7 and cellular RNA. FIG. 19B shows quantification of promoter-labeled Cas9 editing by Sanger sequencing and TIDE analysis. FIG. 19C shows fold knock-in rate of 27 candidate donor designs relative to the GUIDE-seq donor, measured by TIDE. For donor 2 (*), bar represents the mean knock-in of four samples. FIG. 19D shows the sequence of the 34-bp donor (donor 2, SEQ ID NO: 17) encoding an optimized T7 promoter and T7-transcribed barcode sequence. FIG. 19E shows TIDE analysis of genome edit outcomes in 70 / 112 edited samples with R2>0.5 (6 donor only, 5 RNP only, 59 knock- in). Knock-in is defined as +30 bp insertion or greater, and random indel is defined as any other event. Asterisks (*) indicate K562 cell lines selected for the single-cell experiment. FIG. 19F shows frequencies of unedited (none), variable untemplated indels, and donor insertion alleles in four cell types treated with donor+RNP targeting six genome sites. FIG. 19G shows frequency of donor insertion events in 51 / 59 samples treated with donor+RNP exhibiting detectable insertion, and cumulative frequency of in-frame (+0) and frame shift (+1 / +2) events. FIG. 19H shows frequency of donor knock-in and donor-less indel events in 59 samples treated with donor+RNP, labeled by outcome group, depicts homology-free knockin of a phage T7 promoter.
[0065] FIGS. 20A-20H show development of gene-disrupting Cas9 edit labeling with phage T7 promoter according to various embodiments. FIG. 20A shows design variables of 27 candidate donor DNA constructs. FIG. 20B shows comparison of sequences of the GUIDE-seq donor (SEQ ID NO: 18), Cel-seq+ T7 promoter (SEQ ID NO: 19), and the donor for use in single ell edit capture sequencing described herein (SEQ ID NO: 17). GUIDE-seq end modifications (red) and core 18-bp T7 promoter sequence (SEQ ID NO: 1, gold) are indicated. Terminal 5' phosphate (P) and 3' phosphorothioate (*) modifi cations are indicated. FIG. 20C shows TIDE R-squared (A2) values from 112 genomic DNA samples, and 70 samples with A2> 0.5. FIG. 20D shows mean frequencies of unedited (none), untemplated indel, and donor insertion (+30 bp or greater) events for the indicated sample group. FIG.20E shows mechanisms of gene disruption by donor-mediated frame shift events. FIG. 20F shows relative mRNA expression of four chromatin remodeler genes in the absence of editing (none) or the presence of donor+RNP editing, quantified by reverse transcription qPCR (RT- qPCR) and the comparative threshold cycle (Ct) method. Results are represented as a difference of differences in Ct values (ddCt) normalized to unedited samples and reference genes RPL24 and RPS10. Guide RNA indicated. FIG. 20G shows median B2M fluorescence intensity (MFI) of unedited and 7>2A7-targeted B2M donor-edited GM12878 cells, quantified by flow cytometry. FIG. 20H shows indel distribution of edited cells in FIG. 20G, measured by TIDE.
[0066] FIGS. 21 A-21C show bimodal knock-in outcome associates with guide RNA sequence according to various embodiments. FIG. 21 A shows frequency of donor knock-in and random indel events in 49, 59, or 59 knock-in samples labeled by guide RNA (left), cell type (middle) or genome locus (right). FIG. 2 IB shows the frequency of “high indel” or “low indel” outcome, grouped by sample label. FIG. 21C shows positional Shannon entropy (bits) of the protospacer sequences of 10 guide RNAs used in the 27 “high indel” samples, or the 10 guides used in the 32 “low indel” samples, calculated using ggseqlogo.
[0067] FIGS. 22A-22I show the development of in situ transcription of targeted genome edits in fixed cells according to various embodiments. FIG. 22A shows generation of T7 RNA by IVT on purified gDNA containing T7 promoter-labeled genome edits. FIGS. 22B- 22C shows quantity of T7 RNA from IVT reactions in the absence (RNP) or presence (donor+RNP) of donor DNA or in unedited cells (none), measured by reverse transcription quantitative PCR (RT-qPCR) and the comparative threshold cycle (Ct) method. Genome editswere generated at CTLA4 (FIG. 22B) or B2M (FIG. 22C) in the indicated cell type. Sample replicate, assay and guide RNA are indicated. FIG. 22D shows generation of T7 RNA by 1ST on unfixed nuclei containing promoter-labeled genome edits. FIGS. 22E-22F show quantity of T7 RNA from nuclei 1ST reactions, or mock 1ST reactions without T7 RNA polymerase (No pol), targeting the indicated genome edit site (CTLA4 (FIG. 22E) or B2M (FIG. 22F). FIG. 22G shows generation of T7 RNA by 1ST on fixed cells containing T7 promoter-labeled genome edits. FIG. 22H shows quantity of T7 RNA from fixed-cell 1ST reaction supernatants (ambient) or cell pellets (in situ), before and after 1ST reaction optimization. FIG. 221 shows quantity of T7 RNA from replicate fixed-cell 1ST reactions targeting chromatin remodeler genes SMARCA4 and CHD4. For IVT and unfixed 1ST samples (A-D), results are represented as a difference of Ct values (dCt) normalized to two reference assays targeting nontranscribed genomic safe harbor (GSH) sites. For fixed 1ST samples (E,F), results are represented as a difference of differences (ddCt) normalized to reference genes RPL24 and RPS10 and unedited samples.
[0068] FIGS. 23 A-23D show generation of in vitro T7 transcripts at targeted genome edits according to various embodiments. FIG. 23 A shows RNA concentration from T7 or SP6 IVT reactions on unedited primary T-cell genomic DNA (gDNA), in the indicated buffer, measured by UV spectroscopy (A260 / A280). Concentration is normalized to input DNA template concentration. To minimize confounding effects of NTP depletion, short 1-hour IVT reactions were performed. FIG. 23B shows quantity of T7 RNA at T7 promoter-labeled genome edits at CTLA4, generated by IVT on purified primary T-cell gDNA, and measured by RT-qPCR and comparative Ct. Distance and direction of PCR amplicons relative to donor insertion position is indicated. Control assays targeted chromosome sites with homology to T7 promoter (Chr6, Chr7, Chr8), and non-coding sequences at four genome loci, including two non-transcribed “genomic safe harbor” sites. FIGS. 23C-23D show quantity of T7 RNA at the indicated promoter-labeled genome edits at CTLA4 (FIG. 23 C) or B2M (FIG 23D) sites and control sites, using the indicated cell lines and guide RNAs. Assays targeting upstream (5') or downstream (3') of T7 promoter-homologous sequences are indicated. Results are represented as a difference of Ct values (dCt) normalized to GSH assays.
[0069] FIGS. 24A-24F shows generation of in situ T7 transcripts in isolated nuclei and fixed cells according to various embodiments. FIGS. 24A-24D show quantity of T7 RNA and mRNA in 1ST reactions on unfixed nuclei, measured by RT-qPCR and comparative Ct inJurkat cells edited with CTLA4 guide (FIG. 24A), K562 cells edited with CTLA4 guide 2 (FIG. 24B), K562 cells edited with B2M guide 4 (FIG. 24C) or GM12878 cells edited with B2M guide 4 (FIG. 24D). Cell type, target locus, guide RNA, electroporation condition (RNP, Donor+RNP), and polymerase condition (No pol, T7 pol) are indicated. Results are represented as a difference of Ct values (dCt) normalized to GSH assays FIG. 24E shows quantity of T7 RNA in 1ST reactions on paraformaldehyde-fixed cells. Results are represented as either ddCt values normalized to a negative control sample and reference genes RPL24 and RPS10, or as a difference of Ct values (dCt) relative to RPL24 mRNA. FIG. 24F shows quantity of T7 RNA in 1ST reactions on unfixed or fixed nuclei from the same nuclei isolate.
[0070] FIGS. 25A-25B show optimization of in situ transcription in PFA-fixed cells according to various embodiments. FIG. 25 A shows quantity of T7 RNA in fixed-cell 1ST reactions under the indicated reaction conditions: NTP concentration, incubation time, incubation temperature, reaction fraction (ambient supernatant or in situ cell pellet). FIG. 25B shows quantity of T7 RNA in wash supernatant or post- wash cell pellet.
[0071] FIGS. 26A-26B show in situ transcription at chromatin remodeler genes and, specifically, the quantity of T7 RNA in replicate fixed-cell 1ST reactions at chromatin remodeler genome edits in K562 cells according to various embodiments. FIG. 26A shows quantities in unedited cells (none) or a pool of ARIDlA / SMARCA4-edited cells. FIG. 26B shows quantities in unedited cells or a pool of CHD3 / CHD4-edited cells.
[0072] FIGS. 27A-27E show sequencing of in situ T7 transcripts identifies Cas9 genome edits according to various embodiments. FIG. 27A shows a diagram of the bulk RNA-seq experiment. Jurkat and K562 cells were treated with Cas9 RNP and donor targeting CTLA4 and B2M, respectively, and applied to nuclei isolation and T7 1ST, followed by total RNA extraction, RNA-seq library preparation, and paired-end sequencing analysis. FIGS. 27B-27C show sequence alignments at CTLA4 (FIG. 27B) and B2M (FIG. 27C) for treatments with RNP or donor+RNP, without 1ST (No pol) or with 1ST (T7 pol), or donor+RNP. Read pileups (gray) for all samples, and individual forward / reverse-strand reads (blue / red) for donor+RNP with 1ST. Targeted PAM, expected Cas9 cleavage position (vertical line), and unmapped positions (visible bases) are shown. Sequences in FIG. 27B: red reverse strand reads: SEQ ID NOs: 13-15; blue fwd strand reads: SEQ ID NOs: 16-32. Reference sequence (bottom): SEQ ID NO: 33. Sequences in FIG. 27C: red reverse strand reads: SEQ ID NOs:34-39; blue fwd strand reads: SEQ ID NOs: 40-43. Reference sequence (bottom): SEQ ID NO: 44. FIG. 27D shows total and mapped RNA-seq read counts. FIG. 27E shows library fragment size distribution, measured by TapeStation.
[0073] FIGS. 28A-28J show high-throughput single-cell profiling of Cas9 genome edits by single-cell edit capture sequencing according to various embodiments. FIG. 28A shows a diagram of a single-cell edit capture sequencing experiment on 10,000 K562 cells targeting four chromatin remodeler genes with seven guide RNAs. In situ transcription and SPLiT-seq combinatorial barcoding was performed on three pools of unedited, ARID1A / SMARCA4- edited, and CHD3 / CHD4-edited cell lines. Paired-end sequencing reads were analyzed by Split-pipe and custom software suite . FIG. 28B shows coverage of T7 reads called by Sheriff at the seven on-target edit sites (count per million mapped reads, CPM). Expected on- target sites are indicated for each guide RNA (red lines). FIGS. 28C-28E shows sensitivity (FIG. 38C), specificity (FIG. 28D) and false discovery rate (FDR, FIG. 28E) of the singlecell edit capture sequencing platform. Sensitivity was defined as the rate of correct T7 read calls among total reads within 100 bp of the seven expected on-target edit sites. Specificity was defined as the rate of correct non-T7 read calls in the unedited sample. FDR was defined as the rate of non-T7 reads among called T7 reads. FIG. 28F-28H show UMAP clustering of 9,500 K562 cells analyzed by a custom software suite . Cells are colored by total count of unique molecular identifiers (UMIs, FIG. 28F), sample pool (FIG. 28G), or cell cycle score from Scanpy (FIG. 28H). FIG. 281 shows direct single-cell capture of Cas9 edits at all seven on-target sites and 36 off-target sites in 6,230 edited cells, quantified by Sheriff. FIG. 28J shows absolute distance between expected on-target edit sites and called edit sites by Sheriff.
[0074] FIGS. 29A-29D show the generation of a joint single-cell single cell edit capture sequencing library of T7 and endogenous RNA reads according to various embodiments. FIG. 29A shows frequency of donor knock-in and donor-less indel events in six K562 cell lines included in single-cell analysis. FIG. 29B shows diagram of combinatorial addition of barcode (BC) and unique molecular identifier (UMI) sequences to single cell edit capture sequencing cDNA fragments, and final structure of paired-end sequencing library. See FIG. 30 for detailed sequence structure. FIG. 29B shows size distribution of fragments in 500-cell and lOk-cell sequencing libraries. FIG. 29C shows read and cell counts from paired-end sequencing of 500-cell and lOk-cell single cell edit capture sequencing libraries. FIG. 29Dshows summary of Split-pipe quality metrics (total reads, STAR, single cells, read depth, mRNA and genes).
[0075] FIG. 30 shows sequence structure of single cell edit capture sequencing library fragments according to various embodiments. Expected sequences of the library fragments generated by EVERCODE barcoding and Illumina sequencing. Paired-end sequencing run configurations for the 500-cell and lOk-cell libraries. Adapted from Teich Lab resource on SPLiT-seq library sequences
[0076] FIGS. 31A-31H show that single cell edit capture sequencing analysis detects single cell differentially expressed genes and both on- and off-target cas9 edit events according to various embodiments. FIG. 31A shows an overview of the differential expression procedure, which utilized linear mixture modeling combined with a cell grouping procedure to call differentially expressed genes (DEGs) associated with allelic dosage of single cell edit events while accounting for cell state and sequencing depth. FIG. 3 IB shows violin plots of target genes, associating the cell-state and sequencing depth corrected expression of each target gene with the number of single cell cas9 allelic edit events for the associated gene. FIG. 31C is equivalent to FIG. 3 IB, except for off-target genes that were detected by single cell edit capture sequencing. FIG. 3 ID shows an Upsetplot indicating the intersection sizes of DEGs when calling DEGs with respect to the allelic dosage of single cell edit events for both on- and off-target genes. FIG. 3 IE shows a heatmap indicating the proportion of cells with pairs of on- and off-target gene edits. Proportions are conditioned on the gene edit listed on the rows. For example, row one of the heatmap indicates 68% of USP9X edited cells are also CHD4 edited, while row 5 indicates 8% of CHD4 edited cells are also USP9X edited. FIG. 3 IF shows volcano plots of DEGs associated with edits. Edit allele effect sizes are on the x-axes and -loglO Benjamini -Hochberg corrected p-values on the y- axes. FIG. 31G shows gene set over-representation analysis for BAF-associated CRISPR edited genes (on- and off-targets), depicted as a dot-plot of log2 odds-ratios with 95% confidence intervals. Combinations of edited gene DEGs and query gene sets are shown on the y-axis, and the log2 odds-ratio are on the x-axis. The vertical dotted line indicates a log2 odds-ratio expected by random chance. Horizontal dotted lines separate query gene sets. Query gene sets include G2M (Hallmark G2M checkpoint genes), BAF enh (BAF complex associated genes in K562 cells by enhancer regulation), and BAF prom (BAF complex associated genes in K562 cells by promoter regulation). FIG. 31H is equivalent to FIG. 31G,except for NuRD complex associated edited genes (on- and off-targets). Query gene sets included WNT (genes related to WNT -mediated signal transduction), and NuRD (NuRD complex associated genes in K562 cells by promoter regulation). *p < 0.05 by Fisher’s exact test, indicating significant association of edited gene DEGs and the query gene set.
[0077] FIGS. 32A-32E show high-throughput single-cell identification of guide-specific Cas9 on-target and off-target sequences by single cell edit capture sequencing according to various embodiments. FIG. 32A shows genome regions and guide sequence similarity of on- target and off-target edit sites identified by a custom software suite . Name labels are colored by on-target (dark red) or off-target (light red). Bars are colored by causal guide RNA. Intersection of exon (“X”), intron (“I”), or intergenic sequence (“O”) is indicated. FIG. 32B shows aligned guide and target sequences of edit sites associated with CHD3 guide 10 RNA (SEQ ID NO: 60) and corresponding total cell counts. Matches (•) and gaps (-) are indicated, along with the intersecting gene if present, cell count, start position, and strand. Aligned target sequences depicted: ENSG000000287360 (SEQ ID NO: 61), USP9X (SEQ ID NO: 62), - (chr4:88368033 (-)) (SEQ ID NO: 63), FBXO38-DT (SEQ ID NO: 64), CPT1 A (SEQ ID NO: 65), RAP1B (SEQ ID NO: 66), FAM225B (SEQ ID NO: 67), - (chrl9: 10293401 (+)) (SEQ ID NO: 68), ICAM5 (SEQ ID NO: 69) and SMYD3 (SEQ ID NO: 70). FIG. 32C shows absolute distance between called edit sites and guide-targeted PAM sequences for the 36 off-target sites, determined by a custom software suite . FIG. 32D shows frequency of 12 off-target sites of SMARCA4 guide 22 (light red), compared to the on-target site (dark red). FIG. 32E shows intersection of single cell edit capture sequencing detected off-target sites with predicted sites from in silico tools Cas-OFFinder, COSMID, and E-CRISP, across the seven guides.
[0078] FIG. 33 shows genome-wide detection of edit-marking T7 RNA reads according to various embodiments. Per-chromosome coverage of T7 RNA reads containing unmapped (soft-clipped) 5' barcode sequence encoded by T7 promoter donor DNA. On-target chromatin remodeler gene signals are indicated.
[0079] FIGS. 34A-34C show high-throughput single-cell profiling of Cas9 genome edits by single cell edit capture sequencing according to various embodiments. FIG. 34A shows a locus zoom of T7 read coverage at four on-target chromatin remodeler genes. Guide RNA indicated. FIG. 34B shows total count of UMIs, alleles, and cells measured using single-cell edit capture with the indicated guide RNAs. FIG. 34C shows a ternary plot, with each dotindicating a called canonical edit site (total n = 63). Dots are coloured and scaled by the number of cells called with T7 barcoded reads for the edit site. Edit sites are plotted according to the proportion of cells called with the edit in the different treatment samples (unedited, ARID1AISMARCA4 guide treated cells, CHD3ICHD4 guide treated cells). True edit sites are expected to be specific to a particular treatment sample. .
[0080] FIGS. 35A-35I show single cell edit capture sequencing analysis of guide-specific off-target sequences from pooled samples according to various embodiments. FIGS. 35A-35F shows aligned guide and target sequences of the additional six of seven guide RNAs targeting four chromatin remodeler genes: ARID1 guide 5 (SEQ ID NO: 71, FIG. 35A), ARID1 A guide 6 (SEQ ID NO: 75, FIG. 35B), SMARCA4 guide 19 (SEQ ID NO: 77, FIG. 35C), SMARCA4 guide 22 (SEQ ID NO: 87, FIG. 35D), CHD4 guide 12 (SEQ ID NO: 100, FIG. 35E), CHD4 guide 14 (SEQ ID NO: 102, FIG. 35F). Matches (•) and gaps (-) are indicated, along with the intersecting gene if present, cell count, zero-based start position, and strand. Aligned target sequences in FIG. 35 A: ENSG00000266869 (SEQ ID NO: 72), ENSG00000255872 (SEQ ID NO: 73), SORCS2 (SEQ ID NO: 74). Aligned target sequences in FIG. 35B: PDE6C (SEQ ID NO:76). Aligned target sequences in FIG. 35C: - (chr6: 1397175 (+)) (SEQ ID NO: 78), ENSG00000258210 (SEQ ID NO: 79), RAP1GAP2 (SEQ ID NO: 80), LINC02026 (SEQ ID NO: 81), TFDP1 (SEQ ID NO: 82), - (chrl6:81393291 (+)) (SEQ ID NO: 83), KCNJ6 (SEQ ID NO: 84), STXBP3 (SEQ ID NI: 85), CLCN6 (SEQ ID NO: 86). Aligned target sequences in FIG. 35D: ADSS1 (SEQ ID NO: 88), CDC27 (SEQ ID NO: 89), MACROD2 (SEQ ID NO: 90), CNIH3 (SEQ ID NO: 91), - (chr4:59298816) (SEQ ID NO: 92), PCAT4 (SEQ ID NO: 93), - (chr3:31532818 (-)) (SEQ ID NO: 94), RARB (SEQ ID NO: 95), STT3B (SEQ ID NO: 96), MY0M1 (SEQ ID NO: 97), SHANK2 (SEQ ID NO: 98), LINC02343 (SEQ ID NO: 99). Aligned target sequences in FIG.35E: MLLT1 (SEQ ID NO: 101). FIG. 35G shows distribution of distances between canonical edit site positions (i.e. 5' start position of the strongest T7 read variant) and the expected Cas9 cleavage position of target sequences (PAM -3). FIG. 35H shows number of matched bases in global alignment of each guide sequence to each of its associated edit site target sequences, colored by manual labeling as true positive or false positive edit site call. FIG. 351 shows minimum distance between opposite-stranded T7 reads at each edit site.
[0081] FIGS. 36A-36B show sequence homology of guide-specific off-target sequences according to various embodiments. FIG. 36A shows total number of target sequence (n)indicated. Positional Shannon entropy (bits) of the target sequences of the indicated guide, or all seven guides, calculated using ggseqlogo. FIG. 36B shows Shannon entropy of the PAM - 2 position across 43 on-target and off-target sequences, either unweighted or weighted by the observed cell count of each target sequence.
[0082] It is to be understood that the figures are not necessarily drawn to scale, nor are the objects in the figures necessarily drawn to scale in relationship to one another. The figures are depictions that are intended to bring clarity and understanding to various embodiments of apparatuses, systems, and methods disclosed herein. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. Moreover, it should be appreciated that the drawings are not intended to limit the scope of the present teachings in any way.DETAILED DESCRIPTION
[0083] This specification describes various exemplary embodiments of methods and compositions for detecting gene editing events in single cells. The disclosure, however, is not limited to these exemplary embodiments and applications or to the manner in which the exemplary embodiments and applications operate or are described herein. Moreover, the figures may show simplified or partial views, and the dimensions of elements in the figures may be exaggerated or otherwise not in proportion.
[0084] In addition, where reference is made to a list of elements (e.g., elements a, b, c), such reference is intended to include any one of the listed elements by itself, any combination of less than all of the listed elements, and / or a combination of all of the listed elements.Section divisions in the specification are for ease of review only and do not limit any combination of elements discussed.
[0085] It should be understood that any uses of subheadings herein are for organizational purposes, and should not be read to limit the application of those subheaded features to the various embodiments herein. Each and every feature described herein is applicable and usable in all the various embodiments discussed herein and that all features described herein can be used in any contemplated combination, regardless of the specific example embodiments that are described herein. It should further be noted that exemplary descriptionsof specific features are used, largely for informational purposes, and not in any way to limit the design, subfeature, and functionality of the specifically described feature.
[0086] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclatures utilized in connection with, and techniques of, chemistry, biochemistry, molecular biology, pharmacology and toxicology are described herein are those available and commonly used in the art.Definitions:
[0087] As used herein, “substantially” means sufficient to work for the intended purpose. The term “substantially” thus allows for minor, insignificant variations from an absolute or perfect state, dimension, measurement, result, or the like such as would be expected by a person of ordinary skill in the field but that do not appreciably affect overall performance. When used with respect to numerical values or parameters or characteristics that can be expressed as numerical values, “substantially” means within ten percent.
[0088] The term “ones” means more than one.
[0089] As used herein, the term “plurality” can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.
[0090] As used herein, the terms “comprise”, “comprises”, “comprising”, “contain”,“contains”, “containing”, “have”, “having” “include”, “includes”, and “including” and their variants are not intended to be limiting, are inclusive or open-ended and do not exclude additional, unrecited additives, components, integers, elements or method steps. For example, a process, method, system, composition, kit, or apparatus that comprises a list of features is not necessarily limited only to those features but may include other features not expressly listed or inherent to such process, method, system, composition, kit, or apparatus.
[0091] Where values are described as ranges, it will be understood that such disclosure includes the disclosure of all possible sub-ranges within such ranges, as well as specific numerical values that fall within such ranges irrespective of whether a specific numerical value or specific sub-range is expressly stated.
[0092] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclatures utilized in connection with, and techniques of, cell and tissue culture, molecular biology, and protein and oligo- or polynucleotide chemistry and hybridization described herein are those available and commonly used in the art. Standard techniques are used, for example, for nucleic acid purification and preparation, chemical analysis, recombinant nucleic acid, and oligonucleotide synthesis. Enzymatic reactions and purification techniques are performed according to manufacturer's specifications or as commonly accomplished in the art or as described herein. The techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the instant specification. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. 2000). The nomenclatures utilized in connection with, and the laboratory procedures and techniques described herein are those available and commonly used in the art.
[0093] DNA (deoxyribonucleic acid) is a chain of nucleotides consisting of 4 types of nucleotides; A (adenine), T (thymine), C (cytosine), and G (guanine), and that RNA (ribonucleic acid) is comprised of 4 types of nucleotides; A, U (uracil), G, and C. Certain pairs of nucleotides specifically bind to one another in a complementary fashion (called complementary base pairing). That is, adenine (A) pairs with thymine (T) (in the case of RNA, however, adenine (A) pairs with uracil (U)), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand made up of nucleotides that are complementary to those in the first strand, the two strands bind to form a double strand. As used herein, “nucleic acid sequencing data,” “nucleic acid sequencing information,” “nucleic acid sequence,” “gene sequence,” “genetic sequence,” or “fragment sequence,” or “nucleic acid sequencing read” denotes any information or data that is indicative of the order of the nucleotide bases (e.g., adenine, guanine, cytosine, and thymine / uracil) in a molecule (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, fragment, etc.) of DNA or RNA. It should be understood that the present teachings contemplate sequence information obtained using all available varieties of techniques, platforms or technologies, including, but not limited to: capillaryelectrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, electronic signature-based systems, etc.
[0094] A “polynucleotide”, “nucleic acid”, or “oligonucleotide” refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleosidic linkages. Typically, a polynucleotide comprises at least three nucleosides. Usually oligonucleotides range in size from a few monomeric units, e.g. 3-4, to several hundreds of monomeric units. Whenever a polynucleotide such as an oligonucleotide is represented by a sequence of letters, such as “ATGCCTG,” it will be understood that the nucleotides are in 5 '->3' order from left to right and that “A” denotes deoxyadenosine, “C” denotes deoxycytidine, “G” denotes deoxyguanosine, and “T” denotes thymidine, unless otherwise noted. The letters A, C, G, and T may be used to refer to the bases themselves, to nucleosides, or to nucleotides comprising the bases, as is standard in the art.
[0095] The term “barcode,” as used herein, generally refers to a label, or identifier, that conveys or is capable of conveying information about an analyte. A barcode can be part of an analyte. A barcode can be independent of an analyte. A barcode can be a tag attached to an analyte (e.g., nucleic acid molecule) or a combination of the tag in addition to an endogenous characteristic of the analyte (e.g., size of the analyte or end sequence(s)). A barcode may be unique. Barcodes can have a variety of different formats. For example, barcodes can include: polynucleotide barcodes; random nucleic acid and / or amino acid sequences; and synthetic nucleic acid and / or amino acid sequences. A barcode can be attached to an analyte in a reversible or irreversible manner. A barcode can be added to, for example, a fragment of a deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) sample before, during, and / or after sequencing of the sample. Barcodes can allow for identification and / or quantification of individual sequencing-reads.
[0096] The term “subject,” as used herein, generally refers to an animal, such as a mammal (e.g., human) or avian (e.g., bird), or other organism, such as a plant. For example, the subject can be a vertebrate, a mammal, a rodent (e.g., a mouse), a primate, a simian or a human. Animals may include, but are not limited to, farm animals, sport animals, and pets. A subject can be a healthy or asymptomatic individual, an individual that has or is suspected of having a disease (e.g., cancer) or a pre-disposition to the disease, and / or an individual that is in need of therapy or suspected of needing therapy. A subject can be a patient. A subject canbe a microorganism or microbe (e.g., bacteria, fungi, archaea, viruses). The term “non-human animals” includes all vertebrates, e.g., mammals, e.g., rodents, e.g., mice, non-human primates, and other mammals, such as e.g., sheep, dogs, cows, chickens, and non-mammals, such as amphibians, reptiles, etc.
[0097] The terms “adaptor(s)”, “adapter(s)” and “tag(s)” may be used synonymously. An adaptor or tag can be coupled to a polynucleotide sequence to be “tagged” by any approach, including ligation, hybridization, or other approaches.
[0098] The term “sequencing,” as used herein, generally refers to methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides can be, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA). Sequencing can be performed by various systems currently available, such as, without limitation, a sequencing system by Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, or Life Technologies (Ion Torrent®). Alternatively or in addition, sequencing may be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. Such systems may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the systems from a sample provided by the subject. In some examples, such systems provide sequencing reads (also “reads” herein). A read may include a string of nucleic acid bases corresponding to a sequence of a nucleic acid molecule that has been sequenced. In some situations, systems and methods provided herein may be used with proteomic information.
[0099] The terms “coupled,” “linked,” “conjugated,” “associated,” “attached,” “connected” or “fused,” as used herein, may be used interchangeably herein and generally refer to one molecule (e.g., polypeptide, receptor, analyte, etc.) being attached or connected (e.g., chemically bound) to another molecule (e.g., polypeptide, receptor, analyte, etc.).
[0100] The term “analyte,” as used herein, generally refers to a species of interest for detection. An analyte may be biological analyte, such as a nucleic acid molecule or protein. An analyte may be an atom or molecule. An analyte may be a subunit of a larger unit, such as, e.g., a given sequence of a polynucleotide sequence or a sequence as part of a largersequence. An analyte of the present disclosure includes a secreted analyte, a soluble analyte, and / or an extracellular analyte.
[0101] Analytes may be derived from cells and may include at least a portion of a targeting gene locus, a reverse complement thereof, or a derivative thereof. Analytes may include at least a portion of a target or off-target gene locus, a reverse complement thereof, or a derivative thereof.
[0102] The term “insert oligonucleotide,” as used herein, generally means nucleotide molecule, typically double stranded, that is inserted into a gene loci. The term “insert oligonucleotide” can include or exclude oligonucleotides comprising homology arms complementary to the gene locus at a site of a gene break. In some aspects, insert oligonucleotides are configured to be inserted into a gene locus via a non-homologous repair mechanism (e.g., non-homology end joining (NHEJ). In various embodiments, an insert oligonucleotide can comprise a barcode. In various embodiments, an insert oligonucleotide can comprise a bacteriophage promoter sequence operably linked to the barcode.
[0103] As used herein, “operably linked” means that expression of a given nucleic acid is under the control of a component to which it is spatially connected. The term “operably linked” may be bidirectional. That is, a promoter may be “operably linked” to a nucleic acid, such that the promoter directs expression of the nucleic acid. In these cases, the promoter can be positioned 5’ (upstream) or 3’ (downstream) of the nucleic acid under its control. The distance between the promoter and the nucleic acid under its control can be approximately the same as the distance between that promoter and any gene it normally controls. As is known in the art, variation in this distance can be accommodated without loss of promoter function. In other aspects, a nucleic acid (e.g., an endogenous nucleic acid) may be operably linked to an entire gene locus such as an endogenous nucleic acid that is operably linked to a gene locus (e.g., a target gene locus or a non-target gene locus). In this instance, this means that the endogenous nucleic acid is under control of a component in or expressed by the gene locus. For example, the gene locus might comprise a regulatory element that controls expression of the endogenous nucleic acid or it might encode a regulatory protein that acts on a regulatory element that controls expression of the nucleic acid.
[0104] The term “gene locus” as used herein, generally means a specific fixed position on a chromosome where a particular gene or genetic marker is located. As used herein, the genelocus is inclusive of all gene sequences that are involved in the expression of a given gene. These include, but are not necessary limited to: exons, introns, regulatory regions (e.g., promoters, enhancers, repressors), splicing sites and the like. It is understood that any gene edit occurring at a gene locus has the potential to disrupt expression of at least that gene.
[0105] The term “target gene locus,” as used herein, generally means any gene locus comprising a nucleic acid sequence specifically targeted by one or more nucleases provided herein. For instance, when the one or more nucleases comprise a CRISPR endonuclease and are used alongside a gRNA, the “target gene locus” may comprise the target nucleic acid sequence targeted by the spacer region of the gRNA.
[0106] The term “off-target gene locus,” as used herein, generally means any gene locus that does not comprise a nucleic acid a nucleic acid sequence specifically targeted by one or more nucleases provided herein. For instance, when the one or more nucleases comprise a CRISPR endonuclease and are used alongside a gRNA, the “target gene locus” may comprise the target nucleic acid sequence targeted by the spacer region of the gRNA.
[0107] The term “ribonucleoprotein (RNP) complex,” as used herein, generally means a complex formed between RNA and RNA-binding proteins. More specifically, a RNP complex can comprise a targeting gRNA molecule including at least some homology to a spacer or protospacer sequence found within a target gene locus and a Cas endonuclease capable of recognizing a PAM sequence within the target gene locus. In such embodiments, the gRNA can recruit the Cas endonuclease to a cut site within the spacer or protospacer region of the target gene locus.
[0108] The term “guide ribonucleic acid (gRNA),” as used herein, generally means an RNA molecule associated with a targeting RNP complex including at least some homology to a spacer or protospacer sequence found within a target gene locus.
[0109] The term “transfection reagent,” as used herein, generally means any reagent that increases the likelihood of a successful transfection. In various embodiments, a transfection agent can be lipid-based and configured to improve a transfection rate. Non-limiting examples of transfection reagents are lipofectamine and Minis TransiT. Generally, in liposomal transfection, the positive surface charge of the liposomes mediates the interaction of the nucleic acid and the cell membrane, allowing for fusion of the liposome / nucleic acid transfection complex with the negatively charged cell membrane. The transfection complexcan enter the cell through endocytosis. Other useful transfection reagents may include an electrolytic / isotonic buffer.
[0110] The term “transduction reagent” as used herein generally means any reagent that increases the likelihood of a successful transduction (e.g., a viral transduction). In various aspects, the transduction reagent may be optimized for use with certain viral vectors (e.g., lentiviral or adeno-associated viral (AAV) vectors). Other useful transduction reagents may include an electrolytic / isotonic buffer.[OHl] The term “electroporation reagent” as used herein generally means any reagent that increases the likelihood of a successful electroporation. In various embodiments, an electroporation reagent may comprise an exogenous “enhancer” oligonucleotide that increases electroporation delivery efficiency. As a non-limiting example, an Alt-R Cas9 Electroporation Enhancer (INTEGRATED DNA TECHNOLOGIES) is an electroporation reagent that may be used in the methods of the present disclosure.
[0112] As used herein, the term “endogenous” means any component that is native to a cell. Endogenous nucleic acids include any nucleic acids found in or expressed by the cell before any gene editing event such as genomic or mitochondrial DNA, RNA (e.g., mRNA, rRNA, tRNA, miRNA, siRNA, snoRNAs, piRNA, tsRNA, srRNA. , mtDNA, miRNAs, etc). Endogenous proteins include any proteins, peptides, or polypeptides, that exist in or are expressed by a cell before any gene editing event.
[0113] As used herein, the term “exogenous” means any component that is not native to a cell. Exogenous nucleic acids include any nucleic acid that is introduced into the cell and may include, for example, nucleic acids that encode for any of the gene editing components described herein. Exogenous proteins (e.g, an exogenous nuclease) include any proteins that are introduced into a cell or encoded by a nucleic acid that is introduced into a cell (e.g., an exogenous nucleic acid).I. OVERVIEW:
[0114] A longstanding barrier in genome engineering (e.g, CRISPR-Cas9 genome editing) has been the inability to directly connect both on- and off-target edits to their effects on gene expression. The disclosure herein describes single cell edit capture sequencing, an accessible single-cell gene editing sequencing method that can profile the transcriptomiceffects of off-target nuclease genome edits, a longstanding gap in human genome engineering.
[0115] As an example, the RNA-guided Cas9 nuclease is a foundational tool for targeted manipulation of genetic sequence. Grounded in a decade of mechanistic insights, experiment design principles, and analytical methods, Cas9 is the workhorse for effective genome editing in fundamental and translational research. In human biology, two groundbreaking trends are Cas9 gene therapies and CRISPR functional genomics. In therapeutics, two Cas9 gene therapies are now approved to treat blood disorders, and about fifty clinical trials are developing Cas9-edited interventions, notably CAR-T. In fundamental molecular biology, Cas9 gene knockout is a core approach to loss-of-function experiments, from broad functional screens to targeted validation to mechanism interrogation. However, a key obstacle to Cas9 efficacy is “off-targef ’ genome editing-sequence alterations at unintended sites-that can trigger undesirable gene expression changes. Off-target gene perturbations confound cell phenotypes and elevate therapeutic risk. To use Cas9 effectively in biology and medicine, a pressing task is to determine how off-target genome edits affect gene activity.
[0116] Management of Cas9 off-target activity has progressed substantially, but downstream impact of off-target edits on gene expression remains unclear. Recent developments include greater understanding of chromosomal and structural factors, improved predictive modeling of guide off-targeting, and directed evolution of high-specificity Cas9 variants. A key achievement has been unbiased empirical methods that measure Cas9 editing in live cells by high-throughput amplicon sequencing or chromatin immunoprecipitation sequencing. These methods are invaluable tools for uncovering the biological mechanisms of CRISPR off-targets, and for characterizing monitoring off-target events in gene therapy programs. However, there remains a critical gap in methods. Bulk DNA sequencing of off- target edits does not provide functional information on gene expression effects. RNA sequencing approaches, notably single-cell Perturb-seq methods, do measure gene expression. However, these methods infer Cas9 activity from guide sequences, so they are poorly suited to detect and distinguish off-target edits and their associated effects. To bridge this gap, there is a critical need for a method that can quantify transcriptomic consequences of specific off-target edits by Cas9.
[0117] Single-cell edit capture perturbation sequencing, as described herein, is a first-inkind single-cell RNA sequencing (scRNA-seq) method that achieves transcriptomiccharacterization of off-target genome edits through high-throughput joint profiling of nuclease cleavage and cell transcriptomes. In contrast to Perturb-seq methods that read guide sequences, single-cell edit capture perturbation sequencing reads genome sequence at genuine nuclease edit sites, enabling direct association of specific on- and off-target edit events to differential gene expression.
[0118] To capture genome edit events in single cells, the methods described herein leverage bacteriophage promoter transcription, the two-component system of a bacteriophage promoter and RNA polymerase that performs efficient promoter-restricted amplification of DNA sequence. Production of bacteriophage promoter RNA in cell-free in vitro transcription (IVT) reactions is used broadly, from low-input sequencing to mRNA vaccine manufacturing. Recent work has established bacteriophage promoter in situ transcription (1ST), the amplification of bacteriophage promoter RNA from fixed cell genomes. Single cell edit capture sequencing, as described herein, in various embodiments harnesses T7 1ST for unbiased amplification of on-target and off-target genome edit sequence for analysis by scRNA-seq.Single-cell edit capture sequencing
[0119] Single-cell edit capture sequencing has 3 general steps that use off-the-shelf kits and no specialized equipment: (1) a gene editing step comprising knock-in of a barcoded bacteriophage promoter that is inert in human cells to generate promoter-labeled genome edits; (2) a step of generating one or more nucleic acid analytes from each promoter-labeled genome edits; and (3) combinatorial analysis of the nucleic acid analytes alongside markers of endogenous gene expression in the cell. FIG. 8 depicts an exemplary workflow which can be deployed in accordance with various aspects of the disclosure. In some aspects, as shown in FIG. 8, one or more nucleases and exogenous insert oligonucleotide comprising a bacteriophage promoter and a barcode into one or more cells (802). The one or more nucleases can then be used to modify a target gene locus and / or an off-target gene locus in the one or more cells (804). The modification of the target gene locus and / or off-target gene locus may include incorporation of the exogenous insert oligonucleotide (806). Transcription of each edit sites may be performed to generate one or more nucleic acid analytes derived from each edit site (808). Additionally, one or more analytes of endogenous gene expression may be generated (810). Then combinatorial analysis of one or more nucleic acid analytes derived from each edit site and one or more analytes of endogenous gene expression maythen be performed (812) resulting in precision analysis of gene edit effects on gene expression (814). Each of these steps will be described in further detail below.Gene Editing
[0120] In various aspects, a method for genetically modifying one or more cells is provided, the method comprising (a) contacting the one or more cells with one or more gene editing compositions comprising: (i) one or more exogenous nucleases or one or more encoding nucleic acids thereof; and (ii) an exogenous insert oligonucleotide comprising a bacteriophage promoter sequence operably linked to a barcode sequence, or an encoding nucleic acid thereof; and (b) introducing the one or more exogenous nucleases or one or more encoding nucleic acids thereof and the exogenous insert oligonucleotide into the one or more cells, wherein after introduction, the one or more exogenous nucleases cleave one or more endogenous nucleic acids at one or more edit sites, and wherein the exogenous insert oligonucleotide is inserted into at least one edit site.
[0121] FIG. 1 is a graphical illustration of a gene editing system 100 used in the methods of single-cell edit capture sequencing according to various embodiments. For clarity, FIG. 1 specifically depicts genetically editing a single cell, but it should be appreciated that the overall gene editing system is suited for genetically editing (and analyzing) many cells simultaneously. In various embodiments, the system 100 comprises one or more exogenous nucleases 102 or one or more encoding nucleic acids there and an exogenous insert oligonucleotide 104 comprising a bacteriophage promoter sequence 106 operably linked to a barcode sequence 108, or an encoding nucleic acid thereof. In various methods, one or more gene editing compositions comprising the exogenous nucleases 102 and exogenous insert oligonucleotide 104 contact one or more cells 112, thus introducing the exogenous nucleases 102 and exogenous insert oligonucleotide 104 into one or more cells. Once introduced into at least one cell 113, the one or more exogenous nucleases 102 cleave one or more endogenous nucleic acids 114 at one or more edit sites and the exogenous insert oligonucleotide is inserted into one or more of the edit sites 116. In various aspects, the one or more nucleases may be configured to target a specific region in the endogenous nucleic acid (i.e., a “target gene locus”). Therefore, in an aspect, the exogenous nucleotide 104 may be inserted into a target gene locus resulting in an “on-target” edit 118 or it may be inserted into an off-target gene locus resulting in an edit at an “off-target” edit 120. These on and off target edits may then be analyzed in further steps of the method.One or More Nucleic Acid Analytes
[0122] In various aspects, the methods provided herein comprise analyzing one or more nucleic acid analytes derived from the one or more edit sites. FIG. 2 is a graphical illustration of one or more nucleic acid analytes derived from one or more gene edits induced by the gene editing system described in FIG. 1. In various aspects, the method comprises deriving one or more nucleic acid analytes 204 from each edit site 202.
[0123] In various aspects, the one or more nucleic acids analytes 204 each comprise the barcode 108 or a portion thereof and an extension sequence 206 corresponding to a portion of the endogenous nucleic acid near to or adjacent to the edit site. In various aspects, the one or more nucleic acid analytes comprise RNA or a cDNA thereof. Methods of deriving the one or more nucleic acid analytes are described further below.
[0124] FIG. 3 is a graphical illustration of various nucleic acid analytes that may be derived from edits in a single cell 300. In various aspects, the one or more nucleic acid analytes may comprise nucleic acid analytes derived from a single site or nucleic acid analytes derived from more than one edit site. In FIG. 3, two illustrative edit sites are depicted: “on-target edits” 302 or “off-target” edits 304. One or more nucleic acid analytes may be derived from either or both of these edit sites to derive one or more nucleic acid analytes of an on-target edit 306 and / or one or more nucleic acid analytes of an “ off-target’ ’ edit 308. It is also contemplated that more than one gene locus may be targeted by the one or more nucleases (i.e., in a multiplex experiment) so there may be more than one distinct “on- target gene loci” for a given iteration of the method.
[0125] FIG. 4 is another graphical illustration of various nucleic acid analytes that may be derived from edits in a single cell 400. In various aspects, the one or more nucleic acid analytes are derived from more than one “on-target” edits on sister chromosomes. It is appreciated that in diploid cells having pairs of chromosomes that more than one “target nucleic acid” may exist (e.g., on two sister chromosomes). Higher order polyploid cells would have more “target nucleic acids”. Accordingly, in various aspects the gene editing methods herein comprise analyzing one or more nucleic acid analytes that are derived from “on-target” edits on one or more chromosomes in the cell. For example, as shown in FIG. 4, a first set of nucleic analytes 406 may be derived from an insertion at a target site 402 in a first chromosome and a second set of nucleic acid analytes 408 may be derived from ananalogous insertion at a target site 404 in a sister chromosome (e.g., a chromosome that also contains the same target gene locus as the first chromosome). In various aspects, the insertions in the first chromosome and the sister chromosome may both be in a target gene locus, but may not be at precisely the same location in the target gene locus. In these aspects, the gene editing system may induce two or more alleles in the cell at the target gene locus. Likewise, in various aspects, the insertions in the first chromosome and the sister chromosome may both be in the same off- target gene locus, but may not be at precisely the same location in the off-target gene locus. In these aspects, the gene editing system may induce two or more alleles in the cell at the off-target gene locus.
[0126] FIG. 5 is a graphical illustration of various nucleic acid analytes that may be derived from a plurality of cells 500 according to aspects of the present disclosure. In various aspects, the methods of the present disclosure comprise analyzing one or more nucleic acid analytes from one or more edited cells. These one or more nucleic acid analytes may be derived from each edit event in the one or more cells where the exogenous oligonucleotide is successfully inserted at the edit site. In FIG. 5 this is depicted as three separate edits 502, 504, and 506 in a first cell each yielding a set of nucleic acid analytes, and three separate edits 508, 510, and 512 in a second cell, each yielding a set of nucleic acid analytes. It is appreciated that in systems comprising editing more than one cells, that the one or more nucleic acid analytes may be derived from an edit site in a target gene locus, an edit site at an off-target gene locus, or any combination thereof, in one or more of the cells.Methods of Deriving One or More Nucleic Acid Analytes
[0127] As noted above, each nucleic acid analyte comprises the barcode of the exogenous insert oligonucleotide or a portion thereof and an extension sequence corresponding to a portion of the endogenous nucleic acid near to or adjacent to the edit site. In an aspect, the one or more nucleic acid analytes comprise RNA or cDNA thereof. In further aspects, the one or more nucleic acid analytes comprise RNA transcripts of the bacteriophage promoter, or cDNA thereof. These RNA transcripts may be generated by transcribing the exogenous insert oligonucleotide using an RNA polymerase that binds to and acts on the bacteriophage promoter included in the exogenous insert oligonucleotide. Various bacteriophage promoters are described further below and corresponding RNA polymerases would be readily identified by one of skill in the art. In one non-limiting example, the bacteriophage promoter comprises a T7 promoter and the RNA polymerase is a T7 RNA polymerase. In another non-limitingexample, the bacteriophage promoter comprises a T3 promoter and the RNA polymerase is a T3 RNA polymerase. In another non-limiting example, the bacteriophage promoter comprises a SP6 promoter and the RNA polymerase is an SP6 RNA polymerase.
[0128] Although any method of transcription is contemplated, in various aspects, the methods may comprise in situ transcription. FIG. 6 is a graphical illustration of an in situ transcription method that may be used in the methods of the present disclosure. In various aspects, one or more cells 602 that were edited according to the methods described herein above are fixed to yield one or more fixed cells 604. An RNA polymerase corresponding to the bacteriophage promoter used in the exogenous oligonucleotide insert 606 is added to the one or more fixed cells 604. This results in one or more cells 608 and 610 each having a set of RNA transcripts (“nucleic acid analytes”) 612 generated by the RNA polymerase from each edit site 616 in the cells.
[0129] Accordingly, in view of the foregoing, the methods provided herein comprise performing in situ transcription of one or more cells edited according to methods herein to generate the RNA transcripts of the bacteriophage promoter. In some aspects, performing in situ transcription further comprises (i) fixing the one or more cells and (ii) contacting the one or more fixed cells with a transcription composition comprising an RNA polymerase corresponding to the bacteriophage promoter. Exemplary transcription compositions are described further below, but in some aspects, the transcription composition may further comprise nucleoside triphosphates, magnesium, dithiothreitol (DTT), spermidine, inorganic pyrophosphate or any combination thereof.Endogenous Gene Expression
[0130] Gene editing events can have complicated, multifaceted, impacts on endogenous gene expression. For example, an insertion into a coding region of a gene (e.g., into the nucleic acid sequence encoding a protein or nucleic acid product) can directly disrupt transcription and expression of the product. However, insertion into a non-coding region (e.g., insertion into a regulatory sequence like a promoter or repressor and / or insertion into an intron) can also impact gene expression by disrupting its regulation (via the promoter or repressor) or post-transcriptional processing (e.g., by disrupting a splice site in an intron). Insertion into any region of a given gene locus can also have downstream effects on genes that are not located at the gene locus where the insertion occurs. For example, disruptingexpression of a regulatory protein (e.g., a transcription factor) can have downstream effects on the genes it regulates. In various aspects, the methods herein further comprise analyzing changes in endogenous gene expression in edited cells.
[0131] In various aspects, the methods comprise analyzing levels of one or more endogenous RNAs in the one or more cells. In some aspects, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are increased increase relative to a non-edited control cell. In some aspects, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are decreased relative to a non-edited control cell. The endogenous RNA may be coding or non-coding. The endogenous RNA may be messenger RNA (mRNA), ribosomal RNA (rRNA) or transfer RNA (tRNA), for example. The endogenous RNA may be a transcript. The endogenous RNA may be small RNA that are less than 200 nucleic acid bases in length, or large RNA that are greater than 200 nucleic acid bases in length. Small RNAs may include 5.8S ribosomal RNA (rRNA), 5S rRNA, transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNAs), Piwi-interacting RNA (piRNA), tRNA-derived small RNA (tsRNA) and small rDNA-derived RNA (srRNA). The RNA may be doublestranded RNA or single-stranded RNA. The endogenous RNA may be circular RNA. In various aspects, the one or more endogenous RNAs are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus. In various aspects, the one or more endogenous RNAs are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus. In various aspects, the one or more endogenous RNAs are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus and an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus.
[0132] In various aspects, the methods comprise analyzing endogenous protein activity in the one or more cells. In various aspects, the methods comprise detecting a change in activity of one or more endogenous proteins of the one or more cells. For instance, in some aspects, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is increased relative to a non-edited control cell. In other aspects, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is decreased relative to a non-edited control cell. In various aspects, activity of at least one endogenous protein in one or more cells is increased relative to a non-editedcontrol cell and activity of at least one endogenous protein in one or more cells is decreased relative to a non-edited control cell. In various aspects, one or more endogenous proteins are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus. In various aspects, the one or more endogenous proteins are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus. In any of these aspects, detecting the change of activity of the one or more endogenous proteins may comprise detecting a change of an expression level of the one or more endogenous proteins. Methods for measuring protein levels are known in the art (e.g., mass spectrometry, western blots and the like).Methods of Analyzing One or More Nucleic Acid Analytes and Endogenous Gene Expression using Combinatorial Indexing
[0133] FIG. 7 is a graphical illustration of general analytical processes that may be used to analyze the one or more nucleic acid analytes and endogenous gene expression in the cells. For instance, a dish of edited cells 702 may be analyzed using a single cell RNA sequencing method to generate data related to the one or more nucleic acid analytes 704 and endogenous RNAs 706 . Likewise, a single cell proteomics method may be used to generate the data related to endogenous protein activity. In particular aspects, the present disclosure provides for analyzing the gene edit events e.g., as measured by the one or more nucleic acid analytes 704), and a functional impact on gene expression (e.g., as measured by levels of endogenous RNAs 706 and / or endogenous protein activity 708) using combinatorial indexing. In various aspects, the combinatorial indexing can comprise a single cell RNA sequencing procedure. In various aspects, the combinatorial indexing comprises SPLiT-seq, droplet based RNA sequencing, 10X sequencing, sci-RNA-seq, or any combination thereof. In some aspects, the combinatorial indexing comprises a step of barcoding one or more of the nucleic acid analytes (or any endogenous RNA of interest).
[0134] In some instances, barcoding of a nucleic acid molecule may be done using a combinatorial approach. In such instances, one or more nucleic acid molecules (which may be comprised in a cell, e.g., a fixed cell , or cell bead) may be partitioned (e.g., in a first set of partitions, e.g., wells or droplets) with one or more first nucleic acid barcode molecules (optionally coupled to a bead). The first nucleic acid barcode molecules or derivative thereof (e.g., complement, reverse complement) may then be attached to the one or more nucleic acid molecules, thereby generating first barcoded nucleic acid molecules, e.g., using the processesdescribed herein. The first nucleic acid barcode molecules may be partitioned to the first set of partitions such that a nucleic acid barcode molecule, of the first nucleic acid barcode molecules, that is in a partition comprises a barcode sequence that is unique to the partition among the first set of partitions. Each partition may comprise a unique barcode sequence. For example, a set of first nucleic acid barcode molecules partitioned to a first partition in the first set of partitions may each comprise a common barcode sequence that is unique to the first partition among the first set of partitions, and a second set of first nucleic acid barcode molecules partitioned to a second partition in the first set of partitions may each comprise another common barcode sequence that is unique to the second partition among the first set of partitions. Such barcode sequence (unique to the partition) may be useful in determining the cell or partition from which the one or more nucleic acid molecules (or derivatives thereof) originated.
[0135] The first barcoded nucleic acid molecules from multiple partitions of the first set of partitions may be pooled and re-partitioned (e.g., in a second set of partitions, e.g., one or more wells or droplets) with one or more second nucleic acid barcode molecules. The second nucleic acid barcode molecules or derivative thereof may then be attached to the first barcoded nucleic acid molecules, thereby generating second barcoded nucleic acid molecules. As with the first nucleic acid barcode molecules during the first round of partitioning, the second nucleic acid barcode molecules may be partitioned to the second set of partitions such that a nucleic acid barcode molecule, of the second nucleic acid barcode molecules, that is in a partition comprises a barcode sequence that is unique to the partition among the second set of partitions. Such barcode sequence may also be useful in determining the cell or partition from which the one or more nucleic acid molecules or first barcoded nucleic acid molecules originated. The second barcoded nucleic acid molecules may thus comprise two barcode sequences (e.g., from the first nucleic acid barcode molecules and the second nucleic acid barcode molecules).
[0136] Additional barcode sequences may be attached to the second barcoded nucleic acid molecules by repeating the processes any number of times (e.g., in a split-and-pool approach), thereby combinatorically synthesizing unique barcode sequences to barcode the one or more nucleic acid molecules. For example, combinatorial barcoding may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more operations of splitting (e.g., partitioning) and / or pooling (e.g., from the partitions). Additional examples of combinatorial barcoding may alsobe found in International Patent Publication Nos. WO2019 / 165318, which is herein entirely incorporated by reference for all purposes.
[0137] Beneficially, the combinatorial barcode approach may be useful for generating greater barcode diversity, and synthesizing unique barcode sequences on nucleic acid molecules derived from a cell or partition. For example, combinatorial barcoding comprising three operations, each with 100 partitions, may yield up to 106 unique barcode combinations. In some instances, the combinatorial barcode approach may be helpful in determining whether a partition contained only one cell or more than one cell. For instance, the sequences of the first nucleic acid barcode molecule and the second nucleic acid barcode molecule may be used to determine whether a partition comprised more than one cell. For instance, if two nucleic acid molecules comprise different first barcode sequences but the same second barcode sequences, it may be inferred that the second set of partitions comprised two or more cells.
[0138] In some instances, combinatorial barcoding may be achieved in the same compartment. For instance, a unique nucleic acid molecule comprising one or more nucleic acid bases may be attached to a nucleic acid molecule (e.g., a sample or target nucleic acid molecule) in successive operations within a partition (e.g., droplet or well) to generate a first barcoded nucleic acid molecule. A second unique nucleic acid molecule comprising one or more nucleic acid bases may be attached to the first barcoded nucleic acid molecule, thereby generating a second barcoded nucleic acid molecule. In some instances, all the reagents for barcoding and generating combinatorially barcoded molecules may be provided in a single reaction mixture, or the reagents may be provided sequentially.
[0139] In various aspects, the methods herein comprise a SPLiT-seq protocol. SPLiT-seq is described in Rosenberg et al., (Science 2018 Apr 13;360(6385): 176-182) which is incorporated herein by reference in its entirety. SPLiT-seq is a type of combinatorial barcoding procedure optimized for use on fixed cells.
[0140] Once a dataset is obtained comprising a plurality of nucleic acid analytes (derived from gene editing events) and endogenous RNAs, and proteomic data from one or more cells, the data is then analyzed using a combinatorial indexing method as described herein below.
[0141] FIG. 9 provide an overview of a combinatorial indexing method of the present disclosure and specifically shows how to detect single cell CRISPR edit sites, allelic editdosage, and gene expression quantification. In a first step, shown in 1, Fastq files generated from single-cell edit capture sequencing are processed using split-pipe, software from parse biosciences, to generate a bam file. As shown in 2, three kinds of reads are present within this bam file; a)bacteriophage polymerase generated reads with a single-cell edit capture sequencing barcode, evident in the 5’ soft-clipped region of the mapped read, b) reads generated by bacteriophage polymerase that have been fragmented, removing the single-cell edit capture sequencing barcode, and c) reads generated from normal RNA pol II transcription within expressed genic regions. Edit sites will have a high number of reads mapping in both directions across cells (due to single-cell edit capture sequencing donor oligonucleotide inserting in either direction after the nuclease induced double-stranded break) with 5’ soft-clipped reads containing the single-cell edit capture sequencing barcode. As shown in 3, particular-edit-sites are edits occurring at a single cell level, identified by a k-mer match between the single-cell edit capture sequencing barcode and the reads 5’ soft-clipped sequence. If the read maps in the forward direction, the k-mer match should be to the forward version of the barcode. If the read maps in the reverse direction, the k-mer match should be to the reverse-complement of the single-cell edit capture sequencing barcode. Reads called bacteriophage barcoded reads are used to record edited cells, edit location, edit direction and inserted donor sequence. As shown in 4, particular-edits represent slight edit variation around a particular gene location. Canonical-edit-sites are determined by assigning particular-edits to more commonly occurring versions of the edit, with canonical-edits that do not appear to be well supported across multiple cells or by reads in both directions (a characteristic of the single-cell edit capture sequencing donor introducing a bacteriophage pol promoter in either direction, depending on the orientation of the cell particular-edit). Canonical-edit-sites are determined by first constructing a graph of particular-edit locations for each chromosome, and then rank-ordering particular-edits by the frequency that they occur across cells. Progressively, particular-edits from most-to-least frequent are considered as canonical-edit- sites, with all particular-edits within predetermined base-pair radius from the particular-edit assigned to the canonical-edit-site. All particular-edits collapsed to a canonical-edit are then removed from being considered a potential canonical-edit, until all particular-edits are assigned to a canonical-edit-site. As shown in 5, for each canonical-edit-site, the number of edited alleles for that site is then called at a single cell level. The number of edited alleles is determined as the number of particular-edits of a canonical-edit for the respective cell. As shown in 6 for edit-sites that overlap genic regions, non-barcoded t7 reads are removed, to ensure these do not confound gene expression estimation. This is done by removing all readsin-line with the single cell particular-edit, but are past the particular-edit site but within a particular window, and thereby could have been generated from the t7 polymerase but are missing the single-cell edit capture sequencing barcode due to fragmentation. Finally, as shown in 7, the key outputs from the single-cell edit capture sequencing read processing is a cell X canonical-edit-site matrix with 0, 1, or 2 indicating the number of alleles edited for the particular cell at the particular canonical-edit-site, and a cell X gene matrix indicating the number of unique molecular identifiers (UMIs) counted for a particular gene as a measure of single cell gene expression.Advantages for Single Cell Edit Capture Sequencing
[0142] Single cell edit capture sequencing introduces several novel measurements to the genome engineering community: high-throughput, single-cell nuclease cleavage profiles, guide-specific off-target sequence analysis from pooled samples, single-cell edit allele counts, and direct association of on- and off-target edit events to differential transcriptome expression. Single cell edit capture sequencing, as described herein, is a pivotal, first-in-kind gene editing method. It comprises a pioneer sequencing process that finally addresses a critical inability to study side effects of genome engineering on gene expression. Single cell edit capture sequencing captures the functional consequences of total genome edits by delivering single-cell transcriptomes paired with three completely new measures of nuclease activity: (1) direct single-cell capture of genuine on-target edit events, (2) single-cell profiling of nuclease off-targets in high-throughput (thousands of cells), and (3) single-cell edit allele counting and direct analysis. Previous methods have applied the components of single cell edit capture sequencing (unbiased edit labeling, in situ transcription, combinatorial indexing), but no method has integrated these to achieve the unique outputs described herein. For example, Cas9 off-target profiling methods like GUIDE-seq and DISCOVER-seq are sensitive bulk assays of off-target genotypes, but without gene expression or other measure of functional effects. Perturb-seq (pooled CRISPR) methods capture single-cell gene expression changes but cannot report off-target events, and they only infer on-target events from guide RNA barcodes that poorly correlate with true Cas9 activity. Previous reports of bacteriophage promoters like T7 in situ transcription are only capable of deploying T7 within randomly inserted cassettes, not targeted genome edits. Only single cell edit capture sequencing, as described herein, achieves first-ever functional single-cell edit capture anddelivers much needed unbiased observation of the downstream gene perturbations of individual Cas9 edits, whether on-target or off-target.
[0143] Single cell edit capture sequencing is a new paradigm for simplified multi-omics. Compositions and methods provided herein allow for one-pot single-cell indexing of in situ and endogenous transcriptomes. This is a breakthrough simplification for the field of functional multi-omics. The technical achievement is high-performance in situ transcription (1ST) on paraformaldehyde-fixed cells and direct combinatorial barcoding of total transcripts by off-the-shelf, random hexamer-primed SPLiT-seq, without dedicated sequencing preparations. In principle, this protocol can associate single-cell gene expression with any cellular target that can be labeled with a short DNA oligonucleotide (bacteriophage promoter) with unmatched accessibility. An immediate impact will be the functional profiling of other genome features such as precision genome edits and integrated barcodes of spatial phenotypes. One-pot indexing could transform the utility and usability of these emerging techniques that at present deliver either no transcriptomic readout or otherwise use multiple independent sequencing assays at great complexity and cost. Beyond genomic targets, a massive impact will be in pairing this protocol with bacteriophage promoter-labeled antibodies and other peptide binders or substrates, achieving streamlined multi-omic profiling of a vast range of cellular targets. In short, the indexing protocol described herein empowers any standard experimental laboratory to now leverage functional single-cell and spatial multi- omics.
[0144] The magnitude of this protocol advancement is clear by comparing single cell edit capture sequencing, as described herein, to current state-of-the-art methods. Single-cell profiling of in situ and endogenous transcriptomes has been used to characterize either random structural variants or epigenetic factors of prime editing outcomes using T7 reporters, but they do so by alternative protocols: 1ST of methanol -fixed nuclei or cells, followed by indexing with droplet-based Chromium (lOx Genomics) or combinatorial sci-RNA-seq. These current approaches have many shortcomings. First, these methods use multiple sequencing assays due to an incompatibility with randomly primed single-cell indexing, and due to limited T7 1ST performance on methanol-fixed samples. Random priming is likely prohibited due to the susceptibility to library saturation by high-abundance ribosomal and mitochondrial RNA. Thus, these methods instead use specific primers targeting a capture sequence encoded by the donor construct. Targeting priming yields T7 reads that criticallylack a specific genome sequence, using dedicated bulk RNA sequencing of randomly primed T7 reads just to capture missing sequence. Additionally, these methods use a third dedicated sequencing library for transcriptomes, likely due to lower quality and quantity of T7 transcripts, demanding separate single-cell assays. Second, these multi-assay approaches can only be applied to actively dividing cell types amenable to plasmid and viral vectors, like transformed cell lines, because they assay vector-delivered genomic events in sibling daughter cells. Third, these methods depend on custom reagents including a high-complexity donor DNA libraries. Finally, Chromium indexing yields much higher rates of problematic ambient sequences that reduce single-cell dataset quality and depth, a weakness since T7 1ST generates abundant ambient transcripts. In sum, existing methods are burdened by multisequencing workflows, challenging dataset integration, and serious cell type and quality limitations that are prohibitively costly, complex, or unapplicable for many laboratories. In contrast, single-cell edit capture sequencing, as described herein, uses exclusively off-the- shelf reagents, one donor oligonucleotide, one 1ST reaction, and one sequencing workflow that yields a one-pot single-cell library of high-quality in situ and cellular transcriptomes in one sequencing run, and can be applied to sensitive, non-dividing cell types, such as primary cells. This makes the single-cell edit capture sequencing protocol described herein far more accessible and useful to a broad userbase.
[0145] Therefore, single cell edit capture sequencing, as described herein, is an accessible, original single-cell RNA sequencing method that directly associates on-target and off-target genome edits to their transcriptomic effects. For example, as shown in the examples herein, when targeting four chromatin remodeler genes with seven guide RNAs in 9,500 K562 cells, single cell edit capture sequencing, as described herein, identified nearly 12,000 genome edit alleles in over 6,000 edited cells at all seven on-target sites and 36 off- target sites, the majority of which are not predicted by in silico tools. A key finding was that single cell edit capture sequencing, as described herein, associated multiple non-coding off- target edits with significant differential expression of off-target genes. Notably, single cell edit capture sequencing, as described herein, demonstrated transcriptomic characterization of off-target CDC27, a key regulator of cell division. Off-target CDC27 intron X resulting G2- M resulted in decreased CDC27 activity and differential expression of related DNA damage checkpoint genes; the CDC27 pathway is associated with cancer.
[0146] This result demonstrates that non-coding off-target edits can indeed disturb critical cellular pathways. Furthermore, this consequential off-target event was discovered in a set of just seven guide RNAs that were predicted to have high on-target specificity. Therefore, it is plausible that even when using state-of-the-art methods and design principles, Cas9 gene knockout can result in frequent off-target gene dysregulation that has been heretofore been neither observed nor predicted. In many instances, off-target gene perturbations can confound the observed cell phenotype and elevate risk of adverse events in therapeutic applications.
[0147] Current investigations of off-target Cas9 editing largely focus on understanding and managing the causes of off-targets, such as promiscuous guide sequences and complicit Cas9 residues. However, there has been limited focus on studying the consequences of off- target events on the cell transcriptome due to the technical difficulty of doing so. For a given guide RNA, its off-target genome edits remain challenging to precisely predict, vary widely in a cell population, and often co-occur with other on-target and off-target edits. Given this nature, functional off-target characterization has long-awaited a joint single-cell measurement of unbiased Cas9 editing activity and gene expression.
[0148] Single cell edit capture sequencing, as described herein, thus bridges a longstanding critical gap in CRISPR sequencing methods: a lack of method for functional off-target characterization. Standby off-target profiling methods (e.g. GUIDE-seq and DISCOVER-seq) are capable of sensitive, unbiased profiling of Cas9 edit locations, but these technologies do not associate edit events to like gene expression. Functional transcriptomic characterization of specific off-target edit events by single cell edit capture sequencing, as described herein, is a long-awaited, pivotal new ability for the genome engineering community.
[0149] A longstanding barrier in CRISPR-Cas9 genome engineering has been the unclear consequences of off-target genome edits on the cell transcriptome. This Examples and data described herein demonstrated that single cell edit capture sequencing, as described herein, has numerous advantages for the genome engineering community, including (1) high- throughput direct sequencing single-cell genome edits, (2) analysis of guide-specific off- target sequences from pooled samples, (3) single-cell edit allele counting, and (4) association of specific on- and off-target edit events to differential transcriptome expression.
[0150] Single cell edit capture sequencing has many uses, including at least two immediate applications, namely targeted pooled and in vivo ribonucleoprotein CRISPR screens in hard-to-transduce cell types, including primary cells and stem cells, and detection of rare oncogenic gene perturbations by guide protospacers in active clinical trials. The data described in the Examples herein provide evidence of detection of 36 off-target Cas9 edit events across 7 guide RNAs; due to the simultaneous profiling of transcriptomes and Cas9 edits in the same cells, the methodology allowed for the determination of downstream functional effects of these off-target edits.
[0151] Importantly, while the majority of the off-targets were benign, occurring in genes not expressed in the edited cell line (K562 cells), two off-target edits were identified that down-regulated two cancer-associated genes USP9X and CDC27. The functional gene set over enrichment analysis associated the DEGs of these off-target Cas9 edits with downstream functional effects on gene expression.
[0152] Given that there are several genome therapies currently in clinical trials that utilize CRISPR-Cas9 editing, using several separate modalities to characterize off-target edits that do not capture Cas9 edits and gene expression in single cells, single cell edit capture sequencing can be utilized to better screen for functional off-targets when performing guide design for CRISPR therapies. Because single cell edit capture sequencing can be utilized to determine the rarity of the edit event and the functional consequence, it can also be utilized to better select guides that have minimal functional effect and / or rare off-target edits. Hence, in addition to other uses as described herein and as contemplated by one skilled in the art, single cell edit capture sequencing can also be used to replace several methods currently used or mandated by the Food and Drug Administration (FDA) to characterize off-target Cas9 edits in a therapeutic setting, such as in silico off-target prediction, combined with GUTDE-seq and deep sequencing.II. COMPOSITIONS AND METHODS:
[0153] Provided herein are methods for editing cells, processing cells, and analyzing cells for gene editing events and changes in endogenous gene expression. Accordingly, provided herein are gene editing systems, compositions for gene editing and / or cell processing, nucleic acids, expression constructs, edited cells, and kits thereof that all may be used to carry out the methods of the present disclosure (e.g., the single-cell edit capture sequencing) method.Gene Editing Compositions
[0154] As noted above, in a first step of the methods herein, one or more cells are contacted with one or more gene editing compositions comprising (i) one or more exogenous nucleases or one or more encoding nucleic acids thereof and (ii) an exogenous insert oligonucleotide comprising a bacteriophage promoter sequence operably linked to a barcode sequence, or an encoding nucleic acid thereof. Aspects of each of these components are described further below.(i) One or more exogenous nucleases
[0155] In various aspects, the one or more nucleases comprise a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme, or any combination thereof. These are described further below. a) CRISPR Endonuclease System
[0156] The CRISPR-endonuclease system is a naturally occurring defense mechanism in prokaryotes that has been repurposed as an RNA-guided DNA-targeting platform used for gene editing. CRISPR systems include Types I, II, III, IV, V, and VI systems. In some aspects, the CRISPR system is a Type II CRISPR / Cas9 system. In other aspects, the CRISPR system is a Type V CRISPR / Cprf system. CRISPR systems rely on a DNA endonuclease, e.g., Cas9, and two noncoding RNAs — crisprRNA (crRNA) and trans-activating RNA (tracrRNA) — to target the cleavage of DNA.
[0157] The crRNA drives sequence recognition and specificity of the CRISPR- endonuclease complex through Watson-Crick base pairing, typically with a ~20 nucleotide (nt) sequence in the target DNA. Changing the sequence of the 5' 20 nt in the crRNA allows targeting of the CRISPR-endonuclease complex to specific loci. The CRISPR-endonuclease complex only binds DNA sequences that contain a sequence match to the first 20 nt of the single-guide RNA (sgRNA) if the target sequence is followed by a specific short DNA motif (with the sequence NGG) referred to as a protospacer adjacent motif (PAM).
[0158] TracrRNA hybridizes with the 3' end of crRNA to form an RNA-duplex structure that is bound by the endonuclease to form the catalytically active CRISPR-endonuclease complex, which can then cleave the target DNA.
[0159] Usually, the crRNA and TracrRNA are combined to form a single guide RNA molecule (sgRNA) comprising a spacer sequence and a scaffold sequence, as described further below.
[0160] Once the CRISPR-endonuclease complex is bound to DNA at a target site, two independent nuclease domains within the endonuclease each cleave one of the DNA strands three bases upstream of the PAM site, leaving a double-strand break (DSB) where both strands of the DNA terminate in a base pair (a blunt end).
[0161] In various embodiments, the endonuclease is a Cas9 (CRISPR associated protein 9). In various embodiments, the Cas9 endonuclease is the Cas9 can be sourced from Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, or Campylobacter. In various embodiments, the Cas9 endonuclease is from Streptococcus pyogenes, although other Cas9 homologs may be used, e.g., S. aureus Cas9, N. meningitidis Cas9, S. thermophilus CRISPR 1 Cas9, S. thermophilus CRISPR 3 Cas9, or T. denticola Cas9. In various embodiments, the CRISPR endonuclease is Cpfl, e.g., L. bacterium ND2006 Cpfl or Acidaminococcus sp. BV3L6 Cpfl. In various embodiments, the endonuclease is Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslOO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csfl, Csf4, or Cpfl endonuclease. In various embodiments, the Cas endonuclease comprises Cas9, Cas3, CaslO, or Casl2. In various embodiments, wild-type variants may be used. In various embodiments, modified versions (e.g., a homolog thereof, a recombination of the naturally occurring molecule thereof, codon-optimized thereof, or modified versions thereof) of the preceding endonucleases may be used. In various embodiments, the Cas9 comprises a mutation at D10, E762, H840, N854, N863, or D986 with reference to the position numbering of a Streptococcus pyogenes Cas9. In various embodiments, the mutation comprises D10A, E762A, H840A, N854A, N863A or D986A.
[0162] In various embodiments, the endonuclease is Casl2a. In various embodiments, the Casl2a may be sourced from Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW201 l_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, or Porphyromonas macacae.
[0163] In various embodiments, the endonuclease comprises a Cast 2b. In various embodiments, the Cast 2b can be sourced from Alicyclobacillus, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacillus, Candidatus, Desulfatirhabdium, Elusimicrobia, Citrobacter, Methylobacterium, Omnitrophicai, Phycisphaerae, Planctomycetes, Spirochaetes, or Verrucomicrobiaceae. In various embodiments, the Casl2b comprises a mutation at R911, R1000, or R1015 with reference to the position numbering of a Alicyclobacillus acidoterrestris Casl2b.
[0164] In various embodiments, the endonuclease comprises a Casl3. In various embodiments, the Cast 3 can be sourced from Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium or Acidaminococcus.
[0165] The CRISPR nuclease can be linked to at least one nuclear localization signal (NLS). The at least one NLS can be located at or within 50 amino acids of the aminoterminus of the CRISPR nuclease and / or at least one NLS can be located at or within 50 amino acids of the carboxy -terminus of the CRISPR nuclease.
[0166] Exemplary CRISPR / Cas polypeptides include the Cas9 polypeptides as published in Fonfara et al., “Phylogeny of Cas9 determines functional exchangeability of dual-RNA and Cas9 among orthologous type II CRISPR-Cas systems,” Nucleic Acids Research, 2014, 42: 2577-2590. The CRISPR / Cas gene naming system has undergone extensive rewriting sincethe Cas genes were discovered. Fonfara et al. also provides PAM sequences for the Cas9 polypeptides from various species. b) Zinc Finger Nucleases
[0167] Zinc finger nucleases (ZFNs) are modular proteins comprised of an engineered zinc finger DNA binding domain linked to the catalytic domain of the type II endonuclease Fokl. Because FokI functions only as a dimer, a pair of ZFNs can be engineered to bind to cognate target “half-site” sequences on opposite DNA strands and with precise spacing between them to enable the catalytically active Fokl dimer to form. Upon dimerization of the Fokl domain, which itself has no sequence specificity per se, a DNA double-strand break is generated between the ZFN half-sites as the initiating step in genome editing.
[0168] The DNA binding domain of each ZFN is typically comprised of 3-6 zinc fingers of the abundant Cys2-His2 architecture, with each finger primarily recognizing a triplet of nucleotides on one strand of the target DNA sequence, although cross-strand interaction with a fourth nucleotide also can be important. Alteration of the amino acids of a finger in positions that make key contacts with the DNA alters the sequence specificity of a given finger. Thus, a four-finger zinc finger protein will selectively recognize a 12 bp target sequence, where the target sequence is a composite of the triplet preferences contributed by each finger, although triplet preference can be influenced to varying degrees by neighboring fingers. An important aspect of ZFNs is that they can be readily re-targeted to almost any gene address simply by modifying individual fingers. In most applications of ZFNs, proteins of 4-6 fingers are used, recognizing 12-18 bp respectively. Hence, a pair of ZFNs will typically recognize a combined target sequence of 24-36 bp, not including the typical 5-7 bp spacer between half-sites. The binding sites can be separated further with larger spacers, including 15-17 bp. A target sequence of this length is likely to be unique in the human genome, assuming repetitive sequences or gene homologs are excluded during the design process. Nevertheless, the ZFN protein-DNA interactions are not absolute in their specificity so off-target binding and cleavage events do occur, either as a heterodimer between the two ZFNs, or as a homodimer of one or the other of the ZFNs. The latter possibility has been effectively eliminated by engineering the dimerization interface of the Fokl domain to create “plus” and “minus” variants, also known as obligate heterodimer variants, which can only dimerize with each other, and not with themselves. Forcing the obligate heterodimer preventsformation of the homodimer. This has greatly enhanced specificity of ZFNs, as well as any other nuclease that adopts these FokI variants.
[0169] A variety of ZFN-based systems have been described in the art, modifications thereof are regularly reported, and numerous references describe rules and parameters that are used to guide the design of ZFNs; see, e.g., Segal et al., Proc Natl Acad Sci, 1999 96(6):2758-63; Dreier B et al., J Mol Biol., 2000, 303(4):489-502; Liu Q et al., J Biol Chem., 2002, 277(6):3850-6; Dreier et al., J Biol Chem., 2005, 280(42):35588-97; and Dreier et al., J Biol Chem. 2001, 276(31):29466-78. c) Transcription Activator-Like Effector Nucleases (TALENs)
[0170] TALENs represent another format of modular nucleases whereby, as with ZFNs, an engineered DNA binding domain is linked to the FokI nuclease domain, and a pair of TALENs operate in tandem to achieve targeted DNA cleavage. The major difference from ZFNs is the nature of the DNA binding domain and the associated target DNA sequence recognition properties. The TALEN DNA binding domain derives from TALE proteins, which were originally described in the plant bacterial pathogen Xanthomonas sp. TALEs are comprised of tandem arrays of 33-35 amino acid repeats, with each repeat recognizing a single base pair in the target DNA sequence that is typically up to 20 bp in length, giving a total target sequence length of up to 40 bp. Nucleotide specificity of each repeat is determined by the repeat variable diresidue (RVD), which includes just two amino acids at positions 12 and 13. The bases guanine, adenine, cytosine and thymine are predominantly recognized by the four RVDs: Asn-Asn, Asn-Ile, His- Asp and Asn-Gly, respectively. This constitutes a much simpler recognition code than for zinc fingers, and thus represents an advantage over the latter for nuclease design. Nevertheless, as with ZFNs, the protein-DNA interactions of TALENs are not absolute in their specificity, and TALENs have also benefitted from the use of obligate heterodimer variants of the FokI domain to reduce off- target activity.
[0171] Additional variants of the FokI domain have been created that are deactivated in their catalytic function. If one half of either a TALEN or a ZFN pair contains an inactive FokI domain, then only single-strand DNA cleavage (nicking) will occur at the target site, rather than a DSB. The outcome is comparable to the use of CRISPR / Cas9 or CRISPR / Cpfl “nickase” mutants in which one of the Cas9 cleavage domains has been deactivated. DNAnicks can be used to drive genome editing by HDR, but at lower efficiency than with a DSB. The main benefit is that off-target nicks are quickly and accurately repaired, unlike the DSB, which is prone to NHEJ-mediated mis-repair.
[0172] A variety of TALEN-based systems have been described in the art, and modifications thereof are regularly reported; see, e.g., Boch, Science, 2009 326(5959): 1509- 12; Mak et al., Science, 2012, 335(6069):716-9; and Moscou et al., Science, 2009, 326(5959): 1501. The use of TALENs based on the “Golden Gate” platform, or cloning scheme, has been described by multiple groups; see, e.g., Cermak et al., Nucleic Acids Res., 2011, 39(12):e82; Li et al., Nucleic Acids Res., 2011, 39(14):6315-25; Weber et al., PLoS One, 2011, 6(2):el6765; Wang et al., J Genet Genes, 2014, 41(6):339-47; and Cermak T et al., Methods Mol Biol., 2015 1239: 133-59 d) Homing Endonuclease
[0173] Homing endonucleases (HEs) are sequence-specific endonucleases that have long recognition sequences (14-44 base pairs) and cleave DNA with high specificity — often at sites unique in the genome. There are at least six known families of HEs as classified by their structure, including GIY-YIG, His-Cis box, H-N-H, PD-(D / E) K, and Vsr-like that are derived from a broad range of hosts, including eukarya, protists, bacteria, archaea, cyanobacteria and phage. As with ZFNs and TALENs, HEs can be used to create a DSB at a target locus as the initial step in genome editing. In addition, some natural and engineered HEs cut only a single strand of DNA, thereby functioning as site-specific nickases. The large target sequence of HEs and the specificity that they offer have made them attractive candidates to create site-specific DSBs.
[0174] A variety of HE-based systems have been described in the art, and modifications thereof are regularly reported; see, e.g., the reviews by Steentoft et al., Glycobiology, 2014, 24(8):663-80; Belfort and Bonocora, Methods Mol Biol., 2014, 1123: 1-26; and Hafez and Hausner, Genome, 2012, 55(8): 553 -69. e) Transposases
[0175] Transposases are a class of enzyme that normally function by binding to an end of a transposon and catalyzing its movement to another part of a genome. Recently, these enzymes have been engineered for gene editing purposes. For instance, as described in Liu etal., Nature volume 631 : 593-600 (2024), which is incorporated herein by reference in its entirety, transposases can be fused to a Cas nuclease to achieve sequence specific targeted editing. f) Recombinases
[0176] Recombinases, such as serine integrases and tyrosine integrases, catalyze precise rearrangement of DNA through site-specific recombination of small sequences of DNA called attachment (att) sites. Serine recombinases make double strand breaks in DNA forming covalent 5 '-phosphoserine bonds with the DNA backbone before religation, whereas tyrosine recombinases cleave single strands forming covalent 3 '-phosphotyrosine bonds with the DNA backbone and rejoin the strands via a Holliday junction-like intermediate state. g) Hybrid Nucleases
[0177] Also contemplated are nucleases that combine one or more of the classes listed above. For instance, a recombinase may be combined with a zinc finger nuclease as described in: A conditional, zinc-finger-dependent recombinase system for DNA editing. Nat Biotechnol (2024), incorporated herein by reference in its entirety. Another example is combining a recombinase with a CRISPR nuclease as described in Liu et al., Nature volume 631 : 593-600 (2024), which is incorporated herein by reference in its entirety. Another example is the MegaTAL platform and Tev-mTALEN platform which use a fusion of TALE DNA binding domains and catalytically active HEs, taking advantage of both the tunable DNA binding and specificity of the TALE, as well as the cleavage sequence specificity of the HE; see, e.g., Boissel et al., Nucleic Acids Res., 2014, 42: 2591-2601; Kleinstiver et al., G3, 2014, 4: 1155-65; and Boissel and Scharenberg, Methods Mol. Biol., 2015, 1239: 171-96.
[0178] In a further variation, the MegaTev architecture is the fusion of a meganuclease (Mega) with the nuclease domain derived from the GIY-YIG homing endonuclease LTevI (Tev). The two active sites are positioned ~30 bp apart on a DNA substrate and generate two DSBs with non-compatible cohesive ends; see, e.g., Wolfs et al., Nucleic Acids Res., 2014, 42, 8816-29. It is anticipated that other combinations of existing nuclease-based approaches will evolve and be useful in achieving the gene editing methods described herein.(ii) Exogenous insert oligonucleotide
[0179] In various aspects, an exogenous insert oligonucleotide for use in the systems and methods of the present disclosure comprises a bacteriophage promoter sequence operably linked to a barcode sequence. As used herein, the term “operably linked” means that the promoter sequence is positioned such that the barcode sequence will be transcribed by an RNA polymerase acting at the promoter sequence. a) Bacteriophage promoters
[0180] Bacteriophages are viruses that infect and replicate only within the body of bacteria. They are classified based on their morphology and nucleic acid. Bacteriophage promoters, like prokaryotic promoters, are different than eukaryotic promoters because RNA polymerase directly binds the promoter region to initiate transcription. In contrast, RNA polymerase in eukaryotic cells can bind to one or more proteins (e.g., transcription factors) that bind the promoter sequence. Accordingly, in some aspects, a bacteriophage promoter or any other highly expressing non-eukaryotic (e.g., a prokaryotic) promoter may be used. In various aspects, the bacteriophage promoter sequence may comprise a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof. Sequences of these three promoters are provided in Table 1 below, but additional variants are known in the art and contemplated for use herein.Table 1
[0181] In some aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 7 and 103. Insome aspects, the bacteriophage promoter sequence comprises any one of SEQ ID NOs: 1 to 7 and 103. In some aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 3 and 103. In some aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 4 or 5. In some aspects, the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 6 or 7. In some aspects, the bacteriophage promoter sequence comprises any one of SEQ ID NOs: 1 to 3 and 103. In some aspects, the bacteriophage promoter sequence comprises SEQ ID NO: 4 or 5. In some aspects, the bacteriophage promoter sequence comprises SEQ ID NO: 6 or 7. b) Barcode sequence
[0182] In various aspects, the barcode sequence included in the insert oligonucleotide comprises 5 to 10 nucleotides (e.g., 5, 6, 7, 8, 9, or 10 nucleotides) that serve as a unique molecular identifier (UMI) of the barcode. In some aspects, the unique molecular identifier (UMI) of the barcode comprises 7 to 8 nucleotides. In an aspect, the barcode sequence may further comprise one or more functional sequences for coupling to an analyte or analyte tag such as a reporter oligonucleotide. Such functional sequences can include, e.g., a template switch oligonucleotide (TSO) sequence, a primer sequence (e.g., a poly T sequence, or a nucleic acid primer sequence complementary to a target nucleic acid sequence and / or for amplifying a target nucleic acid sequence, a random primer, and a primer sequence for messenger RNA). c) Other components
[0183] In various aspects, the exogenous insert oligonucleotides provided herein may comprise one or more modified nucleotides. As used herein, the term “modified nucleotide” refers to any nucleotide with functionalizing / stabilizing chemical modifications. These modifications may include but are not limited to 5’ phosphate and 5’ or 3’ phosphorothioate linkages, among others. In some aspects, the exogenous insert oligonucleotide may comprise one or two additional nucleotides (modified or not-modified) at the 5’ or 3’ end. For example,in an aspect, the exogenous insert oligonucleotide may comprise a 5’ phosphate group and a phosphorothioate linkage between terminal two or three 3’ nucleotides.
[0184] In view of the foregoing, in some aspects, the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8). In some aspects, the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT) (SEQ ID NO: 9).
[0185] In various aspects, the exogenous insert oligonucleotide is synthetic. In some aspects, the exogenous insert oligonucleotide is encoded by another synthetic nucleic acid as described further below.
[0186] In various embodiments, an exogenous insert oligonucleotide may further comprise a detectable label, such as, e.g., a fluorescent molecule. For example, in various embodiments, an exogenous insert oligonucleotide can comprise a reporter protein, such as a fluorescent protein, that is detectable based on fluorescence, or a nucleic acid molecule encoding a reporter protein. In various embodiments, the fluorescence may be from the reporter protein directly, or may be from activity of the reporter protein on a fluorogenic substrate. In various embodiments, a protein can have binding affinity to a fluorescently labeled compound, thereby resulting in a fluorescent signal. Non-limiting examples of fluorescent proteins include, e.g., green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, and ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, and ZsYellowl), blue fluorescent proteins (e.g., BFP, eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, and T-sapphire), cyan fluorescent proteins (e.g., CFP, eCFP, Cerulean, CyPet, AmCyanl, and Midoriishi-Cyan), red fluorescent proteins (e.g., RFP, mKate, mKate2, mPlum, DsRed, mCherry, mRFPl, DsRed-Express, DsRed2, DsRed2omer, DsRed2omer, DsRed Monomer , HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, and Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, and tdTomato), and any other fluorescent protein whose presence in cells can be detected by flow cytometry methods.(iii) Additional Elements
[0187] In various aspects, the one or more gene editing compositions described herein may comprise one or more additional elements as needed for proper function and improved activity of the one or more nucleases. For instance, if a CRISPR endonuclease is selected for (i), then the one or more gene editing compositions will further comprise one or more gRNAs, described below. In addition, the one or more gene editing compositions may further comprise one or more reagents for transfection, transduction, electroporation, or injection, as described further below. a) gRNAs
[0188] In various aspects, the one or more gene editing compositions herein comprise one or more guide RNAs (gRNAs) that can direct the activities of an associated endonuclease to a specific target site within a polynucleotide.
[0189] Guide RNAs in accordance with embodiments of the disclosure can comprise at least two segments: a spacer or protospacer complementary sequence and a protein binding sequence. In various embodiments, the protein binding sequence allows the gRNAs to associate with the Cas molecules. In various embodiments, the spacer or protospacer complementary sequence enables the RNP complex to associate with a spacer sequence . “Sequence” includes a section or region of a molecule, such as a contiguous stretch of nucleotides in an RNA. Some gRNAs, such as those for Cas9, can comprise two separate RNA molecules: an “activator RNA” (e.g., tracrRNA) and a “targeting RNA” (e.g., CRISPR RNA or crRNA). Other gRNAs are a single RNA molecule (single polynucleotide RNA), which can also be called a “single molecule sgRNA,” a “single guide RNA,” or an “sgRNA.” See, e.g., WO2013 / 176772, WO2014 / 065596, W02014 / 089290, WO2014 / 093622, W02014 / 099750, WO2013 / 142578, and WO2014 / 131833, each of which is incorporated by reference herein in its entirety for all purposes.
[0190] For Cas9, for example, a single guide RNA can comprise a crRNA fused to a tracrRNA (for example, via a linker). For Cpfl, for example, binding to and / or cleavage of a target sequence can be accomplished with only one crRNA. The terms “guide RNA” and “gRNA” include both double molecule (i.e., modular) gRNAs and single molecule gRNAs.
[0191] An exemplary two-molecule gRNA comprises a crRNA-like molecule (“CRISPR RNA” or “targeting RNA” or “crRNA” or “repeating crRNA”) and a corresponding tracrRNA-like molecule (“CRISPR trans-acting RNA” or “Activator RNA” or “tracrRNA”). A crRNA comprises both the DNA (single strand) segment of the DNA gRNA and a stretch of nucleotides (i.e., the crRNA tail) that forms a half of the duplex dsRNA of the segment binding the gRNA protein.
[0192] A corresponding tracrRNA (activator RNA) may comprise a stretch of nucleotides that forms the other half of the duplex dsRNA of the segment linking the gRNA protein. A nucleotide stretch of a crRNA may be complementary to and hybridizes to a nucleotide stretch of a tracrRNA to form the duplex dsRNA of the domain binding the gRNA protein. As such, each crRNA can be said to have a corresponding tracrRNA.
[0193] Guide RNAs may include modifications or sequences that provide additional desirable characteristics (e.g., modified or regulated stability; subcellular targeting; fluorescent marker tracking; a binding site for a protein or protein complex; and the like). Examples of such modifications include, for example, a 5’ cap structure (for example, a 7-methylguanylate (m7G) cap); a 3’ polyadenylated tail (i.e., a 3' poly (A) tail); a riboswitch sequence (for example, to allow regulated stability and / or accessibility regulated by proteins and / or protein complexes); a stability control sequence; a sequence that forms a duplex dsRNA (i.e., a hairpin); a modification or sequence that directs RNA at a subcellular location (for example, nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides tracking (for example, direct conjugation to a fluorescent molecule, conjugation to a portion that facilitates fluorescent detection, a sequence that allows fluorescent detection, and so on); a modification or sequence that provides a binding site for proteins (for example, proteins that act on DNA, including DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and the like); and combinations thereof. Other examples of modifications include engineered stem-loop duplex structures, engineered bump regions, engineered hair clips 3' from the stem-loop duplex structure, or any combination thereof. See, for example, U.S. Patent Publication No. US2015 / 0376586, the disclosure of which is incorporated by reference herein in its entirety for all purposes.
[0194] In some aspects, the one or more gene editing compositions provided herein may comprise more than one gRNA each configured to target a separate target gene locus.i. Ribonucleoprotein (RNP) Complexes
[0195] In various embodiments, compositions described herein include a ribonucleoprotein (RNP) complex comprising an exogenous Cas molecule and an exogenous gRNA molecule. In some aspects, the compositions provided herein may comprise more than one RNP complex targeting more than one target gene locus. b) Reagents
[0196] In various aspects, the one or more gene editing compositions herein may comprise a transfection agent, a transduction reagent and / or an electroporation agent. These reagents assist in introducing the one or more nucleases and the exogenous insert oligonucleotide (along with any other components) into the one or more cells. In an aspect, an electroporation agent that may be used in the gene editing compositions herein is an exogenous enhancer oligonucleotide that increases electroporation delivery efficiency. As a non-limiting example, an Alt-R Cas9 Electroporation Enhancer (INTEGRATED DNA TECHNOLOGIES) may be included. Other electroporation enhancers are known in the art. In an aspect, a transfection agent may comprise an electrolytic / isotonic buffer. A transduction agent may comprise any agent that aids in delivery of a virus into a cell.
[0197] Additional examples of reagents include, but are not limited to: buffers, acidic solution, basic solution, temperature-sensitive enzymes, pH-sensitive enzymes, light-sensitive enzymes, metals, metal ions, magnesium chloride, sodium chloride, manganese, aqueous buffer, mild buffer, ionic buffer, inhibitor, enzyme, protein, polynucleotide, antibodies, saccharides, lipid, oil, salt, ion, detergents, ionic detergents, non-ionic detergents, and oligonucleotides.Transcription Compositions
[0198] In various aspects, methods of the present disclosure comprise contacting one or more cells with a transcription composition comprising an RNA polymerase. Accordingly, in various aspects, a transcription composition is provided herein. The transcription composition comprises an RNA polymerase and may, in various embodiments, comprise one or more further components deliver the RNA polymerase into a cell (e.g., a fixed cell) and / or to allow transcription to proceed.i) RNA polymerases
[0199] RNA polymerases are enzymes responsible for synthesizing RNA from a DNA transcript. In various aspects, the compositions and methods of the disclosure use prokaryotic or bacteriophage RNA polymerases which differ fundamentally from eukaryotic RNA polymerases. Specifically, eukaryotic RNA polymerases can involve one or more protein factors associating with a target gene and induce transcription. These protein factors include transcription factors and the like that directly bind to regulatory regions in the genome (e.g, promoters). In contrast, RNA polymerases from prokaryotes and bacteriophages are capable of directly binding a regulatory element in a nucleic acid (e.g, a promoter).
[0200] Accordingly, in some aspects, the RNA polymerase included in a transcription composition herein comprises a bacteriophage RNA polymerase (or phage polymerase). In an aspect, the RNA polymerase comprises a T7 polymerase. In an aspect, the RNA polymerase comprises a T3 polymerase. In an aspect, the RNA polymerase comprises an SP6 polymerase. Any variant thereof is also contemplated. ii) Further Components
[0201] In various aspects, the transcription composition may comprise one or more components that assist in delivering the RNA polymerase into cells (e.g., fixed cells). These components may comprise, for instance nucleoside triphosphates, magnesium, dithiothreitol (DTT), spermidine, inorganic pyrophosphate or any combination thereof. Any commercial kit for in vitro transcription using a T7, T3 or SP6 RNA polymerase may also be used herein as a transcription composition for use in the methods disclosed herein.Nucleic Acids, Expression Constructs and Vectors
[0202] In various aspects, one or more components in the gene editing compositions (e.g., nucleases, insert oligonucleotides, gRNAs) or the transcription compositions (e.g., RNA polymerase) may be encoded by a nucleic acid. Accordingly, provided herein are nucleic acids encoding one or more of the components described above (e.g., one or more nucleases, one or more exogenous insert oligonucleotides, and / or one or more gRNAs and / or one or more RNA polymerases). Also provided are one or more expression constructs comprising these nucleic acids.
[0203] Expression constructs provided herein comprise a nucleic acid encoding one or more components of the gene editing compositions provided herein (e.g., nucleases, insert oligonucleotides, gRNAs) and one or more regulatory elements configured to regulate the expression of (a). Regulatory elements can include, but are not limited to, promoters and enhancers. The regulatory elements may be cell or organ specific to direct expression in a desired cell type. In some aspects, the one or more regulatory element comprises a mammalian promoter sequence. In some aspects, the mammalian promoter sequence is a human promoter sequence. In various embodiments, the human promoter is a U6 promoter sequence. In various aspects, the expression constructs may further comprise a selective agent, such as for example, an ampicillin resistance gene (amp), to facilitate identification and selection of positively transfected cells.
[0204] In various aspects, additional expression constructs are provided that provide additional components for the MMEJ repair pathway (described below). In some aspects, the expression construct may comprise genes encoding FEN1, Ligase III, MRE11, NBS1, PARP1 and XRCC1 along with any regulatory elements controlling their expression. In some aspects, the regulatory elements comprise a promoter (e.g., a mammalian promoter described above). In some aspects, these expression constructs may further comprise necessary sequences for selection (e.g. an ampicillin sequence).
[0205] In various aspects, the nucleic acids (and associated expression constructs and vectors thereof) may be codon optimized for expression in a desired cell type, such as prokaryotic or eukaryotic cells. As used herein, ‘codon optimization’ can refer to a process of modifying a nucleic acid sequence for enhanced expression in the host cells of interest by replacing at least one codon or more of the native sequence with codons that are more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Various species exhibit particular bias for certain codons of a particular amino acid. As contemplated herein, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the “Codon Usage Database.
[0206] Expression constructs and their included nucleic acids may be further comprised in one or more vectors. The term “vector” as used herein can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded, double- stranded, orpartially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g. circular); nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of polynucleotides known in the art. One type of vector is a “plasmid,” which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein virally-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g. retroviruses, replication defective retroviruses, adenoviruses, replication defective adenoviruses, and adeno-associated viruses). Other vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell. Recombinant expression vectors can include a nucleic acid of the present inventive concept in a form suitable for expression of the nucleic acid in a host cell, can mean that the recombinant expression vectors include one or more regulatory elements, which can be selected on the basis of the host cells to be used for expression, that is operatively-linked to the nucleic acid sequence to be expressed.
[0207] In certain embodiments, conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids in cells, such as prokaryotic cells, eukaryotic cells, plant cells, mammalian cells, or target tissues. Such methods can be used to administer nucleic acids encoding components of an CRISPR-Cas9 system herein to cells in culture, or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g. a transcript of a vector described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle, such as a liposome. Viral vector delivery systems include DNA and RNA viruses, which have either episomal or integrated genomes after delivery to the cell. Any gene therapy method known in the art is contemplated of use herein. Methods of non-viral delivery of nucleic acids include are contemplated herein. Adeno-associated virus (“AAV”) vectors can also be used to transduce cells with target nucleic acids, e.g., in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures.
[0208] In various embodiments, the expression construct is a viral vector. Exemplary viral vectors include lentiviral vectors and adeno-associated viral vectors (AAV). In some aspects, the expression construct is an adeno-associated viral (AAV) vector. AAVs are small viruses which integrate site-specifically into the host genome and can therefore deliver a transgene. Inverted terminal repeats (ITRs) are present flanking the AAV genome and / or the transgene of interest and serve as origins of replication. Also present in the AAV genome arerep and cap proteins which, when transcribed, form capsids which encapsulate the AAV genome for delivery into target cells. Surface receptors on these capsids which confer AAV serotype, which determines which target organs the capsids will primarily bind and thus what cells the AAV will most efficiently infect. There are twelve currently known human AAV serotypes. In various embodiments, any mammalian AAV serotypes can be used herein for delivering the encoding nucleic acids described herein. Adeno-associated viruses are among the most frequently used viruses for gene therapy for several reasons. First, AAVs do not provoke an immune response upon administration to mammals, including humans. Second, AAVs are effectively delivered to target cells, particularly when consideration is given to selecting the appropriate AAV serotype. Finally, AAVs have the ability to infect both dividing and non-dividing cells because the genome can persist in the host cell without integration. This trait makes them an ideal candidate for gene therapy.
[0209] In various embodiments, polynucleotides disclosed herein (e.g., gRNA, Cas9) can be delivered to a cell using at least one AAV vector. An AAV vector typically comprises a protein-based capsid, and a nucleic acid encapsidated by the capsid. The nucleic acid may be, for example, a vector genome comprising a transgene flanked by inverted terminal repeats. The AAV “capsid” is a near-spherical protein shell that comprises individual “capsid proteins” or “subunits.” AAV capsids typically comprise about 60 capsid protein subunits, associated and arranged with T=1 icosahedral symmetry. When an AAV vector is described herein as comprising an AAV capsid protein, it will be understood that the AAV vector comprises a capsid, wherein the capsid comprises one or more AAV capsid proteins (i.e. , subunits). Also described herein are “viral-like particles” or “virus-like particles,” which refers to a capsid that does not comprise any vector genome or nucleic acid comprising a transgene. The virus vectors of the present disclosure can further be “targeted” virus vectors (e.g., having a directed tropism) and / or a “hybrid” parvovirus (i.e., in which the viral TRs and viral capsid are from different parvoviruses) as described in international patent publication WO 00 / 28004 and Chao et al., (2000) Molecular Therapy 2:619. The virus vectors of the present disclosure can further be duplexed parvovirus particles as described in international patent publication WO 01 / 92551 (the disclosure of which is incorporated herein by reference in its entirety). Thus, in various embodiments, double stranded (duplex) genomes can be packaged into the virus capsids of the present inventive concept. Further, the viral capsid or gene elements can contain other modifications, including insertions, deletions and / or substitutions.
[0210] In various embodiments, the isolated nucleic acids encoding a gRNA and / or the nucleaes herein may be packaged into an AAV vector (e.g., a AAV-Cas9 vector). In various embodiments, the AAV vector is a wildtype AAV vector. In various embodiments, the AAV vector contains one or more mutations. In various embodiments, the AAV vector is isolated or derived from an AAV vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11 or any combination thereof.
[0211] Exemplary AAV vectors contain two ITR (inverted terminal repeat) sequences which flank a central sequence region comprising the encoding nucleic acid sequence. In various embodiments, the ITRs are isolated or derived from an AAV vector of serotype AAV1 , AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV 11 or any combination thereof. In various embodiments, the ITRs comprise or consist of full-length and / or wildtype sequences for an AAV serotype. In various embodiments, the ITRs comprise or consist of truncated sequences for an AAV serotype. In various embodiments, the ITRs comprise or consist of elongated sequences for an AAV serotype. In various embodiments, the ITRs comprise or consist of sequences comprising a sequence variation compared to a wildtype sequence for the same AAV serotype. In various embodiments, the sequence variation comprises one or more of a substitution, deletion, insertion, inversion, or transposition. In various embodiments, the ITRs comprise or consist of at least 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 , 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150 base pairs. In various embodiments, the ITRs comprise or consist of 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111 , 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141 , 142, 143, 144, 145, 146, 147, 148, 149 or 150 base pairs. In various embodiments, the ITRs have a length of 110 ± 10 base pairs. In various embodiments, the ITRs have a length of 120 ± 10 base pairs. In various embodiments, the ITRs have a length of 130 ± 10 base pairs. In various embodiments, the ITRs have a length of 140 ± 10 base pairs. In various embodiments, the ITRs have a length of 150 ± 10 base pairs. In various embodiments, the ITRs have a length of 115, 145, or 141 base pairs.
[0212] In various embodiments, the AAV vector may contain one or more nuclear localization signals (NLS). In various embodiments, the AAV-Cas9 vector contains 1, 2, 3, 4,or 5 nuclear localization signals. Exemplary NLS include the c-myc NLS, the SV40 NLS, the hnRNPAI M9 NLS, and the nucleoplasmin NLS.Gene Editing System
[0213] In various aspects, a gene editing system is provided comprising: (a) one or more nucleases or one or more encoding nucleic acids thereof; and (b) an exogenous insert oligonucleotide or encoding nucleic acid thereof, the exogenous insert oligonucleotide comprising a bacteriophage promoter sequence and a barcode sequence. The components of(a) and (b) may be delivered in any gene editing composition provided above. In some aspects, (a) and (b) are delivered in the same gene editing composition. In some aspects, (a) and (b) are provided in different gene editing compositions. Likewise, one or more of (a) and(b) may comprise a nucleic acid encoding the nucleases and / or the exogenous insert oligonucleotide, as provided above.
[0214] In various methods, the gene editing system may be introduced into a cell wherein the one or more nucleases cleave an endogenous nucleic acid. The cleavage event may then be repaired and the exogenous insert oligonucleotide inserted into the cleavage site using a repair mechanism. In various aspects, the cleavage event may comprise a double stranded break which can be repaired using a number of classical repair mechanisms (e.g., non- homologous end joining (NHEJ), microhomology-mediated end joining mechanism (MMEJ) or homology directed repair (HR). In various aspects, for instance, a non-homologous end joining (NHEJ) mechanism can be used to repair a double stranded break caused by the nuclease event. In various embodiments, a microhomology -mediated end joining mechanism (MMEJ) can be used to repair a double stranded break caused a nuclease. The MMEJ pathway is unrelated to the NHEJ pathway, and in various embodiments, the MMEJ pathway produces deletions or insertions which can be much larger than the possible insertions and deletions occurring in the NHEJ pathway. In various embodiments, the NHEJ repair pathway causes base pair alterations such as substitutions, deletions, or insertions. In various embodiments, the alterations downregulate mRNA expression, upregulate mRNA expression, destroy a promoter site such that a polymerase can no longer bind to the targeting locus , and / or a protein encoded by the targeting locus comprises increased or decreased activity. For example, in various embodiments, the alteration causes a frameshift resulting in a nonfunctioning protein. In various embodiments, the alteration can include a newly added stop codon, thereby, shortening the length of the mRNA transcript.
[0215] In various embodiments, an advantage of using the MMEJ can be for an insertion of an insert oligonucleotide using a short homology sequence on each side of the insert oligonucleotide. In contrast, the NHEJ pathway comprises the advantage of the ability to function without sequence homology and can disrupt the sequence, thus leading to a KO of the targeting locus. The NHEJ repair pathway causes random alterations. In the case of MMEJ, it is directed with the insert oligonucleotide including a homology sequence spanning both sides of the cut site. Homology directed repair (e.g., homologous recombination, HR), has the highest fidelity but is hardest to do in non-replicating cells as it only occurs during certain phases of the cell cycle. Therefore, the fact that the gene editing methods herein can be performed with non-homologous methods is an advantage for expanding the types of cells to edit.Methods of Delivering Gene Editing Components
[0216] In various aspects, the methods herein comprise contacting one or more cells with one or more gene editing composition and introducing the one or more components in those gene editing compositions into the cell (e.g., the one or more nucleases or encoding nucleic acid thereof, the exogenous insert oligonucleotide or encoding nucleic acid thereof, and the one or more gRNAs or encoding nucleic acids thereof). Methods of introducing the components into the cells are known in the art and may include, for example, transfection, transduction, electroporation, injection or any combination thereof.
[0217] A variety of delivery systems can be used to introduce the gene editing components into a host cell. In accordance with these embodiments, systems of use for embodiments disclosed herein can include, but are not limited to, yeast systems, lipofection systems, microinjection systems, biolistic systems, virosomes, liposomes, immunoliposomes, polycations, lipid nucleic acid conjugates, virions, artificial virions, viral vectors, electroporation, cell permeable peptides, nanoparticles, nanowires, and / or exosomes. In various embodiments, the expression system and methods comprise a lentiviral capsid-based expression system capable of reducing off-target effects. See, e.g., “Delivering Cas9 / sgRNA ribonucleoprotein (RNP) by lentiviral capsid-based bionanoparticles for efficient ‘hit-and- run’ genome editing” Nucleic Acids Research, Volume 47, Issue 17, 26 September 2019, Page e99, https: / / doi.org / 10.1093 / nar / gkz605, the disclosure of which is incorporated by reference herein in its entirety. In various embodiments, a short-term transfer can be desirableand can include lipofection or electroporation. In various embodiments, lipofection or electroporation produce fewer off-target effects than lentiviral-based expression systems.Gene Edited Cells
[0218] In various aspects, the present disclosure also provides genetically modified cells that are edited by any of the gene editing methods disclosed herein. In particular aspects, a genetically modified cell of the present disclosure comprises one or more nucleic acid insertions in an endogenous nucleic acid, wherein each nucleic acid insertion comprises a bacteriophage promoter sequence and a barcode.
[0219] In various aspects, the nucleic acid insertion comprises the same sequence as any exogenous insert oligonucleotide described herein above. For instance, the nucleic acid insertion may comprise any of the bacteriophage promoter sequences described herein above (e.g., T7, T3, SP6, or any combination thereof). In various aspects, the nucleic acid insertion comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 7. In some aspects, the nucleic acid insertion may comprise any one of SEQ ID NOs: 1 to 7. In some aspects, the nucleic acid insertion may comprise a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 3 and 103. In some aspects, the nucleic acid insertion may comprise a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 4 or 5. In some aspects, the nucleic acid insertion may comprise a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 6 or 7. In some aspects, the nucleic acid insertion may comprise any one of SEQ ID NOs: 1 to 3 and 103. In some aspects, the nucleic acid insertion may comprise SEQ ID NO: 4 or 5. In some aspects, the nucleic acid insertion may comprise SEQ ID NO: 6 or 7.
[0220] Also contemplated are any genetically modified cell comprising a nucleic acid encoding any of the components of the gene editing system provided herein or any expression construct thereof. Also contemplated are any genetically modified cells that have been edited by the gene editing system described herein.
[0221] The cells may be eukaryotic. Eukaryotic cells can be yeast, fungi, algae, plant, animal, or human cells. Eukaryotic cells can be those derived from a particular organism, such as a mammal, including but not limited to human, mouse, rat, rabbit, dog, or non-human mammal including non-human primate. Plant cells can include, without limitation, cells from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen and microspores. In some aspects, the cell is mammalian. In some aspects, the cell is murine or human.
[0222] Other examples of fixation reagents include, for example, organic solvents such as alcohols (e.g., methanol or ethanol), ketones (e.g., acetone), and aldehydes (e.g., paraformaldehyde, formaldehyde (e.g., formalin), or glutaraldehyde). As described herein, cross-linking agents may also be used for fixation including, without limitation, disuccinimidyl suberate (DSS), dimethylsuberimidate (DMS), formalin, and dimethyladipimidate (DMA), dithio-bis(-succinimidyl propionate) (DSP), disuccinimidyl tartrate (DST), and ethylene glycol bis(succinimidyl succinate) (EGS). In some cases, a cross-linking agent may be a cleavable cross-linking agent (e.g., thermally cleavable, photocleavable, etc.). In some cases, more than one fixation reagent can be used in combination when preparing a fixed biological sample.
[0223] In some aspects, the gene edited cells are fixed. The term “fixed” as used herein with regard to cells generally refers to the state of being preserved from decay and / or degradation. “Fixation” generally refers to a process that results in a fixed sample, and in some instances can include contacting the biomolecules within a biological sample with a fixative (or fixation reagent) for some amount of time, whereby the fixative results in covalent bonding interactions such as crosslinks between biomolecules in the sample. A “fixed cell” may generally refer to a cell that has been contacted with a fixation reagent or fixative. For example, a formaldehyde-fixed cell has been contacted with the fixation reagent formaldehyde. “Fixed cells” generally refer to cells that have been in contact with a fixative under conditions sufficient to allow or result in the formation of intra- and inter-molecular covalent crosslinks between biomolecules in the biological sample. Generally, contact of a cell with a fixation reagent (e.g., paraformaldehyde or PF A) results in the formation of intra- and inter-molecular covalent crosslinks between biomolecules in the biological sample. In some cases, provision of the fixation reagent, such as formaldehyde, may result in covalent crosslinks within RNA, DNA, and / or protein molecules. For example, the widely usedfixative reagent, paraformaldehyde or PF A, fixes tissue samples by catalyzing crosslink formation between basic amino acids in proteins, such as lysine and glutamine. Both intramolecular and inter-molecular crosslinks can form in the protein. These crosslinks can preserve protein secondary structure and also eliminate enzymatic activity in the preserved tissue sample. Examples of fixation reagents include but are not limited to aldehyde fixatives (e.g., formaldehyde, also commonly referred to as “paraformaldehyde,” “PF A,” and “formalin”; glutaraldehyde; etc.), imidoesters, NHS (N-Hydroxysuccinimide) esters, and the like.
[0224] Changes to a characteristic or a set of characteristics of a cell or cellular constituents (e.g., incurred upon interaction with one or more fixation agents) may be at least partially reversible (e.g., via rehydration or de-crosslinking). Alternatively, changes to a characteristic or set of characteristics of a cell or cellular constituents (e.g., incurred upon interaction with one or more fixation agents) may be substantially irreversible.III. KITS
[0225] Aspects of the disclosure include kits for performing any of the methods (e.g., gene editing methods) In various embodiments, a kit comprising components allowing for direct introduction of gene editing machinery into a cell using transfection, electroporation, or injection can be provided. In various embodiments, the kit can comprise an insert oligonucleotide comprising bacteriophage promoter and a barcode as provided herein. In various embodiments, the kit may further comprise one or more nucleases as well as any associated components for their activity (e.g., gRNAs for Cas nucleases). In some aspects, the kit may further comprise a transfection, transduction, electroporation, or injection reagent. For example, the kit may comprise an exogenous “enhancer” oligonucleotide that increases electroporation delivery efficiency.
[0226] In various embodiments, all the encoding nucleotide sequences can be expressed from the same construct. In various embodiments, one or more of the encoding nucleotide sequence can be expressed from the same construct. In various embodiments, each of the encoding nucleotide sequences can be expressed from different constructs.V. EXAMPLESExample 1 - Direct Readout of Genome Perturbations and Transcription Outcomes by Single-Cell Edit Capture
[0227] A great gap in biology is how millions of genetic variants with small individual effects control gene expression and disease risk. A key opportunity to address this gap are single-cell CRISPR screens (Perturb-seq) that link hundreds or thousands of perturbations to molecular phenotypes by joint sequencing of cellular nucleic acid and viral vector-integrated guides, or “guide capture” (shown in FIG. 10) . However, Perturb-seq and all CRISPR screens have a key limitation: guide capture poorly correlates with the occurrence of genome perturbations, impairing Perturb-seq effectiveness. This disconnect is caused by variable Cas effector and guide expression levels, variable on-target and off-target guide activity, and viral genome recombination that disrupts guide vector sequences (FIG. 11). As a result, guide capture defines guide-positive cell populations that contain a substantial fraction of confounding unperturbed and off-target cells, and it fails to capture many cells with genuine perturbations. This yields noisy, dampened differential analyses and poor power to study the small effects of non-coding regulatory sequences and variants. To overcome this key limitation of guide capture, we have developed single-cell “edit capture” — direct sequencing of true genome perturbations in single cells — and implement this technology in a new method for joint sequencing of genome-wide edit events and transcriptome mRNA. Single-cell edit capture sequencing, as described herein, has 3 steps that use off-the-shelf kits and no specialized equipment: (1) CRISPR knock-in of a barcoded bacteriophage T7 promoter that is inert in human cells, to generate promoter-labeled genome edits; (2) cell fixation and in vitro transcription by T7 RNA polymerase to generate barcoded edit-marking “fTRNA”; (3) splitpool RNA sequencing and analysis of fTRNA and total cellular mRNA (FIG. 12). In a pilot experiment targeting chromatin remodeler genes in K562 cells, Single-cell edit capture sequencing identifies true on-target and off-target edit events that aid differential expression analysis (FIG. 13). A split-pool RNA sequencing protocol (e.g., SPLiT-seq) was used to derive multiple fTRNA reads (FIG. 14) which were then analyzed to mark edit sites (FIG. 15) resulting in true edit capture in single cells (FIG 16). This directly resulted in off-target capture in single cells (FIG. 17). Single-cell edit capture sequencing can be scaled for arrayed or pooled screens that could potentially characterize all cis-regulatory elements and non-coding variants in any cell type. Overall, as shown in FIG. 18, this technology has many features including the ability to identify true on / off target events using off-the shelf, one library, reagents and analytes and no viruses. Its applications include targeted gene-knockout screens, guide scoring, and gene / cell therapy characterization.Example 2 - Joint profiling of Cas9 edit events and transcriptomes with Single-cell edit capture sequencing
[0228] A longstanding barrier in CRISPR-Cas9 genome engineering has been the inability to directly connect both on- and off-target edits to their effects on gene expression. Here, we present a first-in-kind single-cell edit capture sequencing (Single-cell edit capture sequencing) technology that jointly profiles Cas9 genome edits and transcriptomes by singlecell RNA-seq. The single-cell edit capture sequencing workflow consists of three steps. First, unbiased labeling of Cas9 cleavage events with phage T7 promoter through live-cell electroporation of donor oligonucleotides. Second, cell fixation and T7 in situ transcription to generate edit-reporting T7 RNA. Third, combinatorial indexing and sequencing of in situ and endogenous transcripts with off-the-shelf SPLiT-seq. single-cell edit capture sequencing achieves 40-80% edit labeling at six tested genome loci in three human cell types, including primary T cells. We performed single-cell edit capture sequencing on 9,500 K562 cells, targeting four chromatin remodeler genes with seven guide RNAs. With this data, we demonstrate that single-cell edit capture sequencing achieves: (1) direct single-cell capture of on-target edit events, (2) high-throughput single-cell detection of Cas9 off-target events, (3) single-cell edit allele counting, and (4) association of edit events and gene expression. The simple single-cell edit capture sequencing protocol, as described herein, uses commercial kits, standard laboratory equipment, and can operate without virus, enabling any laboratory to perform these experiments. We envision two immediate impacts of single-cell edit capture sequencing: targeted and in vivo ribonucleoprotein CRISPR screens in hard-to-transduce cell types, and detection of off-target gene perturbations by guide RNAs under active clinical development. Future scaling of single-cell edit capture sequencing with pooled guides will enable powerful Perturb-seq screens, validation and unique design of widely used guide RNA libraries, and high-throughput assessment of CRISPR therapeutics.Example 3 - Homology-free knock-in labels Cas9 edits with T7 promoter
[0229] To develop a method for single-cell edit capture sequencing, we first established the ability to label Cas9 genome edits with T7 promoter sequences using a homology-free knock-in strategy. Knock-in is performed by electroporation of live cells with Cas9-guide RNA ribonucleoprotein (RNP) and a double-stranded donor DNA composed of synthetic annealed oligonucleotides (dsODN). RNP generates genome double-strand breaks, and donor DNA is inserted at break repairs by non-homologous end joining (FIG. 19A). Compared toalternative labeling strategies, the key advantages of homology -free knock-in are the unbiased labeling of on-target and off-target Cas9 edits with the same donor sequence, effective labeling in both dividing and non-dividing cell types, compatibility with sensitive cell types that would be intolerant of larger plasmid or viral donors, and the counteraction of large cytotoxic chromosome rearrangements.
[0230] To establish an effective donor DNA for Cas9 edit labeling, a panel of 27 homology-free donor constructs of length 19-55 bp encoding T7 and SP6 phage promoter sequence variants was made (FIG. 20A). To maximize gene-disrupting frame shifts upon donor insertion in coding sequence, the length of all donors was set to 3n+l to yield +1 shifts upon precise insertion and +2 shifts upon common 1 base pair (bp) templated insertion by Cas9. Donor knock-in efficiency was quantified by TIDE (tracking indels by decomposition (FIG. 19B) and compared labeling performance to a published GUIDE-seq donor that is effective in human primary cells and commonly used for gene therapy characterization (FIG. 20B) Knock-in efficiency varied widely across designs, suggesting that donor inserting is highly sensitive to donor length and composition (FIG. 19C). Over 60% (17 / 27) of donor designs failed to generate detectable knock-in. However, three designs (donors 01, 02, 03) achieved strong knock-in that was 6-fold higher than the GUIDE-seq donor (FIG. 19C). These top-performing donor sequences were composed of optimized 30-bp T7 promoter variants developed for Cel-seq+ and 2-bp sequence ends that match the total length (34 bp), outer bases (5'-GT. . . AT-3'), and base modifications (5' phosphate, two 3' phosphorothioates) of the GUIDE-seq donor (FIG. 20B). The transcribed portion of the donor 02 promoter variant was shown to improve sequencing library amplification, so we selected this design for further single cell edit capture sequencing development (FIG. 19D, FIG. 20B). This evaluation of T7 promoter donors established an optimal design for efficient labeling of Cas9 edit events for T7 transcription.
[0231] To determine that donor knock-in was effective in multiple cell types and genome regions, a total of 112 donor knock-ins with four cell types (GM12878, Jurkat, K562, and primary T cells) and 23 guide RNAs targeting seven genome regions (ARID1A, B2M, CHD3, CHD4, CTLA4, INO80, SMARCA4) were performed. Results are summarized for 70 samples (11 controls, 59 knock-ins) that yielded TIDE results with R2> 0.5 (FIG. 19E, FIG. 20C). As expected, treatment with donor DNA or Cas9 RNP alone resulted in unedited and variable indel sequences, respectively, with some RNP deletions approaching -50 bp. In contrast,treatment with both donor DNA and RNP resulted in consistent +34 bp insertion events, indicating successful labeling of Cas9 edits with full-length donor DNA in a variety of cell and genome contexts (FIG. 19C). High knock-in efficiency of 40-80% was achieved at six of seven genome regions and in all cell types except Jurkat cells (FIG. 19F). In Jurkats up to 20% knock-in and a high proportion of small indels was observed, possibly due to overexpressed end-joining machinery that favors rapid, imprecise indel formation. Excluding Jurkats, donor knock-in decreased the formation of variable indels and larger deletions (FIG. 19E, FIG. 20D). Among 51 samples with detectable donor insertion events, 75% of events were precise +34 bp insertions, with low rates of +30 to +38 bp insertions due to additional small indels, indicating the insertion of full-length T7 promoters at Cas9 genome edits (FIG. 19G). As designed, >90% of donor insertion sizes correspond to frame-shifting events. To confirm that frame shifts result in gene disruption, target gene and protein expression were measured by reverse-transcription qPCR (RT-qPCR) and flow cytometry, respectively. Donor knock-in cells exhibited target transcript knockdown (FIG. 20G). Furthermore, in a sample of cells with 40% donor knock-in alleles and <10% indel alleles, we observed target protein knockdown in 50% of cells (FIG. 20G, FIG. 20H). Together, these results established that homology-free knock-in of our donor DNA achieves strong, versatile T7 promoter labeling of Cas9 edits suitable for gene perturbation.
[0232] Unexpectedly, it was observed that delivery of RNP and donor DNA resulted in two discrete rates of indel formation: a “high indel” outcome with low donor insertion, or a “low indel” outcome with the potential for very high donor insertion (FIG. 19E, FIG. 19H). The low indel group yielded 5-fold higher donor insertion (34.9% vs. 6.8%) and 3-fold lower indels (24.0% vs. 78.8%) (FIG. 20D). To explain this bifurcated editing fate, TIDE results were surveyed for associations with cell type, genome site, and guide RNA. Among 59 knock-in samples, editing outcome did not strictly associate with cell type or genome site (FIG. 21A, FIG. 21B) However, in the subset of 49 samples treated with 12 guide RNAs that were each used in two or more experiments (excluding Jurkat samples), we observed that a given guide always generated the same type of outcome (FIG. 21A, FIG. 21B). We compared guide protospacer sequences within each outcome group and observed limited similarity in high-indel guide sequences (FIG. 21C). However, low-indel guides showed a strong conservation of guanine at protospacer position -2 (FIG. 19H), a feature previously associated with increased Cas9 cleavage efficiency. These observations suggest that certainguide sequences may specify a Cas9 cleavage rate that is optimal for minimizing small indel formation and maximizing homology-free donor insertion.Example 4 - In situ transcription (1ST) uniquely marks Cas9 edit sites with T7 RNA in fixed cells
[0233] Next T7 1ST of promoter-labeled Cas9 cleavage events was systematically developed, starting with IVT on purified genomic DNA (FIGS. 22A-22I). Since previous 1ST methods used T7 to amplify integrated viral genomes, it was first confirmed that phage RNA polymerase transcribes human genomic DNA. In IVT reactions on unedited genomic DNA, both T7 and SP6 RNA polymerases generate in vitro RNA (FIG. 23A), confirming templated polymerization of human genome sequence. Also, RNA yield increased by up to 50% when IVT was performed in a common nuclei buffer for ATAC-seq, confirming phage polymerase activity is compatible with human nuclei.
[0234] Since T7 polymerase activity is strictly promoter-specific, the above results suggested that the human genome contains T7 promoter-like sequences. To test this, the 18- bp core T7 promoter sequence (5'-TAATACGACTCACTATAG-3' (SEQ ID NO: 1)) was aligned to human reference genome hg38 using BLAST (basic local alignment search tool). BLAST revealed 14 sequences on 9 chromosomes with at least 90% sequence homology (16 / 18 bp), including 5 sequences with a preserved terminal guanine at the T7 promoter’s transcription start site (FIG. 20B). To determine whether these sites activate T7 RNA polymerase, the top three sequences (17 / 18 bp homology with preserved terminal G) were selected and downstream T7 RNA was quantified in three cell types by reverse-transcription quantitative PCR (RT-qPCR), using random hexamer primers to capture non-polyadenylated T7 transcripts. It was observed that 2 / 3 promoter-like sequences gave 50-fold induction of T7 RNA compared to five control sites without sequence homology (FIG. 23B). This indicated that predictable promoter-like sites are a likely source of background donor-independent T7 RNA, and suggested the importance that cleavage-specific T7 transcription exceed these background signals.
[0235] Therefore, it was determined that donor-labeled Cas9 edits generate stronger T7 amplification than these background sites. T7 IVT was performed on donor-labeled genomic DNA in four cell types (including primary T cells) and two genome sites (B2M and CTLA4), and T7 RNA was quantified by RT-qPCR (FIG. 22A). Consistent 500-fold induction of T7RNA was observed at donor-labeled Cas9 edit sites (FIG. 22B, FIG. 22C). This was 10-fold higher than the strongest background site identified by BLAST (FIG. 23C). Next, the length and direction of T7 RNA extension was characterized. Most transcripts extended <1 kilobase (kb) from the knock-in site, with a small fraction of transcripts extending >10 kb away (FIG. 23B, FIG. 23C), indicating a concentration of short T7 transcripts at cleavage sites. At background sites, T7 RNA was observed only along the 3' flank of the promoter-like sequence. At Cas9 edits however, T7 RNA extended in both directions with similar abundance, indicating variable orientation of the homology -free donor promoter (FIGS. 23B- 23D). This anti-parallel signature distinguished genuine Cas9 cleavage from background T7 transcription. Together, these IVT results demonstrated that donor-labeled Cas9 edits encode a functional T7 promoter that strongly marks Cas9 edit events with distinct RNA features.
[0236] Next, to determine the performance of T7 RNA generation in intact chromosomes, T7 1ST was performed on isolated nuclei (FIG. 22D). Nuclei isolation was performed on three cell lines with donor-labeled Cas9 edits \. B2M or CTLA4, followed by immediate 1ST, total RNA extraction, and RT-qPCR. Comparable to IVT results, up to 500-fold induction of T7 RNA was observed at donor insertions (FIGS. 22E-22F and FIGS. 24A-24D), indicating effective 1ST on chromosomal DNA. 1ST transcripts were not detectable at the 10 kb range at either loci (FIGS. 24A-24D), suggesting that nuclei 1ST generates even shorter, more localized T7 RNA compared to IVT. Next T7 RNA abundance was compared to that of endogenous mRNA from RPL24, a highly expressed ribosome subunit. Importantly, it was observed that levels of T7 RNA at CTLA4 approached those of RPL24 mRNA, while levels of T7 RNA at B2M were more difficult to compare due to elevated B2M mRNA background (FIGS. 24A-24D) These results established that 1ST can generate strong cleavage-marking T7 RNA.
[0237] Finally, 1ST of Cas9 genome edits was developed in fixed cells (FIG. 22G), reasoning that fixed-cell 1ST would enable the analysis of both T7 and endogenous transcripts by scRNA-seq. Specifically, fixed-cell 1ST that was compatible with SPLiT-seq was developed. SPLiT-seq was selected over alternative fixed-cell scRNA-seq methods because it uses blended random and oligo-dT primers, enabling joint capture of nonpolyadenylated T7 RNA and mRNA into one sequencing library. Additionally, high-quality SPLiT-seq can be conveniently performed with off-the-shelf EVERCODE kits (Parse Biosciences).
[0238] In contrast to previous in situ transcription reports that have used methanol fixation, SPLiT-seq uses paraformaldehyde (PF A) fixation, so it was first determined that T7 1ST reactions are functional in PFA-fixed cells. EVERCODE fixation and immediate T7 1ST was performed of cells with donor-labeled Cas9 cleavage. Both edit-specific and background T7 RNA were observed, indicating that T7 RNA polymerase was active in PFA-fixed cells (FIG. 24E). However, compared to unfixed nuclei of the same cell type and edit site (Jurkat and CTLA4, FIG. 24A), levels of edit-specific T7 RNA were much lower compared to RPL24 mRNA. To determine whether fixation caused the decreased T7 efficiency, a head-to- head 1ST was performed on fixed and unfixed nuclei from the same nuclei isolation. It was observed that the addition of fixation decreased T7 RNA by over 90% (FIG. 24F), indicating that the lower 1ST PFA fixation is an opportunity to improve 1ST efficiency in PFA-fixed samples.
[0239] Next 1ST reaction conditions were systematically optimized for PFA-fixed cells (FIG. 25A). It was found that 1ST efficiency was greatly improved by increasing 1ST incubation temperature (from 37°C to 42°C) and incubation time (from 2 hours to 24 hours). Under optimized 1ST conditions, it was observed that levels of T7 RNA surpassed RPL24 mRNA with minimal increase in background signal (FIG. 22H, FIG. 25A). Importantly, these levels of T7 RNA were observed in the cell pellet fractions of T7 1ST reactions, suggesting that abundant T7 RNA is retained within fixed cell singlets. To determine that unfixed T7 RNA is resistant to wash out, 1ST was performed on fixed cells followed by washing cell pellets and measuring T7 RNA levels in wash supernatant (ambient) and postwash pellet (in situ) fractions. T7 RNA was detected in the ambient fraction, indicating some washout. However, levels of T7 RNA in the in situ fraction remained comparable to ambient levels (relative to background) and equivalent to RPL24 levels (FIG. 25B). This result confirmed that unfixed T7 RNA are retained by fixed cells and suitable for single-cell combinatorial indexing. Finally, to determine that 1ST is efficient at various genome sites, T7 1ST was performed with seven additional guide RNAs targeting four loci encoding chromatin remodeler genes (ARID 1 A. SMARCA4, CHD3. CHD4'). We observed strong T7 1ST at all loci except ARID 1 A (FIG. 221 and FIGS. 26A-26B), at which both edit sites exhibited low T7 1ST despite high donor knock-in efficiency, suggesting that T7 promoter was not limiting and that T7 1ST efficiency can vary between genome sites. Altogether, these results established that in situ transcription of PFA-fixed cells generates abundant edit-marking T7 RNA.Notably, the strong detection T7 RNA by random hexamer-primed RT-qPCR indicated thatrandom primers effectively capture the variable genome sequence of non-polyadenylated T7 RNA, establishing this as a plausible approach for generating a joint scRNA-seq library of Cas9 cleavage and transcriptomes.Example 5 - Sequencing of T7 RNA precisely identifies Cas9 edit sites
[0240] Next it was determined that Cas9 edit capture can be performed by sequencing T7 RNA. Unfixed nuclei 1ST was performed on two cell types (Jurkat and K562) with donor- labeled Cas9 edits at two genome sites (CTLA4 and B2M, respectively), generating a bulk RNA-seq library from total RNA using random hexamer primers (FIG. 27A), and performed shallow paired-end sequencing. A 3' RNA-seq library (i.e. fragments from RNA 3' ends) was made to match the 3' library chemistry of combinatorial scRNA-seq, including SPLiT-seq. In the absence of donor labeling or T7 RNA polymerase (mock 1ST), no reads were observed at the target edit site. However, cells that received both donor labeling and T7 1ST exhibited anti-parallel reads centered at both expected Cas9 edit sites (FIGS. 27B-27C). These reads mapped within 200 bp of edit sites along both 3' flanks, consistent with bi-directional donor insertion and short T7 RNA extension observed by RT-qPCR. Importantly, reads consisted of 5' donor-encoded barcode sequences that were unmapped (soft-clipped), followed by mapped genome sequence that began at the exact position of expected Cas9 cleavage, 3 bp upstream of the PAM (FIGS. 27B-27C). Some reads were without T7 barcodes and began less than 100 bp from the edit site. These were likely reads from T7 cDNA that were truncated during library fragmentation of RNA 5' ends. Given the shallow sequence mapping of this experiment (FIGS. 27D-27E), these results indicated that 3' RNA-seq library chemistry effectively captures short, unfragmented T7 reads that are tightly localized to Cas9 edits. Together, these results established the feasibility of using T7 1ST for high-throughput Cas9 edit capture by 3' RNA-seq, suggesting compatibility with combinatorial scRNA-seq. This is the first identification of Cas9 genome edits by the sequencing of in situ transcripts.Example 6 - Single cell edit capture sequencing generates high-quality scRNA-seq reads
[0241] Single cell edit capture sequencing joint single-cell profiling was performed of Cas9 genome edits and transcriptomes as follows. First , K562 cell lines were prepared with T7 promoter-labeled Cas9 edits by Cas9 RNP electroporation, using seven guide RNAs targeting the coding regions of four chromatin remodeler genes: ARID1A and SMARCA4 of the BAF complex, and CHD3 and CHD4 of the NuRD complex. Three sample pools weregenerated: (1) unedited cells, (2) five cell lines with single knock-in (one guide) or double knock-in (two guide) a ARID 1 A and / or SMARCA4 (four guides total), (3) three cell lines with single knock-in at CHD4 or double knock-in at CHD3 / CHD4 (three guides total) (FIG. 28A). ). Cell lines were proportioned evenly and selected from the “low indel” group and exhibited 75% donor labeling by TIDE (FIG. 19C and FIG. 29A). As a “stress test” of edit capture in the event of lower labeling efficiency, we included one cell line (CHD3 / CHD4 double knock-in) that exhibited 25% donor knock-in.
[0242] Next we performed PFA fixation and optimized T7 1ST on each sample pool, and verified efficient T7 RNA generation. Strong T7 transcripts were observed at all target genes except ARID 1 A (FIGS. 26A-26B). Despite high knock-in efficiency, the pooled sample containing three ARID 1A knock-in cell lines yielded low T7 RNA levels, suggesting a future opportunity to further optimize the lower transcription or capture efficiency of certain genome sites. It was reasoned that ARID 1A would provide another useful stress test of single cell edit capture sequencing edit capture sensitivity, so we moved forward with these cell lines. Additionally, target gene expression was measured in these cell lines individually, which led to an observation of a 50% decrease in mRNA expression in the presence of 75% donor knock-in (FIG. 20F and FIG. 29A), suggesting successful frameshift knockdown. These results established effective pooled samples for pilot single cell edit capture sequencing measurement of single-cell edits and differential gene expression.
[0243] To perform pilot single cell edit capture sequencing, fresh 1ST reactions were prepared on thawed aliquots of fixed cell pools, followed by immediate EVERCODE combinatorial indexing with a kit from Parse Biosciences, following the standard kit protocol (FIG. 28B and FIG. 30). Two libraries were generated, a 500-cell library and a 10,000 (10k) cell library. High-quality fragments of expected size distributions were observed for both libraries (FIG. 29C). It was observed that T7 RNA was present in amplified cDNA (FIG. 29B), indicating successful direct capture of T7 fragments into the combinatorial scRNA-seq library. To assess library quality, deep sequencing (200k reads / cell) of the 500-cell library was performed and analyzed with a standard “split-pipe” pipeline from Parse Biosciences. High-quality reads and cell barcoding were observed that identified 583 cells, 42k gene transcripts / cell and 7.8k genes / cell (FIG. 29D). Next the lOk-cell library at 100k total reads / cell was sequenced to yield a dataset of 20k gene transcripts / cell. Again, high-quality sequences were observed that identified 9.5k cells, 26k transcripts / cell and 6.7k genes / cell.These results established that this protocol generates high-quality scRNA-seq reads with off- the-shelf materials.Example 7 - Single-cell edit capture sequencing directly captures Cas9 genome edits in single cells
[0244] To analyze our data from the single cell edit capture sequencing methods described herein, a custom software suite was developed that identifies single-cell genome edits and associates allelic dosages of edit events to differential gene expression. This platform processes read alignments to quantify both edit sites and unique molecular identifier (UMI) counts per gene (FIG. 28A). The platform detects high-quality T7 reads by the presence of a T7 barcode sequence in the 5' soft-clipped portion of mapped reads, which is encoded in the inserted donor DNA (FIG. 19D). This soft-clipped barcode feature distinguishes edit-marking T7 reads from background T7 reads and endogenous RNA reads. To reduce false-positive T7 read calls, the platform accounts for other sources of clipped reads and barcode-like sequences (e.g. sequencing errors, template switch oligo sequence). With this process, the platform identified deep T7 read pileups precisely at all seven expected on-target edit sites (FIG. 28B). In terms of read processing, this platform performed the following key analysis steps: an unbiased search for high-quality donor-barcoded T7 reads to identify ‘particular’ Cas9 edit events that occur in single cells, followed by calling ‘canonical’ edit sites that aggregate information of particular edit events across cells to filter false positive edit calls that do not have key features of true Cas9 edit events, such as bidirectionality of T7 reads indicating donor-insertion in either orientation (FIG. 29B; see Example 11, below). The platform was used to determine the allelic dosage of the canonical cas9 edit sites using within-cell particular edit variations, and quantify unique molecular identifiers (UMIs) per gene while removing confounding T7 reads (FIG. 29B; see Example 11, below). Linear mixture modeling was then used to call differentially expressed genes with respect to single cell allelic edit events at both on- and off-target edit sites while accounting for cell state and sequencing depth as confounding random-effect variables (FIG. 31A; see Example 11, below).
[0245] The performance of T7 read calling by the platform was estimated using unedited sample reads as a ground truth negative set, and approximated a positive set (i.e. pure T7 reads) by filtering for reads within ± 100 bp of the seven on-target edit sites. The sensitivity, or classification of T7 reads at on-target sites was 95% (FIG. 28C), although this is likely anunderestimate due to confounding mRNA reads. The specificity, or classification of non-T7 reads, was > 99% with a false discovery rate of < 0.2% (i.e. non-T7 reads classified as T7 reads) (FIGS. 28D-28E). This shows single cell edit capture sequencing is a multi-modal single cell method that captures cas9 edits and gene expression in individual cells and that the platform accurately distinguishes T7 reads from non-T7 reads.
[0246] As noted, consistent with bulk RNA-seq results, distinct pileups of barcoded, bidirectional, un-spliced T7 reads were observed at all 7 guide-targeted edit sites, indicating successful sequencing and mapping of edit-recording T7 transcripts (FIG. 28B). Next, the platform generated high-quality single-cell transcriptome counts, removing confounding nonbarcoded T7 reads ± 1 kilobase from barcoded T7 reads. Across samples, Sheriff identified > 45k mean UMIs / cell and > 7.8k mean genes / cell in the 500 cell transcriptomes, and > 27k mean UMIs / cell and > 6.8 mean genes / cell in the 10k cell transcriptomes (FIG. 28F).Standard single-cell analysis and clustering of the 10k cell transcriptomes with Scanpy and Cytocipher identified 11 clusters of cells with distinct transcriptional profiles, with no distinct clustering observed between unedited and edit sample pools (FIGS. 28G-28H). This suggests that Cas9 edits did not result in overt cell state changes, consistent with previous single-cell CRISPR screens. These results established that our single-cell edit capture sequencing workflow generates high-quality single-cell transcriptome data.
[0247] Next, the platform identifies high-confidence Cas9 edit sites by aggregating T7 reads across cells and locating edit-specific features such as divergent opposite-strand reads from bi-directional T7 promoter insertions. Sheriff analysis of the 10k cell library identified the on-target edit sites of all seven guide RNAs, and 36 off-target edit sites in 6,230 edited cells (FIG. 281). High T7 UMI counts were observed at CHD3 (6.2k UMIs), CHD4 (11.6k UMIs) and SMARCA4 (1.5k UMIs) that were consistent with RT-PCR results. Single cell edit capture sequencing also detected Cas9 edits where RT-qPCR was insensitive, with 548 T7 UMIs detected a ARID 1 A. Furthermore, the platform resolved Cas9 edit sites at base-pair resolution, positioning each on-target edit 0-3 bp from the expected Cas9 cut position (FIG.28J).Example 8 - Single cell edit capture sequencing identifies guide-specific off-target sequence profiles from pooled samples.
[0248] In another example, guide-specific off-target sequence profiles were identified using single cell edit capture sequencing in pooled samples.. Specifically, we discovered 36 off-target edit sites that were generated by the seven guide RNAs in the 10k cell library (FIG. 281). These off-targets occurred across 18 chromosomes and varied widely in frequency, from three cells to 1768 cells (0.03%-18.6% of cells within edited sample pools). All off- target sites were within non-coding sequences, and > 80% (30 / 36) were within introns (FIG. 32A). Since non-coding sequences can effect gene regulation, it is challenging to anticipate the downstream effects of these Cas9 off-targets based on edit location alone.
[0249] Since single cell edit capture sequencing captures edit events instead of delivered guides, we analyzed two features of the detected off-target sites to verify the causal guide of each: (1) sequence similarity to a particular guide, and (2) co-occurrence of off-target edit alleles with those of a particular on-target site. In our dataset, each off-target site harbored an adjacent PAM sequence with high sequence similarity to one guide sequence (FIG. 32B, FIGS. 35A-35F), indicating identification of causal guides with high confidence. We observed an average of six off-target edit sites per guide, with a maximum of 13 edit sites for SMARCA4 guide 22. Only CHD4 guide 14 had no detectable off-target edits (FIG. 32A). Despite unique mapping to a single genome site and high predicted specificity, > 85% (6 / 7) of guide RNAs generated off-target edits.
[0250] Most off-target edit sites were detected at a lower frequency than the corresponding on-target site (FIG. 32B, FIGS. 35A-35F). However, SMARCA4 guide 22 generated five off-target edit sites that were detected at a 1.8-34 fold higher frequency than at SMARCA4 (FIG. 32D). The most frequent off-target occurred in an intron of ADSS1 in 1768 cells, compared to 52 cells for the SMARCA4 on-target site. The d / ASiSV off-target sequence had 16 / 20 base matches to SMARCA4 guide 22, all of which occurred in the non-seed region (FIG. 32D, FIGS. 35A-35F).
[0251] Notably, many off-target sequences, even high-frequency sites, had alignments suggestive of bulge formation or non-canonical base pairing to the imputed guide (FIG. 32B, FIGS. 35A-35F). Many mismatches occurred within the “seed” region that strongly influences Cas9 activity, indicating that Cas9 can tolerate a variety of both seed and non-seed guide mismatches.
[0252] We tested whether the off-target sites detected by single cell edit capture sequencing could be predicted in silico using Cas-OFFinder, E-CRISP, and COSMID. Each tool generated a large number of predicted off-target sites, with poor overlap with the observed off-targets (FIG. 32E). Cas-OFFinder identified the most single-cell edit capture sequencing off-targets (64%, 23 / 36), but at the cost of very low specificity (0.05%, 23 / 45,767, FIG. 32E). These results indicate the difficulty of predicting these Cas9 off-target sites in silico, , and the value of de novo edit site detection by experimental methods like single cell edit capture sequencing.Example 9 - Single cell edit capture sequencing identifies gene perturbations of on- target and off-target Cas9 edits
[0253] Single cell edit capture sequencing captures an unprecedented level of detail for single cell edit events, including base-pair resolution of Cas9 edit locations and variation in the donor sequence inserted at the Cas9 edit site. This level of detail allows the user to infer the allelic dosage of the Cas9 edits occurring in each cell along with genome-wide gene expression (see Example 11; FIG. 30). Current methods for capturing CRISPR guides and gene expression in single cells (e.g., MIMOSCA and Mixscape) do not explicitly account for allelic dosage of Cas9 edit events as captured in single cell edit capture sequencing data. Therefore, a novel differential gene expression approach for single cell edit capture sequencing was developed within the software platform described in Example 7 above (FIG. 31A)
[0254] Briefly, the platform’s differential gene expression approach constructs groups of cells with equivalent cell state and sequencing depth, but with varying levels of Cas9 allelic edits with respect to a given Cas9 edited gene. These cell groups are then used as a random effect variable in a linear mixture model to capture the non-linear effects of cell state and sequencing depth as confounding variables. Allelic dosage of a given query Cas9 edited gene was modeled as a fixed effect variable, which allowed for calling differentially expressed genes with respect to the allelic dosage of a given Cas9 edited gene (FIG. 31A). We first applied this approach to test whether on-target Cas9 edits affected the expression of the four targeted chromatin remodeler genes. On-target edits io ARID 1 A, SMARCA4, and CHD3 resulted in significantly decreased gene expression as edit dosage increased (FDR p < 0.05, two-sided Wald test, FIG. 31B), consistent with edit-induced NMD of these genes observed by RT-qPCR and TIDE (FIG. 20F, FIG. 29A). The change in expression of CHD4 was notsignificant (FDR p = 0.79, two-sided Wald test), consistent with modest CHD4 decrease by RT-qPCR and less efficient NMD (FIG. 20F). However, another plausible factor is underestimation of CHD4 edit allele dosage. These results demonstrate that single-cell edit capture sequencing detects the expected effects of on-target Cas9 editing to target gene expression. .
[0255] Importantly, several Cas9 edit events were observed at off-target genes, many of these genes (such as ADSS1) had negligible gene expression. However there were 4 frequently edited off-target genes that had non-negligible expression, including CDC27, USP9X, CNIH3, and FBXO38-DT. While the expression CNIH3 and FBXO38-DT was not significantly affected by the off-target Cas9 edit events, CDC27 and USP9X were significantly knocked-down (FIGS. 31B-31C). Off-target USP9X editing significantly decreased USP9X expression (FDR p = 10'6, two-sided Wald test), a deubiquitinase with diverse cellular roles (FIGS. 31B-31C). The USP9X off-target was generated by the CHD3 guide RNA (FIG. 31B) and occurs 400 bp downstream of the transcription start site within the first intron. The edit site is 2 bp from a candidate enhancer intersecting a REMAP cis- regulatory module and intersects a binding motif and K562 ChlP-seq peak of the transcription factor SP1, knockdown of which has been shown to reduce USP9X expression. Therefore, USP9X expression is likely reduced because the off-target edit disrupts SP1 activity at a cis-regulatory element for this gene. These results show that in addition to the de novo detection of off-target Cas9 edit sites, single-cell edit capture sequencing enables functional effects associated with off-target edits to be characterized, including those that affect non-coding sequences.
[0256] To see if we downstream functional effects could be detected, differentially expressed genes with respect to both on-target Cas9 edits and off-target edits at USP9X and CDC27 were called. The differential expression calling was repeated on 1937 genes largely comprising high variable genes (see Example 11). At a stringent significance threshold (FDR adjusted p<0.001), hundreds of differentially expressed genes (DEGs) associated with each of our detected allelic edit events were observed; including 800 CHD4 DEGs, 699 CHD3 DEGs, 462 ARID 1 A DEGs, 744 SMARCA4 DEGs, 378 CDC27 DEGs, and 97 USP9X DEGs (FIG. 31D). There were numerous overlap combinations between the DEG sets, in some cases these may reflect overlap of edited cells used to call DEGs, such as 60% of CHD3 edited cells also being CHD4 edited reflected in >150 genes in-common between the CHD3and CHD4 DEG sets (FIGS. 31D-31E). However in other cases these DEG intersections reflect independent differential effects, particularly DEGs in-common between the NuRD genes (CHD3 / CHD4) and the BAF genes (ARID1A / SMARCA4) which represent distinct sets of edited cells (FIGS. 31D-31E).
[0257] As noted, hundreds of differentially expressed genes (DEGs) were associated with Cas9 editing at all four on-target genes and at off-target USP9X (FIG. 31F). To determine if the called DEGs were enriched for known functional effects of the edited genes, gene sets representing genes regulated by BAF (both via promoter and enhancer) and NuRD in K562 cells were compiled using independent ENCODE data (see Example 11). Importantly, both USP9X and CDC27 dysregulation have been associated with cancer onset. In the case of CDC27, this association is mediated through encoding a component in the Anaphase Promoting Complex that controls cell cycle division. In the case of USP9X, it has been previously shown to de-ubiquitinate / ?-catenin, stabilizing / ?-catenin peptides that would otherwise be marked for proteasomal degradation. Therefore the ‘Hallmark G2M’ and the ‘WNT signaling’ gene set from a molecular signatures database (mSigDB) were used as a proxy for CDC27 and USP9X downstream effects, respectively (FIGS. 31G-31H).
[0258] Significant over-representation of the BAF, NuRD, G2M, and WNT gene sets were observed in different combinations of the DEG sets, indicating clear and distinct functional gene expression effects of both on- and off-target Cas9 edits detected by single cell edit capture sequencing (FIGS. 31G-31H). Since BAF regulates cell proliferation, we also tested for enrichment with a set of G2 / M cell cycle checkpoint genes. We observed significant intersections of G2 / M genes and BAF -targeted genes among BAF edit DEGs, and likewise for NuRD-targeted genes among NuRD edit DEGs, indicating consistency between observed and expected downstream transcriptome effects. In particular, both up-regulated DEGs associated with CDC27 Cas9 edits were depleted for ‘Hallmark G2M Checkpoint’ genes, while down-regulated CDC27 DEGs were enriched (FIG. 31G), suggesting decreased cell proliferation with CDC27 knock-down in-line with previous studies. ‘WNT Signaling’ genes were over-represented in the overall DEGs associated with the off-target USP9X Cas9 edits, though more strongly in the upregulated DEGs (FIG. 31H). Consistent with the BAF complex as promoting gene expression, both enhancer- and promoter- regulated genes were over-represented in the significant down-regulated DEGs associated with ARID 1 A and SMARCA4 Cas9 edits (FIG. 31G). Likewise G2 / M checkpoint genes were also significantlydownregulated in response to BAF editing (FIG. 31G). In contrast to BAF, the NuRD complex performs transcriptional repression. Consistent with NuRD as a transcriptional repressor, CHD3 and CHD4 down-regulated DEGs were depleted for NuRD target genes (FIG. 31H). Overall, functional single cell gene expression effects consistent with previous literature for both on- and off-target Cas9 edited genes were detected and identified using single cell edit capture sequencing. These results indicate that single cell edit capture sequencing captures the downstream transcriptomic effects of individual Cas9 edits.
[0259] Lastly we examined the downstream effects of the off-target edit at USP9X. Since USP9X regulates protein abundance in myriad cellular processes, the downstream gene expression effects are difficult to anticipate. However, one of the DEGs specific to USP9X editing was NRF1, a transcription factor that regulates proteasome machinery including USP9X. We observed that the set of LSPPX-specific DEGs (i.e. DEGs absent in CHD3 edit DEGs) contained a significant number of NRF1 -targeted genes determined by ChlP-seq (FDR p = 0.014, OR = 2.4, two-sided Fisher’s exact test, FIG. 31H), with 63% (24 / 38) of DEGs regulated by NRF. In contrast, NRF1 target genes were under-represented among CHD3 and CHD4 edit DEGs, suggesting that the NRF 1 effect was specific to USP9X editing (FIG. 31H). In summary, our enrichment analysis revealed that single-cell edit capture sequencing captures downstream gene expression effects of off-target Cas9 edit events.Example 10 - Single cell edit capture sequencing data processing
[0260] FIG. 9 diagrams a data processing method for processing single cell edit capture sequencing data to detect single cell CRISPR cas9 edit sites, allelic edit dosage, and quantify gene expression. Briefly, as shown in panel (1), Fastq files generated from single cell edit capture sequencing were processed using split-pipe, software from Parse biosciences, to generate an annotated bam file. As shown in panel (2), three kinds of reads were present within the annotated bam file; a) t7 polymerase generated reads with a single cell edit capture sequencing barcode, evident in the 5' soft-clipped region of the mapped read, b) reads generated by t7 polymerase that have been fragmented, and hence without single cell edit capture sequencing barcode, and c) reads generated from normal RNA pol II transcription within expressed genic regions as shown in panel. Edit sites were expected to have a high number of reads mapping in both directions across cells (due to donor oligonucleotide inserting in either direction after the CRISPR-cas9 induced double-stranded break) with 5' soft-clipped reads containing the barcode used in the donor olignonucleotide. As shown inpanel (3), particular-edit-sites are edits occurring at a single cell level, identified by a k-mer match between the donor oligonucleotide barcode and the reads 5' soft-clipped sequence. If the read maps in the forward direction, the k-mer match can be to the forward version of the barcode. If the read maps in the reverse direction, the k-mer match can be to the reversecomplement of the single cell edit capture sequencing barcode. Reads called t7 barcoded reads are used to record edited cells, edit location, edit direction and inserted donor sequence. As shown in panel (4), particular-edits represent slight edit variation around a particular genomic location. Canonical-edit-sites were determined by assigning particular-edits to more commonly occurring versions of the edit, with canonical-edits that do not appear to be well supported across multiple cells or by reads in both directions (a characteristic of the donor introducing a t7 pol promoter in either direction, depending on the orientation of the cell particular-edit). Canonical-edit-sites were determined by first constructing a graph of particular-edit locations for each chromosome, and then rank-ordering particular-edits by the frequency that they occur across cells. Progressively, particular-edits from most-to-least frequent were considered as canonical-edit-sites, with all particular-edits within predetermined base-pair radius from the particular-edit assigned to the canonical-edit-site. All particular-edits collapsed to a canonical-edit were then removed from being considered a potential canonical -edit, until all particular-edits were assigned to a canonical-edit-site. As shown in panel (5), for each canonical-edit-site, the number of edited alleles for that site is then called at a single cell level. The number of edited alleles was determined as the number of particular-edits of a canonical-edit for the respective cell, with corrections for sequencing error and variation in t7 polymerase priming. As shown in panel (6), all remaining reads within + / - 500bp from canonical-edit-sites were removed prior to cell / gene UMI counting, since otherwise non-barcoded t7 reads were expected to confound gene expression estimation. As shown in panel (7), the key outputs from the single cell edit capture sequencing read processing is a cell X canonical-edit-site matrix with 0, 1, or 2 indicating the number of alleles edited for the particular cell at the particular canonical-edit-site, and a cell X gene matrix indicating the number of unique molecular identifiers (UMIs) counted for a particular gene as a measure of single cell gene expression.Example 11 - Materials and Methods
[0261] Oligonucleotides: All oligonucleotides were synthesized by Integrated DNATechnologies (IDT).
[0262] Human subjects: Primary human T cells were isolated from de-identified healthy donor blood from the San Diego Blood Bank (San Diego, CA). An institutional review board (IRB) at the Salk Institute for Biological Studies determined that this work did not meet the 45 CFR 46 definition of human subjects research and was therefore exempt from IRB review and informed consent requirements.
[0263] Cell culture: All cells were incubated at 37°C, 5% CO2 and maintained at cell densities of 0.1-0.5 million (M) cells / mL (K562 cells) or 0.2-1.5 M cells / mL (all other cell types). Human cell lines K562, Jurkat (clone E6-1), and GM12878 were cultured in RPMI- 1640 (Gibco #11875-093) supplemented with 10% heat-inactivated fetal bovine serum (FBS, GeminiBio #100-500), 100 U / mL penicillin, and 100 pg / mL streptomycin (Gibco #15140- 122).
[0264] For primary human T-cell culture, peripheral blood mononuclear cells were isolated from healthy donor buffy coats by centrifugation on Ficoll-Paque Plus (Cytiva #17144002) or Lymphoprep (StemCell Technologiesb #07851) density gradient media, and mononuclear fractions were cryopreserved in heat-inactivated FBS with 10% DMSO (Sigma- Aldrich #D8418). Frozen mononuclear cells were thawed and cultured in IMDM with 25 mM HEPES (Gibco #12440-053), supplemented with 10% heat-inactivated FBS, 50 U / mL recombinant human IL-2 (Miltenyi Biotec #130-097-746), 100 U / mL penicillin, and 100 pg / mL streptomycin.
[0265] To activate T cells, 6-well plates were coated with anti-human CD3 antibody (clone OKT3, BioLegend # 317326) by adding 1 mL per well of 1 pg / mL anti-CD3 in Dulbecco’s phosphate-buffered saline without calcium and magnesium (DPBS, Sigma #D8537-500ML), incubating at room temperature for 1 hour, and aspirating. Frozen mononuclear cells were thawed, rested for least 1 hour at 37°C, and seeded in coated plates at 1 M cells / mL in the above T-cell medium supplemented with 1 pg / mL anti-human CD28 antibody (clone CD28.2, BioLegend #302934), then expanded in T-cell medium without anti- CD28.
[0266] Donor DNA design: The 34 base pair (bp) single cell edit capture sequencing donor DNA (donor 02) was generated as follows. An optimized phage T7 promoter sequence was obtained from the Cel-seq+ method. This sequence is composed of a 20-bp core T7 promoter sequence (5'-TAATACGACTCACTATAGGG-3' (SEQ ID NO: 103) and flanking5' and 3' sequences that maximize promoter activity. To this sequence was added 2-bp terminal sequences, 5' phosphate and two 3' phosphorothioate modified bases from the GUIDE-seq donor (FIG. 19B, FIG. 20B). An additional 26 donor constructs were designed by using alternative T7 and SP6 promoter variants and flanking sequences, removing flanking sequence, or incorporating dual promoters and alternative polyA or tracrRNA capture sequences. To maximize gene-disrupting frame shifts upon donor insertion in coding sequence, all donors were made to length 3n+l to yield +1 shifts upon precise insertion and +2 shifts upon common 1-bp templated insertion by Cas9. All donor sequences were ordered with the above GUIDE-seq base modifications as pre-duplexed lyophilized oligonucleotides from IDT, reconstituted to 100 pM in electroporation buffer (MaxCyte #EPB-1), and stored at -20°C.
[0267] Cas9 edit labeling with donor DNA encoding T7 promoter: Unbiased labeling of Cas9 edits with T7 promoter was performed by homology-free knock-in. Live cells were treated with donor 02 DNA and CRISPR-Cas9 ribonucleoprotein (RNP) by electroporation. Single-stranded guide RNAs (sgRNAs) were assembled with 20-base guide (protospacer) sequences, a modified “tracrV2” scaffold sequence, and the common 2' O-methyl 3' phosphorothioate modifications on terminal bases. Guides with specificity scores >50 and efficiency scores >0.3 were selected using GuideScan or GuideScan2. A standard TRAC guide sequence was used. All sgRNAs were synthesized by IDT, reconstituted to 125 pM (4 pg / pL) in electroporation buffer (MaxCyte #EPB-1), and stored at -20°C. RNP was prepared by mixing 1 pL of 10 pg / pL purified S. pyogenes Cas9 protein (IDT #1081059) and 1 pL of 4 pg / pL of sgRNA (2:1 sgRNA:Cas9 molar ratio), incubating at room temperature for 15 minutes, and storing on ice until electroporation. Cells were harvested, washed once with 5- 10 mL electroporation buffer, and resuspended in electroporation buffer to 125 M cells / mL. Per 25 pL electroporation, 5 pL of transfection mix was prepared with 2 pL RNP, 1 pL of 100 pM donor DNA, 1 pL of 100 pM Cas9 enhancer (IDT #1075916), and 1 pL electroporation buffer. Next, 20 pL of cells (2 M) was mixed with 5 pL of transfection mix, transferred to an OC-25x3 process assembly (MaxCyte #SOC-25x3), and immediately transfected with an ExPERT ATx electroporation system (MaxCyte) using cell type-specific instrument protocols provided by MaxCyte. Final electroporation concentrations were 100 M cells / mL, 2.5 pM Cas9, 5 pM sgRNA, 4 pM donor DNA, and 4 pM Cas9 enhancer. After electroporation, cells were immediately seeded at 0.5 M cells / mL (K562) or 1 M cells / mL (allother cell types) in a 6-well plate with warm cell culture medium, expanded, and cryopreserved.
[0268] TIDE quantification of donor-labeled Cas9 edits: On day three postelectroporation, genomic DNA (gDNA) was extracted from knock-in and control cells using a Quick-DNA Microprep kit (Zymo Research #D3020) and eluted in nuclease-free water. PCR amplicons were generated with primer pairs designed by Primer-BLAST with the following custom parameters: size 900-1200, melting temperature 53-55-57 (min-opt-max), database genomes for selected eukaryotes, organism homo sapiens, max size 30, 40-60% GC, GC clamp of 1, max poly-X of 4, max GC in 3' end of 4, human repeat filter, and low complexity filter on. Primer pairs were designed to amplify ~1 kilobase genome regions from -150 bp upstream of knock-in sites to -850 bp downstream. Per 20-pL PCR reaction, -50 ng gDNA was combined with 0.5 pM of each primer, nuclease-free water, and either Phusion Plus (Thermo Scientific #F631S) or Phusion High-Fidelity (Thermo Scientific #F531S) master mix. PCR reactions were run on a SimpliAmp thermal cycler (Applied Biosystems) with the following protocol: 98°C for 2 minutes, 35 cycles of [98°C for 10 seconds, 60°C for 10 seconds, 72°C for 30 seconds], 72°C for 5 minutes, and then 10°C hold (ramp rate 2°C / second). PCR reactions were purified by ExoSAP -IT Express reagent (Applied Biosystems #75001.200.UL) or by NucleoSpin PCR cleanup kit (Takara #740609.25). Sanger sequencing of PCR amplicons was performed by Eton Bioscience (San Diego, CA) or Genewiz (San Diego, CA), and sequence traces (AB1 files) were analyzed by TIDE (tracking indels by decomposition) using a web-based tool with the following parameters: left boundary 50, default decomposition window, indel size range 50 (-50 to +50), and default p- value threshold (0.001).
[0269] Flow cytometry for target protein knockdown: On day 7 post-electroporation, unedited and donor-labeled GM12878 cells were harvested into a 96-well round-bottom plate, washed twice with 200 pL cell staining buffer (BioLegend #420201), spun down at 200 x g for 2 minutes, flick-decanted, gently vortexed, and stained with 50 pL of a 1 :200 dilution of PE anti-human B2M antibody (clone 2M2, BioLegend #316305) in cell staining buffer. An isotype control was prepared by staining unedited cells with an equivalent concentration of diluted PE mouse IgGl K antibody (clone MOPC-21, BioLegend #400114). Cells were stained for in the dark for 15 minutes at room temperature, washed twice with 150-200 pL cell stain buffer, and resuspended in 150 pL cell stain buffer. Forward / side scatter and PEfluorescence intensity were acquired using a BD FACSymphony A3 cell analyzer and HTS attachment. Live singlet events were gated and plotted using R packages flowCore and ggcyto.
[0270] Nuclei isolation: Nuclei buffers were prepared in an RNase-free environment by the Omni-ATAC protocol. A basal buffer of 10 mM Trizma HC1 (Sigma #T2194-100ML), 10 mM NaCl (Sigma #S5150-lL), 3 mM MgCh (Sigma #M 1028-100ML), 0.1% Tween-20 (Sigma #11332465001), and 48.75 mL nuclease-free water (Corning #46000CI) was prepared and stored at 4°C. For each nuclei isolation, fresh lysis and wash buffers were prepared. Lysis buffer was prepared with basal buffer supplemented with 0.1% IGEPAL CA-630 (Sigma #I8896-50ML), 0.01% digitonin (Invitrogen #BN2006) (stock solution dissolved at 65°C), and 1 U / pL RiboLock RNase inhibitor (Thermo Scientific #EO0382). Wash buffer was prepared with basal buffer supplemented with 1 U / pL RiboLock. Fresh buffers were kept on ice until use. To isolate human nuclei, cells were harvested, washed once with DPBS, filtered through 40 pm cell strainer (Fisherbrand #22363547), and counted. All steps and spins were performed on ice or 4°C with ice-cold buffers and low-bind tubes. To lyse cells, 2 M cells were lysed by gently resuspending in 45 pL lysis buffer (pipetting up-down exactly 3 times) and incubating on ice for 3 minutes. Nuclei were immediately washed twice by gently resuspending in 250 pL wash buffer, spinning down at 500 x g for 5 minutes, manually aspirated, resuspended in 100 pL wash buffer, filtered through 40 pm Flowmi cell strainer (SP Bel-Art #H13680-0040), spun down, aspirated, resuspended in 35 pL wash buffer. Nucei were counted and imaged by TC 20 automated cell counter (Bio-Rad). To generate larger quantities of nuclei for formaldehyde fixation, cell and reagent amounts were scaled proportionally. Aliquots of 0.2-0.5 M nuclei were stored at -80°C.
[0271] Cell and nuclei fixation: Cells or freshly isolated nuclei were fixed in an RNase- free environment using EVERCODE Fixation kits (Parse Biosciences #ECF2001, #ECF2003, #ECF2101, #ECF2103) according to kit protocol. Per fixation, 3 M input cells or nuclei were used. Fixation yield was -50% (1.5 M fixed singlets). Aliquots of 0.5 M fixed cells or nuclei were stored at -80°C.
[0272] T7 in vitro and in situ transcription: In vitro transcription (IVT) and in situ transcription (1ST) was performed in an RNase-free environment using the Hi Scribe T7 Quick RNA synthesis kit (New England BioLabs #E2050S). For IVT on purified gDNA, gDNA was extracted from knock-in and control cells using a Quick-DNA Microprep Plus kit(Zymo Research #D4074) and eluted in nuclease-free water. For IVT, 20-pL IVT reactions were prepared with 1 pg gDNA, 10 pL of 2X HiScribe NTP (10 pM final), 2 pL of 10X HiScribe T7 RNA polymerase (or water for controls), and nuclease-free water, and incubated at 37°C for 2 hour. For one experiment testing phage RNA polymerase activity in different buffers , comparable 20-pL IVT reactions were also prepared with the HiScribe SP6 RNA Synthesis kit (New England BioLabs #E2070S), and IVT incubation time was shorted to 1 hour to minimize NTP depletion. For 1ST on unfixed nuclei, 40 pL reactions were prepared by gently mixing 16 pL (~100k) of freshly isolated nuclei with a master mix of 20 pL HiScribe NTP and 4 pL HiScribe T7 polymerase (or water for “No pol” samples), and incubated at 37°C for 2 hours. For single cell edit capture sequencing 1ST on fixed cells (or nuclei), the optimal protocol was to gently add 15 pL (-100 thousand, k) cells in fixation kit storage buffer to a 25-pL master mix of 20 pL HiScribe NTP (10 pM final), 4 pL HiScribe T7 polymerase, and 1 pL of 40 U / pL RiboLock (1 U / pL final), and incubate at 40°C for 16- 24 hours. This method was used on scRNA-seq samples. During 1ST optimization for fixed samples, additional tested methods included exchange of storage buffer with Omni-ATAC buffer (caused cell clumping), increasing NTP concentration to 20 pM (decreased 1ST efficiency), incubating at 37°C or 42°C (less efficient and specific), and incubating for 0.5-4 hours (less efficient) (FIG. 25A).
[0273] Total RNA purification: Total RNA was extracted from IVT and 1ST reactions by TRIzol and column purification. Forty (40) pL IVT and 1ST reactions were treated with 120 pL TRIzol LS reagent (Ambion #10296010). To extract total RNA from 1ST pellet and supernatant fractions separately, 1ST reactions spun down in a mini centrifuge with PCR tube adapter (-2,000 x g for 1 minute) and split into pellet and supernatant fractions. Pellet fractions were treated with 160 pL TRIzol reagent (Ambion #15596026), and 40 pL supernatant fractions were treated with 120 pL TRIzol LS. Total RNA was purified from TRizol lysates using a Direct-zol RNA microprep kit (Zymo Research #R2060), with on- column DNase, and eluted in nuclease-free water. For sequencing, additional 20-pL off- column DNase treatment was performed by mixing 10 pL eluted RNA, 1 pL of 2 U / pL Turbo DNase (Invitrogen #AM2238), 2 pL of 10X Turbo DNase buffer, 7 pL nuclease-free water, incubating at 37°C for 30 minutes, and purifying by RNA Clean and Concentrator kit (Zymo Research #R1013), eluting in nuclease-free water. Purified RNA was quantified by UV spectroscopy (A260 / A280) using a DeNovix DS-11+ spectrophotometer, or by calibrated fluorometry using a Qubit 3 fluorometer (Invitrogen) and RNA HS kit (Invitrogen #Q32852)or RNA BR kit (Invitrogen # QI 0210). For qPCR, purified total RNA was stored -20°C for up to one week. For sequencing, purified total RNA was stored at -80°C.
[0274] Reverse transcription quantitative PCR (RT-qPCR): F or RT -qPCR of T7 RNA, custom primer pairs were designed using Primer-BLAST with the following custom parameters: size 80-120, melting temperature 58-60-62 (min-opt-max), max Tmdifference of 2, database genomes for selected eukaryotes, organism homo sapiens, max size 30, 40-60% GC, GC clamp of 1, max poly-X of 4, max GC in 3' end of 4, human repeat filter, and low complexity filter on. To quantify on-target T7 RNA, primer pairs were designed to target flanking regions of knock-in sites. Control primer pairs were designed targeting several nontargeted intergenic sequences on different chromosomes, including two “genomic safe harbor” gene deserts (GSH1, GHS2). Additional primer pairs were designed flanking three sites with high homology to the core 18-bp T7 promoter discovered using BLAST. For RT- qPCR of cellular mRNA, intron-spanning primer pairs were selected from the literature or pre-designed IDT PrimeTime assays. Two-step RT-qPCR was performed as follows. First, 10-pL first-strand cDNA synthesis reactions were prepared with up to 250 ng total RNA, 4 pL 5X Maxima H Minus master mix or No RT control mix (Thermo Scientific #M1661), and nuclease-free water, and then incubated on a SimpliAmp thermal cycler using the kit- recommended protocol: 25°C for 10 minutes, 50°C for 15 minutes, and 85°C for 5 minutes, and then 10°C hold. Notably, Maxima cDNA synthesis is primed by both random hexamer and oligo-dT, enabling reverse transcription of both non-polyadenylated T7 RNA and endogenous mRNA. For each 20 pL qPCR reaction, 2-25 ng of RNA-equivalent cDNA reaction was mixed with 0.5 pM of each primer and PowerUp SYBR Green master mix (Applied Biosystems #A25742) by combining 10 pL of diluted cDNA reaction with 10 pL of a 2X mix of 1 pM primers and 2X PowerUp mix. Quantitative PCR was run on a QuantStudio 6 Flex Real-Time PCR system (Applied Biosystems, software version 1.3) with the following protocol: 50°C for 2 minutes, 95°C for 2 minutes, 40 cycles of [95°C for 15 seconds, 60°C for 1 minute], 95°C for 15 seconds, 60°C for 1 minute, and then 0.1°C / second ramp to 95°C for 15 seconds. Analysis of software-calculated cycle threshold (Ct) values was performed by the comparative Ct method. Reference assays targeting GSH1 and GSH2 (background signal) were used for IVT and unfixed 1ST samples. For fixed 1ST samples, GSH1 and GSH2 signals were below limit of detection, so RPL24 and RPS10 assays were used as reference. Results are represented as a difference of Ct (dCt, normalized to reference only) or a difference of differences (ddCt, normalized to reference and no knock-in control).
[0275] Bulk RNA sequencing analysis: Four IVT reactions were performed on purified gDNA samples from K562 or Jurkat cells edited at B2M and CTLA4, respectively, with either RNP alone or with RNP and single cell edit capture sequencing donor. An additional two “mock” IVT reactions without T7 RNA polymerase were performed on the two RNP-only samples, and six total RNA samples were purified and quantified. A pooled, ribosomal RNA- depleted RNA-seq library was made using the QIAseq UPXome RNA library kit (Qiagen #334782, #334782) and SimpliAmp thermal cyclers. The pooled library contained 12 barcoded samples, two technical library replicates per total RNA sample. All library preparation steps were performed according to kit instructions, with three modifications. First, cDNa synthesis was performed with 50:50 random hexamer and oligo-dT primers.Second, all bead cleanups were performed with 0.8X QIAseq beads (Qiagen #333923). Third, library amplification was performed with the following thermal cycler protocol: 98°C for 3 minutes, 18 cycles of [98°C for 30 seconds, 55°C for 10 seconds, 72°C for 40 seconds], 72°C for 5 minutes, and then 4°C hold (ramp rate 2°C / second). Library cDNA was analyzed on a TapeStation 4150 (Agilent) and quantified on a Qubit 3 fluorometer using a Qubit dsDNA HS kit (Invitrogen # Q32851). Paired-end sequencing was performed on a MiniSeq sequencer with Mid Output 300-cycle kit (Illumina) at the Salk Institute Next Generation Sequencing (NGS) Core, with this run configuration: 150 / 10 / 10 / 150 cycles (Read 1 / Index 1 / Index 2 / Read 2), 0% PhiX. Sample demultiplexing and adapter / poly sequence trimming was performed using cutadapt. Quality of raw and pre-processed reads was checked by FastQC and MultiQC. Reads were aligned to human reference genome hg38 using BWA-MEM. Aligned reads were post-processed by samtools. For track visualization, bigwig files were generated using deeptools, and track images were generated using the UCSC Track Collection Builder.
[0276] Single cell edit capture sequencing joint Cas9 edit and transcriptome profiling of chromatin remodeler genes: Three K562 cell samples were prepared and frozen: (1) K562 cells treated by electroporation only (unedited), (2) a pool of ARID 1 A MARCA4 knock-in cell lines (each l / 5thof pool), and (3) a pool of three CHD3 / CHD4 knock-in cell lines (each l / 3rdof pool). Aliquots of each cell sample were thawed, and 1ST was performed on fixed cells as described above, with -100 k fixed cells per 40 pL reaction and duplicate reactions for each sample pool. After 1ST incubation for 18 hours, duplicate reactions were pooled, spun down at 200 x g for 10 minutes at 4°C, gently aspirated, resuspended in 40 pL fixation kit storage buffer, and passed through a 40 pm Flowmi cell strainer (SP Bel-Art #1413680- 0040) to ensure a single-cell suspension. Singlets were counted and immediately processedby combinatorial barcoding using an Evercode™ WT Mini v2 kit (Parse Biosciences #ECW02110, #UDI1001), according to kit protocol. A total of 110k input cells were used. Sample proportions were 10% unedited cells, 45% ARJDlA / SMARCA4-edrted cells, and 45% CHD3 / CHD4-edited cells. Barcoded singlets were counted (~20k total) and used as input for a 500-cell sequencing library and a lOk-cell library. Each library preparation yielded ~1 pg amplified cDNA library and -400 ng of sequencing library of the expected quality and size distribution. Paired-end sequencing was performed at the Salk NGS Core. The 500-cell library was run on a NextSeq 2000 sequencer with Pl 300-cycle kit (Illumina), with this run configuration: 198 / 8 / 8 / 86 cycles (Read 1 / Index 1 / Index 2 / Read 2), 5% PhiX. The lOk-cell library was run on a NovaSeq 6000 with SP 200-cycle kit (Illumina), with this run configuration: 98 / 8 / 8 / 124 cycles, 5% PhiX.
[0277] Single cell edit capture sequencing read processing: A graphical overview of the single cell edit capture sequencing read processing pipeline is detailed in FIG. 9. Briefly, the single cell RNA-seq fastq files generated from split-pool barcoded and random hexamer (Parse Biosciences) primed cDNA were processed using split-pipe to generate an annotated bam file (see panel 1 in FIG. 9). T7 barcoded reads were identified by a k-mer match (k=6) to the donor barcode in the 5' soft-clip sequence of mapped reads (See panel 3 in FIG. 9). These reads were considered ‘particular’ edit events, due to occurrence in a particular cell (identified by read cell-barcode ‘CB’ tag), the genomic mapping location of the base pair immediately upstream of the soft-clipped sequence (indicating the base-pair resolution cutsite of the CRISPR-cas9 edit event), and the soft-clip sequence (indicating the donor sequence inserted at the particular edit event). Due to a k-mer match between the donor barcode with the template switching oligo (TSO), a set of observed variations in the TSO sequence from the soft-clipped sequences of the reads was constructed, considered ‘blacklist’ sequences. The donor barcode was matched to each of the blacklisted TSO sequences to generate a list of problematic donor barcode k-mers. If a read had a soft-clipped sequence k- mer match to the donor barcode within the problematic k-mer list, then k-mer matching was performed between the soft-clip sequence and each of the blacklisted sequences. If the read had no other k-mer matches to the donor barcode other than the k-mers in common with the blacklist sequences and had at least one additional k-mer match to any blacklist sequence, then it was not classified as a barcoded t7 read. After processing the mapped reads, a list of particular edit events were identified, called ‘canonical’ edit sites.
[0278] Identification of canonical edit sites: After calling putative T7 barcoded reads to identify particular edit events, these particular edit events were then collapsed to a canonical edit site, representing the most common version of the particular edit sites across cells. By associating particular edit sites with a canonical edit site, this allowed us to further filter edit events for false-positive T7 barcoded read calls, based on metrics for each canonical edit site obtained from aggregating information across cells.
[0279] To determine canonical edit events, first a graph per chromosome using Faiss was constructed, where each node is a particular edit site, and edges between nodes represent the base-pair distance between the particular edit sites (see panel 4 in FIG. 9). Particular edit sites are then processed in order of the number of cells that have the particular edit, denoted as an ordered set / J(see panel 4 in FIG. 9). The first particular edit in , Po, is considered the first canonical edit site. All particular edits within a defined base-pair distance (140bp) from the current canonical edit site (Po for the first iteration) are considered particular edits of the canonical edit site, denoted Pc,o. The ordered set of particular edit sites is then updated, removing the current canonical edit site and the associated particular edit sites, such that P = P - (Pc,o U { Po }). This process is repeated at each iteration, until all particular edit sites are associated with a canonical edit site, resulting in the stopping condition P = 0 (see panel 4 in FIG. 9).
[0280] True positive canonical edit sites were then determined by the following criteria; 1) the canonical edit site is supported by occurrence in at least 3 cells, 2) there are particular edit events that occur in both the forward and reverse direction, indicating donor insertion in either orientation, 3) there is no more than 15 bp between the two closest pairs of forward and reverse particular edit sites (indicating low variation of particular edits around the canonical edit site position), and 4) the canonical edit site does not overlap black-listed genome regions. For criteria 4, blacklisted genome regions consisted of 1) endogenous T7 polymerase promoter sites in the human genome that generate excessive T7 reads but were not clearly edited, 2) endogenous genome locations with sequence similarity to the donor barcode in both the forward and reverse direction in close proximity, which when combined with sequencing error creates soft-clip sequences that resemble donor barcodes, and 3) simple repeat regions of poly-N sites of length greater than 50bp, which were identified to accumulate reads with soft-clipped sequences which resulted in false-positive canonical edit site calls.
[0281] Automated filtering of canonical edit sites as described above identified 63 candidate canonical edit sites. The alignment of called t7 barcoded reads was then manually examined at these edit sites with soft-clip sequence visualization in the Integrative Genome Viewer (IGV). 20 of the 63 sites did not appear to be true edit events, but had k-mer matches to the single cell edit capture sequencing barcode in the 5' soft-clip sequence due to sequencing error resulting in mis-alignment of reads where the genome had sequence similarity to the single cell edit capture sequencing barcode on each strand, resulting in an incorrect edit call. Aside from examining the alignment, these false positives also differed by other metrics when compared to the final 43 called true edit events, including best-matching alignment rate to the set of guide sequences (determined below, FIGS. 35A-35F, FIG. 35H), the distance between the forward and reverse particular edit events (FIG. 351), the frequency of the canonical edit across cells (FIGS. 35H-35I), and the specificity of the edit event for cells that were exposed to the same treatment (FIG. 34C). Therefore, only the manually curated edit sites were kept for further downstream analyses.
[0282] Single cell allelic edit dosage calling: At each called canonical edit site, we then stratify the particular edits associated with the canonical edit by cell barcode (previously corrected for sequencing error by split-pipe). We then perform a pairwise comparison of each of the particular edits, to determine if these particular edits were generated from different alleles. This process considers the orientation of the edit, the location, and donor sequence variation (see panel 4 in FIG. 9). Two edits are considered different if they 1) have a different orientation, and 2) if they have a different reference genome mapping position. If two particular edits have the same orientation and the same reference genome position, we then compare the donor sequence. The T7 polymerase has known variation in priming, which can result in variation in the 5' end of t7 reads. Furthermore, t7 can introduce sequence content at the 5' end of the read, and we observed several sequencing artifacts that could also be present in the 5' end of the read (e.g. TSO). We also observed variation due to erroneous homopolymer calls, for cases where 3 or more base pairs were identical in the donor sequence. Finally, there is also the general sequencing error of mis-called bases.
[0283] To account for these sources of sequencing error, we developed an approach that first determined the longest donor sequences to account for differential t7 priming. We process the particular edits pairwise from smallest to longest. If the edits have different mapping positions or different orientation, and have not been previously classified as sub-edits of another particular edit, they are added as candidate longest donor sequences. If they vary only in donor sequence, we then determine the edit distance between the sequences while accounting for variable sequence length. This is achieved by performing local alignment between the sequences using pairwise2.align.localms with biopython, with matching score +1, mismatch -1, -0.5 for gap open penalty, and -0.5 for gap extension penalty. After the local alignment, we then count the number of mis-matched base pairs until all bases of the shorter donor sequence have been seen. This results in an edit distance of 0 if one donor sequence is a subsequence of the other, and hence accounting for differential t7 priming. If the edit distance is 2 or more, we re-calculate the edit distance between the sequences after replacing instances of the same base repeated 3 or more times with a single instance of the base (to account for homopolymer sequencing error). If the edit distance between the two particular edits is still more than 2 mismatches, we then re-calculate the edit distance by counting the number of mis-matches in the first 10 base pairs of alignment between the sequences from the 3' end of the alignment (to account for sequencing artifacts on the 5' end of the read).
[0284] If the edit distance between the two particular edits (after accounting for homopolymer error and 5' read error) is 2 or less, the longer of the two particular edits is added to the candidate longest edit list (if not previously classified as a sub-edit), and the smaller particular edit is removed from the candidate longest edit list (if present) and added as a classified sub-edit (i.e. read with a truncated donor sequence due to variable t7 priming). If however the two particular edits have more than 3 mismatches after accounting for all aforementioned sequencing errors, then they are both considered candidate longest edits (if not already classified as sub-edits). After comparing all of the particular edit donor sequences from smallest to longest in length, a set of particular edits have been determined which have different orientations, mapping position, or donor sequence variation that cannot be accounted for by multiple sources of sequencing error. Hence the cell allelic edit count was then determined as the number of these particular edits that are likely generated from uniquely edited copies of the genome.
[0285] To call allelic edit dosage per gene, gene start and end positions were intersected with canonical edit sites (defined as canonical edit position + / -140bp). Single cell gene edit dosages were then determined as the sum of the allelic edits for each cell across all intersecting edit sites of a given gene. The number of called allelic edits was capped to thecopy-number for the given gene in K562 cells, determined from a previous study that utilized deep whole genome sequencing of K562 cells.
[0286] Gene expression quantification: To quantify gene expression while correcting for reads generated from the T7 polymerase at CRISPR-cas9 edit sites, we first determined sets of unique molecular identifiers (UMIs) stratified by unique cell barcode and gene using ‘CB’, ‘GN’, and ‘GX’ read tags from the annotated bam file generated by split-pipe (see panel 6 in FIG. 9). We then filtered reads that were + / - 500bp from a canonical edit site across all cells, to remove reads that could have been generated from the T7 polymerase (see panel 6 in FIG. 9). The remaining reads are then used to count the number of unique UMIs per gene per cell, accounting for sequencing error by counting the number of unique UMIs that have more than Ibp hamming distance from one another (see panel 6 in FIG. 9). To account for cases where one or more UMIs have a greater than 1 edit distance from one another but are within 1 base pair of a common UMI, we construct UMI sets where within each set there exists a single edit distance path between the UMIs. The number of these sets is then counted as the number of unique molecules for a given gene and cell, effectively accounting for sequencing error and removing confounding reads from the T7 polymerase (see panels 6 and 7 in FIG. 9). Normalization is then performed per cell, dividing the UMIs per gene by the total UMIs within the cell, multiplying by the median library size observed across cells, and then logscaling with euler’s number as the base after adding a +1 pseudocount (as implemented in scanpy).
[0287] Canonical edit site causal guide inference: To infer the guide associated with an identified canonical edit site, we used a combination of sequence homology to each of the 7 guides used in our experiment, and the co-occurrence of each edit-site with each of the on- target edit sites associated with the guides. To determine the sequence homology, we first identified all PAM sites (identified as sites with ‘NGG’ or ‘NAG’ sequence content) on both the forward and reverse strand within + / - lOObp from each canonical edit site. We then extracted the 20bp of reference genome sequence upstream of each PAM site as candidate target sequences. Each guide sequence was then aligned to each of the candidate target sequences using biopython pairwise2.align.localms, with scoring parameters: +1 for match, - 1 for mismatch, -0.8 for gap, and -0.5 for gap extension. The homology between the guide sequence and candidate target sequences was determined as the number of base-pairs which were aligned. We also scored guides by taking the cosine similarity between the t7 UMIcounts across cells for the guide on-target edit site and each canonical edit site. These cosine similarities per canonical edit site were then scaled by the sum of the cosine similarities across the score on-target guide edit sites. If the homology and cell co-occurrence metrics prioritized the same causal guide for a given canonical edit site, then this candidate guide was called as the causal guide for the canonical edit site. If the canonical edit site had no cooccurrence with an on-target edit site, then the top-matching guide by homology was called as the causal guide. If the homology-based guide inference and coincidence based guide inference disagreed on the casual guide, we then utilized a joint score to determine the highest-scoring guide, that considered both the homology and co-occurrence scores (equation1).Where gscore is the joint guide score, h is the homology between the candidate guide sequence and the canonical edit site, hi is a constant (hi = 15), c is the coincidence score, and ci is also constant (ci = 0.7). The causal guide was taken as the guide with the highest gscore.
[0288] Differential gene expression: To call differentially expressed genes with respect to the number of allelic edit calls, we developed an approach inspired by perturb-seq linear models to call single cell edit effects, but also allowing for non-linear effects of confounding variables similar to mixscape. For a given edited gene, we first determine a set of cells consisting of the unedited control cells and all cells with at least one allelic edit call in the condition associated with the given edited gene (e.g. NuRD KD treated cells if CHD4 is the edited gene). We then determine the set of cells that have the most frequently occurring allelic edit dosage for the given edited gene cellsmode aiieie) . For each cell in cellsmode aiieie, an example cell at each other gene allelic edit dosage is selected that has the most similar cell state and total UMIs as the mode allele cell if there are more than 50 example cells at a given allelic edit dosage. The most similar cell at each allelic dosage is determined as the nearest neighbor cell by manhattan distance to the mode allele cell, comparing z-score standardized UMAP coordinates (cell-state) and z-score normalized natural logarithm of the total UMI counts across genes. This effectively constructs ‘allelic pairings’ of cell groups (g), whereby within each group cells are matched by cell state and total UMIs, but have differing called allelic edit dosages with respect to a given edited gene.
[0289] To call differentially expressed genes with respect to a given edited gene, we use linear mixture modeling, to model the expression of a given query gene i in cell j (yij) as a function of gene e allelic edits in cell j (aej) as a fixed effect variable and allelic pairing group k of cell j gk) as a random effect variable (equation 2).
[0290] Where P is the fixed effect of the number of allelic edits of gene e (aej) on gene i expression (y ), and Ztis the random effect of cell group gk (which is constructed to account for cell state and cell sequencing depth, as detailed above). Differentially expressed genes were called as those with a Benjamini-Hocberg false-discovery rate adjusted p-value for P 0 that was smaller than 0.05 prior to gene set over enrichment analysis (see below).
[0291] The set of edited genes (e) was determined as those with more than 30 cells with 1 or more allelic edits and a mean expression level greater than 0.2. This resulted in selection of our on-target genes (CHD3, CHD4, ARID 1 A, and SMARCA4) and several off-target genes (USP9X, CDC27, IBXO38-DI. and CNIH3) as adequately expressed. We then tested whether each of these expressed edited genes were differentially expressed relative to the number of allelic edits for that gene, to determine if the Cas9 edits were functional. For the on-target genes and the off-target genes that had differential expression relative to the number of allelic edits for that gene (CH D3, CHD4, ARID 1 A, SMARCA4, USP9X, and CDC27), we then tested for differential expression for these edited genes against highly variable genes. Highly variable genes were determined using scanpy, with parameters min mean=0.2, min disp 0.6. and max mean=5. Additionally, we also included genes which were determined to be NuRD or BAF regulated genes in K562 cells (details below), where the min mean=0.03 and min_disp=0 for NuRD targets, and min mean= 0.2 and min disp=0.3 for BAF targets. These additional genes were selected for differential expression analysis to ensure adequate numbers of the BAF and NuRD target genes were tested for robust gene set over enrichment analysis.
[0292] Gene set over-enrichment analysis: We constructed gene sets of NuRD and BAF targets for subsequent gene set over-enrichment analysis within our DEGs called with respect to our NuRD knock-down cells (CHD3 / CHD4 edited cells) and BAF knock-down cells (ARID1A / SMARCA4 edited cells). NuRD targets were determined by downloading ENCODE hg38 IDR thresholded CHD4 ChlP-seq peaks from K562 cells (file accessionENCFF985QBS). These peaks were then intersected with hg38 gene transcription start sites (TSSs) to determine NuRD target genes. For BAF target genes, we downloaded ENCODE hg38 pseudo-replicated ATAC-seq peaks (file accession ENCFF333TAT) and IDR thresholded SMARCA4 ChlP-seq peaks from K562 cells (file accession ENCFF267OGF). The K562 ATAC-seq peaks and the SMARCA4 ChlP-seq peaks were then intersected to obtain K562 open SMARCA4 binding sites. The open SMARCA4 bindings sites were then intersected with gene TSSs to obtain BAF promoter target genes. Since BAF can also act at gene enhancers, we additionally downloaded hgl9 K562 enhancer-gene associations from EnhancerAtlas 2.0. The enhancer hgl9 genomic coordinates were then converted to hg38 using LiftOver, and intersected with the open SMARCA4 peaks. Genes associated with the enhancer genome annotations that overlapped the open SMARCA4 ChlP-seq peaks were then considered BAF enhancer target genes.
[0293] CDC27 and USP9X functional gene sets were ‘Hallmark G2M Checkpoint’ and ‘WNT Signalling’ from mSigDB , as justified in the Results section.
[0294] After using Fisher’s exact test to determine association between the DEGs and each of the functional gene sets determined above, we then used the Benjamini -Hochberg method for false-discovery-rate correction of p-values, and took associations with p<0.05 as significant associations.
[0295] While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.
[0296] In describing the various embodiments, the specification can have presented a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described, and one skilled in the art can readily appreciate that the sequences can be varied and still remain within the spirit and scope of the various embodiments.
Claims
CLAIMS:
1. A method for genetically modifying one or more cells, comprising:(a) contacting the one or more cells with one or more gene editing compositions comprising: (i) one or more exogenous nucleases or one or more encoding nucleic acids thereof; and (ii) an exogenous insert oligonucleotide comprising a bacteriophage promoter sequence operably linked to a barcode sequence, or an encoding nucleic acid thereof; and(b) introducing the one or more exogenous nucleases or one or more encoding nucleic acids thereof and the exogenous insert oligonucleotide into the one or more cells, wherein after introduction, the one or more exogenous nucleases cleave one or more endogenous nucleic acids at one or more edit sites, and wherein the exogenous insert oligonucleotide is inserted into at least one edit site.
2. The method of claim 1, further comprising analyzing one or more nucleic acid analytes derived from the one or more edit sites.
3. The method of claim 2, wherein the one or more nucleic acid analytes comprise nucleic acid analytes derived from a single edit site or nucleic acid analytes derived from more than one edit sites.
4. The method of claim 3, wherein the more than one edit sites comprise edit sites in a single cell and / or more than one cells.
5. The method of any one of claims 2 to 4, wherein the one or more nucleic acid analytes are derived from an edit site in a target gene locus, an edit site at an off-target gene locus, or any combination thereof.
6. The method of any one of claims 2 to 5, wherein each nucleic acid analyte comprises the barcode or a portion thereof and an extension sequence corresponding to a portion of the endogenous nucleic acid near to or adjacent to the edit site.
7. The method of any one of claims 2 to 6, wherein the one or more nucleic acid analytes comprise RNA or a cDNA thereof.
8. The method of claim 7, wherein the one or more nucleic acid analytes comprise RNA transcripts of the bacteriophage promoter, or cDNA thereof.
9. The method of claim 8, further comprising performing in situ transcription or in vitro transcription of the one or more cells or isolated nuclei thereof to generate the RNA transcripts of the bacteriophage promoter.
10. The method of claim 9, wherein performing in situ transcription further comprises (i) fixing the one or more cells and (ii) contacting the one or more fixed cells with a transcription composition comprising an RNA polymerase corresponding to the bacteriophage promoter.
11. The method of claim 10, wherein the transcription composition further comprises nucleoside triphosphates, magnesium, dithiothreitol (DTT), spermidine, inorganic pyrophosphate or any combination thereof.
12. The method of any one of claims 2 to 11, wherein analyzing the one or more nucleic acid analytes comprises determining a sequence or sequences of the one or more nucleic acid analytes, or any reverse complement thereof.
13. The method of any one of claims 1 to 12, wherein the method further comprises analyzing endogenous gene expression in the one or more cells.
14. The method of claim 13, comprising analyzing one or more endogenous RNAs of the one or more cells.
15. The method of claim 14, wherein, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are increased relative to a nonedited control cell.
16. The method of claim 14, wherein, after insertion of the exogenous oligonucleotide, levels of the one or more endogenous RNAs in the cell are decreased relative to a non-edited control cell.
17. The method of any one of claims 14 to 16, wherein the endogenous RNAs comprise mRNA, tRNA, miRNA, rRNA, mtRNA, or any combination thereof.
18. The method of any one of claims 14 to 17, wherein one or more endogenous RNAs are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus.
19. The method of any one of claims 14 to 18, wherein the one or more endogenous RNAs are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus.
20. The method of claim 13, comprising detecting a change in activity of one or more endogenous proteins of the one or more cells.
21. The method of claim 20, wherein, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is increased relative to a non-edited control cell.
22. The method of claim 20, wherein, after insertion of the exogenous oligonucleotide, activity of at least one endogenous protein in one or more cells is decreased relative to a non-edited control cell.
23. The method of any one of claims 20 to 22, wherein detecting the change of activity of the one or more endogenous proteins comprises detecting a change of an expression level of the one or more endogenous proteins.
24. The method of any one of claims 20 to 23, wherein the one or more endogenous proteins are encoded by a target gene locus or an endogenous nucleic acid that is operably linked to the target gene locus.
25. The method of any one of claims 20 to 24, wherein the one or more endogenous proteins are encoded by an off-target gene locus or an endogenous nucleic acid that is operably linked to the off-target gene locus.
26. The method of any one of claims 2 to 25, wherein analyzing the one or more nucleic acid analytes, the one or more endogenous RNAs and / or the one or more endogenous proteins comprises combinatorial indexing.
27. The method of claim 26, wherein the combinatorial indexing comprises single cell RNA sequencing.
28. The method of claim 27, wherein the combinatorial indexing comprises SPLiT-seq, droplet based RNA sequencing, 10X sequencing, sci-RNA-seq, or any combination thereof.
29. The method of claim 27 or 28, wherein the combinatorial indexing comprises SPLiT- Seq.
30. The method of any one of claims 1 to 29, wherein the one or more cells repair the cleavage of the endogenous nucleic acid and inserts the exogenous oligonucleotide at each edit site using a repair mechanism.
31. The method of claim 30, wherein the repair mechanism comprises non-homologous end joining (NHEJ), microhomology mediated end joining (MMEJ), homologous recombination (HR), viral genome integration, transposition, or any combination thereof.
32. The method of claim 31, wherein the repair mechanism comprises non-homologous end joining (NHEJ).
33. The method of any one of claims 1 to 32, wherein the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof.
34. The method of claim 33, wherein the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103.
35. The method of claim 33 or 34, wherein the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8) or (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT) (SEQ ID NO: 9).
36. The method of any one of claims 1 to 35, wherein the one or more exogenous nucleases comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease,a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme.
37. The method of claim 36, wherein the one or more exogenous nucleases comprise a CRISPR associated (Cas) endonuclease and the one or more gene editing compositions further comprises one or more guide RNAs or one or more encoding nucleic acids thereof.
38. The method of claim 37, wherein the one or more gene editing compositions further comprise one or more ribonucleoprotein (RNP) complexes each comprising the Cas endonuclease and at least one gRNA.
39. The method of claim 38, wherein the one or more gene editing composition comprise one or more expression constructs encoding the one or more guide RNAs.
40. The method of any one of claims 37 to 39, wherein the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor.
41. The method of claim 40, wherein the single-strand editing “nickase” editor comprises a prime editor.
42. The method of any one of claims 36 to 40, wherein the Cas endonuclease comprises Cas9, Cas3, CaslO, or Casl2.
43. The method of any one of claims 1 to 42, wherein the one or more gene editing compositions comprise one or more expression constructs comprising the one or more nucleic acids encoding the one or more exogenous nucleases, the exogenous oligonucleotide and / or the one or more gRNAs.
44. The method of claim 43, wherein the expression constructs comprise an expression constructs of any one of claims 76 to 86.
45. The method of any one of claims 1 to 44, wherein introducing the one or more exogenous nucleases or one or more encoding nucleic acids thereof and / or the exogenous insert oligonucleotide into the cell comprises transfection, transduction, electroporation or injection.
46. The method of any one of claims 45, wherein the one or more gene editing compositions further comprises a transfection reagent, a transduction reagent and / or an electroporation reagent.
47. The method of claim 46, wherein the one or more gene editing compositions comprise an electrolytic / isotonic buffer, an exogenous “enhancer” oligonucleotide that increases electroporation delivery efficiency, or any combination thereof.
48. A gene editing system comprising:(a) one or more nucleases or one or more encoding nucleic acids thereof; and(b) an exogenous insert oligonucleotide or encoding nucleic acid thereof, the exogenous insert oligonucleotide comprising a bacteriophage promoter sequence and a barcode sequence.
49. The gene editing system of claim 48, wherein the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof.
50. The gene editing system of claim 49, wherein the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1 to 7 and 103.
51. The gene editing system of any one of claims 48 to 50, wherein the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8).
52. The gene editing system of any one of claims 48 to 51, wherein the insert oligonucleotide comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO: 9).
53. The gene editing system of any one of claims 48 to 52, wherein the exogenous insert oligonucleotide comprises one or more modified nucleotides.
54. The gene editing system of any one of claims 48 to 53, wherein the exogenous insert oligonucleotide is a synthetic oligonucleotide.
55. The gene editing system of any one of claims 48 to 54, wherein the one or more nucleases comprise CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a homing endonuclease, a transposase, an integrase, or a restriction enzyme, or any combination thereof.
56. The gene editing system of claim 55, wherein the one or more exogenous nucleases comprise a CRISPR associated (Cas) endonuclease and the gene editing system further comprises (c) one or more guide RNAs or one or more encoding nucleic acids thereof.
57. The gene editing system of claim 56, wherein the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor.
58. The gene editing system of claim 57, wherein the single-strand editing “nickase” editor comprises a prime editor.
59. The gene editing system of any one of claims 56 to 58, wherein the Cas endonuclease comprises Cas9, Cas3, CaslO, or Casl2.
60. The gene editing system of any one of claims 56 to 59, wherein the gene editing system further comprises one or more ribonucleoprotein (RNP) complexes each comprising the Cas endonuclease and at least one gRNA.
61. The gene editing system of any one of claims 48 to 60, wherein (a), (b), and / or (c) are comprised in one or more gene editing compositions.
62. The gene editing system of claim 61, wherein the one or more gene editing compositions comprise the encoding nucleic acids of (a), (b), and / or (c).
63. The gene editing system of claim 62, wherein the one or more gene editing compositions comprise an expression constructs of any one of claims 76 to 86.
64. The gene editing system of any one of claims 61 to 63, wherein the one or more gene editing compositions further comprise at least one reagent for transfection, transduction, electroporation, or injection.
65. The gene editing system of any one of claims 61 to 64, wherein the one or more gene editing compositions further comprise a cell culture medium.
66. A nucleic acid comprising a bacteriophage promoter operably linked to a barcode sequence.
67. The nucleic acid of claim 66, wherein the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof.
68. The nucleic acid of claim 66 or 67, wherein the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103.
69. The nucleic acid of any one of claims 67 to 68, comprising a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8).
70. The nucleic acid of any one of claims 67 to 69, comprising a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO: 9).
71. The nucleic acid of any one of claims 66 to 70, further comprising one or more modified nucleotides.
72. The nucleic acid of claim 71, wherein the one or more modified nucleotides are located at a 5’ end, a 3’ end, or any combination thereof of the nucleic acid.
73. The nucleic acid of claim 71 or 72 wherein at least one modification on the one or more modified nucleotides comprises a 5’ phosphate and / or a phosphorothioate linkage.
74. The nucleic acid of any one of claims 66 to 73, wherein the nucleic acid comprises DNA.
75. The nucleic acid of any one of claims 66 to 74, wherein the nucleic acid is synthetic.
76. An expression construct encoding for a nucleic acid of any one of claims 66 to 75.
77. The expression construct of claim 76, further comprising one or more regulatory elements configured to regulate the expression of the nucleic acid.
78. The expression construct of claim 76 or 77, further comprising(a) a nucleic acid encoding for a nuclease ; and(b) one or more regulatory elements configured to regulate the expression of (a).
79. The expression construct of claim 78, wherein the nuclease comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme80. The expression construct of claim 79, wherein the nuclease comprises a Cas endonuclease.
81. The expression construct of claim 80, wherein the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor.
82. The expression construct of claim 81, wherein the single-strand editing “nickase” editor comprises a prime editor.
83. The expression construct of any one of claims 80 to 82, wherein the Cas endonuclease comprises a Cas9, Cas3, CaslO, or Casl2 endonuclease.
84. The expression construct of any one of claims 80 to 83, further comprising:(c) a nucleic acid encoding for a guide RNA (gRNA); and(d) one or more regulatory elements configured to regulate the expression of (c).
85. The expression construct of any one of claims 80 to 84, wherein the expression construct is a plasmid or a viral vector.
86. The expression construct of claim 85, wherein the viral vector is a lentiviral vector or an adeno-associated viral vector (AAV).
87. A genetically modified cell, comprising one or more nucleic acid insertions, each nucleic acid insertion comprising a bacteriophage promoter sequence and a barcode.
88. The genetically modified cell of claim 87, wherein the bacteriophage promoter sequence comprises a T7 promoter sequence, a T3 promoter sequence, a SP6 promoter sequence or any combination thereof89. The genetically modified cell of claim 87 or 88, wherein the bacteriophage promoter sequence comprises a nucleic acid sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 1-7 and 103.
90. The genetically modified cell of claim 89, wherein each nucleic acid insertion comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to TAATACGACTCACTATAGnnnnnnn (SEQ ID NO: 8).
91. The genetically modified cell of claim 90, wherein each nucleic acid insertion comprises a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to (GTGAATTTAATACGACTCACTATAGnnnnnnnnAT)(SEQ ID NO: 9).
92. A genetically modified cell comprising the nucleic acid of any one of claims 66 to 75 or the expression construct of any one of claims 76 to 85.
93. A genetically modified cell edited using a gene editing system of any one of claims 48 to 65.
94. The genetically modified cell of any one of claims 87 to 93, wherein the cell is eukaryotic.
95. The genetically modified cell of claim 94, wherein the cell is mammalian.
96. The genetically modified cell of claim 95, wherein the cell is murine or human.
97. A kit comprising a nucleic acid comprising a bacteriophage promoter operably linked to a barcode sequence and at least one container.
98. The kit of claim 97 comprising an expression vector encoding for the nucleic acid.
99. The kit of claim 98, wherein the expression vector further comprises one or more regulatory elements configured to regulate the expression of the nucleic acid.
100. The kit of any one of claims 97 to 99, further comprising one or more nucleases or a nucleic acid encoding for the one or more nucleases.
101. The kit of claim 100, further comprising an expression vector comprising the nucleic acid encoding for the one or more nucleases.
102. The kit of claim 101, wherein the expression vector further comprises one or more regulatory elements configured to regulate the expression of the one or more nucleases.
103. The kit of any one of claims 100 to 102, wherein the one or more nucleases comprise comprises a CRISPR associated (Cas) endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a transposase, an integrase, or a restriction enzyme or any combination thereof.
104. The kit of claim 103, wherein the one or more nucleases comprise comprises a CRISPR associated (Cas) endonuclease.
105. The kit of claim 104, wherein the Cas endonuclease comprises a double strand editing Cas or a single-strand editing Cas “nickase” editor.
106. The kit of claim 105, wherein the single-strand editing “nickase” editor comprises a prime editor.
107. The kit of any one of claims 104 to 106, wherein the Cas endonuclease comprises a Cas9, Cas3, Cas 10, or Cas 12 endonuclease.
108. The kit of any one of claims 104 to 107, further comprising one or more gRNAs or an encoding nucleic acid thereof.
109. The kit of claim 108, comprising an expression vector encoding the one or more gRNAs.
110. The kit of claim 109, wherein the expression vector further comprises one or more regulatory elements configured to regulate the expression of the one or more gRNAs.
111. The kit of any one of claims 97 to 110, wherein each component in the kit is comprised in one or more compositions.
112. The kit of claim 111, wherein the one or more compositions further comprise one or more reagent for transfection, transduction, electroporation, or injection.
Citation Information
Patent Citations
Compositions, methods, modules and instruments for automated nucleic acid-guided nuclease editing in mammalian cells using microcarriers
US11591592B2
Single cell analysis of transposase accessible chromatin
US20180340172A1
Methods and compositions for single cell analysis
US20230304069A1
Compositions and methods for detecting bacterial nucleic acid and diagnosing bacterial vaginosis
WO2020041671A1